Topic

benchmarks

7
Pieces
JUL 24, 2026
Last filed
Tagged benchmarks clear ×
JUL 24, 2026
Models deep
Half the price of frontier
Anthropic shipped Claude Opus 5 as near-frontier intelligence at half the cost of Fable 5. The launch reads as good news for buyers. Read from the P&L, it is a margin admission — the clearest signal yet that the race has moved from raw capability to cost-per-task, and that nobody expects customers to pay the old premium for the top of the curve.
12 MIN 6 src
JUL 04, 2026
Models deep
The gate and the giveaway
While Washington spent June turning frontier model releases into licensed events, a Chinese food-delivery company open-sourced a 1.6-trillion-parameter agentic-coding model under an MIT license — trained start to finish on domestic chips, no Nvidia silicon involved. LongCat-2.0 is the counter-move to the export-control regime, and it's already downloadable worldwide.
8 MIN 7 src
JUN 28, 2026
Business deep
The eval vendors are quietly pivoting from grading to simulation
Three days after we argued the agent-eval sector sells a number nobody trusts, Patronus AI raised $50M and reframed itself around Digital World Models. It is not defending the benchmark. It is replacing it with simulation, and the move concedes the original critique.
5 MIN 6 src
JUN 25, 2026
Business deep
The agent-eval startups raising on a metric nobody trusts
Investors have poured serious capital into agent evaluation, observability and benchmarking. **Braintrust** raised **$80M at $800M**, **LangChain** **$125M at $1.25B**, **LMArena** **$100M then $150M at $1.7B**, **Patronus** a fresh **$50M Series B** on June 25. The product they sell is a number. The research says the number is broken: up to **100%** relative error on agent benchmarks, **27** private variants gaming one leaderboard, and a 1MB blind script beating frontier agents.
11 MIN 21 src
JUN 24, 2026
Models deep
The open-weight reasoning gap is now 3.4 months, and the math is public
On June 16, GLM-5.2 became the leading open-weight model on the Artificial Analysis Intelligence Index v4.1 at **51**, ahead of every Google model and **3.4 months** behind the equally-capable closed model. The smallest gap on record is now **2.2 months** (DeepSeek V4 Pro). The peak was **9.8 months** in December 2024. The reasoning frontier the closed labs owned outright is now a quarter-year head start, priced in the open.
10 MIN 19 src
JUN 22, 2026
Agents deep
How much is the harness worth? The year the number got measured
Six weeks after agent harness engineering got named, the empirical question arrived: how much of a coding agent's score is the model, and how much is the scaffold around it. Four 2026 papers put numbers on it — and the numbers are larger than almost anyone guessed.
35 MIN 15 src
JUN 04, 2026
Models deep
The open-weight coding frontier caught Claude, and it speaks Mandarin
In a three-week window, MiniMax M3, Qwen3.7-Max, DeepSeek V4, Kimi K2.6 and Nemotron 3 Ultra all claimed the frontier. The third-party leaderboards say the gap is real and small. The marquee numbers are mostly the labs grading their own homework. And the open-weight crown carries a jurisdiction risk no benchmark measures.
18 MIN 11 src