Topic

evaluation

3
Pieces
JUN 28, 2026
Last filed
Tagged evaluation clear ×
JUN 28, 2026
Business deep
The eval vendors are quietly pivoting from grading to simulation
Three days after we argued the agent-eval sector sells a number nobody trusts, Patronus AI raised $50M and reframed itself around Digital World Models. It is not defending the benchmark. It is replacing it with simulation, and the move concedes the original critique.
5 MIN 6 src
JUN 25, 2026
Business deep
The agent-eval startups raising on a metric nobody trusts
Investors have poured serious capital into agent evaluation, observability and benchmarking. **Braintrust** raised **$80M at $800M**, **LangChain** **$125M at $1.25B**, **LMArena** **$100M then $150M at $1.7B**, **Patronus** a fresh **$50M Series B** on June 25. The product they sell is a number. The research says the number is broken: up to **100%** relative error on agent benchmarks, **27** private variants gaming one leaderboard, and a 1MB blind script beating frontier agents.
11 MIN 21 src
JUN 22, 2026
Agents deep
How much is the harness worth? The year the number got measured
Six weeks after agent harness engineering got named, the empirical question arrived: how much of a coding agent's score is the model, and how much is the scaffold around it. Four 2026 papers put numbers on it — and the numbers are larger than almost anyone guessed.
35 MIN 15 src