Benchmarks

Published benchmark results, credited to the people who ran them: what was measured, which versions, when, and what the numbers don’t tell you.

Benchmarks · Long contextLong-context recall: what independent benchmarks say "1 million tokens" really buys youOct 9 · 9 min
Benchmarks · Coding agentsSWE-bench, Terminal-Bench and METR: what coding-agent evaluations measure, and what their results don't proveOct 9 · 8 min
All benchmarks · 2

Nothing in this format on this page — try another filter or the next page.