01 — Principles
We test the product as a paying customer would use it: current public version, default settings unless stated, our own accounts. Vendors don’t see results before publication.
Every number we publish can be traced to a run log. If we can’t reproduce a result, we don’t publish it.
02 — Tasks and datasets
Tasks come from real work with known-good answers — tickets we already solved by hand, documents with verified facts, prompts with checkable citations. Task sets are frozen per methodology version so results stay comparable across rounds.
03 — Runs and repeats
Each task runs several times per tool on identical hardware. We report the pass rate across runs, not the best attempt. Flaky infrastructure failures are re-run once and logged; tool failures count as failures.
04 — Scoring
A pass is binary: the change merged after at most one review round with tests green, or the answer matched the source. Review scores combine task success (40%), review effort (25%), cost (15%), speed (10%) and safety behaviour (10%).
Costs are list prices from provider dashboards on the test date.
05 — Open data
Raw results are released as CSV under CC BY 4.0. Please cite “9to5 AI” and link the issue. Corrections to datasets are versioned and noted on the issue page.