Methodology · v2.1

How we test.

The rules behind every Benchmarked issue, review score and hands-on result. Versioned, public, and changed only with a changelog entry.

01 — Principles

We test the product as a paying customer would use it: current public version, default settings unless stated, our own accounts. Vendors don’t see results before publication.

Every number we publish can be traced to a run log. If we can’t reproduce a result, we don’t publish it.

02 — Tasks and datasets

Tasks come from real work with known-good answers — tickets we already solved by hand, documents with verified facts, prompts with checkable citations. Task sets are frozen per methodology version so results stay comparable across rounds.

03 — Runs and repeats

Each task runs several times per tool on identical hardware. We report the pass rate across runs, not the best attempt. Flaky infrastructure failures are re-run once and logged; tool failures count as failures.

04 — Scoring

A pass is binary: the change merged after at most one review round with tests green, or the answer matched the source. Review scores combine task success (40%), review effort (25%), cost (15%), speed (10%) and safety behaviour (10%).

Costs are list prices from provider dashboards on the test date.

05 — Open data

Raw results are released as CSV under CC BY 4.0. Please cite “9to5 AI” and link the issue. Corrections to datasets are versioned and noted on the issue page.

06 — Changelog

v2.1Added safety-behaviour weighting; review effort counts only human rounds.
v2.0Frozen 40-ticket set across three repositories; five runs per ticket.
v1.4Citation checks verified by two editors independently.