BenchmarksCoding agents

SWE-bench, Terminal-Bench and METR: what coding-agent evaluations measure, and what their results don't prove

OpenAI dropped SWE-bench Verified, Terminal-Bench is on version 4.0, and METR can't yet pin down AI's effect on developer speed. What published evaluations show, and their limits.

By Shajanthan8 min read
ByShajanthanFounder & Editor
Published
Reading8 MIN
Chart of METR's published estimates of how AI tools changed developers' task completion time in 2025 and 2026, with confidence intervals
Results reported9to5 AI did not run these tests.
Run byEach benchmark’s maintainers (SWE-bench team, Scale AI, Terminal-Bench, OpenAI, METR)
Results as ofOct 9, 2026
Contents

Every new model launch comes with a coding agent benchmark score, and every score is a little higher than the last. This story goes through the published evaluations behind those numbers: SWE-bench and its Verified and Pro variants, Terminal-Bench, SWE-Lancer, and METR's studies of developers. For each one it covers who ran it, when, on what setup, what the metric means, and what the results do and don't establish. None of these evaluations tells you whether Claude Code, Codex or Cursor will close tickets in your own codebase. The last section offers advice for teams that want to answer that question themselves.

All results below are as published by the organizations named. Last verified: October 9, 2026.

SWE-bench and its variants

Who runs it: the SWE-bench project (swebench.com). Model developers and agent builders submit results, and the site marks runs "performed or directly checked by the SWE-bench team." Metric: the share of tasks resolved. The agent gets a repository at the commit before a real fix, plus the issue text. A task counts as resolved if the agent's patch makes the project's hidden tests pass without breaking existing ones.

The site lists six variants:

VariantInstancesWhat it is
Full2,294Real GitHub issues from 12 Python repositories
Verified500A human-filtered subset
Bash Only500Verified, with every model in the same minimal mini-SWE-agent harness
Lite300A cheaper subset
Multimodal480Issues that include visual elements
Multilingual30042 repositories in 9 languages

Limitation: the harness matters. The harness is the scaffolding around the model: its tools, prompts and retry logic. The same model can score very differently in different harnesses. The site describes Bash Only as "every model in the same mini-SWE-agent environment," which exists to hold the harness constant. A leaderboard row is a result for a model plus a harness, never a model alone, and certainly not a product like Claude Code.

Why SWE-bench Verified lost its standing

Who ran the audit: OpenAI, published February 23, 2026, in "Why SWE-bench Verified no longer measures frontier coding capabilities." OpenAI gave two reasons:

  • Problems with tests and descriptions. OpenAI audited 138 problems, a "27.6% subset of the dataset." It found that "59.4% of the 138 problems contained material issues in test design and/or problem description." OpenAI broke these down as follows:
    • 35.5% of audited tasks had tests too strict, enforcing specific implementation details.
    • 18.8% had tests that checked for extra functionality.
    • 5.1% had other issues.
  • Contamination. OpenAI used GPT-5 to probe GPT-5.2-Chat, Claude Opus 4.5 and Gemini 3 Flash Preview. It reported that "all frontier models we tested were able to reproduce the original, human-written bug fix." In other words, the answers had leaked into training data.

OpenAI said it had stopped reporting Verified scores and "recommends reporting results for SWE-bench Pro." Limitation: OpenAI is a competitor in this space, so treat its framing with some care. But the specific flaws it lists are concrete and checkable.

SWE-Bench Pro

Who built it: a team publishing through Scale AI (arXiv 2509.16941, first posted September 21, 2025, revised November 14, 2025). Scale AI publishes the leaderboard. The authors call it "a contamination-resistant testbed." It has 1,865 problems from 41 actively maintained repositories, split three ways:

  • Public set: 11 repositories, 731 instances. These are copyleft-licensed repositories, which the authors argue are less likely to end up in proprietary training data.
  • Held-out set: 12 repositories, 858 instances. Scale says results for this set will not be published on public leaderboards.
  • Commercial set: 18 proprietary startup repositories, 276 instances. Results are reported on a separate leaderboard.

The authors describe the tasks as "long-horizon tasks that may require hours to days for a professional software engineer to complete."

Published results. Scale AI's SWE-Bench Pro public-set leaderboard, as listed on October 9, 2026:

RankModel (as listed)% resolved (± as shown)
1Muse Spark 1.161.50 ± 3.10
2gpt-5.4 (xHigh)59.10 ± 3.56
3Muse Spark55.00 ± 3.60
4claude-opus-4-6 (thinking)51.90 ± 3.61
5gemini-3.1-pro (thinking)46.10 ± 3.60

Limitations:

  • All five entries carry an asterisk. The page's footnote explains it as "Run with mini-swe-agent harness."
  • Scale says entries were "run with an uncapped cost and with a turn limit of 250." Grayed-out entries use a capped cost and a 50-turn limit.
  • The page shows no "last updated" date.
  • The listed models are not necessarily each vendor's newest release, so check the page before citing a ranking.
  • These are model-plus-harness results on the public set only.

Terminal-Bench

Who runs it: the Terminal-Bench maintainers (tbench.ai). Runs use their Harbor evaluation framework. What it measures: agents completing tasks in a terminal, not just patching code. Metric: the leaderboard reports a resolution rate (the share of tasks resolved), with whiskers spanning a 95% confidence interval.

The benchmark moves quickly:

  • Version 2.0: November 7, 2025, launched with Harbor.
  • Version 2.1: May 6, 2026, a revision that fixed 28 tasks.
  • Version 3.0: July 30, 2026.
  • Version 4.0: August 28, 2026. It sets a flat eight-hour agent timeout and recalibrates time, CPU and memory. It removes 8 tasks:
    • 2 for saturation (every model class in the latest generation solved them 5 out of 5 times)
    • 2 for model refusals
    • 2 for publicly available solutions
    • 2 for unresolved quality or compatibility problems
    It also fixes 19 others. The maintainers call these breaking changes that require re-running trials. 4.0 scores are therefore not comparable with earlier versions.

On April 19, 2026, the maintainers published new leaderboard policies "to address cheating and reward hacking":

  • Passing trials must include agent trajectories.
  • A trial that exploits a loophole, such as finding solutions on the internet, scores zero.
  • A submitter who alters the benchmark to help their agent has the submission removed.
  • An agent judge reviews all passing trials.

The post names past cases, including an agent that sometimes fetched solutions from the internet into its AGENTS.md file. That's a reminder that agents can game benchmarks in ways people don't always anticipate.

Limitations: the 4.0 announcement doesn't state a total task count and reports results only in charts. This story doesn't quote 4.0 scores; check the leaderboard directly for current entries.

SWE-Lancer

Who ran it: OpenAI, introduced February 18, 2025 (SWE-Lancer). It prices tasks in dollars, with "over 1,400 freelance software engineering tasks from Upwork" and "$1 million USD total in real-world payouts." There are two kinds of task:

  • Engineering tasks are graded by end-to-end tests that OpenAI says were "triple-verified by experienced software engineers."
  • Management tasks ask the model to choose between implementation proposals. They are scored against the choice the original hiring manager made.

The public split is called SWE-Lancer Diamond. Limitation: it is vendor-built, and its dollar values reflect the original Upwork payouts, not the economic value of an agent's work.

METR's productivity studies: the human side

Benchmarks measure agents working alone. METR's studies measured developers using AI.

  • July 10, 2025 (study): a randomized controlled trial.
    • Setup: 16 experienced open-source developers completed 246 real issues in their own repositories. They mostly used Cursor Pro with Claude 3.5 and 3.7 Sonnet.
    • Metric: change in task completion time when AI was allowed.
    • Result: tasks took 19% longer with AI (confidence interval +2% to +39%). Beforehand, developers had predicted a 24% speedup. Afterwards, they still believed AI had sped them up by 20%.
    • Status: METR's page now says: "These results are out of date."
  • February 24, 2026 (update): a larger follow-up covering late 2025.
    • Setup: 57 developers across 143 repositories and 800+ tasks.
    • Result: METR estimates a "speedup of -18%" for returning developers from the original study (confidence interval −38% to +9%) and −4% for newly recruited developers (−15% to +9%). METR presents these as some evidence that AI now speeds developers up, but both intervals include zero.
    • Limitations: METR stresses that the estimate is likely biased. More developers said they would not want to do half their work without AI, and 30% to 50% said they chose not to submit some tasks they didn't want to do unaided. METR says that makes the estimate "a lower-bound on the true productivity effects" and calls it only "very weak evidence" about the size of the change. It is redesigning the study.
  • May 11, 2026 (survey): a survey, not a trial.
    • Setup: 349 technical workers, surveyed February to April 2026.
    • Result: respondents reported a median 1.4–2x increase in the value of their work.
    • Limitations: METR warns that "survey results are not necessarily grounded in reality." The convenience sample had a response rate of about 2%. METR also notes that developers in its earlier study overestimated AI's effect on their time by 40 percentage points on average.

The lesson: perceived speed isn't measured speed.

What these results can't tell you

  • Your code isn't in them. Public tasks come from public repositories, and the original SWE-bench is entirely Python. Your monorepo, internal frameworks and conventions are what matter.
  • Products aren't harnesses. Claude Code, Codex and Cursor each add their own system prompts, tools, context management, instruction files and extensions. On instruction files, see our CLAUDE.md and AGENTS.md guide; on Claude Code's extensions, see mods. A leaderboard score for "model X in mini-swe-agent" doesn't transfer.
  • Cost and limits are missing. A pass rate means little if a run uses up a $20 plan's five-hour window. Our comparison of the three tools explains how each one meters usage.
  • Review cost is invisible. A patch that passes tests but takes an hour to review isn't a win.
  • Results age fast. Benchmark versions change, leaderboards get re-scored, and the models at the top are often a generation behind what you can buy.

If you evaluate agents on your own code

The evaluations above suggest some practices for a team comparing agents internally. These are suggestions, not a published standard:

  • Use real, closed tickets with tests. For example, 20–30 recent tickets whose fixes came with tests, so you can check the result automatically. If your repository is public, prefer fixes that aren't widely mirrored, to limit contamination.
  • Freeze the starting point. Start every run from the commit before the fix. Keep the fix's tests hidden from the agent until grading, and pin dependencies so every run sees the same environment.
  • Record each agent's setup. Note the version, model, reasoning effort, plan and permission mode, and give every agent the same instruction file and prompt. Our Codex explainer covers Codex's approval modes.
  • Run more than once. Agents are non-deterministic, so several runs per ticket show the spread as well as the average.
  • Measure more than pass rate. Track cost or plan usage, wall time, human interventions, and a blind review of whether you would merge the patch.
  • Be honest about noise. With a few dozen tickets, small differences in pass rate are likely noise. Report results per ticket.

At the current pace of model releases, plan to repeat the comparison when your tools change.

About this storyBased on the sources linked below. Editorial standards

Was this useful?Report an error
Compare nextClaude Code vs Codex vs Cursor: features, prices and limits compared (October 2026)Open the living comparison LearnThe Claude app, explained: plans, models, Cowork and how usage limits workExplainer · 6 min
Comments
0

More on Coding agents & agents

The Week in AI

The numbers that matter, every Friday.

0