BenchmarksLong context

Long-context recall: what independent benchmarks say "1 million tokens" really buys you

Every frontier API now accepts about a million tokens. Independent evaluations from RULER to ATLAS show how much of that window models actually use.

By Shajanthan9 min read
ByShajanthanFounder & Editor
Published
Reading9 MIN
Diagram of a one-million-token context with five facts placed at different depths, the later ones fading
Results reported9to5 AI did not run these tests.
Run byOpenAI, Fiction.live, Artificial Analysis, Chroma and the model vendors (vendor-reported where noted)
Results as ofOct 8, 2026
Contents

A million-token context window is now standard on frontier APIs. Anthropic's Claude Opus 5.5, Sonnet 5.5 and Haiku 5.5, OpenAI's GPT-6 Astra and GPT-6.1 Sol, and Google's Gemini 3.8 Flash and Gemini 3.1 Pro Preview all accept roughly 1M tokens of input, which is several novels' worth of text. The harder question in long context is how much of that input a model can use reliably: whether it finds the one clause in a contract that matters, connects two facts 400,000 tokens apart, and ignores the passage that looks right but isn't. This story covers how researchers measure that, what the vendors claim, and what independent evaluations have found. Each result below is credited to the team that ran it, with its date, the models it covered, its metric and its limits. Numbers that the model makers published about their own models are labeled vendor-reported.

Window size and recall are different things

The context window is the maximum number of tokens a model can take in one request. Anthropic's docs define it as the model's "working memory" and, unusually for a vendor, add a warning: "As token count grows, accuracy and recall degrade, a phenomenon known as context rot. This makes curating what's in context just as important as how much space is available."

Recall in this sense means whether the model can retrieve and correctly use information that is somewhere in its input. This is not the same thing as retrieval-augmented generation, where a separate system decides what goes into the input in the first place. We cover that distinction in Context windows vs memory vs retrieval (RAG).

How recall is measured

Needle in a haystack, and why it stopped being enough

The test that made long context a public talking point was Greg Kamradt's "needle in a haystack." In November 2023 he hid one out-of-place sentence in a long body of Paul Graham essays, varied the total length and the depth at which the sentence appeared (as a percentage of the document), and asked the model to retrieve it. The resulting heatmaps for GPT-4-128K (run dated November 8, 2023) and Claude 2.1 (November 21, 2023) were widely shared.

The test is easy to understand and easy to game. When the question shares words with the needle, the model can match strings instead of understanding anything. Labs began reporting near-perfect needle scores, and the benchmark stopped separating good models from weak ones.

RULER: more needles, more task types

RULER, from NVIDIA researchers led by Cheng-Ping Hsieh (posted to arXiv on April 9, 2024, revised August 6, 2024), kept the synthetic approach but made it harder. It adds multiple needles, different needle types, multi-hop "variable tracing" and aggregation tasks, across 13 tasks in total. The authors tested 17 models that all claimed at least 32K context and found that while the models scored near-perfectly on vanilla needle tests, "only half of them can maintain satisfactory performance at the length of 32K," judged against a fixed reference score. RULER popularized the idea of an effective context length that can be much shorter than the advertised one. Its limits: the tasks are synthetic, and the models it tested date from 2024.

NoLiMa: remove the word overlap

NoLiMa ("No Literal Matching"), first posted to arXiv on February 7, 2025 by researchers from Adobe Research and LMU Munich and accepted at ICML 2025, attacks the string-matching shortcut directly. Its needles share minimal vocabulary with the question, so the model has to make a latent association. For example, it has to know that a character who lives next to a famous opera house lives in a particular city. The metric is each model's accuracy at a given length compared with its own short-context baseline. In the paper's latest version (July 9, 2025), 11 of 13 models that claimed at least 128K context dropped below half their baseline at just 32K tokens. GPT-4o fell from 99.3% to 69.7%. The models tested date from 2025 or earlier.

MRCR: find the right one of several identical things

Multi-round co-reference resolution (MRCR) is a format Google DeepMind introduced in its September 2024 "Michelangelo" work. OpenAI published its own open version on Hugging Face in April 2025. In OpenAI's version, a long synthetic chat contains 2, 4 or 8 near-identical requests (for example, several poems about tapirs), and the model must return, say, the second one, prefixed with a given hash. Answers are graded by string similarity to the correct response, in length bins from about 4,000 to about 1 million tokens. OpenAI pushed a bugfix in December 2025 that regenerated roughly 15% of the rows.

MRCR is now the benchmark labs most often cite for 1M-token claims. It is harder than a single needle because every distractor looks like the target. It is still synthetic, though, and not the same as reasoning over a messy real document.

Tests built on real text: LongBench v2, Fiction.LiveBench, AA-LCR, ATLAS

  • LongBench v2, from Yushi Bai and colleagues (posted to arXiv on December 19, 2024), has 503 hard multiple-choice questions over contexts of 8,000 to 2 million words, covering documents, dialogue histories, code repositories and structured data. The metric is multiple-choice accuracy. Human experts given 15 minutes scored 53.7%. The best model answering directly scored 50.1%, and OpenAI's o1-preview, which reasons before answering, reached 57.7%. Those results cover late-2024 models.
  • Fiction.LiveBench is run by the Fiction.live writing platform and listed on Epoch AI's benchmark hub, which links reports dated February 21 and March 14, 2025. It asks 36 questions about 30 stories, each tested at a range of lengths from a short summary up to the full text. The questions require tracking characters' knowledge, chronology and implied information, not looking up a sentence. We don't quote scores here: the current leaderboard is published only as an interactive page, and Epoch's listing doesn't define the scoring metric.
  • AA-LCR (Artificial Analysis Long Context Reasoning), published by Artificial Analysis on August 5, 2025, asks 100 human-written questions over 30 document sets that average about 100,000 tokens. Answering requires combining information from several documents. Results are reported as the share of questions answered correctly, and every question has a human-verified answer. At launch, OpenAI's o3 led with 69.3%, ahead of xAI's Grok 4 (68%) and Qwen3 235B 2507 Reasoning (67%). Limits: the documents are text only, the questions were screened against smaller non-frontier models, and the launch results cover mid-2025 models.
  • ATLAS, posted to arXiv on May 27, 2026 by 18 researchers, most of them at Meituan (some also list Fudan University), scores 26 proprietary and open-weight models on 6,438 instances across a grid from 8K to 1M tokens. Its ATLAScore measures the area under each model's score-versus-length curve and combines task categories with a harmonic mean, so a model that is weak in one area is penalized. Each model ran with its provider's recommended settings. Inputs longer than a model's advertised window were truncated in the middle.
    • Up to 128K: Gemini 3.1 Pro Preview (high effort) led with 77.83, ahead of Claude Opus 4.6 (max effort) at 77.10 and GPT-5.5 (xhigh) at 74.63.
    • Up to 1M: Claude Opus 4.6 led with 70.55, ahead of Gemini 3.1 Pro Preview at 68.52 and GPT-5.5 at 67.77.
    • Seven models moved two or more places between the two ranges. The reported margins are 95% confidence intervals of about ±1.5 points, so the top two are close at both lengths.
    • Limits: English only, some components start at 32K rather than 8K, and the study predates the September and October 2026 releases.

Last verified: October 9, 2026 (arXiv, Artificial Analysis and Epoch AI pages)

Chroma's "context rot" study

On July 14, 2025, the vector-database company Chroma published a technical report that coined the phrase now used in Anthropic's own docs. It tested 18 models from Anthropic, OpenAI, Google and Alibaba with 194,480 calls; not every model appears in every experiment. Note that Chroma sells retrieval infrastructure, so it has a commercial interest in this result. Its findings:

  • Performance "grows increasingly unreliable as input length grows," even on simple tasks.
  • Questions that are semantically further from the needle degrade faster as length increases.
  • One topically similar distractor lowers accuracy, and four lower it further. GPT models hallucinated distractor answers most often, while Claude models tended to abstain.
  • Counterintuitively, shuffling the haystack's sentences improved performance compared with coherent text.
  • On its basic needle task, moving the needle across 11 positions made no notable difference.
  • On LongMemEval chat histories of about 113,000 tokens, every model did significantly better when given only the roughly 300 relevant tokens.

What vendors claim for today's 1M-token models (vendor-reported)

Last verified: October 9, 2026

Model (vendor)Context windowLong-context claimNotes
GPT-6 Astra (OpenAI)1,050,000MRCR v2 8-needle: 100.0% at 256K–512K; 96.3% at 512K–1MOpenAI's launch page; compares only with GPT-5.6 Sol (91.5% / 73.8%); OpenAI says its scores are "the maximum at any effort"
Gemini 3.1 Pro Preview (Google)1,048,576 inputMRCR v2 8-needle: 84.9% at 128K average; 26.3% at 1M (pointwise)Model card dated Feb 19, 2026; run at "Thinking High"; Gemini 3 Pro also scored 26.3% at 1M; Claude and GPT rivals listed as "not supported" at 1M
Claude Opus 4.6 (Anthropic, Feb 2026)1MMRCR v2 8-needle 1M: 76% (vs 18.5% for Sonnet 4.5)Anthropic announcement, Feb 5, 2026; no effort setting stated for this result
Claude Opus 5.5 / Sonnet 5.5 / Haiku 5.51M (default, no beta header)No long-context figure on model pagesOpus 5.5's system card lists a long-context section (8.10); Anthropic's docs warn about context rot
Gemini 3.8 Flash (Google)1,048,576 inputNo MRCR figure on its model cardCard published Sep 2, 2026

Three caveats apply to this table. First, these are vendor-reported numbers on benchmarks the vendors chose, and in OpenAI's case built. Second, the rows aren't comparable. "Average up to 128K" and "pointwise at 1M" measure different things, and OpenAI's 512K–1M bin isn't Google's 1M point. Third, independent, like-for-like long-context scores for the newest models from each lab have not been published. ATLAS, the most recent broad independent study, covers the previous generation: there, Gemini 3.1 Pro Preview led up to 128K and Claude Opus 4.6 led up to 1M, and rankings reshuffled between the two ranges.

What it costs to fill the window

Long context also has a price structure that shifts at scale. On Anthropic's current models, a 900K-token request is billed at the same per-token rate as a 9K one, except on Haiku 5.5, which charges more above 100,000 tokens. OpenAI doubles the input and cache rates and charges 1.5x for output on the whole request once input passes 272K tokens. Google's Gemini 3.1 Pro Preview moves from $2 to $4 per million input tokens above 200K. Filling the window once therefore costs about $2 in input on Claude Sonnet 5.5 ($2 per million tokens, no surcharge), but about $4 on GPT-6.1 Sol and Gemini 3.1 Pro Preview, where a 1M-token prompt is billed at the higher long-context rate. Repeated questions over the same corpus get expensive without prompt caching. Token pricing isn't the story covers why list prices mislead, and which Claude model to use covers Anthropic's tiers.

Last verified: October 8, 2026 (official pricing pages)

What this means in practice

  • Treat the advertised window as a ceiling, not a guarantee. Every independent study above found accuracy falling well before the maximum length, especially for indirect questions, multiple similar items, or reasoning across several distant facts.
  • Distractors matter more than length alone. Real document sets are full of near-duplicates: contract versions, repeated boilerplate, similar tickets. That is the condition under which models perform worst.
  • Put only what the task needs into the prompt. Anthropic's documentation and Chroma's results point the same way: curation beats volume. For large corpora, retrieval plus a long window usually beats pasting everything. See our explainer on context, memory and RAG.
  • Ask for quotes and locations. Requiring the model to cite the passage it relied on makes misses visible.
  • Check results on your own material before relying on them. Published benchmarks cover older models and synthetic or curated tasks. For Gemini's document limits specifically, see how Gemini handles a 900-page PDF.

About this storyBased on the sources linked below. Editorial standards

Was this useful?Report an error
Compare nextChatGPT vs Gemini for research: deep research, long documents, prices and privacy comparedOpen the living comparison LearnHow Gemini handles a 900-page PDF: the limits, the costs and what long-context benchmarks showExplainer · 6 minSeriesBenchmarks No. 01 — SWE-bench, Terminal-Bench and METR: what coding-agent evaluations measure, and what their results don't proveOct 9
Comments
0

More on Long context & models

The Week in AI

The numbers that matter, every Friday.

0