Explainer

How Gemini handles a 900-page PDF: the limits, the costs and what long-context benchmarks show

On paper, a 900-page document fits comfortably in Gemini's 1M-token context. Whether it works depends on file size, your plan and what you ask.

By ShajanthanUpdated 6 min read
ByShajanthanFounder & Editor
Published
Reading6 MIN
Illustration comparing a 900-page document stack with a token gauge showing it fills about a quarter of a one-million-token context window
In 20 seconds
  1. The Gemini API accepts PDFs of up to 1,000 pages or 50 MB. Google counts each page as 258 tokens, so a 900-page document is roughly 232,000 tokens, well inside current models' 1M-token windows.
  2. In the Gemini app, only Google AI Pro and Ultra have a 1M-token context. The free tier (32K) and AI Plus (128K) can't hold the whole document at once, and every file is capped at 100 MB.
  3. Google's own documentation warns that recall drops when you ask for many facts at once. Independent studies, including one from May 2026 that covered Gemini 3.1 Pro Preview, found that performance falls as inputs grow.
Contents

Google says its current Gemini models can read about a million tokens at once. Google's own help pages put that at roughly 1,500 pages of text, so a 900-page manual, filing or book should fit with room to spare. Whether it actually works depends on three things: which Gemini you use (the API, the Gemini app or Gemini Notebook), how large the file is in megabytes, and what you ask the model to do with it.

Key terms

  • Token: The unit models count text in. Google's document guide says each PDF page counts as 258 tokens.
  • Context window: How many tokens a model can consider in one request. The current Gemini API models (3.8 Flash, 3.5 Flash-Lite and the 3.1 Pro preview) accept 1,048,576 input tokens and return up to 65,536 output tokens.
  • Needle in a haystack: A test that hides one fact in a long text and asks the model to find it. Google's documentation notes this is "the most basic setup."
  • Context caching: Storing a large input once and reusing it across many questions at a lower per-token price.

The limits, surface by surface

WhereSize limitPage or text limitContext availableNotes
Gemini API, direct PDF50 MB per PDF1,000 pages1,048,576 tokens (3.8 Flash, 3.5 Flash-Lite, 3.1 Pro preview)258 tokens per page
Gemini API, Files API2 GB per file, 20 GB per projectPDF rules above still applySameFree to use; files deleted after 48 hours
Gemini app, no plan100 MB per file, up to 10 files per promptNot stated32K tokens
Gemini app, Google AI PlusSameNot stated128K tokens
Gemini app, Google AI Pro / UltraSameGoogle says "up to 1,500 pages of text"1M tokensDeep Think: 192K
Gemini Notebook (formerly NotebookLM)200 MB per source500,000 words per sourceNot published50 to 600 sources per notebook, depending on plan

Sources: Google AI for Developers (document understanding, Files API, model pages) and Gemini Apps and Gemini Notebook help pages. Last verified: October 8, 2026

Do the math for a 900-page document: 900 × 258 = about 232,000 tokens. That is under a quarter of the API's context window, and under the 1,000-page cap. For many real documents, though, the limit that bites first is 50 MB, not pages. Scanned or image-heavy PDFs can pass 50 MB well before page 900. Splitting the file into two volumes fixes that, and you can send several PDFs in one request as long as the total fits the context window.

In the Gemini app, the tier matters more than the file. A 232,000-token document overflows the free plan's 32K window and AI Plus's 128K, so on those plans Gemini can't consider the whole document at once. Google doesn't say how the app handles the overflow. Only AI Pro ($19.99 a month in the US) and Ultra have the full 1M window. Our ChatGPT vs Gemini for research comparison has the per-plan details.

Gemini Notebook uses a different rule: up to 500,000 words per source. A typical dense page of text runs to a few hundred words, so a 900-page document may land close to that limit. That's our estimate, not Google's. If the notebook rejects or truncates the file, splitting it into a few sources is the documented fix.

How Gemini actually reads a PDF

Google's document guide explains what happens to each page:

  • Pages are rendered as images. Large pages are scaled down to fit 3072 × 3072 pixels, and small ones are scaled up to 768 × 768, keeping the aspect ratio. That is why Gemini can read charts, tables and layout, not just text.
  • On Gemini 3 models, text already embedded in the PDF is extracted and passed along too. Google says "you aren't charged" for those extracted text tokens. The page-image tokens show up in usage reports as image tokens.
  • A media_resolution setting (low, medium or high) changes how much detail each page gets on Gemini 3 models.
  • Only application/pdf gets this visual treatment. Send the same document as plain text, Markdown or HTML and charts and layout are lost.
  • Google's tips: rotate pages to the right orientation and avoid blurry scans.

For large files, Google recommends uploading through the Files API once and referring to the file in later requests. Uploaded files are deleted automatically after 48 hours.

What it costs on the API

These are illustrative input-only estimates, using 232,000 tokens (900 × 258) and Google's published prices. Output and "thinking" tokens are billed on top.

ModelInput price per 1M tokensApprox. input cost to send the whole document once
Gemini 3.8 Flash$0.75 until Dec 31, 2026; $1.50 from Jan 1, 2027about $0.17, rising to about $0.35
Gemini 3.1 Pro (preview)$2 up to 200K tokens; $4 above 200Kabout $0.93 (the whole prompt is over 200K)

Last verified: October 8, 2026 (ai.google.dev pricing)

Two caveats. First, if Google's no-charge rule for extracted text holds, the billed count could differ from 258 × pages; check the usage report after a real request. Second, every new question resends the whole document unless you use context caching, which Google recommends for repeated queries over the same large input. Asking 50 questions about one manual without caching means paying for the document 50 times.

What the long-context evidence says

Fitting a document in the window is not the same as reliably using every page of it. Google is open about this in its own long-context guide:

  • Gemini does well on single-needle tests, but those are "the most basic setup."
  • With multiple facts to retrieve, "the model does not perform with the same accuracy."
  • Google suggests that if you need about 100 separate facts at roughly 99% accuracy, you may need about 100 separate requests.
  • Put your question at the end, after the document.

Vendor results. For its newest model, Gemini 4 Argon, Google reports 84.2% on GraphWalks at 256,000 to 1 million tokens. That is a long-context reasoning test, and on Google's own table OpenAI's GPT-6 Astra scores 71.8% and Anthropic's Claude Opus 5.5 66.8%. These numbers are vendor-reported, and Argon isn't generally available yet; our Argon report has the details. Google's pages for Gemini 3.8 Flash, the model most people will actually use, publish no long-context scores.

Independent results. The most recent broad study is ATLAS, posted to arXiv on May 27, 2026 by researchers mostly at Meituan. It scored 26 models on a grid from 8,000 to 1 million tokens, using each provider's recommended settings. Gemini 3.1 Pro Preview (high effort) ranked first up to 128,000 tokens with an ATLAScore of 77.83, and second up to 1 million tokens with 68.52, behind Claude Opus 4.6. Its score over the full 1M range was about 9 points lower than over the shorter range. ATLAS predates Gemini 3.8 Flash and is English-only.

An older study shows why wording matters. NoLiMa, from researchers at Adobe Research and LMU Munich (ICML 2025), hides facts that share little wording with the question, so the model has to infer the link rather than match words. In the paper's latest version (July 2025), 11 of 13 models fell below half their short-input score at just 32,000 tokens. On the project's leaderboard, last updated July 17, 2025, Gemini 2.5 Flash (without thinking) dropped from 94.4% to 48.4%. NoLiMa covers models from 2025 or earlier, so treat it as evidence of a pattern, not a measure of today's Gemini.

Last verified: October 9, 2026 (arXiv and NoLiMa repository)

Our long-context benchmarks report covers these studies and the vendors' own claims in more detail.

Practical advice for a 900-page document

  1. Pick the right surface. For repeated, precise questions, use the API with the Files API and context caching. For reading and exploring, Gemini Notebook is built around working with a fixed set of sources. In the Gemini app, you need AI Pro or Ultra to fit the whole document.
  2. Check the megabytes first. If the file is over 50 MB (API) or 100 MB (app), split it into volumes along chapter boundaries.
  3. Ask for one thing at a time. Google's own guidance is that accuracy falls as you ask for more facts per request. Break "list every deadline in the contract" into section-by-section questions.
  4. Demand page numbers, then check them. Ask for the page and a short quote for each claim, and spot-check them.
  5. Use retrieval for very large or changing collections. Past a few documents, a retrieval setup (RAG) can be cheaper and easier to audit than resending everything; see context windows vs memory vs RAG.
  6. Put the question last, and keep extra text out of the prompt.

About this storyBased on the sources linked below. Editorial standards

Was this useful?Report an error
Comments
0

More on Google Gemini & apps

The Week in AI

New guides and explainers, every Friday.

0