- The Gemini API accepts PDFs of up to 1,000 pages or 50 MB. Google counts each page as 258 tokens, so a 900-page document is roughly 232,000 tokens, well inside current models' 1M-token windows.
- In the Gemini app, only Google AI Pro and Ultra have a 1M-token context. The free tier (32K) and AI Plus (128K) can't hold the whole document at once, and every file is capped at 100 MB.
- Google's own documentation warns that recall drops when you ask for many facts at once. Independent studies, including one from May 2026 that covered Gemini 3.1 Pro Preview, found that performance falls as inputs grow.
Contents
Google says its current Gemini models can read about a million tokens at once. Google's own help pages put that at roughly 1,500 pages of text, so a 900-page manual, filing or book should fit with room to spare. Whether it actually works depends on three things: which Gemini you use (the API, the Gemini app or Gemini Notebook), how large the file is in megabytes, and what you ask the model to do with it.
Key terms
- Token: The unit models count text in. Google's document guide says each PDF page counts as 258 tokens.
- Context window: How many tokens a model can consider in one request. The current Gemini API models (3.8 Flash, 3.5 Flash-Lite and the 3.1 Pro preview) accept 1,048,576 input tokens and return up to 65,536 output tokens.
- Needle in a haystack: A test that hides one fact in a long text and asks the model to find it. Google's documentation notes this is "the most basic setup."
- Context caching: Storing a large input once and reusing it across many questions at a lower per-token price.
The limits, surface by surface
| Where | Size limit | Page or text limit | Context available | Notes |
|---|---|---|---|---|
| Gemini API, direct PDF | 50 MB per PDF | 1,000 pages | 1,048,576 tokens (3.8 Flash, 3.5 Flash-Lite, 3.1 Pro preview) | 258 tokens per page |
| Gemini API, Files API | 2 GB per file, 20 GB per project | PDF rules above still apply | Same | Free to use; files deleted after 48 hours |
| Gemini app, no plan | 100 MB per file, up to 10 files per prompt | Not stated | 32K tokens | |
| Gemini app, Google AI Plus | Same | Not stated | 128K tokens | |
| Gemini app, Google AI Pro / Ultra | Same | Google says "up to 1,500 pages of text" | 1M tokens | Deep Think: 192K |
| Gemini Notebook (formerly NotebookLM) | 200 MB per source | 500,000 words per source | Not published | 50 to 600 sources per notebook, depending on plan |
Sources: Google AI for Developers (document understanding, Files API, model pages) and Gemini Apps and Gemini Notebook help pages. Last verified: October 8, 2026
Do the math for a 900-page document: 900 × 258 = about 232,000 tokens. That is under a quarter of the API's context window, and under the 1,000-page cap. For many real documents, though, the limit that bites first is 50 MB, not pages. Scanned or image-heavy PDFs can pass 50 MB well before page 900. Splitting the file into two volumes fixes that, and you can send several PDFs in one request as long as the total fits the context window.
In the Gemini app, the tier matters more than the file. A 232,000-token document overflows the free plan's 32K window and AI Plus's 128K, so on those plans Gemini can't consider the whole document at once. Google doesn't say how the app handles the overflow. Only AI Pro ($19.99 a month in the US) and Ultra have the full 1M window. Our ChatGPT vs Gemini for research comparison has the per-plan details.
Gemini Notebook uses a different rule: up to 500,000 words per source. A typical dense page of text runs to a few hundred words, so a 900-page document may land close to that limit. That's our estimate, not Google's. If the notebook rejects or truncates the file, splitting it into a few sources is the documented fix.
How Gemini actually reads a PDF
Google's document guide explains what happens to each page:
- Pages are rendered as images. Large pages are scaled down to fit 3072 × 3072 pixels, and small ones are scaled up to 768 × 768, keeping the aspect ratio. That is why Gemini can read charts, tables and layout, not just text.
- On Gemini 3 models, text already embedded in the PDF is extracted and passed along too. Google says "you aren't charged" for those extracted text tokens. The page-image tokens show up in usage reports as image tokens.
- A
media_resolutionsetting (low, medium or high) changes how much detail each page gets on Gemini 3 models. - Only
application/pdfgets this visual treatment. Send the same document as plain text, Markdown or HTML and charts and layout are lost. - Google's tips: rotate pages to the right orientation and avoid blurry scans.
For large files, Google recommends uploading through the Files API once and referring to the file in later requests. Uploaded files are deleted automatically after 48 hours.
What it costs on the API
These are illustrative input-only estimates, using 232,000 tokens (900 × 258) and Google's published prices. Output and "thinking" tokens are billed on top.
| Model | Input price per 1M tokens | Approx. input cost to send the whole document once |
|---|---|---|
| Gemini 3.8 Flash | $0.75 until Dec 31, 2026; $1.50 from Jan 1, 2027 | about $0.17, rising to about $0.35 |
| Gemini 3.1 Pro (preview) | $2 up to 200K tokens; $4 above 200K | about $0.93 (the whole prompt is over 200K) |
Last verified: October 8, 2026 (ai.google.dev pricing)
Two caveats. First, if Google's no-charge rule for extracted text holds, the billed count could differ from 258 × pages; check the usage report after a real request. Second, every new question resends the whole document unless you use context caching, which Google recommends for repeated queries over the same large input. Asking 50 questions about one manual without caching means paying for the document 50 times.
What the long-context evidence says
Fitting a document in the window is not the same as reliably using every page of it. Google is open about this in its own long-context guide:
- Gemini does well on single-needle tests, but those are "the most basic setup."
- With multiple facts to retrieve, "the model does not perform with the same accuracy."
- Google suggests that if you need about 100 separate facts at roughly 99% accuracy, you may need about 100 separate requests.
- Put your question at the end, after the document.
Vendor results. For its newest model, Gemini 4 Argon, Google reports 84.2% on GraphWalks at 256,000 to 1 million tokens. That is a long-context reasoning test, and on Google's own table OpenAI's GPT-6 Astra scores 71.8% and Anthropic's Claude Opus 5.5 66.8%. These numbers are vendor-reported, and Argon isn't generally available yet; our Argon report has the details. Google's pages for Gemini 3.8 Flash, the model most people will actually use, publish no long-context scores.
Independent results. The most recent broad study is ATLAS, posted to arXiv on May 27, 2026 by researchers mostly at Meituan. It scored 26 models on a grid from 8,000 to 1 million tokens, using each provider's recommended settings. Gemini 3.1 Pro Preview (high effort) ranked first up to 128,000 tokens with an ATLAScore of 77.83, and second up to 1 million tokens with 68.52, behind Claude Opus 4.6. Its score over the full 1M range was about 9 points lower than over the shorter range. ATLAS predates Gemini 3.8 Flash and is English-only.
An older study shows why wording matters. NoLiMa, from researchers at Adobe Research and LMU Munich (ICML 2025), hides facts that share little wording with the question, so the model has to infer the link rather than match words. In the paper's latest version (July 2025), 11 of 13 models fell below half their short-input score at just 32,000 tokens. On the project's leaderboard, last updated July 17, 2025, Gemini 2.5 Flash (without thinking) dropped from 94.4% to 48.4%. NoLiMa covers models from 2025 or earlier, so treat it as evidence of a pattern, not a measure of today's Gemini.
Last verified: October 9, 2026 (arXiv and NoLiMa repository)
Our long-context benchmarks report covers these studies and the vendors' own claims in more detail.
Practical advice for a 900-page document
- Pick the right surface. For repeated, precise questions, use the API with the Files API and context caching. For reading and exploring, Gemini Notebook is built around working with a fixed set of sources. In the Gemini app, you need AI Pro or Ultra to fit the whole document.
- Check the megabytes first. If the file is over 50 MB (API) or 100 MB (app), split it into volumes along chapter boundaries.
- Ask for one thing at a time. Google's own guidance is that accuracy falls as you ask for more facts per request. Break "list every deadline in the contract" into section-by-section questions.
- Demand page numbers, then check them. Ask for the page and a short quote for each claim, and spot-check them.
- Use retrieval for very large or changing collections. Past a few documents, a retrieval setup (RAG) can be cheaper and easier to audit than resending everything; see context windows vs memory vs RAG.
- Put the question last, and keep extra text out of the prompt.
About this storyBased on the sources linked below. Editorial standards




