Explainer

Context windows vs memory vs RAG: what each one actually does

A 1M-token window, a chatbot that "remembers you" and a retrieval pipeline solve different problems. Here's how each works and when to use which.

By ShajanthanUpdated 7 min read
ByShajanthanFounder & Editor
Published
Reading7 MIN
Diagram showing context window, memory and retrieval feeding into a single model prompt
In 20 seconds
  1. The context window is what a model can see in one request; memory is what a product or agent stores and re-injects across sessions; retrieval (RAG) is a system that picks what goes into the window.
  2. ChatGPT, Claude and Gemini all now offer memory, and Anthropic and OpenAI give developers file-based memory for agents. Under the hood, all of it ends up as text placed back in the context window.
  3. Use the window for one bounded body of material, retrieval for large or changing corpora, memory for durable preferences and lessons, and compaction to keep long agent runs inside the window.
Contents

"It has a million-token context," "it remembers you," "it's grounded in your documents": these are three different claims about three different mechanisms, and they're often blurred together. This explainer on long context defines each one precisely, shows where it lives in today's ChatGPT, Claude and Gemini apps and in the Anthropic and OpenAI APIs, and ends with a table for deciding which to use.

The short version: a model only ever "knows" two things, what it learned in training and what is in its context window for this request. Memory and retrieval are both ways of deciding what text to put into that window.

The context window: what the model sees right now

Anthropic's documentation defines the context window as "all the text a language model can reference when generating a response, including the response itself," and calls it the model's "working memory." OpenAI's definition is "the maximum number of tokens that can be used in a single request," which counts input, output and reasoning tokens.

Current flagship APIs accept about 1 million tokens. Claude Opus 5.5, Sonnet 5.5 and Haiku 5.5 take 1M by default with no beta header, OpenAI's GPT-6 models list 1,050,000, and Google's Gemini 3.8 Flash and 3.1 Pro Preview list 1,048,576 input tokens.

Three properties matter more than the headline size:

  • It is stateless. The API forgets everything between requests unless you send it again. Chat apps feel continuous only because they resend the history.
  • It costs money every time. Every token in the window is billed on every call, unless you use prompt caching. Some providers also raise rates above a threshold. OpenAI doubles input rates above 272K tokens, and Google's Gemini 3.1 Pro Preview does the same above 200K.
  • Fuller isn't better. Anthropic's own docs warn that "as token count grows, accuracy and recall degrade, a phenomenon known as context rot." Research on whether a fact's position in the window matters is mixed. The 2023 "Lost in the Middle" paper found performance often highest when relevant information sat at the beginning or end, while Chroma's July 2025 report found no notable variation across 11 needle positions on its basic needle task. Both found that more irrelevant text makes things worse. We explain how this is measured in Long-context recall, explained.

Memory: what persists between sessions

"Memory" means information saved outside the model and fed back into later conversations, as text in the context window. Nothing about the model's weights changes.

In the consumer apps

Last verified: October 8, 2026

AppWhat it storesNotable controls
ChatGPT"Saved memories" (things you ask it to keep, or that it judges useful) and "reference chat history" (relevant past conversations)Settings › Personalization › Memory. Temporary chats "do not create or update memories." Turning off chat-history reference schedules derived information for deletion within 30 days. OpenAI says features vary "by plan, region, platform, and workspace settings."
ClaudeMemory saved "as a set of individual topics as you chat"; each project has its own separate memoryOn by default for Free, Pro and Max. Team and Enterprise owners decide whether it's available. Incognito chats are excluded. Searching past chats requires a paid plan. You can pause or reset memory and edit topics in Settings › Memory.
Gemini"Personal context" (saved info you add or ask it to remember), past-chat referencing, and Personal Intelligence, which connects Gmail, Photos, YouTube and SearchConnecting apps is off by default. Saved info isn't available to users under 18. Google says past-chat referencing is "gradually releasing" on mobile.

Some history matters here, because availability keeps shifting. Anthropic opened Claude's memory to free users in March 2026, per TechRadar. Google launched Personal Intelligence on January 14, 2026 for US AI Pro and Ultra subscribers. Android Authority reported that it expanded to paid users worldwide on April 14, 2026, except in the European Economic Area, Switzerland and the UK. Google brought past-chat memory and chat-history import to the UK on April 29, 2026.

App memory is convenient, but it also changes behavior in ways that are hard to see. An old preference or a wrong inference can quietly shape answers. All three apps let you view and delete what is stored, and they offer a temporary or incognito mode for conversations you don't want remembered.

In the APIs: memory for agents

For developers, "memory" is explicit, file-like storage that the model reads and writes through tools:

  • Anthropic's memory tool (memory_20250818, no beta header, Claude 4 and later) is client-side. Claude asks to view, create, edit, rename or delete files under a /memories directory, and your code decides where those files actually live. Anthropic's docs tell implementers to "validate every path in every command to prevent directory traversal attacks."
  • Claude Managed Agents memory stores (beta header agent-memory-2026-07-22) are hosted, workspace-scoped collections of text documents. They're mounted into an agent session as a directory under /mnt/memory/. A session can attach up to 8 stores. Each memory is capped at 100 kB, and each store at 10,000 memories. Every change creates an immutable version. Versions are kept for 30 days, although the recent versions of a live memory are always kept. Anthropic is unusually direct about the risk: if an agent reads untrusted content, "a successful prompt injection could write malicious content into the store. Later sessions then read that content as trusted memory." It recommends read-only mounts wherever possible.
  • OpenAI doesn't list a dedicated memory tool in its conversation-state guide. State is carried by resending history, by chaining with previous_response_id (stored responses are kept for 30 days by default), or by the Conversations API, whose objects aren't subject to that 30-day limit. In the Agents SDK, Sessions persist conversation history in SQLite, Redis, SQLAlchemy databases or OpenAI's Conversations API. Separately, a sandbox-agent Agent memory feature distills lessons from earlier runs into a MEMORY.md file and a short summary that is injected at the start of the next run.

If you're building your first agent, our Claude Agent SDK guide shows where session state and memory fit.

Retrieval (RAG): choosing what goes into the window

Retrieval-augmented generation, named in a 2020 paper by Patrick Lewis and colleagues, adds a search step before the model answers. A typical pipeline:

  1. Chunk documents into passages. OpenAI's API, for example, lets you set max_chunk_size_tokens and chunk_overlap_tokens per file.
  2. Index each chunk, usually as an embedding vector and often also in a keyword index such as BM25.
  3. Search for the user's question. OpenAI's file search combines "semantic and keyword search," with tunable reciprocal-rank-fusion weights between the two.
  4. Rerank the candidates with a stronger model and keep the top few. OpenAI returns 10 results by default, up to 50.
  5. Generate an answer from those passages, ideally with citations.

Retrieval quality is where these systems succeed or fail. A chunk that says "revenue grew 3%" is useless if it doesn't say which company or which year. Anthropic's 2024 "contextual retrieval" technique prepends a short model-written description to each chunk before indexing. Anthropic reported that this cut failed retrievals by 35%, by 49% when combined with contextual BM25, and by 67% with reranking added. These are vendor-reported numbers from 2024. In the same post, Anthropic said a knowledge base under about 200,000 tokens can simply go in the prompt, with caching, and skip RAG entirely. With 1M-token windows that threshold is higher in principle, but the context-rot caveat still applies.

Hosted options now include OpenAI's file search, which gives the first 1 GB of vector storage free and then charges $0.10 per GB per day. Connectors and tool protocols such as MCP let a model query live systems instead of a static index.

Keeping long sessions inside the window: compaction and context editing

Agents that run for hours fill even a 1M window. Two techniques manage this within a session:

  • Compaction replaces older turns with a model-written summary. Anthropic offers server-side compaction at a token threshold, and since September 14, 2026 an on-demand mode in beta (compact-2026-09-04). OpenAI's Responses API offers context_management with a compact_threshold, plus a standalone /responses/compact endpoint.
  • Context editing removes specific content. Anthropic's clear_tool_uses_20250919 strategy clears old tool results. By default it starts at 100,000 input tokens and keeps the last 3 tool uses. A separate strategy manages thinking blocks. Clearing content invalidates the prompt cache from that point, so it has a cost.

Anthropic suggests pairing these with memory. Compaction keeps the active context small, and the memory tool keeps anything that must survive summarization.

Trade-offs at a glance

Long context windowMemoryRetrieval (RAG)
CostPay for every token on every call; caching helps; surcharges above 200K–272K on some APIsSmall: only stored notes are re-injectedIndexing and storage costs; small prompts per query
LatencyGrows with input sizeLowSearch step adds time; small prompts are fast
RecallDegrades with length and distractors ("context rot")Only as good as what was savedFails silently if the right chunk isn't retrieved
FreshnessWhatever you send this callCan go stale; needs pruningAs fresh as the index or live connector
PrivacyData sent per requestPersistent personal data; needs user controlsCorpus stored in an index; access control per document
Main riskPaying for noisePoisoned or stale memoriesBad chunking and ranking

Which one to use

SituationBest fit
One contract, codebase or report, under a few hundred thousand tokens, asked many questionsLong context with prompt caching
A knowledge base that is large, growing or changes dailyRetrieval (hybrid search and reranking)
User preferences, project conventions, lessons from earlier runsMemory (app memory, memory tool or memory stores)
A long-running agent taskCompaction or context editing, plus memory for anything that must survive
Questions that need exact quotes from huge corporaRetrieval to narrow down, then a long window over the top results
Sensitive data you don't want persistedLong context or temporary/incognito chats, not memory

These approaches combine in practice. A coding agent might retrieve relevant files, load them into a large window, compact old tool output, and write lessons to memory for next time. For what the window-only approach means for one very large document, see how Gemini handles a 900-page PDF.

About this storyBased on the sources linked below. Editorial standards

Was this useful?Report an error
Comments
0

More on Long context & research

The Week in AI

New guides and explainers, every Friday.

0