- Reasoning models spend extra tokens thinking before they answer. Anthropic, OpenAI and Google all bill those tokens at the output rate, even when you can't see them.
- The controls have moved from token budgets to effort levels. On Claude Opus 5.5 you can't switch thinking off at all, so effort and model choice are your cost levers.
- Use high effort for hard, checkable, multi-step work. Use low effort or a small model for lookups, classification and formatting.
Contents
Every flagship model from Anthropic, OpenAI and Google is now a reasoning model. Before it answers, it generates a stream of intermediate "thinking" tokens, and you pay for them at the output-token rate whether or not you ever see them. That changes how you should budget, configure and route requests. It matters more now that the newest models keep thinking switched on by default. That includes Claude Opus 5.5, where it can't be switched off at all, and OpenAI's GPT-6 family (see our GPT-6 launch story).
This explainer covers how reasoning works, how each vendor exposes and bills it as of October 8, 2026, a worked cost example using official list prices, and when the extra spend is worth it.
What a reasoning model actually does
Three ideas sit underneath the marketing.
Chain of thought. In 2022, Google researchers showed that getting a model to write out intermediate steps "significantly improves the ability of large language models to perform complex reasoning" (Wei et al.). Early on this was a prompting trick ("let's think step by step"). Reasoning models build the behavior into training.
Reinforcement learning on verifiable tasks. Labs train models with reinforcement learning (RL) on problems with checkable answers, such as math, code and STEM questions, and reward correct outcomes. When OpenAI introduced o1 on September 12, 2024, it said performance "consistently improves with more reinforcement learning (train-time compute)" and also with more time spent thinking at inference (OpenAI). DeepSeek's R1 paper argued that reasoning "can be incentivized through pure reinforcement learning", with self-reflection and verification behaviors emerging during training (DeepSeek-AI).
Test-time compute. Thinking longer on a hard problem is a way to buy accuracy with inference compute instead of a bigger model. Snell et al. (2024) found that allocating test-time compute adaptively per prompt was more than 4x more efficient than a best-of-N baseline. In a compute-matched comparison it let a smaller model beat one 14x larger on problems where the small model already had some success (arXiv). That caveat matters: extra thinking helps most when the model is already partly capable, and does little for knowledge it simply lacks.
How the big three expose and bill thinking
Last verified: October 8, 2026
| Anthropic (Claude) | OpenAI (GPT-6 family) | Google (Gemini 3.x) | |
|---|---|---|---|
| Main control | output_config.effort: low, medium, high, xhigh, max | reasoning.effort: none → max (varies by model) | thinking_level (values vary by model) |
| Thinking mode | Adaptive: the model decides whether and how much to think | Always reasons on Astra and 6.1 Sol; Luna supports none | Per-model default level |
| Can you see it? | display: "summarized" or "omitted" (omitted is the default on 5.x models) | Summaries only, via summary | Summaries via thinking_summaries |
| Billing | Output tokens, full thinking billed | Output tokens | Output price "including thinking tokens" |
| Usage field | output_tokens_details.thinking_tokens | output_tokens_details.reasoning_tokens | total_thought_tokens |
Anthropic: adaptive thinking and effort
Anthropic has moved its current models from fixed budgets to adaptive thinking: "the model evaluates each request and determines whether to think and how much" (docs). You steer it with the effort parameter. Anthropic's docs describe low as "most efficient" with "some capability reduction", and max as "absolute maximum capability with no constraints on token spending" (effort docs). At lower effort, Claude "may skip thinking entirely on easy inputs" (docs).
The details that break code:
- Manual budgets are gone on new models.
thinking: {type: "enabled", budget_tokens: N}returns a 400 error on Claude 4.7 and later, including Opus 5.5, Sonnet 5.5, Sonnet 5 and Haiku 5.5 (thinking overview). - Opus 5.5 thinking cannot be disabled. Opus 5.5, Fable 5/5.1 and Mythos 5/5.1 reject
type: "disabled". "Thinking can't be turned off on these models," the docs say. The release notes for Opus 5.5's September 22, 2026 launch tell developers to "omit thethinkingfield and control thinking depth with the effort parameter" (release notes). - Other models differ. Sonnet 5.5 rejects
"disabled"but acceptstype: "between_tools"athigheffort or below, which turns off up-front thinking. Haiku 5.5 and Opus 5 accept"disabled"only athighor below. Sonnet 5 accepts it outright (thinking overview). - Defaults vary. Effort defaults to
mediumon Opus 5.5 and Haiku 5.5, and tohighon most other models, including Sonnet 5.5 (effort docs). - Hiding thinking doesn't save money. "You're still charged for the full thinking tokens. Omitting reduces latency, not cost." With summarized display, you're billed for the full thinking, not the summary, so "the billed output token count does not match the count of tokens you see" (thinking overview).
- Thinking carries over. On Opus 4.5 and models numbered 4.6 and higher, earlier turns' thinking blocks stay in context and are billed as input on later turns (docs). Long agent sessions pay for old reasoning repeatedly, which makes prompt caching more valuable.
OpenAI: reasoning effort
OpenAI's reasoning guide lists effort values that "can include none, minimal, low, medium, high, xhigh, and max", with support varying by model (OpenAI). On the current API models page, GPT-6 Astra and GPT-6.1 Sol support low through max, and GPT-6 Luna supports none through max (models). GPT-6.1 Sol defaults to medium.
Reasoning tokens "are not visible via the API". They "occupy space in the model's context window and are billed as output tokens". OpenAI recommends "reserving at least 25,000 tokens for reasoning and outputs" when you start. If you set max_output_tokens too low, the response can come back incomplete before any visible text, and you still pay for the reasoning (OpenAI).
A naming note: OpenAI's API models overview now highlights GPT-6.1 Sol (gpt-6.1-sol), an upgrade shipped in late September 2026 to the GPT-6 Sol released on September 22. The worked example below uses GPT-6.1 Sol (OpenAI).
Google: thinking levels
Gemini 3.x models use thinking_level, and each model has its own set of values and default. Per Google's docs, gemini-3.8-flash offers low, medium and high, defaulting to medium. gemini-3.1-pro-preview defaults to high, and gemini-3.5-flash-lite defaults to minimal (Google). "When thinking is turned on, response pricing is the sum of output tokens and thinking tokens." Google's pricing page labels every output price "including thinking tokens" (pricing). Google's own advice: to cut cost or latency, lower thinking_level rather than shrinking max_output_tokens, which can truncate the answer and still bill the thinking.
A worked cost example
Take one request with 5,000 input tokens and an 800-token visible answer. Then vary only the hidden thinking: none, 4,000 tokens, or 20,000 tokens. The thinking counts are illustrative, not measured: real usage depends on the task, the effort level and the model. Tokenizers also differ between vendors. Prices are standard-tier list prices with no caching or batch discounts.
Last verified: October 8, 2026
| Model | Input / output per 1M tokens | No thinking | 4K thinking | 20K thinking | 20K thinking × 100K requests |
|---|---|---|---|---|---|
| Claude Opus 5.5 | $4 / $20 | $0.036 | $0.116 | $0.436 | $43,600 |
| Claude Sonnet 5.5 | $2 / $10 | $0.018 | $0.058 | $0.218 | $21,800 |
| Claude Haiku 5.5 (≤100K prompt) | $0.10 / $0.50 | $0.0009 | $0.0029 | $0.0109 | $1,090 |
| GPT-6 Astra | $10 / $50 | $0.090 | $0.290 | $1.090 | $109,000 |
| GPT-6.1 Sol | $2 / $10 | $0.018 | $0.058 | $0.218 | $21,800 |
| GPT-6 Luna | $0.10 / $0.50 | $0.0009 | $0.0029 | $0.0109 | $1,090 |
| Gemini 3.8 Flash (to Dec 31, 2026) | $0.75 / $3.75 | $0.0068 | $0.0218 | $0.0818 | $8,175 |
Prices: Anthropic, OpenAI, Google. Gemini 3.8 Flash rises to $1.50/$7.50 on January 1, 2027.
Three takeaways:
- Thinking can dominate the bill. At 20,000 thinking tokens, more than 90% of the Opus 5.5 request cost is thinking. The input barely registers.
- Effort and model choice are worth more than any price cut. Opus 5.5 at 20,000 thinking tokens costs about 12 times as much as the same request answered without thinking. A small model at high effort can still cost less than a frontier model at low effort. That's the argument for routing requests across cheap, mid and frontier models.
- Agents multiply everything. An agent loop makes many calls, and on Opus 4.5 and Claude 4.6+ models earlier thinking is re-billed as input on each later turn. Per-call estimates understate agent costs. We cover this in why token pricing isn't the story.
When to use reasoning, and when not to
Turn effort up for:
- Multi-step problems with a checkable answer: math, algorithmic code, debugging, data reconciliation.
- Planning in long agentic tasks. Anthropic's effort docs point to
xhighfor long-running agentic and coding work. - Cases where a wrong answer costs more than the tokens: contract review, migrations, security analysis.
Turn it down, or use a small model, for:
- Lookups, classification, extraction, reformatting and short rewrites. Google's docs use "Where was DeepMind founded?" as the example of a low-thinking task.
- Latency-sensitive user interfaces. Thinking adds time to first visible token.
- Subagents doing narrow, well-specified steps. Anthropic's docs suggest
lowfor simpler tasks and subagents where speed and cost matter most. - Tasks bottlenecked on knowledge, not reasoning. More thinking can't recall facts the model doesn't have.
A practical recipe:
- Start at the vendor default effort, and log the thinking or reasoning token field on every call.
- Build a small evaluation set from your own tasks. Drop effort one level and measure accuracy and cost.
- Raise
max_tokensormax_output_tokenswhen you raise effort, so answers aren't cut off. - Keep thinking settings stable within a conversation. On Claude, changing thinking configuration invalidates prompt-cache breakpoints.
- Pick the model tier before tuning effort. Our Claude model guide covers the current lineup.
Don't treat the visible reasoning as an audit trail
Even when you can read a model's reasoning, it may not reflect what actually drove the answer. In an April 3, 2025 study, Anthropic gave models hints that influenced their answers. Claude 3.7 Sonnet mentioned the hint only 25% of the time and DeepSeek R1 39% of the time. Unfaithful explanations were often longer than faithful ones (Anthropic). On current Claude models the default display is "omitted", and OpenAI returns only summaries. For builders, the reliable check is external verification: tests, validators and type checkers, not the model's account of itself. OpenAI's September 8, 2026 Navier–Stokes result, explained here, makes the same point at a much larger scale.
About this storyBased on the sources linked below. Editorial standards




