BusinessAnalysis

Token pricing isn't the story: what the official price pages actually tell you

Anthropic and OpenAI now charge identical sticker prices at two tiers. Your bill is decided by tokenizers, thinking tokens, caching, long-context surcharges and tool fees.

ByShajanthanFounder & Editor
Published
Reading8 MIN
Illustration of an AI model price tag with hidden cost multipliers beneath it
What’s new, in 20 seconds
  1. From the official pricing pages: Claude Sonnet 5.5 and GPT-6.1 Sol both cost $2/$10 per million tokens, Claude Haiku 5.5 and GPT-6 Luna both cost $0.10/$0.50, and Claude Fable 5.1 and GPT-6 Astra both cost $10/$50.
  2. The real differences are the multipliers. Anthropic's newer models produce about 30% more tokens for the same text, thinking tokens are billed as output, and long prompts can double the input price.
  3. In our illustrative examples, the same 400K-token request costs $1.72 to $8.60 across flagship and mid-tier models, and prompt caching cuts a 20-turn agent session by roughly 70%.
Contents

As of October 8, 2026, the two biggest frontier model vendors have arrived at nearly the same sticker prices. Anthropic's Claude Sonnet 5.5 and OpenAI's GPT-6.1 Sol both list at $2 per million input tokens and $10 per million output tokens. Claude Haiku 5.5 and GPT-6 Luna both list at $0.10 and $0.50. Claude Fable 5.1 and GPT-6 Astra both list at $10 and $50. If per-token price decided your bill, choosing between them would be a coin toss.

It doesn't. The multipliers around the sticker price decide what you actually pay: tokenizers, reasoning tokens, cache discounts, long-context surcharges, batch discounts and tool fees. They differ a lot between vendors. We pulled every number below from the vendors' own pricing pages and documentation, then worked through what they mean for a real bill.

The price table

US dollars per million tokens (MTok), standard tier, from each vendor's official pricing page.

VendorModelTierInputCached inputOutputLong-context rule
AnthropicClaude Fable 5.1Top$10$0.25$50Flat to 1M
AnthropicClaude Opus 5.5Flagship (default)$4$0.20$20Flat to 1M
AnthropicClaude Sonnet 5.5Mid$2$0.10$10Flat to 1M
AnthropicClaude Haiku 5.5Small$0.10$0.01$0.50>100K prompt: $0.50 / $2.50
OpenAIGPT-6 AstraFlagship$10$1.00$50>272K: $20 / $75
OpenAIGPT-6.1 SolMid$2$0.10$10>272K: $4 / $15
OpenAIGPT-6 LunaSmall$0.10$0.01$0.50>272K: $0.20 / $0.75
GoogleGemini 3.1 Pro (Preview)Flagship (API)$2$0.20 + storage$12>200K: $4 / $18
GoogleGemini 3.8 FlashMid$0.75$0.075 + storage$3.75none listed
GoogleGemini 3.1 Flash-LiteSmall$0.25$0.025 + storage$1.50none listed
xAIGrok 4.7Flagship$2$0.50$6≥200K: $4 / $12

Last verified: October 8, 2026. Full rows with source URLs are in our data file. A few notes on the table:

  • OpenAI's marketing page at openai.com/api/pricing still showed the older GPT-5.6 lineup when we checked. The figures above come from OpenAI's developer pricing page and each model's page.
  • GPT-6 Sol, released before GPT-6.1 Sol, is still available in the API at $2 input, $0.20 cached input and $10 output. Its model page points to GPT-6.1 Sol as "the newer Sol model", so the table lists 6.1. (Last verified: October 9, 2026.)
  • Google's newest model, Gemini 4 Argon, is in limited release and doesn't appear on the Gemini API pricing page. Gemini 3.1 Pro is still labeled "Preview".
  • Gemini 3.8 Flash's prices are listed "through December 31, 2026" and double on January 1, 2027.
  • Mistral's pricing page quotes a Mistral Large price without saying which version, so we left it out.

For the full Claude lineup, including batch and cache-write prices, see our guide to which Claude model to use.

Multiplier 1: the tokenizer

A "token" isn't a fixed unit. Each vendor splits text with its own tokenizer, and the same paragraph can become a different number of tokens. Anthropic is unusually direct about this. Its pricing page says Claude Opus 4.7 and later models use a newer tokenizer that "produces approximately 30% more tokens for the same text", with per-token prices unchanged.

Worked example (illustrative). Take a monthly job of 10 million input tokens and 2 million output tokens, as counted by the older tokenizer.

  • On Claude Sonnet 4.6 ($3/$15): 10 × $3 + 2 × $15 = $60.00.
  • The same text on Sonnet 5.5 ($2/$10) at 30% more tokens: 13 × $2 + 2.6 × $10 = $52.00.

The sticker price fell by a third, but the bill fell 13%. Anthropic also says its newer models finish tasks with fewer tokens, which could offset this. That depends on the workload, so measure it rather than assume it.

Vendors don't publish cross-vendor tokenizer comparisons. Don't assume a "token" from OpenAI and one from Anthropic cover the same amount of your text. Anthropic's token-counting endpoint is free, and counting a representative sample of your own prompts is the only reliable check.

Multiplier 2: thinking tokens are output tokens

Reasoning models "think" before they answer, and you pay for the thinking. OpenAI's reasoning guide says reasoning tokens "are billed as output tokens" and aren't visible through the API. Google labels Gemini output prices "including thinking tokens". On Anthropic's side, Opus 5.5 and Fable 5.1 can't turn thinking off, and the effort setting applies to all output tokens, thinking included.

Output tokens cost five times as much as input on most of these models (six times on Gemini 3.1 Pro and Flash-Lite, three times on Grok 4.7), so thinking dominates the bill on short tasks.

Worked example (illustrative; the thinking volume is our assumption). A classifier on Sonnet 5.5 sees 1,000 input tokens and returns a 50-token label, one million times a month.

  • No thinking: 1M × (1,000 × $2 + 50 × $10) ÷ 1M = $2,500.
  • If each request also produces 2,000 thinking tokens: 1M × (1,000 × $2 + 2,050 × $10) ÷ 1M = $22,500.

That's nine times the cost for the same visible output. The fix is cheap: use the lowest effort that keeps quality acceptable, and send simple jobs to a small model. Our explainer on reasoning models and thinking tokens covers the mechanics.

Multiplier 3: caching works differently at each vendor

Agents and chat apps resend the same long prefix (system prompt, tools, conversation history) on every turn. Caching makes repeated prefixes cheap, but the terms vary:

  • Anthropic charges a premium to write the cache (1.25x base input for 5 minutes, 2x for 1 hour). Reads cost 0.05x base on Opus 5.5 and Sonnet 5.5, 0.025x on Fable 5.1, and 0.1x on Haiku 5.5.
  • OpenAI now lists cache-write prices too (for example, $12.50 on GPT-6 Astra, or 1.25x). Cached input costs 0.1x on Astra and Luna. On GPT-6.1 Sol it costs 0.05x, which OpenAI describes as a 95% discount.
  • Google charges $0.20 per million cached tokens on Gemini 3.1 Pro, plus $4.50 per million tokens per hour of storage.
  • xAI charges $0.50 for cached input on Grok 4.7, or a quarter of the base input price.

Worked example (illustrative; simplified). An agent session runs 20 turns over a 100K-token prefix. The first turn writes the cache. Each later turn reads 90K cached tokens, sends 10K fresh tokens, and outputs 2K tokens. For simplicity this ignores incremental cache writes.

ModelNo cachingWith caching
Claude Opus 5.5$8.80$2.40
Claude Sonnet 5.5$4.40$1.20
GPT-6.1 Sol$4.40$1.20
GPT-6 Astra$22.00$6.86
Gemini 3.1 Pro$4.48$1.40, plus about $0.45 to store 100K tokens for an hour

Caching cuts these sessions by about 70%, which is more than the gap between most competing sticker prices. A cache miss is expensive, though. If you change the system prompt, tool list or earlier messages, the cache can break.

Multiplier 4: long prompts can double the rate

Every vendor in the table advertises context windows of around a million tokens. They charge for using them differently:

  • Anthropic prices the full 1M window at standard rates on Claude 4.6 and later models, except Haiku 5.5. Haiku 5.5 moves the whole request to its higher tier once the prompt passes 100K tokens, at five times the price.
  • OpenAI charges 2x input and cache rates and 1.5x output "for the full request" when the prompt exceeds 272K tokens.
  • Google charges $4/$18 instead of $2/$12 on Gemini 3.1 Pro above 200K tokens.
  • xAI applies long-context rates to all tokens once a prompt reaches 200K tokens.

Worked example (illustrative). One request with a 400K-token prompt and 8K tokens of output:

ModelCost
GPT-6 Astra$8.60 (it would be $4.40 at the base rate)
Claude Fable 5.1$4.40
Claude Opus 5.5$1.76
Gemini 3.1 Pro$1.74
GPT-6.1 Sol$1.72
Grok 4.7$1.70
Claude Sonnet 5.5$0.88
Claude Haiku 5.5$0.22
GPT-6 Luna$0.09

At this prompt size, Fable 5.1 and GPT-6 Astra, which share a sticker price, differ by almost 2x. Below 100K tokens the order changes again: Haiku 5.5 and Luna cost the same, at about $0.013 for a 90K-token prompt with 8K output.

Multiplier 5: batch, speed and residency

  • Batch. Anthropic, OpenAI (Batch and Flex) and Google all take 50% off for asynchronous jobs. xAI lists no batch discount for Grok 4.7.
  • Speed. Faster responses cost more. Anthropic's fast mode for Opus 5.5 costs $8/$40, or 2x. OpenAI's Fast mode is 2x standard.
  • Data residency. Anthropic's US-only inference multiplies token prices by 1.1x. OpenAI adds 10% for regional processing on models released on or after March 5, 2026.

Multiplier 6: tools aren't free

Web search is billed per call on top of tokens: $10 per 1,000 searches at Anthropic and OpenAI, and $5 per 1,000 at xAI. Google's Gemini 3.x models get 5,000 free grounded requests a month, then cost $14 per 1,000. The search results you pull in are usually billed as input tokens too. Tool definitions also cost something: Anthropic documents a fixed system-prompt overhead per request when tools are present (286 tokens on Opus 5.5 with auto tool choice), plus about 4,500 tokens for its computer-use toolset.

Competing explanations: why vendors now talk about cost per task

Both labs increasingly argue on cost per task rather than price per token. OpenAI says GPT-6.1 Sol averaged $5.47 per task on Terminal-Bench Science 0.1 at maximum effort, against $23.21 for Claude Opus 5.5 and $23.80 for GPT-6 Astra. Anthropic says Opus 5.5 costs about 40% less than Opus 5 on typical workloads because it uses fewer tokens. These are vendor-run measurements on benchmarks each vendor chose. They point at something real: tokens per task now varies as much as price per token. But you can't verify them from the outside, and they won't transfer to your workload without testing.

There's a less flattering reading too. Matching sticker prices at round numbers makes the table look like a commodity market, while the terms that actually differentiate (tokenizers, thresholds, cache economics) are buried in footnotes and model pages.

What to do with this

  1. Measure tokens per task, not price per token. Run a sample of real requests and record input, output and thinking tokens from the usage fields.
  2. Recount when you switch models. A tokenizer change can wipe out a price cut.
  3. Set effort deliberately. The default isn't always the cheapest setting that works.
  4. Design for the cache. Keep stable content at the front of the prompt.
  5. Watch thresholds. 100K tokens on Haiku 5.5, 200K on Gemini 3.1 Pro and Grok, 272K on OpenAI.
  6. Batch what can wait. It's an easy 50% at three of the four vendors.
  7. Route. Not every request needs a flagship model. See our explainer on model routing. If you're considering open-weight models, see frontier APIs vs open models.

For the competitive context behind OpenAI's lineup, see our coverage of the GPT-6 launch.

About this storyBased on the sources linked below. Editorial standards

Was this useful?Report an error
TopicFrontier models"Frontier models" are the most capable general-purpose AI models available at a given time, such as the top tiers… 15 stories, 1 comparisons.Open the hub
Comments
0

More on Frontier models & business

The Week in AI

Get the cluster, not just the headline.

0