ModelsAnalysis

Do AI labs really ship 'mini' models first? What the release record since 2024 shows

Across 32 small-vs-flagship pairings from six labs, the small model came first only five times. Here's what really drives the timing, and the price war at the bottom.

ByShajanthanFounder & Editor
Published
Reading9 MIN
Timeline of flagship and small AI model launches by six labs from 2024 to October 2026, showing most small models arriving with or after their flagship
What’s new, in 20 seconds
  1. Of 32 small-vs-flagship pairings since 2024, the small model shipped after the flagship 14 times, alongside it 13 times and before it only 5 times.
  2. Anthropic has never shipped a Haiku first; Google is the only lab that regularly leads with Flash, and even then only in some generations.
  3. What changed in 2026 is the price floor: Claude Haiku 5.5 and GPT-6 Luna both list at $0.10/$0.50 per million tokens.
Contents

On October 7, 2026, Anthropic released Claude Haiku 5.5, its new small model, 15 days after Claude Opus 5.5 and nine days after Sonnet 5.5. That order is typical. A common belief in the industry is that labs ship the cheap "mini" version first to test the waters. The release record doesn't support it. We compiled a dataset of 35 small, mini, Flash, Haiku, nano and Luna releases from OpenAI, Anthropic, Google, Meta, Mistral, Alibaba's Qwen team and Microsoft, each dated from an official source, plus two Anthropic generations that never got a small model. In the 32 cases where a small model can be paired with a flagship, the small model came after the flagship 14 times, alongside it (the same day) 13 times and before it only five times. Three releases have no same-generation flagship to pair with (Phi-4-mini, the 2024 Ministral models and Gemini 3.8 Flash). Some pairings are judgment calls, such as counting Gemini 3.5 Flash-Lite and Llama 4 Scout as "before" flagships that hadn't shipped.

So the premise is mostly wrong. The more useful questions are why small models usually trail or travel with the flagship, why Google sometimes breaks the pattern, and what the 2026 releases say about where the money is.

Last verified: October 8, 2026 (GPT-5.6 date re-checked October 9, 2026). The full dataset, with a source URL for every row, is in data/M3-mini-releases.csv.

The record, lab by lab

Anthropic: Haiku always comes last, or not at all

GenerationFlagship launchHaiku launchGap
Claude 3Opus and Sonnet, Mar 4, 2024Haiku, Mar 13, 2024+9 days
Claude 3.5Sonnet, June 2024Haiku announced Oct 22, 2024about +4 months
Claude 4Opus 4 and Sonnet 4, May 22, 2025none—
Claude 4.5Sonnet 4.5, Sep 29, 2025Haiku 4.5, Oct 15, 2025+16 days (Opus 4.5 followed Nov 24)
Claude 5Sonnet 5, Jun 30, 2026; Opus 5, Jul 24, 2026none—
Claude 5.5Opus 5.5, Sep 22; Sonnet 5.5, Sep 28, 2026Haiku 5.5, Oct 7, 2026+15 days

Anthropic's March 2024 Claude 3 announcement said plainly that "Haiku will be available soon." Two whole generations, Claude 4 and Claude 5, never got a Haiku. Anthropic's small model is something it builds once the larger models exist, not a trial run.

OpenAI: mostly alongside, with one real "mini first"

OpenAI usually launches its small models on the same day as the flagship: o1-mini with o1-preview (September 12, 2024), GPT-4.1 mini and nano with GPT-4.1 (April 14, 2025), o4-mini with o3 (April 16, 2025), GPT-5 mini and nano with GPT-5 (August 7, 2025), and GPT-5.6 Luna with GPT-5.6 Sol and Terra (July 9, 2026, per OpenAI's API changelog).

Two small models came after their flagships. GPT-4o mini (July 18, 2024) followed GPT-4o by about two months. GPT-6 Luna (September 22, 2026) followed GPT-6 Astra (September 3) by 19 days. The clear exception is o3-mini, released January 31, 2025, about 11 weeks before the full o3 shipped on April 16. That is the strongest "mini first" case in OpenAI's record.

Google: the one lab that sometimes leads with Flash

Google is probably where the "mini first" idea comes from. On December 11, 2024, it called an experimental Gemini 2.0 Flash "the first model in the Gemini 2.0 family." 2.0 Pro followed as an experimental model on February 5, 2025. On May 19, 2026, Gemini 3.5 Flash went generally available while Google said 3.5 Pro was "in internal use" and coming the next month. As of October 9, 2026, Google's Gemini API models page still lists no Gemini 3.5 Pro. The newest Pro on that page is 3.1 Pro, in preview. The 3.x Flash line has since moved on to 3.6, 3.7 and 3.8.

In the other generations, Pro came first. Gemini 1.5 Flash followed 1.5 Pro. 2.5 Pro (March 25, 2025) preceded 2.5 Flash (April 17) and 2.5 Flash-Lite (June 17). Gemini 3 Pro (November 18, 2025) preceded 3 Flash (December 17). 3.1 Pro (February 19, 2026) preceded 3.1 Flash-Lite (March 3). Google's newest frontier model, Gemini 4 Argon (September 30, 2026), went to a small group of cyber defenders through its Fairwind Program, with no Gemini 4 Flash announced.

Open-weight labs: usually the whole ladder at once

Open-weight families usually ship every size together, because the point is to give developers a ladder of options. Qwen3 (April 29, 2025) went from 0.6B to 235B in one announcement. Gemma 3 (March 10, 2025) shipped 1B to 27B together, and Gemma 4's E2B and E4B on-device models launched with the 26B and 31B models (March 31, 2026 on Google's releases page, April 2 on its blog). Mistral released Ministral 3 (3B, 8B, 14B) on the same day as Mistral Large 3 (December 2, 2025). OpenAI's gpt-oss-20b came with gpt-oss-120b.

The exceptions run in both directions. Meta's Llama 3.2 1B and 3B (September 25, 2024) came after Llama 3.1. Qwen3.5's 0.8B to 9B models appeared on Hugging Face around March 2, 2026, a few weeks after the February flagship. Meta's Llama 4 is the odd one: Scout and Maverick shipped on April 5, 2025, while the giant Behemoth "teacher" model was still in training.

Why small models usually trail: you need a teacher

The most concrete explanation comes from the labs themselves. Small models are often distilled: trained to imitate a larger model's outputs.

  • Google said Gemini 1.5 Flash was trained "using distillation" from 1.5 Pro.
  • Meta said Llama 3.2 1B and 3B were made by "structured pruning" from Llama 3.1 8B, with "logits from the Llama 3.1 8B and 70B models" used as training targets.
  • Meta said it "codistilled the Llama 4 Maverick model from Llama 4 Behemoth as a teacher model."

If the small model learns from the big one, the big one has to exist first. It doesn't have to be released first, though. Llama 4 Behemoth and Gemini 3.5 Pro both show a lab shipping the student while holding back the teacher. When a "Flash first" launch happens, the likely reason is that the flagship isn't ready to sell yet: it may be too costly to serve, still in safety review, or not yet good enough. OpenAI and Anthropic don't say whether their current small models are distilled. OpenAI says only that GPT-6 Sol and Luna "were trained with methods similar to Astra's." Treat distillation as a plausible general explanation, not a confirmed one for every model.

Why labs make them at all: serving cost, latency, devices and price

Serving cost and latency. A smaller model needs less compute per token and responds faster, which is what high-volume products need. Anthropic says Haiku 5.5 is tuned "for high-volume and latency-sensitive work." Its docs list classification, extraction and routing as typical jobs. That is the backbone of model routing: send the easy 80% of requests to the cheap model and keep the expensive one for hard cases.

Defaults for free users. Small models are what labs can afford to give away. OpenAI's January 2025 o3-mini launch made it "the first time a reasoning model has been made available to free users" in ChatGPT. On October 7, 2026, OpenAI said GPT-6 Luna would power ChatGPT's Free and Go tiers. When a small model becomes the free default, it reaches far more people than the flagship does.

On-device. The smallest models are built for phones and laptops. Apple says its on-device foundation model has about 3 billion parameters, stored at 2 bits per weight. Google says Gemma 4 E2B and E4B "run completely offline" on phones and boards like the Raspberry Pi. Meta launched Llama 3.2 1B and 3B with Qualcomm and MediaTek support from day one. If you want to try one yourself, our guide to choosing a small model for your laptop covers the memory math, and our Ollama guide covers setup. What an NPU actually does explains why your laptop's "AI chip" often sits idle while you do it.

The 2026 story: a price war at the bottom, and Flash moving up

The clearest change in 2026 isn't release order but price.

Small tierLaunchList price per 1M tokens (input / output)
Claude 3 HaikuMar 2024$0.25 / $1.25
GPT-4o miniJul 2024$0.15 / $0.60
Claude Haiku 4.5Oct 2025$1 / $5
GPT-5 nanoAug 2025$0.05 / $0.40
GPT-6 LunaSep 22, 2026$0.10 / $0.50 (prompts up to 272K tokens)
Claude Haiku 5.5Oct 7, 2026$0.10 / $0.50 (prompts up to 100K tokens); $0.50 / $2.50 above
Gemini 3.5 Flash-LiteJul 21, 2026$0.30 / $2.50
Gemini 3.5 FlashMay 19, 2026$1.50 / $9.00

Last verified: October 8, 2026, against the Anthropic, OpenAI and Google pricing pages and launch posts.

Three things stand out.

First, Anthropic and OpenAI now list their small tiers at the same price. GPT-6 Luna and Claude Haiku 5.5 launched 15 days apart at $0.10 in and $0.50 out. Anthropic's Haiku 5.5 page compares the model directly with GPT-6 Luna and claims wins on every benchmark it reports. These are vendor-reported numbers, for example 39.2% vs 16.4% on Terminal-Bench 4.0. The conditions differ, though: Haiku 5.5 costs five times as much once a prompt goes over 100,000 tokens, while OpenAI's Luna price applies up to 272,000 tokens. For long-document work, the "same price" isn't the same. That's one reason token prices alone don't tell you what a job costs.

Second, Haiku went up before it came down. Haiku 4.5 cost four times as much as Claude 3 Haiku, because Anthropic sold it as a near-Sonnet-quality model. Haiku 5.5 reverses that. Anthropic says it costs "around 75% less to run" than Haiku 4.5 on average. Its footnote says this works out to 90% lower for requests up to 100K tokens and 50% lower above that, after accounting for a new tokenizer that uses slightly more tokens per task.

Third, "Flash" no longer means cheap. Gemini 3.5 Flash costs $1.50 in and $9 out. That is closer to Gemini 3.1 Pro Preview ($2 / $12) than to Flash-Lite. Google sells it as a frontier agent model: its launch post says it completes long agentic tasks "often at less than half the cost of other frontier models." Google's low-cost tier is now Flash-Lite. The names have shifted, so compare prices, not labels.

Where the premise does hold

The "mini first" story has some truth in it:

  • Google's 2.0 and 3.5 generations clearly led with Flash. In 3.5's case the Pro model still isn't in the public API almost five months later.
  • o3-mini shipped about 11 weeks before o3.
  • Llama 4 shipped its smaller models while the largest stayed unreleased.

In each case the flagship wasn't ready to sell, rather than the lab choosing to start small. That's an inference from the timing; none of these labs has said publicly why it held a flagship back. When the flagship is ready, the labs mostly ship the whole family together, as OpenAI, Mistral, Qwen and Gemma do, or follow with the small model weeks later, as Anthropic does.

What this means for you

  • Don't wait for a mini to judge a new generation. From Anthropic it may never arrive, as Claude 4 and Claude 5 showed. From OpenAI it usually comes on day one.
  • Re-evaluate your cheap tier every generation. Small-model prices moved from $1/$5 to $0.10/$0.50 at Anthropic in one year. If you route traffic by cost, see our model-routing explainer. If you're choosing among Claude tiers, see which Claude model to use.
  • Read the fine print on context. Tiered pricing above 100K or 272K tokens can change which "cheap" model is actually cheaper for your workload.
  • Confirmed vs uncertain: release dates and prices above are confirmed from official pages. Benchmark comparisons are vendor-reported. Whether a specific current small model was distilled from its flagship is confirmed only where the lab says so: Google for Gemini 1.5 Flash, Meta for Llama 3.2 and Llama 4.
SourcesClaude Platform release notes — Anthropic · Introducing Claude Haiku 5.5 — Anthropic · Pricing — Claude Platform docs · Introducing Claude Haiku 4.5 — Anthropic · Introducing Claude 4 — Anthropic · Introducing the next generation of Claude (Claude 3) — Anthropic · Claude 3 Haiku — Anthropic · Claude 3.5 Sonnet — Anthropic · New Claude 3.5 models and computer use — Anthropic · GPT-4o mini — OpenAI · Hello GPT-4o — OpenAI · Introducing OpenAI o1-preview — OpenAI · OpenAI o3-mini — OpenAI · Introducing o3 and o4-mini — OpenAI · GPT-4.1 in the API — OpenAI · Introducing GPT-5 for developers — OpenAI · Introducing gpt-oss — OpenAI · API changelog — OpenAI · Introducing GPT-6 Sol and Luna — OpenAI · GPT-6 and Intelligent UI for everyone — OpenAI · Gemini 1.5 Flash at I/O 2024 — Google · Introducing Gemini 2.0 — Google · Gemini 3 Flash — Google · Gemini 3.5: frontier intelligence with action — Google · Gemini 4 Argon — Google · Gemini API changelog — Google AI for Developers · Gemini API models — Google AI for Developers · Gemini API pricing — Google AI for Developers · Gemma releases — Google AI for Developers · Gemma 4 — Google · Llama 3.2 — Meta AI · The Llama 4 herd — Meta AI · Un Ministral, des Ministraux — Mistral AI · Introducing Mistral 3 — Mistral AI · Qwen3 — Qwen team · Qwen3.5 collection — Hugging Face · Phi-4-mini-instruct model card — Microsoft on Hugging Face · Apple Intelligence Foundation Language Models 2025 — Apple Machine Learning Research · OpenAI cuts GPT-6 prices in half with Sol and Luna — The Next Web

About this storyBased on the sources linked below. Editorial standards

Was this useful?Report an error
TopicSmall modelsSmall models are language models designed to be cheap and fast enough to run at high volume or on local hardware… 6 stories, 1 guides, 1 comparisons.Open the hub
Comments
0

More on Small models & models

The Week in AI

Get the cluster, not just the headline.

0