DeepSeek-V4.1-Flash: an MIT-licensed open-weights MoE that DeepSeek says beats its own flagship

Released September 10, DeepSeek's new Flash model activates 8B–16B parameters per token, reads 1M tokens and images, and now answers V4-Pro API calls. What's confirmed and what's vendor-reported.

ByShajanthanFounder & Editor
Published
Reading6 MIN
Diagram contrasting DeepSeek-V4.1-Flash's 763 billion published parameters and 552 billion backbone parameters with the 8 and 16 billion parameters active per token
What’s new, in 20 seconds
  1. On September 10, 2026, DeepSeek released DeepSeek-V4.1-Flash: a mixture-of-experts model with a 552B-parameter backbone (763B parameters in the published weights), 8B active parameters for input and 16B for output, a 1M-token context and image input, under the MIT license.
  2. DeepSeek says it beats its own V4-Pro flagship and matches or exceeds some US frontier models on coding-agent benchmarks. The scores are self-reported, and the US comparisons are against models that have since been replaced.
  3. API prices start at $0.15 input and $0.60 output per million tokens off-peak. Self-hosting needs datacenter GPUs, not a laptop.
Contents

On September 10, 2026, DeepSeek released DeepSeek-V4.1-Flash, the most significant open-weights release of the past month. Despite the "Flash" name, it is a very large mixture-of-experts (MoE) model: DeepSeek describes a 552-billion-parameter backbone, and the published weights hold 763 billion parameters. It activates 8 billion parameters per token while reading a prompt and 16 billion while generating. It accepts images as well as text, handles a 1-million-token context, and is published under the MIT license. That is one of the most permissive terms any frontier-scale lab uses; see our comparison of open-weight licenses. DeepSeek also made it the default behind its API: since September 14, requests to the older V4-Pro flagship are answered by V4.1-Flash, at Flash prices.

What shipped

Last verified: October 8, 2026

ItemDeepSeek-V4.1-Flash
ReleasedSeptember 10, 2026
ArchitectureMoE, 1 shared and 384 routed experts per layer, 6 routed experts active per token
Total parameters763B in the published weights (Hugging Face metadata); DeepSeek describes a 552B "backbone" plus a 196B "Engram" memory (see below)
Active parameters8B during prefill (input), 16B during decode (output)
ContextUp to 1M tokens; API max output 384K tokens
ModalitiesText and image in, text out
Pretraining data45 trillion tokens, multimodal
WeightsHugging Face, deepseek-ai/DeepSeek-V4.1-Flash, FP8 safetensors
LicenseMIT (repository and weights)
API model IDdeepseek-flash

Sources: DeepSeek's release note, model card and pricing page.

DeepSeek calls it "the smallest model in our new architecture family, with native visual understanding". That implies larger V4.1 models are coming, and DeepSeek has said a V4.1-Pro will follow, without giving a date.

The architecture changes

The model card describes several departures from earlier DeepSeek models:

  • A "Causal Encoder–Decoder" design. The 40 layers are split into a 20-layer causal encoder and a 20-layer decoder. DeepSeek uses this split to explain why fewer parameters are active while reading input (8B) than while writing output (16B).
  • A smaller KV cache. DeepSeek says the KV cache stores about 890 bytes per token in a 4-bit floating-point format. That is roughly a quarter of the HBM V4-Flash needed, with about one-eighth of the SSD storage. At a 1M-token context, cache size largely determines serving cost, so this matters more than the parameter count.
  • Native vision. Image understanding, which shipped as a separate experimental model in August, is now built in through a DeepSeek-ViT encoder.
  • Adjustable reasoning effort. The card says reasoning effort can be set from 1 to 100.

The parameter count needs a footnote. DeepSeek's release note calls V4.1-Flash a "552B-parameter MoE", and its model card says "552B backbone parameters". The safetensors metadata on Hugging Face, which counts every parameter in the published files, lists 763B. The card separately describes a 196B-parameter "Engram conditional memory", sparsely accessed by token lookup. Backbone plus Engram comes to 748B, so about 15B of the published total isn't itemized by DeepSeek. For memory and hardware planning, 763B is the figure that matters. SiliconANGLE reports that the previous V4-Flash had 284B parameters, so the new model is nearly twice as large.

Benchmarks (DeepSeek-reported)

All numbers below are vendor-reported: they come from DeepSeek's model card, measured at maximum reasoning effort, including the figures DeepSeek gives for competing models.

Last verified: October 8, 2026

BenchmarkV4.1-FlashV4-ProClaude Opus 5†GPT-5.6 Sol
DeepSWE v1.1 (resolved)74.262.774.073.0
Terminal-Bench 2.190.687.989.188.8
Terminal-Bench 4.031.212.451.839.9
AutomationBench54.843.250.345.8
GPQA Diamond90.992.493.494.1
Humanity's Last Exam (no tools)36.8 (39.1*)42.7*56.344.5

*Text-only subset of HLE. DeepSeek reports V4-Pro only on that subset, so the like-for-like comparison is 39.1 vs 42.7. †Labeled "Opus 5.0" in DeepSeek's table. Full rows, including Kimi K3 and GLM-5.3, are in our data file.

How to read this:

  • The claim against its own flagship holds on DeepSeek's agentic coding numbers. V4.1-Flash beats V4-Pro on DeepSWE, Terminal-Bench and AutomationBench, but it trails V4-Pro on GPQA Diamond and on Humanity's Last Exam (39.1 vs 42.7 on the text-only subset).
  • The US comparisons are already dated. The table compares against Claude Opus 5 and GPT-5.6 Sol. On September 22, 2026, twelve days after this release, Anthropic shipped Claude Opus 5.5 and OpenAI shipped GPT-6 Sol. DeepSeek's chart says nothing about how V4.1-Flash compares with either.
  • Harder benchmarks show a gap. On Terminal-Bench 4.0 and Humanity's Last Exam, the older Opus 5 still leads by roughly 20 points.
  • "Tests by multiple parties." DeepSeek's release note says outside tests put V4.1-Flash ahead of V4-Pro on performance, cost, speed and total runtime, but it doesn't name who ran them.

Pricing and availability

V4.1-Flash is live in DeepSeek's API and, according to SiliconANGLE, in its web and mobile apps. DeepSeek charges different rates at peak and off-peak times:

Last verified: October 8, 2026 (per 1M tokens)

Cache-hit inputCache-miss inputOutput
Off-peak$0.003$0.15$0.60
Peak (01:00–04:00 and 06:00–10:00 UTC, weekdays except Chinese public holidays)$0.006$0.30$1.20

DeepSeek retired V4-Flash and its experimental vision model, and calls to either now route to V4.1-Flash. From 04:00 UTC on September 14, 2026, deepseek-v4-pro requests are also served by V4.1-Flash at V4.1-Flash prices "until V4.1-Pro launches". One inconsistency: DeepSeek's pricing page still lists a separate deepseek-v4-pro entry at higher rates. If you call that model ID, check your invoices.

For comparison, those rates are a fraction of current US frontier APIs. OpenAI's GPT-6 Sol lists $2 input and $10 output per million tokens, and Anthropic's Claude Opus 5.5 lists $4 and $20. Price per token isn't the whole cost, though; we explain why in frontier APIs vs open models: what they really cost.

Running it yourself

The weights are on Hugging Face. DeepSeek documents serving with vLLM and SGLang, plus Docker Model Runner. Community quantizations for llama.cpp, Ollama and LM Studio are already listed. DeepSeek does not publish hardware requirements. With 763B parameters in FP8 weights (roughly 0.76 TB at one byte per parameter, before any cache), this is a multi-GPU server model: 8B active parameters cut compute per token, not the memory needed to hold all the experts. Our guide to running open models locally with Ollama covers smaller models that fit on a workstation. The model ships without a standard Jinja chat template; DeepSeek provides a Python reference encoder instead, so check that your serving stack supports it.

Why it matters

  • MIT matters. There's no user cap, no acceptable-use policy attached to the license, and no naming requirement. Compare Meta's Llama licenses, or Qwen's custom license for its 2.4T flagship. Those are the kinds of terms we break down in our license comparison.
  • The cost curve keeps bending. A model DeepSeek says beats its own previous flagship on agentic coding now costs $0.60 per million output tokens off-peak.
  • Regulation applies too. The EU's guidelines presume a model is general-purpose AI if it was trained with more than 10^23 FLOP. DeepSeek hasn't published its training compute, but a common rough estimate (6 × active parameters × training tokens) puts 45 trillion tokens at 8–16B active parameters at about 2–4 × 10^24 FLOP, far above that line. Anyone placing it on the EU market takes on obligations. Open-source licensing exempts some documentation duties, but not if the model is monetized or reaches the systemic-risk threshold. Our EU AI Act explainer covers the rules.

The release has context beyond benchmarks. On the same day, according to SiliconANGLE, Anthropic's threat-intelligence report named DeepSeek among China-based labs it says ran distillation campaigns against Claude. SiliconANGLE reported no response from DeepSeek, and the allegation is Anthropic's.

What else is coming

Last verified: October 9, 2026. Two Western labs announced large open-weights models on October 5 and 6, 2026, but had not released the weights as of October 9:

  • Mistral Large 4 (announced October 6): Mistral says it is a roughly 1-trillion-parameter MoE with 52B active parameters and multimodal input. It is in API preview now at $1.36 and $4.18 per million input and output tokens, and Mistral says the weights will follow "by the end of the month". The license hasn't been stated, and Mistral's Hugging Face page listed no Large 4 repository on October 9. The Register reports that Artificial Analysis places it well behind the top proprietary models.
  • Reflection AI's Beam (announced October 5): TechCrunch reports 501B total and 23B active parameters, text-only, with a 1M context. Reflection says the weights will be released this month. We found no public release of the weights as of October 9.

We'll cover both when the weights and licenses are public.

About this storyBased on the sources linked below. Editorial standards

Was this useful?Report an error
TopicOpen modelsOpen models are AI models whose weights are published for anyone to download and run, usually under a license that… 8 stories, 1 guides, 1 comparisons.Open the hub
Comments
0

More on Open models & open source

The Week in AI

Get the cluster, not just the headline.

0