A year ago, the usual advice was that open models were the cheap option and frontier APIs the expensive, capable one. The official price pages in October 2026 tell a messier story. Hosted open-weight models are inexpensive: OpenAI's open-weight gpt-oss-120b costs $0.15 per million input tokens and $0.60 per million output tokens on both Together AI and Fireworks. But Anthropic's Claude Haiku 5.5 and OpenAI's GPT-6 Luna now list at $0.10 and $0.50. Meanwhile, renting your own GPU to serve an open model costs thousands of dollars a month before it answers a single request.
This comparison covers three ways to buy language-model capability: proprietary frontier APIs, hosted open-weight APIs, and self-hosting open weights on rented GPUs. It uses prices from official pricing pages, keeps assumptions labeled, and lists the costs that don't appear on a price page. The structured version lives on our compare page; this article walks through the reasoning.
Last verified: October 8, 2026
Option 1: frontier APIs
You send requests to Anthropic, OpenAI or Google, and they run the model. You pay per token and get the most capable proprietary models, with no infrastructure to run.
| Model | Input $/MTok | Output $/MTok | Notes |
|---|
| GPT-6 Astra | $10 | $50 | 2x input, 1.5x output above 272K-token prompts |
| Claude Fable 5.1 | $10 | $50 | Flat pricing to 1M tokens |
| Claude Opus 5.5 | $4 | $20 | Anthropic's recommended default |
| Gemini 3.1 Pro (Preview) | $2 | $12 | $4/$18 above 200K tokens |
| Claude Sonnet 5.5 | $2 | $10 | |
| GPT-6.1 Sol | $2 | $10 | |
| Gemini 3.8 Flash | $0.75 | $3.75 | Through Dec 31, 2026; $1.50/$7.50 from Jan 1, 2027 |
| Claude Haiku 5.5 | $0.10 | $0.50 | Prompts up to 100K tokens; $0.50/$2.50 above |
| GPT-6 Luna | $0.10 | $0.50 | |
All three vendors offer about 50% off for batch jobs and discounted cache reads. The details, and why sticker prices mislead, are in token pricing isn't the story. For the Claude lineup in depth, see which Claude model to use.
Option 2: hosted open-weight models
Inference providers run open-weight models on their own GPUs and sell access per token, much like a frontier API. The weights are public, so you can switch providers or later run the same model yourself.
| Model | License (from model card) | Provider | Input $/MTok | Output $/MTok |
|---|
| gpt-oss-120b | Apache 2.0 | Together AI, Fireworks | $0.15 | $0.60 |
| DeepSeek V4.1 Flash | MIT | DeepSeek API (peak), Together AI, Fireworks | $0.30 | $1.20 |
| DeepSeek V4.1 Flash | MIT | DeepSeek API, off-peak | $0.15 | $0.60 |
| DeepSeek V4 Pro (0813) | MIT (DeepSeek-V4-Pro card) | DeepSeek API (peak), Together AI | $1.32 | $3.96 |
| GLM-5.3 | Custom GLM-5.3 license | Together AI, Fireworks | $1.40 | $4.40 |
| Kimi K3 | Custom Kimi K3 License | Fireworks | $3.00 | $15.00 |
Last verified: October 8, 2026. A few notes:
- DeepSeek's first-party API halves prices off-peak. Its peak (full-price) hours are 01:00–04:00 and 06:00–10:00 UTC, Monday to Friday, excluding Chinese public holidays.
- Fireworks bills batch inference at 50% of serverless prices.
- Together AI lists Kimi K3 at $2.70/$13.50 with a "PROMO" tag.
- "Open" doesn't always mean permissive. GLM-5.3 and Kimi K3 use custom licenses with MIT-style grants plus extra conditions. GLM-5.3's license requires "Model as a Service" providers with more than $10 billion in revenue over 12 months to pass a Z.AI security review before commercial use. Kimi K3's requires a separate agreement with Moonshot AI for model-as-a-service providers above $20 million in revenue over 12 months, and a prominent "Kimi K3" credit in products with more than 100 million monthly users or $20 million in monthly revenue. Both have exemptions; read the license files before use. Our guide to open-weight licenses explains the differences.
Two things stand out. First, the price spread among open models (gpt-oss-120b to Kimi K3) is as wide as among frontier models. "Open" isn't a price tier. Second, at the bottom of the market the frontier labs now undercut the hosted open models on list price. Whether Haiku 5.5 or GPT-6 Luna is better than gpt-oss-120b for your task is a separate question that list prices can't answer. Check published independent evaluations and run your own evals. Each vendor tokenizes text differently, so equal list prices don't guarantee equal bills.
Option 3: self-hosting on rented GPUs
You rent GPUs, run an open model with a serving stack such as vLLM, SGLang or TensorRT-LLM, and pay by the hour whether or not traffic arrives.
| Provider | GPU | On-demand $/GPU-hour | Per month (730 h) |
|---|
| RunPod | H100 PCIe | $2.89 | $2,109.70 |
| RunPod | H100 SXM | $3.99 | $2,912.70 |
| Lambda | H100 SXM | $3.99 | $2,912.70 |
| Together AI (GPU clusters) | HGX H100 | $3.99 ($1.99 preemptible) | $2,912.70 |
| RunPod | H200 | $5.29 | $3,861.70 |
| Lambda | B200 | $6.69 | $4,883.70 |
| Together AI (managed dedicated endpoint) | HGX H100 | $5.49 | $4,007.70 |
| Fireworks (managed on-demand deployment) | H100 80GB | $8.00 | $5,840.00 |
Last verified: October 8, 2026; RunPod and Lambda rows re-checked October 9, 2026 (RunPod listed higher H100 SXM and H200 prices on October 9 than on October 8, and its page doesn't say whether they are Secure or Community Cloud rates). The two managed rows include a serving stack. The others are raw compute.
Model size decides how many GPUs you need. OpenAI's Hugging Face card says gpt-oss-120b (117B parameters, 5.1B active) is built to run on "a single 80GB GPU" such as an H100. The larger open models are a different proposition. DeepSeek-V4.1-Flash's published weights total 763B parameters, including what DeepSeek describes as a 552B backbone and a 196B Engram memory. By our arithmetic, at one byte per parameter the weights alone would be about 760 GB, more than the 640 GB on an eight-GPU H100 node. That means H200- or B200-class hardware before you've added any memory for serving.
The arithmetic: when does self-hosting win?
Worked example (illustrative). A product processes 500 million input tokens and 100 million output tokens a month. That averages about 228 tokens per second.
| Option | Monthly cost at list price |
|---|
| GPT-6 Astra or Claude Fable 5.1 | $10,000 |
| Claude Opus 5.5 | $4,000 |
| Kimi K3 (Fireworks) | $3,000 |
| Gemini 3.1 Pro | $2,200 |
| Claude Sonnet 5.5 or GPT-6.1 Sol | $2,000 |
| GLM-5.3 (hosted) | $1,140 |
| Gemini 3.8 Flash (2026 price) | $750 |
| DeepSeek V4.1 Flash (DeepSeek API, peak) | $270 |
| gpt-oss-120b (hosted) | $135 |
| Claude Haiku 5.5 or GPT-6 Luna | $100 |
| Self-hosted gpt-oss-120b, one Lambda H100 | $2,912.70 |
| The same with a second GPU for redundancy | $5,825.40 |
At this volume, self-hosting gpt-oss-120b costs about 20 times more than paying a provider to host the same model, and more than most frontier mid-tier APIs.
The break-even depends on throughput, the number of tokens per second a GPU actually sustains for your workload. No neutral, published per-GPU figure for this model was found. NVIDIA claims up to 1.5 million tokens per second for gpt-oss-120b on a 72-GPU GB200 NVL72 rack (August 2025), which is a vendor peak on newer hardware. So the scenarios below are assumptions, not measurements. They also treat input and output tokens together, which is a simplification: GPUs process input (prefill) much faster than they generate output.
| Assumed sustained throughput per H100 | Cost per million tokens at $3.99/h, 100% busy | At 40% average utilization |
|---|
| 500 tokens/s | $2.22 | $5.54 |
| 2,000 tokens/s | $0.55 | $1.39 |
| 5,000 tokens/s | $0.22 | $0.55 |
| 10,000 tokens/s | $0.11 | $0.28 |
Hosted gpt-oss-120b works out to about $0.225 per million tokens at our 5:1 input-to-output mix. To match it, one H100 at $3.99/hour has to sustain about 4,900 tokens per second around the clock. That's roughly 12.9 billion tokens a month per GPU, more than 20 times our example's volume. Against a mid-tier frontier API ($3.33 per million blended for Sonnet 5.5 or GPT-6.1 Sol at the same mix), break-even falls to about 330 tokens per second. But that compares two different models, and only your evals can say whether the open model is good enough.
The costs that aren't on any price page
Self-hosting adds:
- Engineering time to deploy, tune, monitor and upgrade the serving stack, often the largest line item.
- Idle capacity. You pay for peak capacity all day.
- Redundancy, plus storage and egress.
- Security patching, and evals every time you change a model.
APIs add:
- Rate limits and outages you can't control.
- Model retirements on the vendor's schedule. For example, Anthropic retires Claude Sonnet 4.5 on November 30, 2026.
- Prompt and tokenizer changes between versions.
- Vendor lock-in through proprietary features.
Both add: the cost of measuring quality. A cheaper model that fails 5% more often can cost more once you count retries and human review.
Compliance cuts both ways. Self-hosting keeps data inside your infrastructure. API vendors offer data-residency options at a premium: Anthropic charges 1.1x for US-only inference, and OpenAI adds 10% for regional processing. Check where any hosted provider processes data. If you're in the EU, see what the AI Act requires of businesses that only use model APIs.
Recommendation by use case
- Prototype or low, spiky volume: a frontier API. No fixed costs, and the vendors' most capable models.
- High-volume simple tasks (classification, extraction): price Claude Haiku 5.5, GPT-6 Luna and hosted gpt-oss-120b against each other on your own data. List prices are within a few cents.
- Hardest reasoning and agentic coding: frontier flagship APIs are the default choice, since the vendors position them as their most capable models. Their benchmark claims are mostly vendor-reported, so check current independent evaluations before assuming an open model can or can't substitute.
- Data can't leave your environment: self-host, and accept the fixed costs. Start with a model that fits on one GPU.
- Very high, steady volume (billions of tokens a month) with an in-house platform team: self-hosting can win on cost. Measure throughput first.
- Want an exit from lock-in: start with a hosted open model. The same weights can move to another provider or in-house later.
- Mixed workloads: route between tiers. See model routing.
To try an open model on your own machine before committing, see our guide to running an open model locally with Ollama.
What's confirmed and what isn't
- Confirmed (official pages): every API and GPU price above; the licenses of gpt-oss-120b (Apache 2.0), DeepSeek-V4.1-Flash and DeepSeek-V4-Pro (MIT), and GLM-5.3 and Kimi K3 (custom), from their model cards and license files; model sizes; and the off-peak and batch discounts.
- Assumptions: throughput, utilization, traffic volume and input/output mix.
- Not covered here: quality comparisons between open and frontier models. This comparison is based on official pricing pages, model cards and documentation, checked October 8–9, 2026.