Every price on this page was verified on 2026-08-13 against the vendor's own pricing page: OpenAI, Anthropic, Google, DeepSeek, and Mistral. All figures are USD per 1M tokens, standard non-batch non-cached tier, unless stated. Where a vendor does not publish a number, this page says so rather than guessing — several widely-shared LLM pricing tables currently list models that were retired more than a year ago.
The Thing Most Pricing Comparisons Get Wrong
Every comparison table on the internet, including this one, lists a price per million tokens. That number is only half of a cost, because the other half is how many tokens the vendor's tokenizer produces from your text — and that is not constant across vendors.
Anthropic states plainly in its own pricing documentation that Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text than earlier models. Claude Sonnet 4.6 and earlier use the previous tokenizer.
The consequence is significant and almost universally ignored:
| Comparison | Sticker output price | Tokens for the same text | Effective cost for the same work |
|---|---|---|---|
| Claude Sonnet 4.6 | $15.00 | baseline | $15.00 |
| Claude Sonnet 5 | $10.00 | ~1.3x | ~$13.00 |
| Claude Opus 5 | $25.00 | ~1.3x | ~$32.50 |
Sonnet 5's sticker price is 33% below Sonnet 4.6, but on the same input text the real saving is closer to 13%. The exact increase depends on content and workload shape — Anthropic does not publish a per-content figure — so treat the 1.3x as the vendor's own stated approximation, not a precise multiplier.
Practical rule: compare vendors on measured cost per completed task, not on the rate card. Run the same 20 representative requests through each candidate and compare invoices. Everything below is the rate card; it is necessary but not sufficient.
Cross-Provider Pricing (per 1M tokens)
Frontier tier
| Model | Provider | Input | Cached input | Output |
|---|---|---|---|---|
| GPT-5.5-pro / GPT-5.4-pro | OpenAI | $30.00 | — | $180.00 |
| GPT-5-pro | OpenAI | $15.00 | — | $120.00 |
| Claude Fable 5 | Anthropic | $10.00 | $1.00 | $50.00 |
| GPT-5.6 Sol | OpenAI | $5.00 | $0.50 | $30.00 |
| GPT-5.5 | OpenAI | $5.00 | $0.50 | $30.00 |
| Claude Opus 5 | Anthropic | $5.00 | $0.50 | $25.00 |
| Claude Opus 4.8 | Anthropic | $5.00 | $0.50 | $25.00 |
Mid tier (best balance)
| Model | Provider | Input | Cached input | Output |
|---|---|---|---|---|
| GPT-5.4 | OpenAI | $2.50 | $0.25 | $15.00 |
| Claude Sonnet 4.6 | Anthropic | $3.00 | $0.30 | $15.00 |
| GPT-5.6 Terra | OpenAI | $2.00 | $0.20 | $12.00 |
| Gemini 3.1 Pro | $2.00 | — | $12.00 | |
| Claude Sonnet 5 | Anthropic | $2.00 | $0.20 | $10.00 |
| GPT-5.2 / GPT-5.3-Codex | OpenAI | $1.75 | $0.175 | $14.00 |
| Gemini 3.5 Flash | $1.50 | — | $9.00 | |
| Gemini 3.6 Flash | $1.50 | — | $7.50 | |
| Mistral Medium 3.5 | Mistral | $1.50 | — | $7.50 |
| Gemini 2.5 Pro | $1.25 | — | $10.00 | |
| GPT-5.1 / GPT-5 | OpenAI | $1.25 | $0.125 | $10.00 |
Budget tier (cost-optimised)
| Model | Provider | Input | Cached input | Output |
|---|---|---|---|---|
| Claude Haiku 4.5 | Anthropic | $1.00 | $0.10 | $5.00 |
| GPT-5.4-mini | OpenAI | $0.75 | $0.075 | $4.50 |
| DeepSeek-V4-Pro | DeepSeek | $0.435 | $0.003625 | $0.87 |
| Mistral Large 3 | Mistral | $0.50 | — | $1.50 |
| Gemini 3.5 Flash-Lite | $0.30 | — | $2.50 | |
| Gemini 2.5 Flash | $0.30 | — | $2.50 | |
| Mistral Codestral | Mistral | $0.30 | — | $0.90 |
| GPT-5.4-nano | OpenAI | $0.20 | $0.02 | $1.25 |
| GPT-5.6 Luna | OpenAI | $0.20 | $0.02 | $1.20 |
| GPT-5-mini | OpenAI | $0.25 | $0.025 | $2.00 |
| Mistral Small 4 | Mistral | $0.15 | — | $0.60 |
| DeepSeek-V4-Flash | DeepSeek | $0.14 | $0.0028 | $0.28 |
| Gemini 2.5 Flash-Lite | $0.10 | — | $0.40 | |
| Mistral Ministral 3 (3B) | Mistral | $0.10 | — | $0.10 |
| GPT-5-nano | OpenAI | $0.05 | $0.005 | $0.40 |
Notes on the table. Google tiers several models by context length: Gemini 3.1 Pro charges $2.00/$12.00 up to 200,000 tokens and $4.00/$18.00 above it; Gemini 2.5 Pro charges $1.25/$10.00 up to 200,000 tokens and $2.50/$15.00 above. Google also prices audio input separately on some models ($1.00 on Gemini 2.5 Flash, $0.50 on Gemini 3.1 Flash-Lite). Meta's Llama models have no first-party per-token price because they are open-weight — costs come from whichever host you use, and this page does not quote host prices it has not verified.
<text x="182" y="96" text-anchor="end">Claude Opus 5</text>
<rect x="190" y="84" width="475" height="20" fill="#2813e8"/>
<text x="671" y="99">$25.00</text>
<text x="182" y="130" text-anchor="end">GPT-5.4</text>
<rect x="190" y="118" width="285" height="20" fill="#2813e8"/>
<text x="481" y="133">$15.00</text>
<text x="182" y="164" text-anchor="end">Gemini 3.1 Pro</text>
<rect x="190" y="152" width="228" height="20" fill="#2813e8"/>
<text x="424" y="167">$12.00</text>
<text x="182" y="198" text-anchor="end">Claude Sonnet 5</text>
<rect x="190" y="186" width="190" height="20" fill="#2813e8"/>
<text x="386" y="201">$10.00</text>
<text x="182" y="232" text-anchor="end">Gemini 3.6 Flash</text>
<rect x="190" y="220" width="143" height="20" fill="#2813e8"/>
<text x="339" y="235">$7.50</text>
<text x="182" y="266" text-anchor="end">Claude Haiku 4.5</text>
<rect x="190" y="254" width="95" height="20" fill="#2813e8"/>
<text x="291" y="269">$5.00</text>
<text x="182" y="300" text-anchor="end">GPT-5.6 Luna</text>
<rect x="190" y="288" width="23" height="20" fill="#2813e8"/>
<text x="219" y="303">$1.20</text>
<text x="182" y="334" text-anchor="end">DeepSeek-V4-Flash</text>
<rect x="190" y="322" width="5" height="20" fill="#2813e8"/>
<text x="201" y="337">$0.28</text>
The spread between the most expensive output on this chart ($30.00) and the cheapest ($0.28) is over 100x — and the pro tiers reach $180.00, a 640x spread. That gap, not the headline model name, is where your bill is won or lost.
Two Pricing Changes Worth Knowing About
Pricing pages are not static, and two current vendor statements materially affect planning.
Anthropic cancelled a scheduled Sonnet 5 price rise. Sonnet 5 launched at $2.00/$10.00 as introductory pricing through 2026-08-31, with a scheduled increase to $3.00/$15.00 on 2026-09-01. Anthropic's pricing documentation now states that this is the standard price and the increase will not occur. If you built a budget around Sonnet 5 rising by 50% in September, you can release that reserve.
DeepSeek has announced it will raise prices. DeepSeek's own pricing documentation carries a notice that it plans to raise overall API pricing "in the near future, with a significant increase expected." DeepSeek is currently the price floor at $0.14/$0.28, so anyone architecting around that floor should treat it as temporary and keep a migration path to Gemini Flash-Lite, GPT-5-nano, or Mistral Small 4.
Reasoning Models and the Thinking-Token Trap
Every flagship now has a thinking mode — OpenAI's reasoning effort setting, Claude's extended thinking, Gemini's thinking levels. Those reasoning tokens bill as output tokens, even though you never see them.
A model that thinks for 4,000 tokens before emitting a 200-token answer bills roughly 4,200 output tokens. On GPT-5.6 Sol at $30.00 output that single call costs about $0.13, versus about $0.006 for the visible answer alone — a 21x difference invisible on the rate card.
Practical rules:
- Default to minimal or low effort for extraction, classification, routing, and simple chat.
- Reserve high effort for genuinely hard reasoning, multi-step agentic work, and tricky debugging.
- Measure it. Reasoning token counts appear in the usage object of every provider's response — log them, because they are the line item most likely to blow a budget.
The Discount Levers, With Real Multipliers
These are published, verified, and stack. They matter more than model choice for repetitive workloads.
Prompt caching. Anthropic publishes exact multipliers relative to base input: a 5-minute cache write costs 1.25x, a 1-hour write 2x, and a cache read costs 0.1x. That means a 5-minute cache pays for itself after a single read, and a 1-hour cache after two. OpenAI's published cached-input rates are likewise 10% of standard across the GPT-5.x line — Sol drops from $5.00 to $0.50, Terra from $2.00 to $0.20. Mistral advertises up to 90% off cached input. For any application with a large fixed system prompt or shared document context, this is the single biggest lever available.
Batch processing. Anthropic's Batch API takes 50% off both input and output — Opus 5 falls to $2.50/$12.50, Sonnet 5 to $1.00/$5.00, Haiku 4.5 to $0.50/$2.50. Google shows comparable reductions for Batch and Flex tiers, and Mistral advertises 50%. If a job does not need an answer in the next few seconds, it should not be paying interactive rates.
Context-window pricing. Anthropic includes the full 1M-token window at standard pricing on Claude 4.6 and later — a 900,000-token request bills at the same per-token rate as a 9,000-token one, and caching and batch discounts apply across the whole window. Google instead steps prices above 200,000 tokens. If your workload is genuinely long-context, that structural difference can outweigh the headline rate.
Watch the surcharges. Anthropic applies a 1.1x multiplier when you pin inference to US-only via the data-residency parameter, and a 10% premium for regional or multi-region endpoints on Bedrock and Google Cloud versus global endpoints. Google's Priority tier costs roughly 80% more than standard. Anthropic's Fast mode for Opus 5 and Opus 4.8 is priced at $10.00/$50.00, double the standard rate. None of these appear in comparison tables, and all of them appear on invoices.
Cost Analysis by Use Case
All figures below are computed from the verified rates above.
Chatbot / conversational AI
Typical turn: 800 input, 400 output tokens.
| Model | Cost / 1,000 turns | Monthly (100K turns) |
|---|---|---|
| DeepSeek-V4-Flash | $0.22 | $22 |
| Gemini 2.5 Flash-Lite | $0.24 | $24 |
| GPT-5.6 Luna | $0.64 | $64 |
| Claude Sonnet 5 | $5.60 | $560 |
| GPT-5.4 | $8.00 | $800 |
Document summarisation
Typical: 10,000 input, 500 output tokens.
| Model | Cost / document | Monthly (10K docs) |
|---|---|---|
| DeepSeek-V4-Flash | $0.0015 | $15 |
| Gemini 2.5 Flash | $0.0043 | $43 |
| Claude Haiku 4.5 | $0.0125 | $125 |
| Claude Sonnet 5 | $0.0250 | $250 |
| GPT-5.4 | $0.0325 | $325 |
Summarisation is input-heavy, which is exactly the shape that prompt caching rewards least (the document changes every time) and long-context pricing rewards most. Note how Claude Sonnet 5 beats GPT-5.4 here on raw rates — then remember the tokenizer adjustment above, which narrows that gap.
Code generation
Typical: 2,000 input, 1,000 output tokens.
| Model | Cost / request | Monthly (50K req) |
|---|---|---|
| Mistral Codestral | $0.0015 | $75 |
| Claude Sonnet 5 | $0.0140 | $700 |
| GPT-5.3-Codex | $0.0175 | $875 |
| GPT-5.4 | $0.0200 | $1,000 |
| Claude Opus 5 | $0.0350 | $1,750 |
Code generation is output-heavy, so output rates dominate — which is why the 5-6x input-to-output ratio matters more here than anywhere else.
Managed Platforms: Bedrock, Azure, Vertex
| Platform | Models | Pricing vs direct | Use when |
|---|---|---|---|
| AWS Bedrock | Claude, Llama, Mistral, others | Tracks list price; 10% premium on regional vs global endpoints for Claude 4.5+ | In AWS, need VPC endpoints or compliance |
| Microsoft Foundry / Azure | Claude, GPT-5.x family | Billed in consumption units at $0.01 each; US data-zone deployments carry the 1.1x multiplier | Microsoft stack, content filtering, residency |
| Google Vertex AI | Gemini family, Claude | Tracks Google list; regional and multi-region endpoints carry a 10% premium for Claude | GCP stack, grounding with Search |
Anthropic bills marketplace platforms in Claude Consumption Units, where 100 CCU equals $1.00 of usage at standard rates — the unit is an invoicing wrapper, not a discount or a markup. The per-token rate is rarely why you pick a platform; governance, data residency, and unified billing are. Plan Bedrock spend with our AWS Bedrock Pricing Calculator.
Self-Hosting Is Really a Utilisation Question
Self-hosting open weights (Llama, DeepSeek, Mistral) means renting or owning GPUs and serving with vLLM or llama.cpp. Hourly GPU rates move constantly across providers and regions, and this page does not quote rates it cannot verify against a vendor page — but the structural argument does not depend on the exact number.
The economics hinge on utilisation, because a GPU bills by the hour whether or not it is busy. At single-stream decode, a rented accelerator produces a modest token rate, and dividing an hourly rate by that throughput usually gives a per-token cost worse than an API. But decode is memory-bandwidth-bound and batches extremely well: continuous batching pushes aggregate throughput up by an order of magnitude under concurrent load, and the same hourly cost is then spread across far more tokens, dropping per-token cost to cents.
So the rule is: self-hosting wins on steady high volume, strict data isolation, or fine-tuned models. It loses on bursty or low traffic, where you pay for idle hardware. Don't eyeball it — model it with current quotes for the hardware you can actually get:
- Self-Hosted LLM Cost Calculator — API cost versus own-hardware break-even.
- What LLM Can I Run — GPU detection to model fit.
- LLM Inference Speed Calculator — estimate tokens per second from bandwidth, model, and quantisation.
- LLM GPU Benchmark — real in-browser bandwidth and tokens per second.
- LLM VRAM Calculator — size the memory you need.
For the full hardware, quantisation, and runtime picture, see Running Local AI: The Complete Guide.
Cost Optimisation Strategies
1. Model cascading. Route by complexity — cheap model first, escalate only on low confidence.
def select_model(complexity: str) -> str:
if complexity == "simple":
return "deepseek-v4-flash" # $0.14 / $0.28
elif complexity == "moderate":
return "claude-sonnet-5" # $2.00 / $10.00
return "gpt-5.6-sol" # $5.00 / $30.00
2. Prompt caching. Cache reads run at 0.1x base input. Structure prompts so the stable part comes first and the variable part last, or the cache never hits.
3. Batch processing. 50% off at Anthropic and Mistral, comparable at Google, for anything non-interactive.
4. Cap the reasoning budget. Thinking tokens bill as output — set effort per task tier rather than globally.
5. Measure tokens per task, not tokens per dollar. Tokenizer differences between vendors, and between model generations at the same vendor, mean the rate card is not the whole cost. Anthropic's own ~30% figure for Claude 4.7+ is the clearest example.
6. Watch the surcharge parameters. Data residency, regional endpoints, priority tiers, and fast mode all multiply your rate. Default to global and standard unless a requirement forces otherwise.
Combined, cascading plus caching plus batching routinely cuts 40-70% off a naive "flagship for everything" bill.
Conclusion
LLM costs span more than 100x between tiers, and over 600x once the pro models are included. The discipline:
- Default cheap. DeepSeek-V4-Flash, Gemini Flash-Lite, GPT-5-nano, or Mistral Small 4 for the bulk of traffic — while planning for DeepSeek's announced increase.
- Escalate selectively. GPT-5.6 Sol, Claude Opus 5, or Gemini 3.1 Pro only where quality measurably pays for itself.
- Budget reasoning tokens. They bill as output; turn effort down for easy work.
- Use the levers. Caching at 0.1x input and batching at 50% compound, and they are published, not negotiated.
- Compare on work, not tokens. Tokenizers differ by up to ~30% for identical text — benchmark real requests before switching vendors on price.
- Decide self-hosting with numbers, not vibes — run the break-even calculator.
The best model is the one that delivers the required quality at a cost you can sustain at scale — and the only way to know that cost is to measure it on your own traffic.