AI & Machine Learning

LLM API Pricing Comparison: Every Model's Token Cost (Verified August 2026)

Verified per-million-token input and output prices for OpenAI GPT-5.6, Claude Opus 5, Gemini 3.x, DeepSeek V4, and Mistral — checked against each vendor's own pricing page on 2026-08-13, with the tokenizer difference that makes headline rates misleading.

By Inventive HQ Team

Before you compare model prices, you need to know how many tokens your actual prompts use — that is what every per-million-token rate below gets multiplied by. Paste a real prompt into the token counter and turn these rates into a dollar estimate for your workload.

Loading interactive tool...

Every price on this page was verified on 2026-08-13 against the vendor's own pricing page: OpenAI, Anthropic, Google, DeepSeek, and Mistral. All figures are USD per 1M tokens, standard non-batch non-cached tier, unless stated. Where a vendor does not publish a number, this page says so rather than guessing — several widely-shared LLM pricing tables currently list models that were retired more than a year ago.

The Thing Most Pricing Comparisons Get Wrong

Every comparison table on the internet, including this one, lists a price per million tokens. That number is only half of a cost, because the other half is how many tokens the vendor's tokenizer produces from your text — and that is not constant across vendors.

Anthropic states plainly in its own pricing documentation that Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text than earlier models. Claude Sonnet 4.6 and earlier use the previous tokenizer.

The consequence is significant and almost universally ignored:

ComparisonSticker output priceTokens for the same textEffective cost for the same work
Claude Sonnet 4.6$15.00baseline$15.00
Claude Sonnet 5$10.00~1.3x~$13.00
Claude Opus 5$25.00~1.3x~$32.50

Sonnet 5's sticker price is 33% below Sonnet 4.6, but on the same input text the real saving is closer to 13%. The exact increase depends on content and workload shape — Anthropic does not publish a per-content figure — so treat the 1.3x as the vendor's own stated approximation, not a precise multiplier.

Practical rule: compare vendors on measured cost per completed task, not on the rate card. Run the same 20 representative requests through each candidate and compare invoices. Everything below is the rate card; it is necessary but not sufficient.

Cross-Provider Pricing (per 1M tokens)

Frontier tier

ModelProviderInputCached inputOutput
GPT-5.5-pro / GPT-5.4-proOpenAI$30.00$180.00
GPT-5-proOpenAI$15.00$120.00
Claude Fable 5Anthropic$10.00$1.00$50.00
GPT-5.6 SolOpenAI$5.00$0.50$30.00
GPT-5.5OpenAI$5.00$0.50$30.00
Claude Opus 5Anthropic$5.00$0.50$25.00
Claude Opus 4.8Anthropic$5.00$0.50$25.00

Mid tier (best balance)

ModelProviderInputCached inputOutput
GPT-5.4OpenAI$2.50$0.25$15.00
Claude Sonnet 4.6Anthropic$3.00$0.30$15.00
GPT-5.6 TerraOpenAI$2.00$0.20$12.00
Gemini 3.1 ProGoogle$2.00$12.00
Claude Sonnet 5Anthropic$2.00$0.20$10.00
GPT-5.2 / GPT-5.3-CodexOpenAI$1.75$0.175$14.00
Gemini 3.5 FlashGoogle$1.50$9.00
Gemini 3.6 FlashGoogle$1.50$7.50
Mistral Medium 3.5Mistral$1.50$7.50
Gemini 2.5 ProGoogle$1.25$10.00
GPT-5.1 / GPT-5OpenAI$1.25$0.125$10.00

Budget tier (cost-optimised)

ModelProviderInputCached inputOutput
Claude Haiku 4.5Anthropic$1.00$0.10$5.00
GPT-5.4-miniOpenAI$0.75$0.075$4.50
DeepSeek-V4-ProDeepSeek$0.435$0.003625$0.87
Mistral Large 3Mistral$0.50$1.50
Gemini 3.5 Flash-LiteGoogle$0.30$2.50
Gemini 2.5 FlashGoogle$0.30$2.50
Mistral CodestralMistral$0.30$0.90
GPT-5.4-nanoOpenAI$0.20$0.02$1.25
GPT-5.6 LunaOpenAI$0.20$0.02$1.20
GPT-5-miniOpenAI$0.25$0.025$2.00
Mistral Small 4Mistral$0.15$0.60
DeepSeek-V4-FlashDeepSeek$0.14$0.0028$0.28
Gemini 2.5 Flash-LiteGoogle$0.10$0.40
Mistral Ministral 3 (3B)Mistral$0.10$0.10
GPT-5-nanoOpenAI$0.05$0.005$0.40

Notes on the table. Google tiers several models by context length: Gemini 3.1 Pro charges $2.00/$12.00 up to 200,000 tokens and $4.00/$18.00 above it; Gemini 2.5 Pro charges $1.25/$10.00 up to 200,000 tokens and $2.50/$15.00 above. Google also prices audio input separately on some models ($1.00 on Gemini 2.5 Flash, $0.50 on Gemini 3.1 Flash-Lite). Meta's Llama models have no first-party per-token price because they are open-weight — costs come from whichever host you use, and this page does not quote host prices it has not verified.

Output price per 1M tokens (USD), verified 2026-08-13 Output price per 1M tokens (USD) — a >100x spread $0 $10 $20 $30 GPT-5.6 Sol $30.00
<text x="182" y="96" text-anchor="end">Claude Opus 5</text>
<rect x="190" y="84" width="475" height="20" fill="#2813e8"/>
<text x="671" y="99">$25.00</text>

<text x="182" y="130" text-anchor="end">GPT-5.4</text>
<rect x="190" y="118" width="285" height="20" fill="#2813e8"/>
<text x="481" y="133">$15.00</text>

<text x="182" y="164" text-anchor="end">Gemini 3.1 Pro</text>
<rect x="190" y="152" width="228" height="20" fill="#2813e8"/>
<text x="424" y="167">$12.00</text>

<text x="182" y="198" text-anchor="end">Claude Sonnet 5</text>
<rect x="190" y="186" width="190" height="20" fill="#2813e8"/>
<text x="386" y="201">$10.00</text>

<text x="182" y="232" text-anchor="end">Gemini 3.6 Flash</text>
<rect x="190" y="220" width="143" height="20" fill="#2813e8"/>
<text x="339" y="235">$7.50</text>

<text x="182" y="266" text-anchor="end">Claude Haiku 4.5</text>
<rect x="190" y="254" width="95" height="20" fill="#2813e8"/>
<text x="291" y="269">$5.00</text>

<text x="182" y="300" text-anchor="end">GPT-5.6 Luna</text>
<rect x="190" y="288" width="23" height="20" fill="#2813e8"/>
<text x="219" y="303">$1.20</text>

<text x="182" y="334" text-anchor="end">DeepSeek-V4-Flash</text>
<rect x="190" y="322" width="5" height="20" fill="#2813e8"/>
<text x="201" y="337">$0.28</text>

The spread between the most expensive output on this chart ($30.00) and the cheapest ($0.28) is over 100x — and the pro tiers reach $180.00, a 640x spread. That gap, not the headline model name, is where your bill is won or lost.

Two Pricing Changes Worth Knowing About

Pricing pages are not static, and two current vendor statements materially affect planning.

Anthropic cancelled a scheduled Sonnet 5 price rise. Sonnet 5 launched at $2.00/$10.00 as introductory pricing through 2026-08-31, with a scheduled increase to $3.00/$15.00 on 2026-09-01. Anthropic's pricing documentation now states that this is the standard price and the increase will not occur. If you built a budget around Sonnet 5 rising by 50% in September, you can release that reserve.

DeepSeek has announced it will raise prices. DeepSeek's own pricing documentation carries a notice that it plans to raise overall API pricing "in the near future, with a significant increase expected." DeepSeek is currently the price floor at $0.14/$0.28, so anyone architecting around that floor should treat it as temporary and keep a migration path to Gemini Flash-Lite, GPT-5-nano, or Mistral Small 4.

Reasoning Models and the Thinking-Token Trap

Every flagship now has a thinking mode — OpenAI's reasoning effort setting, Claude's extended thinking, Gemini's thinking levels. Those reasoning tokens bill as output tokens, even though you never see them.

A model that thinks for 4,000 tokens before emitting a 200-token answer bills roughly 4,200 output tokens. On GPT-5.6 Sol at $30.00 output that single call costs about $0.13, versus about $0.006 for the visible answer alone — a 21x difference invisible on the rate card.

Practical rules:

  • Default to minimal or low effort for extraction, classification, routing, and simple chat.
  • Reserve high effort for genuinely hard reasoning, multi-step agentic work, and tricky debugging.
  • Measure it. Reasoning token counts appear in the usage object of every provider's response — log them, because they are the line item most likely to blow a budget.
Advertisement

The Discount Levers, With Real Multipliers

These are published, verified, and stack. They matter more than model choice for repetitive workloads.

Prompt caching. Anthropic publishes exact multipliers relative to base input: a 5-minute cache write costs 1.25x, a 1-hour write 2x, and a cache read costs 0.1x. That means a 5-minute cache pays for itself after a single read, and a 1-hour cache after two. OpenAI's published cached-input rates are likewise 10% of standard across the GPT-5.x line — Sol drops from $5.00 to $0.50, Terra from $2.00 to $0.20. Mistral advertises up to 90% off cached input. For any application with a large fixed system prompt or shared document context, this is the single biggest lever available.

Batch processing. Anthropic's Batch API takes 50% off both input and output — Opus 5 falls to $2.50/$12.50, Sonnet 5 to $1.00/$5.00, Haiku 4.5 to $0.50/$2.50. Google shows comparable reductions for Batch and Flex tiers, and Mistral advertises 50%. If a job does not need an answer in the next few seconds, it should not be paying interactive rates.

Context-window pricing. Anthropic includes the full 1M-token window at standard pricing on Claude 4.6 and later — a 900,000-token request bills at the same per-token rate as a 9,000-token one, and caching and batch discounts apply across the whole window. Google instead steps prices above 200,000 tokens. If your workload is genuinely long-context, that structural difference can outweigh the headline rate.

Watch the surcharges. Anthropic applies a 1.1x multiplier when you pin inference to US-only via the data-residency parameter, and a 10% premium for regional or multi-region endpoints on Bedrock and Google Cloud versus global endpoints. Google's Priority tier costs roughly 80% more than standard. Anthropic's Fast mode for Opus 5 and Opus 4.8 is priced at $10.00/$50.00, double the standard rate. None of these appear in comparison tables, and all of them appear on invoices.

Cost Analysis by Use Case

All figures below are computed from the verified rates above.

Chatbot / conversational AI

Typical turn: 800 input, 400 output tokens.

ModelCost / 1,000 turnsMonthly (100K turns)
DeepSeek-V4-Flash$0.22$22
Gemini 2.5 Flash-Lite$0.24$24
GPT-5.6 Luna$0.64$64
Claude Sonnet 5$5.60$560
GPT-5.4$8.00$800

Document summarisation

Typical: 10,000 input, 500 output tokens.

ModelCost / documentMonthly (10K docs)
DeepSeek-V4-Flash$0.0015$15
Gemini 2.5 Flash$0.0043$43
Claude Haiku 4.5$0.0125$125
Claude Sonnet 5$0.0250$250
GPT-5.4$0.0325$325

Summarisation is input-heavy, which is exactly the shape that prompt caching rewards least (the document changes every time) and long-context pricing rewards most. Note how Claude Sonnet 5 beats GPT-5.4 here on raw rates — then remember the tokenizer adjustment above, which narrows that gap.

Code generation

Typical: 2,000 input, 1,000 output tokens.

ModelCost / requestMonthly (50K req)
Mistral Codestral$0.0015$75
Claude Sonnet 5$0.0140$700
GPT-5.3-Codex$0.0175$875
GPT-5.4$0.0200$1,000
Claude Opus 5$0.0350$1,750

Code generation is output-heavy, so output rates dominate — which is why the 5-6x input-to-output ratio matters more here than anywhere else.

Managed Platforms: Bedrock, Azure, Vertex

PlatformModelsPricing vs directUse when
AWS BedrockClaude, Llama, Mistral, othersTracks list price; 10% premium on regional vs global endpoints for Claude 4.5+In AWS, need VPC endpoints or compliance
Microsoft Foundry / AzureClaude, GPT-5.x familyBilled in consumption units at $0.01 each; US data-zone deployments carry the 1.1x multiplierMicrosoft stack, content filtering, residency
Google Vertex AIGemini family, ClaudeTracks Google list; regional and multi-region endpoints carry a 10% premium for ClaudeGCP stack, grounding with Search

Anthropic bills marketplace platforms in Claude Consumption Units, where 100 CCU equals $1.00 of usage at standard rates — the unit is an invoicing wrapper, not a discount or a markup. The per-token rate is rarely why you pick a platform; governance, data residency, and unified billing are. Plan Bedrock spend with our AWS Bedrock Pricing Calculator.

Self-Hosting Is Really a Utilisation Question

Self-hosting open weights (Llama, DeepSeek, Mistral) means renting or owning GPUs and serving with vLLM or llama.cpp. Hourly GPU rates move constantly across providers and regions, and this page does not quote rates it cannot verify against a vendor page — but the structural argument does not depend on the exact number.

The economics hinge on utilisation, because a GPU bills by the hour whether or not it is busy. At single-stream decode, a rented accelerator produces a modest token rate, and dividing an hourly rate by that throughput usually gives a per-token cost worse than an API. But decode is memory-bandwidth-bound and batches extremely well: continuous batching pushes aggregate throughput up by an order of magnitude under concurrent load, and the same hourly cost is then spread across far more tokens, dropping per-token cost to cents.

So the rule is: self-hosting wins on steady high volume, strict data isolation, or fine-tuned models. It loses on bursty or low traffic, where you pay for idle hardware. Don't eyeball it — model it with current quotes for the hardware you can actually get:

For the full hardware, quantisation, and runtime picture, see Running Local AI: The Complete Guide.

Cost Optimisation Strategies

1. Model cascading. Route by complexity — cheap model first, escalate only on low confidence.

def select_model(complexity: str) -> str:
    if complexity == "simple":
        return "deepseek-v4-flash"   # $0.14 / $0.28
    elif complexity == "moderate":
        return "claude-sonnet-5"     # $2.00 / $10.00
    return "gpt-5.6-sol"             # $5.00 / $30.00

2. Prompt caching. Cache reads run at 0.1x base input. Structure prompts so the stable part comes first and the variable part last, or the cache never hits.

3. Batch processing. 50% off at Anthropic and Mistral, comparable at Google, for anything non-interactive.

4. Cap the reasoning budget. Thinking tokens bill as output — set effort per task tier rather than globally.

5. Measure tokens per task, not tokens per dollar. Tokenizer differences between vendors, and between model generations at the same vendor, mean the rate card is not the whole cost. Anthropic's own ~30% figure for Claude 4.7+ is the clearest example.

6. Watch the surcharge parameters. Data residency, regional endpoints, priority tiers, and fast mode all multiply your rate. Default to global and standard unless a requirement forces otherwise.

Combined, cascading plus caching plus batching routinely cuts 40-70% off a naive "flagship for everything" bill.

Conclusion

LLM costs span more than 100x between tiers, and over 600x once the pro models are included. The discipline:

  1. Default cheap. DeepSeek-V4-Flash, Gemini Flash-Lite, GPT-5-nano, or Mistral Small 4 for the bulk of traffic — while planning for DeepSeek's announced increase.
  2. Escalate selectively. GPT-5.6 Sol, Claude Opus 5, or Gemini 3.1 Pro only where quality measurably pays for itself.
  3. Budget reasoning tokens. They bill as output; turn effort down for easy work.
  4. Use the levers. Caching at 0.1x input and batching at 50% compound, and they are published, not negotiated.
  5. Compare on work, not tokens. Tokenizers differ by up to ~30% for identical text — benchmark real requests before switching vendors on price.
  6. Decide self-hosting with numbers, not vibes — run the break-even calculator.

The best model is the one that delivers the required quality at a cost you can sustain at scale — and the only way to know that cost is to measure it on your own traffic.

Frequently Asked Questions

Which LLM API is cheapest in 2026?

Among first-party APIs verified on 2026-08-13, the cheapest options are GPT-5-nano ($0.05 input / $0.40 output per million tokens), Mistral Ministral 3 3B ($0.10/$0.10), Gemini 2.5 Flash-Lite ($0.10/$0.40), DeepSeek-V4-Flash ($0.14/$0.28), and Mistral Small 4 ($0.15/$0.60). Which one is actually cheapest for you depends on your input-to-output ratio: Ministral 3 wins on output-heavy work, GPT-5-nano on input-heavy work. Note that DeepSeek has publicly stated it plans to raise prices significantly.

Do headline per-token prices actually compare across vendors?

Not directly, and this is the most overlooked cost factor in 2026. Anthropic states that Claude 4.7 and later models use a newer tokenizer that produces roughly 30% more tokens for the same text than earlier models. That means the same document costs about 30% more to process than the sticker price implies relative to a model on an older tokenizer. Compare vendors on cost per unit of work, not cost per token.

Is a flagship model worth the higher cost?

Only for work that needs it. GPT-5.6 Sol ($5.00/$30.00) and Claude Opus 5 ($5.00/$25.00) sit at the top, while Claude Sonnet 5 ($2.00/$10.00), GPT-5.6 Terra ($2.00/$12.00), and Gemini 3.1 Pro ($2.00/$12.00) deliver most of the capability at roughly a third of the output cost. Because output tokens run 5-6x input across nearly every provider, routing the bulk of traffic to a mid tier and reserving the flagship for the hard 10-20% of requests is where the savings are.

What's the difference between input and output token pricing?

Input tokens are read in one parallel pass and are cheap; output tokens are generated one at a time and cost 5-6x more almost everywhere. Claude Opus 5 is exactly 5x ($5 in, $25 out), GPT-5.6 Sol and Terra are 6x. Critically, reasoning or thinking tokens bill as output even though you never see them, so a reasoning-heavy call can cost far more than its visible answer suggests.

How are reasoning / thinking tokens billed?

As output tokens, across every provider — OpenAI's reasoning effort setting, Claude's extended thinking, and Gemini's thinking levels all bill the hidden reasoning trace at the output rate. A model thinking for 4,000 tokens before a 200-token answer bills roughly 4,200 output tokens. On GPT-5.6 Sol at $30 per million output tokens that single call costs about $0.13 versus $0.006 for the visible answer alone. Turn effort down for simple tasks.

How much do prompt caching and batching actually save?

A lot, and they stack. Anthropic prices cache reads at 0.1x base input (a 5-minute cache write costs 1.25x, a 1-hour write 2x), so caching pays for itself after a single read. OpenAI's published cached-input rates are also 10% of standard — GPT-5.6 Sol drops from $5.00 to $0.50. The Batch API takes 50% off both input and output at Anthropic, and Google shows similar reductions for Batch and Flex. Combined, caching plus batching routinely halves a repetitive workload's bill.

Does Claude's 1M context window cost extra?

No. Anthropic states that Claude 4.6 and later models include the full 1M-token context window at standard pricing — a 900,000-token request bills at the same per-token rate as a 9,000-token one, and caching and batch discounts apply across the full window. Google tiers by context instead: Gemini 3.1 Pro charges $2.00 input up to 200,000 tokens and $4.00 above it, with output going from $12.00 to $18.00.

Should I use AWS Bedrock or direct API access?

Managed platforms usually track the model's list price but add convenience and governance. Note two real cost details: Anthropic applies a 10% premium for regional and multi-region endpoints on Bedrock and Google Cloud versus global endpoints, and a 1.1x multiplier on the first-party API when you pin inference to US-only. Use a managed platform when you need enterprise security, data residency, or unified billing; use direct APIs for the newest model versions and the lowest friction.

How do I estimate monthly LLM API costs?

Use (avg input tokens x input price + avg output tokens x output price) x requests per month, all per million tokens. Example on GPT-5.4 at $2.50/$15.00 with 1,000-token inputs and 500-token outputs across 100,000 requests: (1,000 x $2.50 + 500 x $15.00) / 1,000,000 x 100,000 = $1,000/month. Add reasoning tokens to the output side if the model thinks, and adjust for tokenizer differences between vendors. Our LLM token counter does this across many models at once.

Are open-source LLMs really free?

The weights are free to download; serving them is not. Self-hosting means paying for GPU time, engineering, and operations, and the economics hinge almost entirely on utilisation — an idle GPU still bills by the hour, while batched serving spreads that same hourly cost across thousands of concurrent tokens. For low or bursty volume an API almost always wins. Self-hosting wins on steady high volume, strict data isolation, or fine-tuned models, which is exactly what a break-even calculation should decide.

llmapi-pricinggpt-5claudellamageminideepseekai-costs