Compare text-to-speech and speech-to-text API costs: ElevenLabs, OpenAI, Google, Polly, Azure, Deepgram, AssemblyAI. Enter characters or audio hours.
This calculator turns a monthly volume into a monthly bill for every major speech API. For text to speech, enter characters, words or minutes of audio per month and it prices ElevenLabs (pay-as-you-go and every plan), OpenAI, Google Cloud Text-to-Speech, Amazon Polly, Azure AI Speech and Deepgram Aura. For speech to text, enter audio hours or minutes and it prices ElevenLabs Scribe, OpenAI Whisper and the gpt-4o transcribe models, Deepgram Nova-3, AssemblyAI, Google Speech-to-Text, Amazon Transcribe and Azure. The table is sorted cheapest first, the cheapest option is called out, and the URL updates as you type so you can share the exact comparison.
Prices are list prices in US dollars for US regions, taken from each provider’s own pricing page on October 8, 2026, with a link to each source under the table. Where a provider bills by token and publishes no per-character or per-minute figure (OpenAI’s gpt-4o-mini-tts and Google’s Gemini TTS), it is left out rather than estimated.
| Provider and voice type | Price per 1M characters | Always-free allowance |
|---|---|---|
| Google Standard, WaveNet | $4 | 4M characters a month |
| Amazon Polly Standard | $4 | 5M a month listed; not subtracted because AWS does not say it is permanent |
| OpenAI tts-1 | $15 | None |
| Azure Neural | $15 | Separate F0 tier only |
| Deepgram Aura-1 | $15 | One-off $200 signup credit |
| Google Neural2 | $16 | 1M characters a month |
| Amazon Polly Neural | $16 | 1M a month for the first 12 months only |
| Azure Neural HD | $22 | None |
| OpenAI tts-1-hd, Deepgram Aura-2, Polly Generative, Google Chirp 3: HD | $30 | Chirp 3: HD has 1M a month |
| ElevenLabs Flash v2.5 / v4 Turbo | $40 ($0.04 per 1K) | 20,000 characters on the Free plan |
| ElevenLabs v4 / v3 / Multilingual v2 | $80 ($0.08 per 1K) | 10,000 characters on the Free plan |
| Amazon Polly Long-Form, Google Studio | $100, $160 | Studio has 1M a month |
| Provider and model | Price per audio hour |
|---|---|
| AssemblyAI Universal-2 (pre-recorded) and Universal-Streaming | $0.15 |
| OpenAI gpt-4o-mini-transcribe, Google V2 dynamic batch, Azure batch | $0.18 |
| AssemblyAI Universal-3.5 Pro | $0.21 |
| ElevenLabs Scribe v2 | $0.22 |
| Deepgram Nova-3 pre-recorded (monolingual / multilingual) | $0.258 / $0.312 |
| OpenAI gpt-transcribe | $0.27 |
| OpenAI whisper-1 and gpt-4o-transcribe, Amazon Transcribe batch, Azure fast transcription | $0.36 |
| ElevenLabs Scribe v2 Realtime | $0.39 |
| Amazon Transcribe streaming | $0.60 |
| Google Speech-to-Text V2 standard | $0.96, falling to $0.24 above 2M minutes a month |
| Azure real-time | $1.00 |
Every text-to-speech API on this page bills by the character you send, and a character includes spaces, punctuation and newlines. Google and Polly also count SSML markup (Google excludes only <mark> tags), so a heavily tagged script costs more than its spoken text suggests. If you think in words, the calculator assumes 6 characters per English word including the space after it. If you think in finished audio, it assumes about 900 characters per minute, which is 150 spoken words a minute. Use the “count characters from your text” box to measure a real script instead: paste it, set how many times a month you generate something that size, and the count is used directly. The text never leaves your browser.
Speech-to-text is billed by audio duration, not by the length of the transcript. Two details change real bills: AssemblyAI and Google bill each audio channel separately, so a two-channel call recording costs twice as much; and AssemblyAI’s streaming product bills the whole WebSocket session including silence. Amazon Transcribe bills per second with no minimum, and Google rounds each request up to the nearest second.
The pattern holds at every volume: the cheapest voices are not the ones people choose ElevenLabs for. Price is one column; listen to samples of the voices you would actually ship before deciding on it.
ElevenLabs bills API usage in dollars at the same per-1,000-character rate on every plan, and each paid plan’s fee is prepaid usage at that rate: Starter at $6 includes 75,000 Multilingual characters, Creator at $22 includes 275,000, Pro at $99 includes 1,238,000, Scale at $299 includes 3,738,000 and Business at $990 includes 12,375,000. Flash and v4 Turbo cost half as much per character, so each plan covers twice as many characters on those models. For pure API spend, the cheapest choice is whichever plan’s allowance you will actually use up, and the calculator shows every plan so you can see where each one stops being wasted money. Choose a plan for its non-API features (seats, voice cloning, commercial licence) rather than for a discount it does not give.
Both SDKs below read the API key from an environment variable (ELEVENLABS_API_KEY and OPENAI_API_KEY). The ElevenLabs voice ID is the “George” voice used in the ElevenLabs quickstart; swap in any voice ID from your voice library.
ElevenLabs, Python (pip install elevenlabs). convert returns an iterator of audio byte chunks:
import os from elevenlabs.client import ElevenLabs
client = ElevenLabs(api_key=os.getenv("ELEVENLABS_API_KEY")) audio = client.text_to_speech.convert( text="Your order has shipped.", voice_id="JBFqnCBsd6RMkjVDRZzb", model_id="eleven_multilingual_v2", output_format="mp3_44100_128", ) with open("speech.mp3", "wb") as f: for chunk in audio: f.write(chunk)
ElevenLabs, TypeScript (npm install @elevenlabs/elevenlabs-js). textToSpeech.convert returns a web ReadableStream:
import { createWriteStream } from "node:fs"; import { Readable } from "node:stream"; import { pipeline } from "node:stream/promises"; import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
const client = new ElevenLabsClient(); // reads ELEVENLABS_API_KEY const audio = await client.textToSpeech.convert("JBFqnCBsd6RMkjVDRZzb", { text: "Your order has shipped.", modelId: "eleven_multilingual_v2", outputFormat: "mp3_44100_128", }); await pipeline(Readable.fromWeb(audio), createWriteStream("speech.mp3"));
OpenAI, Python (pip install openai):
from openai import OpenAI
client = OpenAI() with client.audio.speech.with_streaming_response.create( model="tts-1", voice="coral", input="Your order has shipped.", ) as response: response.stream_to_file("speech.mp3")
OpenAI, TypeScript (npm install openai):
import fs from "node:fs"; import OpenAI from "openai";
const openai = new OpenAI(); const mp3 = await openai.audio.speech.create({ model: "tts-1", voice: "coral", input: "Your order has shipped.", }); await fs.promises.writeFile("speech.mp3", Buffer.from(await mp3.arrayBuffer()));
Model IDs map to the price rows above: eleven_multilingual_v2, eleven_v3 and eleven_v4 bill at $0.08 per 1,000 characters, eleven_flash_v2_5 and eleven_v4_turbo at $0.04. On OpenAI, tts-1 is $15 per million characters and tts-1-hd is $30; gpt-4o-mini-tts accepts an instructions field for tone but bills per audio token. OpenAI’s usage policy requires telling listeners the voice is AI-generated.
Disclosure: the “Try ElevenLabs” links are referral links, and InventiveHQ may earn a commission if you subscribe. The ranking is by price only and every provider is priced the same way.
As of October 8, 2026: $4 for Google Standard or WaveNet and Amazon Polly Standard voices, $15 for OpenAI tts-1, Azure Neural and Deepgram Aura-1, $16 for Google Neural2 and Polly Neural, $30 for OpenAI tts-1-hd, Deepgram Aura-2, Polly Generative and Google Chirp 3: HD, $40 for ElevenLabs Flash and v4 Turbo, and $80 for ElevenLabs v4, v3 and Multilingual v2.
At list prices the cheapest per hour are AssemblyAI Universal-2 and Universal-Streaming at $0.15, then OpenAI gpt-4o-mini-transcribe, Google V2 dynamic batch and Azure batch at $0.18. Whisper and Amazon Transcribe batch are $0.36 an hour. Check accuracy on your own audio before choosing on price.
In US dollars per 1,000 characters at the same rate on every plan: $0.08 for v4, v3 and Multilingual v2, $0.04 for Flash and v4 Turbo. A paid plan's monthly fee is prepaid usage at that rate, so Pro at $99 covers 1,238,000 Multilingual characters. Scribe v2 transcription is $0.22 an hour.
About 900, assuming 150 spoken words a minute and 6 characters per word including spaces. Fast narration runs higher and slow, expressive reads lower. Paste a real script into the calculator to count characters exactly.
Yes. Providers bill every character sent, including spaces, punctuation and newlines. Google counts all SSML tags except mark, so heavily tagged input costs more than the spoken text alone.
Only always-free monthly allowances, such as Google's 4 million Standard and WaveNet characters, and only when the checkbox is on. Twelve-month trials such as Polly Neural's, and one-off signup credits from Deepgram, AssemblyAI and AWS, are not subtracted.
They were read from each provider's pricing page on October 8, 2026, and each provider's page is linked under the results table. Promotions with an end date are not used. Prices change, so confirm on the provider's page before you commit.
The Try ElevenLabs links are referral links, and InventiveHQ may earn a commission if you subscribe. The table is sorted by price only, and ElevenLabs is priced the same way as every other provider.