One hour of audio can cost $0.04 to transcribe on one speech to text API and $1.44 on another. Same recording, same words, a 36x gap on the invoice. That spread is the reason speech to text API pricing deserves a side-by-side look before you wire anything into production. This article converts every rate into the same unit, ranks the cheapest transcription APIs, and shows what each one costs at 100 and 1,000 audio hours a month. You will also see the add-on fees that quietly double a bill, where low-price models lose accuracy, and how to test three transcription models on PicassoIA without writing a line of code.
💡 Every rate below is a listed pay-as-you-go price from the vendor, not a PicassoIA price. Providers revise them often, so confirm the current number on the vendor's pricing page before you commit.
What Transcription Really Costs
Nearly every vendor bills for audio duration, not for word count or file size. That sounds simple until you compare quotes. Some publish dollars per minute, others per hour, and a few bury the useful number behind a billing rule. The first job is putting everything in one unit.
Per-Minute Versus Per-Hour Rates

The conversion is one line of arithmetic: multiply the per-minute rate by 60 to get the hourly cost. A rate of $0.003 per minute is $0.18 per hour. A rate of $0.024 per minute is $1.44 per hour. Vendors pick whichever unit makes the number look smallest, so do the conversion yourself before comparing anything.
Scale is where the decimals bite. A rate that reads $0.0025 per minute feels like nothing until you notice that a 500-hour archive is 30,000 minutes, and that same archive costs $75 on one provider and $720 on another.
Batch Versus Streaming Pricing
Transcribing a finished file (batch, sometimes called pre-recorded) is the cheap path. Live transcription (streaming) costs more because the provider has to hold GPU capacity open while you talk. The gap is not small:
- AssemblyAI charges $0.15 per hour for batch Universal-2 and $0.45 per hour for the Universal-3.5 Pro real-time model.
- Deepgram lists Nova-3 at roughly $0.0043 per minute for pre-recorded audio and about $0.0077 per minute for streaming, with a limited-time streaming promotion near $0.0048.
- Google Cloud drops to roughly $0.003 to $0.004 per minute on its discounted dynamic batch mode, against $0.016 for standard real-time requests.
If you transcribe recordings after the fact, never pay streaming rates. Upload the file and wait a few seconds. Reserve streaming for live captions, voice agents, and anything where a person waits on the words in real time.
Speech to Text API Pricing Table
Here is every major option converted to the same two units and sorted from cheapest to most expensive. Rates are for the first usage tier, English, no add-ons.
| Provider and model | Per hour | Per minute | Best fit |
|---|
| Groq Whisper Large v3 Turbo | $0.04 | ~$0.0007 | Huge backlogs, raw speed |
| AssemblyAI Universal-2 | $0.15 | $0.0025 | Low-cost batch with extras |
| GPT-4o Mini Transcribe | $0.18 | $0.003 | General purpose, tight budgets |
| AssemblyAI Universal-3.5 Pro | $0.21 | $0.0035 | Tougher audio at low cost |
| Deepgram Nova-3 (pre-recorded) | ~$0.26 | ~$0.0043 | Developer tooling, volume |
| GPT-4o Transcribe | $0.36 | $0.006 | Accents and messy audio |
| Google Cloud Speech-to-Text V2 | $0.96 | $0.016 | Google Cloud shops |
| Azure Speech (standard) | $1.00 | ~$0.0167 | Microsoft ecosystem |
| AWS Transcribe (Tier 1) | $1.44 | $0.024 | AWS-native pipelines |
💡 Whisper large v3 on OpenAI's own API sits at $0.006 per minute, the same as GPT-4o Transcribe. When the price is identical, the newer model is the safer pick for accents and background noise.
Cheapest Picks Under $0.25 an Hour

Four options sit below a quarter per hour of audio. Groq is the outlier: hosted Whisper Large v3 Turbo at $0.04 per hour runs hundreds of times faster than real time, so an hour of audio finishes in well under a minute. The trade-off is a leaner feature set, with fewer built-in extras than the specialist vendors.
AssemblyAI Universal-2 at $0.15 per hour is the cheapest option that still arrives with a full toolbox of optional features, and its Universal-3.5 Pro model adds accuracy for $0.06 more per hour. GPT-4o Mini Transcribe lands at $0.18 per hour and is the sweet spot for students, solo creators, and small teams who want strong accuracy without managing a pile of vendor accounts.
Big Cloud Providers Cost More

Google Cloud, Azure, and AWS charge between $0.96 and $1.44 per hour for standard transcription. That is 6 to 36 times the budget tier. You are not paying for better words. You are paying for compliance certifications, regional data residency, single-sign-on, and the convenience of one invoice that already exists in your finance system.
If your company already runs on one of those clouds and legal requires audio to stay inside it, the premium makes sense. If it does not, the premium is a tax on habit. Google's discounted batch mode is the exception, since it brings the cost back near the mid-tier at about $0.18 to $0.24 per hour.
Monthly Cost at Real Volumes
Per-minute rates only become real when you multiply them by your workload. This table uses 100 and 1,000 hours of audio per month, base rates only.
| Option | 100 hours a month | 1,000 hours a month |
|---|
| Groq Whisper Large v3 Turbo | $4 | $40 |
| AssemblyAI Universal-2 | $15 | $150 |
| GPT-4o Mini Transcribe | $18 | $180 |
| AssemblyAI Universal-3.5 Pro | $21 | $210 |
| Deepgram Nova-3 (pre-recorded) | ~$26 | ~$258 |
| GPT-4o Transcribe | $36 | $360 |
| Google Cloud V2 | $96 | $960 |
| Azure Speech | $100 | $1,000 |
| AWS Transcribe | $144 | $1,440 |
A Podcast Producer at 100 Hours
Picture a network publishing 50 two-hour episodes a month. That is 100 hours of audio. The cheapest option costs $4 and the most expensive $144, a difference of $140 a month. At that volume the invoice is not the problem. A single hour spent fixing a bad transcript costs more than the entire monthly bill on most of these options, so pick on accuracy and workflow first, and use price as the tiebreaker.
A Call Archive at 1,000 Hours
Now scale to a support team that records 1,000 hours of calls a month. The gap between the cheapest and most expensive option becomes $1,400 a month, or $16,800 a year. Mid-tier options between $150 and $360 a month now have real leverage, and that is the range where testing two or three vendors on your own audio pays for itself in a single afternoon.
Hidden Fees That Inflate Bills
The sticker price is the base transcript and nothing else. Real invoices are built from the base rate plus everything you switch on.
Feature Add-Ons Stack Up

Speaker labels, profanity filtering, redaction of personal data, language detection, sentiment tags, and summaries are usually metered separately. Typical charges look like this:
- Speaker labels (diarization): about $0.02 per hour on AssemblyAI, up to roughly $0.12 per hour elsewhere.
- Redaction: around $0.12 per hour in Deepgram's published examples.
- Per-feature fees on Azure: $0.30 per hour for each of language identification, diarization, and pronunciation assessment.
Switch on three or four of those and a $0.15 base rate can double or more. Always price the exact feature set your product needs, not the headline number.
💡 Free tiers help you test before paying. Deepgram offers a $200 starting credit, AssemblyAI gives a free credit allowance, Google includes 60 free minutes a month, and AWS includes 60 free minutes a month for the first year.
Billing rules matter just as much as add-ons. AWS Transcribe bills a minimum of 15 seconds per request, so a voice-command app sending 3-second clips pays for five times the audio it actually sends. Upload limits matter too: OpenAI's transcription endpoints accept files up to about 25 MB, so long recordings need chunking, and overlapping chunks means you pay twice for the overlapped seconds.
Engineering Time Counts Too

Saving $30 a month on transcription is a bad trade if the cheaper vendor costs your developer four extra hours. Check for an SDK in your language, webhooks for finished jobs, sensible rate limits, and predictable error messages before you commit. At a $75 hourly engineering cost, four hours of integration pain equals ten months of a $30 savings.
Cheap Versus Accurate
A price per hour says nothing about how many words come out wrong. Two vendors at the same rate can leave you with very different amounts of cleanup, and cleanup is paid in human time.
Where Cheap Models Stumble

Low-cost models tend to slip in the same places: heavy accents, two people talking over each other, field recordings with wind or machinery, product names, acronyms, and spoken numbers. Each slip becomes an edit.
Run the math on editing. At a $30 hourly rate, an editor costs $0.50 a minute. Every extra 10 minutes of cleanup per audio hour adds $5, which is more than the whole AWS Transcribe bill for that hour. A transcript that costs $0.18 and needs 20 minutes of fixing is more expensive than one that costs $0.36 and needs 5.
The only reliable test is your own audio. Pick three representative files, run each through three models, and count the edits you make. Do this before you decide anything based on a benchmark chart.
Meetings and Speaker Labels

Meetings are the hardest everyday case because the transcript is useless without knowing who said what. Some models include speaker labels in the base price, such as OpenAI's diarization variant of its transcribe model, while others meter them as an add-on. When you compare a meeting-focused workload, add the diarization fee to every option on the shortlist, then compare totals. The cheapest base rate frequently stops being the cheapest total.
How to Use GPT-4o Mini Transcribe

PicassoIA hosts three speech-to-text models you can try straight from the browser: GPT-4o Mini Transcribe, GPT-4o Transcribe, and Gemini 3 Pro. That makes it a quick way to run the three-file test from the previous section before you open a single vendor account.
Step by Step Setup
- Open the GPT-4o Mini Transcribe page on PicassoIA.
- Upload a short audio file, ideally a 3 to 5 minute clip that includes your hardest material: names, numbers, and overlapping voices.
- If the form offers a language field, set it instead of leaving the model to guess.
- Add a short prompt with unusual names or jargon if a prompt field is available.
- Run the transcription and read the output against the audio.
- Repeat with GPT-4o Transcribe and Gemini 3 Pro on the same clip.
- Count the lines you would edit in each result. The lowest count wins your shortlist.
Parameter Tips That Save Money
- Trim silence first. Billing follows duration, so cutting a 10-minute dead intro out of a recording saves 10 minutes of cost on every provider.
- Set the language. Language detection can be an extra fee on some APIs and an error source on short clips.
- Use a vocabulary prompt. A single line listing brand names and technical terms fixes more mistakes than moving to a pricier model.
- Test on identical clips. Compare models on the same file, otherwise differences in audio quality get mistaken for differences in model quality.
Turn Transcripts Into Speech and Music
Transcription is usually the first step in a longer pipeline. Once you have clean text, you can repurpose it: narrate it in a new voice, translate and re-record it, or build a soundtrack around it. PicassoIA keeps those steps in one place.
Re-Voice Transcripts With Natural Speech

Feed a corrected transcript to a text-to-speech model and you get a clean narration track without booking a studio. Speech 2.8 HD targets studio-quality voiceovers, ElevenLabs v3 focuses on expressive delivery, Gemini 3.1 Flash TTS offers 30 voices across more than 70 languages, and Realtime TTS 2 is built for low-latency use. This is the practical loop for podcasters who want a second-language edition of every episode.
Add a Music Bed to the Result
Narration sounds finished once it has a quiet music bed under it. Music 2.6 generates full songs from a text prompt, Lyria 3 Pro builds longer instrumental pieces, and Stable Audio 2.5 is a good fit for short loops and ambient beds. Describe the mood in one sentence, keep the track instrumental, and drop it 15 to 20 decibels under the voice.
Pick Your Cheapest Setup
Here is the shortcut version of everything above:
- Backlog of old recordings, price is everything: Groq Whisper Large v3 Turbo at $0.04 per hour.
- Small team, balanced quality and cost: GPT-4o Mini Transcribe at $0.18 per hour.
- Accents, noise, and tough audio: GPT-4o Transcribe at $0.36 per hour, or AssemblyAI Universal-3.5 Pro at $0.21.
- Meetings with speaker labels: price the diarization add-on first, then compare totals.
- Locked into one cloud by policy: Google, Azure, or AWS, and ask for the discounted batch mode.
Whatever you choose, treat the first month as a test. Log the real invoice, divide it by the hours transcribed, and compare it with the table above. Prices move, and so do your needs. A vendor that wins today on cost can lose next quarter on accuracy or limits, so keep your integration thin enough to swap providers with a one-line change, and keep a 5-minute test clip on hand to re-run whenever a rate changes.
Ready to put it into practice? Open PicassoIA, upload one of your own recordings to GPT-4o Mini Transcribe, then try the same clip on the other transcription models. When the text looks right, narrate it with a text-to-speech voice, add a music bed, and see how far one recording can go. Browse the full catalog at picassoia.com/en/all-models and start experimenting today.