Transcribe audioGenerate speechGenerate music

Speech to Text API Pricing: Cheapest Transcription APIs Compared

A side-by-side look at speech to text API pricing, with every rate converted to cost per hour, from $0.04 to $1.44. See monthly bills at 100 and 1,000 audio hours, the add-on fees that inflate invoices, and how to test accuracy on your own recordings first.

Speech to Text API Pricing: Cheapest Transcription APIs Compared
Cristian Da Conceicao
Founder of Picasso IA

One hour of audio can cost $0.04 to transcribe on one speech to text API and $1.44 on another. Same recording, same words, a 36x gap on the invoice. That spread is the reason speech to text API pricing deserves a side-by-side look before you wire anything into production. This article converts every rate into the same unit, ranks the cheapest transcription APIs, and shows what each one costs at 100 and 1,000 audio hours a month. You will also see the add-on fees that quietly double a bill, where low-price models lose accuracy, and how to test three transcription models on PicassoIA without writing a line of code.

💡 Every rate below is a listed pay-as-you-go price from the vendor, not a PicassoIA price. Providers revise them often, so confirm the current number on the vendor's pricing page before you commit.

What Transcription Really Costs

Nearly every vendor bills for audio duration, not for word count or file size. That sounds simple until you compare quotes. Some publish dollars per minute, others per hour, and a few bury the useful number behind a billing rule. The first job is putting everything in one unit.

Per-Minute Versus Per-Hour Rates

Brushed steel stopwatch next to studio headphones on a dark oak desk

The conversion is one line of arithmetic: multiply the per-minute rate by 60 to get the hourly cost. A rate of $0.003 per minute is $0.18 per hour. A rate of $0.024 per minute is $1.44 per hour. Vendors pick whichever unit makes the number look smallest, so do the conversion yourself before comparing anything.

Scale is where the decimals bite. A rate that reads $0.0025 per minute feels like nothing until you notice that a 500-hour archive is 30,000 minutes, and that same archive costs $75 on one provider and $720 on another.

Batch Versus Streaming Pricing

Transcribing a finished file (batch, sometimes called pre-recorded) is the cheap path. Live transcription (streaming) costs more because the provider has to hold GPU capacity open while you talk. The gap is not small:

  • AssemblyAI charges $0.15 per hour for batch Universal-2 and $0.45 per hour for the Universal-3.5 Pro real-time model.
  • Deepgram lists Nova-3 at roughly $0.0043 per minute for pre-recorded audio and about $0.0077 per minute for streaming, with a limited-time streaming promotion near $0.0048.
  • Google Cloud drops to roughly $0.003 to $0.004 per minute on its discounted dynamic batch mode, against $0.016 for standard real-time requests.

If you transcribe recordings after the fact, never pay streaming rates. Upload the file and wait a few seconds. Reserve streaming for live captions, voice agents, and anything where a person waits on the words in real time.

Speech to Text API Pricing Table

Here is every major option converted to the same two units and sorted from cheapest to most expensive. Rates are for the first usage tier, English, no add-ons.

Provider and modelPer hourPer minuteBest fit
Groq Whisper Large v3 Turbo$0.04~$0.0007Huge backlogs, raw speed
AssemblyAI Universal-2$0.15$0.0025Low-cost batch with extras
GPT-4o Mini Transcribe$0.18$0.003General purpose, tight budgets
AssemblyAI Universal-3.5 Pro$0.21$0.0035Tougher audio at low cost
Deepgram Nova-3 (pre-recorded)~$0.26~$0.0043Developer tooling, volume
GPT-4o Transcribe$0.36$0.006Accents and messy audio
Google Cloud Speech-to-Text V2$0.96$0.016Google Cloud shops
Azure Speech (standard)$1.00~$0.0167Microsoft ecosystem
AWS Transcribe (Tier 1)$1.44$0.024AWS-native pipelines

💡 Whisper large v3 on OpenAI's own API sits at $0.006 per minute, the same as GPT-4o Transcribe. When the price is identical, the newer model is the safer pick for accents and background noise.

Cheapest Picks Under $0.25 an Hour

University student with earbuds studying at a library table

Four options sit below a quarter per hour of audio. Groq is the outlier: hosted Whisper Large v3 Turbo at $0.04 per hour runs hundreds of times faster than real time, so an hour of audio finishes in well under a minute. The trade-off is a leaner feature set, with fewer built-in extras than the specialist vendors.

AssemblyAI Universal-2 at $0.15 per hour is the cheapest option that still arrives with a full toolbox of optional features, and its Universal-3.5 Pro model adds accuracy for $0.06 more per hour. GPT-4o Mini Transcribe lands at $0.18 per hour and is the sweet spot for students, solo creators, and small teams who want strong accuracy without managing a pile of vendor accounts.

Big Cloud Providers Cost More

Long data center aisle lined with black server racks

Google Cloud, Azure, and AWS charge between $0.96 and $1.44 per hour for standard transcription. That is 6 to 36 times the budget tier. You are not paying for better words. You are paying for compliance certifications, regional data residency, single-sign-on, and the convenience of one invoice that already exists in your finance system.

If your company already runs on one of those clouds and legal requires audio to stay inside it, the premium makes sense. If it does not, the premium is a tax on habit. Google's discounted batch mode is the exception, since it brings the cost back near the mid-tier at about $0.18 to $0.24 per hour.

Monthly Cost at Real Volumes

Per-minute rates only become real when you multiply them by your workload. This table uses 100 and 1,000 hours of audio per month, base rates only.

Option100 hours a month1,000 hours a month
Groq Whisper Large v3 Turbo$4$40
AssemblyAI Universal-2$15$150
GPT-4o Mini Transcribe$18$180
AssemblyAI Universal-3.5 Pro$21$210
Deepgram Nova-3 (pre-recorded)~$26~$258
GPT-4o Transcribe$36$360
Google Cloud V2$96$960
Azure Speech$100$1,000
AWS Transcribe$144$1,440

A Podcast Producer at 100 Hours

Picture a network publishing 50 two-hour episodes a month. That is 100 hours of audio. The cheapest option costs $4 and the most expensive $144, a difference of $140 a month. At that volume the invoice is not the problem. A single hour spent fixing a bad transcript costs more than the entire monthly bill on most of these options, so pick on accuracy and workflow first, and use price as the tiebreaker.

A Call Archive at 1,000 Hours

Now scale to a support team that records 1,000 hours of calls a month. The gap between the cheapest and most expensive option becomes $1,400 a month, or $16,800 a year. Mid-tier options between $150 and $360 a month now have real leverage, and that is the range where testing two or three vendors on your own audio pays for itself in a single afternoon.

Hidden Fees That Inflate Bills

The sticker price is the base transcript and nothing else. Real invoices are built from the base rate plus everything you switch on.

Feature Add-Ons Stack Up

Accountant's hands reviewing a printed invoice beside a calculator

Speaker labels, profanity filtering, redaction of personal data, language detection, sentiment tags, and summaries are usually metered separately. Typical charges look like this:

  • Speaker labels (diarization): about $0.02 per hour on AssemblyAI, up to roughly $0.12 per hour elsewhere.
  • Redaction: around $0.12 per hour in Deepgram's published examples.
  • Per-feature fees on Azure: $0.30 per hour for each of language identification, diarization, and pronunciation assessment.

Switch on three or four of those and a $0.15 base rate can double or more. Always price the exact feature set your product needs, not the headline number.

💡 Free tiers help you test before paying. Deepgram offers a $200 starting credit, AssemblyAI gives a free credit allowance, Google includes 60 free minutes a month, and AWS includes 60 free minutes a month for the first year.

Billing rules matter just as much as add-ons. AWS Transcribe bills a minimum of 15 seconds per request, so a voice-command app sending 3-second clips pays for five times the audio it actually sends. Upload limits matter too: OpenAI's transcription endpoints accept files up to about 25 MB, so long recordings need chunking, and overlapping chunks means you pay twice for the overlapped seconds.

Engineering Time Counts Too

Developer at a standing desk working on code in a brick loft

Saving $30 a month on transcription is a bad trade if the cheaper vendor costs your developer four extra hours. Check for an SDK in your language, webhooks for finished jobs, sensible rate limits, and predictable error messages before you commit. At a $75 hourly engineering cost, four hours of integration pain equals ten months of a $30 savings.

Cheap Versus Accurate

A price per hour says nothing about how many words come out wrong. Two vendors at the same rate can leave you with very different amounts of cleanup, and cleanup is paid in human time.

Where Cheap Models Stumble

Radio journalist holding a handheld recorder toward a craftsman in a workshop

Low-cost models tend to slip in the same places: heavy accents, two people talking over each other, field recordings with wind or machinery, product names, acronyms, and spoken numbers. Each slip becomes an edit.

Run the math on editing. At a $30 hourly rate, an editor costs $0.50 a minute. Every extra 10 minutes of cleanup per audio hour adds $5, which is more than the whole AWS Transcribe bill for that hour. A transcript that costs $0.18 and needs 20 minutes of fixing is more expensive than one that costs $0.36 and needs 5.

The only reliable test is your own audio. Pick three representative files, run each through three models, and count the edits you make. Do this before you decide anything based on a benchmark chart.

Meetings and Speaker Labels

Four colleagues around a conference table with a speakerphone

Meetings are the hardest everyday case because the transcript is useless without knowing who said what. Some models include speaker labels in the base price, such as OpenAI's diarization variant of its transcribe model, while others meter them as an add-on. When you compare a meeting-focused workload, add the diarization fee to every option on the shortlist, then compare totals. The cheapest base rate frequently stops being the cheapest total.

How to Use GPT-4o Mini Transcribe

Overhead view of a clean home studio desk with laptop, headphones, and microphone

PicassoIA hosts three speech-to-text models you can try straight from the browser: GPT-4o Mini Transcribe, GPT-4o Transcribe, and Gemini 3 Pro. That makes it a quick way to run the three-file test from the previous section before you open a single vendor account.

Step by Step Setup

  1. Open the GPT-4o Mini Transcribe page on PicassoIA.
  2. Upload a short audio file, ideally a 3 to 5 minute clip that includes your hardest material: names, numbers, and overlapping voices.
  3. If the form offers a language field, set it instead of leaving the model to guess.
  4. Add a short prompt with unusual names or jargon if a prompt field is available.
  5. Run the transcription and read the output against the audio.
  6. Repeat with GPT-4o Transcribe and Gemini 3 Pro on the same clip.
  7. Count the lines you would edit in each result. The lowest count wins your shortlist.

Parameter Tips That Save Money

  • Trim silence first. Billing follows duration, so cutting a 10-minute dead intro out of a recording saves 10 minutes of cost on every provider.
  • Set the language. Language detection can be an extra fee on some APIs and an error source on short clips.
  • Use a vocabulary prompt. A single line listing brand names and technical terms fixes more mistakes than moving to a pricier model.
  • Test on identical clips. Compare models on the same file, otherwise differences in audio quality get mistaken for differences in model quality.

Turn Transcripts Into Speech and Music

Transcription is usually the first step in a longer pipeline. Once you have clean text, you can repurpose it: narrate it in a new voice, translate and re-record it, or build a soundtrack around it. PicassoIA keeps those steps in one place.

Re-Voice Transcripts With Natural Speech

Voice actor recording in a padded booth with a guitar and synthesizer behind glass

Feed a corrected transcript to a text-to-speech model and you get a clean narration track without booking a studio. Speech 2.8 HD targets studio-quality voiceovers, ElevenLabs v3 focuses on expressive delivery, Gemini 3.1 Flash TTS offers 30 voices across more than 70 languages, and Realtime TTS 2 is built for low-latency use. This is the practical loop for podcasters who want a second-language edition of every episode.

Add a Music Bed to the Result

Narration sounds finished once it has a quiet music bed under it. Music 2.6 generates full songs from a text prompt, Lyria 3 Pro builds longer instrumental pieces, and Stable Audio 2.5 is a good fit for short loops and ambient beds. Describe the mood in one sentence, keep the track instrumental, and drop it 15 to 20 decibels under the voice.

Pick Your Cheapest Setup

Here is the shortcut version of everything above:

  • Backlog of old recordings, price is everything: Groq Whisper Large v3 Turbo at $0.04 per hour.
  • Small team, balanced quality and cost: GPT-4o Mini Transcribe at $0.18 per hour.
  • Accents, noise, and tough audio: GPT-4o Transcribe at $0.36 per hour, or AssemblyAI Universal-3.5 Pro at $0.21.
  • Meetings with speaker labels: price the diarization add-on first, then compare totals.
  • Locked into one cloud by policy: Google, Azure, or AWS, and ask for the discounted batch mode.

Whatever you choose, treat the first month as a test. Log the real invoice, divide it by the hours transcribed, and compare it with the table above. Prices move, and so do your needs. A vendor that wins today on cost can lose next quarter on accuracy or limits, so keep your integration thin enough to swap providers with a one-line change, and keep a 5-minute test clip on hand to re-run whenever a rate changes.

Ready to put it into practice? Open PicassoIA, upload one of your own recordings to GPT-4o Mini Transcribe, then try the same clip on the other transcription models. When the text looks right, narrate it with a text-to-speech voice, add a music bed, and see how far one recording can go. Browse the full catalog at picassoia.com/en/all-models and start experimenting today.

Share this article