Large Language ModelsTranscribe audioGenerate images
Audio Transcription MCP: Transcribe Files in Claude and Cursor
An audio transcription MCP server lets Claude and Cursor read your MP3, WAV and M4A files directly. This article shows the config for Claude Desktop, Claude Code and Cursor, the prompts that turn transcripts into notes, and fixes for the errors that block most first setups.
Hand a model a 90 minute MP3 and it has nothing to read. An audio transcription MCP server fixes that. It sits between your chat window or editor and a speech-to-text engine, so you can type "transcribe interview-04.mp3 and pull the three best quotes" and get text back inside Claude or Cursor. No tab switching, no upload forms, no copy and paste between five apps.
This article walks through the exact config for Claude Desktop, Claude Code and Cursor, the prompts that work on messy transcripts, and the fixes for the errors that stop most first attempts. It also shows a no-install route for one-off recordings, using the speech-to-text models on PicassoIA, and a way to turn what people said in a recording into images.
What an Audio Transcription MCP Does
The Model Context Protocol (MCP) is an open standard that lets an AI app call outside tools through a small server process. A transcription server exposes a few tools, usually along the lines of "transcribe this file path" or "split this long recording", and the client decides when to call them. The server does the heavy lifting with a speech-to-text engine and hands plain text back to the conversation.
The Request Flow
Here is what happens when you ask for a transcript:
You type a request that names a local file path.
The client sees the transcription tool in its tool list and asks permission to run it.
The server reads the file and sends it to a cloud API or a local Whisper model.
Text returns as the tool result, and the model carries on: summarizing, extracting quotes, rewriting.
💡 Worth remembering: the model never hears the audio in this setup. It reads the transcript, so any error in a summary traces back to what the engine heard first. Check names, numbers and dates against the source before you publish anything.
Why Chat Alone Falls Short
Pasting a transcript by hand works for one file. It breaks at ten. With a server in place you can point at a whole folder, keep timestamps, and chain steps in a single prompt: transcribe, draft show notes, save to notes.md. Inside Cursor the same tool call can write the result straight into your repository, next to the code or docs it belongs with. Inside Claude Code it can feed a script that renames files by speaker or topic.
Three situations where the setup pays off fast:
Weekly meetings: record, transcribe, and get action items without anyone typing minutes.
Interviews and research calls: pull verbatim quotes with timestamps for a report.
Podcasts and videos: produce show notes, chapter lists and caption drafts from the audio track.
Pick a Transcription Engine
Before you touch a config file, decide where the audio gets processed. That one choice drives cost, privacy and speed more than any setting you will change later.
Cloud API or Local Whisper
Factor
Cloud API
Local Whisper
Setup
Paste a credential into the config
Install a runtime, download a model
Privacy
Audio leaves your machine
Audio stays on disk
Speed
Seconds for most files
Depends on your CPU or GPU
File limits
Upload caps apply (OpenAI's transcription endpoint has long capped uploads at 25 MB)
Limited by disk and memory
Cost
Billed by usage
Free after setup
Best for
Fast turnaround on clean audio
Sensitive recordings and offline work
If the recording is a client call, a medical note or anything under an NDA, pick local. For podcasts and public webinars, a cloud API is usually faster to wire up.
Three Servers Worth Checking
Community projects already exist, and their names say what they do:
mcp-server-whisper: a server built around OpenAI's speech models.
whisper-mcp: local transcription with Whisper models through whisper.cpp, so no audio is uploaded.
fast-whisper MCP servers: wrappers around Faster Whisper for quick local recognition on a GPU or a decent CPU.
Names change and projects go stale. Open the repository, check the date of the last commit, read which tools the server exposes, and copy the install command from its README instead of from this page. The configs below use placeholders for that reason.
💡 Safety check: an MCP server runs as a process with your user permissions. Install only code you have read or trust, and point it at a recordings folder rather than your whole home directory.
Set It Up in Claude
Claude Desktop reads servers from a JSON file. Claude Code adds them with a command. Both start the server as a local process on your machine.
TRANSCRIBE_TOKEN is a stand-in. Use the variable name your server's README asks for. Python servers usually launch with uvx instead of npx, so swap the command and args to match. Then quit Claude Desktop fully, not only the window, and open it again. The tool list shows the new server once the process starts.
Two rules trip people up: every option goes before the server name, and the double dash separates the name from the command that starts the server. Check the result with claude mcp list, inspect it with claude mcp get transcribe, and drop it with claude mcp remove transcribe.
💡 Use --scope project to save the entry in a .mcp.json file your team shares. Never commit a raw credential there. Reference an environment variable such as ${TRANSCRIBE_TOKEN} instead and let each person set their own.
Run a quick test before trusting it with a real file: record ten seconds of yourself on a phone, drop the file in your recordings folder, and ask Claude to transcribe it. A correct transcript confirms the server, the credential and the file path in one go.
Set It Up in Cursor
Cursor reads the same mcpServers shape from a file. Put .cursor/mcp.json inside a project to scope the server to that repo, or ~/.cursor/mcp.json in your home folder to make it available everywhere.
The ${env:NAME} syntax pulls the value from your shell environment, so the file stays safe to commit. Cursor also resolves ${workspaceFolder} inside command, args and env values, which is handy when the server needs a path to a folder of recordings inside the repo.
Call It From Chat
Open the agent chat and check the MCP settings page: the server should show as enabled with its tool names listed. Then ask in plain language:
Transcribe recordings/call-0612.m4a with the transcribe tool and save the text to docs/call-0612.md. Add a five-line summary at the top.
By default Cursor asks you to approve each tool call before it runs. Approve the first one, check the output, and only then consider turning on auto-run for this server.
Prompts That Work on Transcripts
Raw transcripts are long and untidy, so the prompt decides whether you get a wall of text or something you can use. The best prompts name the file, say what to extract, and tell the model what to do when the audio is unclear.
Meeting Notes That Assign Owners
Transcribe standup-0612.m4a. Then list decisions, open questions and action items. For every action item, name the owner and the date they gave. If nobody gave a date, write "no date" instead of guessing.
The last sentence matters. Without it, models tend to fill gaps with a plausible Friday.
Interview Quotes With Timestamps
Transcribe interview-04.mp3 with timestamps. Pick the five strongest quotes, keep each one word for word, and give the timestamp next to it. Mark any quote where the audio sounded unclear.
Verbatim text plus timestamps gives you a way to check each quote against the recording in seconds.
A few habits that save time on every prompt:
Name the path exactly, including the extension.
Ask for verbatim quotes and forbid paraphrase when the wording matters.
Ask for uncertainty flags so unclear passages are marked rather than smoothed over.
Add a glossary of product names, people and jargon, since engines guess spellings.
Work in chunks for recordings longer than an hour: transcribe, summarize each part, then merge.
Fix the Usual Errors
Most failures come down to five causes. This table gets you to the right one fast.
Symptom
Likely cause
Fix
Tool never appears
JSON syntax error or no full restart
Validate the JSON, quit the app fully, reopen
"command not found"
npx or uvx missing from the app's PATH
Use the absolute path to the executable
Authentication error
Wrong or missing credential variable
Match the exact variable name in the README
Timeout on a long file
Upload cap or slow engine
Split the file into 15 minute parts
Misspelled names
The engine guessed the spelling
Add a glossary or a style prompt
Server Missing From the List
Start with the boring checks. A trailing comma breaks the whole file, and a desktop app often does not inherit the PATH you see in your terminal, so npx works in a shell but fails in the app. Paste the absolute path to the executable into command and try again.
Then run the server command by hand in a terminal. If it fails there, it fails in the client too, and the terminal prints the real error. For Claude Desktop, the MCP log files sit in a logs folder: %APPDATA%\Claude\logs on Windows and ~/Library/Logs/Claude on macOS.
Big Files and Timeouts
Both Claude and Cursor can give up on a tool call that runs too long, and cloud engines have upload caps. Shrinking the file solves both. Speech does not need studio quality, so mono at a low sample rate works:
Then ask the model to transcribe part_000.mp3, part_001.mp3 and so on in order, and join the results in one prompt at the end.
Transcribe in the Browser on PicassoIA
For a one-off recording, a server is more setup than the job deserves. PicassoIA runs speech-to-text models in the browser with nothing to install: GPT 4o Transcribe, GPT 4o Mini Transcribe and Gemini 3 Pro. The PicassoIA connector for Claude is built for image and video generation, so transcription happens on the model pages, and you paste the finished text into Claude or Cursor afterward.
Upload your audio file. The model accepts MP3, MP4, MPEG, MPGA, M4A, OGG, WAV and WebM.
Fill in Language with an ISO-639-1 code such as en, es or fr. The model page says supplying it improves both accuracy and latency.
Add an optional Prompt: a short text that sets the style or continues a previous segment. It should match the language of the audio. A line listing speaker names and product terms works well.
Leave Temperature at its default of 0 for the most predictable output.
Run it, then copy the transcript.
💡 Test with a 30 second clip first. If names come out wrong, fix the prompt line before you run the full hour.
Very long audio, transcript plus summary in one request
One audio file up to 8.4 hours, thinking level, system instruction
Gemini 3 Pro takes a text prompt alongside the audio, so you can ask for a transcript, a list of decisions and a three-line summary in one pass. Once you have text, a language model on PicassoIA such as Claude Sonnet 5 can rewrite it into show notes, an email or a changelog.
From Transcript to Image on PicassoIA
A transcript is also a source for visuals. A brainstorm call, a customer interview or a podcast episode is full of concrete scenes that people described out loud, and those scenes make better image prompts than anything you would invent at a blank page.
Write the Image Prompt From Quotes
Ask Claude or Cursor to turn the transcript into five scene descriptions, each with a subject, a setting, the light and a lens. For example:
A small bakery counter at opening time, a baker sliding a tray of loaves toward the camera, warm window light from the left, 50mm lens, shallow depth of field, flour dust in the air.
Paste each description into an image model on PicassoIA: Flux 2 Pro, P-Image or Nano Banana Pro. Run the same prompt through two of them and compare, since each model reads light and texture a little differently. Keep one scene per prompt, name the light direction and the lens, and describe surfaces like wool, wood grain or steam. Those details separate a photograph-like result from a generic one.
Your next recording already contains your next five images. Pick one short clip, transcribe it with the setup above or on PicassoIA, pull out the most vivid scene, and create your own images on Picasso IA today. Change the lens, move the light, and see how far a single spoken sentence can go.