Large Language ModelsTranscribe audioGenerate videos
YouTube MCP Server: Transcripts and Video Search in Claude, and Where It Breaks
A YouTube MCP server lets Claude fetch transcripts, search videos, and pull details on demand. See the three ways these servers work, the Claude Desktop and Claude Code setup, prompts that return quotable answers, and a fallback for videos without captions.
You paste a YouTube link into a chat, ask for the three claims that matter, and Claude answers from the actual transcript instead of guessing from the title. That is the promise of a YouTube MCP server: a small program that sits between Claude and YouTube and hands over transcripts, video details, and search results on request. No more scrubbing through a 90 minute talk to find the one slide about pricing. No more copying captions into a chat window and hoping the formatting survives.
It also has sharp edges. Captions get disabled. Search burns through API quota. Long transcripts eat context. This article shows how these servers work, how to connect one to Claude Desktop or Claude Code, which prompts return answers you can quote, and what to do when a video refuses to give up its text. The last section offers a fallback for videos with no captions at all: transcribe the audio yourself, then hand the text back to Claude.
What a YouTube MCP Server Does
MCP in Plain Terms
The Model Context Protocol, or MCP, is an open standard that Anthropic introduced in late 2024 so AI assistants can call outside tools in one consistent way. An MCP server publishes a list of tools, each with a name, a short description, and typed inputs. A client such as Claude Desktop or Claude Code reads that list. When your request needs one of those tools, Claude calls it and reads whatever comes back.
A YouTube MCP server usually exposes some mix of these tools, though the names change from project to project:
Get transcript: takes a video URL or ID and returns the caption text, often with timestamps.
Get video details: title, channel, description, duration, publish date, and view count.
Search videos: takes a query and returns matching videos.
List channel or playlist videos: useful for batch work on one creator's catalog.
Get comments: included by some servers, handy for checking audience reaction.
Most local servers talk to Claude over stdio. Claude launches the program on your own machine and exchanges messages with it through standard input and output. Nothing is hosted for you, and nothing runs unless the client starts it.
Transcripts vs Video Search
Two capabilities do most of the work, and they solve different problems.
Transcripts answer "what was said in this video." They are cheap, fast, and text only. Claude can quote a sentence, build a timestamped outline, or pull every number a speaker mentioned.
Video search answers "which videos are worth my time." It returns titles, channels, and dates, but it tells you nothing about what is inside until you fetch transcripts for the best candidates. The strongest workflow chains the two: search, shortlist, pull transcripts, compare.
💡 Tip: Ask Claude to show the search shortlist before it fetches a single transcript. You save quota and keep the context window clean.
Three Ways These Servers Work
Not every YouTube MCP server reaches YouTube the same way. The method decides what you install, what it costs, and where it fails.
Caption Scraper Servers
These fetch the caption track YouTube already shows under a video, either through a transcript library or by reading subtitles with yt-dlp. You need no API credential and no Google Cloud project, so setup takes a few minutes. The weakness is obvious: no captions, nothing to fetch. They also tend to get blocked more often when they run from a cloud server than from a home connection.
YouTube Data API Servers
These call the official YouTube Data API v3 and need a credential from the Google Cloud Console. Metadata and search are reliable, but quota is tight. Every project gets 10,000 units per day by default, a search request costs 100 units, and a plain video lookup costs 1. That works out to roughly 100 searches a day before the quota runs dry, and it resets at midnight Pacific Time. The official quota table lists the price of every method.
One more catch: downloading captions through the official API costs 200 units per call and requires edit permission on the video. The API route alone usually cannot fetch transcripts for other creators' videos, so many setups pair the Data API for search with a scraper for transcripts.
Download and Transcribe Servers
These download the audio with yt-dlp, then run speech recognition on it. They are slower and need compute, but they work on videos with no captions. This is the heavy option, and the right pick when accuracy on messy audio matters more than speed.
Approach
Setup effort
Needs API credential
Works without captions
Main limit
Caption scraper
Low
No
No
Missing captions, blocked requests
YouTube Data API
Medium
Yes
No, for other creators' videos
10,000 units per day
Download and transcribe
High
Depends on the server
Yes
Slow, needs compute
Connect It to Claude in Minutes
Before you install anything, pick a server whose source you can read. An MCP server runs code on your machine with your permissions, so check the repository, the maintainer, and the recent commit history first. You also need Node.js for servers launched with npx, or Python with uv for Python based servers.
Claude Desktop Config
Open Settings, then Developer, then Edit Config, or open the file directly. On macOS it lives at ~/Library/Application Support/Claude/claude_desktop_config.json. On Windows it lives at %APPDATA%\Claude\claude_desktop_config.json. Add an entry under mcpServers:
Replace your-youtube-mcp-package with the package name from the server's README, and add any environment variable the README asks for, such as the Data API credential. Then quit Claude Desktop fully and reopen it. The tools menu should now list the YouTube tools.
Claude Code One Liner
Claude Code registers a server with a single command:
claude mcp add youtube -- npx -y your-youtube-mcp-package
Add --scope user to make the server available in every project instead of only the current one. Run /mcp inside Claude Code to check that the connection is live.
💡 Windows tip: on native Windows, servers launched with npx usually need the cmd /c wrapper. Use claude mcp add youtube -- cmd /c npx -y your-youtube-mcp-package.
Test with a short video: "Fetch the transcript for this URL and give me the first five minutes as bullets with timestamps." Claude asks permission the first time it calls a tool. Approve it, then check that the timestamps match the video.
Prompts That Actually Work
Single Video Prompts
Specific requests beat vague ones. Ask for evidence, not impressions:
"Pull the transcript and write a 10 line outline with a timestamp on every line."
"List every price, number, and date the speaker mentions, with the exact sentence it came from."
"Quote each claim the speaker makes that I should verify elsewhere."
"Write five quiz questions from this lecture, with answers and timestamps."
💡 Tip: Asking for timestamps and exact quotes keeps Claude grounded in the text. If an answer has no quote behind it, treat it with suspicion.
Multi Video Prompts
Chaining search and transcripts is where the setup pays off. Try a prompt like this:
"Search YouTube for the five most relevant videos about [topic] and show me titles and channels first. After I pick three, pull their transcripts and compare what each one says about [question] in a table."
Watch the size of what you pull. Spoken English runs near 150 words a minute, and a word costs roughly 1.3 tokens, so these are rough estimates:
Video length
Approx. words
Approx. tokens
10 minutes
1,500
2,000
1 hour
9,000
12,000
3 hours
27,000
36,000
Pulling five one hour talks into a single chat means around 60,000 tokens of raw text before Claude writes a word. Large context windows handle that, but it burns usage fast. Summarize each video in its own request, then ask Claude to merge the summaries.
Real Workflows Worth Copying
Researchers and students. Feed a lecture playlist through the server one video at a time. Ask for a glossary and a list of open questions per lecture, then merge them into a single study sheet. Because every answer carries a timestamp, you can jump back to the exact minute when something looks off.
Creators and podcasters. Pull transcripts from your own back catalog and ask which episodes mention a topic, which hooks repeat, and which ideas deserve a follow up. Long episodes turn into blog drafts, newsletter sections, and short clip scripts without anyone retyping a word.
Product and marketing teams. Watch competitor webinars and product demos without blocking anyone's afternoon. Pull the transcript, ask Claude to list new features, pricing changes, and positioning claims, and drop the result into a shared document. Support teams can turn tutorial videos into help center drafts the same way.
💡 Respect the source: Transcripts belong to the creators. Use them for research, summarize and credit, and do not republish them word for word.
Where It Breaks and What to Do
Missing Captions and Blocked Requests
Most failures trace back to a short list of causes:
Symptom
Likely cause
Fix
"No transcript available"
Captions disabled or not generated yet
Transcribe the audio, see the tutorial below
Works at home, fails on a server
The host IP range is blocked
Run the server locally
Transcript in the wrong language
Server picked the first caption track
Name the language code in your prompt, such as en
Garbled names, no punctuation
Auto-generated captions
Ask Claude to clean the text and flag names it is unsure about
Search stops mid-day
Daily quota spent
Search less, shorten result lists, wait for the reset
Tools missing in Claude
Broken JSON or no restart
Validate the config file, fully restart the client, check the logs
Long Videos and Untrusted Text
Long videos stress the context window. Ask for one summary per 10 minute segment, then a merged outline at the end. If the server returns timestamps in seconds, tell Claude to convert them to minutes and seconds so you can click straight to the spot.
Also treat transcripts as untrusted input. A video can contain spoken or written instructions aimed at an AI model, and a transcript delivers them straight into the chat. Keep tool approvals switched on, and do not run a YouTube session next to tools that can write files, send email, or touch production systems.
When captions are missing or badly garbled, transcribe the audio yourself. Download only audio you have the right to use: your own videos, licensed material, or content whose creator allows it. GPT 4o Transcribe on Picasso IA then turns the recording into text. It accepts mp3, mp4, mpeg, mpga, m4a, ogg, wav, and webm files, so you rarely need to convert anything first.
Fill the language field with an ISO 639 1 code such as en, es, or fr. Naming the language improves both accuracy and speed.
Optionally add a prompt with the spellings you need, for example product names and jargon. It should match the language of the audio.
Leave temperature at its default of 0 for the most precise output.
Run it, then copy the transcript.
💡 Tip: Test a 30 second clip first. If names come out wrong, add them to the prompt field and run the clip again before you commit to the full recording.
Open Claude Sonnet 5 on Picasso IA and paste the transcript into the prompt field. Three settings matter:
effort:low is the default and the fastest, with thinking switched off. Move to medium or high when the transcript is long and the question is subtle.
max_tokens: defaults to 8192, enough for a detailed outline.
system_prompt: set a role once, such as "You are a research assistant who always quotes the speaker and adds a timestamp."
A prompt that works well: "Summarize this transcript in five bullets, list every number mentioned, and quote the two most important sentences." For hour long recordings, split the text into 10 minute chunks and merge the results, the same trick as with the YouTube server.
Turn Notes Into a Video Clip
Once you have a summary, you can turn it into something people actually watch. Ask Claude Sonnet 5 to condense the three strongest points into a short scene description with a clear subject, a camera move, and a lighting note. Then paste that description into a text to video model such as Seedance 2.0 or Veo 3.1 Fast. A short clip makes a strong header animation for a post or a newsletter, and the whole loop from audio to video takes minutes.
Try It on Picasso IA Today
You now have the full loop. A YouTube MCP server gives Claude the transcript and the search results. When captions fail, speech to text fills the gap. Then a language model turns raw text into notes, and a video model turns those notes into a clip.
Start small. Pick one five minute video, transcribe it on Picasso IA, summarize it with Claude Sonnet 5, and animate the best idea with a text to video model. Change one setting at a time, compare the results, and keep the prompts that work. Open Picasso IA, run your first transcription, and see how much of your next research session you can hand off.