Large Language ModelsGenerate videosGenerate images
Veo 3.1 MCP Server: Generate Veo Videos in Claude and Gemini
A Veo 3.1 MCP server lets Claude and Gemini start Google video renders, poll the job, and save the MP4 for you. See the real tool names, config blocks for Claude Desktop, Claude Code and Gemini CLI, per second prices, and prompt patterns that give better clips.
Typing one sentence into Claude and getting a finished video with sound back is no longer a lab trick. With a Veo 3.1 MCP server, it is a normal tool call: your client sends the prompt to Google's video model, waits while the clip renders, and saves the MP4 to a folder you chose. The catch is that nothing works out of the box. You pick a server, wire it into Claude or Gemini, and check the billing before the first render burns through your budget.
This article walks through that setup with real tool names, real config blocks and real per second prices, then shows the no-config alternative on PicassoIA. Everything about third-party servers comes from their public READMEs and listings, so treat it as a reading list, not a test report.
What a Veo MCP Server Does
MCP, the Model Context Protocol, is the open standard that lets an AI client call outside tools. A Veo MCP server is a small program that exposes video generation as a handful of tools. Your client launches it as a local process, reads the tool list, and calls those tools whenever you ask for a clip.
Why not call the Gemini API directly? You can, and a short Python script does the job. An MCP server earns its place when you want the conversation to drive the work: Claude drafts the shot list, picks the settings, starts the render, checks on it, and tells you where the file landed, all in one thread.
The Three Moving Parts
Every setup has the same three layers, whichever server you choose.
The server holds your Google credential, so it is the part that deserves the most scrutiny.
Why Video Needs Async Tools
A text reply arrives in seconds. A video does not. One server's README quotes render times from 11 seconds to 6 minutes depending on the prompt, and another starts the job in the background, hands back a job ID, and tells the client to poll every 15 to 20 seconds. That pattern matters because chat clients give up on tool calls that hang too long.
💡 If a server blocks until the video is finished, expect timeouts on long renders. Prefer servers with one tool that starts a job and a separate tool that checks its status.
Community Servers and Their Tools
Google ships Veo through the Gemini API, and the community wrapped it in several MCP servers. None of them is made by Google or by PicassoIA. The details below come from the alohc server on GitHub, the Gemini Media MCP listing, visualgen mcp and mcp-veo from AceDataCloud.
veo_generate_video, veo_image_to_video, veo_interpolate_video, veo_extend_video, veo_check_job, veo_list_jobs and three utility tools
720p, 1080p or 4K, batches of 1 to 4 videos
visualgen mcp
An image generator plus Veo 3.1 in lite, fast and standard tiers
720p up to 4K, optional image to video
mcp-veo (AceDataCloud)
Text to video, image to video, 1080p upscaling, task tracking
Runs through the AceDataCloud API, not Google directly
Three tool types repeat everywhere: text to video, image to video, and extension. Extension adds roughly 7 seconds to an existing clip, but one README limits it to 720p and 148 seconds in total, and Veo 3.1 Lite does not support extension at all.
The default model in the alohc server is veo-3.1-generate-preview, and the Gemini Media MCP server offers a standard and a fast variant. When a server lets you pick the model by name, write that choice into your prompt so the client does not guess.
Pick One You Can Audit
These programs run on your machine with your Google credential in their environment. Before installing one:
Read the source, since each one is short enough to skim.
Pin a version instead of tracking the latest release.
Create a dedicated Google credential for it, so you can revoke it without breaking anything else.
Set a budget alert on the billing account, because a runaway loop is billed per second of video.
Connect Veo to Claude
Claude reaches MCP servers two ways: Claude Desktop through a JSON config file, and Claude Code through one command. Both launch the same server package.
Claude Desktop Setup
Install uv, the Python runner that provides uvx.
In Claude Desktop, open Settings, then Developer, then Edit Config.
Add your Gemini credential to that same env block, using the variable name your chosen README lists.
Quit Claude Desktop fully and reopen it. The Veo tools should appear in the tools menu.
💡 Output folder variables differ by server. One README uses VIDEO_OUTPUT_DIR, another uses VEO_OUTPUT_DIR. Copy the name from the project you installed.
If the tools do not appear, validate the JSON first, because a single trailing comma breaks the whole file. Then confirm uvx is on your PATH and that you quit the app instead of just closing the window.
Claude Code Setup
One command registers the server for every project you open:
claude mcp add veo -s user -- uvx veo-mcp-server
Pass your credential with the -e flag, then run /mcp to confirm the server shows as connected. After that, a prompt such as "Generate a 6 second vertical clip of rain on a tram window, no music, then tell me where the file is saved" produces a start call, several status checks, and a file path at the end. For your first run, ask for a 4 second clip at 720p, so a wrong config or a vague prompt costs you cents instead of dollars.
Connect Veo to Gemini
You have two routes. Gemini CLI can load the same MCP servers, and the Gemini app can generate Veo clips on its own.
Gemini CLI Config
Gemini CLI reads ~/.gemini/settings.json and expects servers under an mcpServers object, the same shape Claude uses:
Add the credential variable here too, restart the CLI, and run /mcp to list every tool it loaded. Gemini CLI also accepts SSE and HTTP streaming servers, so a hosted server works too, as long as it is one you trust.
If you want a chat model to draft shot lists before you spend anything on rendering, Gemini 3.5 Flash is quick for rough drafts and Gemini 3.1 Pro suits longer scripts.
The Gemini App Without MCP
The Gemini app can also generate Veo videos directly on supported plans. That route suits one-off clips. It gives you no scripting, no batch runs and no file path you control. The moment you need repeatable output, MCP wins.
Prompts That Give Better Clips
Veo 3.1 renders 4, 6 or 8 second clips in 16:9 or 9:16, with audio generated alongside the picture. It accepts a negative prompt, a seed, up to 3 reference images, and a first and last frame. The model's own examples on PicassoIA write dialogue and room sound straight into the prompt, so treat the prompt as a script, not a tag list.
Shot, Subject, Motion, Sound
Build every prompt in four beats:
Shot: lens, angle and distance, such as "low angle, 35mm, slow push in".
Subject: who or what is on screen, with two or three concrete details.
Motion: what changes over the clip, in order.
Sound: ambient noise, music, or quoted dialogue.
Sound is the beat most people skip. The model generates audio by default, so describe the room tone and any spoken line, and write "no music" when you want only natural sound. Where your tool exposes an audio switch, turning it off gives you silent footage to score yourself later.
A working example: Low angle, 35mm, slow push in on a ceramic mug of tea on a rainy windowsill. Steam curls upward, a raindrop slides down the glass behind it. Soft rain, a distant tram bell, no music.
Weak prompt
Strong prompt
A cool city video
Aerial pull back over a rain-slick street market at dusk, vendors lowering their awnings, warm string bulbs reflecting on wet asphalt, distant thunder and murmuring voices
A person talking
Medium close-up of a woman in a wool coat on a train platform, she says "Last one home turns off the lights", station announcement echoing behind her
💡 Put anything you do not want, such as text overlays or handheld shake, in the negative prompt instead of the main prompt.
Start Frame and End Frame
Image to video gives you the most control. Set image as the first frame and last_frame as the last, and the model builds the transition between them. Ideal inputs match the output ratio, about 1280x720 for 16:9 or 720x1280 for 9:16. You can generate those stills with PicassoIA Image, which offers both ratios.
Two limits to remember: reference images only work at 16:9 and 8 seconds, and the last frame is ignored whenever reference images are present.
The Real Cost Per Clip
The MCP server itself costs nothing. Google bills the Gemini API per second of video, audio included, and the same clip can cost 8 times more or less depending on the tier you pick.
Per Second Rates
Published Gemini API rates at the time of writing:
Check Google's pricing page before you budget, since rates change. At 1080p and 4K, clips run 8 seconds.
Pick the Right Tier
Draft on Lite or Fast at 720p. An 8 second Fast draft costs $0.80, and a Lite draft costs $0.40.
Render the winner on the standard Veo 3.1 tier when quality matters most. Lite has no 4K and no extension.
Batch with a ceiling. Twenty 8 second social clips add up to 160 seconds of video: $8.00 on Lite at 720p, $19.20 on Fast at 1080p, $64.00 on standard at 1080p.
💡 Ten 8 second drafts at 720p cost $8 on Fast and $32 on standard. Iterate cheap, then pay once for the final.
One more cost that does not show on the invoice: the alohc README warns that generated videos stay on Google's servers for about 2 days. Make sure your server downloads each file to disk the same day.
Skip the credentials, the config files and the billing account. The model page on PicassoIA exposes the same controls in a browser, and the clip lands in your gallery ready to download.
Optionally add a first frame, a last frame, up to 3 reference images, a negative prompt or a seed.
Generate, preview, and download the MP4.
Setting
Options
Default
Duration
4, 6 or 8 seconds
8
Resolution
720p or 1080p
1080p
Aspect ratio
16:9 or 9:16
16:9
Audio
On or off
On
Veo 3.1 Fast uses the same settings except for reference images, which it does not offer. Choose it for quick drafts: the Fast examples on its page rendered in 45 to 59 seconds, while the standard model's examples took roughly 73 to 114 seconds.
PicassoIA Through MCP
PicassoIA also has a developer API at https://api.picassoia.com/v1 with Replicate-style prediction endpoints and Bearer authentication, plus an MCP connector. Both expose four models: an image model, an image editor, a video model and Seedance 2.5 Lite with audio. Veo 3.1 is not on that list, so for Veo you use the model page, and for scripted video inside Claude you can call PicassoIA Video or Seedance 2.5 Lite. Each account runs up to 5 predictions at once, shared across connections.
Make Your First Clip
Pick one sentence you could film in a single shot. A hand pouring tea, a tram crossing a bridge, a dog waiting at a door. Short, concrete scenes give Veo the clearest target, and they cost the least to repeat.
Three Mistakes to Skip
Forgetting the expiry. Videos left on Google's side disappear after about 2 days, so download them.
Drafting at 4K. Draft cheap, finalize once.
Skipping the status tool. Clients that wait on one long call time out. Use job IDs and poll.
Ready to try it yourself? Open PicassoIA, generate a first frame with PicassoIA Image, then animate it with Veo 3.1. Change one detail per run, compare the results against Seedance 2.5 Lite, and keep the versions you like. Your own images and videos are one prompt away on Picasso IA, and the full model list is waiting at picassoia.com/en/all-models.