Generate videosLarge Language ModelsVisual Effects

Veo 3 API in n8n: Faceless Video Workflow That Runs on Autopilot

Build a faceless video workflow where n8n reads a topic sheet, asks an LLM for scenes, sends each prompt to the Veo 3 API, polls the long-running render, adds narration, and uploads the finished short. Includes node settings, retry rules, and cost tracking.

Veo 3 API in n8n: Faceless Video Workflow That Runs on Autopilot
Cristian Da Conceicao
Founder of Picasso IA

Most faceless channels stall at the same place: the footage. Scripts are quick to write and a synthetic voice costs next to nothing, but sourcing forty believable seconds of video for every upload eats the whole week. Plugging the Veo 3 API into n8n removes that bottleneck. A row in a spreadsheet becomes a script, the script becomes eight-second scenes, the scenes become clips with sound already attached, and a finished vertical video lands in your upload queue while you sleep.

This walkthrough builds that faceless video workflow node by node. You will see which n8n nodes carry the load, how to handle Veo's slow asynchronous responses, where the money leaks, and how to test prompts on Veo 3 at Picasso IA before you hard-wire them into an automation.

Why Veo 3 Fits Faceless Channels

A faceless channel lives on volume and sameness. Viewers return for a format, not a presenter, so the production line matters more than any single clip. That is exactly the kind of problem workflow automation handles well: the same steps, in the same order, with a different topic every run.

Veo 3 suits that line for one practical reason. It is a text-to-video model that renders picture and sound together from a single prompt, and it also accepts a still image as a starting frame, so you can hold one look across every scene of an episode.

Good niches for this format include travel and nature explainers, food and recipe teasers, history stories told with atmospheric footage, product showcases, and ambient loops for sleep or focus. All of them rely on objects, places, and hands rather than a talking head, which is what keeps the format faceless in the first place.

Native audio removes a branch

Older pipelines ran three branches: video, music, and sound effects. Veo 3 generates ambient sound, effects, and spoken lines inside the clip itself. For a nature explainer, a street food montage, or a product teaser, that can be the finished soundtrack. You only add a separate voiceover when you need exact wording or one narrator across a whole series.

What a faceless video really needs

Before you open n8n, write down the output you want. A typical short-form faceless video needs:

  • A hook in the first two seconds, usually a striking visual rather than a sentence
  • Four to six scenes, each short enough to fit one Veo clip
  • Vertical framing at 9:16 for Shorts, Reels, and TikTok
  • One narrator voice or one music bed, never both fighting for space
  • A title, description, and tags ready for the upload step

💡 Tip: Decide the scene count first. Every extra scene adds one more API call, one more wait, and one more chance for something to fail.

Hand-drawn flowchart on grid paper planning an automated faceless video workflow

The Workflow at a Glance

The whole pipeline is ten nodes long. None of them are exotic, and it runs on n8n Cloud or a self-hosted install, with one exception explained in the assembly step.

Stepn8n nodeJob
1Schedule TriggerFires once a day at a fixed hour
2Google SheetsReads the next unused topic row
3HTTP Request or AI AgentAsks an LLM for a scene list in JSON
4CodeTurns each scene into a Veo prompt
5HTTP RequestSubmits the prompt to the Veo 3 endpoint
6Wait + IFPolls until the render is done
7HTTP RequestDownloads the finished MP4
8Execute CommandJoins the clips and adds narration
9YouTubeUploads the final video
10Google SheetsLogs status, link, and cost

The order matters more than the node choice. Build from the top down, run each node once with pinned test data, and only then connect the next one. A broken scene prompt is far cheaper to catch at step four than after three paid clips.

Build the Trigger and Script Nodes

Schedule Trigger and topic sheet

Start with a Schedule Trigger set to one run per day. Next, add a Google Sheets node that reads the first row where the status column is empty. Keep the columns plain: topic, tone, status, video URL, cost. That sheet becomes your content calendar, your queue, and your audit trail in one place.

Close-up of hands typing at a desk while building the first nodes of an automation

Ask an LLM for scenes

Add the writing step. The simplest route is an HTTP Request node, or n8n's AI Agent node, pointed at a chat model such as Claude Sonnet 5, Gemini 3.5 Flash, or GPT 5.4. Ask for strict JSON: an array of scenes, each with a visual description, a camera move, and a one-line sound note.

Structured output matters here. If the model returns loose prose, your Code node will break on the third run, usually at two in the morning.

Writer at a standing desk in a bright loft drafting a video script on a laptop

Shape each scene into a prompt

A Code node loops over the scenes and builds the final prompt for each one. A strong Veo prompt follows the same order every time:

  1. Subject and what it is doing
  2. Setting and time of day
  3. Camera movement and lens feel
  4. Light and colour
  5. Sound you want in the clip

Add one fixed style line to every scene, such as "natural light, handheld, 35mm film look", so the clips match when they are joined. Append a negative prompt that keeps out text overlays, watermarks, and logos.

Here is a finished example for a coffee farming explainer: A weathered farmer's hands pick red cherries from a steep hillside plantation at sunrise, slow dolly-in at waist height, mist rising between the rows, warm backlight, soft rustling leaves and distant birdsong. Natural light, handheld, 35mm film look. Notice that no face is requested. Hands, landscapes, and objects keep the format faceless, and they also avoid the common problem of a person whose face changes from one scene to the next.

Call the Veo 3 API From n8n

This is the heart of the workflow, and it behaves differently from a normal API call.

Configure the HTTP Request node

Veo is exposed through Google's Gemini API as a long-running operation. Your first request does not return a video. It returns an operation name that you check later. Set the node up like this:

  • Method: POST
  • URL: https://generativelanguage.googleapis.com/v1beta/models/veo-3.0-generate-001:predictLongRunning
  • Authentication: a Header Auth credential stored in n8n's credential vault, never pasted into the node itself
  • Body type: JSON

💡 Tip: Model IDs change between preview and stable releases. Confirm the exact ID in Google's current Veo documentation before you publish the workflow.

Send a body shaped like this:

{
  "instances": [
    { "prompt": {{ JSON.stringify($json.veoPrompt) }} }
  ],
  "parameters": {
    "aspectRatio": "9:16",
    "resolution": "720p",
    "negativePrompt": "text overlay, watermark, logo, subtitles"
  }
}

Wrapping the prompt in JSON.stringify stops a stray quotation mark inside a scene from breaking the request. Clips come back at eight seconds on Google's endpoint, and 1080p is tied to widescreen output, so check which resolution your chosen aspect ratio supports before you commit to a vertical format.

Poll with Wait and IF

The response contains a name field. Store it, then build a small loop:

  1. A Wait node pauses for 60 seconds.
  2. An HTTP Request node sends a GET to https://generativelanguage.googleapis.com/v1beta/ followed by the stored name.
  3. An IF node checks whether done is true.
  4. If it is false, the flow returns to a shorter Wait of 20 seconds. If it is true, the flow moves on.

On Veo 3's PicassoIA page, the sample renders finished in roughly 68 to 145 seconds. Treat that as a starting estimate for your own delays. With a 60-second opening wait and 20-second polls, most renders resolve within four or five checks.

Developer's monitor showing a JSON response from a video generation request

Add a counter inside the loop. A Code node increments an attempt number, and a second IF node stops the run after fifteen checks. Without that cap, one stuck render can loop forever and fill your execution history.

Download before the file disappears

When done turns true, the video address sits inside the response under generateVideoResponse.generatedSamples. Add one more HTTP Request node, set the response format to File, send the same credential, and capture the binary data. Google holds generated files only for a short window, so push the MP4 to your own storage (S3, R2, or Drive) in the very next node. Never store the Google link in your sheet and assume it will still work next week.

Red sand timer on a pale wood desk, a visual for waiting while a render finishes

Add Voice and Assemble the Video

When native audio is enough

Test one scene with the built-in track before you add anything. If the ambient sound fits and no spoken line is needed, skip the voice branch entirely. Every branch you cut is one fewer thing that can fail overnight.

Add a narrator with text to speech

For narrated formats, send the script to a speech model. ElevenLabs v3 handles expressive reads, Speech 2.8 HD from MiniMax gives clean studio narration, and Gemini 3.1 Flash TTS is the pick when you need many languages. Each returns an audio file you can save beside the clips.

💡 Tip: Add "no dialogue" to the Veo prompt when you plan to lay a narrator on top, so the clip's own voices never compete with yours.

Condenser microphone with a pop filter inside a small home voiceover booth

Stitch the scenes together

Joining clips is a job for FFmpeg inside an Execute Command node. A basic command reads a list of scenes, lays the narration over them, and writes one file:

ffmpeg -f concat -safe 0 -i scenes.txt -i narration.mp3 \
  -map 0:v -map 1:a -c:v libx264 -c:a aac -shortest episode.mp4

Here is the catch. The Execute Command node works on self-hosted n8n only. On n8n Cloud, send the clip URLs to an external rendering service through an HTTP Request node instead. Either way, the output is one MP4 per episode, ready for the upload step.

Video editor at a two-screen desk working on a multi-track timeline of short clips

Publish and Track Cost

Upload to Shorts, Reels, and TikTok

n8n ships a YouTube node that uploads a binary file with a title, description, and privacy status. For other platforms, call each official posting API through the HTTP Request node, or use a scheduling service that offers a webhook. Set privacy to private for the first few runs and review each video by hand until you trust the output.

Hands holding a smartphone playing a vertical short video at a sunlit cafe table

Log every run in the sheet

Write back the status, final video URL, render attempts, scene count, and estimated cost. Google bills Veo by the second of generated video, and the rates have changed more than once since launch, so store the price per second in one n8n variable or one sheet cell rather than hard-coding it in six nodes. At eight seconds per scene and five scenes, a single episode means 40 billed seconds before you count a single retry.

Retry rules that protect your budget

Failures are normal with generative video, so decide your policy up front:

  • Retry a failed render once with the same prompt
  • On a second failure, ask the LLM to rewrite the prompt, then try one last time
  • After that, mark the row "needs review" and stop
  • Add an Error Trigger workflow that pings you by email or chat
  • Cap daily runs so a loop bug cannot drain your balance

Printed run log with failed lines circled in red pen next to a calculator and invoices

Test Prompts With Veo 3 on PicassoIA

Before you spend API money, prototype each scene prompt by hand. Veo 3 on PicassoIA accepts a prompt, a resolution, an aspect ratio, an optional start image, a negative prompt, and a seed. Here is the routine:

  1. Open the Veo 3 page and sign in.
  2. Paste a scene prompt written in the same order your Code node uses.
  3. Set resolution to 720p for tests and switch to 1080p only for a final render.
  4. Pick 9:16 for vertical shorts or 16:9 for widescreen.
  5. Optionally upload a start image. Ideal frames are 1280x720.
  6. Add a negative prompt such as "text, watermark, logo".
  7. Once a look works, fix the seed so reruns reproduce it.
  8. Generate and wait. Sample renders on the page took about one to two and a half minutes.

Other video models are worth testing against the same prompt:

ModelGood for
Veo 3 FastQuicker drafts of the same look
Veo 3.1The newer generation at up to 1080p
Veo 3.1 FastSpeed with the newer model
Veo 3.1 LiteCheaper test runs
Seedance 2.5 LiteVideo with audio, up to 10 seconds
PicassoIA VideoText or image to video in one place

💡 Tip: PicassoIA also runs a developer API at https://api.picassoia.com/v1 with Replicate-style endpoints: a POST to /v1/models/{owner}/{name}/predictions, then a GET on /v1/predictions/{id} until the job ends. At the time of writing it exposes PicassoIA Video and Seedance 2.5 Lite for video and allows five concurrent predictions per account. Veo 3 is not on that list, which is why the workflow above targets Google's endpoint. The create-then-poll pattern is identical, so the same Wait and IF loop works for both.

Make Your First Clip Today

You do not need the full ten-node pipeline to see whether this idea works for your niche. Pick one topic, write five scene prompts in the order from the routine above, and run them on Picasso IA. Compare a 720p draft with a 1080p render, try a vertical and a widescreen version, and note which prompts hold up across scenes.

When you have three prompts you trust, copy them into your Code node and let n8n take over the repetitive work. Start small, log everything, and add one branch at a time. Open Picasso IA, run your first scene, and see how far one good prompt can carry a faceless channel.

Share this article