Generate videosEdit videosLarge Language Models

AI Video Editing API and SDK: Automate Edits in Your App

Add video editing to your own product through code. See how an AI video editing API works, which edits belong to the API and which to FFmpeg, how to wrap the calls in a small SDK, and how to stay under the concurrency limit. Includes cURL and Node examples.

AI Video Editing API and SDK: Automate Edits in Your App
Cristian Da Conceicao
Founder of Picasso IA

Most video edits do not need a person at a timeline. Swapping the background on 200 product clips, turning a written brief into five vertical cuts, trimming every upload to 15 seconds: these are jobs for code. An AI video editing API lets your app submit that work as HTTP requests and collect finished clips, and an SDK, even a small one you write yourself, keeps those calls tidy. This article shows what such an API can do today, how the PicassoIA developer API works, and how to chain generative edits with plain FFmpeg steps into one automated video pipeline. Where a feature only exists in the web app, the text says so.

What an Editing API Actually Does

Before writing any code, split the word "edit" into two jobs, because they need different tools.

Generative edits versus timeline edits

Timeline edits are deterministic. Cut at 1.0 seconds, join two clips, burn in captions, resize to 9:16: the same input always gives the same output, and FFmpeg does it on your own server. Generative edits are probabilistic. A model re-renders pixels from a prompt, so "make the sofa purple leather" or "animate this photo" can look slightly different on every run unless you fix the seed.

A production pipeline almost always needs both. This is how the common jobs split:

JobTypeWhere it runs
Trim, merge, resizeTimelineFFmpeg on your backend, or Trim Video and Video Merge in the web app
CaptionsTimelineFFmpeg with a transcript, or Autocaption in the web app
Change a color, object or background in footageGenerativeP Video Edit or Lucy Edit 2 in the web app
Erase an objectGenerativeVideo Erase Object in the web app
Animate a still or render a new shotGenerativePicassoIA Video through the API
Edit a still, then re-render the shotGenerativePicassoIA Image Editor Pro plus PicassoIA Video through the API

Where an SDK fits

An SDK is the layer that keeps raw HTTP out of your business logic. It attaches the token, creates jobs, polls for results, retries the right failures, cancels stuck jobs, and limits how many run at once. Because video jobs are asynchronous (you create, you wait, you fetch), almost all the awkward code lives in that waiting. The API page ships examples in Python, Node and cURL. That is enough to build a thin client of your own, which is exactly what the later sections do.

Video editor reviewing a multi-track timeline in a brick-walled studio

What PicassoIA Exposes Today

The developer API lives at https://api.picassoia.com/v1. You authenticate with a Bearer token that starts with pia_sk_, which you create on the API page of picassoia.com (an account can hold up to two). The shape is Replicate-style: you create a prediction, poll it, and read the result.

ActionRequest
Create a jobPOST /v1/models/{owner}/{name}/predictions
Check a jobGET /v1/predictions/{id}
Cancel a jobPOST /v1/predictions/{id}/cancel
List your jobsGET /v1/predictions

Four models behind the API

ModelJobWorth knowing
PicassoIA ImageText to image7 aspect ratios, up to 2 outputs per call
PicassoIA Image Editor ProEdit a still with a promptUp to 3 reference images per edit
PicassoIA VideoText or image to video5 seconds, 24 fps, synchronized audio, 480p or 720p
Seedance 2.5 LiteText or image to video5 or 10 seconds, first and last frame, audio

💡 Plan around this split. As of October 2026, the video editing models in the catalog, such as P Video Edit, Aleph 2 and Lucy Edit 2, run in the web app, not through the API. Use the API to generate and re-render, and the web app for prompt-based edits of existing footage.

Limits to design around

  • 5 concurrent predictions per account, shared across every token and MCP connection.
  • 10 MB request body. Send image and video URLs, never base64 files.
  • 4,000 characters per prompt.
  • 3 hours before a prediction times out.

💡 Access rules and plan requirements change, so confirm them on the API page before you promise a roadmap to your team.

Overhead view of a desk with a laptop, notebook and memory cards

Send Your First Request

Every call below reads one environment variable, PICASSOIA_TOKEN, so the token never lands in your source code.

Create a job with cURL

curl -X POST https://api.picassoia.com/v1/models/picassoia/picassoia-video/predictions \
  -H "Authorization: Bearer $PICASSOIA_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "input": {
      "prompt": "Slow push-in on a ceramic mug of coffee on a sunlit desk, steam rising, soft room tone",
      "image": "https://example.com/first-frame.jpg",
      "resolution": "720p"
    }
  }'

The response hands back a prediction id and a status. The input object follows the model schema. For PicassoIA Video that means a required prompt plus optional image, resolution (480p or 720p, default 720p), aspect_ratio, seed and save_audio. When you pass an image, it becomes the opening frame and the clip inherits its aspect ratio. Every clip runs 5 seconds at 24 fps with synchronized audio, unless you switch save_audio off.

Poll until the clip is ready

Jobs are asynchronous, so the first response is a receipt, not a video. Ask for the prediction every few seconds with GET /v1/predictions/{id} until the status says it succeeded or failed, then read the output URL. Example runs on the model pages finish in roughly 30 seconds to 2 minutes, so a 3 to 5 second polling interval is plenty. If you stop caring about a job, call the cancel endpoint so it does not hold one of your five slots.

Developer's hands typing at a workstation with code on a blurred monitor

Build a Small SDK Wrapper

A wrapper in under 40 lines

Wrap the calls you need in one module. This Node version (18 or later, so fetch is built in) does the job:

const BASE = "https://api.picassoia.com/v1";
const headers = {
  Authorization: `Bearer ${process.env.PICASSOIA_TOKEN}`,
  "Content-Type": "application/json",
};

export async function createPrediction(model, input) {
  const res = await fetch(`${BASE}/models/${model}/predictions`, {
    method: "POST",
    headers,
    body: JSON.stringify({ input }),
  });
  if (!res.ok) throw new Error(`Create failed: ${res.status} ${await res.text()}`);
  return res.json();
}

export async function waitFor(id, { everyMs = 4000, timeoutMs = 10 * 60_000 } = {}) {
  const started = Date.now();
  while (Date.now() - started < timeoutMs) {
    const res = await fetch(`${BASE}/predictions/${id}`, { headers });
    const prediction = await res.json();
    if (prediction.status === "succeeded") return prediction;
    if (prediction.status === "failed" || prediction.status === "canceled") {
      throw new Error(`Prediction ${id} ${prediction.status}`);
    }
    await new Promise((r) => setTimeout(r, everyMs));
  }
  await fetch(`${BASE}/predictions/${id}/cancel`, { method: "POST", headers });
  throw new Error(`Prediction ${id} timed out`);
}

Three habits keep it safe in production. Retry only what can succeed on a second try: network drops and 5xx responses get two or three attempts with growing delays, while 4xx errors such as bad input or a bad token do not, because repeating them changes nothing. Always set a timeout and cancel on expiry, as the code above does, so a stuck job never holds a slot. Log the prediction id next to your own job id, because it is the first thing you will need when something looks wrong. Confirm the exact response fields against the API docs before you ship.

Stay under five jobs at once

The account ceiling is five concurrent predictions, shared across tokens and MCP connections. A batch of 40 clips fired with Promise.all would hit it immediately. Put a small pool in front of the wrapper and hold it at 4, so one slot stays free for manual tests or another tool on the same account:

export function pool(limit = 4) {
  let active = 0;
  const queue = [];
  const next = () => {
    if (active >= limit || queue.length === 0) return;
    active++;
    const { task, resolve, reject } = queue.shift();
    task().then(resolve, reject).finally(() => { active--; next(); });
  };
  return (task) => new Promise((resolve, reject) => { queue.push({ task, resolve, reject }); next(); });
}

const run = pool(4);
const clips = await Promise.all(
  shots.map((shot) => run(async () => {
    const job = await createPrediction("picassoia/picassoia-video", shot);
    return waitFor(job.id);
  }))
);

Developer in profile at a standing desk working on code

Automate Edits Inside Your App

With the wrapper in place, an automated edit is a short chain of steps:

  1. A plan: an LLM turns a brief into a JSON edit list.
  2. Stills: generate or edit images through the API.
  3. Shots: render new clips from those stills.
  4. Timeline work: trim, merge and caption with FFmpeg.
  5. Review and delivery: check each file, then publish.

The next three patterns fill in the middle of that chain.

Edit the first frame, then animate

The API cannot take your footage and apply a text edit to it, but it can get close. Pull a frame from the source clip, edit that still with PicassoIA Image Editor Pro, then render a new shot that starts from the edited frame with PicassoIA Video or Seedance 2.5 Lite.

  1. Grab a frame: ffmpeg -ss 2 -i source.mp4 -frames:v 1 frame.jpg
  2. Upload it to storage so it has a public URL.
  3. Send it in the images array and call it "image 1" in the prompt. The editor takes up to three reference images, and the example edits on its page return in about one to two seconds, so you can reject a bad edit before you spend a render on video.
  4. Pass the edited image as image to a video model with a motion prompt.
const edit = await waitFor((await createPrediction("picassoia/picassoia-image-editor-pro", {
  prompt: "Change the sofa in image 1 to light purple leather. Keep everything else unchanged.",
  images: [frameUrl],
})).id);
const firstFrame = Array.isArray(edit.output) ? edit.output[0] : edit.output;

const clip = await waitFor((await createPrediction("picassoia/picassoia-video", {
  prompt: "Slow push-in toward the sofa, soft window light, a hand places a cushion.",
  image: firstFrame,
  resolution: "720p",
})).id);

This re-renders the shot instead of editing the original pixels, so the motion will differ from your source footage. Treat it as a way to produce a variation, not a frame-accurate retouch. Seedance 2.5 Lite also accepts a last_frame_image and a 10 second duration, which helps when a shot has to land on a specific frame.

Open-plan creative studio seen from above with people editing video

Let an LLM write the edit plan

Hard-coding every edit does not scale. Let a language model turn a plain-language brief into a JSON edit list, then have your code run that list. GPT 5 Structured is built to return clean JSON, and Claude Sonnet 5 and Gemini 3.5 Flash are good drafting partners when you test plans by hand in the web app. In production, call whichever LLM provider your app already uses.

{
  "shots": [
    { "source": "clip_01.mp4", "start": 1.0, "end": 4.5, "caption": "New arrivals" },
    { "generate": "Slow dolly toward a sunlit storefront, shallow depth of field", "seconds": 5 }
  ]
}

Never run a plan blindly. Validate it against a schema, reject unknown fields, clamp durations, and cap the shot count. The model proposes; your code decides.

Product manager adding a sticky note to a workflow whiteboard

Trim, merge, and caption with FFmpeg

Timeline steps stay deterministic and cheap. Run these from Node with child_process or from any job runner:

# trim 3.5 seconds starting at 1.0
ffmpeg -ss 1.0 -t 3.5 -i clip_01.mp4 -c:v libx264 -c:a aac trimmed.mp4

# merge the clips listed in list.txt (same codec, size and frame rate)
ffmpeg -f concat -safe 0 -i list.txt -c copy merged.mp4

# burn captions from an SRT file
ffmpeg -i merged.mp4 -vf subtitles=captions.srt -c:a copy final.mp4

Merging with -c copy only works when every clip shares the same codec, size and frame rate. Clips from PicassoIA Video all come out at 5 seconds and 24 fps, but if you mix them with phone footage, re-encode everything to one spec first. The web app offers the same jobs by hand through Trim Video, Video Merge and Autocaption.

Editor reviewing a vertical clip beside a smartphone on a tripod

Use P Video Edit on PicassoIA

When you need a prompt-based edit on footage you already shot, the tool is P Video Edit. It runs in the PicassoIA web app and accepts a clip up to 15 seconds long. It changes the video by following a plain text instruction, so a request like "change the sky to sunset" or "make the jacket red" needs no timeline.

Step by step in the web app

  1. Open the P Video Edit page and upload your clip (15 seconds maximum).
  2. Write one instruction, for example: Change the material of the sofa to light purple leather. Do not change anything else.
  3. Optional: attach up to four reference images (jpg, jpeg, png or webp) when a color, texture or object has to match exactly.
  4. Switch on Draft for a faster, lower-quality preview before the final render.
  5. Leave Prompt Upsampling on for short instructions. Turn it off when your prompt is already precise.
  6. Keep Save Audio on so the original soundtrack stays in sync.
  7. Run it, review the result, adjust the prompt, and run again. Set a seed if you want to repeat a result exactly.

Example runs on the model page took roughly one to two minutes each.

Prompts that keep the scene intact

Name what must change, then name what must not. One example prompt on the model page follows that pattern: Change only the SUV body paint to yellow. It ends with Keep the environment, lighting and camera movement unchanged. Keep it to one change per run, then chain runs if you need more.

The catalog holds more editors. These are worth a test run on the same clip:

ModelWhat it does
Aleph 2Edit one frame and restyle the full video
Lucy Edit 2Edit any video with a text prompt
Wan 2.7 VideoeditEdit videos by text
LTX 2 RetakeEdit one section of a video
Video Erase ObjectRemove objects from footage
Video Remove BackgroundRemove a background without a green screen
Reframe VideoChange the aspect ratio
Video To SFX v1.5Add realistic sound effects

Freelance video creator checking a clip next to a mirrorless camera

Costs, Failures, and Guardrails

Plan for failures and cost

Treat every prediction as something that can fail. A failed job is final, so resubmit it as a new job, cap retries at two, and store the original input so you can replay it. Meter each job by model and resolution in your own logs: prices and plan rules change, so read the current terms on the API and pricing pages before you quote a customer a per-clip cost. Copy finished files into your own storage right after success, and treat the result URL as a delivery link, not an archive.

Screen inputs before rendering

If users can type prompts or upload images, check them first. Llama Guard 4 12B is a content moderation model you can try in the web app, and the same idea applies with whatever moderation service your stack already uses. Add a per-user rate limit too, so one customer cannot take all five slots.

Run Your Own Edit Today

The fastest way to see the pipeline is to run its pieces by hand. Open PicassoIA, make a still with PicassoIA Image, change one detail with PicassoIA Image Editor Pro, then animate the result with PicassoIA Video. A few minutes of experiments will show you which prompts hold up before you write a line of integration code. Run a second pass on the same still with Seedance 2.5 Lite and compare the motion.

When the results look right, wire the same steps into your app with the wrapper from this article. Browse every model, including the video editors, at picassoia.com/en/all-models, and create your own images today.

Professional at a cafe window table with a laptop and headphones

Share this article