Grok Video Generation API Pricing: Image to Video and Limits
Grok video generation API pricing comes down to the per-second rate, the image input fee and your retry count. See image to video costs at 480p, 720p and 1080p, the 15 second cap, the 10 requests per second limit and the math for a 20 clip budget.
If you are pricing a Grok video generation API integration, the rate on the pricing page is not the number that matters. What matters is what you pay for a clip you will actually publish. xAI bills Grok Imagine video by the second of output, adds a small fee for every image you send in, and caps request volume at 10 requests per second per model. Combine those three facts with your retry rate and you get your true budget.
This page lays out the per-second rates for both xAI video models, the image to video billing rules, the resolution and duration limits, and the math for a real batch of clips. Where the official docs and third-party trackers disagree or go quiet, I say so. Everything here was checked against the xAI video generation docs and public pricing trackers in early October 2026, so confirm the live numbers in your xAI console before you lock a budget.
💡 Short version: the older model starts at $0.05 per second, the newer 1.5 model starts at $0.08 per second, 1080p exists only on 1.5, and a single clip runs 1 to 15 seconds.
What xAI Actually Charges
xAI sells two video models through the same endpoint, POST https://api.x.ai/v1/videos/generations. They share the request shape and the same rate limit, but not the price, the input types or the top resolution.
Rates Per Second
The official model pages list one headline output rate for each model:
Grok Imagine Video (grok-imagine-video): $0.050 per second. Accepts text, image and video input.
Grok Imagine Video 1.5 (grok-imagine-video-1.5): $0.080 per second. Accepts text and image input only.
The official pages do not break those rates down by resolution. A third-party pricing summary published on September 8, 2026 does, and it is the table most developers budget from:
Look at the 480p column: it matches the headline rates exactly. That means the official number is the floor, not the typical bill. A 720p clip on 1.5 costs 75 percent more per second than the headline suggests, and a 1080p clip costs more than three times as much.
💡 Treat the resolution split as reported, not official. xAI's own pages show only the entry rate. Run one 5 second test at each resolution and compare it against your console's usage page before you scale.
The Image Input Fee
Image to video adds a flat charge per source image on top of the output seconds. The reported fees are $0.002 per image on the base model and $0.01 per image on 1.5. On a single clip that rounds to nothing. On 50,000 clips it is $100 versus $500 in image fees alone, so it belongs in your cost model from day one.
How Image to Video Billing Works
Image to video is the mode most teams actually ship. A still photograph gives you control over the subject, the framing and the look before you pay for a single second of motion.
The Image Is the First Frame
When you include the optional image field, xAI uses that picture as the starting frame of the clip. If you do not override the aspect ratio, the output follows the input image's proportions instead of the 16:9 default. The bill for each attempt comes down to one line:
cost per attempt = (seconds × rate for the chosen resolution) + image fee
I found no separate prompt or token fee in the docs, so the text prompt appears to carry no extra charge. Prompts do have a length limit, though, and an oversized one returns an invalid_argument error.
💡 Generate your first frames with a photoreal still model such as PicassoIA Image before you pay for motion. A flawed still becomes a flawed clip, and fixing a still costs a fraction of re-rendering video.
A Minimal Request
xAI's own example for image to video through the Python SDK is short:
response = client.video.generate(
prompt="Slow push-in on the cliff as waves roll in",
model="grok-imagine-video-1.5",
image="https://example.com/cliff.jpg"
)
Under the hood this is the same asynchronous endpoint as text to video. You submit the job, receive a request_id, and poll until the job finishes. Duration, resolution and aspect ratio are set as request parameters, and the next section lists their limits.
Resolution and Duration Limits
Limits decide which ideas the API can produce at all, and they set the ceiling on your per-clip cost. Three settings matter: resolution, duration and aspect ratio.
480p, 720p and 1080p
480p is the default and the cheapest tier. It suits storyboards, drafts and small social thumbnails.
720p is the practical delivery tier for most web and social clips.
1080p is available only on Grok Imagine Video 1.5, and only for text to video and image to video.
Reference to video is capped at 720p.
Because the default is 480p, a request that omits the resolution field returns the cheaper, softer output. Set it on purpose.
One to Fifteen Seconds
The duration parameter accepts 1 to 15 seconds for standard generation. Billing is per second, so cost scales in a straight line: a 10 second clip costs exactly twice a 5 second clip at the same resolution.
Video editing is the exception. It takes no custom duration, keeps the source video's length, and is capped at 8.7 seconds.
Most marketing clips fit in 5 to 6 seconds. Paying for 15 seconds when your edit uses 4 is the most common source of waste.
Supported Aspect Ratios
Seven ratios are supported: 1:1, 16:9, 9:16, 4:3, 3:4, 3:2 and 2:3. Text to video defaults to 16:9, while image to video follows the input image unless you override it. Vertical 9:16 is there for short-form feeds, so you do not need to crop after rendering.
Rate Limits and Throughput
Pricing tells you what a clip costs. Rate limits tell you how fast you can buy them.
Ten Requests Per Second
Both video models list the same ceiling in xAI's docs: 10 requests per second. That works out to 600 submissions per minute. For most products this is not the bottleneck, because video jobs render slowly and you rarely submit more than a handful per second. It starts to matter during bulk backfills, where a naive loop that fires thousands of jobs at once will be throttled.
The model pages list no daily cap, and they do not say how many jobs a single account can render at the same time. Team level quotas can differ, so check the limits page in your console before a large run.
Both models run in us-east-1, us-west-2 and us-saltlake-2, and the Batch API is supported in us-east-1 and us-west-2.
Async Jobs and Expiring URLs
Video generation is asynchronous. You submit a request, receive a request_id, and poll until the status changes. The documented states are pending, done, expired and failed.
Finished videos arrive as temporary URLs. xAI tells developers to download or process the file promptly if they want to keep a copy. In practice:
Poll every few seconds with a gentle backoff instead of hammering the endpoint.
The moment a job reports done, copy the MP4 into your own storage.
Save the request_id next to your own record, so you can trace any clip back to its prompt and source image.
The Real Cost of a Clip
The per-second rate is the price of an attempt. The price of a finished clip is the price of an attempt multiplied by the number of takes you need.
Budget Math for 20 Clips
Here is output cost for single attempts, using the reported per-resolution rates. Add the image fee ($0.002 or $0.01) to each row.
Now take a realistic job: 20 finished clips, 5 seconds each, 720p, one source image per clip. With a single attempt each, the base model costs 20 × $0.352 = $7.04 and 1.5 costs 20 × $0.71 = $14.20. Those are best-case numbers.
Why Retries Decide the Bill
Image to video is not deterministic. Shots with people, hands or fast camera moves tend to need two or three takes, while a static landscape often lands on the first. The same 20 clip job looks like this once you count takes per keeper:
Draft at 480p, finish at 720p. Five 480p drafts of one shot on 1.5 cost 5 × $0.41 = $2.05, and the 720p final adds $0.71, for $2.76. Five takes at 720p cost $3.55. Treat the draft as a prompt check, because a different resolution is a new generation, not a preview of the same frames.
Write one motion per prompt. "Slow push-in while the waves roll" beats a paragraph of competing instructions.
Start from clean stills. Clear subjects, sharp edges and room to move reduce artifacts.
Moderation and Failed Jobs
Responses carry a respect_moderation flag that tells you whether the video passed moderation. Prompts that run too long or break content policy return an invalid_argument error, and a job can also end as failed.
I could not find a statement on whether failed or moderated jobs are billed. Do not assume they are free. Run a small test, watch your usage page, and log every request_id so you can dispute a charge with evidence.
Audio, Editing and Reference Modes
The per-second rate is not the only way you will use these models. Three extra modes change what a clip can do.
Native audio. On PicassoIA, Grok Imagine Video 1.5 generates sound that matches the motion in the frame, so a rainy street comes with rain. The xAI pricing pages list no separate audio charge, which makes sound effectively part of the per-second rate.
Reference to video. Only the base model accepts video as an input type, and that is what reference to video and video extension rely on. If you want a subject to stay consistent from a set of photos, look at Grok Imagine R2V, and remember the 720p cap.
Video editing. Editing keeps the original length and stops at 8.7 seconds, so it suits short fixes and restyles, not long edits.
Run Grok Imagine on PicassoIA
If you want clips without writing API code or opening an xAI billing account, you can run the same model family in the browser. PicassoIA hosts Grok Imagine Video 1.5 as an image to video model with synchronized audio.
Upload a JPG, PNG or WEBP still. This is your first frame.
Describe the motion in one or two sentences.
Choose the duration (up to 5 seconds), the resolution (up to 720p) and the aspect ratio: 16:9, 9:16, 4:3, 1:1, 3:2 or auto-detected from your image.
Generate, preview and download the clip with its audio.
The model page lists 1.6 credits per second of video, so a 5 second clip uses 8 credits. A free trial is available before paid subscription charges apply.
Settings That Save Credits
Keep clips at 5 seconds or less. You pay by the second, and short loops are what most feeds use.
Test a prompt at a lower resolution before you commit to your final.
Reuse a strong still across several prompts instead of generating a new image each time.
Let the aspect ratio follow your image, so the model does not have to crop or invent edges.
Cheaper Models to Compare
Grok is not the only route to a good image to video clip. These PicassoIA models are worth running side by side on the same still:
If you need an HTTP API rather than a browser, PicassoIA also runs a developer API at https://api.picassoia.com/v1 with Replicate style predictions. It currently serves PicassoIA Video and Seedance 2.5 Lite for video, allows 5 concurrent predictions per account, and does not include Grok. Check the access terms on the PicassoIA pricing page before you build on it.
Make Your First Clip Today
Pricing tables only go so far. The fastest way to see what a Grok clip is worth to you is to make one and judge it frame by frame.
Here is a simple plan. Generate a photoreal still with PicassoIA Image. Animate it with Grok Imagine Video 1.5. Run the same still through Seedance 2.5 Lite and compare motion, sound and cost. Ten minutes of testing will tell you more than any rate card.
Open Picasso IA, upload your own photo, and try your first image to video clip. Change one setting at a time, keep the takes you like, and write down what each one cost. That log becomes your real budget, and it is the number that will keep your API bill honest.