Generate imagesGenerate videosVisual Effects

Grok Imagine API Pricing: Moderation, Censorship and Image to Video

Grok Imagine API pricing in plain numbers: $0.02 to $0.05 per image and $0.02 to $0.08 per second of video. See what a five-second image to video clip costs, how xAI moderation blocks requests, why users call it censorship, and which alternatives keep working.

Grok Imagine API Pricing: Moderation, Censorship and Image to Video
Cristian Da Conceicao
Founder of Picasso IA

A five-second clip from the Grok Imagine API costs between $0.10 and $0.40, depending on the video model you call, and a single image costs between $0.02 and $0.05. Those numbers look harmless until you run a few thousand generations, throw away a quarter of them, and find out that a filter blocked the one prompt your client cared about. This article puts the real figures on the table: what xAI charges per image and per second, how an image to video job adds up, how moderation behaves, why so many users call it censorship, and what to do when a request keeps getting refused. If you would rather test an idea before writing any API code, there is also a short tutorial for animating photos with Grok Imagine Video 1.5 on Picasso IA.

What the Grok Imagine API Includes

Grok Imagine is xAI's media family: image generation, image editing and video generation, all reachable through the same REST API and account that power the Grok text models. Image calls return quickly. Video calls are asynchronous: you submit a job, receive a request ID and poll until the status reads done, failed or expired.

The Image Models

xAI currently sells three image tiers:

  • grok-imagine-image: the base model, built for volume and drafts.
  • grok-imagine-image-2.0: the newer model, the tier xAI's model page recommends for images.
  • grok-imagine-image-quality: the highest fidelity tier, meant for hero shots and final art.

A single generation request can return up to 10 images, and an edit request accepts up to 5 source images, supplied as a public URL or a base64 data URI. The same model families are available on Picasso IA as Grok Imagine Image 2 and Grok Imagine Image Quality, which is a convenient way to compare their output before you pay per call.

The Video Models and Modes

Three video models sit on the price list: grok-imagine-video-1.5-lite, grok-imagine-video and grok-imagine-video-1.5. Between them they handle five modes:

  • Text to video: a prompt in, a clip out.
  • Image to video: animate a still with a motion prompt.
  • Reference to video: steer a new clip with one or more reference images.
  • Video editing: change an existing clip with a text prompt.
  • Video extension: continue a clip from its last frame.

Clips run up to 15 seconds, with settings for aspect ratio and resolution. A job starts with a POST to https://api.x.ai/v1/videos/generations and finishes by polling https://api.x.ai/v1/videos/{REQUEST_ID}. If you want to try the same families without code, Picasso IA lists Grok Imagine Video and Grok Imagine R2V.

Developer hands typing on a laptop in a quiet home office

Grok Imagine API Pricing Breakdown

The prices below come from xAI's public model page, checked in October 2026. xAI adjusts rates from time to time, so confirm them before you lock in a budget.

ModelOutputPriceBest for
grok-imagine-imageImage$0.02 per imageDrafts and bulk thumbnails
grok-imagine-image-2.0Image$0.04 per imageEveryday production images
grok-imagine-image-qualityImage$0.05 per imageHero shots and final art
grok-imagine-video-1.5-liteVideo$0.020 per secondPreviews and motion tests
grok-imagine-videoVideo$0.050 per secondSteady mid-tier work
grok-imagine-video-1.5Video$0.080 per secondFinal renders

xAI states that both duration and resolution affect the total cost of a video. Some third-party price trackers list separate 480p and 720p rates for the base video model (for example $0.05 and $0.07 per second), but xAI's own table does not split them. Run a small test batch and read your billing dashboard rather than trusting any blog, this one included.

Overhead view of invoices, a calculator and a notebook of cost columns on a walnut desk

What One Clip Really Costs

Video is billed per second, so the math is a straight multiplication:

Clip lengthLiteBaseVideo 1.5
5 seconds$0.10$0.25$0.40
10 seconds$0.20$0.50$0.80
15 seconds$0.30$0.75$1.20

An image to video job has two parts: the still and the motion. A $0.04 still from grok-imagine-image-2.0 plus a five-second clip on grok-imagine-video-1.5 comes to $0.44. Use the lite model for the motion and the same pair costs $0.14. If you already own the photo, the still costs nothing and you pay for the clip alone, give or take a small input fee that some trackers list at $0.002 per image.

Retries Change the Math

The sticker price is not the real price. Divide it by your success rate. If one generation in four gets thrown away, because of a visual glitch, the wrong motion or a moderation block, a $0.40 clip really costs about $0.53.

Four habits keep that number down:

  • Draft on lite, render on 1.5. Test the motion prompt at $0.10 per five seconds, then spend $0.40 only on the version you plan to publish.
  • Lock the still first. Rerolling a still costs $0.02 to $0.05. Rerolling a clip costs up to 20 times more.
  • Keep clips short. You pay for every second, so ask for five unless the story truly needs more.
  • Log every request ID. When the invoice and your delivered clips disagree, the log tells you why.

💡 Budget tip: Make the lite model your default and promote only the winners to Video 1.5. On a 200-clip month, that habit can cut the bill by almost half.

Monthly Budget Examples

WorkloadSetupMonthly cost
1,000 thumbnailsgrok-imagine-image$20
1,000 hero imagesgrok-imagine-image-quality$50
200 five-second clipsAll on grok-imagine-video-1.5$80
200 lite drafts plus 60 finalsLite, then grok-imagine-video-1.5$44
500 images and 150 clipsImage 2.0 and Video 1.5$80

Scale turns the question from price into throughput. The model table I checked does not publish rate limits, so ask xAI for your account's limits before you schedule a batch of several thousand jobs.

Technician walking through a data center aisle lined with server racks

Image to Video Costs in Practice

Still First, Motion Second

The cheapest workflow treats the still as the real creative decision. Pick a photo with a sharp subject, clean edges and room for the movement you want. A cluttered frame forces the model to invent too much, and invention is where visual glitches come from. Generate or choose three candidate stills, pick the best, and only then spend on motion.

That order also protects you from moderation surprises. The source photo is part of what gets reviewed, so a borderline still can be refused at the clip stage, after you already paid for the image.

Photographer's fingers holding a printed landscape photo in front of a monitor

Prompts That Move the Right Things

A good motion prompt names one subject, one movement and one camera move. Everything else stays still, which is exactly what you want. A simple formula: subject and starting pose, then motion over five seconds, then camera, then light.

  • Weak: "Make it cinematic and epic."
  • Strong: "A woman at a window slowly turns toward the sunlight, gentle push-in from the camera, warm light brightening across her face."

Three mistakes show up again and again:

  1. Too many actions. Asking for a walk, a wave and a turn in five seconds gives you a blur.
  2. Style words instead of motion. "Cinematic" is a look, not a movement.
  3. Contradicting the still. If the photo shows midday, a prompt about candlelight fights the source image.

Filmmaker sketching a camera movement arrow on a paper storyboard

How Grok Imagine Moderation Works

xAI states that generated media is subject to content policy review. Its Acceptable Use Policy asks users to follow the law, avoid harmful use, respect privacy and copyright, respect limits on sexual content involving people, and not try to bypass safeguards. None of that is unusual. What surprises people is how often it shows up in an ordinary workflow.

Content reviewer at a dual monitor workstation checking image thumbnails

What Gets Checked

Secondary reports describe guardrails that scan both the prompt and the generated output. A harmless sentence can still be refused if the result looks wrong, and a risky sentence can be refused before anything renders. For image to video, treat the uploaded still as part of the input.

Commonly reported triggers:

  • Real people in sexual or demeaning scenes
  • Minors in any suggestive context
  • Famous characters and brand logos
  • Graphic violence
  • Prompts that try to talk around the rules

What a Blocked Request Looks Like

Reports describe an invalid_argument error when moderation blocks a video request. The pages I checked do not list moderation error codes, and they do not say whether a blocked generation is billed. Until you have your own answer, treat it as billable: compare the number of delivered clips with the charges on your first invoice.

💡 Handle refusals in code: Catch the error, log the prompt and source image, and never auto-retry the same input. A retry loop on a blocked request is the fastest way to burn budget.

Censorship Complaints and Real Limits

Moderation means rules you can read and predict. Censorship is what it feels like when a rule blocks work that looked harmless. Most complaints about Grok Imagine fall into two buckets.

Likeness and Adult Content Limits

xAI's Acceptable Use Policy prohibits depicting likenesses of real people in a pornographic manner, and that rule applies to every user regardless of subscription level. Photos of real people are therefore the most sensitive input you can send. Non-explicit glamour and swimwear sit in a gray zone where results can vary from one request to the next, which frustrates professionals who need predictable output.

False Positives on Normal Work

Fashion shoots, beach holidays, medical diagrams, boxing matches and historical photographs all get flagged now and then. Safer habits:

  • Describe the movement, not the body.
  • Choose a still with clear context, such as a runway, a studio or a gym.
  • Avoid ambiguous words that read differently out of context.
  • Stop after two refusals and change the input instead of rerolling.

⚠️ Rewording a prompt to slip past a filter breaks xAI's rules, which prohibit attempts to bypass safeguards. Make the legitimate intent of your prompt clearer. Do not try to hide it.

Gallery attendant beside a velvet rope and a partly draped framed photograph

Alternatives When Filters Block You

If your project involves adult but non-explicit themes such as swimwear, glamour or fashion, a strict filter keeps refusing you, and every refusal costs money. Picasso IA hosts models with looser rules. Explicit content stays off the table everywhere, but these options handle non-explicit work:

  1. Seedream 4.5: very realistic output, built-in image editing and results in under 3 seconds. The best first stop.
  2. PicassoIA Image Editor Pro: an image to image editor with unlimited generations on Elite and Infinite plans, results in under a second and a free trial of 3 generations with no credit card. Generating 1,000 images costs nothing extra there, against roughly $100 on typical pay-per-image models.
  3. Qwen Image 2: open source, with very detailed realism.
  4. P-Image: text to image in under 1 second.
  5. Wan 2.2 I2V Fast: turns still images into smooth short videos.
  6. PicassoIA Video: unlimited video generation at up to 720p and 5 seconds.
  7. P-Video: up to 1080p, with a draft mode for instant previews.
  8. Grok Imagine Video: clips up to 15 seconds with no watermarks.
  9. LTX 2.3 Pro: up to 4K at 50 fps with retake and extend editing.

For comparison, 1,000 images at xAI's cheapest tier cost $20, and 1,000 at the quality tier cost $50. An unlimited plan changes that math entirely for high-volume creators.

How to Animate Photos on PicassoIA

Here is a quick test path before you write any API code. Grok Imagine Video 1.5 on Picasso IA takes a still and a motion description, then returns a clip with synchronized audio. The model page lists clips of up to 5 seconds per run, shorter than the 15 seconds the API allows, which is plenty for a motion test.

  1. Open the model page for Grok Imagine Video 1.5.
  2. Upload your still as a JPG, JPEG, PNG or WEBP file. No still yet? Generate one with Grok Imagine Image 2 first.
  3. Write the motion prompt with the formula above: one subject, one movement, one camera move.
  4. Set the options. Leave duration at 5, resolution at 720p and aspect ratio on auto.
  5. Run it and review. The example on the model page took about 33 seconds. Run two or three prompt variations on the same image and keep the best.
  6. Download the clip and drop it into your edit.
ParameterDefaultTip
ImageRequiredSharp subject with room for movement
PromptRequiredOne subject, one movement, one camera move
Duration5Leave at 5 seconds, the longest value listed
Resolution720pDrop to 480p for quick drafts
Aspect ratioAutoPick 9:16 for vertical posts, 16:9 for the web

Want a second opinion on the same still? Run it through these models and compare motion, audio and refusals side by side:

ModelStrength
Seedance 2.0Built-in audio and steady motion
Veo 3.11080p output from text or images
Kling v3 VideoCinematic camera movement
Wan 2.7 I2VAnimates almost any photo

Two colleagues reviewing printed storyboard frames at a studio table

Try It Yourself on Picasso IA

The cheapest way to find out whether Grok Imagine fits your project is to run one still through several models and compare cost, quality and refusals side by side. Open Picasso IA, upload a photo, write a one-movement prompt and render it with Grok Imagine Video 1.5. Then try Seedance 2.0 or Wan 2.7 I2V on the same image. Ten minutes of experiments will tell you more than any pricing page.

Young woman at a cafe window holding a tablet in golden light

Browse the full catalog of image and video models at picassoia.com/en/all-models, then start creating your own images and videos today.

Share this article