Large Language ModelsGenerate imagesGenerate videos
OpenRouter Image Generation in Open WebUI and SillyTavern: What Actually Works
OpenRouter lists image models from OpenAI, Google, ByteDance and Black Forest Labs, but Open WebUI and SillyTavern connect to them differently. See the exact settings, the gateway workaround for Open WebUI, SillyTavern's source and prefix options, cost traps, and a short error checklist.
You already pay for chat models through OpenRouter, so sending image requests through the same account looks like an easy win. No second subscription, no extra dashboard, one balance to watch. The catch is that Open WebUI and SillyTavern reach OpenRouter in very different ways. Open WebUI speaks an OpenAI-style images route, and the documentation I checked never names OpenRouter as an option. SillyTavern lists OpenRouter as a ready-made image source but says little about the setup. This article lays out both paths, marks the places where the documentation goes quiet, and gives you a short checklist for the moment a render comes back blank.
What OpenRouter Offers for Images
OpenRouter sits in front of many providers, and its catalog now includes image models from Google, OpenAI, Black Forest Labs, xAI, ByteDance, Microsoft, Recraft, Krea and Sourceful. You top up one credit balance instead of opening an account at each lab, and you switch models by changing a single string.
How the Image Endpoint Works
OpenRouter's current documentation describes a dedicated image route. You send a POST request to /api/v1/images with your bearer token, and the body carries the fields below.
Field
What it does
model
Model slug, such as bytedance-seed/seedream-4.5
prompt
The text description of the picture
aspect_ratio
A normalized ratio, or auto to let the provider choose
resolution
A tier from 512 up to 4K
size
Shorthand for explicit pixels, like 2048x2048
quality
auto, low, medium or high
output_format
png, jpeg, webp or svg
n
Number of images, from 1 to 10
stream
Sends partial previews as server-sent events
The reply puts each image in a data array as base64 text inside b64_json, next to a media_type such as image/png and a usage block with token counts and cost. You get no hosted link. Anything sitting between OpenRouter and your front end has to decode that text or save it as a file, and that detail explains several of the errors later in this article.
For image-to-image work, the same route accepts input_references, which can be HTTP(S) addresses or base64 data URLs.
đź’ˇ Tip: Older tutorials describe image output through the chat route with a modalities parameter, using ["image", "text"] for models that also write text and ["image"] for image-only models. If a tutorial and the current documentation disagree, trust the documentation and the model page you are actually calling.
Models Worth Trying
Start with a small shortlist instead of scrolling the whole catalog. The IDs below appear in OpenRouter's documentation or in community gateway examples, and the right column points to the same model family on PicassoIA so you can test prompts first.
The catalog changes often. Before you copy any ID, filter OpenRouter's models page by image output and confirm the slug letter for letter.
Before You Connect Anything
Ten minutes of preparation saves an evening of confusing errors. Both front ends need the same two things: a token you can throw away and a clear picture of what a render costs.
Create a Separate Token
Generate a new OpenRouter credential used only for image work. If a front end logs it, a roleplay chat loops, or you paste it into a screenshot by accident, you revoke one token and your regular chat setup keeps running. Name it after the front end, for example one token for Open WebUI and one for SillyTavern, so the activity log tells you which app spent what.
Check Credits and Billing
OpenRouter states that image billing is all-or-nothing. A generation either finishes and is billed in full, or it fails and is not billed. Partial preview images delivered during a streamed request do not create partial charges. That is good news for failed attempts, but it also means a loop of successful renders drains credit at full price, so keep a small balance while you test.
Open WebUI Setup
Open WebUI is the harder of the two, and the reason is not your configuration. It comes down to which request shape each side expects.
Where the Settings Live
Open the Admin Settings and find the Images section. The documentation gives the path as Settings, Admin, Experience, Images, and menu names shift between releases, so search for "Images" if yours differs. Set Image Generation Engine to Default (Open AI). You will see these fields:
API Base URL, the address requests are sent to
API credential, where your token goes
Model, a dropdown or a typed name
Image Size, limited to what the engine allows
The documented size lists for the OpenAI engine are 256x256, 512x512 and 1024x1024 for DALL·E 2, then 1024x1024, 1792x1024 and 1024x1792 for DALL·E 3, and auto, 1024x1024, 1536x1024 and 1024x1536 for the GPT-Image models.
The Route Mismatch Problem
Here is where setups stall. Open WebUI's OpenAI engine sends an OpenAI-style request with fields such as prompt, model, n, size, quality and a response format. OpenRouter's documented image route is its own /api/v1/images, with different fields like aspect_ratio and resolution. The Open WebUI pages I checked describe OpenAI itself, Azure OpenAI, a LiteLLM proxy and an Image Router style service. None of them mention OpenRouter.
⚠️ Heads up: I could not confirm that pointing the OpenAI engine straight at OpenRouter works in every release. Treat the direct route as an experiment. Send one test image, read the exact error, and only then decide whether you need a translation layer.
You have three realistic options:
Test the direct route. Enter https://openrouter.ai/api/v1 as the base URL, paste your token, type a model ID by hand, and generate one image.
Run a translation gateway. A small service accepts OpenAI-style image requests and forwards them to OpenRouter. The next section shows one.
Use another OpenAI-compatible image router. Open WebUI documents this pattern for services that copy the OpenAI syntax.
Using a Gateway in Docker
A community project on Docker Hub called OpenRouter Image Gateway exists for exactly this gap. According to its description it exposes POST /v1/images/generations, GET /v1/models and GET /health. It accepts OpenAI-style parameters (prompt, model, n, size, quality, response_format), forwards your bearer token to OpenRouter, converts pixel sizes into OpenRouter aspect ratios, and returns images as b64_json or as a URL.
With the gateway running, the Open WebUI fields look like this:
Image Generation Engine: Default (Open AI)
API Base URL: http://openrouter-image-gateway:8000/v1
API credential: <your OpenRouter token>
Model: google/gemini-2.5-flash-image
Use http://openrouter-image-gateway:8000/v1 when both containers share a Docker network, and http://localhost:8000/v1 for a local test on the same machine.
⚠️ Heads up: This is third-party software, and it will see your token. Read its source or run it on a machine you control before trusting it with a credential that has real balance behind it. The description also mentions text-to-image only, so do not expect image editing to work through it.
Picking Sizes and Models
Type the model name yourself instead of using the dropdown. Open WebUI's own Image Router instructions say to do exactly that for non-OpenAI providers, because the dropdown lists OpenAI names and will never show google/gemini-2.5-flash-image.
For size, choose the landscape option closest to the shape you want and check the output. When a gateway sits in the middle, it maps your pixel choice to the nearest aspect ratio, so a 1536x1024 request may come back as a clean 3:2 image rather than those exact pixels. That is fine for chat, but check before you build a workflow around exact dimensions.
SillyTavern Setup
SillyTavern takes the opposite approach. Image generation is a built-in extension, and OpenRouter is one entry in its source list.
Choose OpenRouter as the Source
The official documentation lists OpenRouter as a cloud source next to OpenAI, Black Forest Labs, FAL.AI, Google, x.AI, Stability AI and others. Open the Extensions panel, expand Image Generation, pick OpenRouter as the source, enter your token, and select an image model.
Be aware that the documentation gives OpenRouter no dedicated setup section, unlike sources such as Stability AI. Labels and field order can differ between releases, so treat the steps above as a map rather than a script.
Generation Modes You'll Use
SillyTavern builds the prompt for you from the chat, and the mode decides what it describes.
Mode
Slash command
What you get
Yourself
you
Full-body portrait of the current character
Your Face
face
Close-up portrait of the current character
Me
me
Portrait of your user persona
The Whole Story
scene
Visual recap of the chat events
The Last Message
last
Visual recap of the last message
Raw Last Message
raw_last
Last message sent verbatim as the prompt
Background
background
Chat background built from the story context
You can reach these three ways: the Image Generation item in the wand menu, the /sd command followed by a mode or your own free text, or the paintbrush icon on a single message for the raw mode. The command also accepts named arguments, for example negative="blurry, extra fingers".
The scene and last modes suit story-heavy chats, where a single picture can sum up what just happened at the table.
Prefixes That Keep Characters Consistent
Three text boxes decide how stable your pictures look from one render to the next:
Common Prompt Prefix is added before every prompt and sets the overall style.
Character-Specific Prompt Prefix describes one character's looks. It only works in one-to-one chats, not in groups.
Negative Prompt lists what you do not want to see.
A workable starting pair looks like this. Common prefix: candid 35mm photograph, natural window light, fine film grain. Character prefix: woman in her thirties, freckles, loose auburn braid, denim jacket. Keep the character prefix short and physical, and let the style prefix carry the lighting and lens language.
đź’ˇ Tip: The OpenRouter field list I reviewed has no negative prompt field, so that box may do nothing on this source. Describe what you want in positive terms inside the main prompt instead.
Prompts That Work in Both
Whichever front end you use, the model on the other end reads one text prompt. A little structure lifts results in both apps.
Write Photographic Prompts
Build each prompt from five parts: subject, setting, light, lens and texture. Chat-derived prompts from SillyTavern tend to be story summaries, which models handle poorly, so rewrite them into a camera description when the picture matters.
Weak prompt
Stronger prompt
a girl in a tavern
woman in a wool cloak at a candlelit tavern table, 35mm f/1.8, warm side light, wood grain and pewter cups in sharp focus
my room
small attic bedroom at dusk, low-angle shot, soft window light from the right, linen sheets, visible dust in the air
a battle scene
two riders on a muddy road at dawn, 70mm lens, overcast light, wet leather and mud splashes, shallow depth of field
If writing these by hand feels slow, ask a chat model to do the rewriting. Claude Sonnet 5 and Gemini 3.5 Flash both turn a rough scene recap into a camera-ready line in a few seconds.
Match Aspect Ratio to the Job
Pick the shape before you pick the model. A wrong ratio wastes a render.
Job
Suggested ratio
Reason
Character portrait
2:3 or 3:4
Fits a tall frame and a face-and-shoulders crop
Scene recap
3:2 or 16:9
Room for the full setting
Chat background
16:9
Matches a wide screen
Quick Open WebUI test
1:1
Cheapest way to confirm the route works
Costs and Limits That Surprise People
Most surprise bills come from settings, not from prices. The table lists the usual suspects.
Situation
Why it costs more
What to do
Interactive mode in SillyTavern
Messages with a verb like draw or send followed by a noun like photo or picture trigger a render
Switch it off for casual chats
n above 1
One request can return up to 10 images
Keep it at 1 while testing
High resolution tiers
The resolution field runs up to 4K
Start at a lower tier and raise it for finals
Retry loops
Failed runs are not billed, but repeated successes are
Stop after two or three attempts and fix the prompt
SillyTavern's interactive mode deserves a second look. It watches for action verbs such as send, make, draw, paint, render, imagine, create and mail, followed within a few characters by words like pic, picture, image, drawing, painting, photo or photograph. A roleplay line that happens to contain "draw a picture" can spend credit without you pressing anything.
đź’ˇ Tip: Check your OpenRouter activity page after the first ten renders. Compare the cost per image against what you expected, then adjust model and resolution before a long session.
Fixing Common Errors
When an image fails, the cause is usually one of four things: the route, the token, the model slug, or the base64 handling.
Blank Image or Error Toast
Work through this checklist in order:
404 or 405 error. The front end is calling a route the server does not have. Recheck the base URL and whether you need the gateway.
401 error. The token is wrong, was revoked, or the gateway is not forwarding the bearer header.
Model not found. The slug has a typo or a missing prefix. Copy it from OpenRouter instead of typing from memory.
File saves but will not open. The base64 text was stored without decoding. Because OpenRouter returns b64_json, the layer in the middle must turn it into an image file or a URL.
Nothing happens at all. Open your browser console or the server log and read the first red line before changing settings.
Wrong Size or Ratio Returned
OpenRouter thinks in aspect ratios and resolution tiers, while Open WebUI thinks in pixels. A gateway translates between them, and the translation is approximate. If you need an exact shape, call OpenRouter directly with aspect_ratio and resolution, or crop afterward. When the picture looks stretched, check whether the front end is forcing a fixed display box before blaming the model.
Test Prompts Before You Spend Credits
Every wasted render costs real money, so tune your prompts where iteration is cheap. PicassoIA puts GPT Image 2, Seedream 4.5, FLUX.2 Pro and Gemini 2.5 Flash Image in one place, so you can compare the same prompt across makers before you commit one to a SillyTavern prefix or an Open WebUI default.
Here is a quick routine that works:
Open a model page on PicassoIA and paste your stronger, camera-style prompt.
Set the aspect ratio you plan to use in the chat, such as 2:3 for portraits or 16:9 for backgrounds.
Generate two or three variations and note which phrases changed the result most.
Run the same prompt on a second model and keep the one that matches your character best.
Copy the winning wording into your Common Prompt Prefix or your Open WebUI prompt, then switch to OpenRouter for the long session.
Ready to try it? Open PicassoIA, run your three favorite prompts through two models, and see which one earns a place in your setup.