Large Language ModelsGenerate imagesGenerate videos
n8n AI Agent Image Generation: Workflow Example
This n8n AI agent image generation workflow example follows every node from the Chat Trigger to the final image URL: the system message, a sub-workflow that creates and polls predictions, memory for follow-up edits, and a table of fixes for the errors that stall most builds.
Typing a sentence into a chat box and getting a finished photograph back sounds like a small trick, until you wire it up yourself and watch an agent decide on its own what to write, which tool to call and when to stop. That is what an n8n AI agent image generation workflow does. A language model sits in the middle, reads your request, expands it into a detailed prompt, calls an image API as a tool, waits for the render and hands back a link.
This workflow example runs from the first node to the last. You get the node layout, the system message that keeps the agent honest, the HTTP calls that talk to an image model, the polling loop that stops empty results, and the mistakes that waste an afternoon. It works on any n8n instance, cloud or self-hosted, and every model named here has a page on PicassoIA.
What This Workflow Does
The workflow takes a plain request such as "a rainy street market at dusk, shot on 35mm film" and returns a hosted image URL. Between the request and the result, an AI Agent node makes the decisions. It never calls the image API itself. It calls a tool, and the tool does the heavy lifting.
The Idea in Plain Words
Five pieces make up the whole build:
Chat Trigger: receives the message from the n8n chat panel or a public chat link.
AI Agent: reads the message and decides what to do next.
Chat model: the brain behind the agent, attached as a sub-node.
Memory: holds earlier turns, so "make it wider" still makes sense.
Image tool: a sub-workflow that starts the render, waits for it and returns the URL.
That is one conversation loop. The user talks, the agent thinks, the tool renders, the agent answers. Nothing else is needed for a first version, and every later upgrade (approvals, storage, scheduling) bolts onto one of these five pieces.
Why an Agent Beats a Fixed Chain
A fixed chain (trigger, language model, HTTP request, reply) works until someone writes "give me three variations, and make one of them square." The agent can call the tool three times with different arguments, ask a clarifying question, or push back on a request that is too vague to picture. A chain can only do what you drew on the canvas.
Situation
Fixed chain
AI Agent
One simple request
Works
Works
Three variations
Needs a hand-built loop
Calls the tool three times
Vague request
Passes the vague text along
Asks a question or adds detail
Different aspect ratio
Needs a new branch
Passes another argument
Follow-up edits
Forgets the context
Reads its memory
Debugging
Linear and easy
Needs the execution log
💡 The trade-off is real. Agents spend more tokens and behave less predictably than a chain. If every request looks identical, a chain is cheaper and simpler. Pick the agent when requests vary.
What You Need First
Three things: an n8n instance recent enough to include the AI Agent node and its tool sub-nodes, a credential for a chat model, and a credential for an image API. Add patience for one test run and you are ready.
Pick the Brain Model
The agent's model must call tools reliably. A fast model is usually enough, because the agent mostly writes one prompt and picks two arguments. These large language models on PicassoIA are good candidates:
Set the temperature between 0.2 and 0.4. Low randomness means the tool arguments come out in the same shape every time.
Pick the Image Model
Whatever image model sits behind the tool, the agent only sees a name, a description and two inputs. Swapping models later means editing one URL in the sub-workflow. These text to image models make a useful shortlist:
The PicassoIA API exposes picassoia/picassoia-image for text to image and picassoia/picassoia-image-editor-pro for edits. The models in the table run from the web app, which makes them a handy test bench: try a prompt there before you trust the agent to send it.
Building the Workflow in n8n
Step 1: Chat Trigger and Agent
Create a new workflow. Add the Chat Trigger node (listed as When chat message received), then an AI Agent node, and connect the trigger to the agent. Current n8n versions run the agent as a tools agent by default. If your version asks for an agent type, choose Tools Agent.
Under the agent's Chat Model connector, click the plus sign and add the model node that matches your credential. Paste the credential, set the temperature to 0.3 and move on.
Step 2: Write the System Message
Open the AI Agent node, add the System Message option and paste this:
You are an image generation assistant.
When the user asks for an image, call the generate_image tool.
Before calling it, rewrite the request as one detailed photographic prompt of 60 to 120 words: subject, setting, lighting direction, lens and surface textures.
Avoid neon, cartoon and CGI wording unless the user asks for it.
If the request is too vague to picture, ask one short question first.
After the tool returns, reply with the image URL on its own line, then one sentence describing the result.
Never invent a URL. If the tool returns an error, say so and offer one retry.
Every line has a job. The rewrite rule is where image quality comes from. The "never invent a URL" rule exists because an agent will happily produce a plausible link when a tool fails. The retry rule stops endless loops.
💡 In the agent's options, set Max Iterations to 6 or 8. A stuck loop then ends with an error instead of burning tokens all night.
Step 3: Build the Image Tool
Add a Call n8n Workflow Tool sub-node to the agent's Tool connector. Name it generate_image and write the description like a job ticket:
Creates one photographic image from a detailed prompt and returns its public URL.
Inputs: prompt (string, required), aspect_ratio (16:9, 1:1 or 9:16).
Point the tool at a second workflow and build that workflow in this order:
Execute Workflow Trigger with two input fields: prompt and aspect_ratio.
HTTP Request named Create prediction, which starts the render.
Wait node set to 5 seconds.
HTTP Request named Get prediction, which checks the status.
IF node that tests whether the status equals succeeded.
On true: an Edit Fields node that outputs one field called imageUrl.
On false: a second IF that sends failed to a Stop and Error node and every other status back to the Wait node.
Why a sub-workflow instead of a lone HTTP Request Tool? Image APIs are asynchronous. A lone tool would start the render and return an id, forcing the agent to poll on its own, which costs extra turns and tokens. The sub-workflow hides the loop, so the agent sees one tool call and one URL.
Step 4: Add Memory
Attach the Simple Memory sub-node to the agent and set the context window to around eight messages. It uses the chat session id by default, so each visitor gets a separate history. With memory on, "now make the second one wider" works, because the agent can read its earlier prompt and change one detail.
The same agent also answers from a phone. Swap the Chat Trigger for a Telegram or Webhook trigger and the rest of the workflow stays untouched.
Connecting the Image API
The HTTP nodes inside the sub-workflow talk to the PicassoIA API: base address https://api.picassoia.com/v1, a bearer token that starts with pia_sk_, and Replicate-style prediction endpoints. Jobs are asynchronous: you create a prediction, poll it, then read the result. Create your token on the PicassoIA API page and store it in n8n as a Header Auth credential, with the name Authorization and the value Bearer followed by your token.
Create the Prediction
In the Create prediction node, set the method to POST, authentication to Generic Credential Type with Header Auth, and the body type to JSON.
The response includes a prediction id. Input field names differ between models, so confirm them on the API page for the model you call. Prompts can run up to 4,000 characters, which leaves plenty of room for a 120-word photographic description.
Poll Until It Finishes
In Get prediction, send a GET request to /v1/predictions/{{ $('Create prediction').item.json.id }}. The status moves on to succeeded or failed, and the IF node reads {{ $json.status }}. The Edit Fields node then pulls the link out of the response:
Some models return a list and others a single string, hence the check. Add a counter and send the run to Stop and Error after 40 polls, which is about three minutes at five seconds each. Without that ceiling, one stuck job can hold an execution open for hours.
💡 An account can run 5 predictions at once. For batches, put a Loop Over Items node in front of the agent with a batch size of 2 or 3, so the fifth request never lands on a full queue.
Prompts the Agent Writes Better
Let the Agent Expand Short Requests
People type five words. Image models want fifty. The system message bridges that gap, so the agent turns a short request into a full photographic brief.
User: a rainy street market at dusk
Agent prompt:Wide street market at dusk after rain, vendors under striped canvas awnings, wet cobblestones reflecting warm tungsten bulbs, shoppers with umbrellas in the middle distance, soft haze drifting from the left, 35mm lens at f/2, fine film grain, Kodak Portra 400 palette, visible fabric weave and water droplets on fruit crates.
Put these rules in the system message and the agent applies them every time:
Name the lighting direction, such as soft window light from the left.
Paste an expanded prompt the agent wrote in an earlier run.
Choose 16:9 for wide scenes, 1:1 for product tiles or 9:16 for stories.
Run the generation and compare the result with the picture in your head.
Change one detail at a time, lighting first and lens second, then run it again.
When the output matches, copy the winning wording back into the system message.
💡 Change one variable per run. If you rewrite the whole prompt each time, you never find out which edit helped. For short readable text inside a scene, run the same prompt through GPT Image 2 and compare.
Where the Workflow Earns Its Keep
Product Shots on Demand
A shop owner types "ceramic mug on linen, morning light" and gets a draft listing photo back inside one conversation. Add a Google Sheets node after the agent to append the product name, the prompt and the URL to a sheet. The team then holds a log of every render, and nobody has to hunt through chat history for the good one.
Because memory is on, the owner can say "same mug, darker wood table" and the agent reuses the earlier wording with one change. That follow-up habit is what separates an agent from a form.
Blog and Social Batches
Replace the Chat Trigger with a Schedule Trigger. Read a column of topics from a sheet, push each row through Loop Over Items into the agent, and write the returned URL back beside the topic. A week of hero images then arrives while you do something else.
Add a human gate before anything goes public. Send the URLs to Slack or email and wait for a thumbs up. Agents are quick, but a person still catches the odd extra finger or wrong season.
Fixing Common Failures
Symptom
Likely cause
Fix
Agent answers without calling the tool
Vague tool description
Start the description with a verb and name every input
Tool runs but the reply has no URL
Output field name mismatch
Check that the Edit Fields node outputs imageUrl and test the sub-workflow alone
Polling never ends
No failure branch or poll ceiling
Route failed to Stop and Error and cap the loop at 40 polls
401 response
Wrong header or expired token
Recreate the Header Auth credential and keep the Bearer prefix
429 or queue errors
More than 5 jobs at once
Use Loop Over Items with a batch size of 2 or 3
Images look cartoonish
Prompt lacks photographic detail
Add lens, lighting direction and texture rules to the system message
Agent calls the tool again and again
Max Iterations too high
Set it to 6 and keep the single-retry rule
💡 When something breaks, run the sub-workflow alone with pinned input data. If it works on its own, the fault sits in the agent's instructions, not in the HTTP nodes.
Try It on PicassoIA Today
Build the smallest version first: one Chat Trigger, one AI Agent, one image tool. Ten minutes of wiring gives you a chat box that returns photographs, and every upgrade after that is a single node.
Open Flux 2 Pro, GPT Image 2 or Nano Banana Pro in your browser and write your first prompt by hand. Pick a brain from the language models above, grab an API token from the PicassoIA API page, and let the agent write the next hundred prompts for you. Make your own images on PicassoIA, then send the best prompt back into your workflow.