Large Language ModelsGenerate imagesGenerate videos
Create an AI Image Generator App: Stack, API and Costs
A practical plan for building an AI image generator app: the request loop, a stack that ships in a week, how to connect an image API with working Node code, and a monthly budget that shows what each generation really costs before you launch.
An AI image generator app looks like a big product from the outside and a small one from the inside. A person types a sentence, your server hands it to a model, a few seconds later a picture comes back, and you save it somewhere the person can find it again. Everything else is polish: accounts, a gallery, credits, a way to stop abuse. The part that surprises most builders is not the code. It is the bill. Cost per image, multiplied by how many times people press the button, decides whether the app earns money or burns it.
This article lays out a stack you can build in a week, shows how the API calls fit together with working code against the PicassoIA API, and puts honest arithmetic on the monthly cost so you can price the product before launch day.
What You Are Really Building
Strip away the branding and every image generator app does the same job: it turns a text request into a file, and it does that without making the person wait on a frozen screen. Once you see the app as a request loop with a gallery attached, the build gets much smaller.
The Request Loop in Five Steps
The browser sends the prompt and a few options (size, style, number of images) to your back end.
Your back end checks the user, the quota and the prompt, then creates a prediction at the image API.
The API answers right away with an id, because generation is asynchronous and takes seconds, not milliseconds.
Your back end polls that id (or waits for a webhook). When the status reads succeeded, it copies the file into your own storage.
The browser shows the picture and adds it to the user's gallery.
Every feature you will ever add sits on top of those five steps. Credits hook into step 2. Galleries hook into step 5. Video generation is the same loop with a slower step 4.
Decide the Product Shape First
Pick the smallest version that someone would pay for. Each extra input type adds work on the back end.
Longer jobs, bigger files, a video model such as PicassoIA Video
💡 Start with one input box and one button. Features you add later cost less than features you have to rip out after users depend on them.
Pick a Stack That Ships
Boring tools win here. The model does the hard part, so your stack only needs to be reliable, cheap to run and easy to change when a better model appears next month.
Front End Choices
A React framework such as Next.js gets you a prompt page, a gallery and server routes in one project. If most of your users will be on phones, ship it as a progressive web app first and move to React Native or Flutter only when you need deeper camera or file access. The screen itself is simple: a text box, a size picker, a button, a grid.
Back End and Queue
Use Node or Python, whichever your team already writes. The one piece you should not skip is a job queue. Image generation is slow compared with a normal web request, and image APIs limit how many jobs run at once. A queue (Redis with BullMQ, or Celery on Python) lets you accept every click instantly and feed the API at a safe pace.
Keep a single generations table in Postgres with these columns: id, user_id, prompt, model, status, cost_usd, image_url, created_at. That one table powers the gallery, the quota check and your cost reports.
Storage and Delivery
Copy every finished image into your own bucket behind a CDN. Do not depend on a provider's temporary URL staying alive forever. Save the original plus a smaller WebP thumbnail so the gallery loads fast on mobile.
Layer
Simple option
Move up when
Front end
Next.js or plain React
You need native device features
API layer
Node (Fastify) or FastAPI
Traffic needs separate workers
Queue
Redis with BullMQ or Celery
You call more than one provider
Database
Postgres
Reporting gets heavy
Storage
S3-compatible bucket plus CDN
Users sit in many regions
Sign-in
Email links or OAuth
You sell team accounts
Choose an Image API
You can rent a model by the call or run your own GPUs. For a first release the answer is almost always the first one.
Hosted API or Your Own GPU
A hosted API means no drivers, no scaling work, many models behind one interface, and you pay per image. A rented GPU means a fixed hourly bill and a long list of chores: model weights, memory limits, updates, queues, crashes at 3 a.m.
The math decides it. Say a rented GPU costs $1.50 an hour and makes 120 images an hour. If it never sits idle, each picture costs about $0.0125. If it is busy only 20% of the time, you are really making 24 images an hour and each one costs $0.0625. Those are example numbers, but the shape of the result holds: self-hosting only wins with steady, heavy traffic.
Models Worth Wiring In
The text-to-image collection on PicassoIA lists more than 200 models, which is useful because no single model is best at everything. Pick one default and keep one or two alternates behind a setting, so you can switch when quality, speed or price changes.
Most people write prompts like "a dog on a beach". A small language model can expand that into a richer prompt before the image call, and it costs a fraction of a cent. Claude Sonnet 5 and Gemini 3.5 Flash both fit this job, and GPT 5 Structured returns clean JSON when you need the rewrite split into fields such as subject, lighting and lens.
The same family of models can do a second job: checking prompts before they reach the image model. More on that below.
Connect to the PicassoIA API
PicassoIA offers a Replicate-style REST API, so the flow is the one from the request loop: create a prediction, poll it, fetch the result. The endpoints and limits below come from the public API page, and that page is the place to confirm details before you ship.
How to Use PicassoIA Image
Before you write any code, test the model by hand. It takes ten minutes and saves days of guessing.
Type five prompts that match what your users will really write: short ones, long ones, vague ones.
Try the aspect ratios your app will offer, such as 1:1, 16:9 and 9:16.
Generate each prompt several times and note how much the results vary.
Write down which phrasing gave the best photos. That list becomes the template for your prompt rewriter.
💡 If your app edits photos, repeat the same test with PicassoIA Image Editor Pro using real uploads, not stock samples.
Base URL and Authentication
The base URL is https://api.picassoia.com/v1. Every request carries an Authorization: Bearer header holding a secret that starts with pia_sk_. You create it on the account's API page, and an account can hold at most two, so rotate one at a time. Never put that secret in browser or mobile code. It belongs on your server, in an environment variable. Plan requirements and pricing for API access are listed on that page and can change, so read them before you build a budget around it.
5 per account, shared across all secrets and MCP connections
Request body
10 MB
Prompt length
4,000 characters
Job timeout
3 hours
The five-job ceiling shapes your whole architecture, which is why the queue from earlier is not optional.
Create, Poll, Fetch
The endpoints are POST /v1/models/{owner}/{name}/predictions to start a job, GET /v1/predictions/{id} to read it, POST /v1/predictions/{id}/cancel to stop it and GET /v1/predictions to list recent jobs. Here is the whole loop in Node 18 or newer:
const BASE = "https://api.picassoia.com/v1";
const headers = {
Authorization: `Bearer ${process.env.PICASSOIA_TOKEN}`,
"Content-Type": "application/json",
};
async function call(url, options) {
const res = await fetch(url, { headers, ...options });
if (!res.ok) throw new Error(`HTTP ${res.status}`);
return res.json();
}
export async function generate(prompt) {
const created = await call(
`${BASE}/models/picassoia/picassoia-image/predictions`,
{
method: "POST",
body: JSON.stringify({ input: { prompt, aspect_ratio: "16:9" } }),
}
);
let job = created;
while (!["succeeded", "failed", "canceled"].includes(job.status)) {
await new Promise((r) => setTimeout(r, 2000));
job = await call(`${BASE}/predictions/${created.id}`);
}
if (job.status !== "succeeded") throw new Error(job.error ?? job.status);
return Array.isArray(job.output) ? job.output[0] : job.output;
}
Two notes on that code. First, input field names differ from model to model, so read each model page and match them exactly. Second, in production you save created.id to your database before the polling starts, so a server restart never loses a paid job.
What It Really Costs
Costs split into two groups: variable costs that grow with every click, and fixed costs that stay put. Variable costs are the dangerous ones.
Cost Per Image Is the Number
The core formula is short:
Monthly cost = users × generations per user × cost per image + fixed costs
The word to watch is generations. People press the button several times for every picture they keep, so measure generations per kept image from the first day of your beta and budget with that figure, not with the number of saved pictures.
A Worked Monthly Budget
Take 2,000 active users who each make 20 generations a month. That is 40,000 images. The prices below are planning assumptions, so swap in the current rate of the model you choose. Fixed costs for hosting, database and storage are set at $100 a month for this size.
Scenario
Assumed price per image
Image bill
With $100 fixed
Result with $2,700 revenue
Fast draft model
$0.01
$400
$500
+$2,200
Mid-range model
$0.04
$1,600
$1,700
+$1,000
Premium model
$0.08
$3,200
$3,300
-$600
The revenue figure assumes 300 of the 2,000 users pay $9 a month. The premium row loses money even though the app looks healthy, because the free users generate pictures too. Three fixes work well together: cap the free tier, sell credits sized to each model's real price, and send draft requests to the cheap model while reserving the premium one for final renders.
Costs that people forget:
Failed and abandoned generations. You may pay for pictures nobody opens.
Retries. Every automatic retry is another billable call unless the first one clearly failed.
Storage and bandwidth. A gallery of full-size images adds up faster than expected, which is why thumbnails and a CDN matter.
Prompt rewriting and moderation calls. Small per call, real at volume.
Payment fees and app store fees. They come off the top of every sale.
Support time. Someone has to answer "my image looks wrong".
Keep It Safe and Fast
Speed and safety are cheap to add early and painful to add after launch.
Queues, Limits, and Retries
With five concurrent predictions per PicassoIA account, a queue decides how smoothly your app behaves. Assume an image takes about 10 seconds. Five jobs at once then give you roughly 30 images a minute, or 1,800 an hour. The 40,000 images from the budget above average about 55 an hour. Even a rush hour at five times the average, around 280 images, fits with room to spare.
Rules that keep the queue healthy:
Retry only transient errors, such as timeouts and server errors, with a growing delay between attempts. Never retry a request the API rejected for a bad input.
Limit each user to a small number of jobs at once so one person cannot fill all five slots.
Show progress, even a simple "Queued, 3rd in line", so people do not click again.
Use the cancel endpoint when a user leaves, so you stop paying for work nobody will see.
Moderation Before Generation
Check the prompt before it reaches the image model. A safety classifier such as Llama Guard 4 12B reads the text and flags categories you choose to block. It is cheap, fast and keeps your account out of trouble.
If users can upload photos, review the uploads too. Log every refusal with the user id and the reason, because patterns in those logs tell you who is probing your limits.
Two upgrades that pay for themselves:
Cache by recipe. Hash the prompt, model, size and seed. When the same recipe appears again, return the stored file instead of paying for a new one. Prompt templates and example galleries hit the cache constantly.
Add video later. The loop is identical, only slower. PicassoIA Video and Seedance 2.5 Lite are both on the API, and a finished still image can act as the first frame of a clip. Budget video separately, since clips cost more than pictures and their files are bigger.
Your First Week Plan
A small team can ship a private beta in seven days if the scope stays tight.
Day
Task
1
Pick one default model and test 20 realistic prompts by hand
2
Build the back end route that creates and polls a prediction
3
Add storage, thumbnails and the gallery page
4
Add sign-in, a daily quota and cost logging per generation
5
Add prompt moderation and per-user rate limits
6
Add credits or a simple payment link
7
Invite 20 testers and read the logs together
Mistakes that cost the most time: building a custom model pipeline before proving anyone wants the product, hard-coding one model name in twenty places, and forgetting to log the cost of each generation. Fix the last one on day four and every later decision gets easier, because you will see which prompts, users and models drive the bill.
Make Your First Image on PicassoIA
The fastest way to judge a model is to use it. Open PicassoIA Image, type the prompt a real user would type and look at what comes back. Then try the same prompt on two or three other models from the full model list and compare quality, speed and style side by side.
When the results look right, read the PicassoIA API page, create your first secret and run the Node snippet above. One prompt, one prediction, one image saved to your own storage: that is the whole app in miniature, and everything after it is scaling. Start experimenting today, and the first picture your own code creates will tell you more than any plan can.