Large Language ModelsGenerate imagesGenerate videos

Build an AI Image Generator in Python (Step by Step)

Write a Python script that sends a text prompt to an image model, waits for the job, and saves the picture. Build it in four steps, then add a command line, threaded batches, a Flask route, and fixes for the errors you will hit.

Build an AI Image Generator in Python (Step by Step)
Cristian Da Conceicao
Founder of Picasso IA

Most tutorials about AI images stop at "paste a prompt into a website." That works for one picture. It breaks the moment you need fifty product shots, a nightly batch of blog thumbnails, or a small app where visitors type a sentence and get a photograph back. For that you need code, and the good news is that a working image generator in Python fits in roughly 60 lines.

This tutorial builds it in four small steps: send a prompt to a hosted model, wait for the job to finish, download the file, then wrap everything in a command line tool that runs batches. Each step ends with a quick check, so you know it works before you move on.

💡 Quick answer: an AI image generator in Python is a script that sends a text prompt to a model over HTTP, polls until the image is ready, then saves the file. The model runs on someone else's GPU. Your code needs requests, a secret token, and a loop.

Before You Write a Line

Developer's hands typing Python code on a laptop at a wooden desk

You need three things: Python 3.10 or newer, a terminal, and a secret token for the image service. The code below talks to the PicassoIA API over plain HTTP, so there is no GPU driver to install and no model file to download.

Set up the project

Create a folder, a virtual environment, and install one package:

mkdir ai-image-generator && cd ai-image-generator
python -m venv .venv
source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install requests

Next, create a token on the API page and store it in an environment variable instead of your source file. An account can hold up to two tokens at a time.

export PICASSOIA_API_TOKEN="pia_sk_your_token_here"

On Windows PowerShell the same line is $env:PICASSOIA_API_TOKEN = "pia_sk_your_token_here".

⚠️ Never paste the token into your script. Anything committed to git stays in the history, even after you delete it. An environment variable keeps the secret out of the repository.

Pick a model first

The model decides how your images look and how fast they arrive. The script in this tutorial calls picassoia/picassoia-image through the API. The models in the table below are all available in the browser on PicassoIA, which makes them handy for testing prompts before you wire anything into code.

ModelBest use
P ImageFast photorealistic drafts, around a second per image
Flux 2 ProDetailed images from text or reference photos
Flux DevPhotorealistic scenes and portraits
Stable Diffusion 3.5 LargeSharp HD images with strong detail
Imagen 4 FastQuick results in a few seconds
Seedream 4.5High resolution output up to 4K
SDXL Lightning 4StepVery fast previews

How Text Becomes a Picture

A diffusion model starts from random noise and removes it in many small passes, nudged by your prompt at every pass. Each pass is a heavy calculation on a large neural network. That is why image generation wants a powerful GPU, and why the first design choice in your Python project is where that GPU lives.

Rows of black server racks in a quiet data center aisle

Local GPU or hosted API. Running a model on your own machine gives you full control, but it also means a driver stack, tens of gigabytes of weights, and a card with plenty of memory. A hosted API trades some of that control for a setup measured in minutes.

Close-up of a desktop graphics card with copper heat pipes inside an open PC case

FactorLocal GPUHosted API
Setup timeHoursMinutes
HardwareRecent graphics card with 12 GB or more of memoryAny laptop
Model choiceWhatever you downloadWhatever the service lists
ScalingBuy another cardRaise your concurrency
Python dependenciesHeavyrequests only

Why jobs are asynchronous. Generating an image takes anywhere from under a second to about a minute, depending on the model. Holding one HTTP connection open that long is fragile, so the API works in three moves: create a prediction, poll its status, then fetch the output once the status reads succeeded. Your code mirrors those three moves in the next section.

Step by Step: The First Script

Put everything in one file called generate.py. The four steps below build it from top to bottom.

Step 1: Set up the request

The API is Replicate-style: a base URL, a model name in the path, and a Bearer token in the header.

import os
import time
from pathlib import Path

import requests

API_BASE = "https://api.picassoia.com/v1"
MODEL = "picassoia/picassoia-image"
TOKEN = os.environ["PICASSOIA_API_TOKEN"]
HEADERS = {
    "Authorization": f"Bearer {TOKEN}",
    "Content-Type": "application/json",
}

If the variable is missing, Python stops with an error right here, which is exactly what you want. A loud failure at the top of the file beats a confusing 401 later.

Step 2: Send the prompt

def create_prediction(prompt: str) -> str:
    url = f"{API_BASE}/models/{MODEL}/predictions"
    payload = {"input": {"prompt": prompt}}
    response = requests.post(url, headers=HEADERS, json=payload, timeout=30)
    response.raise_for_status()
    return response.json()["id"]

The function returns a job ID, not an image. Prompts can be up to 4,000 characters, which is far more than you will need.

Step 3: Poll until it finishes

def wait_for_result(job_id: str, poll_every: float = 2.0, limit: float = 300.0) -> str:
    url = f"{API_BASE}/predictions/{job_id}"
    deadline = time.time() + limit
    while time.time() < deadline:
        response = requests.get(url, headers=HEADERS, timeout=30)
        response.raise_for_status()
        data = response.json()
        status = data["status"]
        if status == "succeeded":
            output = data["output"]
            return output[0] if isinstance(output, list) else output
        if status in ("failed", "canceled"):
            raise RuntimeError(f"Job {job_id} ended as {status}: {data.get('error')}")
        time.sleep(poll_every)
    raise TimeoutError(f"Job {job_id} did not finish within {limit} seconds")

💡 Tip: the field names follow the Replicate-style format. If a response ever looks different, print response.json() once and adjust the two lines that read status and output. The loop also sets its own deadline, so a stuck job never freezes your script.

Step 4: Save the file

def download(image_url: str, folder: str = "output") -> Path:
    Path(folder).mkdir(exist_ok=True)
    name = image_url.split("/")[-1].split("?")[0] or f"image-{int(time.time())}.png"
    target = Path(folder) / name
    response = requests.get(image_url, timeout=60)
    response.raise_for_status()
    target.write_bytes(response.content)
    return target


if __name__ == "__main__":
    prompt = "A misty mountain lake at sunrise, 35mm photograph, soft natural light, fine film grain"
    job_id = create_prediction(prompt)
    image_url = wait_for_result(job_id)
    print("Saved", download(image_url))

Over-the-shoulder view of a terminal window beside a golden hour beach photograph

Run python generate.py. After a few seconds you should see Saved output/... and a new file in the output folder. Checkpoint: if you can open that image, steps 1 to 4 all work, and everything after this point is packaging.

Turn the Script Into a Tool

Add a command line

Replace the bottom of the file with an argparse entry point so you can pass prompts from the terminal:

import argparse


def main() -> None:
    parser = argparse.ArgumentParser(description="Generate an image from a text prompt")
    parser.add_argument("prompt", help="what the image should show")
    parser.add_argument("--out", default="output", help="folder for saved files")
    args = parser.parse_args()

    image_url = wait_for_result(create_prediction(args.prompt))
    print(download(image_url, args.out))


if __name__ == "__main__":
    main()

Now python generate.py "a red bicycle leaning on a brick wall, 50mm photo" --out bikes does the whole job in one line.

Run batches within the limits

A loop that waits for each image in turn is slow. Threads fix that, because the script spends nearly all its time waiting on the network. The API allows 5 concurrent predictions per account, shared across your tokens and any connected apps, so five workers is the ceiling.

from multiprocessing.pool import ThreadPool


def generate(prompt: str) -> Path:
    return download(wait_for_result(create_prediction(prompt)))


def run_batch(prompts: list[str]) -> list[Path]:
    with ThreadPool(5) as pool:
        return pool.map(generate, prompts)

Photographer inspecting a printed contact sheet of small landscape and portrait frames

If you also use the website while a batch runs, drop to three workers so the two do not fight for the same five slots. One failed prompt raises an error inside pool.map, so wrap generate in a try block that logs the prompt and returns None when you run hundreds at once.

Keep a record. Append each prompt and filename to a log.jsonl file, so you can trace any picture back to the wording that made it. Call this helper inside generate right after download:

import json


def log_result(prompt: str, path: Path) -> None:
    with open("log.jsonl", "a", encoding="utf-8") as log:
        log.write(json.dumps({"prompt": prompt, "file": str(path)}) + "\n")

Six weeks from now, when someone asks which prompt produced the sunrise shot, that file answers in seconds.

Expose it as a web route

Teams rarely want a terminal. A small Flask app lets anyone submit a prompt from a page. Notice that the route returns the job ID immediately and a second route reports progress, so no web request hangs for a minute.

from flask import Flask, jsonify, request

app = Flask(__name__)


@app.post("/generate")
def start():
    prompt = request.get_json()["prompt"]
    return jsonify(id=create_prediction(prompt))


@app.get("/status/<job_id>")
def status(job_id):
    url = f"{API_BASE}/predictions/{job_id}"
    data = requests.get(url, headers=HEADERS, timeout=30).json()
    return jsonify(status=data["status"], output=data.get("output"))

Three colleagues gathered around one monitor showing a grid of photographic thumbnails

Your front end calls /generate, then asks /status/<id> every two seconds until status is succeeded. The token stays on the server, never in browser code.

Write Prompts That Behave

Overhead view of a desk with a laptop, notebook sketches and printed photographs

Build prompts from parts

Vague prompts give vague pictures. A prompt with a subject, a place, a light source, and a lens gives you something repeatable. Put that structure in a function so every image in a batch follows it:

def build_prompt(subject: str, setting: str, light: str, lens: str = "50mm f/1.8") -> str:
    return (
        f"{subject}, {setting}, {light}, shot on a {lens} lens, "
        "natural skin and surface texture, fine film grain, photorealistic"
    )


print(build_prompt("a ceramic mug of black coffee", "on an oak desk", "soft window light from the left"))

Three habits make the results more predictable:

  • One scene per prompt. Two subjects fighting for attention produce muddy pictures.
  • Name the light. "Volumetric morning light from the left" beats "nice lighting."
  • Name the lens. Words like 85mm and f/1.8 push the model toward shallow depth of field.

Brainstorm with a language model

Writing 50 prompts by hand gets dull around prompt twelve. A language model can draft them for you. Claude Sonnet 5 handles structured lists well, and Gemini 3.5 Flash returns quick drafts when you just need volume. Ask for 20 prompts that follow your subject, setting, light, lens structure, one per line, and save the reply as prompts.txt. Then feed the file to your batch runner:

prompts = [line.strip() for line in Path("prompts.txt").read_text().splitlines() if line.strip()]
run_batch(prompts)

How to Use P Image on PicassoIA

Before you spend time on code, test your wording in the browser. P Image answers in about a second, so you can try ten variations in the time one slower model takes to render a single result.

  1. Open the model page and find the prompt box.
  2. Paste a prompt built with the subject, setting, light, lens structure from the previous section.
  3. Pick the 16:9 ratio for article images, or 1:1 for thumbnails.
  4. Generate three variations and change only one detail between them. That tells you which word did the work.
  5. Copy the winning wording into prompts.txt or into your build_prompt call.

💡 If a seed field is available, fix the seed while you compare wording. With the seed locked, any change in the result comes from your prompt and nothing else.

When P Image gives you the right framing but not the right finish, run the same prompt through Flux 2 Pro or Seedream 4.5 and compare the results side by side. Each model has its own habits, and ten minutes of testing saves hours of rewriting later.

Fix the Errors You Will Hit

Authentication errors

A 401 or 403 almost always means the token never reached the request. Check that PICASSOIA_API_TOKEN is set in the same terminal that runs Python, since a variable exported in one window does not exist in another. Also check for stray quotes or a trailing space in the value.

Limits and stuck jobs

SymptomLikely causeFix
429 responseMore than 5 jobs running at onceLower ThreadPool to 3 or 4 and retry after a short sleep
400 responsePrompt over 4,000 characters, or a wrong field nameTrim the prompt and print the error body
Job never finishesHeavy queue or a stalled jobKeep the limit deadline and retry once
failed statusThe model rejected the promptSimplify it and remove unusual symbols

For retries, sleep a little longer after each failure (2, 4, then 8 seconds) and give up after three tries. Hammering the API with instant retries only keeps you over the limit.

Soft or blurry results

Woman examining a large printed landscape photograph at a drafting table

If the framing is right but the detail is soft, the cause is usually resolution, not the prompt. Pick a model that outputs high resolution, such as Seedream 4.5, or run the finished image through a Super Resolution model on PicassoIA to upscale it 2x to 4x. Also check that the prompt names a camera and lens, since "photograph" alone leaves too much open.

Make Your First Image Today

Man in a charcoal sweater holding a tablet with a sunlit forest path photograph in a cafe

You now have a script that turns a sentence into a photograph, a batch runner that respects the five job limit, and a prompt template you can reuse on every project. The fastest next move is a small one: open P Image in your browser, test three prompts, paste the best wording into prompts.txt, and run the batch.

When you want to scale, the PicassoIA API page lists the endpoints, and the full model list shows the text to image, video, and speech models you can try next. The same create, poll, download loop works for video jobs too, so the script you wrote today is the base for a clip generator tomorrow. Pick a prompt, run the script, and see what lands in your output folder.

Share this article