Generate imagesLarge Language ModelsVisual Effects

Gemini API Image Understanding: Image Input and Image to Text in Python

Send a photo to the Gemini API from Python and get text back. This article shows inline bytes, the Files API, and multi-image prompts, then turns the same call into captions, receipt OCR, and bounding boxes, with token math, size limits, and fixes for common errors.

Gemini API Image Understanding: Image Input and Image to Text in Python
Cristian Da Conceicao
Founder of Picasso IA

You have a photo and you need words. A product shot that needs alt text, a receipt that needs a total, a shelf that needs a count. The Gemini API accepts the image as part of the prompt and sends text back, so the whole job fits in one Python call of about ten lines. This article follows the order in which problems actually show up: setup, three ways to send an image, prompts that return usable text, token costs, and the mistakes that burn requests.

One change matters for the code below. Google's current docs show image input through the Interactions API (client.interactions.create) and label the older generateContent method as legacy, while confirming that it stays fully supported. Both versions appear here, so you can paste whichever one matches your project.

What Image Input Actually Does

Woman holding a phone over a printed market photo beside a laptop

Gemini models are multimodal, which means a single request can hold text parts and image parts side by side. You send a photo plus an instruction, and the model answers in text. There is no separate vision endpoint, no preprocessing step, and no OCR library to install first. The image is just another part of the prompt.

Image to Text in One Request

The same call pattern handles very different jobs, depending on the instruction you attach:

  • Captioning: one sentence for a social post or a page description.
  • Alt text: short, literal descriptions for accessibility.
  • Visual questions: "How many chairs are at the table?" or "Is the label facing forward?"
  • OCR: text pulled from receipts, signs, forms, and handwritten notes.
  • Detection: labeled bounding boxes returned as JSON.
  • Comparison: differences between two or more images.

💡 Treat the instruction as the product. The model is the same in every case. The prompt decides whether you get a poem about a photo or a clean JSON total.

Models That Accept Images

Google's models page lists these current IDs, all with image input:

Model IDStatusGoogle's description
gemini-3.8-flashStableMost intelligent Flash model
gemini-3.7-flashStableComplex coding and agentic workflows
gemini-3.6-flashStableGeneral multimodal work
gemini-3.5-flash (Gemini 3.5 Flash)StableHigh-throughput workloads
gemini-3.1-pro-preview (Gemini 3.1 Pro)PreviewComplex problem-solving
gemini-3-flash-preview (Gemini 3 Flash)PreviewMultimodal tasks

Lineups move fast, so check the models page before you pin an ID in production. For image work, a Flash model is the sensible default. Move to a Pro model only when answers on dense scans or tricky scenes come back wrong.

Set Up Your Python Environment

Developer laptop on a walnut desk with a terminal open at dusk

Two minutes of setup now saves an afternoon of confusing import errors later.

Install the SDK

The official package is google-genai. The examples below also use Pydantic for structured output and Pillow for drawing boxes.

pip install -U google-genai pydantic pillow

Do not mix it up with the older google-generativeai package. The new one imports as from google import genai, and every snippet here assumes it.

Create the Client

Generate a credential in Google AI Studio, store it in the environment variable name that Google's setup page specifies, and keep it out of source control. The client reads it automatically:

from google import genai

client = genai.Client()

No arguments, no hardcoded secrets in the script. Every example that follows reuses this client.

Send an Image Three Ways

Pick the method by file size and reuse, not by habit.

Inline Bytes for Small Files

Camera memory card slot with prints of a harbor town behind it

Inline data is the shortest path. You read the file, encode it, and send it with the prompt. The current Interactions API version looks like this:

import base64
from pathlib import Path

image_bytes = Path("street-market.jpg").read_bytes()

interaction = client.interactions.create(
    model="gemini-3.8-flash",
    input=[
        {"type": "text", "text": "Caption this image in one sentence."},
        {
            "type": "image",
            "data": base64.b64encode(image_bytes).decode("utf-8"),
            "mime_type": "image/jpeg",
        },
    ],
)

print(interaction.output_text)

The legacy generateContent version is still valid and slightly shorter, because the SDK handles the encoding:

from google.genai import types

response = client.models.generate_content(
    model="gemini-3.8-flash",
    contents=[
        types.Part.from_bytes(data=image_bytes, mime_type="image/jpeg"),
        "Caption this image in one sentence.",
    ],
)

print(response.text)

Inline data limits the total request (prompt text, system instructions, and image bytes together) to 20 MB. A single phone photo fits easily. A batch of full-resolution scans does not.

Files API for Larger Images

Photographer beside a large mountain lake print holding a hard drive

When the request would pass 20 MB, or when you want to ask several questions about one image, upload it once and reference it by URI:

uploaded = client.files.upload(file="mountain-lake-print.jpg")

interaction = client.interactions.create(
    model="gemini-3.8-flash",
    input=[
        {"type": "text", "text": "Describe the scene and list any visible text."},
        {
            "type": "image",
            "uri": uploaded.uri,
            "mime_type": uploaded.mime_type,
        },
    ],
)

print(interaction.output_text)

Uploaded files are stored temporarily, so treat the Files API as a delivery mechanism rather than an archive. Keep your originals.

Multiple Images in One Prompt

Two nearly identical living room prints with one difference pointed out

Add more image parts to the same input list. Google's docs allow up to 3,600 image files in one request.

before = client.files.upload(file="living-room-before.jpg")
after = client.files.upload(file="living-room-after.jpg")

interaction = client.interactions.create(
    model="gemini-3.8-flash",
    input=[
        {
            "type": "text",
            "text": "The first image is BEFORE and the second is AFTER. "
                    "What is different between them?",
        },
        {"type": "image", "uri": before.uri, "mime_type": before.mime_type},
        {"type": "image", "uri": after.uri, "mime_type": after.mime_type},
    ],
)

print(interaction.output_text)

Say which image is which in the text. The model sees an ordered list, and a plain "compare these" leaves it guessing about roles.

Prompts That Turn Photos Into Text

The call never changes. Only the instruction does.

GoalPrompt patternOutput shape
Caption"Caption this image in one sentence."Plain text
Alt text"Write alt text under 125 characters. Describe only what is visible."Plain text
Visual question"How many red crates are on the left shelf?"Short answer
Extraction"Extract merchant, date, and total."JSON via schema
Detection"Detect all of the prominent items in the image."JSON via schema

Captions and Alt Text

Web editor writing alt text in a notebook next to a monitor

The biggest quality gain comes from constraints. "Describe this image" returns a paragraph. "Write alt text under 125 characters, no opening phrase like 'image of'" returns something you can publish.

prompt = (
    "Write alt text for this photo in under 125 characters. "
    "Describe only what is visible. Do not start with 'image of'."
)

Run that over a folder of photos with a simple loop and you have a first draft of every missing alt attribute on a site. A human still reads the drafts, because a model can misjudge what matters in a scene.

OCR and Receipts

Crumpled receipts and a handwritten list on a cafe counter

Receipts are a good test because they mix printed text, numbers, and creases. Ask for structured output instead of prose. Define the shape with Pydantic and pass its JSON schema through response_format:

from pydantic import BaseModel

class LineItem(BaseModel):
    name: str
    price: float

class Receipt(BaseModel):
    merchant: str
    date: str
    items: list[LineItem]
    total: float

receipt = client.files.upload(file="receipt.jpg")

interaction = client.interactions.create(
    model="gemini-3.8-flash",
    input=[
        {
            "type": "text",
            "text": "Extract the merchant, date, line items, and total.",
        },
        {"type": "image", "uri": receipt.uri, "mime_type": receipt.mime_type},
    ],
    response_format={
        "type": "text",
        "mime_type": "application/json",
        "schema": Receipt.model_json_schema(),
    },
)

data = Receipt.model_validate_json(interaction.output_text)
print(data.merchant, data.total)

If the model returns something that does not fit the schema, model_validate_json raises an error right there instead of letting bad data slide into your database.

Object Detection With Boxes

Grocery aisle with crates of oranges, tomatoes, and peppers

Gemini can return bounding boxes as [ymin, xmin, ymax, xmax], normalized to a 0 to 1000 scale. Ask for them with a schema, then convert to pixels and draw:

from PIL import Image, ImageDraw
from pydantic import BaseModel, Field

class Box(BaseModel):
    box_2d: list[int] = Field(
        description="[ymin, xmin, ymax, xmax] normalized to 0-1000."
    )
    label: str

class Boxes(BaseModel):
    boxes: list[Box]

aisle = client.files.upload(file="grocery-aisle.jpg")

interaction = client.interactions.create(
    model="gemini-3.8-flash",
    input=[
        {"type": "text", "text": "Detect all of the prominent items in the image."},
        {"type": "image", "uri": aisle.uri, "mime_type": aisle.mime_type},
    ],
    response_format={
        "type": "text",
        "mime_type": "application/json",
        "schema": Boxes.model_json_schema(),
    },
)

result = Boxes.model_validate_json(interaction.output_text)

image = Image.open("grocery-aisle.jpg")
width, height = image.size
draw = ImageDraw.Draw(image)

for item in result.boxes:
    ymin, xmin, ymax, xmax = item.box_2d
    left, top = xmin / 1000 * width, ymin / 1000 * height
    right, bottom = xmax / 1000 * width, ymax / 1000 * height
    draw.rectangle((left, top, right, bottom), outline="red", width=4)
    draw.text((left + 6, top + 6), item.label, fill="red")

image.save("grocery-aisle-boxes.jpg")

Segmentation follows the same pattern. The schema adds a mask field holding polygon points, also normalized to 0 to 1000. Google recommends setting the thinking level to minimal for segmentation, since extended reasoning adds latency without improving the polygons.

Limits, Tokens, and Costs

Contact sheet of small frames on a light table with a loupe

Images are billed as tokens, and the count depends on size. Knowing the rule lets you predict a bill before the batch runs.

How Images Count as Tokens

Image situationToken cost
Both dimensions 384 px or smaller258 tokens
Larger imageTiled into 768 x 768 px tiles, 258 tokens per tile
Example: an image that splits into four tiles4 x 258 = 1,032 tokens

The docs also describe a media_resolution setting that caps the maximum number of tokens allocated to each input image. Lower it for bulk captioning where fine detail is irrelevant. Raise it when small print or distant objects matter. Check the current reference for how your SDK version spells the option.

Pricing differs by model, so read Google's pricing page for numbers. The lever you control is the token count, and sending a smaller copy of the file is the cheapest way to pull it down.

Format and Size Limits

LimitValue
Supported formatsPNG, JPEG, WEBP, HEIC, HEIF
Images per requestUp to 3,600 files
Inline request size20 MB total (text, instructions, and bytes)
Bounding box scale0 to 1000, order [ymin, xmin, ymax, xmax]

💡 If a job sends many images with a long prompt, count the inline total before you hit the 20 MB wall. Switching to the Files API mid-project is easy, but doing it early avoids a flaky failure at 2 a.m.

Three Mistakes That Waste Calls

Most failed requests come from the same short list.

Wrong MIME Type

The mime_type must match the real file. Labeling a PNG as image/jpeg or passing an unsupported format produces errors or poor results. Let Python decide instead of typing strings by hand:

import mimetypes

def mime_for(path: str) -> str:
    mime, _ = mimetypes.guess_type(path)
    if mime is None:
        raise ValueError(f"Unknown image type: {path}")
    return mime

Some systems do not know the HEIC type, so add a small manual mapping if you accept iPhone originals.

Boxes in the Wrong Place

If drawn rectangles land in odd spots, check two things. First, the order is [ymin, xmin, ymax, xmax], with the vertical value first. Many people read it as x then y. Second, the numbers sit on a 0 to 1000 scale, not in pixels. Divide by 1000, then multiply by the real width or height.

Free Text Where JSON Belongs

Writing "return JSON" in the prompt works until the model wraps the answer in a code fence or adds a friendly sentence. Pass a schema through response_format and parse with model_validate_json. The contract then lives in code, where a bad response fails loudly and a good one arrives typed.

Try Gemini 3.5 Flash Without Code

Before writing Python, test the prompt in a browser. Gemini 3.5 Flash runs on Picasso IA and takes images directly, so you can tune an instruction in seconds and paste it into your script afterward.

  1. Open the Gemini 3.5 Flash page on Picasso IA.
  2. Attach your photos in the Images field. The model accepts up to 10 images per run, each up to 7 MB.
  3. Type the instruction in the Prompt field, worded exactly as you plan to send it from Python.
  4. Optionally fill System Instruction to fix the role, such as "You write alt text under 125 characters."
  5. Choose a Thinking Level of none, low, or high. Leave it on none for captions and raise it for dense reasoning.
  6. Set Temperature low for extraction and OCR, higher for creative captions.
  7. Run it, compare the answer to what you wanted, and adjust the wording before copying it into code.
FieldWhat it doesStarting value
PromptThe instruction sent with the imagesYour exact production wording
ImagesUp to 10 files, 7 MB eachOne image while testing
System InstructionSets the model's roleOne short sentence
Thinking Levelnone, low, or highnone
TemperatureRandomness from 0 to 20.2 for OCR, 1 for captions
Max Output TokensCaps answer lengthDefault is fine

The limits on Picasso IA differ from the raw API limits above, so treat the page as a prompt lab and the API as the production path. For a second opinion on a hard image, run the same prompt through Gemini 3.1 Pro, Qwen3.7-Plus, which interprets images as well as text, or Granite Vision 4.1 4B, which is built for charts and tables.

Create Your Own Test Images

You do not need a folder of real photos to start. Generate a messy desk, a grocery shelf, a rainy street, or a crumpled receipt with PicassoIA Image or Seedream 4.5, then feed each result to Gemini 3.5 Flash and see what it reads back.

Try three experiments this week. Ask for alt text on five generated photos. Ask for a bounding box around one object in a busy scene. Ask for a JSON total from a receipt image. Each takes minutes, and together they show where the model is sharp and where your prompt needs tightening.

Open Picasso IA, create your first test image, and run your own prompt. Browse every available model at picassoia.com/en/all-models and pair an image generator with a vision model to build your own image to text workflow.

Share this article