Generate imagesVisual EffectsLarge Language Models

GPT API Image Input: Vision Analysis and Detail Settings

The GPT API turns every image into tokens, and the detail field decides how many. This article shows the request format, the low, high, original and auto values, worked token math for tile and patch models, and a step by step test on PicassoIA.

GPT API Image Input: Vision Analysis and Detail Settings
Cristian Da Conceicao
Founder of Picasso IA

One field in a request body decides whether a photo costs 85 tokens or 3,000. That field is detail, it sits inside the image object next to the URL, and most tutorials either skip it or quote numbers that stopped being true two model generations ago. If you send screenshots, receipts, product shots or charts to a GPT model, this setting shapes your bill, your latency and how much of the picture the model can actually read.

This article walks through the request format, the four detail values, the token arithmetic for tile based and patch based models, the cases where low backfires, and the limits worth knowing before you ship. Numbers come from the image input pages of OpenAI's API documentation as they read today, and every worked example shows its arithmetic, so you can check it against the usage block in your own responses. Pages like that change often, which is one more reason to log token counts instead of trusting a table, including the ones below.

Low-angle view across a walnut desk with a developer's hands on a laptop and a printed coastline photograph beside it

How a Photo Becomes Tokens

A GPT model never reads your JPEG as a file. The API resizes it, slices it into small blocks, and turns every block into tokens that sit in the context window next to your text. Those tokens are billed at the model's normal input rate, so a larger or sharper image means a larger bill and more latency. They also compete with your prompt and the answer for the same context window, which matters once a request carries several pictures.

The Request Shape

Images travel inside the user message as content parts. The Responses API uses input_text and input_image parts, while Chat Completions uses text and image_url. In both, detail lives on the image part. Here is a receipt reader on GPT 5.4:

from openai import OpenAI

client = OpenAI()

response = client.responses.create(
    model="gpt-5.4",
    input=[
        {
            "role": "user",
            "content": [
                {"type": "input_text", "text": "List every line item and the total."},
                {
                    "type": "input_image",
                    "image_url": "https://example.com/receipt.jpg",
                    "detail": "original",
                },
            ],
        }
    ],
)

print(response.output_text)
print(response.usage.input_tokens)

The same call on Chat Completions with GPT-4o nests the URL one level deeper:

completion = client.chat.completions.create(
    model="gpt-4o",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Describe this photo in one sentence."},
                {
                    "type": "image_url",
                    "image_url": {"url": "https://example.com/photo.jpg", "detail": "low"},
                },
            ],
        }
    ],
)

print(completion.usage.prompt_tokens)

After every call, read usage.input_tokens on Responses or usage.prompt_tokens on Chat Completions. It is the only number that settles an argument about cost.

Three Ways to Send Pixels

  • Public URL. The simplest option, but OpenAI's servers must be able to fetch it quickly.
  • Base64 data URL. An inline data:image/jpeg;base64,... string works for private files and adds about a third to the payload size.
  • File ID. Upload once through the Files API with the vision purpose, then reference the ID with a file_id field on the image part and reuse it across requests.

Accepted formats are PNG, JPEG, WEBP and non-animated GIF.

The Four Detail Values

The detail field accepts low, high, original and auto. Leave it out and you get auto, which means the model's own default sizing. The names suggest a simple ladder, but what each rung does depends on the model you call.

ValueWhat the docs say it is forWatch for
lowCoarse reading of the pictureNot always cheaper than high on newer models
highStandard high fidelityCapped at 2,500 patches on the GPT 5.4 family
originalLarge, dense, spatially sensitive or computer-use imagesNeeds GPT 5.4 or later, and is not supported on GPT 5.2 or GPT 4.1 mini
autoThe model's default sizingThe value you get when the field is missing

Extreme close-up of a brass loupe magnifying one frame on a printed contact sheet

Low: Coarse and Cheap, Mostly

On the older tile based models, low is a flat fee: 85 tokens on GPT-4o and GPT 4.1, whatever the file size. The model receives a 512 by 512 pixel version, enough to say "a dog on a beach" and not enough to read a street sign. Use it for classification, rough captions, moderation passes and routing decisions where only the gist matters.

High and Auto: The Everyday Default

On tile models, high first fits the image inside a 2,048 by 2,048 square, then scales the shortest side to 768 pixels, then counts 512 pixel tiles. On the GPT 5.x family it works from 32 pixel patches instead, with a 2,048 pixel longest side and a 2,500 patch budget on GPT 5.4 and its smaller siblings. That is plenty for photos, product shots and ordinary screenshots. It starts to hurt with dense documents and tiny interface text, where a few lost pixels turn a 6 into an 8.

Original: When Pixels Matter

original raises the ceiling to 10,000 patches and a 6,000 pixel longest side on GPT 5.4 and its siblings. OpenAI points it at large, dense, spatially sensitive or computer-use images, and at coordinate-sensitive work such as OCR or small object detection. Picture a 4K screenshot where the model must return the position of a button, or a floor plan where one thin line carries meaning. You pay for that sharpness: up to 12,000 image tokens per picture at the 1.2 multiplier.

The practical rule is to earn it. Start a task at high, collect the failures, and move only the failing image types to original. If a field was misread at high and reads correctly at original, the extra tokens bought something. If both settings fail the same way, the problem is the prompt or the source image, and more pixels will not help.

Token Math, Worked Out

Tile Based Models

Tile billing is a base fee plus a fee per 512 pixel tile. GPT-4o and GPT 4.1 charge 85 base and 170 per tile. GPT 5.1 charges 70 and 140. GPT 4o Mini charges 2,833 and 5,667, so moving image work to the mini model inflates the token count instead of shrinking it. Compare total cost, not the per token rate.

ImageDetailModelArithmeticTokens
Any sizelowGPT-4oFlat base fee85
1024 x 1024highGPT-4o4 tiles x 170 + 85765
1920 x 1080highGPT-4oResized to 1365 x 768, 3 x 2 = 6 tiles, 6 x 170 + 851,105
1024 x 1024highGPT 4o Mini4 tiles x 5,667 + 2,83325,501

Straight-on view of a wall of handmade ceramic tiles in off-white, sand and pale blue

Patch Based Models

Newer models count 32 by 32 pixel patches: ceil(width / 32) x ceil(height / 32). When the total exceeds the budget, the image is shrunk until it fits, and the final token count is patches times a model multiplier, rounded up. The multiplier is 1.2 for GPT 5.2, GPT 5.4 and the GPT 5.6 family (Sol, Terra and Luna), and 1.62 for GPT 4.1 mini.

ImageDetailPatchesTokens at the 1.2 multiplier
1920 x 1080low, high or original60 x 34 = 2,0402,448
4000 x 3000highCapped at 2,500Up to 3,000
4000 x 3000lowAbout 3,072 after the 2,048 pixel capAbout 3,700
4000 x 3000originalCapped at 10,000Up to 12,000

The 4000 x 3000 rows are my arithmetic from the published budgets, so treat them as estimates and confirm with usage. The 1080p row is exact because the image already fits every budget.

Scale makes the difference real. Ten thousand product photos at 2,448 tokens each is 24.48 million input tokens before a single word of prompt. The same set at 85 tokens each on GPT-4o low is 850,000, roughly a 29 times difference you would never notice from the code alone.

Aerial view of rectangular farm fields forming a patchwork of green, gold and brown

When Low Costs More Than High

The docs carry a warning that surprises people: low does not always use fewer tokens than high. On GPT 5.4 and its smaller siblings, low allows a 6,144 patch budget while high stops at 2,500, so a large photo can bill more at low. On GPT 5.2 and GPT 4.1 mini, every level shares one sizing rule, a 2,048 pixel longest side and a 6,144 patch budget, so low, high and auto return identical counts and original is unavailable.

Two consequences follow. Never assume the cheap setting is cheap, and expect the meaning of your existing detail values to change when you switch models. A short test settles both:

  1. Pick three representative images: a small one, a 1080p screenshot and a 12 megapixel photo.
  2. Send each one at every detail value the model supports.
  3. Log the input tokens next to the quality of the answer.
  4. Keep the lowest setting that still answers correctly.

💡 If low and high return the same token count on a model, the field does nothing there. Drop it from your code instead of carrying a parameter that suggests savings you are not getting.

Close-up of an antique brass balance scale with silver coins on one pan and a folded receipt on the other

Limits That Bite in Production

Payload and Format Limits

The current docs allow up to 512 MB of total payload and up to 1,500 images per request. Older write-ups quote 50 MB and 500 images, so a library that enforces those numbers may be out of date. A 512 MB request is a latency problem long before it is a limit problem, because base64 inflates the bytes by a third and every image still bills tokens.

Where Vision Still Fails

OpenAI lists the weak spots plainly, and they match what shows up in production:

  • Specialized medical images, such as CT scans, are not a fit.
  • Non-Latin alphabets, such as Japanese or Korean, can underperform.
  • Graphs with varying line styles or colors, where solid, dashed and dotted lines must be told apart, cause mistakes.
  • Precise spatial localization, such as reading chess positions, is unreliable.
  • Object counts come back as approximations.
  • CAPTCHAs are blocked.
  • File names and metadata are never read.

Top-down view of a mid-game wooden chessboard with a hand hovering above a knight

Tiny print is the usual casualty in everyday work. When the part that matters is small, send only that part.

💡 Crop before you send. A 600 x 400 crop of a receipt's total line costs about 300 tokens on GPT 5.4 (19 x 13 = 247 patches, times 1.2). The full 4000 x 3000 page at original can reach 12,000, and the crop usually reads better.

Extreme macro of a crumpled thermal receipt with faded print on a dark slate counter

Mistakes That Waste Tokens

Most overspending comes from a handful of habits:

  • Uploading raw camera files. A 12 megapixel original is shrunk to the model's cap on arrival anyway. Resize to the longest side your detail level allows first (2,048 pixels for most settings, 6,000 on original for GPT 5.4 and its siblings), export a JPEG, and the request uploads faster with nothing lost.
  • Stitching a collage into one image. Six screenshots pasted onto one canvas get shrunk together, so each loses resolution. Send six separate image parts and name them in the text: "Image 1 is the invoice, image 2 is the packing slip." Each part is billed on its own.
  • Asking for everything at once. A prompt that wants a caption, a color list, a defect check and an OCR pass invites shallow answers. One narrow question per call, or one clearly numbered list, returns cleaner output.
  • Never logging usage. Without input token counts per request, a model switch or a new image size can double your bill and nothing will tell you.

Pick a Setting by Task

TaskStart withWhy
Moderation, routing, rough captionslow on tile models, auto on patch modelsThe gist is enough
Product photos, scene descriptionshigh or autoGood detail at moderate cost
Receipts and invoicesCropped high, or original on GPT 5.4 and laterSmall print needs pixels
Charts and dashboardsoriginal, or a specialist readerThin lines carry meaning
Screenshots for interface agentsoriginalCoordinates must be exact

Receipts and Documents

Text is where resizing hurts first. Crop to the region, straighten it, and ask for a fixed JSON shape so a wrong digit is easy to spot. Tell the model to answer unreadable for any field it cannot read, because a model that is allowed to guess will guess, and a confident wrong total is worse than a blank one. If a scan is blurry, repair it before upload with an AI image restoration tool, because no detail setting can invent pixels that were never captured.

Charts and Dense Screenshots

Charts mix thin strokes, small labels and similar colors, which is the exact case OpenAI flags. Send them at original when the model supports it. Ask for the underlying numbers as a table first and the interpretation second, so you can spot-check the values against the picture before you trust any trend the model describes. For tables and charts you extract every day, a specialist such as Granite Vision 4.1 4B is worth a side by side test.

Product Photos at Scale

For catalog work, run the three image test from earlier on your real photos, then fix the setting and the prompt for the whole batch. Ask for the same attributes in the same order every time, such as color, material and visible defects, and the output becomes easy to load into a database. Photograph the items the same way too: a consistent background, distance and light let you use a cheaper setting, because the model no longer has to work around noise. A product shot taken on a plain tabletop under even light often reads fine at high, while the same mug in a cluttered kitchen may not.

A packing station worker photographing a glazed ceramic mug on a white tabletop with a smartphone

How to Use GPT 5.4 on PicassoIA

You do not need code to see what a model reads. GPT 5.4 on PicassoIA accepts images next to a text prompt, and its form exposes the same levers you tune in the API: a system prompt, verbosity, reasoning effort and a completion token limit. That makes it a fast place to settle the questions that come before the settings. Which prompt wording works? Does the model read this kind of picture at all? Is a cheaper model good enough? One caveat: the form has an image input field but no detail switch, so treat results as a test of prompt and model choice, not a measurement of low against high.

Step by Step

  1. Open the GPT 5.4 page and find the Image Input field.
  2. Add your test image. A receipt, a dashboard screenshot or a product photo all work.
  3. Write a narrow prompt: "Return the date and the total as JSON" beats "describe this image".
  4. Add a System Prompt that sets the role and the output format.
  5. Set Verbosity to low for extraction jobs and high when you want a detailed breakdown.
  6. Leave Reasoning Effort at none for plain reading. Raise it for multi-step questions, and raise Max Completion Tokens with it, because high effort can spend the whole budget on reasoning and return an empty answer.
  7. Run the same image on GPT-4o and compare the two outputs.

A woman in a bright home office seen from behind as she drags a landscape photograph into a chat window on a laptop

Vision Models Worth Comparing

Image to text is one of the platform's built-in capabilities, and several language models accept images:

ModelGood for
GPT-4oText and images in one request, JSON output on request
GPT 5.4Screenshots and diagrams with adjustable reasoning effort
Gemini 3.5 FlashFast chat, code and image questions
Qwen3.7-PlusText plus image interpretation
Kimi K2.5Chat that reads text and images
Claude Opus 4.7Coding, sight and reasoning in one model
Granite Vision 4.1 4BChart and table extraction

Send Your Own Test Images

Pick five images from your real workload: one tiny, one 1080p, one 12 megapixel photo, one receipt and one chart. Run them through GPT 5.4 and GPT-4o on Picasso IA, write down which answers were right, and then take the winner to the API and compare token counts at every detail value. An hour of testing like that saves more money than any pricing tweak.

Need test material? Open Picasso IA, generate your own scenes with the text to image models, and turn them into a test set for your vision prompts. You can browse everything available at picassoia.com/en/all-models. Start with one image, ask one narrow question, and see exactly what the model reads.

Share this article