Generate imagesVisual EffectsLarge Language Models
GPT API Image Input: Vision Analysis and Detail Settings
The GPT API turns every image into tokens, and the detail field decides how many. This article shows the request format, the low, high, original and auto values, worked token math for tile and patch models, and a step by step test on PicassoIA.
One field in a request body decides whether a photo costs 85 tokens or 3,000. That field is detail, it sits inside the image object next to the URL, and most tutorials either skip it or quote numbers that stopped being true two model generations ago. If you send screenshots, receipts, product shots or charts to a GPT model, this setting shapes your bill, your latency and how much of the picture the model can actually read.
This article walks through the request format, the four detail values, the token arithmetic for tile based and patch based models, the cases where low backfires, and the limits worth knowing before you ship. Numbers come from the image input pages of OpenAI's API documentation as they read today, and every worked example shows its arithmetic, so you can check it against the usage block in your own responses. Pages like that change often, which is one more reason to log token counts instead of trusting a table, including the ones below.
How a Photo Becomes Tokens
A GPT model never reads your JPEG as a file. The API resizes it, slices it into small blocks, and turns every block into tokens that sit in the context window next to your text. Those tokens are billed at the model's normal input rate, so a larger or sharper image means a larger bill and more latency. They also compete with your prompt and the answer for the same context window, which matters once a request carries several pictures.
The Request Shape
Images travel inside the user message as content parts. The Responses API uses input_text and input_image parts, while Chat Completions uses text and image_url. In both, detail lives on the image part. Here is a receipt reader on GPT 5.4:
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-5.4",
input=[
{
"role": "user",
"content": [
{"type": "input_text", "text": "List every line item and the total."},
{
"type": "input_image",
"image_url": "https://example.com/receipt.jpg",
"detail": "original",
},
],
}
],
)
print(response.output_text)
print(response.usage.input_tokens)
The same call on Chat Completions with GPT-4o nests the URL one level deeper:
After every call, read usage.input_tokens on Responses or usage.prompt_tokens on Chat Completions. It is the only number that settles an argument about cost.
Three Ways to Send Pixels
Public URL. The simplest option, but OpenAI's servers must be able to fetch it quickly.
Base64 data URL. An inline data:image/jpeg;base64,... string works for private files and adds about a third to the payload size.
File ID. Upload once through the Files API with the vision purpose, then reference the ID with a file_id field on the image part and reuse it across requests.
Accepted formats are PNG, JPEG, WEBP and non-animated GIF.
The Four Detail Values
The detail field accepts low, high, original and auto. Leave it out and you get auto, which means the model's own default sizing. The names suggest a simple ladder, but what each rung does depends on the model you call.
On the older tile based models, low is a flat fee: 85 tokens on GPT-4o and GPT 4.1, whatever the file size. The model receives a 512 by 512 pixel version, enough to say "a dog on a beach" and not enough to read a street sign. Use it for classification, rough captions, moderation passes and routing decisions where only the gist matters.
High and Auto: The Everyday Default
On tile models, high first fits the image inside a 2,048 by 2,048 square, then scales the shortest side to 768 pixels, then counts 512 pixel tiles. On the GPT 5.x family it works from 32 pixel patches instead, with a 2,048 pixel longest side and a 2,500 patch budget on GPT 5.4 and its smaller siblings. That is plenty for photos, product shots and ordinary screenshots. It starts to hurt with dense documents and tiny interface text, where a few lost pixels turn a 6 into an 8.
Original: When Pixels Matter
original raises the ceiling to 10,000 patches and a 6,000 pixel longest side on GPT 5.4 and its siblings. OpenAI points it at large, dense, spatially sensitive or computer-use images, and at coordinate-sensitive work such as OCR or small object detection. Picture a 4K screenshot where the model must return the position of a button, or a floor plan where one thin line carries meaning. You pay for that sharpness: up to 12,000 image tokens per picture at the 1.2 multiplier.
The practical rule is to earn it. Start a task at high, collect the failures, and move only the failing image types to original. If a field was misread at high and reads correctly at original, the extra tokens bought something. If both settings fail the same way, the problem is the prompt or the source image, and more pixels will not help.
Token Math, Worked Out
Tile Based Models
Tile billing is a base fee plus a fee per 512 pixel tile. GPT-4o and GPT 4.1 charge 85 base and 170 per tile. GPT 5.1 charges 70 and 140. GPT 4o Mini charges 2,833 and 5,667, so moving image work to the mini model inflates the token count instead of shrinking it. Compare total cost, not the per token rate.
Newer models count 32 by 32 pixel patches: ceil(width / 32) x ceil(height / 32). When the total exceeds the budget, the image is shrunk until it fits, and the final token count is patches times a model multiplier, rounded up. The multiplier is 1.2 for GPT 5.2, GPT 5.4 and the GPT 5.6 family (Sol, Terra and Luna), and 1.62 for GPT 4.1 mini.
Image
Detail
Patches
Tokens at the 1.2 multiplier
1920 x 1080
low, high or original
60 x 34 = 2,040
2,448
4000 x 3000
high
Capped at 2,500
Up to 3,000
4000 x 3000
low
About 3,072 after the 2,048 pixel cap
About 3,700
4000 x 3000
original
Capped at 10,000
Up to 12,000
The 4000 x 3000 rows are my arithmetic from the published budgets, so treat them as estimates and confirm with usage. The 1080p row is exact because the image already fits every budget.
Scale makes the difference real. Ten thousand product photos at 2,448 tokens each is 24.48 million input tokens before a single word of prompt. The same set at 85 tokens each on GPT-4olow is 850,000, roughly a 29 times difference you would never notice from the code alone.
When Low Costs More Than High
The docs carry a warning that surprises people: low does not always use fewer tokens than high. On GPT 5.4 and its smaller siblings, low allows a 6,144 patch budget while high stops at 2,500, so a large photo can bill more at low. On GPT 5.2 and GPT 4.1 mini, every level shares one sizing rule, a 2,048 pixel longest side and a 6,144 patch budget, so low, high and auto return identical counts and original is unavailable.
Two consequences follow. Never assume the cheap setting is cheap, and expect the meaning of your existing detail values to change when you switch models. A short test settles both:
Pick three representative images: a small one, a 1080p screenshot and a 12 megapixel photo.
Send each one at every detail value the model supports.
Log the input tokens next to the quality of the answer.
Keep the lowest setting that still answers correctly.
💡 If low and high return the same token count on a model, the field does nothing there. Drop it from your code instead of carrying a parameter that suggests savings you are not getting.
Limits That Bite in Production
Payload and Format Limits
The current docs allow up to 512 MB of total payload and up to 1,500 images per request. Older write-ups quote 50 MB and 500 images, so a library that enforces those numbers may be out of date. A 512 MB request is a latency problem long before it is a limit problem, because base64 inflates the bytes by a third and every image still bills tokens.
Where Vision Still Fails
OpenAI lists the weak spots plainly, and they match what shows up in production:
Specialized medical images, such as CT scans, are not a fit.
Non-Latin alphabets, such as Japanese or Korean, can underperform.
Graphs with varying line styles or colors, where solid, dashed and dotted lines must be told apart, cause mistakes.
Precise spatial localization, such as reading chess positions, is unreliable.
Object counts come back as approximations.
CAPTCHAs are blocked.
File names and metadata are never read.
Tiny print is the usual casualty in everyday work. When the part that matters is small, send only that part.
💡 Crop before you send. A 600 x 400 crop of a receipt's total line costs about 300 tokens on GPT 5.4 (19 x 13 = 247 patches, times 1.2). The full 4000 x 3000 page at original can reach 12,000, and the crop usually reads better.
Mistakes That Waste Tokens
Most overspending comes from a handful of habits:
Uploading raw camera files. A 12 megapixel original is shrunk to the model's cap on arrival anyway. Resize to the longest side your detail level allows first (2,048 pixels for most settings, 6,000 on original for GPT 5.4 and its siblings), export a JPEG, and the request uploads faster with nothing lost.
Stitching a collage into one image. Six screenshots pasted onto one canvas get shrunk together, so each loses resolution. Send six separate image parts and name them in the text: "Image 1 is the invoice, image 2 is the packing slip." Each part is billed on its own.
Asking for everything at once. A prompt that wants a caption, a color list, a defect check and an OCR pass invites shallow answers. One narrow question per call, or one clearly numbered list, returns cleaner output.
Never logging usage. Without input token counts per request, a model switch or a new image size can double your bill and nothing will tell you.
Text is where resizing hurts first. Crop to the region, straighten it, and ask for a fixed JSON shape so a wrong digit is easy to spot. Tell the model to answer unreadable for any field it cannot read, because a model that is allowed to guess will guess, and a confident wrong total is worse than a blank one. If a scan is blurry, repair it before upload with an AI image restoration tool, because no detail setting can invent pixels that were never captured.
Charts and Dense Screenshots
Charts mix thin strokes, small labels and similar colors, which is the exact case OpenAI flags. Send them at original when the model supports it. Ask for the underlying numbers as a table first and the interpretation second, so you can spot-check the values against the picture before you trust any trend the model describes. For tables and charts you extract every day, a specialist such as Granite Vision 4.1 4B is worth a side by side test.
Product Photos at Scale
For catalog work, run the three image test from earlier on your real photos, then fix the setting and the prompt for the whole batch. Ask for the same attributes in the same order every time, such as color, material and visible defects, and the output becomes easy to load into a database. Photograph the items the same way too: a consistent background, distance and light let you use a cheaper setting, because the model no longer has to work around noise. A product shot taken on a plain tabletop under even light often reads fine at high, while the same mug in a cluttered kitchen may not.
How to Use GPT 5.4 on PicassoIA
You do not need code to see what a model reads. GPT 5.4 on PicassoIA accepts images next to a text prompt, and its form exposes the same levers you tune in the API: a system prompt, verbosity, reasoning effort and a completion token limit. That makes it a fast place to settle the questions that come before the settings. Which prompt wording works? Does the model read this kind of picture at all? Is a cheaper model good enough? One caveat: the form has an image input field but no detail switch, so treat results as a test of prompt and model choice, not a measurement of low against high.
Step by Step
Open the GPT 5.4 page and find the Image Input field.
Add your test image. A receipt, a dashboard screenshot or a product photo all work.
Write a narrow prompt: "Return the date and the total as JSON" beats "describe this image".
Add a System Prompt that sets the role and the output format.
Set Verbosity to low for extraction jobs and high when you want a detailed breakdown.
Leave Reasoning Effort at none for plain reading. Raise it for multi-step questions, and raise Max Completion Tokens with it, because high effort can spend the whole budget on reasoning and return an empty answer.
Run the same image on GPT-4o and compare the two outputs.
Vision Models Worth Comparing
Image to text is one of the platform's built-in capabilities, and several language models accept images:
Pick five images from your real workload: one tiny, one 1080p, one 12 megapixel photo, one receipt and one chart. Run them through GPT 5.4 and GPT-4o on Picasso IA, write down which answers were right, and then take the winner to the API and compare token counts at every detail value. An hour of testing like that saves more money than any pricing tweak.
Need test material? Open Picasso IA, generate your own scenes with the text to image models, and turn them into a test set for your vision prompts. You can browse everything available at picassoia.com/en/all-models. Start with one image, ask one narrow question, and see exactly what the model reads.