Generate imagesVisual EffectsLarge Language Models
GPT Image 2 API Azure: Pricing and Setup on Azure OpenAI
Setting up GPT Image 2 on Azure OpenAI means dealing with token pricing, a 5 images per minute starting quota, base64-only output and deployment names. This article lays out the real cost per render, the Foundry deployment steps, working REST and Python calls, and five mistakes that inflate the bill.
You can call GPT Image 2 from your own Azure subscription, and when you do, the bill, the quota and the security model all follow Azure rules instead of OpenAI's. That changes how you plan for it. Pricing is token based, so a render costs more or less depending on quality and size rather than a flat per-image fee, and the default quota sits at just 5 images per minute until you ask for more. This article walks through what the model costs, how to deploy it in Microsoft Foundry, the exact REST and Python calls, and the mistakes that quietly inflate a bill. If you only want pictures and no cloud setup, there is a browser route near the end.
What GPT Image 2 Offers on Azure
Azure OpenAI, the Azure side of Microsoft Foundry, hosts OpenAI's image models next to its text models. You create a resource, deploy a model under a name you choose, and send requests to an endpoint named after your own resource. Nothing about the request body is exotic. The differences show up in billing, quota, authentication and region choice.
The Model Lineup
Microsoft's documentation lists three generally available image models in the current family, plus an older series that still needs an access application.
The practical takeaway: GPT Image 2 and the 2.5 models need no approval form, so a new subscription can deploy them the same afternoon. The 1 series requires a registration request first.
Who Should Pick Azure
Azure earns its extra setup when at least one of these is true:
Your company already buys cloud services through an Azure agreement, so image spend lands on an existing invoice.
You need to choose the region where requests are processed.
You want Microsoft Entra ID sign-in, role assignments and private networking instead of a shared secret in an environment file.
Quota should be governed per subscription and per deployment, with an admin who can raise it.
Azure is the wrong pick when you are still deciding whether the idea works. A solo builder testing a concept does not need a resource group, a role assignment and a quota request just to see ten renders. Prototype where setup is cheap, then move to Azure once the prompts, the sizes and the monthly volume are settled.
Here is how the three common routes compare.
Azure OpenAI
OpenAI direct
Picasso IA in the browser
Setup
Subscription, resource, deployment
Account and billing
Sign in and type a prompt
Billing unit
Tokens, on an Azure invoice
Tokens, on an OpenAI invoice
Handled by the platform
Sizes
Custom sizes in 16 px steps
Set by OpenAI's own API docs
1:1, 3:2, 2:3
Best for
Production apps with governance needs
Fast prototypes
Drafts, tests, non-developers
The Real Cost of Each Render
There is no sticker price per picture. You pay for tokens: the text of your prompt, any reference images you send, and, by far the largest share, the image the model produces.
Token Rates
OpenAI publishes these standard rates for GPT Image 2, per one million tokens:
Token type
Price per 1M tokens
Text input
$5.00
Cached text input
$1.25
Image input
$8.00
Cached image input
$2.00
Image output
$30.00
Batch processing is listed at half of the standard rate on OpenAI's own pricing page. Third-party Azure price calculators show $5.00 per million input tokens and $30.00 per million output tokens for the Azure deployment, which matches the text input and image output lines above.
💡 Tip: Azure publishes rates by region and deployment type, and those can differ from the public OpenAI table. Confirm the exact numbers in the Azure pricing calculator for your region before you lock a budget.
Worked Cost Examples
A single image's price is its output token count multiplied by $30 and divided by one million. The model decides how many tokens a render uses, and that depends on quality and canvas size. The counts below are assumptions chosen to show the arithmetic, not measurements. Read the usage data returned with each real response to see your true numbers.
Assumed output tokens
Cost at $30 per 1M
Cost for 1,000 images
1,000
$0.03
$30
2,500
$0.075
$75
5,000
$0.15
$150
8,000
$0.24
$240
Add the prompt itself: a 500-token prompt at $5 per million adds $0.0025, which is rounding error next to the image. The output side drives the total, so budget around it.
Cutting the Bill
Draft at low quality. Move to medium or high only after the composition is right.
Match the canvas to the page. A 1536x864 render is plenty for a blog hero. OpenAI's image documentation marks anything above 2560x1440 as experimental, and bigger canvases generally consume more output tokens.
Ask for one image at a time while you iterate. The n parameter accepts 1 to 10, and every extra image is billed.
Reuse prompt prefixes. Cached text input costs $1.25 per million instead of $5.00, which helps when a long style description repeats across thousands of calls.
Check batch options for work that can wait, since batch pricing is half price on OpenAI's side.
Pre-Deployment Checklist
Two things trip people up before they send a single request: where the model is available, and how little default quota they get.
Subscription and Region
You need an active Azure subscription and an Azure OpenAI resource in a region that offers the model. Image generation is available in a subset of regions, so open Microsoft's region availability table, find GPT Image 2, and pick a region close to your users or your data rules. Creating the resource first and checking the table second is the most common reason for a failed deployment.
Quota and Rate Limits
The documented default quota for GPT Image 2 is 5 images per minute. A few consequences follow:
A burst of parallel jobs will hit HTTP 429 responses almost immediately.
Because the quota is counted in images, a single request asking for ten images may exceed it on a fresh deployment. Test your n value before you depend on it.
Plan for retries with exponential backoff, and file a quota increase request early if you expect real traffic.
A small retry wrapper keeps a batch job alive when the limit bites:
import time
from openai import RateLimitError
def generate_with_retry(client, deployment, prompt, attempts=5):
for attempt in range(attempts):
try:
return client.images.generate(
model=deployment, prompt=prompt, size="1536x864", quality="low", n=1
)
except RateLimitError:
if attempt == attempts - 1:
raise
time.sleep(2 ** attempt * 5)
With the default quota, a wait of 5, 10, 20 and 40 seconds between attempts gives the per-minute window time to refill.
💡 Tip: Treat 5 images per minute as a test-drive limit. A product page that renders images on demand will need a raised quota, a queue, or both.
Deploy the Model in Foundry
Portal menus get renamed between releases, so follow the model name rather than a memorized click path.
Create the Resource
In the Azure portal, create an Azure OpenAI resource, or a Microsoft Foundry project, in a region that lists GPT Image 2.
Pick a resource name you can live with. It becomes part of the endpoint: https://<resource-name>.openai.azure.com.
Finish the creation wizard and wait for provisioning to succeed.
Create the Deployment
Open the Foundry portal and go to the deployments page for your resource.
Choose to deploy a base model and search for gpt-image-2.
Select a deployment type, set the capacity, and give the deployment a name.
That name matters. On Azure, the model field in your code is the deployment name, not the public model ID. If you call it images-prod, your code sends images-prod.
Foundry also offers a playground for image models, which is the quickest way to confirm a deployment works before you write code.
Find the Endpoint and Credentials
Your endpoint appears on the resource's overview and on the deployments page. For authentication, Azure lets you choose between a static credential and Microsoft Entra ID. For anything beyond a throwaway test, use Entra ID: assign the Cognitive Services OpenAI User role to your identity, sign in with the Azure CLI, and let the SDK fetch short-lived tokens. Nothing sensitive ends up in your repository that way.
Microsoft's current documentation uses API version 2025-04-01-preview or later. These are the request rules for GPT Image 2:
Size: custom WIDTHxHEIGHT, both edges multiples of 16, neither above 3,840 pixels.
Aspect ratio: between 1:3 and 3:1.
Pixel count: between 655,360 and 8,294,400.
Quality:low, medium or high.
Count:n from 1 to 10.
The REST Request
TOKEN=$(az account get-access-token \
--resource https://cognitiveservices.azure.com \
--query accessToken -o tsv)
curl -X POST "https://<resource-name>.openai.azure.com/openai/deployments/<deployment-name>/images/generations?api-version=2025-04-01-preview" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $TOKEN" \
-d '{
"prompt": "A ceramic mug of coffee on an oak desk, morning window light, 85mm photograph",
"size": "1536x864",
"quality": "medium",
"n": 1
}'
Python With the OpenAI SDK
import base64
from azure.identity import DefaultAzureCredential, get_bearer_token_provider
from openai import AzureOpenAI
token_provider = get_bearer_token_provider(
DefaultAzureCredential(),
"https://cognitiveservices.azure.com/.default",
)
client = AzureOpenAI(
azure_endpoint="https://<resource-name>.openai.azure.com",
azure_ad_token_provider=token_provider,
api_version="2025-04-01-preview",
)
result = client.images.generate(
model="<deployment-name>",
prompt="A ceramic mug of coffee on an oak desk, morning window light, 85mm photograph",
size="1536x864",
quality="medium",
n=1,
)
with open("render.png", "wb") as f:
f.write(base64.b64decode(result.data[0].b64_json))
Save the Base64 Output
Here is the detail that surprises people: the Azure image endpoint for these models returns base64 data only, with no hosted URL option. Every response carries the full image inside the JSON.
That has real consequences for your app:
Decode and write the file to durable storage right away, such as a blob container.
Store the storage URL in your database, not the base64 string.
Resize thumbnails on your side, since the API will not do it for you.
Never regenerate an image you already paid for just because you lost the bytes.
Storage is the cheap part of the pipeline. A few hundred kilobytes per image in blob storage costs a tiny fraction of what the render itself cost, so keeping every output you pay for is almost always the better trade.
Five Mistakes That Waste Money
Leaving quality on high for drafts. Output tokens are the main expense, and high quality asks for the most of them. Draft low, finish high.
Requesting 4K canvases for web pages. If the page shows the image at 1,280 pixels wide, a 3,840 pixel render is money you cannot see.
Retrying 429 errors in a tight loop. Rate-limited calls hammered without backoff waste time and can stack up failed work.
Sending the public model name. The model value must be your deployment name. A mismatch returns an error, and engineers then burn an hour blaming the region.
Not saving the output. Base64 responses vanish with your process. If you did not write the bytes to storage, the only fix is another paid render.
Use GPT Image 2 on PicassoIA
Not every project needs a cloud resource. GPT Image 2 on PicassoIA runs from a web page, which makes it a fast way to test prompts before you spend Azure tokens on them.
Pick Your Settings
Open the model page, type your prompt, and adjust the options you need.
Setting
What it does
Default
Prompt
Describes the image (required)
None
Quality
Low, medium or high detail
auto
Background
Transparent, opaque or automatic
auto
Aspect ratio
1:1, 3:2 or 2:3
1:1
Number of images
1 to 10 variations per run
1
Output format
PNG, JPEG or WebP
webp
Compression
0 to 100 percent
90
Moderation
Content filter level
auto
Input images
Reference images for edits and variations
None
An optional field accepts your own OpenAI credential. Leave it empty and the platform routes the request for you.
💡 Tip: Prompts that name a lens, a light direction and a surface texture produce steadier results. Ask a text model such as Claude Sonnet 5 or GPT 5.6 Terra to draft three prompt variations, then test each one here.
Generate and Download
Press the generate button and wait for the render.
Check text rendering, hands and edges at full size.
Download the file in your chosen format.
Change one setting at a time and run again.
A sensible workflow splits the work in two. Use the browser to settle composition, wording and text placement, because iterations there are quick and you can run up to ten variations at once. Once a prompt reliably produces what you want, copy it into your Azure code and run it at the size and quality your product needs. The prompt carries over unchanged, so every experiment you ran beforehand saves paid calls later.
You now have the full picture of GPT Image 2 on Azure: token pricing that rewards low-quality drafts, a 5 images per minute starting quota, base64-only output, and a deployment name that doubles as the model field. Those four facts decide most budgets and most bugs.
The cheapest way to find a prompt that works is to test it before it touches your Azure invoice. Open GPT Image 2 on Picasso IA, write the product shot, poster or hero image you have in mind, and run five variations at low quality. Pick the winner, then take that exact prompt into your Azure deployment with confidence. Your first render is one prompt away, so go make it.