Generate imagesVisual EffectsLarge Language Models

GPT Image 2 API Azure: Pricing and Setup on Azure OpenAI

Setting up GPT Image 2 on Azure OpenAI means dealing with token pricing, a 5 images per minute starting quota, base64-only output and deployment names. This article lays out the real cost per render, the Foundry deployment steps, working REST and Python calls, and five mistakes that inflate the bill.

GPT Image 2 API Azure: Pricing and Setup on Azure OpenAI
Cristian Da Conceicao
Founder of Picasso IA

You can call GPT Image 2 from your own Azure subscription, and when you do, the bill, the quota and the security model all follow Azure rules instead of OpenAI's. That changes how you plan for it. Pricing is token based, so a render costs more or less depending on quality and size rather than a flat per-image fee, and the default quota sits at just 5 images per minute until you ask for more. This article walks through what the model costs, how to deploy it in Microsoft Foundry, the exact REST and Python calls, and the mistakes that quietly inflate a bill. If you only want pictures and no cloud setup, there is a browser route near the end.

What GPT Image 2 Offers on Azure

Developer hands on a silver laptop in a bright coworking space

Azure OpenAI, the Azure side of Microsoft Foundry, hosts OpenAI's image models next to its text models. You create a resource, deploy a model under a name you choose, and send requests to an endpoint named after your own resource. Nothing about the request body is exotic. The differences show up in billing, quota, authentication and region choice.

The Model Lineup

Microsoft's documentation lists three generally available image models in the current family, plus an older series that still needs an access application.

ModelAzure statusQuality settings
GPT Image 2Generally availablelow, medium, high
GPT Image 2.5 SunburstGenerally availablelow, medium, high, xhigh, max, auto
GPT Image 2.5 FlareGenerally availablelow, medium, high, xhigh, max, auto
GPT Image 1, GPT Image 1.5, GPT Image 1 MiniLimited access previewlow, medium, high

The practical takeaway: GPT Image 2 and the 2.5 models need no approval form, so a new subscription can deploy them the same afternoon. The 1 series requires a registration request first.

Who Should Pick Azure

Azure earns its extra setup when at least one of these is true:

  • Your company already buys cloud services through an Azure agreement, so image spend lands on an existing invoice.
  • You need to choose the region where requests are processed.
  • You want Microsoft Entra ID sign-in, role assignments and private networking instead of a shared secret in an environment file.
  • Quota should be governed per subscription and per deployment, with an admin who can raise it.

Azure is the wrong pick when you are still deciding whether the idea works. A solo builder testing a concept does not need a resource group, a role assignment and a quota request just to see ten renders. Prototype where setup is cheap, then move to Azure once the prompts, the sizes and the monthly volume are settled.

Here is how the three common routes compare.

Azure OpenAIOpenAI directPicasso IA in the browser
SetupSubscription, resource, deploymentAccount and billingSign in and type a prompt
Billing unitTokens, on an Azure invoiceTokens, on an OpenAI invoiceHandled by the platform
SizesCustom sizes in 16 px stepsSet by OpenAI's own API docs1:1, 3:2, 2:3
Best forProduction apps with governance needsFast prototypesDrafts, tests, non-developers

The Real Cost of Each Render

Overhead desk with a printed statement, calculator and cup of tea

There is no sticker price per picture. You pay for tokens: the text of your prompt, any reference images you send, and, by far the largest share, the image the model produces.

Token Rates

OpenAI publishes these standard rates for GPT Image 2, per one million tokens:

Token typePrice per 1M tokens
Text input$5.00
Cached text input$1.25
Image input$8.00
Cached image input$2.00
Image output$30.00

Batch processing is listed at half of the standard rate on OpenAI's own pricing page. Third-party Azure price calculators show $5.00 per million input tokens and $30.00 per million output tokens for the Azure deployment, which matches the text input and image output lines above.

💡 Tip: Azure publishes rates by region and deployment type, and those can differ from the public OpenAI table. Confirm the exact numbers in the Azure pricing calculator for your region before you lock a budget.

Worked Cost Examples

A single image's price is its output token count multiplied by $30 and divided by one million. The model decides how many tokens a render uses, and that depends on quality and canvas size. The counts below are assumptions chosen to show the arithmetic, not measurements. Read the usage data returned with each real response to see your true numbers.

Assumed output tokensCost at $30 per 1MCost for 1,000 images
1,000$0.03$30
2,500$0.075$75
5,000$0.15$150
8,000$0.24$240

Add the prompt itself: a 500-token prompt at $5 per million adds $0.0025, which is rounding error next to the image. The output side drives the total, so budget around it.

Cutting the Bill

  • Draft at low quality. Move to medium or high only after the composition is right.
  • Match the canvas to the page. A 1536x864 render is plenty for a blog hero. OpenAI's image documentation marks anything above 2560x1440 as experimental, and bigger canvases generally consume more output tokens.
  • Ask for one image at a time while you iterate. The n parameter accepts 1 to 10, and every extra image is billed.
  • Reuse prompt prefixes. Cached text input costs $1.25 per million instead of $5.00, which helps when a long style description repeats across thousands of calls.
  • Check batch options for work that can wait, since batch pricing is half price on OpenAI's side.

Pre-Deployment Checklist

Rows of server racks along a clean data center aisle

Two things trip people up before they send a single request: where the model is available, and how little default quota they get.

Subscription and Region

You need an active Azure subscription and an Azure OpenAI resource in a region that offers the model. Image generation is available in a subset of regions, so open Microsoft's region availability table, find GPT Image 2, and pick a region close to your users or your data rules. Creating the resource first and checking the table second is the most common reason for a failed deployment.

Quota and Rate Limits

The documented default quota for GPT Image 2 is 5 images per minute. A few consequences follow:

  • A burst of parallel jobs will hit HTTP 429 responses almost immediately.
  • Because the quota is counted in images, a single request asking for ten images may exceed it on a fresh deployment. Test your n value before you depend on it.
  • Plan for retries with exponential backoff, and file a quota increase request early if you expect real traffic.

A small retry wrapper keeps a batch job alive when the limit bites:

import time
from openai import RateLimitError

def generate_with_retry(client, deployment, prompt, attempts=5):
    for attempt in range(attempts):
        try:
            return client.images.generate(
                model=deployment, prompt=prompt, size="1536x864", quality="low", n=1
            )
        except RateLimitError:
            if attempt == attempts - 1:
                raise
            time.sleep(2 ** attempt * 5)

With the default quota, a wait of 5, 10, 20 and 40 seconds between attempts gives the per-minute window time to refill.

💡 Tip: Treat 5 images per minute as a test-drive limit. A product page that renders images on demand will need a raised quota, a queue, or both.

Deploy the Model in Foundry

Cloud engineer at a standing desk with two monitors

Portal menus get renamed between releases, so follow the model name rather than a memorized click path.

Create the Resource

  1. In the Azure portal, create an Azure OpenAI resource, or a Microsoft Foundry project, in a region that lists GPT Image 2.
  2. Pick a resource name you can live with. It becomes part of the endpoint: https://<resource-name>.openai.azure.com.
  3. Finish the creation wizard and wait for provisioning to succeed.

Create the Deployment

  1. Open the Foundry portal and go to the deployments page for your resource.
  2. Choose to deploy a base model and search for gpt-image-2.
  3. Select a deployment type, set the capacity, and give the deployment a name.

That name matters. On Azure, the model field in your code is the deployment name, not the public model ID. If you call it images-prod, your code sends images-prod.

Foundry also offers a playground for image models, which is the quickest way to confirm a deployment works before you write code.

Find the Endpoint and Credentials

Your endpoint appears on the resource's overview and on the deployments page. For authentication, Azure lets you choose between a static credential and Microsoft Entra ID. For anything beyond a throwaway test, use Entra ID: assign the Cognitive Services OpenAI User role to your identity, sign in with the Azure CLI, and let the SDK fetch short-lived tokens. Nothing sensitive ends up in your repository that way.

Call the API From Code

Hands typing on a laptop at a kitchen table in morning light

The image endpoint follows this pattern:

https://<resource-name>.openai.azure.com/openai/deployments/<deployment-name>/images/generations?api-version=2025-04-01-preview

Microsoft's current documentation uses API version 2025-04-01-preview or later. These are the request rules for GPT Image 2:

  • Size: custom WIDTHxHEIGHT, both edges multiples of 16, neither above 3,840 pixels.
  • Aspect ratio: between 1:3 and 3:1.
  • Pixel count: between 655,360 and 8,294,400.
  • Quality: low, medium or high.
  • Count: n from 1 to 10.

The REST Request

TOKEN=$(az account get-access-token \
  --resource https://cognitiveservices.azure.com \
  --query accessToken -o tsv)

curl -X POST "https://<resource-name>.openai.azure.com/openai/deployments/<deployment-name>/images/generations?api-version=2025-04-01-preview" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $TOKEN" \
  -d '{
    "prompt": "A ceramic mug of coffee on an oak desk, morning window light, 85mm photograph",
    "size": "1536x864",
    "quality": "medium",
    "n": 1
  }'

Python With the OpenAI SDK

import base64
from azure.identity import DefaultAzureCredential, get_bearer_token_provider
from openai import AzureOpenAI

token_provider = get_bearer_token_provider(
    DefaultAzureCredential(),
    "https://cognitiveservices.azure.com/.default",
)

client = AzureOpenAI(
    azure_endpoint="https://<resource-name>.openai.azure.com",
    azure_ad_token_provider=token_provider,
    api_version="2025-04-01-preview",
)

result = client.images.generate(
    model="<deployment-name>",
    prompt="A ceramic mug of coffee on an oak desk, morning window light, 85mm photograph",
    size="1536x864",
    quality="medium",
    n=1,
)

with open("render.png", "wb") as f:
    f.write(base64.b64decode(result.data[0].b64_json))

Save the Base64 Output

Printed photographs arranged in a grid on a design studio table

Here is the detail that surprises people: the Azure image endpoint for these models returns base64 data only, with no hosted URL option. Every response carries the full image inside the JSON.

That has real consequences for your app:

  • Decode and write the file to durable storage right away, such as a blob container.
  • Store the storage URL in your database, not the base64 string.
  • Resize thumbnails on your side, since the API will not do it for you.
  • Never regenerate an image you already paid for just because you lost the bytes.

Storage is the cheap part of the pipeline. A few hundred kilobytes per image in blob storage costs a tiny fraction of what the render itself cost, so keeping every output you pay for is almost always the better trade.

Five Mistakes That Waste Money

Man frowning at a laptop beside a stack of paper receipts

  1. Leaving quality on high for drafts. Output tokens are the main expense, and high quality asks for the most of them. Draft low, finish high.
  2. Requesting 4K canvases for web pages. If the page shows the image at 1,280 pixels wide, a 3,840 pixel render is money you cannot see.
  3. Retrying 429 errors in a tight loop. Rate-limited calls hammered without backoff waste time and can stack up failed work.
  4. Sending the public model name. The model value must be your deployment name. A mismatch returns an error, and engineers then burn an hour blaming the region.
  5. Not saving the output. Base64 responses vanish with your process. If you did not write the bytes to storage, the only fix is another paid render.

Use GPT Image 2 on PicassoIA

Designer choosing an image from a gallery on a large monitor

Not every project needs a cloud resource. GPT Image 2 on PicassoIA runs from a web page, which makes it a fast way to test prompts before you spend Azure tokens on them.

Pick Your Settings

Open the model page, type your prompt, and adjust the options you need.

SettingWhat it doesDefault
PromptDescribes the image (required)None
QualityLow, medium or high detailauto
BackgroundTransparent, opaque or automaticauto
Aspect ratio1:1, 3:2 or 2:31:1
Number of images1 to 10 variations per run1
Output formatPNG, JPEG or WebPwebp
Compression0 to 100 percent90
ModerationContent filter levelauto
Input imagesReference images for edits and variationsNone

An optional field accepts your own OpenAI credential. Leave it empty and the platform routes the request for you.

💡 Tip: Prompts that name a lens, a light direction and a surface texture produce steadier results. Ask a text model such as Claude Sonnet 5 or GPT 5.6 Terra to draft three prompt variations, then test each one here.

Generate and Download

  1. Press the generate button and wait for the render.
  2. Check text rendering, hands and edges at full size.
  3. Download the file in your chosen format.
  4. Change one setting at a time and run again.

A sensible workflow splits the work in two. Use the browser to settle composition, wording and text placement, because iterations there are quick and you can run up to ten variations at once. Once a prompt reliably produces what you want, copy it into your Azure code and run it at the size and quality your product needs. The prompt carries over unchanged, so every experiment you ran beforehand saves paid calls later.

Want to compare engines? Try GPT Image 2.5 Flare, Seedream 4.5 or Nano Banana 2 Lite with the same prompt and keep whichever looks best.

Make Your First Image Today

Woman smiling at a tablet by a sunlit window

You now have the full picture of GPT Image 2 on Azure: token pricing that rewards low-quality drafts, a 5 images per minute starting quota, base64-only output, and a deployment name that doubles as the model field. Those four facts decide most budgets and most bugs.

The cheapest way to find a prompt that works is to test it before it touches your Azure invoice. Open GPT Image 2 on Picasso IA, write the product shot, poster or hero image you have in mind, and run five variations at low quality. Pick the winner, then take that exact prompt into your Azure deployment with confidence. Your first render is one prompt away, so go make it.

Share this article