Large Language ModelsGenerate imagesGenerate videos

GPT-6 API Cost: Pricing, Access and Free Options

GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens, while Sol and Luna cost far less. See every surcharge, a monthly bill for 100,000 requests, who gets API access, and which free options hold up.

GPT-6 API Cost: Pricing, Access and Free Options
Cristian Da Conceicao
Founder of Picasso IA

Ten dollars in, fifty dollars out. That is what a million tokens of GPT-6 Astra costs through the API, and it is double the list rate of the flagship that came before it. If you were hoping for a cheap upgrade, that headline number stings. The catch is that GPT-6 is not one price. It is a three-tier family, and the smallest model costs about one percent of the biggest.

This article puts every number in one place: per-token rates for Astra, Sol and Luna, the surcharges that quietly inflate invoices, what 100,000 real requests cost, who can actually call the API, and which free options hold up. Prices come from public pricing trackers in early October, so treat them as a snapshot and confirm on OpenAI's own pricing page before you commit a budget.

What GPT-6 Costs Per Token

Astra was unveiled as a limited preview on September 3 and opened to the public the next day. Sol and Luna followed on September 22. All three share a context window of roughly one million tokens with up to 128K tokens of output, although trackers disagree on the exact figure (1.05M versus 1.1M).

Astra, the Flagship Price

Astra is built for long, agentic jobs: computer use, browser automation, big coding tasks. You pay for that ambition.

Billing linePrice per 1M tokens
Input$10.00
Cached input$1.00
Output$50.00
Input above 272K tokens$20.00
Output above 272K tokens$75.00
Batch or Flex (input / output)$5.00 / $25.00

Output is the expensive half. Every token the model writes costs five times what a token it reads costs, and reasoning-heavy models write a lot, including thinking tokens you never see in the final answer. One early third-party test put Astra at roughly $167 per finished task on aggregate benchmarks, which shows how fast agentic runs add up.

A woman's hands pressing calculator buttons beside a fanned stack of printed invoices on an oak desk

Sol and Luna for Smaller Budgets

Here is the whole family side by side.

ModelInputCached inputOutputWhere it fits
GPT-6 Astra$10.00$1.00$50.00Agents, hard coding, long research runs
GPT-6 Sol$2.00$0.10$10.00Daily coding, writing, support replies
GPT-6 Luna$0.10$0.01$0.50Routing, tagging, short answers at volume

Sol costs one fifth of Astra on every line. Luna costs one hundredth. Think of three coffee cups: same drink, very different bill.

Token rates are abstract, so here is a rough translation. One token is about three quarters of an English word. A single dollar of Astra output buys around 15,000 words, close to a 30-page report. The same dollar buys about 75,000 words on Sol and roughly 1.5 million words on Luna. Reading is cheaper than writing on every tier, which is why summarizing a long document costs far less than drafting one.

Three ceramic coffee cups in small, medium and extra large sizes on a cafe counter

The previous generation is not automatically cheaper. According to one pricing tracker, GPT 5.6 Sol sits at $4 input and $20 output during a promotional window that runs until November 21, which makes GPT-6 Sol the cheaper line item under the same tier name. GPT 5.6 Terra lists at $2 and $12, and GPT 5.6 Luna at $0.20 and $1.20, double the rate of GPT-6 Luna.

💡 Quick read: if your workload does not need agents or deep reasoning, the Sol tier is the new default to price first. Astra is the exception, not the baseline.

So when is Astra worth five times the price of Sol? When a failed run costs more than the price gap: an agent that burns an hour of engineer time, or a legal draft that would otherwise need three revision rounds. The break-even is simple. If Sol needs more than five attempts to match one good Astra result, Astra wins on price alone. If Sol gets there in two or three, you are overpaying.

Surcharges That Inflate the Bill

Long Prompts Double the Rate

Cross 272K input tokens and the whole request is repriced, not only the extra tokens. Input jumps from $10 to $20 per million, and output from $50 to $75.

Picture a 300K-token contract bundle with a 2K-token summary as the answer:

  • At the long-context rate: $6.00 for input plus $0.15 for output, so $6.15.
  • Trimmed to 270K tokens at the standard rate: $2.70 plus $0.10, so $2.80.

Thirty thousand extra tokens more than doubled the price of the same request.

A man holding an enormously thick stack of printed pages in a quiet library aisle

💡 Tip: keep a token counter in your pipeline and split big documents just under the threshold instead of sending one giant prompt.

Fast Mode and Regional Uplift

Two more multipliers hide in the fine print:

  • Fast mode costs about double the standard rate ($20 input, $100 output on Astra) in exchange for roughly 2.5x the speed. It is not available with EU data residency.
  • Data-residency endpoints add 10% on top of whatever tier you chose.

Stack them and an innocent-looking Astra call can cost more than twice its list price. Decide on purpose, not by default.

Fast mode earns its premium when a human is waiting on the answer, like an interactive coding assistant or a live support chat. It is wasted on overnight jobs where nobody checks the clock. Those belong in the batch queue at half price.

Real Monthly Costs in Dollars

One Support Ticket, Priced Out

Take a support bot where each ticket uses 2,000 input tokens and 500 output tokens. On Astra that is $0.02 for the prompt plus $0.025 for the answer, so $0.045 per ticket. Small number, until you multiply it.

A bearded founder in a denim jacket reviewing a printed monthly statement at a cafe window

The Same Workload on Three Tiers

Here is that exact ticket at 100,000 tickets a month:

ModelPer ticket100,000 ticketsWith Batch discount
GPT-6 Astra$0.045$4,500$2,250
GPT-6 Sol$0.009$900$450
GPT-6 Luna$0.00045$45$22.50

Caching adds another lever. If 1,500 of those 2,000 input tokens are the same system prompt and examples every time, Astra drops to $0.0315 per ticket, or $3,150 a month. That saves $1,350 without touching quality.

💡 Rule of thumb: on Astra, a 500-token answer ($0.025) costs more than a 2,000-token prompt ($0.02). Output length decides the invoice.

Support tickets are only one shape of workload. Here are three more, priced at standard rates with no discounts, so you can find the one closest to your own project:

WorkloadTokens per requestRequests a monthLunaSolAstra
Solo developer chat app1,000 in, 400 out5,000$1.50$30$150
Content team drafts3,000 in, 2,500 out300$0.47$9.30$46.50
Coding agent runs200,000 in, 30,000 out440$15.40$308$1,540

The table prices tokens, not quality. A Luna agent will not match an Astra agent, so read the last row as a ceiling and a floor, not a recommendation. Notice how the coding agent dwarfs everything else: only 440 runs, yet $1,540 on Astra. Caching 150K of each 200K prompt brings that down to about $946.

Who Can Call the API

Where the API Starts

Astra is not available on the API free tier. You need at least a Tier 1 developer account, which means a paid account with billing set up. OpenAI had not published detailed rate limits or throughput guarantees for Astra at launch, so check the limits page in your dashboard once access is granted.

A hand sliding a white access card across a wooden reception desk to another hand

ChatGPT Plans Are Separate

A ChatGPT subscription does not give you API credits. The two are billed separately, and the right choice depends on how you work. If you mostly chat, a Plus plan at $20 is usually simpler than paying per token. If you are building a product, you need the API no matter which plan you hold, because customers cannot use your personal subscription. Inside the app, Astra access looks like this:

PlanMonthly priceAstra access
Free$0No
Go$8No
Plus$20Yes
ProFrom $100Yes, with expanded limits
EnterprisePer seatYes, switched off by default

Luna had not reached free ChatGPT users yet when this was written, so even the cheapest tier is not a free lunch inside the app.

Before your first API call, run through this short list:

  1. Add a billing method and a small prepaid balance.
  2. Set a monthly spend cap in the dashboard before you write any code.
  3. Create a project-scoped API credential and store it in an environment variable, never in the repository.
  4. Build in development on Luna, then switch the model name once the pipeline works.
  5. Log input, cached and output token counts for every request so the bill never surprises you.

Free Options That Actually Work

Why Astra Has No Free Tier

No trial allocation has been published for Astra, and at $50 per million output tokens that is not a surprise. A single long agentic run would burn through any realistic free credit within minutes.

A university student studying with a laptop and notebook at a long library table

Where Free Usage Still Exists

You still have real routes to test ideas without a big bill:

  • Luna as the cheap entry point. Five dollars of credit buys about 50 million input tokens or 10 million output tokens. That is thousands of short conversations.
  • Open-weight models. GPT OSS 120B and GPT OSS 20B are OpenAI's open-weight releases. If you own the hardware, running them yourself costs no per-token fee.
  • Rival models in the browser. Try Claude Sonnet 5, Gemini 3.5 Flash or GPT 5.4 with your real prompts before you decide where to spend.

A practical free workflow looks like this. Prototype the prompt in a browser until the answers are good, lock it, then count the tokens of a typical request. Price it yourself with one formula: (input tokens x input rate + output tokens x output rate) / 1,000,000. Run that for Luna, Sol and Astra, and you will know your real monthly range before a single paid call goes out.

Try GPT 5.6 Sol on PicassoIA

One honest note: GPT-6 itself is not on the PicassoIA model list yet. The closest relatives are the GPT 5.6 family: GPT 5.6 Sol, GPT 5.6 Terra and GPT 5.6 Luna. Sol is free to try online with no setup, which makes it a cheap rehearsal before you commit API budget.

Step by Step on PicassoIA

  1. Open the GPT 5.6 Sol page.
  2. Paste a real task from your product into Prompt, not a "hello".
  3. Add a System Prompt to set the role and tone.
  4. Pick a Reasoning effort from none to xhigh. Start at none and raise it only when answers feel shallow.
  5. Set Verbosity to low, medium or high.
  6. Raise Max Completion Tokens if you choose a high reasoning level, otherwise thinking can use the whole budget and leave an empty reply.
  7. Upload a screenshot or diagram in Image Input if the task needs one.
  8. Run it, then repeat the same prompt on GPT 5.6 Luna and compare quality against price.

Over-the-shoulder view of a young creator browsing on a laptop in a bright shared workspace

Settings That Move Your Cost

Reasoning tokens are billed as output tokens on reasoning models, so the dials on that page map straight onto an invoice:

SettingWhat it doesCost effect
Reasoning effort: none or lowFast replies, little thinkingLowest
Reasoning effort: high or xhighDeeper, slower answersMore output tokens
Verbosity: lowShort, direct answersFewer output tokens
Max Completion TokensHard cap on reply lengthUpper limit per request

Ways to Cut the Bill

A hand dropping a gold coin into a glass jar half full of coins on a kitchen table

Route by Difficulty

Send easy work to Luna, middle work to Sol, and only the hardest 5% to Astra. With a made-up split of 80% Luna, 15% Sol and 5% Astra, the 100,000-ticket workload costs about $396 instead of $4,500.

Cache and Batch

Cached input is 90% cheaper on Astra and Luna, and 95% cheaper on Sol. Batch and Flex processing halve the standard rate for anything that can wait.

A logistics worker scanning rows of identical cardboard parcels in a bright warehouse

Cap the Output

Set a maximum token limit and ask for short formats such as bullets or compact JSON. With output priced at five times input, every sentence you skip is money kept.

Three mistakes that inflate bills:

  • Resending the whole chat every turn. A 10-turn conversation that adds 300 tokens per turn sends 16,500 input tokens in total, not 3,000. Summarize old turns or cache the shared prefix.
  • Leaving reasoning on high for easy jobs. Thinking tokens bill as output, the expensive side of the rate card.
  • Testing on the flagship. Build and debug on Luna, then switch the model name for the final run.

💡 Checklist: route by difficulty, cache the repeated prefix, batch whatever can wait, cap the reply length, and re-check prices monthly because tiers and promotions change.

Your Turn to Create

Text is only half of a product budget. Most launches also need pictures: hero shots, thumbnails, product mockups. Try creating your own images with Picasso IA and see how far a prompt can go. Seedream 4.5 is great for sharp 4K detail, P-Image is quick for drafts, and GPT Image 2 follows long prompts closely.

Pick a scene from your own project, write three prompts with different camera angles, and run them side by side. Browse the full catalog at picassoia.com/en/all-models and start experimenting today.

Share this article