Ten dollars in, fifty dollars out. That is what a million tokens of GPT-6 Astra costs through the API, and it is double the list rate of the flagship that came before it. If you were hoping for a cheap upgrade, that headline number stings. The catch is that GPT-6 is not one price. It is a three-tier family, and the smallest model costs about one percent of the biggest.
This article puts every number in one place: per-token rates for Astra, Sol and Luna, the surcharges that quietly inflate invoices, what 100,000 real requests cost, who can actually call the API, and which free options hold up. Prices come from public pricing trackers in early October, so treat them as a snapshot and confirm on OpenAI's own pricing page before you commit a budget.
What GPT-6 Costs Per Token
Astra was unveiled as a limited preview on September 3 and opened to the public the next day. Sol and Luna followed on September 22. All three share a context window of roughly one million tokens with up to 128K tokens of output, although trackers disagree on the exact figure (1.05M versus 1.1M).
Astra, the Flagship Price
Astra is built for long, agentic jobs: computer use, browser automation, big coding tasks. You pay for that ambition.
| Billing line | Price per 1M tokens |
|---|
| Input | $10.00 |
| Cached input | $1.00 |
| Output | $50.00 |
| Input above 272K tokens | $20.00 |
| Output above 272K tokens | $75.00 |
| Batch or Flex (input / output) | $5.00 / $25.00 |
Output is the expensive half. Every token the model writes costs five times what a token it reads costs, and reasoning-heavy models write a lot, including thinking tokens you never see in the final answer. One early third-party test put Astra at roughly $167 per finished task on aggregate benchmarks, which shows how fast agentic runs add up.

Sol and Luna for Smaller Budgets
Here is the whole family side by side.
| Model | Input | Cached input | Output | Where it fits |
|---|
| GPT-6 Astra | $10.00 | $1.00 | $50.00 | Agents, hard coding, long research runs |
| GPT-6 Sol | $2.00 | $0.10 | $10.00 | Daily coding, writing, support replies |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | Routing, tagging, short answers at volume |
Sol costs one fifth of Astra on every line. Luna costs one hundredth. Think of three coffee cups: same drink, very different bill.
Token rates are abstract, so here is a rough translation. One token is about three quarters of an English word. A single dollar of Astra output buys around 15,000 words, close to a 30-page report. The same dollar buys about 75,000 words on Sol and roughly 1.5 million words on Luna. Reading is cheaper than writing on every tier, which is why summarizing a long document costs far less than drafting one.

The previous generation is not automatically cheaper. According to one pricing tracker, GPT 5.6 Sol sits at $4 input and $20 output during a promotional window that runs until November 21, which makes GPT-6 Sol the cheaper line item under the same tier name. GPT 5.6 Terra lists at $2 and $12, and GPT 5.6 Luna at $0.20 and $1.20, double the rate of GPT-6 Luna.
💡 Quick read: if your workload does not need agents or deep reasoning, the Sol tier is the new default to price first. Astra is the exception, not the baseline.
So when is Astra worth five times the price of Sol? When a failed run costs more than the price gap: an agent that burns an hour of engineer time, or a legal draft that would otherwise need three revision rounds. The break-even is simple. If Sol needs more than five attempts to match one good Astra result, Astra wins on price alone. If Sol gets there in two or three, you are overpaying.
Surcharges That Inflate the Bill
Long Prompts Double the Rate
Cross 272K input tokens and the whole request is repriced, not only the extra tokens. Input jumps from $10 to $20 per million, and output from $50 to $75.
Picture a 300K-token contract bundle with a 2K-token summary as the answer:
- At the long-context rate: $6.00 for input plus $0.15 for output, so $6.15.
- Trimmed to 270K tokens at the standard rate: $2.70 plus $0.10, so $2.80.
Thirty thousand extra tokens more than doubled the price of the same request.

💡 Tip: keep a token counter in your pipeline and split big documents just under the threshold instead of sending one giant prompt.
Fast Mode and Regional Uplift
Two more multipliers hide in the fine print:
- Fast mode costs about double the standard rate ($20 input, $100 output on Astra) in exchange for roughly 2.5x the speed. It is not available with EU data residency.
- Data-residency endpoints add 10% on top of whatever tier you chose.
Stack them and an innocent-looking Astra call can cost more than twice its list price. Decide on purpose, not by default.
Fast mode earns its premium when a human is waiting on the answer, like an interactive coding assistant or a live support chat. It is wasted on overnight jobs where nobody checks the clock. Those belong in the batch queue at half price.
Real Monthly Costs in Dollars
One Support Ticket, Priced Out
Take a support bot where each ticket uses 2,000 input tokens and 500 output tokens. On Astra that is $0.02 for the prompt plus $0.025 for the answer, so $0.045 per ticket. Small number, until you multiply it.

The Same Workload on Three Tiers
Here is that exact ticket at 100,000 tickets a month:
| Model | Per ticket | 100,000 tickets | With Batch discount |
|---|
| GPT-6 Astra | $0.045 | $4,500 | $2,250 |
| GPT-6 Sol | $0.009 | $900 | $450 |
| GPT-6 Luna | $0.00045 | $45 | $22.50 |
Caching adds another lever. If 1,500 of those 2,000 input tokens are the same system prompt and examples every time, Astra drops to $0.0315 per ticket, or $3,150 a month. That saves $1,350 without touching quality.
💡 Rule of thumb: on Astra, a 500-token answer ($0.025) costs more than a 2,000-token prompt ($0.02). Output length decides the invoice.
Support tickets are only one shape of workload. Here are three more, priced at standard rates with no discounts, so you can find the one closest to your own project:
| Workload | Tokens per request | Requests a month | Luna | Sol | Astra |
|---|
| Solo developer chat app | 1,000 in, 400 out | 5,000 | $1.50 | $30 | $150 |
| Content team drafts | 3,000 in, 2,500 out | 300 | $0.47 | $9.30 | $46.50 |
| Coding agent runs | 200,000 in, 30,000 out | 440 | $15.40 | $308 | $1,540 |
The table prices tokens, not quality. A Luna agent will not match an Astra agent, so read the last row as a ceiling and a floor, not a recommendation. Notice how the coding agent dwarfs everything else: only 440 runs, yet $1,540 on Astra. Caching 150K of each 200K prompt brings that down to about $946.
Who Can Call the API
Where the API Starts
Astra is not available on the API free tier. You need at least a Tier 1 developer account, which means a paid account with billing set up. OpenAI had not published detailed rate limits or throughput guarantees for Astra at launch, so check the limits page in your dashboard once access is granted.

ChatGPT Plans Are Separate
A ChatGPT subscription does not give you API credits. The two are billed separately, and the right choice depends on how you work. If you mostly chat, a Plus plan at $20 is usually simpler than paying per token. If you are building a product, you need the API no matter which plan you hold, because customers cannot use your personal subscription. Inside the app, Astra access looks like this:
| Plan | Monthly price | Astra access |
|---|
| Free | $0 | No |
| Go | $8 | No |
| Plus | $20 | Yes |
| Pro | From $100 | Yes, with expanded limits |
| Enterprise | Per seat | Yes, switched off by default |
Luna had not reached free ChatGPT users yet when this was written, so even the cheapest tier is not a free lunch inside the app.
Before your first API call, run through this short list:
- Add a billing method and a small prepaid balance.
- Set a monthly spend cap in the dashboard before you write any code.
- Create a project-scoped API credential and store it in an environment variable, never in the repository.
- Build in development on Luna, then switch the model name once the pipeline works.
- Log input, cached and output token counts for every request so the bill never surprises you.
Free Options That Actually Work
Why Astra Has No Free Tier
No trial allocation has been published for Astra, and at $50 per million output tokens that is not a surprise. A single long agentic run would burn through any realistic free credit within minutes.

Where Free Usage Still Exists
You still have real routes to test ideas without a big bill:
- Luna as the cheap entry point. Five dollars of credit buys about 50 million input tokens or 10 million output tokens. That is thousands of short conversations.
- Open-weight models. GPT OSS 120B and GPT OSS 20B are OpenAI's open-weight releases. If you own the hardware, running them yourself costs no per-token fee.
- Rival models in the browser. Try Claude Sonnet 5, Gemini 3.5 Flash or GPT 5.4 with your real prompts before you decide where to spend.
A practical free workflow looks like this. Prototype the prompt in a browser until the answers are good, lock it, then count the tokens of a typical request. Price it yourself with one formula: (input tokens x input rate + output tokens x output rate) / 1,000,000. Run that for Luna, Sol and Astra, and you will know your real monthly range before a single paid call goes out.
Try GPT 5.6 Sol on PicassoIA
One honest note: GPT-6 itself is not on the PicassoIA model list yet. The closest relatives are the GPT 5.6 family: GPT 5.6 Sol, GPT 5.6 Terra and GPT 5.6 Luna. Sol is free to try online with no setup, which makes it a cheap rehearsal before you commit API budget.
Step by Step on PicassoIA
- Open the GPT 5.6 Sol page.
- Paste a real task from your product into Prompt, not a "hello".
- Add a System Prompt to set the role and tone.
- Pick a Reasoning effort from none to xhigh. Start at none and raise it only when answers feel shallow.
- Set Verbosity to low, medium or high.
- Raise Max Completion Tokens if you choose a high reasoning level, otherwise thinking can use the whole budget and leave an empty reply.
- Upload a screenshot or diagram in Image Input if the task needs one.
- Run it, then repeat the same prompt on GPT 5.6 Luna and compare quality against price.

Settings That Move Your Cost
Reasoning tokens are billed as output tokens on reasoning models, so the dials on that page map straight onto an invoice:
| Setting | What it does | Cost effect |
|---|
| Reasoning effort: none or low | Fast replies, little thinking | Lowest |
| Reasoning effort: high or xhigh | Deeper, slower answers | More output tokens |
| Verbosity: low | Short, direct answers | Fewer output tokens |
| Max Completion Tokens | Hard cap on reply length | Upper limit per request |
Ways to Cut the Bill

Route by Difficulty
Send easy work to Luna, middle work to Sol, and only the hardest 5% to Astra. With a made-up split of 80% Luna, 15% Sol and 5% Astra, the 100,000-ticket workload costs about $396 instead of $4,500.
Cache and Batch
Cached input is 90% cheaper on Astra and Luna, and 95% cheaper on Sol. Batch and Flex processing halve the standard rate for anything that can wait.

Cap the Output
Set a maximum token limit and ask for short formats such as bullets or compact JSON. With output priced at five times input, every sentence you skip is money kept.
Three mistakes that inflate bills:
- Resending the whole chat every turn. A 10-turn conversation that adds 300 tokens per turn sends 16,500 input tokens in total, not 3,000. Summarize old turns or cache the shared prefix.
- Leaving reasoning on high for easy jobs. Thinking tokens bill as output, the expensive side of the rate card.
- Testing on the flagship. Build and debug on Luna, then switch the model name for the final run.
💡 Checklist: route by difficulty, cache the repeated prefix, batch whatever can wait, cap the reply length, and re-check prices monthly because tiers and promotions change.
Your Turn to Create
Text is only half of a product budget. Most launches also need pictures: hero shots, thumbnails, product mockups. Try creating your own images with Picasso IA and see how far a prompt can go. Seedream 4.5 is great for sharp 4K detail, P-Image is quick for drafts, and GPT Image 2 follows long prompts closely.
Pick a scene from your own project, write three prompts with different camera angles, and run them side by side. Browse the full catalog at picassoia.com/en/all-models and start experimenting today.