Claude has two price tags, and they do not measure the same thing. The API charges for every token you send and receive. A Claude subscription charges one flat monthly fee for a capped amount of chat and coding time. Pick the wrong one and you either pay for tokens you barely use, or you hit a usage wall in the middle of a deadline. This article puts real numbers on both sides, using Anthropic's published rates as of October 2026, so you can see where the per token meter wins and where a flat plan comes out ahead.

How Per Token Billing Works
Every API request is metered. You pay for the text going in, the text coming out, and a few optional extras, all measured in tokens. Prices are quoted per million tokens, written MTok, so $2 / MTok means $2 for every million tokens. There is no monthly fee, no seat count, and no cap: the bill follows the traffic.
What Counts as a Token
A token is a chunk of text, usually a short word or a piece of a longer one. Anthropic's rule of thumb is about 4 characters or 0.75 words of English per token, so a 2,500 word article is roughly 3,300 tokens. Code, non-English text, and heavy formatting usually take more.

One detail matters when you compare old and new rate cards. Claude 4.7 and later models use a newer tokenizer that produces about 30% more tokens for the same text than Sonnet 4.6 and earlier models. A lower price per token does not always mean a lower bill if your prompt now splits into more tokens.
Why Output Costs More
Output tokens cost five times as much as input tokens across the current lineup. Sonnet 5.5 charges $2 per million input tokens and $10 per million output tokens. Opus 5.5 charges $4 and $20. Generating text is the expensive step, while reading a prompt is cheap.
That ratio decides how you save money. Long documents sent in cost little. Long answers coming back cost a lot. A prompt asking for three bullet points is cheaper than one that invites a 2,000 word essay.
How Chat History Inflates Bills
The API is stateless, so each new turn resends the whole conversation and you pay for all of it as input again. Imagine a chat where you type 500 tokens and the model answers with 500 tokens, for ten turns. You wrote only 5,000 tokens of fresh text, yet the input meter adds up to 50,000 tokens across the ten calls, because turn 10 resends everything from turns 1 to 9.
On Sonnet 5.5 that conversation costs $0.10 for input plus $0.05 for the 5,000 output tokens, so $0.15 in total. Small on its own. Multiply it by a hundred conversations a day and the history becomes the biggest line on the invoice, which is exactly where caching helps.

Claude API Prices by Model
Current Per Million Token Rates
These are the rates from Anthropic's pricing page, in US dollars per million tokens. The last column is the price of a prompt cache read, which is the cheapest way to resend text you have already sent.

| Model | Input | Output | Cache read |
|---|
| Claude Haiku 5.5 (prompts up to 100K tokens) | $0.10 | $0.50 | $0.01 |
| Claude Haiku 4.5 | $1 | $5 | $0.10 |
| Claude Sonnet 5.5 | $2 | $10 | $0.10 |
| Claude Sonnet 4.6 | $3 | $15 | $0.30 |
| Claude Opus 5.5 | $4 | $20 | $0.20 |
| Claude Fable 5.1 | $10 | $50 | $0.25 |
Haiku 5.5 is priced by prompt length: prompts over 100,000 tokens pay $0.50 input and $2.50 output.
💡 Newer is not always pricier. Opus 5.5 at $4 / $20 undercuts the Opus 5 and Opus 4.8 generation at $5 / $25, and Sonnet 5.5 at $2 / $10 undercuts Claude Sonnet 4.6 at $3 / $15. Check the rate card before assuming the bigger number belongs to the newer model.
What One Request Really Costs
Take a typical request: 3,000 input tokens (a system prompt plus a pasted document) and 800 output tokens (a solid answer). Here is the math per request and per thousand requests, using the same token counts on every row to keep the comparison simple.
| Model | Cost per request | Cost per 1,000 requests |
|---|
| Claude Haiku 5.5 | $0.0007 | $0.70 |
| Claude Haiku 4.5 | $0.007 | $7.00 |
| Claude Sonnet 5.5 | $0.014 | $14.00 |
| Claude Sonnet 4.6 | $0.021 | $21.00 |
| Claude Opus 5.5 | $0.028 | $28.00 |
| Claude Fable 5.1 | $0.070 | $70.00 |
The spread between the cheapest and the priciest row is 100 times. Model choice moves the bill more than any plan decision does. Because newer models tokenize the same text into more tokens, add up to 30% to the 4.7 and later rows in real use.
Caching and Batch Discounts

Two discounts do most of the saving work, and they stack:
- Prompt caching. Writing a prompt prefix to the 5 minute cache costs 1.25x the base input price, and the 1 hour cache costs 2x. Every read after that costs 0.1x. On Opus 5.5 and Sonnet 5.5 a read costs just 0.05x, and on Fable 5.1 it costs 0.025x. The 5 minute cache pays for itself after a single read.
- Batch API. Asynchronous jobs get 50% off both input and output tokens. Sonnet 5.5 drops to $1 / $5 and Opus 5.5 drops to $2 / $10.
Caching helps less than people expect when answers are long, because output is not discounted. A coding session with 50,000 input tokens and 15,000 output tokens costs $0.50 on Opus 5.5. Serve 40,000 of those input tokens from cache and it falls to about $0.35, a 30% saving, because the $0.30 of output stays the same.
What the Subscriptions Cost
Free, Pro, and Max
A subscription swaps the meter for a ceiling. You pay a fixed price and Claude limits how much you can use inside rolling five hour windows, with the exact number of messages depending on how long and complex each one is.

| Plan | Price | What you get |
|---|
| Free | $0 | Chat on web, desktop and mobile, with web search, file creation, code execution and memory |
| Pro | $20 per month, or $17 per month billed annually | More usage than Free, plus Claude Code, Projects and additional models |
| Max 5x | $100 per month | Five times the usage of Pro, higher output limits and early access to new features |
| Max 20x | $200 per month | Twenty times the usage of Pro |
Team and Enterprise Seats

Team standard seats cost $20 per seat per month on annual billing, or $25 billed monthly. Premium seats cost $100 per seat per month on annual billing ($125 monthly) and give five times the usage of a standard seat. Enterprise starts at $20 per seat per month on annual billing, with usage charged at API rates, and adds SCIM, audit logs, custom data retention and role based access.
Why the API Is Billed Separately
A Claude subscription does not include API credits. Anthropic bills the API separately from its plans, so a $20 Pro subscription will not pay for an app you ship to customers. That needs API access with its own billing. Individual and Team subscribers can buy extra usage credits at standard API rates once plan limits run out, which works as a safety valve for busy weeks.
Where the Break Even Sits
The simplest test is to divide the plan price by your cost per request. At $20 for Pro and the 3,000 input and 800 output request from earlier, the break even looks like this:

| Model | Cost per request | Requests to reach $20 | Per day over 30 days |
|---|
| Claude Haiku 5.5 | $0.0007 | 28,571 | about 952 |
| Claude Sonnet 5.5 | $0.014 | 1,429 | about 48 |
| Claude Opus 5.5 | $0.028 | 714 | about 24 |
| Claude Fable 5.1 | $0.070 | 286 | about 10 |
This compares cost only. A plan also bundles a chat interface, Projects and app connections, while the API gives you raw access and nothing else. You need an app or tool that calls it.
Light Chat Users
If you ask a handful of questions a day, pay per token is almost always cheaper. Thirty Sonnet 5.5 requests a day for a month comes to about $12.60, and the same habit on Haiku 5.5 costs pennies. The flat plan only wins on price once your volume passes the break even, or when you want the interface and extras more than the lowest bill.
Heavy Coding Agents

Agentic coding flips the answer. Use the token counts from Anthropic's own one hour session example: 50,000 input tokens and 15,000 output tokens. On Opus 5.5 that is $0.50 per session with no caching. Twenty sessions in a working day is $10, or about $220 over 22 working days, more than the $200 Max 20x plan.
Caching narrows the gap. At about $0.35 per session, the same twenty sessions a day cost roughly $153 a month. Either way, a power user lands well above the $100 Max 5x plan, which is the workload the flat tiers were built for. The trade-off is that plan usage caps still apply.
Hidden Costs Worth Budgeting
Token math is not the whole invoice. These extras are easy to miss:
- Web search: $10 per 1,000 searches, plus the tokens the results add to your prompt. Web fetch has no extra charge beyond tokens.
- US only inference: Setting
inference_geo to "us" applies a 1.1x multiplier to every token category on Claude 4.6 and later models.
- Fast mode: Opus 5.5 in fast mode costs $8 input and $40 output per MTok, double the standard rate.
- Managed Agents: Tokens are billed at normal rates, plus $0.08 per session hour while the session is running.
- Tool overhead: Any request with tools adds a built-in tool prompt, about 286 tokens on the 5.5 generation, before your own tool definitions.
- Tokenizer shift: Up to 30% more tokens for the same text on Claude 4.7 and later models.
💡 Budget tip: Track cache_read_input_tokens in API responses. If it stays at zero across repeated requests, something in your prompt prefix changes every call and you are paying full price for text you could be reading from cache.
Use Claude Sonnet 5 on PicassoIA
If you want to try Claude without opening an API account or watching a token meter, Claude Sonnet 5 is available on Picasso IA. It is the previous Sonnet generation, priced on the API at the same $2 / $10 as Sonnet 5.5, and its page lists multi-step coding, tool use and image input among its features.
Step by Step Setup
- Open the model page. Go to Claude Sonnet 5 on Picasso IA.
- Write your prompt. The Prompt field is the only required input. Paste the stack trace, the document or the question.
- Set the effort level. The default is
low, which turns thinking off for the fastest and cheapest replies. Move to medium or high for harder coding work, and xhigh or max for bugs that touch several files.
- Cap the output. Max Tokens defaults to 8,192. Lower it when you want short answers, since output is the expensive side of any Claude price list.
- Add a system prompt. This is optional. Fix a role or coding style once and reuse it across a project.
- Attach an image. This is optional too. A screenshot of an error or a wireframe sketch works as context.
- Run it and refine. Read the answer, tighten the prompt, and run it again.
Parameter Tips
| Parameter | Default | Change it when |
|---|
| Effort | low | The answer misses details, or the bug spans several files |
| Max Tokens | 8192 | You want shorter replies or longer blocks of code |
| System Prompt | empty | You need a fixed role, tone or coding style |
| Max Image Resolution | 0.5 megapixels | Small text in a screenshot gets lost after scaling |
💡 Run the same prompt on Claude Sonnet 4.6 and Claude 4.5 Haiku to see the quality gap before you pick a model for your API budget. For the hardest coding problems, try Claude Fable 5.
Picking the Right Setup
Match the option to the workload:
- Chat, writing and research for yourself: Start on Free. Move to Pro at $20 when you hit limits.
- Daily coding agents: Max 5x or Max 20x, because heavy sessions cost more on the meter than the flat price.
- A product or automation: The API, using Haiku 5.5 or Sonnet 5.5, with prompt caching and batch jobs wherever latency allows.
- A team: Team seats for people, plus the API for anything you ship.
- Both: Many people hold a plan for personal use and API access for automation, since the two are billed separately.
The cheapest choice is the one that matches how you actually work, so spend a week counting requests before you commit.
Your Next Move on Picasso IA
Pricing math is easier once you have something to build, and Picasso IA is a quick place to try both text and images. Open Claude Sonnet 5 to draft prompts, then turn those prompts into pictures with an image model such as GPT Image 2, Flux 2 Pro or P-Image, the model behind the photos in this article.
Pick a scene, describe the light and the lens, and press generate. Then change one detail and run it again. Experimenting costs a few minutes, and it shows you how much a good prompt changes the result. Browse the full catalog at picassoia.com/en/all-models and create your first image today.