Anthropic shipped Claude Opus 5.5 on September 22, 2026, five months after Claude Opus 4.7 arrived on April 16. The surprise is the price tag: it went down. Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, while Opus 4.7 still sits at $5 and $25. A cheaper model with a newer knowledge cutoff turns the upgrade question into a practical one.
Upgrades still carry costs that never show up on an invoice. Some API calls that ran fine on 4.7 now return a 400 error on 5.5, and your evals deserve a rerun before you change a model ID in production. This comparison sticks to Anthropic's own pricing and model pages plus third-party benchmark indexes, and it says plainly where the evidence is thin.

💡 Short version: if you pay API bills every month, Opus 5.5 is a 20% to 31% discount on the same workloads, with a few breaking changes to check first. If you use Claude through a hosted chat page, there is no bill to cut and no code to migrate, so the decision is simpler.
The Quick Verdict
Here is the answer before the details. Opus 5.5 is the better model on paper and the cheaper one on the invoice. Anthropic's own model overview now tells developers to start with Opus 5.5 for most workloads and to reserve Fable 5.1 for demanding reasoning and long-horizon agent work. The only solid reason to hold back is compatibility.

Switch Now If This Is You
- You run agent loops with heavy prompt caching. Cache reads fall from $0.50 to $0.20 per million tokens, the biggest saving in the whole comparison.
- You send tool-heavy requests. The built-in tool-use system prompt shrinks from 675 tokens to 286 tokens per request.
- You want fast mode. Opus 5.5 supports it at premium pricing, while Opus 4.7 rejects the request outright.
- You need fresher knowledge. Opus 5.5 lists a reliable knowledge cutoff of June 2026.
- You are starting a new project. Nothing needs migrating, so there is little reason to begin on the older model.
Wait If This Is You
Hold off if your code relies on any of the following, because Opus 5.5 returns a 400 error for each one:
- Turning thinking off with
thinking: {"type": "disabled"}, or setting a manual budget_tokens.
- Forcing a specific tool with
tool_choice.
- Using the
computer_20251124 computer-use tool.
- Editing thinking blocks between turns.
Teams in audited or regulated workflows who cannot rerun their evals this quarter should also stay put. Opus 4.7 remains available as a legacy model, and a calm migration beats a rushed one.
Opus 4.7 and 5.5 Side by Side
Most of the specs line up, which is what makes the differences easy to spot. Both models offer a 1M token context window and 128K tokens of maximum output. Both support up to 300K output tokens on the Batch API with a beta header. The money and the tooling are where they split.
| Spec | Opus 4.7 | Opus 5.5 |
|---|
| Release date | April 16, 2026 | September 22, 2026 |
| Input price per 1M tokens | $5 | $4 |
| Output price per 1M tokens | $25 | $20 |
| Cache read per 1M tokens | $0.50 | $0.20 |
| 5 minute cache write | $6.25 | $5 |
| 1 hour cache write | $10 | $8 |
| Batch input / output | $2.50 / $12.50 | $2 / $10 |
| Context window | 1M tokens | 1M tokens |
| Max output | 128K tokens | 128K tokens |
| Fast mode | Not available | $8 input / $40 output |
| Tool-use system prompt | 675 tokens | 286 tokens |

Where Each Version Sits
Opus 4.7 and 5.5 are not neighbors. Opus 4.8 and Opus 5 shipped in between, and both kept the $5 and $25 price. Moving from 4.7 to 5.5 therefore bundles several generations of change into one migration, which is why the breaking changes later in this article matter.
Anthropic's current lineup has four tiers: Fable 5.1 at $10 and $50, Opus 5.5 at $4 and $20, Sonnet 5.5 at $2 and $10, and Haiku 5.5 starting at $0.10 and $0.50. Opus 5.5 is also the cheapest Opus since the $5 and $25 generation began. If you want to try the neighbors in the family today, PicassoIA hosts Claude Opus 4.7, Claude Opus 4.6, Claude Sonnet 5 and Claude Fable 5.
What Changed in Opus 5.5
Four changes carry most of the weight. Three of them save money. One of them can break a build.
Lower Prices, Cheaper Cache Reads
The 20% base price cut is the headline, but the cache change matters more for real workloads. Cache reads on Opus 5.5 are billed at 5% of the input price instead of 10%, which is $0.20 per million tokens against $0.50 on Opus 4.7. That is a 60% drop. Agents that re-read a long system prompt and a growing transcript on every step spend most of their input budget on cache reads, so they feel the change first.

Adaptive Thinking, Medium Default
Opus 5.5 runs adaptive thinking that is always on, steered by an effort setting whose default is medium. The model decides how long to think, and you nudge it with effort instead of a manual token budget. Opus 4.7 brought in the xhigh effort level between high and max, so effort controls are not new. What is new is that 5.5 removes the option to switch thinking off.
Medium as the default is a cost decision as much as a quality decision. Launch reporting says that Opus 5.5 at default effort beat Opus 5 at max effort on Terminal-Bench 4.0 for about a fifth of the cost per attempt. That comparison is against Opus 5, not 4.7, so treat it as a direction rather than a promise for your workload.
In practice, set effort explicitly instead of trusting the default. Start at medium, move to high only for tasks where a wrong answer is expensive, such as security reviews or multi-file refactors, and log the token count at each level so you can see what the extra thinking costs.

Fast Mode Arrives
Fast mode is a research-preview option that trades money for speed. On Opus 5.5 it costs $8 per million input tokens and $40 per million output tokens, exactly double the standard rate. Opus 4.7 does not support it at all, and a request with speed: "fast" returns an error. Fast mode is available on the Claude API only, not on partner cloud platforms, and it cannot be combined with the Batch API.
Leaner Tool Overhead
Every request that includes tools carries a hidden system prompt. On Opus 4.7 it costs 675 tokens with automatic tool choice. On Opus 5.5 it costs 286. Across 1,000 tool-enabled requests, that is 675,000 tokens against 286,000, or about $3.38 against $1.14 at each model's input price. It is small money per call, but a busy agent makes a lot of calls.
Benchmarks: Real Gains or Noise?
Benchmarks are where upgrade posts get slippery, so here is the honest version. Anthropic's launch numbers for Opus 5.5 compare it with Opus 5, and the figures published for Opus 4.7 in April come from older versions of the test suites. We could not verify a clean, same-version head-to-head. The best neutral evidence is a third-party index that scores both models under one methodology.

Third-Party Index Scores
The LLM Stats comparison page scores both models like this:
| Index | Opus 4.7 | Opus 5.5 | Gap |
|---|
| Overall score | 43.5 | 60.4 | +16.9 |
| Reasoning | 45.2 | 59.0 | +13.8 |
| Coding | 35.2 | 51.0 | +15.8 |
| Vision | 32.8 | 41.3 | +8.5 |
| Tool use | 24.7 | 32.8 | +8.1 |
| Multimodal | 31.4 | 40.9 | +9.5 |
| Math | 39.0 | 41.8 | +2.8 |
That page ranks Opus 5.5 first overall and Opus 4.7 forty-eighth. The gap is widest in coding and reasoning, while math barely moves. Aggregator numbers shift as new results come in, so read them as a shape, not a decimal.
What Launch Numbers Measure
At launch, Opus 4.7 was reported at 64.3% on SWE-bench Pro, up from 53.4% on Opus 4.6, and at 87.6% on SWE-bench Verified. It also raised image input from 1.15 to 3.75 megapixels. Opus 5.5, at max effort, was reported at 66.4% on Terminal-Bench 4.0, 81.8% on OSWorld 2.0, 57.8% on CursorBench 4.0 and 1,846 Elo on GDPval-AA v2.1.
Those are different suites, different versions and different effort settings. Do not subtract one from the other.
💡 Test it yourself: pick 20 real tasks from your own backlog, run them on both models, and score the results blind. Whichever model wins on your tasks is the better model for you, whatever a leaderboard says.
The Real Cost of Switching
Price per token is easy to compare. Price per finished task is what you actually pay. Both models use the same tokenizer, so identical text produces identical token counts. Output length can still differ, because adaptive thinking decides how much to write, so measure a sample before you trust any projection.
Four Sample Workloads
These figures use the published rates and assume equal token counts. They leave out cache write costs to keep the math readable.
| Workload | Opus 4.7 | Opus 5.5 | Saving |
|---|
| 10M input, 2M output, no caching | $100 | $80 | 20% |
| 50M input (40M cache reads), 5M output | $195 | $148 | 24% |
| Batch: 20M input, 4M output | $100 | $80 | 20% |
| Agent loop: 200M input (190M cache reads), 8M output | $345 | $238 | 31% |
The pattern is clear. Plain traffic saves the headline 20%. The more of your input arrives as cache reads, the bigger the discount grows, because the cache rate fell by 60% while the base rate fell by 20%.
To read the table for your own traffic, pull last month's usage report, split the input tokens into fresh and cache reads, and apply the same formula: fresh input times the input rate, plus cache reads times the cache rate, plus output times the output rate. If you never turned caching on, do that first. A cache read on Opus 4.7 already costs 90% less than a fresh input token, so the savings from caching can dwarf the savings from switching models.

Batch Runs and Fast Mode
The Batch API takes 50% off both models, so batch prices land at $2 and $10 on Opus 5.5 against $2.50 and $12.50 on Opus 4.7. That keeps the 20% gap for overnight jobs. Fast mode works the other way around. At $8 and $40 it costs 60% more than Opus 4.7's standard rate, so use it only on the slice of traffic where a person is waiting on the answer.
Migration Pitfalls to Check First
Most upgrades fail on small things nobody tested. Run this checklist against a staging copy of your app before you change a model ID.

Calls That Now Fail
- Thinking disabled or a manual budget.
thinking: {"type": "disabled"} and budget_tokens both return a 400 invalid_request_error.
- Forced tool choice. A
tool_choice that forces any tool or a named tool returns a 400.
- Edited thinking blocks. Thinking blocks are bound to the conversation state, so edits return a 400.
- The old computer-use tool.
computer_20251124 is rejected. Switch to computer_toolset_20260801.
Each one is a one-line fix, but each one will take down a request path if you meet it for the first time in production.
Tokenizer and Thinking Blocks
The good news first: there is no tokenizer change. The current tokenizer arrived with Opus 4.7 and produces roughly 30% more tokens than pre-4.7 models for the same text. Opus 5.5 uses the same one. If you already moved from 4.6 to 4.7, your token counts hold steady. If you are still on 4.6, budget for that 30% on top of everything above.
The quieter issue is behavioral. Text the model writes between tool calls now arrives inside thinking blocks instead of text blocks, so a UI that streams progress notes from text blocks can go blank until you adjust the thinking.display setting. Thinking blocks are also readable only by Opus 5.5, Fable 5.1 and Mythos 5.1, so a mid-conversation fallback to a different model needs its own plan.
A safe rollout takes four steps. First, replay a day of real traffic against Opus 5.5 in staging and log every 400 error. Second, compare cost per finished task, not cost per token, across at least 200 requests. Third, send 5% of live traffic to the new model for a week while Opus 4.7 handles the rest. Fourth, raise the share only if error rates and user feedback hold. If anything regresses, the model ID is the only line you need to change back.
How to Use Opus 4.7 on PicassoIA
You can test the older side of this comparison without writing a line of code. Claude Opus 4.7 runs straight from its model page, with no API credentials, installs or configuration. At the time of writing, the catalog lists Opus 4.7, Opus 4.6, Sonnet 5 and Fable 5 from Anthropic, but not Opus 5.5, so this section walks through the 4.7 half of your test.

Step by Step
- Open the Claude Opus 4.7 page in the large language models collection.
- Type your task in the Prompt field. Paste code, a contract clause or a long report.
- Add a System Prompt if you want a fixed role, such as "You are a senior reviewer. Answer in short bullet points."
- Optionally upload an Image, such as a screenshot, chart or diagram. The model reads it alongside your text. Max Image Resolution defaults to 0.5 megapixels, so raise it for dense screenshots.
- Set Max Tokens. The default is 8,192, and you can lower it when you want shorter answers.
- Run the prompt, read the answer, then continue the conversation in the same thread.
Prompts That Test the Gap
Build a small test set and reuse it on every model you compare. Three prompts span most of the ground:
- A bug hunt. Paste a real file with a known defect and ask for a line-by-line review.
- A long summary. Paste a 5,000 word document and ask for the ten points a busy manager needs.
- A screenshot question. Upload a dense dashboard and ask what changed between two charts.
Write down the time, the cost and a 1 to 5 score for each. Ten minutes of this beats an hour of reading benchmark threads.
Make Your Own Images Next
Language models write the plan. Picasso IA turns it into pictures. Ask Opus for five photorealistic image prompts with lens, lighting and composition spelled out, then paste them into a text-to-image model and compare the results side by side. Good starting points are Seedream 4.5, Flux 2 Pro and PicassoIA Image.
Once a still frame looks right, you can set it in motion with Seedance 2.0 or Kling v3 Video. Pick a prompt, run it, and keep the version you like best. Your first image is a few seconds away, so try Picasso IA today and see what your own ideas look like.