Large Language ModelsGenerate imagesGenerate videos
Claude Agent SDK Pricing: Subscription, Credits and Costs
Claude Agent SDK usage can draw from plan limits, from the monthly API credits added to Max and Team plans on October 7, or from pay-per-token billing. This article sets out every price, the credit amounts, cache savings and a worked cost for one agent run.
The Claude Agent SDK has no price of its own. The bill comes from the model behind it, and as of October 8, 2026 that bill can reach you in three different ways: through your subscription limits, through the new monthly API credits on Max and Team plans, or through plain pay-per-token billing. Pick the wrong route and you either hit a wall in the middle of a run or pay more than you had to. This article sets out each option with the numbers from Anthropic's own pricing and help pages, then works out what a single agent run costs on three different models.
How the SDK Gets Billed
The Agent SDK is a Python and TypeScript library that puts Claude Code's agent loop, built-in tools, sessions and hooks inside your own program. Anthropic's documentation describes it as a library that runs the Claude Code binary in a process you operate. Its pricing page lists no SDK fee, so what you pay for is model usage: input tokens, output tokens and cache activity.
Three Ways to Pay
Where those tokens get charged depends on how the SDK signs in:
Subscription limits. The SDK uses your Claude login and draws from the same allowance as your chats and interactive coding. There is no dollar meter, only limits that reset.
Monthly API credits. Max and Team plans now include a dollar amount each billing cycle that the SDK spends at normal API prices.
Pay-per-token API. You buy credits in the Claude Console, create an API credential, and every token is billed at list price.
Route
Billed in
Best for
Watch out for
Subscription limits
Plan allowance
Personal scripts and nightly jobs
Sharing one pool with your chats
Monthly credits
Dollars that expire
Prototypes and side projects
No rollover at the end of the cycle
Pay-per-token
Dollars you buy
Products with real users
No cap unless you set one
๐ก Cloud providers are a fourth route. The SDK can run on Amazon Bedrock, Google Cloud or Microsoft Foundry. Those bills come from the cloud provider, and the Max and Team credits do not apply there.
Why Old Articles Disagree
Search results are full of conflicting claims because the rules moved three times this year:
May 13: Anthropic announced that from June 15 the Agent SDK, claude -p, Claude Code GitHub Actions and third-party apps built on the SDK would move to a separate monthly credit billed at full API rates: $20 on Pro, $100 on Max 5x and $200 on Max 20x.
June 15: Anthropic paused that change on the day it was due. Agent SDK usage kept drawing from subscription limits, and there was no credit to claim.
October 7: Max and Team plans gained monthly API credits that include the SDK. Pro did not get one.
If an article says Pro subscribers receive a $20 SDK credit, it is describing the plan that was paused.
Plan Limits or Monthly Credits
Running on Plan Limits
Anthropic's help article for the SDK says usage that goes through your Claude login keeps counting against your subscription limits, the same as claude -p. That is the right route for personal automation: a script on your laptop, a nightly job on your own machine. A subscription costs $20 a month for Pro, $100 for Max 5x and $200 for Max 20x, and you pay nothing per token.
The catch is sharing. The SDK draws from the same pool as every chat and every interactive coding session you run, so a long agent loop can leave you short of allowance for the rest of your day.
The Max and Team Credits
On October 7, 2026, Anthropic added monthly API credits to Max and Team plans. They apply to the Claude API, Claude Managed Agents, the Claude Agent SDK and the Console playground.
Plan
Monthly credit
Max 5x
$100
Max 20x
$200
Team, Standard seat
$20 per seat
Team, Premium seat
$100 per seat
Free, Pro and Enterprise
Not eligible
A credit is spent at the API prices in the next section, so $100 means $100 of tokens. Interactive Claude Code and extra usage inside the Claude apps are excluded, and the credits leave your plan limits unchanged.
Claiming and Spending Credits
Credits are not automatic. Follow these steps:
On claude.ai, open Settings > Billing on a Max plan, or Organization settings > Billing on a Team plan.
In the API credits section, pick the Claude Console organization that should receive them and accept the program terms.
Find the balance in the Console under Settings > Billing, in the Promotional credits section.
Create an API credential in that organization and point the SDK at it.
A few rules decide how far the money goes:
New subscribers can claim after 7 days on an eligible plan.
Each plan links to one Console organization, and you cannot change the link yourself afterwards, so choose the organization you will build in.
Included credits are spent before any credits you bought, and no payment method is needed on Claude Platform.
Unused credits expire at the end of each billing cycle (monthly on annual plans) and never roll over.
When the balance reaches zero and nothing else is available, requests stop until the next batch arrives. Usage is never charged to your Claude plan. The error reads Your credit balance is too low to access the Anthropic API.
Token Prices Model by Model
The Current Price Table
These are list prices per million tokens from Anthropic's pricing page, checked on October 8, 2026.
Claude Haiku 5.5 costs more when a prompt passes 100,000 tokens: $0.50 for input and $2.50 for output.
๐ก Claude Sonnet 5 launched with $2 and $10 as introductory pricing through August 31, 2026. Anthropic's page now says that is the standard price and the planned September 1 rise to $3 and $15 will not happen. Any page still listing that rise is out of date.
Models from version 4.7 onward use a newer tokenizer that produces roughly 30% more tokens for the same text than the 4.6 generation and earlier. Compare the cost of a finished task, not only the price per token.
Cache Reads Cost Less
The SDK uses prompt caching on its own. A 5-minute cache write costs 1.25 times the base input price, a 1-hour write costs 2 times, and a cache read costs 0.1 times. Opus 5.5 and Sonnet 5.5 read at 0.05 times, and Fable 5.1 at 0.025 times. An agent loop resends a long, stable prefix on every step, so most of its input tokens end up as cheap reads.
With an API credential the cache lives for 5 minutes by default. If your runs sit further apart than that, set the ENABLE_PROMPT_CACHING_1H environment variable. The 1-hour option costs more to write and pays off when later runs read the same prefix.
Hidden Multipliers to Watch
Tokens are not the whole bill. Check these before you budget:
Web search: $10 per 1,000 searches on the Claude API, on top of tokens. Web fetch adds nothing beyond the tokens it brings in.
US-only inference: a 1.1x multiplier on every token category when you set inference_geo: "us" on Claude 4.6 and later models.
Regional cloud endpoints: on Bedrock and Google Cloud, regional and multi-region endpoints carry a 10% premium over global ones.
Fast mode: Opus 5.5 runs at $8 input and $40 output per million tokens in research preview, double the standard rate.
Batch API: 50% off input and output, but it suits offline jobs sent to the API directly, not an interactive SDK loop.
What One Agent Run Costs
The Setup
Picture a bug-fix run: the agent reads a few files, edits code, runs the tests and repeats for about 30 steps. Assume 600,000 input tokens across those steps, made up of 20,000 uncached tokens, 40,000 written to the 5-minute cache and 540,000 read back from it, plus 30,000 output tokens. These are illustrative numbers, not a measurement, but they have the shape of a cached agent loop.
The Result
On Sonnet 5.5 the arithmetic is:
20,000 uncached input tokens ร $2 per million = $0.04
40,000 cache-write tokens ร $2.50 per million = $0.10
540,000 cache-read tokens ร $0.10 per million = $0.054
30,000 output tokens ร $10 per million = $0.30
That adds up to $0.494. Run the same shape on three models and one $100 credit buys very different amounts of work:
Model
Cost per run
Runs from a $100 credit
Claude Haiku 5.5
about $0.03
more than 3,600
Claude Sonnet 5.5
about $0.49
about 200
Claude Opus 5.5
about $0.99
about 100
Two lessons stand out. First, caching saves roughly two thirds: the same Sonnet 5.5 run with no cache would cost $1.50. Second, output is the expensive part. Those 30,000 output tokens are about 5% of the volume and 61% of the cost. A Max 20x credit simply doubles every number in the last column.
Tracking Spend Inside the SDK
Reading the Cost Estimate
Every query() call ends with a result message that carries total_cost_usd, the estimated cost of that call. The modelUsage field breaks it down per model, including cache reads and writes. Three details trip people up:
total_cost_usd and modelUsage include subagent work, but usage leaves it out, so usage undercounts as soon as subagents run.
Parallel tool calls produce several assistant messages that share one ID. Count each ID once.
The output_tokens on each step is a placeholder. Read the real output count from the result message.
๐ก total_cost_usd is a client-side estimate built from a price table bundled with the SDK. Anthropic says to use the Usage and Cost API or the Usage page in the Console for authoritative billing, and never to bill your own customers from the estimate.
import { query } from "@anthropic-ai/claude-agent-sdk";
for await (const message of query({
prompt: "Fix the failing test in src/auth.ts",
options: { maxBudgetUsd: 2 },
})) {
if (message.type === "result") {
console.log(message.subtype, message.total_cost_usd);
}
}
Capping a Run
The maxBudgetUsd option in TypeScript, or max_budget_usd in Python, stops a call once its own spend crosses the limit. It counts only that call: spend restored from a resumed session does not count against it, and a /clear starts the budget over. The result then arrives with the subtype error_max_budget_usd.
For a ceiling across a whole project, give it its own Console workspace and set that workspace's spend limit. Every credential and workspace in a linked organization draws from the same credit balance, so a workspace limit stops one noisy job from eating the month's allowance.
Teams, Products and Managed Agents
How Team Pools Work
Team credits are $20 per Standard seat and $100 per Premium seat, pooled into one balance capped at $500 a month. A team with three Standard seats and two Premium seats receives $260 a month, and adding a Premium seat lifts the next month's credit to $360. The amount is set by the seats on the plan at the start of each billing month. Linking an organization also moves it to at least the Start usage tier, or the Build tier when the monthly credit is above $200.
Building for Customers
If you ship a product, plan on API billing. Anthropic does not allow third-party developers to offer claude.ai login or its rate limits in their products, including agents built on the SDK, unless Anthropic approved it in advance. The monthly credits are meant for building and running your own applications and agents, so keep customer traffic on pay-per-token billing, where you control spend limits per workspace.
Managed Agents Compared
Claude Managed Agents is the hosted alternative: Anthropic runs the agent loop in a cloud sandbox it manages. It bills the same token rates plus $0.08 per session-hour, counted only while the session is running. Idle time does not count. A one-hour session on Sonnet 5.5 with 50,000 input tokens and 15,000 output tokens costs $0.10 + $0.15 + $0.08, or $0.33. With the SDK there is no runtime line, but the machine it runs on is yours to pay for. The Max and Team credits work on both.
Before you spend credits tuning an agent, draft and test its system prompt in a chat model. Claude Sonnet 5 is on PicassoIA, built for multi-step coding and tool-use tasks, which makes it a handy place to try a prompt before wiring it into the SDK.
Open the Claude Sonnet 5 page in the Large Language Models collection.
Paste the task into Prompt, then put the role and rules into System Prompt.
Set Effort. The default, low, turns thinking off for the fastest and cheapest replies. Move up through medium, high, xhigh and max for a bug that touches several files.
Leave Max Tokens at 8,192 unless you expect a longer answer.
Attach a screenshot in Image if the task starts from an error screen. Max Image Resolution defaults to 0.5 megapixels, and an image counts as width ร height รท 750 input tokens.
Run it, adjust the wording, and copy the version that works into your SDK options.
๐ก The effort ladder also moves your SDK bill. Higher effort means more output tokens, and output is the costly side of every price table above.
Every photo in this article came from P Image, a text-to-image model on PicassoIA that returns a finished picture in under a second. It accepts 16:9, 1:1, 9:16 and custom sizes, offers prompt upsampling for short ideas, and takes a seed so a result you like can be reproduced exactly.
Try it now. Describe one scene in a single sentence, such as a desk, a window and late afternoon light. Run it, change a single detail, and run it again. Ten tries takes a few minutes, and by the end you will have a feel for how camera, light and texture words shape a photo.
PicassoIA also offers 91 text-to-image models and 87 text-to-video models, plus music, speech and language models, all in the full model list. Pick one, type a prompt and make something of your own today.