Claude Code pricing depends on the billing path: API use is charged by token consumption, while Pro and Max subscribers have usage included in their subscription, according to Claude Code’s cost guide. For an API task, estimate the selected model, input and cache volume, output, and thinking tokens before work begins; then control the task by matching model choice to complexity and accounting for context, subagents, and cache behavior.
Set up the estimate
Start with the billing path you will actually use, then classify the task by its main cost drivers.
- API: Claude Code charges by API token consumption. Use the current model rates and estimate each relevant token category separately, as explained in the Claude Code cost documentation.
- Free: The Claude pricing page lists Claude Code as not included in the Free plan.
- Pro: The same pricing page lists Claude Code as included in Pro. Pro is listed at $17 per month with the annual subscription discount and $200 billed up front, or $20 if billed monthly.
- Max: The pricing page lists Claude Code as included in Max 5x and Max 20x. Max is listed from $100 per month.
An API estimate needs the published list rates for the selected model. The Claude pricing page lists these rates per MTok:
| Model | Cache read | Cache write | Input | Output |
|---|---|---|---|---|
| Opus 5.5 | $0.20 | $5 | $4 | $20 |
| Sonnet 5.5 | $0.20 | $2.50 | $2 | $10 |
| Haiku 4.5 | $0.10 | $1.25 | $1 | $5 |
Keep cache reads, cache writes, uncached input, and output visible in the estimate. The Claude Platform prompt-caching documentation says that 1-hour cache-write tokens cost 2 times the base input-token price, while cache-read tokens cost 0.1 times the base input-token price. The pricing table provides separate model-specific read and write rates, so use the published rates that match the model and token category being estimated.
A useful calculation structure is:
estimated API cost =
uncached input token volume × uncached input rate
+ cache-write token volume × cache-write rate
+ cache-read token volume × cache-read rate
+ output token volume × output rate
Apply rates on the same per-MTok basis as the pricing table. Do not invent a blended token rate when cache use is present. Likewise, an LLM cost calculator is only useful here if it can account for cache reads and cache writes when they occur.
Thinking belongs in the output side of the estimate. Anthropic’s cost guidance says thinking tokens are billed as output tokens and that the default thinking budget can be tens of thousands of tokens per request, depending on the model. The visible answer length is therefore not enough to represent the full billed output volume.
Control Claude Code pricing by task type
Routine coding and architecture. Anthropic’s cost guide says Sonnet handles most coding tasks well and costs less than Opus, while Opus should be reserved for complex architectural decisions or multi-step reasoning. Use that as routing guidance rather than assuming every task needs the model with the highest rate.
In our setup, mechanical batch work and first drafts go to cheaper models, while architecture decisions, root-cause analysis, and verification stay with the strongest model. We also keep planning, file writes, and deploys with the main session. That is our operating practice, not a documented Claude Code default.
A switch to unrelated work. Start a fresh context with the documented command:
/clear
The Claude Code cost guide recommends using /clear when switching to unrelated work because stale context wastes tokens on every subsequent message. Treat the command as a task boundary: use it when the previous context no longer serves the next job.
Verbose operations. Delegating verbose work to a subagent keeps that output in the subagent’s context while a summary returns to the main conversation. The cost documentation also makes the important limitation explicit: the subagent’s own requests still draw on usage. A cleaner main transcript does not make the delegated work free.
In our setup, we send noisy searches and independent parallel work to subagents, but we do not delegate a check we could run inline. The summary should answer the assigned question; it should not conceal whether the underlying work was actually verified.
Plan-mode agent teams. Agent teams use approximately 7x more tokens than standard sessions when teammates run in plan mode because each teammate maintains its own context window and runs as a separate Claude instance, according to the cost guide. Classify this as its own high-cost task type rather than treating a team run as one ordinary session.
Background work. Background processes typically consume under $0.04 per session even without active interaction, as reported in the same cost guide. That figure is a typical cost, not a ceiling or a substitute for checking a particular run.
Check it worked
Use the built-in usage command:
/usage
The Session block in /usage shows API token usage and is intended for API users. For Max and Pro subscribers, the displayed session cost is not relevant to billing because their usage is included in the subscription, as explained in Claude Code’s cost guide.
Claude Code computes the dollar figure locally from token counts at list price unless an organization’s modelPricing table is in effect. Use that display to compare runs handled through the same billing path, not to treat a subscription session as an API invoice.
Before the task, record the billing path, task type, selected model, and expected cost drivers. Afterward, run /usage and compare actual token usage with the estimate. If cache or thinking categories matter separately, the cost documentation does not say that /usage exposes every category as a separate field. Check the current usage reporting available to the account and leave unsupported breakdowns unresolved rather than guessing.
In our setup, a new batch starts with a representative, high-risk item that goes through the full production chain and an audit before the rest runs. For cost control, that item should include the model, context pattern, and delegation pattern you expect to repeat. We do not treat an agent summary or successful command as proof that the spend is understood; done means checked.
Anthropic reports an enterprise average of about $13 per developer per active day and $150-250 per developer per month, with 90% of users below $30 per active day in the Claude Code cost documentation. Use those figures as a reported enterprise reference, not as a task estimate, budget cap, or expected result for your own workload.
Where it breaks
Cache assumptions break after a pause. The first message after a break longer than the cache lifetime misses the cache and reprocesses the full context. The cost guide gives the cache lifetime as one hour on a subscription and five minutes by default with an API key or cloud provider. Include that possibility in the task boundary instead of assuming every continuation behaves like the previous message.
Hard-coded rates go stale. Reopen the current pricing page and cost guide before relying on a forecast. List prices, included subscription usage, and organization pricing serve different purposes; one blended number cannot represent all of them accurately.
A lower model rate does not guarantee a cheaper task. Context growth, long output, thinking, and delegated work can outweigh the apparent advantage of a lower per-token rate. Compare complete task estimates and actual usage rather than ranking models in isolation.
Automation can invent controls that do not exist. The cost documentation does not specify a universal model-routing flag, a task-wide thinking cap, or a file path and format for the modelPricing table. Check the current product documentation before adding those controls to a script, and do not infer syntax from unrelated configuration keys.
Treat each estimate as a boundary for the next run. If actual token use follows the expected pattern, keep the routing. If it does not, identify the context, output, cache, or subagent driver before loosening the control.