Claude Code does not always charge you per token: billing depends on whether you signed in with a Claude plan or use an API key. With API billing, prompt caching can lower the charge for repeated prompt text, but cache writes cost extra and cached content still occupies context-window space.
How is Claude Code token usage metered?
There are two different billing routes:
- Claude plan: Eligible plan seats, including Claude Pro, include Claude Code subject to plan usage limits. This is not ordinarily a per-token invoice; practical capacity varies with factors such as conversation length and complexity, model, and features. Check Anthropic’s plan page for current inclusions.
- API key: Usage is charged per token to the relevant API account or provider, using the applicable model and pricing. Anthropic says
/costdisplays token and dollar usage for the current session when using API billing. See Claude Code usage guidance.
Do not apply API cache-price multipliers to plan usage limits. They describe API token pricing; the cited sources do not give a universal dollar conversion for subscription usage.
How much does Claude Code cost per token?
There is no single Claude Code price per token. For API billing, the amount depends on the selected model’s base input and output rates, the number of uncached input and output tokens, cache-write and cache-read tokens, provider, and applicable pricing modifiers. Anthropic’s live pricing page is the place to check model rates; they can change, so an undated model-price figure is not a reliable estimate.
Anthropic’s current standard API pricing describes these prompt-cache rates as multipliers of the model’s base input price:
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match#1 Best Overall
| Token type | API price relative to base input |
|---|---|
| Ordinary input | 1× |
| Five-minute cache write | 1.25× |
| One-hour cache write | 2× |
| Cache read | 0.1× |
These are Anthropic’s published multipliers for the cited standard tier, not a complete bill estimate or a promise of savings. A cache write costs more than ordinary input; savings depend on whether later requests read the cached prefix often enough to offset that initial write cost. See Anthropic API pricing.
What is Claude Code’s cache TTL?
TTL means “time to live”: how long a prompt-cache entry remains available for reuse. Anthropic’s prompt-caching documentation gives a default minimum cache lifetime of five minutes and an optional one-hour lifetime. Using an entry refreshes its TTL. The one-hour option is useful when requests are likely to be separated by more than five minutes, but its API cache-write multiplier is higher. See Anthropic prompt-caching documentation.
Rank #2
When does the cache timer start?
The timer starts at the beginning of the request that writes or reads the cache entry, not when the response finishes. For example, if a response takes four minutes under a five-minute TTL, roughly one minute remains for reuse afterward. A long generation therefore uses up some of the cache window.
Does Claude Code use a 5-minute or 1-hour cache?
Anthropic documents both TTL choices: five minutes is the default minimum, while one hour is an available extended lifetime. Which one is appropriate depends on the expected gap between requests. The price tradeoff applies to API billing: a five-minute write costs 1.25× base input, a one-hour write costs 2×, and a cache read costs 0.1× in the cited standard tier. The figures are multipliers, not dollar totals. Details are in the TTL documentation and API pricing.
Rank #3
How does caching affect CLAUDE.md?
Anthropic’s Enterprise guidance describes prompt caching for CLAUDE.md: the first request in a session pays the file’s full input price, and subsequent turns within roughly five minutes can read it from cache at the lower cache-read rate. If the file changes, its content-addressed cached version is invalidated, so the changed content must be sent at full input price on the next request. Keeping the file concise remains useful because caching does not reduce the context space it occupies. See Anthropic’s CLAUDE.md guidance.
Does prompt caching make Claude Code free?
No. With API billing, cache reads are charged at a reduced rate, while the initial cache write has its own price. Cached text also remains part of the context Claude Code carries; caching changes billing treatment for repeated prompt prefixes, not their context-window occupancy. With a Claude plan, usage is governed by plan limits rather than the API multipliers. Anthropic explains the context point in its Claude Code usage guidance.
Quick Recap
Best Value
Rank #4
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




