Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
World desk3 min

How Claude Code Token Pricing and Cache TTL Work

Claude Code may use plan limits or API token billing. Here’s how cache TTL timing and API cache pricing affect repeated prompts.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Claude Code does not always charge you per token: billing depends on whether you signed in with a Claude plan or use an API key. With API billing, prompt caching can lower the charge for repeated prompt text, but cache writes cost extra and cached content still occupies context-window space.

How is Claude Code token usage metered?

There are two different billing routes:

  • Claude plan: Eligible plan seats, including Claude Pro, include Claude Code subject to plan usage limits. This is not ordinarily a per-token invoice; practical capacity varies with factors such as conversation length and complexity, model, and features. Check Anthropic’s plan page for current inclusions.
  • API key: Usage is charged per token to the relevant API account or provider, using the applicable model and pricing. Anthropic says /cost displays token and dollar usage for the current session when using API billing. See Claude Code usage guidance.

Do not apply API cache-price multipliers to plan usage limits. They describe API token pricing; the cited sources do not give a universal dollar conversion for subscription usage.

How much does Claude Code cost per token?

There is no single Claude Code price per token. For API billing, the amount depends on the selected model’s base input and output rates, the number of uncached input and output tokens, cache-write and cache-read tokens, provider, and applicable pricing modifiers. Anthropic’s live pricing page is the place to check model rates; they can change, so an undated model-price figure is not a reliable estimate.

Anthropic’s current standard API pricing describes these prompt-cache rates as multipliers of the model’s base input price:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Token type API price relative to base input
Ordinary input 1×
Five-minute cache write 1.25×
One-hour cache write 2×
Cache read 0.1×

These are Anthropic’s published multipliers for the cited standard tier, not a complete bill estimate or a promise of savings. A cache write costs more than ordinary input; savings depend on whether later requests read the cached prefix often enough to offset that initial write cost. See Anthropic API pricing.

What is Claude Code’s cache TTL?

TTL means “time to live”: how long a prompt-cache entry remains available for reuse. Anthropic’s prompt-caching documentation gives a default minimum cache lifetime of five minutes and an optional one-hour lifetime. Using an entry refreshes its TTL. The one-hour option is useful when requests are likely to be separated by more than five minutes, but its API cache-write multiplier is higher. See Anthropic prompt-caching documentation.

When does the cache timer start?

The timer starts at the beginning of the request that writes or reads the cache entry, not when the response finishes. For example, if a response takes four minutes under a five-minute TTL, roughly one minute remains for reuse afterward. A long generation therefore uses up some of the cache window.

Does Claude Code use a 5-minute or 1-hour cache?

Anthropic documents both TTL choices: five minutes is the default minimum, while one hour is an available extended lifetime. Which one is appropriate depends on the expected gap between requests. The price tradeoff applies to API billing: a five-minute write costs 1.25× base input, a one-hour write costs 2×, and a cache read costs 0.1× in the cited standard tier. The figures are multipliers, not dollar totals. Details are in the TTL documentation and API pricing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How does caching affect CLAUDE.md?

Anthropic’s Enterprise guidance describes prompt caching for CLAUDE.md: the first request in a session pays the file’s full input price, and subsequent turns within roughly five minutes can read it from cache at the lower cache-read rate. If the file changes, its content-addressed cached version is invalidated, so the changed content must be sent at full input price on the next request. Keeping the file concise remains useful because caching does not reduce the context space it occupies. See Anthropic’s CLAUDE.md guidance.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Does prompt caching make Claude Code free?

No. With API billing, cache reads are charged at a reduced rate, while the initial cache write has its own price. Cached text also remains part of the context Claude Code carries; caching changes billing treatment for repeated prompt prefixes, not their context-window occupancy. With a Claude plan, usage is governed by plan limits rather than the API multipliers. Anthropic explains the context point in its Claude Code usage guidance.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.