Claude Code’s /usage view separates session tokens into input, output, cache reads, and cache writes, grouped by model. Input can include tool definitions and results as well as your messages; cache reads and writes are distinct input-side operations. The displayed session cost is an estimate—not the authoritative API bill. For API billing, check the Usage page in the Claude Console.
What the token categories mean
Claude Code sends a request to a model, receives a response, and may reuse parts of the prompt through caching. The categories in /usage distinguish these kinds of usage rather than treating every token alike. Anthropic describes the categories and session display in its Claude Code cost guide.
As an Amazon Associate I earn from qualifying purchases.
Input tokens
Input is the material sent to the model. It is not limited to the latest message you typed. In a coding session it can include instructions and conversation context, along with tool definitions, tool calls, and tool results. Anthropic’s API pricing documentation notes that tool requests are priced on total input sent, including the tools parameter and tool_use and tool_result blocks.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Output tokens
Output is the content generated by the model. API pricing treats input and output as separate categories with separate rates, so an output-token count should not be read as an input count or priced as though the two were interchangeable.
#1 Best Overall
Cache-read and cache-write tokens
A cache write stores prompt content for reuse; a cache read retrieves previously cached content in a later request. They are input-side usage categories, not extra output tokens. Writes and reads have distinct pricing from base input: Anthropic’s current general API pricing rules list five-minute cache writes at 1.25× base input and one-hour writes at 2×, while cache reads are 0.1× base input for most listed models. Model-specific exceptions and other pricing modifiers apply, and rates can change, so consult the live pricing page for the model and cache duration in question.
Cached tokens are not simply “free”: a cache read has its own price, and creating a cache entry can cost more than base input. These multipliers describe the general API rules, not a guarantee that every model or provider route uses the same effective rate.
Rank #2
How to inspect usage in Claude Code
-
In an active Claude Code session, run
/usage. The/costcommand is an alias. In the Session block, inspect the input, output, cache-read, and cache-write counts shown by model.Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy. -
To see how much of the active context window is being used, run
/context. The command visualizes current context consumption, including context-heavy tools and capacity warnings. It answers a different question from/usage, which reports session usage and cost. See the Claude Code command reference for current command details and version-dependent features.
In supported versions, Claude Code also reports prompt-cache statistics such as cache-hit share, misses, and warm or cold status. The cost guide says this line is based on cache-token fields returned by the API and covers the main conversation, not subagents. Because command features evolve, check the current command reference if these statistics are not visible in your version.
Why the displayed cost may differ from your bill
Claude Code calculates its displayed API session cost locally from token counts and list prices, unless an organization-managed modelPricing table applies. Anthropic labels this figure an estimate and directs API users to the Claude Console Usage page for authoritative billing. The CLI usage documentation also says the --max-budget-usd limit is enforced against a client-side estimate that may differ from the bill.
Rank #4
Account type and authentication route affect what the number means. The Session cost block is intended for API users. Pro and Max subscribers receive usage through their subscription, so that session figure is not a measure of a separate API bill. For gateway-routed requests, the gateway credential and upstream provider determine who is charged; Anthropic says an active gateway credential replaces the subscription login for those requests, with per-token billing to the owner of the forwarded credential. See the LLM gateway documentation.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →When investigating a difference, compare the same model and session, distinguish input and output from cache reads and writes, and verify the account or gateway route. Then compare Claude Code’s local estimate with the provider’s billing record rather than treating the estimate as an invoice.
Best Value
What you can and cannot infer from a token count
Character count or word count alone cannot reliably reproduce the token total for a Claude Code request. The reported usage reflects the complete request and response, which may include tool schemas, tool activity, conversation context, and cached content. Use the session or API usage fields for the actual request instead of applying a universal words-to-tokens conversion.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




