Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
World desk6 min

Why Agentic Systems Should Care About Cache-Hit Pricing

Agent loops can resend the same prompt prefix across many model calls. Cache-hit pricing can reduce repeated input costs, but write rates, retention windows and real hit rates determine the savings.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Agentic systems often resend the same prompt prefix—such as instructions, tool definitions and conversation history—on many model calls. When a provider can reuse that prefix from cache, those repeated input tokens may cost less and require less prefill work. Across a multi-step agent loop, the savings can add up, but only when the next request actually hits a valid cache entry.

What a cache hit saves in an agent loop

A prompt prefix is the beginning portion of a request that stays the same across calls. A cache hit means the provider recognizes that matching prefix and reuses its processed state instead of processing those prefix tokens from scratch. The new user message, any changed content and the model’s generated output still have to be handled; caching does not make the entire future request free.

As an Amazon Associate I earn from qualifying purchases.

This distinction matters for agents because a single task can involve several model calls: the agent reasons, invokes a tool, receives its result, and calls the model again. If each request repeats a long, unchanged prefix, cache-hit pricing can reduce the cost of that repeated input. The benefit compounds over repeated reads, rather than applying automatically to every token or every call.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reuse can also avoid much of the repeated prefill computation for the cached portion, which may help latency. But a lower list price for cached input is not proof that a particular workflow will be faster or cheaper: provider routing, cache availability and the agent’s timing all affect realized results.

How write costs and read discounts change the math

A cache write is the initial processing and storage of a prefix; a cache read is a later reuse of it. Compare the write premium with the discounted reads that follow. Looking only at a headline read discount misses the cost of creating the entry.

OpenAI API pricing

As described in OpenAI’s prompt-caching guide and API pricing page, GPT-5.6 and later have cache writes priced at 1.25× the standard uncached input rate. Subsequent cache reads are 0.1× on most such models and 0.05× on GPT-6.1 Sol. These are model-specific multipliers, not a universal OpenAI rate; check the pricing page for the exact model and platform before using a dollar amount.

For an illustrative workload where each request reuses the same full prefix, at a 0.1× read rate one write plus one full read costs 1.35× one ordinary input pass, compared with 2× for two ordinary passes. One write plus nine full reads costs 2.15× one ordinary pass, compared with 10× for ten ordinary passes. This arithmetic applies to the reused prefix only; new input, generated output and other charges remain separate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

OpenAI’s September 22, 2026 GPT-6 prompt-caching announcement says eligible shared prefixes reused within a 30-minute window can receive discounts of up to 90% on cached input tokens. “Up to” is material: it is an announcement about eligible prefixes, not a promise that an agent will realize that saving across all its tokens or calls.

Anthropic Claude API pricing

Anthropic’s Claude Platform pricing documentation lists a 5-minute cache write at 1.25× base input price, a one-hour cache write at 2×, and a cache read generally at 0.1× base input price, with model-specific exceptions. At that general read rate, Anthropic says the 5-minute write premium is paid back after one cache read and the one-hour write premium after two. These comparisons concern token rates; ordinary uncached input, output and platform charges still apply. Services hosted through partner platforms such as Amazon Bedrock or Google Cloud may have their own pricing.

API pricing example Cache write Cache read Break-even guidance
OpenAI API, GPT-5.6 and later, as documented in the current guide 1.25× standard uncached input rate 0.1× for most models in this generation; 0.05× for GPT-6.1 Sol Depends on the model’s read rate and repeated reads; use exact model rates for a workload calculation.
Anthropic Claude API, as documented in its pricing page 5-minute write: 1.25× base input; one-hour write: 2× base input Generally 0.1× base input, with model-specific exceptions At the general read rate, Anthropic says one read pays back the 5-minute write premium and two reads pay back the one-hour write premium.

The table compares provider-documented token-rate mechanics, not end-to-end costs or guaranteed savings. For an agent, calculate total workload cost using cache writes, cache reads, uncached and new input, output, and any platform-specific charges.

Why an agent can miss the cache between steps

Agent loops are not necessarily continuous. After a model call, the system may run a tool, wait for a human approval, or pause for another reason before sending the follow-up request. If that gap exceeds the provider’s cache lifetime, the next call may need a fresh write instead of a cache read.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A July 2026 preprint by Maxim Khailo, “Keeping the Cache Warm Pays: Keepalive Economics for Agentic Workloads”, analyzes this “think, act, wait” pattern and evaluates periodic keepalives. It is one researcher’s analysis, not an official provider recommendation or a universally validated operating rule. Whether keeping a cache warm makes economic sense depends on the workload’s pauses and the provider’s write, read and retention terms.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What determines whether a prefix is reusable

A large prompt alone does not guarantee a hit. Providers impose their own requirements for prefix matching, eligible content, minimum lengths, cache locations and retention. A request that changes content near the beginning can prevent reuse of later content in that prefix, and a cache entry may not be available on the route that handles the next call.

OpenAI API behavior

OpenAI’s guide describes machine-local cache states and notes that cache location, routing, cache lifetime and traffic can affect reuse. It recommends keeping conversation history and tool definitions stable. For GPT-5.6 and later, the guide documents explicit cache breakpoints and a 30-minute retention control, with at least 30 minutes after the latest write or reuse for that generation. Older OpenAI models have different behavior, minimum lengths and retention options, so those newer rules should not be assumed for every model.

Prompt design for agent workloads

  • Keep reusable instructions, reference material and stable tool definitions at the beginning of the prompt.
  • Where the API allows it, put changing per-turn content later so it does not alter the shared prefix.
  • Append to conversation history rather than rewriting its earlier turns, when that matches the application’s context strategy.
  • Check how truncating or rebuilding history affects the prefix; removing or changing earlier content can forfeit reuse.

How to evaluate cache-hit pricing for your agent

  1. Identify the exact API and model. Record the provider, model, platform and date of the pricing terms. Do not apply one model’s cache multipliers to another model or to a partner-hosted service.
  2. Estimate repeated-prefix reads. For representative tasks, measure how many follow-up calls reuse the same prefix and how many tokens remain unchanged. New user input and generated output are not cached prefix tokens.
  3. Compare write premium with likely reads. Use the provider’s write and read rates to calculate the cost of the initial write and subsequent reads against ordinary input processing. Include cases where the agent gets only one read or none.
  4. Compare retention with real pauses. Measure the time from one model call to its follow-up after tool runs and approvals. A cache lifetime that is shorter than common pauses may produce misses even when the prompt is otherwise stable.
  5. Inspect actual usage and billed cost. Use provider usage details or dashboards to check cached-token counts and cost on representative agent runs. A maximum advertised discount does not establish the workload’s realized hit rate.

OpenAI’s announcement attributes to GitHub Chief Product Officer Mario Rodriguez a reduction of more than 50% in the share of prompt tokens requiring fresh processing across billions of requests to OpenAI models, relative to GitHub’s previous baseline. That is GitHub’s reported company result, not an independent cross-provider statistic or a forecast for other agents. Rodriguez described prompt caching as playing a critical role in helping GitHub Copilot deliver fast, efficient experiences at scale.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.