Skip to content
12 min read

When Cache Reads Get Cheap Enough, You Are Paying for Cache Writes

Cached input now costs as little as 2.5 percent of base input. On a read-heavy agent session that moves the input bill onto cache writes, which prompt design controls.

Antonio J. del Águila

Knaisoma

On 22 September, Anthropic released Claude Opus 5.5 and OpenAI released GPT-6 Sol and GPT-6 Luna, all cheaper than the models they replace. If you priced the change from the input column, you got one of two very different answers without being told which.

GPT-5.6 Sol to GPT-6 Sol is the straightforward case. Input, cached input and output all halved, from $4, $0.40 and $20 per million tokens to $2, $0.20 and $10. If your usage does not change, halving last month’s bill is exactly right. Claude Opus 5 to Opus 5.5 is not that case: base input and output each fell 20 percent, while the price of a cache read fell 60 percent, from $0.50 to $0.20 per million tokens. A workload whose input arrives mostly as cache reads got a much larger cut than the input column advertises, and a workload that never caches got the smaller one.

Which of those you are is not something the price list can tell you, because it depends on what share of your input arrives as cache reads. On a multi-turn agent that share is high. Simon Willison, writing the same day about the new pricing landscape, observed that for longer agentic conversations more than ninety percent of input tokens are processed at cached token prices. That is one practitioner’s observation rather than a survey, but any multi-turn agent produces the same shape.

What follows from a deep read discount is worth sitting with. Anthropic now prices cache hits on Opus 5.5 at 5 percent of base input and on Fable 5.1 at 2.5 percent, against the 10 percent multiplier that applies to its other models and to OpenAI’s cached input across the GPT-6 family. Base input still sets the scale of the whole bill, because a cache write is itself a multiple of it. What moves is which part of the prompt the money is attached to.

Where a read-heavy session actually spends

Take an illustrative scenario rather than a customer story: a support triage agent with a 30,000 token stable prefix holding its system prompt, tool definitions and policy extracts, running twelve turns, where each turn adds about 1,500 tokens of tool results and user text and the model replies with about 400. Assume for now that the cache never expires mid-session, and that each turn caches the conversation as it stands.

On those assumptions the model processes 485,400 input tokens and produces 4,800 output tokens. Of the input, 434,500 tokens are cache reads and 50,900 are cache writes, so reads are 89.5 percent of the tokens, matching the shape Willison describes. Priced on Opus 5.5, the writes cost $0.2545, the reads $0.0869 and the output $0.096. The session lands at $0.44 against $2.04 uncached, so caching removes about four fifths of the bill and whether to use it is not in question. Look at the split, though: ten and a half percent of the input tokens carry three quarters of the input spend.

89.5%

Share of input tokens billed as cache reads

Illustrative 12-turn session

74.5%

Share of input spend that is cache writes

Same session, Opus 5.5 pricing

0.025x

Lowest published cache read multiplier

Anthropic pricing, September 2026

A write is charged at 1.25 times base input for the five minute cache and twice base input for the one hour cache, so on Opus 5.5 each write token costs twenty-five to forty times what a read token costs. Prefix size matters, as it always did. What the deep read discount adds is that prefix turnover now matters just as much, and turnover is a property of how you built the agent rather than of the model you sent it to.

The three ways a prefix gets rewritten

The first is expiry. Anthropic’s default cache lives five minutes, refreshed on each hit; OpenAI’s, on GPT-5.6 and later, remains eligible for reuse for thirty minutes after its most recent write or reuse. A six-minute gap between two turns of the same conversation is unremarkable when a person approves a step or a tool call queues behind a slow system, and on the shorter window it costs a rewrite.

The cost of that rewrite is not the whole write price. Tokens you were going to pay for as reads are instead paid for as writes, so the incremental charge is the previously cached span multiplied by the difference between the two rates, which on Opus 5.5 is $4.80 per million. Early in the illustrative session the reusable span is about 31,900 tokens and a miss costs about $0.15; by the last turn the span has grown to 50,900 tokens and the same miss costs about $0.24. Against a $0.44 session, one stall is expensive and a habit of them is a budget line.

A sharper version of this catches teams running long reasoning turns. Anthropic states that the cache lifetime is measured from the start of the request that writes or reads the entry, not from the end of its response, so if a response takes four minutes to stream, the follow-up must start within about one minute of that response completing. A single deep reasoning turn can consume most of a five-minute window before the next request is composed.

The second is invalidation. The cache is an exact prefix match over tools, then system, then messages, and a change to the tool definitions invalidates the entire cache, not merely the tool that changed. Teams that hot-reload a tool registry, append a per-request identifier to the system prompt, or let a compliance banner carry a timestamp are buying a full rewrite on every call, and the resulting bill looks as if caching had been switched off.

The third produces no signal at all. Every vendor sets a minimum cacheable length, and it varies more than people expect: Anthropic’s is 512 tokens on its newest models but 4,096 on Opus 4.5, 4.6 and Haiku 4.5, OpenAI’s is 1,024 visible input tokens on GPT-5.6 and later, and Google’s implicit cache needs 4,096 tokens on Gemini 3.8 Flash. Anthropic’s documentation is explicit about what happens below the threshold: shorter prompts cannot be cached even when marked, and no error is returned. A prompt that caches cleanly on one model in a routing tier can silently stop caching on the smaller model beside it.

Three vendors, three ways to pay for retention

The mechanism has converged and the read discount is broadly similar. How you pay to keep a cache alive has not converged at all, and that is the part that changes when you switch.

Cost to populateRetention
Anthropic, Claude API1.25x input, 2x for the hour5 minutes, or 1 hour
OpenAI, GPT-5.6 and later1.25x input30 minutes, rolling
Google, Gemini explicit cachesinput, plus storageyour TTL, billed hourly

Reads are the boring column, which is why they are not in the table: they run at 0.1x input across all three vendors’ mainstream models, with Anthropic’s 0.05x on Opus 5.5 and 0.025x on Fable 5.1 as the exceptions. The opt-in differs too. Anthropic wants a cache_control marker on the request, OpenAI caches by default on supported models, and Google runs two modes: an implicit cache that is on by default from Gemini 2.5 onward and passes savings through automatically, and explicit cache objects you create and hold yourself.

Only the explicit mode carries the storage charge, and that row is the one to read twice. Google prices explicit context caching at $0.075 per million tokens against $0.75 for input on Gemini 3.8 Flash, plus $0.50 per million tokens per hour of storage through the end of 2026, with both roughly doubling in January. For a single 30,000 token prefix the storage is $0.015 an hour, which is nothing. For a thousand tenants each holding their own prefix through an eight-hour working day it is $120 a day, charged whether or not a request arrives. Choosing the implicit mode avoids that charge and the control it buys. Either way, a retention design that is free under one cost model is a standing charge under another, and no pricing page will tell you which, because none is describing your workload.

The control surface moves as well. Google’s documentation notes that the Interactions API supports implicit caching only, so explicit cache objects are not available on every surface you might build against. Plans that treat a vendor change as a configuration swap discover this after the migration, when the cache behaviour they tuned for is not on offer.

What a cache hit does not buy you

Two limits matter before anyone promises a capacity benefit alongside the cost one. OpenAI’s guide states that cached prompts still count toward tokens-per-minute rate limits, so a workload that halves its bill through caching has bought no throughput headroom. If your constraint is rate limits rather than spend, caching is the wrong lever.

The second is scope. Anthropic’s caches are isolated per workspace within an organization, which quietly punishes a common piece of housekeeping. Splitting traffic across workspaces so that each team or environment gets its own billing line also splits the cache, so four workspaces sharing one 30,000 token system prompt pay four times to populate it, and pay again on every expiry, for a prefix that is byte-identical across all of them. That is a governance decision with a runtime price, worth making deliberately rather than inheriting from how the accounts were set up.

The decision this actually produces

A decision tree. Start by asking whether the stable prefix reaches the model's minimum cacheable length. If it does not, the workload is billed at full input price with no error, so the fix is to consolidate the prefix or move to a model with a lower minimum. If it does, ask whether consecutive requests arrive inside the retention window. If they do not, the workload is paying repeated writes, so extend retention or close the gaps. If they do, ask whether the prefix is byte-identical across requests. If it is not, tool or header churn is invalidating the cache and that must be fixed first. If it is, ask whether each tenant needs a distinct prefix. If so, model the cost per tenant rather than per request, because write premiums and storage charges multiply by tenant. Once all four hold, comparing models on a replay of real traffic gives a trustworthy answer.

flowchart TD
  A[Price a workload] --> B{Above the<br/>minimum?}
  B -->|No| B1[No cache,<br/>no error]
  B -->|Yes| C{Inside the<br/>window?}
  C -->|No| C1[Repeated<br/>rewrites]
  C -->|Yes| D{Identical<br/>prefix?}
  D -->|No| D1[Tool or header<br/>churn]
  D -->|Yes| E{One prefix<br/>per tenant?}
  E -->|Yes| E1[Price per<br/>tenant]
  E -->|No| F[Replay real traffic]
The four checks before a model comparison means anything: does the prefix clear the model minimum, do requests land inside the retention window, is the prefix identical every time, and does each tenant need its own. Three are settled by your prompt and pacing; the first is partly the model's.

The retention branch is worth working through with numbers, because the answer is closer than it looks. On the illustrative session the one hour cache costs $0.59 with no misses, against $0.44 on the five minute cache, because its write premium is 2x rather than 1.25x. A single miss on the short cache costs $0.15 to $0.24 depending on how much history had accumulated, so one stall roughly cancels the difference and two settle it in favour of the longer window. The useful form is the formula, not the figure: weigh the extra premium you pay on every write against the reusable span times the gap between the write and read rates, times the stalls you actually observe. Measure the stalls before choosing.

Why the price columns cannot be compared directly anyway

There is a further problem with the spreadsheet, and it is not about caching. A token is not a fixed quantity of text. Anthropic states that its models from 4.7 onward use a newer tokenizer producing approximately 30 percent more tokens for the same text than the one its earlier models use, with the exact increase depending on content and workload shape. Migrating across that boundary changes the token count of every document you send, in a direction that partly offsets a headline price cut, and the published rate alone will not say by how much. The lesson is narrower than a conspiracy and more useful: a per-million-token price is comparable between two models only once you know how each one counts.

This is why the only comparison that survives contact with a real bill is a replay. Take a representative slice of production traffic, run it against each candidate with your actual prompt structure and pacing, and read the cached and uncached token counts out of the usage block. The dollar figures here were computed from published prices on a session the author constructed to make the structure visible. They are not a ranking of the models.

The broader point outlasts this price war. Vendors compete on the number buyers compare, and buyers have been comparing base input. As the terms around that number multiply, a read multiplier here, a retention window there, a storage charge somewhere else, the published price becomes a worse proxy for what a workload costs. Treat cache terms as part of the model’s interface, alongside its context window and its tool-calling behaviour, and re-derive the number when you switch.

Further reading: our earlier piece on long-context models and the retrieval decision covers prefix stability from the retrieval-architecture side, and Anthropic’s prompt caching documentation sets out the invalidation rules in full.

Model spend that grows faster than usage usually has its answer somewhere in the prompt, and finding it is unglamorous work that few teams have a spare week for. We can help: instrumenting token and cache accounting per workload, restructuring agent prompts and tool registries so that caches actually hold, and building the replay harness that turns a vendor comparison into a measured decision. Talk with us about your AI running costs.

AI Agentic AI LLMOps Architecture
Share:

Stay updated

Get insights on engineering transformation delivered to your inbox.

Newsletter coming soon.