Reuse Before Recompute
Treat prompt-prefix caching as part of context architecture, not merely a provider billing feature.
A team running a long debugging session finally acts on what the Chapter 2 instrument showed. Their traces show 20,000 tokens of duplicate file reads and obsolete tool output, so they add a cleanup pass that strips the duplicates before every request. Token counts fall by a sixth. The next invoice rises. Latency does not improve either. Nothing about the model changed and nothing about the task changed. The cleanup was real, the savings were arithmetic, and the bill disagrees, because the bill never measured tokens. It measured computation, and the cleanup destroyed the reuse that had been quietly subsidising every turn.
Chapter 8 ended with recomputation waste: identical prefixes processed from scratch each request because no interface remembers prior work. This chapter is about the interfaces that do remember, and about the layout discipline they demand. Prompt-prefix caching is usually presented as a billing feature. For this book it is something more consequential: the first mechanism where the shape of the context, not just its size, determines what the session costs.
Two caches, held apart
Chapter 8 drew the line between the model’s KV state and the provider’s prompt cache. Everything below concerns the second: a serving-layer mechanism in which a new request whose rendered leading tokens match a previously computed prefix reuses that work instead of redoing it. The two are related, OpenAI’s documentation describes its prompt cache directly as preserved KV tensors, but they have different owners, granularity and controls. This chapter says “prompt cache” for the serving concept and “KV state” for the model concept, always, and treats any sentence where either reading is possible as defective.
The primitive: repeated context need not mean repeated computation
An agent session is a study in repetition. A hundred turns may all begin with the same system instructions, the same tool definitions, the same project rules, the same early history, diverging only in the latest observations and the current request. The naive computational model processes everything every time:
request 1 process A B C D
request 2 process A B C D E
request 3 process A B C D E F
With reusable prefix computation, the shared work happens once:
request 1 compute A B C D
request 2 reuse A B C D, compute E
request 3 reuse A B C D E, compute F
The chapter’s primitive follows: repeated context does not necessarily imply repeated computation, but only while the reusable part stays structurally compatible with what the cache holds. Compatibility is the entire subject. Everything below describes what preserves it, what breaks it, and what it costs to create it in the first place.
The geometry of a reusable prefix
Take two requests, R1 = A B C D E and R2 = A B C D F. Their reusable prefix is A B C D. At the first divergence, E against F, reuse stops, because nothing about the matched prefix certifies the unmatched suffix. The geometry generalises:
shared prefix
↓
candidate reusable computation
first divergence
↓
new computation required
Two qualifications keep this honest across providers. First, no provider promises literal token-by-token global matching as its implementation; each documents its own units of reuse, breakpoints, prefix units, lookup windows, and the chapter treats those as separate mechanisms sharing one geometry rather than one mechanism with different labels. Second, reuse is attempted, never guaranteed: routing, expiry, and best-effort service mean an identical prefix can still miss, a fact the lifetime section makes operational. The geometry describes what can be reused. Provider mechanics describe what is.
Open serving systems show the same geometry without any commercial API attached. SGLang’s RadixAttention organises prompts in a radix tree so that requests sharing prefixes share KV-state computation, which is exactly the candidate-reuse shape above implemented as scheduler data structure. The book cites it as the open counterpart that suggests the principle is architectural rather than vendor-specific, and draws no inference whatsoever about what any commercial provider runs internally.
Stable first, volatile later, within semantic law
Chapter 6 established stable-versus-dynamic placement as an ordering distinction and left the economic payoff uncollected. This chapter collects it. A cache-friendly bundle has a characteristic shape:
STABLE
system and product instructions
tool definitions
project rules
stable shared reference material
↓
SEMI-STABLE
conversation history
previous tool calls and results
↓
DYNAMIC
current state
current observation
current user message
Stable material leads because every request reuses it; volatile material trails because each request recomputes from the divergence anyway. The providers’ own guidance converges on this shape from different directions. OpenAI’s documentation advises placing stable developer instructions and shared reference material first, pushing timestamps and user-specific content late or into later messages. Anthropic’s breakpoint guidance says to mark the last block whose prefix stays identical across the requests meant to share a cache. Kimi’s documentation says the same in plain terms: stable prompts, tool definitions and reference material at the front, content that changes every turn at the end. Three vendors, one geometry.
The shape is a hypothesis to test, not a licence to reorder arbitrarily, because authority and semantics constrain layout first. Moving a project rule after the observations it governs may improve prefix stability while breaking instruction precedence; burying the current request early to protect a history cache may satisfy the cache and starve the task. Chapter 6’s canonical layout already encodes authority-before-data and framing-before-evidence, and cache placement optimises strictly within the layouts those constraints leave legal. The reconciliation, recorded here as policy:
1. correctness and authority requirements
2. behavioural ordering requirements
3. freshness requirements
4. cache-friendly placement where still free to choose
Caching never outranks correctness. A layout that buys hits with misplaced governance is not optimised. It is corrupt.
Append versus rewrite
The largest structural choice a harness makes is whether history grows by addition or by revision, and caching prices the two oppositely. Append-only growth preserves every earlier prefix:
turn 1 A B
turn 2 A B C
turn 3 A B C D
turn 4 A B C D E
Each turn’s request extends the last, so the reusable prefix grows monotonically and every earlier computation stays valid. Historical rewrite does the opposite:
turn 1 A B
turn 2 A B C
turn 3 A SUMMARY(B C)
turn 4 A SUMMARY2(B C D) E
The rewritten request is shorter at every step past the first edit, and every rewrite moves the divergence point backward, orphaning the cached computation of everything after it. This is the economic conflict Chapters 10 through 12 inherit and must respect: deletion and summarisation buy token reduction with cache invalidation, and neither side of the trade is visible from token counts alone. The chapter does not resolve the conflict. It prices it, and the price arrives in the next section.
Mutation radius
A small edit early can cause a large recomputation footprint, and the concept deserves a name because token-diff thinking misses it completely. Call it mutation radius: how much downstream prefix reuse is lost when an earlier part of the rendered context changes. The term earns its place the first time a reader sees a one-line timestamp edit invalidate 80,000 tokens of cached prefix, an event every provider’s documentation describes in its own vocabulary. OpenAI’s gotcha pages show extending a message, switching breakpoint modes, or rewriting developer content orphaning everything after the change. Anthropic’s timestamp-trap example shows a per-request block defeating an entire static prefix because no write ever accumulated behind it. Kimi’s documentation states the general rule: when any part of a prefix changes, the content after that position cannot be reused.
The useful content of the term is one asymmetry:
small edit size
≠
small recomputation footprint
positioned early, the first dominates the second by orders of magnitude. Every later chapter that mutates history, and all of them do, must report its mutation radius alongside its token savings, or its savings are unaudited. That reporting requirement is Chapter 9’s main bequest to Chapter 10.
Writes, reads, lifetimes: the amortisation curve
A reusable prefix is not free to create, and the economics turn on who pays for creation. Conceptually every provider implements some version of the same curve:
total cost
= initial write or computation
+ Σ later cache reads
+ Σ uncached suffix computation
against
total uncached cost
= Σ full-request computation
A cacheable prefix repays its setup cost only through enough reuse before expiry, which makes caching an amortisation problem rather than a discount. OpenAI’s current caching semantics, verified in September 2026, make the cleanest worked example the book has yet had. Cache writes cost 1.25 times the ordinary input rate; subsequent reads cost 0.1 times. Writing a prefix once and fully reusing it once costs 1.35 times its ordinary input cost against 2 times without caching; across ten requests, one write plus nine full reads costs 2.15 times against 10 times. The break-even is nearly immediate for genuinely repeated prefixes and never arrives for one-shot content, which is why explicit-only mode exists: content unlikely to be reused should never incur the write charge at all.
The four documented implementations differ in controls while rhyming in economics. As the September 2026 documentation the chapter cites describes them:
| Provider | Write cost | Read cost | Lifetime | Controls |
|---|---|---|---|---|
| OpenAI, current models | 1.25 times ordinary input | 0.1 times | Thirty-minute minimum, from latest write or reuse | Explicit mode; up to four writes per request |
| Anthropic | 1.25 times (five-minute default) or 2 times (one-hour option) | A tenth, with per-model exceptions | Five minutes or one hour, refreshed free on reuse | Up to four breakpoints; telemetry separates created, read, and post-breakpoint tokens |
| DeepSeek | No separately documented charge | Separate hit pricing; hit and miss tokens reported per request | Best effort, hours to days | Automatic prefix units with full-match rules |
| Kimi | Same as ordinary input for the five-minute lifetime, twice that for the one-hour option | One tenth of ordinary input | Five minutes by default, or one hour; fixed at first write | Automatic by default; layout advice is the main lever |
Four implementations, one curve: pay to create, discount to reuse, expire eventually.
Lifetime is the curve’s third term and the easiest to forget. Anthropic measures entry life from the start of the writing or reading request, with free refresh on reuse, so a four-minute response leaves roughly one minute for the follow-up under the default TTL. OpenAI’s current models offer a thirty-minute minimum TTL from latest write or reuse. DeepSeek promises nothing beyond automatic clearing after some hours to some days; Kimi offers five minutes or one hour, locked at the first write. A miss therefore proves nothing about layout on its own: identical prefixes miss across expiry, across routing changes, across regions, across model or configuration changes. The experiment below records intervals and settings for exactly this reason. Structurally reusable and currently available are different claims, and only the second decides the bill.
The minimum-length trap
Some providers decline to cache short prefixes at all, which breaks the naive rule that fewer tokens cost less. OpenAI’s current floor is 1,024 tokens; Anthropic’s minima run from 512 to 4,096 by model; Kimi caches in blocks and cannot write a portion smaller than one block. Below the floor, a stable prefix earns zero reuse however perfectly arranged. Just above it, the same material amortises across the session. The documentation works the resulting break-even explicitly: with a 1,024-token floor, tenth-rate reads, and 1.25-rate writes, a prefix needs on the order of a hundred-plus tokens of stable content with realistic reuse before expansion beats brevity, and the tiniest prefixes never benefit at any request count.
The chapter recommends nothing padded. Useless text injected to cross a threshold buys cache eligibility with interference, the exact trade Chapter 5 forbids. Where the docs advise expansion, they mean useful stable material, examples, reference content, calibration text, or else shortening what cannot earn reuse. The trap matters here as pure first principles: it is the cleanest case in the book where the token-minimal bundle is not the cost-minimal bundle, stated with numbers rather than slogans.
Breakpoints are context decisions
Where providers expose explicit breakpoints, caching policy becomes context management by another name. Each breakpoint declares that the prefix behind it is stable enough to pay for saving, and each omission declares the suffix too volatile to bother writing. The decision variables are now visible: expected reuse count, prefix size, volatility, lifetime, write cost, read discount, latency sensitivity. OpenAI’s explicit mode makes the economics literal, up to four writes per request, unmarked suffixes processed uncached with no write charge. Anthropic’s model makes the failure mode literal instead: writes happen only at breakpoints, reads search backward through a twenty-block window for prior writes, so a breakpoint placed on per-request content writes expensively every turn and reads never. The general lesson needs no vendor: breakpoint placement is admission policy for computation, and it belongs beside the admission policy for content.
Tools and rules as cache examples
Two standing costs from earlier chapters return with second properties. Tool definitions, Chapter 3’s standing tax, are often the most reusable bytes in the bundle: forty schemas unchanged across hundreds of requests form a prefix worth writing once and reading forever. Mutate one schema per request, a timestamp, a dynamic identifier, session-specific text, early enough in the order, and the mutation radius wipes out everything downstream. OpenAI’s guidance is explicit because the failure is common: keep definitions, ordering, and schemas stable, disable tools per request by choice rather than by removal, defer loading so discovered tools append late instead of churning early. None of this is Chapter 17’s tool design. It is the cache ledger’s view of the same objects.
Project instruction files behave identically. AGENTS.md and CLAUDE.md change rarely, which makes them excellent prefix material, until someone embeds the current time, a session identifier, or live branch status inside the early standing block. Chapter 3’s hidden context becomes economically observable here: bytes nobody typed and nobody reviews, invalidating reuse on every turn through pure volatility. The fix costs nothing behavioural, render dynamic values late or not at all, and it is invisible without the telemetry this chapter’s experiment records.
Freshness, ordering, retention: three reconciliations
Cache stability collides with three earlier constraints, and each collision resolves the same way: correctness first, cache second.
Freshness first. A stable stale snapshot, a file state the repository has moved past, reuses beautifully and answers wrongly. Cache stability is not information validity, and no hit rate justifies retaining what is incorrect. Chapter 20 will own validity machinery; this chapter establishes only the precedence, because a compiler optimising for hits will otherwise learn to prefer the stale. Correctness dominates cache optimisation, and any future objective function must encode that ordering, not discover it.
Ordering second. Chapter 6 showed that sequence changes behaviour through authority, salience, and adjacency, none of which consult the cache. The reconciliation from the stable-first section stands: semantic and authority constraints define the legal layouts, and cache stability chooses among them. “Put all stable tokens first” as an unconditional rule would happily demote a governing instruction beneath volatile observations. The legal-layouts framing is what lets Chapters 6 and 9 agree instead of competing.
Retention third. Chapter 7’s classes describe what information requires, not what reuses well, and the two dimensions cross freely. PIN plus highly volatile is important and poorly reusable: a live deployment lock consulted every turn but rewritten every turn. REFETCHABLE plus stable is eminently cacheable: project reference material re-read rarely but reused constantly while present. Retention semantics and cacheability are independent axes, and the chapter introduces no new taxonomy field for the second because the existing telemetry already measures it: realised read ratios per span are cacheability observed, no label required. If a candidate builder ever needs a cacheability predictor rather than a measurement, that field must earn its place against observed ratios. It has not yet.
Cacheability is an economic property, not a relevance judgement
That sentence is worth isolating because the whole chapter compresses into it. Nothing about a span’s hit rate says anything about whether the span should be there. A cached irrelevance is still an irrelevance; it merely costs less to be irrelevant. Conversely, an uncacheable necessity, the genuinely novel observation each turn must carry, is not waste because it misses. Teams that optimise the hit-rate dashboard instead of the bundle will keep cheap filler and starve live signal, repeating at the economic level the exact error Chapter 5 diagnosed at the behavioural level. The instrument reports reuse. Judgement about membership stays with content, retention, and ordering policy.
What caching cannot solve
The boundary needs stating without softening, because caching success feels like context success. Prefix reuse reduces recomputation, latency, and input-processing cost. It does not touch window occupancy: cached tokens still count against capacity on the providers that count them, and the bundle still fills. It does not touch interference, position sensitivity, staleness, authority conflict, or bloat. A million cached irrelevant tokens could still be bad context even if processing them were free. Hence the chapter’s cleanest sentence, kept verbatim:
Caching can make bad context cheaper.
That is why caching and selection are complements rather than substitutes, and why the runtime chapters survive every cache improvement. Selection decides what deserves to be there. Caching decides what repeated presence costs. Neither answers the other’s question.
Proposed experiment
The question. How much does a small early edit cost in recomputation, compared with the same edit late or a rewrite of history?
The design, in brief. Measure money and computation, not intelligence, using a trivial deterministic task. Build one large stable prefix (instructions, tools, reference material, frozen history) with a small dynamic suffix, and run five conditions with identical semantics: stable append-only growth; a small inert change near the beginning; the same-sized change after the cacheable prefix; an older segment replaced by a fixed-length synthetic summary, built mechanically so that Chapter 11’s behavioural questions stay out; and a churn control that guarantees cold computation. Repeat the core run at immediate, short-delay, near-expiry and post-expiry intervals.
The measurement that matters. Per request: visible input, cache-write, cache-read and uncached tokens, time to first token, request interval, first-divergence position, and cost computed independently of the provider’s own figure. Where a provider hides a metric it is recorded as hidden. No composite score across providers, because their reporting units differ.
What would change the book. If an early edit does not cost more than a late one, mutation radius is not a real effect on these providers and the constraint pruning inherits loosens. Nothing here has been run.
The constraint pruning inherits
Return to the session that opened the chapter: 120,000 tokens in context, 80,000 of them cached cheaply, and 20,000 of them the duplicate file reads and obsolete tool output. Caching reduced their computational cost. It changed nothing about their membership. The coming deletion chapter must therefore optimise something closer to this than to bare token counts:
benefit of removal
≈ reduced occupancy
+ reduced uncached processing
+ reduced interference
− information-loss risk
− cache-reuse loss
− mutation and rewrite cost
The equation is deliberately approximate, a list of variables with signs, not a scoring formula. Its cache terms are this chapter’s bequest: every deletion carries a reuse price, every rewrite a mutation radius, and any pruning result reported without both is unaudited. With that ledger open, the book may finally start removing things:
Which context can we remove without destroying information we still need, and without accidentally losing more cache value than the removal saves?
References
- OpenAI. “Prompt caching.” Official documentation, verified 25 September 2026 for the newest model generation. Current-model write and read economics, breakpoints, TTL, routing, telemetry, minimum-length trap. https://developers.openai.com/api/docs/guides/prompt-caching
- Anthropic. “Prompt caching.” Official documentation, verified September 2026. Automatic and explicit breakpoints, write/read pricing, TTL options, minima, lookback, invalidation hierarchy. https://platform.claude.com/docs/en/build-with-claude/prompt-caching
- Anthropic. “Context windows.” Official documentation, verified September 2026. Cached prefixes still occupy the window; cached computation is not freed capacity. https://platform.claude.com/docs/en/build-with-claude/context-windows
- DeepSeek. “Context Caching.” Official documentation, verified September 2026. Disk-backed prefix units, full-match rules, hit/miss telemetry, best-effort lifetime. https://api-docs.deepseek.com/guides/kv_cache
- Moonshot AI. “Use the Context Caching Feature of Kimi API.” Official documentation, verified September 2026. Automatic caching with five-minute or one-hour lifetimes, block-based writes, stability-first layout. https://platform.kimi.ai/docs/guide/use-context-caching-feature-of-kimi-api
- Zheng, L., Yin, L., Xie, Z., et al. “SGLang: Efficient Execution of Structured Language Model Programs.” Preprint, arXiv:2312.07104. RadixAttention prefix reuse in open serving systems. https://arxiv.org/abs/2312.07104