← Context From First Principles

The Context Window Is a Budget

Treat the model window as a constrained allocation problem rather than a bucket to fill.

Chapter 3 ended with an observation: every hidden layer costs something. A team running a coding agent on a model with a million-token window read that sentence and shrugged. Their sessions rarely passed 150,000 tokens. Nothing was being truncated, nothing was overflowing, and the window, they reasoned, had settled the matter. Then two things happened. First, a long debugging session died at turn eleven with the model stopping mid-repair, not because the input had crossed a million tokens but because the input plus the output the model still needed no longer fit together. Second, the invoice arrived, and the “plenty of headroom” sessions turned out to be the most expensive ones they had ever run. The window had not settled anything. It had merely set the outer wall of a much smaller room they were actually living in.

This chapter builds that room. The hidden stack of layers becomes a set of competing claims on bounded capacity, and the engineering question changes accordingly: not how to fill the window, but what deserves a share of the budget.

What an advertised limit actually means

The confusion starts with the number itself. A “1M context window” sounds like permission to supply a million tokens of input. It is not, and the providers’ own documentation says so in different dialects that must not be assumed identical.

OpenAI’s conversation-state documentation describes the context window as covering input, output, and reasoning tokens in a single request: the total includes what you send and what the model generates, including its intermediate reasoning. The current model catalogue makes the split concrete. The flagship GPT-6 family, Astra, Sol, and Luna, advertises a 1.05M-token context window alongside a 128K maximum output. A million-token input on such a model leaves no room for a reply; the usable input budget was never a million tokens. It was a million minus whatever the response requires.

Anthropic’s context-window documentation draws the boundary in the same place with different machinery. Everything in the request counts: the system prompt, every message including tool results, images, and documents, and the tool definitions themselves. Everything the model generates counts too, including extended thinking tokens. Named current models carry a 1M-token window with up to 128K output tokens per request; others, including Sonnet 4.5, carry 200K. Overflow behaviour is specified rather than left to folklore: an input that alone exceeds the window is rejected outright, while on recent models an input that fits but leaves insufficient room for the requested output is accepted and may halt mid-generation with an explicit stop reason. The lesson generalises beyond one vendor: the advertised figure is the size of the whole computation, not the size of your input allowance.

DeepSeek’s model documentation supplies a third dialect. Its V4-generation models list a 1M context length with a 384K maximum output, alongside separate cache-hit and cache-miss input pricing. The accounting moral is the same, and the pricing columns add a second one that this chapter will return to: tokens that occupy identical capacity can cost very different amounts.

Moonshot’s Kimi documentation shows the range within a single provider’s catalogue. The K3 flagship offers a 1M-token window while the K2.6 general-purpose and K2.7 Code models offer 256K. Same vendor, same API shapes, different outer walls. Any harness logic that assumes “the window” as a constant is already wrong across the models it might call.

Qwen’s case adds a final distinction the later architecture chapter will inherit: native context length varies by model size. The Qwen3 generation ships 32K native windows on its small dense models and 128K on larger dense and mixture-of-experts variants, with hybrid thinking modes whose thinking budget the user controls. A thinking budget is window space spent before the answer begins, which makes it a reservation problem of exactly the kind this chapter describes, visible here in product form.

The table below compresses these dialects into one accounting comparison. It is pinned to documentation verified in September 2026 and will date; its purpose is not ranking but demonstrating that no two providers mean precisely the same thing by their headline number.

Provider / model generationHeadline figure (Sep 2026 docs)What the figure coversStated output reservation
OpenAI GPT-6 family1.05M context windowInput, output, and reasoning tokens per request128K max output
Anthropic current models1M or 200K by modelSystem prompt, messages incl. tool results and media, tool definitions, output incl. thinkingUp to 128K output on 1M models
DeepSeek V4 generation1M context lengthInput and output tokens384K max output
Moonshot Kimi1M (K3) or 256K (K2.6, K2.7 Code)Per-request context, model-dependentModel-dependent
Qwen3 generation32K–128K native by model sizeInput and training-extended contextThinking budget user-controlled

Two conclusions follow that the rest of the chapter depends on. First, there is no universal unit called “the context window” across providers; there are per-model accounting rules, and a budget built for one is miscalibrated for another. Second, in every dialect, the headline number overstates the input allowance, because output, reasoning, and runtime-added material all draw from the same pool. The outer wall is real, but nobody lives at the outer wall.

Three budgets, not one

With the headline number deflated, the chapter can state its central model. Context engineering works under three distinct budgets that must not be conflated:

Hard capacity is the provider and model limit: the outer constraint beyond which the request is rejected or truncated. It is set by parties outside your control and changes with model versions.

Usable budget is the portion you are actually willing to allocate to input material after reserving headroom for output, reasoning, runtime additions, and safety margin. It is a policy decision, not a provider fact, and it is always smaller than hard capacity, often much smaller.

Economic budget is what the chosen context costs in money and latency. It is governed by token counts transformed through provider pricing, caching behaviour, and the latency characteristics of long inputs. Identical token counts can have very different economic footprints.

A useful accounting picture of the usable budget looks like this. Treat it as an engineering model for reasoning about allocation, not as a description of any provider’s API semantics:

total available capacity
        │
        ├── standing context
        │     system / harness / tools / project rules
        │
        ├── task context
        │     current request / current files / evidence
        │
        ├── accumulated context
        │     history / previous tool results / summaries
        │
        └── reserved headroom
              output / reasoning / runtime needs

Each branch has different dynamics, which is why the decomposition earns its place. Standing context is paid once per session in content but once per invocation in tokens: it recurs. Task context is paid per task and is where relevance is highest. Accumulated context grows without bound unless something stops it, which makes it the branch that eventually breaks every unmanaged budget. Reserved headroom is spent on nothing visible and protects everything: it is the capacity you deliberately leave empty. Roughly mapped onto the Chapter 2 instrument, standing context is the stable instruction, tool-definition, and project-rule categories; task context is the current request, files, and evidence admitted for this task; accumulated context is history, tool results, and summaries. The mapping is approximate, which is why capture keeps the raw parts and the branches are derived from them afterwards; the two views must reconcile, since branch totals are sums over captured parts.

A worked budget: one session, three turns

The accounting model becomes concrete with numbers. Consider a repair trajectory on a model with 200,000 tokens of hard capacity, against which the harness sets a usable ceiling of 120,000 input tokens, reserving the rest for output, reasoning, and growth. The figures below are synthetic, labelled as illustration the way Chapter 2 labels its capture table; a real budget table carries measured counts from the observer.

Horizontal stacked bars for turns 1, 6 and 12 of a synthetic session. Standing context stays at 8,400 tokens while task and accumulated context grow, so total input rises from 22,600 to 68,800 to 113,800 tokens, reaching 95 per cent of a 120,000-token usable ceiling while still well inside a 200,000-token hard capacity.
One session against two limits. Standing context never changes, accumulated context starts at zero and ends as the largest share, and the usable ceiling is nearly reached while hard capacity still looks abundant. Illustrative synthetic example from this chapter, not an experimental result.

Three readings follow directly. First, standing context never changed yet was paid twelve times: 100,800 tokens for 8,400 tokens of content, the recurring tax made visible. Second, accumulated context, zero at the start, is the majority by turn twelve; any budget policy that scrutinises only what is admitted per turn misses the branch doing the damage. Third, the session approaches the usable ceiling while hard capacity looks abundant: at 114,000 input tokens against a 200,000 window, the transcript of a capacity story would report “57 per cent full” while the budget story reports 95 per cent committed with the trajectory still growing. The usable ceiling, not the headline, is the number the harness should watch, and the instrument of Chapter 2 is what watches it.

Note what the illustration does not claim. It does not say turn twelve behaved badly; behaviour is Chapter 5’s subject. It says the accounting has a shape, that the shape is predictable from the branch dynamics, and that a team tracking only the headline number cannot see it coming.

Standing costs and marginal costs

The accounting picture implies the chapter’s most practical distinction: between what context costs to have and what it costs to add.

A standing cost recurs every invocation for as long as its source is configured. Tool schemas, project instruction files, stable system material. A tool definition costing 500 tokens sounds cheap in isolation. Across a forty-turn agent trajectory it costs 20,000 tokens of input, every token competing with task material and every token billed. Standing costs are the recurring tax of capability: each tool added, each project rule appended, each harness preamble extended is a charge levied on all future invocations, usually approved by someone who considered only a single turn.

A marginal cost is paid when new material enters: another file read, another tool result, another conversation turn, another retrieved document. Marginal costs feel like the real spending because each admission is a visible decision, but they differ in kind. A 10,000-token file loaded once for the task it serves is expensive and possibly justified; the same 10,000 tokens re-sent unchanged for thirty turns is a standing cost wearing marginal clothing. The instrument from Chapter 2 exists precisely to tell these apart: stability and repetition, computed from consecutive captures, separate the recurring tax from genuine accumulation, and no budget policy is sound without that separation.

The distinction prepares later chapters without teaching them. Cache-aware layout (Chapter 9) is largely the art of keeping standing costs stable and early so their price is paid once. Tool-context management (Chapter 17) is the art of questioning whether every standing tool definition earns its recurring tax, a question Anthropic’s documentation now addresses directly with deferred tool loading. Pruning and externalisation (Chapters 10 and 13) operate on accumulated context, the branch whose marginal costs compound. Each of those chapters inherits this vocabulary; none of them is needed yet.

Per-source accounting: enforcing the budget

A budget nobody tracks is a wish. The enforcement mechanism is per-source accounting: every token in the rendered bundle attributed to the subsystem that placed it, as far as the capture of Chapter 2 allows it to be attributed. The ledger lines map directly onto ownership. Standing lines belong to the harness team: system material, tool schemas, project-rule injection. Nobody approves those lines per turn, which is why they need periodic review against behaviour rather than per-turn scrutiny. Accumulated lines belong to the trajectory: history growth, tool-result volume, summary churn. Those need per-session watching, because they are the lines that break budgets mid-task. Task lines belong to the current admission decision and are judged case by case.

The discipline this enables is differential, not uniform. A ten per cent overrun driven by standing costs calls for configuration surgery: remove a tool, shorten a preamble, defer a schema. The same overrun driven by accumulated context calls for runtime machinery: pruning, externalisation, compaction. Treating both as “the context got big” prescribes the wrong cure, and the ledger is what prevents that conflation. When later chapters propose their mechanisms, each will name the ledger lines it moves. A mechanism that cannot name its lines has not understood its costs.

Why systems leave the window deliberately unfilled

If hard capacity is the outer wall, why not build right up to it? Because several claimants on the pool cannot be measured in advance, and prudent systems reserve against all of them:

  • Output headroom. The response must go somewhere. Agentic turns that call tools need room for multiple response-and-result cycles, not one short answer.
  • Reasoning allowance. Where the model thinks before answering, thinking tokens draw from the same pool. Qwen’s user-controlled thinking budget makes the trade explicit; elsewhere it is implicit but equally real.
  • Runtime-added material. Harnesses inject content after your assembly: budget tags, warnings, formatted tool results, compaction output. Anthropic’s context-awareness feature, which injects remaining-capacity warnings into the request, is itself a consumer of the capacity it reports on.
  • Tokenisation variance. Token counts are estimates until the provider’s tokeniser runs, and even then a count is declared, estimated or provider-reported, three different things (Chapter 2). A bundle measured at 98 per cent of capacity with one tokeniser may overflow under another.
  • Future turns. In a multi-turn trajectory, filling the window now borrows against later turns. History, tool results, and summaries have not arrived yet but certainly will.
  • Safety margin. Providers reject or halt at the boundary with varying grace. Margin converts a cliff into a slope.

The usable budget is therefore hard capacity minus these reservations, and the reservations are scenario-dependent: a single classification call needs little headroom, while an open-ended repair trajectory needs a great deal. “What fraction of the window should we use?” has no universal answer, which is exactly why it is an engineering decision rather than a configuration default. A team that sets its ceiling at 60 per cent of hard capacity for agent trajectories is not wasting 40 per cent. It is pricing the future.

Token count is not cost

The third budget needs its boundary marked, briefly, because the confusion is expensive. Two bundles with identical token counts can differ severalfold in price and latency. Cached prefixes are billed at reduced rates on the providers that offer caching; uncached identical tokens are billed in full. Longer inputs increase time to first token regardless of price. Output tokens are priced above input tokens in every catalogue surveyed. None of this changes what fits in the window, and all of it changes what the window costs to use.

Anthropic’s documentation states the relationship crisply: cached prompt prefixes still occupy the context window, so caching changes what you pay for tokens, not whether they count. Capacity budget and economic budget are different objects governed by different rules, and optimising one while ignoring the other produces the familiar twin failures: the bundle that fits but bankrupts, and the bundle that economises itself into uselessness. The full mechanics of caching belong to Chapter 9. Here the point is only that the budget has three faces, and a decision recorded against one must be checked against the other two.

The allocation problem

The chapter’s main intellectual move can now be stated plainly:

The engineering problem is not how to fill the window, but what deserves a share of the budget.

In schematic form:

candidate information
        ↓
competes for
        ↓
bounded budget

It is tempting to formalise this as utility per token and reach for optimisation machinery. In its simplest sketch, each candidate item i would carry an estimated value v(i) and a token price p(i), and assembly would maximise value under the usable budget. The sketch is useful as a way of thinking and dangerous as a blueprint, because every interaction in the list above breaks one of its assumptions. A concrete case: two files where the second defines the terms the first uses. Admitted together they are worth more than the sum of their separate values; admitted separately the first is near worthless. No per-item score survives that dependency, and real bundles are dense with such couplings: instructions that scope tool use, examples that disambiguate rules, history that gives a retrieved file its meaning. Chapter 22 returns to this problem and shows why ranking cannot solve it, and Chapter 23 builds machinery that treats dependencies, authority, position, and fidelity as first-class inputs rather than corrections to a score. This chapter’s contribution is making that machinery necessary: once allocation is the problem, naive scoring is visibly insufficient.

Dependency is one interaction among several. Instructions carry different authority, so identical tokens do not have identical standing. Position changes effect, so the same item has different value in different slots. Representations differ in density and fidelity, so token count is a poor proxy for information content. Some information must remain exact while other information may degrade, so items have different loss tolerances. Some context creates dependencies, where admitting an item commits future budget to its consequences. And cache layout changes the economics, so the price of an item depends on its neighbours. A knapsack with interacting, authority-weighted, position-sensitive, fidelity-graded items whose prices depend on ordering is not a knapsack at all. It is the rest of this book.

What Chapter 4 establishes is only that the allocation problem is unavoidable. Every layer from Chapter 3’s stack is a claimant. The budget is smaller than the headline. The claimants interact. From here, filling the window to the brim looks less like thoroughness and more like abdication: letting every claimant take what it wants and hoping the result works.

Proposed experiment

The question. Does cost scale with admitted material the way the budget model predicts, and at what volume do the reservations for output, reasoning and future turns come under pressure?

The design, in brief. Fix the model, the settings and a task with a small critical evidence set at a fixed position. Add irrelevant material in five conditions: critical evidence only; plus small, medium and large irrelevant context; and plus a large amount of plausible distractors.

The measurement that matters. Rendered input tokens by category, reconciled against the total; utilisation against hard capacity; and, for each condition, how much headroom remained for the answer and for a projected ten further turns at the observed growth rate. Outcome is recorded as a covariate for reading cost, not as the verdict; whether added material helps or harms belongs to Chapter 5.

What would change the book. If category totals cannot be reconciled with the rendered total, or if standing costs cannot be told apart from marginal ones by byte stability, the budget vocabulary has not earned its place. Nothing here has been run. The full design is in the companion repository’s experiment designs.

The question that remains

Suppose the accounting is done, the reservations are honoured, and the request sits comfortably below the hard limit with headroom to spare. The budget view is satisfied. Is the extra context then harmless? Is a bundle that fits necessarily a bundle that works?

The next chapter shows why the answer is no:

Suppose the request is still comfortably below the hard limit. Is more context harmless?

References