Context Is a Bottleneck
Remembering more than fits forces assembly, not just selection.
Chapter 10 took selection as far as a frozen run has taken it: which memories the present work needs, admitted by an explicit, traceable policy at a fixed budget.
This chapter shows selection is not enough. Admission assumes the admitted memories fit. Increasingly they do not — and the Chapter 10 run puts a number on the shortfall.
The ledger oracle reaches perfect required-evidence recall on a mean of 587 estimated tokens. The best non-oracle Chapter 10 condition spends 1,161 to reach 0.902 recall. Rendered with the source headers and validity marks the reader actually sees, those become roughly 805 against 1,505. The gap survives rendering.
The auditable oracle defined in this chapter — content plus the disagreements, licence bearers, and marks required by this book’s audit contract — stands near 696 counted tokens, still roughly 40% below the practical bundle.
The oracle gap shows substantial headroom under a correct frame and a good candidate pool, while the practical bundle’s distractor admission sits near 0.45. It does not establish that every token above the oracle minimum is useless.
This chapter therefore holds the selected set fixed and asks a different question: what happens between admission and the exact context handed to the model?
The system now holds events, beliefs, provenance chains, validity intervals, intentions, derived loops, rules, and prior failures. The obvious solution — put all the relevant memory into the prompt — fails in ways that are structural rather than incidental.
Context is a bottleneck. What passes through it must be assembled, not merely selected.
Why “put it all in” fails
Take the compatibility-release task, with everything the book’s machinery now considers relevant:
- The migration decision, with its support chain.
- The backup debt — since completed, so labelled historical rather than seated as live.
- The docs scope condition.
- The fixture warning — likewise historical.
- The reversibility rule.
- The caller migrations, with their trap annotations.
- The contracted facade that must not be touched.
Each item has a defensible reason to be selected. Together they exceed every fixed budget this chapter tests.
And the failure modes of stuffing them all in are several distinct problems — not one problem called “too long.”
Finite window and cost. The hard limit is obvious: context windows end, and tokens cost money and latency on every task. But the soft limits bite first. Every additional memory raises the price of every future task that carries it, so a policy of admitting everything relevant taxes the system continuously for the possibility of need.
Distraction and dilution. As admitted context grows, the task itself occupies a shrinking fraction of what the model must process. The mechanism here is stated carefully as hypothesis, not fact: beyond some load, task-critical instructions compete with background memories for influence over the output, and the spec’s distraction cost exists to measure the effect rather than assume it. No broad claims about attention mechanisms appear here; what needs measurement is behavioural — does constraint adherence degrade as admitted context grows at fixed relevance?
Duplication. The decision record, the runbook echo, and the August restatement all carry the same belief in different words. Admitting all three spends budget triplicating one fact while crowding out the fixture warning. Deduplication is not retrieval’s job — each copy is genuinely similar to the query — it is an assembly operation over the admitted set.
Contradiction in context. Historical background and current belief disagree by design: the March SQLite passages sit beside the July migration. Admitted raw, without their temporal resolution attached, they present the model with a contradiction the system already resolved and ask it to resolve it again, under task pressure, from wording. The Chapter 8 resolution must travel with the memory, not stay behind in the store.
Staleness smuggled in. A memory admitted for background — the spent 2023 discussion — carries claims that were true then and false now. Without validity marking, background becomes misinformation. The assembly must label what selection admitted: current, historical, superseded, uncertain.
Lost task focus. The subtlest failure: a context dominated by history reframes the task as summarization. The model answers what happened fluently and never performs what is needed. The task statement drowns in its own background.
Each failure motivates a different assembly operation, which is why the chapter separates them instead of prescribing “compression” as a single fix.
Assembly is where the bottleneck actually bites:
flowchart LR
SEL["Selected evidence<br/>(more than fits)"] --> ASM{{"hard token budget"}}
ASM -->|kept| DEC["Decisive evidence<br/>preserved"]
ASM -->|"dropped, each with a reason"| CUT["Everything else"]
DEC --> READER[Reader]
style CUT fill:#ffffff,stroke:#c0392b,color:#c0392b,stroke-dasharray:3 3
When the admitted set exceeds the fixed budget, something must be dropped. The important question is whether the system records what it dropped and why.
Retrieval is not assembly
The separation this chapter establishes, matching the spec’s failure-attribution categories:
Retrieval asks: what candidates might matter?
Context assembly asks: what exact representation of memory should this model receive for this task?
Selection ends with a ranked set. Assembly turns that set into a prompt: which items survive, in what order, grouped how, deduplicated against what, compressed to what degree, cited to what depth, with uncertainty marked where. Two systems with identical selection and different assembly can behave differently on the same task, which makes assembly a layer with its own measurements rather than a formatting detail.
Possible operations, separated here so their effects can be measured independently:
- select — the upstream policy output, established in Chapter 10 and constrained by Chapter 13’s frame gate;
- order — task-critical first, background last, traps flagged rather than buried;
- group — the support chain travels with its belief; the loop travels with its triple;
- deduplicate — echo documents collapse to one representative with a count, not three admissions;
- compress — shorten while preserving provenance; durable compression belongs to Chapter 15 rather than this assembly layer;
- cite — every assembled claim traceable to store and artifact, the Chapter 7 discipline at the point of use;
- expose uncertainty — calibrated marks where state is genuinely unresolved;
- preserve provenance — chains and triples survive assembly rather than dissolving into prose.
The list stops short deliberately. Consolidation, durable compression and forgetting now meet as one long-term growth problem in Chapter 15; outcome adaptation and reusable procedures meet at the memory-learning boundary in Chapter 16.
The experiment
Selection is frozen at Chapter 10’s C5 admitted set — eleven tasks, the same bundles the frozen run recorded.
C5 is the input rather than C6, deliberately. C6 adds coalescing and precision falls from 0.474 to 0.464 at equal recall, so C6 belongs in the comparison, not in the foundation.
Conditions vary only what happens after admission: raw render, reordering, extractive dedup, grouping, validity marking, budget dropping with a reason recorded per drop, and all of them composed. Against those run three references — the ledger-content oracle, the auditable oracle with disagreements and licence bearers restored, and frozen C6 — plus empty and random-drop controls.
The counted-token budget sweeps from 384 through 1536 estimated tokens, plus an unconstrained condition; rendered context can be larger because reader-visible marks add tokens.
The frozen run carries zero model calls for the evidence suite, plus a bounded 55-call reader probe through the frozen Chapter 10 reader and key-claim scorer.1
At unconstrained budget, over the eleven tasks:
| Condition | Required recall | Counted tokens | Rendered tokens | Contradiction kept | Licence kept |
|---|---|---|---|---|---|
A0 raw | 0.902 | 1,174 | 1,505 | 1.00 | 1.00 |
A2 dedup | 0.902 | 1,156 | 1,480 | 1.00 | 1.00 |
A6 composed | 0.902 | 1,156 | 1,747 | 1.00 | 1.00 |
CO-content | 1.000 | 587 | 805 | 0.00 | 0.40 |
CO-auditable | 1.000 | 696 | 1,060 | 1.00 | 1.00 |
C6 (Ch 10) | 0.902 | 1,161 | 1,644 | 1.00 | 1.00 |
Three facts stand out.
Dedup saves eighteen counted tokens. Genuine echo redundancy is nearly absent from the admitted sets, because selection already declines to admit most of it. The bottleneck is not repetition.
The content oracle is not a fair ceiling. It drops contradiction entirely (0.00) and most licences (0.40). The auditable oracle restores them for 109 tokens.
Marks are not free in the render the model sees. The composed policy’s rendered cost (1,747) exceeds raw (1,505), even as its counted cost falls. Audit metadata and reader-visible tokens are different budgets — and the trace keeps the full lineage either way.
The budget sweep is where assembly earns or loses. Required recall against budget:
| Budget | A6 composed | Random drop | CO-auditable |
|---|---|---|---|
| 384 | 0.58 | 0.58 | 1.00 |
| 512 | 0.58 | 0.61 | 1.00 |
| 768 | 0.72 | 0.67 | 1.00 |
| 1024 | 0.82 | 0.79 | 1.00 |
| 1280 | 0.85 | 0.89 | 1.00 |
| 1536 | 0.90 | 0.90 | 1.00 |
At 768 tokens, the composed policy preserves structures that random dropping does not: contradiction preservation is 1.00 against 0.57, and licence preservation 1.00 against 0.38.
But the guarantee costs recall under pressure. At 512, random dropping keeps more required evidence (0.61 against 0.58) — because the policy spends scarce budget protecting disagreement and licences that the ledger does not score as required.
That trade is visible rather than hidden. Every drop carries its tier, its saving, and its cover.
One case at 768 shows both sides. A required architecture-spine unit falls with no cover at all. A prior open-loop result unit falls covered by the dissent it shares a group with — and that dissent itself survives.
The reader probe sharpens the picture without settling Chapter 12’s question. Mean key-claim coverage through the frozen reader, forbidden-claim rate zero everywhere:
| Condition | Counted tokens | Key-claim coverage |
|---|---|---|
A0 full | 1,174 | 0.879 |
A6 at 768 | 659 | 0.955 |
A6 full | 1,156 | 0.894 |
CO-auditable | 696 | 0.879 |
C6 | 1,161 | 0.939 |
A smaller assembled context (659 tokens) scores higher on key-claim coverage than the full raw context (1,174) — 0.955 against 0.879, and ahead of the frozen Chapter 10 C6 condition at 0.939 — while required-evidence recall at that budget is only 0.72. The two metrics therefore diverge: the ledger counts required units that the reader does not need in exactly that form on every task, and one task’s coverage moves with ordering alone. Chapter 12 supplies the behavioural interpretation; the headline here stays bounded to assembly, evidence preservation, and cost.
Sufficiency belongs to the reader too
Minimum sufficient context is not purely a property of evidence. It is a relation among task, evidence, representation, and reader. A strong-reader re-render reaches key-claim coverage 1.000 on raw, full assembled, and auditable contexts alike, and 0.955 at the compact 768-token budget.2 In this probe, assembly therefore acts primarily as an efficiency and bounded-context mechanism rather than a capability unlock. One stronger reader does not establish a universal scaling law: sufficiency remains reader-relative, and assembly should be evaluated against its intended reader.
Policy revision follows the Chapter 10 and 11 discipline. The tempting repair — treat contradiction as ordinary budget weight when a long dissent crowds out supporting evidence — is proposed as an immutable drop-policy version, replayed over the suite at budget 512, and rejected: it saves 9.2 tokens and even nudges recall, but contradiction preservation collapses from 1.00 to 0.14. A cheaper context bought with lost disagreement stays out.
Book result. Assembly earns a distinct layer with a Type B profile and a Type C headline. Simple extractive dedup saves only 18 of 1,174 counted tokens, so repetition is not the main bottleneck on these admitted sets. Under a hard budget, the composed policy preserves contradiction and licences where random dropping does not, and at 768 tokens it scores higher than the full raw context on the existing key-claim scorer. The fair oracle gap is real but narrower than Chapter 10 suggested: 696 auditable tokens against 1,156 assembled (1,174 selected), with required recall 1.00 against 0.90. Demotion clauses: controlled project fixtures, one primary reader, estimated token counts, C5-admitted evidence only, no generative compression, which Chapter 15 owns.
Where the middle of the book lands
Fourteen chapters now form an investigation progression rather than a runtime stack:
retrieval
→ reconstruction
→ belief
→ unfinished intention
→ situational relevance
→ active context
Look at what the system can now do. It retrieves history. It reconstructs what happened. It justifies its beliefs and tracks them through time. It maintains what remains undone, including consequences nobody stated. It selects what the present task needs by explicit policy, and assembles that for a bottlenecked context.
What it has not yet earned is the durable half of the bottleneck problem. Assembly can discard, reorder, group, or shorten material for one execution, but it does not decide what should survive across tasks in a smaller form.
The system has not yet established a long-term policy for durable compression or forgetting, nor a mechanism by which evaluated outcomes should alter what is retained or become reusable procedure.
A system can now remember too much. The next problem is deciding what to preserve, what to compress, and eventually what to forget.
Three boundaries hold for what comes next. This chapter shows that a smaller assembled context can preserve contradiction and licences while scoring at least as well on the existing answer-level probe, even though some ledger-required evidence is dropped; richer downstream project behaviour — constraint adherence, failed-approach avoidance, open-work continuation — is measured by Chapter 12’s separate instrument. Wrong-frame behaviour under a tight budget is Chapter 13’s territory; the headline experiment uses declared, correct WorkFrames throughout. And nothing here rewrites the durable store: every compact representation is derived, ephemeral, and rebuildable, while long-term storage compression remains Chapter 15’s problem.
What long-context research adds to assembly
LongBench, RULER, and InfiniteBench measure long-context behaviour across tasks and lengths rather than inferring it from an advertised window. Lost in the Middle (arXiv preprint 2023) makes position an experimental variable: the same evidence can have different effects depending on where it appears.
An assembler should therefore record inclusion, position, representation level, token cost, and attached validity/provenance marks. The assembly trace can then distinguish exclusion, dilution, placement, and representation failures.
The papers show that long inputs can remain difficult; they do not prove that grouping, deduplication, or validity marking repairs them. That is why the experiment above isolates assembly operations under matched budgets and reports performance against load instead of treating them as unconditional improvements.
Research foundations
That evidence supports treating context as a limited behavioural resource, but it does not establish this chapter’s particular assembly operations. The matched-budget experiment above provides that evidence at fixture scale. Two recent results sharpen the design without supplying the mechanism.
SARA (Jin et al., ACL 2026) optimises RAG under fixed token budgets by pairing a small set of text passages with compressed semantic vectors, reporting answer gains from the hybrid.
The question it leaves this book is whether structured project memory — support groups, validity intervals, triple licences — already supplies the coverage information SARA buys with vectors. This chapter keeps every representation extractive and inspectable, rather than finding out the other way.
Adaptive-k (Taguchi et al., EMNLP 2025) shows fixed passage counts either waste tokens or omit evidence, and selects a query-specific count by a single-pass score threshold.
That is adaptive sizing — the complement of this chapter’s fixed-budget fitting. The sweep above measures quality against budget first. Adaptive budgets stay reserved until a run earns them.
References
- LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding (2023).
- Lost in the Middle: How Language Models Use Long Contexts (2024, TACL; arXiv preprint 2023).
- RULER: What’s the Real Context Size of Your Long-Context Language Models? (2024).
- InfiniteBench: Extending Long Context Evaluation Beyond 100K Tokens (2024).
- SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression, ACL 2026.
- Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive-k, EMNLP 2025.
Frozen run
ch14-20260920T174259Z-context-assemblyfromsolution/context_frames/assembly.py(assembly-v1, drop policydrop-preference-outside-first-v1); reader probe through the frozen Chapter 10 reader (llama3.1:8b). ↩︎Strong-reader re-render
sr3-ch14-20260920T231031Z-muse: frozen renders, no reassembly, budget held fixed. ↩︎