← Context From First Principles

Bring Back Only What You Need

Use retrieval as a context-admission mechanism without turning this book into Retrieval From First Principles.

Chapter 13 left the model holding a tidy live context and a shelf of external artifacts:

Incident 17 — lock-order inversion.
Evidence: artifact://incident-17

plus five more references of the same shape. Then the current task changes. The user asks why the deadlock appeared only under concurrent import. The incident artifact probably matters. So might the benchmark report, the migration plan, and the compiler run. The naive response is to load everything, which recreates within one turn the exact occupancy problem Chapter 13 solved. Leaving was only half the problem. Existence plus recoverability does not mean admission. This chapter decides what comes back.

Retrieved is not admitted

The chapter’s central equation, protected throughout:

retrieved
   ≠
admitted
   ≠
used
   ≠
helpful

Four different claims. An artifact surfaced by some search is a candidate. A candidate inserted into the bundle is admitted. Admitted material the model actually references is used. And helpful means the task improved because of it, which no earlier stage guarantees. The pipeline this chapter owns is the middle of that chain:

    flowchart TD
    X["External information"] --> G["Candidate generation<br/>other mechanisms, not taught here"]
    G --> C["Candidates"]
    C --> A("Admission<br/>this chapter")
    A -->|reject| R["Left out"]
    A -->|admit| B["Context bundle"]
    A -->|expand| B
    B --> M["Model"]
    M --> H["Behaviour<br/>measured downstream"]
    classDef outside stroke-dasharray: 5 4
    class G,H outside
  

Everything above the candidate set is some other mechanism’s business: exact identifier lookup, file lookup, text search, BM25, embedding retrieval, database query, graph lookup, tool search, agent navigation, memory recall. The book does not teach these, compare them, or rank them. They all end at the same boundary, a candidate set, and Chapter 14 begins there. Everything below admission is behaviour, measured downstream. The chapter lives in the admission box: which candidates enter, in what representation, at what cost, and whether the entering material helped.

Candidates are not context

Candidate generation identifies information that may be relevant to a computation. Admission determines which representation of those candidates is actually allowed into the context bundle.

A candidate, in the sense the rest of the book uses, is not a document. It is one representation of one piece of information, together with what a later decision needs to know about it: what it costs, whether it is eligible, what it depends on. Chapter 12 showed that one content can exist in several forms. Each form is a separate candidate, and admitting one of them is a different act from deciding the content is wanted. The word matters because the mechanism of Chapter 23 takes exactly this: a pool of candidates that somebody else built, and no say in how they were found.

The distinction matters because each stage fails differently. Imagine ten external artifacts, three containing information the current task requires. The generator returns seven: the three required plus four irrelevant. Candidate generation achieved perfect recall. Admitting all seven may still fail the computation through bloat, interference, position pressure, and cost. Retrieval recall is not admission precision, and a mediocre generator feeding a strong admission policy can outperform a strong generator feeding no policy at all. The primary experiment therefore freezes the candidate set deliberately: construct candidate pools directly, hold generation constant, and vary only admission. If the chapter changed the retriever and the gate at once, the result would be uninterpretable.

A later pipeline can fail at either stage, and the vocabulary keeps them apart. A candidate miss means the required artifact never entered the candidate set; later retrieval work owns that. An admission miss means the required artifact sat in the candidates and was rejected or never expanded; context engineering owns that. The frozen-candidate experiment studies the second failure with the first held fixed, and an optional ecological extension runs one simple real generator afterwards to check the admission policy still behaves under imperfect candidates. That extension is subordinate. The chapter does not depend on it.

The cost of admitting everything

The experiment’s preload-everything condition is the baseline every other policy must beat, and the reasons it can lose are now measurable rather than atmospheric. Candidate-level precision understates the damage. One relevant 10-token artifact plus one irrelevant 20,000-token artifact is 50 per cent candidate precision and near-total token waste. The chapter’s most important metric is therefore token-level: irrelevant admitted tokens counted alongside required admitted tokens wherever oracle labels make the split meaningful.

That irrelevant tokens merely waste money would be a budgeting complaint. The evidence says they can actively harm. Amiraz and colleagues formalise the distracting effect of an irrelevant passage with respect to a query and a model, measured as the probability that the model fails to abstain when the passage cannot answer the question, and show the scores correlate strongly across seven tested models from 3 to 70 billion parameters. Their headline behavioural result is the one this chapter imports: on the Natural Questions benchmark, adding a hard distracting passage alongside the gold passage drops answer accuracy by 6 to 11 points depending on the model, and the degradation persists at 70-billion scale. The figure is from the full paper, not its abstract. Two of their secondary findings sharpen the point for admission design. Higher-ranked irrelevant passages are more likely to distract, and adding a reranking stage amplifies the effect: the passages that fool the pipeline are also the passages that mislead the generator. Fixed top-k admission over a strong ranker is therefore not a neutral default. It preferentially admits exactly the material most likely to distract. Their population is question answering over retrieved passages, not agentic artifact admission, and the chapter claims nothing beyond the narrow finding: retrieved-but-irrelevant passages can degrade generation, and how much depends on how distracting the passage is, not merely on its irrelevance.

Over-admission can thus resurrect everything Chapter 5 buried: bloat, interference, and position pressure without ever touching the hard context limit. Externalisation solved residency. Admission decides whether the solution survives contact with the next task.

Relevance is not utility

Similarity scoring ends at the candidate boundary, and the chapter states the boundary without re-teaching embedding geometry. A document can be highly similar to the task and still be redundant, outdated, lower-authority, too broad, or simply not required; a dissimilar artifact can hold the one exact constraint the task needs. Retrieval score is not context utility. Top-ranked is not authoritative either: a project note is not a user instruction, and Chapter 14 preserves provenance and authority metadata through admission without adjudicating conflicts. Chapter 19 owns precedence. This chapter keeps the metadata Chapter 19 will need, including which source each admitted section came from, so no chunk arrives as anonymous prose.

Utility is also marginal, not per-document. Two independently correct documents can repeat each other, consume capacity, shift positions, and increase reasoning burden. Consider a migration task with three retrieved candidates: the migration plan stating tenant 042 needs manual cutover, a progress note restating the same cutover requirement, and the compiler run showing the current build state. The first document carries the decision. The second adds no new fact the task needs; admitting it anyway spends tokens, shifts the plan’s position, and doubles the model’s opportunity to misread the tenant identifier. The experiment tests one sufficient source against three redundant ones without building a major new axis, to establish the working definition:

Admission utility is task-relative: the behavioural value of adding a representation to the current bundle, net of the context resources and interference it consumes.

No scalar score is constructed from this. Task success, tokens, latency, cost, and interference stay separate dimensions for the eventual compiler to trade off. Admission is marginal context utility, not document relevance alone.

How much of a candidate should enter

A candidate artifact is not atomic. A 30,000-token investigation may owe the current task one decision, one exact identifier, or one evidence section, and the admission choices range across identifier only, title plus metadata, semantic anchor, specific section, selected evidence span, and full artifact. That range is admission granularity, and it is not Chapter 12 fidelity wearing new clothes:

FIDELITY
How detailed is this representation?

GRANULARITY
How much of the candidate does this representation cover?

A full-fidelity hundred-line section and a low-fidelity whole-document summary differ on both axes at once. The experiment manipulates one at a time: a granularity arm compares full artifact against relevant section against short typed extraction under matched informational requirements, asking whether admitting the whole candidate imposes unnecessary cost relative to the needed part. Where the external artifact already carries Chapter 12’s tiered representations, admission consumes them, choosing which level enters for this task without rebuilding the fidelity policy. And Chapter 7’s retention semantics ride along: a candidate holding exact constraint material admits it exactly, because no admission pressure compresses anything below its legal fidelity.

Admission also composes with what is already resident, and the composed states deserve names because the experiment will produce them constantly. A resident anchor plus an admitted section is the normal working shape: eighty tokens of orientation that never left, joined by the two hundred lines the current question actually needs. A resident anchor plus an admitted full artifact is the escalation shape, reached only through EXPAND after the section proved insufficient. Two admitted sections from the same artifact are one source admitted twice, and the token ledger should show it that way rather than as two independent wins. The admission record keeps these compositions visible so evaluation can distinguish a policy that leans on resident anchors well from one that re-admits what the bundle already knew.

Progressive disclosure

The natural strategy these pieces compose into is progressive disclosure: identity first, then anchor, then section, then full artifact, each step admitted only when the current material proves insufficient.

IDENTITY
artifact name / title / metadata
        ↓
ANCHOR
short semantic description
        ↓
SECTION
targeted material
        ↓
FULL
complete artifact

Chapter 12 pre-generated multiple resident representations and selected among them over time. Chapter 14 progressively admits more information from a non-resident candidate. The direction of travel is opposite — fading out versus drawing in — and the chapter keeps the terms apart for exactly that reason.

Disclosure is not free, and the chapter prices both sides. Lean starts buy lower initial occupancy and less irrelevant admission; they pay additional tool calls, latency, extra reasoning steps, and two characteristic failures. The first is premature stopping: the model sees enough to believe it understands and never loads the necessary detail. The fixture builds this deliberately, a candidate whose anchor reads sufficient while its detail section carries the decisive qualification — an incident anchor recording the fix with a detail section confining it to the import path while the export path still deadlocks. Required expansion available but not requested is scored as an admission miss inside progressive disclosure, plainly named and unbranded. The second failure is wandering: expanding A, B, and C, searching D, reading E, after the correct artifact was already identified, at the cost of latency, calls, and context. Anthropic’s context-engineering guidance names this risk directly, warning that agents navigating on demand can chase dead ends or miss key information without good heuristics. The chapter takes that as engineering evidence for the failure mode, not as law about its frequency.

Whether the model or the runtime drives disclosure is an architecture choice the chapter compares without crowning. In system-selected admission a deterministic gate chooses candidates and expansions; in agent-selected admission the model chooses what to inspect next. The experiment runs one of each, which also sets up the handoff Chapter 15 owns: an agent that retrieves one misleading candidate may form the hypothesis that sends it after supporting material, a self-reinforcing loop this chapter names and leaves alone.

Eager admission and just-in-time admission inherit the same even-handed treatment. Eager preloading buys low round-trip latency and immediate evidence at the price of over-admission; just-in-time starts lean at the price of steps, misses, and wandering. The trade is sharpest where the task’s information needs are predictable: a debugging task that always needs the failing test output plus the suspect module rewards eager admission of both, while an exploratory task with six plausible artifacts rewards starting from identities. Anthropic’s guidance describes production hybrids, stable or critical material preloaded with the rest explored on demand, and the chapter derives the same shape from its own costs before citing it: preload what every near-future computation needs, disclose the rest. Final assembly belongs to Chapters 22 and 23.

Fixed top-k gets one short section because it is really a policy wearing a parameter’s clothes. Retrieving five and inserting five couples ranking and admission into a single number with no theory behind it. The right count varies with artifact size, task complexity, candidate redundancy, evidence diversity, context pressure, and retrieval confidence, and k knows none of these. The chapter builds no dynamic optimiser. It establishes the coupling and moves on.

Where admission lives

The decisions of this chapter do not all live in the same place, and the word admission has been used for more than one of them.

Deciding which candidates from a pool enter a bundle, in which form, under a budget and a set of rules, is an assembly decision, and Chapter 23 gives the mechanism that makes it. It is deterministic and works on a pool it is handed.

Deciding whether to fetch more is not that. Progressive disclosure, expanding an anchor into a section when the section proves necessary, is behaviour of the system around the assembler: it looks at what the model did, notices that the material was insufficient, and asks for a richer representation. So is deferring a candidate until there is evidence it is worth the tokens. Those are runtime behaviours. The compiler has no such decisions. It admits or it does not, once, for one computation, and anything that wants a second look calls it again with a different pool.

Three identity rules keep the pipeline honest wherever the decisions live. Stable artifact identities from Chapter 13 survive from candidate generation through expansion to evaluation, so results attribute to artifacts and not to fuzzy phrases. Admitted sections keep their source: an artifact with a named section, never anonymous prose. And fragment identity is not artifact identity: five chunks retrieved from one artifact are one source admitted five times, and counts that hide that mislead. Recovery integrity also comes before admission quality: did the reference resolve to the intended artifact and version is asked before was this the right candidate to admit. A failed lookup is a recovery failure, never a ranking failure.

Ordering, cache position, freshness and scope are each somebody else’s, as earlier chapters said and later ones repeat. Memory candidates and tool outputs enter as candidates like any others.

Proposed experiments

The question. Over a frozen candidate pool, does an admission policy beat loading everything, and beat a fixed top-k, once irrelevant tokens are counted?

The design, in brief. Hold the candidate pool, its ranking, the model and the resident baseline fixed, and vary only admission. A pool of a dozen to twenty artifacts holds two required items, two helpful but unnecessary, four plausible distractors, and irrelevant material, with hidden labels no policy sees. The conditions are: no external admission; preload everything; fixed top-k chosen in advance; an admission gate over the same pool; identifier-first progressive disclosure with targeted expansion; and an oracle that admits the minimum sufficient part. A granularity arm compares full artifact, relevant section and short typed extraction. Three fixtures carry the traps: a large plausible distractor, a small decisive candidate of low similarity, and a sufficient-looking anchor over a detail that reverses the conclusion.

The measurements that matter. Admission recall and precision against the oracle, and token-level accounting: required, distractor and irrelevant admitted tokens per task, so a policy cannot hide one expensive mistake in a favourable average. The cost of the admission machinery itself is counted, since a gate that saves 8,000 tokens by spending 6,000 reasoning about candidates has a net to report. Available but not admitted stays distinct from admitted but unused.

What would change the book. If preload-all matches the gate under realistic budgets, or top-k matches an elaborate gate, or irrelevant candidates cause no harm on the tested models, the gate is deleted and the chapter’s surviving contribution is the accounting that showed it. Nothing here has been run. The corpus of real sessions is empty, so how often ordinary work surfaces large externalisable artifacts, or re-reads them, is unmeasured.

What the agent creates next

Admission assumed a universe of candidates the agent moves through. The next chapter removes that comfort: agents do not merely retrieve context, they manufacture it. Plans, hypotheses, tool observations, summaries, and intermediate artifacts continuously create new future candidates, including the misleading ones whose retrieval can steer everything after. Retrieval generates possibilities and admission creates context; what happens when the agent starts generating the possibilities themselves is the problem Chapter 15 inherits.

References

  • Anthropic Applied AI team (Rajasekaran, Dixon, Ryan, Hadfield, et al.). “Effective context engineering for AI agents.” First-party engineering essay, September 2025, verified September 2026. Embedding and pre-inference retrieval; just-in-time context with lightweight identifiers; agentic navigation with progressive disclosure; hybrid preload plus runtime exploration; dead-end exploration risk. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
  • Jones, A., Kelly, C. “Code execution with MCP: Building more efficient agents.” First-party engineering essay, Anthropic, November 2025, verified September 2026. Progressive disclosure via filesystem-presented tools and search_tools with detail-level selection; filtering data before it enters context; keeping large intermediate results out of the model context. https://www.anthropic.com/engineering/code-execution-with-mcp
  • Amiraz, C., Cuconasu, F., Filice, S., Karnin, Z. “The Distracting Effect: Understanding Irrelevant Passages in RAG.” Preprint, arXiv:2505.06914, May 2025 (arXiv page lists a related 2025 ACL proceedings DOI). Distracting-effect measure with cross-model robustness; on Natural Questions, hard distracting passages cut accuracy 6–11 points even alongside the gold passage (full paper, Section 4.4); higher-ranked irrelevant passages distract more. https://arxiv.org/abs/2505.06914
  • boxpositron. “WithContext MCP Server.” Third-party implementation, MIT licence, verified September 2026. Search/read workflow with project-scoped notes; explicit write-out and read-back separation. https://github.com/boxpositron/with-context-mcp