← Context From First Principles

Assemble for the Task

Why ranking candidates and filling the budget does not produce a bundle, what a mechanism would have to do instead, and why the right output is sometimes a refusal.

The usable input budget is 4,000 tokens. The pile of things that must be present reads: the standing instructions, 1,100 tokens; the exact constraints, 1,400; the schema of a tool the task cannot proceed without, 900; the minimum evidence for the task, 1,200. That is 4,600 tokens against a ceiling of 4,000, and no ranking, weighting or cleverness changes the arithmetic.

A relevance packer would delete 400 tokens of whichever item scored lowest. Perhaps that is the second half of the exact-constraint block. Perhaps it is the tool schema’s required parameters. Either way it would emit a plausible bundle that quietly breaks the rules it was built to serve, and the failure would surface three turns later as an invalid call or a guessed identifier.

The correct output of this computation is not a bundle. It is a notice that names the 600-token excess. Deleting exact constraints to fit a budget is not optimisation. It is fabrication with good formatting.

Everything in that pile was eligible. Nothing could be assembled. And the first question, before any question about which candidates are best, was whether assembly was possible at all.

Five things people call the context

The previous twenty-one chapters have built up a large vocabulary for what happens to information before a model sees it. By now the word context is doing too much work. Five different objects sit between the information a system has and the request it sends, and they are not the same.

ObjectWhat it isWho or what decides it
Available informationeverything the system could reachthe world
Candidatesrepresentations of some of it, each with its cost, its eligibility judgements and its dependenciesthe systems that retrieve, remember, act and generate (Chapters 13 to 18)
The admitted setthe candidates that are allowed inthis chapter’s subject
The ordered sequencethe admitted set, arrangedChapter 6’s subject, applied
The rendered textthat sequence as exact bytes, with the labels and separators the format needsthe renderer

The last two together are what the rest of the book calls a ContextBundle: the exact, ordered material for one computation, and its text.

Each step can go wrong independently, and each step can be right while the next one is wrong. A perfect candidate pool can be admitted badly. A correct admitted set can be arranged so that a rule ends up behind the noise. A correct arrangement can render to more tokens than the budget allowed.

Two more objects appear after this chapter. The render is wrapped and delivered into a running session, and something else observes what actually arrived. Those are the subject of Chapter 25. For now the question is narrower: given a pool of candidates, a budget and a task, how does a system decide what goes in?

The obvious answer

The obvious answer is a ranking. Score every candidate for relevance to the task, sort, and take from the top until the budget is used.

function topK(candidates: Candidate[], budget: number): Candidate[] {
  const ranked = [...candidates].sort((a, b) => b.relevance - a.relevance);
  const chosen: Candidate[] = [];
  let used = 0;
  for (const c of ranked) {
    if (used + c.tokens > budget) continue;
    chosen.push(c);
    used += c.tokens;
  }
  return chosen;
}

It is short, it is fast, and it is how a great deal of production retrieval works. It also assumes four things, none of which the earlier chapters allow.

It assumes that candidates are independent, so that the value of a set is the sum of the values of its members. It assumes that relevance captures everything about whether an item belongs. It assumes that a full budget is the goal. And it assumes that any set of candidates it returns is a legitimate answer.

The rest of this chapter takes those assumptions one at a time, on the running situation: an agent asked to investigate a failing database migration in a repository.

Legality cannot be bought

Two items in the pool are wrong for this task. One is a configuration file from a different project. It scores 0.99 for relevance, because it looks almost exactly like the file the task needs. The other is an observation of which backend the project uses, taken at an earlier revision. It exists in two versions: a compact one of 100 tokens, relevance 0.9, and a fresh full one of 600 tokens, relevance 0.75.

A ranking takes the configuration file first. It also takes the stale compact observation, because it is cheaper and scores higher than the fresh one. The two mistakes have the same cause. Being in the wrong project, and being out of date, are not points on the scale of relevance. Chapters 19 to 21 argued that authority, freshness and scope are separate questions, with their own evidence and their own owners. Putting them into the same score does not make them commensurable; it only hides them.

The obvious repair is a blended score: relevance, minus a penalty for being out of scope, minus another for being stale, and so on. That repair fails in a particular way. Any blend lets a sufficiently relevant illegal item outrank a moderately relevant legal one once the weights drift, and weights always drift. A rule that must hold in every case should not depend on how large a number happens to be.

What is wanted is a different kind of statement. Some candidates are not eligible, and being ineligible means being excluded, whatever else is true of them.

Cost is not local

A reference to the earlier incident costs twenty tokens to write down. It is cheap, and a ranking that divides value by cost adores it.

But a reference is only useful if the model can follow it, and following it needs the definition of the tool that resolves it. That definition is 650 tokens. The reference costs 670 tokens to use, and if two references share the resolver the second one costs only its own tokens.

Chapter 4 raised this with a small example of two files, the second of which defined the terms of the first. The example was enough to show that a per-item score could not be right. It is also enough to show what would be. The unit that costs tokens is not the candidate. It is the candidate together with everything it needs, less whatever is already in the bundle.

The consequence is worse than a mispriced item. Static value-per-token orderings misfire whenever dependencies are shared, because the price of an item depends on which other items were chosen first. A candidate that looks like a bargain at 100 tokens but needs a 900-token dependency costs 1,000. A slightly less valuable one at a flat 500 costs half as much. Ranked by local price the first wins. Priced as the bundle would pay for it, the second does.

Pieces that travel together

Some material is misleading in part. The claim that the migration completed is 120 tokens. Its qualification, except for one tenant, is 80. A bundle containing the first without the second is worse than one containing neither: it is a confident statement that is false.

A ranking sees two candidates. It may well score the claim higher than the qualification, since the claim resembles the task more closely, and admit one and not the other. Nothing in a per-item score can express that one item is only safe in the company of another.

The same thing happens with disagreement. If a decision record and a README disagree about the database, Chapter 19 said the honest response was to preserve the conflict: both claims, each with its source, and a marker saying they conflict. Admitting one side, or two sides without the marker, gives the model either a confident half-truth or two confident and incompatible facts.

One content, several forms

The incident report exists in four forms: the full text at 900 tokens, a compact summary at 300, an anchor at 90, and the bare reference at 20 with its resolver. All four say something about the same incident.

To a ranking they are four candidates. Rank them by relevance and it may admit the full text and the compact form, or the anchor and the reference. Every extra form spends budget to say what the bundle already says.

There is a stricter problem underneath. The forms are not interchangeable, because Chapter 11 and Chapter 12 showed that some content must survive in exact form. If the rule about generated files is offered as a paraphrase, no amount of relevance makes the paraphrase adequate. A representation below the floor for its content is not a cheaper version of the same thing. It is a different, worse thing, and the system needs to know it may not be used.

Not everything is a candidate for trimming

The opening pile is the general case. Some material must be present or the computation is unsound. Chapter 16 stated the principle for memory: the task’s non-negotiables come first, and historical material competes only for what remains. Chapter 7 showed why no single importance score can carry it. A constraint that must be obeyed can look unimportant to a similarity measure; a paragraph that closely resembles the task can be entirely optional.

A ranking cannot say this. It can only sort. Put the non-negotiable material at the top by hand, and the ranking has become a different mechanism with a special case bolted on. Then the special case has to say what happens when the non-negotiable material does not fit, and the only honest answer is the one the chapter opened with.

A budget is a ceiling

The fourth assumption is the quietest. A ranking fills the budget because the budget is there. If two long notes about subjects the task never touches score just above nothing, a ranking with 2,700 tokens to spare will still take them, on the theory that some relevance is better than none.

Chapter 5 gave the reason not to. Distractors mislead and volume dilutes, and both effects grow with what is added. The window is a ceiling, not a target. A bundle that stops when nothing further earns its place is not wasting capacity. It is declining to spend it on material that costs more than it returns.

This is also the least measurable of the five objections, so it is worth being plain about what is and is not known. How much a given amount of extra material hurts a particular model on a particular task is exactly the kind of question the book’s later chapters can only pose. What can be said now is structural. If nothing ever earns exclusion by being unhelpful, then the only remaining reason to stop is that the budget has run out, and that treats the budget as the objective.

What the failures have in common

None of the five is exotic. Each is a situation the earlier chapters met on its own: wrong-world material, stale material, references and their resolvers, claims and their qualifications, forms and floors, mandatory constraints, distractors. What they share is that a per-item score is the wrong kind of object for them.

Legality is a yes or no, not a number. Cost belongs to sets, not to items. Some pieces belong together. Some alternatives exclude each other. Some things must be present and some may be absent. And some requests cannot be met.

There is external work on the part of this problem that does look like ranking. Selecting a subset of a long text under a strict token budget is a constrained optimisation, and recent work on it treats it that way and finds that the best method depends on the task and the budget. Coding-agent practice arrives at the same guidance from the other side: curate the smallest set of high-signal material for each task, and note that smallest does not mean short. Both are useful. Neither addresses what a ranking cannot express: eligibility, dependencies, groups, alternatives and refusal. A knapsack is one sub-problem of assembly, and it misleads as a description of the whole.

How often real candidate pools contain these traps is a different question. The situations above were built to be clear. Nothing in the book yet measures how frequently ordinary sessions present a stale cheap alternative, an unresolvable reference or a mandatory overflow, and until real captures exist that remains unmeasured. The mechanism can be shown to work on them. It cannot be shown to matter in general.

What a mechanism would have to do

Set the five failures side by side and the requirements follow, in something close to the order in which they would be applied.

It has to decide eligibility before it compares anything, and treat an ineligible candidate as absent, not as a low score. It has to treat some candidates as non-negotiable and say what happens when they cannot all fit. It has to price a candidate as the whole set it needs, net of what is already in. It has to admit some pieces together or not at all, and one form of each content, never two. It has to stop when nothing earns admission, not when the budget is gone. It has to order what it admits deliberately, render it exactly, and check that the rendered cost still fits, because the number that decided admission is not the number the model will be charged.

And it has to be able to refuse. The correct answer to some requests is that no legal bundle exists, and that answer has to be produced as reliably and as legibly as a bundle would have been, with the reason.

Two further properties are not about the answer but about trusting it. The mechanism should be deterministic, so that the same inputs always give the same bundle, because a system that changes its mind cannot be debugged. And it should leave a record of what it decided and why, so that a person can audit a rejection without guessing.

That is a specification, not yet a system. It has a familiar shape. A program is checked against rules before it is translated; a database refuses a write that violates a constraint; a type checker rejects an assignment. In each case an input is validated against explicit rules, constraints are hard, and the output is either a well-formed artefact or an error with a reason. The comparison will be useful, and it has to be used carefully, because a context assembler translates nothing and promises no optimality. The next chapter uses it only where it explains.

What comes next

We can name what a bundle needs. We can see why ranking cannot give it. We do not yet have the machine that can, and the machine has to be small enough to inspect, because its whole value is that it can be checked.

What does a mechanism look like that can validate first, price whole, keep pieces together and say no?

References

  • Qureshi, K., Martin, G., Peng, Y. “Budget-Aware Routing for Long Clinical Text.” Peer-reviewed, Findings of ACL 2026. Knapsack-constrained subset selection with a submodular relevance-coverage-diversity objective; routing by budget regime with task-dependent optima. Used for the constrained-subset framing and the no-universal-selector finding; clinical objectives never imported. https://aclanthology.org/2026.findings-acl.2114/
  • Ghulyani, M., Singh, A., Bharadwaj, K., et al. “PACMS: Submodular Context Selection as a Pluggable Engine for LLM Agents.” Preprint, arXiv:2606.20047 v2, September 2026. Unified turns/memory/tool-output pool with facility-location selection against recency truncation. Cited for the pooled-assembly direction; objective unadopted, effect sizes untouched. https://arxiv.org/abs/2606.20047
  • Anthropic Applied AI team. “Effective context engineering for AI agents.” First-party engineering essay, September 2025, verified September 2026. Iterative minimal high-signal curation with minimal explicitly not meaning short. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents