← Context From First Principles

The Context Compiler

The mechanism that turns a request, candidate representations and a policy into an exact ordered bundle, or an explicit refusal, and the rules it applies in the order it applies them.

Here is the whole interface.

const output = compileContext(request, candidates, policy);

if (output.result.success) {
  const bundle = output.bundle;      // the exact, ordered ContextBundle
  const trace = output.result.trace; // why each candidate was or was not admitted
} else {
  const reason = output.result.failure?.reason; // a named refusal, never a guess
}

A request, a list of candidates and a policy go in. Either a bundle and its decision trace come out, or a refusal and the record of what had been decided when the refusal came. Everything else in this chapter is what happens inside that call, in the order it happens.

The call reads no files. It asks no model anything. It uses no network, no clock and no random number. Whatever it knows about the world arrived in its three arguments, and the same three arguments give the same answer every time.

That is a narrow machine, and the narrowness is the point. The chapters before this one described everything that happens upstream: how retrieval, memory, tools and the agent itself produce material; how that material is reduced, given several forms, checked for authority, freshness and scope. The previous chapter showed why a ranked list cannot turn the result into a bundle. This chapter shows the mechanism that can, and it does so one rule at a time, because each rule exists to answer a failure the reader has already met.

The task

The numbers in this chapter come from small synthetic cases that ship with the compiler. They are demonstrations of a mechanism, not measurements of anything, and anyone with the package can rerun them. Their subject is the book’s running situation.

An agent has been asked to investigate a failing database migration in a repository. Upstream systems have produced a pool of material: a standing rule that generated files must never be edited; the task itself; a report on an earlier incident, available in several forms; an observation of which backend the project used, taken at an earlier revision; a configuration file that belongs to a different project; a claim that the migration completed, with a qualification that it did not complete for one tenant; two documents that disagree about the database. There is more material than the budget allows and some of it must not be used at all.

The compiler is handed all of it and has to decide what this one computation receives.

What the compiler is given

The request says what is being compiled and under what limits. It carries a task identifier, the usable token budget, the active scope (which project this computation belongs to), the identifiers of any candidates the task names explicitly, the name of the policy the caller expects, and a creation time. The creation time is supplied by the caller. The compiler never reads a clock, which is one reason the same request can be replayed a year later and produce the same bundle.

The usable budget is the one Chapter 4 defined: hard capacity, minus room for the answer, minus everything else the runtime has reserved. The compiler does not work it out. It treats it as a ceiling it must not exceed.

The candidates are the material. A candidate is not a document. It is one representation of one piece of information, together with everything the compiler needs to decide whether that representation may be used:

What it carriesWhat it saysEarned in
Identitya candidate identifier; a content identity naming the information it represents; a representation identifier naming the form12, 13, 18
Forma form rank, and a minimum rank below which this content is not allowed to fall (the floor)7, 11, 12
Costa token count and the source of that count (declared, estimated)2, 4
Requirementone of four classes: mandatory, required, preferred, discretionary7, 16
Eligibilitythree yes-or-no judgements, each with a reason: scope, freshness, authority19, 20, 21
Dependenciescandidates that must accompany this one4, 13, 17
Groupa group identifier and whether the group is required11, 19
Signalsa relevance value; coverage keys naming what the candidate helps cover5, 14
Placementan order role: instruction, task, state, evidence, support or tool6
Payloadthe source kind, a source reference, a kind, and the content itself3

Two of those rows need a word before going on. The form rank is a number the builder assigns; the compiler only compares it with the floor. In the compiler’s own test cases it runs from zero for a bare reference to three for the full text, with an anchor and a compact form between, and that is a convention, not something the compiler knows. And the token count is whatever the builder declared. The compiler trusts it.

The policy is small on purpose. It holds four values: the minimum relevance a candidate needs to be admitted when it adds nothing new; how mandatory content chooses among its forms; the order in which roles appear; and what to do when the rendered bundle overruns the budget (drop the last discretionary item and try again). The second has one supported setting, “cheapest”, and the compiler refuses any other value instead of quietly ignoring it. A policy field that the code does not read is documentation, not policy, and an earlier version of the compiler had exactly that defect. The policy has a name and a version. If the request expects one policy and the caller supplies another, the compiler refuses to run.

Malformed input gets the same treatment. Duplicate candidate identifiers, a negative token count, a dependency on a candidate that is not in the pool, a required group that spans two requirement classes: each is an error thrown before any decision is made. That is different from a compile failure, which is a legitimate answer to a well-formed request. The distinction will matter at the end of the chapter.

Eligibility is an input

One fact about that table needs stating plainly, because it is easy to read past.

The compiler does not discover truth about scope, freshness or authority. It receives eligibility judgements as inputs and enforces them during assembly.

Whether a configuration file belongs to this project, whether an observation still describes the repository, whether a piece of text is allowed to instruct: those were the questions of Chapters 19 to 21, and answering them is the job of whatever builds the candidates. The compiler has no way to check. If a builder marks a stale observation as fresh, the compiler will admit it and record, faithfully, that it passed the freshness gate.

That is a limitation, and it is also the reason the gate can be absolute. A rule that has to guess cannot be strict. A rule that is handed a verdict can be.

One piece of information, several candidates

Take the earlier incident. Upstream, someone has prepared four forms of it:

FormRankTokensNeeds
full text3900nothing
compact summary2300nothing
anchor190nothing
bare reference020a resolver definition of 150 tokens

All four carry the same content identity. To the compiler they are alternatives, not four items. It will admit at most one, and it will never admit two forms of the same content, because the second would spend budget to say something the bundle already says.

This is the split Chapter 12 introduced, and Chapter 18 generalised: information is not the same as its representation. The compiler makes it operational. It also makes clear where the compiler stops. It chooses among forms it is given. It does not write a compact form when only a full one exists, and it does not shorten the full one to fit. Producing the forms is upstream work.

The floor is what stops a cheap form being used where it is not safe. If the rule about generated files is offered only as a paraphrase, and its floor says the exact wording is required, then no legal representation of it exists. The compiler treats that as a fact about the request, not as a problem to be smoothed over.

Non-negotiable, needed, wanted, spare

Importance is not one number, and Chapter 7 already showed why a single score cannot carry it. The compiler expresses the difference as four requirement classes, and each class is handled differently.

ClassMeaningHandledIf it cannot be admitted
Mandatorygoverning constraints; the task itselffirst, in its cheapest legal formthe compilation fails
Requiredneeded for this computation to be soundsecond; required groups before single itemsthe compilation fails
Preferredsupports the taskthird, by ranking within the classit is skipped
Discretionarymight help; sparelast, by ranking; the first thing dropped if the render overrunsit is skipped

A request can also name candidates by identifier. A named candidate is promoted to required (a mandatory one stays mandatory), and if it is not in the pool at all the compilation fails at once, before any other work, with a reason that says exactly which candidate is missing.

This is where the common instruction, rank everything by relevance and take the top of the list, gives way. A low-scoring constraint that must be obeyed outranks a highly similar paragraph that merely might help, and no amount of similarity changes that. Relevance still matters, but only inside the two lower classes, deciding the order in which optional material is considered. It never decides whether a constraint is present.

Two axes, not one. Chapter 7’s retention classes (pin, compressible, externalisable, refetchable, discardable) and these requirement classes both sound like importance, and it would be tidy to merge them. They should not be merged, because they answer different questions. A retention class says what may be done to an item when space is short. A requirement class says whether this compilation must include it.

The rule about generated files is pinned: whatever happens to it, its wording must survive exactly. In a task that edits code it is also mandatory. In a task that asks only for a summary of a design note, the same rule could be discretionary. It would still have to be exact if it appears, but the compilation would not fail without it. A decision record that is externalisable may leave the window in general and still be required by a task that names it.

The two meet in one place. A retention class becomes the floor on the candidates a builder emits for the item: a pinned item is offered only in exact form. The compiler enforces the floor and never sees the retention class.

Before any candidate is compared with another, each one is asked four questions in a fixed order. Is it in scope? Is it fresh? Is it authorised? Does its form meet its floor? This is the whole of the gate, abridged from the engine:

function hardGate(c: ContextCandidate): Gate {
  if (!c.scopeEligible)       return reject("scope_ineligible", c.scopeReason);
  if (!c.freshnessEligible)   return reject("freshness_ineligible", c.freshnessReason);
  if (!c.authorityEligible)   return reject("authority_ineligible", c.authorityReason);
  if (c.formRank < c.minRank) return reject("illegal_representation", "below floor");
  return eligible();
}

There is no score in it. A candidate that fails any test is rejected, the reason is recorded, and nothing that happens later can reverse it. Hard illegality cannot be repaired by a high relevance score.

Two cases from the synthetic corpus show what that means. In the first, a configuration file from a different project scores 0.99 for relevance: it is nearly the same text as the right file. It is rejected at the gate, however closely it resembles what the task needs, and the trace records that the item belongs to another project. In the second, an observation of the project’s backend exists in two versions. The stale one is a compact form of 100 tokens, relevance 0.9. The fresh one is a full form of 600 tokens, relevance 0.75. The stale one is cheaper and scores higher. It is rejected, the fresh one is admitted, and the trace says why.

A ranking would have taken both of the losers. The compiler cannot, because the deciding facts are not on the scale.

Hard gatesScoring signals
Membersscope, freshness, authority, representation floorrelevance, coverage, cost, identifier (as a tie-break)
Effectremove a candidate absolutelyorder the candidates that remain
Can be outweighed?neverby definition
Recorded asrejected, with a reasonadmitted or not, with a position

The price of a reference

A reference to the earlier incident costs twenty tokens to write down. Using it costs more, because the model cannot follow a reference without the definition of the tool that resolves it, and in the case at hand that definition is 650 tokens.

    flowchart LR
    R1["Reference to incident<br/>20 tokens"] --> T["Resolver definition<br/>650 tokens"]
    R2["Second reference<br/>25 tokens"] --> T
  

The compiler computes what it calls the dependency closure: the candidate, everything it depends on, everything those depend on, and so on. What the candidate costs to admit is the cost of the part of that set that is not already in the bundle. The reference costs 670 tokens, not 20. Two references that share a 300-token resolver cost 345 tokens together, because the second one pays only for itself. Two candidates that depend on each other are admitted together or not at all, and the compiler does not loop.

A dependency that is itself illegal blocks everything that needs it. A discretionary candidate in that position is skipped. A required one causes the compilation to fail, and the diagnostic names the dependency.

This is the sense in which local cost is not true cost. Chapter 4 argued it from a small example, and Chapter 17 priced the tool definition that is the commonest dependency of all. The compiler prices the whole set, not the item.

It does not promise to price the set cheaply. In the same case, the full text of the incident was available at 400 tokens, less than the 670 the reference needs. The reference won because, inside the discretionary class, candidates are ordered by relevance first and cost only breaks ties, and the reference scored 0.9 against 0.7 for the full text. At the tightest budget, the compiler spent 670 tokens to obtain what 400 would have bought.

That is a property of the ranking rule, not an accident. The compiler is not an optimiser. It is a legality checker with a deterministic way of choosing among legal options, and nothing yet shows that its way of choosing is the best one.

Pieces that travel together

Some material is misleading when only part of it is present. The claim that the migration completed is 120 tokens. Its qualification, except for one tenant, is 80. The first without the second is worse than neither.

A candidate can therefore carry a group identifier and be marked as part of a required group. Members of a required group are admitted together or not at all. They must sit in the same requirement class, and the compiler refuses input in which they do not. If the group cannot fit, the compilation fails with a reason that names the group. If the budget later overruns and something must be dropped, dropping a member drops the whole group.

The same mechanism handles disagreement. When a decision record and a README disagree about the database, the compiler does not decide which is right. That question belongs to authority policy (Chapter 19), and by the time the pool reaches the compiler it has an answer: preserve the conflict. The two claims and a marker saying they conflict form a required group of 60, 60 and 40 tokens. They enter together or not at all, so the model sees a disagreement that has been labelled as one, not two confident and incompatible facts.

Choosing a form

When several forms of one required or mandatory item are legal, the compiler admits the one with the least marginal cost, counting dependencies, and breaks ties by identifier. The policy names this as its mandatory-form setting, and the corpus shows what it does.

The incident was required, with the four forms above. Reference plus resolver costs 170; the anchor costs 90; the compact form 300; the full text 900. The compiler admits the anchor. It does so at a budget of 550 and at a budget of 2,000, because a richer form is not admitted merely because it fits.

Whether a reader does as well with the anchor as with the full text is a behavioural question, and the book has not answered it. Chapter 26 returns to what the small runs did and did not show. For now the important thing is that it is a choice, written down where it can be argued with, and that a budget with room to spare is not treated as an instruction to spend it.

Inside the preferred and discretionary classes the forms compete as separate candidates by relevance, as the previous section showed. Only the content identity keeps two of them from both being admitted.

A ceiling, not a target

After the mandatory and required material is in, the compiler admits preferred and then discretionary candidates one at a time. Each round it looks at the units that still fit and are still legal, discards those that would add nothing, and admits the best of the rest by relevance, then by how much new coverage it brings, then by cost. It stops when no unit qualifies.

A candidate adds nothing if its relevance is below the policy’s threshold and it covers no key that the bundle does not already cover. The threshold in the corpus is 0.3. That number is a policy value, not a finding.

The consequence is the one Chapter 5 argued for. In a synthetic case with a budget of 3,000 tokens, the compiler produced a bundle of 265. One useful note of 150 tokens (relevance 0.8) went in. Two long notes of 400 and 500 tokens (relevance 0.15 and 0.2), both about subjects the task did not touch, were rejected as adding nothing, with 2,700 tokens still unspent. Unused budget is not waste. A window is a ceiling.

Saying no

Sometimes no legal bundle exists. The compiler can say so, and the ways it can say it are named.

ReasonWhen
INSUFFICIENT_BUDGETmandatory or required material cannot fit, or the rendered bundle overruns and nothing may legally be dropped
UNSATISFIED_DEPENDENCYa mandatory or required candidate depends on something ineligible or colliding
NO_LEGAL_REPRESENTATIONevery form of a mandatory item is below its floor
UNRESOLVED_REQUIRED_GROUPa required group cannot be admitted whole
REQUIRED_SOURCE_UNAVAILABLEthe request names a candidate that is not in the pool
REQUIRED_INELIGIBLEa mandatory item has no form that passes the gates

Three cases from the corpus show them. Two mandatory items, of 2,000 and 2,500 tokens, meet a budget of 4,000, and then of 4,400. The first fits; the second does not; the compiler stops with the second one named, and reports that 2,000 tokens were spent. A mandatory item offered only in a form below its floor produces the no-legal-representation refusal at every budget, however generous. A request that names a candidate absent from the pool produces the required-source refusal before anything else is considered.

A failure is a complete result. It carries the reason, the identifiers that blocked it, how much budget had been used, a diagnostic sentence, and the trace of the decisions reached before the refusal. It does not carry a bundle, because there is no legal one to carry.

Engineers already accept this behaviour elsewhere. A compiler rejects an invalid program. A database rejects a write that violates a constraint. A type checker rejects an assignment. In each case the refusal is the useful behaviour, because the alternative is an artefact that looks valid and is not. A context compiler that silently truncates a mandatory instruction to make it fit has done something worse than fail: it has produced a bundle that reads as legal.

Recovery happens outside the compiler. A larger window, a smaller task, a leaner tool surface, moving evidence out of the window, or a person deciding: each belongs to a different owner, and none of them is something the compiler can do by relaxing its own rules.

Order, render, check

Admission decides which candidates enter. It does not decide where they sit. The compiler orders the admitted candidates by role, in the order the policy lists (instruction, task, state, evidence, support, tool), and within a role by identifier. Relevance plays no part. The role order is a policy default, in the sense Chapter 6 gave that word: a choice, not a finding.

Then it renders. The bundle holds items, each with an identifier, a source, a kind, its content and a token count. The renderer is deliberately plain: a header, the item contents separated by a fixed marker, a closing line.

[CONTEXT BUNDLE]
id: dependency-trap-req-bundle
items: 4
---
Application constraint (synthetic): never modify generated files.
---
Investigate the migration failure (synthetic).
---
Incident syn-17 resolved (synthetic reference).
---
Artifact resolver capability (synthetic): reads external artifacts.
[/CONTEXT BUNDLE]

The number that decides admission is not the number that reaches the model. Admission adds up the token counts the candidates declared. The rendered text also carries a header, a separator between items and a label for each item, and the compiler counts those with a rough word-based estimate, because it has no tokeniser. In the case above the declared counts sum to 770 and the rendered bundle costs 789.

Usually the difference is harmless. Sometimes it decides the outcome. In another synthetic case the admitted candidates summed to exactly the budget of 250. Rendered, they cost 265. The compiler noticed, dropped the discretionary item, which the policy names as the first thing to go, and rendered again. The final bundle cost 111. If no discretionary item had been left to drop, the compilation would have failed.

Finally the compiler checks its own output before it releases it: the rendered cost against the budget, the layout against the item order, every item against the admitted set. Only then does it return the bundle and the trace.

The count the compiler trusts is the count it was given, and a judgement is only as good as its units. In the project’s own calibration on real sessions, a word-based estimate ran well below what a provider reported for the same requests, on one session with one model. A bundle that fits by these numbers fits in the units the candidates declared. It is not, on that evidence alone, a bundle that fits by a provider’s count.

    flowchart TD
    IN["Request, candidates, policy"] --> V["Validate input<br/>and named sources"]
    V --> G["Hard gates<br/>scope, freshness, authority, floor"]
    G --> M["Mandatory<br/>cheapest legal form"]
    M --> R["Required<br/>groups first, then single items"]
    R --> P["Preferred, then discretionary<br/>ranked, and only if they earn it"]
    P --> O["Order by role"]
    O --> X["Render, check exact cost,<br/>drop discretionary if over"]
    X --> OK["Bundle and trace"]
    V -.-> F["Failure and trace"]
    M -.-> F
    R -.-> F
    X -.-> F
  

What it does not do

The list matters as much as the mechanism, because the word compiler invites a larger claim than the evidence allows.

It does not derive scope, freshness or authority. It does not build forms of a candidate, prune history, compact anything, move information out of the window or bring it back. It does not know what the model will find easy to read. It does not adapt to a provider’s cache. It does not count tokens exactly. It does not optimise. And it does not call a model or ask anyone.

Each of those is the subject of an earlier chapter, and each remains work for the system that builds the candidates. The compiler was built to be the point at which their outputs are assembled, checked and, if necessary, refused.

The name is a comparison, and comparisons should be used where they explain and dropped where they do not. A conventional compiler translates a program from one language to another under explicit rules. This one does not translate anything. What it shares with one is a set of properties: explicit validation, hard constraints, a defined output format, a refusal when the input is invalid, and an output that is the same every time. Where the analogy would suggest optimisation, or a source language, or a proof of correctness, it stops.

Where it runs

The same code is offered three ways: as a library, as a command-line tool that reads and writes JSON, and as a set of explicit tools inside the coding agent used in this book’s experiments. The core has no dependency on that agent. Installing the tools does not change what the agent sends to a model. It gives the agent operations it can call, and nothing compiles until something asks it to.

That separation is deliberate. The compiler answers one question, what exactly this computation should receive. Whether that answer reaches the model is a different question, and it belongs to a different piece of software.

What the rules amount to

Each rule above can be stated in a sentence. Illegal candidates are removed before any comparison. Non-negotiable material is admitted first and its absence is a failure. A candidate’s cost includes what it needs. Some pieces come together or not at all. One content, one form. A ceiling is not a target. When no legal bundle exists, the compiler says so.

Sentences are cheap. What matters is whether they are true of the code, and how anyone else would know. The compiler produces a trace of every decision it made and an assertion that the same inputs always give the same output, and both are claims that can be checked. What checking them establishes, and what it cannot, is where the next chapter begins.

References