
Context From First Principles
Context as a compiled, observable artefact: see what a model actually receives, learn what capacity, content, representation, time, authority and scope each demand, then build a deterministic Context Compiler, deliver its output to a runtime, observe that it arrived, and test whether it helped.
A context window tells us how much a model can receive. Context engineering decides what it should receive, in what form, in what order, and what must be left out.
Imagine an AI coding agent working inside a mature project.
It can see the conversation. The repository contains thousands of files. The project has instructions, architecture notes, tests, issue history, tool definitions, previous attempts, summaries, memories, search results, plans, and state accumulated over hours of work.
The model can technically accept a very large prompt.
What should the system send?
One implementation keeps adding material until the window is nearly full.
Another identifies the current task, preserves what must survive, removes stale and duplicated material, externalises bulky state, recalls only what becomes relevant, orders the surviving information deliberately, and records exactly what the model received.
Both systems may use the same model.
Both may have access to the same information.
They do not have the same context.
That difference is the subject of this book.
The window is not the context
A production AI system can draw from far more information than one model call should contain:
- system and developer instructions,
- conversation history,
- project rules,
- files,
- retrieved passages,
- tool definitions and results,
- execution traces,
- plans and hypotheses,
- summaries,
- memories,
- external documents,
- policies,
- cached state,
- and artifacts generated by the agent itself.
A context window defines capacity.
It does not decide which information deserves admission.
It does not decide what can safely be removed.
It does not decide whether a summary preserved the fact that later becomes decisive.
It does not decide whether an old plan is stale.
It does not decide which source wins when two sources disagree.
And it does not tell us whether the resulting context improved what the model did.
Those are context-engineering problems.
A stricter definition
This book keeps several ideas separate.
Available information is everything the system could potentially access.
Session state is information persisted across turns or actions, whether or not it is currently shown to the model.
Context is the information made available to a model for a particular computation.
Context window is the capacity that bounds the invocation.
Context selection decides which candidate information enters.
Context representation decides the form in which admitted information appears.
Context assembly turns admitted information into an ordered, policy-compliant bundle.
That separation matters because a larger window solves only one problem.
A system can have abundant capacity and still build poor context.
It can include the wrong evidence.
It can preserve stale instructions.
It can bury decisive information inside thousands of merely related tokens.
It can compress away a contradiction.
It can retrieve something useful and fail to admit it.
It can admit the right information in a representation the model does not use effectively.
It can let an untrusted source appear more authoritative than a project rule.
So the central question becomes:
Given more potentially useful information than a model can or should consume, what should it see right now?
Four dimensions instead of one
The book begins with three dimensions that are often collapsed:
capacity content representation
how much? which bits? in what form?
Then it adds time:
time
when is this information valid, useful, or worth keeping live?
Context grows.
It goes stale.
It is duplicated.
It is summarised.
It leaves the live window.
It may be externalised into files, caches, memory, databases, or derived state.
Later, some of it returns.
The agent also creates future context through plans, notes, summaries, tool traces, critiques, and intermediate artifacts.
Context is therefore not a static prompt.
It is a changing working set.
Seven questions instead of a bag of tricks
The investigation is organised around seven practical questions:
- What did the model actually receive?
- What information matters to this task?
- What must survive exactly, and what may be transformed?
- What can leave the live window?
- What should come back, and when?
- How should the surviving information be ordered and represented?
- Did the resulting context improve behaviour enough to justify the machinery?
The sequence matters.
Before changing context, observe it.
Before compressing information, decide what must survive.
Before adding retrieval, memory, or external storage, identify why information left the live window.
Before optimising token count, check whether the cheaper context still supports the same behaviour.
And before adding a sophisticated mechanism to the architecture, demonstrate a failure that the simpler system could not solve.
The book is an experiment
The method is deliberately conservative.
Build the strongest simple baseline.
Measure it.
Stress it.
Find the failure.
Diagnose where the failure occurred.
Add only the mechanism that the diagnosis requires.
Then rerun the same cases and check that the repair did not break what already worked.
observe → baseline → failure → smallest intervention
→ controlled test → ablation → keep, revise or remove
Context engineering is full of plausible mechanisms.
Longer windows sound useful.
Caching sounds useful.
Pruning sounds useful.
Summaries sound useful.
Retrieval sounds useful.
Memory injection sounds useful.
Structured state, tiered fidelity, externalisation, graphs, and policy layers can all sound useful.
But usefulness is conditional.
A mechanism belongs in the final architecture only if it repairs a demonstrated problem at an acceptable quality, cost, latency, and complexity price.
Negative results stay in the argument when they remove unnecessary machinery.
The system the book arrives at is smaller than the investigation that produced it, and much of the investigation is argued from the literature and not yet built.
How the evidence is presented
When this book reports something we measured, it says what we found, in plain language, and what population the finding covers. It does not print commit hashes, run identifiers or software versions in the text. Each finding has a short name, and the identifiers, code revisions and published artifacts behind it are listed in the evidence register in the companion repository.
If a later experiment changes what we believe, the text will say what we thought earlier and what we think now. The register keeps the full history.
More is not automatically better
One of the book’s central claims is intentionally simple:
Making more information available to a model does not guarantee better behaviour.
Additional context can help.
It can also distract.
Duplicate evidence consumes budget without adding information.
Old plans can compete with current plans.
Failed attempts can look like instructions.
A retrieved passage can be topically relevant but operationally useless.
Important content can become harder to notice when surrounded by a larger amount of merely related content.
Ordering can change which instructions or evidence dominate.
A summary can preserve the theme while deleting the exception that matters.
So token count is not a proxy for useful context.
The important test is behavioural:
Hold the task and model fixed. Change the context. Does the model behave differently, and is the difference useful?
That becomes the eventual test for every major mechanism in the book.
The context window is a budget
A context window is not merely a size limit.
It is a budget.
Every admitted item consumes some combination of tokens, attention, latency, provider cost, cacheability, and competition with other information.
Inclusion therefore has an opportunity cost.
Adding one item may force another out.
Preserving raw evidence may cost more tokens than a summary but protect a distinction that matters.
Inlining a large tool result may simplify access but crowd out source material.
Keeping a stable prefix may improve cache economics while constraining how volatile material is arranged.
The engineering problem is not to maximise occupancy.
It is to maximise useful behaviour under a bounded budget.
Remove before you compress
When context becomes too large, summarisation is an attractive first response.
This book treats it as a later response.
Before rewriting information, ask whether some information should remain live at all.
Duplicate content may be removable.
Completed work may be retired.
Failed branches may be externalised.
Stale observations may be invalidated.
Bulky artifacts may be replaced by stable references.
Only then does lossy compression become the obvious next step.
That ordering matters.
Removing irrelevant material can be lossless for the current task.
Compression is a claim about what can disappear without changing future behaviour.
That claim should be measured.
Compression is information selection
A shorter summary is not automatically a better representation.
Compression can preserve the gist while deleting:
- a contradiction,
- a numerical constraint,
- an exception,
- an unresolved question,
- a failed attempt that must not be repeated,
- or the evidence that licences a conclusion.
So the question is not simply:
How many tokens did we save?
It is also:
Which information survived, which information disappeared, and did the loss change behaviour?
Some information may tolerate aggressive reduction.
Some may need structure.
Some may need a pointer back to the source.
Some may need to remain verbatim.
Not all tokens are equal.
Context can leave without being forgotten
A bounded model input does not require a bounded system.
Information can leave the live context while remaining available elsewhere.
That creates an important distinction:
Live context is temporary. Accessible state can be durable.
A system may externalise bulky or lower-priority information into files, stores, caches, indexes, memory systems, or other representations.
But externalisation creates a second problem.
Once information has left the working set, how does the system know when to bring it back?
That is why retrieval appears in this book.
Not as the definition of context.
Not as the definition of memory.
But as one possible admission mechanism for information that is no longer live.
The question is whether the right material returns at the right moment, in the right form, under the available budget.
Memory is a context source
This book follows Memory From First Principles, but the two problems are different.
Memory asks how retained past experience remains capable of changing present behaviour.
Context asks which information and state are actually made available to one execution.
A memory system may retain years of project history.
A context system decides whether any of it should enter this model call.
Memory can therefore be a source of context.
So can retrieval.
So can the conversation.
So can tools.
So can files.
So can policy.
So can state the agent generated five seconds ago.
Context engineering sits at the boundary where those sources compete for admission.
The agent writes its own future context
Agentic systems have an unusual property: the model is both a consumer and a producer of future context.
A plan written now may become input later.
A summary may replace the messages from which it was derived.
A hypothesis may persist across tool calls.
A critique may alter the next attempt.
A tool trace may become evidence.
An intermediate artifact may become a dependency.
context
↓
model action
↓
generated artifact / state
↓
future context candidate
↓
later model action
Poor context can therefore reproduce itself.
A mistaken summary can survive longer than the raw evidence it replaced.
A speculative plan can acquire authority through repetition.
A stale derived artifact can remain in circulation after the source changed.
Generated context needs the same scrutiny as retrieved context.
Tools produce context
Tool use is often described as something outside the prompt.
From the model’s perspective, it creates context.
Tool definitions consume tokens.
Arguments encode intent.
Outputs may be large, noisy, or partially relevant.
Errors become state.
Repeated observations accumulate.
A long-running agent may spend more context on tool interaction than on the original request.
The system may therefore need to decide which tools to expose, how much schema detail to include, which outputs remain live, what can be reduced, and when the source should be reread rather than trusting an earlier summary.
Tools do not merely act on the world.
They reshape what the model knows about the world.
Representation is part of context
The same information can appear as raw prose, quoted evidence, a table, typed state, a graph edge, a structured record, a summary, a diff, or a compact reference.
Those forms are not interchangeable.
Representation changes token cost.
It changes what relationships are explicit.
It changes how easily provenance survives.
It changes what can be omitted accidentally.
And it may change what the model notices.
Context engineering is therefore not finished when the correct source has been selected.
The representation itself has to earn its place.
Authority, freshness, and scope
Once several sources are assembled together, new failures appear.
A project rule may conflict with an old conversation.
A tool result may conflict with a cached summary.
A memory may preserve a decision that has since been reversed.
A user may change direction.
A dependency version may move.
A plan may complete.
A retrieved document may be relevant to the topic but belong to the wrong project or authority domain.
The context system needs enough provenance and policy to ask:
Where did this claim come from?
Is it raw evidence or a derived interpretation?
Is it still valid?
What scope does it belong to?
Which source has standing when two claims conflict?
Can the original source be reopened?
Relevance is not enough.
Context also needs authority, freshness, and boundaries.
From information to a context bundle
By the later chapters, the central chain becomes explicit:
available information
↓
representations of it, and judgements about each
↓
candidates
↓
Context Compiler ← request and policy
↓
ContextBundle + DecisionTrace (or a named refusal)
↓
render
↓
runtime adapter injects it
↓
an independent observer reads what was assembled
↓
reconciliation of what was intended with what arrived
↓
model behaviour
↓
evaluation
Every arrow can fail.
Useful information may never become a candidate.
A good candidate may be rejected.
A stale candidate may be admitted, because someone said it was fresh.
The right source may be represented badly.
A legal bundle may never reach the model, or may reach it twice.
The model may ignore information that survived every earlier stage.
The final answer may even look correct while depending on the wrong evidence.
That is why the book measures the path rather than scoring only the final output, and why it keeps separate the different kinds of evidence each stage allows.
The Context Compiler
The middle of that chain is a Context Compiler, and the name is meant narrowly.
A compiler does not dump its entire input universe into the target.
It takes a defined input, enforces constraints, and emits a specific artefact for a specific execution, or refuses with a reason.
Concretely:
request (budget, active scope)
+ candidates (representations with costs, eligibility
judgements, requirements, dependencies)
+ policy
↓
CONTEXT COMPILER
↓
ContextBundle + DecisionTrace or CompileFailure
The compiler decides membership, representation and order for one computation.
It is deterministic.
It enforces scope, freshness, authority and representation floors as hard gates, and it may refuse.
It does not discover scope, freshness or authority.
It does not write the summaries or judge whether an answer was useful.
Those judgements are supplied to it, and the hard open problem in the book is how to supply them trustworthily. The final chapter says so plainly.
The compiler is not assumed at the beginning of the book.
Each of its rules is introduced as the answer to a failure the earlier chapters showed, and each mechanism the book discusses but has not built waits for its own evidence.
The current investigation
The book has twenty-seven chapters in five parts.
Part I, see the thing (Chapters 1 to 7). Define context, observe what a model receives, and learn what capacity, order and the different natures of information demand.
Part II, capacity and reduction (Chapters 8 to 12). The main responses to a bounded window: a larger one, caching, pruning, compression, and fidelity that fades gradually.
Part III, context outside and around the window (Chapters 13 to 18). Externalisation and recall, the agent’s own generated state, memory, tools, and representation.
Part IV, govern (Chapters 19 to 21). Authority, freshness and scope: the three judgements that decide whether a candidate may be considered at all.
Part V, compile, deliver, prove (Chapters 22 to 27). Why ranking is not assembly, the compiler that assembles, what a compilation does and does not prove, whether the result reached the model, whether it helped, and what the whole system looks like.
The evidence is of four kinds that the book keeps apart: structural evidence about the bundle as an artefact, transport evidence that it reached an observable boundary, behavioural evidence about what the model did, and ecological evidence about how often any of this matters in real use. The last is currently empty, and the final chapter says how far each of the others goes.
Most of the mechanisms in Chapters 8 to 21 are argued from the literature and have not yet had their own experiment. The book records them as conditional and not as earned.
Two things are being built
The book builds a context runtime, but it also builds the instrument required to distrust that runtime.
The instrument begins before the advanced architecture does. It observes, read-only and independently of whatever injects context:
- the actual payload assembled for the model,
- how many tokens each source consumed, and of what kind those counts are,
- where each item appeared,
- what the model did,
- and what changes when context is removed, restored, reordered, compressed, or replaced.
It also has limits, and the book states them. The observer sees what the host assembles for the model, and does not see authority, scope, repository state, provider-side rewriting or cache decisions.
The system and the instrument grow together, and the instrument has to be shown alive in each run that matters. A behavioural wave that skipped that check produced twenty-four runs and no valid observation.
When the book claims that a context mechanism helps, that claim is tied to behaviour and not to token count, and is limited to what the evidence covers.
What you should be able to do after reading
The aim is practical. By the end you should be able to look at the context of your own system, build small tools that improve it, and design the context for a task deliberately rather than by accumulation. The chapter notebooks are a place to start building.
You should also be able to ask much harder questions than how large is the context window?
What did the model actually receive?
What information was available but excluded?
Why did one item enter and another stay out?
Which content had to survive exactly?
Which content was transformed?
What did the transformation lose?
What left the live context but remained recoverable?
Why did a particular item return?
Which context was generated by the agent itself?
What became stale?
What crossed a scope boundary?
Which source had authority when two items disagreed?
How much did the final bundle cost?
Did changing the context change the behaviour?
Did that change help?
And which parts of the final architecture were measured results rather than attractive ideas?
Those questions turn context engineering from prompt craft into systems engineering.
The larger idea
The deeper argument is not that prompts should be shorter.
It is not that prompts should be longer.
It is not that every system needs retrieval, memory, summaries, or a giant context window.
It is that context should be treated as a controlled compilation of a larger information world into the bounded working state of one execution.
That compilation has to decide what enters.
It has to decide what stays out.
It has to preserve distinctions that matter.
It has to keep stale or out-of-scope information from acquiring accidental influence.
It has to move information out of the live window without making it irrecoverable.
It has to bring information back without flooding the model.
It has to preserve provenance and authority when sources compete.
It has to respect a budget without mistaking fewer tokens for better context.
And finally, it has to be delivered to the model and tested for whether it made the model behave better.
The destination is a system that can answer a more demanding question:
What is the smallest, safest, most useful representation of what this model needs to know right now?
That answer will sometimes be a raw source passage.
Sometimes a structured record.
Sometimes a tool result.
Sometimes a memory.
Sometimes a compact reference.
Sometimes a summary.
Sometimes an instruction.
Sometimes nothing at all.
The value is not in giving the model everything.
It is in giving it the right information, in the right form, with the right standing, at the right moment, and being able to show why.
The book is candid about how much of that is established. Building a bundle deterministically under explicit rules, and checking that it reached the model, are shown. Building trustworthy candidates in the first place, and showing that any of it matters on real workloads, are open.
Context engineering begins when a larger window stops being an answer.
Continue with What Context Means.
Chapters
What Context Means
Separate context from prompts, context windows, session state, memory, retrieval, and the larger universe of information a system could access.
The Measurement Instrument
Observe the real payload sent to a model before attempting to improve it: what a read-only observer can and cannot see, how tokens are counted, and why the instrument has to be qualified before it is trusted.
The Context You Didn't Type
Dissect the hidden context stack behind a production AI assistant, using Claude as the most inspectable case.
The Context Window Is a Budget
Treat the model window as a constrained allocation problem rather than a bucket to fill.
More Is Not Better
Measure context bloat, interference, distraction, and position-sensitive failure before proposing a cure.
Order Changes Meaning
Show that context is not a set: position, grouping, precedence, and adjacency affect what the model does.
Not All Tokens Are Equal
Classify information by retention requirements before pruning or compression.
Make the Window Bigger
Compare architectural approaches to long context without confusing capacity with context management.
Reuse Before Recompute
Treat prompt-prefix caching as part of context architecture, not merely a provider billing feature.
Remove What No Longer Matters
Prune duplicate, stale, failed, or completed material before reaching for lossy summarization.
Compression Is Loss
Treat summarization and compaction as irreversible information-selection operations that require measurement.
Let Old Context Fade
Explore progressive fidelity and tiered representations instead of cliff-edge truncation or one-shot compaction.
Externalize the Working Set
Move durable or bulky information out of the live context while preserving a usable reference.
Bring Back Only What You Need
Use retrieval as a context-admission mechanism without turning this book into Retrieval From First Principles.
The Agent Writes Its Own Context
Treat plans, hypotheses, summaries, tool traces, critiques, and intermediate artifacts as self-generated future inputs.
Memory Is a Context Source
Connect the Memory book to Context without collapsing the two problems.
Tools Produce Context
Treat tool definitions, calls, outputs, errors, and observations as first-class context with their own budget and policy.
Representation Is Part of Context
Compare raw prose with structured records, tables, typed state, graphs, and compact references.
Who Gets to Be Right?
Resolve authority, provenance, contradiction, and prompt injection when context sources disagree.
Context Goes Stale
Make freshness, version, validity, and invalidation explicit in context assembly.
Keep Context in the Right World
Use scope, isolation, and compartment boundaries to prevent cross-project and cross-task contamination.
Assemble for the Task
Why ranking candidates and filling the budget does not produce a bundle, what a mechanism would have to do instead, and why the right output is sometimes a refusal.
The Context Compiler
The mechanism that turns a request, candidate representations and a policy into an exact ordered bundle, or an explicit refusal, and the rules it applies in the order it applies them.
What a Compilation Proves
Read a decision trace, check what determinism buys, and separate what the compiler's construction evidence establishes from what it cannot.
Did It Reach the Model?
A correct compile is not a correct delivery. How to show that the bundle the compiler produced is the context the model received, and what that does and does not establish.
Did the Context Help?
Evaluate context by what the model does with it: influence versus utility, a matched design that changes only the bundle, and what a small set of behavioural runs does and does not show.
The Complete Context System
The whole chain from available information to behavioural evaluation, what each part decides, what the project has shown about each, and where the difficult work now lies.