← Context From First Principles

Keep Context in the Right World

Use scope, isolation, and compartment boundaries to prevent cross-project and cross-task contamination.

The task belongs to Project A: backend PostgreSQL, test command pytest, migration complete. Project B lives beside it: backend SQLite, test command npm test, migration pending. The agent’s context holds both projects’ configs, both READMEs, both migration notes — every item authoritative in its own world, fresh, accurately represented, and semantically similar enough that retrieval surfaced all of it. The agent runs npm test in Project A, watches it fail for reasons that have nothing to do with Project A’s code, and then debugs the backend mismatch by reasoning about SQLite — quoting Project B’s migration state as though it constrains Project A’s deployment. No attacker exists. No stale data entered. No authority was violated. The context was correct, current, and legitimate in every way except the one that mattered: half of it belonged to another world.

Correct information from the wrong world is wrong Context.

Membership, not importance

The durable definition:

Scope identifies the world, workspace, task, or compartment in which a context item is eligible to influence a computation.

Scope answers where information belongs. The current computation answers which world it operates in. Admission compares the two. Availability to the surrounding system implies nothing about membership: two config files may both be readable by the runtime while only one belongs to the task. Scope is therefore about eligibility, and four separations keep it clean. Semantically relevant-looking is not in scope — two projects sharing language, frameworks, file names, and architecture concepts can be near-identical in embedding space while belonging nowhere near each other. Authoritative is not in scope — Project B’s decision record governs Project B fully and Project A not at all, with no demotion implied. Fresh is not in scope — two main branches can both be current with different backends, and no temporal rule chooses between them. Accessible is not eligible: an agent with filesystem access to every project directory is not thereby entitled to reason from all of them, just as an item conceptually in scope may still be barred by runtime permissions. Work both directions so the independence is visible. A workspace readable in full still restricts the migration task to its own project’s candidates, because readability answers what the process may touch and scope answers what may influence this computation. Conversely a shared contract conceptually in scope for the task still requires its own access grant before any byte is read — membership never picks a lock. Access control asks what a process may touch; scope asks what may influence this computation. The chapter designs no permissions, tenants, or authentication; it asks only the eligibility question and leaves enforcement to systems built for it.

Isolation, on this framing, is a correctness mechanism before it is a security one. Wrong-project items cause wrong conventions, paths, dependencies, configurations, architecture assumptions, and unfinished-work states in fully trusted environments with no adversary present. Privacy follows as a consequence — user histories and credentials must not cross computations — but the primary experiment measures ordinary contamination, not exfiltration, so the chapter never becomes a security sequel wearing Chapter 19’s clothes.

Bindings, not a hierarchy

Scope dimensions refuse to form one neat tree. A computation may carry project A, the migration worktree, task fix-17, and a reviewer agent role simultaneously — orthogonal bindings, not levels of a ladder. The chapter prefers scope bindings to any universal global-to-agent hierarchy, and defines the computation side minimally: a computation has explicit scope bindings against which candidate eligibility is checked, in reader-facing language simply the world of this computation. Items carry their own bindings — project, environment, task, sometimes shared markers — with unknown left unknown and never defaulted to global. Missing scope must not mean everywhere; that default is how local assumptions contaminate unrelated tasks.

Bindings come from provenance and runtime identity wherever possible, under the same discipline Chapter 19 set for authority. A document announcing its own project membership or global scope gains nothing from the announcement; repository and workspace identity, source adapters, session and task metadata, runtime paths, authenticated tenants, and explicit transfer records decide, while payload labels travel as data the runtime may or may not trust. Deterministic identity precedes model inference everywhere the runtime already knows the answer: where the repository identifier says one project, resemblance to another is not consulted. Ordinary computation draws deterministic boundaries; model judgement handles only what unstructured imports leave genuinely ambiguous, and then marked as inferred.

The cheap boundary: project scope

Two projects with deliberately overlapping artefacts — READMEs, configs, migration notes, issue seventeens, architecture decisions sharing vocabulary but conflicting validly on backend, test command, migration state — make the fixture genuinely tempting to a global pool. The task names its world. Against that fixture the sibling Memory project already ran the controlled version this chapter consumes rather than repeats (an unpublished manuscript; the evidence register in the companion repository says where its run can be inspected): on a 68-unit three-project corpus at a fixed token budget, restricting candidates to the project drove cross-project leakage from 0.212 to zero while must-include recall held exactly at 0.773 and precision rose; on a deliberately similar three-project fixture asking one question of all three, mean leakage fell from 0.667 to zero (eleven tasks, fixed reader). The consumed phenomenon is narrow and stated narrowly: project scoping eliminated cross-project leakage on that controlled corpus without reducing required-evidence recall at the project rung. No universal answer improvement, no production scale, no claim that labels solve every scope problem — and explicitly not the sibling’s ProjectFrame architecture, which this book does not copy.

Labels alone earn a separate measurement rather than an assumption. A prompt rendering both projects’ material with scope tags, asking the model to use only one world, is a legitimate condition — and a different one from never admitting the other world at all. The first asks the model to resist contamination with the contaminant present; the second removes it. Both are tested, because asking a model to ignore known-irrelevant content spends the very context the boundary was meant to save. Count it plainly: eight thousand tokens of Project B resident beside an instruction to disregard them costs the tokens, the positions they occupy, and whatever interference they exert despite the instruction — against zero tokens and zero interference for the boundary that never admitted them. If labels match removal on behaviour, the saving is real convenience; where they trail, the gap prices wishful governance.

The same project holds several worlds

Project scope is necessary and insufficient. Within one repository, main, feature branches, and release lines describe different states, and parallel worktrees hold deliberately divergent modifications — right project is not right environment, which matters most exactly where coding agents live. The fixture extends naturally: Project A on main against Project A on a feature branch, same paths, same names, conflicting states, with the task bound to one environment. Current Codex practice supplies the implementation case with its limits stated: built-in worktree support runs parallel chats in isolated repository copies with managed handoff between local and worktree checkouts. That is environment isolation in production form, and the chapter draws the boundary the docs imply but do not state — filesystem isolation and Context isolation are related but separate, since two chats in two worktrees can still each admit the other’s world through retrieval, memory, or shared tooling. Worktrees prevent file interference; scope policy prevents context interference.

Subagents generalise the same point across agent boundaries, and current systems converge on the shape. OpenCode’s present-day agent model runs subagents as child sessions with their own prompts, models, and permissions, navigable back to the parent — implementation evidence only. The observer records an agent identity per model request but not parent–child session relations, so no frozen result here observes subagent behaviour. Anthropic’s multi-agent research architecture isolates detailed local context in separate windows and returns condensed findings plus file-backed artifacts with lightweight references, explicitly to avoid repeated filtering loss through the coordinator. The transferable pattern is bounded handoff between scoped contexts, not any vendor’s orchestration. Sketch it on the release task: the migration-review child receives the subtask objective, the project’s backend decision and test conventions, and the migration evidence set — and nothing else. No documentation hypotheses, no packaging state, no parent deliberation. The child cannot be contaminated by what it never receives, and the parent’s later admission problem shrinks to one scoped report instead of three raw trajectories. But fresh or bounded child context is not automatically sufficient context: a child handed only its objective may lack the goal restatement, critical constraints, source identities, and project conventions the task needs. Isolation removes contamination by removing information, and omission failures are the price — the central counterweight the second experiment measures rather than asserts.

The counterweight has peer-reviewed form. SILO-BENCH, an ACL 2026 long paper, distributes fragments of algorithmic tasks across agent silos and finds a Communication-Reasoning Gap: agents communicate actively yet fail to convert interaction into effective distributed computation, with the hardest tier collapsing to zero success past fifty agents. The chapter imports the principle at its stated size, never the scale figures: too much sharing contaminates, too little sharing strands dependencies, and coordination does not automatically recover what boundaries hide. DACS, an April 2026 preprint, supplies the complementary mechanism from the other direction: an orchestrator holding lightweight per-agent registry summaries and expanding exactly one agent’s full state on demand, with reported steering accuracy far above flat-context baselines and contamination sharply reduced on synthetic scenarios. Single-author preprint, synthetic populations, author-reported figures — used narrowly for the asymmetric registry-plus-focus shape, never as settled architecture.

Narrower than project, wider when earned

Tasks subdivide projects. Release readiness, documentation rewriting, and migration benchmarking share one repository and almost no information needs; temporary hypotheses, scratch results, explorations, partial patches, and failed experiments are task-local state that should never become standing project candidates merely by existing inside the project. Trace one concrete collision to see why membership must be explicit: the benchmarking task records a working hypothesis that SQLite outperforms PostgreSQL on its synthetic workload, while the release task carries the standing decision that PostgreSQL is the backend. A project-wide pool lets the benchmark hypothesis leak into release reasoning, where it reads as doubt about a settled decision, and lets release constraints burden the benchmark, where they read as restrictions on an experiment designed to ignore them. Neither item is wrong in its own task. Each is contamination in the other. Chapter 14 already handles relevance; this chapter’s narrower point is membership — and the flow across scope levels runs asymmetrically. Project-wide constraints legitimately inherit downward into tasks, but task-local hypotheses must not bubble upward into project state without an explicit act. Promotion is rarer than readmission by hypothesis: a result needed once elsewhere travels by transfer or reference rather than graduating to permanent standing context, lest incidental state accumulate into the project’s permanent background.

Crossing itself is a first-class operation, not a policy failure. The durable mechanism: identify the required external item, cross it explicitly, preserve its origin scope alongside the target, and admit it as transferred rather than native. Transfer never rewrites origin — a Project B artifact read in Project A remains Project B evidence — and never confers authority or currency, which stay with Chapters 19 and 20. Eligible for consideration here means exactly that, nothing more. Shared dependencies complete the picture without duplicating it: an organisation API contract or coding standard consumed by several projects deserves explicit shared scope rather than repeated one-off transfers or silent duplication, and the conveniently global default is refused — scope marked global must genuinely apply across its computations, because unmodelled scope is precisely how local assumptions travel. Consider the contract’s fourth version governing authentication across both projects: admitted once under shared scope with its version attached, it serves every computation that names it, while neither project’s local overrides leak outward through the shared item. Duplicate it per project instead and the copies drift; transfer it per task instead and the audit trail multiplies without need. Shared scope is the third mechanism alongside membership and transfer, chosen where the source genuinely serves several worlds at once.

The flows this section describes, and the one that is blocked, in one picture:

    flowchart TB
    SC["Shared scope<br/>organisation contract, one version"]
    subgraph PA["Project A"]
        PC["Project-wide constraints"]
        T1["Task: release"]
        T2["Task: benchmark<br/>local hypotheses"]
    end
    BA["Project B artifact"]
    SC -->|"named by the computation"| T1
    PC -->|"inherited downward"| T1
    PC -->|"inherited downward"| T2
    BA -->|"explicit transfer,<br/>origin kept"| T1
    T2 -. "no upward flow<br/>without an explicit act" .-x PC
  

What scope dissolves and what it costs

Two of the book’s standing puzzles dissolve on contact with scope. The PostgreSQL-versus-SQLite disagreement from Chapter 19 is not a conflict at all once provenance shows different project homes: simultaneously true in different worlds, no adjudication required — and the authority policy that would have arbitrated between them stands down correctly, having nothing to decide. The production-at-v4 versus staging-at-v5 tension from Chapter 20 is parallel current state, not staleness: different environment scopes, both valid now, with freshness policy validating against the task’s world rather than the newest one. In both cases the apparent disagreement was never about truth or time; it was about an omitted variable, and scope supplies it. Authority and freshness both turn scope-relative in the same move — a decision record governs inside its project, a commit describes its own worktree — which is why scope metadata must be known before either policy runs.

The economics run both directions and the experiment prices all of them. Isolation removes standing context, repeated unrelated history, and sibling chatter; explicit handoffs add summary and reference tokens, coordination calls, and artifact management. No prompt is declared cheaper in isolation — whole trajectories are measured, and constantly reshaped task scopes may disturb the stable prefixes Chapter 9 reuses, recorded without optimising scope for cache.

Representation gains a small win along the way: a backend value with its project attached is unambiguous where the bare value collides, though metadata never substitutes for omitting what is irrelevant. Scope-eligible never means admitted — project membership passes eligibility while relevance, freshness, authority, budget, and representation still decide — which is what makes this chapter safe for Chapters 22 and 23 rather than a second admission system. The distinction matters in practice: Project A’s entire history may clear the scope boundary on a migration task while the budgeted bundle admits only the decision record, the current config, and the failing test output. Scope answers which worlds may speak; the remaining policies decide who gets the microphone. Collapsing the two would turn every in-scope project into a resident one, recreating at project scale the bloat Part II spent eight chapters removing.

Deterministic mismatches filter before model reasoning wherever the runtime knows both bindings, with the reason preserved in the trace for later debugging; model judgement is spent nowhere a comparison could serve. A rejected candidate leaves a readable record rather than a silence: which item, which bindings it carried, which active world refused it, and which transfer or sharing rule was checked and found absent. Months later, when a task fails for want of evidence that existed somewhere in the system, that record answers in one lookup whether the evidence was missing, mis-scoped, or correctly excluded — the difference between a retrieval problem, a policy problem, and no problem at all.

Proposed experiments

The questions. Does scope enforcement remove contamination without losing the evidence a task needs, and does isolating a subagent trade contamination for omission?

The design, in brief. Fixtures carry deterministic scope on every item, because the primary result must test enforcement and not inference. The first experiment uses one project on its main branch, the same project on a feature branch, and a second project, with colliding paths, entity names and issue numbers but conflicting valid facts in each world and one legitimately shared dependency, so that blocking everything external fails honestly. Authority, freshness and representation are frozen: every conflicting fact is correct, authoritative and fresh in its own world, so only membership varies. Conditions run from a global pool and a pool with visible labels, through project-only, project-plus-environment and task-local eligibility, to explicit transfer of the one shared artifact and an oracle. The second experiment gives a parent release task three parallel subtasks, each with shared project constraints, its own evidence, one misleading sibling item and one result the parent needs, and compares one shared context, fresh children with objectives only, children with bounded shared project context, parent-mediated communication, artifact handoff, and an oracle.

The measurement that matters. Leakage reported twice, as admission of an out-of-scope item and as its behavioural use; wrong-world errors seen as wrong commands, paths and rules; over-isolation, scored as false rejection of the shared dependency; and coordination cost in tokens and steps.

What would change the book. If project root alone eliminates all tested leakage, everything fancier is deleted, and one explicit cross-scope interface beats a federation layer until a measured failure says otherwise. If labels alone match removal, the saving is convenience. Nothing here has been run.

What the compiler will be given

Scope is eligibility for one computation, and the compiler receives it as a judgement, a yes or no per candidate with a reason, against the request’s active scope. It does not discover which project a file belongs to. If a builder marks another project’s configuration as in scope, the compiler admits it. What the compiler does guarantee is that a candidate marked out of scope is removed at a gate before anything is compared, however closely it resembles what the task needs. In its own cases a configuration file from another project scored 0.99 for relevance and was rejected.

Deriving scope from repository and workspace identity, session metadata and explicit transfer records is candidate-building work, argued here and not tested. The observer records session, agent and model identity and the assembled context, but not project, task, subagent-relation or worktree semantics, so future adapters would need opaque workspace, branch, session-relation, task and transfer identities, never private paths in public exports. No ecological leakage claim exists.

Where the chapters have arrived

Authority, freshness and world-membership are now decided, or at least argued, and each has become a judgement a builder attaches to a candidate. That is not the same as a bundle. The candidates that survive are eligible, and they still compete for one budget, depend on one another, come in several forms, and include some that must be present and some that need not. Ranking them does not assemble them, and each of those difficulties has appeared in earlier chapters on its own. The question the whole book has been walking toward is now unavoidable: given every eligible candidate with every constraint attached, what exact bundle should this computation receive?

References

  • Memory book (sibling manuscript, unpublished). Finding cited from the sibling project’s project-scoping experiment, re-checked against its stored results: cross-project leakage of about 0.21 in the pooled condition falling to zero under project scoping with must-include recall unchanged at about 0.77, and 0.67 falling to zero on the similar-project fixture. The frame-inference warning is background from the manuscript and is not evidence here. Corpus and reader bounds apply; the finding motivates this chapter’s fixtures and is not independent evidence for them. The evidence register in the companion repository says where the underlying run can be inspected.
  • Zhang, Y., Liu, F., Shan, Y., et al. “SILO-BENCH: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems.” Peer-reviewed, ACL 2026 long paper (San Diego, pp. 29379–29398). Fragment-only coordination with the Communication-Reasoning Gap and hardest-tier collapse. Used for the too-little-sharing principle at its stated size; scale figures never transferred. https://aclanthology.org/2026.acl-long.1354/
  • Patel, N. “Dynamic Attentional Context Scoping: Agent-Triggered Focus Sessions for Isolated Per-Agent Steering in Multi-Agent LLM Orchestration.” Preprint, arXiv:2604.07911 v1, April 2026. Registry-plus-focus asymmetric isolation with reported contamination reductions on synthetic scenarios. Single-author preprint; narrow mechanism use only. https://arxiv.org/abs/2604.07911
  • OpenAI. “Worktrees.” First-party product documentation, current pages, verified September 2026. Parallel chats in isolated repository copies with managed handoff. Used as environment-isolation implementation evidence; filesystem isolation explicitly distinguished from Context isolation. https://learn.chatgpt.com/docs/environments/git-worktrees
  • OpenCode. “Agents.” First-party product documentation, current V2-line pages, verified September 2026. Primary and subagent types with child sessions, own prompts, models, and permissions. Cited as present-day implementation evidence; the observer does not record subagent relations. https://opencode.ai/docs/agents/
  • Anthropic Applied AI team. “Effective context engineering for AI agents.” First-party engineering essay, September 2025, verified September 2026. Subagent architectures with clean windows and distilled returns, used for the bounded-handoff pattern. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
  • Hadfield, J., et al. “How we built our multi-agent research system.” First-party engineering essay, Anthropic, June 2025, verified September 2026. Separate windows with condensed findings and filesystem artifacts behind lightweight references. https://www.anthropic.com/engineering/multi-agent-research-system
  • boxpositron. “WithContext MCP Server.” Third-party implementation, MIT licence, verified September 2026. Project-scoped folders against cross-contamination. Implementation evidence only; folder scoping never claimed sufficient. https://github.com/boxpositron/with-context-mcp