Memory Is a Context Source
Connect the Memory book to Context without collapsing the two problems.
The current request says, verbatim: “Do not modify the existing migration.” The admitted context contains months of migration history — previous plans, old decisions, the full archaeology of the project — alongside that one-sentence instruction. The model answers the history instead of obeying the instruction, and helpfully proposes the modification it was told not to make. Every remembered item was valid and correctly stored, retrieved, and rendered. The failure is none of those stages; the failure is that durable past information entered a computation it should never have influenced. That is the entire subject of this chapter, and restraint is its hardest case.
Memory is not the prompt
The two books define their objects differently, and the difference is the chapter:
Context is the information made available to a model for a particular computation.
Memory is when retained past experience changes what the system does now.
The first is about a single invocation. The second is a behavioural claim across time, tested counterfactually: take the history away, run the same task, and watch whether behaviour differs and whether the difference is an improvement. Two consequences follow immediately. A memory system may hold thousands of durable items, none of which is currently in context. And a context bundle may hold the user request, system instructions, tool definitions, fresh observations, retrieved documents, agent-generated state, and memory candidates — most of which are not memory. So:
memory stored
≠
memory in context
The sibling book’s durable principle transfers verbatim, because it is the same boundary from the other side:
Memory is durable. Context is selected.
Memory preserves information because it could matter later. Context admits information because it matters enough for this computation now. Those are different policies, run by different layers, measured by different outcomes. Durability protects future availability. It does not grant permanent context-window residency — a store holding an architecture decision, an old incident, a preference, a failed approach, a benchmark, a superseded configuration, and an unfinished obligation has said nothing about which of them belongs in the next invocation.
Where the Context book begins
PAST EXPERIENCE
↓
MEMORY SYSTEM
↓
durable memory store
↓
memory candidate generation
↓
MEMORY CANDIDATES
↓
context admission
↓
CONTEXT BUNDLE
↓
MODEL
↓
BEHAVIOUR
The Context book begins primarily at MEMORY CANDIDATES. Storage, consolidation, derived state, temporal validity, provenance, belief, framing, compression, and forgetting live on the Memory side and are not re-taught here. The chapter assumes the memory system can supply an item with identity, representation, provenance, validity metadata, and scope, and asks exactly one question:
Should this memory candidate influence this computation?
That question extends the book’s running equation by one stage on the left:
remembered
≠
retrieved
≠
admitted
≠
used
≠
helpful
The sibling book’s evidence ladder already expresses the same idea — past preserved, available, retrieved, judged relevant, admitted to bounded context, used, behaviour changed, behaviour improved — and this chapter reuses it selectively rather than re-deriving it. The Context-specific cut is the middle three rungs: candidate, admission, current context. Everything the memory system does before candidacy is consumed as interface. Everything the model does after admission is behaviour, measured downstream.
The states involved compare cleanly enough to tabulate, which is worth doing once so later chapters can point at rows instead of renegotiating meanings:
| State | Durable? | Currently resident? | Memory? | Context? |
|---|---|---|---|---|
| Old project decision in the store | yes | no | yes | no |
| Active task plan in the window | no | yes | no | yes |
| Current tool result in the bundle | no | yes | no | yes |
| Memory candidate awaiting admission | yes | no | yes | not yet |
| Admitted memory in the bundle | yes | yes | by origin | yes |
The last row is the chapter’s entire mechanism in miniature: an item can be memory by origin and context by current position, and neither property implies the other.
Memory plugs into the common interface
Chapter 14 built one admission architecture with a single candidate type. Memory enters through it like every other source:
retrieval ───────┐
memory ──────────┤
tools ───────────┤
external files ──┤→ ContextCandidate → admission
agent state ─────┘
There is no parallel memory-context pipeline in this book. Building one would undo Chapters 14 and 15 in a single diagram. A memory candidate carries source_kind = memory with the common fields — source identity, representation, token cost, provenance, scope, validity metadata — and the generic admission policy decides. No MemoryContextCandidate subtype unless implementation later proves memory-specific fields cannot fit the common interface. The desired end state, held explicitly as a falsifiable preference, is a memory-source adapter feeding the common candidate type into generic admission, with no memory-specific admission policy at all.
That flatness forbids one architectural mistake by construction. Memory says X, therefore show X is never valid. The structure is always memory says X, therefore candidate X, therefore is X appropriate now — because memory items arrive with all the ordinary defects of candidates: irrelevant, redundant, superseded, wrong, too broad, too expensive, lower-authority, out of scope. Some of those properties are the memory system’s responsibility to mark. The rest are admission’s to judge. And memory carries no privilege of origin. A remembered assistant suggestion is not a user instruction, a project policy, or a decision record until Chapter 19 says otherwise; Chapter 16 preserves the producer metadata Chapter 19 will need and adjudicates nothing.
Working state ends where reuse begins
Chapter 15 asked what state an agent creates during its trajectory. This chapter asks what survives it. The boundary is functional, never a clock reading: working state maintains the current trajectory — goal, plan, active hypotheses, progress, open step, recent summary — while memory retains past experience for potential reuse beyond the immediate active trajectory. A progress.md that survives overnight because the task continues tomorrow is continued working state, not long-term memory. A six-month-old decision consulted by a new task is memory, however briefly it is read. The pragmatic questions are whether this belongs to the currently active task state, or whether it is retained past experience that must be selected because it may matter again. Overlap is expected; state graduates from one to the other when the task ends and some of it is retained for later tasks. The book does not draw a sharper ontological line than the system needs, and it certainly does not taxonomise by hours elapsed.
That graduation is where Chapters 15 and 16 meet: agent-generated state, persisted past its trajectory, retained past its task, becomes a memory candidate. Concretely, the migration task ends; its progress file is discarded with the working state, but the decision record — PostgreSQL selected for the event store, with the SQLite attempt and its failure attached — is retained. Weeks later a new task asks about store latency, and that record arrives as a candidate with source, time, and status intact. Nothing about the record changed at graduation except its role: from maintaining a finished trajectory to offering a reusable past. The retention and consolidation policy deciding what graduates stays in the Memory book. This chapter starts when the graduate arrives as a candidate.
What the sibling project measured, used sparingly
The Memory book is an unpublished sibling manuscript, so this chapter uses it in two separate ways. Its concepts, that memory is durable and context is selected, motivate the design here and are background, not evidence. Its experiments are cited as findings, each with its limits, and the evidence register in the companion repository says where their runs can be inspected. Neither counts as independent evidence for this book’s own claims. The sibling project asked a plain question: if a task depends on remembered project history, does supplying the right memory change what the reader does, and does supplying the wrong memory hurt? It held the task, the reader and the prompt fixed and varied only the supplied memory. Assembled, structured memory roughly doubled task success over no memory, from about a quarter to about a half, and beat strong retrieval alone. Removing the one decisive memory pulled success back down, and restoring it brought success back up. Deliberately wrong memory did the most damage: it drove harmful actions, such as acting on a stale record or deleting something the project still relied on. Adding volume without the decisive item did not recover the benefit, so the effect came from information, not from size. And a task that could be answered from the present request alone was made worse by supplied memory, because the reader trusted the store over the request in front of it. These results come from nine controlled tasks and five transfer tasks on synthetic fixtures with small readers, and the chapter claims nothing beyond those bounds.
Two findings from the sibling project’s transfer runs matter more than any headline number. The first is that undifferentiated history is reader-dependent. Handing the small primary reader the complete history left it worse off than handing it nothing, while a substantially stronger reader did well on the same history. That stronger reader is also reported to have gained from assembled memory on tasks whose decisive facts were arbitrary project history; we could not trace that particular result to a stored summary, so we treat it as reported rather than verified, and it carries no weight in this chapter’s argument. Preservation alone does not guarantee the best context, and no universal bundle follows: the minimum sufficient memory context depends on the task, the memory, the representation and the reader. The second finding is that selecting context well and behaving well are different things. Conditioning selection on a project frame improved the selection metrics but did not improve downstream behaviour by itself on either reader, and on the stronger reader a frame-selected bundle even beat the assembled one, reversing the order seen on the small reader. Assembly policy is itself reader-relative, and this chapter measures selection quality and behavioural quality while equating neither.
Two external results reinforce the same boundary from outside either book. Mem2ActBench, an ACL 2026 long paper, builds its benchmark on the explicit gap between passive fact recall and active memory use: 2,029 synthesised multi-turn sessions yielding 400 tool-use tasks, 91.3 per cent judged strongly memory-dependent by human evaluation, with seven tested memory frameworks still inadequate at grounding tool parameters from memory. MemoryArena, a September 2026 preprint, couples acquisition to action across interdependent multi-session tasks and reports that agents near saturation on recall-style long-context benchmarks perform poorly once memory must guide later decisions. Both are used narrowly, for the single point each earns: remembering information is not using memory to act, and recall scores do not certify agentic memory. Their populations, metrics, and limits stay in their papers.
The sibling book’s warnings transfer alongside its numbers. Derived state may be worth persisting for cost, but persistence does not automatically earn epistemic authority: a persistent interpretation preserves reusable mistakes as faithfully as reusable understanding, so durability is not authority and persistence is not current relevance. The chapter does not reopen graph-memory mechanics to say so; it inherits the sentence. Likewise the assembly finding that sufficiency is reader-relative arrives as a constraint on every bundle the Context book will ever recommend.
The restraint machinery
A valid memory can still deserve rejection. “Project uses PostgreSQL for its event store” may be correct, current, and well-provenanced — and the right admission decision on a task reading “Rewrite this paragraph for clarity” is DO NOT ADMIT. Sometimes the correct memory context is empty, and for the negative-control tasks in the experiment, many candidates in with nothing out is the best outcome. Restraint is a capability, not a retrieval failure, which is why the legal outcomes always include the null path.
Restraint has a sharper edge than irrelevance. Historical relevance does not grant priority over present-task requirements: the rich migration history must not outrank “Do not modify the existing migration.” The current task is not merely another memory to be weighed symmetrically. Memory is historical candidate input; the present request defines the computation. Operationally the asymmetry means the task pins its non-negotiables first — instructions, constraints, fresh observations — and memory competes only for the remainder under the budget, never the reverse. Chapter 23’s compiler expresses that asymmetry as classes of requirement, and Chapter 19 formalises whose claims win. Here the failure is exposed and the metadata preserved.
Metadata discipline is what makes restraint implementable rather than aspirational. A candidate stripped to anonymous prose can fuse what the store had kept apart: “SQLite was used before July” plus “PostgreSQL is active” becomes an apparent contradiction the moment the historical/current marking is lost. Work through what correct handling requires. The store supplies both items with validity intervals; candidacy carries both forward because the task mentions the event store; admission, seeing a rewrite task with no historical dependence, rejects both — or, on a migration task, admits the PostgreSQL item at full representation with its current marking and the SQLite item as a one-line historical anchor so the record shows what was abandoned without reasoning from it. Strip the markings and the same two admissions produce either a contradiction or a confident regression. Admission preserves whatever behaviour requires — source, time, status, scope, authority, representation identity — and nothing it does not. Representation itself is an admission choice under Chapter 12’s rules: raw evidence, structured state, summary, anchor, or full artifact are different admittable forms of one candidate, which is one more reason the store and the bundle stay separate systems. Scope metadata travels for Chapter 21, authority metadata for Chapter 19, temporal markings for Chapter 20. None of those chapters is built here; each receives its inputs intact.
Two attributions keep the layers honest. If the required memory never reaches the candidate pool, that is a memory candidate miss, not an admission failure — Chapter 14’s retrieval boundary restated. If the correct memory was stored, retrieved, admitted, and represented faithfully but the model ignores it, that is not automatically a storage failure. And Context never becomes the memory truth-maintenance system: outdated, contradicted, or invalid items are the memory layer’s to expose and repair, with Chapter 20 handling freshness only at the Context boundary. Where memory and current evidence visibly conflict — the store says PostgreSQL was selected, the tool output says production runs SQLite — both keep source and status so later governance can decide, and this chapter decides nothing between them.
Proposed experiments
The questions. Does admitting memory change behaviour, does wrong memory harm, and can the system leave history out when the task does not need it?
The design, in brief. Freeze the memory store first, so that no extraction, consolidation, retrieval or trust mechanism varies while admission is under test: required memories, helpful but unnecessary ones, irrelevant valid ones, redundant ones and historical background. Tasks depend on arbitrary project history that no model recovers from pretraining. The first experiment compares no memory, all memory, a frozen candidate pool admitted whole, selected admission over that pool, an oracle minimal memory, and a wrong-memory positive control kept outside the headline. Negative controls are mandatory: rewrite a sentence, echo the current identifier, transform data fully present now, because a policy that cannot leave history out has failed restraint. The second experiment runs one clearly decisive memory admitted, removed and restored.
The measurement that matters. Behaviour scored as behaviour, not mention: the correct tool, the grounded parameter, the preserved constraint, the avoided failed approach. Influence and improvement are recorded separately, since wrong memory moves behaviour the wrong way.
What would change the book. If all memory matches selected memory at realistic budgets, or the generic admission of Chapter 14 suffices over a memory adapter’s output, this chapter earns no mechanism of its own, which is a good result. Nothing here has been run.
What travels forward
Deliberately little. Memory arrives as candidates with a source kind and provenance: the source item, the source system, validity and status, scope. There is no memory hierarchy, no separate candidate type, and no extraction, consolidation, graph or temporal store inside the compiler, which consumes the outputs of a memory system and must never grow its own long-term belief store. A future adapter may consume candidates from a memory service, but this chapter defines an interface, not an integration.
What remains is the last source of candidates, and the contrast the chapter was built to set up. Memory contributes durable historical candidates that must earn present influence. Tools contribute on both sides of execution at once: standing capability descriptions occupying context before any call, and fresh observations arriving after it. One source asks what the past may still decide. The other asks what the present affords and what it just produced. The next chapter takes the second of them.
References
- Memory book (sibling manuscript, unpublished). Conceptual background, not evidence for this book: the behavioural definition of memory and its evidence ladder, the assembly layer and reader-relative sufficiency, and the final architecture (“Memory is durable. Context is selected”, the strong-retrieval baseline, the derived-state authority warning). Findings cited, each within the sibling project’s fixture, reader and task limits and none independent evidence for this book: the controlled memory-dependent tasks read by a small local reader and by a stronger reader. The evidence register in the companion repository lists where the underlying runs can be inspected. One further reported result could not be traced to a stored summary and is not used.
- Shen, Y., Li, K., Zhou, W., Hu, S. “Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents.” Peer-reviewed, ACL 2026 long paper (64th Annual Meeting, San Diego, pp. 8173–8190). Passive recall versus active memory use; 400 tool-use tasks, 91.3% human-judged strongly memory-dependent; tested frameworks inadequate at parameter grounding. https://aclanthology.org/2026.acl-long.370/
- He, Z., Wang, Y., Zhi, C., et al. “MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks.” Preprint, arXiv:2602.16313 v2, September 2026 (page comments ICML 2026). Acquisition-to-action coupling across interdependent sessions; recall-saturated agents performing poorly in the agentic setting. https://arxiv.org/abs/2602.16313
- Anthropic Applied AI team (Rajasekaran, Dixon, Ryan, Hadfield, et al.). “Effective context engineering for AI agents.” First-party engineering essay, September 2025, verified September 2026. Structured notes persisted outside the window and pulled back later, used only as boundary illustration for persist-then-reintroduce. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents