← Memory From First Principles

What Should Memory Keep?

A long-lived memory must control what continues to compete for present use without destroying the history future work may need.

Chapter 14 ended with an unexpectedly useful result. Once the candidate set was already good, elaborate assembly did not uncover a large reservoir of duplicate text waiting to be removed. Extractive deduplication saved only eighteen of 1,174 counted tokens on the frozen C5 bundles. The larger problem was competition: many individually defensible memories still occupied the same scarce working context, and the system had to decide what should continue to matter for the present task.

Chapter 12 made the cost of getting that wrong behavioural rather than theoretical.

On matched tasks with the small development reader, a system given the full history scored 0.048. A system given no memory at all scored 0.226. Ordinary retrieval scored 0.393.

More of the past, supplied indiscriminately, was worse than none of it — for that reader.

A substantially stronger reader turns the same unfiltered history into genuinely useful context (0.571), and selection still improves on it further.

So growth is not only a storage or retrieval problem. Once retained history competes for bounded present use, it also becomes a competition problem — and Chapter 12 shows that competition can affect behaviour.

Its severity depends partly on the reader. The small development reader is harmed by large undifferentiated history, while the substantially stronger transfer reader exploits the same history far better. On these fixtures, both readers still do better under structured selection.

That observation changes the next question. The problem is no longer three separate topics called consolidation, compression, and forgetting. It is one long-term problem:

As history accumulates, what should remain able to compete for present behaviour?

A memory system can preserve years of history without making every part of that history equally available on every task. It can keep raw evidence while deriving cheaper representations. It can archive a superseded belief without erasing the period in which it was true. It can stop an echo from competing as if it were independent evidence. It can retain a failure for the one future task where that old failure suddenly matters again.

The important distinction is therefore not keep versus delete. It is a four-stage distinction:

preserve
    ↓
represent
    ↓
make available
    ↓
admit to the present context

A raw artifact may survive forever while its derived representation changes, its retrieval availability falls, and its chance of entering a particular ContextBundle approaches zero. Those are different operations with different failure modes. Treating all of them as “forgetting” hides the engineering problem rather than solving it.

The store is not the context

Chapter 8 showed that historical truth and current truth must coexist. “The service used SQLite in March” remains true after PostgreSQL becomes current. Deleting the SQLite history to protect current answers would improve one query by corrupting another.

Chapter 14 showed the complementary point: preserving an item does not imply seating it in the current model context. The store may contain thousands of memories. The model receives a bounded working set selected for the current goal.

These results produce a useful invariant:

Exclusion from the present context is not forgetting, and preservation in the store is not a promise of present influence.

That invariant suggests a conservative growth strategy. The system does not need to destroy history merely because history grows. Before irreversible deletion is considered, it can change which representations compete and under what conditions.

The chapter therefore begins with a conservative bias: preserve first, suppress reversibly, delete only for reasons stronger than convenience.

Some deployments require deletion for reasons independent of memory quality — privacy policy, legal compliance, secret rotation, or user-requested erasure. Those requirements do not establish that age or low retrieval frequency is a sound memory-decay rule.

Three operations, one problem

Consolidation, compression, and forgetting remain distinct operations, but they belong inside one growth problem.

Consolidation changes what the system believes

Suppose several migration episodes all end the same way: destructive schema change arrives before compatible code and the system fails. A memory layer may derive a higher-level claim:

For this project, compatibility should precede destructive
schema changes.

That is not merely a shorter representation. It is a new claim inferred from several experiences.

The danger is epistemic. The system may overgeneralise. One episode may have had a different cause. A later case may succeed safely without compatibility mode. The pattern may hold only for one platform or one era. Several echoes of the same incident may masquerade as independent evidence.

A consolidated claim therefore needs everything the book has already learned to demand from derived state:

statement
support set
scope
validity
exceptions
provenance
derivation time
status

and, most importantly, a way back to the raw episodes.

The raw history remains canonical. A consolidated memory is a hypothesis about that history, not a rewrite of it.

Compression changes representation

Compression is different. A compressed form should not assert a new truth merely because it is shorter.

The useful boundary comes from Chapter 14:

execution-time assembly
    decides what this model receives now

durable compression
    changes how long-lived memory is represented or served

The execution-time problem is already measured. The durable problem remains open.

A project may preserve every source artifact while serving cheaper derived forms for common queries:

raw artifact
    ↓
event / claim
    ↓
compact current-state view
    ↓
task-specific context representation

The cheaper levels answer fewer questions. The raw source remains available when the cheap form is insufficient.

Compression needs an operational test:

Compression fails when information it removed was necessary for the later behaviour the system was supposed to produce.

That definition is task-relative. The same summary can be adequate for “where did we discuss the migration?” and inadequate for “what constraint makes this migration safe?”

Forgetting changes availability

The third operation asks which retained memories should continue to compete strongly for future use.

For this architecture, the useful first interpretation of forgetting is:

Forgetting is availability management before it is deletion.

Several forms fit under that heading:

retrieval suppression
historical-only archival
supersession-aware demotion
context exclusion
replacement by a derived view
hard deletion

They are not interchangeable.

A superseded SQLite decision should be suppressed for a current architecture task but available for a historical query. A verbose meeting transcript may become archival once its decision and evidence have been extracted, while remaining available for audit. A known failed approach may be old and rarely retrieved but extremely valuable when the same risky situation returns.

Age alone is therefore an unsafe default forgetting policy for this design. Usage frequency has the same failure mode: a critical disaster memory may be exactly the item the system has not needed for two years.

Old does not mean useless

A useful control for any forgetting policy is an old failure.

If incident-026 has not been relevant for hundreds of tasks, a usage-based policy may want to suppress it. Then another schema migration arrives. The memory is suddenly one of the highest-value pieces of history in the store.

That does not justify a permanent global importance number. Chapter 10 showed why static importance is the wrong shape. Importance is conditional on present work.

A safer candidate policy is more modest: protect some memory classes from naive decay until evidence shows that suppression is safe. Candidates include unresolved constraints, destructive-action warnings, provenance keystones, evidence for current decisions, memories explicitly pinned by project policy, failures whose recurrence cost is high, and historical state required for audit.

These are policy categories, not a scalar ranking of the past.

Confidence and utility stay separate

A memory can be high-confidence and low-utility, or moderate-confidence and high-utility. The old build server hostname may be known with certainty and almost never matter. A tentative warning about a destructive migration may be uncertain and still deserve retrieval when the relevant task appears.

Truth-confidence asks how well a claim is supported. Utility asks how useful a memory is likely to be for the present work. They are different questions, and one must not silently rewrite the other.

That separation becomes essential in Chapter 16, because outcomes can legitimately change expected usefulness without making a claim more true.

Store it, derive it, or rebuild it?

Consolidation introduces a design choice earlier chapters have already encountered.

Chapter 9 compared maintained open-loop state with deriving status from history. Quality matched on the controlled fixture, while the maintained view bought faster listing at the cost of staleness. The same choice appears here:

raw episodes
    ↓
derive pattern when needed

versus:

raw episodes
    ↓
maintain consolidated pattern
    ↓
keep it current as evidence changes

The second is faster to reuse. It also creates a maintenance obligation.

Every stored consolidated claim must be reconsidered when a supporting episode is corrected, an exception arrives, a source is invalidated, the environment changes, or the claim’s scope expires.

A derived view that cannot be re-earned becomes precisely the stale memory the chapter was meant to control.

The architecture therefore favours rebuildable derived memory. Persistence is an optimisation. Canonical history remains the source record from which derived views can be challenged and rebuilt.

A growth experiment, not three mechanism demos

This chapter has no dedicated growth run. Its growth comparison is a hypothesis, and is labelled as one.

Consolidation, compression, and forgetting could each carry their own experiment, but a common scaling comparison would test them under the same pressure: increase the size and age of one memory world, then ask which policy preserves useful behaviour as competition rises.

The proposed sweep would grow that memory world across orders of magnitude while preserving critical old memories, current decisions, historical truths, echoes, completed work, stale summaries, and ordinary background.

Six candidate policies would compete directly — KEEP-ALL, SUPERSESSION-AWARE, ECHO-COLLAPSE, CONSOLIDATED, ON-DEMAND, and COMBINED — on the frozen Chapter 12 controlled-outcome harness.

Alongside behaviour, the sweep would measure historical recoverability, stale-memory admission, critical-old-memory retention, cost, provenance preservation, false generalisation, and recovery after exceptions.

A decisive control would be a deliberately old catastrophic memory. A policy that improves mean cost while suppressing that memory would fail the control even if its average efficiency improved.

Book hypothesis. Long-term memory quality will depend more on controlling competition than on irreversible deletion. Supersession-aware availability and rebuildable derived views should capture much of the value of forgetting and consolidation while preserving historical recoverability. Automatic consolidation should survive only if it improves downstream behaviour beyond raw episodes or on-demand derivation without introducing harmful generalisation.

That remains a hypothesis in the absence of the scaling experiment.

What the existing runs already constrain

Earlier experiments already constrain what a viable growth mechanism may do, even without a dedicated run:

ChapterConstraint on growth
7Repetition is not corroboration; consolidation cannot count echoes as independent evidence.
8Superseded history keeps historical value; “not current” never means “delete”.
9Maintained derived state buys speed at the price of staleness without re-verification.
10Relevance is conditional on project and work state; no permanent importance value.
11Derived state needs an explicit licence and a conservative gate, with crisp-fixture and no-reader caveats.
14Simple deduplication found little in already-selected context; availability is a policy over distinct evidence, not repetition removal.

Together these earlier results constrain any future growth policy toward a conservative design:

A minimal candidate for long-lived memory

A first growth-policy implementation could begin with one new record: why an item is available or suppressed:

memory item
availability state
policy version
effective time
reason
source / derivation

In this design, suppression does not delete an item, consolidation does not destroy its episodes, and a compact serving representation points back to the source that can reconstruct detail.

The system should be able to answer why did this memory not appear? with the same seriousness Chapter 10 brought to why did this memory enter the context? That makes long-term forgetting auditable instead of mysterious.

Where this chapter stops

Growth management is still memory management. The system is deciding how retained history remains available.

The next step is qualitatively different.

Suppose a migration succeeds. Should the memories involved become more likely to influence the next migration? Suppose a sequence of actions works repeatedly. Should the system extract a reusable procedure?

Those operations do not merely decide what part of the past remains available.

If evaluated outcomes are allowed to change the mechanism or policy that shapes future behaviour, the system has crossed from memory management into learning.

That boundary is the subject of Chapter 16.

Research foundations

Several existing systems motivate higher-level memory without resolving the book’s safety boundary. Generative Agents derives reflections from accumulated observations; ExpeL extracts transferable lessons from trajectories; MemoryBank updates long-term records; RAPTOR explores hierarchical representations for retrieval. These systems support the possibility that raw episodes need not be the only useful representation.

The chapter’s stricter requirement comes from the architecture already built: any higher-level memory remains derived, versioned, scoped, and traceable to sources. A cheaper representation is valuable only while the distinctions needed for future behaviour remain recoverable.

References