← Memory From First Principles

What Did We Leave Unfinished?

The past creates requirements on the future. Memory must preserve those requirements until later events satisfy, cancel, or supersede them.

Chapter 8 left the system with a way to maintain belief through time. That system is a good historian. This chapter shows where it is a poor colleague: it can reconstruct the past, but it cannot yet tell a closed chapter from an open obligation.

The migration is finished. The work is not.

Extend the running migration history past the July decision:

Thursday 11 July, adr-007:
Decision: move the event store to PostgreSQL.

Monday 15 July, commit-112:
On the migration branch, application writes use PostgreSQL in staging.
Production still uses SQLite.

Wednesday 17 July, session-051:
"Importer is green on PostgreSQL. Backups still target
SQLite — need to move those before release."

In the canonical history this opens intent-401, which stands open at the 23 August cut and completes with commit-118 on 27 August, three days before release-024 ships on 30 August.

Ask the Chapter 8 system what the migration state is. It answers correctly: the store moved to PostgreSQL, the decision is current, the SQLite era is historical. Every Question 1–4 metric passes. Now hand it a present task at the 23 August cut: prepare the release. It proceeds even though the backup obligation is still open and the benchmark fixtures still encode SQLite assumptions. At the same time, it may falsely report the deployment documentation as unfinished because the completion commit never linked itself to the earlier obligation.

The failure now runs in both directions. The system can omit genuinely unfinished work, and it can keep silently completed work open forever. The migration was an event with future consequences, and the system recorded the event while losing track of those consequences’ changing status. This is the book’s fifth question — what did we leave unfinished? — and the naive view conflates the two things it must separate:

An obligation has a small set of stored states, and later evidence can move it out of OPEN. Remaining unresolved is not another transition; it is the query-time result when no valid close has been found. The state machine makes that distinction visible:

    stateDiagram-v2
    [*] --> OPEN
    OPEN --> SATISFIED
    OPEN --> CANCELLED
    OPEN --> SUPERSEDED
  

Each state-changing arrow has a distinct evidence requirement. Remaining OPEN requires absence reasoning rather than another state transition.

Underneath the state machine is the distinction the naive view conflates:

past event

versus:

past event with unresolved future consequences

The historian and the colleague

The distinction is worth stating plainly because it organises the whole chapter. The historian knows the migration happened. The colleague knows the migration happened, but backups still need to move before release. The difference is not more historical accuracy. It is that one system carries an expected future transition forward while the other discards it at the moment of recording.

That makes Question 5 genuinely different from Questions 1–4, and it yields a measurable condition the harness can plant: histories over which every reconstructive check passes while the assistance check fails. The E9-K experiment constructs exactly that case — opening present, closing present, Chapter 8 belief correct — and shows the mention-based reader still reporting the backup task open forever after its completion. Perfect history, failed colleague, on the record rather than as rhetoric.

Chapter 8 remembers what happened

The substrate is reused, not rebuilt. Chapter 8 implemented the actual trajectory:

state
→ event
→ state

with an append-only log, deterministic replay, and bitemporal queries. Chapter 9 adds a second object alongside it:

state
→ expectation
→ expected future state

No second history store exists. Conceptually, the append-only event log feeds the temporal resolver, and the resolver feeds expected-transition status: actual events satisfy, cancel, or supersede expectations. Expectations themselves are derived state — produced from historical events by derivation the system records — never written back into the raw history as if they had been observed.

What was supposed to happen next?

Given the July remark, the system should hold something like:

current:
backup.backend = SQLite

expectation (intent-401):
before release-024:
backup.backend → PostgreSQL

The expectation is not yet an event. It describes a required future transition that has not yet been observed. That is a new class of memory object, and the chapter’s central definition follows:

An open loop is an established expectation about a future state transition for which no satisfactory closing transition has yet been observed.

An event says what happened. An expectation says what has not happened yet but is supposed to. An open loop is the gap between an established expectation and the observed trajectory. The definition covers tasks, promises, required follow-ups, deferred changes, and expected state transitions without prematurely creating five separate ontologies — and it draws the boundary the experiment needs: a missing event alone means little, while an expected event still missing is the entire subject.

An event is not an expectation

One boundary governs the whole experimental design. Chapter 9 does not ask whether a sentence was really a commitment. That question — suggestion versus promise, musing versus obligation — was deferred rather than answered: commitment extraction was archived unresolved, and Chapter 11’s S2 gate rejects inferred expectations instead of admitting them. For the primary experiment here, openings are oracle-labelled or ledger-established: the benchmark states outright that a genuine obligation opened, and Chapter 9 asks only whether the system maintains its status correctly as later history arrives.

Chapter 9 does not infer commitment; every result here is conditional on oracle-labelled openings.

The separation is what makes the result interpretable. If status resolution fails even with perfect openings, the status mechanism itself has failed. If it succeeds, the success is cleanly conditional — status maintenance earned, extraction still owed.

Open is not a keyword

The serious baseline is mention-based detection, implemented plausibly rather than as a straw target. It matches task-like patterns — TODO, need to, should, must, before release — with deduplication, in two strengths: mention-only, and mention plus nearby closure words.

Measured on the canonical texts, it finds four of five genuine mentions at precision 0.8. The two errors point in opposite directions.

A session says we should probably use Redis. That gets flagged open, although the team never accepted it.

The SQLite tuning obligation is phrased as an issue rather than a task — tune SQLite write performance. That is missed entirely.

So mention-based detection is neither fully precise nor complete. It flags what was never promised, and misses what was never task-worded.

That measured profile, not an assumption, is why the chapter needs a status representation.

The failures separate four responsibilities. Never-accepted mentions require commitment evidence, which this experiment isolates with oracle labels; silent completions require cross-artifact linkage; completed-then-mentioned-again requires maintained status against stale-task errors; cancellation and supersession require Chapter 8’s temporal machinery applied to expectations rather than beliefs.

The smallest useful state machine

The representation that carries the distinction is small by design:

OPEN
 │
 ├── completion evidence ─────► SATISFIED
 │
 ├── cancellation evidence ───► CANCELLED
 │
 ├── superseding event ───────► SUPERSEDED
 │
 └── no valid close ──────────► remains OPEN

UNRESOLVED is not a fifth stored state. It is computed: open at query time with no valid closing transition — the same query-time temporal semantics Chapter 8 uses for belief, now over expectations. UNKNOWN covers insufficient evidence or history, and a deadline is a derived flag (OPEN with deadline_passed) rather than a further state, keeping the machine at four.

The expectation record itself is deliberately spare: identifier, subject, a declarative desired transition (backup.backend equals PostgreSQL — serializable, versioned, never executable code), opening timestamp, opening evidence with derivation provenance, optional deadline or trigger bounds, scope. No owner, priority, points, or workflow state. Those would turn it into a tracker; the chapter wants memory.

Satisfaction is deliberately the general term rather than completion: an expectation can be satisfied by an observed configuration state even where no explicit task-completion message exists. Cancellation ends the obligation without fulfilment (staging backups ruled out of scope). Supersession retires it because the world moved on (the issue-041 SQLite tuning task after the adr-007 migration). The tuning and backup expectations share vocabulary but never merge: superseded-by-migration and satisfied-by-migration are different transitions on different objects, and the suite checks both.

The same obligation, four futures

The decisive experiment holds the opening fixed and varies only the subsequent trajectory. From one expectation — move backups to PostgreSQL — four histories follow: a migration commit (satisfied), an out-of-scope decision (cancelled), a replacement architecture (superseded), and unrelated work followed by release (open). The frozen E9-B run resolves all four correctly, with the closing evidence and tier recorded per case. The opening content stays fixed; only the later trajectory varies, and status changes exactly as labelled. That is the Chapter 8 permutation discipline applied to absence.

Frozen run: ch9-20260920T000629Z-open-loops (Type A, conditional on oracle openings).

The release scenario then exercises the mechanism across six expectations in one project history: the event-store migration complete, the backup migration satisfied by commit-118, the docs update satisfied, the fixtures updated by refactor, the tuning task superseded, and the staging migration cancelled. Each carries its own evidence; project-level completion does not imply that every consequence has closed.

Unresolved is computed

Status takes a time parameter, for the same reason belief does.

The backup expectation is correctly open at two earlier standpoints and satisfied at a later one. One function, three standpoints, verified by the history trace in the frozen artifacts.

Late-arriving completion reuses Chapter 8’s bitemporality. A migration that happened on one date but was learned about nine days later yields two correct answers for the day after it happened — actually satisfied, believed open — without rewriting what the system believed at the time.

A deadline obligation behaves differently again. Token rotation before a fixed date stays open with the flag clear before the bound, and open with deadline_passed after it. Once the bound has passed, a bounded obligation can be resolved against a sufficiently complete finite trace in a way an indefinite “eventually” promise cannot; the run checks both sides.

The hardest evidence is missing evidence

Chapter 8 works with positive evidence: events occurred. Chapter 9 must reason about expected events not observed — and “not found” does not imply “did not happen.” The chapter’s position, repeated wherever absence is scored, is:

Open status is strongest when positive opening evidence is combined with a sufficiently complete search for closing evidence.

The mechanism before any probabilities is the search footprint: which channels were checked (sessions, commits, diffs, issues), through what time, how many candidate closures examined, what gaps encountered, whether current state was consulted. An open verdict on the full canonical history reads “searched 4/4 expected evidence channels”; the same verdict with channels missing reads INCOMPLETE_SEARCH. The E9-F run shows the footprint changing auditability while the status stays identical — a legitimate outcome the metrics record rather than punish.

Sequence gaps from Chapter 8 propagate directly: a skipped sequence number where the completion could have hidden forces UNKNOWN with incomplete-search confidence, never confident openness. The E9-G run plants exactly that omission and passes. Numeric confidence is refused throughout; calibrated classes (CORROBORATED_OPEN, INCOMPLETE_SEARCH, UNKNOWN) carry the uncertainty the record actually leaves.

The footprint design starts with the simpler measurable requirement. A probabilistic absence calculus — the odds of openness given nothing found in searched set S — is not implemented. In v0.1, every open verdict instead carries the scope of the search that produced it. The frozen runs show that the footprint distinguishes complete from incomplete searches while leaving status unchanged when appropriate, so “no completion found” remains an inspectable search claim rather than a claim that nothing happened.

Completion may happen somewhere else

Fulfilment routinely lives in another artifact type than the promise, so closure evidence is tiered by strength and each tier is measured separately:

  1. state transition — the desired state observed (diff, config, state event);
  2. explicit — authoritative complete/cancel/supersede markers;
  3. semantic — commit text overlapping the opening statement;
  4. textual — a “done”-style claim, weakest.

Two rules keep the tiers honest.

Mentions never close mentions. Candidacy is restricted to work artifacts, so a remark cannot satisfy itself.

Closings must follow openings in event time.

Two runs show the tiers working. One links a session promise to a commit whose message shares no task vocabulary with it — the semantic tier, correctly. Another closes the fixtures debt through a refactor whose message mentions neither fixtures nor completion — the state-transition tier, where both mention baselines see nothing at all.

The adversary runs alongside. A “backup cleanup” commit that moves no backend must not close anything, and it does not.

Fixtures were checked for accidental lexical overlap before freezing: weak overlap by construction, never verbatim repetition.

Maintain it or derive it?

The chapter’s architectural comparison runs three conditions over identical histories:

  • Full-history derivation. Status recomputed from scratch at query time, with no stored state.
  • Maintained expectations. An incremental open list.
  • Maintained plus snapshot. The same, with a current-state snapshot as extra corroboration.

Quality agrees exactly between derivation and maintenance on every canonical expectation. So the representation earns its keep, while a separately maintained store is not required for correctness at fixture scale.

What differs is cost and risk.

Listing from the maintained projection avoids replaying the full history; its cost scales with the maintained open set instead. Derivation replays history per query — about 22ms at ten thousand events, linear throughout.

But the maintained list goes stale when closure evidence lands between refreshes: three of three post-closure cycles stale-open in the frozen simulation. Mandatory re-verification repaired all three.

The snapshot condition changed confidence metadata only, never status. History-only status stands without the perception leg, which remains Chapter 11’s contract to build.

That is the Type-B-shaped nuance inside a Type-A chapter, reported rather than smoothed over.

Open lists accumulate mistakes

The re-verification finding is the chapter’s strongest systems result, because the error dynamics are structural.

Belief errors repeat per query. Status errors persist across runs — a wrongly-open item stays open until something re-examines it.

The frozen simulation shows it plainly. The never-rechecked list is stale on every cycle after delayed evidence arrives. Periodic re-verification repairs all three stale post-closure cycles in that simulation.

Re-verification carries its own provenance — expectation, last-checked time, footprint, result. So “open and freshly checked” stays distinguishable from “open but unexamined for months”, recorded as metadata rather than as another state.

Human prospective-memory research reports the same shape from the other side: completed intentions keep residual activation and interfere until actively stood down. That is cited as analogy for the maintenance requirement, not as evidence for the mechanism, which stands on the ledger numbers.

Did open-loop memory earn its place?

Pre-registered types resolve as Type A on fixture evidence: given oracle-labelled openings, expectation status repairs the mention baseline’s planted false-positive, silent-completion, and supersession cases, with explicit uncertainty where history is gapped and recorded search footprints elsewhere. Canonical status accuracy is 1.0 across six expectations; the H-family, cross-artifact, side-effect, bitemporal, deadline, and second-domain (flag disablement after incident) dimensions all pass on the frozen fixtures.

The demotion clauses travel with the verdict.

Openings are oracle-labelled, so nothing here establishes commitment extraction. Every result in this chapter is conditional on that.

Maintenance and derivation agree on quality. The maintained store therefore has no correctness advantage on these fixtures; its case is list-query cost, offset by the need for re-verification.

Snapshot corroboration adds no status information at fixture scale, so Chapter 11’s perception leg stays a corroboration contract rather than a prerequisite.

Cross-artifact linkage is deterministic on fixtures with weak but non-zero overlap. Learned linkers and silent real-world completions remain open problems.

And because full-history derivation matches maintenance, the chapter’s durable contribution may turn out to be the expectation representation plus on-demand resolution, rather than the store itself. The experiment was designed to permit exactly that simplification.

If a principle survives, it is this: memory does not only need to preserve what happened; it must sometimes represent what was supposed to happen next. And the more precise form: an open loop is not a missing event but a missing event relative to an established expectation. Both conclusions are bounded by the fixture scope of the frozen runs.

What we still do not know

Given a valid expectation, the system can now maintain its status. Two questions follow.

The first is where expectations come from: what language and action create a genuine future-directed commitment rather than a suggestion, prediction, wish, or idea. That question is real and unresolved: the intention-record plan was archived without a run, and no experiment has established commitment extraction.

The second is larger, and the accumulated machinery now makes it answerable. Seven layers can say a great deal about history. Whether any of it improves the one thing memory exists to do — put the right past in front of present work — is what the next chapter puts on trial, against the Chapter 3 baseline:

What Matters Right Now?

The chapter leaves commitment-evidence extraction with its quarantine rules, probabilistic absence modelling, richer obligation logic (conditions, conflicts, delegation), reopened obligations, numeric confidence, and the full snapshot contract unresolved. Chapter 11 still takes derived expectations from state mismatch under its own abstention discipline.

Research foundations

Prospective memory research provides the structural vocabulary: the retrospective/prospective cut, and the event-, time-, and activity-based trigger taxonomy (Einstein and McDaniel; Zuber and colleagues). Work on intention deactivation motivates the standing re-verification machinery (Walser, Fischer and Goschke).

Formal methods sharpen the obligation shape. Liveness (Lamport) captures the idea that some commitments assert a future event must eventually occur. Runtime verification (Leucker and Schallhart) supplies the monitor metaphor: status over a finite trace is a verdict of satisfied, violated, or inconclusive.

Software-engineering empirics ground the baselines. Self-admitted technical debt is common and often never removed (Potdar and Shihab). Obsolete-TODO detection needs comment–change–message triples, with word-overlap baselines performing worst (Gao and colleagues). Cross-artifact linkage is inherently approximate (Fischer, Pinzger and Gall).

Agent-memory systems show persistent state influencing plans without maintaining actionable obligations (Generative Agents; MemGPT; LoCoMo). PM-Bench provides an external comparison on a similar problem shape: its authors report difficulty with deferred intentions, rescheduling, and cancellations, with ledger-style scaffolds improving precision (Liu and Gabriel, author-reported numbers).

Deontic logic (von Wright) is noted and deferred beyond the active-obligation fragment.

References

  • Gilles O. Einstein and Mark A. McDaniel, Normal Aging and Prospective Memory (1990).
  • Sascha Zuber and colleagues, Remembering Future Intentions (2024).
  • Moritz Walser, Rico Fischer and Thomas Goschke, The Failure of Deactivating Intentions (2012).
  • Leslie Lamport, Proving the Correctness of Multiprocess Programs (1977).
  • Martin Leucker and Christian Schallhart, A Brief Account of Runtime Verification (2009).
  • Aniket Potdar and Emad Shihab, An Exploratory Study on Self-Admitted Technical Debt (2014).
  • Zhipeng Gao and colleagues, Automating the Removal of Obsolete TODO Comments (2021).
  • Michael Fischer, Martin Pinzger and Harald Gall, Populating a Release History Database (2003).
  • Georg H. von Wright, Deontic Logic (1951).
  • Genglin Liu and Saadia Gabriel, PM-Bench: Evaluating Prospective Memory in LLM Agents (2026).
  • Park and colleagues, Generative Agents (2023).
  • Charles Packer and colleagues, MemGPT (2023).
  • Adyasha Maharana and colleagues, LoCoMo (2024).