← Memory From First Principles

Is It Still True?

Some information lives in the ordered transitions between remembered states. A ZeroMQ-distributed history, a durable event log, and a permutation experiment test exactly when order changes meaning.

The same events, a different story

Consider three remembered events from the migration history:

A = benchmark detects SQLite contention
B = team decides to adopt PostgreSQL
C = PostgreSQL deployment completes

As an unordered set, {A, B, C}, the system knows all three happened.

As a trajectory, A → B → C, it knows something stronger. The benchmark preceded the decision, and the decision preceded the deployment.

Now hold the contents constant and change only the order, to B → A → C. Nothing was added. Nothing was removed.

Yet the benchmark can no longer have motivated the decision. It may corroborate the decision, strengthen confidence in it, or justify keeping it. But a later measurement cannot have been the reason for an earlier choice.

This chapter tests one hypothesis: a memory system is not only a collection of remembered states; for some questions, the ordered transitions between those states are themselves part of what must be remembered. A set of events records what occurred. A trajectory records how one state became another. The experiment asks whether that difference is measurable: hold the memories constant, change only their temporal structure, and check whether the system changes its conclusion exactly when it should — and remains invariant when ordering should not matter.

The hypothesis is stated as a prediction, not a fact. Everything that follows either earns it or narrows it.

The chapter’s central distinction is between historical truth and current truth. A claim can be historically true at one moment and currently false at another, and the same record must preserve both facts without overwriting the earlier truth with the later one.

Two independent axes, and one question that reads them both:

    flowchart TB
    subgraph VALID["valid time — when it held in the world"]
      direction LR
      VA["A holds"] --> VB["B holds"]
    end
    subgraph RECORD["record time — when the system learned it"]
      direction LR
      RA["A recorded"] --> RB["B recorded"]
    end
    VALID --> Q{{"one question,<br/>two standpoints"}}
    RECORD --> Q
    Q --> HIST["What was true then?<br/><b>A</b>"]
    Q --> CURR["What is true now?<br/><b>B</b>"]
    style CURR fill:#3978c5,color:#fff
  

Both answers are correct. They differ only in which standpoint the question was asked from.

Time is not metadata

Treating time as annotation — attach timestamps to memories, sort by them, prefer the newest — fails in two directions at once.

First, timestamps without transition semantics cannot answer historical questions. The project used SQLite in April and PostgreSQL after the July cutover; both statements are true of different times. Sorting passages by date retrieves the right documents for either question, but retrieval is not resolution. Something must still decide which retrieved state held at the requested time, and a timestamp sort is a ranking, not a model of change.

Second, timestamp order alone cannot carry authority.

The Redis history makes the point. session-040 floats the idea. session-044 observes the working set fits in memory. adr-009 decides against introducing Redis.

The decision wins because it is a decision — not because it is last.

Now reverse the pattern: a standing decision, followed by later discussion that revisits it without reopening it. Newest-first ranking now prefers chatter over commitment.

When something was said and what it did have to be interpreted together.

The temporal model therefore keeps six distinctions explicit:

  • History versus position.
  • Supersession versus contradiction.
  • Correction versus revision.
  • Validity as an interval.
  • Recency as a baseline rather than a solution.
  • Uncertainty where conflict is genuinely unresolved.

What they have lacked is machinery — an explicit representation of transitions, tested by permutation.

A set is not a trajectory

The chapter’s representation uses four objects.

State is what currently holds within some scope: event-store.backend = SQLite. Event is something that may change, establish, revise, record, or describe state: a benchmark, a decision, a deployment completing. Transition binds the two: state-before --event--> state-after. Trajectory is an ordered or partially ordered sequence of transitions:

S0 --E1--> S1 --E2--> S2 --E3--> S3

A weak memory retains the endpoint: PostgreSQL is current. A stronger temporal memory preserves how the system got there:

SQLite current
    ↓ contention observed (session-014)
SQLite [problem known]
    ↓ PostgreSQL selected (adr-007)
SQLite [PostgreSQL planned]
    ↓ deployment completed
PostgreSQL current

The second representation answers more questions: what was current in April, what changed, what happened immediately before the transition, which evidence could have informed the decision. Whether that extra machinery is worth its cost is the experiment’s question, not this section’s assumption.

When order changes meaning

For some events, applying A then B reaches a different state than applying B then A. PROPOSE → ACCEPT is not ACCEPT → PROPOSE. DECIDE → REVOKE leaves no standing commitment; REVOKE → DECIDE leaves a live one. CLAIM → CORRECTION revises understanding; CORRECTION → CLAIM reads as a claim made after its own correction, which is incoherent. TASK_CREATED → TASK_COMPLETED closes a loop that the reverse order never opened.

The word for this property is non-commutativity, used here only as shorthand: order changes the resulting state. The chapter does not develop the algebra further because the experiment needs the phenomenon, not the formalism.

The migration fixture exhibits it directly. In the canonical order, the June benchmark is an admissible antecedent of the July decision: it precedes the decision, so it could have motivated it. In the permuted order, where the decision predates the benchmark, the same benchmark with the same contents is admissible only as corroboration. The frozen E8-B run checks exactly this pair and passes: the admissibility verdict changes while the event set is identical. Contents held constant, structure changed, conclusion changed — the predicted effect, measured rather than asserted.

When order should not matter

Sensitivity without discrimination is superstition. Swapping two independent events — a README update and an unrelated CSS fix — must leave the migration conclusions untouched. A system that becomes hypersensitive to ordering, twitching at every permutation, has not learned time; it has learned to hallucinate significance in sequence numbers.

The suite therefore scores both directions as first-class metrics. Order-sensitive accuracy asks whether the answer changed correctly when a labelled permutation should change it. Order-invariant stability asks whether the answer stayed put when a labelled permutation should change nothing. The E8-C run swaps the independent pair and checks three subjects under all five resolver conditions; the temporal resolver is stable on every one. A useful temporal memory is sensitive to order without being superstitious about it.

Before does not mean because

Temporal order rules some causal explanations out more easily than it proves others. If the benchmark postdates the decision, it cannot have been the decision’s antecedent reason — that exclusion is solid. But if the benchmark predates the decision, precedence alone does not establish that it caused the decision. The decision record, its stated rationale, and the evidence lineage must do that work.

This restraint shapes the Chapter 7 integration.

Chapter 7 records which evidence supports which claim. This chapter asks one more question of every support edge: is that evidence temporally admissible for the role it is claimed to play?

An earlier benchmark can support the proposition that SQLite had contention, and it can support a decision as an antecedent. A later benchmark offered as the reason for an earlier decision fails admissibility — while remaining perfectly good corroboration.

The implemented check distinguishes three relation families:

  • MOTIVATED_BY requires precedence.
  • CORROBORATED_BY permits later evidence.
  • CORRECTED_BY requires the correction to follow its target in record time.

Unknown relation kinds abstain rather than invent constraints. The suite verifies both halves: a later benchmark is rejected as antecedent, and accepted as corroboration.

Support says why a claim should be believed. Time says whether the story about when is coherent. Neither collapses into the other.

There is more than one clock

Real histories conflate at least five meanings of time, and the experiment keeps them distinct even where it does not implement all five as independent axes:

  • event time: when the represented event occurred;
  • valid time: when a state actually held;
  • record time: when the memory system learned it;
  • decision time: when a commitment was made;
  • effective time: when that commitment began changing the relevant state.

For questions that distinguish what was true from what was known, the minimum model used here is bitemporal: valid time crossed with record time. Suppose the 22 July cutover reaches the memory system only in August. Then two questions with different correct answers become askable for 23 July: what was actually true (PostgreSQL), and what did the system know at the time (SQLite, if the cutover was still unrecorded). A single temporal axis cannot express both standpoints as the same state history. The query interface therefore carries an explicit standpoint:

state.bitemporal(subject, valid_at=..., known_at=...)

Given what was known by August, what was valid in July — versus what was believed at the time. The E8-E run exercises the same shape with a synthetic late-arrival cutover and passes: the retroactive answer reports PostgreSQL, while the historical-standpoint answer refuses to pretend the system knew it then.

Two related separations matter.

First, a decision and a production state are different claims. From 11 July, new event-store work targets PostgreSQL; production remains on SQLite until the 22 July deployment. The frozen future-effective probe uses a separate synthetic case to check both sides of the same kind of boundary, and passes.

Second, questions of agency — current for whom — are deliberately deferred. The v0.1 standpoint is the project’s. Agent and task standpoints are designed as extension points, not implemented machinery. The chapter claims no more multi-agent semantics than it runs.

Distributed histories do not arrive neatly

With several independent publishers, ordering ambiguity is not a thought experiment but the normal case. The demonstration uses three publishers — benchmark, decisions, deployment — each with its own sequence numbers, exchanging messages through a ZeroMQ forwarder and into a recorder. The recorder preserves both the semantic order and the arrival order, because their difference is the demo’s entire point.

ZeroMQ is transport, not memory. Publishers send through a forwarder into a recorder over PUB/SUB sockets, while an in-memory transport serves the deterministic suite so temporal semantics never depend on socket behaviour. Because publishers have no cross-publisher ordering guarantee, arrival order is only an observation about the network. The memory layer reconstructs history from the append-only log described next, not from socket arrival order.

The live run makes the separation visible. Semantic order in the frozen capture:

ch8-benchmark → ch8-decision → ch8-deploy

Arrival order at the recorder, under controlled send delays:

ch8-deploy → ch8-benchmark → ch8-decision

Resolved temporal order after replay:

ch8-benchmark → ch8-decision → ch8-deploy

Naive arrival-order processing would reconstruct the transition sequence incorrectly, replaying the deployment before the decision that authorised it. The temporal engine recovers the semantic order from event times and explicit causal parents.

One honest qualification belongs here. In this fixture, the arrival-order resolver still lands on PostgreSQL as the current value, because the decision event itself carries that value.

So it reaches the right value through the wrong trajectory — replaying the deployment before the decision that authorised it. That misattributes the transition even where the endpoint happens to coincide.

The failure class is arrival order confused with event order. It shows up in the reconstructed rationale, not in the final string.

The durable log

The durable layer is deliberately unglamorous: newline-delimited JSON over an append-only log, indexed in memory. No database server, no model endpoint, nothing that cannot be inspected with standard tools.

Each line pairs a versioned event envelope — identity, kind, subject and fluent value, the three time axes, explicit relations, schema version — with strictly observational recorder metadata.

Read that metadata carefully. received_at and ingest_seq record when this system saw the message. They never substitute for when the event occurred, or when its state became effective.

Out-of-order arrival has to remain ingestible, so a referential check that fails — a causal parent not yet seen — defers to health reporting rather than rejecting the write. That is deliberate. A transport that refused disorder could never observe it.

Corrections append. They never edit an earlier line. What was originally recorded and what was later learned coexist, and that is what makes the bitemporal queries possible at all.

Health reporting is structural, and the chapter refuses to oversell it: event and source counts, sequence gaps, duplicate identifiers and sequences, unknown references, temporal-causality violations, projection version and drift, last ingest. A clean report means the history is well-formed. It does not mean the history is true.

Trajectory reconstruction

The temporal engine is a deterministic reducer with no model calls:

new_state = reduce(old_state, event)

Replay applies it in temporal order — causal parents first, then event time, then a stable identifier order for genuinely concurrent events. The rules are generic over subjects, keys, values, validity, and relations; no rule names PostgreSQL, SQLite, or any fixture literal. Facts establish state, and observations enrich the record without moving it. Decisions record intent without changing current state until their effective transition arrives, revisions supersede going forward, corrections append revised understanding while preserving the original record, and unknown kinds are recorded without being applied.

Two architectural choices are implemented and priced, rather than merely discussed.

Current belief is a view over the log, and the suite runs both maintenance policies. A computed resolver replays at query time — always fresh, cost growing with history. A materialized projection is maintained during ingest — cheap lookup, with invalidation complexity.

Late arrivals trigger replay from the insertion point, and a rebuild command regenerates the projection from the log. The run verifies that all three resulting digests match, including across a correction.

The measured costs, on synthetic histories up to ten thousand events:

  • Replay time grows linearly — about two milliseconds at ten thousand events.
  • Current-state queries land in the low milliseconds under either policy.
  • Disordered ingest pays suffix-replay proportional to the disorder distance, going quadratic in the pathological all-disordered case.

The result is straightforward. Materialized belief wins for current-state lookups when ingest is mostly ordered. Its advantage is a performance optimisation with a measured price — not a semantic improvement.

No reader or language model participates in any of this. A model may later render temporal answers into prose, downstream of semantics it must not reinterpret.

Partial order is represented rather than papered over.

Where two events share a timestamp and no causal edge connects them, the resolver reports their relative order as unknown. The suite scores an invented total order as a failure.

That matters more than it sounds. With several publishers and imperfect clocks, A and B both preceded C; their order is unknown is a correct answer, and the machinery must be allowed to give it.

Two pieces of heavier machinery were researched and deferred. Vector clocks: source sequences plus explicit causal parents already resolve every ambiguity the controlled fixtures contain. And the full thirteen-relation interval network: the reasoning here uses a small subset of Allen’s relations — before, meets, overlaps, during, equals — which is enough for validity-interval overlap and bounded-unknown answers.

Five temporal conditions

The experiment compares five conditions over identical frozen logs — the benchmark ladder’s ablation discipline applied to time:

  • T0 bag ignores every temporal field and answers from event presence alone;
  • T1 arrival resolves by newest arrival, the naive recency baseline;
  • T1b event-time sorts by event time and takes the latest, repairing transport disorder without modelling transitions;
  • T2 ordered replays event-time order through the reducer’s transition semantics;
  • T3 temporal adds valid/record standpoints, effective dates, and supersession/correction semantics.

Event-time sorting is deliberately isolated as its own condition: if it solved everything, the chapter’s conclusion would be that richer trajectories were unnecessary, and the design invites that verdict.

How the claim is tested

The E8 suite (E8-A through E8-J, plus a second-domain permutation) ran deterministically with zero model calls; both the emulated-disorder and live captures are frozen with manifests recording code, schema, fixture, reducer and projection versions, transport mode, and log digests. Each dimension is reported separately — there is no single temporal score:

DimensionT0T1T1bT2T3
Current-state accuracy1.01.01.01.01.0
Historical-state accuracy (April)0.00.00.01.01.0

Caption: the table above carries the headline contrast (current-state versus April historical accuracy). The remaining per-dimension checks — admissibility flip under permutation, stability under independent swaps, arrival-disorder resolution, actual-versus-known standpoints, current-versus-planned across the effective boundary, correction with pointer intact, refused total order, gap detection, and digest equivalence — pass under the temporal resolver and are recorded in the frozen run manifest.

Recency is treated fairly here, and on clean monotonic histories it wins. Every condition identifies PostgreSQL as current.

Reaching into the past is where it breaks. Without transition semantics and a valid-at standpoint, even perfectly sorted timestamps get the historical question wrong — bag, arrival, and event-time conditions all score zero on it.

The suite adds a strategic-vagueness control. Where the log determines an exact date, bounded-unknown hedging is penalised, so uncertainty scores only where the record actually leaves it.

A second domain — a feature-flag incident history — confirms the reducer operates on generic subjects rather than fixture literals. Throughout, precedence is never read as proof of contribution.

What did time actually buy us?

The pre-registered verdict resolves as Type A on fixture evidence. The ordered and temporal conditions solve the permutation, correction, late-arrival, and effective-time cases that the bag and recency approaches fail, while remaining invariant on irrelevant swaps.

The scope is narrow and worth stating: a deterministic ledger, nine to twelve events per case, no language-model reader, no real-corpus transfer.

The other verdict types resolve as follows.

Type B (timestamps suffice) is rejected for historical and standpoint queries. It is explicitly sustained for clean current-state queries — recency is the right tool where change is monotonic and honestly recorded.

Type C (bitemporal machinery only for specialist cases) describes part of the result fairly. Late knowledge and future-effective scheduling are specialist shapes, and the architecture selects temporal machinery by failure class rather than applying it everywhere.

Type D (transition semantics too brittle) is not supported on these fixtures: the reducer’s generic rules cover two domains with no fixture-specific answer rules. Two domains are not enough to dismiss the brittleness risk, which is why the ontology stays small.

Type E — materialized state improves current-state performance without changing the semantics — is supported here. Its staleness is managed by derivation from the log plus rebuild equivalence, and the suffix-replay cost is measured above.

If a broader principle survives, it is this:

For some memory tasks, preserving events without preserving their temporal relationships is insufficient. The same event set can imply different state, different admissibility of explanations, and different expectations under different valid orderings.

Time, in state-changing histories, is part of the memory representation — not an annotation on it.

And a complementary discipline follows:

A memory is sometimes not an item but a transition. What the system must retain is how one state became another.

Both conclusions are bounded by the fixture scope of the permutation experiment.

What remains unfinished

The temporal layer hands Chapter 9 a richer state model than the one this chapter inherited. Question 5 — what did we leave unfinished — requires distinguishing intentions created from transitions completed: a planned migration with an effective date that arrived without its deployment is detectable only because the log records both the commitment and the absence of its expected transition. Once memory can reconstruct how state changed through time, the next problem is detecting expected transitions that never happened.

The chapter leaves agent- and task-relative standpoints, full vector-clock causality, the thirteen-relation interval network, write-path integration with the Chapter 4 graph, any new routing capability, and real-corpus validation of its empirical conclusions unresolved.

One analogy is worth naming and then setting aside. Current reasoning systems process trajectories — process supervision, multi-step tool loops, reflection traces. That shows capable systems increasingly reason over intermediate states. It establishes nothing about long-term memory architecture.

This chapter’s case had to be earned here, on the ledger, one permutation at a time.

Research foundations

Temporal question answering supports treating truth as time-indexed: TimeQA requires reasoning over facts holding at different times, SituatedQA makes temporal context part of correctness, StreamingQA evaluates adaptation to arriving knowledge, and LongMemEval tests updates across sessions. The lookup key in all four is effectively topic plus temporal standpoint, which this chapter makes explicit as valid and record time.

The deeper foundations are older, and each one contributes a specific piece:

Transport behaviour follows the ZeroMQ publish-subscribe specification, the ZeroMQ Guide, and the pyzmq documentation.

References

  • Leslie Lamport, Time, Clocks, and the Ordering of Events in a Distributed System (1978).
  • James F. Allen, Maintaining Knowledge about Temporal Intervals (1983).
  • Robert Kowalski and Marek Sergot, A Logic-Based Calculus of Events (1986).
  • Richard T. Snodgrass and Ilsoo Ahn, A Taxonomy of Time in Databases (1985).
  • ZeroMQ RFC 29: Publish-Subscribe; ZeroMQ Guide; pyzmq documentation (docs).
  • TimeQA: A Dataset for Answering Time-Sensitive Questions (2021).
  • SituatedQA: Incorporating Extra-Linguistic Contexts into QA (2021).
  • StreamingQA: A Benchmark for Adaptation to New Knowledge over Time (2022).
  • LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory (2024).
  • Bhuwan Dhingra et al., Time-Aware Language Models as Temporal Knowledge Bases (2022).