Pathways Through Memory
Explore how activation can propagate through a memory graph so that one remembered thing leads to another, and measure whether associative retrieval improves on static graph search and conventional RAG.
Chapter 4 set out to give the system a map. Instead of reconstructing meaning from raw passages at every query, it keeps interpretation: entities, relationships, claims, communities, each traceable back to the artifact it came from. Whether that map pays for itself is Chapter 4’s own question and Chapter 4’s own experiment; nothing here assumes its verdict.
This chapter asks a question that arises either way. Suppose the map exists. A map is a structure to search, but associative recall suggests a different mechanism: one memory evokes another, which evokes another, until useful evidence appears several steps from the original cue.
The question is whether that difference is mechanically real, and whether it buys anything:
Once memories have persistent representations and relationships, how should activation move through them so that one memory can evoke another?
The working name for this is associative memory. The chapter’s job is to build it, to compare it fairly against the systems the book already has, and to report what happens even if the answer is that it was not worth the trouble.
The pipeline is a small set of stages with one moving dial — how far activation propagates:
flowchart LR
C[Query cue] --> S[Seed memories]
S --> P[Bounded propagation]
P --> N[Activated neighbourhood]
N --> E[Selected evidence]
F1["too few seeds"]:::fail -.-> S
F2["unbounded spread"]:::fail -.-> P
F3["drifts from relevance"]:::fail -.-> N
classDef fail fill:#ffffff,stroke:#c0392b,color:#c0392b,stroke-dasharray:3 3
Each stage has its own way of going wrong, marked in red. Nothing downstream can repair a stage that already failed.
A graph is not yet remembering
Here is the shape of the book’s progression so far, and the step this chapter proposes.
Chapter 3
raw history → retrieval → context → answer
Chapter 4
raw history → persistent derived graph → graph-aware retrieval → answer
Chapter 5
query/cue → seed activation → propagate through the graph
→ retrieve associated memories → answer
Chapters 3 and 4 both treat retrieval as a lookup: given a query, find the items most like it. In Chapter 3 the items are passages and the likeness is embedding distance. In Chapter 4 the items are graph elements and the likeness may include a neighbourhood, but the neighbourhood is still selected by a query-to-item comparison. In both, the relationship structure is something the retriever consults.
The alternative is that the structure is something recall travels through. Activation enters the graph at whatever the cue touches, and then moves — attenuating with distance, competing with other branches, stopping when it runs out. What comes back is not the set of nodes nearest the query but the region the cue lit up.
That is a real mechanical difference and it should have measurable consequences. The rest of the chapter is about finding out which ones.
Memories lead to memories
Someone thinks cat.
What follows is not a ranked list of cats. It is a particular cat, then an incident involving that cat, then the house where the incident happened, then a person associated with the house, then something that person once said. By the fourth step the memory has almost nothing to do with cats. It is not similar to the cue. It is reachable from the cue, along a chain in which each link was, at the time, the obvious next thing.
A project history has the same shape. The question
What was behind the budget problem in the migration?
may need an artifact that never mentions the migration, never mentions a budget, and shares no vocabulary with the question — because the route runs through a person, to a project, to a supplier, to an invoice, and the invoice is the answer.
Two things are worth extracting from the intuition before it gets over-used.
The first is that a useful memory can have weak direct similarity to the cue and strong reachability from it. If that is common in real histories, a retriever that scores items only by direct similarity has a systematic blind spot: the missing signal lies in the path connecting memories, not only in the representation of each item.
The second is that the path depends on the cue, not only on the entity.
Take one person in this book’s history. a.silva appears in the event-store migration, in the removal of a legacy import domain, and in a backfill failure that left orphaned rows. Three unrelated regions of the project, one name.
So a query about a.silva and the backend should not light up the same neighbourhood as a query about a.silva and foreign-key checks.
The classical statement is Tulving and Thomson’s encoding specificity principle: what can be retrieved depends jointly on what was stored and on the cue present at retrieval — not on the stored trace alone.
Recast as engineering: the relevant neighbourhood is conditioned by the active cue. Entity identity is not enough to fix it.
Both of these are testable. Neither is established by finding the intuition appealing.
From similarity to pathways
The mechanism has a long history, and the honest version of that history is a warning as much as an invitation.
Collins and Loftus described semantic memory as a network of concepts joined by links of varying strength, in which processing a concept makes it a source of activation that spreads outward in parallel, weakening with distance and with time. Where activation from two sources intersects, relatedness is detected. The theory was explicit that activation must attenuate: without decay the whole network lights up and the signal is gone.
Information retrieval tried this, and its verdict is the useful part.
Crestani’s survey reviews two decades of spreading-activation models over semantic networks, and asks critically whether the technique ever earned its place.
The vocabulary the field settled on is itself the finding. The usable variant is constrained spreading activation — hedged with limits on how far activation may travel, how many neighbours it may reach, which paths are eligible, and how weak it may get before it stops.
The historical lesson is not that spreading activation is useless. It is that unconstrained propagation can retrieve broadly without preserving relevance.
That matters here because Chapter 3’s development work already leaves source precision as a residual weakness. Associative expansion is, structurally, a way of retrieving more. A mechanism that recovers useful indirect memories while tripling the irrelevant ones has not repaired that weakness; it has moved the failure somewhere less visible.
So the chapter’s prediction, registered before the runs: the constraints will matter more than the propagation rule.
Two independent pieces of prior work point the same way.
The MemORAI authors — a 2026 preprint, so author-reported and unrefereed — measure their query-conditioned edge weighting at about two points of turn-level recall@10 and their subgraph scoping at about thirteen, crediting topic segmentation with most of the system’s overall performance.
The SYNAPSE authors, in a refereed paper, report that removing lateral inhibition cost about one point of average F1. Removing decay collapsed temporal reasoning. Removing the fan term collapsed open-domain performance.
In both systems the elaborate mechanism was the cheap contributor, and the blunt constraint was the expensive one.
The graph becomes active
The graph this chapter needs is smaller than the one Chapter 4 produces. It needs nodes with descriptions, edges with relation labels, and source provenance on both:
nodes entities, claims, and the source artifacts themselves
edges a relation label, a weight, and the artifacts that evidence it
provenance every node and edge traceable to raw history
Everything else — communities, community reports, claim typing — is Chapter 4’s business, and Chapter 5 consumes the abstraction rather than the package behind it. The adapter takes a structured-memory snapshot and produces a memory graph; if the backend changes, the adapter changes and nothing above it does.
Onto that graph the chapter adds one thing that Chapter 4 does not have: a per-query quantity.
activate the seeded nodes
↓
send activation over eligible edges
↓
activate neighbours to different degrees
↓
propagate again
↓
decay, suppress weak paths
↓
select a bounded memory subgraph
Activation is never written to the graph. It belongs to this act of remembering and disappears when the query does. The graph records persistent interpretations derived from the project’s history; activation records what this cue reached. Keeping those in separate objects is not fastidiousness — it is what makes the same graph answer two different questions differently, and it is what stops one query’s traversal from becoming a fact about the corpus.
The cue has to get in somewhere
Before anything propagates, the cue has to become a set of seeded nodes, and a bad seed can misdirect every mechanism downstream. This is easy to leave implicit and expensive to get wrong: the HippoRAG 2 authors report that changing seeding alone — from named-entity matching to matching the whole query against extracted triples — moved multi-hop recall by over twelve per cent, which is larger than most differences between propagation methods.
Four seeders are implemented, measured separately from propagation for exactly that reason (interfaces are in the repository). Each answers a different suspicion:
- Lexical. Weighted term overlap. This handles the identifiers a project history is full of —
adr-007,incident-026— better than any embedding. - Embedding. Compares a cue vector against node-description vectors, reusing the Chapter 3 providers so embedding identity means the same thing in both chapters.
- Hybrid. Blends the two.
- Oracle. Reads the evaluator’s ledger directly.
Oracle is not a deployable system. It is a diagnostic. When oracle seeding fixes a case that hybrid seeding fails, the fault was at the entrance to the graph — not on the journey through it.
Two ways to move
Two propagation mechanisms dominate the modern literature, and they are different enough that implementing only one would have decided the question by construction.
Personalized PageRank leaves the graph alone and changes where the random walk restarts. Jeh and Widom’s formulation biases the stationary distribution toward a preference vector; Haveliwala’s topic-sensitive variant had already shown that a global importance measure becomes query-relative by changing the reset distribution rather than the graph. The reset vector is exactly the seed-activation interface this chapter already needs, which is why HippoRAG and HippoRAG 2 both use it: seed from the query, run one pass, read off importance.
Its limitation is specific and worth naming. The Personalized PageRank score itself does not encode hop count. It can report that a memory is reachable and important, but the number of graph steps between the seeds and that memory has to be computed separately rather than read from the stationary score.
Spreading activation keeps the hop structure. Each step, every active node retains a fraction of its own activation and receives what its neighbours send:
$$ a_i(t+1) \;=\; \underbrace{r \, a_i(t)}_{\text{retained}} \;+\; \underbrace{\sum_{j} \frac{s \, w_{ji} \, a_j(t)}{\operatorname{fan}(j)}}_{\text{received from neighbours}} $$Here $r$ is the fraction of its own activation a node keeps, $s$ the fraction of what is transmitted, $w_{ji}$ the weight of the edge from $j$ to $i$, and $\operatorname{fan}(j)$ the number of edges leaving $j$ — the term that stops a hub from flooding the graph.
The implementation follows the form SYNAPSE reports, with their published settings as the starting point rather than tuned values: half the activation retained, four-fifths of what is transmitted passed along, and division by the sender’s fan (loop details are in the repository).
Termination does not depend on remembering where the walk has been. Activation decays multiplicatively and the threshold cuts it off, so a cycle dies out instead of circulating. The hop and node budgets exist as well, and the tests check that each of the three can be the binding constraint, because a mechanism whose only safety property is a hop limit is a mechanism that has not thought about cycles.
A third strategy is included as the control: bounded-hop neighbourhood expansion, which is roughly what a static graph memory already does. If that matches spreading activation, propagation has bought nothing.
Context changes the path
The fourth strategy makes propagation itself query-dependent. In plain propagation, activation leaves a node the same way whatever the cue was. Cue-conditioned propagation recomputes the edge weight per query from the agreement between the cue and the edge’s own description, so two questions about the same person can travel differently. This is the mechanism MemORAI calls dynamic weighted PageRank.
One design decision inside it matters more than it looks. Conditioning attenuates edges. It never removes them.
A cue-irrelevant edge is scaled toward a floor, not severed. Severing would make the graph itself query-dependent, and the whole architecture rests on a separation: the graph asserts what the history contains, retrieval expresses what this question prefers. Those two must not become the same object.
The blending is implemented so that at zero conditioning strength it reduces exactly to unconditioned propagation. That is what makes the ablation a one-line change rather than a second implementation.
Association is not truth
This is the chapter’s most important distinction and it costs nothing to get right at the start, so it is worth stating before any results.
There are at least two entirely different things an edge can be carrying. The first is a claim about the world, grounded in evidence:
adr-007 SUPERSEDES adr-003
evt-205 SUPPORTED_BY evt-203
event store CURRENT_BACKEND PostgreSQL
The second is how readily one memory should evoke another during recall:
event store → evt-205 0.75
evt-205 → evt-203 0.80
evt-203 → session-019 0.70
These are not the same quantity and they must not be summed into one weight. A route can be heavily travelled and lead to something false; a proposition can be certainly true and sit on a route nobody needs. The principle the rest of the book inherits is:
Association strength is retrieval priority, not epistemic confidence.
It follows that a frequently traversed pathway can be useful without making the underlying proposition any more true, and that nothing may raise a claim’s confidence because a route to it was popular. The implementation keeps four quantities in four places and never collapses them:
| quantity | means | lives on |
|---|---|---|
relation | what the edge asserts about the world | the edge |
association | retrieval priority | the edge |
evidence_confidence | epistemic support for the proposition | the edge |
activation | how strongly this cue lit the node | the per-query state |
Query-relative relevance, source authority, and recency are further distinct quantities. They enter later chapters; the point here is that the single-weight temptation is available at every step and is refused at every step.
Follow the pathway
If memory is going to reach an artifact through four hops, the system has to be able to say how. The trace is not decoration; it is the difference between a retrieval you can debug and one you can only accept.
cue: What still had to move before the first release after the event-store migration?
seeds: source:release-024(0.49), source:session-051(0.40),
source:migration-run-103(0.32), claim:intent-401(0.26)
intent-401 backup obligation --RAISED--> commit-112 (w=0.60, delivered=0.045)
intent-401 backup obligation --ABOUT--> backup configuration (w=0.70, delivered=0.053)
intent-401 backup obligation --COMPLETED_BY--> commit-118 (w=0.70, delivered=0.053)
A trace answers one question: why did this memory reach my context?
That is not the same as why should I believe what it says? The book keeps them apart under a name: retrieval causality.
The path is evidence about how retrieval behaved. It is not evidence for the claim at the end of it.
Chapter 7 builds the structure that answers the second question. Conflating the two here would let a well-travelled route masquerade as support for whatever it happens to lead to — which is precisely the error this chapter’s later sections are about.
The trace renders with that disclaimer attached, because a path is persuasive-looking in a way that invites the confusion.
Implementation
The package sits above Chapter 4’s abstraction and beside Chapter 3’s baseline rather than replacing either. Its layout — configuration, pipeline, deterministic fixture graph and cue set, seeders, propagation strategies, activation state, pathway traces with versioned weights, the instrument bridge, and health checks — is described in the repository.
ORIGINAL EVIDENCE
↓
Chapter 3 RAG substrate
↓
Chapter 4 persistent graph
↓
Chapter 5 associative activation
↓
selected memory subgraph
↓
context
↓
reader / action
Fallback runs the other way. When associative retrieval is uncertain, static graph retrieval remains available; when that is uncertain, strong RAG remains available; underneath all of it the raw source artifacts remain canonical and reachable. No layer here is permitted to become the only way to reach history.
Seeding, retrieval over the graph, and reading admitted memories with their paths are three calls (see the repository).
Before any of this touches an LLM-extracted graph, it runs on a fixture small enough to reason about by hand: 73 nodes, 101 edges, built directly from the project’s canonical running examples. Every node corresponds to an artifact the rest of the book already uses, and no identifier is invented.
The fixture is deliberately unfriendly. It contains:
- A dominant hub. The event-store decision touches thirteen neighbours, against an average degree of 2.8.
- Derived echoes.
runbook-006restatesadr-007without independently supporting it;staging-log-018does the same toincident-026. - A same-symptom, different-cause trap.
incident-106is a timeout caused by network saturation, sitting right next to the nine-minute schema outage. - A superseded procedure, plus weakly connected truths of degree one or two.
- Three actors who each appear in three unrelated regions — and cycles.
Twenty cues probe it, labelled by what they are testing: cases where direct similarity should already succeed, cases where the evidence is two or more edges away, cases where the same entity appears under different active contexts, adversarial cases built around each trap, and one unanswerable case where abstention is correct. A deterministic demonstration runs with no model server, no database, and no network (see the repository).
Measure it
Chapter 2 owns scoring, and nothing here reimplements a scorer. An associative retrieval is converted into the instrument’s SystemOutput and handed to score_task exactly as the Chapter 3 baseline is. What the chapter adds is a small set of path-quality observations that only make sense for a system that traverses, kept alongside the instrument’s metrics rather than blended into them:
Path recall — did any activated path reach the required evidence, whether or not it survived selection? Path precision — how much of what propagation touched was relevant? Expansion factor — memories activated per memory admitted. Useful-hop distance — how far the evidence was. Activation concentration — whether recall settled on a coherent region or smeared. Hub capture and distractor rate — whether the loudest or the labelled-wrong memories won.
The most useful of these is the gap between path recall and source recall, because it helps localise which layer failed. High path recall with low source recall means propagation reached the evidence but selection discarded it. Low path recall means the traversal itself failed to reach the required evidence. Without the distinction, both read as “the system missed it”.
One constraint is non-negotiable in the comparison. Associative search may explore broadly inside the graph, but what it hands to a reader stays bounded — at most eight memories under every condition, with the exploration cost reported separately. A mechanism cannot win this comparison by injecting more history.
What the fixture runs show
The suite is frozen as run ch5-20260919-e5. Its scope has to be stated plainly before its numbers are read: these are retrieval-level results on a synthetic graph. No reader runs, no answers are generated, and nothing here is a comparison of Chapter 3 against Chapter 4 against Chapter 5 on task accuracy. Chapter 12 later compares strong RAG with the selected, assembled pipeline, but it does not isolate associative propagation; the direct three-way comparison remains unrun.
Twenty cues, matched context budget, mean over cases:
| condition | source recall | source precision | path recall | expansion | context tokens | edges walked |
|---|---|---|---|---|---|---|
| seed-only (no propagation) | 0.850 | 0.463 | 0.850 | 1.00 | 25.8 | 0 |
| direct neighbourhood | 0.892 | 0.283 | 1.000 | 5.92 | 43.5 | 2,448 |
| Personalized PageRank | 0.850 | 0.268 | 0.967 | 1.49 | 44.0 | 4,040 |
| spreading activation | 0.875 | 0.350 | 0.900 | 1.08 | 41.0 | 2,035 |
| cue-conditioned | 0.900 | 0.446 | 0.900 | 1.00 | 31.4 | 1,361 |
| shuffled-edge control | 0.675 | 0.219 | 0.800 | 1.27 | 41.4 | 2,142 |
Four things in that table are worth more than the headline.
The edges carry information. The shuffled-edge control rewires the graph while preserving its degree distribution — a control borrowed from the temporal-shuffle test in Dury’s 2026 preprint on associative memory. Recall falls from 0.875 to 0.675, below even the no-propagation floor. On this fixture, propagation over the shuffled structure is worse than not propagating at all. That supports the narrower conclusion needed here: the relationships carry useful retrieval signal beyond merely touching more nodes.
Propagation’s gain over seeding alone is real but small. Cue-conditioned propagation recovers five points of source recall over the floor. That is the entire headline effect, and it costs about six context tokens per cue.
Precision behaves exactly as the historical warning predicted — except under conditioning. Unconditioned spreading trades eleven points of precision for two and a half points of recall. Plain PageRank loses twenty points of precision and gains nothing. Only cue conditioning gains recall while roughly holding precision, at 0.446 against the floor’s 0.463. Conditioning is the mechanism that prevents the recall gain from being accompanied by the much larger precision losses seen in the other propagation strategies.
Direct neighbourhood expansion is the clearest illustration of the failure mode. It has perfect path recall — one-hop expansion reaches everything — and path precision of 0.055, with an expansion factor of 5.9. It finds all the evidence and buries it.
On the eleven cases where evidence sits two or more edges away, cue-conditioned propagation recovers nine points of indirect recall over seeding alone, at a three-point precision cost.
And it leaves the direct cases exactly where it found them — identical recall and precision to no propagation at all.
That second fact matters because a multi-hop capability that degrades the cases direct similarity already solves would not be a clean improvement. It would be a trade.
Unconditioned spreading and PageRank recover nothing on the indirect cases. Whatever multi-hop capability this fixture has, conditioning is where it comes from.
Why activation must decay
Ablating one constraint at a time gives the clearest result in the chapter (full figures are in the frozen run record).
Removing the activation threshold multiplies the nodes touched by 5.7 and collapses path precision from 0.348 to 0.042 without recovering a single additional piece of evidence. Removing fan division does the same thing less dramatically. Removing decay entirely — no retention, full transmission — costs over a third of recall, because activation that never attenuates never concentrates anywhere, and the selection step has nothing to rank by.
The unconstrained condition is the Crestani warning reproduced on a 73-node graph: 11,509 edge traversals, 8.9 memories activated per memory admitted, path precision of 0.032, and worse recall than the constrained version. Removing every guard does not retrieve more of what was wanted. It retrieves more of everything, and the ranking drowns.
There is also a counter-intuitive result, worth reporting rather than smoothing away.
Deeper propagation improves precision. At five hops instead of three, precision rises from 0.350 to 0.479, context tokens fall from 41.0 to 33.5, and recall is unchanged.
The reason is visible in the traces. Decay applies at every step, so nodes that stop receiving activation fall below the threshold and drop out. Deep propagation concentrates rather than accumulating. The surviving active set at four hops is smaller than at one.
This is the opposite of the usual intuition that more hops means more noise. It is also a property of this fixture at this scale. Whether it survives on a graph two orders of magnitude larger is not something these runs can say.
Why memories compete
Lateral inhibition is the mechanism the chapter expected to matter most and the one the evidence supports least.
Suppressing all but the strongest seven nodes on each frontier changed precision from 0.352 to 0.350 — effectively unchanged at this fixture scale, and in the wrong direction. It changed nothing on the adversarial subset either: hub capture and distractor rate were identical with inhibition on and off. The same asymmetry appears in the SYNAPSE authors’ own ablations, where removing lateral inhibition cost about one point of average F1 while removing decay cost tens.
The honest reading is that on a graph of this size, the threshold and the fan term have already done the work inhibition was supposed to do. Competition between pathways is real — the fixture is built so that a.silva has four branches competing for one budget — but the competition is resolved by attenuation before any explicit suppression step gets to act.
Fan division, by contrast, earns its place, and the way it earns it is a genuine trade. On the adversarial cases, removing it raises recall to 1.000 and drops precision from 0.517 to 0.410. Dividing a hub’s output by its degree is not free: it makes the loud node quieter, which occasionally silences something that was loud and right.
That trade has a name in the literature. The SYNAPSE authors report a failure they call cognitive tunnelling, where inhibition suppresses a minor but correct detail in favour of a hub. The measurement that would catch it here is hub capture paired with source recall, and it is reported for every condition rather than assumed away.
The cue chooses the branch — but not where expected
The context-conditioning experiment gives a result that complicates the chapter’s own story.
Three cues about a.silva — the event-store proposal, the corpus_import caller migration, the foreign-key failure — each recovered their own branch completely. Recall was 1.000 on every branch cue under every strategy. The cue, not the entity, chose the region. That part of the intuition holds.
But the branch was chosen at seeding, not during propagation. The overlap between the admitted source sets for different cues about the same actor was 0.234 for plain spreading, 0.333 for cue-conditioned propagation, and 0.405 for direct expansion. Conditioning did not separate the branches better than plain propagation did; it separated them slightly worse, because attenuating cue-irrelevant edges keeps activation nearer the seeds and the seeds were already cue-specific.
This lines up with what MemORAI’s ablations reported: their query-focused subgraph scoping was worth roughly seven times what their query-conditioned edge weighting was worth. The finding this chapter can add is a sharper version of the same thing. Where the cue enters the graph matters more than how the cue steers the walk. Conditioning’s contribution shows up in precision on multi-hop cases, not in branch separation — and the intuition about entity-plus-context was right about the phenomenon and wrong about the mechanism that produces it.
When Chapter 4 gets it wrong
Derived structure is fallible structure, and a false relation is not a hypothetical. The false-edge experiment injects the most damaging plausible one: a fabricated SUPPORTED_BY edge wiring the event-store decision hub directly to incident-106, the timeout that looks like the schema outage but was caused by network saturation. It is exactly the mistake an extractor makes when two passages share vocabulary.
In aggregate the effect is small: path precision moves from 0.348 to 0.335. On the cue it was built to derail, it is concrete. Asked what caused the nine-minute outage during a migration, the clean graph admits incident-026, incident-106, and procedure-123. The corrupted graph admits those and adr-007 — the event-store decision, pulled into an answer about an outage it had nothing to do with, by one edge that no source evidences.
Two observations follow. A well-constrained propagation absorbs a single false edge without catastrophe, which is mildly reassuring. And the graph health check flags the edge before any query runs, because it has no source provenance — which is more useful than the absorption, and is the reason the provenance requirement is structural rather than advisory.
Can pathways learn?
Everything above is static. The graph does not change when it is used, and the first result had to be interpretable without learned weights, which is why it is.
The next question is whether experience using memory should change future traversal. A route that repeatedly led somewhere useful might reasonably become easier to travel; a route that led nowhere might fade.
cue → pathway → useful result → strengthen the route
cue → pathway → irrelevant → weaken the route
route unused for a long period → decay toward its base
This is staged deliberately. Stage A asks whether propagation itself helps. Only if it does is Stage B — adaptive pathways — worth pursuing, and Stage B stays behind an explicit flag that is off by default.
The implementation makes one structural commitment before it makes any behavioural one. Adaptive weights are an append-only overlay on an immutable base:
association state = base graph + ordered update log
Nothing mutates the base. Replaying the log reproduces the current weights, and truncating the log rolls back to any earlier point.
So the question why is this pathway strong? has a literal answer: the list of updates that made it so, each with its reason, its cue, and the outcome it came from. Both properties are asserted in the run record, not claimed in prose.
The alternative would be silently mutating weights on a graph that is nominally rebuildable. That leaves the derived state neither rebuildable nor auditable — a worse position than not adapting at all.
On the fixture, three related cues were run in sequence, each judged against the ledger, and the resulting weights re-evaluated on the whole cue set. Ten updates were recorded: nine weakenings and one strengthening. Recall moved from 0.875 to 0.900 and precision from 0.350 to 0.368.
That is a small measured improvement on this fixture, produced by almost entirely negative feedback. The single strengthening is worth noticing for a reason that is not about its size: the routes that reached correct evidence were mostly the seeded ones, and a seeded memory has no pathway to credit. There was less to reinforce than the mechanism assumed.
The danger of reinforcing mistakes
The reason Stage B is flagged rather than default is a failure mode that deserves the space.
Suppose a wrong association exists. If retrieval frequency strengthened pathways, then the wrong association would be retrieved, and being retrieved would strengthen it, and being stronger would cause it to be retrieved more:
incorrect association
→ retrieved more often
→ appears useful because it is retrieved
→ strengthened further
The broader failure mode has empirical support.
The ACL 2026 study of experience-following behaviour shows that agents can produce highly similar outputs when a retrieved record’s input resembles the current one. That provides evidence for the risk that inaccuracies in stored experience can propagate into later behaviour — and that even previously correct executions can sometimes mislead.
A 2026 preprint, MemEvoBench, calls the aggregate phenomenon memory misevolution: behavioural drift from repeated exposure to misleading information. On its authors’ own unrefereed measurements, it reports substantial safety degradation under biased memory updates, with prompt-level defences insufficient.
The experiment makes the loop visible rather than claiming to have solved it. The same corrupted graph, the same cue repeated eight times, two update policies:
cue repeated 8 times: What caused the nine-minute outage during a migration?
the graph carries one fabricated SUPPORTED_BY edge, weight 0.60
frequency-only 0.65 → 0.70 → 0.75 → 0.80 → 0.85 → 0.90 → 0.95 → 0.95
health flags: unprovenanced-edges, frequency-reinforcement
outcome-gated 0.55 → 0.50 → 0.45 → 0.40 → 0.35 → 0.30 → 0.30 → 0.30
Frequency-driven reinforcement drives a fabricated edge to the ceiling in six rounds. Outcome-gated reinforcement drives it down. The safeguards that produce the difference are four, and each is a refusal rather than a heuristic:
- Outcomes, not frequency. Only a recorded task outcome moves a weight. Being retrieved moves nothing.
- Evidence gates. An edge with no source provenance is never strengthened, however well the task went; nor is one whose evidence confidence is below a floor. The fabricated edge fails the first gate and the derived echoes fail the second.
- Bounds and decay. Weights are clamped and unused routes drift back toward their base.
- Append-only records. Every change is logged with its reason and can be replayed or rolled back.
And then the result that matters most, which is not the one the safeguards were designed for. Both policies admitted the wrong memory in all eight rounds. The gates stop the weight from running away. They do not remove the false edge, and they do not stop it from putting adr-007 into an answer about an outage. The safeguard addresses the feedback loop, not the fault.
That distinction is the one to carry forward. A memory system that cannot reinforce its mistakes is not thereby a memory system without mistakes.
Co-occurrence is not usefulness
One tempting shortcut is ruled out explicitly. Two memories appearing together often is not a reason to strengthen the route between them.
Frequent co-occurrence can mean duplication, boilerplate, one source quoting another, a popular distractor, or a misconception restated. The fixture contains two instances on purpose. runbook-006 restates adr-007 for operators; staging-log-018 restates incident-026. Both co-occur constantly with what they echo. Neither is independent corroboration of anything, and the ledger records exactly that — their derivation edges point at what they echo.
The evidence-confidence gate is what encodes the refusal: the echo edges carry deliberately low evidence confidence, and the learner refuses to strengthen them even when a route through them produced a good outcome. The test for that refusal is explicit, because it is the kind of property that decays silently.
The general form is that three different relationships hide behind the same statistic:
co-occurrence two memories appear together
useful transition one memory usefully leads to the other
independent support two memories independently evidence a claim
The book’s position is that repetition is not truth, and that a system which strengthens on repetition has confused the first of these with the third.
Did associative memory earn its place?
Four outcomes were registered in advance. The result maps to the first, weakly, with a large piece of the third.
Type A — associative retrieval clearly helps. Partially supported, at a modest size. Cue-conditioned propagation improved source recall from 0.850 to 0.900 overall — roughly holding precision — at about six extra context tokens per cue, and recovered nine points of indirect recall (0.773 to 0.864) at a three-point precision cost, without degrading the cases that direct similarity already solved. The shuffled-edge control supports the conclusion that graph structure contributes to the gain rather than expansion alone.
Type B — static graph and RAG are already enough. Not supported, but close enough to matter. The no-propagation floor scores 0.850 recall at the best precision of any condition. Everything this chapter builds is competing for five points of recall.
Type C — recall improves, precision worsens. Supported for every mechanism except cue conditioning. Plain spreading activation, Personalized PageRank, and direct neighbourhood expansion all degrade precision, two of them without recovering any indirect recall. The residual precision weakness the Chapter 3 runs already show is made worse by naive propagation, exactly as the historical record predicted.
Type D — adaptive pathways become unstable. Demonstrated under the refused policy and contained under the proposed one. Frequency-driven reinforcement produced a runaway weight on a fabricated edge in six rounds. Outcome-gating with evidence gates reversed it. Neither prevented the wrong memory from being admitted.
So the chapter earns a conditional, narrow result at the retrieval layer.
On this evidence, the retained candidate is cue-conditioned propagation with an activation threshold, fan division, and a bounded hop count. Where associative retrieval is used, it sits above the static graph and below the context assembler, with static graph retrieval and strong RAG still available beneath it. The chapter does not establish that every remembering system needs this layer.
Lateral inhibition does not earn its place on this evidence. It is retained only as an ablatable option.
Adaptive pathways do not enter the architecture at all. The mechanism exists behind a flag, so that Chapter 16 can take up the question with the failure mode already visible.
The fair summary of the numbers is that the constraints mattered more than the propagation rule, which is what the prior work predicted and what this chapter now has its own evidence for.
An analogy, used carefully
The vocabulary in this chapter is borrowed: spreading activation, decay, inhibition, the fan effect, reinforcement. The borrowing is deliberate and it has limits worth stating rather than leaving to inference.
The mechanisms here are inspired by semantic networks, by cue-dependent recall, by hippocampal indexing ideas, and by reinforcement analogies. Each borrowing supplied a hypothesis that turned out to be implementable and testable: decay came from Collins and Loftus and proved essential on this fixture; division by fan came from Anderson and Reder’s account of the fan effect and exposed a real precision/recall trade; cue conditioning came from encoding specificity and was the only tested propagation mechanism that gained recall while keeping precision close to the no-propagation floor.
What is not claimed: that any of this is how human memory works, or that the graph is a neural network.
Neural networks generally learn distributed weights through optimisation over a loss. This system has explicit memory nodes, interpretable links, and weights that change only through logged, gated, replayable updates.
The Hebbian-sounding idea — that repeatedly useful co-activation strengthens future access — is a statement about routes through an inspectable structure. It is not a claim of biological equivalence.
Where a borrowed term corresponds to a real algorithmic mechanism, it is used. Where it would only be decoration, it is avoided.
The analogy is a source of hypotheses. The evidence is what decides, and in this chapter it decided against one of the borrowed mechanisms.
What prior work contributes
HippoRAG established the design this chapter’s PageRank strategy follows: an open knowledge graph over passages, seeded from the query, ranked by Personalized PageRank in a single pass rather than an iterative retrieve-and-read loop.
Its successor, HippoRAG 2, supplied two things adopted directly here. Passage nodes bound to concept nodes, so a path can always terminate on evidence rather than on derived interpretation. And the finding that seeding method alone moves multi-hop recall substantially.
Their honest note matters too: performance on complex associative tasks degrades with corpus growth at a rate similar to plain retrieval. That is why this chapter measures expansion factor rather than recall alone.
A-MEM is the most useful counter-example in the literature.
It builds an associative note graph with LLM-generated links — and then retrieves by embedding similarity over the notes. The links change what notes say, through a memory-evolution step. They are never traversed.
That is a genuine alternative hypothesis: the value of associations lies in re-encoding, not in propagation. It is why the no-propagation floor is a first-class condition in this chapter rather than a formality.
Its evolution mechanism, which rewrites the contents of stored notes, is not adopted here. This book instead keeps canonical sources unchanged and treats evolving interpretations as derived state that should remain traceable and rebuildable.
SYNAPSE supplied the mechanism inventory: dual-trigger seeding, the fan term, temporal decay, lateral inhibition, a firing threshold, and a hybrid fusion of similarity, activation, and structural importance. It also supplied the ablation design and, in its own numbers, the first hint that inhibition would underperform its billing. Its reported cognitive-tunnelling limitation is what the hub-capture metric exists to detect.
MemORAI supplied cue-conditioned edge weighting and turn-level provenance stored on the edge. Its ablations are the most useful table in the recent literature for this chapter’s purposes, because they measure the propagation refinement against the structural decisions and find the refinement much the smaller term.
MRAgent proposes a third position — a deliberately shallow cue-tag-content structure in which an LLM explores and prunes paths during access rather than arithmetic propagating through them. Think-on-Graph occupies the far end of the same axis, with the model running beam search over the graph. Both are more selective and far more expensive per query than anything implemented here; the chapter tests whether cheap arithmetic propagation suffices before reaching for them.
From the classical side, four contributions shaped this chapter:
- Collins and Loftus supplied the mechanism, and the warning that it must attenuate.
- Tulving and Thomson supplied the argument that retrieval is a function of trace and cue jointly.
- Anderson and Reder supplied the fan effect.
- Crestani’s survey supplied the historical record that unconstrained spreading activation is hard to control.
A fifth is deliberately held back. Anderson and Schooler’s rational analysis — that the probability a memory will be needed tracks frequency and recency in the environment — is the strongest classical argument for need-driven access priority. It is left to Chapter 16, because using retrieval frequency as a proxy for need is exactly the error this chapter’s reinforcement section is about.
The experience-following study and MemEvoBench supply the modern evidence that the reinforcement failure is real rather than theoretical, and the first of them supplies the constructive half as well: future task evaluations can serve as quality labels for stored memory, which is the argument for outcome-gating rather than frequency-gating.
The full matrix, with what was adopted and what was refused for each system, is kept in the project’s research notes rather than reproduced here.
What remains unsolved
The residuals from this chapter are specific enough to hand on.
- The direct comparison has not been run. These are retrieval-level results on a synthetic 73-node graph. Chapter 12 later supplies a behavioural comparison between strong RAG and the selected, assembled pipeline, but not a Chapter 3 versus Chapter 4 versus Chapter 5 ablation under a reader and matched budgets. An oracle-path condition belongs in that still-open comparison, to separate remaining traversal failures from reasoning failures.
- Scale, seeding, and selection bound every number here. The fixture is small and hand-built, so every property — including deeper propagation improving precision — is a property of this graph at this scale, and the adapter for a real Chapter 4 snapshot has not been run against a real index. Branch separation happened at seeding rather than during the walk, cue conditioning matches terms naively, and the selection side is untouched: a memory reached at two hops can still be ranked out by a hub, which is a ranking problem for the chapters that own relevance and assembly.
- Contained is not removed, and access is not truth. The safeguards stop reinforcement from amplifying an extraction error but neither detect nor remove it. Association decay — a preference about access — is carefully not the same thing as a claim ceasing to hold; validity belongs to Chapter 8, long-term availability to Chapter 15.
Activation can travel through a memory graph, and travelling through it recovers evidence that similarity alone does not reach — modestly, and only when the cue constrains the route.
But look at what the system now returns: a set of memories, each reached by a route. It still cannot say why any of them should be believed. A path is not a justification.
With several ways to reach memory now on the table, the next problem is choosing between them for the situation at hand. Beyond that choice lies the harder demand — the structure that answers why a claim deserves belief.
References
- A Spreading-Activation Theory of Semantic Processing — Collins and Loftus, Psychological Review 82(6), 1975.
- Encoding Specificity and Retrieval Processes in Episodic Memory — Tulving and Thomson, Psychological Review 80(5), 1973.
- The Fan Effect: New Results and New Theories — Anderson and Reder, Journal of Experimental Psychology: General 128(2), 1999.
- Reflections of the Environment in Memory — Anderson and Schooler, Psychological Science 2(6), 1991.
- Application of Spreading Activation Techniques in Information Retrieval — Crestani, Artificial Intelligence Review 11, 1997.
- Topic-Sensitive PageRank — Haveliwala, WWW 2002.
- Scaling Personalized Web Search — Jeh and Widom, WWW 2003.
- HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models — Gutiérrez et al., NeurIPS 2024.
- From RAG to Memory: Non-Parametric Continual Learning for Large Language Models — Gutiérrez et al., ICML 2025.
- A-MEM: Agentic Memory for LLM Agents — Xu et al., NeurIPS 2025.
- Synapse: Empowering LLM Agents with Episodic-Semantic Memory via Spreading Activation — Jiang et al., Findings of ACL 2026.
- How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior — Xiong et al., ACL 2026.
- Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph — Sun et al., ICLR 2024.
- MemORAI: Memory Organization and Retrieval via Adaptive Graph Intelligence for LLM Conversational Agents — Van et al., arXiv 2605.01386, 2026 (preprint).
- Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents — Ji, Li, and Hooi, arXiv 2606.06036, 2026 (preprint).
- MemEvoBench: Benchmarking Safety Risks from Memory Misevolution in LLM Agents — Xie et al., arXiv 2604.15774, 2026 (preprint).
- Zep: A Temporal Knowledge Graph Architecture for Agent Memory — Rasmussen et al., arXiv 2501.13956, 2025 (vendor-reported).
- Predictive Associative Memory: Retrieval Beyond Similarity Through Temporal Co-occurrence — Dury, arXiv 2602.11322, 2026 (preprint).