← Memory From First Principles

From Retrieval to Persistent Understanding

Strong RAG already understands retrieved evidence at query time. This chapter asks what changes when some of that understanding is preserved as reusable state, builds a persistent derived graph, and measures what it costs.

Chapter 3 built the strongest conventional retrieval the book could assemble, and left it standing. That result constrains this chapter before it starts.

The usual argument for structured memory goes: retrieval finds text, but it does not understand it. About the system Chapter 3 actually built, that is simply false. A hybrid retriever, a cross-encoder reranker, and a capable reader already do a great deal of interpreting.

Watch what that reader does. It tells proposals from decisions when the passages allow it. It follows rationale across artifacts. It abstains when the evidence runs out.

That is query-time interpretation. It is performed for one query, then thrown away immediately afterwards.

The question is therefore sharper than the usual argument for structured memory:

What do we gain by preserving AI-derived understanding as reusable state rather than reconstructing it from raw evidence independently on every query?

The answer has two halves.

The architectural half: a persistent derived graph, built with Microsoft GraphRAG — inspectable, versioned, rebuildable, with raw history still canonical underneath it.

The experimental half is less flattering. On this workload, strong conventional RAG stays extremely competitive. The graph produces no general quality win, and it costs a great deal. Its value turns out to be conditional, and its mistakes, unlike a reader’s, do not go away.

RAG already understands

The strawman version of this chapter says that retrieval matches words while graphs understand meaning. Nothing in the book’s own evidence supports that division.

Chapter 3’s reader takes up to six passages and reasons across them jointly. The reranker scores query and passage together before the reader ever sees them.

Watch it work. Asked what event store should new services use?, the baseline answers PostgreSQL and cites the decision record — from exactly the same raw passages any graph condition would see.

Interpretation happened. It happened at query time, inside one inference, and it left no trace.

That observation reframes the entire enterprise. The transition this book studies is not from dumb matching to smart comprehension. It is from transient interpretation to persistent interpretation:

RAG interprets history for this query. Persistent graph memory preserves some of that interpretation for future queries.

Everything the graph stores — an entity, a relationship, a claim, a community summary — is an assertion of one form: an AI system read these passages and concluded this.

That cuts two ways.

It amortises reading. The tenth question about the migration does not have to re-derive that PostgreSQL replaced SQLite, because a relationship already says so.

It also fossilises error. If the extraction was wrong, the tenth question inherits the mistake without re-reading the evidence that would have caught it.

Reusable intelligence and reusable mistakes are the same mechanism. The experiment measures both.

What disappears after the answer?

Consider what the baseline knows at two moments. During inference on why was PostgreSQL chosen?, its context contains the contention report, the benchmark, the incident, and the decision record, and its activations transiently encode the connection between them. One second later the answer is delivered and that connection is gone. The passages remain in the index. The understanding does not.

Three consequences follow.

Every query pays full price. Retrieval, reranking, and reading all start from raw text, no matter how many times that same connection was derived before.

Nothing accumulates. Twenty questions about the migration leave the system exactly as knowledgeable, in any stored sense, as zero questions did.

Consistency is accidental. Ask what we decided, why, who proposed the alternative, and what was rejected. Whether those four answers agree depends on the reader reconstructing the same distinctions four separate times.

A persistent derived memory is the hypothesis that storing some interpretation is worth its cost. Note what the hypothesis does not say. It does not say retrieval is weak — Chapter 3 forbids that reading. It does not say every interpretation should be stored — query-relative judgements belong to the query, as the three-layers section shows. It says that some interpretations are stable enough across queries to earn storage, and that the book should measure whether that storage pays.

Make understanding persistent

The architecture under test keeps two stores with different epistemic status:

    flowchart TD
    SRC[original sources] --> RAG[conventional RAG]
    SRC --> IDX[GraphRAG indexing]
    IDX --> DER[derived graph]
    RAG --> READER[reader]
    DER --> READER
    READER --> ANS[answer]
  

Raw sources are canonical. The derived graph is a hypothesis about history, not a rewrite of it.

So if the graph is wrong, the repair is to fix or rebuild the derived state. Never to edit the source. The source stays the ground the system falls back to and the auditor inspects.

That invariant — canonical history underneath, derived memory above — governs every later chapter. Association traverses the graph but cites the sources. Routing chooses graph modes but keeps raw retrieval available. Evidence lineage treats derivation as something to verify rather than trust.

Derived memory is a hypothesis about history, not a rewrite of history.

The comparison this chapter runs is therefore:

source history
    ↓
retrieval → context → reader interprets evidence for this query → answer

against:

source history
    ↓
AI interpretation during indexing
    ↓
persistent derived structure (entities, relationships, claims, communities)
    ↓
query over raw + derived state → answer

The comparison holds the corpus, reader model, and final context budget fixed. Its intended contrast is whether machine-derived interpretation persists between queries, although the graph pipeline also introduces its own indexing, chunking, and query machinery. Those implementation differences have to remain visible when the results are interpreted.

Raw evidence and derived state

Three layers need distinct names because they fail differently:

Layer 1 — raw source. What actually existed: a session transcript, a commit entry, a decision record, a benchmark note. Layer 1 is canonical historical evidence. It can itself be incomplete, mistaken, or misleading, but it is not the memory system’s derived interpretation of what happened.

Layer 2 — persistent derived interpretation. What the indexing system inferred: this span mentions PostgreSQL; PostgreSQL replaced SQLite here; these entities belong to one community; this claim is supported by those units. Layer 2 is reusable and versioned. It is also where extraction errors live permanently until rebuilt.

Layer 3 — query-relative use. What matters for the current question. The same artifact is decision evidence for what did we decide?, historical background for what runs in production now?, and irrelevant for the caching question. Layer 3 is computed per query and must not be stored as permanent truth, or one question’s framing becomes every later question’s prejudice.

One boundary matters enough to name, and Chapter 7 enforces it.

Layer 2 records derivation provenance: this graph object came from those text units, in those artifacts.

It does not record evidential support: whether those sources actually license the claim.

A relationship row reading PostgreSQL → SQLite is a stored assertion about what the extractor inferred. It is not a support edge. Graph connectivity must never be read as truth, and a source mapping must never be cited as justification.

This chapter answers where did this derived representation come from? Chapter 7 answers why does that evidence license this claim?

Why a graph?

Persistent interpretation could take many forms: summaries, timelines, tables, embeddings with metadata. The graph form earns its place through three properties, stated here as established background from the survey literature (Peng et al., ACM TOIS 2025; Zhang et al., preprint 2025), not as findings of this book.

First, project history is relational. Decisions supersede proposals, incidents motivate migrations, benchmarks replicate one another, runbooks echo decisions. Plain passage retrieval does not necessarily represent those connections as reusable objects; a graph can store them explicitly as relationships with descriptions and weights.

Second, relations compose across artifacts. No single passage states the contention observed in session-014, confirmed by the benchmark in session-019, caused the importer failure in incident-021, which motivated the decision in adr-007. That chain spans four artifacts. Relationships plus communities give the system somewhere to put multi-hop structure that no chunk contains whole.

Third, communities summarise at scale. When the corpus grows beyond any context window, per-passage retrieval degrades into sampling; community reports offer precomputed corpus-wide summaries organised by topic rather than by file. Whether that helps on real questions is what the Global condition was built to test — though on the fixture suite its home territory went unmeasured (see below).

The cost of these benefits must be stated alongside the mechanism, because the book’s later verdict depends on it: every benefit above is purchased with LLM inference at index time, stored state to maintain, and a new class of persistent error.

Persistent derived memory with GraphRAG

The implementation is Microsoft GraphRAG behind a backend abstraction the book owns. The abstraction matters more than the package: memory is defined here as persistent AI-derived structured understanding, and GraphRAG is the experimental implementation. A later system — incremental, lighter, hand-built — replaces the backend, not the book’s definition.

What the package contributes, verified against the installed version rather than the marketing text: documents, text units, entities, relationships, optional claim covariates, Leiden-hierarchy communities, community reports, and embeddings, all persisted as a parquet index.

Four query modes ship with it:

  • Basic — vector search over text units. The closest internal analogue of conventional retrieval.
  • Local — entity-centred, combining graph neighbourhoods, text units, and community context.
  • Global — map/reduce synthesis over community reports.
  • DRIFT — a primer over community reports, then iterative local follow-ups.

The chapter describes only the modes it actually exercised. DRIFT’s status is reported honestly below.

Source handling preserves the context that interpretation depends on.

Consider one sentence: PostgreSQL is the right choice. It means something different in a chat proposal, an assistant reply, a decision record, a benchmark note, and a commit message. Same words, five different kinds of claim.

So the source adapters normalise every input into a shared envelope — artifact identifier, source type, path, timestamp, actors, content, hash, metadata — reusing Chapter 3’s source identity rather than inventing competing identifiers.

Source type tells the system what kind of historical object it has. It does not make the content correct. That distinction governs the whole pipeline: nothing in the pipeline treats a decision record as true merely for being a decision record.

Provenance mapping walks every derived object back toward canonical artifacts through text units and input documents. Anything unreachable is recorded as an orphan with its broken link named, usable as a hypothesis but never as evidence. Health checks report counts, mapping coverage, near-duplicate entities, indexing failures, and build freshness — structural health, explicitly not semantic correctness. Full inspection commands are described in the repository.

Look inside the derived memory

Book result. The frozen index over the Chapter 3 fixture corpus (ch3-fixture-v0.1, corpus hash 04aac354) contains 20 documents, 20 text units, 39 entities, 65 relationships, 56 claims, 6 communities, and 5 community reports. Indexing ran 999 seconds of wall time, 55 LLM responses, and roughly 128,000 tokens, using ministral-3:8b for extraction and bge-m3 for embeddings. One community (community 2) has no report: the summariser’s output failed schema validation, recorded in the index report as an extraction failure. Build counts and provenance below are read directly from the committed index, not estimated.

The content is recognisably the project’s history, reorganised. SQLITE (degree 17) and POSTGRESQL (degree 8) anchor the migration; ADR-007, SESSION-044, SESSION-040, and the benchmark entities carry the decision and its evidence. A first community report summarises the event-store migration from SQLite to PostgreSQL through contention and benchmarking; another summarises the Redis rejection. So far, so much like a competent reading of the corpus — which is precisely the point. The graph stores what a good reader would conclude, so that later queries need not conclude it again.

Then the mess, which this chapter shows rather than tidies away.

Entity resolution failed in at least three places:

  • J. LINDQVIST and J. LINQVIST are the same person, split by a one-letter typo. The typo node owns three relationships of its own.
  • EVENT-STORE and EVENT STORE are the same system under two spellings, stored as two nodes of different declared kinds.
  • A. NOVAK and A.NOVAK repeat the spacing-split pattern.

These are not query-time mistakes that vanish on the next question. They are stored, versioned, queryable mistakes. Every downstream consumer inherits them — association in Chapter 5, routing in Chapter 6, lineage in Chapter 7 — until someone rebuilds the index.

The chapter’s central failure exhibit is not a wrong answer. It is a wrong node.

Provenance, by contrast, is complete:

Book result. All 166 derived objects (39 entities, 65 relationships, 56 claims, 6 communities) map back to canonical source artifacts. The orphan rate is zero in every kind. Source mapping tells us where each derived object came from; per the Layer 2 boundary, that is derivation provenance, not evidential support.

Health verdict: structurally unhealthy for exactly one reason — the missing community-2 report. The sole indexing failure is named, logged, and preserved rather than repaired silently, because a rebuilt-silent index would break the versioning invariant the next section states.

One source, several questions

The canonical history runs Monday to Thursday: session-031 proposes moving the event store to PostgreSQL, session-033 prefers SQLite, adr-007 decides PostgreSQL. Six related questions address the same history, and this section investigates reuse across them rather than failure.

What did we decide?
Why?
Who proposed the alternative?
What was rejected?
What evidence supported the decision?
What systems are connected to it?

Strong RAG answers each by reconstructing the distinctions fresh: retrieving the three passages, reading headers and status lines, and assigning roles per query. The persistent alternative can draw from stored structure: entities for the people and systems, relationships for proposal and supersession, claims for the contention evidence, the community summary for the rationale. The empirical question is whether stored distinctions are more consistent across the six questions than six independent reconstructions — and what happens on the seventh question, where the stored distinctions are wrong.

Cross-query consistency of this form is unmeasured in the current runs; the runner records per-case answers but infers nothing across them, and the chapter labels consistency an explicit pending obligation rather than smuggling agreement across query modes into a metric. The claim it does test is narrower: whether graph-backed conditions answer the individual questions as well as strong RAG, at what cost, and with what new errors.

Basic Search is vector search over the index’s text units followed by the same reader that serves the baseline — the closest internal comparison to conventional retrieval, differing mainly in chunking and in what text the index holds. It is therefore the graph condition most comparable with Chapter 3; any divergence must first be attributed to the index and retrieval pipeline rather than to graph structure itself.

Local Search centres on entities: resolve the question’s entities, gather their neighbourhoods, and combine graph context with text units and community summaries. This is the mode whose design best matches decision and relational questions — what did we decide about X?, which incidents influenced the decision? — because those questions name entities whose stored neighbourhoods should contain the answer’s parts.

Global Search never retrieves passages. It maps over community reports and reduces the partial syntheses into an answer.

Its natural territory is corpus-wide synthesis — what major architectural themes emerged? — where no single passage is the answer and the work is compression across topics.

Judged on an exact local lookup, it is the wrong tool by construction. The chapter refuses to score it there as though the result meant something. The fair test routes global questions to Global Search and local questions to local modes, then reports each mode on its own territory, with the baseline measured everywhere.

DRIFT

DRIFT Search combines a community-report primer with iterative local follow-ups. Its status in this chapter is: unmeasured on the fixture suite — the frozen comparison contains no DRIFT cells, and Chapter 6 independently records DRIFT as unmeasured in its routing matrix. The chapter does not estimate DRIFT quality from documentation, and leaves DRIFT explicitly pending. Time and cost are experimental variables, and an unmeasured mode is reported as unmeasured.

The experiment

Book hypothesis (pre-registered). Preserving interpretation as a persistent derived graph will match strong RAG on local decision questions, improve relational and synthesis questions where structure composes across artifacts, keep per-answer provenance mappable to raw sources, and cost substantially more in indexing and query latency.

Four verdict types were registered before any analysis:

  • A — persistent structure earns a core role, through material gains at acceptable cost.
  • B — strong RAG remains sufficient; graph memory stays optional.
  • C — structure enables later mechanisms without improving direct question answering.
  • D — persistent derived state is net harmful; graph memory is demoted to experimental.

The verdict may combine categories.

The frozen comparison runs 14 tasks over the Chapter 3 fixture corpus, across seven families — locate, decision, provenance, temporal, use, relational, global. Every condition shares the same reader.

The graph conditions add one GraphRAG query per task. The derived synthesis takes at most a third of the final context, with raw passages filling the rest. Source-recall scoring covers only directly admitted raw evidence, so derived lineage can never inflate a recall score.

One coverage defect must be stated up front. Global ran 8 of the 14 tasks. Its six missing cells include both global-family and both relational-family tasks — which is to say, its home territory went untested. The analysis does not forgive that.

Two fairness qualifications must be stated before any number is read. First, total final context is matched but its composition is not: graph conditions surrender a third of raw budget to the derived block, so part of any recall gap measures that budget split rather than retrieval quality. Second, the unsupported-source scorer penalises graph conditions differentially for a citation-format difference rather than a grounding difference. Both are recorded as instrument obligations: alias lists and citation normalisation need repair before any rerun counts as publication-grade.

What happened

Mechanical scores first, exactly as the frozen artifacts record them (decision exactness over the 9 applicable cases; source recall over all tasks; context tokens estimated; latencies measured wall time):

conditiontaskssource recalldecision exactnessmean context tokensmean latency
Chapter 3 best141.0000.889 (8/9)580~7 s reader-side¹
Graph Basic140.8190.889 (8/9)131751.5 s
Graph Local140.8190.667 (6/9)125092.7 s
Graph Global80.8120.800 (4/5)1525196.1 s (max 585 s)

¹ The baseline cases record component latencies (retrieval plus roughly 5 s generation) rather than end-to-end totals; graph conditions record end-to-end totals including the GraphRAG call. Latency ratios are therefore approximate and, if anything, flatter the graph conditions’ overhead.

The mechanical decision gaps are dominated by scorer normalisation, not answer quality: missing alias lists fail correct paraphrases of the causal mechanism, and a word-order variant fails on a substantively correct answer. Full cell detail is committed alongside the frozen run as analysis, not as a modification of frozen data.

Adjudicated rescoring under one stated rule — rescue only cells whose answers contain the expected mechanism in paraphrase the alias list omits, or near-verbatim word-order variants; no rescue for vague or mechanism-free answers — gives best 8/9, Basic 9/9, Local 8/9, Global 4/5 on its covered tasks. The adjudication is committed alongside the frozen run as analysis, not as a modification of frozen data. The largest gap between any two full-coverage conditions is one case in nine, too small on this fixture to support a victory claim in any direction.

Per-family detail sharpens the picture without changing it.

The tested temporal questions are solved in every full-coverage condition — full recall, all current and historical answers correct, including the stale-summary trap. The baseline’s reader handles supersession straight from raw passages, and nothing in the graph improves or degrades that result on these cases.

One relational case is the run’s single genuine point for stored structure. The baseline misses it thinly; both Basic and Local answer it with the mechanism stated explicitly.

One synthesis case goes the other way. Local’s answer is genuinely vaguer than the baseline’s, naming performance issues without the contention, benchmark, or incident specifics.

Global’s synthesis territory went untested because of the missing cells, not because the scores judged it.

Graph structure is not automatically better retrieval

The honest summary: strong conventional RAG is extremely difficult to beat, and the graph conditions do not beat it.

After adjudication, every full-coverage condition answers eight or nine of nine decision cases correctly. The baseline does it with the smallest context, the lowest latency, no indexing phase, and no stored state to maintain or correct.

What the graph conditions pay for that parity:

  • Basic matches the baseline’s quality at roughly 2.3× the context and 7× the latency, plus a 999-second index build.
  • Local costs roughly 12× the latency, for one genuine miss more.
  • Global costs roughly 25×, with incomplete coverage.

This is a useful result, not a failed experiment. It is the result the book’s discipline demands: the rival was built strong, the reader and total final context budget were held fixed, and the new mechanism was not permitted to win by facing a weakened opponent or by hiding its costs. The composition of that budget still differs between conditions, as the fairness qualification above records. A chapter that had forced GraphRAG to win here would have taught less than this one does.

Where structure may still help

Three places remain open, each narrower than a general quality claim.

First, the relational case above: composing influences across five artifacts is the shape of question stored relationships exist for, and both graph conditions gave the mechanism explicitly where the baseline named identifiers. One case proves nothing; a relational family with ledger-derived scoring would prove something, and building it is a recorded obligation.

Second, corpus-wide synthesis is untested, not refuted. Global Search ran none of the synthesis questions. Its value proposition — precomputed community summaries when no passage is the answer — cannot be evaluated on a fixture whose questions a six-passage reader already answers. The small global-query family the design calls for does not exist yet; the chapter specifies it (a handful of theme-level questions, ledger-derived where possible, frozen-judge multi-dimensional scoring where synthesis requires judgement) rather than pretending the current suite covers it.

Third, downstream reuse is concrete. Chapter 5’s propagation ran over the derived graph through the snapshot adapter; Chapter 7 mapped derived graph artifacts into lineage; Chapter 6 routed to graph modes selectively. None of that is evidence that the graph answers questions better. It shows that persistent structure can support later mechanisms — a distinct claim the verdict keeps separate from direct question-answering quality.

Persistent mistakes

Book result. Persistent understanding creates reusable intelligence and reusable mistakes. The frozen index exhibits both: complete source mapping alongside the stored splits exhibited above — a typo-node person split, a spelling-split system, and a spacing-split participant. A reader’s misreading disappears with the query; these misreadings are stored state, inherited by every later mechanism until a rebuild corrects them.

Three properties make stored errors structurally different from transient ones. They are silent: nothing in a later answer marks which parts came from a typo node. They propagate downstream: association, routing, and lineage can all consume the graph as input. They are expensive to fix in the current backend: correction means re-extraction and community rebuild, not a better query-time prompt. The missing community-2 report is the same lesson in miniature — a generation failure persisted as structural absence, detected only because the health check names it.

The mitigation the architecture actually implements is boundary, not prevention. Raw sources stay canonical and reachable, so any derived claim can be re-derived or distrusted; provenance mapping exposes orphans before they are used as evidence; version metadata (corpus hash, package version, models, prompts, chunking, community settings, timestamp, code commit) identifies each derived build so errors attach to versions rather than to history itself. What the architecture does not implement is verification of derived content — that is Chapter 7’s work, and this chapter’s errors are its motivating exhibits.

Keep the original history

The rebuild rule follows from the epistemic status. Source corpus canonical, derived graph rebuildable, versions explicit.

Change the GraphRAG version, the prompts, the extraction model, the embedding model, or the source corpus, and you have a new derived-memory version. Experimental state is never silently overwritten. Every run manifest records corpus version and hash, package version, models, chunking, community level, and prompt hashes.

Incremental update is the open problem. When history grows, what must be rebuilt — and what happens to communities and provenance? The current backend answers only by full rebuild. LightRAG’s incremental union design (Guo et al., EMNLP Findings 2025) is recorded here as the contrast that makes that cost visible.

Update economics, not update machinery, is this chapter’s finding.

The intended fallback runs downward toward the canonical evidence: when graph-derived memory is uncertain or insufficient, the system can return to strong Chapter 3 retrieval over raw sources. That path is what later allows a routing mistake to remain recoverable rather than turning a derived interpretation into the only available account.

What did GraphRAG actually buy us?

Costs, measured where measurable and estimated nowhere:

  • Index build. 999 seconds of wall time, 55 LLM responses, roughly 128,000 tokens — for 20 small documents. Indexing amortises only over query volumes this fixture never approaches.
  • Query latency. Basic ~51 s, Local ~93 s, Global ~196 s mean with a 585 s maximum. The baseline’s reader-side cost is near 7 s.
  • Context. 2.2–2.6× the baseline’s tokens, for decision quality that is equal or indistinguishable.
  • Storage. A full parquet index plus embedding stores, holding edge lists a flat file could carry.
  • Model calls per graph query. Uncounted by the package’s public API, and recorded as unavailable rather than estimated.

Against that price, the frozen run shows no general gain on direct question answering: adjudicated decision accuracy differs by at most one case among full-coverage conditions, the tested temporal cases are solved everywhere, and the one relational bright spot is a single case. Downstream reuse exists — traversal, selective routing, and lineage mapping all consume the graph — but architectural reuse is not empirical quality, and the chapter does not convert it into a score.

Did persistent understanding earn its place?

Developmental verdict: Type C with a Type B core.

On direct question answering over this workload, strong RAG remains sufficient. The graph adds nothing in quality, at substantial cost. That is Type B.

The graph nevertheless supplies a persistent substrate — entities, relationships, claims, and communities with complete source mapping — that associative retrieval, selective routing, and evidence lineage genuinely consume. That is Type C.

Nothing supports Type A or Type D. Local’s vague synthesis answer and the stored entity errors are costs, not disqualifiers, and fallback and rebuild contain them.

This upgrades to a book result once four obligations are met: repaired scorer normalisation, a relational question family, a tested synthesis family, and cross-query consistency measurement.

The principle that survives any single package:

The transition studied here is not from retrieval to understanding. It is from transient interpretation to persistent derived state: some interpretation of the past survives the query and becomes available to influence future remembering.

Whether preserving that interpretation is worth what it costs is then an empirical question per workload — asked here, answered conditionally, and re-asked by every later layer that inherits the graph.

The next question

A graph is static. Its relationships sit there until a query arrives.

And retrieval over them is still fundamentally a lookup operation — entity-centred or community-aware, but driven by selecting stored material for a query.

Remembering may do something else. A cue touches one memory. Activation moves along the relations. The memory you needed arrives three steps later, by a route no similarity computation planned.

The map exists now, with its virtues and its typo nodes. The next chapter asks how recall should move through it — and whether moving buys anything that lookup, however well built, cannot.

References