The RAG Baseline
Build conventional retrieval-augmented memory, inspect what reaches the reader, and establish the baseline every later mechanism must beat.
The simplest answer to the memory problem is also the one every more elaborate system has to beat: keep the project history, retrieve the parts that look relevant, and let a capable model read them.
That approach already has many of the ingredients people casually call memory. The past is retained. A query selects evidence from it. A reader interprets that evidence in the context of the present task. If the resulting behaviour improves because the right part of the past was recovered, the system has done something useful without maintaining a separate persistent account of what the project believes, what changed, or what should be remembered next.
So before building richer memory machinery, we need a serious baseline:
How far can a well-built conventional RAG system take us before additional memory machinery earns its cost?
This cannot be answered with a deliberately weak retriever. If a later mechanism only wins because the baseline used poor chunking, weak search, no reranking, an artificially small context, or an incapable reader, then the comparison tells us very little about memory. Conventional retrieval has to be treated as a strong rival.
The companion implementation therefore combines PostgreSQL, lexical and vector search, rank fusion, a conventional reranker, explicit context assembly, and a capable reader. It preserves stage traces so that a wrong answer can be followed from source collection through retrieval, ranking, admission, and interpretation.
Chapter 2’s Memory Measurement Instrument supplies the tasks and judges the outputs. This chapter does not redefine that instrument. It gives the instrument a system worth testing.
What counts as conventional RAG?
The boundary is easiest to describe by asking what the system keeps between queries. Here it keeps source records, chunks, their ordinary metadata, and search indexes. It does not maintain its own interpreted account of the project’s decisions or current beliefs.
A source path, file type, content hash, chunk position, or date present in an artifact is ordinary retrieval metadata. A field declaring that one statement is the current authoritative decision is an interpretation. If the source itself says a proposal was rejected, the baseline can read that statement. It does not receive the evaluator’s rejection label as a hidden shortcut.
The implemented parser recognises filenames such as adr-007.md as decision-record files. That is a convention visible in the source collection; it does not establish whether the contents are accepted, superseded, disputed, or applicable to this query. In the current retrieval queries, that file-type field is not an authority boost.
Project artifacts
↓ discover, parse, retain provenance, chunk
PostgreSQL
├── source records and chunk text
├── full-text index
└── pgvector embeddings and optional ANN index
↓
Query → lexical candidates + dense candidates
↓
rank fusion
↓
query–passage reranking
↓
deduplicate and admit context
↓
reader → answer + citations
RELATE — the companion embeddings book’s relation-aware system — sits outside this baseline, as do knowledge graphs, persistent inferred relationships, belief state, consolidation, and learned memory policies. Their possible benefits belong to later comparisons. Ordinary retrieval engineering remains inside: better encoders, better lexical matching, better rerankers, and better use of a bounded context are legitimate improvements to the rival those comparisons face.
This is a functional baseline rather than a permanent commitment to one model. Each published comparison freezes an implementation. A later comparison may use a stronger conventional system, with a new version and new runs. A mechanism that only beats an obsolete retriever has a much narrower claim than one that survives that refresh.
Follow one piece of history
The dates now expose a distinction the baseline has to preserve.
The project history contains an earlier SQLite setup, a contention report, a benchmark, a later PostgreSQL decision, and subsequent work carried out under that decision. Months later, a contributor asks for a second service with its own event log. The repository contains both the evidence that matters and a great deal that does not.
The baseline’s job is not to remember a date because the book told it to. Its job is to retrieve enough of this history for the reader to reconstruct the relevant state.
| Memory condition | What the reader sees | What the answer must do |
|---|---|---|
| No supplied history | Present task only | Guess or abstain |
| Full history | Every retained passage | Find the decisive evidence among everything else |
| Strong RAG | Ranked relevant candidates | Distinguish proposals, evidence, decisions, and later state |
| Assembled context | Selected evidence within a budget | Preserve what is decisive while excluding or compressing the rest |
This is a walkthrough of the canonical history, not a claimed retrieval trace. The example defines what useful retrieval would have to make available without pretending that an experiment has already established where each source ranks.
Each question asks retrieval for something different:
- Where did we discuss the backend? The sessions and the decision record are useful places to read.
- Why PostgreSQL? The reader needs the contention evidence, and its connection to the decision.
- What runs in production on 15 July? The decision alone is not enough. New work targets PostgreSQL, but production still runs SQLite until 22 July.
- Build the second service. The system has to recover the decision that governs new work — not merely find an older passage containing the same technology names.
For this example, no special temporal representation is required if the relevant sentences reach a capable reader together. Whether ordinary retrieval delivers those sentences, whether context assembly preserves them, and whether the reader uses them correctly are separate questions. That separation is the chapter’s organising principle.
The walkthrough also shows why the corpus matters. Ingesting the entire manuscript would expose explanations of the test. Ingesting the hidden ledger would expose the answers. The retriever receives only the designated history directory; tasks, expected sources, and scoring labels stay with the instrument.
Build the historical substrate
The companion implementation discovers supported files and turns each readable, non-empty file into a source record. The current reader accepts UTF-8 Markdown, text, several source-code formats, JSON, YAML, TOML, and log files. These are text inputs; accepting a Python file does not imply understanding its syntax, and accepting a log export does not imply a live Git connector.
The source identifier is the path relative to the selected corpus root. Each source has a content hash, a file-type label, and a timestamp when the parser recognises one. Chunks retain their source identifier, ordinal, character span, text, and content hash. The stored row also records the chunker and embedding versions.
Those fields let us ask a concrete diagnostic question: which part of which source produced the passage the reader saw? Without that route back to the source, a plausible answer cannot be inspected and an incorrect answer cannot be repaired with confidence.
The PostgreSQL schema keeps source identity, chunk text with its search representations, and schema metadata in three tables, with chunks keyed to their sources and full-text and vector search available over chunk text. These facilities are documented in the pgvector project.
One database keeps the evidence and its search representations inspectable together. It also keeps the teaching problem manageable: the reader can follow a source from file to row to candidate without crossing several storage services.
The authoritative history still needs its own retention policy. This index is a rebuildable search view. Removing a file from a live checkout and deleting its indexed chunks does not preserve the old file. A project that needs historical versions must supply those versions as artifacts; an index of today’s checkout is not automatically an archive of the project.
The chunk is an engineering decision
Suppose a passage ends immediately after the sentence proposing Redis. The rejection appears in the next passage. A retriever may find the proposal perfectly and still omit the outcome. Alternatively, putting an entire session into one chunk may keep both statements together while spending most of the context on unrelated work.
Chunking decides what travels together. It changes the evidence available to ranking and the price of admitting that evidence to the reader. The implementation offers three policies: fixed character windows with overlap (is simple local coverage sufficient?), sentence packing around a target size (does preserving sentences improve usable evidence?), and Markdown heading boundaries (does document structure keep the right context together?).
The default is sentence packing with a target of 2,000 characters and 300 characters of overlap. Those are configuration values, not token counts or empirically established optima. Markdown-section chunks can exceed the target size, and the current code has no syntax-aware code chunker.
Overlap can preserve a qualification across a boundary, but it also repeats text in storage and in candidate lists. Deduplicating exact repeated passages does not remove every partial overlap. A useful comparison measures evidence coverage after context assembly as well as before it. If two overlapping chunks occupy two admission slots but supply one fact, source recall alone will miss the waste.
The development driver can compare the three policies. The existing short fixture does not establish which one handles long sessions best. A discriminating experiment needs decisions split from their rationale, headings that help and headings that mislead, and artifacts large enough for the policies to make different boundaries. Identical scores on documents that largely fit into one chunk would say little about chunking quality.
What the embedding contributes
The first retrieval path maps each passage into a vector and maps the query into a compatible space. The Dense Passage Retrieval paper establishes the independently encoded query-and-passage approach and evaluates it on open-domain question answering. That is background for the mechanism, not evidence of its performance on this project’s history.
For non-zero vectors, cosine similarity is:
$$ \operatorname{cosine}(q, p) \;=\; \frac{q \cdot p}{\lVert q \rVert \, \lVert p \rVert} $$It compares direction while removing magnitude. On unit-normalised vectors, dot product and cosine rank identically; squared Euclidean distance is 2 − 2 × cosine, so minimising that distance gives the same ranking as maximising cosine. Changing among equivalent scoring rules cannot recover a missing distinction.
The useful lesson from Embeddings From First Principles is that a retrieval result belongs to a representation and a comparison rule. A high similarity score is not a probability that a passage is true, authoritative, current, or sufficient to answer the question. It is also not proof that the encoder lacks those distinctions. A scalar comparison may fail to expose information that another reader can recover from the text.
The Redis sequence makes this concrete. session-040 proposes Redis. session-044 judges it unnecessary. adr-009 records the rejection.
All three concern the same technology and the same purpose. Finding all three would be useful. Treating whichever one ranks first as the decision would be a different operation, and a poorly justified one.
The provider interface makes the encoder replaceable. The existing driver names BGE-M3, Nomic Embed Text, Mixedbread, MiniLM (available through two providers), and Qwen3 Embedding as candidates. The configured default is BGE-M3. The available development comparison does not justify calling it the best model for the book’s workload, and this chapter makes no such claim.
Model choice includes more than a name: revision, dimension, input preparation, query/document conventions, and runtime matter. Candidate selection needs held-out retrieval tasks, Recall@k, ranking metrics, latency, and storage measurements. Neither model size nor a public leaderboard supplies the answer for this corpus.
Exact words still matter
The second path searches the words. A query containing adr-007, a function name, an error code, or a rare version string carries information that a semantic representation may not preserve strongly enough. Lexical retrieval deserves an independent condition in the experiment.
The current implementation builds English full-text representations with to_tsvector, combines extracted query terms with OR, and ranks matches using ts_rank_cd (implementation details are in the repository).
The OR permits partial matches: a passage need not contain every word of a natural-language question. PostgreSQL’s text-search configuration normalises words before matching; its ranking functions order the resulting matches. This is PostgreSQL full-text ranking, not an implementation of BM25. See the PostgreSQL text-search documentation.
There is a limit worth making visible. The current query parser extracts alphanumeric terms, and the indexed representation contains chunk text. That filenames exist in the source table does not mean exact identifier preservation or filename search is fully implemented. A failure on a punctuated identifier would first earn work on this retrieval path. It would not establish a need for a new memory representation.
Combine candidates before choosing context
The lexical and dense paths produce scores on different scales. Adding a text-search score to a cosine score would give those scales an arbitrary influence. The baseline instead uses reciprocal rank fusion: a passage receives a rank-based contribution from each list in which it appears, and those contributions are summed (implementation details are in the repository).
Here k is the fusion constant, configured as 60; it is not the number of passages admitted to context. A passage near the top of both lists receives two contributions. A passage found only by one path can still survive. The method combines rankings without pretending their original scores are comparable.
Fusion is not verification. Two retrievers can agree on a stale summary. Their agreement says that the passage deserves consideration, not that its claim should govern the answer.
The configured retrieval paths each request up to 30 candidates. Their union can exceed 30; the current candidate_k setting does not itself truncate that union. That distinction belongs in cost accounting. A comparison which gives hybrid search twice as many candidates as lexical search has held the per-path budget fixed, not the total candidate budget.
Let a second model read the candidates
A cross-encoder scores the query and passage together, allowing interaction between their tokens before producing a relevance score. This is the conventional second stage described in the Sentence Transformers documentation. It costs a model evaluation for each pair, which is why it operates on candidates rather than every passage in a large corpus.
The companion uses cross-encoder/ms-marco-MiniLM-L-6-v2 and retains up to eight reranked passages. Its current input preparation takes the first 2,000 characters of each candidate. A qualification beyond that cut cannot influence its score, even if the full passage would later reach the reader. Truncation is part of the experiment, not an invisible library detail.
It would be inaccurate to describe this whole system as one independent cosine comparison. Reranking introduces joint query–passage scoring. The reader then sees several passages together and can reason across them. Conventional RAG already includes substantial interpretation at query time.
Reranking earns its place if it improves the evidence admitted under a fixed context allowance enough to justify its latency. A configured reranker is a credible candidate component, not a guarantee of improvement. The ladder retains the condition without it so that the experiment can remove it if necessary.
Retrieval is a proposal; context is the decision
A source can exist in the database, enter the candidate pool, survive fusion, and still never reach the reader. These are different events. The context boundary makes them observable.
ContextTrace records admitted passages, duplicate drops, budget drops, character count, an estimated token count, and source coverage. Assembly walks the ranked list, suppresses passages whose text is identical after whitespace and case normalisation, and applies passage and character limits. It records source diversity but does not enforce a diversity quota.
The ordinary configuration admits at most six passages with a nominal 6,000-character allowance. There are three qualifications in the current implementation. The first passage is admitted even if it exceeds the character allowance. Source labels and separators are added during rendering and are not included in the admitted-text count. Reported tokens are estimated as characters divided by four.
Consequently, this is not yet an exact token-budget implementation. Comparing later systems at a claimed identical token limit requires counting the complete rendered input with the reader’s tokenizer and enforcing the limit consistently. Until then, the experiment can report character allowances and estimated tokens, with those limitations attached.
Suppose the contention benchmark enters the candidate set but loses its place to three overlapping extracts from an old setup guide. Candidate recall can be perfect while the explanation lacks its evidence. The appropriate repair concerns admission, deduplication, or ranking. A new persistent representation has not yet earned credit.
Position effects also deserve testing. Lost in the Middle tested multi-document question answering and key-value retrieval and found performance depended on where relevant information appeared. That motivates testing order and irrelevant context with the reader used here. It does not establish the size of those effects for this implementation.
From passages to an answer
The core path in Baseline.ask is short enough to read directly:
retrieval = retriever.retrieve(query)
context = assemble(
retrieval.reranked, context_override or self.config.context
)
answer = self.generator.answer(query, context)
The surrounding method returns all three objects. A final answer is accompanied by the evidence search and the actual context from which it was produced.
The reader receives source and chunk identifiers alongside the text. Its task is to answer from the supplied project history, identify its evidence, distinguish proposals from decisions, and acknowledge when the supplied history is insufficient. Those are ordinary reading requirements. The baseline is allowed to use a capable reader that meets them.
Citation checks need care. A named file may exist without supporting the claim attributed to it. The current instrument’s unsupported-source check detects unknown identifiers; it does not establish that every sentence follows from the cited passage. Claim-level support and a complete rationale chain require additional evaluation.
The question from Chapter 1 remains decisive. Naming PostgreSQL in an answer can pass a mechanical state check. It does not establish that the generated code uses PostgreSQL, or that it avoids the failed SQLite configuration.
Establishing that needs a downstream task with an inspectable outcome, run against a matched condition with no supplied history. One caution about that condition: “no memory” means no retrieved project history, not empty model weights. The base model still has everything it learned in training.
Keep the index healthy
A wrong answer caused by a missing chunk is not evidence against ordinary retrieval. Before measuring usefulness, the baseline needs an integrity check.
The current refresh cycle discovers files, compares hashes, indexes changed sources, removes sources no longer seen, creates the vector index where supported, and checks for orphan chunks and mixed embedding versions. The health report counts sources and chunks, checks duplicate hashes and orphan records, lists embedding versions, and checks index presence.
Those are useful diagnostics. They are not yet a complete self-correcting refresh system. Code inspection exposes three consequential gaps. Unchanged source hashes skip reprocessing even when the chunker or encoder changes. Updating a source and replacing its chunks are not enclosed in one transaction. A file that cannot be parsed is not added to the seen set, so the later removal step can treat its old indexed record as deleted.
Each has a concrete consequence. A new encoder can leave old vectors behind. An interrupted update can leave the source hash ahead of its chunks. A temporary read failure can remove previously searchable history. Detecting some of these problems afterwards does not make the update atomic or self-repairing.
A reliable refresh has to compare representation identity as well as source content, commit each replacement consistently, and distinguish absence from failed observation. The present implementation is a working companion with these limits, not a finished production reliability claim.
Health and memory quality remain separate even after those repairs. A database can contain exactly the intended rows and still retrieve the wrong evidence. Conversely, a good answer on a tiny fixture can conceal an index defect. The health checker answers whether the substrate passes its implemented checks; Chapter 2 asks whether the resulting history use helps.
Connect the instrument
evaluation.py adapts the baseline’s output into Chapter 2’s SystemOutput and calls the instrument’s score_task. It does not introduce a competing definition of correctness. Saved cases retain the answer, abstention flag, retrieved, admitted and cited source identifiers, estimated context tokens, stage timings, and observations.
| Condition | What it isolates |
|---|---|
| No supplied history | What the reader answers without project evidence. |
| Lexical | What word matching contributes. |
| Dense | What the embedding path contributes. |
| Hybrid | What combining candidate paths contributes. |
| Hybrid plus reranker | What the second-stage ranking contributes. |
| Best configured conventional system | The reference condition for later mechanisms. |
| Full history, where it fits | Whether selection helps compared with reading everything. |
| Oracle evidence | What the reader does when given evaluator-selected sources. |
In the current driver, best and hybrid-reranked use the same configured pipeline. They are two names for one configuration, not two independent architectural advances. The oracle selects whole expected source files and then assembles them; it is not guaranteed to be a minimal sufficient span oracle. Full history is a separate resource condition whose actual input size must be reported.
For a reader-strength experiment, each reader gets the same assembled evidence. If a stronger reader resolves the supposed representation failure, the book has found a reader limitation. For a budget experiment, the reader and retrieval configuration stay fixed while context changes. If admitting more ordinary history solves the problem at acceptable cost, that is a legitimate result too.
What the existing runs establish — and what they do not
Development artifacts exist for embedding comparisons, chunking comparisons, the eight-condition ladder, reader comparisons, and context sweeps. At this stage they do not supply an admissible comparative book result: the history is a small hand-authored fixture rather than the controlled-world generator or a separately adjudicated real corpus, and there is no genuine unfinished-work evaluation or executed downstream service-building test. Known fixture, scoring-stage, and labelling defects are recorded in the research notes for versioned follow-ups.
Chapter 12 later compares this strong-RAG condition with no memory and with the selected, assembled pipeline on behavioural tasks. That result establishes RAG as a strong rival; it does not retrospectively validate every embedding, chunking, or reranking choice in this chapter, nor isolate Chapters 4 and 5. The narrower status here is therefore implemented baseline; component-level comparative verdict still open.
Reading a failure without inventing its cause
When the answer is wrong, the investigation follows the evidence through the system:
| Observation | First place to investigate |
|---|---|
| Necessary evidence is absent from the exposed corpus. | Collection, parsing, or task answerability. |
| Evidence is indexed but absent from candidates. | Lexical matching, embedding, chunking, filters, or ANN search. |
| Evidence is a candidate but ranks below distractors. | Fusion and reranking. |
| It ranks adequately but is not admitted. | Duplication, truncation, and context limits. |
| It reaches context but the answer misreads it. | Interpretation by the reader. |
| The answer states the constraint but the action ignores it. | Downstream evidence use. |
| A stale source displaces the relevant current source. | The stage of displacement, then temporal interpretation. |
| Extra context appears to make the answer worse. | A matched removal or ordering experiment. |
| The output is defensible but the scorer rejects it. | Labels, normalisation, and the evaluation contract. |
The trace cannot directly reveal whether the model internally understood a passage and ignored it. That diagnosis needs observable evidence — for example, a correct explanation followed by contradictory code — and a controlled comparison. A confident attribution is not made from the final wrong answer alone.
Persistent representation becomes a candidate when ordinary retrieval and reading leave a repeatable, consequential gap after these alternatives have been tested. Even then, the new mechanism must show that it repairs the diagnosed failure under stated costs. The experiment can find that it merely moves the same mistake into an extraction stage.
Reproducibility
The companion runs against PostgreSQL with local embedding and generation models plus a reranker; the exact environment, commands, and evaluation driver are described in the repository. Reproducing a reported result requires more than those commands: immutable model and corpus identities, the effective configuration of each condition, a pinned grader, and traces sufficient to reconstruct admission and scoring.
Why this system is a serious rival
The baseline can ingest a new kind of project note without first inventing an ontology for it. Its search representations can be rebuilt from source text. A better encoder can improve candidate retrieval, a better reranker can improve ordering, and a better reader can recover distinctions directly from the evidence. Each component can be inspected and replaced independently.
A derived memory representation takes on additional responsibilities. It must interpret the source, retain important qualifications, react to updates, and preserve enough provenance to correct itself. Those responsibilities may pay for themselves. They are costs the conventional baseline does not incur in the same form.
This makes reader improvement a serious challenge to the rest of the book. The case for a persistent interpretation cannot rest on a distinction that a stronger reader recovers cheaply from the same passages. Nor can an elaborate system claim victory merely because it receives more context, more calls, or an easier version of the task. Later comparisons report accuracy alongside ingestion cost, query latency, storage, context use, and maintenance work.
The baseline also remains valuable when another mechanism wins. Raw source retrieval is the route back to the historical record when a derived interpretation is incomplete, stale, or disputed. It lets the system inspect what was actually written and lets a human audit the answer. Chapter 6 later tests explicit routing between memory mechanisms, including fallback to raw evidence, without assuming that such a router must become a permanent accuracy layer. The evidential value of the fallback exists independently of that later result.
What the next chapters will test
Later chapters will test whether richer derived structures improve on this baseline.
A useful result may show that structure helps only when evidence is scattered widely. It may show that a better reranker is sufficient. It may show that the reader, the context allowance, or the evaluator was the real bottleneck. Each outcome changes what additional machinery, if any, the book has earned.
Before sophisticated memory comes a credible account of how far retrieval can take us. If a later mechanism cannot improve on this baseline enough to justify its cost, we keep the baseline. The investigation has still produced an answer.