<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Retrieval on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/retrieval/</link><description>Recent content in Retrieval on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Wed, 23 Sep 2026 05:00:14 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/retrieval/index.xml" rel="self" type="application/rss+xml"/><item><title>What Remembering Means</title><link>https://aibussin.com/books/memory/01-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:01 +0100</pubDate><guid>https://aibussin.com/books/memory/01-chapter/</guid><description>&lt;p&gt;A small team spends a week arguing about where to keep its event log.&lt;/p&gt;&#10;&lt;p&gt;They start with SQLite. It works at first, then slows to a crawl once several writers hit it at once. So they move the event store to PostgreSQL and write down why. Within a month nobody thinks about it any more. It is just how the system works.&lt;/p&gt;&#10;&lt;p&gt;A few months later, a new contributor asks the team&amp;rsquo;s AI assistant to scaffold a second service, with its own event log.&lt;/p&gt;</description></item><item><title>From Similarity to Search</title><link>https://aibussin.com/books/embeddings-from-first-principles/09-chapter/</link><pubDate>Mon, 07 Sep 2026 11:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/09-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part III — Retrieval Is an Experiment&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="retrieval-in-four-lines"&gt;Retrieval in four lines&lt;/h2&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# assume embed() returns L2-normalized vectors, so dot product == cosine (Chapter 4)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;query_vec &lt;span style="color:#f92672"&gt;=&lt;/span&gt; embed(query) &lt;span style="color:#75715e"&gt;# one unit vector&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;scores &lt;span style="color:#f92672"&gt;=&lt;/span&gt; corpus_vecs &lt;span style="color:#f92672"&gt;@&lt;/span&gt; query_vec &lt;span style="color:#75715e"&gt;# a cosine score for every document&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;order &lt;span style="color:#f92672"&gt;=&lt;/span&gt; np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;argsort(&lt;span style="color:#f92672"&gt;-&lt;/span&gt;scores) &lt;span style="color:#75715e"&gt;# rank all n documents, best first&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;results &lt;span style="color:#f92672"&gt;=&lt;/span&gt; [corpus[i] &lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; i &lt;span style="color:#f92672"&gt;in&lt;/span&gt; order[:k]] &lt;span style="color:#75715e"&gt;# keep the top k&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That is the whole primitive for the simple policy shown here. A vector database may add several different kinds of machinery around it: an ANN index or quantization may &lt;strong&gt;approximate&lt;/strong&gt; the same ranking more cheaply; sharding may distribute the computation; filtering may change the eligible candidate set; persistence and replication may make the system operable. Do not collapse all of those into &amp;ldquo;faster search&amp;rdquo; — some preserve the target policy, some approximate it, and some redefine it.&lt;/p&gt;</description></item><item><title>Bring Back Only What You Need</title><link>https://aibussin.com/books/context/14-chapter/</link><pubDate>Wed, 23 Sep 2026 05:00:14 +0100</pubDate><guid>https://aibussin.com/books/context/14-chapter/</guid><description>&lt;p&gt;Chapter 13 left the model holding a tidy live context and a shelf of external artifacts:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Incident 17 — lock-order inversion.&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Evidence: artifact://incident-17&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;plus five more references of the same shape. Then the current task changes. The user asks why the deadlock appeared only under concurrent import. The incident artifact probably matters. So might the benchmark report, the migration plan, and the compiler run. The naive response is to load everything, which recreates within one turn the exact occupancy problem Chapter 13 solved. Leaving was only half the problem. Existence plus recoverability does not mean admission. This chapter decides what comes back.&lt;/p&gt;</description></item><item><title>Context Is a Bottleneck</title><link>https://aibussin.com/books/memory/14-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:14 +0100</pubDate><guid>https://aibussin.com/books/memory/14-chapter/</guid><description>&lt;p&gt;Chapter 10 took selection as far as a frozen run has taken it: which memories the present work needs, admitted by an explicit, traceable policy at a fixed budget.&lt;/p&gt;&#10;&lt;p&gt;This chapter shows selection is not enough. Admission assumes the admitted memories &lt;em&gt;fit&lt;/em&gt;. Increasingly they do not — and the Chapter 10 run puts a number on the shortfall.&lt;/p&gt;&#10;&lt;p&gt;The ledger oracle reaches perfect required-evidence recall on a mean of 587 estimated tokens. The best non-oracle Chapter 10 condition spends 1,161 to reach 0.902 recall. Rendered with the source headers and validity marks the reader actually sees, those become roughly 805 against 1,505. The gap survives rendering.&lt;/p&gt;</description></item><item><title>Search, Memory and Long Context</title><link>https://aibussin.com/books/dspy-from-first-principles/16-chapter/</link><pubDate>Fri, 28 Aug 2026 11:10:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/16-chapter/</guid><description>&lt;p&gt;Chapter 15 measured an agent that had every advantage. Two files, one defect, a bounded tool surface, a budget it used well. It found both implicated files, held both halves of the bug in context, and still stopped one step short. That was a reasoning failure, and Chapter 17 exists to attack it.&lt;/p&gt;&#10;&lt;p&gt;This chapter is about the problem that sits &lt;em&gt;before&lt;/em&gt; the reasoning problem. The fixture had two files. A real repository has thousands, and a bounded tool surface does not tell the program which of them to read. Something has to decide what evidence the program even looks at, and that decision is not one mechanism. It is at least five, and they are routinely collapsed into a single word.&lt;/p&gt;</description></item><item><title>Retrieval Is Not Geometry</title><link>https://aibussin.com/books/embeddings-from-first-principles/22-chapter/</link><pubDate>Wed, 09 Sep 2026 01:50:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/22-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VI — Crossing Embedding Spaces · Did it find the counterpart, or recreate the neighborhood?&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-result-that-looks-like-success"&gt;The result that looks like success&lt;/h2&gt;&#10;&lt;p&gt;Fit a bridge, translate a held-out source vector, and ask the simplest possible question: does the exact native target counterpart show up nearby? On the Wave 6 benchmark, for a paired ridge bridge translating &lt;code&gt;mxbai-embed-large-v1&lt;/code&gt; into &lt;code&gt;Qwen3-Embedding-8B&lt;/code&gt;, the answer is about as good as that question can get. Across 589 held-out sentences, using the same fitted bridge and the same translated vectors throughout:&lt;/p&gt;</description></item><item><title>RELATE: Searching Embeddings by Relation, Not Just Similarity</title><link>https://aibussin.com/post/relate/</link><pubDate>Wed, 05 Aug 2026 23:24:45 +0100</pubDate><guid>https://aibussin.com/post/relate/</guid><description>&lt;p&gt;Embeddings are everywhere in modern AI.&lt;/p&gt;&#10;&lt;p&gt;They power semantic search, retrieval-augmented generation, recommendations, clustering, duplicate detection, code search, memory systems, and many of the mechanisms through which an AI system decides what information is relevant.&lt;/p&gt;&#10;&lt;p&gt;Yet most systems interrogate embeddings in essentially the same way:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Take two vectors and calculate cosine similarity.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;That is useful. But it also makes a strong assumption.&lt;/p&gt;&#10;&lt;p&gt;It assumes that the information we care about is expressed directly through the default geometry of the embedding space.&lt;/p&gt;</description></item></channel></rss>