<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Embeddings From First Principles on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/books/embeddings-from-first-principles/</link><description>Recent content in Embeddings From First Principles on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Wed, 09 Sep 2026 02:00:00 +0000</lastBuildDate><atom:link href="https://aibussin.com/books/embeddings-from-first-principles/index.xml" rel="self" type="application/rss+xml"/><item><title>What Is an Embedding?</title><link>https://aibussin.com/books/embeddings-from-first-principles/01-chapter/</link><pubDate>Mon, 07 Sep 2026 09:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/01-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part I — A Vector Is Not Meaning&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="three-words-and-a-list-of-numbers"&gt;Three words and a list of numbers&lt;/h2&gt;&#10;&lt;p&gt;Take three words:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cat&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;dog&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;airplane&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Turn them into numbers. Any embedding API will do it. You get three arrays, each maybe 384 or 768 or 1,536 floats long:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cat [ 0.021, -0.114, 0.062, ... ]&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;dog [ 0.019, -0.098, 0.071, ... ]&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;airplane [-0.087, 0.203, -0.041, ... ]&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Compute the angle between &lt;code&gt;cat&lt;/code&gt; and &lt;code&gt;dog&lt;/code&gt;. It is small. Compute the angle between &lt;code&gt;cat&lt;/code&gt; and &lt;code&gt;airplane&lt;/code&gt;. It is larger. Something about &amp;ldquo;cats and dogs are both pets&amp;rdquo; appears to have survived the trip into number-space.&lt;/p&gt;</description></item><item><title>Meaning Becomes Geometry</title><link>https://aibussin.com/books/embeddings-from-first-principles/02-chapter/</link><pubDate>Mon, 07 Sep 2026 09:10:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/02-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part I — A Vector Is Not Meaning&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="a-space-you-can-draw"&gt;A space you can draw&lt;/h2&gt;&#10;&lt;p&gt;Give six words two coordinates each, by hand:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; x = monarchy / institutional power y = gendered association (-1 … +1)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;king ( 0.9, 0.6)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;queen ( 0.9, -0.6)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;duke ( 0.7, 0.5)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;duchess ( 0.7, -0.5)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;apple (-0.7, 0.1)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;orange (-0.7, -0.1)&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Now several semantic questions have a geometric form:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;em&gt;Which words form a group?&lt;/em&gt; → &lt;code&gt;apple&lt;/code&gt; and &lt;code&gt;orange&lt;/code&gt; sit almost on top of each other, about 0.2 apart, and far from the four titles. The titles form a looser group of their own near &lt;code&gt;x ≈ 0.8&lt;/code&gt;. Grouping shows up as density.&lt;/li&gt;&#10;&lt;li&gt;&lt;em&gt;What distinguishes &lt;code&gt;king&lt;/code&gt; from &lt;code&gt;queen&lt;/code&gt;?&lt;/em&gt; → the vector &lt;code&gt;king − queen ≈ (0, 1.2)&lt;/code&gt;. It points almost purely along &lt;code&gt;y&lt;/code&gt;. The difference between them, in this space, is one attribute. (&lt;code&gt;duke − duchess ≈ (0, 1.0)&lt;/code&gt; points the same way, but only because we placed the points that way; a learned space owes us no such consistency.)&lt;/li&gt;&#10;&lt;li&gt;&lt;em&gt;Is &lt;code&gt;apple&lt;/code&gt; oriented like &lt;code&gt;king&lt;/code&gt;?&lt;/em&gt; → the angle between them is obtuse and their cosine is negative. Fruit and titles point into different half-planes. Orientation separates the two categories cleanly, even though, as we are about to see, it is noisier &lt;em&gt;within&lt;/em&gt; a category.&lt;/li&gt;&#10;&lt;li&gt;&lt;em&gt;Are &lt;code&gt;king&lt;/code&gt; and &lt;code&gt;queen&lt;/code&gt; related?&lt;/em&gt; → obviously: they are the two monarchs. But their Euclidean distance is 1.2, while &lt;code&gt;king&lt;/code&gt; sits 0.22 from &lt;code&gt;duke&lt;/code&gt; and about 1.12 from &lt;code&gt;duchess&lt;/code&gt;. Rank &lt;code&gt;king&lt;/code&gt;&amp;rsquo;s neighbors by distance and &lt;code&gt;queen&lt;/code&gt; comes third, behind &lt;code&gt;duke&lt;/code&gt; and &lt;code&gt;duchess&lt;/code&gt;. The gender axis — which a search for &amp;ldquo;royal titles&amp;rdquo; need not care about at all — has quietly reordered the neighborhood.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;That last point is the one to carry forward.&lt;/p&gt;</description></item><item><title>Learning an Embedding Space</title><link>https://aibussin.com/books/embeddings-from-first-principles/03-chapter/</link><pubDate>Mon, 07 Sep 2026 09:20:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/03-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part I — A Vector Is Not Meaning&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="refusing-the-handout"&gt;Refusing the handout&lt;/h2&gt;&#10;&lt;p&gt;For two chapters the vectors were given. Chapter 1 took them from an API and asked what was in them; Chapter 2 placed some by hand and treated the rest as the output of a model we agreed not to open. Chapter 2 closed on the question it left unanswered: &lt;strong&gt;where does the geometry come from?&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;This chapter builds a small space from nothing but a corpus and a counting rule. A usable geometry appears — related words land near each other without anyone labeling an axis &amp;ldquo;animal&amp;rdquo; or &amp;ldquo;drink.&amp;rdquo; Then comes the more important half: we inspect the construction closely enough to predict, from the mechanism alone, what it cannot preserve. When the measured results arrive, the failures are not surprises. They are consequences of the path the information took.&lt;/p&gt;</description></item><item><title>Similarity Is a Decision</title><link>https://aibussin.com/books/embeddings-from-first-principles/04-chapter/</link><pubDate>Mon, 07 Sep 2026 09:30:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/04-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part I — A Vector Is Not Meaning&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="two-vectors-four-answers"&gt;Two vectors, four answers&lt;/h2&gt;&#10;&lt;p&gt;Here are three document vectors, kept to three dimensions so every number is checkable:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;x = ( 2.0, 0.0, 0.0 ) a short doc, one strong topic&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;y = ( 6.0, 0.1, 0.0 ) a long doc, same topic, much more of it&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;z = ( 0.0, 2.0, 0.0 ) a short doc, a different topic&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Before computing anything, look at the geometry. &lt;code&gt;x&lt;/code&gt; and &lt;code&gt;y&lt;/code&gt; point in almost exactly the same direction but have very different lengths (norm 2 versus about 6). &lt;code&gt;x&lt;/code&gt; and &lt;code&gt;z&lt;/code&gt; have the same length but point ninety degrees apart. So &amp;ldquo;is &lt;code&gt;x&lt;/code&gt; more like &lt;code&gt;y&lt;/code&gt; or &lt;code&gt;z&lt;/code&gt;?&amp;rdquo; comes down to a prior question: &lt;strong&gt;which kind of difference should count — a difference in direction, or a difference in size?&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Dimensions Do Not Mean What You Think</title><link>https://aibussin.com/books/embeddings-from-first-principles/05-chapter/</link><pubDate>Mon, 07 Sep 2026 10:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/05-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part II — Inside the Space&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="what-does-dimension-173-mean"&gt;What does dimension 173 mean?&lt;/h2&gt;&#10;&lt;p&gt;Pull the 173rd coordinate of every vector in your corpus and sort. You get a list of documents ordered by something. Once in a while a coordinate is weakly readable — &amp;ldquo;this one runs high for questions&amp;rdquo; — but usually the sorted list has no nameable theme. Dimension 173 is not &amp;ldquo;formality,&amp;rdquo; not &amp;ldquo;sentiment,&amp;rdquo; not &amp;ldquo;is about sports.&amp;rdquo;&lt;/p&gt;&#10;&lt;p&gt;And yet the vectors work. Similarity search returns sensible results, clusters fall out, nearest neighbors are reasonable. The usable structure is there. It is just not sitting on the individual axes.&lt;/p&gt;</description></item><item><title>Neighborhoods and Hubs</title><link>https://aibussin.com/books/embeddings-from-first-principles/06-chapter/</link><pubDate>Mon, 07 Sep 2026 10:10:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/06-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part II — Inside the Space&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-question-a-vector-cannot-answer-alone"&gt;The question a vector cannot answer alone&lt;/h2&gt;&#10;&lt;p&gt;Hand someone a single embedding vector — &lt;code&gt;[0.13, -0.82, ...]&lt;/code&gt; — and ask what it means. They cannot say. Chapter 5 is why: the basis is not identifiable, the scale is model-specific, no coordinate names a concept.&lt;/p&gt;&#10;&lt;p&gt;Hand them the same vector &lt;em&gt;and&lt;/em&gt; its ten nearest neighbors — all customer-support apologies, all in the same register — and they can usually name the topic, the register, and roughly what the point is about.&lt;/p&gt;</description></item><item><title>How Many Dimensions Does Meaning Need?</title><link>https://aibussin.com/books/embeddings-from-first-principles/07-chapter/</link><pubDate>Mon, 07 Sep 2026 10:20:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/07-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part II — Inside the Space&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="a-representation-with-several-sizes"&gt;A representation with several sizes&lt;/h2&gt;&#10;&lt;p&gt;&lt;code&gt;all-mpnet-base-v2&lt;/code&gt; emits vectors of length 768. Embed the 1,173 RELATE items with it and ask several different questions about the size of the resulting cloud, and you do not get one answer:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;nominal dimension 768&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;centered matrix-rank ceiling 768 (= min(1,173−1, 768))&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;95%-of-variance dimension 205&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;participation ratio 67.0&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;entropy effective rank 387.4&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;intrinsic-dimension estimate (MLE, k=10) 6.81&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;intrinsic-dimension estimate (TwoNN) 4.23&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The Wave 2 dimensionality artifact records the last five numerical summaries plus nominal dimension; it does &lt;strong&gt;not&lt;/strong&gt; persist an observed matrix-rank value, so the line above reports only the algebraic ceiling. Same vectors, same corpus: the reported or implied notions of size still range from about 4 to 768 — almost two orders of magnitude.&lt;/p&gt;</description></item><item><title>The Shape of an Embedding Space</title><link>https://aibussin.com/books/embeddings-from-first-principles/08-chapter/</link><pubDate>Mon, 07 Sep 2026 10:30:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/08-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part II — Inside the Space&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="same-corpus-five-geometries"&gt;Same corpus, five geometries&lt;/h2&gt;&#10;&lt;p&gt;Freeze the input. Embed exactly the same 1,173 RELATE items with five models and, before scoring any task, measure the geometry each one imposes on identical text.&lt;/p&gt;&#10;&lt;p&gt;The evaluation harness requests L2-normalized embeddings (&lt;code&gt;normalize_embeddings=True&lt;/code&gt;), so every vector analyzed in this chapter lies on the unit sphere, and cosine is purely a statement about direction.&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;em&gt;MEASURED on RELATE v0.1, Wave 2 row 2.8 — artifact &lt;code&gt;experiments/embeddings-from-first-principles/wave2/artifacts/shape-comparison.json&lt;/code&gt;. All descriptors computed from one run on the frozen corpus.&lt;/em&gt;&lt;/p&gt;</description></item><item><title>From Similarity to Search</title><link>https://aibussin.com/books/embeddings-from-first-principles/09-chapter/</link><pubDate>Mon, 07 Sep 2026 11:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/09-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part III — Retrieval Is an Experiment&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="retrieval-in-four-lines"&gt;Retrieval in four lines&lt;/h2&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#75715e"&gt;# assume embed() returns L2-normalized vectors, so dot product == cosine (Chapter 4)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;query_vec &lt;span style="color:#f92672"&gt;=&lt;/span&gt; embed(query) &lt;span style="color:#75715e"&gt;# one unit vector&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;scores &lt;span style="color:#f92672"&gt;=&lt;/span&gt; corpus_vecs &lt;span style="color:#f92672"&gt;@&lt;/span&gt; query_vec &lt;span style="color:#75715e"&gt;# a cosine score for every document&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;order &lt;span style="color:#f92672"&gt;=&lt;/span&gt; np&lt;span style="color:#f92672"&gt;.&lt;/span&gt;argsort(&lt;span style="color:#f92672"&gt;-&lt;/span&gt;scores) &lt;span style="color:#75715e"&gt;# rank all n documents, best first&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;results &lt;span style="color:#f92672"&gt;=&lt;/span&gt; [corpus[i] &lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; i &lt;span style="color:#f92672"&gt;in&lt;/span&gt; order[:k]] &lt;span style="color:#75715e"&gt;# keep the top k&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That is the whole primitive for the simple policy shown here. A vector database may add several different kinds of machinery around it: an ANN index or quantization may &lt;strong&gt;approximate&lt;/strong&gt; the same ranking more cheaply; sharding may distribute the computation; filtering may change the eligible candidate set; persistence and replication may make the system operable. Do not collapse all of those into &amp;ldquo;faster search&amp;rdquo; — some preserve the target policy, some approximate it, and some redefine it.&lt;/p&gt;</description></item><item><title>The Nearest Neighbor Can Be Wrong</title><link>https://aibussin.com/books/embeddings-from-first-principles/10-chapter/</link><pubDate>Mon, 07 Sep 2026 11:10:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/10-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part III — Retrieval Is an Experiment&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-result-that-is-closest-and-wrong"&gt;The result that is closest and wrong&lt;/h2&gt;&#10;&lt;p&gt;&lt;em&gt;Illustrative example — not a RELATE measurement.&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;Query:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Is Dublin the capital of Ireland?&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Imagine the passage that ranks first is this one:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Dublin is not, and has never been, the capital of Ireland — that&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;distinction belongs to the older seat of government at Tara.&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Fluent, on-topic, confidently phrased, and among the geometrically closest things in the corpus — and false. A model handed this passage as context may repeat its claim. The retrieval system did its job: it found a near vector. &amp;ldquo;Near&amp;rdquo; was not &amp;ldquo;correct.&amp;rdquo;&lt;/p&gt;</description></item><item><title>Hard Negatives</title><link>https://aibussin.com/books/embeddings-from-first-principles/11-chapter/</link><pubDate>Mon, 07 Sep 2026 11:20:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/11-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part III — Retrieval Is an Experiment&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="same-model-four-negative-sets-four-verdicts"&gt;Same model, four negative sets, four verdicts&lt;/h2&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;em&gt;MEASURED on RELATE v0.1, Wave 1 row 1.8 — artifact &lt;code&gt;experiments/embeddings-from-first-principles/wave1/artifacts/margin-collapse.json&lt;/code&gt;. Model &lt;code&gt;all-mpnet-base-v2&lt;/code&gt;, unchanged across every row below.&lt;/em&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;negative-selection rule mean margin selected-positive win rate&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;same-domain random, up to 5 per query 0.4738 1.0000&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;in-model selector, as implemented 0.2094 0.9814&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;structured perturbations 0.0932 0.9331&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;BM25 lexical top-5 0.0595 0.7621&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The model did not change. The 269-query set did not change. The reference-positive rule did not change. Only the rule that picked which passages the positive had to beat changed — and the mean margin ranged from 0.4738 to 0.0595, a nearly eightfold difference between the widest and narrowest measured conditions.&lt;/p&gt;</description></item><item><title>Retrieval Is a Policy</title><link>https://aibussin.com/books/embeddings-from-first-principles/12-chapter/</link><pubDate>Mon, 07 Sep 2026 11:30:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/12-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part III — Retrieval Is an Experiment&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="just-retrieve-the-relevant-documents"&gt;&amp;ldquo;Just retrieve the relevant documents&amp;rdquo;&lt;/h2&gt;&#10;&lt;p&gt;There is no such operation. Chapter 11 closed on a single insight: a score never travels alone — it arrives with a negative-selection rule that is part of the measurement. This chapter widens the frame. A retrieval &lt;em&gt;result&lt;/em&gt; never travels alone either. It arrives wrapped in a query construction, a representation, a candidate-eligibility rule, a scoring and fusion policy, a selection rule, an execution strategy, and a context-assembly step that decides what a downstream model actually sees.&lt;/p&gt;</description></item><item><title>How Do You Evaluate an Embedding?</title><link>https://aibussin.com/books/embeddings-from-first-principles/13-chapter/</link><pubDate>Mon, 07 Sep 2026 12:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/13-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part IV — Measuring the Representation&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="same-encoders-same-item-pool-different-winner"&gt;Same encoders, same item pool, different winner&lt;/h2&gt;&#10;&lt;p&gt;Three encoders. One item pool. One typed-pair pool. One metric implementation. Three relevance definitions, applied consistently. The encoder models and item-side corpus are unchanged between the two releases; what expands is the query population, from 269 queries to 411.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;RELATE v0.1 (269 queries): all-mpnet-base-v2 wins every relevance definition&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;RELATE v0.2 (411 queries): bge-large-en-v1.5 wins every relevance definition&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That is not a hypothetical. It is Wave 1 row 1.9, run twice against two frozen releases, and it is this chapter&amp;rsquo;s demonstration in miniature before the demonstration itself. &lt;strong&gt;The models did not change. The evaluation workload did. The winner changed.&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Calibration</title><link>https://aibussin.com/books/embeddings-from-first-principles/14-chapter/</link><pubDate>Mon, 07 Sep 2026 12:10:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/14-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part IV — Measuring the Representation&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="what-is-081"&gt;What is 0.81?&lt;/h2&gt;&#10;&lt;p&gt;&lt;em&gt;Illustrative scenario — invented values for intuition, not a measured result.&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;A pipeline decides two documents are &amp;ldquo;duplicates&amp;rdquo; if their cosine exceeds 0.8. Someone picked 0.8 because it looked reasonable. Imagine the two populations that number is actually sitting between:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;distribution of cosine for KNOWN duplicate pairs: mean ~0.79&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;distribution of cosine for KNOWN non-duplicate pairs: mean ~0.71&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;If those two populations overlap the way this sketch implies, the threshold sits inside both: a true duplicate at 0.78 is rejected, an unrelated pair at 0.82 is accepted. This chapter&amp;rsquo;s actual experiment, later, replaces the sketch with measured positive and negative score samples and derived discrimination/operating statistics. The artifact does not persist the full distributions themselves, and its 85.6% figure is the coverage of a specific quantile-derived band rather than a generic measure of distributional overlap.&lt;/p&gt;</description></item><item><title>Is Similarity One-Dimensional?</title><link>https://aibussin.com/books/embeddings-from-first-principles/15-chapter/</link><pubDate>Mon, 07 Sep 2026 12:20:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/15-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part IV — Measuring the Representation&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="two-pairs-same-cosine-different-situations"&gt;Two pairs, same cosine, different situations&lt;/h2&gt;&#10;&lt;p&gt;&lt;em&gt;Illustrative vignette — constructed for intuition, not a measured RELATE example.&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;Consider what a retrieval system typically logs when it claims a match: one number.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Pair 1: cos = 0.78&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; the query&amp;#39;s nearest neighbor scores 0.78; the 2nd scores 0.44&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; the region is sparse&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; → the winner stands alone&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Pair 2: cos = 0.78&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; the nearest neighbor scores 0.78; the 2nd, 3rd, 4th score 0.77, 0.76, 0.76&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; the region is crowded, several near-ties within a few hundredths of the winner&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; → a coin toss dressed as a match&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The scalar is identical. The situations, as constructed, are not: in Pair 1 the winner has no close competition; in Pair 2 three runners-up sit within &lt;code&gt;0.01&lt;/code&gt;–&lt;code&gt;0.02&lt;/code&gt; score points of the winner and the local neighborhood is dense with candidates that could plausibly have won instead. Whatever separates a decisive match from a fragile one in this sketch was available at retrieval time. Cosine, by construction, cannot express it — not because cosine is broken, but because it was never asked the question.&lt;/p&gt;</description></item><item><title>Change the Model, Change the Universe</title><link>https://aibussin.com/books/embeddings-from-first-principles/16-chapter/</link><pubDate>Mon, 07 Sep 2026 13:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/16-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part V — Embedding Spaces Are Not Universal&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-same-sentence-three-universes"&gt;The same sentence, three universes&lt;/h2&gt;&#10;&lt;p&gt;Embed one sentence with three models:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model A (384-d): [ ... ]&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model B (768-d): [ ... ]&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model C (768-d): [ ... ]&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A and B differ in length, so nobody expects to compare them coordinate-wise. But B and C are both 768-dimensional. It is tempting to line up their vectors and read something off the alignment directly — to compute &lt;code&gt;cos(B_sentence, C_sentence)&lt;/code&gt; and interpret whatever number comes back.&lt;/p&gt;</description></item><item><title>Versioning the Space</title><link>https://aibussin.com/books/embeddings-from-first-principles/17-chapter/</link><pubDate>Mon, 07 Sep 2026 13:10:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/17-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part V — Embedding Spaces Are Not Universal&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-upgrade-that-broke-search-quietly"&gt;The upgrade that broke search quietly&lt;/h2&gt;&#10;&lt;p&gt;&lt;em&gt;Illustrative scenario — not a measured incident.&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;Production runs embedding model v1. A better v2 ships. Someone updates the client library. New documents get v2 vectors; the existing documents are still carrying v1 vectors. Nothing errors. Queries are embedded with v2 and scored directly against a mix of v1 and v2 document vectors.&lt;/p&gt;&#10;&lt;p&gt;No exception is raised, because none of this is mechanically invalid — a v2 query and a v1 document, at the same dimension, dot-product into a perfectly ordinary float. What actually breaks is quieter than an error: the index now contains vectors from two declared spaces while every query is produced in only one of them, so a portion of the comparisons the system runs no longer satisfy the representation contract either side assumed. Retrieval keeps returning ten plausible-looking results. Nothing about the response shape announces that some of those ten were scored by an unsupported cross-space comparison. That is exactly the danger: silent retrieval-contract violation, not silent catastrophe.&lt;/p&gt;</description></item><item><title>Can One Embedding Space Be Translated Into Another?</title><link>https://aibussin.com/books/embeddings-from-first-principles/18-chapter/</link><pubDate>Mon, 07 Sep 2026 14:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/18-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VI — Crossing Embedding Spaces · Can a map be learned at all?&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-question-stated-carefully"&gt;The question, stated carefully&lt;/h2&gt;&#10;&lt;p&gt;We have two encoders, &lt;code&gt;E_A&lt;/code&gt; and &lt;code&gt;E_B&lt;/code&gt;, and a set of objects &lt;code&gt;x&lt;/code&gt;. Each object has two representations: &lt;code&gt;E_A(x)&lt;/code&gt; in space A, &lt;code&gt;E_B(x)&lt;/code&gt; in space B. Chapter 16 established that these are distinct declared spaces with no coordinate correspondence assumed by default. When dimensions permit, raw cross-space arithmetic is numerically possible, but without an alignment contract it has no authorized cross-space semantics. That is not the same claim as &amp;ldquo;the spaces are unrelated.&amp;rdquo; Chapter 16 also measured, between several genuinely independent encoders, substantial neighborhood overlap and high linear CKA — real structural agreement, existing side by side with the absence of any authorized coordinate correspondence. Those two facts coexisting is exactly what makes this chapter&amp;rsquo;s question worth asking rather than settled in advance.&lt;/p&gt;</description></item><item><title>Alignment</title><link>https://aibussin.com/books/embeddings-from-first-principles/19-chapter/</link><pubDate>Mon, 07 Sep 2026 14:10:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/19-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VI — Crossing Embedding Spaces · Which map family survives held-out data?&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="what-chapter-18-left-open"&gt;What Chapter 18 left open&lt;/h2&gt;&#10;&lt;p&gt;Chapter 18 fit the simplest serious bridge — affine ridge regression — from &lt;code&gt;all-MiniLM-L6-v2&lt;/code&gt; into &lt;code&gt;all-mpnet-base-v2&lt;/code&gt;, and evaluated it on entity families the map never saw. The result was not a verdict. It was a profile, uneven enough to make one number impossible to trust on its own:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;coordinate reconstruction 0.5860&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;10-NN neighborhood overlap 0.7128&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;retrieval nDCG@10 ratio 0.7987&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;rank-triplet agreement 0.7529&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;calibration-transfer score 0.7882&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;relation-profile Pearson correlation 0.8673&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;structured hard-negative margin ratio 0.2342&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;held-out/train reconstruction ratio 0.6480&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The chapter&amp;rsquo;s lesson was not &amp;ldquo;ridge works.&amp;rdquo; It was that &lt;code&gt;≈&lt;/code&gt; is a preservation contract — a named property, measured on held-out data, not a single number a map either clears or fails. That leaves an open question this chapter exists to answer:&lt;/p&gt;</description></item><item><title>The Embedding Bridge</title><link>https://aibussin.com/books/embeddings-from-first-principles/20-chapter/</link><pubDate>Mon, 07 Sep 2026 14:20:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/20-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VI — Crossing Embedding Spaces · When may a system actually use one?&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="a-map-is-not-a-bridge"&gt;A map is not a bridge&lt;/h2&gt;&#10;&lt;p&gt;Chapter 19 produced several fitted candidates — a Procrustes pipeline with its dimensionality adapter, an affine ridge map, a small MLP — each with its own held-out preservation profile, and none of them a declared winner. Chapter 19 ended deliberately unresolved: &lt;code&gt;status: measured_not_authorized&lt;/code&gt;. Deploying any one of those transformations as &amp;ldquo;space A and space B are now compatible&amp;rdquo; is the mistake this chapter exists to prevent.&lt;/p&gt;</description></item><item><title>Did the Bridge Preserve the Space?</title><link>https://aibussin.com/books/embeddings-from-first-principles/21-chapter/</link><pubDate>Mon, 07 Sep 2026 14:30:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/21-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VI — Crossing Embedding Spaces · What did each measurement actually certify?&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="one-bridge-eight-measurements"&gt;One bridge, eight measurements&lt;/h2&gt;&#10;&lt;p&gt;Chapter 20 built the machinery that turns preservation evidence into a scoped authorization — a requirement, a measured value, and a decision, all cited together. What it deliberately did not do is ask how much any one of those measured values is actually entitled to say. That is this chapter&amp;rsquo;s entire job.&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;A preservation metric is a sensor. Before trusting its reading, ask which failure modes the sensor is even capable of seeing.&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Retrieval Is Not Geometry</title><link>https://aibussin.com/books/embeddings-from-first-principles/22-chapter/</link><pubDate>Wed, 09 Sep 2026 01:50:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/22-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VI — Crossing Embedding Spaces · Did it find the counterpart, or recreate the neighborhood?&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-result-that-looks-like-success"&gt;The result that looks like success&lt;/h2&gt;&#10;&lt;p&gt;Fit a bridge, translate a held-out source vector, and ask the simplest possible question: does the exact native target counterpart show up nearby? On the Wave 6 benchmark, for a paired ridge bridge translating &lt;code&gt;mxbai-embed-large-v1&lt;/code&gt; into &lt;code&gt;Qwen3-Embedding-8B&lt;/code&gt;, the answer is about as good as that question can get. Across 589 held-out sentences, using the same fitted bridge and the same translated vectors throughout:&lt;/p&gt;</description></item><item><title>What Should a Translation Preserve?</title><link>https://aibussin.com/books/embeddings-from-first-principles/23-chapter/</link><pubDate>Wed, 09 Sep 2026 02:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/23-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VI — Crossing Embedding Spaces · Whose geometry counts as success?&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="preserve-the-geometry-is-not-a-complete-instruction"&gt;&amp;ldquo;Preserve the geometry&amp;rdquo; is not a complete instruction&lt;/h2&gt;&#10;&lt;p&gt;A bridge translates vectors from space A into space B. The natural instinct is to say:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Preserve the source geometry while you translate.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;That sounds obviously correct. If two source points are close, keep them close. If two source points are far apart, keep them far apart. Preserve pairwise cosine, distances, neighborhoods.&lt;/p&gt;</description></item><item><title>Can a Smaller Representation Preserve a Larger One?</title><link>https://aibussin.com/books/embeddings-from-first-principles/24-chapter/</link><pubDate>Mon, 07 Sep 2026 15:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/24-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VII — What Survives Transformation&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="two-vectors-for-one-document"&gt;Two vectors for one document&lt;/h2&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;full_doc → E(full_doc) one vector, dimension d&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;summary(doc) → E(summary(doc)) one vector, dimension d&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Both vectors have exactly the same dimension. What became smaller is the &lt;strong&gt;text&lt;/strong&gt; — the number of words available to encode — not the vector. One embedding was produced from a full document; the other from a much shorter compression of it, passed through the identical embedding pipeline. This is not the dimensionality reduction Chapter 7 covered, and it is not the cross-space translation Part VI built — both the full document and its compression are embedded natively, by the same encoder, into the same space. The question is:&lt;/p&gt;</description></item><item><title>From Deltas to Operators</title><link>https://aibussin.com/books/embeddings-from-first-principles/25-chapter/</link><pubDate>Mon, 07 Sep 2026 15:10:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/25-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VII — What Survives Transformation&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-other-thing-a-subtraction-might-mean"&gt;The other thing a subtraction might mean&lt;/h2&gt;&#10;&lt;p&gt;&lt;code&gt;king − man + woman ≈ queen&lt;/code&gt; is the famous demonstration that a &lt;em&gt;direction&lt;/em&gt; in embedding space can correspond to a semantic relation. It is also, as Chapter 2 noted, partly curated and works best locally — and it has a long list of documented problems: the offset method&amp;rsquo;s success is entangled with plain cosine-neighborhood structure, so a &amp;ldquo;the direction transfers&amp;rdquo; result has to beat the baseline of &lt;em&gt;ignoring the offset and returning the nearest neighbour of the source word&lt;/em&gt; (&lt;a href="https://aclanthology.org/W16-2503/"&gt;Linzen, 2016&lt;/a&gt;); and performance varies wildly by relation type (&lt;a href="https://aclanthology.org/S17-1017/"&gt;Rogers, Drozd &amp;amp; Li, 2017&lt;/a&gt;).&lt;/p&gt;</description></item><item><title>Building an Embedding Runtime</title><link>https://aibussin.com/books/embeddings-from-first-principles/26-chapter/</link><pubDate>Mon, 07 Sep 2026 16:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/26-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VIII — Embeddings Become Infrastructure&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="why-an-embedding-runtime-became-necessary"&gt;Why an embedding runtime became necessary&lt;/h2&gt;&#10;&lt;p&gt;An embedding system used to look like this:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;text → vector → cosine → top-k&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Twenty-five chapters ago that was a reasonable sketch. It no longer is, because along the way this book kept asking questions that sketch has no place to answer:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Which exact pipeline — model, revision, normalization, prefix, truncation — produced this vector? (Chapters 1, 17)&lt;/li&gt;&#10;&lt;li&gt;Can two spaces even be compared, and does comparing them license mixing their vectors? (Chapters 16, 20)&lt;/li&gt;&#10;&lt;li&gt;Does a fitted map between two spaces preserve retrieval, or fine relations, or a calibrated threshold — and are those three different questions? (Chapters 18–21)&lt;/li&gt;&#10;&lt;li&gt;Does a compressed or whitened representation still support the task that mattered, or only the geometry that was easy to check? (Chapters 7, 8, 24)&lt;/li&gt;&#10;&lt;li&gt;Does a semantic edit admit a reusable vector operator at all, and how would we know before trusting one? (Chapter 25)&lt;/li&gt;&#10;&lt;li&gt;What happens after a chain of these — a translation, then a compression, then an edit? Does whatever passed the first hop still hold?&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;None of these questions has a script-sized answer. Each one produced a companion component — a &lt;code&gt;space_record&lt;/code&gt;, a &lt;code&gt;neighborhood_report&lt;/code&gt;, a &lt;code&gt;calibration_record&lt;/code&gt;, a &lt;code&gt;bridge&lt;/code&gt;, a &lt;code&gt;preservation_profile&lt;/code&gt;, a &lt;code&gt;compression_record&lt;/code&gt;, a &lt;code&gt;transformation_record&lt;/code&gt; — because each is a measurement, and a measurement needs somewhere to live that is not the next paragraph of prose. Once there are seven or eight of these, kept as separate scripts, a system built from them accumulates a specific failure mode: nothing enforces that a translated vector still carries its bridge&amp;rsquo;s scope, or that a compressed one still carries its retention knee, or that anyone checks before mixing two spaces that were never compared. The instruments exist. Nothing makes using them the path of least resistance.&lt;/p&gt;</description></item></channel></rss>