<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Embeddings on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/embeddings/</link><description>Recent content in Embeddings on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Fri, 25 Sep 2026 02:45:00 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/embeddings/index.xml" rel="self" type="application/rss+xml"/><item><title>What Is an Embedding?</title><link>https://aibussin.com/books/embeddings-from-first-principles/01-chapter/</link><pubDate>Mon, 07 Sep 2026 09:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/01-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part I — A Vector Is Not Meaning&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="three-words-and-a-list-of-numbers"&gt;Three words and a list of numbers&lt;/h2&gt;&#10;&lt;p&gt;Take three words:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cat&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;dog&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;airplane&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Turn them into numbers. Any embedding API will do it. You get three arrays, each maybe 384 or 768 or 1,536 floats long:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cat [ 0.021, -0.114, 0.062, ... ]&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;dog [ 0.019, -0.098, 0.071, ... ]&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;airplane [-0.087, 0.203, -0.041, ... ]&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Compute the angle between &lt;code&gt;cat&lt;/code&gt; and &lt;code&gt;dog&lt;/code&gt;. It is small. Compute the angle between &lt;code&gt;cat&lt;/code&gt; and &lt;code&gt;airplane&lt;/code&gt;. It is larger. Something about &amp;ldquo;cats and dogs are both pets&amp;rdquo; appears to have survived the trip into number-space.&lt;/p&gt;</description></item><item><title>Meaning Becomes Geometry</title><link>https://aibussin.com/books/embeddings-from-first-principles/02-chapter/</link><pubDate>Mon, 07 Sep 2026 09:10:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/02-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part I — A Vector Is Not Meaning&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="a-space-you-can-draw"&gt;A space you can draw&lt;/h2&gt;&#10;&lt;p&gt;Give six words two coordinates each, by hand:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; x = monarchy / institutional power y = gendered association (-1 … +1)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;king ( 0.9, 0.6)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;queen ( 0.9, -0.6)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;duke ( 0.7, 0.5)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;duchess ( 0.7, -0.5)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;apple (-0.7, 0.1)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;orange (-0.7, -0.1)&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Now several semantic questions have a geometric form:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;em&gt;Which words form a group?&lt;/em&gt; → &lt;code&gt;apple&lt;/code&gt; and &lt;code&gt;orange&lt;/code&gt; sit almost on top of each other, about 0.2 apart, and far from the four titles. The titles form a looser group of their own near &lt;code&gt;x ≈ 0.8&lt;/code&gt;. Grouping shows up as density.&lt;/li&gt;&#10;&lt;li&gt;&lt;em&gt;What distinguishes &lt;code&gt;king&lt;/code&gt; from &lt;code&gt;queen&lt;/code&gt;?&lt;/em&gt; → the vector &lt;code&gt;king − queen ≈ (0, 1.2)&lt;/code&gt;. It points almost purely along &lt;code&gt;y&lt;/code&gt;. The difference between them, in this space, is one attribute. (&lt;code&gt;duke − duchess ≈ (0, 1.0)&lt;/code&gt; points the same way, but only because we placed the points that way; a learned space owes us no such consistency.)&lt;/li&gt;&#10;&lt;li&gt;&lt;em&gt;Is &lt;code&gt;apple&lt;/code&gt; oriented like &lt;code&gt;king&lt;/code&gt;?&lt;/em&gt; → the angle between them is obtuse and their cosine is negative. Fruit and titles point into different half-planes. Orientation separates the two categories cleanly, even though, as we are about to see, it is noisier &lt;em&gt;within&lt;/em&gt; a category.&lt;/li&gt;&#10;&lt;li&gt;&lt;em&gt;Are &lt;code&gt;king&lt;/code&gt; and &lt;code&gt;queen&lt;/code&gt; related?&lt;/em&gt; → obviously: they are the two monarchs. But their Euclidean distance is 1.2, while &lt;code&gt;king&lt;/code&gt; sits 0.22 from &lt;code&gt;duke&lt;/code&gt; and about 1.12 from &lt;code&gt;duchess&lt;/code&gt;. Rank &lt;code&gt;king&lt;/code&gt;&amp;rsquo;s neighbors by distance and &lt;code&gt;queen&lt;/code&gt; comes third, behind &lt;code&gt;duke&lt;/code&gt; and &lt;code&gt;duchess&lt;/code&gt;. The gender axis — which a search for &amp;ldquo;royal titles&amp;rdquo; need not care about at all — has quietly reordered the neighborhood.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;That last point is the one to carry forward.&lt;/p&gt;</description></item><item><title>Learning an Embedding Space</title><link>https://aibussin.com/books/embeddings-from-first-principles/03-chapter/</link><pubDate>Mon, 07 Sep 2026 09:20:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/03-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part I — A Vector Is Not Meaning&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="refusing-the-handout"&gt;Refusing the handout&lt;/h2&gt;&#10;&lt;p&gt;For two chapters the vectors were given. Chapter 1 took them from an API and asked what was in them; Chapter 2 placed some by hand and treated the rest as the output of a model we agreed not to open. Chapter 2 closed on the question it left unanswered: &lt;strong&gt;where does the geometry come from?&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;This chapter builds a small space from nothing but a corpus and a counting rule. A usable geometry appears — related words land near each other without anyone labeling an axis &amp;ldquo;animal&amp;rdquo; or &amp;ldquo;drink.&amp;rdquo; Then comes the more important half: we inspect the construction closely enough to predict, from the mechanism alone, what it cannot preserve. When the measured results arrive, the failures are not surprises. They are consequences of the path the information took.&lt;/p&gt;</description></item><item><title>Similarity Is a Decision</title><link>https://aibussin.com/books/embeddings-from-first-principles/04-chapter/</link><pubDate>Mon, 07 Sep 2026 09:30:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/04-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part I — A Vector Is Not Meaning&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="two-vectors-four-answers"&gt;Two vectors, four answers&lt;/h2&gt;&#10;&lt;p&gt;Here are three document vectors, kept to three dimensions so every number is checkable:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;x = ( 2.0, 0.0, 0.0 ) a short doc, one strong topic&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;y = ( 6.0, 0.1, 0.0 ) a long doc, same topic, much more of it&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;z = ( 0.0, 2.0, 0.0 ) a short doc, a different topic&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Before computing anything, look at the geometry. &lt;code&gt;x&lt;/code&gt; and &lt;code&gt;y&lt;/code&gt; point in almost exactly the same direction but have very different lengths (norm 2 versus about 6). &lt;code&gt;x&lt;/code&gt; and &lt;code&gt;z&lt;/code&gt; have the same length but point ninety degrees apart. So &amp;ldquo;is &lt;code&gt;x&lt;/code&gt; more like &lt;code&gt;y&lt;/code&gt; or &lt;code&gt;z&lt;/code&gt;?&amp;rdquo; comes down to a prior question: &lt;strong&gt;which kind of difference should count — a difference in direction, or a difference in size?&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Dimensions Do Not Mean What You Think</title><link>https://aibussin.com/books/embeddings-from-first-principles/05-chapter/</link><pubDate>Mon, 07 Sep 2026 10:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/05-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part II — Inside the Space&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="what-does-dimension-173-mean"&gt;What does dimension 173 mean?&lt;/h2&gt;&#10;&lt;p&gt;Pull the 173rd coordinate of every vector in your corpus and sort. You get a list of documents ordered by something. Once in a while a coordinate is weakly readable — &amp;ldquo;this one runs high for questions&amp;rdquo; — but usually the sorted list has no nameable theme. Dimension 173 is not &amp;ldquo;formality,&amp;rdquo; not &amp;ldquo;sentiment,&amp;rdquo; not &amp;ldquo;is about sports.&amp;rdquo;&lt;/p&gt;&#10;&lt;p&gt;And yet the vectors work. Similarity search returns sensible results, clusters fall out, nearest neighbors are reasonable. The usable structure is there. It is just not sitting on the individual axes.&lt;/p&gt;</description></item><item><title>Neighborhoods and Hubs</title><link>https://aibussin.com/books/embeddings-from-first-principles/06-chapter/</link><pubDate>Mon, 07 Sep 2026 10:10:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/06-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part II — Inside the Space&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-question-a-vector-cannot-answer-alone"&gt;The question a vector cannot answer alone&lt;/h2&gt;&#10;&lt;p&gt;Hand someone a single embedding vector — &lt;code&gt;[0.13, -0.82, ...]&lt;/code&gt; — and ask what it means. They cannot say. Chapter 5 is why: the basis is not identifiable, the scale is model-specific, no coordinate names a concept.&lt;/p&gt;&#10;&lt;p&gt;Hand them the same vector &lt;em&gt;and&lt;/em&gt; its ten nearest neighbors — all customer-support apologies, all in the same register — and they can usually name the topic, the register, and roughly what the point is about.&lt;/p&gt;</description></item><item><title>How Many Dimensions Does Meaning Need?</title><link>https://aibussin.com/books/embeddings-from-first-principles/07-chapter/</link><pubDate>Mon, 07 Sep 2026 10:20:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/07-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part II — Inside the Space&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="a-representation-with-several-sizes"&gt;A representation with several sizes&lt;/h2&gt;&#10;&lt;p&gt;&lt;code&gt;all-mpnet-base-v2&lt;/code&gt; emits vectors of length 768. Embed the 1,173 RELATE items with it and ask several different questions about the size of the resulting cloud, and you do not get one answer:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;nominal dimension 768&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;centered matrix-rank ceiling 768 (= min(1,173−1, 768))&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;95%-of-variance dimension 205&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;participation ratio 67.0&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;entropy effective rank 387.4&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;intrinsic-dimension estimate (MLE, k=10) 6.81&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;intrinsic-dimension estimate (TwoNN) 4.23&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The Wave 2 dimensionality artifact records the last five numerical summaries plus nominal dimension; it does &lt;strong&gt;not&lt;/strong&gt; persist an observed matrix-rank value, so the line above reports only the algebraic ceiling. Same vectors, same corpus: the reported or implied notions of size still range from about 4 to 768 — almost two orders of magnitude.&lt;/p&gt;</description></item><item><title>Similarity Is Not Discovery</title><link>https://aibussin.com/books/language/08-chapter/</link><pubDate>Fri, 25 Sep 2026 02:45:00 +0100</pubDate><guid>https://aibussin.com/books/language/08-chapter/</guid><description>&lt;p&gt;The sidecar of Chapter 7 has an embarrassing regular guest. A researcher reading a paper on retrieval evaluation opens the side panel and finds, ranked first, a paper with nearly identical vocabulary — same benchmarks named, same metrics discussed, same dataset family. She opens it. It contributes nothing: same conclusions, weaker experiments, no new evidence, no disagreement, no extension. Topically it is the closest thing in the corpus to what she is reading. Informationally it is empty. The sidecar did its retrieval job perfectly and its discovery job not at all.&lt;/p&gt;</description></item><item><title>Feature Space: What Does a Linear Model Actually See?</title><link>https://aibussin.com/books/pytorch-from-first-principles/09-chapter/</link><pubDate>Fri, 28 Aug 2026 15:30:00 +0100</pubDate><guid>https://aibussin.com/books/pytorch-from-first-principles/09-chapter/</guid><description>&lt;p&gt;Here is a sequence classifier. Eight examples, 128 positions each, 768 features per position, and a linear head that produces one score.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;x &lt;span style="color:#f92672"&gt;=&lt;/span&gt; torch&lt;span style="color:#f92672"&gt;.&lt;/span&gt;randn(&lt;span style="color:#ae81ff"&gt;8&lt;/span&gt;, &lt;span style="color:#ae81ff"&gt;128&lt;/span&gt;, &lt;span style="color:#ae81ff"&gt;768&lt;/span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;classifier &lt;span style="color:#f92672"&gt;=&lt;/span&gt; nn&lt;span style="color:#f92672"&gt;.&lt;/span&gt;Linear(&lt;span style="color:#ae81ff"&gt;768&lt;/span&gt;, &lt;span style="color:#ae81ff"&gt;1&lt;/span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;scores &lt;span style="color:#f92672"&gt;=&lt;/span&gt; classifier(x)&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;input (8, 128, 768)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Linear(768,1) -&amp;gt; (8, 128, 1)&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;One thousand and twenty-four scores where eight were expected. Nothing raised, nothing is non-finite, and the shape is entirely predictable once you know the rule. The layer did exactly what the tensor asked of it: &lt;code&gt;nn.Linear&lt;/code&gt; transforms the last axis and preserves every axis before it, so it produced one score for every &lt;code&gt;(example, position)&lt;/code&gt; pair — 128 scores per example, computed independently.&lt;/p&gt;</description></item><item><title>Hard Negatives</title><link>https://aibussin.com/books/embeddings-from-first-principles/11-chapter/</link><pubDate>Mon, 07 Sep 2026 11:20:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/11-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part III — Retrieval Is an Experiment&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="same-model-four-negative-sets-four-verdicts"&gt;Same model, four negative sets, four verdicts&lt;/h2&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;em&gt;MEASURED on RELATE v0.1, Wave 1 row 1.8 — artifact &lt;code&gt;experiments/embeddings-from-first-principles/wave1/artifacts/margin-collapse.json&lt;/code&gt;. Model &lt;code&gt;all-mpnet-base-v2&lt;/code&gt;, unchanged across every row below.&lt;/em&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;negative-selection rule mean margin selected-positive win rate&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;same-domain random, up to 5 per query 0.4738 1.0000&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;in-model selector, as implemented 0.2094 0.9814&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;structured perturbations 0.0932 0.9331&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;BM25 lexical top-5 0.0595 0.7621&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The model did not change. The 269-query set did not change. The reference-positive rule did not change. Only the rule that picked which passages the positive had to beat changed — and the mean margin ranged from 0.4738 to 0.0595, a nearly eightfold difference between the widest and narrowest measured conditions.&lt;/p&gt;</description></item><item><title>Retrieval Is a Policy</title><link>https://aibussin.com/books/embeddings-from-first-principles/12-chapter/</link><pubDate>Mon, 07 Sep 2026 11:30:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/12-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part III — Retrieval Is an Experiment&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="just-retrieve-the-relevant-documents"&gt;&amp;ldquo;Just retrieve the relevant documents&amp;rdquo;&lt;/h2&gt;&#10;&lt;p&gt;There is no such operation. Chapter 11 closed on a single insight: a score never travels alone — it arrives with a negative-selection rule that is part of the measurement. This chapter widens the frame. A retrieval &lt;em&gt;result&lt;/em&gt; never travels alone either. It arrives wrapped in a query construction, a representation, a candidate-eligibility rule, a scoring and fusion policy, a selection rule, an execution strategy, and a context-assembly step that decides what a downstream model actually sees.&lt;/p&gt;</description></item><item><title>Calibration</title><link>https://aibussin.com/books/embeddings-from-first-principles/14-chapter/</link><pubDate>Mon, 07 Sep 2026 12:10:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/14-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part IV — Measuring the Representation&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="what-is-081"&gt;What is 0.81?&lt;/h2&gt;&#10;&lt;p&gt;&lt;em&gt;Illustrative scenario — invented values for intuition, not a measured result.&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;A pipeline decides two documents are &amp;ldquo;duplicates&amp;rdquo; if their cosine exceeds 0.8. Someone picked 0.8 because it looked reasonable. Imagine the two populations that number is actually sitting between:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;distribution of cosine for KNOWN duplicate pairs: mean ~0.79&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;distribution of cosine for KNOWN non-duplicate pairs: mean ~0.71&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;If those two populations overlap the way this sketch implies, the threshold sits inside both: a true duplicate at 0.78 is rejected, an unrelated pair at 0.82 is accepted. This chapter&amp;rsquo;s actual experiment, later, replaces the sketch with measured positive and negative score samples and derived discrimination/operating statistics. The artifact does not persist the full distributions themselves, and its 85.6% figure is the coverage of a specific quantile-derived band rather than a generic measure of distributional overlap.&lt;/p&gt;</description></item><item><title>Is Similarity One-Dimensional?</title><link>https://aibussin.com/books/embeddings-from-first-principles/15-chapter/</link><pubDate>Mon, 07 Sep 2026 12:20:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/15-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part IV — Measuring the Representation&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="two-pairs-same-cosine-different-situations"&gt;Two pairs, same cosine, different situations&lt;/h2&gt;&#10;&lt;p&gt;&lt;em&gt;Illustrative vignette — constructed for intuition, not a measured RELATE example.&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;Consider what a retrieval system typically logs when it claims a match: one number.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Pair 1: cos = 0.78&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; the query&amp;#39;s nearest neighbor scores 0.78; the 2nd scores 0.44&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; the region is sparse&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; → the winner stands alone&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Pair 2: cos = 0.78&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; the nearest neighbor scores 0.78; the 2nd, 3rd, 4th score 0.77, 0.76, 0.76&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; the region is crowded, several near-ties within a few hundredths of the winner&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; → a coin toss dressed as a match&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The scalar is identical. The situations, as constructed, are not: in Pair 1 the winner has no close competition; in Pair 2 three runners-up sit within &lt;code&gt;0.01&lt;/code&gt;–&lt;code&gt;0.02&lt;/code&gt; score points of the winner and the local neighborhood is dense with candidates that could plausibly have won instead. Whatever separates a decisive match from a fragile one in this sketch was available at retrieval time. Cosine, by construction, cannot express it — not because cosine is broken, but because it was never asked the question.&lt;/p&gt;</description></item><item><title>Change the Model, Change the Universe</title><link>https://aibussin.com/books/embeddings-from-first-principles/16-chapter/</link><pubDate>Mon, 07 Sep 2026 13:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/16-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part V — Embedding Spaces Are Not Universal&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-same-sentence-three-universes"&gt;The same sentence, three universes&lt;/h2&gt;&#10;&lt;p&gt;Embed one sentence with three models:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model A (384-d): [ ... ]&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model B (768-d): [ ... ]&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model C (768-d): [ ... ]&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A and B differ in length, so nobody expects to compare them coordinate-wise. But B and C are both 768-dimensional. It is tempting to line up their vectors and read something off the alignment directly — to compute &lt;code&gt;cos(B_sentence, C_sentence)&lt;/code&gt; and interpret whatever number comes back.&lt;/p&gt;</description></item><item><title>Can One Embedding Space Be Translated Into Another?</title><link>https://aibussin.com/books/embeddings-from-first-principles/18-chapter/</link><pubDate>Mon, 07 Sep 2026 14:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/18-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VI — Crossing Embedding Spaces · Can a map be learned at all?&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-question-stated-carefully"&gt;The question, stated carefully&lt;/h2&gt;&#10;&lt;p&gt;We have two encoders, &lt;code&gt;E_A&lt;/code&gt; and &lt;code&gt;E_B&lt;/code&gt;, and a set of objects &lt;code&gt;x&lt;/code&gt;. Each object has two representations: &lt;code&gt;E_A(x)&lt;/code&gt; in space A, &lt;code&gt;E_B(x)&lt;/code&gt; in space B. Chapter 16 established that these are distinct declared spaces with no coordinate correspondence assumed by default. When dimensions permit, raw cross-space arithmetic is numerically possible, but without an alignment contract it has no authorized cross-space semantics. That is not the same claim as &amp;ldquo;the spaces are unrelated.&amp;rdquo; Chapter 16 also measured, between several genuinely independent encoders, substantial neighborhood overlap and high linear CKA — real structural agreement, existing side by side with the absence of any authorized coordinate correspondence. Those two facts coexisting is exactly what makes this chapter&amp;rsquo;s question worth asking rather than settled in advance.&lt;/p&gt;</description></item><item><title>Retrieval Is Not Geometry</title><link>https://aibussin.com/books/embeddings-from-first-principles/22-chapter/</link><pubDate>Wed, 09 Sep 2026 01:50:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/22-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VI — Crossing Embedding Spaces · Did it find the counterpart, or recreate the neighborhood?&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-result-that-looks-like-success"&gt;The result that looks like success&lt;/h2&gt;&#10;&lt;p&gt;Fit a bridge, translate a held-out source vector, and ask the simplest possible question: does the exact native target counterpart show up nearby? On the Wave 6 benchmark, for a paired ridge bridge translating &lt;code&gt;mxbai-embed-large-v1&lt;/code&gt; into &lt;code&gt;Qwen3-Embedding-8B&lt;/code&gt;, the answer is about as good as that question can get. Across 589 held-out sentences, using the same fitted bridge and the same translated vectors throughout:&lt;/p&gt;</description></item><item><title>What Should a Translation Preserve?</title><link>https://aibussin.com/books/embeddings-from-first-principles/23-chapter/</link><pubDate>Wed, 09 Sep 2026 02:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/23-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VI — Crossing Embedding Spaces · Whose geometry counts as success?&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="preserve-the-geometry-is-not-a-complete-instruction"&gt;&amp;ldquo;Preserve the geometry&amp;rdquo; is not a complete instruction&lt;/h2&gt;&#10;&lt;p&gt;A bridge translates vectors from space A into space B. The natural instinct is to say:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Preserve the source geometry while you translate.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;That sounds obviously correct. If two source points are close, keep them close. If two source points are far apart, keep them far apart. Preserve pairwise cosine, distances, neighborhoods.&lt;/p&gt;</description></item><item><title>Can a Smaller Representation Preserve a Larger One?</title><link>https://aibussin.com/books/embeddings-from-first-principles/24-chapter/</link><pubDate>Mon, 07 Sep 2026 15:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/24-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VII — What Survives Transformation&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="two-vectors-for-one-document"&gt;Two vectors for one document&lt;/h2&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;full_doc → E(full_doc) one vector, dimension d&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;summary(doc) → E(summary(doc)) one vector, dimension d&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Both vectors have exactly the same dimension. What became smaller is the &lt;strong&gt;text&lt;/strong&gt; — the number of words available to encode — not the vector. One embedding was produced from a full document; the other from a much shorter compression of it, passed through the identical embedding pipeline. This is not the dimensionality reduction Chapter 7 covered, and it is not the cross-space translation Part VI built — both the full document and its compression are embedded natively, by the same encoder, into the same space. The question is:&lt;/p&gt;</description></item><item><title>From Deltas to Operators</title><link>https://aibussin.com/books/embeddings-from-first-principles/25-chapter/</link><pubDate>Mon, 07 Sep 2026 15:10:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/25-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VII — What Survives Transformation&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-other-thing-a-subtraction-might-mean"&gt;The other thing a subtraction might mean&lt;/h2&gt;&#10;&lt;p&gt;&lt;code&gt;king − man + woman ≈ queen&lt;/code&gt; is the famous demonstration that a &lt;em&gt;direction&lt;/em&gt; in embedding space can correspond to a semantic relation. It is also, as Chapter 2 noted, partly curated and works best locally — and it has a long list of documented problems: the offset method&amp;rsquo;s success is entangled with plain cosine-neighborhood structure, so a &amp;ldquo;the direction transfers&amp;rdquo; result has to beat the baseline of &lt;em&gt;ignoring the offset and returning the nearest neighbour of the source word&lt;/em&gt; (&lt;a href="https://aclanthology.org/W16-2503/"&gt;Linzen, 2016&lt;/a&gt;); and performance varies wildly by relation type (&lt;a href="https://aclanthology.org/S17-1017/"&gt;Rogers, Drozd &amp;amp; Li, 2017&lt;/a&gt;).&lt;/p&gt;</description></item><item><title>Building an Embedding Runtime</title><link>https://aibussin.com/books/embeddings-from-first-principles/26-chapter/</link><pubDate>Mon, 07 Sep 2026 16:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/26-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VIII — Embeddings Become Infrastructure&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="why-an-embedding-runtime-became-necessary"&gt;Why an embedding runtime became necessary&lt;/h2&gt;&#10;&lt;p&gt;An embedding system used to look like this:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;text → vector → cosine → top-k&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Twenty-five chapters ago that was a reasonable sketch. It no longer is, because along the way this book kept asking questions that sketch has no place to answer:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Which exact pipeline — model, revision, normalization, prefix, truncation — produced this vector? (Chapters 1, 17)&lt;/li&gt;&#10;&lt;li&gt;Can two spaces even be compared, and does comparing them license mixing their vectors? (Chapters 16, 20)&lt;/li&gt;&#10;&lt;li&gt;Does a fitted map between two spaces preserve retrieval, or fine relations, or a calibrated threshold — and are those three different questions? (Chapters 18–21)&lt;/li&gt;&#10;&lt;li&gt;Does a compressed or whitened representation still support the task that mattered, or only the geometry that was easy to check? (Chapters 7, 8, 24)&lt;/li&gt;&#10;&lt;li&gt;Does a semantic edit admit a reusable vector operator at all, and how would we know before trusting one? (Chapter 25)&lt;/li&gt;&#10;&lt;li&gt;What happens after a chain of these — a translation, then a compression, then an edit? Does whatever passed the first hop still hold?&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;None of these questions has a script-sized answer. Each one produced a companion component — a &lt;code&gt;space_record&lt;/code&gt;, a &lt;code&gt;neighborhood_report&lt;/code&gt;, a &lt;code&gt;calibration_record&lt;/code&gt;, a &lt;code&gt;bridge&lt;/code&gt;, a &lt;code&gt;preservation_profile&lt;/code&gt;, a &lt;code&gt;compression_record&lt;/code&gt;, a &lt;code&gt;transformation_record&lt;/code&gt; — because each is a measurement, and a measurement needs somewhere to live that is not the next paragraph of prose. Once there are seven or eight of these, kept as separate scripts, a system built from them accumulates a specific failure mode: nothing enforces that a translated vector still carries its bridge&amp;rsquo;s scope, or that a compressed one still carries its retention knee, or that anyone checks before mixing two spaces that were never compared. The instruments exist. Nothing makes using them the path of least resistance.&lt;/p&gt;</description></item><item><title>Preference Rankers — Learning Which Answer Is Better</title><link>https://aibussin.com/books/models-from-first-principles/09-chapter/</link><pubDate>Sat, 22 Aug 2026 01:18:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/09-chapter/</guid><description>&lt;h1 id="preference-rankers--learning-which-answer-is-better"&gt;Preference Rankers — Learning Which Answer Is Better&lt;/h1&gt;&#10;&lt;p&gt;So far in &lt;strong&gt;Models From First Principles&lt;/strong&gt;, we have mostly trained models by telling them what the answer should be.&lt;/p&gt;&#10;&lt;p&gt;MR.Q took a context and a response and produced a number:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;context + response&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 0.82&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That number might mean quality.&lt;/p&gt;&#10;&lt;p&gt;Or usefulness.&lt;/p&gt;&#10;&lt;p&gt;Or reward.&lt;/p&gt;&#10;&lt;p&gt;But there is an awkward question hiding underneath that architecture:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Where did the 0.82 come from?&lt;/p&gt;</description></item><item><title>MR.Q — Building a Neural Quality Model From Two Embeddings</title><link>https://aibussin.com/books/models-from-first-principles/02-chapter/</link><pubDate>Sat, 08 Aug 2026 14:33:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/02-chapter/</guid><description>&lt;h1 id="mrq--building-a-neural-quality-model-from-two-embeddings"&gt;MR.Q — Building a Neural Quality Model From Two Embeddings&lt;/h1&gt;&#10;&lt;p&gt;In the previous post, we established the core idea behind this series:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;A complicated model becomes understandable when you recursively decompose it into smaller models, blocks, layers and tensor operations.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;Now we build the first real model.&lt;/p&gt;&#10;&lt;p&gt;Not a transformer.&lt;/p&gt;&#10;&lt;p&gt;Not a giant language model.&lt;/p&gt;&#10;&lt;p&gt;Not an agent.&lt;/p&gt;&#10;&lt;p&gt;A scorer.&lt;/p&gt;&#10;&lt;p&gt;We will take two embeddings:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;context embedding&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;response embedding&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;combine them, encode the relationship between them, and predict one scalar:&lt;/p&gt;</description></item><item><title>RELATE: Searching Embeddings by Relation, Not Just Similarity</title><link>https://aibussin.com/post/relate/</link><pubDate>Wed, 05 Aug 2026 23:24:45 +0100</pubDate><guid>https://aibussin.com/post/relate/</guid><description>&lt;p&gt;Embeddings are everywhere in modern AI.&lt;/p&gt;&#10;&lt;p&gt;They power semantic search, retrieval-augmented generation, recommendations, clustering, duplicate detection, code search, memory systems, and many of the mechanisms through which an AI system decides what information is relevant.&lt;/p&gt;&#10;&lt;p&gt;Yet most systems interrogate embeddings in essentially the same way:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Take two vectors and calculate cosine similarity.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;That is useful. But it also makes a strong assumption.&lt;/p&gt;&#10;&lt;p&gt;It assumes that the information we care about is expressed directly through the default geometry of the embedding space.&lt;/p&gt;</description></item></channel></rss>