<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Cosine Similarity on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/cosine-similarity/</link><description>Recent content in Cosine Similarity on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Mon, 07 Sep 2026 09:30:00 +0000</lastBuildDate><atom:link href="https://aibussin.com/tags/cosine-similarity/index.xml" rel="self" type="application/rss+xml"/><item><title>Meaning Becomes Geometry</title><link>https://aibussin.com/books/embeddings-from-first-principles/02-chapter/</link><pubDate>Mon, 07 Sep 2026 09:10:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/02-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part I — A Vector Is Not Meaning&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="a-space-you-can-draw"&gt;A space you can draw&lt;/h2&gt;&#10;&lt;p&gt;Give six words two coordinates each, by hand:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; x = monarchy / institutional power y = gendered association (-1 … +1)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;king ( 0.9, 0.6)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;queen ( 0.9, -0.6)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;duke ( 0.7, 0.5)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;duchess ( 0.7, -0.5)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;apple (-0.7, 0.1)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;orange (-0.7, -0.1)&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Now several semantic questions have a geometric form:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&lt;em&gt;Which words form a group?&lt;/em&gt; → &lt;code&gt;apple&lt;/code&gt; and &lt;code&gt;orange&lt;/code&gt; sit almost on top of each other, about 0.2 apart, and far from the four titles. The titles form a looser group of their own near &lt;code&gt;x ≈ 0.8&lt;/code&gt;. Grouping shows up as density.&lt;/li&gt;&#10;&lt;li&gt;&lt;em&gt;What distinguishes &lt;code&gt;king&lt;/code&gt; from &lt;code&gt;queen&lt;/code&gt;?&lt;/em&gt; → the vector &lt;code&gt;king − queen ≈ (0, 1.2)&lt;/code&gt;. It points almost purely along &lt;code&gt;y&lt;/code&gt;. The difference between them, in this space, is one attribute. (&lt;code&gt;duke − duchess ≈ (0, 1.0)&lt;/code&gt; points the same way, but only because we placed the points that way; a learned space owes us no such consistency.)&lt;/li&gt;&#10;&lt;li&gt;&lt;em&gt;Is &lt;code&gt;apple&lt;/code&gt; oriented like &lt;code&gt;king&lt;/code&gt;?&lt;/em&gt; → the angle between them is obtuse and their cosine is negative. Fruit and titles point into different half-planes. Orientation separates the two categories cleanly, even though, as we are about to see, it is noisier &lt;em&gt;within&lt;/em&gt; a category.&lt;/li&gt;&#10;&lt;li&gt;&lt;em&gt;Are &lt;code&gt;king&lt;/code&gt; and &lt;code&gt;queen&lt;/code&gt; related?&lt;/em&gt; → obviously: they are the two monarchs. But their Euclidean distance is 1.2, while &lt;code&gt;king&lt;/code&gt; sits 0.22 from &lt;code&gt;duke&lt;/code&gt; and about 1.12 from &lt;code&gt;duchess&lt;/code&gt;. Rank &lt;code&gt;king&lt;/code&gt;&amp;rsquo;s neighbors by distance and &lt;code&gt;queen&lt;/code&gt; comes third, behind &lt;code&gt;duke&lt;/code&gt; and &lt;code&gt;duchess&lt;/code&gt;. The gender axis — which a search for &amp;ldquo;royal titles&amp;rdquo; need not care about at all — has quietly reordered the neighborhood.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;That last point is the one to carry forward.&lt;/p&gt;</description></item><item><title>Similarity Is a Decision</title><link>https://aibussin.com/books/embeddings-from-first-principles/04-chapter/</link><pubDate>Mon, 07 Sep 2026 09:30:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/04-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part I — A Vector Is Not Meaning&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="two-vectors-four-answers"&gt;Two vectors, four answers&lt;/h2&gt;&#10;&lt;p&gt;Here are three document vectors, kept to three dimensions so every number is checkable:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;x = ( 2.0, 0.0, 0.0 ) a short doc, one strong topic&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;y = ( 6.0, 0.1, 0.0 ) a long doc, same topic, much more of it&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;z = ( 0.0, 2.0, 0.0 ) a short doc, a different topic&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Before computing anything, look at the geometry. &lt;code&gt;x&lt;/code&gt; and &lt;code&gt;y&lt;/code&gt; point in almost exactly the same direction but have very different lengths (norm 2 versus about 6). &lt;code&gt;x&lt;/code&gt; and &lt;code&gt;z&lt;/code&gt; have the same length but point ninety degrees apart. So &amp;ldquo;is &lt;code&gt;x&lt;/code&gt; more like &lt;code&gt;y&lt;/code&gt; or &lt;code&gt;z&lt;/code&gt;?&amp;rdquo; comes down to a prior question: &lt;strong&gt;which kind of difference should count — a difference in direction, or a difference in size?&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Feature Space: What Does a Linear Model Actually See?</title><link>https://aibussin.com/books/pytorch-from-first-principles/09-chapter/</link><pubDate>Fri, 28 Aug 2026 15:30:00 +0100</pubDate><guid>https://aibussin.com/books/pytorch-from-first-principles/09-chapter/</guid><description>&lt;p&gt;Here is a sequence classifier. Eight examples, 128 positions each, 768 features per position, and a linear head that produces one score.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;x &lt;span style="color:#f92672"&gt;=&lt;/span&gt; torch&lt;span style="color:#f92672"&gt;.&lt;/span&gt;randn(&lt;span style="color:#ae81ff"&gt;8&lt;/span&gt;, &lt;span style="color:#ae81ff"&gt;128&lt;/span&gt;, &lt;span style="color:#ae81ff"&gt;768&lt;/span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;classifier &lt;span style="color:#f92672"&gt;=&lt;/span&gt; nn&lt;span style="color:#f92672"&gt;.&lt;/span&gt;Linear(&lt;span style="color:#ae81ff"&gt;768&lt;/span&gt;, &lt;span style="color:#ae81ff"&gt;1&lt;/span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;scores &lt;span style="color:#f92672"&gt;=&lt;/span&gt; classifier(x)&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;input (8, 128, 768)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Linear(768,1) -&amp;gt; (8, 128, 1)&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;One thousand and twenty-four scores where eight were expected. Nothing raised, nothing is non-finite, and the shape is entirely predictable once you know the rule. The layer did exactly what the tensor asked of it: &lt;code&gt;nn.Linear&lt;/code&gt; transforms the last axis and preserves every axis before it, so it produced one score for every &lt;code&gt;(example, position)&lt;/code&gt; pair — 128 scores per example, computed independently.&lt;/p&gt;</description></item><item><title>RELATE: Searching Embeddings by Relation, Not Just Similarity</title><link>https://aibussin.com/post/relate/</link><pubDate>Wed, 05 Aug 2026 23:24:45 +0100</pubDate><guid>https://aibussin.com/post/relate/</guid><description>&lt;p&gt;Embeddings are everywhere in modern AI.&lt;/p&gt;&#10;&lt;p&gt;They power semantic search, retrieval-augmented generation, recommendations, clustering, duplicate detection, code search, memory systems, and many of the mechanisms through which an AI system decides what information is relevant.&lt;/p&gt;&#10;&lt;p&gt;Yet most systems interrogate embeddings in essentially the same way:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Take two vectors and calculate cosine similarity.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;That is useful. But it also makes a strong assumption.&lt;/p&gt;&#10;&lt;p&gt;It assumes that the information we care about is expressed directly through the default geometry of the embedding space.&lt;/p&gt;</description></item></channel></rss>