<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Attention on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/attention/</link><description>Recent content in Attention on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Fri, 25 Sep 2026 02:45:00 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/attention/index.xml" rel="self" type="application/rss+xml"/><item><title>Make the Window Bigger</title><link>https://aibussin.com/books/context/08-chapter/</link><pubDate>Wed, 23 Sep 2026 05:00:08 +0100</pubDate><guid>https://aibussin.com/books/context/08-chapter/</guid><description>&lt;p&gt;Seven chapters have treated the model as fixed ground and asked how the surrounding system should manage information against it. The natural objection has been waiting since Chapter 4: if context is scarce and difficult to manage, why not simply make the window much larger? A team that moves its agent onto a million-token model watches several selection problems relax at once. Histories that needed pruning now fit. Files that needed triage can all be included. The compaction schedule gets quieter. For a while, context engineering looks like a stopgap awaiting cheaper capacity.&lt;/p&gt;</description></item><item><title>Designing Your Digital Lens</title><link>https://aibussin.com/books/agent-architectures/09-chapter/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/agent-architectures/09-chapter/</guid><description>&lt;h3 id="ai-as-an-attention-interface"&gt;AI as an Attention Interface&lt;/h3&gt;&#10;&lt;p&gt;The previous chapter looked inward: reflection, routines, personal notes, and companion roles. This chapter looks outward.&lt;/p&gt;&#10;&lt;p&gt;Most of digital life is not scarce. It is excessive. Feeds update constantly. Notifications arrive from systems with different incentives. Search returns more than you can read. News, entertainment, work, advertising, research, and social pressure all arrive through the same devices.&lt;/p&gt;&#10;&lt;p&gt;A &lt;strong&gt;digital lens&lt;/strong&gt; is an agentic layer that helps filter, prioritize, summarize, or reshape that flow according to a purpose you choose.&lt;/p&gt;</description></item><item><title>Attention: Which Position Is Comparing With Which?</title><link>https://aibussin.com/books/pytorch-from-first-principles/10-chapter/</link><pubDate>Sat, 29 Aug 2026 10:00:00 +0100</pubDate><guid>https://aibussin.com/books/pytorch-from-first-principles/10-chapter/</guid><description>&lt;p&gt;Here is a tensor of sequence representations and the line that splits it into heads.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;B, T, E, Nh &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;1&lt;/span&gt;, &lt;span style="color:#ae81ff"&gt;4&lt;/span&gt;, &lt;span style="color:#ae81ff"&gt;8&lt;/span&gt;, &lt;span style="color:#ae81ff"&gt;2&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Dh &lt;span style="color:#f92672"&gt;=&lt;/span&gt; E &lt;span style="color:#f92672"&gt;//&lt;/span&gt; Nh &lt;span style="color:#75715e"&gt;# 4&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;x &lt;span style="color:#f92672"&gt;=&lt;/span&gt; torch&lt;span style="color:#f92672"&gt;.&lt;/span&gt;arange(B &lt;span style="color:#f92672"&gt;*&lt;/span&gt; T &lt;span style="color:#f92672"&gt;*&lt;/span&gt; E)&lt;span style="color:#f92672"&gt;.&lt;/span&gt;reshape(B, T, E)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;heads &lt;span style="color:#f92672"&gt;=&lt;/span&gt; x&lt;span style="color:#f92672"&gt;.&lt;/span&gt;reshape(B, Nh, T, Dh)&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;x: (1, 4, 8)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;heads: (1, 2, 4, 4)&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That is exactly the shape multi-head attention wants: batch, heads, positions, head dimension. Nothing raised. Every shape assertion passes.&lt;/p&gt;&#10;&lt;p&gt;Now the same split written the other way:&lt;/p&gt;</description></item><item><title>Too Much Information</title><link>https://aibussin.com/books/language/11-chapter/</link><pubDate>Fri, 25 Sep 2026 02:45:00 +0100</pubDate><guid>https://aibussin.com/books/language/11-chapter/</guid><description>&lt;p&gt;The Semantic Browser works. On any page you open, it chooses a representation, enriches it, types its relationships, and gates everything through preservation evidence. Now count what it never touches: the forty hours of conference video in your subscriptions, the papers published this week in your field, the podcasts queued behind the podcasts, the threads, the advisories, the second-order citations of everything you read. The browser perfects the encounter with information already in front of you. It does nothing for the information you never reach — which is nearly all of it.&lt;/p&gt;</description></item><item><title>Assembly: A Language Model You Can Interrogate</title><link>https://aibussin.com/books/pytorch-from-first-principles/15-chapter/</link><pubDate>Sun, 30 Aug 2026 10:00:00 +0100</pubDate><guid>https://aibussin.com/books/pytorch-from-first-principles/15-chapter/</guid><description>&lt;p&gt;Here are two training runs of the same small language model on the same data. The only difference is one line in the batching function.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;CORRECT y = data[i+1 : i+block+1] the next-token shift&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;MISALIGNED y = data[i : i+block] no shift&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;step correct val misaligned val&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 0 3.6695 2.6469&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 250 1.1038 0.0007&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 500 0.4730 0.0004&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 750 0.3940 0.0003&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;1000 0.3690 0.0002&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;1250 0.3581 0.0002&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;1500 0.3558 0.0001&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The misaligned run&amp;rsquo;s loss falls to &lt;code&gt;0.0001&lt;/code&gt; — three thousand times lower than the correct run&amp;rsquo;s &lt;code&gt;0.36&lt;/code&gt;. By every number on the dashboard it is the best training run you have ever seen. Nothing raised. The shapes are identical: &lt;code&gt;x&lt;/code&gt; is &lt;code&gt;[B, T]&lt;/code&gt;, &lt;code&gt;y&lt;/code&gt; is &lt;code&gt;[B, T]&lt;/code&gt;, the loss is a finite scalar with a &lt;code&gt;grad_fn&lt;/code&gt;.&lt;/p&gt;</description></item><item><title>Appendix A: PyTorch Diagnostic Field Guide</title><link>https://aibussin.com/books/pytorch-from-first-principles/appendix-a/</link><pubDate>Sun, 30 Aug 2026 12:00:00 +0100</pubDate><guid>https://aibussin.com/books/pytorch-from-first-principles/appendix-a/</guid><description>&lt;p&gt;This appendix introduces no new PyTorch mechanism.&lt;/p&gt;&#10;&lt;p&gt;Everything here was earned earlier in the book by building something small, breaking it deliberately, measuring what changed, and locating the first place where reality stopped matching the intended computation.&lt;/p&gt;&#10;&lt;p&gt;The purpose of this appendix is different.&lt;/p&gt;&#10;&lt;p&gt;When a real model is failing, you usually do not need another explanation of autograd, broadcasting, attention or CUDA. You need to answer a narrower question:&lt;/p&gt;</description></item><item><title>Inside Tiny — Residual Blocks, Attention and Sparse Autoencoders</title><link>https://aibussin.com/books/models-from-first-principles/07-chapter/</link><pubDate>Sat, 08 Aug 2026 15:00:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/07-chapter/</guid><description>&lt;h1 id="inside-tiny-residual-blocks-attention-and-sparse-autoencoders"&gt;Inside Tiny: Residual Blocks, Attention and Sparse Autoencoders&lt;/h1&gt;&#10;&lt;p&gt;In the previous post, we built a compact recursive model around one idea:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;context + candidate + latent state&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; projection&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; reusable core&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; proposed update&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; z ← z + α · update&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; repeat&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That architecture looked more sophisticated than MR.Q, EBT or SICQL because it introduced recurrence.&lt;/p&gt;&#10;&lt;p&gt;But the central idea of this series is that a model stops looking mysterious when we keep opening it.&lt;/p&gt;</description></item><item><title>PyTorch Attention Shapes: Q, K, V, Multi-Head Attention Masks and Transformer Dimension Errors</title><link>https://aibussin.com/post/pytorch-zero-to-hero-07/</link><pubDate>Sat, 08 Aug 2026 13:30:00 +0100</pubDate><guid>https://aibussin.com/post/pytorch-zero-to-hero-07/</guid><description>&lt;h2 id="pytorch-zero-to-hero--step-07"&gt;PyTorch: Zero to Hero — Step 07&lt;/h2&gt;&#10;&lt;p&gt;Attention code is where tensor-shape mistakes stop being annoying and start becoming architectural.&lt;/p&gt;&#10;&lt;p&gt;A CNN usually makes its dimensional assumptions fairly obvious. Attention does not.&lt;/p&gt;&#10;&lt;p&gt;A tensor that starts as:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;(batch, sequence, embedding)&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;is projected into Q, K and V, split into heads, transposed, multiplied, masked, normalized, multiplied again, transposed again, concatenated and projected back to the embedding dimension.&lt;/p&gt;&#10;&lt;p&gt;A single bad &lt;code&gt;view&lt;/code&gt;, &lt;code&gt;transpose&lt;/code&gt;, mask shape or head calculation can produce anything from an immediate runtime error to a model that trains while attending to the wrong tokens.&lt;/p&gt;</description></item></channel></rss>