<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Models From First Principles on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/books/models-from-first-principles/</link><description>Recent content in Models From First Principles on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Sat, 05 Sep 2026 18:54:00 +0100</lastBuildDate><atom:link href="https://aibussin.com/books/models-from-first-principles/index.xml" rel="self" type="application/rss+xml"/><item><title>From Vectors to Symbols — The Binding Problem Inside Neural Networks</title><link>https://aibussin.com/books/models-from-first-principles/11-chapter/</link><pubDate>Sat, 05 Sep 2026 18:54:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/11-chapter/</guid><description>&lt;h1 id="from-vectors-to-symbols--the-binding-problem-inside-neural-networks"&gt;From Vectors to Symbols — The Binding Problem Inside Neural Networks&lt;/h1&gt;&#10;&lt;p&gt;The previous chapter changed our scale of analysis.&lt;/p&gt;&#10;&lt;p&gt;We stopped asking only what architecture a model had and started asking where the computation lived.&lt;/p&gt;&#10;&lt;p&gt;Was capability coming from new parameters?&lt;/p&gt;&#10;&lt;p&gt;More recurrent steps?&lt;/p&gt;&#10;&lt;p&gt;More sampled trajectories?&lt;/p&gt;&#10;&lt;p&gt;A selector?&lt;/p&gt;&#10;&lt;p&gt;A verifier?&lt;/p&gt;&#10;&lt;p&gt;Different post-training?&lt;/p&gt;&#10;&lt;p&gt;That wider view exposed a second question.&lt;/p&gt;&#10;&lt;p&gt;Suppose the computation works.&lt;/p&gt;&#10;&lt;p&gt;Suppose a neural network solves arithmetic, manipulates code, follows grammatical structure, or reasons over logical relations.&lt;/p&gt;</description></item><item><title>Reasoning Is More Than Architecture — Where Extra Computation Lives</title><link>https://aibussin.com/books/models-from-first-principles/10-chapter/</link><pubDate>Tue, 01 Sep 2026 10:41:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/10-chapter/</guid><description>&lt;h1 id="reasoning-is-more-than-architecture--where-extra-computation-lives"&gt;Reasoning Is More Than Architecture — Where Extra Computation Lives&lt;/h1&gt;&#10;&lt;p&gt;So far in &lt;strong&gt;Models From First Principles&lt;/strong&gt;, we have changed several different things and called all of them model design.&lt;/p&gt;&#10;&lt;p&gt;We changed what a model predicts.&lt;/p&gt;&#10;&lt;p&gt;MR.Q produced one learned quality score.&lt;/p&gt;&#10;&lt;p&gt;EBT added Q, V, policy, and advantage.&lt;/p&gt;&#10;&lt;p&gt;SICQL turned those outputs into explicit model components.&lt;/p&gt;&#10;&lt;p&gt;Then we changed how computation unfolds.&lt;/p&gt;&#10;&lt;p&gt;HRM introduced recurrent state operating at different timescales.&lt;/p&gt;</description></item><item><title>Preference Rankers — Learning Which Answer Is Better</title><link>https://aibussin.com/books/models-from-first-principles/09-chapter/</link><pubDate>Sat, 22 Aug 2026 01:18:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/09-chapter/</guid><description>&lt;h1 id="preference-rankers--learning-which-answer-is-better"&gt;Preference Rankers — Learning Which Answer Is Better&lt;/h1&gt;&#10;&lt;p&gt;So far in &lt;strong&gt;Models From First Principles&lt;/strong&gt;, we have mostly trained models by telling them what the answer should be.&lt;/p&gt;&#10;&lt;p&gt;MR.Q took a context and a response and produced a number:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;context + response&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; 0.82&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That number might mean quality.&lt;/p&gt;&#10;&lt;p&gt;Or usefulness.&lt;/p&gt;&#10;&lt;p&gt;Or reward.&lt;/p&gt;&#10;&lt;p&gt;But there is an awkward question hiding underneath that architecture:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Where did the 0.82 come from?&lt;/p&gt;</description></item><item><title>Which Model Should You Use? MR.Q, EBT, SICQL, HRM, Tiny and PACS Compared</title><link>https://aibussin.com/books/models-from-first-principles/15-chapter/</link><pubDate>Sat, 08 Aug 2026 15:11:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/15-chapter/</guid><description>&lt;h1 id="which-model-should-you-use-mrq-ebt-sicql-hrm-tiny-and-pacs-compared"&gt;Which Model Should You Use? MR.Q, EBT, SICQL, HRM, Tiny and PACS Compared&lt;/h1&gt;&#10;&lt;p&gt;This is the final post in &lt;strong&gt;Models From First Principles&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;The earlier posts asked a sequence of architectural questions:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;how do we score a context-response pair?&lt;/li&gt;&#10;&lt;li&gt;when is one scalar no longer enough?&lt;/li&gt;&#10;&lt;li&gt;when should Q, V and policy become explicit components?&lt;/li&gt;&#10;&lt;li&gt;when is one forward pass insufficient?&lt;/li&gt;&#10;&lt;li&gt;when does recurrence help?&lt;/li&gt;&#10;&lt;li&gt;when does hierarchy help?&lt;/li&gt;&#10;&lt;li&gt;when is a smaller recursive model a better trade-off?&lt;/li&gt;&#10;&lt;li&gt;when should attention or a sparse autoencoder be added?&lt;/li&gt;&#10;&lt;li&gt;when should we change the optimizer rather than the model?&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;This post asks the question that matters when building a real system:&lt;/p&gt;</description></item><item><title>PACS — Building an Optimizer From Gradient Statistics</title><link>https://aibussin.com/books/models-from-first-principles/08-chapter/</link><pubDate>Sat, 08 Aug 2026 15:05:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/08-chapter/</guid><description>&lt;h1 id="pacs--building-an-optimizer-from-gradient-statistics"&gt;PACS — Building an Optimizer From Gradient Statistics&lt;/h1&gt;&#10;&lt;p&gt;So far in &lt;strong&gt;Models From First Principles&lt;/strong&gt;, every post has asked some version of the same question:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;What should the model compute?&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;MR.Q gave us a scalar quality estimate.&lt;/p&gt;&#10;&lt;p&gt;EBT split one shared representation into Q, V and Policy.&lt;/p&gt;&#10;&lt;p&gt;SICQL made those heads explicit, replaceable components.&lt;/p&gt;&#10;&lt;p&gt;HRM introduced repeated computation over fast and slow latent states.&lt;/p&gt;&#10;&lt;p&gt;Tiny compressed iterative refinement into one recursive latent state.&lt;/p&gt;</description></item><item><title>Inside Tiny — Residual Blocks, Attention and Sparse Autoencoders</title><link>https://aibussin.com/books/models-from-first-principles/07-chapter/</link><pubDate>Sat, 08 Aug 2026 15:00:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/07-chapter/</guid><description>&lt;h1 id="inside-tiny-residual-blocks-attention-and-sparse-autoencoders"&gt;Inside Tiny: Residual Blocks, Attention and Sparse Autoencoders&lt;/h1&gt;&#10;&lt;p&gt;In the previous post, we built a compact recursive model around one idea:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;context + candidate + latent state&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; projection&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; reusable core&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; proposed update&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; z ← z + α · update&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; repeat&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That architecture looked more sophisticated than MR.Q, EBT or SICQL because it introduced recurrence.&lt;/p&gt;&#10;&lt;p&gt;But the central idea of this series is that a model stops looking mysterious when we keep opening it.&lt;/p&gt;</description></item><item><title>Tiny — Recursive Reasoning With a Small Neural Network</title><link>https://aibussin.com/books/models-from-first-principles/06-chapter/</link><pubDate>Sat, 08 Aug 2026 14:55:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/06-chapter/</guid><description>&lt;h1 id="tiny--recursive-reasoning-with-a-small-neural-network"&gt;Tiny — Recursive Reasoning With a Small Neural Network&lt;/h1&gt;&#10;&lt;p&gt;The previous post introduced a much more ambitious architecture.&lt;/p&gt;&#10;&lt;p&gt;Instead of taking one representation and predicting from it once, the &lt;strong&gt;Hierarchical Reasoning Model&lt;/strong&gt; repeatedly updated two latent states:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;input&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;low-level state&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓ ↓ ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;high-level state&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;repeat&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That gave us something genuinely new:&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;computation could continue without adding a new set of parameters for every step.&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>HRM — Hierarchical Reasoning With Fast and Slow Recurrent State</title><link>https://aibussin.com/books/models-from-first-principles/05-chapter/</link><pubDate>Sat, 08 Aug 2026 14:48:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/05-chapter/</guid><description>&lt;h1 id="hrm--hierarchical-reasoning-with-fast-and-slow-recurrent-state"&gt;HRM — Hierarchical Reasoning With Fast and Slow Recurrent State&lt;/h1&gt;&#10;&lt;p&gt;The previous models in this series were mostly &lt;strong&gt;one-pass models&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;MR.Q took two embeddings and produced one score.&lt;/p&gt;&#10;&lt;p&gt;EBT kept the same basic structure but added several heads.&lt;/p&gt;&#10;&lt;p&gt;SICQL made those heads explicit components.&lt;/p&gt;&#10;&lt;p&gt;The architecture grew, but the shape of the computation was still familiar:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;input&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;encoder&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;representation&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;heads&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;outputs&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;HRM changes the question.&lt;/p&gt;</description></item><item><title>SICQL — Building a Model From Q, V and Policy Networks</title><link>https://aibussin.com/books/models-from-first-principles/04-chapter/</link><pubDate>Sat, 08 Aug 2026 14:44:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/04-chapter/</guid><description>&lt;h1 id="sicql--building-a-model-from-q-v-and-policy-networks"&gt;SICQL — Building a Model From Q, V and Policy Networks&lt;/h1&gt;&#10;&lt;p&gt;In the previous post we took the MR.Q idea and expanded it into something richer.&lt;/p&gt;&#10;&lt;p&gt;Instead of asking one question of a shared representation, EBT asked several:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;How good is this state-action pair? -&amp;gt; Q&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;How good is the state more generally? -&amp;gt; V&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;What action should be preferred? -&amp;gt; Policy&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;How much better is Q than V? -&amp;gt; Advantage&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That already gave us a more expressive system.&lt;/p&gt;</description></item><item><title>EBT — From One Score to Q, V, Policy and Advantage</title><link>https://aibussin.com/books/models-from-first-principles/03-chapter/</link><pubDate>Sat, 08 Aug 2026 14:39:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/03-chapter/</guid><description>&lt;h1 id="ebt--from-one-score-to-q-v-policy-and-advantage"&gt;EBT — From One Score to Q, V, Policy and Advantage&lt;/h1&gt;&#10;&lt;p&gt;In the previous post we built MR.Q: a small model that takes a context embedding and a response embedding, combines them, and predicts one scalar.&lt;/p&gt;&#10;&lt;p&gt;That architecture is useful because it is brutally simple:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;context embedding&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; +&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;response embedding&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; encoder&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; representation z&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; predictor&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; Q value&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;But one scalar eventually becomes restrictive.&lt;/p&gt;</description></item><item><title>MR.Q — Building a Neural Quality Model From Two Embeddings</title><link>https://aibussin.com/books/models-from-first-principles/02-chapter/</link><pubDate>Sat, 08 Aug 2026 14:33:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/02-chapter/</guid><description>&lt;h1 id="mrq--building-a-neural-quality-model-from-two-embeddings"&gt;MR.Q — Building a Neural Quality Model From Two Embeddings&lt;/h1&gt;&#10;&lt;p&gt;In the previous post, we established the core idea behind this series:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;A complicated model becomes understandable when you recursively decompose it into smaller models, blocks, layers and tensor operations.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;Now we build the first real model.&lt;/p&gt;&#10;&lt;p&gt;Not a transformer.&lt;/p&gt;&#10;&lt;p&gt;Not a giant language model.&lt;/p&gt;&#10;&lt;p&gt;Not an agent.&lt;/p&gt;&#10;&lt;p&gt;A scorer.&lt;/p&gt;&#10;&lt;p&gt;We will take two embeddings:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;context embedding&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;response embedding&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;combine them, encode the relationship between them, and predict one scalar:&lt;/p&gt;</description></item><item><title>The Model Inside the Model</title><link>https://aibussin.com/books/models-from-first-principles/01-chapter/</link><pubDate>Sat, 08 Aug 2026 14:27:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/01-chapter/</guid><description>&lt;h1 id="the-model-inside-the-model"&gt;The Model Inside the Model&lt;/h1&gt;&#10;&lt;p&gt;This is the first post in &lt;strong&gt;Models From First Principles&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;The previous &lt;strong&gt;PyTorch From First Principles&lt;/strong&gt; series worked from the bottom up.&lt;/p&gt;&#10;&lt;p&gt;We started with tensors.&lt;/p&gt;&#10;&lt;p&gt;Then gradients.&lt;/p&gt;&#10;&lt;p&gt;Then &lt;code&gt;nn.Module&lt;/code&gt;.&lt;/p&gt;&#10;&lt;p&gt;Then data pipelines, convolution, attention, debugging, performance and finally a small GPT-style language model built from scratch.&lt;/p&gt;&#10;&lt;p&gt;That series answered:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;What are the pieces?&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;This series asks a different question:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;What happens when we start composing those pieces into increasingly sophisticated models?&lt;/p&gt;</description></item><item><title/><link>https://aibussin.com/books/models-from-first-principles/_work/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/models-from-first-principles/_work/</guid><description>&lt;h1 id="models-from-first-principles--working-ledger"&gt;Models From First Principles — Working Ledger&lt;/h1&gt;&#10;&lt;h2 id="book-intention"&gt;Book intention&lt;/h2&gt;&#10;&lt;p&gt;Understand modern AI models by opening the abstractions and building the mechanisms in layers, while keeping the mathematics, architecture, optimization, inference, training, and working code connected.&lt;/p&gt;&#10;&lt;p&gt;The review question is not merely whether each chapter is well written. It is whether each chapter does the job it is supposed to do &lt;strong&gt;at this point in the sequence&lt;/strong&gt; and whether the sequence leaves a useful model or mechanism missing.&lt;/p&gt;</description></item></channel></rss>