
Models From First Principles
Learn to read modern AI models by decomposing them into smaller mechanisms, then rebuilding quality scorers, value and policy heads, recurrent reasoning systems, optimizers, and preference rankers in PyTorch.
Apply this book: open the related solution paths →
A complicated model is usually a collection of smaller models.
Those smaller models are collections of blocks.
The blocks are collections of layers.
And eventually the layers become tensor operations that we can inspect, execute, test, and understand.
That is the starting point of Models From First Principles.
The book is not organized around memorizing architecture names. It is organized around a method:
Open the abstraction until every important mechanism becomes visible.
We begin with one of the smallest useful learned models in the book: two embeddings go in and one quality score comes out.
Then we add complexity only when the simpler architecture exposes a limitation.
One score becomes several related quantities.
Anonymous output branches become explicit component models.
One-pass computation becomes recurrent state.
Hierarchical recurrence is compressed into a smaller recursive model.
That recursive model is opened to expose the residual blocks, attention, sparse representation, and output heads inside it.
Then we move underneath architecture itself and examine optimization.
Finally, we ask what changes when an absolute score is difficult to specify but a relative preference is much easier to provide.
Then we zoom out. Modern reasoning systems reveal that architecture is only one place capability can change. The same parameters can be given more recurrent steps, more sampled trajectories, stronger selection or verification, different post-training, or distilled behaviour. The final technical chapter asks where the extra computation actually lives.
The recurring question throughout the book is:
What is this model actually made of, what computation does each part perform, and what additional capability earns the additional complexity?
The method
The basic method is recursive decomposition:
model
↓
submodels
↓
blocks
↓
layers
↓
tensor operations
↓
mathematics
But understanding a model is not only about moving downward.
We also rebuild upward:
tensor operations
↓
small blocks
↓
learned components
↓
model architecture
↓
training behaviour
↓
system-level decision
That movement in both directions matters.
If we only look at the top-level architecture, names such as quality model, policy head, hierarchical reasoning model, attention block, or preference ranker can hide the mechanism.
If we only look at individual tensor operations, we can lose sight of why the pieces were assembled in the first place.
The aim of this book is to keep both views connected.
What this book is designed to teach
By working through the models, you will learn how to read an unfamiliar architecture as a set of explicit computational choices.
You will learn how to:
- decompose a model recursively into smaller models, blocks, layers, and tensor operations;
- build a neural quality model from a context embedding and a candidate embedding;
- understand what a learned scalar score can represent and where a single score becomes insufficient;
- separate Q, V, policy, and advantage into related but distinct decision quantities;
- turn output heads into explicit, independently inspectable model components;
- reason about architectures composed from several learned submodels;
- understand recurrent computation as repeated reuse of the same learned mechanism;
- work with latent state that changes over several computational steps;
- compare fast and slow recurrent state in a hierarchical architecture;
- simplify a hierarchical recurrent design into a smaller recursive model;
- open a recursive model and inspect residual blocks, normalization, attention, sparse autoencoders, and output heads separately;
- understand why architecture and computation depth are different things;
- follow the path from loss to gradients to optimizer state to parameter updates;
- build an optimizer from gradient statistics rather than treating
optimizer.step()as magic; - understand why relative preference labels can be easier to obtain than absolute quality targets;
- build pairwise preference models that learn which candidate should be preferred;
- distinguish model architecture from inference strategy, selection, verification, and post-training;
- understand self-consistency, Best-of-N, test-time scaling, reinforcement learning, and distillation as different places to spend computation; and
- choose a model by the decision it needs to support rather than by how sophisticated the architecture sounds.
The goal is not to leave you with one preferred architecture.
It is to leave you with a way to read, decompose, build, compare, and question models.
One progression, built model by model
The current book develops that method through the following sequence:
01 The Model Inside the Model
Establish the first-principles method: a model becomes understandable
when we recursively open its abstractions and inspect the mechanisms inside.
02 MR.Q — Building a Neural Quality Model From Two Embeddings
Start with a compact learned scorer:
context + candidate → representation → quality value.
03 EBT — From One Score to Q, V, Policy and Advantage
Ask what happens when one scalar is not enough.
Use one representation to support several related decision quantities.
04 SICQL — Building a Model From Q, V and Policy Networks
Make those heads first-class components.
Treat a larger model as a composition of smaller specialized models.
05 HRM — Hierarchical Reasoning With Fast and Slow Recurrent State
Move beyond one-pass computation.
Introduce repeated state updates operating at different timescales.
06 Tiny — Recursive Reasoning With a Small Neural Network
Keep repeated computation while reducing the architecture.
Reuse a small learned core to iteratively update one latent state.
07 Inside Tiny — Residual Blocks, Attention and Sparse Autoencoders
Open the recursive model itself.
Trace its behaviour into recognizable internal mechanisms.
08 PACS — Building an Optimizer From Gradient Statistics
Move underneath architecture into learning dynamics.
Build parameter updates from gradient statistics and optimizer state.
09 Preference Rankers — Learning Which Answer Is Better
Replace difficult absolute targets with relative comparisons.
Learn from pairs: which candidate should be preferred?
10 Reasoning Is More Than Architecture — Where Extra Computation Lives
Zoom out from individual models to the full reasoning system.
Separate architecture, recurrent compute, sampling, selection,
verification, post-training, and distillation.
Final Which Model Should You Use?
Compare the model family as a decision path rather than a leaderboard.
Choose the simplest mechanism that solves the problem you actually have.
The sequence is not intended to be a ladder from weak models to strong models.
It is a sequence of architectural pressures.
Need one learned score?
↓
MR.Q
One score no longer enough?
↓
EBT
Need explicit specialized components?
↓
SICQL
Need computation to continue over time?
↓
HRM
Need recurrence with less machinery?
↓
Tiny
Need to understand what Tiny contains?
↓
residual + attention + sparse components
Is optimization now the important mechanism?
↓
PACS
Are absolute targets hard but comparisons easy?
↓
Preference Ranker
Need to know whether more capability requires
new parameters, more inference, better selection,
or different training?
↓
Reasoning Is More Than Architecture
Each step should earn its place.
From one score to a model family
The first architectural move is deliberately small.
MR.Q asks for one learned quantity:
context embedding
+
candidate embedding
↓
encoder
↓
representation
↓
predictor
↓
quality
That gives us something concrete enough to inspect end to end.
But a single scalar cannot answer every useful decision question.
EBT therefore asks several questions of a shared representation:
How good is this state-action pair? → Q
How good is the state generally? → V
What action should be preferred? → Policy
How much better is Q than V? → Advantage
SICQL takes the next step.
Instead of treating these as anonymous branches inside one class, it makes them explicit components.
That leads to one of the book’s central architectural ideas:
A model can itself be built from models.
Once the components are explicit, they become easier to inspect, test, replace, compare, and reason about independently.
From one pass to repeated computation
The early models have a familiar shape:
input
↓
encoder
↓
representation
↓
heads
↓
output
HRM changes the shape of the computation.
The representation is no longer produced once and immediately consumed.
State evolves.
A low-level state can update several times before a higher-level state changes.
The architecture introduces a distinction between parameters and computation: the same learned mechanism can be reused over several steps without allocating a new parameter set for every step.
Tiny then asks an important engineering question:
How much of the useful repeated-computation idea survives if we remove some of the machinery?
Its core loop is deliberately compact:
context + candidate + latent state
↓
projection
↓
reusable core
↓
proposed update
↓
updated latent state
↓
repeat
The point is not that recurrence automatically produces better reasoning.
The point is that recurrence gives us another architectural resource: computation depth through reuse.
Open the model again
Once Tiny exists, the first-principles method starts again.
We open it.
Tiny
├── state fusion
├── residual block
├── optional attention
├── sparse autoencoder
└── output heads
Each apparently sophisticated subsystem becomes another object that can be decomposed.
A residual block becomes normalization, linear transformations, activation, and addition.
Attention becomes projections, similarity, weighting, and recombination.
A sparse autoencoder becomes an encoder, a constrained latent representation, and a decoder.
The abstraction is useful.
But the abstraction is not the explanation.
That distinction is one of the most durable skills in the book.
Architecture is only half the model
A model definition does not learn by itself.
Eventually every architecture reaches:
loss
↓
backward()
↓
gradients
↓
optimizer state
↓
parameter update
PACS moves the book underneath the forward pass.
Instead of treating the optimizer as infrastructure outside the model, we inspect its state and update rule directly.
Gradient statistics accumulate.
Those statistics influence scale and direction.
Parameters move.
The same first-principles question applies:
What state is being kept, what transformation is being performed, and why should that transformation help learning?
Once you can answer those questions, a custom optimizer becomes another mechanism rather than another black box.
Sometimes comparison is easier than scoring
The preference-ranking chapter introduces a different kind of supervision.
Absolute labels can be awkward.
If someone asks:
How good is this answer from 0.0 to 1.0?
the target may feel arbitrary.
But if the question becomes:
Which answer would you rather keep?
the supervision can become much more natural.
That changes the learning problem:
candidate A
candidate B
↓
comparison model
↓
preference
This connects the earlier quality-model idea to pairwise ranking and preference learning.
The model no longer needs to discover an externally meaningful absolute number.
It needs to learn a boundary between alternatives.
Reasoning is more than architecture
By this point the book has changed architecture, recurrence, optimization, and supervision. The next step is to separate those mechanisms from the inference system around them.
A fixed model can still be used in very different ways:
single sample
more recurrent steps
multiple sampled trajectories
majority vote
Best-of-N
verifier-guided selection
partial-trajectory search
And training can change the same model family again through preference objectives, reinforcement learning, or distillation.
This produces a broader first-principles question:
Where does the extra computation live?
The answer may be in parameters, recurrent state, runtime sampling, selection, verification, or training. These mechanisms should be measured separately rather than hidden inside the label reasoning model.
How to read an unfamiliar model
After working through the book, an unfamiliar architecture should become less intimidating.
Instead of beginning with the model’s name, begin with questions such as:
What are the inputs?
What tensors represent them?
What state persists across computation?
What is projected or encoded?
Which blocks are reused?
Which components have independent parameters?
Where does information branch?
What does each head predict?
What is the loss?
Where do the targets come from?
What gradients are produced?
What state does the optimizer maintain?
How are the parameters updated?
Is the model predicting an absolute quantity
or comparing alternatives?
Which component is actually necessary for the decision?
These questions are more durable than the name of any individual architecture.
They also make model debugging more precise.
Instead of saying:
The model does not work.
you can ask whether the problem is in the representation, the target, the head, the recurrent update, the state schedule, the sparse bottleneck, the optimizer, or the supervision itself.
The central engineering principle
A more complicated model is not automatically a better model.
Every additional component introduces another assumption:
another head
another state
another recurrence
another attention block
another latent code
another optimizer statistic
another training objective
Those assumptions cost complexity, compute, debugging effort, and experimental uncertainty.
The final comparison therefore returns to a deliberately conservative rule:
Use the simplest model that solves the decision you actually have.
MR.Q is not merely an early step on the way to HRM.
Tiny is not automatically better than SICQL.
A custom optimizer is not inherently better than a standard optimizer.
A preference ranker is not the answer when reliable scalar targets already exist.
The architecture should follow the problem.
Who this book is for
This book is for readers who are comfortable with basic Python and want to become more confident reading and constructing neural models in PyTorch.
It is particularly useful if model code still sometimes feels like a wall of named abstractions.
You do not need to memorize every architecture in advance.
The book works by repeatedly taking something that initially looks large and opening it until the important mechanism is visible.
It is especially relevant to:
- developers building learned scoring or ranking components;
- engineers working with embeddings and model heads;
- readers interested in value, policy, and quality models;
- people exploring recurrent or recursive reasoning architectures;
- readers who want to distinguish model architecture from test-time scaling and post-training;
- developers who want to understand attention and sparse representations inside larger systems;
- anyone who has used optimizers without yet implementing one; and
- builders working with preference data or comparative evaluation.
What this book does not try to cover
This is not an encyclopedia of modern machine learning.
It does not attempt to catalogue every transformer variant, diffusion architecture, multimodal model, reinforcement-learning algorithm, or optimization method.
It is also not a claim that the architectures in the book form a universal progression.
They are worked examples chosen to expose different model-design questions:
scoring
multi-head prediction
component composition
recurrent state
hierarchical computation
recursive computation
attention
sparse representation
optimization
preference learning
inference-time scaling
selection and verification
post-training and distillation
model selection
The purpose is not to memorize the examples.
The purpose is to acquire the method.
The promise
By the end of Models From First Principles, you should be able to look at an unfamiliar model and begin taking it apart.
You should be able to trace a named architecture into components, the components into layers, the layers into tensor operations, and then rebuild the larger purpose from those operations.
You should also be able to ask the question that matters before adding another mechanism:
What limitation in the simpler model requires this additional complexity?
Once you can answer that question, model architecture becomes less about names and more about decisions.
The mystery gets smaller.
The mechanism becomes visible.
And the model inside the model becomes something you can actually reason about.