Assembly: A Language Model You Can Interrogate
Assemble a small GPT-style language model and interrogate it with the book's full diagnostic stack—tokenizer and target contracts, attention interventions, parameter accounting, learning and performance ledgers, compilation, generation, and paired evaluation—so the finished model remains inspectable rather than magical.
Here are two training runs of the same small language model on the same data. The only difference is one line in the batching function.
CORRECT y = data[i+1 : i+block+1] the next-token shift
MISALIGNED y = data[i : i+block] no shift
step correct val misaligned val
0 3.6695 2.6469
250 1.1038 0.0007
500 0.4730 0.0004
750 0.3940 0.0003
1000 0.3690 0.0002
1250 0.3581 0.0002
1500 0.3558 0.0001
The misaligned run’s loss falls to 0.0001 — three thousand times lower than the correct run’s 0.36. By every number on the dashboard it is the best training run you have ever seen. Nothing raised. The shapes are identical: x is [B, T], y is [B, T], the loss is a finite scalar with a grad_fn.
Then you generate from each:
correct: 'The planet measures the ancient planet because measures it near the water.
A sudden market builds a hidden market again. When that planet counts slowly, '
misaligned: 'The '
The misaligned model learned the identity task. With y[:, t] == x[:, t], the answer is already present at the same position: the residual stream contains the embedding of token t, and causal self-attention is also allowed to attend to that position itself. The network therefore has an extremely cheap path from the visible token to the target token. The loss rewarded that wrong task the whole way down. What catches it is not a better loss curve — it is one line, checked before training starts:
assert torch.equal(x[:, 1:], y[:, :-1]) # does the target still describe the input?
That is Chapter 11’s first link, and it is the entire chapter in miniature.
The loss can be actively misleading. A metric that improves is not evidence that the model is learning the thing you meant — and a model large enough to hide its mistakes needs every instrument you have.
Where we are
Fourteen chapters, fourteen instruments. Chapter 2 asked which tensor first became wrong rather than first became illegal. Chapter 3 asked where the gradient path stops existing. Chapter 5 asked what PyTorch thinks belongs to your model. Chapter 6 asked where the training loop is actually waiting. Chapter 7 asked what the model actually sees. Chapter 8 derived convolutional geometry into a shape ledger. Chapter 9 made feature space observable. Chapter 10 asked which position is comparing with which. Chapter 11 proved the learning chain one link at a time. Chapter 12 localized performance to a phase and an operator. Chapter 13 found the first compiler assumption that stopped holding. Chapter 14 asked whether a difference between two runs is real.
Each of those was earned from one constructed failure against one small system.
This chapter builds the first system too large to hold in your head — a GPT-style language model, about eight hundred thousand parameters, forty-odd tensors, a forward pass that goes text → tokens → embeddings → attention → residual blocks → logits → loss, and a second execution mode for generation. No single instrument certifies it. But every instrument the book built is now a question you ask of it:
tokenizer + batching does the target still describe the input? Chapter 11
embeddings what does PyTorch think belongs to the model? Chapter 5
attention can a future token change an earlier output? Chapter 10
block assembly where does the shape ledger first diverge? Chapter 8
first training steps which link in the learning chain is broken? Chapter 11
throughput and memory which phase, at which sequence length? Chapter 12
compilation did it help, and did correctness hold? Chapter 13
comparison is this run really different from that one? Chapter 14
The reader who finishes this chapter should be able to say which instrument answers which question about a transformer, and reach for the right one without being told. That is the book’s real deliverable. The model is the demonstration.
The environment
Every number, trace and generated sample in this chapter came from running the committed code under experiments/ch15/:
Python 3.11.4
PyTorch 2.6.0+cu118
GPU NVIDIA GeForce RTX 2060, compute capability 7.5
The model is deliberately small: n_layer=4, n_head=4, n_embd=128, block_size=64, character-level vocabulary. It has 804,224 parameters and trains to convergence in under a minute on this GPU; the whole experiment suite, including the seed sweeps, is about twenty minutes of compute. A larger model would not teach anything the instruments do not already show at this size.
The data is a self-contained synthetic corpus — 200,000 characters of English-like sentences generated from a small fixed word list and five templates:
That garden repeats that small garden because repeats it by the door. Near the water,
that hidden shadow carries that shadow. When a river holds again, a golden river holds.
There is no download and no copyrighted text: the corpus is reproducible from a seed, and it has real learnable structure — word boundaries, template grammar, common bigrams, punctuation — so a trained model produces recognisable (if nonsensical) output, which is what makes the generation experiments meaningful. experiments/ch15/corpus.py builds it.
The pipeline
text
↓ tokenizer characters ↔ integer IDs
token IDs [B, T]
↓ batching x = window; y = window shifted by one
↓ token + position embedding
[B, T, C]
↓ transformer blocks pre-norm attention + MLP, residual, ×n_layer
↓ final LayerNorm
↓ lm_head [B, T, vocab] logits
↓ cross entropy vs the shifted targets
loss → backward → AdamW → repeat
↓
generation feed the model its own output, one token at a time
Two boundaries in that pipeline are genuinely new and get most of this chapter’s attention: text → token IDs, and training → generation. Everything between them is composition of parts the book already built.
Data: the tokenizer
A character tokenizer is a pair of dictionaries. What matters is that it round-trips and that its IDs are in range.
class CharTokenizer:
def __init__(self, text):
self.chars = sorted(set(text))
self.stoi = {c: i for i, c in enumerate(self.chars)}
self.itos = {i: c for c, i in self.stoi.items()}
self.vocab_size = len(self.chars)
def encode(self, s): return [self.stoi[c] for c in s]
def decode(self, ids): return "".join(self.itos[int(i)] for i in ids)
round-trip decode(encode(s)) == s True for the full corpus
full corpus tensor: dtype int64 min 0 max 36 vocab 37
The round-trip check is Chapter 7’s contract idea at the input boundary: prove the transform is lossless before trusting anything downstream. The dtype check matters because token IDs are indices, not continuous values. nn.Embedding expects an integer index tensor; passing floating-point IDs is a boundary-contract error and should be diagnosed there, before the model’s representation is discussed.
The failure the tokenizer causes when it is wrong is an out-of-range ID reaching the embedding table:
emb = nn.Embedding(vocab_size, 8)
emb(torch.tensor([[0, 1, vocab_size]])) # one past the end
IndexError: index out of range in self
That is the exact error people paste into a search box. The mechanism is a contract mismatch between the IDs produced by the data/tokenizer side and the rows the embedding table actually owns: a tokenizer fit on a different corpus, a stale checkpoint with a smaller vocabulary, or a special-token ID the model was never sized for can all create it. A tokenizer fit on a 2,000-character sample of this corpus, for instance, is missing the character S and raises KeyError: 'S' the moment it meets a capital-S sentence in the full text. Verify the tokenizer round trip, the ID range, and the model’s vocabulary size at the boundary, and this class of failure never reaches training.
Data: batching and the alignment contract
def get_batch(source, batch_size, block_size):
starts = torch.randint(0, len(source) - block_size - 1, (batch_size,))
x = torch.stack([source[i : i + block_size] for i in starts])
y = torch.stack([source[i + 1 : i + block_size + 1] for i in starts])
return x, y
Every position of x predicts the next token, so y is x shifted left by one. The contract:
assert torch.equal(x[:, 1:], y[:, :-1])
This is the check from the opening. It is worth restating why it is the single most valuable line in the build: the misalignment failure is invisible in the loss (the loss improves), invisible in the shapes (both [B, T]), invisible in the gradients (they are finite and flow everywhere), and it produces a model that scores brilliantly and generates nothing. The only place it is visible is here, in the relationship between the input and the target — Chapter 11’s TASK boundary, the first thing to prove and the first thing people skip.
The model: embeddings and positions
self.token_embedding = nn.Embedding(vocab_size, n_embd) # [V, C]
self.position_embedding = nn.Embedding(block_size, n_embd) # [block, C]
...
x = self.drop(self.token_embedding(idx) + self.position_embedding(pos))
Content-based self-attention without a positional signal or order-dependent mask is permutation-equivariant. This model’s causal mask already introduces an ordering constraint — earlier positions cannot see later ones — but it does not supply the explicit learned position identity that the model uses here. That comes from the position embedding. idx is [B, T]; token_embedding(idx) is [B, T, C]; position_embedding(arange(T)) is [T, C] and broadcasts across the batch. This is Chapter 2’s broadcasting, used deliberately.
Both embedding tables are nn.Module attributes, so they register — they appear in the model’s registered parameter/state structure and, once the optimizer is built over model.parameters(), its trainable parameters can be owned by the optimizer. Chapter 5’s question — does PyTorch think this belongs to the model? — is answered here by construction. Weight tying later makes object identity important again: sharing two names is not the same thing as owning two independent parameters, and tying after optimizer construction would create exactly the kind of stale-parameter mismatch Chapter 11 taught us to audit.
The model: attention
Chapter 10 built attention carefully: the head split as a partition of the feature axis, the [B, H, T, T] score matrix, the softmax over keys, the causal mask, and — crucially — the four levels at which “correct” can mean four different things (shape, axis semantics, numerical invariants, behavior). This chapter does not re-derive any of that. It packages it:
def split_heads(x, n_head):
B, T, C = x.shape
return x.view(B, T, n_head, C // n_head).transpose(1, 2) # [B, H, T, D]
def merge_heads(x):
B, H, T, D = x.shape
return x.transpose(1, 2).contiguous().view(B, T, H * D) # [B, T, C]
class CausalSelfAttention(nn.Module):
def forward(self, x):
q, k, v = self.qkv(x).chunk(3, dim=-1)
q, k, v = (split_heads(t, self.n_head) for t in (q, k, v))
y = F.scaled_dot_product_attention(q, k, v, is_causal=True,
dropout_p=self.dropout if self.training else 0.0)
return self.resid_dropout(self.proj(merge_heads(y)))
F.scaled_dot_product_attention with is_causal=True is the whole mechanism — Chapter 10 verified it against a hand-written implementation. The two things that are still worth checking at model scale are Level 2 (does the head split preserve the data?) and Level 4 (is it actually causal?), and neither is visible in a shape.
Level 2 — the head split round trip. split_heads followed by merge_heads must return the original tensor exactly. The wrong version — q.reshape(B, H, T, D) instead of q.view(B, T, H, D).transpose(1, 2) — produces a tensor of identical shape that has cut the sequence into head-sized blocks instead of partitioning the feature axis. The shape ledger below catches it; a shape assertion does not.
Level 4 — causality by intervention. A triangular mask is a claim. The evidence is behavioral: put the model in evaluation mode, change only the tokens after position i, rerun, and require the outputs at positions 0..i to agree within a stated numerical tolerance while some suffix output actually changes.
On the correct model, after training:
max |Δ logits| at positions 0..i : 0.000e+00 the past did not move
max |Δ logits| at positions i+1.. : varies the future did
This run produced exact zero on the prefix; the stronger portable contract is that the prefix delta stays within the tolerance appropriate to the backend and dtype. The second number matters too: if the suffix delta were also zero, the test could be passing because the intervention had no observable effect.
Now set is_causal=False and train the same model:
causal non-causal
val loss (1000 steps) 0.369 0.023 16x lower -- looks extraordinary
past Δ after training 0.000e+00 1.03e+01 the past moved by 10 logits
causal sample: 'The planet measures the ancient planet because measures it by the door.'
non-causal sample: 'The h , hh h ,, ,,, , , h h h h h h h h h h h h h h'
The non-causal model can attend from position t to position t+1, whose input token is exactly the next-token target for position t (except at the boundary). The loss rewards that leakage all the way down. Inspecting the mask can suggest the bug; the future-token intervention is the behavioral evidence that proves the model actually uses information it should not have — Chapter 10’s Level 4, at full scale.
The model: MLP, block, assembly
The rest is composition. The per-token MLP (Linear → GELU → Linear → Dropout, hidden 4C) operates independently at every position and preserves [B, T, C]. The pre-norm block wraps attention and MLP in residual connections:
def forward(self, x):
x = x + self.attn(self.ln1(x))
x = x + self.mlp(self.ln2(x))
return x
Every block preserves [B, T, C], which is the invariant that makes stacking them trivial. The full model adds the embeddings, the block stack, a final LayerNorm, and the lm_head projection to vocabulary logits.
To verify the assembly, run Chapter 8’s shape ledger through the whole forward pass — a row per stage, each carrying the shape and the invariant that stage owes:
SHAPE LEDGER B=2 T=16 C=128 H=4 V=37
token_embedding (2, 16, 128) (B,T,C)=(2,16,128) ok
position_embedding (16, 128) (T,C)=(16,128) ok
tok + pos (2, 16, 128) (B,T,C) by broadcast ok
head split (2, 4, 16, 32) merge(split(q)) == q (round trip) ok
block[0..3] (2, 16, 128) (B,T,C) preserved ok
lm_head (2, 16, 37) (B,T,V)=(2,16,37) ok
flatten for CE (32, 37) (B*T,V)=(32,37) ok
first divergence: none
Now build the model with the wrong head reshape and run the same ledger:
head split (2, 4, 16, 32) merge(split(q)) == q (round trip) X <-- FIRST DIVERGENCE
block[0..3] (2, 16, 128) (B,T,C) preserved ok
lm_head (2, 16, 37) (B,T,V)=(2,16,37) ok
first divergence: head split
The broken split’s output shape is (2, 4, 16, 32) — identical to the correct one. Every downstream row still reports ok. A forward pass would not raise, and the model would train to a mediocre loss and never say why. The round-trip row is the one that sees it, because it checks the data, not the dimensions. This is Chapter 8’s first-divergence rule and Chapter 10’s Level 2, composed.
Parameter accounting
def param_count(model):
seen = set()
total = 0
for p in model.parameters():
if id(p) in seen:
continue
seen.add(id(p))
total += p.numel()
return total
THIS MODEL n_layer=4 n_head=4 n_embd=128 vocab=37 total 808,960
mlp 526,848 65.1%
attention 262,144 32.4%
embeddings (tok + pos) 12,928 1.6%
lm_head (output projection) 4,736 0.6%
layernorm 2,304 0.3%
There is a widely repeated claim that the embeddings dominate a small language model. At subword scale that is true — the same accounting at GPT-2-small scale:
n_layer=12 n_embd=768 vocab=50257 total 124,356,864
mlp 56,623,104 45.5%
embeddings (tok + pos) 39,383,808 31.7% <- the token table alone is 31%
attention 28,311,552 22.8%
— the token embedding alone is about 38.6M parameters, roughly 31% of the canonical 124.4M tied GPT-2-small count. But at character scale the vocabulary is 37, the token table is only 37 × 128 = 4,736 values, and the MLPs dominate. The important variable is the vocabulary-to-model-width/body ratio, not a universal rule that “embeddings dominate.” Knowing which regime you are in tells you how significant the next section can be.
Weight tying
The input embedding maps a token ID to a vector; the output lm_head maps a vector back to a distribution over tokens. Both are [vocab, n_embd] matrices about the same vocabulary, and a standard design shares them:
self.lm_head.weight = self.token_embedding.weight
Verify the tie took — Chapter 5’s discipline, because an assignment that runs is not proof of anything:
tied: lm_head.weight is token_embedding.weight -> True
untied: same object -> False
config params parameter tensors shares storage
tied 804,224 44 True
untied 808,960 45 False
At the GPT-2-small configuration above, an untied output projection would add another 50,257 × 768 = 38,597,376 parameters. Tying avoids that matrix: about 31% of the canonical tied model’s parameter count, or about 24% of the hypothetical untied total. Here it removes only 4,736, about 0.6% of the untied model. So the question for this model is not primarily memory; it is whether sharing the input and output representation changes behavior enough to resolve. A paired seed sweep (four seeds, 2,000 steps each, same seed on both sides):
seed tied untied untied - tied
0 0.3525 0.3520 -0.0005
1 0.3517 0.3519 +0.0002
2 0.3554 0.3527 -0.0027
3 0.3518 0.3515 -0.0003
mean untied - tied: -0.0008 stdev 0.0011 (1/4 favour tied)
The mean per-seed difference is -0.0008 (untied - tied) with a standard deviation of 0.0011, and the four paired deltas do not establish a stable advantage in either direction. The six-seed baseline variation below is also much larger than the observed mean shift, but that range is context, not a statistical equivalence threshold. The honest result is not resolved with four paired seeds. The parameter saving is exact; the representational effect is an empirical question, and this experiment is not strong enough to call it better, worse, or equivalent.
Training: prove it is happening
Chapter 11 turned “is the model learning?” into a chain of checkable consequences on one fixed batch:
TASK → OBJECTIVE → DEPENDENCY → GRADIENT → OWNERSHIP → UPDATE → CAPABILITY
Run its learning ledger on the transformer, with two capstone-scale additions. OBJECTIVE gains an initialization sanity check: a uniform predictor has cross-entropy ln(vocab_size), so for this deliberately small-logit initialization an initial loss on the same scale is expected. That is a reference, not a universal invariant for every language-model initialization. And CAPABILITY is overfit one tiny batch, which for a model this size is the strongest single end-to-end proof that the training machinery can fit those examples.
LEARNING LEDGER TinyGPT batch (8, 32)
TASK alignment x[:,1:] == y[:,:-1] ok
id range [0, 36] of 37 ok
OBJECTIVE initial loss 3.668 ln(vocab) = 3.611 ratio 1.02 ok
one-example CE manual 3.52942 vs 3.52942 ok
GRADIENT present 44/44 ok
finite 44/44 ok
OWNERSHIP owned 44/44 ok
UPDATE moved 44/44 ok
first divergence: none
CAPABILITY -- overfit one tiny batch (8 x 32):
step 0 loss 3.6557
step 100 loss 0.0230
step 200 loss 0.0083
step 300 loss 0.0045
step 399 loss 0.0029
Every link holds; the model drives one batch to near-zero loss in 400 steps. Now the fault the initial-loss check exists to catch. Tie the weights but skip the small-scale initialization — leave the embedding table at PyTorch’s default N(0, 1), which the tied lm_head then inherits:
TASK alignment x[:,1:] == y[:,:-1] ok
OBJECTIVE initial loss 119.463 ln(vocab) = 3.611 ratio 33.08 <-- FIRST DIVERGENCE
GRADIENT present 44/44 ok
OWNERSHIP owned 44/44 ok
UPDATE moved 44/44 ok
first divergence: OBJECTIVE/initial loss
The chain passes TASK and first looks wrong at the OBJECTIVE measurement: with the tied output matrix inheriting the embedding table’s N(0, 1) scale, this run produces enormous, spiky logits and a loss thirty-three times the uniform-predictor reference. The downstream mechanical checks still report health — gradients exist, the optimizer owns the parameters, and one step moves them. That proves the system can update from this starting point; it does not prove that training will recover cleanly or quantify how badly it will behave without running that experiment. The ledger has already done its job: it names the earliest surprising consequence before a falling loss can be mistaken for success.
Performance
Chapter 12’s performance ledger: decompose the step, then sweep the axes that matter.
WORKLOAD TinyGPT 4L/4H/128d | training step | batch 64 x 64 | CUDA
PHASE DECOMPOSITION (median ms, synchronized)
forward 3.428 31.5%
loss 0.165 1.5%
backward 6.176 56.8%
optimizer 1.104 10.2%
sum 10.872
Backward dominates at 57% — the opposite of Chapter 12’s shallow MLP, where the AdamW step over 30 million parameters was the largest phase. Here the model has only about 800,000 parameters, so the optimizer is comparatively cheap, while backpropagation through the four transformer blocks dominates the measured step. This phase decomposition does not by itself prove that scaled_dot_product_attention is the dominant operator inside backward; that would require the profiler or another operator-level experiment. The Chapter 12 lesson holds: localize before optimizing, and do not localize more finely than the evidence supports.
The batch-size and sequence-length sweeps:
BATCH SIZE SWEEP (block_size 64)
batch ms/step tokens/s peak MiB
16 9.98 102,601 77.1
32 9.74 210,347 113.6
64 11.04 370,930 187.2
128 17.04 480,788 337.1
256 31.19 525,278 634.6
SEQUENCE LENGTH SWEEP (batch 32) -- the transformer-specific axis
block ms/step tokens/s peak MiB
16 10.09 50,721 56.2
32 10.14 101,025 77.0
64 9.72 210,728 113.6
128 11.47 357,221 187.3
256 20.65 396,659 337.3
Batch size and sequence length both raise throughput with diminishing returns and both raise peak memory — but not at the same rate. Doubling the batch from 128 to 256 roughly doubles the measured peak; doubling the sequence length from 128 to 256 raises peak memory and step time by about 1.8× in this run. Chapter 12 measured a manual attention implementation that explicitly materialized [B, H, T, T] score/probability tensors and therefore exposed quadratic memory growth directly. The SDPA backend selected in this experiment is memory-efficient and does not materialize those same full tensors, so measured peak allocation grows much more gently over this range. The logical dense attention relation is still pairwise in sequence length; physical allocation depends on the selected kernel.
Compilation
Chapter 13’s question, asked once: did torch.compile help this model, and did correctness hold?
WORKLOAD TinyGPT training step | batch 64 x 64 | float32 | CUDA
config first call steady ms vs eager
eager - 9.37 1.00x
compiled ~60 s cold 8.12 1.15x
~2 s warm
correctness: max |Δ logits| eager vs compiled, same weights = 1.9e-06
A modest steady-state speedup — about 1.15× in this measurement, with other runs reaching roughly 1.3× — comes with a first-call cost of about a minute on a cold Inductor cache and about two seconds when the cache is warm. The compiled logits differ from eager by at most 1.9e-06 for the tested input and weights, so the correctness check passes at that stated tolerance; the measurement alone does not need to assign the difference to a particular fusion/reordering mechanism.
The amortization is easy to calculate. The measured steady-state saving is 9.37 - 8.12 = 1.25 ms/step. A 60 s cold compile therefore needs roughly 60,000 / 1.25 ≈ 48,000 reused steps to break even. A 2 s warm-cache first call needs roughly 1,600. The canonical 3,000-step run easily amortizes the warm cost but not the cold one. Whether compilation is worth taking depends on cache state and workload lifetime, not merely on the existence of a steady-state speedup.
That is the whole compiler section. Chapter 13 owns graph breaks, guards, recompilation and the cold-versus-warm cache split; if the compiled model were slower or kept recompiling, the investigation would move there. Here it is one row of the ledger.
Generation is a different program
Training and generation run the same weights through different code. Training is one forward pass over a fixed batch with targets. Generation is a loop that feeds the model its own output, and it has failure modes training never exercises.
Sampling strategy, same trained model, same seed:
greedy (argmax) 'The machine crosses the machine, because the gentle machine crosses at
dawn. That machine carries that machine, because that gentle machine...'
temperature 0.5 'The hidden market crosses the hidden market again. The golden window
counts the golden window without warning. That signal builds that...'
temperature 1.0 'The quiet shadow remembers the quiet shadow by the door. When one
garden watches without warning, one restless garden watches...'
temperature 1.5 'The quiet small window builds Quietly. Quietly, that restless window
answers that window. For years, this narrow forest thout warning...'
temp 1.0 + top_k 10 'The hidden market crosses the hidden market by the door. When one
garden watches without warning, one restless garden watches...'
Greedy is deterministic and, in this sample, falls into a strongly repetitive continuation (“the machine crosses the machine…”). Temperature 0.5 sharpens toward high-probability continuations; temperature 1.5 flattens the distribution enough that this sample contains misspellings (thout warning) and odd capitalization; top_k removes all but the highest-ranked candidates before sampling. None of these settings is universally “correct” — they alter the sampling distribution, and the useful tradeoff depends on the generation task.
Dropout left active during generation. model.eval() is not optional here — it is what turns dropout off. Generate with the model in train() mode and every call gives a different answer, because dropout’s stochasticity is layered on top of the sampling’s. This is Chapter 11’s eval()-versus-no_grad() distinction arriving as a generation bug: torch.no_grad() saves memory but does not touch dropout; model.eval() is the one that matters. TinyGPT.generate calls self.eval() and restores the previous mode in a finally — the pattern from Chapter 10’s attention module.
Generating past the context length. The model has a block_size-row position table. Feed it a longer sequence and it raises at the contract:
ValueError: sequence length 84 exceeds block_size 64
The generate loop truncates with idx[:, -block_size:] every step; delete that line and the context eventually grows past block_size and hits this. The explicit raise in forward is Chapter 8’s principle — fail at the contract with a message that names the problem, not forty lines later inside an embedding lookup.
The cost structure. Naive generation re-encodes the entire visible prefix at every step:
new tokens wall s ms/token
32 0.102 3.183
64 0.219 3.418
128 0.413 3.229
256 0.863 3.370
512 1.730 3.379
Milliseconds-per-token is roughly flat here because block_size is 64 and truncation caps the visible context: once the window is full, each call recomputes a transformer forward over at most 64 positions. Below that cap, naive generation reruns the entire visible prefix for every new token. In dense self-attention, that forward contains O(T²) attention interactions at visible length T (plus O(T) per-token projections/MLPs), so the per-generated-token cost grows with context before the window saturates.
A KV cache changes the computation rather than merely making the same loop faster: each layer stores past keys and values, computes the new token’s projections once, and attends the new query over the cached T keys/values, making the attention work for the new token O(T) rather than recomputing a full T × T attention pass. It also introduces mutable generation state, cache memory, eviction/window semantics and new correctness contracts. That is why it is deliberately outside this build: establish the stateless baseline first.
The run record, and this model’s baseline variation
Before comparing any two training runs of this model — tied versus untied, a new learning rate, a refactor — Chapter 14 says to know how much two runs of the identical configuration differ from the seed alone.
ONE RUN RECORD
{
"seed": 0, "config_hash": "0ead31df3032", "params": 804224,
"torch": "2.6.0+cu118", "device": "cuda",
"final_train_loss": 0.3487, "final_val_loss": 0.3525,
"peak_mib": 174.9, "wall_s": 31.54
}
BASELINE VARIATION TinyGPT 4L/4H/128d, 2000 steps, 6 seeds
seed train val
0 0.3487 0.3525
1 0.3475 0.3517
2 0.3505 0.3554
3 0.3459 0.3518
4 0.3484 0.3493
5 0.3476 0.3515
val loss median 0.3518 stdev 0.0018 spread 0.0061
wall time median 28.8 s spread 9.6 s (machine contention, not the model)
Across these six identical-configuration seeds, validation loss spans 0.0061 with a standard deviation of 0.0018. That empirical variation gives scale to the comparisons: the misalignment, causality and initialization failures move the metric enormously relative to it, so their effects are not subtle. Weight tying moves the paired mean by only 0.0008, and four paired seeds do not resolve that effect. Do not promote the six-seed min/max spread into a universal detection threshold — Chapter 14 showed that paired comparisons can sometimes resolve effects smaller than an unpaired range. The honest statement here is narrower: with the paired evidence collected, weight tying remains not resolved.
We have removed most of the magic
Look back at what the model is:
embedding lookup
+ learned positional representation
+ repeated (pre-norm attention, pre-norm MLP) residual blocks
+ final layer normalization
+ linear projection to vocabulary logits
+ cross entropy against the next-token shift
+ gradient descent with AdamW
There is no pipeline() call, no pretrained checkpoint, no hidden service. Every tensor in the forward pass is one you can name, shape, and trace to its source.
That does not make frontier systems simple. Scale changes the engineering entirely: distributed training across thousands of accelerators, data pipelines measured in trillions of tokens, custom attention kernels, learning-rate schedules and optimizer variants, architectural refinements, quantization, and inference systems with their own deep stack. Those layers introduce failure modes that this small model does not contain.
But the core path is no longer opaque. When a larger system does something unexpected, the questions this book taught remain useful — what tensor is this, which contract should hold, where did the first consequence diverge? At scale those questions are joined by new distributed and systems questions, and the experiments become more expensive, but the discipline of demanding evidence survives.
The LLM-era programmer
An assistant can generate every class in this chapter. It can write the tokenizer, the attention module, the training loop, the sampling function, the checkpoint code. It can generate a confident, fluent explanation of why your loss is not decreasing. Typing the code is no longer the valuable part.
Understanding the runtime is. When the generated system misbehaves, the useful questions are:
Does the target still describe the input? (alignment)
Are the token IDs in range? (tokenizer)
Can a future token change an earlier output? (causality)
Where does the shape ledger first diverge? (assembly)
Is the initial loss plausible for the declared initialization, with ln(vocab) as the uniform reference? (initialization)
Do gradients exist and are they finite? (the chain)
Does the optimizer own the parameters? (ownership)
Did the parameters actually move? (update)
Can the model overfit one tiny batch? (capability)
Which phase, at which sequence length, is the cost? (performance)
Did compilation help, and did correctness hold? (compilation)
Is this run really different from the last one? (comparison)
Source code can suggest answers to many of those questions, but it cannot establish the runtime facts by inspection alone. The book’s instruments turn the claims into observations: actual shapes, actual gradients, actual parameter identities, actual intervention deltas, actual timings. That is why so much of this book was about debugging: in a workflow where code is generated in seconds, evidence about what that code actually did is the scarce resource.
Using AI on a generated language model
A transformer is exactly the artifact an assistant will most confidently produce and most confidently misdiagnose. The prompt below refuses the explanation and demands the chain.
Here is a generated GPT-style language model that trains but generates poorly
(or trains suspiciously well, or is slow). Do not propose a fix or explain the
loss yet.
Prove the chain, in order, and stop at the first link without evidence:
1. TOKENIZER decode(encode(s)) == s on real text; every ID in [0, vocab).
2. ALIGNMENT torch.equal(x[:, 1:], y[:, :-1]) on a real batch.
3. CAUSALITY put the model in eval mode; change only tokens after position i.
Report the max delta on the prefix AND suffix. The prefix must
remain within a stated numerical tolerance and the suffix should
change, or the intervention is not informative.
4. SHAPE a per-stage ledger from [B,T] to logits, with the head-split
round trip merge(split(q)) == q. Name the first divergence.
5. INIT compare initial loss with the initialization contract; use
ln(vocab_size) as the uniform-predictor reference, not a universal
required value.
6. REGISTRATION every expected parameter in named_parameters(); note any tied.
7. GRADIENTS present and finite for every parameter after one backward.
8. OWNERSHIP the optimizer's parameter ids vs the model's.
9. UPDATE snapshot, step, confirm every parameter moved.
10. CAPABILITY can it overfit one tiny fixed batch to near-zero loss?
Only after all ten hold: is the loss actually good (compare to ln(vocab) and to
a validation set), and where does the step spend its time?
Do not suggest architecture changes, hyperparameter tuning, more data, a better
tokenizer, mixed precision or torch.compile until a specific link has failed.
The clause that earns its place is the third: report the max delta on the prefix and on the suffix. An assistant asked to “check causality” will inspect the mask and pronounce it triangular. The intervention is falsifiable — a non-zero prefix delta is a fact, and a zero suffix delta means the test proved nothing — and it is the check that would have caught this chapter’s most convincing failure.
An assistant can write the whole model and can inspect the batching code, but source inspection is not proof that the runtime targets are aligned. Ask it to prove the chain, not merely to explain the loss.
What you should now be able to answer
“The loss is dropping fast — training is working.” Suspiciously fast? On this corpus the correct model bottoms out near ln(vocab)/10; a loss heading for 0.001 means the model has found something cheaper to predict than the next token. Check torch.equal(x[:, 1:], y[:, :-1]) and check causality by intervention.
“Validation loss is excellent.” Better than the correct model’s? A non-causal model in this chapter reached 0.023 against the causal model’s 0.37. The metric rewards a model that can see the answer. Run the future-token intervention before celebrating.
“The model runs, so the shapes are right.” The wrong head reshape produces the right final shape and every downstream shape. Trace the ledger; the row that catches it is the head-split round trip, not any shape assertion.
“Generation produces garbage, so the model didn’t train.” Or: dropout is still active (model.eval()?), the context is not truncated to block_size, or the sampling distribution has been collapsed or distorted by an extreme temperature / top_k choice. top_k=1 is effectively greedy selection, not evidence of a broken model. Generation is a separate program with separate bugs; check them before re-training.
“It’s slow.” Which phase, at which sequence length, measured how? The performance ledger decomposes the step; the sequence-length sweep is the transformer-specific axis. “Slow” without a phase and a shape is not a diagnosis.
“I’ll add AMP and compile it.” Against which baseline, and did correctness hold? Chapter 12’s contract and Chapter 13’s correctness check apply unchanged. A faster wrong answer is not an optimization.
“The LLM wrote my training loop, and it looks right.” Then prove the chain: alignment, initial loss, gradients, ownership, movement, overfit-one-batch. “Looks right” is the state every failure in this chapter was in.
Exercises
The misalignment failure. Reproduce the opening: train with
y = data[i : i+block]. Record both loss curves and a generated sample from each. Then explain, from the causal mask, exactly why predicting tokentfrom tokentis nearly free.Tokenizer forensics. Build a tokenizer on a truncated sample of a corpus, then encode the full corpus with it. Catch the
KeyError. Then feed an out-of-range ID tonn.Embeddingand catch theIndexError. Write the two-line check that would have prevented both.Causality by intervention. Take a trained causal model. Change only the tokens after position
iand confirm the outputs at0..ido not move. Then setis_causal=False, retrain, and show the validation loss improving while the intervention fails. Report both deltas.The shape ledger. Implement the per-stage ledger for the full model. Break the head split with
reshape(B, H, T, D), confirm the output shape is unchanged, and show the round-trip row catching it. Then break the merge instead and show the same.Parameter accounting at two scales. Count this model’s parameters by module. Then compute the same breakdown for a config with
vocab_size=32000,n_embd=512. Identify the vocabulary size at which the embedding table overtakes one transformer block, and explain what that means for weight tying.Weight tying, measured. Run a paired seed sweep of tied versus untied on your workload. Report the per-seed deltas, their mean and spread, and the baseline seed variation for context. State whether the behavioral effect is resolved; do not use the baseline min/max range as an equivalence threshold. Then decide whether you would tie the weights, separating the exact parameter saving from the uncertain behavioral effect.
The learning ledger. Run the chain on a fresh model. Then introduce, one at a time: a detached residual stream, an optimizer built before a submodule is replaced, and a
10xlearning rate. For each, name the first link the ledger reports.Generation cost. Measure milliseconds-per-token as a function of visible context length below
block_size, then show what happens after the sliding window saturates. Explain why naive full-prefix dense attention recomputesO(T²)pairwise work per forward before the cap. Then estimate how a KV cache changes the new-token attention work toO(T)while adding cache memory and state.This model’s baseline variation. Run the identical configuration at eight seeds. Report the validation-loss distribution. Then make a change you believe helps and evaluate it at the same eight seeds. Use the paired deltas to judge the effect, with the baseline distribution as context rather than treating its min/max spread as a formal threshold.
The whole pipeline, on an unfamiliar model. Take a GPT implementation you did not write. Without changing it, run every instrument in this chapter in order and produce a one-line verdict for each: tokenizer, alignment, causality, shape ledger, initialization, learning ledger, performance ledger, generation modes. Identify the first that fails, or certify that none did.
The final checklist
Organised by instrument, and every item is one this chapter demonstrated.
TASK (Chapter 11)
[ ] decode(encode(text)) == text
[ ] every token ID in [0, vocab_size)
[ ] torch.equal(x[:, 1:], y[:, :-1]) the next-token shift
[ ] train/validation split is a real holdout
MODEL STRUCTURE (Chapters 5, 8, 10)
[ ] n_embd % n_head == 0
[ ] split_heads then merge_heads round-trips exactly
[ ] every block preserves [B, T, C]; logits are [B, T, vocab]
[ ] every expected parameter is in named_parameters(); tied weights verified
[ ] future-token intervention leaves the prefix unchanged within tolerance and changes the suffix
OBJECTIVE AND LEARNING (Chapter 11)
[ ] initial loss is plausible for the declared initialization; ln(vocab_size) is the uniform reference
[ ] loss is finite; one-example cross entropy reproduces by hand
[ ] gradients present and finite for every parameter
[ ] optimizer owns every trainable parameter
[ ] parameters move after a step
[ ] the model overfits one tiny batch
PERFORMANCE (Chapters 12, 13)
[ ] the step is decomposed into phases, synchronized
[ ] throughput and peak memory measured, not assumed
[ ] batch-size and sequence-length behaviour measured
[ ] compilation benchmarked warm, correctness checked
GENERATION (this chapter)
[ ] greedy works before sampling is added
[ ] context truncated to block_size every step
[ ] model.eval() -- dropout is off
[ ] temperature > 0; top-k bounded by vocab size; probs finite
COMPARABILITY (Chapter 14)
[ ] this model's seed-only baseline variation is measured
[ ] comparisons use shared seeds when pairing is valid
[ ] the run record carries config, tokenizer, seed, environment
What this book built, and what it did not
Fifteen chapters, starting from one scalar parameter and ending with a language model you can take apart. The through-line was never the model — it was the practice of making hidden structure visible and finding the first place it diverges from what you intended.
What the book deliberately left out, and where each now attaches:
- KV caching and fast inference — attach to this chapter’s generation loop: the naive full-prefix forward recomputes dense attention over the visible context, while a cache changes new-token attention to reuse past keys/values and perform
O(T)attention work against them. - Distributed training — attaches to Chapter 12’s phase decomposition and Chapter 5’s parameter ownership, now across processes.
- Better tokenization (BPE, byte-level) — attaches to the text→IDs boundary and Chapter 7’s representation contract; the model does not care how the IDs were made.
- Quantization and mixed precision beyond AMP — attach to Chapter 12’s dtype and memory accounting.
- Serving, batching, streaming — attach to Chapter 6’s producer/consumer model and Chapter 12’s latency-versus-throughput distinction.
- Evaluation at scale — attaches to Chapter 14’s baseline-variation measurement and eval-path fingerprint.
None of those was required to expose the core mechanisms in this first-principles build. At production scale several become foundational engineering concerns in their own right. The advantage is that they now attach to a system you can already inspect, so you can approach them asking what state was added, which contract changed, and what evidence would show whether the intervention helped — rather than as isolated API calls to be pasted in and hoped over.
Final thought
The most useful skill in PyTorch is not remembering function names. It is being able to look at a running model and ask:
What tensor is this?
What shape should it have, and what does each dimension mean?
Where did this value come from?
Does the gradient reach this parameter, and did the optimizer move it?
Is the target still describing the input?
What did this optimization actually improve?
Is this run really better, or does it just look that way?
If you can answer those, the abstractions stop being magic. And when an assistant writes the next thousand lines for you, you still know how to find out whether they work.
That was the whole point. The GPT-style model matters, but it was never the destination. The destination is the moment a model does something you did not expect and you know, without guessing, how to find out why.