← PyTorch From First Principles

Training: Which Link in the Learning Chain Is Broken?

Diagnose models that run but do not learn by walking TASK → OBJECTIVE → DEPENDENCY → GRADIENT → OPTIMIZER OWNERSHIP → PARAMETER UPDATE → CAPABILITY, using one fixed batch, exact parameter identity, controlled SGD, and tiny-batch overfitting to find the first missing consequence of learning.

Here is an ordinary piece of transfer-learning code. A small model, a frozen feature extractor, an optimizer, and a new classification head for the new task.

model = TinyClassifier()

for p in model.features.parameters():
    p.requires_grad_(False)

optimizer = torch.optim.AdamW(model.parameters(), lr=1e-2, weight_decay=0.0)

model.head = nn.Linear(32, 2)          # adaptation, performed later

Train it for two hundred steps on a balanced, perfectly aligned, entirely learnable binary task:

step=   0 loss=0.683606 acc=0.7144
step=  50 loss=0.683606 acc=0.7144
step= 100 loss=0.683606 acc=0.7144
step= 150 loss=0.683606 acc=0.7144
step= 200 loss=0.683606 acc=0.7144

Not approximately constant. Identical to six decimal places, two hundred times.

Nothing raised. Now look at the evidence a normal debugging session would collect:

loss                      0.683606
loss.requires_grad        True
loss.grad_fn              NllLossBackward0
head.weight  grad is None=False finite=True norm=1.555253e-01
head.bias    grad is None=False finite=True norm=4.327927e-03
optimizer.step() returned without raising

The loss is finite. It has a grad_fn. backward() populated a finite, nonzero gradient on the head. optimizer.step() ran. Every check a programmer normally reaches for reports health, and the model is completely stationary:

features.0.weight    delta=0.000000e+00
features.0.bias      delta=0.000000e+00
features.2.weight    delta=0.000000e+00
features.2.bias      delta=0.000000e+00
head.weight          delta=0.000000e+00
head.bias            delta=0.000000e+00

One check tells you why, and it is not a check about gradients:

features.0.weight    requires_grad=False in_optimizer=True
features.0.bias      requires_grad=False in_optimizer=True
features.2.weight    requires_grad=False in_optimizer=True
features.2.bias      requires_grad=False in_optimizer=True
head.weight          requires_grad=True  in_optimizer=False
head.bias            requires_grad=True  in_optimizer=False

The optimizer was constructed one line too early. It holds the four frozen feature tensors, which have no gradient and are skipped, and the two parameter objects belonging to the old head, which no longer participates in the forward pass. The new head is registered on the model, listed by named_parameters(), saved by state_dict(), reached by autograd, and absent from the optimizer. So optimizer.step() faithfully updates nothing.

Two statements are doing all the damage here, and both of them are things people say without noticing:

loss.backward() proves that a derivative was computed. It does not prove that the value receiving that derivative will ever change.

optimizer.step() proves that an optimizer method ran. It does not prove that it owned the parameters you intended to train.

Where we are

Chapter 10 finished attention with a very specific and very limited certificate. It established that an implementation has correct external shapes, correct axis semantics, correct head rearrangement, correct mask semantics, the right softmax axis, verified causal behavior, agreement with PyTorch’s own reference implementation, and finite gradients reaching every expected parameter.

Then it stopped, deliberately, and handed this chapter the harder case:

forward runs
loss is finite
backward runs
gradients exist
optimizer.step() runs

and the model is still useless

The opening above is exactly that state. So is a model whose loss sits at log(num_classes) forever, a model whose accuracy never leaves the majority baseline, a model that trains beautifully on garbage, and a model that becomes NaN on step three.

Chapter 3 gave us the machinery for one part of this: which parameters have a path to the loss. Chapter 5 gave us another: which structure actually owns a tensor object. Chapter 7 gave us a third: whether the representation the model sees means what we intended. This chapter is where those become one investigation, because a training system is not one operation that either works or does not.

What is the first consequence of learning that fails to appear?

The environment

Every number, table and trace in this chapter came from executing the code shown, with fixed seeds, in this environment:

PyTorch      2.13.0
Python       3.12.3
OS           Linux, CPU only

There is no CUDA device here, so this chapter makes no claims about mixed precision or device-specific numerics. Those belong to Chapter 12 and are left there.

The technique: prove the chain one boundary at a time

Every chapter has added an investigation method. Chapter 2 asked for the first wrong tensor rather than the first illegal one. Chapter 3 asked where the gradient path stops existing. Chapter 5 asked which structure disagreed about ownership. Chapter 7 asked where a sample stopped satisfying its contract. Chapter 8 asked where derived and observed geometry first diverged. Chapter 10 asked which attention invariant failed first.

Chapter 11 adds:

Trace one fixed batch from the task contract to parameter movement. Stop at the first expected consequence that fails to occur.

The reason this works is that learning is not an event. It is a chain of causes, and every arrow in it implies evidence that can be measured on one batch:

sample + target
      ↓        does y still describe x?
forward computation
      ↓        are the outputs in the form the objective expects?
objective / loss
      ↓        is this scalar the loss we intended?
autograd dependency
      ↓        does that loss depend on the parameters we mean to train?
gradient on intended parameters
      ↓        did those parameters receive finite derivatives?
optimizer ownership
      ↓        does the optimizer hold those exact objects?
parameter update
      ↓        did those exact objects move?
changed model behavior
      ↓        can repeated controlled steps fit a tiny fixed problem?
loss reduction on a controlled problem

The opening failure satisfies every one of those boundaries up to ownership, and fails there. A detached hidden layer gets as far as gradients. A shuffled target relation fails at the very first one and satisfies every other. Those three failures produce the same headline symptom — a loss that will not move usefully — and the same reflex, which is to change the learning rate.

The learning rate is not on that list. It appears later in this chapter, once every boundary above it has evidence, because tuning an optimizer against a broken task contract is the single most common way to waste a week.

The important shift is from:

“My model isn’t learning. What hyperparameter should I change?”

to:

“Which link in the learning chain have I actually proved?”

The controlled problem

Debugging without a baseline is guessing with extra steps, so this chapter builds one small system and keeps it for the whole chapter. Every deliberate failure below differs from it by exactly one change.

import torch
import torch.nn as nn
import torch.nn.functional as F

RULE_W = torch.tensor([1.0, 0.75])

def known_rule(x):
    return ((x @ RULE_W) > 0).long()

def make_data(n=2048, seed=0):
    g = torch.Generator().manual_seed(seed)
    x = torch.randn(n, 2, generator=g)
    return x, known_rule(x)

class TinyClassifier(nn.Module):
    def __init__(self, hidden=32, num_classes=2):
        super().__init__()
        self.features = nn.Sequential(
            nn.Linear(2, hidden), nn.ReLU(),
            nn.Linear(hidden, hidden), nn.ReLU(),
        )
        self.head = nn.Linear(hidden, num_classes)

    def forward(self, x):
        return self.head(self.features(x))

def build(seed=42, **kw):
    torch.manual_seed(seed)
    return TinyClassifier(**kw)

The task is linearly separable by a rule we wrote down: class 1 when x0 + 0.75*x1 > 0. That single property is what makes the chapter possible. Because the rule is known, target alignment is checkable exactly rather than by eye. Because the problem is easy, a failure to learn is unambiguous rather than a matter of patience. And because the whole thing runs on CPU in under a second, every claim below could be re-derived by a reader in a terminal.

The model is deliberately split into features and head, because that is the structure in which the opening bug actually occurs in real code.

The healthy reference

train class counts: [1039, 1009]
majority baseline : 0.50732421875
chance loss ln(2) : 0.6931471824645996
val class counts  : [258, 254] val majority 0.50390625
model = build(42)
opt = torch.optim.AdamW(model.parameters(), lr=1e-2, weight_decay=0.0)

for step in range(301):
    opt.zero_grad(set_to_none=True)
    loss = F.cross_entropy(model(x), y)
    loss.backward()
    opt.step()
 step  train loss  train acc  val loss  val acc
    0      0.6933     0.4976    0.6938   0.4922
   50      0.0158     0.9990    0.0121   0.9961
  100      0.0079     0.9995    0.0083   0.9961
  150      0.0045     0.9995    0.0075   0.9961
  200      0.0028     0.9995    0.0060   0.9961
  250      0.0019     1.0000    0.0049   0.9980
  300      0.0014     1.0000    0.0042   0.9980

Three numbers here are worth keeping, because later comparisons are meaningless without them. The initial loss is 0.6933, which is close to ln(2): the cross-entropy loss produced by equal probability on two classes. An arbitrary untrained classifier is not guaranteed to emit uniform probabilities; the near-match here tells us that this particular initialization begins close to that reference. The majority baseline is 0.5073, which is what “always predict the larger class” achieves on this dataset. And a healthy run reaches 0.0014 and perfect training accuracy in three hundred full-batch steps.

That last number matters most. When a broken variant plateaus at 0.68, the interesting fact is not that 0.68 is a large number. It is that this exact model, on this exact data, with this exact optimizer, reached 0.0014.

A note on the baseline, because it is routinely over-read. The majority baseline tells you what a metric looks like when the model has learned nothing useful. It does not tell you why the model is sitting there, and a model can beat it while learning nothing at all — the opening failure sat at 0.7144 accuracy, well above 0.5073, purely because a random projection followed by a random linear head happened to correlate with a linear rule. The baseline is a symptom threshold. It is not a diagnosis.

Boundary A: does the target still describe the input?

Chapter 7 established that preprocessing is part of the model’s definition, and that a representation can be perfectly legal and completely wrong. It also established the method: state the contract, then check it, independently of whether optimization succeeds. This chapter reuses that method rather than reteaching the five transform categories.

For a classification batch, the contract has a structural part and a semantic part.

def task_report(x, y, num_classes, rule=None):
    print(f"  x            shape={tuple(x.shape)} dtype={x.dtype} "
          f"finite={torch.isfinite(x).all().item()}")
    print(f"               range=[{x.min().item():+.3f}, {x.max().item():+.3f}] "
          f"mean={x.mean().item():+.4f} std={x.std().item():.4f}")
    print(f"  y            shape={tuple(y.shape)} dtype={y.dtype} "
          f"min={y.min().item()} max={y.max().item()}")
    counts = torch.bincount(y, minlength=num_classes)
    print(f"  class counts {counts.tolist()}  majority baseline "
          f"{(counts.max() / counts.sum()).item():.4f}")
    print(f"  aligned rows x.shape[0]==y.shape[0] {x.shape[0] == y.shape[0]}")
    if rule is not None:
        print(f"  known-rule agreement {(rule(x) == y).float().mean().item():.4f}")

Now destroy the alignment in the way real code destroys it — a permutation applied to one array and not the other:

perm = torch.randperm(len(x), generator=g)
x_bad = x[perm]
y_bad = y                 # forgot to permute
HEALTHY
  x            shape=(2048, 2) dtype=torch.float32 finite=True
               range=[-4.094, +4.101] mean=-0.0096 std=0.9934
  y            shape=(2048,) dtype=torch.int64 min=0 max=1
  class counts [1039, 1009]  majority baseline 0.5073
  aligned rows x.shape[0]==y.shape[0] True
  known-rule agreement 1.0000

SHUFFLED INPUTS, LABELS LEFT ALONE
  x            shape=(2048, 2) dtype=torch.float32 finite=True
               range=[-4.094, +4.101] mean=-0.0096 std=0.9934
  y            shape=(2048,) dtype=torch.int64 min=0 max=1
  class counts [1039, 1009]  majority baseline 0.5073
  aligned rows x.shape[0]==y.shape[0] True
  known-rule agreement 0.5000

Every reported structural statistic is unchanged. The shape, dtype, range, class counts and row count cannot reveal a permutation because no value or example was removed; the displayed mean and standard deviation are unchanged as well. What changed is the pairing, and exactly one row can see it: agreement with the known rule fell to 0.5000, which is chance on this balanced two-class problem.

Now train on it and watch what a healthy mechanism does with a broken problem:

all trainable params owned: True
step-0 total grad norm 1.6257e-02  all finite True  min delta 1.414e-02 max delta 2.985e-01
step=   0 loss=0.6936 train acc=0.4976
step= 100 loss=0.6625 train acc=0.5918
step= 200 loss=0.6501 train acc=0.6045
step= 300 loss=0.6410 train acc=0.6167
held-out acc on the real task: 0.568359375

The loss is falling. Training accuracy is rising. Gradients are finite, the optimizer owns everything, every parameter moves. This is a mechanically healthy training system fitting an arbitrary input-label pairing. After 300 steps it has reached only 0.6167 training accuracy, so the experiment does not yet establish that all 2,048 random pairings were memorized; it establishes something more important here: optimization progress can coexist with a broken task relation. Held-out accuracy on the actual task is 0.568, barely above the 0.504 baseline.

That is why the task contract comes first:

A healthy training mechanism can optimize a broken learning problem and can look as though it is making progress while doing it.

Real datasets rarely come with a known_rule. But most of them come with something you can check cheaply: a class distribution you can predict, a handful of examples you can look at side by side, an invariant the label must satisfy, a subset whose answer you know. Chapter 7’s argument applies unchanged — a probe you build on purpose is the only instrument that sees this category, because none of the tensor metadata can.

The mirror image of misalignment is leakage: a feature that contains the answer. It fails at the same boundary and is caught by the same discipline, but it usually produces suspiciously good results rather than a model that will not learn, so it is not this chapter’s subject.

Boundary B: is this scalar the objective you intended?

The loss is a function, and a function is code, and code can be wrong while running perfectly. The cheapest possible check is to reproduce it by hand for one example.

logits = model(x[:1])
target = y[:1]

loss = F.cross_entropy(logits, target)
manual = -F.log_softmax(logits, dim=-1)[0, target.item()]
logits             [[0.091391 0.048356]]
target             0
F.cross_entropy    0.67186153
-log_softmax[t]    0.67186153
allclose           True

Then the batch version, which also pins down the reduction:

batch F.cross_entropy        0.69213694
mean of per-example manual   0.69213694
reduction is 'mean'          True

Twelve lines that convert “I am using cross entropy” into “I know what number this function produced and why.” When it disagrees, you have found a boundary. When it agrees, you have ruled out one whole class of confusion for the price of one forward pass.

Chapter 7 already established the nuance about targets, and it is worth carrying rather than flattening. F.cross_entropy accepts two different target representations, and both are legitimate:

index-target loss    0.69213694
prob-target loss     0.69213694
smoothed prob target 0.69224751
label_smoothing=0.1  0.69224751

Class-index targets and one-hot probability targets produce exactly the same number here, and hand-built 0.95/0.05 soft targets reproduce label_smoothing=0.1 exactly, which for two classes is what smoothing by 0.1 means. So “cross entropy requires integer class indices” is not true. What is universal for this function is the other side:

The model’s contribution to cross entropy must be logits. The target may be indices or per-class probabilities.

The anchor failure: softmax before cross entropy

probs = logits.softmax(dim=-1)
loss = F.cross_entropy(probs, y)

This is the mistake everyone has made once. It runs. It produces a finite scalar. It has a grad_fn. Gradients flow, the optimizer owns everything, parameters move. Run it from an identical initialization against the correct version:

logits (correct)
   step      loss   grad norm  train acc
      0    0.6933  3.0339e-01     0.4976
      1    0.6467  3.0642e-01     0.9395
     50    0.0158  1.2085e-02     0.9990
    100    0.0079  5.2352e-03     0.9995
    200    0.0028  1.8874e-03     0.9995
    300    0.0014  4.1497e-03     1.0000

pre-softmaxed
   step      loss   grad norm  train acc
      0    0.6932  1.5088e-01     0.4976
      1    0.6695  1.5858e-01     0.9390
     50    0.3220  5.8824e-03     0.9995
    100    0.3178  3.1225e-03     0.9995
    200    0.3149  1.8243e-03     0.9995
    300    0.3141  1.0929e-03     1.0000

Read that carefully, because the naive story is wrong. The broken version learns. It reaches 100% training accuracy. What it does not do is reduce its loss below 0.3141, and it never will, because that is the floor of the objective it is actually optimizing:

perfect probs [1,0], target 0 -> 0.31326165795326233
-log(e^1/(e^1+e^0)) =           0.31326165795326233

Cross entropy applies a log-softmax to whatever it is given. Given probabilities in [0, 1], the best achievable input is a one-hot vector, whose log-softmax at the correct class is -log(e / (e + 1)), or about 0.3133. The loss can approach that and stop. A programmer watching this run sees a loss that plateaus at a suspiciously specific number and starts changing learning rates, when the correct observation is that the objective has a floor it is already at.

A finite, differentiable, decreasing loss is not evidence that it is the loss you intended.

Note also what this does to the gradient: the pre-softmax run has roughly half the gradient norm at step 0, because the softmax squashes the model’s output range before the loss ever sees it. Two objectives, same model, same initialization, same data, different gradients, and no error anywhere.

Boundary C: which parameters does the loss actually depend on?

The usual check is:

print("loss.requires_grad:", loss.requires_grad)
print("loss.grad_fn:", loss.grad_fn)

Keep it, and be precise about what it establishes. loss.requires_grad becomes True as soon as any single tensor on the path requires gradients. Chapter 3 made this concrete: a loss with a healthy grad_fn is evidence that something is connected, and no evidence at all about what.

Change one line in the model:

class DetachedClassifier(TinyClassifier):
    def forward(self, x):
        h = self.features(x)
        h = h.detach()          # the one changed line
        return self.head(h)
loss                 0.693254
loss.requires_grad   True
loss.grad_fn         NllLossBackward0

parameter            req_grad          grad  in_opt        delta
features.0.weight    True              None    True   0.0000e+00
features.0.bias      True              None    True   0.0000e+00
features.2.weight    True              None    True   0.0000e+00
features.2.bias      True              None    True   0.0000e+00
head.weight          True        1.5752e-01    True   7.8739e-02
head.bias            True        5.3857e-03    True   1.4142e-02

Both graph checks pass. Four of six parameters have requires_grad=True, are in the optimizer, and receive no gradient at all. This is Chapter 3’s question, asked at model scale:

Which intended parameter first stops having a path to this loss?

And here is what makes it dangerous rather than merely broken:

step=   0 loss=0.6845 acc=0.5649
step= 100 loss=0.3047 acc=0.9438
step= 200 loss=0.2122 acc=0.9590
step= 300 loss=0.1676 acc=0.9702

The model still reaches 97%. A linear head on a fixed random projection is enough for this task, so the failure does not announce itself as a failure. It announces itself as a model that is slightly worse than it should be — 0.1676 against the reference run’s 0.0014 — which is the kind of gap that gets attributed to hyperparameters and never investigated. The feature extractor is not dead weight: it still participates in every forward pass and supplies the fixed representation the head uses. What is dead is its learning path. Four trainable parameter tensors are carried through the model but can never adapt, and the only visible symptom may be a merely disappointing number.

This is a partial graph break, not “autograd is broken”, and the distinction is the whole diagnosis. Everything downstream of the detach is intact, which is exactly why editing anything downstream cannot help.

None is not zero, and zero is not None

Those two states look similar in a print statement and mean completely different things.

Here is a controlled way to produce the second. Drive one layer’s pre-activations far negative so its ReLU output is identically zero:

model = build(42)
with torch.no_grad():
    model.features[0].bias.fill_(-100.0)
features.0 preactivation max : -96.8088
post-ReLU zero fraction      : 1.0

parameter                  grad         norm       max|g|
features.0.weight        tensor   0.0000e+00   0.0000e+00
features.0.bias          tensor   0.0000e+00   0.0000e+00
features.2.weight        tensor   0.0000e+00   0.0000e+00
features.2.bias          tensor   1.4861e-02   7.9112e-03
head.weight              tensor   1.5895e-02   4.5727e-03
head.bias                tensor   3.6695e-02   2.5947e-02

Every one of those six parameters has a gradient tensor. Three of those tensors are exactly zero. The mechanism is local and readable: the ReLU derivative is zero wherever its input was negative, so nothing flows back into features.0; and features.2.weight’s derivative is proportional to its input, which is all zeros, so it is zero too. features.2.bias is unaffected, because a bias derivative does not depend on the layer’s input. The model can still change — through biases only — which is why this produces a slow, confusing, partial kind of learning rather than a clean stop.

Side by side with the detached model:

detached model  features.0.weight.grad is None: True
dead-relu model features.0.weight.grad is None: False

Same headline symptom, different boundary, different repair.

Two things to avoid concluding from this. First, a zero gradient does not mean “dead ReLU” — that is one mechanism among several. Second, and more important, do not attach absolute numerical thresholds to gradient magnitudes. There is no architecture-independent “healthy gradient norm”, because the number depends on the loss scale, the parameter shapes, the depth, the initialization and the batch size. What is diagnostic is comparison: how a norm changes over steps, how it differs across layers, how it compares against a known-good run, and where non-finite values first appear.

Boundary D: gradient evidence

The useful gradient report is not a wall of statistics. It distinguishes four states per parameter, and where each state occurs in the model is the diagnosis:

missing        grad is None
zero           grad tensor exists, every element is zero
finite/nonzero grad tensor exists and carries signal
non-finite     grad tensor contains NaN or Inf

Total gradient norm is worth computing, as one row rather than a section:

total = torch.stack([p.grad.norm() for p in model.parameters()
                     if p.grad is not None]).norm()

On the healthy reference at step 0 that is 3.4720e-01. That number means nothing on its own. It becomes evidence when the same measurement on the same model with a pre-softmaxed loss reads 1.7188e-01, or when it goes from 3.47e-01 to 5.94e+05 in one step, or when a layer near the input reads 0.0 while a layer near the output reads 1.6e-02.

One rule about the instrument itself, which the rest of the chapter depends on:

Observation and intervention are different operations. A diagnostic pass must not clip gradients, change the model’s mode, rebuild the optimizer, step a scheduler or alter regularization.

clip_grad_norm_ in particular does not belong in a debugging step, because it modifies the thing being measured. When it is used deliberately, know exactly what it returns:

signature: (parameters, max_norm, norm_type=2.0, error_if_nonfinite=False, foreach=None)
total grad norm before clipping : 0.347198
value returned by the function  : 0.347198
total grad norm after clipping  : 0.100000
returned value is the PRE-clip norm: True

The returned value is the norm before clipping. Logging it is genuinely useful; logging it while believing it describes the gradients the optimizer then consumed is not. And “clipping made the NaN go away” is a workaround, not a diagnosis.

Boundary E: does the optimizer own those exact objects?

Chapter 5 established the mechanism: an optimizer is a list of tensor objects captured at construction time, and membership is by object identity — not by name, not by shape, not by value. This chapter turns that into a two-way audit, because there are two ways for the two structures to disagree and only one of them is usually checked.

def optimizer_param_ids(optimizer):
    return {id(p) for g in optimizer.param_groups for p in g["params"]}

def ownership_audit(model, optimizer):
    """Two-way membership by object identity, in both directions."""
    opt_ids = optimizer_param_ids(optimizer)
    model_ids = {id(p) for p in model.parameters()}

    missing = [n for n, p in model.named_parameters()
               if p.requires_grad and id(p) not in opt_ids]
    stale = [p for g in optimizer.param_groups for p in g["params"]
             if id(p) not in model_ids]
    return missing, stale

missing is the fatal direction: trainable parameters the optimizer will never touch. stale is the diagnostic direction: tensors the optimizer still holds that the model no longer reaches. On the opening failure, missing is [head.weight, head.bias] and stale has two entries. Together, those facts strongly suggest model surgery after optimizer construction: the model has new trainable objects while the optimizer still owns unreachable old ones. The object identities establish the mismatch; the matching head-like shapes are supporting context, not a unique signature.

The subtlety worth executing

The optimizer’s zero_grad clears gradients for the parameters the optimizer owns. The new head is not one of them, so nothing in the training loop ever clears its gradient:

iteration   head.weight.grad norm   features.0.weight.grad
        1                0.155525   None
        2                0.311051   None
        3                0.466576   None
        4                0.622101   None
        5                0.777627   None

ratio to first iteration: 5.000010504061109

Exactly linear accumulation. The gradient is identical every iteration, because the parameters never move, so five backward passes deposit exactly five times the first gradient into a buffer nothing will ever read. And:

optimizer.zero_grad(set_to_none=True) then inspect:
  head.weight.grad is None : False

Chapter 1 established that PyTorch accumulates into .grad and that something has to clear it. Chapter 5 established that optimizer.zero_grad() is that something, and that it can only clear what it was given. Here both facts combine into a parameter that silently grows a gradient buffer for the entire run and never uses it.

For reference, zero_grad on this version defaults to set_to_none=True:

(self, set_to_none: bool = True) -> None
after default zero_grad, grad is None: True
after set_to_none=False, grad: tensor([[0., 0.], [0., 0.]])

That default matters for the previous section: after a default zero_grad, an owned parameter that receives no gradient has .grad is None, not a zero tensor. So “None” can mean “was cleared and never refilled” as well as “has no path to the loss”, and telling those apart is a question about when you looked.

The repair

Change one thing — construct the optimizer after the surgery:

model.head = nn.Linear(32, 2)

optimizer = torch.optim.AdamW(
    [p for p in model.parameters() if p.requires_grad],
    lr=1e-2, weight_decay=0.0,
)
features.0.weight    requires_grad=False in_optimizer=False
features.0.bias      requires_grad=False in_optimizer=False
features.2.weight    requires_grad=False in_optimizer=False
features.2.bias      requires_grad=False in_optimizer=False
head.weight          requires_grad=True  in_optimizer=True
head.bias            requires_grad=True  in_optimizer=True

features.0.weight    delta=0.000000e+00
features.0.bias      delta=0.000000e+00
features.2.weight    delta=0.000000e+00
features.2.bias      delta=0.000000e+00
head.weight          delta=7.873873e-02
head.bias            delta=1.414209e-02

step=   0 loss=0.675015 acc=0.7344
step=  50 loss=0.406795 acc=0.9297
step= 100 loss=0.302602 acc=0.9463
step= 150 loss=0.246815 acc=0.9570
step= 200 loss=0.210920 acc=0.9609
val acc 0.97265625

The frozen features still have delta=0, which is correct — that was the declared intent. The head moves, the loss falls, and held-out accuracy reaches 0.973. That is the difference between a stationary model and a working linear probe.

Two things changed in that construction and only one of them was necessary. Moving the optimizer after the surgery is the repair. Filtering on requires_grad is a separate, deliberate tightening: handing the frozen tensors to the optimizer would also have worked because parameters with .grad is None are skipped. But then optimizer membership would no longer mean “this is a parameter I intend to train.” Keeping the optimizer’s contents aligned with the declared trainable set makes later ownership reports easier to interpret.

Boundary F: did those objects actually move?

A gradient is a claim about a derivative. A parameter delta is a claim about the model. They are different measurements and the second one is the one that matters:

before = {n: p.detach().clone() for n, p in model.named_parameters()}
optimizer.step()
deltas = {n: (p.detach() - before[n]).norm().item()
          for n, p in model.named_parameters()}

A parameter can have a perfectly healthy gradient and a delta of exactly zero if it is not owned by the optimizer, if its group’s learning rate is zero, or if the optimizer skipped it for its own reasons. The delta is the evidence.

But movement is not automatically the update you think it is. Stateful optimizers make the relationship between the current gradient and the current step indirect: momentum carries earlier gradients forward, Adam and AdamW maintain running moment estimates, and AdamW’s decoupled weight decay contributes a term that does not come from the gradient at all. So two questions have to stay separate:

DID THE PARAMETER MOVE?
    measurable under the real optimizer, on any step

DID THIS GRADIENT PRODUCE EXACTLY THE UPDATE I EXPECT?
    needs a controlled optimizer

Chapter 1, at model scale

Chapter 1 built training out of one line:

w_new = w - lr * grad

That was one scalar parameter and a hand-written loop. It is worth proving that the same equation still describes an actual torch.optim step on an actual neural network, because once you have seen it hold you stop treating the optimizer as an oracle.

The experiment has to be controlled, so it uses a copy of the model and a fresh plain SGD with every update modifier disabled:

probe = copy.deepcopy(model)
lr = 0.1
sgd = torch.optim.SGD(probe.parameters(), lr=lr,
                      momentum=0.0, weight_decay=0.0)

for p in probe.parameters():
    p.grad = None
F.cross_entropy(probe(xb), yb).backward()

before = {n: p.detach().clone() for n, p in probe.named_parameters()}
grads = {n: p.grad.detach().clone() for n, p in probe.named_parameters()}
sgd.step()
parameter             max|observed - (-lr*g)|    ||delta||     lr*||g||
features.0.weight                   2.724e-08   7.7473e-03   7.7473e-03
features.0.bias                     2.725e-08   2.6625e-03   2.6625e-03
features.2.weight                   7.422e-09   2.7025e-02   2.7025e-02
features.2.bias                     6.636e-09   5.4703e-03   5.4703e-03
head.weight                         7.218e-09   1.7827e-02   1.7827e-02
head.bias                           2.794e-09   7.7633e-03   7.7633e-03
allclose over every parameter: True

w_new = w - lr * grad, elementwise, across 1,218 parameters in six tensors, to floating-point tolerance. The scalar loop from Chapter 1 was not merely toy arithmetic: under these deliberately restricted SGD settings, it is exactly what this optimizer does.

Now the same measurement under AdamW, from the same state, on its first step:

features.0.weight    ||delta||=7.9999e-01  lr*||g||=7.7473e-03   ratio=103.260
features.0.bias      ||delta||=5.6562e-01  lr*||g||=2.6625e-03   ratio=212.441
features.2.weight    ||delta||=2.9306e+00  lr*||g||=2.7025e-02   ratio=108.439
features.2.bias      ||delta||=5.5677e-01  lr*||g||=5.4703e-03   ratio=101.782
head.weight          ||delta||=7.8740e-01  lr*||g||=1.7827e-02   ratio=44.170
head.bias            ||delta||=1.4142e-01  lr*||g||=7.7633e-03   ratio=18.217

Between 18 and 212 times lr * ||grad||, and a different ratio for every tensor. That is not a bug; it is what an adaptive optimizer is for, and it is precisely why the equality above has to be stated with its conditions attached:

Δθ = -lr · grad is a property of the controlled plain-SGD experiment, not a universal optimizer invariant. Do not try to force AdamW into it — reproduce one step with fresh SGD instead.

This is also worth remembering when reading Chapter 7’s opening again: the same representation bug was catastrophic under SGD and nearly invisible under Adam. Adaptive update normalization changes what a gradient does to a parameter, which changes what a bug looks like from the outside.

One instrument: the learning ledger

Chapter 7 has a stage report. Chapter 8 has a shape ledger. Chapter 9 has a feature-space inspector. Chapter 10 has an attention ledger. This chapter’s instrument walks the learning chain over one fixed batch and names the first boundary whose evidence contradicts the intended training system.

The design constraints are the ones argued for above, with one important warning: this is not a read-only report. It calls backward() and optimizer.step(), so it mutates gradients, parameter values and optimizer state. Run it on a disposable copy/checkpoint, or treat the step as a deliberate sacrificial probe that you intend to keep. Within that probe it performs no clipping, scheduler step, mode change or regularization edit. It clears .grad on every model parameter before backward() so the gradient rows describe this batch rather than an earlier accumulation, but it reports any pre-existing gradients first so accidental carry-over remains visible.

The default expected set below is every parameter with requires_grad=True. That is correct for this dense controlled model. In conditional, sparse or mixture-style models, a trainable parameter may legitimately not participate in a particular batch; in that case the expected set must come from the model’s execution contract rather than from requires_grad alone.

def learning_ledger(model, optimizer, x, y, num_classes,
                    rule=None, loss_fn=F.cross_entropy):
    """One controlled, state-mutating probe over one fixed batch.

    Calls backward() and optimizer.step(), so it changes gradients, model
    parameters and optimizer state. It performs no clipping, scheduler step,
    mode change or regularization edit. Gradients are cleared to None before
    backward so the reported derivatives belong to this batch.
    """
    rows, first = [], None

    def row(section, label, text, ok=None, detail=None):
        nonlocal first
        rows.append([section, label, text, ok, detail])
        if ok is False and first is None:
            first = f"{section}/{label}"

    named = dict(model.named_parameters())
    expected = [n for n, p in named.items() if p.requires_grad]

    # TASK ------------------------------------------------------------
    row("TASK", "rows", f"x {x.shape[0]} / y {y.shape[0]}",
        x.shape[0] == y.shape[0])
    row("TASK", "types", f"x {x.dtype} / y {y.dtype}", y.dtype == torch.long)
    row("TASK", "finite", f"x finite {torch.isfinite(x).all().item()}",
        bool(torch.isfinite(x).all()))
    counts = torch.bincount(y, minlength=num_classes)
    row("TASK", "classes",
        f"counts {counts.tolist()}  majority baseline "
        f"{(counts.max() / counts.sum()).item():.3f}",
        bool(y.min() >= 0 and y.max() < num_classes))
    if rule is not None:
        agreement = (rule(x) == y).float().mean().item()
        row("TASK", "known rule", f"agreement {agreement:.3f}", agreement > 0.99)

    # FORWARD ---------------------------------------------------------
    row("FORWARD", "mode", f"model.training = {model.training}")
    logits = model(x)
    finite = bool(torch.isfinite(logits).all())
    row("FORWARD", "logits",
        f"{tuple(logits.shape)} {logits.dtype}  finite {finite}",
        finite and logits.shape[0] == x.shape[0])
    out = logits.detach()
    row("FORWARD", "scale", f"mean {out.mean():+.4f}  std {out.std():.4f}  "
                            f"max|.| {out.abs().max():.4f}")

    # LOSS ------------------------------------------------------------
    loss = loss_fn(logits, y)
    row("LOSS", "value", f"{loss.item():.6f}", bool(torch.isfinite(loss)))
    if loss_fn is F.cross_entropy and y.dtype == torch.long:
        manual = -F.log_softmax(logits[:1], dim=-1)[0, y[0]]
        one = F.cross_entropy(logits[:1], y[:1])
        row("LOSS", "one example",
            f"manual {manual.item():.6f} vs reported {one.item():.6f}",
            bool(torch.allclose(manual, one, atol=1e-6)))
    row("LOSS", "graph",
        f"requires_grad {loss.requires_grad}  grad_fn "
        f"{type(loss.grad_fn).__name__ if loss.grad_fn else None}",
        loss.requires_grad and loss.grad_fn is not None)

    # GRADIENT --------------------------------------------------------
    carried = [n for n in expected if named[n].grad is not None]
    row("GRADIENT", "pre-existing",
        f"{len(carried)}/{len(expected)} already held a gradient before backward",
        not carried)
    for p in model.parameters():
        p.grad = None
    loss.backward()

    absent = [n for n in expected if named[n].grad is None]
    live = [n for n in expected if named[n].grad is not None]
    nonfinite = [n for n in live if not torch.isfinite(named[n].grad).all()]
    zero = [n for n in live if not (named[n].grad != 0).any()]

    row("GRADIENT", "expected", f"{len(expected)} trainable parameters")
    row("GRADIENT", "present", f"{len(live)}/{len(expected)}", not absent,
        None if not absent else "None: " + ", ".join(absent))
    row("GRADIENT", "finite", f"{len(live) - len(nonfinite)}/{len(live)}",
        not nonfinite,
        None if not nonfinite else "non-finite: " + ", ".join(nonfinite))
    row("GRADIENT", "nonzero", f"{len(live) - len(zero)}/{len(live)}", None,
        None if not zero else "exactly zero: " + ", ".join(zero))
    if live:
        total = torch.stack([named[n].grad.norm() for n in live]).norm().item()
        row("GRADIENT", "total norm", f"{total:.4e}")

    # OPTIMIZER -------------------------------------------------------
    missing, stale = ownership_audit(model, optimizer)
    groups = "  ".join(
        f"[{i}] lr {g['lr']:g} wd {g.get('weight_decay', 0):g} n {len(g['params'])}"
        for i, g in enumerate(optimizer.param_groups))
    row("OPTIMIZER", "groups", f"{type(optimizer).__name__}  {groups}")
    row("OPTIMIZER", "owned",
        f"{len(expected) - len(missing)}/{len(expected)} trainable model parameters",
        not missing, None if not missing else "missing: " + ", ".join(missing))
    row("OPTIMIZER", "stale",
        f"{len(stale)} optimizer tensors no longer reachable from the model",
        not stale)

    # UPDATE ----------------------------------------------------------
    before = {n: p.detach().clone() for n, p in named.items()}
    optimizer.step()
    deltas = {n: (named[n].detach() - before[n]).norm().item() for n in expected}
    still = [n for n, d in deltas.items() if d == 0.0]
    row("UPDATE", "moved",
        f"{len(deltas) - len(still)}/{len(deltas)} expected parameters",
        not still, None if not still else "unchanged: " + ", ".join(still))
    if deltas:
        row("UPDATE", "delta range",
            f"{min(deltas.values()):.3e} .. {max(deltas.values()):.3e}")

    width = max(len(t) for _, _, t, _, _ in rows)
    print(f"LEARNING LEDGER   {type(model).__name__}   {type(optimizer).__name__}"
          f"   batch {tuple(x.shape)}")
    previous = None
    for section, label, text, ok, detail in rows:
        head = "" if section == previous else section
        previous = section
        mark = ""
        if ok is False:
            mark = ("  <-- FIRST DIVERGENCE" if f"{section}/{label}" == first
                    else "  fail")
        print(f"  {head:10s}{label:14s}{text:{width}s}{mark}")
        if detail:
            print(f"  {'':10s}{'':14s}{detail}")
    print(f"  first divergence: {first or 'none'}")
    return {"first_divergence": first, "loss": loss.item(), "deltas": deltas}

On the healthy reference:

LEARNING LEDGER   TinyClassifier   AdamW   batch (256, 2)
  TASK      rows          x 256 / y 256
            types         x torch.float32 / y torch.int64
            finite        x finite True
            classes       counts [115, 141]  majority baseline 0.551
            known rule    agreement 1.000
  FORWARD   mode          model.training = True
            logits        (256, 2) torch.float32  finite True
            scale         mean +0.1037  std 0.0457  max|.| 0.3531
  LOSS      value         0.694159
            one example   manual 0.671861 vs reported 0.671861
            graph         requires_grad True  grad_fn NllLossBackward0
  GRADIENT  pre-existing  0/6 already held a gradient before backward
            expected      6 trainable parameters
            present       6/6
            finite        6/6
            nonzero       6/6
            total norm    3.4720e-01
  OPTIMIZER groups        AdamW  [0] lr 0.01 wd 0 n 6
            owned         6/6 trainable model parameters
            stale         0 optimizer tensors no longer reachable from the model
  UPDATE    moved         6/6 expected parameters
            delta range   1.414e-02 .. 2.931e-01
  first divergence: none

Twenty-two rows, one batch, and a claim about every boundary from the task contract to parameter movement. That is what “the training step is mechanically sound” should mean before anyone touches a hyperparameter.

Four broken systems, one instrument

Each of the systems below differs from the healthy reference by exactly one change, and each is run through the same function on the same batch. The later listings are abridged to the rows that differ; every omitted row is identical to the healthy run above.

Stale optimizer membership

  TASK      rows          x 256 / y 256
            types         x torch.float32 / y torch.int64
            finite        x finite True
            classes       counts [115, 141]  majority baseline 0.551
            known rule    agreement 1.000
  FORWARD   mode          model.training = True
            logits        (256, 2) torch.float32  finite True
            scale         mean -0.0539  std 0.0429  max|.| 0.1612
  LOSS      value         0.685485
            one example   manual 0.675722 vs reported 0.675722
            graph         requires_grad True  grad_fn NllLossBackward0
  GRADIENT  pre-existing  0/2 already held a gradient before backward
            expected      2 trainable parameters
            present       2/2
            finite        2/2
            nonzero       2/2
            total norm    1.9480e-01
  OPTIMIZER groups        AdamW  [0] lr 0.01 wd 0 n 6
            owned         0/2 trainable model parameters          <-- FIRST DIVERGENCE
                          missing: head.weight, head.bias
            stale         2 optimizer tensors no longer reachable from the model  fail
  UPDATE    moved         0/2 expected parameters                   fail
                          unchanged: head.weight, head.bias
            delta range   0.000e+00 .. 0.000e+00
  first divergence: OPTIMIZER/owned

Everything above OPTIMIZER reports health, including the gradient section, which is where most people stop looking. Two rows are worth reading closely. expected 2 trainable parameters is the ledger taking the freeze at its word: requires_grad=False is a declared intent, so frozen parameters are not counted as expected to move. And groups ... n 6 against owned 0/2 is the whole bug in two numbers — the optimizer holds six tensors and none of them is one of the two the model wants trained.

The UPDATE row is real but downstream. It is a consequence of the ownership failure, not an independent finding, which is exactly the discipline Chapters 8 and 10 established: fix the first divergence, then rerun.

Detached hidden representation

Every TASK, FORWARD and LOSS row is identical to the healthy run, so the listing starts where it stops being identical:

  GRADIENT  pre-existing  0/6 already held a gradient before backward
            expected      6 trainable parameters
            present       2/6                                      <-- FIRST DIVERGENCE
                          None: features.0.weight, features.0.bias,
                                features.2.weight, features.2.bias
            finite        2/2
            nonzero       2/2
            total norm    1.9444e-01
  OPTIMIZER groups        AdamW  [0] lr 0.01 wd 0 n 6
            owned         6/6 trainable model parameters
            stale         0 optimizer tensors no longer reachable from the model
  UPDATE    moved         2/6 expected parameters                   fail
                          unchanged: features.0.weight, features.0.bias,
                                     features.2.weight, features.2.bias
  first divergence: GRADIENT/present

The optimizer section is clean. Ownership is perfect. The failure is one boundary earlier, and the UPDATE row shows the same four names for a completely different reason: not “the optimizer does not hold them” but “they had no gradient to step with.”

Two failures, nearly identical UPDATE rows, opposite repairs. That is the argument for walking the chain in order rather than starting from the symptom.

Destroyed target alignment

  TASK      rows          x 256 / y 256
            types         x torch.float32 / y torch.int64
            finite        x finite True
            classes       counts [115, 141]  majority baseline 0.551
            known rule    agreement 0.434                          <-- FIRST DIVERGENCE
  FORWARD   mode          model.training = True
            logits        (256, 2) torch.float32  finite True
  LOSS      value         0.695320
            one example   manual 0.685739 vs reported 0.685739
            graph         requires_grad True  grad_fn NllLossBackward0
  GRADIENT  present       6/6
            finite        6/6
            nonzero       6/6
            total norm    1.1394e-01
  OPTIMIZER owned         6/6 trainable model parameters
            stale         0 optimizer tensors no longer reachable from the model
  UPDATE    moved         6/6 expected parameters
            delta range   1.414e-02 .. 2.960e-01
  first divergence: TASK/known rule

Every other row in the ledger reports health, under one broken row. If the ledger started at FORWARD — which is where a debugging session usually starts, because the model is the interesting part — it would report a flawless training system optimizing nonsense.

The failure the ledger cannot see

Run it with the pre-softmaxed objective:

  LOSS      value         0.693597
            graph         requires_grad True  grad_fn NllLossBackward0
  GRADIENT  present       6/6
            finite        6/6
            nonzero       6/6
            total norm    1.7188e-01
  OPTIMIZER owned         6/6 trainable model parameters
  UPDATE    moved         6/6 expected parameters
  first divergence: none

first divergence: none, on a model that will plateau at 0.3141 forever. This is not a defect to patch; it is the instrument being honest about its reach. Notice what it did not do: the one example row is absent, because the loss function is not F.cross_entropy and the ledger declines to verify a formula it does not know. It reports what it can measure and makes no claim about what it cannot.

Chapter 10 made the same point about the attention ledger reaching Levels 1 through 3 and not Level 4. Here the boundary is:

The ledger proves that the machinery is intact. Only you can say whether the objective it is optimizing is the one you meant.

The failure matrix

Every row below is an executed experiment, not an expectation:

FailureTaskLossGradientOptimizerDeltaTiny batch
stale replaced headpasspasspass, 2/2 finitefail, 0/2 ownedfail, 0/2 movedflat at 0.68758
detached hidden statepasspassfail, 2/6 presentpassfail, 2/6 moved0.149 vs 0.002
shuffled target relationfail, 0.434passpasspasspassmemorizes
softmax before CEpassfloor at 0.3133passpasspassstops at 0.3136
lr = 1e4passpass at step 0pass at step 0passmoved, wrong scalediverges to NaN

Four of those five have a first divergence the ledger names on its own. The learning rate is the exception and is worth being precise about: on step 0 every ledger row passes, including UPDATE/moved, because the parameters do move. What is wrong is the size of the movement relative to the parameters, which the delta range row records but does not judge. That failure needs the escalation instruments below, and it is the reason the ledger reports evidence rather than verdicts wherever a threshold would have to be invented.

Boundary G: can this system fit a tiny fixed batch?

Once every boundary above is clean for a single step, one question remains, and it is the only one that tests them together:

Can this model, this objective and this update path reduce this loss on these exact examples?

The test is not “train and see”. It is a controlled environment in which learning should be trivially easy, so that failure is informative:

def tiny_batch_capability(model, x_small, y_small, optimizer=None,
                          steps=500, lr=0.5):
    """Controlled capability test on a fixed batch.

    optimizer=None copies the model and builds fresh plain SGD over its
    trainable parameters, answering 'can this MODEL fit these examples under
    this simplified optimizer?'. Passing an optimizer uses and mutates the
    supplied model/optimizer, answering 'can this TRAINING SYSTEM fit these
    examples?'. The caller is responsible for any model-level regularization
    such as dropout and for any augmentation applied before this function.
    """
    was_training = model.training
    if optimizer is None:
        model = copy.deepcopy(model)
        optimizer = torch.optim.SGD(
            [p for p in model.parameters() if p.requires_grad],
            lr=lr, momentum=0.0, weight_decay=0.0)
    model.train()
    initial = None
    for step in range(steps + 1):
        logits = model(x_small)
        loss = F.cross_entropy(logits, y_small)
        if step == 0:
            initial = loss.item()
        optimizer.zero_grad(set_to_none=True)
        loss.backward()
        optimizer.step()
    model.eval()
    with torch.no_grad():
        logits = model(x_small)
        final = F.cross_entropy(logits, y_small).item()
        acc = (logits.argmax(-1) == y_small).float().mean().item()
    model.train(was_training)
    return initial, final, acc

That optimizer argument is not a convenience. It is the point:

system                         optimizer      initial     final  accuracy
healthy                        as built       0.69209   0.00227    1.0000
healthy                        fresh SGD      0.69209   0.00227    1.0000
frozen feats + new head        as built       0.68758   0.68758    0.7188
frozen feats + new head        fresh SGD      0.68758   0.14840    0.9688
detached representation        as built       0.69209   0.14906    0.9688
healthy, RANDOM labels         fresh SGD      0.69527   0.00080    1.0000

Look at the two rows for the stale-head system. With the real optimizer, the loss does not move at all: 0.68758 to 0.68758. With a fresh one, it drops to 0.148 and fits 31 of 32 examples. The model was always capable. The training system was not. A tiny-batch test that quietly rebuilds the optimizer has repaired the bug it was supposed to detect, and will report that everything is fine.

So the version of the test that people usually write — new model, new optimizer, small batch — answers a question about capacity. It is a real question, and it is not the question you have when a specific training run will not learn.

What the tiny-batch test proves

Passing it, with the real training system, supports all of:

the representation reaches the model
the objective can be reduced on these examples
the graph delivers useful gradients
the optimizer can change the model
the model has enough capacity to fit this batch under this setup

Now the last row of that table, which is why this must not become a ritual. Replace the labels with coin flips:

random labels agreement with known rule: 0.4375

  step      loss  accuracy
     0   0.69527    0.3750
    25   0.51209    0.7500
   100   0.40176    0.7812
   500   0.11932    1.0000
  2000   0.00080    1.0000

Perfect memorization of 32 examples whose labels mean nothing. The tiny-batch test passed with flying colors on data that has no learnable structure whatsoever.

Passing the tiny-batch test proves that this model, objective and update path can fit these particular examples. It does not prove that the examples encode the task you intended.

Which is exactly why it sits at the end of the chain rather than the beginning. The task contract is what rules that out, and nothing downstream of it can.

Failing it does not identify a cause either:

it does not tell you whether the problem is capacity, the objective,
the graph, ownership, the update scale or the regularization

It tells you that something in the local system is still wrong and that generalization is not yet the question. The chain above it is what localizes it.

Regularization is a confounder during forensics

If the tiny-batch test fails, the first move is not to change the model. It is to remove the causal paths that were added deliberately. Here is the full recipe on 32 examples the model can memorize perfectly: dropout 0.5, weight decay 0.1, input noise 0.5, label smoothing 0.1.

  step     0  train loss  0.74043   eval loss  0.69209  eval acc 0.5000
  step   100  train loss  0.41364   eval loss  0.26129  eval acc 0.9375
  step   500  train loss  0.37314   eval loss  0.25478  eval acc 0.9688
  step  1000  train loss  0.45278   eval loss  0.26875  eval acc 0.9688
  step  2000  train loss  0.40803   eval loss  0.24468  eval acc 0.9688

Two thousand steps on 32 examples and it will not go below 0.24. Nothing is broken. Remove the confounders one at a time:

 no input noise
  step  2000  train loss  0.25678   eval loss  0.07379  eval acc 1.0000
 no noise, no smoothing
  step  2000  train loss  0.08353   eval loss  0.01601  eval acc 1.0000
 + weight_decay 0
  step  2000  train loss  0.02566   eval loss  0.00067  eval acc 1.0000
 + dropout 0
  step  2000  train loss  0.00000   eval loss  0.00000  eval acc 1.0000

Each of those is an intervention that removes exactly one source of variation, and every one of them moves the number. None of these mechanisms is bad. Each adds a causal path between the parameters and the loss, and during forensics every extra path is a confounder.

Build the smallest training system that should obviously learn. Then add back one mechanism at a time.

While simplifying, get the mode right, and be precise about what it does. model.train() and model.eval() set the module hierarchy’s training-mode flag. Modules such as dropout and batch normalization — and any custom module that consults self.training — may change behavior in response. These calls do not enable or disable autograd:

train mode, two identical calls agree: False
eval  mode, two identical calls agree: True

in eval mode:
  loss.requires_grad True
  loss.grad_fn       NllLossBackward0
  gradients present  4 / 4

in train mode inside torch.no_grad():
  loss.requires_grad False
  loss.grad_fn       None

eval() does not disable autograd; a backward pass in eval mode populates every gradient. torch.no_grad() does not disable dropout. They solve different problems and neither substitutes for the other.

Escalation: what to inspect when every boundary passed

Everything so far is cheap — one batch, one step. If all of it is clean and the tiny-batch test still fails, the next instruments cost more and are worth reaching for in this order.

Activation statistics

Forward hooks observe module boundaries. That is a real limitation worth stating: a hook on nn.Linear sees that module’s output, not arbitrary tensor operations written inline in a forward method, so a functional-style model is largely invisible to them.

def activation_report(model, x, types=(nn.Linear, nn.ReLU)):
    rows, handles = [], []

    def make(name):
        def hook(module, inputs, output):
            if not torch.is_tensor(output):
                return
            o = output.detach()
            rows.append((name, type(module).__name__, tuple(o.shape),
                         bool(torch.isfinite(o).all()),
                         o.mean().item(), o.std().item(),
                         o.min().item(), o.max().item(),
                         (o == 0).float().mean().item(),
                         o.std(dim=0).mean().item()))
        return hook

    for name, module in model.named_modules():
        if isinstance(module, types):
            handles.append(module.register_forward_hook(make(name)))
    try:
        with torch.no_grad():
            model(x)
    finally:
        for h in handles:
            h.remove()
    return rows

It returns the rows rather than printing them; the formatting below is a separate concern. The try/finally is not decoration. A hook left attached keeps running on every subsequent forward pass, including the ones in your real training loop.

HEALTHY (untrained)
  module      type    shape      fin        mean      std      min      max  zero%  batch std
  features.0  Linear  (256, 32)  True     -0.019    0.725   -2.675    3.478   0.0%     0.5386
  features.1  ReLU    (256, 32)  True      0.280    0.430    0.000    3.478  51.8%     0.3051
  features.2  Linear  (256, 32)  True     -0.059    0.313   -2.263    1.246   0.0%     0.2134
  features.3  ReLU    (256, 32)  True      0.090    0.137    0.000    1.246  52.9%     0.0930
  head        Linear  (256, 2)   True      0.104    0.046    0.022    0.353   0.0%     0.0449

HEALTHY (after 300 steps of the reference run)
  module      type    shape      fin        mean      std      min      max  zero%  batch std
  features.0  Linear  (256, 32)  True      0.141    1.097   -6.173    5.953   0.0%     0.9362
  features.1  ReLU    (256, 32)  True      0.498    0.700    0.000    5.953  44.5%     0.5866
  features.2  Linear  (256, 32)  True      0.395    5.769  -56.815   46.287   0.0%     4.8293
  features.3  ReLU    (256, 32)  True      2.152    3.823    0.000   46.287  53.8%     3.0047
  head        Linear  (256, 2)   True      0.098   54.267 -211.192  208.711   0.0%    53.3090

DEAD FIRST LAYER (features.0.bias = -100)
  module      type    shape      fin        mean      std      min      max  zero%  batch std
  features.0  Linear  (256, 32)  True    -99.986    0.571 -102.768  -96.809   0.0%     0.5386
  features.1  ReLU    (256, 32)  True      0.000    0.000    0.000    0.000 100.0%     0.0000
  features.2  Linear  (256, 32)  True     -0.014    0.116   -0.170    0.176   0.0%     0.0000
  features.3  ReLU    (256, 32)  True      0.044    0.063    0.000    0.176  53.1%     0.0000
  head        Linear  (256, 2)   True      0.102    0.067    0.035    0.168   0.0%     0.0000

The zero% column is a trap and this table is the proof. A ReLU is designed to produce zeros: the healthy model reads 51.8% and 52.9%, and after 300 steps of successful training it reads 44.5% and 53.8%. In the dead model, features.3 reads 53.1% — completely normal — while the network has already collapsed two layers earlier. There is no useful “dead ReLU percentage” threshold.

The column that identifies the failure is the last one: standard deviation across the batch, which asks whether the layer’s output depends on which example went in. It reads 0.0000 from features.1 onward. The logits are literally identical for all 256 examples, which is a model that cannot discriminate by construction, and it is a much stronger statement than any zero fraction.

Update scale, which is not gradient scale

An absurd learning rate is worth executing because the trace contradicts the story people tell about it:

step         loss    grad norm   param norm   delta norm logits finite
   0   6.9416e-01   3.4720e-01   5.3009e+00   3.4720e+03          True
   1   9.7960e+07   5.9419e+05   3.4720e+03   5.9419e+09          True
   2   1.2313e+26   5.3972e+17   5.9419e+09          inf          True
   3          nan          nan          inf          nan         False

At step 0 the gradient norm is 3.47e-01 — the same value the healthy reference reports, because it is the same model on the same batch. Nothing about the gradient is large. What is large is lr * grad: a parameter tensor with norm 5.3 receives an update with norm 3472. The explosion at steps 1 and 2 is a consequence of the first update, not its cause.

“The gradients exploded” and “the update destroyed the parameters” are different diagnoses. Report gradient norm, parameter norm and delta norm as separate columns, because the first step distinguishes them and every later step does not.

The first non-finite boundary

Do not debug the final NaN. Find the first tensor that became non-finite. Continue the run above and look at the step where the loss first reports nan:

parameters entering step 3:
  features.0.weight    finite=True  max|.|=7.6892e+20
  features.0.bias      finite=True  max|.|=1.2599e+21
  features.2.weight    finite=True  max|.|=2.0724e+21
  features.2.bias      finite=True  max|.|=7.5346e+12
  head.weight          finite=True  max|.|=8.0718e+20
  head.bias            finite=True  max|.|=1.5646e+03

  input          finite=True  inf=    0 max|.|=4.1015e+00
  features.0     finite=True  inf=    0 max|.|=4.3012e+21
  features.1     finite=True  inf=    0 max|.|=2.2211e+21
  features.2     finite=False inf=  783 max|.|=inf
  features.3     finite=True  inf=    0 max|.|=4.8522e+34
  head           finite=False inf=  208 max|.|=inf
  first non-finite stage: features.2

Every parameter is finite. The first non-finite value appears inside the forward pass, at features.2, where a matrix of magnitude 2e21 meets activations of magnitude 2e21 and overflows float32.

Then look at features.3. It is finite again. All 783 infinities were negative, and ReLU(-inf) is 0 — the ReLU swallowed the evidence, and it reappears two stages later in the logits. A check placed only on the loss would have reported a NaN three stages downstream of the first overflow, which is itself three steps downstream of the update that caused it. Six boundaries between the symptom and the cause, and every one of them is cheap to look at once you decide to look at the first one instead of the last.

def assert_finite(name, t):
    if not torch.isfinite(t).all():
        raise RuntimeError(f"{name} first became non-finite")

Explicit staged checks like that are usually enough for forward-pass problems. When the non-finite value first appears during backward() rather than forward(), Chapter 3 covers torch.autograd.detect_anomaly(), which associates the failing backward computation with the forward operation that created it. It is a debugging tool with real overhead; do not leave it enabled.

The learning rate, and what the optimizer actually holds

Now — after task, objective, dependency, gradient, ownership and movement all have evidence — a learning-rate comparison means something. The rule is that every candidate starts from the same model state, the same batch and the same number of steps:

base = build(42)
state = copy.deepcopy(base.state_dict())

for lr in [1e-5, 1e-4, 1e-3, 1e-2, 1e-1, 1.0, 10.0]:
    m = TinyClassifier()
    m.load_state_dict(state)
    ...
        lr   loss @0  loss @100  acc @100  total delta
     1e-05    0.6942     0.6940    0.4883   3.4794e-04
     1e-04    0.6942     0.6930    0.5156   3.4668e-03
     1e-03    0.6942     0.6825    0.7812   3.4065e-02
     1e-02    0.6942     0.5925    0.7227   3.1063e-01
     1e-01    0.6942     0.0949    1.0000   2.1698e+00
     1e+00    0.6942     0.0162    1.0000   4.3333e+00
     1e+01    0.6942     1.4909    0.5039   2.2257e+02

The identical loss @0 in every row is the evidence that this is a controlled comparison. Comparing different random initializations and calling it a learning-rate experiment produces a table that looks exactly like this one and means nothing. Chapter 14 owns the general question of when two runs are comparable at all; here the requirement is narrow — same state, same data, same steps.

The total delta column is the interpretable one. At lr=1e-5 the entire model moved by 3.5e-04 in a hundred steps. That is a training run that is not stuck; it is a training run that is too slow to see, and the two require different responses.

Finally, do not read the learning rate off the configuration file, and do not assume param_groups[0] speaks for the optimizer:

group 0: lr=0.0001 weight_decay=0 params=4
group 1: lr=0.1 weight_decay=0.01 params=2
after step 0: lrs = [1e-05, 0.010000000000000002]
after step 1: lrs = [1.0000000000000002e-06, 0.0010000000000000002]
after step 2: lrs = [1.0000000000000002e-07, 0.00010000000000000003]

Two groups with a thousand-fold difference in learning rate and different weight decay, and a StepLR stepped once per batch instead of once per epoch has taken the feature learning rate from 1e-4 to 1e-7 in three steps. After the first scheduler step, param_groups[0]["lr"] correctly reports 1e-05 for group 0. The mistake is treating that one value as the optimizer’s learning rate: it says nothing about the head group, which is simultaneously at 0.01.

Inspect the optimizer’s current state, group by group. Not the learning rate you remember configuring.

Loss, accuracy and what counts as learning

One more measurement trap, because it sends people down the wrong path regularly. Watch the reference model under plain SGD for forty steps:

step      loss  accuracy  mean margin
   0   0.69416    0.4766     -0.00160
   5   0.68821    0.7266      0.01023
  10   0.68253    0.7773      0.02178
  20   0.67173    0.7109      0.04443
  30   0.66136    0.6836      0.06704
  40   0.65142    0.6680      0.08948

Loss falls monotonically. The mean margin between the correct and incorrect logit rises monotonically. Accuracy goes up, then down, from 0.7773 to 0.6680. Nothing is wrong. Accuracy is a step function of the argmax, so it changes only when an example crosses the decision boundary, and early in training a model can improve its average confidence while a handful of examples move the wrong way.

Loss and accuracy answer different questions. A flat accuracy is not evidence of a flat model, and a falling accuracy over a few steps is not evidence of a broken one. Read the loss and a task-appropriate metric together.

The same caution applies to individual steps. Do not assert that every optimizer.step() must reduce the current batch’s loss — a stochastic or adaptive optimizer can take steps that do not, for reasons that have nothing to do with a bug. The one-step invariant established in this chapter is parameter movement under a controlled update; useful loss reduction is a multi-step capability claim. Keep them separate.

Where this chapter stops

Two boundaries are deliberately left closed.

If a model trains correctly in full precision and fails only under mixed precision, this chapter’s job is already done: establish that the same controlled problem learns in float32, then hand the mixed-precision execution path to Chapter 12, which owns autocast, gradient scaling and their interaction with performance. This environment has no CUDA device, so no claim is made about them here.

If the model learns but a new run is worse than an old one, that is a different question with a different method. Chapter 11 uses fixed seeds, a fixed batch and a frozen initial state because a failure that cannot be reproduced cannot be investigated. It does not follow that two full training runs are comparable, and Chapter 14 owns that: run-to-run variation, what changed in the configuration or environment, and whether a validation difference is a regression or ordinary run-to-run variation.

When a failure is intermittent, preserve the evidence rather than trying to reproduce it from memory:

if not torch.isfinite(loss):
    torch.save({"x": x_batch.detach().cpu(), "y": y_batch.detach().cpu(),
                "model": model.state_dict(),
                "optimizer": optimizer.state_dict(), "step": step},
               "failing_batch.pt")
    raise RuntimeError("non-finite loss")

A saved batch converts a rare failure into a deterministic test case, which is the state in which every technique in this chapter applies.

Using AI on a model that will not learn

An assistant asked “why is my model not learning” will produce a plausible list, because a plausible list is what the question invites. Every item on it will be a real cause of a real failure somewhere. None of them will be evidence about your system.

The prompt below works because it refuses the repair and asks for the localization. It is long deliberately: the structure is the value.

My PyTorch model runs but does not learn.

Do not rewrite the model and do not suggest hyperparameters yet.

I will provide evidence from one fixed batch.

Build a learning-chain table with these boundaries:

1. TASK
   - what does one input represent?
   - what does its target mean?
   - what evidence says they are still aligned?

2. OBJECTIVE
   - logits shape and range
   - exact loss function and target representation
   - manual one-example verification

3. DEPENDENCY
   - loss.requires_grad and grad_fn
   - the full list of parameters I expect to be trainable

4. GRADIENT
   - for each expected parameter:
     grad None / exactly zero / finite / non-finite, and grad norm

5. OPTIMIZER OWNERSHIP
   - trainable model parameter ids
   - optimizer parameter ids
   - model parameters missing from the optimizer
   - optimizer parameters no longer in the model

6. UPDATE
   - parameter norm before
   - parameter delta after one step
   - the current learning rate for EVERY parameter group

7. CAPABILITY
   - tiny fixed-batch initial loss
   - loss after controlled repeated steps with the REAL optimizer
   - final tiny-batch accuracy

Identify the FIRST boundary whose evidence contradicts the intended
training system.

For that boundary:
- give me the two or three most plausible mechanisms;
- propose the smallest experiment that distinguishes them;
- say which measurement should change if the diagnosis is correct;
- say which measurement must stay unchanged.

Do not propose architecture changes, learning-rate tuning, gradient
clipping, mixed precision, more data or another optimizer until every
earlier boundary has passed.

Three clauses earn their place. Asking for the first boundary rather than a list prevents a plausible downstream fix from being applied to an upstream cause — the UPDATE row failed in two of this chapter’s four ledger runs, with nearly identical text and two completely different causes. Asking what must stay unchanged is what makes the proposed experiment falsifiable rather than merely confirmatory. And asking for the tiny-batch test with the real optimizer closes the loophole that a fresh optimizer silently repairs the exact bug being investigated.

The last paragraph is the one to keep. An assistant given a symptom will reach for the interventions at the end of the chain, because those are the ones most discussed in its training data. Learning rate, optimizer choice, clipping and architecture are all downstream of six boundaries that are cheaper to check and more likely to be broken.

Ask AI to locate the first missing consequence before asking it to repair training.

The learning-chain debugging sequence

 1. Freeze one known batch and a known initial state. Make the failure
    repeatable before trying to explain it.

 2. Prove the TASK. Does each target still describe its input? Use a probe
    you built on purpose; no tensor statistic can see this.

 3. Prove the OBJECTIVE. Are the model's outputs in the form this loss
    expects? Reproduce the loss by hand for one example.

 4. Prove the DEPENDENCY. Which of the parameters you intend to train does
    this loss actually reach? loss.grad_fn is not that evidence.

 5. Prove the GRADIENT. Are the expected gradients present and finite?
    Distinguish None from exactly zero from non-finite, and note where in
    the model each occurs.

 6. Prove OWNERSHIP. Does the optimizer hold those exact parameter objects?
    Check both directions: missing from the optimizer, and stale in it.

 7. Prove the UPDATE. Snapshot, step, measure deltas. Read the current
    learning rate and weight decay for every parameter group.

 8. If the update semantics are still unclear, reproduce one step on a copy
    with fresh SGD (momentum 0, weight decay 0) and verify d(theta) = -lr * grad.

 9. Prove CAPABILITY. Strip augmentation, dropout, weight decay, smoothing
    and schedulers, then overfit a tiny fixed batch — with the REAL
    optimizer, not a fresh one.

10. If that still fails, escalate: per-layer activation statistics including
    across-batch variation, gradient versus update scale, and the first
    non-finite stage.

11. Only once the controlled system learns should you reintroduce
    regularization, augmentation, schedulers, mixed precision and the full
    dataset — one at a time.

12. Change one thing. Rerun from the first boundary that change could affect.

Steps 2 through 4 catch what no gradient inspection can. Steps 5 through 7 catch the opening failure. Step 9 catches what nothing before it can, and proves less than people think it does.

What you should now be able to answer

Here is a ledger-style summary of the kind you will be handed. It is a constructed scenario, not an experiment from this chapter. Work through it before reading on.

loss 0.6931 at step 0, 0.6929 at step 500
accuracy 0.507 throughout
all 14 parameters: requires_grad True, grad present, grad finite
total grad norm 4.1e-06
optimizer: AdamW, 1 group, lr 3e-4, 14 params, 0 missing, 0 stale
all 14 parameters: nonzero delta after one step
tiny batch (32 examples, real optimizer, 2000 steps): 0.6931 -> 0.6902

Which boundary is the first one without support? None of A through F. Task, objective, dependency, gradient, ownership and update all report evidence consistent with a healthy system. The first unsupported link is capability: the model cannot fit 32 examples. That is where the investigation continues, and everything above it has been ruled out cheaply.

Is 4.1e-06 a small gradient norm? Unanswerable as stated, and this is the trap. There is no architecture-independent scale. It becomes evidence when compared: against the same measurement on a known-good run, against the norms at other layers, or against its own trajectory over steps.

The loss sits at 0.6931 and accuracy at 0.507. What does that tell you? That the model is producing near-uniform predictions on a balanced two-class problem, since ln(2) = 0.6931 and the majority baseline is 0.507. It is a precise description of the symptom and says nothing about the cause. A stalled optimizer, a broken graph, a misaligned target relation and a learning rate three orders of magnitude too small all pass through this exact state.

Every parameter has a nonzero delta. Does that prove the optimizer is working correctly? No. It proves the objects moved. Under AdamW the delta ranged from 18 to 212 times lr * ||grad|| in this chapter’s measurements, and decoupled weight decay contributes movement that does not come from the gradient. To check that a specific gradient produces a specific update, use a copy and a fresh plain SGD.

The tiny-batch test passes when you rebuild the optimizer. Is the system fixed? No — you have changed the system under test. That is precisely how the opening failure hides: 0.68758 unchanged with the real optimizer, 0.148 with a fresh one. The gap between those two numbers is the diagnosis.

Training loss falls and validation loss does not. Is that this chapter’s problem? No. Every boundary here is about whether the mechanism works, and it does. A train/validation gap is a question about generalization and about whether the two measurements are comparable, which belongs with Chapters 7 and 14.

The loss plateaus near 0.3133 and accuracy is 100%. What happened? In this chapter’s two-class experiment, pre-softmaxing before cross entropy produces exactly that limiting floor, so it is a strong clue to inspect the objective boundary. It is not a unique fingerprint in an unfamiliar system: verify what tensor is actually being passed to the loss before naming the cause.

A NaN appears in the loss at step 40. Where do you look? Not at step 40’s loss. Trace the forward pass stage by stage for the first tensor that is non-finite, then check whether the parameters entering that step were already enormous, then check the update scale on the step before. This chapter’s example had three stages and three steps between the cause and the NaN — and a ReLU that zeroed the infinities in between, so the non-finite value briefly disappeared and then came back.

What single piece of evidence separates “not learning” from “learning too slowly”? There is no universal single measurement. Parameter movement separates a stationary model from one that is changing, and a controlled loss trajectory tells you whether that movement is useful. In this chapter’s lr=1e-5 run, the model moved by only 3.5e-04 over 100 steps while every mechanical boundary passed and the loss changed only slightly; together those observations support “learning too slowly.” A zero delta localizes a different kind of failure, but nonzero movement alone does not prove useful learning.

Exercises

  1. Gradient but no update. Reproduce the opening: freeze the feature extractor, construct the optimizer, then replace the head. Prove that the new head appears in named_parameters(), requires gradients, receives a finite gradient, is absent from the optimizer, and has a parameter delta of exactly zero. Then measure how its .grad behaves over five iterations and explain why optimizer.zero_grad() did not clear it. Repair only the optimizer construction.

  2. Chapter 1 at model scale. On one fixed batch, using a copy of the model and a fresh SGD(lr, momentum=0, weight_decay=0), verify that every parameter satisfies Δθ = -lr · grad to floating-point tolerance. Then repeat with momentum=0.9, and with AdamW, and record the ratio of ||Δθ|| to lr·||grad|| for each parameter. Explain why the equality is not expected to survive either change.

  3. Detach one edge. Insert detach() between two layers. Identify the first parameter whose gradient becomes None, and confirm that loss.requires_grad and loss.grad_fn are both healthy throughout. Then measure the final loss against the reference run and argue why this failure is harder to notice than a complete stop.

  4. Zero is not None. Build a controlled dead-ReLU case. For each parameter, record whether .grad is None, an exactly-zero tensor, or nonzero, and explain each result from the local derivative. Then explain why one bias in the dead region still receives a gradient.

  5. Destroy the alignment. Permute the inputs without permuting the labels. Show that shape, dtype, range, mean, standard deviation and class counts are all unchanged, that the known-rule probe is the only check that fails, and that the full training system remains mechanically healthy. Report held-out accuracy on the real task and compare it with the trivial baseline.

  6. Wrong objective, right mechanism. Compare cross entropy on logits with cross entropy on already-softmaxed probabilities from an identical initialization. Report the loss trajectory, the step-0 gradient norm and the final training accuracy for both. Derive the floor of the broken objective analytically for C classes, and verify it numerically for C = 2 and C = 10.

  7. Two tiny-batch tests. Run the capability test on a broken system twice: once with the real optimizer and once with a fresh one. Explain why the two answers differ and which question each one answers. Then run it on labels drawn from randint and report how many steps memorization takes.

  8. Simplify one mechanism at a time. Take a model with dropout, weight decay, label smoothing and input noise that will not overfit 32 examples. Remove one mechanism per run, always from the same initial state, and record the final loss. Identify which single mechanism accounts for most of the gap, and state what that does and does not tell you about the full training recipe.

  9. One-batch forensics on an unfamiliar model. Run the learning ledger against a small model you did not write. Identify the first unsupported link before changing anything. Then deliberately introduce a second, later failure, rerun, and confirm the ledger still reports the earlier one.

Next: it learns, but is it using the machine?

Learning is no longer a thing that either happens or does not. It is a chain of observable consequences, and every link leaves evidence on a single batch:

TASK        the target still describes the input        known-rule probe
OBJECTIVE   the scalar is the loss we intended          one-example derivation
DEPENDENCY  the loss reaches the intended parameters    grad present, not None
GRADIENT    those parameters received real derivatives  None / zero / finite
OWNERSHIP   the optimizer holds those exact objects     two-way id() audit
UPDATE      those exact objects moved                   snapshot and delta
CAPABILITY  repeated controlled steps fit a tiny batch  real optimizer, no regularization

Five deliberate failures stressed different parts of that chain, but they were not all discoverable by the same instrument. The stale head satisfied everything up to ownership and never moved. The detached layer reached the gradient boundary, where four expected gradients disappeared even though the fixed feature representation still supported partial learning. The shuffled targets failed at the task boundary and then optimized nonsense competently. The pre-softmaxed objective passed the generic mechanical ledger because deciding whether a differentiable scalar is the intended objective requires an objective-specific check. The absurd learning rate also passed the one-step mechanical ledger; its failure appeared in update scale and only later as non-finite values. That division of labor is the point: the chain tells you which evidence to demand, not that one generic checker can certify every link.

Not one of them raised an exception.

Learning is a chain of observable consequences. A correct target should produce the intended loss; that loss should reach the intended parameters; those parameters should receive finite gradients; the optimizer should own those exact objects; the step should move them; and repeated controlled steps should change the model’s behavior. Find the first consequence that fails, and debug there.

We can now establish that a known sample means what its target says, that the intended objective depends on the intended parameters, that gradients reach them, that the optimizer owns those exact objects, and that those objects move. On a controlled copy with plain SGD we can go further and verify the update equation itself. Finally, repeated controlled steps can establish that the training system is capable of fitting a small batch.

That still says nothing about whether the machine is being used well. A training loop can be entirely correct and waste most of a GPU. It can fit in memory and spend half its time waiting on synchronization. It can be slower after compilation than before it. None of the instruments in this chapter would notice, because all of them measure whether the right thing happened and none of them measures what it cost.

The next chapter starts from a model whose learning we have proved, and asks a completely different question: where is the time going, where is the memory going, and why is the device waiting? Its governing idea is as simple as this chapter’s and just as easy to skip:

Measure first. Optimize second. Measure again.