← PyTorch From First Principles

Regressions: Did the Model Change, or the Measurement?

Establish whether a regression exists before explaining it by measuring baseline variation, using paired comparisons when valid, fingerprinting the evaluation path, controlling randomness and determinism, bisecting interacting changes, and recording runs so model changes can be separated from measurement changes.

Here is a refactor. The old training script set the seed at the top; the new one moves that line below model construction, where it reads more naturally, and draws the batch-shuffle permutation from a local torch.Generator() instead of the global RNG, so the shuffling is “self-contained.” The training procedure is mathematically the same and the model architecture is unchanged. But the concrete initialized parameters are no longer coupled to the old run, because model construction now consumes randomness before manual_seed is applied. Run it once:

baseline validation loss   0.9726
after the refactor          0.9980

The refactor cost 0.026 validation loss. That is the sentence an engineer writes in the pull request, and it is not yet supported by anything.

Run the unchanged baseline at twelve different training seeds:

seed  0  val_loss 0.9726        seed  6  val_loss 0.9757
seed  1  val_loss 0.9847        seed  7  val_loss 0.9598
seed  2  val_loss 0.9762        seed  8  val_loss 0.9871
seed  3  val_loss 0.9778        seed  9  val_loss 0.9771
seed  4  val_loss 0.9690        seed 10  val_loss 0.9706
seed  5  val_loss 0.9731        seed 11  val_loss 0.9754

median 0.9756   stdev 0.0071   min 0.9598   max 0.9871   spread 0.0273

The spread of runs that differ only in their seed is 0.027. The refactor’s 0.026 is inside it. Run the refactored code at twelve seeds and it produces median 0.982, range [0.959, 0.999] — a distribution that overlaps the baseline’s almost entirely. Moving the manual_seed line changed which random numbers reached the model’s initialization, so every run is different; whether the distribution shifted at all is not resolved by twelve runs. The “regression” was a single sample from a noisy process, reported as a fact. There was a number, and there was a story, and the story arrived first.

Now the other half. Take one trained model — fixed, saved, not retrained — and evaluate it four ways, changing only the evaluation code:

  evaluation variant                           val loss   val acc
  canonical                                      0.9726    0.6760
  model left in train() mode (dropout on)        1.0711    0.6290
  a different 500-example validation subset       1.0348    0.6580
  inputs scaled x1.15 on the eval path only       1.0881    0.6330

The model is byte-for-byte identical in every row. Every number moved — some by four times the seed noise — and nothing regressed. The meaning of the number changed.

A metric that moved is not yet a finding. It is a symptom with at least three explanations, and only one of them is about the model.

Where we are

Chapters 11, 12 and 13 all investigate inside a single run. Is the model learning — which link in the chain is broken? Where is the time and memory going? What did the compiler capture, and what invalidated it? Every one of those questions has an answer you can extract from one execution.

This chapter is the first whose unit of analysis is a pair of runs. That is its entire reason to exist. A regression is not a property of one run; it is a claimed difference between two, and the claim is only as good as the comparison behind it.

Chapter 13 handed this over precisely: compilation adds dimensions along which two runs can differ invisibly — a different sequence of shapes seen, a warm cache in one run and cold in the other, the recompile limit reached this week and not last week. None of that raises. None of it shows in a unit test. The question this chapter answers:

When this run is worse than that one — did the model change, did the measurement change, or did nothing change and the difference is noise?

Three explanations, one symptom:

The technique

Every chapter earns an investigation method. Chapter 11: prove the chain one boundary at a time. Chapter 12: measure, localize, intervene, remeasure. Chapter 13: find the first assumption that stopped holding.

Chapter 14 adds:

Establish the noise floor before attributing a cause. Before explaining a difference, prove there is one.

The hidden structure this makes visible is the causal chain behind any reported number:

code  +  configuration  +  data  +  randomness  +  environment  +  measurement
                                    ↓
                                 result

A regression means something in that chain changed enough to move the result. The investigation is to find which link. The link that people reach for first is code, because that is the one they were working on. The links that are actually broken are usually measurement (the benchmark changed) or randomness (nothing changed; it is a different seed). And measurement is a link like any other — a changed evaluation path and a changed model produce exactly the same symptom.

The environment these numbers came from

Every number, table and trace in this chapter was produced by running the code shown, on one machine:

Python   3.11.4
PyTorch  2.6.0+cu118
OS       Windows 11
CPU      24 cores; torch.set_num_threads(1) for every training run in the chapter
GPU      present but used only for the cross-device comparison in E3

This chapter is almost entirely CPU work, and cheaply so. The controlled workload trains in about 0.7 seconds; a twelve-seed sweep is under fifteen seconds. Every experiment below, including the development runs that tuned the workload, is roughly fifteen minutes of CPU in total. There is no honest excuse for an unmeasured number in a chapter about not trusting unmeasured numbers.

One thread is the canonical setting, chosen deliberately: on one thread, with the data fixed and the seed fixed, two runs of this workload are bit-identical, which makes every later difference attributable to something. It is also, for a model this small, the fastest option.

One controlled workload

A single synthetic classification task, used for every experiment.

DEFAULT_CONFIG = {
    "n_train": 3000, "n_val": 1000, "in_dim": 32, "hidden": 32, "classes": 4,
    "noise": 0.30,          # 30% of labels randomised: val loss has a real floor
    "data_seed": 0,         # the DATA is fixed; only the training seed varies
    "epochs": 20, "batch_size": 128, "lr": 0.02,
    "weight_decay": 3e-3, "dropout": 0.3, "optimizer": "adamw",
}

The task is a random linear rule with 30% label noise, so no model can drive validation loss to zero — there is a floor to sit at, around 0.97. The model is a three-layer MLP. The data is generated once from data_seed and never varies; the seed passed to a run controls only initialization, batch order and the dropout mask. That separation is what makes “run it again at a different seed” a clean experiment.

The shared instruments are in explab.py: train_run(seed, config) returning a run record, seed_sweep, a median/spread summariser, paired and unpaired comparison functions, and an eval-path fingerprint. Every table in this chapter is reproducible from them.

E1 — The noise floor

This is the first experiment, before any change is evaluated, and skipping it is why most reported regressions cannot be argued about.

QUESTION   how much do two runs of the IDENTICAL configuration differ, from the
           training seed alone?
METHOD     run DEFAULT_CONFIG at 12 seeds; report the distribution.
  identical config, 12 seeds
    val loss    median 0.9756   stdev 0.0071   min 0.9598   max 0.9871   spread 0.0273
    train loss  median 0.8911   stdev 0.0098   spread 0.0370
    val acc     median 0.6795   stdev 0.0065   spread 0.0200
    examples/s  median 196143                  spread 75794

Across these twelve observed seeds, validation loss spans 0.0273 and has standard deviation 0.0071. That is an empirical scale for ordinary seed-to-seed variation in this experiment, not a universal detection threshold. A single-run difference of 0.026 is therefore not compelling evidence of a regression: it is the same size as variation already produced by the unchanged configuration. Establishing a smaller systematic shift requires repeated runs — preferably paired when the pairing remains meaningful.

This baseline distribution is the yardstick for the rest of the chapter. It is a property of this workload under this protocol — a bigger model, a real dataset, a different environment or a different metric needs its own baseline distribution. The throughput spread is enormous here — nearly 40% — because this desktop’s CPU load swung during the sweep; that is the observed variation of whole-run throughput under this measurement protocol, not a threshold for every performance benchmark on the machine. Chapter 12 deliberately used tighter timing windows and repeated measurements for that reason. The validation-loss path, by contrast, is bit-identical when seed and execution conditions are held fixed in E2; across E1 the variation comes from deliberately changing the training seed.

When a candidate’s single-run delta is comparable to ordinary baseline variation, the regression has not been established. Compare distributions or paired deltas before explaining it.

E2 — Same seed, same code, twice

What does setting the seed actually buy?

  run A (1 thread) vs run B (1 thread), same seed:
    val curves bit-identical: True
    final val loss  A 0.972560   B 0.972560

On one thread, with the data, code and seed fixed, these two executions produce bit-identical validation curves. That is a strong reproducibility result for this controlled workload under these recorded conditions, and it is what makes later interventions interpretable.

Now change one thing — the thread count — and nothing else:

E3 — Nondeterminism you can execute

The thread-count difference in E2 has one mechanism, and it is worth stating plainly because it explains a large fraction of what people call flakiness.

Floating-point addition is not associative.

  (a + b) + c = 0.0        with a = 1.0, b = 1e16, c = -1e16
  a + (b + c) = 1.0

That is not a pathology invented for this chapter; it is a consequence of finite-precision arithmetic. Floating-point sums can change when their accumulation order changes. Sum the same 200,000 float32 values three ways:

  sequential, forward order  : -730.639587
  sequential, backward order : -730.646729
  pairwise (torch.sum)        : -730.646667
  max gap: 7.14e-03

A gap of 7e-3 from accumulation order alone. Put that effect inside training, where reductions feed later floating-point operations, and a small numerical difference can propagate. Change the thread count:

torch.use_deterministic_algorithms(True)

This forces deterministic implementations where they exist and raises where they do not:

torch.use_deterministic_algorithms(True)
x = torch.randn(2, 3, 8, 8, device="cuda", requires_grad=True)
grid = torch.rand(2, 8, 8, 2, device="cuda") * 2 - 1
torch.nn.functional.grid_sample(x, grid, align_corners=False).sum().backward()
RuntimeError: grid_sampler_2d_backward_cuda does not have a deterministic
implementation, but you set 'torch.use_deterministic_algorithms(True)'. You can
turn off determinism just for this operation, or you can use the 'warn_only=True'
option, if that's acceptable for your application.

That is the feature working. Deterministic algorithms can also change performance, and some CUDA matrix operations require CUBLAS_WORKSPACE_CONFIG (:4096:8 or :16:8) for deterministic execution or PyTorch raises.

What the setting does not provide is universal bit identity across environments:

E4 — Worker seeding

DataLoader workers are Chapter 6’s subject. The reproducibility question is narrower: with num_workers > 0, which randomness in the pipeline is tied to the run’s seed, and which floats free? Put a draw from four different RNGs inside __getitem__, run the loader twice from the same torch.manual_seed(0), and compare:

                              num_workers=0   num_workers=2   with the repair
  data order (shuffle)            reproduces      reproduces      reproduces
  torch.rand in __getitem__       reproduces      reproduces      reproduces
  np.random.rand (legacy global)  DOES NOT        reproduces      reproduces
  random.random                   DOES NOT        reproduces      reproduces
  np.random.default_rng() object  DOES NOT        DOES NOT        DOES NOT

The surprising result is version- and implementation-specific, so keep its boundary visible. In this PyTorch 2.6 experiment, the multi-process worker path caused the legacy NumPy global RNG and Python’s random stream to reproduce, while the single-process num_workers=0 path did not. An independently created np.random.default_rng() object remained outside that mechanism in every case.

For portable code, do not rely on this observed worker-startup side effect. Pass an explicit generator= to the DataLoader; seed external libraries explicitly in worker_init_fn when workers exist; and, when num_workers=0, seed those external RNGs in the main process because worker_init_fn is never called. A default_rng() object must itself be constructed or reseeded from a controlled seed.

Reproducibility of pipeline randomness depends on the exact RNG object that produces it. Control that object explicitly rather than assuming the DataLoader will discover it.

E5 — Paired comparison

E1 gave the noise floor. When a change’s real effect is comparable to it, how the two configurations are compared decides whether the effect is visible at all.

Configuration B lowers the validation set’s label noise from 0.30 to 0.29 — a “the data got slightly cleaner” change, the kind a dataset revision makes. Run A and B, twelve seeds each.

Unpaired — summarize the two sweeps without using the seed correspondence:

Paired — same seed on both sides, compare per-seed:

  seed   0:  0.9726 -> 0.9456   -0.0270
  seed   1:  0.9847 -> 0.9532   -0.0315
  seed   2:  0.9762 -> 0.9692   -0.0070
  seed   3:  0.9778 -> 0.9353   -0.0425
  seed   4:  0.9690 -> 0.9587   -0.0102
  seed   5:  0.9731 -> 0.9864   +0.0133
  seed   6:  0.9757 -> 0.9606   -0.0151
  seed   7:  0.9598 -> 0.9541   -0.0057
  seed   8:  0.9871 -> 0.9605   -0.0265
  seed   9:  0.9771 -> 0.9605   -0.0166
  seed  10:  0.9706 -> 0.9268   -0.0437
  seed  11:  0.9754 -> 0.9450   -0.0304
  mean per-seed delta: -0.0203   stdev of deltas: 0.0166
  11/12 seeds improved   |mean| / stdev = 1.2

Eleven of twelve paired deltas are negative. Using the same seed on both sides creates a common-random-number comparison: when the two runs consume randomness in corresponding ways, seed-specific difficulty is partly shared and the variance of the difference can be much smaller than the variance of the raw outcomes.

This chapter deliberately stops short of turning that pattern into a formal significance claim. |mean| / stdev = 1.2 is a descriptive effect-to-spread ratio, not a hypothesis test or confidence interval. The evidence is directional and consistent enough to justify the next experiment; if the decision matters, predeclare an inferential criterion or collect more paired seeds rather than converting one descriptive ratio into “proof.”

Pairing is strongest when a shared seed induces comparable randomness on both sides. A refactor that changes RNG consumption — an extra draw, a reordered stochastic pipeline, a different sampler — can weaken that coupling even though the seed numbers still match. In that case inspect the paired-delta variance itself; if the coupling disappeared, the paired design may buy little over an unpaired comparison.

E6 — The broken benchmark

The MEASUREMENT-CHANGED branch. Take one trained model, fixed, and evaluate it while changing only the evaluation path:

  evaluation variant                        val loss   val acc   fingerprint    == canonical?
  canonical                                   0.9726    0.6760   7176cf3fce66   -
  model in train() mode (dropout on)          1.0711    0.6290   f8dce02c6c6b   False
  different validation subset (500 of 1000)   1.0348    0.6580   0e6c39db2124   False
  inputs scaled x1.15 on eval path only       1.0881    0.6330   9fb90744a142   False
  cross_entropy reduction='sum'            1074.8778      -      22aacc4965b5   False

Every reported metric moved, three by much more than the baseline seed-to-seed variation and the last by a factor of a thousand. The last value is simply the mean loss summed over one thousand examples, but without the aggregation contract it looks like the model exploded. The model parameters are identical in every row; the evaluation procedure is not.

The defence is an eval-path fingerprint over the components this experiment has chosen to make part of the measurement contract:

Leaving evaluation in train() mode is the specific case the previous chapter’s semantics — eval() versus no_grad() — exist to prevent; here it is a cause whose symptom is a 0.10 metric shift, and the fingerprint is what surfaces it.

E7 — A real regression, diagnosed

Now the MODEL-CHANGED branch, with a genuine cause: the learning rate is ten times too high.

Detect — seed sweeps:

  baseline    median 0.9744   range [0.9598, 0.9847]
  lr x10      median 1.3900   range [1.3880, 1.4082]
  shift = +0.4156   noise floor spread = 0.027   ->  15x the noise floor

The ranges do not come close to overlapping; the shift is fifteen times the noise floor. Detection is trivial. The work is the explanation.

Explain 1 — learning curves and gradient norms, one seed:

  epoch          0       3       6       9      12      15      18
  baseline   1.016   0.996   0.969   0.965   1.003   0.978   0.984
  lr x10     1.393   1.400   1.388   1.392   1.399   1.393   1.401
  baseline gradient norm: median 0.510   max 0.983
  lr x10   gradient norm: median 0.084   max 6.581

The broken run’s validation loss is flat from the first epoch and its gradient-norm distribution is dramatically different from baseline: much smaller in the median with occasional large spikes. Those observations tell us when the trajectories separated and give us a numerical symptom to pursue. They do not, by themselves, prove that “most units saturated” or that the step size “destroyed the parameters.” To make that mechanistic claim, inspect the corresponding activations/parameter updates or run a targeted intervention such as restoring the learning rate from the same initialization. The chapter’s own rule applies here too: do not let a plausible story outrun the measurement.

Explain 2 — a different cause, per-example evidence. The training labels are shifted by one in the loader, so every input is paired with the next input’s label:

  baseline val loss  0.9726        shifted val loss  1.4034   (+0.4308)
  examples correct before, wrong after: 528 / 1000
  shifted-model accuracy: 0.2320   (baseline 0.6760)

Accuracy 0.232 is close to the 0.25 chance rate on four classes; being slightly below chance on one finite validation set is not evidence by itself that the model learned an anti-rule. What localizes this failure is the known intervention — labels were shifted — together with the large aggregate degradation and the per-example flips. The model has been trained against a broken input-target relation, so the data pipeline is the first boundary to inspect. This is Chapter 11’s shuffled-target failure arriving as a regression.

The aggregate metric detects a change. Curves, per-example failures and numerical signals narrow the mechanism — but none should be asked to prove more than it measured.

E8 — Curves, not endpoints

Two configurations, baseline versus a lower learning rate run for twice as long. Their final validation losses differ by less than the baseline single-run spread:

  A: 1.02 0.99 1.00 1.00 1.00 0.97 0.97 0.98 0.97 0.96 0.96 0.98 1.00 0.98 ...
     min 0.9648 at epoch 10, final 0.9726
  B: 1.13 0.98 0.97 0.97 0.95 0.95 0.96 0.96 0.96 0.95 0.95 0.95 0.96 0.95 0.94 0.94 ...
     min 0.9433 at epoch 22, final 0.9824

B descends to 0.943 by epoch 22 and then rises again. Across eight shared seeds, B’s lowest recorded validation loss is lower than A’s in all eight, by about 0.015 on average. That is interesting trajectory evidence, but comparing post-hoc minima deserves care: B runs for twice as many epochs and therefore gets more opportunities to produce a low validation point. If the real training procedure uses early stopping or best-checkpoint selection, define that checkpoint policy before comparing the configurations and evaluate the policy across seeds. The chapter’s durable finding is that equal-looking endpoints can hide materially different trajectories.

“When did the runs diverge?” is a question the final number cannot answer. If deployment selects a best checkpoint, make the checkpoint-selection rule part of the experiment rather than choosing the minimum after seeing the curves.

E9 — Bisection

“Change one thing at a time” is sound advice and useless after someone has already changed twelve. The candidate branch here makes six plausible improvements at once, and it regressed badly:

  baseline (0 changes)  val loss median 0.9747
  candidate (6 changes) val loss median 2.6609   (+1.69, 62x noise floor)

Order the six changes and binary-search on how many of them, applied as a prefix, are enough to cross the regression criterion:

When both the code and the data changed

If a commit changed the model and a dataset revision landed the same week, the four cells of a 2 × 2 experiment separate main effects from interaction:

E10 — A policy guardrail from baseline variation

A regression gate needs a decision boundary, but twelve baseline seeds do not magically turn one formula into a confidence bound. This chapter uses a deliberately conservative policy threshold derived from the observed baseline variability:

  new healthy run (seed 999):    val_loss 0.9924   threshold 1.0574   pass=True   margin +0.065
  new run with lr 10x (seed 999): val_loss 1.3902   threshold 1.0574   pass=False  margin -0.333

The illustrative test:

The run record

Every recent chapter has a reusable instrument: Chapter 7’s stage report, Chapter 10’s attention ledger, Chapter 11’s learning ledger, Chapter 12’s performance ledger, Chapter 13’s compilation ledger. Chapter 14’s is the run record — and it is the one that makes all the others comparable across time.

The comparison is run at a shared set of seeds so it can report the distribution and the per-seed delta, and it reports across several dimensions because a change routinely helps one and hurts another — the multi-dimensional habit from Chapter 12’s performance ledger:

RUN candidate-lr10x   BASELINE main-baseline   (10 shared seeds)

IDENTITY
  config hash      587cf3ca1f03  ->  91b44992b14c   CHANGED
  eval fingerprint 7176cf3fce66  ->  7176cf3fce66   SAME
  env hash         0838b7da7ddc  ->  0838b7da7ddc   SAME

CORRECTNESS   (median [min..max])
  val loss   0.9760 [0.9598..0.9871]  ->  1.3900 [1.3880..1.4082]
  val acc    0.6785  ->  0.2460
  paired val-loss delta: mean +0.4173  stdev 0.0121  (10/10 worse)
  -> 15x the noise floor

COST   (median)
  examples/sec  94178  ->  95149   (+1.0%)

DECLARED BEFORE THE RUN
  hypothesis: raising the learning rate speeds convergence without
              materially worsening validation loss
  success:    val-loss delta within the noise floor
  result:     NOT MET

Read top to bottom, the report narrows the comparison before debugging starts. The eval fingerprint is SAME, so none of the recorded evaluation-contract fields changed. The env hash is SAME, so none of the fields included in that environment record changed. The config hash changed, and the paired delta is enormous with every seed moving the same way. That combination is strong evidence that the candidate configuration changed model behavior rather than merely changing the recorded benchmark or environment.

A failed experiment recorded this way is not a wasted run. “Raising the learning rate destroyed convergence” is a real, durable fact about this system, worth more than a run quietly deleted because it did not confirm an intuition.

A worked pass

Put the sequence together on one regression. A commit reorganised the data loading, and validation loss on a single run went from 0.97 to 1.40.

1 — Baseline variation. The baseline configuration at eight seeds: median 0.974, range [0.960, 0.985], spread 0.025. This is the empirical reference for this version of the workload; if the data, evaluation contract or environment changes materially, re-establish it.

2 — Is the difference plausibly ordinary seed variation? 1.40 - 0.97 = 0.43, about 17× the observed baseline range. That is far too large to dismiss as the ordinary variation seen in step 1, so the investigation continues.

3 — Eval-path fingerprint. a4c1... == a4c1... on both runs. The recorded evaluation indices, preprocessing, aggregation and eval() state are identical. That rules out changes in those recorded fields; it does not prove no unrecorded measurement code changed.

4 — Environment. Same env_hash: the recorded Python/PyTorch/thread settings match. No recorded environment difference explains the shift.

5 — Pair it. Run the new code at the same eight seeds as the baseline. Every seed lands near 1.41; the paired delta is +0.44 ± 0.01, 8/8 worse. The effect is large and consistent.

6 — Now suspect model/data behavior. The recorded measurement and environment contracts match and the repeated comparison is unambiguous, so something in the changed training system genuinely moved the behavior.

7 — Curves. The new run’s validation loss is flat from epoch 1. It shows no useful improvement under this training recipe.

8 — Individual failures. Validation accuracy is 0.23, close to the four-class chance rate of 0.25. That number alone does not prove the model learned a coherent wrong rule. It tells us the degraded model is near chance on the real validation task, so per-example and data-contract evidence matter next.

9 — Numerical signals. Gradient norms remain finite, so there is no simple “backward is broken” explanation. That narrows the search; it does not identify the cause.

10 — Isolate. The commit diff has one change that touches label handling: an enumerate that now starts at 1. Reverting only that line restores 0.97.

11 — Test. Add an assertion that the training set’s known-rule agreement is 1.0 before the loop starts — Chapter 11’s contract promoted to a guardrail — plus the separately calibrated metric gate from E10 as a backstop.

Steps 3 and 4 are cheap places where the investigation could have ended as measurement or environment drift, and steps 1 and 2 determine whether repeated model-level diagnosis is worth paying for. The model/data investigation comes later, after the comparison itself has survived those checks.

Forensic mode versus everyday mode

Chapter 12 kept forensic instrumentation (hooks, the profiler, anomaly detection) separate from the production workload because the instruments cost time. The run record is the exception: it costs almost nothing — a few hashes and a JSON write — and its entire value is being there before you knew you would need it. Write it on every run. The forensic escalation in this chapter is not more instrumentation; it is more seeds — and the ten-second sweep is the reason that is affordable.

Using AI on a suspected regression

Given two numbers and a diff, an assistant will produce a fluent causal story — a learning-rate effect, a precision effect, a data effect — and it will do this whether or not the difference is real, because narrative explanation is what the request pattern-matches to. That fluency is precisely the hazard, and it is the same hazard this chapter teaches you to resist in yourself.

The prompt below refuses the explanation and demands the comparison.

Run B looks worse than run A. Do not propose a cause yet.

First reconstruct the COMPARISON:
    1. What exactly are the two runs -- code version, config, data version, seed?
  2. What is the baseline distribution of run A's configuration across several seeds?
     If I have not measured it, say so and stop -- that is the first missing
     piece of evidence.
  3. Is B - A large relative to that baseline variation? If A and B are single
     runs, the honest answer may be "not resolved"; a finite observed range is
     context, not a formal significance threshold.
  4. Were the seeds paired (same seeds on both sides), and did the code preserve
     enough RNG coupling for pairing to reduce variance? Report the paired deltas.
  5. Is the recorded evaluation contract identical -- same eval indices,
     preprocessing, aggregation, and model.eval() state? Compare fingerprints,
     while remembering they cover only the fields included in the hash.
  6. What else changed: code, config, data, environment (PyTorch / device /
     thread count), and compilation (shapes seen, recompiles, cache state)?
  7. Was the success criterion written before or after the result was seen?

Identify the FIRST link in {code, config, data, randomness, environment,
measurement} that lacks evidence, and stop there.

Do not propose a mechanism -- learning rate, precision, normalization, data --
until (a) repeated evidence establishes a difference worth explaining and
(b) the recorded evaluation contract has been checked.

Three clauses carry the weight. Demanding the seed sweep before anything else puts the noise floor first, where the chapter’s argument says it belongs. Asking whether the seeds were paired is what separates “cannot resolve” from “resolved with a sensitive comparison.” And forbidding a mechanism until the difference is established and the eval path is verified enforces the order: rule out NOTHING-CHANGED and MEASUREMENT-CHANGED before spending a day on MODEL-CHANGED.

Ask AI whether there is a difference before asking it to explain one.

The regression-debugging sequence

The early steps are intentionally cheap: many apparent regressions can be resolved as ordinary variation or a changed measurement contract before model-level forensics begin.

What you should now be able to answer

“Validation loss went from 0.97 to 1.01. The change hurt us.” The unchanged workload produced a 0.027 range across twelve seeds, so a 0.04 single-run gap is large enough to investigate but not enough to justify a causal claim by itself. Run both configurations at shared seeds and inspect the paired deltas.

“I set the seed, so the run is reproducible.” A seed controls one source of randomness. E2 showed this workload stopped being bit-identical when the thread configuration changed, and E4 showed that independent RNG objects require their own control.

“It’s nondeterministic, that’s just how GPUs are.” Floating-point non-associativity explains why changing reduction order can change numerical results, including many parallel reductions and atomic update patterns. cuDNN benchmarking is a different mechanism: benchmark noise can select a different algorithm on a later run. Name the actual source instead of calling every difference “GPU nondeterminism.”

“I changed twelve things and it’s better.” Which one? Bisect only when the good/bad predicate behaves suitably, then confirm the candidate alone; interactions can make a clean prefix boundary misleading.

“Accuracy dropped from 68% to 66%.” The unchanged workload showed a two-point min-to-max accuracy range across twelve seeds. That is context, not a formal threshold. Repeat the comparison — preferably paired — before naming a regression; if it is real, inspect which examples flipped.

“Both runs ended at 0.97, so they’re equivalent.” Similar endpoints do not establish equivalence. E8 showed trajectories can differ substantially before converging to similar final values. Compare curves and apply a predeclared checkpoint-selection policy if best checkpoints matter.

“A 5% drop should fail CI.” A round number is not automatically better than an empirical rule, but neither is median + 3 × observed range a statistical law. Calibrate the gate from baseline behavior, then validate its false-alarm rate on held-out healthy runs.

“The compiled version regressed.” Same eval fingerprint? Same shapes seen, same recompile count, same cache state (Chapter 13)? Compilation is a link in the causal chain, and it changes between runs more easily than the code does.

Exercises

These build on the current chapter’s “five deliberate regressions” idea — each one is a different meaning of “the new version is worse,” and the exercise is to tell which.

  1. Baseline variation. For a training workload of your own, run the unchanged configuration across multiple seeds. Report validation-loss, accuracy and throughput distributions. State what the observed variability tells you — and explicitly state why a finite min/max range does not define the smallest defensible single-run regression.

  2. Bit-identical, and where it stops. Run the same seed twice and confirm whether the full curve matches. Then change, one at a time: thread count, device, and PyTorch minor version if available. For each, report whether the run is bit-identical and how large any difference is relative to the baseline variation.

  3. Which RNG. Put a draw from torch, NumPy legacy RNG, Python random, and an independent np.random.default_rng() inside Dataset.__getitem__. Test num_workers=0 and 2. Then build an explicit seeding strategy for the main process, worker processes, and independent generator object rather than relying on worker-startup side effects.

  4. Unpaired versus paired. Make a subtle controlled change. Compare it once as two unpaired distributions and once using shared seeds. Report the actual paired-delta variance and whether pairing improved precision; do not design the exercise so a particular conclusion is guaranteed.

  5. The broken benchmark, five ways. Take one trained model and evaluate it under five deliberately changed measurement contracts. Record the fields that your fingerprint covers, and name at least one plausible evaluation change that your fingerprint would currently miss.

  6. Detect, then explain. Introduce a learning-rate intervention large enough to produce a reproducible effect. First establish the shift across seeds. Then use curves, updates, activations or gradients to test a specific mechanism rather than inferring one from gradient norm alone.

  7. Curves and checkpoint policy. Find two configurations with similar final metrics but different trajectories. Compare them under a checkpoint-selection or early-stopping rule declared before examining the test run; explain why post-hoc minima from unequal numbers of epochs are biased evidence.

  8. Bisect a diff. Construct a config change with six simultaneous edits, one of which is sufficient to regress the metric. Bisect an ordered predicate, confirm the suspected change alone, then add one interaction that violates monotonicity and show how the simple bisection story changes.

  9. A policy gate. Use baseline runs to propose a regression threshold, then test it on additional healthy seeds that were not used to choose the threshold. Report the observed false alarms before calling the gate calibrated.

  10. The run record, later. Save a run record and comparison report for one experiment. Reconstruct later what changed, what was declared beforehand, what evidence supported the verdict, and which causal fields your record failed to capture.

Next: the capstone

The book’s investigative toolkit is now complete. Tensor geometry and the first wrong value. The gradient path and where it breaks. Registered structure and the four things that can disagree about what the model owns. The input pipeline and where it waits. The representation contract and where a legal tensor stops being a meaningful one. Convolutional and attention geometry, derived rather than guessed. The learning chain, one boundary at a time. Performance, localized to a phase and an operator. Compilation, and the assumption that stopped holding. And now comparability — whether a difference between two runs is real, and which link in the chain produced it.

The capstone builds a small GPT-style language model from scratch. Its premise is that success is no longer the script finished without an exception. Every instrument in this book becomes a question that can be asked of that model as it is built: what does this tensor represent, and what should its shape be? Does the loss reach every parameter? Is the input contract intact? Where does the step spend its time? Did compilation help? And — the question this chapter added — when the next change lands, is the model actually better, or does it just look that way?

The capstone owns the model. This chapter’s job was to make sure that when it is built, you can tell whether it works.

A result is not a number. It is a comparison between two runs that share enough structure to support an explanation. Establish the difference before explaining it, and check the measurement before blaming the model.