← PyTorch From First Principles

Compilation: Which Assumption Stopped Holding?

Treat torch.compile as a runtime contract: distinguish correctness from capture, stability, and value, inspect graph breaks, guards, cached variants, dynamic shapes, recompilation, and cold versus warm cost, then find the first compiler assumption that stopped holding.

Here is an inference service. It loads a model, compiles it once, and serves requests whose sequence length varies from request to request — the ordinary situation for anything that takes text.

model = load_model().eval()
compiled = torch.compile(model, dynamic=False)   # "dynamic shapes are slow, force static"

for request in stream:
    with torch.no_grad():
        output = compiled(request.tokens)         # [1, T], T varies per request

Thirteen requests. Every compiled output agrees with the eager model to a maximum absolute difference of 7e-7. Nothing raises.

   req  seqlen     eager (ms)    compiled
      0      60      112.45          11.42 s  <- compiling
      1      63        2.08           8.49 s  <- compiling
      2      69        1.95           8.61 s  <- compiling
      3      16        4.42           5.92 s  <- compiling
      4      19        6.03           8.91 s  <- compiling
      5      75        2.63           8.40 s  <- compiling
      6      19        2.40            2.1 ms   <- reuse (seq length 19 seen at request 4)
      7      55        2.03           8.68 s  <- compiling
      8      25        2.07           8.39 s  <- compiling
      9      35        1.95            7.8 ms
     10      37        5.15            1.6 ms
     11      66        3.08            1.9 ms
     12      52        1.80            5.1 ms

  eager    total:    0.148 s
  compiled total:   68.838 s     (465x)
  output agreement with eager (max |Δ|): 7.15e-07

The eager service answered all thirteen requests in 0.15 seconds. The compiled service took sixty-nine seconds — four hundred and sixty-five times slower — and produced identical answers.

Read the trace. The first eight distinct sequence lengths each produced a cached static specialization, with an eight-to-eleven-second compile stall. Request 6 reused the already-seen T=19 specialization in milliseconds. At request 9, the ninth distinct length (T=35) failed all eight cached guards, emitted the cache-limit warning, and did not create a ninth specialization. The remaining previously unseen lengths then ran eagerly in milliseconds. A later request matching one of the eight existing guard sets could still reuse its compiled graph. The only explanation of that change in policy was a single line nobody had asked to see:

torch._dynamo.convert_frame: torch._dynamo hit config.cache_size_limit (8)
   last reason: 0/0: tensor 'L['idx']' size mismatch at index 1. expected 60, actual 35

(This run used a cold Inductor cache. With the on-disk cache warm from an earlier run, the per-shape stalls shrink to about a second and the slowdown drops to 90× — still dramatically slower than eager, still every output correct.)

Every request was answered correctly. torch.compile was called. Kernels were generated. And the service is now paying the full price of compilation while receiving, for most requests, none of the benefit.

A compiled program that produces correct numbers has proved only that it is correct. It has not proved that anything was captured, that a graph was reused, or that the workload improved.

Where we are

Chapter 12 treated torch.compile as one optimization candidate among several. It asked one question: did compilation help this workload? — measured eager steady state against compiled steady state under a held-constant contract, found the call roughly neutral on a GEMM-bound model and a large speedup on a fusable pointwise chain, and, when a compiled result looked wrong, stopped and handed off.

This chapter is that handoff. Chapter 12’s method — measure, localize, intervene, remeasure — runs out of questions the moment the compiler itself is the thing behaving strangely: the first call takes seconds you cannot account for, a new input shape pauses to compile again, the logs mention recompilation and you cannot tell what invalidated the graph, graph breaks appear inside forward and fragment the program, one backend works and another fails.

Chapter 12 asked: did compilation help?

Chapter 13 asks: what did the compiler capture, what assumption invalidated the graph, and why did it compile again?

The idea the chapter is built on:

A compiled PyTorch program is still your Python program. But the runtime must decide which parts of it can become graphs, and when those graphs are still valid. Compilation is a contract. Debugging it means finding the first assumption the runtime could not keep.

That contract has a structure, and every layer of it can be inspected:

your Python function
      │   what could the tracer capture as a graph?
      ▼
captured graphs, and the graph breaks between them
      │   what did each graph assume in order to specialize?
      ▼
guards  (shape, dtype, device, some Python values)
      │   did those assumptions survive the next call?
      ▼
graph reuse   ·   recompilation   ·   silent fallback to eager

Every chapter in this book makes one hidden structure visible. This one makes the compiler’s.

Boundaries

  • Chapter 12 owns benchmark hygiene: warmup, synchronization, the measurement contract, latency versus throughput versus peak memory versus cold-start cost, the profiler, forensic versus performance mode. This chapter uses those instruments — where a compiler experiment needs a benchmark, it is the Chapter 12 harness — and does not reteach them. Chapter 13’s new contribution to the cost story is narrow: recompilation turns a one-time cold cost into a recurring one.
  • Chapter 14 owns run-to-run comparability, noise thresholds and regression policy. This chapter may say “record the evidence” and no more.
  • Chapter 15 is the capstone. No GPT is built here.

The five levels

Chapter 10 measured attention correctness at four levels because “correct” meant four different things. A compiled function is the same: “it works” can mean any of five things, and the failures in this chapter are best sorted by which level first catches them.

LEVEL 1  RUNS       compiled code executes without raising
LEVEL 2  CORRECT    output agrees with eager at a stated tolerance
LEVEL 3  CAPTURED   the hot region became one graph, not a chain of breaks
LEVEL 4  STABLE     guards hold in steady state; no recurring recompilation
LEVEL 5  WORTH IT   the workload metric improved, after amortizing compile cost

The opening failure reaches Level 2 — it runs, it is correct — and fails Level 4 and Level 5 spectacularly. A model with a print in the wrong place reaches Level 2 and fails Level 3. Every level is cheap to mistake for the one above it.

The environment

Every number, log line and trace in this chapter was produced by running the code shown, on one machine:

Python        3.11.4
PyTorch       2.6.0+cu118        CUDA runtime 11.8
GPU           NVIDIA GeForce RTX 2060, compute capability 7.5, 12 GiB
torch.compile Inductor backend; triton 3.2.0 (triton-windows); MSVC toolchain
config        torch._dynamo.config.cache_size_limit == 8
              torch._dynamo.config.accumulated_cache_size_limit == 256

Three environment facts shape the chapter.

The core Dynamo mechanisms are not CUDA-only. Graph capture, guards, recompilation and dynamic-shape specialization can all be investigated on CPU as well as CUDA. But do not turn that into “everything behaves the same”: device is itself part of the guarded state, backend/kernel coverage differs by device, and exact graph breaks, compile cost and generated code can change with backend and PyTorch version. The CUDA timings in this chapter therefore belong to this machine even when the diagnostic method transfers.

First-call timings on this machine depend strongly on the on-disk Inductor cache. A previous compilation can leave reusable artifacts on disk, so a fresh process may pay much less than a genuinely cold cache. The cache key reflects more than “structure and shapes” — generated code, configuration, environment and input specialization all matter. In the measured opening failure, roughly 8 seconds per new shape fell to about 1 second with relevant disk artifacts warm, and the 465× slowdown fell to about 90×. Where cold cost matters, the experiment uses a fresh TORCHINDUCTOR_CACHE_DIR and says so.

torch._dynamo.reset() / torch.compiler.reset() is called between experiments to clear process-local compiler state. That prevents one experiment’s in-memory Dynamo cache and compile counters from silently changing the next one. It does not delete Inductor’s filesystem cache, which is why cold-cache experiments need a separate cache directory or an explicit disk-cache reset.

One naming point this book cannot skip. torch.compiler.* is the public API: torch.compile, torch.compiler.reset, torch.compiler.disable, torch.compiler.set_stance, torch.compiler.list_backends. The instruments with the best diagnostic value in this chapter — torch._dynamo.explain, torch._dynamo.mark_dynamic, torch._dynamo.config — live under torch._dynamo, which is not a stable public surface. They are worth using and worth knowing are internal.

The controlled workload

One workload, reused across every experiment: a small transformer over [B, T, E] — token and position embeddings, a stack of pre-norm self-attention-plus-MLP blocks, an output projection.

class Block(nn.Module):
    def forward(self, x):
        B, T, E = x.shape
        h = self.ln1(x)
        q, k, v = self.qkv(h).split(self.dim, dim=-1)
        q, k, v = (t.view(B, T, self.heads, E // self.heads).transpose(1, 2) for t in (q, k, v))
        a = F.scaled_dot_product_attention(q, k, v, is_causal=True)
        x = x + self.proj(a.transpose(1, 2).reshape(B, T, E))
        return x + self.mlp(self.ln2(x))

It is chosen for two properties: it captures cleanly as one graph, so a broken capture is obviously broken; and the sequence length T is a natural axis of variation, so recompilation experiments do not need a contrived shape.

The shared instruments are in complab.py: hard_reset(), a capture_logs(**categories) context manager that turns on specific TORCH_LOGS categories and captures what they emit, an explain_summary() wrapper, and the Chapter 12 timing helper.

The cheapest bisection: run it eager first

Before any compiler question, one measurement splits the search space in half.

eager_out = model(x)                      # does the ordinary program work?
compiled = torch.compile(model)
torch.testing.assert_close(eager_out, compiled(x), rtol=1e-4, atol=1e-4)
HEALTHY MODEL
  eager forward: ok, output (8, 64, 512)
  assert_close(eager, compiled, rtol=1e-4, atol=1e-4): PASS
  max |Δ| eager vs compiled: 1.43e-06

The tolerance is deliberate and stated. In this experiment Inductor’s generated code differs from eager by 1.4e-6, consistent with a different floating-point evaluation schedule. A healthy compile may be bitwise identical on some workloads and merely numerically close on others. Use assert_close with a tolerance justified by the computation; require bitwise equality only when exact equality is actually part of the contract.

Now a model with an ordinary bug — a reshape that assumes a fixed sequence length:

def forward(self, x):        # x: [B, T, 256]
    return self.lin(x).reshape(x.shape[0], 64, x.shape[2])   # only valid when T == 64
  eager, T=48:    RuntimeError: shape '[8, 64, 256]' is invalid for input of size 98304
  compiled, T=48: TorchRuntimeError: Failed running call_method reshape(*(FakeTensor(..., size=(8, 48, 256)), 8, 64, 256), **{})

Same root cause, one wrapped in compiler framing. If you had only run the compiled version, “torch.compile is broken” is exactly the wrong conclusion.

Fails eager → an ordinary model problem, and the compiler is a distraction. Works eager, fails compiled → a compiler-path problem, and now the search is worth doing.

Everything after this assumes the eager program is correct.

What was captured: graph breaks

torch.compile traces your function into a graph of tensor operations. When it reaches Python it cannot put in the graph — a data-dependent branch, a call into an unsupported library, a value pulled out of a tensor — it breaks: it compiles the graph so far, returns to the Python interpreter to run the unsupported part, and resumes tracing after it.

The controlled model captures cleanly. torch._dynamo.explain reports what happened:

torch._dynamo.explain(model, tokens):
  graph_count        1
  graph_break_count  0
  break_reasons      []

One graph, no breaks — the whole forward, embeddings through output projection, became a single region the compiler can optimize across. That is Level 3 passing, and it is the baseline every failure below is measured against.

Now the canonical break — a branch on a tensor’s data:

def with_break(x):
    y = torch.sin(x) @ x
    if x.sum().item() > 0:          # tensor -> Python float -> control flow
        y = y * 2
    return torch.cos(y)

Three instruments see it. Reach for them in this order.

torch._dynamo.explain — a structured account, no log parsing:

torch._dynamo.explain(with_break):
  graph_count        2
  graph_break_count  1
  break_reasons      ['Tensor.item']

TORCH_LOGS=graph_breaks (or capture_logs(graph_breaks=True)) — the exact location and reason, in the compiler’s own words:

torch._dynamo.symbolic_convert.__graph_breaks: Graph break in user code at e3_graph_breaks.py:24
Reason: Unsupported: Tensor.item
  if x.sum().item() > 0:              # tensor -> Python value -> control flow

fullgraph=True — turns the silent break into a raised exception, which is what you want when you are hunting one:

torch.compile(with_break, fullgraph=True)(x)
  -> Unsupported: Tensor.item

fullgraph=True is useful as a diagnostic because it converts a graph break into an error and requires one captured graph for the region. It is also a legitimate production or library-testing constraint when single-graph capture is the requirement. What it is not is a universal performance fix: ordinary fullgraph=False code can be correct and fast with intentional breaks.

A graph break is not a correctness failure

The broken function computes the same intended result as the unbroken version. A graph break creates a boundary where Dynamo returns to Python and cannot optimize across the two compiled regions. The cost depends on what happens at that boundary. In this particular example, .item() also pulls a tensor value into Python and can synchronize CUDA work with the host; other graph breaks need not have that same synchronization cost.

Measure it. Compare the compiled broken function against a compiled version with the branch rewritten to stay in the graph (torch.where), and against eager:

eager (with break)          median 0.184 ms
compiled, with break        median 0.323 ms
compiled, no break          median 0.184 ms

The compiled unbroken version matches eager on this tiny workload — compilation was neutral here to begin with. But the compiled broken version is 75% slower than either. The break did not just remove the benefit; the machinery of breaking and resuming made it worse than not compiling at all.

That number is not universal — on other runs the gap was 40–50% — but the direction is stable, and so is the lesson: a break’s cost is a measurement, not an assumption. Sometimes it is nothing. That is a finding too.

Break location matters more than break count

Two models, the controlled transformer, one data-dependent branch each. In one it sits before the block stack; in the other it sits between two blocks.

TORCH_LOGS=graph_breaks confirms both have exactly one break:

  no break               breaks=0
  break before blocks    breaks=1   at forward line 26
  break mid-stack        breaks=1   at forward line 40

The steady-state cost is not the same:

  no break              1.867 ms    1.00x the no-break graph    0.60x eager
  break before blocks   1.856 ms    0.99x the no-break graph    0.59x eager
  break mid-stack       2.359 ms    1.26x the no-break graph    0.76x eager

The break before the blocks has no resolved cost in this benchmark: 1.856 ms against 1.867 ms for the no-break graph, a difference too small to call. The expensive block stack still captures as one graph behind it. The break in the middle splits that stack into two graphs and returns to the interpreter halfway through every forward pass — 26% slower here, eroding much of the compilation speedup.

Same count. Opposite cost. “Minimise the number of graph breaks” is the wrong goal.

Ask where the break is, how often that code path runs, and what it measurably cost. A break in setup and a break in the hot loop are different findings.

The instruments disagree, and one is right

For the mid-stack break, torch._dynamo.explain reported:

  break mid-stack        graphs=1  breaks=0

Zero breaks — while TORCH_LOGS=graph_breaks recorded a break at the expected source line and fullgraph=True rejected the same construct. In this PyTorch 2.6 experiment, torch._dynamo.explain therefore undercounted. Do not promote any one diagnostic to universal “ground truth”: use explain for the structured summary, the graph-break log for the tracer’s emitted location/reason, and fullgraph=True to test whether this exact call can be captured as one graph. Agreement between independent instruments is stronger than authority assigned to one of them.

Do not report a construct as “causing a graph break” from memory or from one instrument. The tracer’s coverage changes between releases; show the evidence, name the version.

The assumptions: guards and recompilation

When the compiler specializes a graph, it records the facts about the inputs it relied on — as guards. A guard is a cheap runtime check: this tensor is 2-D, this size, this dtype, on this device; this Python value is still 4. Before reusing a compiled graph, the runtime checks its guards. If one fails, the graph is not valid for these inputs, and PyTorch compiles another one.

dynamic=False — the “force static, it’s faster” setting from the opening — makes shape a guard on every dimension. Call the compiled model across a sequence of lengths and capture TORCH_LOGS=recompiles:

call sequence T = [32, 32, 32, 48, 48, 32, 64, 48, 80, 96, 112]

  call  T     wall (s)   note
     0   32      8.393   compile (new shape)
     1   32      0.001   reuse
     2   32      0.001   reuse
     3   48      6.135   compile (new shape)
     4   48      0.001   reuse
     5   32      0.001   reuse       <- guard still holds, graph 0 reused
     6   64      6.359   compile (new shape)
     ...
  distinct shapes seen: 6   recompiles logged: 5

Every new length compiles a fresh graph (6–8 s); every repeat reuses one in a millisecond. The recompile log says exactly which assumption failed:

Recompiling function forward in complab.py:211
    triggered by the following guard failure(s):
    - 0/0: tensor 'L['idx']' size mismatch at index 1. expected 32, actual 64
    - 0/1: tensor 'L['idx']' size mismatch at index 1. expected 48, actual 64

Read those two lines carefully. On the call with T=64, the recompile log shows the runtime rejecting the T=32 and T=48 cached specializations before creating the T=64 one. Cached entries are guard-checked until a valid specialization is found or compilation is needed. That guard lookup has a cost, but this experiment did not isolate it: the opening compile stalls were not monotonically getting slower, and seconds of code generation dwarfed guard evaluation. The log proves which assumptions failed; it does not prove that guard scanning caused the compile-time curve.

Guards are not only about shape

Feed the same compiled function int32 tokens instead of int64:

Recompiling function forward in complab.py:211
    triggered by the following guard failure(s):
    - 0/0: tensor 'L['idx']' dtype mismatch. expected Long, actual Int

A dtype guard. Device is guarded the same way, and so are some Python values the graph closed over — a config flag, an integer that reached a .view(). A recompile whose guard-failure text is not a shape is a different investigation and a different fix.

Past the limit: the silent fallback

torch._dynamo.config.cache_size_limit is 8 in this PyTorch 2.6 environment. Once this code object has eight cached specializations and another input would require compilation, Dynamo stops creating new variants and runs that unmatched call eagerly. Previously compiled variants remain cached and can still be reused when their guards pass. In the opening, request 9 (T=35) was the first unmatched call after eight variants existed; its recompile attempt hit the limit instead of producing graph number nine:

torch._dynamo hit config.cache_size_limit (8)

This failure is easy to miss because correctness survives while execution becomes a mixture of policies. Inputs matching one of the eight cached guard sets can still use compiled code; new unmatched inputs run eagerly because no more variants are being created. The service has already paid for eight compilations, yet much of a shape-diverse request stream may receive no compiled steady-state benefit.

Making it loud

torch.compiler.set_stance turns “is this recompiling in steady state?” from a log-reading exercise into an assertion. Wrap the region you believe is stable:

compiled(warmup_batch)                                  # compile once

with torch.compiler.set_stance("fail_on_recompile"):
    for batch in steady_state_loader:
        compiled(batch)
  same shape (T=32): reused, no error
  new shape (T=48): RuntimeError: Detected recompile when torch.compile stance is 'fail_on_recompile'

If it raises, a guard you thought was stable is failing. The other stances are diagnostic too: "force_eager" ignores every torch.compile in the process (bisect compiler bugs against eager without editing call sites), "eager_on_recompile" runs eagerly instead of paying to recompile.

The cost that recurs

Chapter 12 established that compile cost is not one number. Measured on this chapter’s model: on a cold Inductor cache the first compiled call took about 13 s; in a fresh process reusing that cache, about 3.4 s. Chapter 13’s addition is what recompilation does to that.

A one-time cold cost can be amortized. Recompilation is different: when no cached guard set accepts an input and the limit still permits another specialization, a fresh compile cost is paid again. A guard failure alone does not imply a compile — another cached specialization may match later — and the limit-hit fallback itself does not pay full code-generation cost. In the opening, the service paid for eight static specializations, reused the T=19 one once, then the ninth distinct shape hit the limit and ran eagerly in 7.8 ms.

A cold compile is a deposit. A recompiling loop is a subscription.

Dynamic shapes: a response to proved variation

The opening failure has an obvious-looking fix: stop specializing on shape. But “add dynamic=True” is a flag reach, and this book does not do flag reaches. Measure all four options on the same shape-varying workload — 13 calls, 10 distinct sequence lengths:


The results are not what "dynamic is better" predicts.

**`dynamic=False`** is the opening failure: eight static specializations are created, the ninth compile attempt hits the cache limit, and roughly `30` seconds of the `30.4`-second sweep are spent in the compile-heavy prefix. Calling that column "nine compiles" would overcount the specialization that was refused.

**`dynamic=None` — the default** — produced two compile attempts in this PyTorch 2.6 experiment: an initial specialization, then one recompile after shape variation appeared. After that the workload emitted no further recompile logs and reused cached compiled code across the tested lengths; its steady state (`1.322` ms) was the fastest of the four. This is consistent with the default policy of detecting dynamism after a recompile. It is **not** a guarantee that every shape-varying program will generalize after exactly one recompile: operations and guards can still force specialization. The important finding is narrower — the default handled this workload's sequence-length variation, while forcing `dynamic=False` prevented it from doing so.

**`dynamic=True`** skips the specialization pass — one compile, slightly slower steady state than letting it specialize first.

**`mark_dynamic(x, 1)`** — marking exactly the sequence dimension dynamic — compiles once, fewest seconds in the sweep, but its steady state (`1.918` ms) came out no better than eager's `1.776` ms. On this model, that precise intervention did not pay.

For this workload, the transformer's compiled steady state is in the same small range across the dynamic policies (`1.3`–`2.0` ms against eager's `1.8`), while their compile costs differ dramatically. The defensible verdict is therefore specific: **`dynamic=False` created the recompilation failure for this measured shape distribution, and the default handled that variation without manual help.** That is evidence about this workload, not a rule that the default resolves every dynamic-shape program.

The workflow, when recompilation *is* real:

```text
prove the recompilation is caused by shape variation   (recompile log: "size mismatch")
        ↓
try the default first; it may already generalize
        ↓
if not, mark exactly the varying dimension dynamic
        ↓
measure compile count AND steady-state time
        ↓
keep it only if the real workload improved

Static or dynamic is a question about your data

If production always runs one shape, a fully specialized graph is ideal and dynamic flexibility buys nothing. If production receives dozens of sequence lengths, specialization is the opening failure. The compiler question is never “are dynamic shapes better?” It is what distribution of shapes does my actual program see? — which is a measurement, taken from the real request stream, not a guess.

The observer effect

You have a compiled function that captures as one graph. You want to see an intermediate value, so you add a print:

if i == D // 2:
    print("  [debug] mid mean:", float(x.mean()))
  graph breaks (TORCH_LOGS=graph_breaks):
    clean (debug=False)      breaks = 0
    same class, debug=True   breaks = 1      (Unsupported: Tensor.item, at the print)

  steady-state compiled forward:
    clean / debug=False          0.164 ms    1.00x clean    0.09x eager
    debug=True (print inside)     1.987 ms   12.12x clean    1.06x eager

The clean function was 11× faster than eager — a fusable chain, exactly the workload compilation is best at. The print — through float(x.mean()), which pulls a Python number out of a tensor — broke the graph at that line and erased the entire speedup. The compiled function with the print is 12× slower than the compiled function without it, and no faster than eager.

The two are the same class. Only a Python flag differs. debug=True is a different graph from the one that runs in production, and the thing you were trying to observe is not the thing that was running.

The repairs, in order of preference:

Return the intermediate and inspect it outside the compiled region.

def forward(self, x):
    ...
        return x, mid          # stays capturable; extra outputs can still affect liveness/memory
  inspecting the RETURNED intermediate outside the graph: mid mean 1.6448  (0 breaks, clean-graph speed)

Gate the diagnostic on a flag that is False in the measured configuration — Dynamo specializes on that Python value and folds the dead branch out of that specialization. Changing the flag is itself a change in guarded program state and may trigger a different specialization; benchmark the same flag value you intend to ship.

Mark a diagnostic helper @torch.compiler.disable so the tracer does not try to capture it — though note this still breaks the graph at the call site, so it removes the “compiler chokes on my print” problem, not the cost. Use it for setup-time diagnostics, not hot-path ones.

This is Chapter 12’s forensic-versus-performance-mode distinction, one layer down. There, an enabled profiler changed the number you measured. Here, a debugging print changes the graph you compiled.

The instrumentation is part of the system. Measure the version you ship, and keep diagnostics out of the compiled region.

Narrowing the compiled region

You do not have to compile the whole training step. Compare compiling only the model’s forward against compiling the entire step — zero_grad, forward, loss, backward, optimizer.step:

  config                  graph breaks   steady ms   first call
  eager                              -       6.235        -
  compile(forward)                   0       5.353       4.4 s
  compile(whole step)                4       4.632       3.9 s

Compiling the whole step was about 13% faster than compiling only the forward in this run (4.632 vs 5.353 ms), while also producing four graph breaks. That measurement establishes an important point by itself: fewer breaks did not imply the lowest step time. It does not establish that the wider wrapper won because it “captured the backward” or fused a particular optimizer operation — compiling a model for training already involves AOTAutograd’s backward machinery, and this experiment did not profile the source of the extra speedup. What the wider wrapper definitely changes is the region Dynamo must reason about: loss, optimizer calls and more Python orchestration are now inside the compiled function.

Narrowing the region is therefore a debugging intervention, not a universal speed rule. It often makes compiler behavior easier to inspect because fewer Python mechanisms sit inside the boundary, but the performance decision still belongs to the measured workload.

Backend bisection

torch.compile is a stack. TorchDynamo observes Python execution and captures compilable tensor regions into FX graphs. The "aot_eager" path adds AOTAutograd transformations such as functionalization and, when autograd is required, ahead-of-time handling/partitioning of forward and backward computation. The default Inductor backend then lowers compiled graphs into generated kernels or calls to optimized libraries — Triton is one important GPU code-generation path, while CPU lowering commonly emits C++/OpenMP code. A failure or numerical change can enter at different layers, and the backend= argument gives a useful ablation:

torch.compile(model, backend="eager")       # TorchDynamo capture only, then run eagerly
torch.compile(model, backend="aot_eager")   # + AOTAutograd, then run eagerly
torch.compile(model)                         # + TorchInductor (the default)
  backend        first call (s)         loss   max|Δ| vs eager
  eager                   0.380     6.601730          0.00e+00
  aot_eager               2.267     6.601730          0.00e+00
  inductor                7.732     6.601730          9.54e-07

On the controlled model there is no failure to localize, but the table still demonstrates the ablation. First-call cost grows as more compiler machinery is enabled (0.4 → 2.3 → 7.7 s). In this experiment, the eager and aot_eager outputs were bitwise-equal at the reported precision, while the first measured difference (9.5e-7) appeared with Inductor. That narrows a hypothetical numerical investigation toward the Inductor/code-generation stage; it does not prove that every numerical difference in another program must originate there.

The interpretation is a search-space reduction, not a theorem:

fails at backend="eager"      -> Dynamo capture / program interaction
ok eager, fails aot_eager     -> AOTAutograd (functionalization, the joint forward-backward graph)
ok aot_eager, fails inductor  -> Inductor lowering / kernel generation

For large models where raw TORCH_LOGS output is overwhelming, TORCH_TRACE plus the tlparse tool renders the compilation as a browsable structure — compile frames, breaks, recompiles — so you can find the expensive region before reading a single log line. Structure first, then zoom.

Compilation success is not optimization success

This table combines measurements from four configurations, but the columns are not all populated by the same experiment. In particular, e10_ledger_table.py measures output agreement for the two fixed-shape compiled rows; for the shape-varying row it records the entire ten-call loop wall time and recompile logs, and does not compute a max-difference or peak-memory value. Keep those cells honest:

The compilation ledger

The reusable instrument, assembled now that the pieces are earned. It records what a compiler result must expose so the reader can ask which assumption is at risk.

COMPILATION LEDGER

region        what is inside torch.compile
capture       frames compiled / graphs produced / graph breaks + reasons  (explain, then TORCH_LOGS)
assumptions   the guards that specialize each graph: shape, dtype, device, Python values
stability     calls / recompiles / distinct shapes seen / cache_size_limit reached?
cost          cold first call (fresh cache) / warm first call (fresh process) / steady-state ms
correctness   agreement with eager at a stated tolerance
verdict       did the compiled version improve the Chapter 12 contract metric, after amortizing cost?

Filled for the opening failure:

It runs. It is correct. It captured cleanly. The contract broke at Level 4, and the fix is one argument: dynamic=False → the default.

And for the debugging print from the observer-effect section, the ledger fails one level earlier:

region        fusable chain, torch.compile(model)  with print(x.mean().item()) inside
capture       2 graphs, 1 break: "Unsupported: Tensor.item" at the print statement  (Level 3 FAILS)
assumptions   n/a -- the question stops here
stability     n/a
cost          irrelevant -- the graph you measured is not the graph you ship
correctness   output still matches eager
verdict       Level 3 FAILS: the print fragments the fusable region; 12x slower than without it

Same instrument, a different first failing level, a different fix — remove the diagnostic from the compiled region.

A worked pass: the opening failure, end to end

The five levels are a procedure. Run them on the opening failure, in order, stopping at the first one without evidence.

Level 1 — RUNS. The compiled service answers every request without raising. Pass.

Level 2 — CORRECT. max |Δ| against eager is 7e-7, inside the chapter’s stated rtol=1e-4, atol=1e-4 comparison contract. Pass. For these requests, the compiler path has not introduced a numerically significant disagreement under that contract.

Level 3 — CAPTURED. torch._dynamo.explain(model, request.tokens) reports graph_count=1, graph_break_count=0. fullgraph=True does not raise. The forward captured as one graph. Pass. The problem is not a fragmented program.

Level 4 — STABLE. TORCH_LOGS=recompiles on one epoch of requests:

Recompiling function forward ...
    triggered by: tensor 'L['idx']' size mismatch at index 1. expected 60, actual 63
Recompiling function forward ...
    triggered by: tensor 'L['idx']' size mismatch at index 1. expected 63, actual 69
... (eight times)
torch._dynamo hit config.cache_size_limit (8)

Fail, here. The guard-failure text names the mechanism precisely: size mismatch at index 1 — the sequence-length dimension of the token tensor. Not dtype, not a Python value. Shape, on the axis that varies per request.

Stop the diagnosis at Level 4. The opening already measured Level 5 end to end (465× slower), but after the limit there is no single “compiled steady state” for unseen shapes: calls matching cached guards may still use compiled variants, while unmatched shapes run eagerly. The stability failure is sufficient to explain why the service cannot amortize compilation over this request stream.

The smallest experiment that tests the diagnosis. If the cause is shape specialization, then removing shape from the guards should remove the recompiles. Change one thing — dynamic=False to the default torch.compile(model) — and re-run the same epoch under TORCH_LOGS=recompiles:

Recompiling function forward ...
    triggered by: tensor 'L['idx']' size mismatch ...
(one recompile, then silence)

Two successful compiled variants are enough for this experiment — an initial specialization and a generalized variant after the first shape change. After that, the log goes quiet: subsequent requests reuse whichever cached compiled variant has guards that accept them. The diagnosis holds.

Re-measure against the Chapter 12 contract. Eager serves the 13 requests in 0.15 s. Under the default policy, the service pays for two compiled variants and then emits no further recompilation logs for this shape distribution; subsequent calls reuse cached compiled code at roughly 1.3 ms/request against eager’s 1.8. On a sufficiently long-lived service that steady-state edge can amortize the startup cost. Over only 13 requests, compilation cost dominates. Which returns the decision to Chapter 12’s question — is this workload long enough to amortize the compile? — now that Chapter 13 has removed the repeated specialization failure.

The reflex the measurement replaced: “compiled is slower, add dynamic=True.” That would also have worked — one compile instead of two — but it would have been a guess, and on the dynamic experiment’s numbers it was slightly slower in steady state than letting the default specialize first. The guard-failure text is what turned the guess into a diagnosis.

Using AI on a compiler problem

Asked why torch.compile is slow, an assistant reaches for the flag surface — dynamic, fullgraph, mode=, a different backend — because that is what its training data is full of. Flags are downstream of four cheaper questions. The prompt below refuses the flags and demands the evidence.

My torch.compile'd model is slower than eager (or recompiles, or behaves oddly).

Do not suggest dynamic=True, fullgraph=True, mode=, a different backend,
mark_dynamic, region narrowing, or abandoning compilation yet. Not until a
measurement identifies a mechanism one of those could affect.

Reconstruct the compilation evidence from what I give you:
  1. REGION   - exactly what is inside torch.compile
  2. CORRECT  - eager output vs compiled output, max abs difference, at a
                stated tolerance
    3. CAPTURED - torch._dynamo.explain: graph count, break count, break reasons;
                and TORCH_LOGS=graph_breaks output. If they disagree, preserve
                both observations and use fullgraph=True on the same call to test
                whether single-graph capture actually succeeds.
  4. STABLE   - TORCH_LOGS=recompiles: how many recompiles, the guard-failure
                text for each, how many distinct shapes the workload produced,
                and whether cache_size_limit was reached
  5. COST     - first call on a cold Inductor cache, first call in a fresh
                process with a warm cache, steady-state ms
  6. CONTRACT - the Chapter 12 measurement contract the compiled version is
                supposed to improve, and the eager baseline under it

Identify the FIRST level that lacks evidence:
  RUNS -> CORRECT -> CAPTURED -> STABLE -> WORTH IT

Stop there. If CORRECT fails, it may be an eager bug the compiler surfaced -
check eager alone. If STABLE fails, tell me which guard is failing and on what
input property, and what the smallest change to the workload or the compile
call would be to test that it is the cause.

Three clauses carry the weight. Forbidding flags until a measurement justifies one stops a plausible downstream fix from being applied to an upstream cause. Asking for the guard-failure text forces the diagnosis to be specific: “it recompiles” is not actionable; “size mismatch at index 1” names the axis and points at the input contract. And asking which level lacks evidence keeps the investigation in order. fullgraph=True, for example, is a useful capture constraint/diagnostic; it does not repair a slow graph merely by being enabled.

Ask AI what the compiler captured before asking it which compiler flag to set.

The compiler-debugging sequence

 1. Prove eager correctness. If the model is wrong eager, the compiler is a
    distraction. Fix the model.

   2. Compare eager and compiled output with assert_close at a stated tolerance.
    A healthy compile may be bitwise equal or may differ slightly because of a
    different floating-point schedule. The tolerance belongs to the computation.

  3. Ask what was captured before asking how fast it is. Use
    torch._dynamo.explain for a structured summary and TORCH_LOGS=graph_breaks
    for emitted source locations/reasons. If they disagree, reproduce the same
    input under fullgraph=True to test whether single-graph capture is possible;
    do not declare one diagnostic universally infallible.

  4. Judge graph breaks by location, frequency and measured cost, not count.
    A break before the hot region may have no resolvable cost; a break inside
    it can fragment optimization and execute Python on every call.

  5. Measure what a break costs. Compare against the same function with the
    data-dependent construct removed. A break with no measured cost is a finding.

  6. Enable one log category at a time. graph_breaks, then recompiles, then
    guards. The goal is to answer one question, not to drown in output.

  7. Read the guard-failure text. "size mismatch at index N" identifies a shape
    assumption on that tensor axis; "dtype mismatch" or a Python-value guard is
    a different mechanism and needs a different intervention.

  8. Count the shapes the real workload produces before reaching for dynamic
    shapes. Distinguish compile attempts from successfully cached variants.
    Try the default first, then mark only the dimension whose variation is proved,
    and measure both compile behavior and steady-state time.

  9. Check whether cache_size_limit was hit. After the limit, existing compiled
    variants remain reusable, but unmatched calls run eagerly because no new
    specialization is created. Use fail_on_recompile earlier when "no steady-state
    recompile is acceptable" should be an assertion rather than a log surprise.

10. Bisect the backend when behavior differs: eager, then aot_eager, then
    inductor. The first that changes localizes the layer. Search-space
    reduction, not proof.

11. Narrow the compiled region to the numerical hot path. Keep logging, I/O,
    metrics and orchestration outside it.

12. Keep diagnostics out of the compiled region. A print inside it changes the
    graph you are measuring.

13. Re-measure the uninstrumented workload against Chapter 12's baseline, under
    the same contract.

14. Record the compilation ledger. Region, capture, assumptions, stability,
    cost, verdict.

What you should now be able to answer

"torch.compile doesn’t work." Works eager? If not, it is a model bug the compiler surfaced. Works eager, fails compiled? Bisect the backend: eager, aot_eager, inductor. “Fails at aot_eager” is a bug report; “doesn’t work” is not.

“There’s a graph break, I need to remove it.” Where is it — TORCH_LOGS=graph_breaks gives the emitted source location. How often does that path run? What did it measurably cost against a version with the break removed? The pre-block break measured 1.856 ms against 1.867 ms for no break — no resolved penalty in this run — while the same construct mid-stack cost 26%. Removing an unmeasured or negligible break before fixing an expensive one is wasted effort.

“It recompiles a lot.” TORCH_LOGS=recompiles gives the guard-failure text. “size mismatch at index 1” means that tensor’s sequence-length axis is violating a cached shape assumption — how many distinct values does the real workload produce? In this PyTorch 2.6 experiment the default specialized once and then generalized after one recompile. Do not assume that exact pattern for every model; verify whether the log actually goes quiet. If dynamic=False is forcing a new static variant for a genuinely variable axis, that setting is the first thing to test.

“I’ll set dynamic=True.” Is the recompilation actually caused by shape variation, or by a dtype, device or Python-value guard? If it is shape, test the default first. In this chapter the default was fastest at 1.322 ms; dynamic=True was still faster than eager (1.415 vs 1.776 ms) but slower than the default, while mark_dynamic measured 1.918 ms and was slower than eager. “More dynamic” was not a monotonic performance improvement.

“It compiled, so compilation worked.” It reached Level 1. Was the hot region captured as one graph (Level 3)? Do the guards hold in steady state, or is it recompiling (Level 4)? Did the metric improve after amortizing the first call (Level 5)? The opening failure passed Levels 1 and 2 and was 465× slower.

“The prints show something weird.” The prints changed what could be captured. float(x.mean()) inside a compiled region breaks the graph there. Return the value and inspect it outside.

“Compile is slower than eager.” Which cost? Cold Inductor cache first call (8–11 s/shape in the opening), warm-cache first call (~1 s/shape), recurring recompilation (the 465×), or genuinely slower steady state (a Chapter-13-then-stop situation)? They have different fixes.

"torch._dynamo.explain says zero breaks but it’s slow." explain undercounted the mid-stack case in this PyTorch 2.6 experiment. Cross-check the same call with TORCH_LOGS=graph_breaks for the tracer’s emitted location/reason and with fullgraph=True to test whether one-graph capture succeeds. If those instruments all support clean capture, move on to guards, recompilation and ordinary workload economics rather than treating the summary as proof by itself.

“It was fast last week and slow this week, same code.” Compilation is a new axis of run-to-run difference: a different sequence of input shapes, a cold cache versus a warm one, the cache_size_limit reached in one run and not the other. Capture TORCH_LOGS=recompiles for both runs and compare the compile counts and guard failures. This is where Chapter 13 hands off to Chapter 14.

Exercises

  1. Reproduce the cache-limit fallback. Compile the controlled model with dynamic=False and serve more than cache_size_limit distinct sequence lengths. Capture TORCH_LOGS=recompiles. Distinguish (a) a successful cached specialization, (b) reuse of an existing specialization, (c) the recompile attempt that hits the limit, and (d) an unmatched post-limit eager call. Record the warning and verify that a previously seen shape can still reuse its cached compiled variant after the limit.

  2. The bisection. Take a model with an ordinary shape bug and show it fails under eager as well as compiled execution. Then fix the bug and compare eager with compiled using assert_close under a stated tolerance. Also test torch.equal; if it differs, explain why numerical closeness is the right contract here, and if it happens to match exactly, explain why exact equality is not guaranteed by compilation in general.

  3. Three instruments, one break. Put x.sum().item() in a branch. Record what torch._dynamo.explain, TORCH_LOGS=graph_breaks, and fullgraph=True each report. Then construct a case where the summary and log disagree. Reconcile the evidence: which instrument gives a structured count, which gives the emitted source reason, and what does the fullgraph=True result establish for that exact call?

  4. Break location. Insert one identical data-dependent branch (a) before your model’s main compute and (b) in the middle of it. Confirm both are one break. Measure the steady-state cost of each. Explain the difference in terms of what the compiler could fuse.

  5. Guard forensics. Compile a function, then call it with: a new shape, a new dtype, a new device (if you have one), and a changed Python constant that reaches a tensor op. For each, capture the guard-failure text and classify it. Which ones would dynamic=True fix, and which would it not?

  6. Dynamic shapes, measured. Sweep a varying dimension through dynamic=False, dynamic=None, dynamic=True, and mark_dynamic on that dimension. Record recompile log entries, whether the cache limit was hit, successful cached variants where you can inspect them, sweep wall time and steady-state ms. State which policy you would ship and why — and whether compilation is worth it for this workload at all.

  7. The observer effect. Compile a fusable function and measure its speedup over eager. Add print(x.mean().item()) in the middle. Re-measure. Then apply all three repairs (return the value, gate on a flag, @torch.compiler.disable) and report which recover the speedup and which only silence the tracer.

  8. Narrow the region. Compile a full training step, then compile only the forward. Record graph count, break count, first-call cost and steady-state time for both. Decide which you would use and defend it on grounds other than raw speed.

  9. The break-even. For your model on a fixed shape, measure eager steady state, compiled steady state, and the cold and warm first-call costs. Compute the number of calls at which compilation pays for itself in each cache state. Then describe a workload where it never does.

  10. The ledger on someone else’s code. Take a torch.compile call from a repository. Without changing it, fill in the compilation ledger: region, capture, assumptions, stability, cost, verdict. Identify the first level that lacks evidence.

Next: when the difference is between runs

Chapter 13 leaves the compiler inspectable. Graph breaks have locations and costs. Recompilation has a guard-failure reason you can read. The cache limit and its silent fallback can be made loud. The compiled version can be compared to eager honestly, at every level from “runs” to “worth it”.

But this chapter also created a new problem. Compilation adds a whole dimension along which two runs of the same code can differ invisibly: a different sequence of shapes seen, a different set of guards installed, a different graph selected, a warm cache in one run and cold in the other, the cache_size_limit reached this week and not last week. None of that shows up in a unit test. None of it raises.

Chapter 14 is about exactly that class of bug — the kind that is not a failure inside one run, but a difference between runs. Nothing crashes. The graph compiles. The model trains. The tests pass. And this week’s version is worse than last week’s.

Chapter 13 can name compilation as one more reason a workload needs that discipline — and then stop, without teaching comparability, noise thresholds or CI, because those belong to the next chapter.

A compiled program is a bet that a set of assumptions will keep holding. Debugging it means finding the first assumption that did not — and asking whether the graph that replaced it was worth building.