- What Does It Mean to Debug?
- What Are We Actually Doing?
- The First Divergence
- The Tensor: What Is Actually Flowing Through the Loop?
- Evidence Before Explanation
- Autograd: What Did PyTorch Record, and Where Does the Gradient Stop?
- The Debugging Stack
- The Network: What Is It Without nn.Module?
- Reading Python Exceptions
- nn.Module: What Does PyTorch Think Belongs to Your Model?
- Inspect State, Don't Guess
- DataLoader: Where Is the Training Loop Actually Waiting?
- Debug the Boundary
- Transforms: What Does the Model Actually See?
- Assertions, Invariants, and Contracts
- CNN Geometry: What Shape Reaches the Next Layer?
- Environment Bugs
- Feature Space: What Does a Linear Model Actually See?
- The Notebook Is Not the Program You See
- Attention: Which Position Is Comparing With Which?
- Hidden Notebook State
- Training: Which Link in the Learning Chain Is Broken?
- Reproducible Notebooks
- Debug the Data Before the Model
- Compilation: Which Assumption Stopped Holding?
- Shapes, Types, Devices, and Tensors
- Regressions: Did the Model Change, or the Measurement?
- When Training Goes Wrong
- Assembly: A Language Model You Can Interrogate
- Debugging Evaluation
- Appendix A: PyTorch Diagnostic Field Guide
- Debugging What You Cannot See
- Is the Model Actually the Problem?
- Inspect the Actual Model Input
- Context Windows and Truncation
- Sampling Is Part of the Program
- Internal Signals
- Representation and Behavioral Diffs
- AI as Builder, Designer, Researcher, and Reviewer
- Debugging Intent
- Debugging Context for Coding Agents
- Debugging AI-Generated Designs
- Debugging AI Research
- Debugging Coding Agents
- Appendix 13: AI Hallucination Check Prompt
- Treat Prompts as Programs
- Minimize the Prompt
- Retrieval Is a Pipeline
- Retriever Failure or Generator Failure?
- Debugging Hallucinations
- The Model's Explanation Is Not a Trace
- An Agent Is a Trajectory
- Trace the Agent
- Agent Failure Taxonomy
- Loops, Thrashing, and Retry Storms
- Time Travel, Replay, and Forking
- Causal Replay
- Trajectory Diff
- Multi-Agent Systems
- Can One AI Debug Another?
- The AI Crash Dump
- Diagnostic AI Invariants
- From Symptom to Hypotheses
- Discriminating Experiments
- How Do You Know the Diagnosis Is Right?
- AIDebugBench
- Debug the Debugger
- AI Observability
- From Production Failure to Regression
- Runtime Invariants and Guardrails
- Debugging Cost and Latency
- Debugging in Production
- The Ten-Minute Debug
- The One-Hour Investigation
- The Full AI Incident Investigation
- The Debugging AI Toolkit
- Advanced Agents From First Principles 13: How Do You Debug an Agent That Made the Wrong Decision? Add Trajectory Observability
- Agents From First Principles 06: AI Agent Chooses the Wrong Tool? Design Better Tool Interfaces, Schemas and Routing
- PyTorch Model Not Learning? A Systematic Debugging Guide
- PyTorch Attention Shapes: Q, K, V, Multi-Head Attention Masks and Transformer Dimension Errors
- PyTorch CNN Shape Errors: Conv2d Output Sizes, Channels, Flatten Bugs and How to Debug Them
- PyTorch nn.Module Explained: Missing Parameters, state_dict, Buffers and Registration Bugs
- PyTorch Autograd Debugging: requires_grad, detach, backward() and NaN Gradients
- PyTorch Tensor Shapes: Broadcasting, Reshape, View, Permute and the Errors That Waste Your Time
- Debugging Jupyter Notebooks in VS Code