<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Performance on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/performance/</link><description>Recent content in Performance on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Wed, 02 Sep 2026 15:30:00 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/performance/index.xml" rel="self" type="application/rss+xml"/><item><title>DataLoader: Where Is the Training Loop Actually Waiting?</title><link>https://aibussin.com/books/pytorch-from-first-principles/06-chapter/</link><pubDate>Sat, 08 Aug 2026 13:21:00 +0100</pubDate><guid>https://aibussin.com/books/pytorch-from-first-principles/06-chapter/</guid><description>&lt;p&gt;Here are two controlled training pipelines. They use the same batch size, the same machine, and the same synthetic post-batch workload. Each sample also carries the same nominal two-millisecond cost: in one pipeline that cost is waiting, while in the other it is fixed CPU work.&lt;/p&gt;&#10;&lt;p&gt;The tensor construction around that controlled cost is the same in both cases. What changes is the resource those two milliseconds consume.&lt;/p&gt;&#10;&lt;p&gt;Both are given the same treatment — raise &lt;code&gt;num_workers&lt;/code&gt; from 0 to 8 — and measured the same way.&lt;/p&gt;</description></item><item><title>Cold Starts, Warm Runs and Real Latency</title><link>https://aibussin.com/books/browser-ai-from-first-principles/11-chapter/</link><pubDate>Wed, 02 Sep 2026 15:30:00 +0100</pubDate><guid>https://aibussin.com/books/browser-ai-from-first-principles/11-chapter/</guid><description>&lt;p&gt;“Local AI is faster” is not a measurement.&lt;/p&gt;&#10;&lt;p&gt;It compresses several different waiting periods into one adjective. A browser-native feature can avoid network round trips and still make a user wait for model acquisition, process startup, session creation, context ingestion or slow generation.&lt;/p&gt;&#10;&lt;p&gt;To understand latency, we have to take the lifecycle apart.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="1-one-duration-hides-several-clocks"&gt;1. One duration hides several clocks&lt;/h2&gt;&#10;&lt;p&gt;For one interaction, useful timestamps include:&lt;/p&gt;&#10;&lt;table&gt;&#10;&#9;&lt;thead&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Phase&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Starts&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Ends&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;User-visible?&lt;/th&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/thead&gt;&#10;&#9;&lt;tbody&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Capability inspection&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;&lt;code&gt;availability()&lt;/code&gt; call&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;state returned&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Usually not&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Acquisition&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;session request requires assets&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;assets ready&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Yes, if it blocks&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Session creation&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;&lt;code&gt;create()&lt;/code&gt; called&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;session returned&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Often&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Prompt startup&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;prompt called&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;first chunk&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Yes&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Generation&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;first chunk&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;final chunk&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Yes&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Validation&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;output complete&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;feature accepted/rejected&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;Sometimes&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;p&gt;The perceived wait for a streaming feature is often dominated by time to first output:&lt;/p&gt;</description></item><item><title>Performance: What Is the Machine Waiting For?</title><link>https://aibussin.com/books/pytorch-from-first-principles/12-chapter/</link><pubDate>Sat, 29 Aug 2026 16:00:00 +0100</pubDate><guid>https://aibussin.com/books/pytorch-from-first-principles/12-chapter/</guid><description>&lt;p&gt;Here is a benchmark. A model, a batch, a loop, a timer. Nothing exotic.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model&lt;span style="color:#f92672"&gt;.&lt;/span&gt;train()&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;t0 &lt;span style="color:#f92672"&gt;=&lt;/span&gt; time&lt;span style="color:#f92672"&gt;.&lt;/span&gt;perf_counter()&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;for&lt;/span&gt; _ &lt;span style="color:#f92672"&gt;in&lt;/span&gt; range(&lt;span style="color:#ae81ff"&gt;20&lt;/span&gt;):&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; optimizer&lt;span style="color:#f92672"&gt;.&lt;/span&gt;zero_grad(set_to_none&lt;span style="color:#f92672"&gt;=&lt;/span&gt;&lt;span style="color:#66d9ef"&gt;True&lt;/span&gt;)&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; loss &lt;span style="color:#f92672"&gt;=&lt;/span&gt; F&lt;span style="color:#f92672"&gt;.&lt;/span&gt;cross_entropy(model(x), y) &lt;span style="color:#75715e"&gt;# MLP, batch 256&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; loss&lt;span style="color:#f92672"&gt;.&lt;/span&gt;backward()&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; optimizer&lt;span style="color:#f92672"&gt;.&lt;/span&gt;step()&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ms_per_step &lt;span style="color:#f92672"&gt;=&lt;/span&gt; (time&lt;span style="color:#f92672"&gt;.&lt;/span&gt;perf_counter() &lt;span style="color:#f92672"&gt;-&lt;/span&gt; t0) &lt;span style="color:#f92672"&gt;/&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;20&lt;/span&gt; &lt;span style="color:#f92672"&gt;*&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;1000&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;naive loop of 20: 9.491 ms/step reported + 132.0 ms still queued at the final sync -&amp;gt; really 16.093 ms/step&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;naive loop of 100: 14.692 ms/step reported + 138.8 ms still queued at the final sync -&amp;gt; really 16.080 ms/step&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;naive loop of 400: 15.741 ms/step reported + 132.1 ms still queued at the final sync -&amp;gt; really 16.071 ms/step&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;synchronized + warmed, median of 400 : 16.430 ms/step &amp;lt;- the real number&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;one step, host-visible return only : 4.866 ms/step&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;first training step ever : 20.890 ms&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Read the first line again. The loop ran twenty iterations and reported &lt;code&gt;9.491&lt;/code&gt; ms per step. Then a single &lt;code&gt;torch.cuda.synchronize()&lt;/code&gt; immediately afterward blocked for &lt;code&gt;132&lt;/code&gt; more milliseconds — roughly eight more steps&amp;rsquo; worth of work that the GPU had not finished when the timer stopped. Loop over a hundred iterations instead and the same code reports &lt;code&gt;14.7&lt;/code&gt;. Over four hundred, &lt;code&gt;15.7&lt;/code&gt;. Time one step in isolation and it looks like &lt;code&gt;4.9&lt;/code&gt;.&lt;/p&gt;</description></item><item><title>Appendix A: PyTorch Diagnostic Field Guide</title><link>https://aibussin.com/books/pytorch-from-first-principles/appendix-a/</link><pubDate>Sun, 30 Aug 2026 12:00:00 +0100</pubDate><guid>https://aibussin.com/books/pytorch-from-first-principles/appendix-a/</guid><description>&lt;p&gt;This appendix introduces no new PyTorch mechanism.&lt;/p&gt;&#10;&lt;p&gt;Everything here was earned earlier in the book by building something small, breaking it deliberately, measuring what changed, and locating the first place where reality stopped matching the intended computation.&lt;/p&gt;&#10;&lt;p&gt;The purpose of this appendix is different.&lt;/p&gt;&#10;&lt;p&gt;When a real model is failing, you usually do not need another explanation of autograd, broadcasting, attention or CUDA. You need to answer a narrower question:&lt;/p&gt;</description></item><item><title>Profile Before You Optimize</title><link>https://aibussin.com/books/cellular-automata-from-first-principles/50-chapter/</link><pubDate>Mon, 10 Aug 2026 20:52:00 +0100</pubDate><guid>https://aibussin.com/books/cellular-automata-from-first-principles/50-chapter/</guid><description>&lt;p&gt;By now we have built dozens of automata.&lt;/p&gt;&#10;&lt;p&gt;The natural temptation is to make them faster.&lt;/p&gt;&#10;&lt;p&gt;The wrong first question is:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Which optimization trick should I use?&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;The right first question is:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Where is the time actually going?&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;Performance work begins with measurement, and with the doctrine Part VI starts from:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;Optimize the verified implementation, not the first implementation that happens to render something plausible.&lt;/strong&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;Running this book&amp;rsquo;s own earlier code turned up undefined helpers, wrong signatures and a misaligned FFT, all inside chapters that produced plausible figures. Speed means nothing until equivalence is established.&lt;/p&gt;</description></item><item><title>PyTorch Performance Debugging: CUDA OOM, Slow Training, GPU Utilization and torch.compile</title><link>https://aibussin.com/post/pytorch-zero-to-hero-09/</link><pubDate>Sat, 08 Aug 2026 13:56:00 +0100</pubDate><guid>https://aibussin.com/post/pytorch-zero-to-hero-09/</guid><description>&lt;h2 id="pytorch-zero-to-hero--step-09"&gt;PyTorch: Zero to Hero — Step 09&lt;/h2&gt;&#10;&lt;p&gt;At this point in the series, the model runs.&lt;/p&gt;&#10;&lt;p&gt;That does not mean it runs well.&lt;/p&gt;&#10;&lt;p&gt;A training loop can be correct and still waste most of the machine.&lt;/p&gt;&#10;&lt;p&gt;A model can fit in memory and still spend half its time waiting on synchronization.&lt;/p&gt;&#10;&lt;p&gt;A &lt;code&gt;torch.compile&lt;/code&gt; call can make code faster, slower, or simply move the bottleneck somewhere else.&lt;/p&gt;&#10;&lt;p&gt;A CUDA out-of-memory error can be caused by the model, the optimizer, activations, fragmentation, a leaked reference, a larger batch, a longer sequence, or an innocent-looking tensor that was kept alive by Python.&lt;/p&gt;</description></item><item><title>PyTorch DataLoader Performance: num_workers, pin_memory, Prefetching and Why Your GPU Is Waiting</title><link>https://aibussin.com/post/pytorch-zero-to-hero-05/</link><pubDate>Sat, 08 Aug 2026 13:21:00 +0100</pubDate><guid>https://aibussin.com/post/pytorch-zero-to-hero-05/</guid><description>&lt;h2 id="pytorch-zero-to-hero--step-05"&gt;PyTorch: Zero to Hero — Step 05&lt;/h2&gt;&#10;&lt;p&gt;A fast model with a slow input pipeline is still a slow training system.&lt;/p&gt;&#10;&lt;p&gt;One of the most common PyTorch performance failures looks like this:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;GPU utilization: 20% → 95% → 10% → 90% → 15%&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The model is not necessarily slow.&lt;/p&gt;&#10;&lt;p&gt;The GPU may simply be waiting for the next batch.&lt;/p&gt;&#10;&lt;p&gt;This article is about finding out &lt;strong&gt;where the wait is happening&lt;/strong&gt;.&lt;/p&gt;</description></item></channel></rss>