<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Pin_memory on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/pin_memory/</link><description>Recent content in Pin_memory on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Sat, 08 Aug 2026 13:21:00 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/pin_memory/index.xml" rel="self" type="application/rss+xml"/><item><title>DataLoader: Where Is the Training Loop Actually Waiting?</title><link>https://aibussin.com/books/pytorch-from-first-principles/06-chapter/</link><pubDate>Sat, 08 Aug 2026 13:21:00 +0100</pubDate><guid>https://aibussin.com/books/pytorch-from-first-principles/06-chapter/</guid><description>&lt;p&gt;Here are two controlled training pipelines. They use the same batch size, the same machine, and the same synthetic post-batch workload. Each sample also carries the same nominal two-millisecond cost: in one pipeline that cost is waiting, while in the other it is fixed CPU work.&lt;/p&gt;&#10;&lt;p&gt;The tensor construction around that controlled cost is the same in both cases. What changes is the resource those two milliseconds consume.&lt;/p&gt;&#10;&lt;p&gt;Both are given the same treatment — raise &lt;code&gt;num_workers&lt;/code&gt; from 0 to 8 — and measured the same way.&lt;/p&gt;</description></item><item><title>PyTorch DataLoader Performance: num_workers, pin_memory, Prefetching and Why Your GPU Is Waiting</title><link>https://aibussin.com/post/pytorch-zero-to-hero-05/</link><pubDate>Sat, 08 Aug 2026 13:21:00 +0100</pubDate><guid>https://aibussin.com/post/pytorch-zero-to-hero-05/</guid><description>&lt;h2 id="pytorch-zero-to-hero--step-05"&gt;PyTorch: Zero to Hero — Step 05&lt;/h2&gt;&#10;&lt;p&gt;A fast model with a slow input pipeline is still a slow training system.&lt;/p&gt;&#10;&lt;p&gt;One of the most common PyTorch performance failures looks like this:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;GPU utilization: 20% → 95% → 10% → 90% → 15%&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The model is not necessarily slow.&lt;/p&gt;&#10;&lt;p&gt;The GPU may simply be waiting for the next batch.&lt;/p&gt;&#10;&lt;p&gt;This article is about finding out &lt;strong&gt;where the wait is happening&lt;/strong&gt;.&lt;/p&gt;</description></item></channel></rss>