<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Verification on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/verification/</link><description>Recent content in Verification on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Mon, 14 Sep 2026 05:00:21 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/verification/index.xml" rel="self" type="application/rss+xml"/><item><title>If There's Any Doubt, It's Deterministic</title><link>https://aibussin.com/books/applied-ai/03-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:03 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/03-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 1 — Where You Stand&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="seven-operations-one-of-them-intelligent"&gt;Seven operations, one of them intelligent&lt;/h2&gt;&#10;&lt;p&gt;Take the paragraph review from Chapter 1 and write down everything the process actually has to do. Not the prompt — the process.&lt;/p&gt;&#10;&lt;ol&gt;&#10;&lt;li&gt;Select the paragraph to review.&lt;/li&gt;&#10;&lt;li&gt;Decide whether it has changed since the last review.&lt;/li&gt;&#10;&lt;li&gt;Decide whether there is budget left for a call.&lt;/li&gt;&#10;&lt;li&gt;Identify claims in the paragraph that need supporting evidence.&lt;/li&gt;&#10;&lt;li&gt;Extract the flagged sentence and its position.&lt;/li&gt;&#10;&lt;li&gt;Check that a cited source resolves and that its year matches the bibliography.&lt;/li&gt;&#10;&lt;li&gt;Decide whether to write the change to the file.&lt;/li&gt;&#10;&lt;/ol&gt;&#10;&lt;p&gt;Six of those have exactly one correct answer, computable without a model. Selection is an index lookup and change detection is a hash comparison. Budget is arithmetic; extraction is parsing. Source checking resolves the request and compares it against the bibliography. The write decision evaluates policy.&lt;/p&gt;</description></item><item><title>Scrum Built the Training Set</title><link>https://aibussin.com/books/applied-ai/04-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:04 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/04-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 1 — Where You Stand&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Why checkable work became unusually favorable terrain&lt;/strong&gt;&lt;/p&gt;&#10;&lt;h2 id="an-uncomfortable-piece-of-bookkeeping"&gt;An uncomfortable piece of bookkeeping&lt;/h2&gt;&#10;&lt;p&gt;Software engineering became an unusually fertile early target for language-model automation. The usual explanation is that code is logical and therefore tractable for a machine.&lt;/p&gt;&#10;&lt;p&gt;That explanation points at only part of what made software favorable. The account this chapter defends is about the surrounding work records and checks.&lt;/p&gt;&#10;&lt;p&gt;Over decades, software teams increasingly worked through issue trackers, version control, code review, automated tests, and iterative planning methods. Not every team used Scrum, and none of these practices guarantees the others, but together they often left behind unusually structured records:&lt;/p&gt;</description></item><item><title>Evidence and Verification</title><link>https://aibussin.com/books/agents-from-first-principles/10-chapter/</link><pubDate>Sat, 08 Aug 2026 17:31:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/10-chapter/</guid><description>&lt;p&gt;The agent says:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Done.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;That is a claim.&lt;/p&gt;&#10;&lt;p&gt;It is not evidence.&lt;/p&gt;&#10;&lt;p&gt;Every mechanism in this book so far has made the agent better at deciding what to do, and none of them establishes that the user&amp;rsquo;s goal was achieved. A planner can produce a coherent plan for the wrong problem. A tool can return exit code zero without producing the intended effect. A search can select the highest-scoring branch when every branch is wrong. A memory system can retrieve a perfectly relevant fact that stopped being true in March.&lt;/p&gt;</description></item><item><title>Building the Complete Agent</title><link>https://aibussin.com/books/agents-from-first-principles/11-chapter/</link><pubDate>Wed, 26 Aug 2026 10:00:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/11-chapter/</guid><description>&lt;p&gt;Every mechanism in this book was argued against a problem chosen to isolate it.&lt;/p&gt;&#10;&lt;p&gt;That isolation was deliberate, and it was also a form of protection. The memory chapter picked a task where recall was the bottleneck, held everything else still, and measured the one thing it came to measure. The result is a clean explanation and a weak claim. Nothing in it establishes that the same retrieval policy behaves when a search controller is expanding forty nodes, or when a verifier insists that every piece of evidence carry a state identity the search controller has never heard of.&lt;/p&gt;</description></item><item><title>From Measurements to Policy</title><link>https://aibussin.com/books/hallucination-from-first-principles/12-chapter/</link><pubDate>Sun, 30 Aug 2026 18:37:00 +0100</pubDate><guid>https://aibussin.com/books/hallucination-from-first-principles/12-chapter/</guid><description>&lt;p&gt;Chapter 11 ended with a typed reliability record.&lt;/p&gt;&#10;&lt;p&gt;It might say:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;containment = PASS&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;relation_fidelity = PASS&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;sensitivity = PASS&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;epistemic_adequacy = ANSWERABLE&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;provenance = UNVERIFIED&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That record describes the candidate.&lt;/p&gt;&#10;&lt;p&gt;It still does not authorize the candidate.&lt;/p&gt;&#10;&lt;p&gt;To make the distinction concrete, we built a deliberately flawed reference policy that checked containment, structural fidelity, and answerability but forgot to require verified provenance for a high-risk action.&lt;/p&gt;&#10;&lt;p&gt;Running the same immutable record through two policy versions produced:&lt;/p&gt;</description></item><item><title>Research: From Paper to Pipeline with AI</title><link>https://aibussin.com/books/freestyle-cognition/12-chapter/</link><pubDate>Fri, 28 Aug 2026 10:00:00 +0100</pubDate><guid>https://aibussin.com/books/freestyle-cognition/12-chapter/</guid><description>&lt;blockquote&gt;&#10;&lt;p&gt;&amp;ldquo;Research is formalized curiosity. It is poking and prying with a purpose.&amp;rdquo;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;&#10;&lt;p&gt;There is still something powerful about picking up a research paper and asking:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;What can I build from this?&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That question turns reading into motion. A paper stops being a sealed object. It becomes a source of ideas, mechanisms, tests, questions, and possible systems.&lt;/p&gt;&#10;&lt;p&gt;But the modern version of this workflow is more disciplined than the old one.&lt;/p&gt;</description></item><item><title>Verification, Repair, and Rejection</title><link>https://aibussin.com/books/hallucination-from-first-principles/13-chapter/</link><pubDate>Sun, 30 Aug 2026 18:38:00 +0100</pubDate><guid>https://aibussin.com/books/hallucination-from-first-principles/13-chapter/</guid><description>&lt;p&gt;Chapter 12 gave the system a control plane.&lt;/p&gt;&#10;&lt;p&gt;It can now say:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;commitment = HOLD&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;next_action = VERIFY&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;or:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;commitment = HOLD&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;next_action = RETRIEVE&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;or:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;commitment = HOLD&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;next_action = REFINE&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;or:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;commitment = DENY&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;response_mode = REJECTION&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Those are policy decisions.&lt;/p&gt;&#10;&lt;p&gt;They are not yet recovery implementations.&lt;/p&gt;&#10;&lt;p&gt;The obvious implementation is dangerously tempting:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;The answer failed a check.&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Ask the model to fix it.&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;But the model that produced the first unsupported statement can produce a second unsupported statement while &amp;ldquo;correcting&amp;rdquo; it.&lt;/p&gt;</description></item><item><title>Machine Coding: Building Software with AI</title><link>https://aibussin.com/books/freestyle-cognition/13-chapter/</link><pubDate>Fri, 28 Aug 2026 10:00:00 +0100</pubDate><guid>https://aibussin.com/books/freestyle-cognition/13-chapter/</guid><description>&lt;blockquote&gt;&#10;&lt;p&gt;&amp;ldquo;First, solve the problem. Then, write the code.&amp;rdquo;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;&#10;&lt;p&gt;Machine coding is no longer just asking an AI to write files for you.&lt;/p&gt;&#10;&lt;p&gt;That was the early version. You described an app, the model produced a pile of code, you copied it into a folder, ran it, pasted the errors back, and hoped the loop converged.&lt;/p&gt;&#10;&lt;p&gt;That workflow taught us something important: conversation can turn into software. But it also hid the real lesson.&lt;/p&gt;</description></item><item><title>Building Systems That Distrust Their Models</title><link>https://aibussin.com/books/hallucination-from-first-principles/15-chapter/</link><pubDate>Sun, 30 Aug 2026 21:01:00 +0100</pubDate><guid>https://aibussin.com/books/hallucination-from-first-principles/15-chapter/</guid><description>&lt;p&gt;The first chapter began with a simple observation:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;A language model can produce a fluent answer&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;without possessing a mechanism that proves the answer is true.&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Fourteen chapters later, that fact has not changed.&lt;/p&gt;&#10;&lt;p&gt;The model can still:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;invent&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;misbind&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;misattribute&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ignore decisive context&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;answer without enough evidence&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;accept bad retrieval&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;repair one error by creating another&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;repeat its own stored mistake&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The final architecture does not make those possibilities disappear.&lt;/p&gt;</description></item><item><title>The Agent Cannot Grade Its Own Homework</title><link>https://aibussin.com/books/applied-ai/21-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:21 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/21-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 4 — Make It Safe and Verifiable&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Independence, adequacy, and binding&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;A model saying its work is correct is not verification. Neither, on its own, is running a separate test. A test can run separately and check the wrong thing. It can check the right thing against a version of the artifact that has since changed. And a PASS that nobody can tie to the exact state being accepted is only another claim.&lt;/p&gt;</description></item><item><title>Build a Production AI Agent From First Principles: The Complete Reference Architecture</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/44-chapter/</link><pubDate>Sun, 09 Aug 2026 16:09:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/44-chapter/</guid><description>&lt;h1 id="build-a-production-ai-agent-from-first-principles-the-complete-reference-architecture"&gt;Build a Production AI Agent From First Principles: The Complete Reference Architecture&lt;/h1&gt;&#10;&lt;p&gt;We have spent this series adding mechanisms only when a specific failure demanded them.&lt;/p&gt;&#10;&lt;p&gt;We started with a model call.&lt;/p&gt;&#10;&lt;p&gt;Then we added candidate generation, critique, planning, tool use, memory, search and verification.&lt;/p&gt;&#10;&lt;p&gt;Then the system stopped looking like a clever prompt.&lt;/p&gt;&#10;&lt;p&gt;It started looking like software.&lt;/p&gt;&#10;&lt;p&gt;Then distributed systems problems arrived:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;duplicate work,&lt;/li&gt;&#10;&lt;li&gt;retries,&lt;/li&gt;&#10;&lt;li&gt;leases,&lt;/li&gt;&#10;&lt;li&gt;fencing,&lt;/li&gt;&#10;&lt;li&gt;backpressure,&lt;/li&gt;&#10;&lt;li&gt;dependency failure,&lt;/li&gt;&#10;&lt;li&gt;behavioral drift,&lt;/li&gt;&#10;&lt;li&gt;release compatibility,&lt;/li&gt;&#10;&lt;li&gt;replay,&lt;/li&gt;&#10;&lt;li&gt;incident forensics,&lt;/li&gt;&#10;&lt;li&gt;SLOs,&lt;/li&gt;&#10;&lt;li&gt;authority,&lt;/li&gt;&#10;&lt;li&gt;competence,&lt;/li&gt;&#10;&lt;li&gt;capability acquisition,&lt;/li&gt;&#10;&lt;li&gt;placement,&lt;/li&gt;&#10;&lt;li&gt;handoff,&lt;/li&gt;&#10;&lt;li&gt;stale state,&lt;/li&gt;&#10;&lt;li&gt;stale intent,&lt;/li&gt;&#10;&lt;li&gt;commitments,&lt;/li&gt;&#10;&lt;li&gt;durable workflows,&lt;/li&gt;&#10;&lt;li&gt;transaction recovery,&lt;/li&gt;&#10;&lt;li&gt;trust boundaries,&lt;/li&gt;&#10;&lt;li&gt;multi-agent coordination,&lt;/li&gt;&#10;&lt;li&gt;and finally an explicit control plane.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;At this point the architecture is complete enough that adding another isolated mechanism would make the series worse rather than better.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 19: Can Your Agent Explore in Parallel Without Creating Chaos? Use Speculative Execution and Early Cancellation</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/19-chapter/</link><pubDate>Sun, 09 Aug 2026 11:10:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/19-chapter/</guid><description>&lt;h1 id="advanced-agents-from-first-principles-19-can-your-agent-explore-in-parallel-without-creating-chaos-use-speculative-execution-and-early-cancellation"&gt;Advanced Agents From First Principles 19: Can Your Agent Explore in Parallel Without Creating Chaos? Use Speculative Execution and Early Cancellation&lt;/h1&gt;&#10;&lt;p&gt;A production agent often has more than one useful thing it could do next.&lt;/p&gt;&#10;&lt;p&gt;It could:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;inspect repository state,&lt;/li&gt;&#10;&lt;li&gt;run a targeted test,&lt;/li&gt;&#10;&lt;li&gt;retrieve documentation,&lt;/li&gt;&#10;&lt;li&gt;ask a second model to critique a candidate,&lt;/li&gt;&#10;&lt;li&gt;generate an alternative implementation,&lt;/li&gt;&#10;&lt;li&gt;probe an API,&lt;/li&gt;&#10;&lt;li&gt;inspect a deployment,&lt;/li&gt;&#10;&lt;li&gt;or verify an invariant.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;If those actions are independent, executing them one by one can be needlessly slow.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 18: What Should Your Agent Observe Next? Use Expected Value of Information</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/18-chapter/</link><pubDate>Sun, 09 Aug 2026 11:06:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/18-chapter/</guid><description>&lt;h1 id="what-should-your-agent-observe-next"&gt;What Should Your Agent Observe Next?&lt;/h1&gt;&#10;&lt;p&gt;Your agent is uncertain.&lt;/p&gt;&#10;&lt;p&gt;That does not tell you what to do.&lt;/p&gt;&#10;&lt;p&gt;In the previous post we split uncertainty into operational categories:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;interpretation uncertainty,&lt;/li&gt;&#10;&lt;li&gt;evidence uncertainty,&lt;/li&gt;&#10;&lt;li&gt;route uncertainty,&lt;/li&gt;&#10;&lt;li&gt;state uncertainty,&lt;/li&gt;&#10;&lt;li&gt;tool uncertainty,&lt;/li&gt;&#10;&lt;li&gt;candidate uncertainty,&lt;/li&gt;&#10;&lt;li&gt;verification uncertainty.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;That is already better than one generic confidence score.&lt;/p&gt;&#10;&lt;p&gt;But it still leaves a harder question:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;Which piece of information is worth buying next?&lt;/strong&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;Suppose a coding agent is trying to fix a failing test.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 17: What Is Your Agent Actually Uncertain About?</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/17-chapter/</link><pubDate>Sun, 09 Aug 2026 11:02:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/17-chapter/</guid><description>&lt;h1 id="what-is-your-agent-actually-uncertain-about"&gt;What Is Your Agent Actually Uncertain About?&lt;/h1&gt;&#10;&lt;p&gt;An agent reaches a difficult point in a task.&lt;/p&gt;&#10;&lt;p&gt;It is not sure what to do next.&lt;/p&gt;&#10;&lt;p&gt;A common implementation responds like this:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uncertain&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;call the model again&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;still uncertain&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;call a stronger model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;still uncertain&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;search more&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That is not a reasoning strategy.&lt;/p&gt;&#10;&lt;p&gt;It is a spending strategy.&lt;/p&gt;&#10;&lt;p&gt;The system is using more computation without identifying what information is actually missing.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 16: Where Should an Agent Spend Its Compute? Build a Dynamic Budget Scheduler</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/16-chapter/</link><pubDate>Sun, 09 Aug 2026 10:53:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/16-chapter/</guid><description>&lt;h1 id="where-should-an-agent-spend-its-compute"&gt;Where Should an Agent Spend Its Compute?&lt;/h1&gt;&#10;&lt;p&gt;A production agent has a budget whether you designed one or not.&lt;/p&gt;&#10;&lt;p&gt;Every model call costs something.&lt;/p&gt;&#10;&lt;p&gt;Every search node costs something.&lt;/p&gt;&#10;&lt;p&gt;Every tool invocation costs something.&lt;/p&gt;&#10;&lt;p&gt;Every verifier costs something.&lt;/p&gt;&#10;&lt;p&gt;Every retry adds latency.&lt;/p&gt;&#10;&lt;p&gt;Every escalation to a stronger model spends money and time that could have been used somewhere else.&lt;/p&gt;&#10;&lt;p&gt;The naive architecture gives every subsystem its own fixed limit:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;MAX_STEPS &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;20&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;MAX_SEARCH_NODES &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;32&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;MAX_CRITIC_CALLS &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;3&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;MAX_RETRIES &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;4&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;MAX_VERIFIER_CALLS &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;2&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That looks safe.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 15: How Do You Optimize an Agent Policy Without Turning It Into Another Black Box?</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/15-chapter/</link><pubDate>Sun, 09 Aug 2026 10:49:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/15-chapter/</guid><description>&lt;h1 id="how-do-you-optimize-an-agent-policy-without-turning-it-into-another-black-box"&gt;How Do You Optimize an Agent Policy Without Turning It Into Another Black Box?&lt;/h1&gt;&#10;&lt;p&gt;By now our advanced agent can do a lot.&lt;/p&gt;&#10;&lt;p&gt;It can:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;route tasks to different models or specialists,&lt;/li&gt;&#10;&lt;li&gt;decide whether to search,&lt;/li&gt;&#10;&lt;li&gt;choose a search budget,&lt;/li&gt;&#10;&lt;li&gt;decide when to escalate,&lt;/li&gt;&#10;&lt;li&gt;invoke critics,&lt;/li&gt;&#10;&lt;li&gt;retry or recover,&lt;/li&gt;&#10;&lt;li&gt;stop when evidence is strong enough,&lt;/li&gt;&#10;&lt;li&gt;learn from verified production trajectories,&lt;/li&gt;&#10;&lt;li&gt;and trace the decisions that produced each outcome.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;That creates a new problem.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 14: Can Your Agent Learn From Its Own Trajectories Without Learning the Wrong Lessons?</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/14-chapter/</link><pubDate>Sun, 09 Aug 2026 10:33:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/14-chapter/</guid><description>&lt;p&gt;An advanced agent now leaves behind something extremely valuable:&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;evidence.&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;Not merely chat history.&lt;/p&gt;&#10;&lt;p&gt;Not merely model outputs.&lt;/p&gt;&#10;&lt;p&gt;Not merely traces.&lt;/p&gt;&#10;&lt;p&gt;A sufficiently instrumented system can record:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;what state it was in,&lt;/li&gt;&#10;&lt;li&gt;what alternatives it considered,&lt;/li&gt;&#10;&lt;li&gt;which route it selected,&lt;/li&gt;&#10;&lt;li&gt;what branches it pruned,&lt;/li&gt;&#10;&lt;li&gt;which model or specialist it escalated to,&lt;/li&gt;&#10;&lt;li&gt;which tools it called,&lt;/li&gt;&#10;&lt;li&gt;which critic changed the answer,&lt;/li&gt;&#10;&lt;li&gt;what verification evidence was produced,&lt;/li&gt;&#10;&lt;li&gt;how much compute was spent,&lt;/li&gt;&#10;&lt;li&gt;and whether the final result actually passed.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;That immediately suggests a tempting idea:&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 12: Is Your Advanced Agent Actually Better? Benchmark It Under Equal Budgets</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/12-chapter/</link><pubDate>Sun, 09 Aug 2026 10:10:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/12-chapter/</guid><description>&lt;h1 id="is-your-advanced-agent-actually-better-benchmark-it-under-equal-budgets"&gt;Is Your Advanced Agent Actually Better? Benchmark It Under Equal Budgets&lt;/h1&gt;&#10;&lt;p&gt;You replace one model call with eight.&lt;/p&gt;&#10;&lt;p&gt;Success rises from 62% to 74%.&lt;/p&gt;&#10;&lt;p&gt;Great.&lt;/p&gt;&#10;&lt;p&gt;Except the new system used:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;eight times the inference,&lt;/li&gt;&#10;&lt;li&gt;three extra judges,&lt;/li&gt;&#10;&lt;li&gt;two rounds of critique,&lt;/li&gt;&#10;&lt;li&gt;a larger context,&lt;/li&gt;&#10;&lt;li&gt;a stronger verifier,&lt;/li&gt;&#10;&lt;li&gt;and several times the latency.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Did the architecture improve?&lt;/p&gt;&#10;&lt;p&gt;Or did you just buy more attempts?&lt;/p&gt;&#10;&lt;p&gt;This is one of the easiest mistakes to make in advanced agent engineering.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 11: Which Advanced Agent Architecture Should You Use? A Practical Selection Guide</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/11-chapter/</link><pubDate>Sun, 09 Aug 2026 09:20:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/11-chapter/</guid><description>&lt;h1 id="which-advanced-agent-architecture-should-you-use"&gt;Which Advanced Agent Architecture Should You Use?&lt;/h1&gt;&#10;&lt;p&gt;You now have too many options.&lt;/p&gt;&#10;&lt;p&gt;That is a better problem than having none.&lt;/p&gt;&#10;&lt;p&gt;But it is still a problem.&lt;/p&gt;&#10;&lt;p&gt;You can add:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;self-consistency,&lt;/li&gt;&#10;&lt;li&gt;Tree of Thoughts,&lt;/li&gt;&#10;&lt;li&gt;beam search,&lt;/li&gt;&#10;&lt;li&gt;Monte Carlo Tree Search,&lt;/li&gt;&#10;&lt;li&gt;evolutionary search,&lt;/li&gt;&#10;&lt;li&gt;specialist routing,&lt;/li&gt;&#10;&lt;li&gt;planner/executor/critic separation,&lt;/li&gt;&#10;&lt;li&gt;multi-agent debate,&lt;/li&gt;&#10;&lt;li&gt;adaptive policies,&lt;/li&gt;&#10;&lt;li&gt;learning from previous runs,&lt;/li&gt;&#10;&lt;li&gt;or a mixture-of-agents runtime that chooses among several of them.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The temptation is to combine everything.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 08: Is Your Agent Spending the Same Compute on Every Task? Build Adaptive Agents That Escalate Only When Needed</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/08-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/08-chapter/</guid><description>Build adaptive agent runtimes that start cheap, measure uncertainty and failure, and escalate selectively into deeper reasoning, more samples, search, stronger models or specialist review only when the evidence justifies it.</description></item><item><title>Advanced Agents From First Principles 26: Why Did the Agent Fail? Build an Incident Forensics Pipeline</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/26-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/26-chapter/</guid><description>A practical incident-forensics workflow for advanced agents: reconstruct the run, find the earliest divergence, distinguish root cause from downstream symptoms, measure blast radius, and prove that a remediation would have prevented the incident.</description></item><item><title>Advanced Agents From First Principles 27: How Reliable Does an Agent Need to Be? Define SLOs and Error Budgets</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/27-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/27-chapter/</guid><description>A practical reliability framework for advanced agents: define verified-success SLOs, false-success ceilings, UNKNOWN budgets, latency and cost targets, then use error-budget burn to decide when to ship capability and when to stop and harden the system.</description></item><item><title>Advanced Agents From First Principles 28: Where Should You Spend the Next Engineering Hour? Prioritize Reliability by Risk and Expected Return</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/28-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/28-chapter/</guid><description>A practical framework for deciding where to spend the next engineering hour in an advanced-agent system: rank remediation by expected reduction in verified reliability loss, severity, recurrence, blast radius, confidence and implementation cost.</description></item><item><title>Advanced Agents From First Principles 29: When Should an Agent Stop and Ask a Human? Design Authority Boundaries and Escalation</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/29-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/29-chapter/</guid><description>A practical architecture for agent authority boundaries: decide what an agent may do autonomously, when it must escalate, what evidence a human reviewer needs, and how to avoid turning human approval into rubber-stamping.</description></item><item><title>Advanced Agents From First Principles 30: Is This Task Outside Your Agent’s Competence? Build Competence Envelopes and OOD Detection</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/30-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/30-chapter/</guid><description>A practical framework for competence envelopes in production agents: distinguish uncertainty from lack of validated competence, detect out-of-distribution tasks, contract authority when evidence is weak, and expand autonomy only through measured evidence.</description></item><item><title>Advanced Agents From First Principles 31: How Can an Agent Learn New Capabilities Without Expanding Its Own Authority? Use Sandboxed Capability Acquisition</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/31-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/31-chapter/</guid><description>A practical architecture for sandboxed capability acquisition: let agents explore tasks outside their validated competence envelope, accumulate externally verified evidence, and propose capability expansion without ever granting themselves production authority.</description></item><item><title>Advanced Agents From First Principles 32: Which Capabilities Are Actually Worth Building? Design a Capability Portfolio</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/32-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/32-chapter/</guid><description>A practical framework for deciding which agent capabilities are worth acquiring: rank missing capabilities by user value, verifier availability, reliability risk, acquisition cost, maintenance burden, and the quality of human or deterministic alternatives.</description></item><item><title>Advanced Agents From First Principles 33: Which Shared Components Actually Unlock More Capability? Build a Capability Dependency Graph</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/33-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/33-chapter/</guid><description>A practical capability-dependency architecture for advanced agents: identify shared primitives that unlock many capabilities, quantify leverage, expose correlated-failure hotspots, and invest in platform components without creating hidden systemic risk.</description></item><item><title>Advanced Agents From First Principles 34: Where Should This Task Actually Run? Build Capability-Aware Placement Across Models, Providers and Resource Pools</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/34-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/34-chapter/</guid><description>A practical placement architecture for advanced agents: route work across local and frontier models, providers, regions, GPUs, browser pools and specialist runtimes using demonstrated competence, verifier availability, policy constraints, health, cost and latency rather than model prestige.</description></item><item><title>Advanced Agents From First Principles 35: How Do You Move a Running Agent Between Workers Without Losing Meaning? Build Portable Execution State and Safe Handoff</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/35-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/35-chapter/</guid><description>A practical architecture for portable agent execution state: checkpoint long-running runs, transfer ownership safely across workers and providers, preserve evidence and authority, and reject migrations that cannot be proven compatible.</description></item><item><title>Advanced Agents From First Principles 36: Is Your Agent Acting on Stale State? Build Temporal Consistency, Freshness Budgets and Conflict Detection</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/36-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/36-chapter/</guid><description>A practical architecture for keeping long-running agents from acting on stale assumptions: classify state by freshness, track version vectors, detect conflicts, revalidate before consequential actions, and force replanning when the world has changed underneath the run.</description></item><item><title>Advanced Agents From First Principles 37: Is Your Agent Still Solving the Right Task? Build Intent Versioning, Supersession and Cancellation</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/37-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/37-chapter/</guid><description>A production architecture for intent versioning, supersession and cancellation in long-running agents: stop obsolete work, preserve committed effects, reconcile partial actions, and prevent stale goals from retaining authority.</description></item><item><title>Advanced Agents From First Principles 39: How Do You Make an Agent Survive for Days? Build Durable Long-Running Workflows</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/39-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/39-chapter/</guid><description>A practical architecture for long-running agents: keep workflow state durable while treating models and workers as disposable, with explicit waits, retries, timers, human approvals, checkpoints, commitments, cancellation and replay.</description></item><item><title>Advanced Agents From First Principles 41: What Should Your Agent Trust? Build Explicit Security and Trust Boundaries</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/41-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/41-chapter/</guid><description>A practical security architecture for production agents: separate data from authority, classify trust, scope credentials and capabilities, preserve provenance, isolate generated code, resist prompt injection, and keep security-critical decisions outside model control.</description></item><item><title>Advanced Agents From First Principles 42: How Do Multiple Agents Coordinate Without Becoming a Distributed Argument?</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/42-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/42-chapter/</guid><description>A practical architecture for multi-agent coordination: explicit ownership, delegation, contracts, commitment transfer, shared intent, evidence provenance, conflict handling, deadlock prevention, and independent verification instead of agents merely chatting until they agree.</description></item><item><title>Advanced Agents From First Principles 43: Who Controls the Agent? Build an Explicit Agent Control Plane</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/43-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/43-chapter/</guid><description>A production-agent architecture that separates control-plane policy from execution-plane reasoning: intent, competence, authority, placement, budgets, reliability, releases, security and escalation remain enforceable outside the model.</description></item><item><title>You Probably Don't Need All of This: Build the Minimum Production Agent Architecture</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/45-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/45-chapter/</guid><description>&lt;h1 id="you-probably-dont-need-all-of-this"&gt;You Probably Don&amp;rsquo;t Need All of This&lt;/h1&gt;&#10;&lt;p&gt;Over the previous forty-five steps, we built almost every major mechanism you might need in a serious agent platform.&lt;/p&gt;&#10;&lt;p&gt;Search.&lt;/p&gt;&#10;&lt;p&gt;Critique.&lt;/p&gt;&#10;&lt;p&gt;Planning.&lt;/p&gt;&#10;&lt;p&gt;Memory.&lt;/p&gt;&#10;&lt;p&gt;Verification.&lt;/p&gt;&#10;&lt;p&gt;Distributed execution.&lt;/p&gt;&#10;&lt;p&gt;Leases.&lt;/p&gt;&#10;&lt;p&gt;Fencing.&lt;/p&gt;&#10;&lt;p&gt;Backpressure.&lt;/p&gt;&#10;&lt;p&gt;Behavioral releases.&lt;/p&gt;&#10;&lt;p&gt;Replay.&lt;/p&gt;&#10;&lt;p&gt;Incident forensics.&lt;/p&gt;&#10;&lt;p&gt;SLOs.&lt;/p&gt;&#10;&lt;p&gt;Competence envelopes.&lt;/p&gt;&#10;&lt;p&gt;Authority boundaries.&lt;/p&gt;&#10;&lt;p&gt;Capability portfolios.&lt;/p&gt;&#10;&lt;p&gt;Placement.&lt;/p&gt;&#10;&lt;p&gt;Portable execution state.&lt;/p&gt;&#10;&lt;p&gt;Temporal consistency.&lt;/p&gt;&#10;&lt;p&gt;Intent versioning.&lt;/p&gt;&#10;&lt;p&gt;Commitments.&lt;/p&gt;&#10;&lt;p&gt;Durable workflows.&lt;/p&gt;&#10;&lt;p&gt;Transaction recovery.&lt;/p&gt;&#10;&lt;p&gt;Security boundaries.&lt;/p&gt;&#10;&lt;p&gt;Multi-agent coordination.&lt;/p&gt;&#10;&lt;p&gt;An explicit control plane.&lt;/p&gt;&#10;&lt;p&gt;And finally, in Step 44, we assembled those ideas into a complete reference architecture for a production AI agent.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 10: Are You Combining Every Agent Technique Into One Monster? Build a Mixture-of-Agents Runtime</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/10-chapter/</link><pubDate>Sun, 09 Aug 2026 00:25:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/10-chapter/</guid><description>&lt;p&gt;You have a working agent.&lt;/p&gt;&#10;&lt;p&gt;Then you add retrieval.&lt;/p&gt;&#10;&lt;p&gt;Then memory.&lt;/p&gt;&#10;&lt;p&gt;Then Best-of-N.&lt;/p&gt;&#10;&lt;p&gt;Then critique and revision.&lt;/p&gt;&#10;&lt;p&gt;Then Tree of Thoughts.&lt;/p&gt;&#10;&lt;p&gt;Then MCTS.&lt;/p&gt;&#10;&lt;p&gt;Then specialist models.&lt;/p&gt;&#10;&lt;p&gt;Then adversarial review.&lt;/p&gt;&#10;&lt;p&gt;Then a planner, executor, critic and verifier.&lt;/p&gt;&#10;&lt;p&gt;Then a stronger model for hard cases.&lt;/p&gt;&#10;&lt;p&gt;Eventually the architecture starts to look like this:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;request&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;planner&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;retrieval&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;reasoning&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Best-of-N&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;critic&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Tree of Thoughts&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;MCTS&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;frontier model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;second critic&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;verifier&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Every mechanism may have been individually reasonable.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 09: Can Your Agent Actually Learn From Previous Runs?</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/09-chapter/</link><pubDate>Sun, 09 Aug 2026 00:19:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/09-chapter/</guid><description>&lt;h1 id="can-your-agent-actually-learn-from-previous-runs"&gt;Can Your Agent Actually Learn From Previous Runs?&lt;/h1&gt;&#10;&lt;p&gt;A production agent can execute the same class of task hundreds or thousands of times.&lt;/p&gt;&#10;&lt;p&gt;It can see the same failure repeatedly.&lt;/p&gt;&#10;&lt;p&gt;It can discover the same workaround repeatedly.&lt;/p&gt;&#10;&lt;p&gt;It can call the same expensive model repeatedly.&lt;/p&gt;&#10;&lt;p&gt;And still behave as if every task is the first one it has ever seen.&lt;/p&gt;&#10;&lt;p&gt;That is not necessarily a memory problem.&lt;/p&gt;&#10;&lt;p&gt;It may already have excellent memory.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 06: Does One Agent Plan, Execute and Judge Its Own Work? Build a Planner-Executor-Critic Architecture</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/06-chapter/</link><pubDate>Sat, 08 Aug 2026 23:49:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/06-chapter/</guid><description>&lt;h1 id="does-one-agent-plan-execute-and-judge-its-own-work-build-a-planner-executor-critic-architecture"&gt;Does One Agent Plan, Execute and Judge Its Own Work? Build a Planner-Executor-Critic Architecture&lt;/h1&gt;&#10;&lt;p&gt;A single model can often do all of these things:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;understand a task,&lt;/li&gt;&#10;&lt;li&gt;decide what to do,&lt;/li&gt;&#10;&lt;li&gt;execute a tool call,&lt;/li&gt;&#10;&lt;li&gt;inspect the result,&lt;/li&gt;&#10;&lt;li&gt;critique its own work,&lt;/li&gt;&#10;&lt;li&gt;decide whether it succeeded,&lt;/li&gt;&#10;&lt;li&gt;and produce the final answer.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;That is convenient.&lt;/p&gt;&#10;&lt;p&gt;It is also a dangerous concentration of responsibilities.&lt;/p&gt;&#10;&lt;p&gt;If the same component creates the plan, executes it, explains why the result is good, and decides whether the job is complete, then failures become difficult to localize.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 02: Why Does My Reasoning Agent Give a Different Answer Every Time? Use Self-Consistency Without Confusing Consensus With Truth</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/02-chapter/</link><pubDate>Sat, 08 Aug 2026 22:44:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/02-chapter/</guid><description>&lt;p&gt;A reasoning agent gives you one answer.&lt;/p&gt;&#10;&lt;p&gt;You run it again.&lt;/p&gt;&#10;&lt;p&gt;It gives you another.&lt;/p&gt;&#10;&lt;p&gt;You change nothing important:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;same task,&lt;/li&gt;&#10;&lt;li&gt;same tools,&lt;/li&gt;&#10;&lt;li&gt;same model family,&lt;/li&gt;&#10;&lt;li&gt;same broad context.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Yet the result changes.&lt;/p&gt;&#10;&lt;p&gt;That is not necessarily a bug.&lt;/p&gt;&#10;&lt;p&gt;A probabilistic model is allowed to produce more than one plausible trajectory.&lt;/p&gt;&#10;&lt;p&gt;The engineering question is different:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;How should an agent system use that variation?&lt;/strong&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;One common answer is &lt;strong&gt;self-consistency&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;Generate several independent reasoning trajectories.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 01: Does Your AI Agent Fail on Complex Reasoning Tasks? Treat Chain of Thought as Computation, Not Proof</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/01-chapter/</link><pubDate>Sat, 08 Aug 2026 22:35:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/01-chapter/</guid><description>&lt;p&gt;Most developers first encounter chain of thought as a prompting trick:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Think step by step.&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That framing is too shallow for agent engineering.&lt;/p&gt;&#10;&lt;p&gt;For an advanced agent, the useful idea is not that the model should produce a long explanation. The useful idea is that a difficult task may benefit from &lt;strong&gt;intermediate computational state&lt;/strong&gt; before the system commits to an action or answer.&lt;/p&gt;&#10;&lt;p&gt;That is a very different claim.&lt;/p&gt;&#10;&lt;p&gt;A reasoning trace can help a system decompose a problem, preserve intermediate conclusions, identify missing information, decide what to verify next, and expose places where search or tools should be used.&lt;/p&gt;</description></item><item><title>Agents From First Principles 09: AI Agent Says It Worked When It Didn’t? Verify the Result Outside the LLM</title><link>https://aibussin.com/post/agents-from-first-principles-09/</link><pubDate>Sat, 08 Aug 2026 17:31:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-09/</guid><description>&lt;p&gt;An AI agent says:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Done. The task is complete.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;That sentence is almost worthless.&lt;/p&gt;&#10;&lt;p&gt;The agent may have:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;edited the wrong file,&lt;/li&gt;&#10;&lt;li&gt;changed the right file incorrectly,&lt;/li&gt;&#10;&lt;li&gt;skipped part of the request,&lt;/li&gt;&#10;&lt;li&gt;broken another subsystem,&lt;/li&gt;&#10;&lt;li&gt;failed to save its work,&lt;/li&gt;&#10;&lt;li&gt;misread a tool result,&lt;/li&gt;&#10;&lt;li&gt;passed a stale test,&lt;/li&gt;&#10;&lt;li&gt;inspected the wrong environment,&lt;/li&gt;&#10;&lt;li&gt;or simply decided that its own answer looked convincing.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The central problem is simple:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;The system that produced the answer should not be the only system deciding whether the answer is correct.&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 07: Do Your Agents Agree Too Easily? Use Adversarial Review and Multi-Agent Debate Without Confusing Debate With Truth</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/07-chapter/</link><pubDate>Sat, 08 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/07-chapter/</guid><description>&lt;p&gt;A multi-agent system can look sophisticated while every agent quietly repeats the same mistake.&lt;/p&gt;&#10;&lt;p&gt;That is one of the most dangerous failure modes in advanced agent architectures.&lt;/p&gt;&#10;&lt;p&gt;You ask one model to solve the problem.&lt;/p&gt;&#10;&lt;p&gt;Then you ask a second model to review it.&lt;/p&gt;&#10;&lt;p&gt;Then a third model judges the disagreement.&lt;/p&gt;&#10;&lt;p&gt;Three calls later, the system sounds more confident than before.&lt;/p&gt;&#10;&lt;p&gt;But if all three agents share the same blind spot, the extra machinery has not created independent evidence.&lt;/p&gt;</description></item><item><title>Codex Manager: Building a Prompt-State Runtime for Hackathon-Grade Code Optimization</title><link>https://aibussin.com/post/codex/</link><pubDate>Sat, 16 May 2026 20:15:59 +0100</pubDate><guid>https://aibussin.com/post/codex/</guid><description>&lt;h2 id="tldr"&gt;TL;DR&lt;/h2&gt;&#10;&lt;p&gt;Codex Manager uses AI to generate code as an artifact, then tests that artifact, diagnoses what happened, and repairs the prompt state that produced it. The code is not the thing being optimized directly. The prompt state is.&lt;/p&gt;&#10;&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;&#10;&lt;p&gt;&lt;a href="https://huggingface.co/humanitys-last-hackathon?utm_source=chatgpt.com"&gt;Humanity’s Last Hackathon&lt;/a&gt; framed the challenge as a test of &lt;strong&gt;context, not code&lt;/strong&gt;: the task was hard enough that the real question was not whether someone could hand-write one clever kernel, but whether they could build a system that used AI effectively under changing constraints.&lt;/p&gt;</description></item><item><title>A Memory Gate for AI: Policy-Bounded Acceptance in the Executable Cognitive Kernel</title><link>https://aibussin.com/post/verify/</link><pubDate>Tue, 17 Mar 2026 09:58:14 +0000</pubDate><guid>https://aibussin.com/post/verify/</guid><description>&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;&#10;&lt;p&gt;Dynamic AI systems face a hidden failure mode: they can learn from their own mistakes.&#10;If every output is allowed into memory, stochastic errors do not stay local they accumulate.&lt;/p&gt;&#10;&lt;p&gt;In earlier posts, I argued that AI systems should not be trusted to enforce their own correctness.&lt;/p&gt;&#10;&lt;p&gt;Modern models are stochastic. They produce correct outputs, partially correct outputs, and completely incorrect outputs, but they do not reliably distinguish between them. That means a system that stores everything it generates will eventually learn from its own mistakes.&lt;/p&gt;</description></item><item><title>Episteme: Distilling Knowledge into AI</title><link>https://aibussin.com/post/episteme/</link><pubDate>Fri, 03 Oct 2025 11:24:58 +0100</pubDate><guid>https://aibussin.com/post/episteme/</guid><description>&lt;h2 id="-summary"&gt;🚀 Summary&lt;/h2&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;When you can measure what you are speaking about… you know something about it; but when you cannot measure it… your knowledge is of a meagre and unsatisfactory kind. &lt;em&gt;Lord Kelvin&lt;/em&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;Remember that time you spent an hour with an AI, and in one perfect response, it solved a problem you&amp;rsquo;d been stuck on for weeks?&lt;/strong&gt; Where is that answer now? Lost in a scroll of chat history, a fleeting moment of brilliance that vanished as quickly as it appeared. This post is about how to make that moment permanent, and turn it into an intelligence that amplifies everything you do.&lt;/p&gt;</description></item></channel></rss>