<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Agent Debugging on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/agent-debugging/</link><description>Recent content in Agent Debugging on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Sat, 08 Aug 2026 17:26:00 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/agent-debugging/index.xml" rel="self" type="application/rss+xml"/><item><title>The Action Boundary</title><link>https://aibussin.com/books/agents-from-first-principles/02-chapter/</link><pubDate>Sat, 08 Aug 2026 15:46:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/02-chapter/</guid><description>&lt;p&gt;The previous chapter ended with a division of responsibility:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;The model proposes. The runtime decides what may execute. The environment supplies observations about what happened.&lt;/strong&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;This chapter takes the first clause seriously. What does it actually mean for a model to &lt;strong&gt;propose an action&lt;/strong&gt;?&lt;/p&gt;&#10;&lt;p&gt;Suppose the model emits the sentence &lt;em&gt;&amp;ldquo;Search the documentation for the latest PyTorch optimizer API.&amp;rdquo;&lt;/em&gt; A human reads that and understands the intention immediately. A runtime cannot act on it at all, because it needs answers to six questions the sentence does not contain: which action, which arguments, whether those arguments are well formed, whether their values are meaningful, whether this action is permitted in this run, and whether its execution preconditions currently hold.&lt;/p&gt;</description></item><item><title>Candidate Generation and Selection</title><link>https://aibussin.com/books/agents-from-first-principles/03-chapter/</link><pubDate>Sat, 08 Aug 2026 15:54:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/03-chapter/</guid><description>&lt;p&gt;The previous chapter built a boundary that stops arbitrary model output from acquiring execution authority without explicit checks. It made one class of failure inspectable and enforceable, and it is silent about another.&lt;/p&gt;&#10;&lt;p&gt;Suppose the model is asked to solve a coding problem. One run produces the right patch. The next produces a plausible but incomplete one. A third produces something better again. Nothing is malformed, nothing violates the action schema, and every one of them would pass the boundary we just built. Validity and quality are different properties. A proposal can be completely valid and still be a poor choice, which relocates the uncertainty rather than removing it:&lt;/p&gt;</description></item><item><title>Critique, Revision, and Acceptance</title><link>https://aibussin.com/books/agents-from-first-principles/04-chapter/</link><pubDate>Sat, 08 Aug 2026 16:23:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/04-chapter/</guid><description>&lt;p&gt;Selection can only choose among the candidates it is given. A strong evaluator may recognize that every available candidate is poor, but it cannot select a correct answer that the generator never produced.&lt;/p&gt;&#10;&lt;p&gt;That limitation becomes practical because the candidates usually come from one model answering one prompt, and such samples correlate. When four candidates share the same misreading of the evidence, ranking them can still produce a confident winner while leaving the shared defect untouched.&lt;/p&gt;</description></item><item><title>Planning and Execution</title><link>https://aibussin.com/books/agents-from-first-principles/05-chapter/</link><pubDate>Sat, 08 Aug 2026 16:35:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/05-chapter/</guid><description>&lt;p&gt;The critique loop works on one thing at a time. It assumes the task already exists as a candidate we can hold, inspect and improve.&lt;/p&gt;&#10;&lt;p&gt;Some tasks have no draft to hold.&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Inspect a project, reproduce the failing test, find the cause, patch the code, rerun the relevant tests, and report what changed.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;No single narrow action responsibly completes that goal, and the actions constrain each other. Patching before diagnosis is guesswork wearing the costume of work, and reporting success before observing a passing test is a claim about the world that nothing in the run supports. Worse, if execution reveals that an assumption was wrong, the rest of the route may be wrong too. A runtime choosing each action independently can react to the new observation, but without an explicit representation of the intended route it has nothing concrete to compare the changed world against.&lt;/p&gt;</description></item><item><title>Runtime State, Progress, and Termination</title><link>https://aibussin.com/books/agents-from-first-principles/06-chapter/</link><pubDate>Sat, 08 Aug 2026 16:56:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/06-chapter/</guid><description>&lt;p&gt;A plan is a claim about the future. It says that reproducing the failure, then diagnosing it, then patching, then testing, is a route from here to a working system. Making that claim explicit was worth the machinery. But it is written before any of the work happens, and execution may invalidate one of its assumptions almost immediately.&lt;/p&gt;&#10;&lt;p&gt;Execution produces something else entirely: a trajectory. Actions went out, observations came back, and some of what came back may contradict the intended route. The patch step failed because a dependency is missing. The test step is not merely late; it is unreachable.&lt;/p&gt;</description></item><item><title>Capabilities and Routing</title><link>https://aibussin.com/books/agents-from-first-principles/07-chapter/</link><pubDate>Sat, 08 Aug 2026 17:09:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/07-chapter/</guid><description>&lt;p&gt;Every mechanism built so far has taken the action set as given. The action boundary validates a proposal against a fixed list of permitted action types. Candidate generation samples several proposals from the same list. The runtime-state chapter watches what happens after execution and decides whether to continue. All of them assume that somebody, somewhere, already decided which capabilities the policy could choose from.&lt;/p&gt;&#10;&lt;p&gt;So far we have treated that decision as an input. This chapter makes it explicit, because the capability surface changes the decision problem the policy has to solve and therefore belongs inside the agent architecture itself.&lt;/p&gt;</description></item><item><title>Trajectory Search</title><link>https://aibussin.com/books/agents-from-first-principles/09-chapter/</link><pubDate>Sat, 08 Aug 2026 17:26:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/09-chapter/</guid><description>&lt;p&gt;An agent can make every local decision look reasonable and still lose the task.&lt;/p&gt;&#10;&lt;p&gt;A coding agent sees that &lt;code&gt;test_checkout_redirect&lt;/code&gt; is failing, concludes there is an implementation bug in &lt;code&gt;checkout.py&lt;/code&gt;, and then behaves impeccably for twenty steps: it reads the file, edits it, runs the tests, repairs the new failures its edit introduced, rewrites the patch, and runs the tests again. Every one of those steps is defensible given the step before it.&lt;/p&gt;</description></item><item><title>Agents From First Principles 08: AI Agent Picks the First Solution? Add Search Instead of One-Shot Generation</title><link>https://aibussin.com/post/agents-from-first-principles-08/</link><pubDate>Sat, 08 Aug 2026 17:26:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-08/</guid><description>&lt;p&gt;An AI agent often fails for a surprisingly ordinary reason:&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;it commits too early.&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;It finds one plausible next action, follows it, and then spends the rest of the run trying to make that first choice work.&lt;/p&gt;&#10;&lt;p&gt;That can look intelligent because the agent keeps reasoning, calling tools, revising plans, and explaining itself.&lt;/p&gt;&#10;&lt;p&gt;But underneath, the trajectory may be almost completely determined by an early mistake.&lt;/p&gt;&#10;&lt;p&gt;A coding agent chooses the wrong implementation strategy and spends twenty tool calls repairing it.&lt;/p&gt;</description></item><item><title>Agents From First Principles 05: AI Agent Gets Stuck in a Loop? Add State, Feedback and Stopping Conditions</title><link>https://aibussin.com/post/agents-from-first-principles-05/</link><pubDate>Sat, 08 Aug 2026 16:56:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-05/</guid><description>&lt;p&gt;An AI agent that keeps calling the same tool, revisiting the same page, rewriting the same file, or repeatedly saying “I’ll try again” is not displaying persistence.&lt;/p&gt;&#10;&lt;p&gt;It is displaying a control-flow bug.&lt;/p&gt;&#10;&lt;p&gt;This is one of the most common failure modes in agent software because the basic loop is deceptively simple:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;observe&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;decide&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;act&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;observe&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;repeat&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The problem is hidden inside the final word.&lt;/p&gt;</description></item><item><title>Agents From First Principles 04: AI Agent Fails on Multi-Step Tasks? Separate Planning From Execution</title><link>https://aibussin.com/post/agents-from-first-principles-04/</link><pubDate>Sat, 08 Aug 2026 16:35:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-04/</guid><description>&lt;p&gt;A surprising number of agent failures are not really model failures.&lt;/p&gt;&#10;&lt;p&gt;The model may be perfectly capable of writing each individual step. The failure happens because the system tries to decide &lt;strong&gt;what to do&lt;/strong&gt; and &lt;strong&gt;do it&lt;/strong&gt; at the same time.&lt;/p&gt;&#10;&lt;p&gt;That works for simple tasks:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;question&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;answer&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It becomes fragile when success depends on several ordered actions:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;goal&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;step 1&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;step 2&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;step 3&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;verification&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A useful next step in agent design is therefore to separate two jobs:&lt;/p&gt;</description></item><item><title>Agents From First Principles 03: AI Agent Keeps Making the Same Mistake? Add a Critique-and-Revision Loop</title><link>https://aibussin.com/post/agents-from-first-principles-03/</link><pubDate>Sat, 08 Aug 2026 16:23:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-03/</guid><description>&lt;p&gt;An AI agent can fail in a particularly frustrating way: it produces an answer that is almost right, you ask it to improve the answer, and it produces another answer with the same underlying defect.&lt;/p&gt;&#10;&lt;p&gt;Sometimes the wording changes. Sometimes it adds more explanation. Sometimes it becomes longer and more confident. But the important mistake survives.&lt;/p&gt;&#10;&lt;p&gt;That usually means the system is doing this:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;prompt&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;answer&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;or this:&lt;/p&gt;</description></item><item><title>Agents From First Principles 02: AI Agent Gives Inconsistent Answers? Generate Multiple Candidates and Rank Them</title><link>https://aibussin.com/post/agents-from-first-principles-02/</link><pubDate>Sat, 08 Aug 2026 15:54:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-02/</guid><description>&lt;p&gt;One of the first things you notice when you build anything around a large language model is that the same prompt does not always produce the same quality of answer.&lt;/p&gt;&#10;&lt;p&gt;Sometimes the first response is excellent.&lt;/p&gt;&#10;&lt;p&gt;Sometimes it is merely acceptable.&lt;/p&gt;&#10;&lt;p&gt;Sometimes it misses the point entirely.&lt;/p&gt;&#10;&lt;p&gt;That creates a very common agent-engineering question:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;If the model is inconsistent, should the agent trust the first answer it gets?&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;Often, no.&lt;/p&gt;</description></item><item><title>AI Agent Returning Invalid Tool Calls? How to Validate LLM Actions</title><link>https://aibussin.com/post/agents-from-first-principles-01/</link><pubDate>Sat, 08 Aug 2026 15:46:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-01/</guid><description>&lt;p&gt;If your agent sometimes invents a tool name, omits a required argument, returns malformed JSON, or produces an action that looks plausible but cannot actually be executed, the problem is usually not &amp;ldquo;the agent is dumb.&amp;rdquo;&lt;/p&gt;&#10;&lt;p&gt;The problem is that &lt;strong&gt;raw language-model output has been allowed to cross directly into execution&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;That boundary is too weak.&lt;/p&gt;&#10;&lt;p&gt;The simplest useful agent architecture is not:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;prompt&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;execute whatever came back&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It is:&lt;/p&gt;</description></item></channel></rss>