<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Agentic AI on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/agentic-ai/</link><description>Recent content in Agentic AI on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Wed, 26 Aug 2026 10:00:00 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/agentic-ai/index.xml" rel="self" type="application/rss+xml"/><item><title>Evidence and Verification</title><link>https://aibussin.com/books/agents-from-first-principles/10-chapter/</link><pubDate>Sat, 08 Aug 2026 17:31:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/10-chapter/</guid><description>&lt;p&gt;The agent says:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Done.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;That is a claim.&lt;/p&gt;&#10;&lt;p&gt;It is not evidence.&lt;/p&gt;&#10;&lt;p&gt;Every mechanism in this book so far has made the agent better at deciding what to do, and none of them establishes that the user&amp;rsquo;s goal was achieved. A planner can produce a coherent plan for the wrong problem. A tool can return exit code zero without producing the intended effect. A search can select the highest-scoring branch when every branch is wrong. A memory system can retrieve a perfectly relevant fact that stopped being true in March.&lt;/p&gt;</description></item><item><title>Building the Complete Agent</title><link>https://aibussin.com/books/agents-from-first-principles/11-chapter/</link><pubDate>Wed, 26 Aug 2026 10:00:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/11-chapter/</guid><description>&lt;p&gt;Every mechanism in this book was argued against a problem chosen to isolate it.&lt;/p&gt;&#10;&lt;p&gt;That isolation was deliberate, and it was also a form of protection. The memory chapter picked a task where recall was the bottleneck, held everything else still, and measured the one thing it came to measure. The result is a clean explanation and a weak claim. Nothing in it establishes that the same retrieval policy behaves when a search controller is expanding forty nodes, or when a verifier insists that every piece of evidence carry a state identity the search controller has never heard of.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 19: Can Your Agent Explore in Parallel Without Creating Chaos? Use Speculative Execution and Early Cancellation</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/19-chapter/</link><pubDate>Sun, 09 Aug 2026 11:10:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/19-chapter/</guid><description>&lt;h1 id="advanced-agents-from-first-principles-19-can-your-agent-explore-in-parallel-without-creating-chaos-use-speculative-execution-and-early-cancellation"&gt;Advanced Agents From First Principles 19: Can Your Agent Explore in Parallel Without Creating Chaos? Use Speculative Execution and Early Cancellation&lt;/h1&gt;&#10;&lt;p&gt;A production agent often has more than one useful thing it could do next.&lt;/p&gt;&#10;&lt;p&gt;It could:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;inspect repository state,&lt;/li&gt;&#10;&lt;li&gt;run a targeted test,&lt;/li&gt;&#10;&lt;li&gt;retrieve documentation,&lt;/li&gt;&#10;&lt;li&gt;ask a second model to critique a candidate,&lt;/li&gt;&#10;&lt;li&gt;generate an alternative implementation,&lt;/li&gt;&#10;&lt;li&gt;probe an API,&lt;/li&gt;&#10;&lt;li&gt;inspect a deployment,&lt;/li&gt;&#10;&lt;li&gt;or verify an invariant.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;If those actions are independent, executing them one by one can be needlessly slow.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 06: Does One Agent Plan, Execute and Judge Its Own Work? Build a Planner-Executor-Critic Architecture</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/06-chapter/</link><pubDate>Sat, 08 Aug 2026 23:49:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/06-chapter/</guid><description>&lt;h1 id="does-one-agent-plan-execute-and-judge-its-own-work-build-a-planner-executor-critic-architecture"&gt;Does One Agent Plan, Execute and Judge Its Own Work? Build a Planner-Executor-Critic Architecture&lt;/h1&gt;&#10;&lt;p&gt;A single model can often do all of these things:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;understand a task,&lt;/li&gt;&#10;&lt;li&gt;decide what to do,&lt;/li&gt;&#10;&lt;li&gt;execute a tool call,&lt;/li&gt;&#10;&lt;li&gt;inspect the result,&lt;/li&gt;&#10;&lt;li&gt;critique its own work,&lt;/li&gt;&#10;&lt;li&gt;decide whether it succeeded,&lt;/li&gt;&#10;&lt;li&gt;and produce the final answer.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;That is convenient.&lt;/p&gt;&#10;&lt;p&gt;It is also a dangerous concentration of responsibilities.&lt;/p&gt;&#10;&lt;p&gt;If the same component creates the plan, executes it, explains why the result is good, and decides whether the job is complete, then failures become difficult to localize.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 05: Is One Model Doing Everything? Build a Mixture of Experts at the Agent Level</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/05-chapter/</link><pubDate>Sat, 08 Aug 2026 23:41:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/05-chapter/</guid><description>&lt;p&gt;A common agent architecture starts simply:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;request&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;action&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That simplicity is valuable.&lt;/p&gt;&#10;&lt;p&gt;It should be your default.&lt;/p&gt;&#10;&lt;p&gt;But eventually you may notice something strange.&lt;/p&gt;&#10;&lt;p&gt;The same model is being asked to do everything:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;classify the task,&lt;/li&gt;&#10;&lt;li&gt;search documentation,&lt;/li&gt;&#10;&lt;li&gt;reason about code,&lt;/li&gt;&#10;&lt;li&gt;write SQL,&lt;/li&gt;&#10;&lt;li&gt;review a patch,&lt;/li&gt;&#10;&lt;li&gt;summarize logs,&lt;/li&gt;&#10;&lt;li&gt;judge another model,&lt;/li&gt;&#10;&lt;li&gt;decide whether a deployment is safe,&lt;/li&gt;&#10;&lt;li&gt;and answer simple questions that did not require an expensive model in the first place.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;At that point the problem may no longer be:&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 04: Does Your Agent Prune Good Ideas Too Early? Use Monte Carlo Tree Search for Long-Horizon Reasoning</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/04-chapter/</link><pubDate>Sat, 08 Aug 2026 23:37:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/04-chapter/</guid><description>&lt;p&gt;A common failure in search-based agents is easy to miss.&lt;/p&gt;&#10;&lt;p&gt;The agent generates several plausible branches.&lt;/p&gt;&#10;&lt;p&gt;It scores them.&lt;/p&gt;&#10;&lt;p&gt;One branch looks weak.&lt;/p&gt;&#10;&lt;p&gt;So the runtime prunes it.&lt;/p&gt;&#10;&lt;p&gt;Later, you discover that the discarded branch was the only one that could have reached the correct solution.&lt;/p&gt;&#10;&lt;p&gt;The problem was not generation.&lt;/p&gt;&#10;&lt;p&gt;The problem was not necessarily the model.&lt;/p&gt;&#10;&lt;p&gt;The problem was &lt;strong&gt;search allocation&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;The agent spent too much compute exploiting what looked good early and too little compute exploring alternatives whose value only became visible later.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 03: Does Your Agent Commit to a Bad Reasoning Path Too Early? Build a Tree of Thoughts</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/03-chapter/</link><pubDate>Sat, 08 Aug 2026 23:25:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/03-chapter/</guid><description>&lt;p&gt;A reasoning agent can fail even when every individual step looks plausible.&lt;/p&gt;&#10;&lt;p&gt;The problem is often not that the model cannot produce a good line of reasoning.&lt;/p&gt;&#10;&lt;p&gt;The problem is that it commits too early.&lt;/p&gt;&#10;&lt;p&gt;It chooses one interpretation, one hypothesis, one plan, or one next step and then spends the rest of the run trying to make that decision work.&lt;/p&gt;&#10;&lt;p&gt;That gives us a common failure pattern:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;problem&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;first plausible thought&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;second thought conditioned on the first&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;third thought conditioned on both&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;...&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;confident answer built on an early mistake&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;If the first branch was wrong, every later step inherits the error.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 02: Why Does My Reasoning Agent Give a Different Answer Every Time? Use Self-Consistency Without Confusing Consensus With Truth</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/02-chapter/</link><pubDate>Sat, 08 Aug 2026 22:44:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/02-chapter/</guid><description>&lt;p&gt;A reasoning agent gives you one answer.&lt;/p&gt;&#10;&lt;p&gt;You run it again.&lt;/p&gt;&#10;&lt;p&gt;It gives you another.&lt;/p&gt;&#10;&lt;p&gt;You change nothing important:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;same task,&lt;/li&gt;&#10;&lt;li&gt;same tools,&lt;/li&gt;&#10;&lt;li&gt;same model family,&lt;/li&gt;&#10;&lt;li&gt;same broad context.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Yet the result changes.&lt;/p&gt;&#10;&lt;p&gt;That is not necessarily a bug.&lt;/p&gt;&#10;&lt;p&gt;A probabilistic model is allowed to produce more than one plausible trajectory.&lt;/p&gt;&#10;&lt;p&gt;The engineering question is different:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;How should an agent system use that variation?&lt;/strong&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;One common answer is &lt;strong&gt;self-consistency&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;Generate several independent reasoning trajectories.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 01: Does Your AI Agent Fail on Complex Reasoning Tasks? Treat Chain of Thought as Computation, Not Proof</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/01-chapter/</link><pubDate>Sat, 08 Aug 2026 22:35:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/01-chapter/</guid><description>&lt;p&gt;Most developers first encounter chain of thought as a prompting trick:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Think step by step.&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That framing is too shallow for agent engineering.&lt;/p&gt;&#10;&lt;p&gt;For an advanced agent, the useful idea is not that the model should produce a long explanation. The useful idea is that a difficult task may benefit from &lt;strong&gt;intermediate computational state&lt;/strong&gt; before the system commits to an action or answer.&lt;/p&gt;&#10;&lt;p&gt;That is a very different claim.&lt;/p&gt;&#10;&lt;p&gt;A reasoning trace can help a system decompose a problem, preserve intermediate conclusions, identify missing information, decide what to verify next, and expose places where search or tools should be used.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 00: When Should You Use an Advanced Agent Architecture?</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/00-chapter/</link><pubDate>Sat, 08 Aug 2026 22:27:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/00-chapter/</guid><description>&lt;h1 id="advanced-agents-from-first-principles-00-when-should-you-use-an-advanced-agent-architecture"&gt;Advanced Agents From First Principles 00: When Should You Use an Advanced Agent Architecture?&lt;/h1&gt;&#10;&lt;p&gt;You built an agent.&lt;/p&gt;&#10;&lt;p&gt;It can:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;call tools,&lt;/li&gt;&#10;&lt;li&gt;maintain state,&lt;/li&gt;&#10;&lt;li&gt;plan,&lt;/li&gt;&#10;&lt;li&gt;revise its own work,&lt;/li&gt;&#10;&lt;li&gt;search over alternatives,&lt;/li&gt;&#10;&lt;li&gt;remember useful information,&lt;/li&gt;&#10;&lt;li&gt;and verify whether the requested outcome actually happened.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Now the temptation begins.&lt;/p&gt;&#10;&lt;p&gt;You add another model.&lt;/p&gt;&#10;&lt;p&gt;Then a critic.&lt;/p&gt;&#10;&lt;p&gt;Then a planner.&lt;/p&gt;&#10;&lt;p&gt;Then a judge.&lt;/p&gt;&#10;&lt;p&gt;Then a router.&lt;/p&gt;&#10;&lt;p&gt;Then three specialist agents.&lt;/p&gt;&#10;&lt;p&gt;Then a tree search.&lt;/p&gt;</description></item><item><title>Agents From First Principles 09: AI Agent Says It Worked When It Didn’t? Verify the Result Outside the LLM</title><link>https://aibussin.com/post/agents-from-first-principles-09/</link><pubDate>Sat, 08 Aug 2026 17:31:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-09/</guid><description>&lt;p&gt;An AI agent says:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Done. The task is complete.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;That sentence is almost worthless.&lt;/p&gt;&#10;&lt;p&gt;The agent may have:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;edited the wrong file,&lt;/li&gt;&#10;&lt;li&gt;changed the right file incorrectly,&lt;/li&gt;&#10;&lt;li&gt;skipped part of the request,&lt;/li&gt;&#10;&lt;li&gt;broken another subsystem,&lt;/li&gt;&#10;&lt;li&gt;failed to save its work,&lt;/li&gt;&#10;&lt;li&gt;misread a tool result,&lt;/li&gt;&#10;&lt;li&gt;passed a stale test,&lt;/li&gt;&#10;&lt;li&gt;inspected the wrong environment,&lt;/li&gt;&#10;&lt;li&gt;or simply decided that its own answer looked convincing.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The central problem is simple:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;The system that produced the answer should not be the only system deciding whether the answer is correct.&lt;/strong&gt;&lt;/p&gt;</description></item></channel></rss>