<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Ai Agents on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/ai-agents/</link><description>Recent content in Ai Agents on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Wed, 26 Aug 2026 10:00:00 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/ai-agents/index.xml" rel="self" type="application/rss+xml"/><item><title>What Is an Agent, Really?</title><link>https://aibussin.com/books/agents-from-first-principles/01-chapter/</link><pubDate>Sat, 08 Aug 2026 15:40:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/01-chapter/</guid><description>&lt;p&gt;This book is about building systems &lt;strong&gt;around&lt;/strong&gt; models. Before we build planners, memory, tool routing, critics, search and verifiers, we need an answer to a question that turns out to be harder than it looks:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;What is an agent?&lt;/strong&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;The word is currently applied to almost everything: a single LLM call, a chatbot, a fixed pipeline, a tool-using loop, and any five model calls with class names ending in &lt;code&gt;Agent&lt;/code&gt;. Those systems may all be useful. But when one word covers all of them, it stops telling us anything about the computation, and we lose the ability to say which mechanism is doing the work.&lt;/p&gt;</description></item><item><title>The Action Boundary</title><link>https://aibussin.com/books/agents-from-first-principles/02-chapter/</link><pubDate>Sat, 08 Aug 2026 15:46:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/02-chapter/</guid><description>&lt;p&gt;The previous chapter ended with a division of responsibility:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;The model proposes. The runtime decides what may execute. The environment supplies observations about what happened.&lt;/strong&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;This chapter takes the first clause seriously. What does it actually mean for a model to &lt;strong&gt;propose an action&lt;/strong&gt;?&lt;/p&gt;&#10;&lt;p&gt;Suppose the model emits the sentence &lt;em&gt;&amp;ldquo;Search the documentation for the latest PyTorch optimizer API.&amp;rdquo;&lt;/em&gt; A human reads that and understands the intention immediately. A runtime cannot act on it at all, because it needs answers to six questions the sentence does not contain: which action, which arguments, whether those arguments are well formed, whether their values are meaningful, whether this action is permitted in this run, and whether its execution preconditions currently hold.&lt;/p&gt;</description></item><item><title>Candidate Generation and Selection</title><link>https://aibussin.com/books/agents-from-first-principles/03-chapter/</link><pubDate>Sat, 08 Aug 2026 15:54:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/03-chapter/</guid><description>&lt;p&gt;The previous chapter built a boundary that stops arbitrary model output from acquiring execution authority without explicit checks. It made one class of failure inspectable and enforceable, and it is silent about another.&lt;/p&gt;&#10;&lt;p&gt;Suppose the model is asked to solve a coding problem. One run produces the right patch. The next produces a plausible but incomplete one. A third produces something better again. Nothing is malformed, nothing violates the action schema, and every one of them would pass the boundary we just built. Validity and quality are different properties. A proposal can be completely valid and still be a poor choice, which relocates the uncertainty rather than removing it:&lt;/p&gt;</description></item><item><title>Critique, Revision, and Acceptance</title><link>https://aibussin.com/books/agents-from-first-principles/04-chapter/</link><pubDate>Sat, 08 Aug 2026 16:23:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/04-chapter/</guid><description>&lt;p&gt;Selection can only choose among the candidates it is given. A strong evaluator may recognize that every available candidate is poor, but it cannot select a correct answer that the generator never produced.&lt;/p&gt;&#10;&lt;p&gt;That limitation becomes practical because the candidates usually come from one model answering one prompt, and such samples correlate. When four candidates share the same misreading of the evidence, ranking them can still produce a confident winner while leaving the shared defect untouched.&lt;/p&gt;</description></item><item><title>Planning and Execution</title><link>https://aibussin.com/books/agents-from-first-principles/05-chapter/</link><pubDate>Sat, 08 Aug 2026 16:35:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/05-chapter/</guid><description>&lt;p&gt;The critique loop works on one thing at a time. It assumes the task already exists as a candidate we can hold, inspect and improve.&lt;/p&gt;&#10;&lt;p&gt;Some tasks have no draft to hold.&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Inspect a project, reproduce the failing test, find the cause, patch the code, rerun the relevant tests, and report what changed.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;No single narrow action responsibly completes that goal, and the actions constrain each other. Patching before diagnosis is guesswork wearing the costume of work, and reporting success before observing a passing test is a claim about the world that nothing in the run supports. Worse, if execution reveals that an assumption was wrong, the rest of the route may be wrong too. A runtime choosing each action independently can react to the new observation, but without an explicit representation of the intended route it has nothing concrete to compare the changed world against.&lt;/p&gt;</description></item><item><title>Runtime State, Progress, and Termination</title><link>https://aibussin.com/books/agents-from-first-principles/06-chapter/</link><pubDate>Sat, 08 Aug 2026 16:56:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/06-chapter/</guid><description>&lt;p&gt;A plan is a claim about the future. It says that reproducing the failure, then diagnosing it, then patching, then testing, is a route from here to a working system. Making that claim explicit was worth the machinery. But it is written before any of the work happens, and execution may invalidate one of its assumptions almost immediately.&lt;/p&gt;&#10;&lt;p&gt;Execution produces something else entirely: a trajectory. Actions went out, observations came back, and some of what came back may contradict the intended route. The patch step failed because a dependency is missing. The test step is not merely late; it is unreachable.&lt;/p&gt;</description></item><item><title>Capabilities and Routing</title><link>https://aibussin.com/books/agents-from-first-principles/07-chapter/</link><pubDate>Sat, 08 Aug 2026 17:09:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/07-chapter/</guid><description>&lt;p&gt;Every mechanism built so far has taken the action set as given. The action boundary validates a proposal against a fixed list of permitted action types. Candidate generation samples several proposals from the same list. The runtime-state chapter watches what happens after execution and decides whether to continue. All of them assume that somebody, somewhere, already decided which capabilities the policy could choose from.&lt;/p&gt;&#10;&lt;p&gt;So far we have treated that decision as an input. This chapter makes it explicit, because the capability surface changes the decision problem the policy has to solve and therefore belongs inside the agent architecture itself.&lt;/p&gt;</description></item><item><title>Memory and Selective Recall</title><link>https://aibussin.com/books/agents-from-first-principles/08-chapter/</link><pubDate>Sat, 08 Aug 2026 17:14:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/08-chapter/</guid><description>&lt;p&gt;The capability boundary gave the agent a defined action space and a rule for which capabilities are eligible at each step. Runtime state gave it an explicit working representation of what this run has established so far and how it reached that point. Together they govern the current execution, but neither gives information from an earlier run a controlled way to influence this one. There is a family of failures that requires exactly that.&lt;/p&gt;</description></item><item><title>Trajectory Search</title><link>https://aibussin.com/books/agents-from-first-principles/09-chapter/</link><pubDate>Sat, 08 Aug 2026 17:26:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/09-chapter/</guid><description>&lt;p&gt;An agent can make every local decision look reasonable and still lose the task.&lt;/p&gt;&#10;&lt;p&gt;A coding agent sees that &lt;code&gt;test_checkout_redirect&lt;/code&gt; is failing, concludes there is an implementation bug in &lt;code&gt;checkout.py&lt;/code&gt;, and then behaves impeccably for twenty steps: it reads the file, edits it, runs the tests, repairs the new failures its edit introduced, rewrites the patch, and runs the tests again. Every one of those steps is defensible given the step before it.&lt;/p&gt;</description></item><item><title>Evidence and Verification</title><link>https://aibussin.com/books/agents-from-first-principles/10-chapter/</link><pubDate>Sat, 08 Aug 2026 17:31:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/10-chapter/</guid><description>&lt;p&gt;The agent says:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Done.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;That is a claim.&lt;/p&gt;&#10;&lt;p&gt;It is not evidence.&lt;/p&gt;&#10;&lt;p&gt;Every mechanism in this book so far has made the agent better at deciding what to do, and none of them establishes that the user&amp;rsquo;s goal was achieved. A planner can produce a coherent plan for the wrong problem. A tool can return exit code zero without producing the intended effect. A search can select the highest-scoring branch when every branch is wrong. A memory system can retrieve a perfectly relevant fact that stopped being true in March.&lt;/p&gt;</description></item><item><title>Building the Complete Agent</title><link>https://aibussin.com/books/agents-from-first-principles/11-chapter/</link><pubDate>Wed, 26 Aug 2026 10:00:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/11-chapter/</guid><description>&lt;p&gt;Every mechanism in this book was argued against a problem chosen to isolate it.&lt;/p&gt;&#10;&lt;p&gt;That isolation was deliberate, and it was also a form of protection. The memory chapter picked a task where recall was the bottleneck, held everything else still, and measured the one thing it came to measure. The result is a clean explanation and a weak claim. Nothing in it establishes that the same retrieval policy behaves when a search controller is expanding forty nodes, or when a verifier insists that every piece of evidence carry a state identity the search controller has never heard of.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 19: Can Your Agent Explore in Parallel Without Creating Chaos? Use Speculative Execution and Early Cancellation</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/19-chapter/</link><pubDate>Sun, 09 Aug 2026 11:10:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/19-chapter/</guid><description>&lt;h1 id="advanced-agents-from-first-principles-19-can-your-agent-explore-in-parallel-without-creating-chaos-use-speculative-execution-and-early-cancellation"&gt;Advanced Agents From First Principles 19: Can Your Agent Explore in Parallel Without Creating Chaos? Use Speculative Execution and Early Cancellation&lt;/h1&gt;&#10;&lt;p&gt;A production agent often has more than one useful thing it could do next.&lt;/p&gt;&#10;&lt;p&gt;It could:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;inspect repository state,&lt;/li&gt;&#10;&lt;li&gt;run a targeted test,&lt;/li&gt;&#10;&lt;li&gt;retrieve documentation,&lt;/li&gt;&#10;&lt;li&gt;ask a second model to critique a candidate,&lt;/li&gt;&#10;&lt;li&gt;generate an alternative implementation,&lt;/li&gt;&#10;&lt;li&gt;probe an API,&lt;/li&gt;&#10;&lt;li&gt;inspect a deployment,&lt;/li&gt;&#10;&lt;li&gt;or verify an invariant.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;If those actions are independent, executing them one by one can be needlessly slow.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 26: Why Did the Agent Fail? Build an Incident Forensics Pipeline</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/26-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/26-chapter/</guid><description>A practical incident-forensics workflow for advanced agents: reconstruct the run, find the earliest divergence, distinguish root cause from downstream symptoms, measure blast radius, and prove that a remediation would have prevented the incident.</description></item><item><title>Advanced Agents From First Principles 27: How Reliable Does an Agent Need to Be? Define SLOs and Error Budgets</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/27-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/27-chapter/</guid><description>A practical reliability framework for advanced agents: define verified-success SLOs, false-success ceilings, UNKNOWN budgets, latency and cost targets, then use error-budget burn to decide when to ship capability and when to stop and harden the system.</description></item><item><title>Advanced Agents From First Principles 28: Where Should You Spend the Next Engineering Hour? Prioritize Reliability by Risk and Expected Return</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/28-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/28-chapter/</guid><description>A practical framework for deciding where to spend the next engineering hour in an advanced-agent system: rank remediation by expected reduction in verified reliability loss, severity, recurrence, blast radius, confidence and implementation cost.</description></item><item><title>Advanced Agents From First Principles 29: When Should an Agent Stop and Ask a Human? Design Authority Boundaries and Escalation</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/29-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/29-chapter/</guid><description>A practical architecture for agent authority boundaries: decide what an agent may do autonomously, when it must escalate, what evidence a human reviewer needs, and how to avoid turning human approval into rubber-stamping.</description></item><item><title>Advanced Agents From First Principles 30: Is This Task Outside Your Agent’s Competence? Build Competence Envelopes and OOD Detection</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/30-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/30-chapter/</guid><description>A practical framework for competence envelopes in production agents: distinguish uncertainty from lack of validated competence, detect out-of-distribution tasks, contract authority when evidence is weak, and expand autonomy only through measured evidence.</description></item><item><title>Advanced Agents From First Principles 31: How Can an Agent Learn New Capabilities Without Expanding Its Own Authority? Use Sandboxed Capability Acquisition</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/31-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/31-chapter/</guid><description>A practical architecture for sandboxed capability acquisition: let agents explore tasks outside their validated competence envelope, accumulate externally verified evidence, and propose capability expansion without ever granting themselves production authority.</description></item><item><title>Advanced Agents From First Principles 32: Which Capabilities Are Actually Worth Building? Design a Capability Portfolio</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/32-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/32-chapter/</guid><description>A practical framework for deciding which agent capabilities are worth acquiring: rank missing capabilities by user value, verifier availability, reliability risk, acquisition cost, maintenance burden, and the quality of human or deterministic alternatives.</description></item><item><title>Advanced Agents From First Principles 33: Which Shared Components Actually Unlock More Capability? Build a Capability Dependency Graph</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/33-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/33-chapter/</guid><description>A practical capability-dependency architecture for advanced agents: identify shared primitives that unlock many capabilities, quantify leverage, expose correlated-failure hotspots, and invest in platform components without creating hidden systemic risk.</description></item><item><title>Advanced Agents From First Principles 34: Where Should This Task Actually Run? Build Capability-Aware Placement Across Models, Providers and Resource Pools</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/34-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/34-chapter/</guid><description>A practical placement architecture for advanced agents: route work across local and frontier models, providers, regions, GPUs, browser pools and specialist runtimes using demonstrated competence, verifier availability, policy constraints, health, cost and latency rather than model prestige.</description></item><item><title>Advanced Agents From First Principles 35: How Do You Move a Running Agent Between Workers Without Losing Meaning? Build Portable Execution State and Safe Handoff</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/35-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/35-chapter/</guid><description>A practical architecture for portable agent execution state: checkpoint long-running runs, transfer ownership safely across workers and providers, preserve evidence and authority, and reject migrations that cannot be proven compatible.</description></item><item><title>Advanced Agents From First Principles 36: Is Your Agent Acting on Stale State? Build Temporal Consistency, Freshness Budgets and Conflict Detection</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/36-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/36-chapter/</guid><description>A practical architecture for keeping long-running agents from acting on stale assumptions: classify state by freshness, track version vectors, detect conflicts, revalidate before consequential actions, and force replanning when the world has changed underneath the run.</description></item><item><title>Advanced Agents From First Principles 37: Is Your Agent Still Solving the Right Task? Build Intent Versioning, Supersession and Cancellation</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/37-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/37-chapter/</guid><description>A production architecture for intent versioning, supersession and cancellation in long-running agents: stop obsolete work, preserve committed effects, reconcile partial actions, and prevent stale goals from retaining authority.</description></item><item><title>Advanced Agents From First Principles 38: A Plan Is Not a Commitment — Model Goals, Commitments and Executable Work</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/38-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/38-chapter/</guid><description>A practical architecture for separating goals, plans, commitments, tasks and actions in long-running agents so replanning, cancellation, handoff and external obligations remain correct.</description></item><item><title>Advanced Agents From First Principles 39: How Do You Make an Agent Survive for Days? Build Durable Long-Running Workflows</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/39-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/39-chapter/</guid><description>A practical architecture for long-running agents: keep workflow state durable while treating models and workers as disposable, with explicit waits, retries, timers, human approvals, checkpoints, commitments, cancellation and replay.</description></item><item><title>Advanced Agents From First Principles 40: Your Agent Changed the World. What Happens When Step Two Fails? Build Transactions, Compensation and Reconciliation</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/40-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/40-chapter/</guid><description>A practical architecture for agent workflows that span systems without a global transaction: classify side effects, prepare carefully, commit with identity, verify externally, compensate when possible, reconcile ambiguity, and never pretend rollback is free.</description></item><item><title>Advanced Agents From First Principles 41: What Should Your Agent Trust? Build Explicit Security and Trust Boundaries</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/41-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/41-chapter/</guid><description>A practical security architecture for production agents: separate data from authority, classify trust, scope credentials and capabilities, preserve provenance, isolate generated code, resist prompt injection, and keep security-critical decisions outside model control.</description></item><item><title>Advanced Agents From First Principles 42: How Do Multiple Agents Coordinate Without Becoming a Distributed Argument?</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/42-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/42-chapter/</guid><description>A practical architecture for multi-agent coordination: explicit ownership, delegation, contracts, commitment transfer, shared intent, evidence provenance, conflict handling, deadlock prevention, and independent verification instead of agents merely chatting until they agree.</description></item><item><title>Advanced Agents From First Principles 43: Who Controls the Agent? Build an Explicit Agent Control Plane</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/43-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/43-chapter/</guid><description>A production-agent architecture that separates control-plane policy from execution-plane reasoning: intent, competence, authority, placement, budgets, reliability, releases, security and escalation remain enforceable outside the model.</description></item><item><title>You Probably Don't Need All of This: Build the Minimum Production Agent Architecture</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/45-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/45-chapter/</guid><description>&lt;h1 id="you-probably-dont-need-all-of-this"&gt;You Probably Don&amp;rsquo;t Need All of This&lt;/h1&gt;&#10;&lt;p&gt;Over the previous forty-five steps, we built almost every major mechanism you might need in a serious agent platform.&lt;/p&gt;&#10;&lt;p&gt;Search.&lt;/p&gt;&#10;&lt;p&gt;Critique.&lt;/p&gt;&#10;&lt;p&gt;Planning.&lt;/p&gt;&#10;&lt;p&gt;Memory.&lt;/p&gt;&#10;&lt;p&gt;Verification.&lt;/p&gt;&#10;&lt;p&gt;Distributed execution.&lt;/p&gt;&#10;&lt;p&gt;Leases.&lt;/p&gt;&#10;&lt;p&gt;Fencing.&lt;/p&gt;&#10;&lt;p&gt;Backpressure.&lt;/p&gt;&#10;&lt;p&gt;Behavioral releases.&lt;/p&gt;&#10;&lt;p&gt;Replay.&lt;/p&gt;&#10;&lt;p&gt;Incident forensics.&lt;/p&gt;&#10;&lt;p&gt;SLOs.&lt;/p&gt;&#10;&lt;p&gt;Competence envelopes.&lt;/p&gt;&#10;&lt;p&gt;Authority boundaries.&lt;/p&gt;&#10;&lt;p&gt;Capability portfolios.&lt;/p&gt;&#10;&lt;p&gt;Placement.&lt;/p&gt;&#10;&lt;p&gt;Portable execution state.&lt;/p&gt;&#10;&lt;p&gt;Temporal consistency.&lt;/p&gt;&#10;&lt;p&gt;Intent versioning.&lt;/p&gt;&#10;&lt;p&gt;Commitments.&lt;/p&gt;&#10;&lt;p&gt;Durable workflows.&lt;/p&gt;&#10;&lt;p&gt;Transaction recovery.&lt;/p&gt;&#10;&lt;p&gt;Security boundaries.&lt;/p&gt;&#10;&lt;p&gt;Multi-agent coordination.&lt;/p&gt;&#10;&lt;p&gt;An explicit control plane.&lt;/p&gt;&#10;&lt;p&gt;And finally, in Step 44, we assembled those ideas into a complete reference architecture for a production AI agent.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 06: Does One Agent Plan, Execute and Judge Its Own Work? Build a Planner-Executor-Critic Architecture</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/06-chapter/</link><pubDate>Sat, 08 Aug 2026 23:49:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/06-chapter/</guid><description>&lt;h1 id="does-one-agent-plan-execute-and-judge-its-own-work-build-a-planner-executor-critic-architecture"&gt;Does One Agent Plan, Execute and Judge Its Own Work? Build a Planner-Executor-Critic Architecture&lt;/h1&gt;&#10;&lt;p&gt;A single model can often do all of these things:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;understand a task,&lt;/li&gt;&#10;&lt;li&gt;decide what to do,&lt;/li&gt;&#10;&lt;li&gt;execute a tool call,&lt;/li&gt;&#10;&lt;li&gt;inspect the result,&lt;/li&gt;&#10;&lt;li&gt;critique its own work,&lt;/li&gt;&#10;&lt;li&gt;decide whether it succeeded,&lt;/li&gt;&#10;&lt;li&gt;and produce the final answer.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;That is convenient.&lt;/p&gt;&#10;&lt;p&gt;It is also a dangerous concentration of responsibilities.&lt;/p&gt;&#10;&lt;p&gt;If the same component creates the plan, executes it, explains why the result is good, and decides whether the job is complete, then failures become difficult to localize.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 05: Is One Model Doing Everything? Build a Mixture of Experts at the Agent Level</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/05-chapter/</link><pubDate>Sat, 08 Aug 2026 23:41:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/05-chapter/</guid><description>&lt;p&gt;A common agent architecture starts simply:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;request&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;action&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That simplicity is valuable.&lt;/p&gt;&#10;&lt;p&gt;It should be your default.&lt;/p&gt;&#10;&lt;p&gt;But eventually you may notice something strange.&lt;/p&gt;&#10;&lt;p&gt;The same model is being asked to do everything:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;classify the task,&lt;/li&gt;&#10;&lt;li&gt;search documentation,&lt;/li&gt;&#10;&lt;li&gt;reason about code,&lt;/li&gt;&#10;&lt;li&gt;write SQL,&lt;/li&gt;&#10;&lt;li&gt;review a patch,&lt;/li&gt;&#10;&lt;li&gt;summarize logs,&lt;/li&gt;&#10;&lt;li&gt;judge another model,&lt;/li&gt;&#10;&lt;li&gt;decide whether a deployment is safe,&lt;/li&gt;&#10;&lt;li&gt;and answer simple questions that did not require an expensive model in the first place.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;At that point the problem may no longer be:&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 04: Does Your Agent Prune Good Ideas Too Early? Use Monte Carlo Tree Search for Long-Horizon Reasoning</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/04-chapter/</link><pubDate>Sat, 08 Aug 2026 23:37:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/04-chapter/</guid><description>&lt;p&gt;A common failure in search-based agents is easy to miss.&lt;/p&gt;&#10;&lt;p&gt;The agent generates several plausible branches.&lt;/p&gt;&#10;&lt;p&gt;It scores them.&lt;/p&gt;&#10;&lt;p&gt;One branch looks weak.&lt;/p&gt;&#10;&lt;p&gt;So the runtime prunes it.&lt;/p&gt;&#10;&lt;p&gt;Later, you discover that the discarded branch was the only one that could have reached the correct solution.&lt;/p&gt;&#10;&lt;p&gt;The problem was not generation.&lt;/p&gt;&#10;&lt;p&gt;The problem was not necessarily the model.&lt;/p&gt;&#10;&lt;p&gt;The problem was &lt;strong&gt;search allocation&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;The agent spent too much compute exploiting what looked good early and too little compute exploring alternatives whose value only became visible later.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 03: Does Your Agent Commit to a Bad Reasoning Path Too Early? Build a Tree of Thoughts</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/03-chapter/</link><pubDate>Sat, 08 Aug 2026 23:25:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/03-chapter/</guid><description>&lt;p&gt;A reasoning agent can fail even when every individual step looks plausible.&lt;/p&gt;&#10;&lt;p&gt;The problem is often not that the model cannot produce a good line of reasoning.&lt;/p&gt;&#10;&lt;p&gt;The problem is that it commits too early.&lt;/p&gt;&#10;&lt;p&gt;It chooses one interpretation, one hypothesis, one plan, or one next step and then spends the rest of the run trying to make that decision work.&lt;/p&gt;&#10;&lt;p&gt;That gives us a common failure pattern:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;problem&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;first plausible thought&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;second thought conditioned on the first&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;third thought conditioned on both&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;...&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;confident answer built on an early mistake&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;If the first branch was wrong, every later step inherits the error.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 02: Why Does My Reasoning Agent Give a Different Answer Every Time? Use Self-Consistency Without Confusing Consensus With Truth</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/02-chapter/</link><pubDate>Sat, 08 Aug 2026 22:44:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/02-chapter/</guid><description>&lt;p&gt;A reasoning agent gives you one answer.&lt;/p&gt;&#10;&lt;p&gt;You run it again.&lt;/p&gt;&#10;&lt;p&gt;It gives you another.&lt;/p&gt;&#10;&lt;p&gt;You change nothing important:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;same task,&lt;/li&gt;&#10;&lt;li&gt;same tools,&lt;/li&gt;&#10;&lt;li&gt;same model family,&lt;/li&gt;&#10;&lt;li&gt;same broad context.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Yet the result changes.&lt;/p&gt;&#10;&lt;p&gt;That is not necessarily a bug.&lt;/p&gt;&#10;&lt;p&gt;A probabilistic model is allowed to produce more than one plausible trajectory.&lt;/p&gt;&#10;&lt;p&gt;The engineering question is different:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;How should an agent system use that variation?&lt;/strong&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;One common answer is &lt;strong&gt;self-consistency&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;Generate several independent reasoning trajectories.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 01: Does Your AI Agent Fail on Complex Reasoning Tasks? Treat Chain of Thought as Computation, Not Proof</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/01-chapter/</link><pubDate>Sat, 08 Aug 2026 22:35:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/01-chapter/</guid><description>&lt;p&gt;Most developers first encounter chain of thought as a prompting trick:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Think step by step.&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That framing is too shallow for agent engineering.&lt;/p&gt;&#10;&lt;p&gt;For an advanced agent, the useful idea is not that the model should produce a long explanation. The useful idea is that a difficult task may benefit from &lt;strong&gt;intermediate computational state&lt;/strong&gt; before the system commits to an action or answer.&lt;/p&gt;&#10;&lt;p&gt;That is a very different claim.&lt;/p&gt;&#10;&lt;p&gt;A reasoning trace can help a system decompose a problem, preserve intermediate conclusions, identify missing information, decide what to verify next, and expose places where search or tools should be used.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 00: When Should You Use an Advanced Agent Architecture?</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/00-chapter/</link><pubDate>Sat, 08 Aug 2026 22:27:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/00-chapter/</guid><description>&lt;h1 id="advanced-agents-from-first-principles-00-when-should-you-use-an-advanced-agent-architecture"&gt;Advanced Agents From First Principles 00: When Should You Use an Advanced Agent Architecture?&lt;/h1&gt;&#10;&lt;p&gt;You built an agent.&lt;/p&gt;&#10;&lt;p&gt;It can:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;call tools,&lt;/li&gt;&#10;&lt;li&gt;maintain state,&lt;/li&gt;&#10;&lt;li&gt;plan,&lt;/li&gt;&#10;&lt;li&gt;revise its own work,&lt;/li&gt;&#10;&lt;li&gt;search over alternatives,&lt;/li&gt;&#10;&lt;li&gt;remember useful information,&lt;/li&gt;&#10;&lt;li&gt;and verify whether the requested outcome actually happened.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Now the temptation begins.&lt;/p&gt;&#10;&lt;p&gt;You add another model.&lt;/p&gt;&#10;&lt;p&gt;Then a critic.&lt;/p&gt;&#10;&lt;p&gt;Then a planner.&lt;/p&gt;&#10;&lt;p&gt;Then a judge.&lt;/p&gt;&#10;&lt;p&gt;Then a router.&lt;/p&gt;&#10;&lt;p&gt;Then three specialist agents.&lt;/p&gt;&#10;&lt;p&gt;Then a tree search.&lt;/p&gt;</description></item><item><title>Agents From First Principles 09: AI Agent Says It Worked When It Didn’t? Verify the Result Outside the LLM</title><link>https://aibussin.com/post/agents-from-first-principles-09/</link><pubDate>Sat, 08 Aug 2026 17:31:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-09/</guid><description>&lt;p&gt;An AI agent says:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Done. The task is complete.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;That sentence is almost worthless.&lt;/p&gt;&#10;&lt;p&gt;The agent may have:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;edited the wrong file,&lt;/li&gt;&#10;&lt;li&gt;changed the right file incorrectly,&lt;/li&gt;&#10;&lt;li&gt;skipped part of the request,&lt;/li&gt;&#10;&lt;li&gt;broken another subsystem,&lt;/li&gt;&#10;&lt;li&gt;failed to save its work,&lt;/li&gt;&#10;&lt;li&gt;misread a tool result,&lt;/li&gt;&#10;&lt;li&gt;passed a stale test,&lt;/li&gt;&#10;&lt;li&gt;inspected the wrong environment,&lt;/li&gt;&#10;&lt;li&gt;or simply decided that its own answer looked convincing.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The central problem is simple:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;The system that produced the answer should not be the only system deciding whether the answer is correct.&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Agents From First Principles 08: AI Agent Picks the First Solution? Add Search Instead of One-Shot Generation</title><link>https://aibussin.com/post/agents-from-first-principles-08/</link><pubDate>Sat, 08 Aug 2026 17:26:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-08/</guid><description>&lt;p&gt;An AI agent often fails for a surprisingly ordinary reason:&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;it commits too early.&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;It finds one plausible next action, follows it, and then spends the rest of the run trying to make that first choice work.&lt;/p&gt;&#10;&lt;p&gt;That can look intelligent because the agent keeps reasoning, calling tools, revising plans, and explaining itself.&lt;/p&gt;&#10;&lt;p&gt;But underneath, the trajectory may be almost completely determined by an early mistake.&lt;/p&gt;&#10;&lt;p&gt;A coding agent chooses the wrong implementation strategy and spends twenty tool calls repairing it.&lt;/p&gt;</description></item><item><title>Agents From First Principles 07: AI Agent Forgets Previous Work? Add Working, Semantic and Episodic Memory</title><link>https://aibussin.com/post/agents-from-first-principles-07/</link><pubDate>Sat, 08 Aug 2026 17:14:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-07/</guid><description>&lt;h1 id="ai-agent-forgets-previous-work-add-working-semantic-and-episodic-memory"&gt;AI Agent Forgets Previous Work? Add Working, Semantic and Episodic Memory&lt;/h1&gt;&#10;&lt;p&gt;An agent can use the right model, call the right tools, execute the right plan, and still behave as if nothing that happened five minutes ago matters.&lt;/p&gt;&#10;&lt;p&gt;You see the symptoms quickly:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;it re-reads files it already inspected;&lt;/li&gt;&#10;&lt;li&gt;it repeats research it already completed;&lt;/li&gt;&#10;&lt;li&gt;it asks for information the user already supplied;&lt;/li&gt;&#10;&lt;li&gt;it forgets why a previous approach failed;&lt;/li&gt;&#10;&lt;li&gt;it loses decisions made earlier in a long task;&lt;/li&gt;&#10;&lt;li&gt;it treats every new run as if the system has never seen the problem before;&lt;/li&gt;&#10;&lt;li&gt;it retrieves an old answer and treats it as current truth;&lt;/li&gt;&#10;&lt;li&gt;it fills the prompt with so much history that the useful information is buried.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The usual response is:&lt;/p&gt;</description></item><item><title>Agents From First Principles 06: AI Agent Chooses the Wrong Tool? Design Better Tool Interfaces, Schemas and Routing</title><link>https://aibussin.com/post/agents-from-first-principles-06/</link><pubDate>Sat, 08 Aug 2026 17:09:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-06/</guid><description>&lt;p&gt;An agent can have a perfectly capable model and still behave badly because its tools are badly designed.&lt;/p&gt;&#10;&lt;p&gt;This is one of the most common agent failures in production:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;user goal&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;agent&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;wrong tool&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;wrong action&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The model may understand the task.&lt;/p&gt;&#10;&lt;p&gt;The agent may have enough context.&lt;/p&gt;&#10;&lt;p&gt;The problem is that the action space is ambiguous.&lt;/p&gt;&#10;&lt;p&gt;If two tools overlap, their descriptions are vague, their schemas are huge, or their results are difficult to interpret, the model has to guess.&lt;/p&gt;</description></item><item><title>Agents From First Principles 05: AI Agent Gets Stuck in a Loop? Add State, Feedback and Stopping Conditions</title><link>https://aibussin.com/post/agents-from-first-principles-05/</link><pubDate>Sat, 08 Aug 2026 16:56:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-05/</guid><description>&lt;p&gt;An AI agent that keeps calling the same tool, revisiting the same page, rewriting the same file, or repeatedly saying “I’ll try again” is not displaying persistence.&lt;/p&gt;&#10;&lt;p&gt;It is displaying a control-flow bug.&lt;/p&gt;&#10;&lt;p&gt;This is one of the most common failure modes in agent software because the basic loop is deceptively simple:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;observe&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;decide&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;act&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;observe&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;repeat&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The problem is hidden inside the final word.&lt;/p&gt;</description></item><item><title>Agents From First Principles 04: AI Agent Fails on Multi-Step Tasks? Separate Planning From Execution</title><link>https://aibussin.com/post/agents-from-first-principles-04/</link><pubDate>Sat, 08 Aug 2026 16:35:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-04/</guid><description>&lt;p&gt;A surprising number of agent failures are not really model failures.&lt;/p&gt;&#10;&lt;p&gt;The model may be perfectly capable of writing each individual step. The failure happens because the system tries to decide &lt;strong&gt;what to do&lt;/strong&gt; and &lt;strong&gt;do it&lt;/strong&gt; at the same time.&lt;/p&gt;&#10;&lt;p&gt;That works for simple tasks:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;question&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;answer&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It becomes fragile when success depends on several ordered actions:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;goal&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;step 1&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;step 2&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;step 3&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;verification&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A useful next step in agent design is therefore to separate two jobs:&lt;/p&gt;</description></item><item><title>Agents From First Principles 03: AI Agent Keeps Making the Same Mistake? Add a Critique-and-Revision Loop</title><link>https://aibussin.com/post/agents-from-first-principles-03/</link><pubDate>Sat, 08 Aug 2026 16:23:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-03/</guid><description>&lt;p&gt;An AI agent can fail in a particularly frustrating way: it produces an answer that is almost right, you ask it to improve the answer, and it produces another answer with the same underlying defect.&lt;/p&gt;&#10;&lt;p&gt;Sometimes the wording changes. Sometimes it adds more explanation. Sometimes it becomes longer and more confident. But the important mistake survives.&lt;/p&gt;&#10;&lt;p&gt;That usually means the system is doing this:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;prompt&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;answer&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;or this:&lt;/p&gt;</description></item><item><title>Agents From First Principles 02: AI Agent Gives Inconsistent Answers? Generate Multiple Candidates and Rank Them</title><link>https://aibussin.com/post/agents-from-first-principles-02/</link><pubDate>Sat, 08 Aug 2026 15:54:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-02/</guid><description>&lt;p&gt;One of the first things you notice when you build anything around a large language model is that the same prompt does not always produce the same quality of answer.&lt;/p&gt;&#10;&lt;p&gt;Sometimes the first response is excellent.&lt;/p&gt;&#10;&lt;p&gt;Sometimes it is merely acceptable.&lt;/p&gt;&#10;&lt;p&gt;Sometimes it misses the point entirely.&lt;/p&gt;&#10;&lt;p&gt;That creates a very common agent-engineering question:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;If the model is inconsistent, should the agent trust the first answer it gets?&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;Often, no.&lt;/p&gt;</description></item><item><title>AI Agent Returning Invalid Tool Calls? How to Validate LLM Actions</title><link>https://aibussin.com/post/agents-from-first-principles-01/</link><pubDate>Sat, 08 Aug 2026 15:46:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-01/</guid><description>&lt;p&gt;If your agent sometimes invents a tool name, omits a required argument, returns malformed JSON, or produces an action that looks plausible but cannot actually be executed, the problem is usually not &amp;ldquo;the agent is dumb.&amp;rdquo;&lt;/p&gt;&#10;&lt;p&gt;The problem is that &lt;strong&gt;raw language-model output has been allowed to cross directly into execution&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;That boundary is too weak.&lt;/p&gt;&#10;&lt;p&gt;The simplest useful agent architecture is not:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;prompt&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;execute whatever came back&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It is:&lt;/p&gt;</description></item><item><title>Agents From First Principles 00: What Is an Agent, Really?</title><link>https://aibussin.com/post/agents-from-first-principles-00/</link><pubDate>Sat, 08 Aug 2026 15:40:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-00/</guid><description>&lt;h1 id="what-is-an-agent-really"&gt;What Is an Agent, Really?&lt;/h1&gt;&#10;&lt;p&gt;This is the first post in &lt;strong&gt;Agents From First Principles&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;It follows two earlier series.&lt;/p&gt;&#10;&lt;p&gt;In &lt;strong&gt;PyTorch: Zero to Hero&lt;/strong&gt;, we worked upward from tensors, autograd and neural-network building blocks until we could build a small language model ourselves.&lt;/p&gt;&#10;&lt;p&gt;In &lt;strong&gt;Models From First Principles&lt;/strong&gt;, we moved one level higher. We looked at how learned components can be composed into scorers, value models, policy heads, recurrent models, hierarchical models and compact recursive systems.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 07: Do Your Agents Agree Too Easily? Use Adversarial Review and Multi-Agent Debate Without Confusing Debate With Truth</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/07-chapter/</link><pubDate>Sat, 08 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/07-chapter/</guid><description>&lt;p&gt;A multi-agent system can look sophisticated while every agent quietly repeats the same mistake.&lt;/p&gt;&#10;&lt;p&gt;That is one of the most dangerous failure modes in advanced agent architectures.&lt;/p&gt;&#10;&lt;p&gt;You ask one model to solve the problem.&lt;/p&gt;&#10;&lt;p&gt;Then you ask a second model to review it.&lt;/p&gt;&#10;&lt;p&gt;Then a third model judges the disagreement.&lt;/p&gt;&#10;&lt;p&gt;Three calls later, the system sounds more confident than before.&lt;/p&gt;&#10;&lt;p&gt;But if all three agents share the same blind spot, the extra machinery has not created independent evidence.&lt;/p&gt;</description></item><item><title>The Moment: Intelligence beyond context</title><link>https://aibussin.com/post/moment/</link><pubDate>Fri, 19 Jun 2026 14:49:56 +0100</pubDate><guid>https://aibussin.com/post/moment/</guid><description>&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;Most AI systems answer and move on. The next step is different: preserve the reasoning state, replay it, measure whether it improves, and keep only what survives verification.&lt;/strong&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;&#10;&lt;p&gt;Most AI workflows still treat intelligence as a single pass.&lt;/p&gt;&#10;&lt;p&gt;You ask a question.&#10;The model answers.&#10;Maybe you ask it to try again.&#10;Maybe you add more context.&#10;Maybe you save something to memory.&lt;/p&gt;&#10;&lt;p&gt;But the basic shape remains the same:&lt;/p&gt;</description></item></channel></rss>