<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Agent Architecture on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/agent-architecture/</link><description>Recent content in Agent Architecture on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Sun, 30 Aug 2026 21:01:00 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/agent-architecture/index.xml" rel="self" type="application/rss+xml"/><item><title>Planning and Execution</title><link>https://aibussin.com/books/agents-from-first-principles/05-chapter/</link><pubDate>Sat, 08 Aug 2026 16:35:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/05-chapter/</guid><description>&lt;p&gt;The critique loop works on one thing at a time. It assumes the task already exists as a candidate we can hold, inspect and improve.&lt;/p&gt;&#10;&lt;p&gt;Some tasks have no draft to hold.&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Inspect a project, reproduce the failing test, find the cause, patch the code, rerun the relevant tests, and report what changed.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;No single narrow action responsibly completes that goal, and the actions constrain each other. Patching before diagnosis is guesswork wearing the costume of work, and reporting success before observing a passing test is a claim about the world that nothing in the run supports. Worse, if execution reveals that an assumption was wrong, the rest of the route may be wrong too. A runtime choosing each action independently can react to the new observation, but without an explicit representation of the intended route it has nothing concrete to compare the changed world against.&lt;/p&gt;</description></item><item><title>Architecting Agent-Based Systems</title><link>https://aibussin.com/books/agent-architectures/06-chapter/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/agent-architectures/06-chapter/</guid><description>&lt;p&gt;The previous chapters built up one agentic workflow: roles, tools, memory, reflection, revision, and versioning. This chapter changes scale.&lt;/p&gt;&#10;&lt;p&gt;An agent-based system is not one large assistant with a bigger prompt. It is a collection of components that divide responsibility, communicate through explicit channels, use shared or separate state, and preserve enough identity that you can tell which part did what.&lt;/p&gt;&#10;&lt;p&gt;The design question is:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Which responsibilities belong together,&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;which should be separated,&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;and how should the separated parts coordinate?&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Architecture begins when those boundaries have to be chosen deliberately.&lt;/p&gt;</description></item><item><title>Memory and Selective Recall</title><link>https://aibussin.com/books/agents-from-first-principles/08-chapter/</link><pubDate>Sat, 08 Aug 2026 17:14:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/08-chapter/</guid><description>&lt;p&gt;The capability boundary gave the agent a defined action space and a rule for which capabilities are eligible at each step. Runtime state gave it an explicit working representation of what this run has established so far and how it reached that point. Together they govern the current execution, but neither gives information from an earlier run a controlled way to influence this one. There is a family of failures that requires exactly that.&lt;/p&gt;</description></item><item><title>Building the Complete Agent</title><link>https://aibussin.com/books/agents-from-first-principles/11-chapter/</link><pubDate>Wed, 26 Aug 2026 10:00:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/11-chapter/</guid><description>&lt;p&gt;Every mechanism in this book was argued against a problem chosen to isolate it.&lt;/p&gt;&#10;&lt;p&gt;That isolation was deliberate, and it was also a form of protection. The memory chapter picked a task where recall was the bottleneck, held everything else still, and measured the one thing it came to measure. The result is a clean explanation and a weak claim. Nothing in it establishes that the same retrieval policy behaves when a search controller is expanding forty nodes, or when a verifier insists that every piece of evidence carry a state identity the search controller has never heard of.&lt;/p&gt;</description></item><item><title>Building Systems That Distrust Their Models</title><link>https://aibussin.com/books/hallucination-from-first-principles/15-chapter/</link><pubDate>Sun, 30 Aug 2026 21:01:00 +0100</pubDate><guid>https://aibussin.com/books/hallucination-from-first-principles/15-chapter/</guid><description>&lt;p&gt;The first chapter began with a simple observation:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;A language model can produce a fluent answer&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;without possessing a mechanism that proves the answer is true.&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Fourteen chapters later, that fact has not changed.&lt;/p&gt;&#10;&lt;p&gt;The model can still:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;invent&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;misbind&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;misattribute&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ignore decisive context&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;answer without enough evidence&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;accept bad retrieval&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;repair one error by creating another&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;repeat its own stored mistake&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The final architecture does not make those possibilities disappear.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 23: Your Infrastructure Is Healthy. Why Is the Agent Getting Worse? Detect Behavioral Drift and Roll Back Safely</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/23-chapter/</link><pubDate>Sun, 09 Aug 2026 12:00:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/23-chapter/</guid><description>&lt;h1 id="your-infrastructure-is-healthy-why-is-the-agent-getting-worse"&gt;Your Infrastructure Is Healthy. Why Is the Agent Getting Worse?&lt;/h1&gt;&#10;&lt;p&gt;Your dashboards are green.&lt;/p&gt;&#10;&lt;p&gt;The model endpoint is responding.&lt;/p&gt;&#10;&lt;p&gt;The browser workers are alive.&lt;/p&gt;&#10;&lt;p&gt;The database is healthy.&lt;/p&gt;&#10;&lt;p&gt;The queue is draining.&lt;/p&gt;&#10;&lt;p&gt;The verifier service is up.&lt;/p&gt;&#10;&lt;p&gt;Latency has not exploded.&lt;/p&gt;&#10;&lt;p&gt;There are no obvious exceptions.&lt;/p&gt;&#10;&lt;p&gt;And yet the agent is getting worse.&lt;/p&gt;&#10;&lt;p&gt;It fixes fewer bugs.&lt;/p&gt;&#10;&lt;p&gt;It retrieves weaker evidence.&lt;/p&gt;&#10;&lt;p&gt;It escalates to expensive models more often.&lt;/p&gt;&#10;&lt;p&gt;It chooses the wrong tools more frequently.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 21: What Happens When Too Many Agents Compete for the Same Resources? Add Admission Control, Quotas and Backpressure</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/21-chapter/</link><pubDate>Sun, 09 Aug 2026 11:21:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/21-chapter/</guid><description>&lt;h1 id="what-happens-when-too-many-agents-compete-for-the-same-resources"&gt;What Happens When Too Many Agents Compete for the Same Resources?&lt;/h1&gt;&#10;&lt;p&gt;A single agent can look healthy in isolation.&lt;/p&gt;&#10;&lt;p&gt;It gets a request.&lt;/p&gt;&#10;&lt;p&gt;It calls a model.&lt;/p&gt;&#10;&lt;p&gt;It launches a few search branches.&lt;/p&gt;&#10;&lt;p&gt;It opens a browser.&lt;/p&gt;&#10;&lt;p&gt;It runs tests.&lt;/p&gt;&#10;&lt;p&gt;It asks a verifier to check the result.&lt;/p&gt;&#10;&lt;p&gt;Everything works.&lt;/p&gt;&#10;&lt;p&gt;Then production traffic arrives.&lt;/p&gt;&#10;&lt;p&gt;Ten agents start at once.&lt;/p&gt;&#10;&lt;p&gt;Then fifty.&lt;/p&gt;&#10;&lt;p&gt;Then five hundred.&lt;/p&gt;&#10;&lt;p&gt;Now every agent still has a perfectly reasonable local plan.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 20: Can Your Agent Coordinate Across Machines Without Duplicating Work? Use Leases, Idempotency and Fencing</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/20-chapter/</link><pubDate>Sun, 09 Aug 2026 11:13:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/20-chapter/</guid><description>&lt;h1 id="can-your-agent-coordinate-across-machines-without-duplicating-work"&gt;Can Your Agent Coordinate Across Machines Without Duplicating Work?&lt;/h1&gt;&#10;&lt;p&gt;A single-process agent can already be complicated.&lt;/p&gt;&#10;&lt;p&gt;It can plan.&lt;/p&gt;&#10;&lt;p&gt;It can search.&lt;/p&gt;&#10;&lt;p&gt;It can launch speculative branches.&lt;/p&gt;&#10;&lt;p&gt;It can cancel losing work.&lt;/p&gt;&#10;&lt;p&gt;It can verify outcomes.&lt;/p&gt;&#10;&lt;p&gt;Then you move that work onto multiple workers.&lt;/p&gt;&#10;&lt;p&gt;Now a new class of failure appears.&lt;/p&gt;&#10;&lt;p&gt;Two workers both believe they own the same task.&lt;/p&gt;&#10;&lt;p&gt;One worker pauses for thirty seconds.&lt;/p&gt;&#10;&lt;p&gt;Another worker assumes it died and takes over.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 19: Can Your Agent Explore in Parallel Without Creating Chaos? Use Speculative Execution and Early Cancellation</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/19-chapter/</link><pubDate>Sun, 09 Aug 2026 11:10:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/19-chapter/</guid><description>&lt;h1 id="advanced-agents-from-first-principles-19-can-your-agent-explore-in-parallel-without-creating-chaos-use-speculative-execution-and-early-cancellation"&gt;Advanced Agents From First Principles 19: Can Your Agent Explore in Parallel Without Creating Chaos? Use Speculative Execution and Early Cancellation&lt;/h1&gt;&#10;&lt;p&gt;A production agent often has more than one useful thing it could do next.&lt;/p&gt;&#10;&lt;p&gt;It could:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;inspect repository state,&lt;/li&gt;&#10;&lt;li&gt;run a targeted test,&lt;/li&gt;&#10;&lt;li&gt;retrieve documentation,&lt;/li&gt;&#10;&lt;li&gt;ask a second model to critique a candidate,&lt;/li&gt;&#10;&lt;li&gt;generate an alternative implementation,&lt;/li&gt;&#10;&lt;li&gt;probe an API,&lt;/li&gt;&#10;&lt;li&gt;inspect a deployment,&lt;/li&gt;&#10;&lt;li&gt;or verify an invariant.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;If those actions are independent, executing them one by one can be needlessly slow.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 18: What Should Your Agent Observe Next? Use Expected Value of Information</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/18-chapter/</link><pubDate>Sun, 09 Aug 2026 11:06:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/18-chapter/</guid><description>&lt;h1 id="what-should-your-agent-observe-next"&gt;What Should Your Agent Observe Next?&lt;/h1&gt;&#10;&lt;p&gt;Your agent is uncertain.&lt;/p&gt;&#10;&lt;p&gt;That does not tell you what to do.&lt;/p&gt;&#10;&lt;p&gt;In the previous post we split uncertainty into operational categories:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;interpretation uncertainty,&lt;/li&gt;&#10;&lt;li&gt;evidence uncertainty,&lt;/li&gt;&#10;&lt;li&gt;route uncertainty,&lt;/li&gt;&#10;&lt;li&gt;state uncertainty,&lt;/li&gt;&#10;&lt;li&gt;tool uncertainty,&lt;/li&gt;&#10;&lt;li&gt;candidate uncertainty,&lt;/li&gt;&#10;&lt;li&gt;verification uncertainty.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;That is already better than one generic confidence score.&lt;/p&gt;&#10;&lt;p&gt;But it still leaves a harder question:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;Which piece of information is worth buying next?&lt;/strong&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;Suppose a coding agent is trying to fix a failing test.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 17: What Is Your Agent Actually Uncertain About?</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/17-chapter/</link><pubDate>Sun, 09 Aug 2026 11:02:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/17-chapter/</guid><description>&lt;h1 id="what-is-your-agent-actually-uncertain-about"&gt;What Is Your Agent Actually Uncertain About?&lt;/h1&gt;&#10;&lt;p&gt;An agent reaches a difficult point in a task.&lt;/p&gt;&#10;&lt;p&gt;It is not sure what to do next.&lt;/p&gt;&#10;&lt;p&gt;A common implementation responds like this:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uncertain&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;call the model again&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;still uncertain&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;call a stronger model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;still uncertain&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;search more&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That is not a reasoning strategy.&lt;/p&gt;&#10;&lt;p&gt;It is a spending strategy.&lt;/p&gt;&#10;&lt;p&gt;The system is using more computation without identifying what information is actually missing.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 16: Where Should an Agent Spend Its Compute? Build a Dynamic Budget Scheduler</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/16-chapter/</link><pubDate>Sun, 09 Aug 2026 10:53:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/16-chapter/</guid><description>&lt;h1 id="where-should-an-agent-spend-its-compute"&gt;Where Should an Agent Spend Its Compute?&lt;/h1&gt;&#10;&lt;p&gt;A production agent has a budget whether you designed one or not.&lt;/p&gt;&#10;&lt;p&gt;Every model call costs something.&lt;/p&gt;&#10;&lt;p&gt;Every search node costs something.&lt;/p&gt;&#10;&lt;p&gt;Every tool invocation costs something.&lt;/p&gt;&#10;&lt;p&gt;Every verifier costs something.&lt;/p&gt;&#10;&lt;p&gt;Every retry adds latency.&lt;/p&gt;&#10;&lt;p&gt;Every escalation to a stronger model spends money and time that could have been used somewhere else.&lt;/p&gt;&#10;&lt;p&gt;The naive architecture gives every subsystem its own fixed limit:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;MAX_STEPS &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;20&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;MAX_SEARCH_NODES &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;32&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;MAX_CRITIC_CALLS &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;3&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;MAX_RETRIES &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;4&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;MAX_VERIFIER_CALLS &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#ae81ff"&gt;2&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That looks safe.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 15: How Do You Optimize an Agent Policy Without Turning It Into Another Black Box?</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/15-chapter/</link><pubDate>Sun, 09 Aug 2026 10:49:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/15-chapter/</guid><description>&lt;h1 id="how-do-you-optimize-an-agent-policy-without-turning-it-into-another-black-box"&gt;How Do You Optimize an Agent Policy Without Turning It Into Another Black Box?&lt;/h1&gt;&#10;&lt;p&gt;By now our advanced agent can do a lot.&lt;/p&gt;&#10;&lt;p&gt;It can:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;route tasks to different models or specialists,&lt;/li&gt;&#10;&lt;li&gt;decide whether to search,&lt;/li&gt;&#10;&lt;li&gt;choose a search budget,&lt;/li&gt;&#10;&lt;li&gt;decide when to escalate,&lt;/li&gt;&#10;&lt;li&gt;invoke critics,&lt;/li&gt;&#10;&lt;li&gt;retry or recover,&lt;/li&gt;&#10;&lt;li&gt;stop when evidence is strong enough,&lt;/li&gt;&#10;&lt;li&gt;learn from verified production trajectories,&lt;/li&gt;&#10;&lt;li&gt;and trace the decisions that produced each outcome.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;That creates a new problem.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 14: Can Your Agent Learn From Its Own Trajectories Without Learning the Wrong Lessons?</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/14-chapter/</link><pubDate>Sun, 09 Aug 2026 10:33:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/14-chapter/</guid><description>&lt;p&gt;An advanced agent now leaves behind something extremely valuable:&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;evidence.&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;Not merely chat history.&lt;/p&gt;&#10;&lt;p&gt;Not merely model outputs.&lt;/p&gt;&#10;&lt;p&gt;Not merely traces.&lt;/p&gt;&#10;&lt;p&gt;A sufficiently instrumented system can record:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;what state it was in,&lt;/li&gt;&#10;&lt;li&gt;what alternatives it considered,&lt;/li&gt;&#10;&lt;li&gt;which route it selected,&lt;/li&gt;&#10;&lt;li&gt;what branches it pruned,&lt;/li&gt;&#10;&lt;li&gt;which model or specialist it escalated to,&lt;/li&gt;&#10;&lt;li&gt;which tools it called,&lt;/li&gt;&#10;&lt;li&gt;which critic changed the answer,&lt;/li&gt;&#10;&lt;li&gt;what verification evidence was produced,&lt;/li&gt;&#10;&lt;li&gt;how much compute was spent,&lt;/li&gt;&#10;&lt;li&gt;and whether the final result actually passed.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;That immediately suggests a tempting idea:&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 13: How Do You Debug an Agent That Made the Wrong Decision? Add Trajectory Observability</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/13-chapter/</link><pubDate>Sun, 09 Aug 2026 10:18:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/13-chapter/</guid><description>&lt;p&gt;An advanced agent fails.&lt;/p&gt;&#10;&lt;p&gt;You look at the final answer.&lt;/p&gt;&#10;&lt;p&gt;It is wrong.&lt;/p&gt;&#10;&lt;p&gt;So you inspect the prompt.&lt;/p&gt;&#10;&lt;p&gt;The prompt looks reasonable.&lt;/p&gt;&#10;&lt;p&gt;You inspect the model response.&lt;/p&gt;&#10;&lt;p&gt;That also looks reasonable.&lt;/p&gt;&#10;&lt;p&gt;But somewhere between the original request and the final result the system:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;chose the wrong specialist,&lt;/li&gt;&#10;&lt;li&gt;pruned the branch that contained the right solution,&lt;/li&gt;&#10;&lt;li&gt;trusted a critic that was wrong,&lt;/li&gt;&#10;&lt;li&gt;escalated to an expensive model unnecessarily,&lt;/li&gt;&#10;&lt;li&gt;failed to escalate when it should have,&lt;/li&gt;&#10;&lt;li&gt;retrieved stale memory,&lt;/li&gt;&#10;&lt;li&gt;spent most of its budget exploring duplicates,&lt;/li&gt;&#10;&lt;li&gt;accepted a weak verifier signal,&lt;/li&gt;&#10;&lt;li&gt;retried the same strategy under a different name,&lt;/li&gt;&#10;&lt;li&gt;or transformed a local success into a global failure.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The final answer does not tell you which one happened.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 11: Which Advanced Agent Architecture Should You Use? A Practical Selection Guide</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/11-chapter/</link><pubDate>Sun, 09 Aug 2026 09:20:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/11-chapter/</guid><description>&lt;h1 id="which-advanced-agent-architecture-should-you-use"&gt;Which Advanced Agent Architecture Should You Use?&lt;/h1&gt;&#10;&lt;p&gt;You now have too many options.&lt;/p&gt;&#10;&lt;p&gt;That is a better problem than having none.&lt;/p&gt;&#10;&lt;p&gt;But it is still a problem.&lt;/p&gt;&#10;&lt;p&gt;You can add:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;self-consistency,&lt;/li&gt;&#10;&lt;li&gt;Tree of Thoughts,&lt;/li&gt;&#10;&lt;li&gt;beam search,&lt;/li&gt;&#10;&lt;li&gt;Monte Carlo Tree Search,&lt;/li&gt;&#10;&lt;li&gt;evolutionary search,&lt;/li&gt;&#10;&lt;li&gt;specialist routing,&lt;/li&gt;&#10;&lt;li&gt;planner/executor/critic separation,&lt;/li&gt;&#10;&lt;li&gt;multi-agent debate,&lt;/li&gt;&#10;&lt;li&gt;adaptive policies,&lt;/li&gt;&#10;&lt;li&gt;learning from previous runs,&lt;/li&gt;&#10;&lt;li&gt;or a mixture-of-agents runtime that chooses among several of them.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The temptation is to combine everything.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 29: When Should an Agent Stop and Ask a Human? Design Authority Boundaries and Escalation</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/29-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/29-chapter/</guid><description>A practical architecture for agent authority boundaries: decide what an agent may do autonomously, when it must escalate, what evidence a human reviewer needs, and how to avoid turning human approval into rubber-stamping.</description></item><item><title>Advanced Agents From First Principles 32: Which Capabilities Are Actually Worth Building? Design a Capability Portfolio</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/32-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/32-chapter/</guid><description>A practical framework for deciding which agent capabilities are worth acquiring: rank missing capabilities by user value, verifier availability, reliability risk, acquisition cost, maintenance burden, and the quality of human or deterministic alternatives.</description></item><item><title>Advanced Agents From First Principles 33: Which Shared Components Actually Unlock More Capability? Build a Capability Dependency Graph</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/33-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/33-chapter/</guid><description>A practical capability-dependency architecture for advanced agents: identify shared primitives that unlock many capabilities, quantify leverage, expose correlated-failure hotspots, and invest in platform components without creating hidden systemic risk.</description></item><item><title>Advanced Agents From First Principles 34: Where Should This Task Actually Run? Build Capability-Aware Placement Across Models, Providers and Resource Pools</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/34-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/34-chapter/</guid><description>A practical placement architecture for advanced agents: route work across local and frontier models, providers, regions, GPUs, browser pools and specialist runtimes using demonstrated competence, verifier availability, policy constraints, health, cost and latency rather than model prestige.</description></item><item><title>Advanced Agents From First Principles 38: A Plan Is Not a Commitment — Model Goals, Commitments and Executable Work</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/38-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/38-chapter/</guid><description>A practical architecture for separating goals, plans, commitments, tasks and actions in long-running agents so replanning, cancellation, handoff and external obligations remain correct.</description></item><item><title>Advanced Agents From First Principles 42: How Do Multiple Agents Coordinate Without Becoming a Distributed Argument?</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/42-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/42-chapter/</guid><description>A practical architecture for multi-agent coordination: explicit ownership, delegation, contracts, commitment transfer, shared intent, evidence provenance, conflict handling, deadlock prevention, and independent verification instead of agents merely chatting until they agree.</description></item><item><title>Advanced Agents From First Principles 43: Who Controls the Agent? Build an Explicit Agent Control Plane</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/43-chapter/</link><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/43-chapter/</guid><description>A production-agent architecture that separates control-plane policy from execution-plane reasoning: intent, competence, authority, placement, budgets, reliability, releases, security and escalation remain enforceable outside the model.</description></item><item><title>Advanced Agents From First Principles 10: Are You Combining Every Agent Technique Into One Monster? Build a Mixture-of-Agents Runtime</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/10-chapter/</link><pubDate>Sun, 09 Aug 2026 00:25:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/10-chapter/</guid><description>&lt;p&gt;You have a working agent.&lt;/p&gt;&#10;&lt;p&gt;Then you add retrieval.&lt;/p&gt;&#10;&lt;p&gt;Then memory.&lt;/p&gt;&#10;&lt;p&gt;Then Best-of-N.&lt;/p&gt;&#10;&lt;p&gt;Then critique and revision.&lt;/p&gt;&#10;&lt;p&gt;Then Tree of Thoughts.&lt;/p&gt;&#10;&lt;p&gt;Then MCTS.&lt;/p&gt;&#10;&lt;p&gt;Then specialist models.&lt;/p&gt;&#10;&lt;p&gt;Then adversarial review.&lt;/p&gt;&#10;&lt;p&gt;Then a planner, executor, critic and verifier.&lt;/p&gt;&#10;&lt;p&gt;Then a stronger model for hard cases.&lt;/p&gt;&#10;&lt;p&gt;Eventually the architecture starts to look like this:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;request&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;planner&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;retrieval&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;reasoning&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Best-of-N&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;critic&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Tree of Thoughts&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;MCTS&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;frontier model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;second critic&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;verifier&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Every mechanism may have been individually reasonable.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 09: Can Your Agent Actually Learn From Previous Runs?</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/09-chapter/</link><pubDate>Sun, 09 Aug 2026 00:19:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/09-chapter/</guid><description>&lt;h1 id="can-your-agent-actually-learn-from-previous-runs"&gt;Can Your Agent Actually Learn From Previous Runs?&lt;/h1&gt;&#10;&lt;p&gt;A production agent can execute the same class of task hundreds or thousands of times.&lt;/p&gt;&#10;&lt;p&gt;It can see the same failure repeatedly.&lt;/p&gt;&#10;&lt;p&gt;It can discover the same workaround repeatedly.&lt;/p&gt;&#10;&lt;p&gt;It can call the same expensive model repeatedly.&lt;/p&gt;&#10;&lt;p&gt;And still behave as if every task is the first one it has ever seen.&lt;/p&gt;&#10;&lt;p&gt;That is not necessarily a memory problem.&lt;/p&gt;&#10;&lt;p&gt;It may already have excellent memory.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 00: When Should You Use an Advanced Agent Architecture?</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/00-chapter/</link><pubDate>Sat, 08 Aug 2026 22:27:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/00-chapter/</guid><description>&lt;h1 id="advanced-agents-from-first-principles-00-when-should-you-use-an-advanced-agent-architecture"&gt;Advanced Agents From First Principles 00: When Should You Use an Advanced Agent Architecture?&lt;/h1&gt;&#10;&lt;p&gt;You built an agent.&lt;/p&gt;&#10;&lt;p&gt;It can:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;call tools,&lt;/li&gt;&#10;&lt;li&gt;maintain state,&lt;/li&gt;&#10;&lt;li&gt;plan,&lt;/li&gt;&#10;&lt;li&gt;revise its own work,&lt;/li&gt;&#10;&lt;li&gt;search over alternatives,&lt;/li&gt;&#10;&lt;li&gt;remember useful information,&lt;/li&gt;&#10;&lt;li&gt;and verify whether the requested outcome actually happened.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Now the temptation begins.&lt;/p&gt;&#10;&lt;p&gt;You add another model.&lt;/p&gt;&#10;&lt;p&gt;Then a critic.&lt;/p&gt;&#10;&lt;p&gt;Then a planner.&lt;/p&gt;&#10;&lt;p&gt;Then a judge.&lt;/p&gt;&#10;&lt;p&gt;Then a router.&lt;/p&gt;&#10;&lt;p&gt;Then three specialist agents.&lt;/p&gt;&#10;&lt;p&gt;Then a tree search.&lt;/p&gt;</description></item><item><title>Agents From First Principles 07: AI Agent Forgets Previous Work? Add Working, Semantic and Episodic Memory</title><link>https://aibussin.com/post/agents-from-first-principles-07/</link><pubDate>Sat, 08 Aug 2026 17:14:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-07/</guid><description>&lt;h1 id="ai-agent-forgets-previous-work-add-working-semantic-and-episodic-memory"&gt;AI Agent Forgets Previous Work? Add Working, Semantic and Episodic Memory&lt;/h1&gt;&#10;&lt;p&gt;An agent can use the right model, call the right tools, execute the right plan, and still behave as if nothing that happened five minutes ago matters.&lt;/p&gt;&#10;&lt;p&gt;You see the symptoms quickly:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;it re-reads files it already inspected;&lt;/li&gt;&#10;&lt;li&gt;it repeats research it already completed;&lt;/li&gt;&#10;&lt;li&gt;it asks for information the user already supplied;&lt;/li&gt;&#10;&lt;li&gt;it forgets why a previous approach failed;&lt;/li&gt;&#10;&lt;li&gt;it loses decisions made earlier in a long task;&lt;/li&gt;&#10;&lt;li&gt;it treats every new run as if the system has never seen the problem before;&lt;/li&gt;&#10;&lt;li&gt;it retrieves an old answer and treats it as current truth;&lt;/li&gt;&#10;&lt;li&gt;it fills the prompt with so much history that the useful information is buried.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The usual response is:&lt;/p&gt;</description></item><item><title>Agents From First Principles 04: AI Agent Fails on Multi-Step Tasks? Separate Planning From Execution</title><link>https://aibussin.com/post/agents-from-first-principles-04/</link><pubDate>Sat, 08 Aug 2026 16:35:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-04/</guid><description>&lt;p&gt;A surprising number of agent failures are not really model failures.&lt;/p&gt;&#10;&lt;p&gt;The model may be perfectly capable of writing each individual step. The failure happens because the system tries to decide &lt;strong&gt;what to do&lt;/strong&gt; and &lt;strong&gt;do it&lt;/strong&gt; at the same time.&lt;/p&gt;&#10;&lt;p&gt;That works for simple tasks:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;question&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;answer&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It becomes fragile when success depends on several ordered actions:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;goal&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;step 1&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;step 2&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;step 3&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;verification&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A useful next step in agent design is therefore to separate two jobs:&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 07: Do Your Agents Agree Too Easily? Use Adversarial Review and Multi-Agent Debate Without Confusing Debate With Truth</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/07-chapter/</link><pubDate>Sat, 08 Aug 2026 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/07-chapter/</guid><description>&lt;p&gt;A multi-agent system can look sophisticated while every agent quietly repeats the same mistake.&lt;/p&gt;&#10;&lt;p&gt;That is one of the most dangerous failure modes in advanced agent architectures.&lt;/p&gt;&#10;&lt;p&gt;You ask one model to solve the problem.&lt;/p&gt;&#10;&lt;p&gt;Then you ask a second model to review it.&lt;/p&gt;&#10;&lt;p&gt;Then a third model judges the disagreement.&lt;/p&gt;&#10;&lt;p&gt;Three calls later, the system sounds more confident than before.&lt;/p&gt;&#10;&lt;p&gt;But if all three agents share the same blind spot, the extra machinery has not created independent evidence.&lt;/p&gt;</description></item><item><title>A Memory Gate for AI: Policy-Bounded Acceptance in the Executable Cognitive Kernel</title><link>https://aibussin.com/post/verify/</link><pubDate>Tue, 17 Mar 2026 09:58:14 +0000</pubDate><guid>https://aibussin.com/post/verify/</guid><description>&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;&#10;&lt;p&gt;Dynamic AI systems face a hidden failure mode: they can learn from their own mistakes.&#10;If every output is allowed into memory, stochastic errors do not stay local they accumulate.&lt;/p&gt;&#10;&lt;p&gt;In earlier posts, I argued that AI systems should not be trusted to enforce their own correctness.&lt;/p&gt;&#10;&lt;p&gt;Modern models are stochastic. They produce correct outputs, partially correct outputs, and completely incorrect outputs, but they do not reliably distinguish between them. That means a system that stores everything it generates will eventually learn from its own mistakes.&lt;/p&gt;</description></item><item><title>Intelligence Through Execution: The Executable Cognitive Kernel</title><link>https://aibussin.com/post/eck/</link><pubDate>Tue, 10 Mar 2026 21:58:14 +0000</pubDate><guid>https://aibussin.com/post/eck/</guid><description>&lt;h2 id="-summary"&gt;🧭 Summary&lt;/h2&gt;&#10;&lt;p&gt;Most modern AI systems treat intelligence as something stored inside a model.&lt;/p&gt;&#10;&lt;p&gt;A neural network is trained on massive datasets, its weights are adjusted, and those weights become the system’s knowledge. When the model produces an output, we interpret that output as the result of the intelligence encoded inside those parameters.&lt;/p&gt;&#10;&lt;p&gt;But this perspective has a limitation.&lt;/p&gt;&#10;&lt;p&gt;Once training is complete, the model is largely static. It does not improve through its own actions, and it does not adapt based on the outcome of its behavior unless we retrain it.&lt;/p&gt;</description></item><item><title>Self-Improving AI: A System That Learns, Validates, and Retrains Itself</title><link>https://aibussin.com/post/rivals/</link><pubDate>Mon, 30 Jun 2025 10:13:03 +0100</pubDate><guid>https://aibussin.com/post/rivals/</guid><description>&lt;h2 id="-the-static-ai-trap"&gt;🤖 &lt;strong&gt;The Static AI Trap&lt;/strong&gt;&lt;/h2&gt;&#10;&lt;p&gt;Today’s AI systems are frozen in time: trained once, deployed forever. Yet the real world never stops evolving. Goals shift overnight. New research upends old truths. Context transforms without warning.&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;What if your AI could wake up?&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;In this post, we engineer an intelligence that &lt;strong&gt;teaches itself&lt;/strong&gt; a system that continuously learns from the web, audits its own judgments, and retrains itself when confidence wavers.&lt;/p&gt;</description></item><item><title>Thoughts of Algorithms</title><link>https://aibussin.com/post/thoughts/</link><pubDate>Mon, 23 Jun 2025 11:10:59 +0100</pubDate><guid>https://aibussin.com/post/thoughts/</guid><description>&lt;blockquote&gt;&#10;&lt;p&gt;How a self-evolving AI learns to reflect, score, and rewrite its own reasoning&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;h2 id="-summary"&gt;🧪 Summary&lt;/h2&gt;&#10;&lt;p&gt;What if an AI could think not just solve problems, but reevaluate its beliefs in the face of new information?&lt;/p&gt;&#10;&lt;p&gt;In this post, we introduce a system that does exactly that. At the core of our pipeline is a lightweight scoring model called MR.Q, responsible for evaluating ideas and choosing the best ones. But when it encounters a new domain, a new goal, or a shift in task format, it doesn’t freeze it adapts.&lt;/p&gt;</description></item><item><title>Document Intelligence: Turning Documents into Structured Knowledge</title><link>https://aibussin.com/post/docs/</link><pubDate>Tue, 17 Jun 2025 23:31:13 +0100</pubDate><guid>https://aibussin.com/post/docs/</guid><description>&lt;h2 id="-summary"&gt;📖 Summary&lt;/h2&gt;&#10;&lt;p&gt;Imagine drowning in a sea of research papers, each holding a fragment of the knowledge you need for your next breakthrough. How does an AI system, striving for self-improvement, navigate this information overload to find precisely what it needs? This is the core challenge our Document Intelligence pipeline addresses, transforming chaotic documents into organized, searchable knowledge.&lt;/p&gt;&#10;&lt;p&gt;In this post we combine insights from &lt;a href="https://arxiv.org/pdf/2505.21497" target="_blank" class="paper-badge"&#10; style="display: inline-block; padding: 6px 10px; background: #f3f4f6; border-left: 4px solid #3b82f6; border-radius: 4px; margin: 4px 0; text-decoration: none; color: #1f2937;"&gt;&#10; &lt;strong&gt;Paper2Poster&lt;/strong&gt;: Towards Multimodal Poster Automation from Scientific Papers&#10;&lt;/a&gt; and&#10;&lt;a href="https://arxiv.org/abs/2506.10952" target="_blank" class="paper-badge"&#10; style="display: inline-block; padding: 6px 10px; background: #f3f4f6; border-left: 4px solid #3b82f6; border-radius: 4px; margin: 4px 0; text-decoration: none; color: #1f2937;"&gt;&#10; &lt;strong&gt;Domain2Vec&lt;/strong&gt;: Vectorizing Datasets to Find the Optimal Data Mixture without Training&#10;&lt;/a&gt; to build an AI document profiler that transforms unstructured papers into structured, searchable knowledge graphs.&lt;/p&gt;</description></item><item><title>Learning to Learn: A LATS-Based Framework for Self-Aware AI Pipelines</title><link>https://aibussin.com/post/lats/</link><pubDate>Thu, 12 Jun 2025 09:23:46 +0100</pubDate><guid>https://aibussin.com/post/lats/</guid><description>&lt;h2 id="-summary"&gt;📖 Summary&lt;/h2&gt;&#10;&lt;p&gt;In this post, we introduce the LATSAgent, an implementation of &lt;a href="https://arxiv.org/pdf/2310.04406" target="_blank" class="paper-badge"&#10; style="display: inline-block; padding: 6px 10px; background: #f3f4f6; border-left: 4px solid #3b82f6; border-radius: 4px; margin: 4px 0; text-decoration: none; color: #1f2937;"&gt;&#10; &lt;strong&gt;LATS&lt;/strong&gt;: Language Agent Tree Search Unifies Reasoning..&#10;&lt;/a&gt; within the &lt;a href="https://github.com/ernanhughes/co-ai"&gt;stephanie&lt;/a&gt; framework. Unlike prior agents that followed a single reasoning chain, this agent explores multiple reasoning paths in parallel, evaluates them using multidimensional scoring, and learns symbolic refinements over time. This is our most complete integration yet of search, simulation, scoring, and symbolic tuning bringing together all of our previous work on sharpening, pipeline reflection, and symbolic rules into a unified, intelligent reasoning loop.&lt;/p&gt;</description></item><item><title>Programming Intelligence: Using Symbolic Rules to Steer and Evolve AI</title><link>https://aibussin.com/post/symbolic/</link><pubDate>Wed, 04 Jun 2025 20:57:20 +0100</pubDate><guid>https://aibussin.com/post/symbolic/</guid><description>&lt;h2 id="-summary"&gt;🧪 Summary&lt;/h2&gt;&#10;&lt;p&gt;&amp;ldquo;What if AI systems could learn how to improve themselves not just at the level of weights or prompts, but at the level of strategy itself? In this post, we show how to build such a system, powered by symbolic rules and reflection.&lt;/p&gt;&#10;&lt;p&gt;The paper &lt;a href="https://arxiv.org/pdf/2406.18532v1" target="_blank" class="paper-badge"&#10; style="display: inline-block; padding: 6px 10px; background: #f3f4f6; border-left: 4px solid #3b82f6; border-radius: 4px; margin: 4px 0; text-decoration: none; color: #1f2937;"&gt;&#10; &lt;strong&gt;Symbolic Agents&lt;/strong&gt;: Symbolic Learning Enables Self-Evolving Agents&#10;&lt;/a&gt; introduces a framework where &lt;strong&gt;symbolic rules&lt;/strong&gt; guide, evaluate, and evolve agent behavior.&lt;/p&gt;</description></item><item><title>Adaptive Reasoning with ARM: Teaching AI the Right Way to Think</title><link>https://aibussin.com/post/arm/</link><pubDate>Wed, 28 May 2025 22:22:46 +0100</pubDate><guid>https://aibussin.com/post/arm/</guid><description>&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;&#10;&lt;p&gt;Chain-of-thought is powerful, but which chain? Short explanations work for easy tasks, long reflections help on hard ones, and code sometimes beats them both. What if your model could adaptively pick the best strategy, per task, and improve as it learns?&lt;/p&gt;&#10;&lt;p&gt;The &lt;code&gt;Adaptive Reasoning Model&lt;/code&gt; &lt;strong&gt;(ARM)&lt;/strong&gt; is a framework for teaching language models how to choose the right reasoning format direct answers, chain-of-thoughts, or code depending on the task. It works by evaluating responses, scoring them based on rarity, conciseness, and difficulty alignment, and then updating model behavior over time.&lt;/p&gt;</description></item></channel></rss>