<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>LLM Agents on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/llm-agents/</link><description>Recent content in LLM Agents on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Sun, 30 Aug 2026 18:39:00 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/llm-agents/index.xml" rel="self" type="application/rss+xml"/><item><title>What Is an Agent, Really?</title><link>https://aibussin.com/books/agents-from-first-principles/01-chapter/</link><pubDate>Sat, 08 Aug 2026 15:40:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/01-chapter/</guid><description>&lt;p&gt;This book is about building systems &lt;strong&gt;around&lt;/strong&gt; models. Before we build planners, memory, tool routing, critics, search and verifiers, we need an answer to a question that turns out to be harder than it looks:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;What is an agent?&lt;/strong&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;The word is currently applied to almost everything: a single LLM call, a chatbot, a fixed pipeline, a tool-using loop, and any five model calls with class names ending in &lt;code&gt;Agent&lt;/code&gt;. Those systems may all be useful. But when one word covers all of them, it stops telling us anything about the computation, and we lose the ability to say which mechanism is doing the work.&lt;/p&gt;</description></item><item><title>Introduction to LLM Agents</title><link>https://aibussin.com/books/agent-architectures/01-chapter/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/agent-architectures/01-chapter/</guid><description>&lt;h3 id="what-is-an-llm-agent"&gt;What Is an LLM Agent?&lt;/h3&gt;&#10;&lt;p&gt;An &lt;strong&gt;LLM agent&lt;/strong&gt; is a software system built around a language model that can pursue a goal through context, state, tool use, feedback, and repeated interaction.&lt;/p&gt;&#10;&lt;p&gt;The model is important, but it is not the whole agent. The larger system supplies the prompt, chooses what context the model sees, exposes tools, records state, retrieves memory, validates actions, handles errors, and decides when the work is finished. When people say that an agent &amp;ldquo;remembers,&amp;rdquo; &amp;ldquo;plans,&amp;rdquo; or &amp;ldquo;uses a tool,&amp;rdquo; the precise mechanism usually belongs to this surrounding runtime, not to the model by itself.&lt;/p&gt;</description></item><item><title>Methodologies and Core Patterns</title><link>https://aibussin.com/books/agent-architectures/02-chapter/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/agent-architectures/02-chapter/</guid><description>&lt;h3 id="five-core-shifts-in-the-ai-human-paradigm"&gt;Five Core Shifts in the AI-Human Paradigm&lt;/h3&gt;&#10;&lt;p&gt;Chapter 1 introduced agents as systems built around language models, not as models alone. This chapter looks at the working patterns that make those systems useful.&lt;/p&gt;&#10;&lt;p&gt;The patterns are not complicated. They are easy to miss because they do not look like traditional programming. You converse, inspect, revise, compare, preserve, and try again. The loop is simple, but its consequences are large.&lt;/p&gt;&#10;&lt;p&gt;Five shifts sit underneath that loop.&lt;/p&gt;</description></item><item><title>Candidate Generation and Selection</title><link>https://aibussin.com/books/agents-from-first-principles/03-chapter/</link><pubDate>Sat, 08 Aug 2026 15:54:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/03-chapter/</guid><description>&lt;p&gt;The previous chapter built a boundary that stops arbitrary model output from acquiring execution authority without explicit checks. It made one class of failure inspectable and enforceable, and it is silent about another.&lt;/p&gt;&#10;&lt;p&gt;Suppose the model is asked to solve a coding problem. One run produces the right patch. The next produces a plausible but incomplete one. A third produces something better again. Nothing is malformed, nothing violates the action schema, and every one of them would pass the boundary we just built. Validity and quality are different properties. A proposal can be completely valid and still be a poor choice, which relocates the uncertainty rather than removing it:&lt;/p&gt;</description></item><item><title>Critique, Revision, and Acceptance</title><link>https://aibussin.com/books/agents-from-first-principles/04-chapter/</link><pubDate>Sat, 08 Aug 2026 16:23:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/04-chapter/</guid><description>&lt;p&gt;Selection can only choose among the candidates it is given. A strong evaluator may recognize that every available candidate is poor, but it cannot select a correct answer that the generator never produced.&lt;/p&gt;&#10;&lt;p&gt;That limitation becomes practical because the candidates usually come from one model answering one prompt, and such samples correlate. When four candidates share the same misreading of the evidence, ranking them can still produce a confident winner while leaving the shared defect untouched.&lt;/p&gt;</description></item><item><title>Planning and Execution</title><link>https://aibussin.com/books/agents-from-first-principles/05-chapter/</link><pubDate>Sat, 08 Aug 2026 16:35:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/05-chapter/</guid><description>&lt;p&gt;The critique loop works on one thing at a time. It assumes the task already exists as a candidate we can hold, inspect and improve.&lt;/p&gt;&#10;&lt;p&gt;Some tasks have no draft to hold.&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Inspect a project, reproduce the failing test, find the cause, patch the code, rerun the relevant tests, and report what changed.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;No single narrow action responsibly completes that goal, and the actions constrain each other. Patching before diagnosis is guesswork wearing the costume of work, and reporting success before observing a passing test is a claim about the world that nothing in the run supports. Worse, if execution reveals that an assumption was wrong, the rest of the route may be wrong too. A runtime choosing each action independently can react to the new observation, but without an explicit representation of the intended route it has nothing concrete to compare the changed world against.&lt;/p&gt;</description></item><item><title>Runtime State, Progress, and Termination</title><link>https://aibussin.com/books/agents-from-first-principles/06-chapter/</link><pubDate>Sat, 08 Aug 2026 16:56:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/06-chapter/</guid><description>&lt;p&gt;A plan is a claim about the future. It says that reproducing the failure, then diagnosing it, then patching, then testing, is a route from here to a working system. Making that claim explicit was worth the machinery. But it is written before any of the work happens, and execution may invalidate one of its assumptions almost immediately.&lt;/p&gt;&#10;&lt;p&gt;Execution produces something else entirely: a trajectory. Actions went out, observations came back, and some of what came back may contradict the intended route. The patch step failed because a dependency is missing. The test step is not merely late; it is unreachable.&lt;/p&gt;</description></item><item><title>Memory and Selective Recall</title><link>https://aibussin.com/books/agents-from-first-principles/08-chapter/</link><pubDate>Sat, 08 Aug 2026 17:14:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/08-chapter/</guid><description>&lt;p&gt;The capability boundary gave the agent a defined action space and a rule for which capabilities are eligible at each step. Runtime state gave it an explicit working representation of what this run has established so far and how it reached that point. Together they govern the current execution, but neither gives information from an earlier run a controlled way to influence this one. There is a family of failures that requires exactly that.&lt;/p&gt;</description></item><item><title>Trajectory Search</title><link>https://aibussin.com/books/agents-from-first-principles/09-chapter/</link><pubDate>Sat, 08 Aug 2026 17:26:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/09-chapter/</guid><description>&lt;p&gt;An agent can make every local decision look reasonable and still lose the task.&lt;/p&gt;&#10;&lt;p&gt;A coding agent sees that &lt;code&gt;test_checkout_redirect&lt;/code&gt; is failing, concludes there is an implementation bug in &lt;code&gt;checkout.py&lt;/code&gt;, and then behaves impeccably for twenty steps: it reads the file, edits it, runs the tests, repairs the new failures its edit introduced, rewrites the patch, and runs the tests again. Every one of those steps is defensible given the step before it.&lt;/p&gt;</description></item><item><title>The Memory Contamination Problem</title><link>https://aibussin.com/books/hallucination-from-first-principles/14-chapter/</link><pubDate>Sun, 30 Aug 2026 18:39:00 +0100</pubDate><guid>https://aibussin.com/books/hallucination-from-first-principles/14-chapter/</guid><description>&lt;p&gt;Chapter 13 ended with a strict recovery invariant:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;repair&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;→ new candidate state&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;→ re-measure&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;→ re-authorize&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;and with one additional rule:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;candidate in HOLD&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;→ do not persist as trusted factual memory&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That second rule changes the scale of the problem.&lt;/p&gt;&#10;&lt;p&gt;A hallucinated sentence displayed once is transient.&lt;/p&gt;&#10;&lt;p&gt;The same sentence written into persistent state can survive the conversation that created it.&lt;/p&gt;&#10;&lt;p&gt;It can then be:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;retrieved&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;summarized&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;copied&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;cited&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;used as agent experience&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;used to fill a user profile&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;inserted into a vector index&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;fed into another model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;used as a future repair source&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;and eventually appear to the system as if it came from somewhere else.&lt;/p&gt;</description></item><item><title>References and Supporting Papers</title><link>https://aibussin.com/books/agent-architectures/90-chapter/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://aibussin.com/books/agent-architectures/90-chapter/</guid><description>&lt;p&gt;This section collects papers, essays, and project references related to the themes in the book. It should be treated as a starting point for further reading, not as a fully audited citation apparatus. Some entries need source verification before a formal publication pass.&lt;/p&gt;&#10;&lt;h2 id="chapter-1-introduction-to-llm-agents"&gt;Chapter 1: Introduction to LLM Agents&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Vaswani et al. (2017). &lt;em&gt;Attention is All You Need&lt;/em&gt;. Introduced the transformer architecture foundational to modern LLMs.&lt;/li&gt;&#10;&lt;li&gt;Brown et al. (2020). &lt;em&gt;Language Models are Few-Shot Learners&lt;/em&gt; (GPT-3). Demonstrates general capabilities of LLMs as zero/few-shot learners.&lt;/li&gt;&#10;&lt;li&gt;OpenAI (2023). &lt;em&gt;Introducing Function Calling&lt;/em&gt;. Relevant to tool-calling interfaces and structured model outputs.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="chapter-2-methodologies-and-core-patterns"&gt;Chapter 2: Methodologies and Core Patterns&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Yao et al. (2022). &lt;em&gt;ReAct: Synergizing Reasoning and Acting in Language Models&lt;/em&gt;. Relevant to reasoning-and-acting loops.&lt;/li&gt;&#10;&lt;li&gt;Jiang et al. (2023). &lt;em&gt;Active-Prompt: Prompt Engineering with Chain-of-Thought Reasoning&lt;/em&gt;. Related to prompt refinement and idea iteration.&lt;/li&gt;&#10;&lt;li&gt;McLuhan, M. (1964). &lt;em&gt;Understanding Media: The Extensions of Man&lt;/em&gt;. &amp;ldquo;The medium is the message&amp;rdquo; section reference.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="chapter-3-the-architecture-of-agent-behavior"&gt;Chapter 3: The Architecture of Agent Behavior&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Shinn et al. (2023). &lt;em&gt;Reflexion: Language Agents with Verbal Reinforcement Learning&lt;/em&gt;. Relevant to reflection and feedback loops.&lt;/li&gt;&#10;&lt;li&gt;Liu et al. (2023). &lt;em&gt;ToolLLM: Facilitating Tool Learning with Language Models&lt;/em&gt;. Basis for tool-augmented agent capabilities.&lt;/li&gt;&#10;&lt;li&gt;Rajani et al. (2019). &lt;em&gt;Explain Yourself! Leveraging Language Models for Commonsense Reasoning&lt;/em&gt;. Relevant background for explanation and reasoning traces.&lt;/li&gt;&#10;&lt;li&gt;Microsoft (2023). &lt;em&gt;AutoGen: Enabling Next gen LLM Applications&lt;/em&gt;. Practical implementation of agent roles and multi-agent coordination.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="chapter-4-designing-your-first-agent-on-your-phone"&gt;Chapter 4: Designing Your First Agent (On Your Phone)&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;OpenAI Community &amp;amp; Prompt Engineering Guides (2022–2023). Prompt design as an accessible interface to agent behaviors.&lt;/li&gt;&#10;&lt;li&gt;Qin et al. (2023). &lt;em&gt;ToolBench: Towards Empowering Large Language Models with In-Context Tool Learning&lt;/em&gt;. Basis for simulating tools with prompts.&lt;/li&gt;&#10;&lt;li&gt;Shinn et al. (2023). &lt;em&gt;Reflexion&lt;/em&gt;. Related to critique and revision patterns.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="chapter-5-the-thinking-agent"&gt;Chapter 5: The Thinking Agent&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Madaan et al. (2023). &lt;em&gt;Self-Refine: Iterative Refinement with Self-Feedback&lt;/em&gt;. The model-as-critic structure.&lt;/li&gt;&#10;&lt;li&gt;Liu et al. (2023). &lt;em&gt;Reviewer LLMs&lt;/em&gt;. Citation needs verification before publication.&lt;/li&gt;&#10;&lt;li&gt;Bai et al. (2022). &lt;em&gt;Training a Helpful and Harmless Assistant with RLHF&lt;/em&gt;. Introduces reward feedback loops and output alignment strategies.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="chapter-6-architecting-agent-based-systems"&gt;Chapter 6: Architecting Agent-Based Systems&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Wu et al. (2023). &lt;em&gt;AgentVerse: Facilitating Multi-Agent Collaboration&lt;/em&gt;. Details centralized vs decentralized agent architectures.&lt;/li&gt;&#10;&lt;li&gt;Zhang et al. (2023). &lt;em&gt;CAMEL: Communicative Agents for Mind Exploration of Large Scale Language Model Society&lt;/em&gt;. Supports multi-agent dialogue frameworks.&lt;/li&gt;&#10;&lt;li&gt;Zeng et al. (2022). &lt;em&gt;A Survey of Multi-Agent Systems&lt;/em&gt;. Gives academic grounding to MAS coordination techniques.&lt;/li&gt;&#10;&lt;li&gt;Patil et al. (2023). &lt;em&gt;Gorilla: Large Language Model Connected with Massive APIs&lt;/em&gt;. Relevant to tool and API use.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="chapter-7-a-new-way-of-working-with-technology"&gt;Chapter 7: A New Way of Working With Technology&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Schick et al. (2023). &lt;em&gt;Toolformer: Language Models Can Teach Themselves to Use Tools&lt;/em&gt;. &lt;a href="https://arxiv.org/abs/2302.04761"&gt;arXiv:2302.04761&lt;/a&gt;&lt;br&gt;&#10;→ Relevant to tool-use interfaces.&lt;/li&gt;&#10;&lt;li&gt;Paranjape et al. (2023). &lt;em&gt;DSPy: Compiling Declarative Language Model Programs&lt;/em&gt;. &lt;a href="https://arxiv.org/abs/2310.01848"&gt;arXiv:2310.01848&lt;/a&gt;&lt;br&gt;&#10;→ Relevant to declarative language-model programs.&lt;/li&gt;&#10;&lt;li&gt;Yao et al. (2022). &lt;em&gt;ReAct: Synergizing Reasoning and Acting in Language Models&lt;/em&gt;. &lt;a href="https://arxiv.org/abs/2210.03629"&gt;arXiv:2210.03629&lt;/a&gt;&lt;br&gt;&#10;→ Relevant to reasoning-and-acting loops.&lt;/li&gt;&#10;&lt;li&gt;Wu et al. (2023). &lt;em&gt;AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation Frameworks&lt;/em&gt;. &lt;a href="https://arxiv.org/abs/2309.11455"&gt;arXiv:2309.11455&lt;/a&gt;&lt;br&gt;&#10;→ Demonstrates modular, conversation-first interactions among agents.&lt;/li&gt;&#10;&lt;li&gt;Shinn et al. (2023). &lt;em&gt;Reflexion: Language Agents with Verbal Reinforcement Learning&lt;/em&gt;. &lt;a href="https://arxiv.org/abs/2303.11366"&gt;arXiv:2303.11366&lt;/a&gt;&lt;br&gt;&#10;→ Relevant to reflection and revision patterns.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="chapter-8-companion-agents"&gt;Chapter 8: Companion Agents&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Shinn et al. (2023). &lt;em&gt;Reflexion: Language Agents with Verbal Reinforcement Learning&lt;/em&gt;.&lt;br&gt;&#10;→ Related to reflection loops; does not by itself establish companion-agent memory.&lt;/li&gt;&#10;&lt;li&gt;Liu et al. (2023). &lt;em&gt;CAMEL: Communicative Agents for Mind Exploration of Large Scale Language Model Society&lt;/em&gt;.&lt;br&gt;&#10;→ Relevant to role assignment in multi-agent simulations.&lt;/li&gt;&#10;&lt;li&gt;Paranjape et al. (2023). &lt;em&gt;DSPy&lt;/em&gt;.&lt;br&gt;&#10;→ Relevant to modular language-model program design.&lt;/li&gt;&#10;&lt;li&gt;Wu et al. (2023). &lt;em&gt;AutoGen&lt;/em&gt;.&lt;br&gt;&#10;→ Foundation for prompt-based team construction and multi-role behavior.&lt;/li&gt;&#10;&lt;li&gt;Yao et al. (2024). &lt;em&gt;MARS: A Multi-Agent Framework Incorporating Socratic Guidance for Automated Prompt Optimization&lt;/em&gt;.&lt;br&gt;&#10;→ Citation needs verification before publication.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="chapter-9-designing-your-digital-lens"&gt;Chapter 9: Designing Your Digital Lens&lt;/h2&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Yao et al. (2022). &lt;em&gt;ReAct&lt;/em&gt;.&lt;br&gt;&#10;→ Relevant to task-contextual reasoning.&lt;/li&gt;&#10;&lt;li&gt;Paranjape et al. (2023). &lt;em&gt;DSPy&lt;/em&gt;.&lt;br&gt;&#10;→ Relevant to declarative interfaces; filtering claim needs verification.&lt;/li&gt;&#10;&lt;li&gt;Wu et al. (2023). &lt;em&gt;AutoGen&lt;/em&gt;.&lt;br&gt;&#10;→ Conversation-driven interface for lens behavior.&lt;/li&gt;&#10;&lt;li&gt;Schick et al. (2023). &lt;em&gt;Toolformer&lt;/em&gt;.&lt;br&gt;&#10;→ Relevant to tool-use learning; content-filtering application is an extrapolation.&lt;/li&gt;&#10;&lt;li&gt;Shinn et al. (2023). &lt;em&gt;Reflexion&lt;/em&gt;.&lt;br&gt;&#10;→ Relevant to feedback-driven revision loops.&lt;/li&gt;&#10;&lt;li&gt;Yao et al. (2024). &lt;em&gt;MCTS-RAG: Enhance Retrieval-Augmented Generation with Monte Carlo Tree Search&lt;/em&gt;.&lt;br&gt;&#10;→ Retrieval/planning relevance needs verification before publication.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The following references are candidates for Chapter 10&amp;rsquo;s discussion of freestyle cognition, research automation, and AI-assisted prototyping. Verify each title, author list, date, and relevance before final publication.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 06: Does One Agent Plan, Execute and Judge Its Own Work? Build a Planner-Executor-Critic Architecture</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/06-chapter/</link><pubDate>Sat, 08 Aug 2026 23:49:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/06-chapter/</guid><description>&lt;h1 id="does-one-agent-plan-execute-and-judge-its-own-work-build-a-planner-executor-critic-architecture"&gt;Does One Agent Plan, Execute and Judge Its Own Work? Build a Planner-Executor-Critic Architecture&lt;/h1&gt;&#10;&lt;p&gt;A single model can often do all of these things:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;understand a task,&lt;/li&gt;&#10;&lt;li&gt;decide what to do,&lt;/li&gt;&#10;&lt;li&gt;execute a tool call,&lt;/li&gt;&#10;&lt;li&gt;inspect the result,&lt;/li&gt;&#10;&lt;li&gt;critique its own work,&lt;/li&gt;&#10;&lt;li&gt;decide whether it succeeded,&lt;/li&gt;&#10;&lt;li&gt;and produce the final answer.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;That is convenient.&lt;/p&gt;&#10;&lt;p&gt;It is also a dangerous concentration of responsibilities.&lt;/p&gt;&#10;&lt;p&gt;If the same component creates the plan, executes it, explains why the result is good, and decides whether the job is complete, then failures become difficult to localize.&lt;/p&gt;</description></item><item><title>Agents From First Principles 08: AI Agent Picks the First Solution? Add Search Instead of One-Shot Generation</title><link>https://aibussin.com/post/agents-from-first-principles-08/</link><pubDate>Sat, 08 Aug 2026 17:26:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-08/</guid><description>&lt;p&gt;An AI agent often fails for a surprisingly ordinary reason:&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;it commits too early.&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;It finds one plausible next action, follows it, and then spends the rest of the run trying to make that first choice work.&lt;/p&gt;&#10;&lt;p&gt;That can look intelligent because the agent keeps reasoning, calling tools, revising plans, and explaining itself.&lt;/p&gt;&#10;&lt;p&gt;But underneath, the trajectory may be almost completely determined by an early mistake.&lt;/p&gt;&#10;&lt;p&gt;A coding agent chooses the wrong implementation strategy and spends twenty tool calls repairing it.&lt;/p&gt;</description></item><item><title>Agents From First Principles 07: AI Agent Forgets Previous Work? Add Working, Semantic and Episodic Memory</title><link>https://aibussin.com/post/agents-from-first-principles-07/</link><pubDate>Sat, 08 Aug 2026 17:14:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-07/</guid><description>&lt;h1 id="ai-agent-forgets-previous-work-add-working-semantic-and-episodic-memory"&gt;AI Agent Forgets Previous Work? Add Working, Semantic and Episodic Memory&lt;/h1&gt;&#10;&lt;p&gt;An agent can use the right model, call the right tools, execute the right plan, and still behave as if nothing that happened five minutes ago matters.&lt;/p&gt;&#10;&lt;p&gt;You see the symptoms quickly:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;it re-reads files it already inspected;&lt;/li&gt;&#10;&lt;li&gt;it repeats research it already completed;&lt;/li&gt;&#10;&lt;li&gt;it asks for information the user already supplied;&lt;/li&gt;&#10;&lt;li&gt;it forgets why a previous approach failed;&lt;/li&gt;&#10;&lt;li&gt;it loses decisions made earlier in a long task;&lt;/li&gt;&#10;&lt;li&gt;it treats every new run as if the system has never seen the problem before;&lt;/li&gt;&#10;&lt;li&gt;it retrieves an old answer and treats it as current truth;&lt;/li&gt;&#10;&lt;li&gt;it fills the prompt with so much history that the useful information is buried.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The usual response is:&lt;/p&gt;</description></item><item><title>Agents From First Principles 05: AI Agent Gets Stuck in a Loop? Add State, Feedback and Stopping Conditions</title><link>https://aibussin.com/post/agents-from-first-principles-05/</link><pubDate>Sat, 08 Aug 2026 16:56:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-05/</guid><description>&lt;p&gt;An AI agent that keeps calling the same tool, revisiting the same page, rewriting the same file, or repeatedly saying “I’ll try again” is not displaying persistence.&lt;/p&gt;&#10;&lt;p&gt;It is displaying a control-flow bug.&lt;/p&gt;&#10;&lt;p&gt;This is one of the most common failure modes in agent software because the basic loop is deceptively simple:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;observe&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;decide&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;act&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;observe&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;repeat&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The problem is hidden inside the final word.&lt;/p&gt;</description></item><item><title>Agents From First Principles 04: AI Agent Fails on Multi-Step Tasks? Separate Planning From Execution</title><link>https://aibussin.com/post/agents-from-first-principles-04/</link><pubDate>Sat, 08 Aug 2026 16:35:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-04/</guid><description>&lt;p&gt;A surprising number of agent failures are not really model failures.&lt;/p&gt;&#10;&lt;p&gt;The model may be perfectly capable of writing each individual step. The failure happens because the system tries to decide &lt;strong&gt;what to do&lt;/strong&gt; and &lt;strong&gt;do it&lt;/strong&gt; at the same time.&lt;/p&gt;&#10;&lt;p&gt;That works for simple tasks:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;question&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;answer&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;It becomes fragile when success depends on several ordered actions:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;goal&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;step 1&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;step 2&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;step 3&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;verification&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A useful next step in agent design is therefore to separate two jobs:&lt;/p&gt;</description></item><item><title>Agents From First Principles 03: AI Agent Keeps Making the Same Mistake? Add a Critique-and-Revision Loop</title><link>https://aibussin.com/post/agents-from-first-principles-03/</link><pubDate>Sat, 08 Aug 2026 16:23:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-03/</guid><description>&lt;p&gt;An AI agent can fail in a particularly frustrating way: it produces an answer that is almost right, you ask it to improve the answer, and it produces another answer with the same underlying defect.&lt;/p&gt;&#10;&lt;p&gt;Sometimes the wording changes. Sometimes it adds more explanation. Sometimes it becomes longer and more confident. But the important mistake survives.&lt;/p&gt;&#10;&lt;p&gt;That usually means the system is doing this:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;prompt&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;answer&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;or this:&lt;/p&gt;</description></item><item><title>Agents From First Principles 02: AI Agent Gives Inconsistent Answers? Generate Multiple Candidates and Rank Them</title><link>https://aibussin.com/post/agents-from-first-principles-02/</link><pubDate>Sat, 08 Aug 2026 15:54:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-02/</guid><description>&lt;p&gt;One of the first things you notice when you build anything around a large language model is that the same prompt does not always produce the same quality of answer.&lt;/p&gt;&#10;&lt;p&gt;Sometimes the first response is excellent.&lt;/p&gt;&#10;&lt;p&gt;Sometimes it is merely acceptable.&lt;/p&gt;&#10;&lt;p&gt;Sometimes it misses the point entirely.&lt;/p&gt;&#10;&lt;p&gt;That creates a very common agent-engineering question:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;If the model is inconsistent, should the agent trust the first answer it gets?&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;Often, no.&lt;/p&gt;</description></item><item><title>Agents From First Principles 00: What Is an Agent, Really?</title><link>https://aibussin.com/post/agents-from-first-principles-00/</link><pubDate>Sat, 08 Aug 2026 15:40:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-00/</guid><description>&lt;h1 id="what-is-an-agent-really"&gt;What Is an Agent, Really?&lt;/h1&gt;&#10;&lt;p&gt;This is the first post in &lt;strong&gt;Agents From First Principles&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;It follows two earlier series.&lt;/p&gt;&#10;&lt;p&gt;In &lt;strong&gt;PyTorch: Zero to Hero&lt;/strong&gt;, we worked upward from tensors, autograd and neural-network building blocks until we could build a small language model ourselves.&lt;/p&gt;&#10;&lt;p&gt;In &lt;strong&gt;Models From First Principles&lt;/strong&gt;, we moved one level higher. We looked at how learned components can be composed into scorers, value models, policy heads, recurrent models, hierarchical models and compact recursive systems.&lt;/p&gt;</description></item><item><title>Intelligence Through Execution: The Executable Cognitive Kernel</title><link>https://aibussin.com/post/eck/</link><pubDate>Tue, 10 Mar 2026 21:58:14 +0000</pubDate><guid>https://aibussin.com/post/eck/</guid><description>&lt;h2 id="-summary"&gt;🧭 Summary&lt;/h2&gt;&#10;&lt;p&gt;Most modern AI systems treat intelligence as something stored inside a model.&lt;/p&gt;&#10;&lt;p&gt;A neural network is trained on massive datasets, its weights are adjusted, and those weights become the system’s knowledge. When the model produces an output, we interpret that output as the result of the intelligence encoded inside those parameters.&lt;/p&gt;&#10;&lt;p&gt;But this perspective has a limitation.&lt;/p&gt;&#10;&lt;p&gt;Once training is complete, the model is largely static. It does not improve through its own actions, and it does not adapt based on the outcome of its behavior unless we retrain it.&lt;/p&gt;</description></item></channel></rss>