<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Agent Evaluation on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/agent-evaluation/</link><description>Recent content in Agent Evaluation on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Sun, 09 Aug 2026 10:10:00 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/agent-evaluation/index.xml" rel="self" type="application/rss+xml"/><item><title>Candidate Generation and Selection</title><link>https://aibussin.com/books/agents-from-first-principles/03-chapter/</link><pubDate>Sat, 08 Aug 2026 15:54:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/03-chapter/</guid><description>&lt;p&gt;The previous chapter built a boundary that stops arbitrary model output from acquiring execution authority without explicit checks. It made one class of failure inspectable and enforceable, and it is silent about another.&lt;/p&gt;&#10;&lt;p&gt;Suppose the model is asked to solve a coding problem. One run produces the right patch. The next produces a plausible but incomplete one. A third produces something better again. Nothing is malformed, nothing violates the action schema, and every one of them would pass the boundary we just built. Validity and quality are different properties. A proposal can be completely valid and still be a poor choice, which relocates the uncertainty rather than removing it:&lt;/p&gt;</description></item><item><title>Evidence and Verification</title><link>https://aibussin.com/books/agents-from-first-principles/10-chapter/</link><pubDate>Sat, 08 Aug 2026 17:31:00 +0100</pubDate><guid>https://aibussin.com/books/agents-from-first-principles/10-chapter/</guid><description>&lt;p&gt;The agent says:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Done.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;That is a claim.&lt;/p&gt;&#10;&lt;p&gt;It is not evidence.&lt;/p&gt;&#10;&lt;p&gt;Every mechanism in this book so far has made the agent better at deciding what to do, and none of them establishes that the user&amp;rsquo;s goal was achieved. A planner can produce a coherent plan for the wrong problem. A tool can return exit code zero without producing the intended effect. A search can select the highest-scoring branch when every branch is wrong. A memory system can retrieve a perfectly relevant fact that stopped being true in March.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 12: Is Your Advanced Agent Actually Better? Benchmark It Under Equal Budgets</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/12-chapter/</link><pubDate>Sun, 09 Aug 2026 10:10:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/12-chapter/</guid><description>&lt;h1 id="is-your-advanced-agent-actually-better-benchmark-it-under-equal-budgets"&gt;Is Your Advanced Agent Actually Better? Benchmark It Under Equal Budgets&lt;/h1&gt;&#10;&lt;p&gt;You replace one model call with eight.&lt;/p&gt;&#10;&lt;p&gt;Success rises from 62% to 74%.&lt;/p&gt;&#10;&lt;p&gt;Great.&lt;/p&gt;&#10;&lt;p&gt;Except the new system used:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;eight times the inference,&lt;/li&gt;&#10;&lt;li&gt;three extra judges,&lt;/li&gt;&#10;&lt;li&gt;two rounds of critique,&lt;/li&gt;&#10;&lt;li&gt;a larger context,&lt;/li&gt;&#10;&lt;li&gt;a stronger verifier,&lt;/li&gt;&#10;&lt;li&gt;and several times the latency.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Did the architecture improve?&lt;/p&gt;&#10;&lt;p&gt;Or did you just buy more attempts?&lt;/p&gt;&#10;&lt;p&gt;This is one of the easiest mistakes to make in advanced agent engineering.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 11: Which Advanced Agent Architecture Should You Use? A Practical Selection Guide</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/11-chapter/</link><pubDate>Sun, 09 Aug 2026 09:20:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/11-chapter/</guid><description>&lt;h1 id="which-advanced-agent-architecture-should-you-use"&gt;Which Advanced Agent Architecture Should You Use?&lt;/h1&gt;&#10;&lt;p&gt;You now have too many options.&lt;/p&gt;&#10;&lt;p&gt;That is a better problem than having none.&lt;/p&gt;&#10;&lt;p&gt;But it is still a problem.&lt;/p&gt;&#10;&lt;p&gt;You can add:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;self-consistency,&lt;/li&gt;&#10;&lt;li&gt;Tree of Thoughts,&lt;/li&gt;&#10;&lt;li&gt;beam search,&lt;/li&gt;&#10;&lt;li&gt;Monte Carlo Tree Search,&lt;/li&gt;&#10;&lt;li&gt;evolutionary search,&lt;/li&gt;&#10;&lt;li&gt;specialist routing,&lt;/li&gt;&#10;&lt;li&gt;planner/executor/critic separation,&lt;/li&gt;&#10;&lt;li&gt;multi-agent debate,&lt;/li&gt;&#10;&lt;li&gt;adaptive policies,&lt;/li&gt;&#10;&lt;li&gt;learning from previous runs,&lt;/li&gt;&#10;&lt;li&gt;or a mixture-of-agents runtime that chooses among several of them.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The temptation is to combine everything.&lt;/p&gt;</description></item><item><title>Agents From First Principles 02: AI Agent Gives Inconsistent Answers? Generate Multiple Candidates and Rank Them</title><link>https://aibussin.com/post/agents-from-first-principles-02/</link><pubDate>Sat, 08 Aug 2026 15:54:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-02/</guid><description>&lt;p&gt;One of the first things you notice when you build anything around a large language model is that the same prompt does not always produce the same quality of answer.&lt;/p&gt;&#10;&lt;p&gt;Sometimes the first response is excellent.&lt;/p&gt;&#10;&lt;p&gt;Sometimes it is merely acceptable.&lt;/p&gt;&#10;&lt;p&gt;Sometimes it misses the point entirely.&lt;/p&gt;&#10;&lt;p&gt;That creates a very common agent-engineering question:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;If the model is inconsistent, should the agent trust the first answer it gets?&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;Often, no.&lt;/p&gt;</description></item></channel></rss>