<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Evaluation on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/evaluation/</link><description>Recent content in Evaluation on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Fri, 25 Sep 2026 02:45:00 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/evaluation/index.xml" rel="self" type="application/rss+xml"/><item><title>The Measurement Instrument</title><link>https://aibussin.com/books/memory/02-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:02 +0100</pubDate><guid>https://aibussin.com/books/memory/02-chapter/</guid><description>&lt;p&gt;Chapter 1 ended with a test. Does the past change what the system does now, and does it change it for the better?&lt;/p&gt;&#10;&lt;p&gt;That is easy to state and hard to run. This chapter builds the thing that runs it.&lt;/p&gt;&#10;&lt;p&gt;Two words will recur for the rest of the book, and they must not blur together. The &lt;strong&gt;memory system&lt;/strong&gt; is the thing that remembers. The &lt;strong&gt;instrument&lt;/strong&gt; is the apparatus that decides whether it does. They are separate products, and this book builds both.&lt;/p&gt;</description></item><item><title>Meat Proxy</title><link>https://aibussin.com/books/applied-ai/05-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:05 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/05-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 1 — Where You Stand&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Review degrades because the system succeeds&lt;/strong&gt;&lt;/p&gt;&#10;&lt;h2 id="ten-weeks"&gt;Ten weeks&lt;/h2&gt;&#10;&lt;p&gt;Week one, the model drafts a paragraph review and you read every line of it. You disagree with two of its four flags, check a source yourself, rewrite one sentence. The whole thing takes twenty minutes and you feel you have earned the output.&lt;/p&gt;&#10;&lt;p&gt;Week three, the reviews have been good. You read them properly but you stop re-checking the sources it says it checked. Nothing bad happens.&lt;/p&gt;</description></item><item><title>The Memory Nexus</title><link>https://aibussin.com/books/memory/06-chapter/</link><pubDate>Sat, 19 Sep 2026 05:00:06 +0100</pubDate><guid>https://aibussin.com/books/memory/06-chapter/</guid><description>&lt;p&gt;Chapters 3 through 5 leave the system with an embarrassment of options. Chapter 3 built hybrid retrieval over raw history with reranking. Chapter 4 added a persistent derived graph with several query modes. Chapter 5 added cue-conditioned associative propagation over that graph. Each chapter earned its mechanism conditionally, and each left the cheaper layers available underneath. The question none of them answers is the one a deployed system meets first:&lt;/p&gt;</description></item><item><title>Intelligence in the Wrong Direction</title><link>https://aibussin.com/books/applied-ai/07-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:07 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/07-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 1 — Where You Stand&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Capability without an objective is magnitude without direction&lt;/strong&gt;&lt;/p&gt;&#10;&lt;h2 id="the-thing-the-marketing-does-not-cover"&gt;The thing the marketing does not cover&lt;/h2&gt;&#10;&lt;p&gt;Every vendor will sell you more capability. None of them can tell you whether that capability moves your process toward its objective without a measurement you supply.&lt;/p&gt;&#10;&lt;p&gt;A stronger model can execute the wrong objective more effectively just as it can execute the right one more effectively. Capability alone does not tell you whether the gap between what you wanted and what you got narrowed or widened. That is an empirical question.&lt;/p&gt;</description></item><item><title>Examples Are Experimental Data</title><link>https://aibussin.com/books/dspy-from-first-principles/07-chapter/</link><pubDate>Fri, 28 Aug 2026 10:30:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/07-chapter/</guid><description>&lt;p&gt;Chapter 6 gave the program a dependency boundary and, more importantly, evidence that identical recorded program state can still produce different outputs. We can now say what the program is, what ran it, and why a small score delta cannot be interpreted without repeated measurement.&lt;/p&gt;&#10;&lt;p&gt;What we still cannot do is say whether the program is any good, because we have nothing to measure it against.&lt;/p&gt;&#10;&lt;p&gt;Examples in DSPy are not decoration around the program. They are experimental data with several possible roles. Optimization consumes that evidence: training data may shape candidate state, development data may select among candidates, and holdout data must remain outside both processes if it is to support an independent claim.&lt;/p&gt;</description></item><item><title>You Cannot Optimize What You Cannot Measure</title><link>https://aibussin.com/books/dspy-from-first-principles/08-chapter/</link><pubDate>Fri, 28 Aug 2026 10:35:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/08-chapter/</guid><description>&lt;p&gt;Chapter 7 produced material. This chapter produces evidence.&lt;/p&gt;&#10;&lt;p&gt;That makes optimization possible. It does not yet make optimization safe.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;no metric&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; → no systematic optimization&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;metric&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; → behavior can be optimized&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;wrong metric&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; → the wrong behavior can be optimized systematically&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This chapter establishes the surface on which search can operate. Chapter 9 asks what that surface rewards when the metric becomes the target.&lt;/p&gt;&#10;&lt;pre class="mermaid"&gt;&#10; flowchart TD&#10; P[frozen program] --&amp;gt; E[evaluate]&#10; C[frozen cases] --&amp;gt; E&#10; M[metric] --&amp;gt; E&#10; E --&amp;gt; PC[per-case results]&#10; PC --&amp;gt; AG[aggregate result]&#10; PC --&amp;gt; FI[failure inspection]&#10; &lt;/pre&gt;&#10; &lt;p&gt;No optimization happens here. Note that the per-case results feed &lt;em&gt;two&lt;/em&gt; readings — the aggregate and the failure inspection — and this chapter&amp;rsquo;s finding comes from the second. The question is narrow:&lt;/p&gt;</description></item><item><title>What Survived the Transformation?</title><link>https://aibussin.com/books/language/09-chapter/</link><pubDate>Fri, 25 Sep 2026 02:45:00 +0100</pubDate><guid>https://aibussin.com/books/language/09-chapter/</guid><description>&lt;p&gt;Part I has spent five chapters transforming information. Visual previews compress pages into triage cards. The router adds diagrams, tables, and timelines — or declines. The semantic record stabilises structure. The sidecar brings in surrounding material. Relation readout classifies how that material bears on the source. Every one of these operations can fail silently: the output looks fluent, looks professional, looks &lt;em&gt;right&lt;/em&gt;, while the content has shifted underneath. A diagram drops the exception. A summary upgrades a correlation to a cause. A sidecar panel restates the source&amp;rsquo;s speculative conclusion as its headline finding. Nothing in the rendering announces the damage.&lt;/p&gt;</description></item><item><title>When the Metric Becomes the Target</title><link>https://aibussin.com/books/dspy-from-first-principles/09-chapter/</link><pubDate>Fri, 28 Aug 2026 10:40:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/09-chapter/</guid><description>&lt;p&gt;Chapter 8 built a metric and found a defect in it by reading a table. Six cases where the correct answer is capped at 0.65, discovered by grouping per-case scores and noticing that a whole category sat on an exact number.&lt;/p&gt;&#10;&lt;p&gt;That was luck dressed as diligence. This chapter looks on purpose.&lt;/p&gt;&#10;&lt;p&gt;The urgency comes from what happens next. A DSPy optimizer does not want your program to be good. It wants your number to go up, and it will find every route to that outcome, including the ones you did not intend to leave open.&lt;/p&gt;</description></item><item><title>A Revolver, Not a Foundation</title><link>https://aibussin.com/books/applied-ai/10-chapter/</link><pubDate>Mon, 14 Sep 2026 05:00:10 +0100</pubDate><guid>https://aibussin.com/books/applied-ai/10-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part 2 — Get the Model Out of the Chat Box&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;The model will change; the frame should survive it&lt;/strong&gt;&lt;/p&gt;&#10;&lt;h2 id="the-book-you-could-write-better-next-year"&gt;The book you could write better next year&lt;/h2&gt;&#10;&lt;p&gt;Write a book with AI this year and you will be able to write the same book better next year. And better again the year after. The same is true of the code, the research synthesis, the review process, the classifier — anything where a model does part of the work.&lt;/p&gt;</description></item><item><title>The Nearest Neighbor Can Be Wrong</title><link>https://aibussin.com/books/embeddings-from-first-principles/10-chapter/</link><pubDate>Mon, 07 Sep 2026 11:10:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/10-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part III — Retrieval Is an Experiment&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="the-result-that-is-closest-and-wrong"&gt;The result that is closest and wrong&lt;/h2&gt;&#10;&lt;p&gt;&lt;em&gt;Illustrative example — not a RELATE measurement.&lt;/em&gt;&lt;/p&gt;&#10;&lt;p&gt;Query:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Is Dublin the capital of Ireland?&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Imagine the passage that ranks first is this one:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;Dublin is not, and has never been, the capital of Ireland — that&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;distinction belongs to the older seat of government at Tara.&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Fluent, on-topic, confidently phrased, and among the geometrically closest things in the corpus — and false. A model handed this passage as context may repeat its claim. The retrieval system did its job: it found a near vector. &amp;ldquo;Near&amp;rdquo; was not &amp;ldquo;correct.&amp;rdquo;&lt;/p&gt;</description></item><item><title>Optimize the Instructions</title><link>https://aibussin.com/books/dspy-from-first-principles/12-chapter/</link><pubDate>Fri, 28 Aug 2026 10:55:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/12-chapter/</guid><description>&lt;p&gt;Chapter 11&amp;rsquo;s optimizer could only select demonstrations. It never touched an instruction, and it never looked at the development set — which is why upgrading its metric changed nothing.&lt;/p&gt;&#10;&lt;p&gt;MIPROv2 does both.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;instruction candidates&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;+&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;demonstration candidates&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;+&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;metric&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;+&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;development data&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;candidate program&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;This is a strictly larger search over a strictly larger space, evaluated against evidence the previous optimizer could not see. If chapter 11&amp;rsquo;s result was an artifact of a weak mechanism, this is where it gets corrected.&lt;/p&gt;</description></item><item><title>A Resolved Promise Is Not a Correct Answer</title><link>https://aibussin.com/books/browser-ai-from-first-principles/13-chapter/</link><pubDate>Wed, 02 Sep 2026 16:00:00 +0100</pubDate><guid>https://aibussin.com/books/browser-ai-from-first-principles/13-chapter/</guid><description>&lt;p&gt;The following code has only one definition of success:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-javascript" data-lang="javascript"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&lt;span style="color:#66d9ef"&gt;try&lt;/span&gt; {&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#66d9ef"&gt;const&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;output&lt;/span&gt; &lt;span style="color:#f92672"&gt;=&lt;/span&gt; &lt;span style="color:#66d9ef"&gt;await&lt;/span&gt; &lt;span style="color:#a6e22e"&gt;session&lt;/span&gt;.&lt;span style="color:#a6e22e"&gt;prompt&lt;/span&gt;(&lt;span style="color:#a6e22e"&gt;input&lt;/span&gt;);&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#a6e22e"&gt;show&lt;/span&gt;(&lt;span style="color:#a6e22e"&gt;output&lt;/span&gt;);&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;} &lt;span style="color:#66d9ef"&gt;catch&lt;/span&gt; (&lt;span style="color:#a6e22e"&gt;error&lt;/span&gt;) {&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; &lt;span style="color:#a6e22e"&gt;showFailure&lt;/span&gt;(&lt;span style="color:#a6e22e"&gt;error&lt;/span&gt;);&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;}&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;If the promise resolves, the interface displays the result.&lt;/p&gt;&#10;&lt;p&gt;That is sufficient for a deterministic lookup whose return value is authoritative. It is not sufficient for a probabilistic system that can return fluent, well-formed and incorrect text.&lt;/p&gt;&#10;&lt;p&gt;A resolved promise proves that an operation completed according to the runtime contract. It does not prove that the feature did its job.&lt;/p&gt;</description></item><item><title>Evaluate the Feature, Not the Demo</title><link>https://aibussin.com/books/browser-ai-from-first-principles/14-chapter/</link><pubDate>Wed, 02 Sep 2026 16:15:00 +0100</pubDate><guid>https://aibussin.com/books/browser-ai-from-first-principles/14-chapter/</guid><description>&lt;p&gt;A demo asks whether one carefully chosen example can work.&lt;/p&gt;&#10;&lt;p&gt;A feature evaluation asks how often the product contract holds across the inputs, states and failures users will actually encounter.&lt;/p&gt;&#10;&lt;p&gt;Both have value. Confusing them is the problem.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="1-the-unit-of-evaluation-is-the-feature"&gt;1. The unit of evaluation is the feature&lt;/h2&gt;&#10;&lt;p&gt;“Evaluate Gemma 4” is too broad for our application. “Evaluate this prompt” is too narrow.&lt;/p&gt;&#10;&lt;p&gt;The useful unit is something a user attempts:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;summarize selected documentation&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;rewrite permission copy without changing authority&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;detect language and decide whether to translate&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;produce structured arguments for a read-only tool&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Each feature includes input preparation, capability selection, model output, validation, fallback and UI behavior.&lt;/p&gt;</description></item><item><title>Appendix: The Evidence Ledger</title><link>https://aibussin.com/books/hallucination-from-first-principles/16-appendix/</link><pubDate>Mon, 31 Aug 2026 09:00:00 +0100</pubDate><guid>https://aibussin.com/books/hallucination-from-first-principles/16-appendix/</guid><description>&lt;p&gt;The book argues that a claim should be governed by its provenance, its evidential support, and an explicit statement of what it does not establish. This appendix applies that discipline to the book itself.&lt;/p&gt;&#10;&lt;p&gt;Every original empirical number produced for this book is listed here with three things: where it was measured, what its provenance status is, and what it does not establish.&lt;/p&gt;&#10;&lt;p&gt;Two facts are visible immediately from the ledger:&lt;/p&gt;</description></item><item><title>Don't Let the Optimizer Cheat</title><link>https://aibussin.com/books/dspy-from-first-principles/18-chapter/</link><pubDate>Fri, 28 Aug 2026 11:15:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/18-chapter/</guid><description>&lt;h2 id="opening--the-optimizer-doesnt-have-to-cheat-intentionally"&gt;Opening — The Optimizer Doesn&amp;rsquo;t Have to Cheat Intentionally&lt;/h2&gt;&#10;&lt;p&gt;In Chapters 15 and 16 we expanded the program&amp;rsquo;s access. The agent could search repositories, retrieve memories, and adaptively select evidence. Chapter 17 then added search over reasoning states, showing that tree search can allocate inference-time compute toward promising paths.&lt;/p&gt;&#10;&lt;p&gt;Now the question becomes urgent: once a program can search, remember, and adaptively select both evidence and reasoning, how do we prove that it did not obtain information it was never supposed to see?&lt;/p&gt;</description></item><item><title>Tool Choice Is a Behavioral Problem</title><link>https://aibussin.com/books/browser-ai-from-first-principles/19-chapter/</link><pubDate>Wed, 02 Sep 2026 17:30:00 +0100</pubDate><guid>https://aibussin.com/books/browser-ai-from-first-principles/19-chapter/</guid><description>&lt;p&gt;Chapter 18 built a tool that is correct by construction: its contract is tested, its results are bounded, its lifecycle is traced. None of that determines whether an agent uses it well.&lt;/p&gt;&#10;&lt;p&gt;Once a model can see several capabilities, the outcome depends on a behavioral pipeline:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;discovery → selection → argument construction → admission → execution → result use&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A single &amp;ldquo;task passed&amp;rdquo; score cannot tell you which stage failed. Chapter 13 made this argument for model output: a resolved promise is not a correct answer. The same layering applies to tool use, with more stages and higher stakes.&lt;/p&gt;</description></item><item><title>Did the Bridge Preserve the Space?</title><link>https://aibussin.com/books/embeddings-from-first-principles/21-chapter/</link><pubDate>Mon, 07 Sep 2026 14:30:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/21-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VI — Crossing Embedding Spaces · What did each measurement actually certify?&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="one-bridge-eight-measurements"&gt;One bridge, eight measurements&lt;/h2&gt;&#10;&lt;p&gt;Chapter 20 built the machinery that turns preservation evidence into a scoped authorization — a requirement, a measured value, and a decision, all cited together. What it deliberately did not do is ask how much any one of those measured values is actually entitled to say. That is this chapter&amp;rsquo;s entire job.&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;A preservation metric is a sensor. Before trusting its reading, ask which failure modes the sensor is even capable of seeing.&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Can a Smaller Representation Preserve a Larger One?</title><link>https://aibussin.com/books/embeddings-from-first-principles/24-chapter/</link><pubDate>Mon, 07 Sep 2026 15:00:00 +0000</pubDate><guid>https://aibussin.com/books/embeddings-from-first-principles/24-chapter/</guid><description>&lt;p&gt;&lt;em&gt;Part VII — What Survives Transformation&lt;/em&gt;&lt;/p&gt;&#10;&lt;h2 id="two-vectors-for-one-document"&gt;Two vectors for one document&lt;/h2&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;full_doc → E(full_doc) one vector, dimension d&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;summary(doc) → E(summary(doc)) one vector, dimension d&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Both vectors have exactly the same dimension. What became smaller is the &lt;strong&gt;text&lt;/strong&gt; — the number of words available to encode — not the vector. One embedding was produced from a full document; the other from a much shorter compression of it, passed through the identical embedding pipeline. This is not the dimensionality reduction Chapter 7 covered, and it is not the cross-space translation Part VI built — both the full document and its compression are embedded natively, by the same encoder, into the same space. The question is:&lt;/p&gt;</description></item><item><title>Did the Context Help?</title><link>https://aibussin.com/books/context/26-chapter/</link><pubDate>Wed, 23 Sep 2026 05:00:23 +0100</pubDate><guid>https://aibussin.com/books/context/26-chapter/</guid><description>&lt;p&gt;Two prerequisites have been met, on small cases. The compiler builds a legal bundle and can show its working (Chapters 23 and 24). The bundle can be delivered to a running model and independently seen to have arrived (Chapter 25). Neither says anything about the model.&lt;/p&gt;&#10;&lt;p&gt;This chapter asks a different question. Given that the intended context was constructed and delivered, did it improve what the model did?&lt;/p&gt;&#10;&lt;p&gt;The book now has three kinds of evidence, and this is the third.&lt;/p&gt;</description></item><item><title>How Much of You Can an AI Reproduce?</title><link>https://aibussin.com/books/language/30-chapter/</link><pubDate>Fri, 25 Sep 2026 02:45:00 +0100</pubDate><guid>https://aibussin.com/books/language/30-chapter/</guid><description>&lt;p&gt;Part III taught a system to adapt to a person without reproducing them. Now the boundary gets pushed deliberately: given enough evidence — interviews, policies, artifacts, interaction history — how much observable behaviour can an AI reproduce, and how is that claim measured? Prediction, not identity. This chapter measures the first and refuses the second so completely that Chapter 31 inherits a clean, hungry question.&lt;/p&gt;&#10;&lt;p&gt;The first discipline is vocabulary. &amp;ldquo;Reproduce&amp;rdquo; must not float between style, preferences, values, decisions, and identity. The chapter fixes a &lt;strong&gt;reproduction surface&lt;/strong&gt; of observable targets — language/expression, declared preferences, factual self-knowledge, value judgements, task choices, scenario decisions, multi-step behaviour — governed by one rule:&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 23: Your Infrastructure Is Healthy. Why Is the Agent Getting Worse? Detect Behavioral Drift and Roll Back Safely</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/23-chapter/</link><pubDate>Sun, 09 Aug 2026 12:00:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/23-chapter/</guid><description>&lt;h1 id="your-infrastructure-is-healthy-why-is-the-agent-getting-worse"&gt;Your Infrastructure Is Healthy. Why Is the Agent Getting Worse?&lt;/h1&gt;&#10;&lt;p&gt;Your dashboards are green.&lt;/p&gt;&#10;&lt;p&gt;The model endpoint is responding.&lt;/p&gt;&#10;&lt;p&gt;The browser workers are alive.&lt;/p&gt;&#10;&lt;p&gt;The database is healthy.&lt;/p&gt;&#10;&lt;p&gt;The queue is draining.&lt;/p&gt;&#10;&lt;p&gt;The verifier service is up.&lt;/p&gt;&#10;&lt;p&gt;Latency has not exploded.&lt;/p&gt;&#10;&lt;p&gt;There are no obvious exceptions.&lt;/p&gt;&#10;&lt;p&gt;And yet the agent is getting worse.&lt;/p&gt;&#10;&lt;p&gt;It fixes fewer bugs.&lt;/p&gt;&#10;&lt;p&gt;It retrieves weaker evidence.&lt;/p&gt;&#10;&lt;p&gt;It escalates to expensive models more often.&lt;/p&gt;&#10;&lt;p&gt;It chooses the wrong tools more frequently.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 17: What Is Your Agent Actually Uncertain About?</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/17-chapter/</link><pubDate>Sun, 09 Aug 2026 11:02:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/17-chapter/</guid><description>&lt;h1 id="what-is-your-agent-actually-uncertain-about"&gt;What Is Your Agent Actually Uncertain About?&lt;/h1&gt;&#10;&lt;p&gt;An agent reaches a difficult point in a task.&lt;/p&gt;&#10;&lt;p&gt;It is not sure what to do next.&lt;/p&gt;&#10;&lt;p&gt;A common implementation responds like this:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;uncertain&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;call the model again&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;still uncertain&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;call a stronger model&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;still uncertain&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;search more&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That is not a reasoning strategy.&lt;/p&gt;&#10;&lt;p&gt;It is a spending strategy.&lt;/p&gt;&#10;&lt;p&gt;The system is using more computation without identifying what information is actually missing.&lt;/p&gt;</description></item><item><title>Advanced Agents From First Principles 14: Can Your Agent Learn From Its Own Trajectories Without Learning the Wrong Lessons?</title><link>https://aibussin.com/books/advanced-agents-from-first-principles/14-chapter/</link><pubDate>Sun, 09 Aug 2026 10:33:00 +0100</pubDate><guid>https://aibussin.com/books/advanced-agents-from-first-principles/14-chapter/</guid><description>&lt;p&gt;An advanced agent now leaves behind something extremely valuable:&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;evidence.&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;Not merely chat history.&lt;/p&gt;&#10;&lt;p&gt;Not merely model outputs.&lt;/p&gt;&#10;&lt;p&gt;Not merely traces.&lt;/p&gt;&#10;&lt;p&gt;A sufficiently instrumented system can record:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;what state it was in,&lt;/li&gt;&#10;&lt;li&gt;what alternatives it considered,&lt;/li&gt;&#10;&lt;li&gt;which route it selected,&lt;/li&gt;&#10;&lt;li&gt;what branches it pruned,&lt;/li&gt;&#10;&lt;li&gt;which model or specialist it escalated to,&lt;/li&gt;&#10;&lt;li&gt;which tools it called,&lt;/li&gt;&#10;&lt;li&gt;which critic changed the answer,&lt;/li&gt;&#10;&lt;li&gt;what verification evidence was produced,&lt;/li&gt;&#10;&lt;li&gt;how much compute was spent,&lt;/li&gt;&#10;&lt;li&gt;and whether the final result actually passed.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;That immediately suggests a tempting idea:&lt;/p&gt;</description></item><item><title>Agents From First Principles 09: AI Agent Says It Worked When It Didn’t? Verify the Result Outside the LLM</title><link>https://aibussin.com/post/agents-from-first-principles-09/</link><pubDate>Sat, 08 Aug 2026 17:31:00 +0100</pubDate><guid>https://aibussin.com/post/agents-from-first-principles-09/</guid><description>&lt;p&gt;An AI agent says:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Done. The task is complete.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;That sentence is almost worthless.&lt;/p&gt;&#10;&lt;p&gt;The agent may have:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;edited the wrong file,&lt;/li&gt;&#10;&lt;li&gt;changed the right file incorrectly,&lt;/li&gt;&#10;&lt;li&gt;skipped part of the request,&lt;/li&gt;&#10;&lt;li&gt;broken another subsystem,&lt;/li&gt;&#10;&lt;li&gt;failed to save its work,&lt;/li&gt;&#10;&lt;li&gt;misread a tool result,&lt;/li&gt;&#10;&lt;li&gt;passed a stale test,&lt;/li&gt;&#10;&lt;li&gt;inspected the wrong environment,&lt;/li&gt;&#10;&lt;li&gt;or simply decided that its own answer looked convincing.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;The central problem is simple:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;The system that produced the answer should not be the only system deciding whether the answer is correct.&lt;/strong&gt;&lt;/p&gt;</description></item></channel></rss>