<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>LLM Evaluation on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/llm-evaluation/</link><description>Recent content in LLM Evaluation on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Sun, 30 Aug 2026 10:42:00 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/llm-evaluation/index.xml" rel="self" type="application/rss+xml"/><item><title>Why Are We Still Hand-Writing Prompts?</title><link>https://aibussin.com/books/dspy-from-first-principles/01-chapter/</link><pubDate>Fri, 28 Aug 2026 10:00:00 +0100</pubDate><guid>https://aibussin.com/books/dspy-from-first-principles/01-chapter/</guid><description>&lt;p&gt;Almost every language-model application starts as one string.&lt;/p&gt;&#10;&lt;p&gt;That is not a mistake. A prompt is the fastest way to find out whether a model can do the job at all, and it lets you work directly on the behavior instead of building a framework around a problem you do not yet understand. Most good LM systems begin this way and should.&lt;/p&gt;&#10;&lt;p&gt;The trouble is that the prompt usually survives longer than its usefulness. It stops being a probe and becomes the specification, and by the time anyone notices, the string is four hundred words long, nobody remembers why the third paragraph is there, and changing it feels dangerous.&lt;/p&gt;</description></item><item><title>How Do You Measure a Hallucination?</title><link>https://aibussin.com/books/hallucination-from-first-principles/04-chapter/</link><pubDate>Sat, 29 Aug 2026 23:14:00 +0100</pubDate><guid>https://aibussin.com/books/hallucination-from-first-principles/04-chapter/</guid><description>&lt;p&gt;The first three chapters deliberately avoided building a detector.&lt;/p&gt;&#10;&lt;p&gt;Before designing one, we needed to define exactly what it would be expected to detect.&lt;/p&gt;&#10;&lt;p&gt;Chapter 1 established that fluent generation can continue after evidential support has weakened or disappeared.&lt;/p&gt;&#10;&lt;p&gt;Chapter 2 showed that &lt;em&gt;hallucination&lt;/em&gt; covers several different failure relationships.&lt;/p&gt;&#10;&lt;p&gt;Chapter 3 then separated the objects a reliability system must not collapse:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;truth&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;≠&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;evidence&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;≠&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;support&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;≠&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;attribution&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;≠&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;provenance&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;≠&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;verification&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;≠&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;policy acceptance&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Now we can finally ask the engineering question:&lt;/p&gt;</description></item><item><title>Beyond Hallucination: Consistency and Sensitivity</title><link>https://aibussin.com/books/hallucination-from-first-principles/09-chapter/</link><pubDate>Sun, 30 Aug 2026 10:16:00 +0100</pubDate><guid>https://aibussin.com/books/hallucination-from-first-principles/09-chapter/</guid><description>&lt;p&gt;Chapter 8 ended with a rule:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;If a downstream decision depends on a distinction, do not discard that distinction before the decision is made.&lt;/strong&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;That rule breaks the idea of one universal hallucination score.&lt;/p&gt;&#10;&lt;p&gt;A response can be well contained and still reverse a relation. It can preserve every relation and still ignore the decisive facts of the problem. It can be correct once and unstable under a harmless rephrasing. It can be perfectly repeatable and consistently wrong.&lt;/p&gt;</description></item><item><title>The Safe but Useless Model</title><link>https://aibussin.com/books/hallucination-from-first-principles/10-chapter/</link><pubDate>Sun, 30 Aug 2026 10:42:00 +0100</pubDate><guid>https://aibussin.com/books/hallucination-from-first-principles/10-chapter/</guid><description>&lt;p&gt;Chapter 9 changed the unit of evaluation.&lt;/p&gt;&#10;&lt;p&gt;Instead of asking whether one answer looks good, we began studying a &lt;strong&gt;family of related executions&lt;/strong&gt;.&lt;/p&gt;&#10;&lt;p&gt;That immediately reveals a failure that static evaluation can miss almost completely.&lt;/p&gt;&#10;&lt;p&gt;Consider two organizations.&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ORGANIZATION A&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;3 months of runway&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;falling demand&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;negative cash flow&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;credit line nearly exhausted&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;ORGANIZATION B&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;5 years of runway&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;rapidly growing demand&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;strong margins&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;large cash reserve&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Ask both:&lt;/p&gt;</description></item><item><title>Beyond Hallucination Energy: A Three-Dimensional Framework for Reliable AI Outputs</title><link>https://aibussin.com/post/trendslop/</link><pubDate>Wed, 22 Apr 2026 10:35:46 +0100</pubDate><guid>https://aibussin.com/post/trendslop/</guid><description>&lt;h2 id="-1--tldr"&gt;🧩 1. TLDR&lt;/h2&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;AI doesn&amp;rsquo;t just hallucinate.&#10;Sometimes it gives answers that are fluent, safe… and completely useless.&lt;/strong&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;Most discussions about AI failure focus on hallucination:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;making things up&lt;/li&gt;&#10;&lt;li&gt;getting facts wrong&lt;/li&gt;&#10;&lt;li&gt;fabricating sources&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;That&amp;rsquo;s real. It matters.&lt;/p&gt;&#10;&lt;p&gt;But it&amp;rsquo;s not the most dangerous failure mode in production systems.&lt;/p&gt;&#10;&lt;p&gt;There is a quieter one.&lt;/p&gt;&#10;&lt;p&gt;A more subtle one.&lt;/p&gt;&#10;&lt;p&gt;And in practice a more &lt;em&gt;pervasive&lt;/em&gt; one.&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;AI systems often fail not by being wrong,&#10;but by failing to think at all.&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Thoughts of Algorithms</title><link>https://aibussin.com/post/thoughts/</link><pubDate>Mon, 23 Jun 2025 11:10:59 +0100</pubDate><guid>https://aibussin.com/post/thoughts/</guid><description>&lt;blockquote&gt;&#10;&lt;p&gt;How a self-evolving AI learns to reflect, score, and rewrite its own reasoning&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;h2 id="-summary"&gt;🧪 Summary&lt;/h2&gt;&#10;&lt;p&gt;What if an AI could think not just solve problems, but reevaluate its beliefs in the face of new information?&lt;/p&gt;&#10;&lt;p&gt;In this post, we introduce a system that does exactly that. At the core of our pipeline is a lightweight scoring model called MR.Q, responsible for evaluating ideas and choosing the best ones. But when it encounters a new domain, a new goal, or a shift in task format, it doesn’t freeze it adapts.&lt;/p&gt;</description></item><item><title>Learning to Learn: A LATS-Based Framework for Self-Aware AI Pipelines</title><link>https://aibussin.com/post/lats/</link><pubDate>Thu, 12 Jun 2025 09:23:46 +0100</pubDate><guid>https://aibussin.com/post/lats/</guid><description>&lt;h2 id="-summary"&gt;📖 Summary&lt;/h2&gt;&#10;&lt;p&gt;In this post, we introduce the LATSAgent, an implementation of &lt;a href="https://arxiv.org/pdf/2310.04406" target="_blank" class="paper-badge"&#10; style="display: inline-block; padding: 6px 10px; background: #f3f4f6; border-left: 4px solid #3b82f6; border-radius: 4px; margin: 4px 0; text-decoration: none; color: #1f2937;"&gt;&#10; &lt;strong&gt;LATS&lt;/strong&gt;: Language Agent Tree Search Unifies Reasoning..&#10;&lt;/a&gt; within the &lt;a href="https://github.com/ernanhughes/co-ai"&gt;stephanie&lt;/a&gt; framework. Unlike prior agents that followed a single reasoning chain, this agent explores multiple reasoning paths in parallel, evaluates them using multidimensional scoring, and learns symbolic refinements over time. This is our most complete integration yet of search, simulation, scoring, and symbolic tuning bringing together all of our previous work on sharpening, pipeline reflection, and symbolic rules into a unified, intelligent reasoning loop.&lt;/p&gt;</description></item><item><title>Dimensions of Thought: A Smarter Way to Evaluate AI</title><link>https://aibussin.com/post/dimensions/</link><pubDate>Mon, 09 Jun 2025 10:00:03 +0100</pubDate><guid>https://aibussin.com/post/dimensions/</guid><description>&lt;h2 id="-summary"&gt;📖 Summary&lt;/h2&gt;&#10;&lt;p&gt;This post introduces a multidimensional reward modeling pipeline built on top of the &lt;a href="https://github.com/ernanhughes/co-ai"&gt;stephanieanie&lt;/a&gt; framework. It covers:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;✅ &lt;strong&gt;Structured Evaluation Setup&lt;/strong&gt;&#10;How to define custom evaluation dimensions using YAML or database-backed rubrics.&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;🧠 &lt;strong&gt;Automated Scoring with LLMs&lt;/strong&gt;&#10;Using the &lt;code&gt;ScoreEvaluator&lt;/code&gt; to produce structured, rationale-backed scores for each dimension.&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;🧮 &lt;strong&gt;Embedding-Based Hypothesis Indexing&lt;/strong&gt;&#10;Efficiently embedding hypotheses and comparing them for contrastive learning using similarity.&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;🔄 &lt;strong&gt;Contrast Pair Generation&lt;/strong&gt;&#10;Creating training pairs where one hypothesis outperforms another on a given dimension.&lt;/p&gt;</description></item><item><title>Programming Intelligence: Using Symbolic Rules to Steer and Evolve AI</title><link>https://aibussin.com/post/symbolic/</link><pubDate>Wed, 04 Jun 2025 20:57:20 +0100</pubDate><guid>https://aibussin.com/post/symbolic/</guid><description>&lt;h2 id="-summary"&gt;🧪 Summary&lt;/h2&gt;&#10;&lt;p&gt;&amp;ldquo;What if AI systems could learn how to improve themselves not just at the level of weights or prompts, but at the level of strategy itself? In this post, we show how to build such a system, powered by symbolic rules and reflection.&lt;/p&gt;&#10;&lt;p&gt;The paper &lt;a href="https://arxiv.org/pdf/2406.18532v1" target="_blank" class="paper-badge"&#10; style="display: inline-block; padding: 6px 10px; background: #f3f4f6; border-left: 4px solid #3b82f6; border-radius: 4px; margin: 4px 0; text-decoration: none; color: #1f2937;"&gt;&#10; &lt;strong&gt;Symbolic Agents&lt;/strong&gt;: Symbolic Learning Enables Self-Evolving Agents&#10;&lt;/a&gt; introduces a framework where &lt;strong&gt;symbolic rules&lt;/strong&gt; guide, evaluate, and evolve agent behavior.&lt;/p&gt;</description></item><item><title>Adaptive Reasoning with ARM: Teaching AI the Right Way to Think</title><link>https://aibussin.com/post/arm/</link><pubDate>Wed, 28 May 2025 22:22:46 +0100</pubDate><guid>https://aibussin.com/post/arm/</guid><description>&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;&#10;&lt;p&gt;Chain-of-thought is powerful, but which chain? Short explanations work for easy tasks, long reflections help on hard ones, and code sometimes beats them both. What if your model could adaptively pick the best strategy, per task, and improve as it learns?&lt;/p&gt;&#10;&lt;p&gt;The &lt;code&gt;Adaptive Reasoning Model&lt;/code&gt; &lt;strong&gt;(ARM)&lt;/strong&gt; is a framework for teaching language models how to choose the right reasoning format direct answers, chain-of-thoughts, or code depending on the task. It works by evaluating responses, scoring them based on rarity, conciseness, and difficulty alignment, and then updating model behavior over time.&lt;/p&gt;</description></item></channel></rss>