AIBussinBuild useful systems with modern AI
  • Home
  • Books
  • Solutions
  • Toolkit
  • Prompts
  • Articles
  • About

Evaluation

  • The Measurement Instrument
  • Meat Proxy
  • The Memory Nexus
  • Intelligence in the Wrong Direction
  • Examples Are Experimental Data
  • You Cannot Optimize What You Cannot Measure
  • What Survived the Transformation?
  • When the Metric Becomes the Target
  • A Revolver, Not a Foundation
  • The Nearest Neighbor Can Be Wrong
  • Optimize the Instructions
  • A Resolved Promise Is Not a Correct Answer
  • Evaluate the Feature, Not the Demo
  • Appendix: The Evidence Ledger
  • Don't Let the Optimizer Cheat
  • Tool Choice Is a Behavioral Problem
  • Did the Bridge Preserve the Space?
  • Can a Smaller Representation Preserve a Larger One?
  • Did the Context Help?
  • How Much of You Can an AI Reproduce?
  • Advanced Agents From First Principles 23: Your Infrastructure Is Healthy. Why Is the Agent Getting Worse? Detect Behavioral Drift and Roll Back Safely
  • Advanced Agents From First Principles 17: What Is Your Agent Actually Uncertain About?
  • Advanced Agents From First Principles 14: Can Your Agent Learn From Its Own Trajectories Without Learning the Wrong Lessons?
  • Agents From First Principles 09: AI Agent Says It Worked When It Didn’t? Verify the Result Outside the LLM
AIBussin

Build useful systems with modern AI.

Books, solutions, experiments and working tools for practical AI.

© 2026 Ernan Hughes
  • Books
  • Solutions
  • Toolkit
  • Prompts
  • Articles
  • About
  • Programmer.ie
  • ZeroModel.org
  • GitHub