<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>MR.Q on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/mr.q/</link><description>Recent content in MR.Q on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Fri, 10 Jul 2026 11:32:58 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/mr.q/index.xml" rel="self" type="application/rss+xml"/><item><title>What Does a Preference Know About the Future?</title><link>https://aibussin.com/post/future/</link><pubDate>Fri, 10 Jul 2026 11:32:58 +0100</pubDate><guid>https://aibussin.com/post/future/</guid><description>&lt;h2 id="we-trained-a-model-on-editorial-choices-to-see-whether-it-learned-what-happened-next"&gt;We Trained a Model on Editorial Choices to See Whether It Learned What Happened Next&lt;/h2&gt;&#10;&lt;p&gt;Most preference-learning systems use a choice to change the future.&lt;/p&gt;&#10;&lt;p&gt;A model produces two responses. A human selects one. The chosen response becomes positive evidence, the rejected response becomes negative evidence, and training makes outputs resembling the chosen response more likely.&lt;/p&gt;&#10;&lt;p&gt;The preference acts as an instruction:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Produce more things like this.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;I wanted to know whether the same choice could also function as evidence.&lt;/p&gt;</description></item><item><title>Compiling Thought: Building a Prompt Compiler for Self-Improving AI</title><link>https://aibussin.com/post/compiler/</link><pubDate>Thu, 26 Jun 2025 13:31:41 +0100</pubDate><guid>https://aibussin.com/post/compiler/</guid><description>&lt;p&gt;&lt;strong&gt;How to design a pipeline that turns vague goals into smart prompts&lt;/strong&gt;&lt;/p&gt;&#10;&lt;h2 id="-summary"&gt;🧪 Summary&lt;/h2&gt;&#10;&lt;p&gt;Why spend hours engineering prompts when AI can optimize its own instructions. This blog post introduces a novel approach toward creating a self-improving AI by treating prompts as programs. Traditional AI systems often rely on static instructions rigid and limited in adaptability. Here, we present a different perspective: viewing the Large Language Model (LLM) as a &lt;strong&gt;prompt compiler&lt;/strong&gt; capable of dynamically transforming raw instructions into optimized prompts through iterative cycles of decomposition, evaluation, and intelligent reassembly.&lt;/p&gt;</description></item><item><title>Thoughts of Algorithms</title><link>https://aibussin.com/post/thoughts/</link><pubDate>Mon, 23 Jun 2025 11:10:59 +0100</pubDate><guid>https://aibussin.com/post/thoughts/</guid><description>&lt;blockquote&gt;&#10;&lt;p&gt;How a self-evolving AI learns to reflect, score, and rewrite its own reasoning&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;h2 id="-summary"&gt;🧪 Summary&lt;/h2&gt;&#10;&lt;p&gt;What if an AI could think not just solve problems, but reevaluate its beliefs in the face of new information?&lt;/p&gt;&#10;&lt;p&gt;In this post, we introduce a system that does exactly that. At the core of our pipeline is a lightweight scoring model called MR.Q, responsible for evaluating ideas and choosing the best ones. But when it encounters a new domain, a new goal, or a shift in task format, it doesn’t freeze it adapts.&lt;/p&gt;</description></item><item><title>Document Intelligence: Turning Documents into Structured Knowledge</title><link>https://aibussin.com/post/docs/</link><pubDate>Tue, 17 Jun 2025 23:31:13 +0100</pubDate><guid>https://aibussin.com/post/docs/</guid><description>&lt;h2 id="-summary"&gt;📖 Summary&lt;/h2&gt;&#10;&lt;p&gt;Imagine drowning in a sea of research papers, each holding a fragment of the knowledge you need for your next breakthrough. How does an AI system, striving for self-improvement, navigate this information overload to find precisely what it needs? This is the core challenge our Document Intelligence pipeline addresses, transforming chaotic documents into organized, searchable knowledge.&lt;/p&gt;&#10;&lt;p&gt;In this post we combine insights from &lt;a href="https://arxiv.org/pdf/2505.21497" target="_blank" class="paper-badge"&#10; style="display: inline-block; padding: 6px 10px; background: #f3f4f6; border-left: 4px solid #3b82f6; border-radius: 4px; margin: 4px 0; text-decoration: none; color: #1f2937;"&gt;&#10; &lt;strong&gt;Paper2Poster&lt;/strong&gt;: Towards Multimodal Poster Automation from Scientific Papers&#10;&lt;/a&gt; and&#10;&lt;a href="https://arxiv.org/abs/2506.10952" target="_blank" class="paper-badge"&#10; style="display: inline-block; padding: 6px 10px; background: #f3f4f6; border-left: 4px solid #3b82f6; border-radius: 4px; margin: 4px 0; text-decoration: none; color: #1f2937;"&gt;&#10; &lt;strong&gt;Domain2Vec&lt;/strong&gt;: Vectorizing Datasets to Find the Optimal Data Mixture without Training&#10;&lt;/a&gt; to build an AI document profiler that transforms unstructured papers into structured, searchable knowledge graphs.&lt;/p&gt;</description></item><item><title>Learning to Learn: A LATS-Based Framework for Self-Aware AI Pipelines</title><link>https://aibussin.com/post/lats/</link><pubDate>Thu, 12 Jun 2025 09:23:46 +0100</pubDate><guid>https://aibussin.com/post/lats/</guid><description>&lt;h2 id="-summary"&gt;📖 Summary&lt;/h2&gt;&#10;&lt;p&gt;In this post, we introduce the LATSAgent, an implementation of &lt;a href="https://arxiv.org/pdf/2310.04406" target="_blank" class="paper-badge"&#10; style="display: inline-block; padding: 6px 10px; background: #f3f4f6; border-left: 4px solid #3b82f6; border-radius: 4px; margin: 4px 0; text-decoration: none; color: #1f2937;"&gt;&#10; &lt;strong&gt;LATS&lt;/strong&gt;: Language Agent Tree Search Unifies Reasoning..&#10;&lt;/a&gt; within the &lt;a href="https://github.com/ernanhughes/co-ai"&gt;stephanie&lt;/a&gt; framework. Unlike prior agents that followed a single reasoning chain, this agent explores multiple reasoning paths in parallel, evaluates them using multidimensional scoring, and learns symbolic refinements over time. This is our most complete integration yet of search, simulation, scoring, and symbolic tuning bringing together all of our previous work on sharpening, pipeline reflection, and symbolic rules into a unified, intelligent reasoning loop.&lt;/p&gt;</description></item><item><title>Dimensions of Thought: A Smarter Way to Evaluate AI</title><link>https://aibussin.com/post/dimensions/</link><pubDate>Mon, 09 Jun 2025 10:00:03 +0100</pubDate><guid>https://aibussin.com/post/dimensions/</guid><description>&lt;h2 id="-summary"&gt;📖 Summary&lt;/h2&gt;&#10;&lt;p&gt;This post introduces a multidimensional reward modeling pipeline built on top of the &lt;a href="https://github.com/ernanhughes/co-ai"&gt;stephanieanie&lt;/a&gt; framework. It covers:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;✅ &lt;strong&gt;Structured Evaluation Setup&lt;/strong&gt;&#10;How to define custom evaluation dimensions using YAML or database-backed rubrics.&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;🧠 &lt;strong&gt;Automated Scoring with LLMs&lt;/strong&gt;&#10;Using the &lt;code&gt;ScoreEvaluator&lt;/code&gt; to produce structured, rationale-backed scores for each dimension.&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;🧮 &lt;strong&gt;Embedding-Based Hypothesis Indexing&lt;/strong&gt;&#10;Efficiently embedding hypotheses and comparing them for contrastive learning using similarity.&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;🔄 &lt;strong&gt;Contrast Pair Generation&lt;/strong&gt;&#10;Creating training pairs where one hypothesis outperforms another on a given dimension.&lt;/p&gt;</description></item><item><title>Programming Intelligence: Using Symbolic Rules to Steer and Evolve AI</title><link>https://aibussin.com/post/symbolic/</link><pubDate>Wed, 04 Jun 2025 20:57:20 +0100</pubDate><guid>https://aibussin.com/post/symbolic/</guid><description>&lt;h2 id="-summary"&gt;🧪 Summary&lt;/h2&gt;&#10;&lt;p&gt;&amp;ldquo;What if AI systems could learn how to improve themselves not just at the level of weights or prompts, but at the level of strategy itself? In this post, we show how to build such a system, powered by symbolic rules and reflection.&lt;/p&gt;&#10;&lt;p&gt;The paper &lt;a href="https://arxiv.org/pdf/2406.18532v1" target="_blank" class="paper-badge"&#10; style="display: inline-block; padding: 6px 10px; background: #f3f4f6; border-left: 4px solid #3b82f6; border-radius: 4px; margin: 4px 0; text-decoration: none; color: #1f2937;"&gt;&#10; &lt;strong&gt;Symbolic Agents&lt;/strong&gt;: Symbolic Learning Enables Self-Evolving Agents&#10;&lt;/a&gt; introduces a framework where &lt;strong&gt;symbolic rules&lt;/strong&gt; guide, evaluate, and evolve agent behavior.&lt;/p&gt;</description></item><item><title>Adaptive Reasoning with ARM: Teaching AI the Right Way to Think</title><link>https://aibussin.com/post/arm/</link><pubDate>Wed, 28 May 2025 22:22:46 +0100</pubDate><guid>https://aibussin.com/post/arm/</guid><description>&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;&#10;&lt;p&gt;Chain-of-thought is powerful, but which chain? Short explanations work for easy tasks, long reflections help on hard ones, and code sometimes beats them both. What if your model could adaptively pick the best strategy, per task, and improve as it learns?&lt;/p&gt;&#10;&lt;p&gt;The &lt;code&gt;Adaptive Reasoning Model&lt;/code&gt; &lt;strong&gt;(ARM)&lt;/strong&gt; is a framework for teaching language models how to choose the right reasoning format direct answers, chain-of-thoughts, or code depending on the task. It works by evaluating responses, scoring them based on rarity, conciseness, and difficulty alignment, and then updating model behavior over time.&lt;/p&gt;</description></item><item><title>General Reasoner: The smarter Local Agent</title><link>https://aibussin.com/post/general/</link><pubDate>Thu, 22 May 2025 21:41:54 +0100</pubDate><guid>https://aibussin.com/post/general/</guid><description>&lt;h2 id="-summary"&gt;🔧 Summary&lt;/h2&gt;&#10;&lt;p&gt;The &lt;a href="https://arxiv.org/abs/2505.14652"&gt;General Reasoner&lt;/a&gt; paper shows how we can train LLMs to reason across domains using diverse data and a generative verifier. In this post, I walk through our open-source implementation showing how we built a modular reasoning agent capable of generating multiple hypotheses, evaluating them with an LLM-based judge, and selecting the best answer.&lt;/p&gt;&#10;&lt;hr&gt;&#10;&lt;h2 id="-what-we-built"&gt;🧠 What We Built&lt;/h2&gt;&#10;&lt;p&gt;We built a &lt;code&gt;GeneralReasonerAgent&lt;/code&gt; that:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Dynamically generates multiple hypotheses using different &lt;strong&gt;reasoning strategies&lt;/strong&gt; (e.g., &lt;code&gt;cot&lt;/code&gt;, &lt;code&gt;debate&lt;/code&gt;, &lt;code&gt;verify_then_answer&lt;/code&gt;, etc.)&lt;/li&gt;&#10;&lt;li&gt;Evaluates each pair of hypotheses using either a &lt;strong&gt;local LLM judge&lt;/strong&gt; or our custom &lt;strong&gt;MR.Q evaluator&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;Classifies the winning hypothesis using &lt;strong&gt;rubric dimensions&lt;/strong&gt;&lt;/li&gt;&#10;&lt;li&gt;Logs structured results to a PostgreSQL-backed system&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;All of this was integrated with our existing stephanie framework, which includes:&lt;/p&gt;</description></item><item><title>Self-Improving Agents: Applying the Sharpening Framework to Local LLMs</title><link>https://aibussin.com/post/sharpen/</link><pubDate>Tue, 20 May 2025 09:23:16 +0100</pubDate><guid>https://aibussin.com/post/sharpen/</guid><description>&lt;blockquote&gt;&#10;&lt;p&gt;This is the second post in a 100-part series, where we take breakthrough AI papers and turn them into working code building the next generation of AI, one idea at a time.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;h2 id="-summary"&gt;🔧 Summary&lt;/h2&gt;&#10;&lt;p&gt;In my previous post, I introduced &lt;code&gt;stephanie&lt;/code&gt; a &lt;strong&gt;modular implementation of the AI co-scientist concept&lt;/strong&gt;, inspired by DeepMind’s recent paper &lt;em&gt;&lt;a href="https://arxiv.org/abs/2502.18864"&gt;Towards an AI Co-Scientist&lt;/a&gt;&lt;/em&gt;.&lt;/p&gt;&#10;&lt;p&gt;But now, we’re going deeper.&lt;/p&gt;&#10;&lt;p&gt;This isn’t just about &lt;strong&gt;running prompts through an agent system&lt;/strong&gt; it’s about building something radically different:&lt;/p&gt;</description></item></channel></rss>