<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Reinforcement Learning on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/reinforcement-learning/</link><description>Recent content in Reinforcement Learning on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Tue, 01 Sep 2026 10:41:00 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/reinforcement-learning/index.xml" rel="self" type="application/rss+xml"/><item><title>Reasoning Is More Than Architecture — Where Extra Computation Lives</title><link>https://aibussin.com/books/models-from-first-principles/10-chapter/</link><pubDate>Tue, 01 Sep 2026 10:41:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/10-chapter/</guid><description>&lt;h1 id="reasoning-is-more-than-architecture--where-extra-computation-lives"&gt;Reasoning Is More Than Architecture — Where Extra Computation Lives&lt;/h1&gt;&#10;&lt;p&gt;So far in &lt;strong&gt;Models From First Principles&lt;/strong&gt;, we have changed several different things and called all of them model design.&lt;/p&gt;&#10;&lt;p&gt;We changed what a model predicts.&lt;/p&gt;&#10;&lt;p&gt;MR.Q produced one learned quality score.&lt;/p&gt;&#10;&lt;p&gt;EBT added Q, V, policy, and advantage.&lt;/p&gt;&#10;&lt;p&gt;SICQL turned those outputs into explicit model components.&lt;/p&gt;&#10;&lt;p&gt;Then we changed how computation unfolds.&lt;/p&gt;&#10;&lt;p&gt;HRM introduced recurrent state operating at different timescales.&lt;/p&gt;</description></item><item><title>EBT — From One Score to Q, V, Policy and Advantage</title><link>https://aibussin.com/books/models-from-first-principles/03-chapter/</link><pubDate>Sat, 08 Aug 2026 14:39:00 +0100</pubDate><guid>https://aibussin.com/books/models-from-first-principles/03-chapter/</guid><description>&lt;h1 id="ebt--from-one-score-to-q-v-policy-and-advantage"&gt;EBT — From One Score to Q, V, Policy and Advantage&lt;/h1&gt;&#10;&lt;p&gt;In the previous post we built MR.Q: a small model that takes a context embedding and a response embedding, combines them, and predicts one scalar.&lt;/p&gt;&#10;&lt;p&gt;That architecture is useful because it is brutally simple:&lt;/p&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4;-webkit-text-size-adjust:none;"&gt;&lt;code class="language-text" data-lang="text"&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;context embedding&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; +&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt;response embedding&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; encoder&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; representation z&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; predictor&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; ↓&#10;&lt;/span&gt;&lt;/span&gt;&lt;span style="display:flex;"&gt;&lt;span&gt; Q value&#10;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;But one scalar eventually becomes restrictive.&lt;/p&gt;</description></item><item><title>Intelligence Through Execution: The Executable Cognitive Kernel</title><link>https://aibussin.com/post/eck/</link><pubDate>Tue, 10 Mar 2026 21:58:14 +0000</pubDate><guid>https://aibussin.com/post/eck/</guid><description>&lt;h2 id="-summary"&gt;🧭 Summary&lt;/h2&gt;&#10;&lt;p&gt;Most modern AI systems treat intelligence as something stored inside a model.&lt;/p&gt;&#10;&lt;p&gt;A neural network is trained on massive datasets, its weights are adjusted, and those weights become the system’s knowledge. When the model produces an output, we interpret that output as the result of the intelligence encoded inside those parameters.&lt;/p&gt;&#10;&lt;p&gt;But this perspective has a limitation.&lt;/p&gt;&#10;&lt;p&gt;Once training is complete, the model is largely static. It does not improve through its own actions, and it does not adapt based on the outcome of its behavior unless we retrain it.&lt;/p&gt;</description></item><item><title>Search–Solve–Prove: building a place for thoughts to develop</title><link>https://aibussin.com/post/ssp/</link><pubDate>Sun, 02 Nov 2025 01:13:06 +0000</pubDate><guid>https://aibussin.com/post/ssp/</guid><description>&lt;h2 id="-summary"&gt;🌌 Summary&lt;/h2&gt;&#10;&lt;p&gt;What if you could &lt;strong&gt;see an AI think&lt;/strong&gt; not just the final answer, but the whole stream of reasoning: every search, every dead end, every moment of insight? We’re building exactly that: a visible, measurable thought process we call &lt;strong&gt;the Jitter&lt;/strong&gt;. This post &lt;strong&gt;the first in a series&lt;/strong&gt; shows how we’re creating the &lt;strong&gt;habitat&lt;/strong&gt; where that digital thought stream can live and grow.&lt;/p&gt;&#10;&lt;p&gt;We’ll draw on ideas from:&lt;/p&gt;</description></item><item><title>Self-Improving AI: A System That Learns, Validates, and Retrains Itself</title><link>https://aibussin.com/post/rivals/</link><pubDate>Mon, 30 Jun 2025 10:13:03 +0100</pubDate><guid>https://aibussin.com/post/rivals/</guid><description>&lt;h2 id="-the-static-ai-trap"&gt;🤖 &lt;strong&gt;The Static AI Trap&lt;/strong&gt;&lt;/h2&gt;&#10;&lt;p&gt;Today’s AI systems are frozen in time: trained once, deployed forever. Yet the real world never stops evolving. Goals shift overnight. New research upends old truths. Context transforms without warning.&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;What if your AI could wake up?&lt;/strong&gt;&lt;/p&gt;&#10;&lt;p&gt;In this post, we engineer an intelligence that &lt;strong&gt;teaches itself&lt;/strong&gt; a system that continuously learns from the web, audits its own judgments, and retrains itself when confidence wavers.&lt;/p&gt;</description></item><item><title>Document Intelligence: Turning Documents into Structured Knowledge</title><link>https://aibussin.com/post/docs/</link><pubDate>Tue, 17 Jun 2025 23:31:13 +0100</pubDate><guid>https://aibussin.com/post/docs/</guid><description>&lt;h2 id="-summary"&gt;📖 Summary&lt;/h2&gt;&#10;&lt;p&gt;Imagine drowning in a sea of research papers, each holding a fragment of the knowledge you need for your next breakthrough. How does an AI system, striving for self-improvement, navigate this information overload to find precisely what it needs? This is the core challenge our Document Intelligence pipeline addresses, transforming chaotic documents into organized, searchable knowledge.&lt;/p&gt;&#10;&lt;p&gt;In this post we combine insights from &lt;a href="https://arxiv.org/pdf/2505.21497" target="_blank" class="paper-badge"&#10; style="display: inline-block; padding: 6px 10px; background: #f3f4f6; border-left: 4px solid #3b82f6; border-radius: 4px; margin: 4px 0; text-decoration: none; color: #1f2937;"&gt;&#10; &lt;strong&gt;Paper2Poster&lt;/strong&gt;: Towards Multimodal Poster Automation from Scientific Papers&#10;&lt;/a&gt; and&#10;&lt;a href="https://arxiv.org/abs/2506.10952" target="_blank" class="paper-badge"&#10; style="display: inline-block; padding: 6px 10px; background: #f3f4f6; border-left: 4px solid #3b82f6; border-radius: 4px; margin: 4px 0; text-decoration: none; color: #1f2937;"&gt;&#10; &lt;strong&gt;Domain2Vec&lt;/strong&gt;: Vectorizing Datasets to Find the Optimal Data Mixture without Training&#10;&lt;/a&gt; to build an AI document profiler that transforms unstructured papers into structured, searchable knowledge graphs.&lt;/p&gt;</description></item><item><title>Adaptive Reasoning with ARM: Teaching AI the Right Way to Think</title><link>https://aibussin.com/post/arm/</link><pubDate>Wed, 28 May 2025 22:22:46 +0100</pubDate><guid>https://aibussin.com/post/arm/</guid><description>&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;&#10;&lt;p&gt;Chain-of-thought is powerful, but which chain? Short explanations work for easy tasks, long reflections help on hard ones, and code sometimes beats them both. What if your model could adaptively pick the best strategy, per task, and improve as it learns?&lt;/p&gt;&#10;&lt;p&gt;The &lt;code&gt;Adaptive Reasoning Model&lt;/code&gt; &lt;strong&gt;(ARM)&lt;/strong&gt; is a framework for teaching language models how to choose the right reasoning format direct answers, chain-of-thoughts, or code depending on the task. It works by evaluating responses, scoring them based on rarity, conciseness, and difficulty alignment, and then updating model behavior over time.&lt;/p&gt;</description></item><item><title>MR.Q: Model-Based Representations for Model-Free Trading</title><link>https://aibussin.com/post/mrq/</link><pubDate>Tue, 18 Mar 2025 13:05:34 +0000</pubDate><guid>https://aibussin.com/post/mrq/</guid><description>&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;&#10;&lt;p&gt;Model-free reinforcement learning learns a policy directly from experience, but it can struggle to discover useful representations from sparse or noisy rewards. Model-based reinforcement learning receives a denser training signal by learning how states, actions, rewards, and termination relate to one another, but it often pays for that knowledge through planning complexity and model error.&lt;/p&gt;&#10;&lt;a href="https://arxiv.org/abs/2501.16142" target="_blank" class="paper-badge"&#10; style="display: inline-block; padding: 6px 10px; background: #f3f4f6; border-left: 4px solid #3b82f6; border-radius: 4px; margin: 4px 0; text-decoration: none; color: #1f2937;"&gt;&#10; &lt;strong&gt;MR.Q&lt;/strong&gt;: MR.Q&#10;&lt;/a&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Can a model-free agent keep the representation-learning benefits of a learned model without using that model to plan?&lt;/p&gt;</description></item></channel></rss>