<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>MRQ on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/mrq/</link><description>Recent content in MRQ on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Wed, 21 May 2025 18:36:51 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/mrq/index.xml" rel="self" type="application/rss+xml"/><item><title>Building a Self-Improving Chain-of-Thought Agent: Local LLMs Meet the CoT Encyclopedia</title><link>https://aibussin.com/post/cot/</link><pubDate>Wed, 21 May 2025 18:36:51 +0100</pubDate><guid>https://aibussin.com/post/cot/</guid><description>&lt;blockquote&gt;&#10;&lt;p&gt;Most AI systems generate answers. Ours examines how they think. This isn’t just prompt engineering this is structured reasoning at scale.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;h2 id="-summary"&gt;🔧 Summary&lt;/h2&gt;&#10;&lt;p&gt;Large Language Models are transforming every field, yet their internal reasoning remains a formidable black box. We can get brilliant outputs, but without understanding how those conclusions were reached, we&amp;rsquo;re left guessing how to improve, debug, or even trust them. This opacity limits our ability to build truly reliable and self-improving AI systems.&lt;/p&gt;</description></item><item><title>MR.Q: Model-Based Representations for Model-Free Trading</title><link>https://aibussin.com/post/mrq/</link><pubDate>Tue, 18 Mar 2025 13:05:34 +0000</pubDate><guid>https://aibussin.com/post/mrq/</guid><description>&lt;h2 id="introduction"&gt;Introduction&lt;/h2&gt;&#10;&lt;p&gt;Model-free reinforcement learning learns a policy directly from experience, but it can struggle to discover useful representations from sparse or noisy rewards. Model-based reinforcement learning receives a denser training signal by learning how states, actions, rewards, and termination relate to one another, but it often pays for that knowledge through planning complexity and model error.&lt;/p&gt;&#10;&lt;a href="https://arxiv.org/abs/2501.16142" target="_blank" class="paper-badge"&#10; style="display: inline-block; padding: 6px 10px; background: #f3f4f6; border-left: 4px solid #3b82f6; border-radius: 4px; margin: 4px 0; text-decoration: none; color: #1f2937;"&gt;&#10; &lt;strong&gt;MR.Q&lt;/strong&gt;: MR.Q&#10;&lt;/a&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Can a model-free agent keep the representation-learning benefits of a learned model without using that model to plan?&lt;/p&gt;</description></item></channel></rss>