<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Preference Learning on AIBussin — AI applications, systems and books</title><link>https://aibussin.com/tags/preference-learning/</link><description>Recent content in Preference Learning on AIBussin — AI applications, systems and books</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Fri, 25 Sep 2026 02:45:00 +0100</lastBuildDate><atom:link href="https://aibussin.com/tags/preference-learning/index.xml" rel="self" type="application/rss+xml"/><item><title>The Interface Learns You</title><link>https://aibussin.com/books/language/23-chapter/</link><pubDate>Fri, 25 Sep 2026 02:45:00 +0100</pubDate><guid>https://aibussin.com/books/language/23-chapter/</guid><description>&lt;p&gt;Chapter 21 gave the reader controls. Chapter 22 forbade turning their settings into a personality. Between those poles sits an unsolved accumulation problem: every interaction produces evidence — a correction here, a repeated choice there, a form that keeps performing better on one task — and with nowhere to accumulate, each encounter starts ignorant. Configuring every dimension manually is burdensome; silently inferring from clicks confuses habit and prominence with desire. This chapter learns from what actually happens, under Chapter 22&amp;rsquo;s evidentiary rule executed mechanically: no update without conditioning context, no type ever.&lt;/p&gt;</description></item><item><title>What Does a Preference Know About the Future?</title><link>https://aibussin.com/post/future/</link><pubDate>Fri, 10 Jul 2026 11:32:58 +0100</pubDate><guid>https://aibussin.com/post/future/</guid><description>&lt;h2 id="we-trained-a-model-on-editorial-choices-to-see-whether-it-learned-what-happened-next"&gt;We Trained a Model on Editorial Choices to See Whether It Learned What Happened Next&lt;/h2&gt;&#10;&lt;p&gt;Most preference-learning systems use a choice to change the future.&lt;/p&gt;&#10;&lt;p&gt;A model produces two responses. A human selects one. The chosen response becomes positive evidence, the rejected response becomes negative evidence, and training makes outputs resembling the chosen response more likely.&lt;/p&gt;&#10;&lt;p&gt;The preference acts as an instruction:&lt;/p&gt;&#10;&lt;blockquote&gt;&#10;&lt;p&gt;Produce more things like this.&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;p&gt;I wanted to know whether the same choice could also function as evidence.&lt;/p&gt;</description></item><item><title>The Preference Was Only the Beginning</title><link>https://aibussin.com/post/preferences/</link><pubDate>Tue, 07 Jul 2026 00:00:00 +0100</pubDate><guid>https://aibussin.com/post/preferences/</guid><description>&lt;blockquote&gt;&#10;&lt;p&gt;&lt;strong&gt;A preference is not only a label on what just happened. When the decision belongs to a continuing trajectory, it can also be evidence about what happens next.&lt;/strong&gt;&lt;/p&gt;&#10;&lt;/blockquote&gt;&#10;&lt;h2 id="abstract"&gt;Abstract&lt;/h2&gt;&#10;&lt;p&gt;Most preference-learning systems stop at the choice.&lt;/p&gt;&#10;&lt;p&gt;A model produces two responses. A human selects one. The chosen response becomes positive evidence, the rejected response becomes negative evidence, and the training system moves on.&lt;/p&gt;&#10;&lt;p&gt;The work itself usually continues.&lt;/p&gt;&#10;&lt;p&gt;The selected answer may later be revised, partially retained, contradicted or abandoned. The rejected alternative may reveal a constraint that remains active long after the immediate decision. The preference is therefore not necessarily the outcome. It may be an event inside a longer trajectory.&lt;/p&gt;</description></item><item><title>Document Intelligence: Turning Documents into Structured Knowledge</title><link>https://aibussin.com/post/docs/</link><pubDate>Tue, 17 Jun 2025 23:31:13 +0100</pubDate><guid>https://aibussin.com/post/docs/</guid><description>&lt;h2 id="-summary"&gt;📖 Summary&lt;/h2&gt;&#10;&lt;p&gt;Imagine drowning in a sea of research papers, each holding a fragment of the knowledge you need for your next breakthrough. How does an AI system, striving for self-improvement, navigate this information overload to find precisely what it needs? This is the core challenge our Document Intelligence pipeline addresses, transforming chaotic documents into organized, searchable knowledge.&lt;/p&gt;&#10;&lt;p&gt;In this post we combine insights from &lt;a href="https://arxiv.org/pdf/2505.21497" target="_blank" class="paper-badge"&#10; style="display: inline-block; padding: 6px 10px; background: #f3f4f6; border-left: 4px solid #3b82f6; border-radius: 4px; margin: 4px 0; text-decoration: none; color: #1f2937;"&gt;&#10; &lt;strong&gt;Paper2Poster&lt;/strong&gt;: Towards Multimodal Poster Automation from Scientific Papers&#10;&lt;/a&gt; and&#10;&lt;a href="https://arxiv.org/abs/2506.10952" target="_blank" class="paper-badge"&#10; style="display: inline-block; padding: 6px 10px; background: #f3f4f6; border-left: 4px solid #3b82f6; border-radius: 4px; margin: 4px 0; text-decoration: none; color: #1f2937;"&gt;&#10; &lt;strong&gt;Domain2Vec&lt;/strong&gt;: Vectorizing Datasets to Find the Optimal Data Mixture without Training&#10;&lt;/a&gt; to build an AI document profiler that transforms unstructured papers into structured, searchable knowledge graphs.&lt;/p&gt;</description></item><item><title>Adaptive Reasoning with ARM: Teaching AI the Right Way to Think</title><link>https://aibussin.com/post/arm/</link><pubDate>Wed, 28 May 2025 22:22:46 +0100</pubDate><guid>https://aibussin.com/post/arm/</guid><description>&lt;h2 id="summary"&gt;Summary&lt;/h2&gt;&#10;&lt;p&gt;Chain-of-thought is powerful, but which chain? Short explanations work for easy tasks, long reflections help on hard ones, and code sometimes beats them both. What if your model could adaptively pick the best strategy, per task, and improve as it learns?&lt;/p&gt;&#10;&lt;p&gt;The &lt;code&gt;Adaptive Reasoning Model&lt;/code&gt; &lt;strong&gt;(ARM)&lt;/strong&gt; is a framework for teaching language models how to choose the right reasoning format direct answers, chain-of-thoughts, or code depending on the task. It works by evaluating responses, scoring them based on rarity, conciseness, and difficulty alignment, and then updating model behavior over time.&lt;/p&gt;</description></item></channel></rss>