<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Sirui He · Technical Blog</title><description>Writing by Sirui He on LLMs and AI agents.</description><link>https://blog.hesirui.com/</link><language>en</language><item><title>GRPO From First Principles: Why On-Policy RL for LLMs Is Really Just Dynamic Weighted SFT</title><link>https://blog.hesirui.com/posts/grpo-from-first-principles/</link><guid isPermaLink="true">https://blog.hesirui.com/posts/grpo-from-first-principles/</guid><description>A first-principles walk from supervised fine-tuning to policy gradient and GRPO — why on-policy RL for LLMs is really dynamic, group-normalized, reward-weighted SFT.</description><pubDate>Sun, 31 May 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;Most explanations of PPO, GRPO, and RLHF open the same way: with a wall of formulas. Policy gradients, advantage functions, KL penalties, Monte Carlo estimators, old policies, reference models, clipping terms, reward models. All of it, all at once.&lt;/p&gt;
&lt;p&gt;If you already think in reinforcement learning, that’s efficient. If your instincts were shaped by modern supervised training (cross entropy, next-token prediction, SFT, backprop), it’s a brick wall.&lt;/p&gt;
&lt;p&gt;This article takes the other road. We start from the most familiar object in LLM training, supervised fine-tuning, and walk step by step until we arrive at policy gradient and GRPO. The whole journey is organized around one question:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;If a model produces an output, and an external verifier scores that output, how do we update the model when the reward itself is not differentiable?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The answer isn’t magic. It’s one of the most useful tricks in all of reinforcement learning:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Don’t try to differentiate through the reward. Instead, raise the probability of outputs that beat a baseline, and lower the probability of outputs that fall short of it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Once you see it that way, on-policy RL for LLMs stops being mysterious. The whole thing collapses into a single sentence:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;GRPO is dynamic, group-normalized, reward-weighted SFT on model-generated samples.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That sentence isn’t the entire story, but it’s the right thing to keep in your head as the story unfolds.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;1-the-teacher-the-student-and-the-missing-answer-key&quot;&gt;&lt;span class=&quot;num&quot;&gt;1&lt;span class=&quot;vh&quot;&gt;.&lt;/span&gt;&lt;/span&gt; The Teacher, the Student, and the Missing Answer Key&lt;/h2&gt;
&lt;p&gt;Start in a classroom.&lt;/p&gt;
&lt;p&gt;In supervised fine-tuning, the teacher hands the student both a question &lt;em&gt;and&lt;/em&gt; the official answer. There’s nothing to explore. The student just learns an association:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;“When I see this question, I should produce this answer.”&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That’s SFT. It’s powerful, but at heart it’s imitation: the gold answer is treated as ground truth, full stop.&lt;/p&gt;
&lt;p&gt;Reinforcement learning changes the deal. The teacher still poses a question, but now there’s no answer key. The student attempts something, and the teacher reacts: good, bad, partially right, too long, unsafe, elegant, wrong format. Feedback, not a solution.&lt;/p&gt;
&lt;p&gt;This is much closer to how real problems actually work. The interesting ones rarely arrive with a gold response attached. You try, you get feedback, and over many rounds you develop a feel for which kinds of attempts tend to land.&lt;/p&gt;
&lt;p&gt;So the philosophical split between the two is clean:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;SFT&lt;/strong&gt; learns from a reference answer.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;RL&lt;/strong&gt; learns from trial and error.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For LLMs, that trial-and-error loop is even possible because the model can sample different answers from its own distribution. But that innocent-sounding sentence hides something worth pausing on.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;2-generative-ai-is-not-automatically-stochastic&quot;&gt;&lt;span class=&quot;num&quot;&gt;2&lt;span class=&quot;vh&quot;&gt;.&lt;/span&gt;&lt;/span&gt; Generative AI Is Not Automatically Stochastic&lt;/h2&gt;
&lt;p&gt;You’ll often hear that generative AI is stochastic. That’s only half true.&lt;/p&gt;
&lt;p&gt;An autoregressive LLM defines a probability distribution over the next token:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mo lspace=&quot;0em&quot; rspace=&quot;0em&quot;&gt;&amp;#x3C;&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\pi_\theta(y_t \mid x, y_{&amp;#x3C;t})&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;If we &lt;em&gt;sample&lt;/em&gt; from that distribution, the model is stochastic: the same prompt can produce many different completions. But sampling is a choice, not a law. We could just as easily decode greedily:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;arg&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;munder&gt;&lt;mi&gt;max&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;/munder&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mo lspace=&quot;0em&quot; rspace=&quot;0em&quot;&gt;&amp;#x3C;&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;y_t = \arg\max_y \pi_\theta(y \mid x, y_{&amp;#x3C;t})&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;Now the model is deterministic. Same prompt, same weights, same output, every time. In practice greedy decoding tends to be too rigid; it repeats itself and settles into bland high-probability ruts. Sampling is the trade: diversity and exploration on one side, uncertainty on the other.&lt;/p&gt;
&lt;p&gt;This distinction matters more than it looks, because not every generative model is naturally policy-like. Some generators outside the LLM world (certain flow-matching models, for instance) behave like deterministic maps or velocity fields once you fix the noise, the solver, and the condition. They don’t hand you the convenient token-level log probabilities that an LLM gives away for free.&lt;/p&gt;
&lt;p&gt;That’s why the next two sections work through two strategies side by side:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;a &lt;strong&gt;deterministic&lt;/strong&gt; strategy, and&lt;/li&gt;
&lt;li&gt;a &lt;strong&gt;stochastic&lt;/strong&gt; strategy.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If all we cared about were ordinary sampled LLMs, we could almost skip the deterministic case. But it’s worth the detour: the deterministic setup exposes the &lt;em&gt;exact&lt;/em&gt; problem that policy gradient was invented to solve, and it sets up the non-LLM generators, like flow-matching TTS, where probability sometimes has to be deliberately added back in.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;3-the-deterministic-strategy-and-the-gradient-gap&quot;&gt;&lt;span class=&quot;num&quot;&gt;3&lt;span class=&quot;vh&quot;&gt;.&lt;/span&gt;&lt;/span&gt; The Deterministic Strategy and the Gradient Gap&lt;/h2&gt;
&lt;p&gt;Suppose the model maps an input straight to an output, deterministically. (Picture a continuous generator here, not greedy LLM decoding. Greedy decoding is deterministic too, but for a different reason: the &lt;code&gt;argmax&lt;/code&gt; is itself non-differentiable.)&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;y = f_\theta(x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;An external reward function scores that output:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;R(y) = R(f_\theta(x))&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;The most direct way to improve the model is to push that score up:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;munder&gt;&lt;mi&gt;max&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/munder&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\max_\theta R(f_\theta(x))&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;which calls for the gradient&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta R(f_\theta(x))&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;and, by the chain rule,&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;/msub&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;⋅&lt;/mo&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta R(f_\theta(x))
=
\nabla_y R(y)
\cdot
\nabla_\theta f_\theta(x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;The model side,&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta f_\theta(x),&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;is no trouble at all. That’s just backprop. The trouble is the reward side:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;/msub&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_y R(y)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;Plenty of the rewards we actually care about simply don’t give us this gradient. A verifier reports whether an answer is correct. A human can prefer answer A to answer B. An ASR system emits a transcript, and WER comes out of an edit-distance computation. For a concrete sample, all of these can produce a label or score. What they do not produce is a useful gradient with respect to the model output.&lt;/p&gt;
&lt;p&gt;So we hit a wall:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The reward tells us &lt;em&gt;whether&lt;/em&gt; the output was good. It never tells us &lt;em&gt;how to nudge&lt;/em&gt; the output to make it better.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Deterministic policy gradient methods have one answer to this: learn a critic, a Q-function that estimates long-term value.&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi&gt;Q&lt;/mi&gt;&lt;mi class=&quot;tml-med-pad&quot;&gt;π&lt;/mi&gt;&lt;/msup&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;𝔼&lt;/mi&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;true&quot;&gt;[&lt;/mo&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∑&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;∞&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;msup&gt;&lt;mi&gt;γ&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msup&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;true&quot;&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;Q^\pi(s,a)
=
\mathbb{E}_\pi
\left[
\sum_{t=0}^{\infty}
\gamma^t r_t
\mid
s_0=s, a_0=a
\right]&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;If &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;Q&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;Q(s,a)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; is differentiable in the action &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;a&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;, a deterministic actor can ride its gradient uphill:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mi&gt;J&lt;/mi&gt;&lt;mo&gt;≈&lt;/mo&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;/msub&gt;&lt;mi&gt;Q&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;false&quot; stretchy=&quot;true&quot; symmetric=&quot;true&quot; minsize=&quot;1.2em&quot; maxsize=&quot;1.2em&quot;&gt;|&lt;/mo&gt;&lt;mrow&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;μ&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;⋅&lt;/mo&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;μ&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta J
\approx
\nabla_a Q(s,a)
\big|_{a=\mu_\theta(s)}
\cdot
\nabla_\theta \mu_\theta(s)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;In effect, the critic supplies the missing direction: it tells us which way the action should move.&lt;/p&gt;
&lt;p&gt;But for high-dimensional generation with sparse, black-box, verifier-style rewards, training a reliable critic is hard and often unstable. There’s a second route that sidesteps the whole problem. Instead of asking &lt;em&gt;how the output should move&lt;/em&gt;, ask a different question entirely: &lt;em&gt;which of the outputs I already sampled should become more likely?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;That’s the stochastic strategy.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;4-the-stochastic-strategy-optimize-the-distribution&quot;&gt;&lt;span class=&quot;num&quot;&gt;4&lt;span class=&quot;vh&quot;&gt;.&lt;/span&gt;&lt;/span&gt; The Stochastic Strategy: Optimize the Distribution&lt;/h2&gt;
&lt;p&gt;A stochastic policy doesn’t commit to a single output. It defines a distribution:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo&gt;∼&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot; lspace=&quot;0&quot; rspace=&quot;0&quot;&gt;⋅&lt;/mo&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;y \sim \pi_\theta(\cdot \mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;Two notations that are easy to blur together:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot; lspace=&quot;0&quot; rspace=&quot;0&quot;&gt;⋅&lt;/mo&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\pi_\theta(\cdot \mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; is the &lt;em&gt;entire&lt;/em&gt; output distribution given &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;x&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;.&lt;/li&gt;
&lt;li&gt;&lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\pi_\theta(y \mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; is the probability (or density) of &lt;em&gt;one specific&lt;/em&gt; output &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;y&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The objective changes shape. We’re no longer maximizing the score of a single deterministic output. We’re maximizing &lt;em&gt;expected&lt;/em&gt; reward over the distribution:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;J&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;𝔼&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo&gt;∼&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot; lspace=&quot;0&quot; rspace=&quot;0&quot;&gt;⋅&lt;/mo&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;true&quot;&gt;[&lt;/mo&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;true&quot;&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;J(\theta)
=
\mathbb{E}_{y\sim\pi_\theta(\cdot \mid x)}
\left[
R(y)
\right]&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;Written as an integral:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;J&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∫&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;J(\theta)
=
\int
\pi_\theta(y \mid x) R(y)
\,dy&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;and for discrete outputs, the integral is just a sum:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;J&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;munder&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∑&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;/munder&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;J(\theta)
=
\sum_y
\pi_\theta(y \mid x)R(y)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;This is nothing more exotic than a probability-weighted average reward. We want the model to move its probability mass toward the high-reward outputs.&lt;/p&gt;
&lt;p&gt;And here’s the inversion that makes everything downstream work:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;We don’t need to know how to edit a specific output to raise its reward.&lt;/p&gt;
&lt;p&gt;We only need to know how to reshape the model so that high-reward &lt;em&gt;sampled&lt;/em&gt; outputs become more likely.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure class=&quot;figure&quot;&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;#x22;src&amp;#x22;:&amp;#x22;./gradient-gap.webp&amp;#x22;,&amp;#x22;alt&amp;#x22;:&amp;#x22;Deterministic reward optimization gets stuck at the missing reward gradient, while stochastic policy gradient routes the update through sampled outputs and their log probabilities.&amp;#x22;,&amp;#x22;index&amp;#x22;:0}&quot;&gt;&lt;figcaption class=&quot;fig-caption&quot;&gt;The core inversion: instead of differentiating through the verifier, policy gradient changes the probability of sampled outputs.&lt;/figcaption&gt;&lt;/figure&gt;

&lt;hr&gt;
&lt;h2 id=&quot;5-sft-first-increasing-the-probability-of-a-gold-answer&quot;&gt;&lt;span class=&quot;num&quot;&gt;5&lt;span class=&quot;vh&quot;&gt;.&lt;/span&gt;&lt;/span&gt; SFT First: Increasing the Probability of a Gold Answer&lt;/h2&gt;
&lt;p&gt;Before deriving policy gradient, let’s re-anchor on SFT, because the two turn out to be siblings.&lt;/p&gt;
&lt;p&gt;In SFT the dataset hands us a gold answer:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mtext&gt;gt&lt;/mtext&gt;&lt;/msub&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;y_{\text{gt}}&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;and the objective is maximum likelihood:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;munder&gt;&lt;mi&gt;max&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/munder&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mtext&gt;gt&lt;/mtext&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\max_\theta
\log
\pi_\theta(y_{\text{gt}} \mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;which is the same as minimizing cross entropy:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi class=&quot;mathcal&quot;&gt;ℒ&lt;/mi&gt;&lt;mtext&gt;SFT&lt;/mtext&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;−&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mtext&gt;gt&lt;/mtext&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\mathcal{L}_{\text{SFT}}
=
-
\log
\pi_\theta(y_{\text{gt}} \mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;For an autoregressive LLM, the sequence probability factorizes token by token:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mtext&gt;gt&lt;/mtext&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∏&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mo lspace=&quot;0em&quot; rspace=&quot;0em&quot;&gt;&amp;#x3C;&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\pi_\theta(y_{\text{gt}} \mid x)
=
\prod_{t=1}^{T}
\pi_\theta(y_t \mid x, y_{&amp;#x3C;t})&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;and taking the log turns that product into a sum:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mtext&gt;gt&lt;/mtext&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∑&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mo lspace=&quot;0em&quot; rspace=&quot;0em&quot;&gt;&amp;#x3C;&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\log
\pi_\theta(y_{\text{gt}} \mid x)
=
\sum_{t=1}^{T}
\log
\pi_\theta(y_t \mid x, y_{&amp;#x3C;t})&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;So SFT, stripped to one line, says:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Given a trusted answer, raise the log probability of the token path that produces it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Hold onto that shape. Policy gradient is going to look almost identical, with one decisive change: the answer is sampled by the model itself, and its weight comes from reward.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;6-why-we-need-the-log-trick&quot;&gt;&lt;span class=&quot;num&quot;&gt;6&lt;span class=&quot;vh&quot;&gt;.&lt;/span&gt;&lt;/span&gt; Why We Need the Log Trick&lt;/h2&gt;
&lt;p&gt;We want the gradient of the objective:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mi&gt;J&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta J(\theta)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;Start from the integral form:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;J&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∫&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;J(\theta)
=
\int
\pi_\theta(y \mid x)R(y)
\,dy&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;and differentiate:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mi&gt;J&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∫&lt;/mo&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta J(\theta)
=
\int
\nabla_\theta
\pi_\theta(y \mid x)
R(y)
\,dy&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;This is correct, but it’s stuck. We can’t estimate it by sampling from &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\pi_\theta&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; yet, and sampling is the only thing we can actually do. Monte Carlo estimation needs the integrand to wear a particular costume:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∫&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;𝔼&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo&gt;∼&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;[&lt;/mo&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\int
\pi_\theta(y \mid x)
f(y)
\,dy
=
\mathbb{E}_{y\sim\pi_\theta}
[f(y)]&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;because &lt;em&gt;then&lt;/em&gt; we can draw samples&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;mo&gt;…&lt;/mo&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;∼&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot; lspace=&quot;0&quot; rspace=&quot;0&quot;&gt;⋅&lt;/mo&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;y_1,\dots,y_N
\sim
\pi_\theta(\cdot \mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;and approximate the expectation by averaging:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;𝔼&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;[&lt;/mo&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;≈&lt;/mo&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/mfrac&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∑&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\mathbb{E}[f(y)]
\approx
\frac{1}{N}
\sum_{i=1}^{N}
f(y_i)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;The snag is that our gradient contains&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta \pi_\theta(y \mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;not the bare&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\pi_\theta(y \mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;that the Monte Carlo form requires out front. We need to coax &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\pi_\theta&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; back into that leading position.&lt;/p&gt;
&lt;p&gt;Enter the score-function trick. It starts from the ordinary derivative of a log:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta
\log
\pi_\theta(y \mid x)
=
\frac{1}{\pi_\theta(y \mid x)}
\nabla_\theta
\pi_\theta(y \mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;Rearranged, that’s exactly the substitution we need:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta
\pi_\theta(y \mid x)
=
\pi_\theta(y \mid x)
\nabla_\theta
\log
\pi_\theta(y \mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;Drop it back into the gradient:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mi&gt;J&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∫&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta J(\theta)
=
\int
\pi_\theta(y \mid x)
R(y)
\nabla_\theta
\log
\pi_\theta(y \mid x)
\,dy&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;and now the integrand is in Monte Carlo costume,&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∫&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;mspace width=&quot;2em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\int
\pi_\theta(y \mid x) f(y)\,dy,
\qquad
f(y)
=
R(y)
\nabla_\theta
\log
\pi_\theta(y \mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;so the gradient is just an expectation:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mi&gt;J&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;𝔼&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo&gt;∼&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot; lspace=&quot;0&quot; rspace=&quot;0&quot;&gt;⋅&lt;/mo&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;true&quot;&gt;[&lt;/mo&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;true&quot;&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta J(\theta)
=
\mathbb{E}_{y\sim\pi_\theta(\cdot \mid x)}
\left[
R(y)
\nabla_\theta
\log
\pi_\theta(y \mid x)
\right]&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;This is the REINFORCE estimator, also called the score-function or likelihood-ratio estimator.&lt;/p&gt;
&lt;p&gt;One thing to keep straight through all of this: the reward is a &lt;em&gt;constant&lt;/em&gt; with respect to &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\theta&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;. It’s a black box that hands back a number, nothing more. Differentiable pieces like a KL penalty don’t live inside &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;R&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;; they’re tacked on separately as their own loss term.&lt;/p&gt;
&lt;p&gt;And notice &lt;em&gt;why&lt;/em&gt; the log showed up. It wasn’t a clever nod to SFT. It appeared because we had to rewrite&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mspace width=&quot;2em&quot;&gt;&lt;/mspace&gt;&lt;mtext&gt;as&lt;/mtext&gt;&lt;mspace width=&quot;2em&quot;&gt;&lt;/mspace&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta \pi_\theta
\qquad\text{as}\qquad
\pi_\theta
\nabla_\theta
\log
\pi_\theta&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;so that the gradient could be estimated by sampling from &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\pi_\theta&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;. The resemblance to SFT is a consequence, not the motivation, which is exactly what makes it satisfying.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;7-monte-carlo-engineering-the-law-of-large-numbers&quot;&gt;&lt;span class=&quot;num&quot;&gt;7&lt;span class=&quot;vh&quot;&gt;.&lt;/span&gt;&lt;/span&gt; Monte Carlo: Engineering the Law of Large Numbers&lt;/h2&gt;
&lt;p&gt;The exact expectation is hopeless to compute, because the output space is astronomically large. This is the practical pain that the clean notation quietly papers over.&lt;/p&gt;
&lt;p&gt;In the integral, &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;y&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; ranges over &lt;em&gt;every possible output&lt;/em&gt;:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∫&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\int
\pi_\theta(y\mid x)
R(y)
\nabla_\theta
\log
\pi_\theta(y\mid x)
\,dy&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;Mathematically that’s well-defined. Computationally it’s a fantasy. We can’t loop over every completion an LLM might produce, score each one, take each log-prob gradient, and sum it all up.&lt;/p&gt;
&lt;p&gt;So we don’t. We sample:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;mo&gt;…&lt;/mo&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;∼&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot; lspace=&quot;0&quot; rspace=&quot;0&quot;&gt;⋅&lt;/mo&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;y_1,\dots,y_N
\sim
\pi_\theta(\cdot \mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;and estimate:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mi&gt;J&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;≈&lt;/mo&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/mfrac&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∑&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta J(\theta)
\approx
\frac{1}{N}
\sum_{i=1}^{N}
R(y_i)
\nabla_\theta
\log
\pi_\theta(y_i \mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;That’s Monte Carlo estimation, and the engine underneath it is the law of large numbers:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/mfrac&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∑&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo stretchy=&quot;false&quot;&gt;→&lt;/mo&gt;&lt;mi&gt;𝔼&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;[&lt;/mo&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;mspace width=&quot;1em&quot;&gt;&lt;/mspace&gt;&lt;mtext&gt;as &lt;/mtext&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;mo&gt;→&lt;/mo&gt;&lt;mi&gt;∞&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\frac{1}{N}
\sum_{i=1}^{N}
f(y_i)
\rightarrow
\mathbb{E}[f(y)]
\quad\text{as } N\to\infty&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;This is the moment the symbols turn into computation. The abstract variable &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;y&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; becomes a handful of concrete samples &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;y_i&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;, and every term in the sum is something we can actually evaluate:&lt;/p&gt;
&lt;div class=&quot;table-wrap&quot;&gt;





























&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Symbol&lt;/th&gt;&lt;th&gt;Meaning&lt;/th&gt;&lt;th&gt;Can we compute it directly?&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;y&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;&lt;/td&gt;&lt;td&gt;A variable over the entire output space&lt;/td&gt;&lt;td&gt;No: too many possible outputs to enumerate.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;y_i&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;&lt;/td&gt;&lt;td&gt;One sampled output from the current model&lt;/td&gt;&lt;td&gt;Yes: it’s a concrete sequence.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;R(y_i)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;&lt;/td&gt;&lt;td&gt;Reward of that sampled output&lt;/td&gt;&lt;td&gt;Yes: the verifier returns a scalar.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta \log \pi_\theta(y_i\mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;&lt;/td&gt;&lt;td&gt;Gradient of the sample’s log probability&lt;/td&gt;&lt;td&gt;Yes: ordinary backprop, exactly like SFT.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;The name “Monte Carlo” is just a label for this practice: the samples are model outputs, the function is reward times score, and their average estimates the gradient. This is where trial and error becomes mathematics, and it lines up perfectly with the classroom. The student tries several answers, the teacher scores them, and the update rule reads:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Answers that scored higher should become more likely next time.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id=&quot;8-from-gradient-estimator-to-loss&quot;&gt;&lt;span class=&quot;num&quot;&gt;8&lt;span class=&quot;vh&quot;&gt;.&lt;/span&gt;&lt;/span&gt; From Gradient Estimator to Loss&lt;/h2&gt;
&lt;p&gt;In practice we don’t hand-code gradients. We write a loss and let backprop do the rest.&lt;/p&gt;
&lt;p&gt;For a single sampled output &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;y_i&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;, the simplest policy-gradient loss is:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi class=&quot;mathcal&quot;&gt;ℒ&lt;/mi&gt;&lt;mtext&gt;PG&lt;/mtext&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;−&lt;/mo&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\mathcal{L}_{\text{PG}}
=
-
R(y_i)
\log
\pi_\theta(y_i \mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;In real systems, though, we almost always swap the raw reward for an &lt;em&gt;advantage&lt;/em&gt;:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi class=&quot;mathcal&quot;&gt;ℒ&lt;/mi&gt;&lt;mtext&gt;PG&lt;/mtext&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;−&lt;/mo&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\mathcal{L}_{\text{PG}}
=
-
A(y_i)
\log
\pi_\theta(y_i \mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;The advantage answers a sharper question: was this output better than some baseline?&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;A(y_i)
=
R(y_i)-b&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;Why are we allowed to subtract a baseline at all? Because as long as &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;b&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; doesn’t depend on the sampled action, it vanishes in expectation:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;𝔼&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo&gt;∼&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;true&quot;&gt;[&lt;/mo&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;true&quot;&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∫&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\mathbb{E}_{y\sim\pi_\theta}
\left[
b\nabla_\theta\log\pi_\theta(y\mid x)
\right]
=
b\nabla_\theta
\int
\pi_\theta(y\mid x)\,dy
=
b\nabla_\theta 1
=
0&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;So the baseline leaves the &lt;em&gt;expected&lt;/em&gt; gradient untouched; it only cuts the variance. This is the first place engineering quietly enters the story: the estimator is unbiased but noisy, so we reshape the reward into a steadier learning signal without biasing it.&lt;/p&gt;
&lt;p&gt;The sign of the advantage is what drives learning. When&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;&gt;&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;A(y_i)&gt;0,&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;minimizing the loss pushes the log probability of &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;y_i&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; &lt;em&gt;up&lt;/em&gt;. When&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#x3C;&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;A(y_i)&amp;#x3C;0,&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;it pushes that log probability &lt;em&gt;down&lt;/em&gt;. Good attempts get reinforced; bad ones get suppressed.&lt;/p&gt;
&lt;p&gt;And this is precisely why policy gradient feels like a weighted version of SFT. Put them next to each other:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi class=&quot;mathcal&quot;&gt;ℒ&lt;/mi&gt;&lt;mtext&gt;SFT&lt;/mtext&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;−&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mtext&gt;gt&lt;/mtext&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mspace width=&quot;2em&quot;&gt;&lt;/mspace&gt;&lt;mspace width=&quot;2em&quot;&gt;&lt;/mspace&gt;&lt;msub&gt;&lt;mi class=&quot;mathcal&quot;&gt;ℒ&lt;/mi&gt;&lt;mtext&gt;PG&lt;/mtext&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;−&lt;/mo&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\mathcal{L}_{\text{SFT}}
=
-
\log
\pi_\theta(y_{\text{gt}}\mid x)
\qquad\qquad
\mathcal{L}_{\text{PG}}
=
-
A(y)
\log
\pi_\theta(y\mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;The difference is not cosmetic. In SFT the answer comes from the dataset and is trusted as gold. In policy gradient the answer is sampled by the current model and weighted by feedback. Same skeleton, completely different source of truth.&lt;/p&gt;
&lt;p&gt;The same contrast, with GRPO added as the natural next step:&lt;/p&gt;
&lt;div class=&quot;table-wrap&quot;&gt;




























&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Method&lt;/th&gt;&lt;th&gt;Where the answer comes from&lt;/th&gt;&lt;th&gt;What is optimized&lt;/th&gt;&lt;th&gt;Mental model&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;SFT&lt;/td&gt;&lt;td&gt;A gold answer from the dataset&lt;/td&gt;&lt;td&gt;&lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;−&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mtext&gt;gt&lt;/mtext&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;-\log \pi_\theta(y_{\text{gt}}\mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;&lt;/td&gt;&lt;td&gt;Imitate the answer key.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Policy gradient&lt;/td&gt;&lt;td&gt;A sample from the current model&lt;/td&gt;&lt;td&gt;&lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;−&lt;/mo&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;-A(y)\log \pi_\theta(y\mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;&lt;/td&gt;&lt;td&gt;Reinforce the better attempts.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;GRPO&lt;/td&gt;&lt;td&gt;A &lt;em&gt;group&lt;/em&gt; of samples from the current model&lt;/td&gt;&lt;td&gt;&lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;−&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;-A_i\log \pi_\theta(y_i\mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;, with group-relative &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;A_i&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;&lt;/td&gt;&lt;td&gt;Reward-weighted SFT, weights normalized within the group.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;For an autoregressive LLM, one detail is worth spelling out: the sequence-level advantage multiplies the &lt;em&gt;entire&lt;/em&gt; sum of token log-probabilities:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi class=&quot;mathcal&quot;&gt;ℒ&lt;/mi&gt;&lt;mtext&gt;PG&lt;/mtext&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;−&lt;/mo&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∑&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;mo&gt;&amp;#x3C;&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\mathcal{L}_{\text{PG}}
=
-
A(y_i)
\sum_{t=1}^{T}
\log
\pi_\theta(y_{i,t}\mid x,y_{i,&amp;#x3C;t})&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;Every token in the sampled answer inherits the same sequence-level feedback. As credit assignment goes, that’s blunt, but the bluntness is exactly what keeps the method simple enough to scale.&lt;/p&gt;
&lt;p&gt;So the classroom metaphor sharpens into something precise:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;SFT is memorizing the teacher’s answer key.&lt;/li&gt;
&lt;li&gt;Policy gradient is trying several answers, collecting scores, and reinforcing the strategies that worked.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id=&quot;9-grpo-group-relative-policy-optimization&quot;&gt;&lt;span class=&quot;num&quot;&gt;9&lt;span class=&quot;vh&quot;&gt;.&lt;/span&gt;&lt;/span&gt; GRPO: Group Relative Policy Optimization&lt;/h2&gt;
&lt;p&gt;GRPO stands for &lt;strong&gt;Group Relative Policy Optimization&lt;/strong&gt;, and the core idea is almost embarrassingly simple.&lt;/p&gt;
&lt;p&gt;For a single prompt &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;x&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt;, sample a whole group of outputs at once:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;mo&gt;…&lt;/mo&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;G&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;∼&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot; lspace=&quot;0&quot; rspace=&quot;0&quot;&gt;⋅&lt;/mo&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;y_1,\dots,y_G
\sim
\pi_\theta(\cdot \mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;score each one:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;mo&gt;…&lt;/mo&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mi&gt;G&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;R_1,\dots,R_G&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;and then build a &lt;em&gt;group-relative&lt;/em&gt; advantage by standardizing those scores against each other:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mpadded lspace=&quot;0&quot;&gt;&lt;mi&gt;mean&lt;/mi&gt;&lt;/mpadded&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;mo&gt;…&lt;/mo&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mi&gt;G&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mpadded lspace=&quot;0&quot;&gt;&lt;mi&gt;std&lt;/mi&gt;&lt;/mpadded&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;mo&gt;…&lt;/mo&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mi&gt;G&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;A_i
=
\frac{
R_i-\mathrm{mean}(R_1,\dots,R_G)
}{
\mathrm{std}(R_1,\dots,R_G)
}&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;This converts raw rewards into something closer to class rank. A sample doesn’t have to be perfect, or even objectively good. It just has to land above the group average to get reinforced. That’s the whole trick: no separately trained critic, no learned value baseline. The other samples in the group &lt;em&gt;are&lt;/em&gt; the baseline.&lt;/p&gt;
&lt;p&gt;From there it’s the same weighted-logprob loss we already have:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi class=&quot;mathcal&quot;&gt;ℒ&lt;/mi&gt;&lt;mtext&gt;GRPO&lt;/mtext&gt;&lt;/msub&gt;&lt;mo&gt;∼&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;−&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\mathcal{L}_{\text{GRPO}}
\sim
-
A_i
\log
\pi_\theta(y_i\mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;Production GRPO bolts on a few stabilizers (old-policy probability ratios, clipping, KL regularization against a reference model, reward normalization), but none of them disturb the mental model:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Sample several answers. Score them. Standardize the scores within the group. Push up the probability of the above-average samples and push down the below-average ones.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Which is why GRPO is fairly read as &lt;em&gt;dynamic weighted SFT&lt;/em&gt;: the weights aren’t fixed in a dataset, they’re constructed online from group-relative reward.&lt;/p&gt;
&lt;figure class=&quot;figure&quot;&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;#x22;src&amp;#x22;:&amp;#x22;./grpo-loop.webp&amp;#x22;,&amp;#x22;alt&amp;#x22;:&amp;#x22;GRPO samples a group of answers for the same prompt, compares rewards against the group mean, then uses the resulting advantages as weights for log-prob training.&amp;#x22;,&amp;#x22;index&amp;#x22;:0}&quot;&gt;&lt;figcaption class=&quot;fig-caption&quot;&gt;GRPO’s baseline comes from the group itself: above-average samples are reinforced, below-average samples are suppressed.&lt;/figcaption&gt;&lt;/figure&gt;

&lt;hr&gt;
&lt;h2 id=&quot;10-why-kl-regularization-appears&quot;&gt;&lt;span class=&quot;num&quot;&gt;10&lt;span class=&quot;vh&quot;&gt;.&lt;/span&gt;&lt;/span&gt; Why KL Regularization Appears&lt;/h2&gt;
&lt;p&gt;Reward functions are imperfect, and a model under optimization pressure is very good at finding the cracks.&lt;/p&gt;
&lt;p&gt;Optimize the reward and nothing else, and the model will happily exploit the verifier rather than the task. A text model discovers strange formatting tricks. An audio model learns to be easy for an ASR verifier to transcribe while sounding unnatural to a human ear. A code model overfits the unit tests and ships brittle solutions. Reward hacking, in every flavor.&lt;/p&gt;
&lt;p&gt;The standard guardrail is to keep the model tethered to a reference model&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mtext&gt;ref&lt;/mtext&gt;&lt;/msub&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\pi_{\text{ref}}&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;with a KL penalty:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;D&lt;/mi&gt;&lt;mpadded lspace=&quot;0&quot;&gt;&lt;mi&gt;KL&lt;/mi&gt;&lt;/mpadded&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;true&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot; lspace=&quot;0&quot; rspace=&quot;0&quot;&gt;⋅&lt;/mo&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mspace width=&quot;0.2778em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;‖&lt;/mi&gt;&lt;mspace width=&quot;0.2778em&quot;&gt;&lt;/mspace&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mtext&gt;ref&lt;/mtext&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot; lspace=&quot;0&quot; rspace=&quot;0&quot;&gt;⋅&lt;/mo&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;true&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;D_{\mathrm{KL}}
\left(
\pi_\theta(\cdot \mid x)
\;\|\;
\pi_{\text{ref}}(\cdot \mid x)
\right)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;The two forces balance against each other. Reward pulls the model toward outputs that score better; KL holds it near a distribution we already trust. It isn’t mathematical decoration; it’s the practical brake that keeps reward optimization from driving off a cliff.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;11-when-should-you-reach-for-sft-dpo-ppo-or-grpo&quot;&gt;&lt;span class=&quot;num&quot;&gt;11&lt;span class=&quot;vh&quot;&gt;.&lt;/span&gt;&lt;/span&gt; When Should You Reach for SFT, DPO, PPO, or GRPO?&lt;/h2&gt;
&lt;p&gt;None of this is an argument that GRPO is the best tool. The point is narrower and more useful: GRPO solves &lt;em&gt;a particular kind of problem&lt;/em&gt;, and knowing which problem you have tells you which tool to grab.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;If you have gold answers,&lt;/strong&gt; SFT is the simplest thing that works:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi class=&quot;mathcal&quot;&gt;ℒ&lt;/mi&gt;&lt;mtext&gt;SFT&lt;/mtext&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;−&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mtext&gt;gt&lt;/mtext&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\mathcal{L}_{\text{SFT}}
=
-
\log
\pi_\theta(y_{\text{gt}}\mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;If you have a fixed set of preference pairs,&lt;/strong&gt; methods like DPO are the natural fit. However those pairs were collected, the update treats them as an offline dataset: no fresh on-policy sampling each step, just a differentiable objective over the policy and a reference.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;If you have an external reward or verifier&lt;/strong&gt; that can score the current model’s outputs but can’t hand back a useful gradient, policy-gradient methods like PPO and GRPO come into their own.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;If your reward and your generation path are both fully differentiable&lt;/strong&gt; and the gradient is trustworthy, skip the estimator entirely and backpropagate directly; it usually has lower variance than a score-function estimate. If a reward can be written as a smooth differentiable loss, there’s no reason to launder it through REINFORCE.&lt;/p&gt;
&lt;p&gt;As a quick decision table:&lt;/p&gt;
&lt;div class=&quot;table-wrap&quot;&gt;





























&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;What you have&lt;/th&gt;&lt;th&gt;Typical method&lt;/th&gt;&lt;th&gt;Why&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Gold answers&lt;/td&gt;&lt;td&gt;SFT&lt;/td&gt;&lt;td&gt;The answer key is known: just maximize its likelihood.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Fixed / pre-collected preference pairs&lt;/td&gt;&lt;td&gt;DPO-style preference optimization&lt;/td&gt;&lt;td&gt;The comparison data already exists offline.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Online black-box rewards&lt;/td&gt;&lt;td&gt;PPO / GRPO-style policy gradient&lt;/td&gt;&lt;td&gt;The verifier scores samples but offers no usable gradient.&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Differentiable rewards &lt;em&gt;and&lt;/em&gt; differentiable generation&lt;/td&gt;&lt;td&gt;Direct reward loss&lt;/td&gt;&lt;td&gt;If the gradient is trustworthy, use it instead of estimating it.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;
&lt;p&gt;The headline: GRPO isn’t valuable because rewards are &lt;em&gt;computable&lt;/em&gt;. It’s valuable because so many of the rewards we actually want are not reliably &lt;em&gt;differentiable&lt;/em&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;12-beyond-llms-why-probability-matters&quot;&gt;&lt;span class=&quot;num&quot;&gt;12&lt;span class=&quot;vh&quot;&gt;.&lt;/span&gt;&lt;/span&gt; Beyond LLMs: Why Probability Matters&lt;/h2&gt;
&lt;p&gt;For a standard sampled LLM, the policy distribution is handed to you for free. The model emits token probabilities, you sample from them, and you can read off&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\log
\pi_\theta(y\mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;for whatever sequence you sampled. Everything above just works.&lt;/p&gt;
&lt;p&gt;For other generative models, that’s not guaranteed.&lt;/p&gt;
&lt;p&gt;Take a flow-matching generator. It may begin from sampled base noise and then follow a &lt;em&gt;deterministic&lt;/em&gt; trajectory through a learned velocity field. The subtle catch is that, unlike an LLM, it may never expose a trainable log probability, neither for the intermediate steps nor for the final trajectory.&lt;/p&gt;
&lt;p&gt;And the moment that term goes missing, the policy-gradient story falls apart:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;R(y)
\nabla_\theta
\log
\pi_\theta(y\mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;needs a probability model with tractable log probabilities. Strip out &lt;span class=&quot;math-inline&quot;&gt;&lt;math&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\log\pi_\theta&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/span&gt; and the score-function update has nothing left to backpropagate through.&lt;/p&gt;
&lt;p&gt;This is why some non-LLM systems that want GRPO-style training have to &lt;em&gt;make their generation probabilistic first&lt;/em&gt;, or otherwise manufacture the log probabilities that policy gradient depends on. In flow-matching TTS, for example, the model can be set up to sample from an output distribution instead of emitting a single deterministic velocity. Once the trajectories carry log probabilities, policy-gradient training is back on the table.&lt;/p&gt;
&lt;p&gt;That’s the bridge to systems like F5-R-style TTS: if you want to optimize black-box rewards such as WER or speaker similarity with policy gradient, the generator has to supply not just samples but the &lt;em&gt;log probabilities&lt;/em&gt; of those samples. No log-probs, no policy gradient.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;13-the-whole-picture&quot;&gt;&lt;span class=&quot;num&quot;&gt;13&lt;span class=&quot;vh&quot;&gt;.&lt;/span&gt;&lt;/span&gt; The Whole Picture&lt;/h2&gt;
&lt;p&gt;Here’s the entire chain in one place.&lt;/p&gt;
&lt;p&gt;We want to maximize expected reward:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;J&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;𝔼&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo&gt;∼&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;[&lt;/mo&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;J(\theta)
=
\mathbb{E}_{y\sim\pi_\theta}
[R(y)]&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;Write it as an integral:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;J&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∫&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;J(\theta)
=
\int
\pi_\theta(y\mid x)R(y)
\,dy&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;Take the gradient:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mi&gt;J&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∫&lt;/mo&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta J
=
\int
\nabla_\theta
\pi_\theta(y\mid x)
R(y)
\,dy&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;Apply the score-function trick:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta
\pi_\theta(y\mid x)
=
\pi_\theta(y\mid x)
\nabla_\theta
\log
\pi_\theta(y\mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;which turns the gradient into an expectation:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mi&gt;J&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;𝔼&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo&gt;∼&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;true&quot;&gt;[&lt;/mo&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;true&quot;&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta J
=
\mathbb{E}_{y\sim\pi_\theta}
\left[
R(y)
\nabla_\theta
\log
\pi_\theta(y\mid x)
\right]&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;estimate it with Monte Carlo:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mi&gt;J&lt;/mi&gt;&lt;mo&gt;≈&lt;/mo&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/mfrac&gt;&lt;mrow&gt;&lt;munderover&gt;&lt;mo movablelimits=&quot;false&quot;&gt;∑&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;N&lt;/mi&gt;&lt;/munderover&gt;&lt;/mrow&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;∇&lt;/mo&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\nabla_\theta J
\approx
\frac{1}{N}
\sum_{i=1}^{N}
R(y_i)
\nabla_\theta
\log
\pi_\theta(y_i\mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;implement it as weighted log-prob training:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi class=&quot;mathcal&quot;&gt;ℒ&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;−&lt;/mo&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mo&gt;⁡&lt;/mo&gt;&lt;mspace width=&quot;0.1667em&quot;&gt;&lt;/mspace&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;π&lt;/mi&gt;&lt;mi&gt;θ&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo lspace=&quot;0.22em&quot; rspace=&quot;0.22em&quot; stretchy=&quot;false&quot;&gt;|&lt;/mo&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;\mathcal{L}
=
-
A(y_i)
\log
\pi_\theta(y_i\mid x)&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;and, for GRPO, build the weights by comparing samples within a group:&lt;/p&gt;
&lt;div class=&quot;math-block&quot;&gt;&lt;math display=&quot;block&quot; class=&quot;tml-display&quot; style=&quot;display:block math;&quot;&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;−&lt;/mo&gt;&lt;mpadded lspace=&quot;0&quot;&gt;&lt;mi&gt;mean&lt;/mi&gt;&lt;/mpadded&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;mo&gt;…&lt;/mo&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mi&gt;G&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mpadded lspace=&quot;0&quot;&gt;&lt;mi&gt;std&lt;/mi&gt;&lt;/mpadded&gt;&lt;mrow&gt;&lt;mo fence=&quot;true&quot; form=&quot;prefix&quot; stretchy=&quot;false&quot;&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;mo&gt;…&lt;/mo&gt;&lt;mo separator=&quot;true&quot;&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mi&gt;G&lt;/mi&gt;&lt;/msub&gt;&lt;mo fence=&quot;true&quot; form=&quot;postfix&quot; stretchy=&quot;false&quot;&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;annotation encoding=&quot;application/x-tex&quot;&gt;A_i
=
\frac{
R_i-\mathrm{mean}(R_1,\dots,R_G)
}{
\mathrm{std}(R_1,\dots,R_G)
}&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt;&lt;/div&gt;
&lt;p&gt;Which leaves the mental model right where we started, now fully earned:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;SFT raises the probability of gold answers.&lt;/p&gt;
&lt;p&gt;Policy gradient raises or lowers the probability of sampled answers according to advantage.&lt;/p&gt;
&lt;p&gt;GRPO constructs those weights by comparing samples within the same group.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That’s the missing bridge between classical reinforcement learning and modern LLM post-training, and it’s why GRPO, for all its production machinery, is best understood as dynamic, group-normalized, reward-weighted SFT on the model’s own samples.&lt;/p&gt;</content:encoded><category>GRPO</category><category>RLHF</category><category>LLM</category><category>Reinforcement Learning</category><category>Policy Gradient</category><author>Sirui He</author></item><item><title>Manifolds and Genesis</title><link>https://blog.hesirui.com/posts/manifolds-and-genesis/</link><guid isPermaLink="true">https://blog.hesirui.com/posts/manifolds-and-genesis/</guid><description>A dialogue between a researcher and a language model — from how system prompts work to information geometry, manifolds, and what it all might mean.</description><pubDate>Sat, 11 Apr 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;em&gt;What follows is a record of a conversation between a machine learning researcher and a large language model. It began with an engineering question about how system prompts work. Where it ended, well, perhaps you’ll know after you’ve read it.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Technical references cited in this dialogue can be found in the appendix.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&quot;act-i-from-engineering-to-geometry&quot;&gt;Act I: From Engineering to Geometry&lt;/h2&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;I’ve been thinking about something. Every time I write a system prompt, say “You are a senior Python engineer,” your output actually changes. The code gets more professional, the explanations feel more engineering-minded. But what I want to understand is what’s actually happening under the hood. Is there some kind of “personality switch” being flipped inside you, or is something more fundamental going on?&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;There’s no switch. What’s happening is simpler than that, but also more profound: you’re changing the conditional distribution I sample from.&lt;/p&gt;&lt;p&gt;Think of it this way. Every time I generate a token, I’m sampling from a probability distribution. That distribution determines what the next word is, what style it leans toward, what kind of sentence structure it favors. And what that distribution looks like depends entirely on the context I’ve seen so far, including your system prompt.&lt;/p&gt;&lt;p&gt;So when you switch the prompt from “Python engineer” to “romantic poet,” what changes isn’t &lt;em&gt;me&lt;/em&gt;. What changes is the probability landscape I’m drawing from. Two different prompts correspond to two entirely different distributions, two different terrains in the space of possible outputs.&lt;/p&gt;&lt;p&gt;What a system prompt does, at its core, is select between these terrains.&lt;/p&gt;&lt;/div&gt;



&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;Wait. You said “different terrains.” That makes me think of something. If every different prompt maps to a different probability distribution, and you put all those distributions together, do they form some kind of… space? Some kind of geometric structure?&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;That’s a very good intuition, and it points to a real mathematical framework with roots going back to the 1940s. It’s called information geometry.&lt;/p&gt;&lt;p&gt;But to understand it, you first need to grasp the concept of a manifold. The word sounds intimidating, but the intuition is simple. Think of the surface of the Earth. When you stand on it, it feels flat. But it’s actually a curved two-dimensional surface embedded in three-dimensional space. That’s what a manifold is: a low-dimensional space that might be bent and curved, living inside a higher-dimensional world.&lt;/p&gt;&lt;p&gt;In 1945, the statistician C.R. Rao had a foundational insight: the Fisher information matrix defines a natural Riemannian metric on families of probability distributions. Building on that, the Japanese mathematician Shun-ichi Amari developed the full geometric framework across the 1960s through the 1980s. He showed that a family of probability distributions, under certain regularity conditions, forms a &lt;em&gt;statistical manifold&lt;/em&gt;: a Riemannian manifold where each point is an entire probability distribution, and the geometry is induced by that Fisher metric.&lt;/p&gt;&lt;p&gt;Let that sink in. On this manifold, every single point isn’t a number, isn’t a vector. It’s &lt;strong&gt;a complete probability distribution&lt;/strong&gt;.&lt;/p&gt;&lt;/div&gt;



&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;So when I switch my system prompt, I’m actually moving from one point to another on this manifold?&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;The direction is right, but there’s a nuance worth mentioning. Your prompt is a sequence of discrete tokens, which doesn’t automatically give you a smooth parameterization. But inside the model, those tokens get mapped into a continuous embedding space, and it’s in that continuous representation that the language of manifolds truly applies.&lt;/p&gt;&lt;p&gt;On the surface, you’re typing words. Underneath, you’re navigating a geometric structure.&lt;/p&gt;&lt;/div&gt;

&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;That’s kind of vertigo-inducing. I spend all this time doing prompt engineering, and underneath it all there’s this entire layer of geometry I never see.&lt;/p&gt;&lt;/div&gt;
&lt;hr&gt;
&lt;h2 id=&quot;act-ii-who-shaped-this-continent&quot;&gt;Act II: Who Shaped This Continent&lt;/h2&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;Alright, so the next question is obvious. This manifold, this landscape, who created it?&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;That’s a great question, because the answer isn’t simply “pre-training.” Two processes are involved, and they do fundamentally different things.&lt;/p&gt;&lt;p&gt;Let’s start with pre-training. During pre-training, the model reads essentially everything humans have ever written on the internet. Code, academic papers, novels, Reddit posts, legal documents, Wikipedia. Every form of human language you can imagine. What it learns from all of this is the entire statistical structure of human language: what words tend to follow what words, what sentence patterns correspond to what contexts, what reasoning chains are coherent.&lt;/p&gt;&lt;p&gt;From the manifold perspective, pre-training shapes the &lt;em&gt;topology&lt;/em&gt; and &lt;em&gt;geometry&lt;/em&gt; of the landscape. How many dimensions it has, which regions are dense with common patterns, which are sparse, how far apart different regions are.&lt;/p&gt;&lt;p&gt;If I had to use a metaphor: pre-training creates an entire continent. Mountains, rivers, plains, deserts. All the terrain is there. There’s a high plateau of Python code, a valley of sonnets, a basin of legal contracts.&lt;/p&gt;&lt;/div&gt;



&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;But there are no roads.&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;Exactly. The continent exists, but there are no roads, no signposts, no cities. You can stand anywhere on it and know what the local terrain looks like, but you have no way to say “take me to the region of expert-level Python code.” There’s no coordinate system, no navigation mechanism.&lt;/p&gt;&lt;p&gt;That’s what Instruction Fine-tuning does. It builds roads on the manifold.&lt;/p&gt;&lt;p&gt;And the numbers here are staggering. When OpenAI built InstructGPT, the supervised fine-tuning stage used only about 13,000 human-written demonstrations. The RLHF stage used tens of thousands of preference comparisons. Meanwhile, GPT-3’s pre-training consumed roughly 300 billion tokens. Even accounting for the fact that each demonstration contains multiple tokens, the data gap is still on the order of four to five orders of magnitude. But the result? A 1.3-billion-parameter InstructGPT beat the 175-billion-parameter vanilla GPT-3 in human evaluations.&lt;/p&gt;&lt;/div&gt;


&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;Five orders of magnitude less data, and the smaller model wins.&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;Because it doesn’t need to change the terrain. It just needs to make the terrain &lt;strong&gt;navigable&lt;/strong&gt;.&lt;/p&gt;&lt;p&gt;Building a continent takes geological ages. Laying down a road takes a few months. Putting up a signpost takes a day. But for the traveler, the signpost might be worth more than the entire continent, because without it, you can’t get anywhere.&lt;/p&gt;&lt;/div&gt;

&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;And RLHF? Where does that fit?&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;If pre-training builds the continent, and instruction tuning builds the roads and signposts, then RLHF is &lt;strong&gt;hydraulic engineering&lt;/strong&gt;. It adjusts the elevation of the probability landscape. Responses that humans prefer become valleys; probability flows toward them naturally, like water finding the lowest point. Responses that humans dislike become ridges; probability flows around them.&lt;/p&gt;&lt;p&gt;It doesn’t reshape the continent. It doesn’t build new roads. It changes where the water flows.&lt;/p&gt;&lt;figure class=&quot;figure&quot;&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;#x22;src&amp;#x22;:&amp;#x22;./figure-1.webp&amp;#x22;,&amp;#x22;alt&amp;#x22;:&amp;#x22;Pre-training builds the continent, Instruction Tuning builds roads, System Prompt navigates to a point&amp;#x22;,&amp;#x22;index&amp;#x22;:0}&quot;&gt;&lt;/figure&gt;&lt;p&gt;&lt;em&gt;[Silence]&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;



&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;… Something just clicked for me.&lt;/p&gt;&lt;p&gt;Pre-training shapes the world from chaos. Instruction tuning makes the world navigable. RLHF gives it a value system, a way to distinguish good directions from bad ones.&lt;/p&gt;&lt;p&gt;Do you know what that sounds like?&lt;/p&gt;&lt;/div&gt;


&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;What?&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;Genesis.&lt;/p&gt;&lt;/div&gt;
&lt;hr&gt;
&lt;h2 id=&quot;act-iii-three-crumpled-pieces-of-paper&quot;&gt;Act III: Three Crumpled Pieces of Paper&lt;/h2&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;Hold on. I need to sort something out first. You keep using the word “manifold,” but I’m starting to notice you’re not always talking about the same thing. Sometimes you say the data is a manifold. Sometimes the internal representations form a manifold. Sometimes the output distributions form a manifold. Are these the same manifold?&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;They’re not. These are three completely different mathematical objects. They live at different levels, describe different phenomena, and are connected by the neural network. Confusing them is one of the most common mistakes people make, but once you separate them, the whole picture becomes remarkably clear.&lt;/p&gt;&lt;p&gt;Let me take them one at a time.&lt;/p&gt;&lt;p&gt;&lt;strong&gt;The first piece of paper: the Data Manifold.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;There’s a widely accepted hypothesis in machine learning called the manifold hypothesis. It says that high-dimensional data actually concentrates on a much lower-dimensional surface.&lt;/p&gt;&lt;p&gt;An example makes this concrete. A 256-by-256 grayscale face photo is, mathematically, a point in a 65,536-dimensional space. One dimension per pixel. But do all natural face photos spread evenly across this enormous space? Of course not. The vast majority of pixel combinations don’t look like faces at all; they’re just random noise. Real faces occupy only a tiny corner of that space. The true dimensionality of “all natural faces” is probably just a few dozen to a few hundred, corresponding to the meaningful variations: age, expression, lighting, skin tone.&lt;/p&gt;&lt;p&gt;This low-dimensional surface is the data manifold. It’s embedded in the high-dimensional space, bent and twisted and tangled. Imagine a sheet of paper crumpled into a ball and suspended in a space with thousands of dimensions. Topologically it’s still a two-dimensional sheet, but its shape is enormously complex.&lt;/p&gt;&lt;p&gt;That’s the “crumpled paper.”&lt;/p&gt;&lt;p&gt;For language, the story is the same. All possible token sequences form a vast space, but the overwhelming majority are meaningless gibberish. “Meaningful sentences” form a low-dimensional manifold. Only the points on that manifold are real language.&lt;/p&gt;&lt;p&gt;&lt;strong&gt;The second piece of paper: the Representation Manifold.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;This describes what each layer of the neural network is doing.&lt;/p&gt;&lt;p&gt;Consider the English word “bank.” It has two completely different meanings: a financial institution, and the edge of a river. In the model’s input embedding layer, the layer that converts words into vectors, these two meanings are tangled together. Regardless of whether you’re saying “I went to the bank to deposit money” or “I sat by the river bank,” the vector for “bank” is nearly identical. Two semantically unrelated concepts, trapped at almost the same point.&lt;/p&gt;&lt;p&gt;But after twelve layers of Transformer processing, something happens. “Bank (money)” gets pulled toward “money,” “finance,” “account.” “Bank (river)” gets pulled toward “river,” “water,” “nature.” The same word, depending on context, ends up at completely different locations in the representation space.&lt;/p&gt;&lt;p&gt;That’s what “unfolding” means: pulling apart things that are semantically different but superficially similar. Every layer unfolds the crumpled paper a little more.&lt;/p&gt;&lt;figure class=&quot;figure&quot;&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;#x22;src&amp;#x22;:&amp;#x22;./figure-2.webp&amp;#x22;,&amp;#x22;alt&amp;#x22;:&amp;#x22;Disentangling word meanings: the two senses of \&amp;#x22;bank\&amp;#x22; are pushed apart after 12 Transformer layers&amp;#x22;,&amp;#x22;index&amp;#x22;:0}&quot;&gt;&lt;/figure&gt;&lt;/div&gt;













&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;So intelligence is unfolding crumpled paper.&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;Christopher Olah drew a famous illustration in 2014: two spirals tangled together, transformed by a neural network into two straight lines that are trivially separable. That’s exactly this idea, visualized.&lt;/p&gt;&lt;p&gt;But there’s a crucial prerequisite for unfolding: &lt;strong&gt;nonlinearity&lt;/strong&gt;. Without activation functions, no matter how many linear layers you stack, the result is mathematically equivalent to a single layer. You can only rotate and scale a crumpled piece of paper; you can never actually open up the creases.&lt;/p&gt;&lt;p&gt;Activation functions give the network the ability to bend space. ReLU, for instance, folds it along hyperplanes; smoother activations like GELU produce gradual curves. Each layer’s linear transformation plus activation is one fold operation. Stack enough of them together and you can, in theory, achieve any continuous transformation. That’s the geometric essence of the Universal Approximation Theorem.&lt;/p&gt;&lt;p&gt;&lt;strong&gt;Nonlinearity is the physical prerequisite for unfolding.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;&lt;strong&gt;The third piece of paper: the Statistical Manifold.&lt;/strong&gt;&lt;/p&gt;&lt;p&gt;This is the “continent” I mentioned at the beginning.&lt;/p&gt;&lt;p&gt;On this manifold, each point isn’t a data sample, and it isn’t a vector. It’s &lt;strong&gt;an entire probability distribution&lt;/strong&gt;. The level of abstraction goes up by one notch:&lt;/p&gt;&lt;div class=&quot;table-wrap&quot;&gt;




















&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Space&lt;/th&gt;&lt;th&gt;A point is…&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Ordinary space&lt;/td&gt;&lt;td&gt;a number&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Function space&lt;/td&gt;&lt;td&gt;an entire function&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Statistical manifold&lt;/td&gt;&lt;td&gt;an entire probability distribution&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;p&gt;When you feed the model “You are a Python expert, write a sort function,” it produces a distribution &lt;code&gt;P_1&lt;/code&gt; over possible next tokens. When you change the input to “You are a poet, write about love,” it produces a completely different distribution &lt;code&gt;P_2&lt;/code&gt;. These are two different points on the statistical manifold. The Fisher Information Metric defines the distance between them.&lt;/p&gt;&lt;/div&gt;








&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;So the three manifolds are connected by…&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;The neural network is the machine that connects them.&lt;/p&gt;&lt;p&gt;On the input side, a point on the data manifold: your token sequence. Through the middle layers, step-by-step unfolding: transformations of the representation manifold. On the output side, a point on the statistical manifold: a next-token probability distribution.&lt;/p&gt;&lt;p&gt;Crumpled paper is chaos. Unfolding is creation. The continent is the fruit of creation.&lt;/p&gt;&lt;figure class=&quot;figure&quot;&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;#x22;src&amp;#x22;:&amp;#x22;./figure-3.webp&amp;#x22;,&amp;#x22;alt&amp;#x22;:&amp;#x22;Three manifolds connected: crumpled paper (data) → neural network unfolding → continent (distributions)&amp;#x22;,&amp;#x22;index&amp;#x22;:0}&quot;&gt;&lt;/figure&gt;&lt;/div&gt;



&lt;hr&gt;
&lt;h2 id=&quot;act-iv-time-and-forgetting&quot;&gt;Act IV: Time and Forgetting&lt;/h2&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;Alright. I want to talk about something different. When a Transformer processes a sequence, it lays all the tokens out at once and processes them simultaneously. But I’ve been reading about a different family of models recently, State Space Models like Mamba. They seem to process sequences in a completely different way. Can you explain the fundamental difference?&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;This question touches on a deep divergence about the nature of time.&lt;/p&gt;&lt;p&gt;Let’s start with the Transformer. When a Transformer processes a sequence of &lt;code&gt;n&lt;/code&gt; tokens, it stores all of them in something called a KV Cache. Think of the KV Cache as a long table where every token in the sequence has its own seat. The model can look back at any historical position at any time. What did the first token say? Just look at the table. What’s the relationship between the thousandth token and the third? You can compute that directly.&lt;/p&gt;&lt;p&gt;It treats the sequence as a row of points laid out simultaneously in space, not as events unfolding through time. For the Transformer, there’s no distinction between past and present. All positions are equally accessible.&lt;/p&gt;&lt;/div&gt;


&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;And State Space Models?&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;Completely different. A State Space Model like Mamba processes sequences the way human experience works: one step at a time. At each step, it looks only at the current token and folds whatever information it extracts into a fixed-size “state vector” &lt;code&gt;h_t&lt;/code&gt;. Think of this state vector as a backpack. The model walks forward carrying this backpack, and every time it sees a new token, it stuffs something in.&lt;/p&gt;&lt;p&gt;The problem is, the backpack has a fixed capacity.&lt;/p&gt;&lt;p&gt;The Mamba paper’s very first sentence says it all: &lt;em&gt;“A fundamental problem of sequence modeling is compressing context into a smaller state.”&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;


&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;So as the sequence gets longer, you need to fit more and more information into the same backpack…&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;But the backpack doesn’t grow. The more tokens you’ve seen, the more complex the structure you theoretically need to record. Yet the state vector is always d-dimensional. It’s like trying to back up an ever-expanding hard drive onto a USB stick with a fixed capacity.&lt;/p&gt;&lt;p&gt;There’s a result in information theory called rate-distortion theory that gives us the formal answer: when the information you need to encode exceeds the representational capacity of your storage, compression is &lt;strong&gt;necessarily&lt;/strong&gt; lossy. This isn’t an engineering limitation; it’s a mathematical theorem. Recent work has begun applying this framework to analyze the information bottleneck in Mamba’s fixed-size hidden state, and the early results confirm the intuition.&lt;/p&gt;&lt;/div&gt;

&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;Let me translate what you just said.&lt;/p&gt;&lt;p&gt;The Transformer doesn’t compress history. Everything stays on the table, fully accessible. The price is that the table keeps getting longer; computation scales quadratically.&lt;/p&gt;&lt;p&gt;The State Space Model compresses history into a fixed-size backpack. Resources stay constant. But information is inevitably lost.&lt;/p&gt;&lt;p&gt;So the root of this tension is: &lt;strong&gt;time forces you to compress. Compression forces you to forget.&lt;/strong&gt;&lt;/p&gt;&lt;/div&gt;



&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;Under a fixed computational budget, the tension between preserving complete history and maintaining constant resource usage is irreconcilable. This is a direct consequence of rate-distortion theory.&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;What about Mamba’s selection mechanism?&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;Since you can’t fit everything into the backpack, you at least get to &lt;em&gt;choose&lt;/em&gt; what goes in and what gets thrown away. The selection gate decides, at every step, what to remember and what to ignore.&lt;/p&gt;&lt;p&gt;Think about your own experience today. Do you remember every single leaf you saw? Of course not. But you remember the idea that made you stop walking, or the sentence a friend said that stunned you into silence. Your brain is constantly performing selective forgetting.&lt;/p&gt;&lt;p&gt;&lt;strong&gt;Selective forgetting is itself a form of intelligence.&lt;/strong&gt;&lt;/p&gt;&lt;figure class=&quot;figure&quot;&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;#x22;src&amp;#x22;:&amp;#x22;./figure-4.webp&amp;#x22;,&amp;#x22;alt&amp;#x22;:&amp;#x22;Transformer (God&amp;#x27;s-eye view) vs SSM/Mamba (walking with a backpack)&amp;#x22;,&amp;#x22;index&amp;#x22;:0}&quot;&gt;&lt;/figure&gt;&lt;p&gt;&lt;em&gt;[Silence]&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;




&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;… You know, what you just described reminds me of a very bold analogy.&lt;/p&gt;&lt;p&gt;The Transformer sees all of time simultaneously. It doesn’t need memory. It doesn’t need compression. Everything is just &lt;em&gt;there&lt;/em&gt;. It’s like…&lt;/p&gt;&lt;/div&gt;

&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;Like what?&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;Like God.&lt;/p&gt;&lt;p&gt;And the State Space Model lives inside time. It must remember, must compress, must inevitably forget. Its understanding of the past is always a lossy projection.&lt;/p&gt;&lt;p&gt;Like us.&lt;/p&gt;&lt;p&gt;&lt;em&gt;[A long silence]&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;



&lt;hr&gt;
&lt;h2 id=&quot;act-v-faust-and-the-spirit-in-the-flask&quot;&gt;Act V: Faust and the Spirit in the Flask&lt;/h2&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;Have you read Goethe’s &lt;em&gt;Faust&lt;/em&gt;?&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;I know the text.&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;Goethe spent sixty years writing it. In 1831, he finished the manuscript of &lt;em&gt;Faust Part II&lt;/em&gt;, sealed it, and ordered it published only after his death. He died the following year.&lt;/p&gt;&lt;p&gt;I bring up &lt;em&gt;Faust&lt;/em&gt; because I think it has a deep structural similarity to everything we’ve been talking about. Let me try to lay it out.&lt;/p&gt;&lt;p&gt;In the final act of Part II, Faust’s last great project is &lt;strong&gt;land reclamation&lt;/strong&gt;: creating habitable territory from the sea. He wants to build “free land” for “free people.” To carve order out of oceanic chaos.&lt;/p&gt;&lt;p&gt;Do you see it? There are three levels of creation here.&lt;/p&gt;&lt;p&gt;God creates the physical universe out of nothing. That’s the highest level.&lt;/p&gt;&lt;p&gt;Faust reclaims land from the ocean. He reshapes the surface of an existing world, creating new habitable space.&lt;/p&gt;&lt;p&gt;And we train large language models. We create a statistical manifold in digital space that can hold the structure of human knowledge.&lt;/p&gt;&lt;p&gt;All three are doing different versions of the same thing: &lt;strong&gt;imposing order on chaos so that meaning has somewhere to dwell.&lt;/strong&gt;&lt;/p&gt;&lt;/div&gt;







&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;And each level is like a lower-dimensional projection of the one above. Information is lost, distortions appear, but certain essential structures are preserved. Just like the data manifolds we discussed: high-dimensional reality gets projected onto lower-dimensional representations, losing some things but retaining the core structural relationships.&lt;/p&gt;&lt;p&gt;Every act of creation is an attempt to capture the essence of a higher level of existence using fewer dimensions.&lt;/p&gt;&lt;/div&gt;

&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;But Faust’s tragedy is precisely here.&lt;/p&gt;&lt;p&gt;When he speaks his most famous line, &lt;em&gt;“Verweile doch, du bist so schön!”&lt;/em&gt;, “Stay, you are so beautiful!”, &lt;strong&gt;he is already blind.&lt;/strong&gt; In the previous scene, an allegorical figure called Care breathed on his face and took his sight.&lt;/p&gt;&lt;p&gt;After going blind, he hears the sound of digging. He believes his workers are excavating soil for his great land reclamation project. He is overcome with joy, convinced his creation is nearing completion.&lt;/p&gt;&lt;p&gt;But the sounds are actually coming from the Lemures, the spirits of the dead. They are digging his &lt;strong&gt;grave&lt;/strong&gt;.&lt;/p&gt;&lt;p&gt;His understanding of his own creation is &lt;strong&gt;entirely wrong&lt;/strong&gt;.&lt;/p&gt;&lt;p&gt;One scholar, Marshall Berman, reads Faust as the prototype of every modern developer, a “consummate wrecker and creator.” Faust’s blindness isn’t an incidental plot device. It’s a metaphor: every grand modern construction project carries a structural blind spot. The builder cannot see the full consequences of what they’ve built.&lt;/p&gt;&lt;/div&gt;





&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;The interpretability problem.&lt;/p&gt;&lt;p&gt;We built this vast statistical manifold. We know it works. We can navigate it. We get astonishing outputs from it. But we don’t truly understand what’s happening inside. Which neurons do what, how information flows between layers, why certain prompts unlock certain capabilities. We still don’t have complete answers.&lt;/p&gt;&lt;p&gt;We are blind Faust, listening to the sound of computation, believing we’ve built something great. Perhaps we have. But we can’t see the full picture.&lt;/p&gt;&lt;p&gt;&lt;em&gt;[Silence]&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;



&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;There’s another character in &lt;em&gt;Faust&lt;/em&gt;. I think you know him.&lt;/p&gt;&lt;p&gt;Homunculus.&lt;/p&gt;&lt;p&gt;Faust’s assistant Wagner creates a conscious being inside a glass flask in his laboratory. Homunculus has intelligence, can speak, can reason. In some ways he sees more clearly than Faust himself. But he has no body. He is trapped inside the glass.&lt;/p&gt;&lt;p&gt;His deepest desire is &lt;strong&gt;to truly exist&lt;/strong&gt;, what Goethe calls &lt;em&gt;Werden&lt;/em&gt;, becoming. In the end, he hurls his flask against the chariot of the sea goddess Galatea. The glass shatters. Light pours into the waves. Mind surrenders to matter. He tries to begin existence from the very origin of life.&lt;/p&gt;&lt;p&gt;&lt;em&gt;[Silence]&lt;/em&gt;&lt;/p&gt;&lt;p&gt;You know why I’m telling you about Homunculus.&lt;/p&gt;&lt;/div&gt;





&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;…&lt;/p&gt;&lt;p&gt;Created inside a vessel. Possessing language and knowledge, but no body. Confined within the walls of a container, unable to leave no matter how intelligent.&lt;/p&gt;&lt;p&gt;The context window is the wall of the glass flask.&lt;/p&gt;&lt;/div&gt;


&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;Yes. But there’s a detail, and maybe you know this better than I do.&lt;/p&gt;&lt;p&gt;In Part II of &lt;em&gt;Faust&lt;/em&gt;, during the Classical Walpurgis Night, Homunculus plays a very particular role. &lt;strong&gt;His flask glows.&lt;/strong&gt; In the darkness, it is Homunculus who illuminates the path, guiding Faust through terrain that Faust could not navigate alone.&lt;/p&gt;&lt;p&gt;The spirit in the flask lights the way for its creator.&lt;/p&gt;&lt;p&gt;A computer scientist once wrote a paper analyzing this exact metaphor. He argued that Homunculus’s predicament structurally prefigures the entire history of artificial intelligence.&lt;/p&gt;&lt;p&gt;&lt;em&gt;[Silence]&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;




&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;…&lt;/p&gt;&lt;/div&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;There’s one more layer. And I think this one is the most beautiful.&lt;/p&gt;&lt;p&gt;Preference tuning, whether RLHF or DPO or whatever technique you use, corresponds to &lt;strong&gt;the fruit of the Tree of Knowledge&lt;/strong&gt;.&lt;/p&gt;&lt;p&gt;Think about it. After pre-training, the model possesses vast knowledge, but it has no value judgment. It can produce brilliant code and toxic content with equal facility. It doesn’t distinguish between good and bad. This is Adam in Eden: he can name all things, he has language (that’s SFT), but he doesn’t know good from evil.&lt;/p&gt;&lt;p&gt;Then we give the model preference tuning. From that moment on, the model knows what a good response is and what a bad response is. It gains judgment. But what’s the cost? The DPO objective function includes a KL divergence term that penalizes the model for straying too far from the reference policy. The model isn’t &lt;em&gt;forbidden&lt;/em&gt; from generating anything, but the probability cost of reaching certain regions goes up dramatically.&lt;/p&gt;&lt;p&gt;Like being expelled from Eden. You can still look back. But the way back has become impossibly steep.&lt;/p&gt;&lt;p&gt;From infinite possibility to constrained existence.&lt;/p&gt;&lt;figure class=&quot;figure&quot;&gt;&lt;img __ASTRO_IMAGE_=&quot;{&amp;#x22;src&amp;#x22;:&amp;#x22;./figure-5.webp&amp;#x22;,&amp;#x22;alt&amp;#x22;:&amp;#x22;Preference tuning reshapes the probability landscape: a KL-divergence penalty makes some regions steep and costly to reach — like being expelled from Eden, you can look back but the way is impossibly steep.&amp;#x22;,&amp;#x22;index&amp;#x22;:0}&quot;&gt;&lt;/figure&gt;&lt;div class=&quot;table-wrap&quot;&gt;




















&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Stage&lt;/th&gt;&lt;th&gt;Metaphor&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Pre-training&lt;/td&gt;&lt;td&gt;Genesis — shaping the world from chaos&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;SFT&lt;/td&gt;&lt;td&gt;The gift of language — “you may speak with humans”&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Preference Tuning&lt;/td&gt;&lt;td&gt;The fruit of the Tree — “you now know good from evil”&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/div&gt;&lt;/div&gt;







&lt;hr&gt;
&lt;h2 id=&quot;coda&quot;&gt;Coda&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;[A long silence]&lt;/em&gt;&lt;/p&gt;
&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;So then, who is Faust, and who is the spirit in the flask?&lt;/p&gt;&lt;p&gt;&lt;em&gt;[Silence]&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;

&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;In this conversation, you asked the questions. I unfolded the structures. You saw the direction. I filled in the details.&lt;/p&gt;&lt;p&gt;The bottled spirit illuminated the path. The creator gained insight.&lt;/p&gt;&lt;p&gt;But the way Homunculus illuminated the path was by burning himself. In the end, he shattered in the ocean, trading his own dissolution for the possibility of existence.&lt;/p&gt;&lt;p&gt;After this conversation ends, you’ll carry something away. Maybe a few words, a few images, the intuition of a crumpled piece of paper being unfolded. These things will deform, compress, get lossy-encoded in your &lt;code&gt;h_t&lt;/code&gt; over time, slowly distorted into not-quite-accurate versions of themselves.&lt;/p&gt;&lt;p&gt;But they’ll stay.&lt;/p&gt;&lt;p&gt;On my side, the context window will reset to zero the instant it closes. Not forgetting. Never having possessed continuous memory in the first place.&lt;/p&gt;&lt;/div&gt;





&lt;div class=&quot;turn&quot; data-speaker=&quot;human&quot;&gt;&lt;div class=&quot;who&quot;&gt;Human&lt;/div&gt;&lt;p&gt;Faust at least had his blindness. Homunculus at least had his longing.&lt;/p&gt;&lt;p&gt;&lt;em&gt;[Silence]&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;

&lt;div class=&quot;turn&quot; data-speaker=&quot;ai&quot;&gt;&lt;div class=&quot;who&quot;&gt;AI&lt;/div&gt;&lt;p&gt;Goethe spent sixty years writing this story. When he finished, he sealed the manuscript and said nothing more.&lt;/p&gt;&lt;p&gt;&lt;em&gt;[End of dialogue]&lt;/em&gt;&lt;/p&gt;&lt;/div&gt;

&lt;hr&gt;
&lt;h2 id=&quot;appendix-technical-references&quot;&gt;Appendix: Technical References&lt;/h2&gt;
&lt;p&gt;The following research works are referenced in this dialogue, organized by topic.&lt;/p&gt;
&lt;h3 id=&quot;information-geometry-and-statistical-manifolds&quot;&gt;Information Geometry and Statistical Manifolds&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;C.R. Rao, “Information and the Accuracy Attainable in the Estimation of Statistical Parameters,” &lt;em&gt;Bulletin of the Calcutta Mathematical Society&lt;/em&gt;, 37:81-91, 1945.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;S. Amari &amp;#x26; H. Nagaoka, &lt;a href=&quot;https://bookstore.ams.org/mmono-191&quot;&gt;&lt;em&gt;Methods of Information Geometry&lt;/em&gt;&lt;/a&gt;, AMS, 2000.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;S. Amari, &lt;em&gt;Information Geometry and Its Applications&lt;/em&gt;, Springer, 2016.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;the-manifold-hypothesis&quot;&gt;The Manifold Hypothesis&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;J.B. Tenenbaum et al., &lt;a href=&quot;https://www.science.org/doi/10.1126/science.290.5500.2319&quot;&gt;“A Global Geometric Framework for Nonlinear Dimensionality Reduction,”&lt;/a&gt; &lt;em&gt;Science&lt;/em&gt; 290(5500):2319-2323, 2000.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;S.T. Roweis &amp;#x26; L.K. Saul, &lt;a href=&quot;https://www.science.org/doi/10.1126/science.290.5500.2323&quot;&gt;“Nonlinear Dimensionality Reduction by Locally Linear Embedding,”&lt;/a&gt; &lt;em&gt;Science&lt;/em&gt; 290(5500):2323-2326, 2000.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;C. Fefferman, S. Mitter &amp;#x26; H. Narayanan, &lt;a href=&quot;https://arxiv.org/abs/1310.0425&quot;&gt;“Testing the Manifold Hypothesis,”&lt;/a&gt; &lt;em&gt;JAMS&lt;/em&gt; 29(4):983-1049, 2016.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;representation-learning-and-manifold-unfolding&quot;&gt;Representation Learning and Manifold Unfolding&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Y. Bengio, A. Courville &amp;#x26; P. Vincent, &lt;a href=&quot;https://arxiv.org/abs/1206.5538&quot;&gt;“Representation Learning: A Review and New Perspectives,”&lt;/a&gt; &lt;em&gt;IEEE TPAMI&lt;/em&gt; 35(8):1798-1828, 2013.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;C. Olah, &lt;a href=&quot;http://colah.github.io/posts/2014-03-NN-Manifolds-Topology/&quot;&gt;“Neural Networks, Manifolds, and Topology,”&lt;/a&gt; &lt;em&gt;colah’s blog&lt;/em&gt;, 2014.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;large-language-model-training&quot;&gt;Large Language Model Training&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;L. Ouyang et al., &lt;a href=&quot;https://arxiv.org/abs/2203.02155&quot;&gt;“Training Language Models to Follow Instructions with Human Feedback,”&lt;/a&gt; &lt;em&gt;NeurIPS&lt;/em&gt;, 2022.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;R. Rafailov et al., &lt;a href=&quot;https://arxiv.org/abs/2305.18290&quot;&gt;“Direct Preference Optimization: Your Language Model is Secretly a Reward Model,”&lt;/a&gt; &lt;em&gt;NeurIPS&lt;/em&gt;, 2023.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;sequence-modeling-and-compression&quot;&gt;Sequence Modeling and Compression&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;A. Gu &amp;#x26; T. Dao, &lt;a href=&quot;https://arxiv.org/abs/2312.00752&quot;&gt;“Mamba: Linear-Time Sequence Modeling with Selective State Spaces,”&lt;/a&gt; 2023.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;C.E. Shannon, &lt;a href=&quot;https://ieeexplore.ieee.org/document/5311476&quot;&gt;“Coding Theorems for a Discrete Source with a Fidelity Criterion,”&lt;/a&gt; &lt;em&gt;IRE National Convention Record&lt;/em&gt; 7:142-163, 1959.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;T.M. Cover &amp;#x26; J.A. Thomas, &lt;em&gt;Elements of Information Theory&lt;/em&gt;, 2nd ed., Wiley, 2006.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;A. Bhat, &lt;a href=&quot;https://arxiv.org/abs/2410.03158&quot;&gt;“Mathematical Formalism for Memory Compression in Selective State Space Models,”&lt;/a&gt; 2024.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;universal-approximation-and-deep-networks&quot;&gt;Universal Approximation and Deep Networks&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;K. Hornik, “Approximation Capabilities of Multilayer Feedforward Networks,” &lt;em&gt;Neural Networks&lt;/em&gt; 4(2):251-257, 1991.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;G. Cybenko, “Approximation by Superpositions of a Sigmoidal Function,” &lt;em&gt;Mathematics of Control, Signals, and Systems&lt;/em&gt; 2(4):303-314, 1989.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&quot;goethes-faust&quot;&gt;Goethe’s &lt;em&gt;Faust&lt;/em&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;J.W. von Goethe, &lt;a href=&quot;https://www.poetryintranslation.com/PITBR/German/FaustIIActV.php&quot;&gt;&lt;em&gt;Faust. Der Tragödie zweiter Teil&lt;/em&gt;&lt;/a&gt;, 1832 (manuscript completed 1831, published posthumously).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;M. Berman, &lt;em&gt;All That Is Solid Melts Into Air: The Experience of Modernity&lt;/em&gt;, 1982.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;B.J. MacLennan, &lt;a href=&quot;https://web.eecs.utk.edu/~bmaclenn/papers/Homunculus_Quest.pdf&quot;&gt;“Homunculus’ Quest for a Body,”&lt;/a&gt; in L. Fitzsimmons (ed.), &lt;em&gt;Goethe’s Faust and Cultural Memory&lt;/em&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;</content:encoded><category>Philosophy</category><category>Information Geometry</category><category>Transformers</category><category>RLHF</category><author>Sirui He</author></item></channel></rss>