GRPO
Writing1 post
GRPO From First PrinciplesWhy On-Policy RL for LLMs Is Really Just Dynamic Weighted SFT
A first-principles walk from supervised fine-tuning to policy gradient and GRPO — why on-policy RL for LLMs is really dynamic, group-normalized, reward-weighted SFT.