RLHF
Writing2 posts
GRPO From First PrinciplesWhy On-Policy RL for LLMs Is Really Just Dynamic Weighted SFT
A first-principles walk from supervised fine-tuning to policy gradient and GRPO — why on-policy RL for LLMs is really dynamic, group-normalized, reward-weighted SFT.
Manifolds and Genesis
A dialogue between a researcher and a language model — from how system prompts work to information geometry, manifolds, and what it all might mean.