SFT vs. RL: What Changes Inside the Model?
In modern generative AI post-training, two fundamental paradigms dominate: Supervised Fine-Tuning (SFT) and Reinforcement Learning with Verifiable Rewards (RLVR / GRPO).
While practitioners often trea
gfactor.hashnode.dev7 min read