ARAleksei Romanovingfactor.hashnode.dev·11h ago · 7 min readSFT vs. RL: What Changes Inside the Model?In modern generative AI post-training, two fundamental paradigms dominate: Supervised Fine-Tuning (SFT) and Reinforcement Learning with Verifiable Rewards (RLVR / GRPO). While practitioners often trea00
ARAleksei Romanovingfactor.hashnode.dev·1d ago · 10 min readLatent-GRPO: Reinforcement Learning in Continuous Thought SpaceWhen you sit down to solve a complex puzzle or plan three moves ahead in chess, do you narrate every synaptic firing to yourself in full, grammatically correct English sentences? Of course not. Human 00
ARAleksei Romanovingfactor.hashnode.dev·3d ago · 13 min readHigh-Throughput LLM Inference & Training: A Deep Dive into vLLMEditor's Note: Originally published on the g factor engineering blog. All benchmarks and telemetry in this article were conducted on dedicated NVIDIA H100 and H200 clusters on gft-studio. If you have00