Latent-GRPO: Reinforcement Learning in Continuous Thought Space
When you sit down to solve a complex puzzle or plan three moves ahead in chess, do you narrate every synaptic firing to yourself in full, grammatically correct English sentences? Of course not. Human
gfactor.hashnode.dev10 min read