GRPO: How Language Models Learn to Reason
Do you remember the times when we used to make LLMs count the occurrence of a specific letter in a word, like "How many r's in strawberry?"
Back then, LLMs used to get it wrong a lot of times, but now
swaritshukla.hashnode.dev5 min read