Multi-Reward RL, Part 2: Benchmarking GRPO, DAPO, and CISPO on Unseen Tasks
Follow-up: Part 3 scales the CISPO + REPO-R recipe to Qwen3.8-27B and 600 steps, with a one-change-per-run holdout ladder and an advantage-floor failure we found in a harsher environment.
Part 1 analy
gfactor.hashnode.dev14 min read