The pitch that you can match GRPO and PPO by evolving prompts with reflective critiques, no weight access needed, is a big deal for teams who can never touch the model internals. The Pareto selection is the interesting part, since it is what keeps prompt evolution from collapsing onto a single overfit candidate. Have you run GEPA on a real compound pipeline yet, and did the evolved prompts stay stable when the base model was swapped underneath them?