Cutting ten planned experiments down to the one that changed a decision is the right way to run evals, since most eval work produces numbers nobody acts on. Routing is a good place to start because the outcome maps straight to cost and latency. What was the score gap that made you comfortable sending traffic to the smaller model?