Model choice is a full eval pass, not a vibe check, and the only way to know which model works for your task is to run the same golden set against both. I have seen the "winner" flip between models when the eval set changed from final-answer scoring to trajectory scoring, which is the axis that actually predicts production reliability.