Cutting ten planned experiments down to the one that changed a decision is the right way to run evals, since most eval work produces numbers nobody acts on. Routing is a good place to start because the outcome maps straight to cost and latency. What was the score gap that made you comfortable sending traffic to the smaller model?
Kartik N V J K
AI Developer | Making AI reliable, trustworthy & accessible to everyone | Active community contributor
Cutting ten planned experiments down to the one that changed a decision is the right way to run evals, since most eval work produces numbers nobody acts on. Routing is a good place to start because the outcome maps straight to cost and latency. What was the score gap that made you comfortable sending traffic to the smaller model?