Biasing to the stronger model on close scores is the right asymmetry, since over-spending on Opus just costs money while under-routing a hard debug costs you a wrong answer you trust. The part I would instrument is exactly the misroutes you list at the end: keyword matching drifts, and without logging routing decisions against outcomes you cannot tell if the router is actually saving anything. Are you tracking how often the deep tier gets a prompt Haiku would have nailed?