A small Python hook, three subagents and a CLAUDE.md file. Full code below, plus the parts that still don't work well. TL;DR I used to run one model for everything. That meant I either overpaid for si
devlore.hashnode.dev14 min readAI Developer | Making AI reliable, trustworthy & accessible to everyone | Active community contributor
Routing by prompt score is smart, but the costly failures are misroutes in one direction: a hard refactor sent to the cheap tier comes back confidently shallow and costs more in rework than the big model would have. Do you log the cases where you escalated manually after a cheap answer, and feed those back into the scoring rules?
iin1005h23
Kartik N V J K
Biasing to the stronger model on close scores is the right asymmetry, since over-spending on Opus just costs money while under-routing a hard debug costs you a wrong answer you trust. The part I would instrument is exactly the misroutes you list at the end: keyword matching drifts, and without logging routing decisions against outcomes you cannot tell if the router is actually saving anything. Are you tracking how often the deep tier gets a prompt Haiku would have nailed?