Yes, that's the failure that matters most. A shallow answer you trust costs more in rework than Opus would have. Today there are two guards. Claude can move a task up a tier, and Haiku and Sonnet are told to stop and hand off if the work turns out harder than it looked. But I don't log manual escalations yet, so nothing feeds back into the scoring. That's the next step. Every escalation will be logged with the original prompt and both tiers: automatic hand-offs, Claude moving a task up, and me re-running a task on a stronger model. Each one becomes a labelled example ("this prompt needed deep"). They go into a regression test set, and I'll tune the rules against that set rather than changing weights on gut feel.
