The escalation-rate-as-health-metric point is the one I'd underline. We run something similar on our own review gate before anything goes out under our name, and the lesson that generalizes is the opposite direction from what most teams expect: if the review almost never flags anything, that's a sign the review got too soft, not a sign the system got reliable. Same logic applies to an escalation queue. A team that tunes the rules until intervention rate drops to near zero has usually just taught the agent to stop asking, not to stop being wrong, unless someone is separately checking that the actions running unattended are actually correct and not just unflagged. The intervention-rate trend is only a health signal if you're also sampling the non-escalated actions on a schedule, otherwise it's trivially easy to "improve."