Sharp point: more detectors without release discipline just adds noise you learn to ignore. Tying eval gates to the release itself did more for my reliability than any new metric. Are you gating deploys on the eval suite, or running it after the fact?