Your eval dashboard has 30 metrics. When one "moves," that is usually arithmetic, not a regression.
Here is the ritual. You ship a prompt change, rerun the eval suite, and open the dashboard. Thirty numbers sit there: faithfulness, answer relevance, context precision, toxicity, latency-adjusted qual