The "every request returned 200 while the agent hallucinated tracking numbers" story is the whole case for eval-in-the-loop, since uptime tells you nothing about correctness. I ran into the mirror image on Bedrock, where the built-in eval graded my agent green while only looking at about a quarter of what mattered: kartiknvjk.hashnode.dev/aws-bedrock-s-built-in-ev… . Are you scoring the regression set on every prompt change, or only on model swaps?