The "every request returned 200 while the agent hallucinated tracking numbers" story is the whole case for eval-in-the-loop, since uptime tells you nothing about correctness. I ran into the mirror image on Bedrock, where the built-in eval graded my agent green while only looking at about a quarter of what mattered: kartiknvjk.hashnode.dev/aws-bedrock-s-built-in-ev… . Are you scoring the regression set on every prompt change, or only on model swaps?
The "every request returned 200 while the agent hallucinated tracking numbers" story is the whole case for eval-in-the-loop, since uptime tells you nothing about correctness. I ran into the mirror image on Bedrock, where the built-in eval graded my agent green while only looking at about a quarter of what mattered: kartiknvjk.hashnode.dev/aws-bedrock-s-built-in-ev… . Are you scoring the regression set on every prompt change, or only on model swaps?