Testing AI-Dependent Code: Mocking LLMs, Evaluating Outputs, and Avoiding Flaky Tests
The test checked whether the extracted invoice total matched the expected value. It passed nine times out of ten. On the tenth run, the model formatted the number differently — "1,234.56" instead of "
blog.madhav.dev6 min read