The schema-versus-value distinction is useful, but there’s an interesting gap between “valid output” and “correct output.” A response can satisfy the schema perfectly while extracting the wrong invoice total, misclassifying a document, or producing a summary that misses a critical fact. I’d treat those as separate testing layers: deterministic tests verify the contract and how the application handles the response, while a smaller set of evaluation cases checks semantic correctness against known expectations. That also makes flaky tests easier to avoid without weakening coverage. The goal isn't to make every assertion deterministic; it's to make each test responsible for a property that can actually be judged reliably.