The line about validation being a mandatory second layer resonates, especially your Goal Health example where a good prompt still let the model invent numbers. I have found the same split: prompt quality reduces the error rate, but only a schema check plus fallback actually bounds the blast radius. Do you run those output validators inline per request, or as a separate offline eval pass on sampled traffic?