The failure you list that gets under-measured is retrieval of the wrong-but-plausible document, since generation stays fluent and the answer looks right while resting on the wrong source. I've found that separating retrieval accuracy from answer accuracy in eval is the only way to catch it, because end-to-end correctness hides which half broke. Do you score retrieval hit rate independently, or judge the final answer only?