The framing that "it didn't crash" tells you nothing about whether retrieval was relevant is the exact gap most RAG builds ignore until answers start drifting. Precision and recall at K get the attention, but MRR and NDCG are usually what expose a reranker that looks fine on averages while burying the best chunk at position four. Which of these do you actually gate on in practice, or do you track them for diagnosis and let end-to-end answer quality be the real bar?