The point about hallucinated context being scarier than a hallucinated answer is the one I keep coming back to. When the model fabricates the inputs and downstream steps treat them as ground truth, the failure is silent and much harder to trace. Your BM25-plus-metadata-filter split for exact identifiers matches what I have seen too. How do you decide the cutoff between the structured residue and the genuinely semantic query in practice?