The distinction between finding relevant evidence and preserving sufficient evidence for reasoning is an important one here. A retrieval system can return highly relevant chunks and still fail if the answer depends on a definition, exception, or relationship that wasn't included in the final evidence bundle. I’d be interested in evaluating “evidence completeness” alongside retrieval precision and recall: not just whether the expected passage was retrieved, but whether all dependencies required to justify the answer survived selection and packing. That metric could also make the RAG-versus-long-context decision more concrete. If a task repeatedly requires broad evidence relationships, long context or structured parent expansion may be the better baseline even when top-k retrieval looks strong.