A golden set tied to chunk IDs becomes stale as soon as the parser or chunker changes, even if the underlying evidence is unchanged. I prefer labeling source spans or evidence claims, then mapping those labels onto each index version before scoring. It is also useful to separate topical relevance from answer support: a chunk can be about the right subject yet lack the exact fact needed for generation. Reporting metrics by query slice and with hard negatives makes improvements much harder to hide inside one aggregate NDCG number.