Provenance travelling with the content from the moment it enters the pipeline is the part most teams skip, and it breaks citations later. I grade retrieval before generation too, but the retry cap matters as much as the grader. I wrote about keeping that RAG gate honest here: kartiknvjk.hashnode.dev/how-i-stopped-my-rag-ci-g…. Are you scoring relevance with an LLM judge or something cheaper?