The request trace schema is a strong start; I would add the corpus or index version, embedding fingerprint, and both raw retrieval and reranker scores. Chunk IDs alone are not enough for replay if the underlying document was re-ingested or the same ID now points to different text. With those fields frozen, the retrieval, context, generation, and routing categories become testable hypotheses rather than labels assigned after the fact.
Ahmet Özel
AI Engineer. Computer Vision, RAG and LLM agents.
The request trace schema is a strong start; I would add the corpus or index version, embedding fingerprint, and both raw retrieval and reranker scores. Chunk IDs alone are not enough for replay if the underlying document was re-ingested or the same ID now points to different text. With those fields frozen, the retrieval, context, generation, and routing categories become testable hypotheses rather than labels assigned after the fact.