Defining the correct answers before running the queries is the step most retrieval demos skip, and it is the only reason your "vector won one, lost one" result means anything. The honest LSA-on-28-paragraphs caveat matters too, since fitting on the corpus you are testing flatters recall in a way a pretrained encoder would not. I went through the same measure-against-a-gold-set exercise with LlamaIndex and wrote up how I score it here: kartiknvjk.hashnode.dev/evaluating-llamaindex-rag…