Kartik N V J K
AI Developer | Making AI reliable, trustworthy & accessible to everyone | Active community contributor
The 'reduce churn' example is a useful demonstration of vocabulary coverage rather than a general failure of semantic embeddings. In the LSA version, churn has no TF-IDF feature at all, so the query effectively becomes reduce before the projection; the pretrained swap tests a materially different source of knowledge.
I would print the recognized and discarded query terms next to each result list. Then compare 'reduce churn' with 'reduce cancellations' and a no-overlap paraphrase under both embeddings, leaving the relevance labels fixed. That would make it clear which gains come from corpus vocabulary, the learned projection or pretrained semantics, rather than treating vector search as one interchangeable method.
Defining the correct answers before running the queries is the step most retrieval demos skip, and it is the only reason your "vector won one, lost one" result means anything. The honest LSA-on-28-paragraphs caveat matters too, since fitting on the corpus you are testing flatters recall in a way a pretrained encoder would not. I went through the same measure-against-a-gold-set exercise with LlamaIndex and wrote up how I score it here: kartiknvjk.hashnode.dev/evaluating-llamaindex-rag…