The 'reduce churn' example is a useful demonstration of vocabulary coverage rather than a general failure of semantic embeddings. In the LSA version, churn has no TF-IDF feature at all, so the query effectively becomes reduce before the projection; the pretrained swap tests a materially different source of knowledge.
I would print the recognized and discarded query terms next to each result list. Then compare 'reduce churn' with 'reduce cancellations' and a no-overlap paraphrase under both embeddings, leaving the relevance labels fixed. That would make it clear which gains come from corpus vocabulary, the learned projection or pretrained semantics, rather than treating vector search as one interchangeable method.