The line about two posts arguing opposite points scoring nearly identical in embedding space is the real failure, and it is why cosine similarity alone kept surfacing my stale notes. A reranker fixed relevance for me too, but the per-candidate latency forced me to cap the candidate set first. What candidate count did the BGE cross-encoder stay worth it at before the latency hurt the response?