Exactly - retrieval without a refusal is the quiet failure mode: the model treats the top hit as an answer no matter how weak the match is. My rule of thumb is two gates, not one: a minimum similarity score below which nothing is eligible, and a margin check (the top hit must beat the runner-up by a set gap, so "lots of mediocre matches" doesn't sneak through). A concrete starting point for cosine on normalized embeddings: require ~0.75+ to answer, treat 0.6–0.75 as "retrieve but hedge," and refuse below that - then tune on your own eval set, because the right threshold is corpus-specific.
When nothing clears the bar, the honest move is a refusal with a next step, not a guess: "I don't have a confident source for that," plus what I do have (closest topics) or a handoff. A visible "no answer" is a feature —-it's the line between a bot that's confidently wrong and one people actually trust.