"Restricting the LLM's decision to things the system already knows how to handle" is a good way to put it, and it changes what you measure as well.
Once retrieval constrains the output space, the interesting metric stops being answer quality and becomes whether the right constraint was selected. That is a classification problem with a confusion matrix, not something you need a judge model for, which makes the eval far cheaper and far more stable.
The failure that survives is the query whose correct handling is not in the taxonomy at all. That one is invisible unless you deliberately log the no-match cases and read them, since from the inside it looks identical to a query the system handled.
Retrieval used to constrain what the model can output, rather than to feed it context, is the part worth pulling out - it turns an open generation problem into a closed one over a known taxonomy, and hallucinating a food category that does not exist in the catalogue simply stops being possible. Different failure surface entirely. The without onions half of that query is the interesting bit, because negation is where pure embedding similarity is weakest: not onions and onions sit close together in vector space since they appear in near-identical sentences. Mapping it to a structured exclusion filter is the only reliable handling, and it is a good argument for why an entity-and-constraint layer earns its place next to semantic search rather than being replaced by a bigger model. Worth noting the latency budget too - a search box has maybe a couple of hundred milliseconds before it feels broken, so an LLM in the hot path usually means the parse is cached per query pattern rather than run per keystroke.