The cost point hits home. I’ve seen repeated prompts quietly eat through API budgets. Semantic caching helps, but only if you measure the hit rate first.
One thing I'd add. Hit rate on its own can flatter you. A loose threshold looks
like a great result until you check false hits next to it. I measure duplicate
rate in shadow mode before building anything, then track both numbers together.