The cost section resonates: retrieval quality and cost pull in opposite directions once you scale the index. I found reranking let me pull fewer candidates up front without losing recall, which dropped both latency and spend. Are you caching embeddings at the query level, or only on the document side?