Stop Caching LLM Responses. Cache the Thinking Instead.
One of the biggest surprises I had while working on RAG systems wasn't retrieval.
It was what happened after retrieval.
Most conversations around RAG optimization focus on:
Better embeddings
Better
coalent.hashnode.dev2 min read