The messy-data point is huge. A strong next step is to make provenance portable: every retrieved chunk should carry a compact proof of source, freshness, and transformation history. Then retrieval quality becomes inspectable, not just a leaderboard score.
The distinction between RAG and the information system around it is the part that stands out. Once enterprise data has different freshness, ownership, permissions, and structure, retrieval quality becomes an architectural property, not just a model or vector-database problem.
One thing I’d add is that provenance should travel with the data through every transformation, not just appear at the final citation layer. If a chunk can’t tell you which source version, extraction path, permission context, and indexing state produced it, debugging a bad answer becomes much harder. In production AI work at IT Path Solutions, that kind of lineage is often just as important as improving retrieval itself.
The “smallest evidence set sufficient for the task” principle is also important. Sending more context to a stronger model can hide weaknesses in ingestion and retrieval for a while, but it makes cost, latency, and conflicting evidence harder to control. Good RAG architecture is ultimately about knowing when to retrieve, when to query the source directly, and when not to involve the LLM at all.
Spot on. The retrieval pipeline is trivial compared to the endless friction of untangling legacy data models and access controls.
Julian Neagu
500+ AI tools shipped solo. Founder of VisionVix.
The part about RAG being a search problem, not just a vector DB problem, hits home. Exact IDs and dates can matter more than semantic similarity. Learned this the hard way.