Framing the lineage through the papers is a great way in, and the Self-RAG idea, training the model to decide whether retrieval is even needed, is the pivot I point people to, since always-retrieve wastes budget and can hurt on questions the model already knows. The through-line you draw, that the gains came from better retrieval rather than bigger models, has held up in everything I've shipped. Do you think the next jump is in retrieval quality or in the model knowing when to trust what it retrieved?