One thing I’d be curious to see measured is how stable a chunking strategy remains as the query distribution changes. A setup can perform well against a fixed evaluation set while degrading when users start asking questions that cross section boundaries or require context from neighboring chunks. Evaluating retrieval against evolving query patterns, rather than treating the benchmark as static, could reveal failures that Recall@K alone might miss.