Filling the window with the right tokens beats filling it with more tokens, every time I have measured it. Past a point, extra context actively hurt because the model latched onto the wrong passage. Are you ranking by relevance before packing the window, or trimming after retrieval?