Excessive Agency climbing to number three matches what I'm seeing, the risk moved from what the model says to what it's allowed to do. Blast radius is the right lens, a wrong answer is cheap but a wrong action with tool access is not. How are you scoping agent permissions to keep that radius small?