The strongest point here is separating “can the agent be manipulated?” from “can that manipulation actually cross an authority boundary?” That distinction makes the security model much easier to reason about. A model will eventually produce a bad tool call; the important invariant is that the call remains unauthorized regardless of how it was generated.
In production agentic systems at IT Path Solutions, I’d also treat the tool gateway as a first-class security boundary, not just middleware. Every decision should leave an auditable trail of principal, requested action, resource, policy result, and resulting state change. That makes the authority layer testable independently and gives you something concrete to investigate when an agent behaves unexpectedly.
The CI-level direct-call tests are especially valuable because they verify the actual security invariant without depending on prompt behavior.