The interesting shift here is that agent security is less about whether the model can be trusted and more about where the enforcement boundary lives. A prompt or project rule can describe what the agent should avoid, but the stronger controls are the ones that still hold when the agent ignores or misinterprets the instruction. That distinction becomes much more important once agents can call tools and modify a repository instead of only generating text.