One thing this highlights is that the risk isn't just the model it's the combination of autonomy + tools + long-running execution. A model generating bad text is one thing. An agent that can plan, use tools, and modify its environment introduces a completely different class of engineering problems. That shifts the focus toward runtime controls, least privilege, and validating actions rather than assuming prompts alone are enough