The Nuru example is what makes this land - most governance writeups stop at "define the policy" and never show what happens when it's tested against something real. One thing I'd push on in Layer 3: the risk tier itself usually gets estimated by the agent doing the acting, and that's the same failure mode as the rest of the piece in disguise. A misclassified action doesn't throw an error, it just skips escalation quietly instead of loudly, which is exactly the silent-failure pattern you're warning about, just one layer up. What's worked better for us is keeping the escalation trigger deterministic and outside the model's own judgment - hardcoded on the resource/verb pair (any write to payment or medical fields, any delete) rather than trusting a confidence score the agent generated about its own action.
The Nuru example is what makes this land - most governance writeups stop at "define the policy" and never show what happens when it's tested against something real. One thing I'd push on in Layer 3: the risk tier itself usually gets estimated by the agent doing the acting, and that's the same failure mode as the rest of the piece in disguise. A misclassified action doesn't throw an error, it just skips escalation quietly instead of loudly, which is exactly the silent-failure pattern you're warning about, just one layer up. What's worked better for us is keeping the escalation trigger deterministic and outside the model's own judgment - hardcoded on the resource/verb pair (any write to payment or medical fields, any delete) rather than trusting a confidence score the agent generated about its own action.