"Let AI gather context, reason through the incident, propose a safe plan, and pause for human approval" is exactly the right design target. We built FlowTux around the same principle — it reads the codebase to reason about root cause, but it only ever acts through a fixed, auditable allow-list of commands, so the pause-for-approval boundary is enforced by design, not by convention.