Really thorough methodology, especially separating the client-side watch timeout from the controller's own ProgressDeadlineExceeded, that distinction trips people up constantly. One thing worth asking about the recovery step: reapplying the same healthy spec created a brand new revision each time (2 to 4, 4 to 6, 6 to 8) instead of landing back on an existing one. In a setup where you're triggering an automated rollback off ProgressDeadlineExceeded, does kubectl rollout undo without an explicit --to-revision reliably land on the last genuinely healthy spec, or does it just walk back to the most recent distinct ReplicaSet regardless of whether that one was itself a different failure? Two failures in a row would make that ambiguity matter a lot more than the single-fault case this lab walks through.
Really thorough methodology, especially separating the client-side watch timeout from the controller's own ProgressDeadlineExceeded, that distinction trips people up constantly. One thing worth asking about the recovery step: reapplying the same healthy spec created a brand new revision each time (2 to 4, 4 to 6, 6 to 8) instead of landing back on an existing one. In a setup where you're triggering an automated rollback off ProgressDeadlineExceeded, does kubectl rollout undo without an explicit --to-revision reliably land on the last genuinely healthy spec, or does it just walk back to the most recent distinct ReplicaSet regardless of whether that one was itself a different failure? Two failures in a row would make that ambiguity matter a lot more than the single-fault case this lab walks through.