the part that's easy to skip past is step 1, "hold it long enough to observe normal load." we treat that observation window as the actual gate, not just a waiting period -- pull evaluation telemetry for the whole window and confirm every evaluation landed on the variant you think won, not just that nobody complained.
a kill-switch flipped in the provider dashboard during an incident and never flipped back produces exactly the silent-default failure this post describes, except it looks like a clean removal, because the code default happens to match whatever got served that week by accident. the four-state inventory is the right fix for that -- "evaluated only through default or error paths" catches it -- but only if someone's actually looking at states three and four instead of just expired vs still-active.