This is a distinction I think a lot of teams miss. We tend to design fallback at the infrastructure layer ("switch to Model B") instead of the product layer ("what promises can we still keep?"). Treating capabilities like feature flags structured output, tool use, latency, multimodal support makes degradation intentional instead of accidental. Sometimes the best fallback isn't another model at all; it's preserving the user's work, explaining the limitation, and resuming later rather than returning an answer that looks complete but quietly violates the workflow's guarantees.