The fraud example makes delayed labels an important part of the monitoring story. A model can appear healthy or unhealthy simply because recent transactions have not had enough time to receive a confirmed outcome. I would compare quality on cohorts with the same label maturity, while monitoring feature and prediction distributions separately for faster warning signals.
That also changes the retraining trigger. A distribution shift can justify investigation without proving that a replacement model is better. The release gate should compare the current and candidate models on a mature, time-separated cohort using the operating threshold and error costs that the service actually uses.