The fraud example also exposes a monitoring delay: confirmed outcomes can arrive much later than predictions. A precision dashboard that mixes fully reviewed older transactions with mostly unresolved recent ones can look like the model improved when only the label coverage changed.
I would group performance by prediction-time cohort and report how many outcomes are still pending in each cohort. Keep model version and decision threshold with every prediction so a later label can be attributed correctly. Input drift and latency remain useful immediate signals, while outcome metrics become trustworthy only when their observation window and label completeness are explicit.