The fraud example also exposes a monitoring delay: confirmed outcomes can arrive much later than predictions. A precision dashboard that mixes fully reviewed older transactions with mostly unresolved recent ones can look like the model improved when only the label coverage changed.
I would group performance by prediction-time cohort and report how many outcomes are still pending in each cohort. Keep model version and decision threshold with every prediction so a later label can be attributed correctly. Input drift and latency remain useful immediate signals, while outcome metrics become trustworthy only when their observation window and label completeness are explicit.
The fraud example also exposes a monitoring delay: confirmed outcomes can arrive much later than predictions. A precision dashboard that mixes fully reviewed older transactions with mostly unresolved recent ones can look like the model improved when only the label coverage changed.
I would group performance by prediction-time cohort and report how many outcomes are still pending in each cohort. Keep model version and decision threshold with every prediction so a later label can be attributed correctly. Input drift and latency remain useful immediate signals, while outcome metrics become trustworthy only when their observation window and label completeness are explicit.