The Problem Most pipeline observability today is a wall of thresholds. Row count dropped below X. Job ran longer than Y minutes. Null percentage exceeded Z. Someone picked those numbers months ago, of
tech4nirvana.com8 min read
Consistency and continuous learning make all the difference. Thanks for sharing this valuable breakdown.
Kartik N V J K
AI Developer | Making AI reliable, trustworthy & accessible to everyone | Active community contributor
Median/MAD over mean plus stddev is the right call for pipeline metrics, since one giant batch or a retry storm wrecks a z-score but barely moves the median. The same-weekday baseline is the detail I'd steal, because a Monday reprocessing job looks nothing like a Sunday and averaging across the week hides both. How do you handle holidays and one-off backfills, do they poison the trailing-8 window, or do you exclude them explicitly?