The drift check is the part I'd tighten before trusting it in production. Running a KS test at alpha=0.05 on every feature every hour means if you're watching, say, twenty features, you should expect roughly one false positive an hour from pure chance, before anything in the world actually changed. That's a fast route to alert fatigue and a team that starts ignoring the drift page entirely, which defeats the whole point of the section. The fix isn't a stricter alpha per test, since that just trades false positives for missed real drift, it's either correcting for the number of features being tested (Bonferroni or Benjamini-Hochberg) or requiring persistence, only alerting once the same feature flags on two or three consecutive hourly windows, which filters out one-off noise without dulling sensitivity to a real, sustained shift.
The drift check is the part I'd tighten before trusting it in production. Running a KS test at alpha=0.05 on every feature every hour means if you're watching, say, twenty features, you should expect roughly one false positive an hour from pure chance, before anything in the world actually changed. That's a fast route to alert fatigue and a team that starts ignoring the drift page entirely, which defeats the whole point of the section. The fix isn't a stricter alpha per test, since that just trades false positives for missed real drift, it's either correcting for the number of features being tested (Bonferroni or Benjamini-Hochberg) or requiring persistence, only alerting once the same feature flags on two or three consecutive hourly windows, which filters out one-off noise without dulling sensitivity to a real, sustained shift.