The separate fetch and evaluation budgets make the fallback ladder measurable instead of a hidden latency shortcut. I would also slice calibration by fallback reason, because remote-feature timeouts may be concentrated in particular users, exchanges or traffic conditions rather than resemble random missing features.
Training with a randomly absent feature group could miss that selection effect. A useful shadow check would replay the observed timeout subset through the reduced model and compare its predicted rates with mature outcomes, while keeping the full-model result as a counterfactual only where those features actually arrived. That would connect your fallback-rate dashboard to the price-quality risk of the specific traffic falling down the ladder.