The goal-filter example makes the selection problem clear. I would freeze the wallet list using a discovery window that ends before the replay window. For event ordering, keep both exchange and receive timestamps: one reconstructs market state; the other says what the bot knew. Otherwise sorting can repair a fill calculation while giving the simulated bot an event it had not yet received. Do your replays retain both clocks?
Arnold Holm
Challenge Pre-Check: know before you risk. Your rules, tested against prop limits. Historical simulation, not a forecast.