The goal-filter example makes the selection problem clear. I would freeze the wallet list using a discovery window that ends before the replay window. For event ordering, keep both exchange and receive timestamps: one reconstructs market state; the other says what the bot knew. Otherwise sorting can repair a fill calculation while giving the simulated bot an event it had not yet received. Do your replays retain both clocks?
Yes, the recorder keeps the exchange timestamp next to our receive time, that's where the ~17 ms lag figure comes from. And the wallet list in the honest test was frozen on Sep 17-29, before the replay week.
That keeps wallet selection separate from the replay week. I would keep one late-arrival case: an event with an earlier exchange time reaches the bot after its decision. The replay should preserve that decision, then let the late event affect the next one, rather than move it into the earlier input set.
Arnold Holm
Challenge Pre-Check: know before you risk. Your rules, tested against prop limits. Historical simulation, not a forecast.
The goal-filter example makes the selection problem clear. I would freeze the wallet list using a discovery window that ends before the replay window. For event ordering, keep both exchange and receive timestamps: one reconstructs market state; the other says what the bot knew. Otherwise sorting can repair a fill calculation while giving the simulated bot an event it had not yet received. Do your replays retain both clocks?