Good question, and it sent me back to the source (master at 8019dc5). Short answer: the snapshots are budgeted, and the post should have said so. They don't live in the draft context at all. common_context_params_to_llama sets n_rs_seq = speculative.need_n_rs_seq(), which is --spec-draft-n-max for MTP, EAGLE3, DFlash and DSpark. The target model's llama_memory_recurrent constructor then allocates rs_size x (1 + n_rs_seq) state rows up front, with rs_size = n_seq_max. fit.cpp measures the main model by creating a context and reading llama_get_memory_breakdown, so those rows land in its context term and scale with n_max and -np. The draft is measured separately with n_rs_seq = 0, so a failed draft measurement doesn't drop them, and they don't grow with prompt length. What fit doesn't cover is --ctx-checkpoints, which are host-side copies. I haven't verified this on hardware; running with --verbose and comparing fit's projected memory with the "RS buffer size" log line would confirm it. Thanks for pushing on this.
