The fit trap you describe is specifically about the draft's own KV-cache context measurement failing when the draft needs the main model's context. What isn't clear to me is whether the recurrent-state snapshot cost, the 449 MB from three copies in the Qwen3.8 example, is something fit.cpp ever accounts for on any model, or whether that term sits entirely outside its budget regardless of whether the KV-cache measurement itself succeeds or fails. If it's the latter, manually setting -c to work around the trap fixes the KV-cache half of the problem but still leaves the snapshot memory unbudgeted, so a hybrid model with a higher --spec-draft-n-max could still OOM on a long prompt after doing everything the fix section says, just from a smaller and harder to spot amount. Did you check whether that snapshot memory shows up anywhere in fit's own calculation, even for a draft whose context measurement succeeds?
Kartik N V J K
AI Developer | Making AI reliable, trustworthy & accessible to everyone | Active community contributor
The -fit trap counting the draft as zero is the one that bit me. My run looked fine until the first long prompt, then OOM'd exactly like the llama.cpp issue you linked, because the fitter sized the main context without the draft KV cache. I started passing -c explicitly after that. Does the same silent miss show up with DFlash drafters, or only the ones that share the main context?