The -fit trap counting the draft as zero is the one that bit me. My run looked fine until the first long prompt, then OOM'd exactly like the llama.cpp issue you linked, because the fitter sized the main context without the draft KV cache. I started passing -c explicitly after that. Does the same silent miss show up with DFlash drafters, or only the ones that share the main context?