Ddonald991234inmodelvram.hashnode.dev·3d ago · 4 min readHow much extra VRAM speculative decoding needs in llama.cpp (MTP, DFlash, draft models)Speculative decoding makes local LLMs noticeably faster, but it is not free: the draft (an MTP head, a DFlash drafter or a small draft model) needs its own VRAM. People regularly run out of memory aft24KM