Running a A 27B Model in 3.5 GB VRAM
Bonsai-27B-gguf is Qwen3.6-27B with every weight stored as a single bit. The file is 3.5 GB and it fits on an 8 GB GPU.
FP16 weights for a 27B model are about 54 GB. A conventional 4-bit GGUF is about
mamonu.hashnode.dev10 min read