Understanding AirLLM: How to Run a 70B Model on a 4GB GPU
I went down a rabbit hole trying to understand AirLLM — a library that claims to run 70B+ parameter models on GPUs with as little as 4-6GB of VRAM. What started as "what is this tool" turned into a mu
mujtaba08.hashnode.dev16 min read