LoRA & DoRA: The Math, Memory, and Trade-offs
Every engineer who has ever tried to fine-tune a modern 27B or 70B parameter model knows the rude awakening of GPU memory arithmetic.
You look at the raw model weights and think: “27 billion parameter
gfactor.hashnode.dev10 min read