How Much Do Attention Heads and Activation Functions Really Matter?
TL;DR
I tested two architectural choices in a small GPT model while keeping the rest of the setup stable:
Attention heads: 1, 4, and 8 heads under fixed-token and fixed-time budgets.
Activation func
curious-pm.hashnode.dev7 min read