Do Bigger GPT Models Always Perform Better?
TL;DR: What I tested and observed
Experiment 1 — Different sizes, fixed tokens: I increased model width, depth, and total parameter count while giving every model the same 200,192 training tokens. In
curious-pm.hashnode.dev11 min read