How Much Context Does a Small GPT Model Really Need?
TL;DR
I ran three context-length experiments using the same compact GPT model (Github):
Experiment 1 — Fixed training-token budget: When every model processed approximately 200,000 tokens, the 64-tok
curious-pm.hashnode.dev12 min read