25 LLM architecture blocks, side by side, in runnable PyTorch
GPT-2 to Kimi Linear is seven years of architecture research, and almost all of it fits in
about twenty lines per model. Below are 25 decoder blocks — GPT-2, OPT, Llama 2/3/4, Gemma
2/3, Qwen 2.5/3/3-