Part 4 of 4 , Scaling LLM Training. Code for the series: github.com/rocks-saka/Scaling-llm-training The previous three posts got us a long way on a single axis: map the memory, recompute activations,
sakshityagi.hashnode.dev3 min read
No responses yet.