RS
Thank you for your breakdown! While this article was just a projection of my current knowledge, I’ll definitely highlight this distinction in a future update. Thanks for adding so much value to the post!" Today, I got to learn something new! 'The distinction between the prefill phase (compute-bound, parallelized prompt processing) and the decoding phase (memory-bandwidth bound, memory-bound autoregressive token generation) is where so many benchmarking comparisons fall flat'. Looking forward to learn more.
