Batching and the Idle GPU: Serving LLMs Faster
Part 2 of 4 Serving LLMs in Production.
Part 1 left us with an uncomfortable fact: when the model writes its answer the slow, one-word-at-a-time decode phase the GPU's powerful math units sit mostly i
sakshityagi.hashnode.dev6 min read