Anton R Gordon on Prefill vs. Decode: Finding the Real Bottleneck in LLM Inference
By Anton R Gordon
When an LLM endpoint feels slow, the instinct is usually to blame the model.
The model is too large.
The GPU is too slow.
The context window is too long.
We need more GPUs.
But those
antonrgordon.hashnode.dev12 min read