Making LLM Decode Cheaper: Quantization and More
Part 3 of 4 Serving LLMs in Production.
Part 2 filled the idle GPU by serving many users at once. But we still haven't touched the core problem from Part 1: every single word the model writes drags th
sakshityagi.hashnode.dev6 min read