I Built a 9× Faster LLM Memory Allocator in C++ — and Worked Out When It Pays Off
I spent a weekend porting the memory manager of an LLM inference engine from Python to C++: the block table, the free list, the reference counts, behind pybind11.
It works. All 35,635 recorded operati
seng-wei-chieh.hashnode.dev18 min read