C++/CUDA engineer building LLM inference and GPU code from scratch: verbum.cpp, Lattice, RAAG. 2 CUDA PRs merged into llama.cpp. I write about inference, GPU performance, and measuring things honestly.
Available for
Software engineering roles, technical collaborations, and discussions on systems architecture & compiler tooling.