I work on efficiency × reliability for foundation models, mostly quantization and what it quietly breaks
Collaborations on LLM efficiency and reliability, open-source GPU kernel work, paper discussions, and 1:1 mentoring. Always happy to talk quantization, inference systems, or compilers.