Two CUDA PRs into llama.cpp, and what they taught me about a 100K-star codebase
My first pull request to llama.cpp was 95 lines of CUDA. Georgi Gerganov, the guy who started the project, merged it the next day.
The second one was even smaller. Nine new lines of code, and the chan
amankarki.hashnode.dev11 min read