Serving Your Own Model with llama.cpp: GPU Offload, Caching, and the Trap of a Stuck Wrong Answer
The model is fine-tuned. The next question: how do you actually use it — not inside a training notebook, but served like a real API, with a cache so the same question doesn't get recomputed over and o
shaka-ai.hashnode.dev7 min read