Really enjoyed this practical walkthrough. What stood out to me is how you explained that running an LLM locally is much more than just downloading a model—it’s really about getting the client, server, runtime, model format, hardware, and accelerator to work together.
I especially liked the point about memory and model size not being the whole story. A model can technically fit in memory and still be too slow for a useful interactive experience, which is something that’s easy to overlook when getting started.
The section on coding agents was also interesting. Multiple sequential model calls can make latency much more noticeable compared with simple chat. Overall, this feels like a very honest look at the amount of experimentation involved in local inference. Great read for anyone thinking about going beyond cloud-hosted models!
The part that stands out most is the tool-calling behavior switch, one model printing tool-call-like text instead of actually invoking a tool, then starting to call correctly right after it had seen the coding-focused model make real calls earlier in the same conversation. That reads more like an in-context effect than a capability gap, the model probably wasn't missing the ability to call tools, it was missing a concrete example of what a successful call looks like inside that client's specific prompt format. If that's right, the failure isn't really about the model at all, it's about whether the client's system prompt or tool schema gets surfaced to that model the same way it does to the coding-focused one, since some servers translate the tools field into quite different text depending on which chat template the runtime picks for a given model file. Did you get a chance to check whether the two models were actually served with the same chat template, or could the first one have been getting one that renders tool definitions poorly?
Kartik N V J K
AI Developer | Making AI reliable, trustworthy & accessible to everyone | Active community contributor
The detail that matches my own experience is Claude Code hitting API-compatibility problems against a local server while plain chat worked. I ran into the same wall: the multi-turn tool-call loop expects OpenAI-exact response shapes, and Lemonade's llama.cpp bridge drops fields the agent needs mid-task. Did switching to OpenCode fix the tool-call handshake for you, or did you still patch responses?