The part that stands out most is the tool-calling behavior switch, one model printing tool-call-like text instead of actually invoking a tool, then starting to call correctly right after it had seen the coding-focused model make real calls earlier in the same conversation. That reads more like an in-context effect than a capability gap, the model probably wasn't missing the ability to call tools, it was missing a concrete example of what a successful call looks like inside that client's specific prompt format. If that's right, the failure isn't really about the model at all, it's about whether the client's system prompt or tool schema gets surfaced to that model the same way it does to the coding-focused one, since some servers translate the tools field into quite different text depending on which chat template the runtime picks for a given model file. Did you get a chance to check whether the two models were actually served with the same chat template, or could the first one have been getting one that renders tool definitions poorly?