That split holds up well in practice, and the switch between them is close to free since both speak an OpenAI-compatible API — swapping the base URL doesn't mean rewriting your client. One gotcha if you're doing quality comparisons between local and cloud: Ollama's default model tags are Q4_K_M, so "the cloud model is smarter" is sometimes just "I compared a 4-bit quant to an 8-bit one." Worth pulling a q8 tag before judging.