Running a Coding Agent on a Local GLM 4.7 with vLLM: The Setup Nobody Documents
Giving a coding agent its own brain for zero API cost is genuinely possible now. GLM 4.7 Flash serves on two consumer RTX 3090s, and a real agent framework can point at it instead of a cloud endpoint.
joshgreen.hashnode.dev7 min read