One GPU, Two LLMs
What I was aiming simple in theory: deploy two open-weight LLMs behind a custom gateway, on Kubernetes, infrastructure as code. The contraint that made it interesting was a hard cost cap. One spot GPU
orhunkupeli.hashnode.dev6 min read