Making GPU Failure Invisible: A Zero-Cost Fallback for LLM Inference on Kubernetes
I will be giving a formal talk on this implementation soon. More to come in a highly deep dive technical blog followed with the talk
Most GPU-aware routing demos assume you have GPUs. I wanted to kno
blog.raeveen.dev4 min read