KDkalyan dahakeinaekshant.hashnode.dev·1d ago · 4 min readI Built a Prefix-Cache-Aware LLM Gateway in Go — And Reduced TTFT by 90%Most backend engineers treat LLM inference like standard microservices. You stand up an Nginx or AWS ALB round-robin load balancer, spin up multiple GPU workers running vLLM, Triton, or Ollama, and di00
KDkalyan dahakeinaekshant.hashnode.dev·May 10 · 6 min readBeginner-Friendly MLOps Project: Reproducibility with Git, DVC & AWS S3If you’re starting MLOps, don’t begin with Kubernetes or fancy pipelines. Start with this question: “Can I recreate my ML project exactly after 6 months?” If the answer is NO, then your ML system is00
KDkalyan dahakeinaekshant.hashnode.dev·Apr 1 · 3 min readMastering Go Concurrency: From Slices to Worker Pools (Interview + Real-World Guide)If you're preparing for backend roles, Go concurrency is not optional — it's expected. This post covers the most important concepts you must know: Slices & memory behavior Goroutines & channels Wor00
KDkalyan dahakeinaekshant.hashnode.dev·Jan 13 · 3 min read📱 Don’t Throw it Away! How I Turned My Old Phone into a Web ServerWe all have that one old Android phone lying in a drawer. It might be too slow for modern apps, but it’s actually a powerful, low-energy computer. Instead of letting it collect dust, I decided to turn mine into a live web server for my portfolio. If ...00
KDkalyan dahakeinaekshant.hashnode.dev·Dec 24, 2025 · 4 min readFrom PDFs to Vectors: Building a Production-Grade LLM Ingestion Pipeline on AWS (Handling Data Drift the Right Way)Large Language Models are only as good as the data you feed them.In real systems, that data never stays static — new documents arrive continuously, formats change, and pipelines must adapt automatically. In this blog, I’ll walk through how I built an...00