SDE I – Site Reliability Engineering at Dexcom, working on infrastructure that supports medical-grade, life-critical systems.
I specialize in building observable, self-healing, and automated infrastructure — the kind that pages you less and recovers faster.
What I work with at Dexcom:
Kubernetes · GCP · Pulumi · Python · Linux · CI/CD · GitOps · Datadog
What I've shipped independently:
→ Lumina — AI-powered observability platform ingesting 10K+ logs/hour across 4 microservices, reducing MTTR by 82% with OpenAI-driven RCA and GitOps auto-rollback via ArgoCD + Helm.
→ Self-Healing AWS Infrastructure — event-driven remediation system where Prometheus monitors EC2 health and triggers automated recovery via Lambda + Terraform, eliminating manual intervention.
→ Service Mesh Platform — Kubernetes + Istio setup with canary deployments (90/10 split), cutting incident detection from 10 min to 3 min.
Also writing about it:
· 16+ articles on Kubernetes, AWS, and SRE on Hashnode — debugging guides, cost engineering, production patterns.
Stack: Kubernetes · Docker · Terraform · AWS · GCP · ArgoCD · Helm · Prometheus · Grafana · Elasticsearch · Istio · Python · Go · Linux
📩 davesaurabh59@gmail.com