MTMRIDUL TIWARIinmriduliti.hashnode.dev·1d ago · 7 min readWhen Kafka Hits 100% Disk and the Volume Won't GrowThe alert didn't come from consumer lag. It came from disk — two of three brokers on our production Kafka cluster reporting /data at 100%, with about 20K free on an 850G EBS volume. That's not "we sho00
MTMRIDUL TIWARIinmriduliti.hashnode.dev·Aug 22 · 8 min readWhy our liveness probe timed out even though /healthCheck never touched RedisThe alert looked like a dependency outage. Readiness and liveness on our Java service were failing in bursts — context deadline exceeded (Client.Timeout exceeded while awaiting headers) — and around t00
MTMRIDUL TIWARIinmriduliti.hashnode.dev·Aug 16 · 6 min readOne Helm template change broke every Argo appThe Slack thread started the way these things usually do: three people pasting the same Argo CD error within a few minutes of each other. Sync failed. Not one app — several. All on the same shared EKS00
MTMRIDUL TIWARIinmriduliti.hashnode.dev·Aug 1 · 9 min readEight Jenkins masters, one runbook, and the day AL2023 refused to behave like Amazon Linux 2The alert wasn't a pager — it was a 404 during dnf install jenkins on the first clone target. We were mid-wave on a Jenkins master migration: eight production controllers, each moving from an old host00
MTMRIDUL TIWARIinmriduliti.hashnode.dev·Jul 25 · 7 min readThe cron was fine — our log lifecycle was eating 51 GB of ghost diskThe third pager this week landed at 2:47 AM with the same subject line: root filesystem at 97%. I'd already verified the hourly S3 upload cron twice. sudo crontab -l showed the job firing at :30 every00