MHMuhammad Hassaan Javedinblog.infraforge.agency·6d ago · 23 min readKarpenter picked an instance family our Savings Plan never coveredSavings Plan coverage across the fleet went from 100% in June to about 41% two weeks into July, and not one thing on the platform team's dashboards moved. Node count flat. Pod count flat. p99 latency 00
MHMuhammad Hassaan Javedinblog.infraforge.agency·Aug 18 · 13 min readHow a do-not-disrupt annotation broke Karpenter consolidationKarpenter launched 2,800 nodes over one weekend and consolidated only 180 of them, and our EC2 on-demand bill jumped 2.2x. The cause was one line added to a shared internal Helm chart on Friday aftern10
MHMuhammad Hassaan Javedinblog.infraforge.agency·Aug 13 · 11 min readHow we recovered a k3s cluster after its client certs expired119 ephemeral preview namespaces should have been reaped between 2am and 6am. None were. The teardown cron had been failing on x509: certificate has expired or not yet valid for six hours before anyon10
MHMuhammad Hassaan Javedinblog.infraforge.agency·Aug 13 · 11 min readWhy terraform plan wants to destroy 5 live failover resourcesBy 08:47 the terraform plan output was up on the shared screen and nobody wanted to be the one to type apply. The summary line read 'Plan: 5 to add, 0 to change, 5 to destroy.' Every one of the 5 dest00
MHMuhammad Hassaan Javedinblog.infraforge.agency·Aug 10 · 12 min readHow to recover pods a ConfigMap hook race left with empty envIf a Helm rollback of your service left a subset of pods in CrashLoopBackOff with empty database credentials while the rest keep serving, you are looking at a ConfigMap deletion race, not a bad rollba00