MHMuhammad Hassaan Javedinblog.infraforge.agency·20h ago · 12 min readHow to recover pods a ConfigMap hook race left with empty envIf a Helm rollback of your service left a subset of pods in CrashLoopBackOff with empty database credentials while the rest keep serving, you are looking at a ConfigMap deletion race, not a bad rollba00
MHMuhammad Hassaan Javedinblog.infraforge.agency·23h ago · 11 min readWorker drops jobs intermittently: startup race, env drift, schema skewIf your worker is dropping jobs intermittently, sometimes crashing at boot with 'Error 111 connecting to redis:6379. Connection refused' and sometimes producing an empty result.json with no error at a00
MHMuhammad Hassaan Javedinblog.infraforge.agency·Jul 16 · 8 min readHow a missing S3 gateway endpoint route quadrupled our NAT billThe finance lead asked why AWS charged us $2,100 for NAT gateway data processing last month. Our normal was around $400. Nothing in the release calendar explained it: no new services, no traffic bump 10
MHMuhammad Hassaan Javedinblog.infraforge.agency·Jul 16 · 12 min readHow one stuck PDB doubled our EKS autoscaler bill in five daysThe AWS bill for EKS Compute went from $4,200 to $9,800 in five days. Same cluster, no launch, no traffic event, no team asking for headroom. kubectl get nodes returned 60 m5.4xlarge instances where t00
MHMuhammad Hassaan Javedinblog.infraforge.agency·Jul 13 · 11 min readRecovering a status page from a half-finished schema migrationThe log line was 'database schema version 23 is dirty, refusing to start' and the pod exited immediately after printing it. The team had already tried a Helm rollback to the previous chart version. Th00