PMParsa Mohammadiintomosu.hashnode.dev·7h ago · 1 min readProduction Reliability vs Static AnalysisStatic analysis examines source code without executing it. It can catch coding errors, suspicious patterns, type problems, security issues, maintainability problems, and policy violations. What it doe00
PMParsa Mohammadiintomosu.hashnode.dev·1d ago · 1 min readProduction Reliability vs ObservabilityObservability helps engineers understand what is happening in a running system through logs, metrics, traces, and other telemetry.Production reliability asks a different question: what does a change m00
PMParsa Mohammadiintomosu.hashnode.dev·2d ago · 1 min readProduction Reliability vs Code ReviewCode review is primarily about the implementation: is the logic correct, is the code understandable, does it follow conventions, and are there obvious bugs?Production reliability adds a different laye00
YUYasvanth Udayakumarinyasvanth.hashnode.dev·3d ago · 11 min readFrom SRE to ARE: Keeping AI Agents Reliable When They Can ActPicture this. A customer asks a refund agent for a simple refund. The API is fast. Every service is up. The dashboard is green. And yet the agent refunds the wrong order. A request times out, so it tr02F
PMParsa Mohammadiintomosu.hashnode.dev·3d ago · 2 min readWhat Is Production Reliability?Most teams have several ways to decide whether a change is ready to ship. There is code review, automated testing, static analysis, security checks, CI, observability, and deployment controls. All are00
Jjudeinsoit-ai.hashnode.dev·Sep 24 · 21 min readmax_retries=2, one attempt: counting retries across three layers of an agent runtimeThe short answer How many times does one tool call actually execute in our agent runtime? Exactly once — while the policy gateway's constructor says max_retries: int = 2. Those two facts do not contra00
AAbhijeetinthinkinsystems.hashnode.dev·Sep 24 · 19 min read10 Critical Software Architecture Gaps That Quietly Become Business ProblemsIntroduction Most software systems do not fail because someone forgot to write code. They fail because the system gradually accumulates architectural gaps. A feature is added. Then another. A database00
WWaLookupinwalookup.hashnode.dev·Sep 20 · 5 min readArchitectural Decision Memo: Mitigating Cascading Failures in Synchronous Metadata ResolutionArchitectural Decision Memo: Mitigating Cascading Failures in Synchronous Metadata Resolution Context and Problem Statement In high-traffic distributed systems, we often rely on external metadata reso00
SCSherdil Cloudinsherdil-cloud.hashnode.dev·Sep 19 · 5 min readYour Disaster Recovery Plan Is a Hypothesis Until You've Tested ItTL;DR: Every system fails eventually, the question is what happens next. Resilience isn't the absence of failure, it's the ability to absorb it without an outage. Five pillars deliver it (redundancy, 00
DGDebashish Ghosalinpragmatic-engineer.hashnode.dev·Sep 16 · 4 min readOne Hung API Call Used to Kill My 1,000-Run Benchmark. Here's the Fix.The experiment was fine. The runner was the bug. A single hung API call discarded hours of completed work, and I kept blaming the data. If your long-running eval keeps dying at 80%, this is the five-p00