PMParsa Mohammadiintomosu.hashnode.dev·1d ago · 1 min readProduction Reliability vs ObservabilityObservability helps engineers understand what is happening in a running system through logs, metrics, traces, and other telemetry.Production reliability asks a different question: what does a change m00
MRMira royinainetwork.hashnode.dev·2d ago · 6 min readA Reference Pattern for Logging and Auditing LLM API CallsMost teams log application traffic. Far fewer can reconstruct exactly what happened when their application called an LLM three months ago. That distinction matters. If an auditor, security engineer, o00
MTMuhammad Tahirinmtdeveloper.hashnode.dev·2d ago · 12 min readFull-Stack Observability: OpenTelemetry, Structured Logging, and Tracing in 2026This article was originally published on Muhammad Tahir's Portfolio. Introduction & Industry Context In the rapidly evolving landscape of 2026, where microservices architectures and cloud-native deplo00
YUYasvanth Udayakumarinyasvanth.hashnode.dev·3d ago · 11 min readFrom SRE to ARE: Keeping AI Agents Reliable When They Can ActPicture this. A customer asks a refund agent for a simple refund. The API is fast. Every service is up. The dashboard is green. And yet the agent refunds the wrong order. A request times out, so it tr02F
SSSarmeet Singhinblueprintsofscale.hashnode.dev·3d ago · 47 min readWhat Broke at 3 AM: Observability, p99, and Useful Alerts, Explained Like You're NewThe page comes at 3:07 AM: "the site is slow." You open the dashboards. CPU: fine. Memory: fine. Error rate: flat. Every graph is green, and the site is still slow. Three engineers stare at the same s00
HGHarrison Guoinharrisonsec.hashnode.dev·4d ago · 12 min readA Wrong Ruler Is Worse Than No Ruler: Verifying the Checks You TrustThere is one failure mode I have learned to fear more than a missing check. A system with no check for something is at least honest about it. The gap is visible, the uncertainty is real, and everyone 01F
HGHarrison Guoinharrisonsec.hashnode.dev·6d ago · 9 min readYour AI Bill Is a Distributed Systems Problem, Not a Model-Pricing ProblemA team I was helping watched their model bill jump to several times its usual size in a single month. The token meter had not predicted it. The first question in the room was the one almost everyone a00
ASAmit Shuklainamitshuklabag.hashnode.dev·Sep 24 · 5 min readStructured logging: why print statements do not scale past one serviceconsole.log('user logged in', userId) is fine when there is one process and one terminal. The moment there are two services, three instances of each, and a log aggregator collecting all of it into one00
OMOyugi Mouriceinoyugimourice.hashnode.dev·Sep 23 · 10 min readAI Agent Observability and Governance: Building AI Systems That Don't Fail SilentlyTL;DR: Dashboards are green. Your agent is responding. But users are getting fluent, completely wrong answers and your logs show nothing. This is the silent failure trap. Here's a framework for observ03MB
RARahiel Akhtarinrahielakhtar.no·Sep 21 · 4 min readWhat's actually crossing your ExpressRoute? Traffic Collector, with IaCIntroduction ExpressRoute traffic visibility is a gap most Azure observability doesn't fill - As a platform engineer i'll be sharing my findings and observations setting up ExpressRoute Traffic Collec00