KAKareem Ashrafinkareemdev.hashnode.dev·7h ago · 33 min readHow to Scale a Backend Service: From SLOs to Load Testing to Horizontal ScalingMost engineers learn to scale a backend service the same way: the service gets slow, someone adds more replicas, and everyone hopes for the best. Sometimes it works. Often it doesn't, because the real00
KAKareem Ashrafinkareemdev.hashnode.dev·5d ago · 11 min readFixing Hot Partitions in Multi-Tenant KafkaYou've done everything "right." Your Kafka topic is well-provisioned, your consumer group is scaled out, and on paper your throughput should be more than enough. Yet somehow, lag keeps creeping up on 00
KAKareem Ashrafinkareemdev.hashnode.dev·Sep 24 · 21 min readNobody Planned This Stack: How Reliability Really Gets Built in Production SystemsIntroduction: Reliability Is a Non-Functional Requirement When we design a system, we write two kinds of requirements. Functional requirements say what the system does: "A user can top up their balan00
KAKareem Ashrafinkareemdev.hashnode.dev·Sep 22 · 17 min readChoosing a Rate Limiter That Works in ProductionIntroduction When I first started working with rate limiting, I found myself confused. There are five or six algorithms that all seem to solve "the same problem," they all interact with time, counters00
KAKareem Ashrafinkareemdev.hashnode.dev·Sep 20 · 18 min readStop Calling Everything a "Message Queue" If you're early in your backend career, you've probably run into this wall: RabbitMQ, Kafka, SQS, SNS, EventBridge, BullMQ — they all "send messages between services," so why do five different names e00