The five-stage ladder is a useful way to frame this: check the quick wins (replicas, prefix-cache routing, queue behavior) before tearing into kernels or NCCL configs. Using kv-cache utilization as a fork between tuning and scaling is a detail I will remember.