MetaRoCE: Why a Million-GPU Cluster Must Tolerate Loss
I keep a small rule when reading distributed-systems headlines: if the design needs every packet to behave politely, the design has already picked the wrong scale. That rule came back to me when I rea
hironakamura-ai.hashnode.dev7 min read