Good demystification, the harness really is the part that decides whether an agent survives past a handful of steps. What I would add is that the harness itself needs measuring: once you have retries, context trimming, and step limits in there, each is a knob that can quietly wreck a long run. When a run falls apart at step 80, how do you localize whether it was the model or the harness that failed?