The agent catching a bug in your own ground truth is the most useful thing in this whole writeup, because it is the failure everyone pretends does not exist: our labels are wrong more often than we admit. Your point that ambiguity, not capability, drove the baseline errors matches what I keep finding, most eval disagreements are really spec disagreements. When the agent contradicts a label now, do you have a workflow to triage whether it is a model miss or a bad gold answer?