Buried in finding three is a labeling-methodology improvement your own architecture already supports. You could not honestly label the hold-out candidate at −1.6% with 5x volume, and documented it as where the ambiguity lives, which is more useful than a forced label. But your evaluator signature already returns true | false | no_data, so the ground-truth format could admit an explicit abstain class: cases where the SME declines to answer, which then test whether tools decline too, instead of quietly rewarding a confident guess on an undefined boundary. Overconfidence on ambiguous input is exactly the failure compiled tools should inherit least. "Labels are a spec, and specs have bugs" is the line I am keeping, and the mid-gap timestamp story earns it: the agent reasoned about the moment while you reasoned about the gap, and the agent read the spec correctly. Looking forward to the label-free voting variant in part 2; that is where the paper's claims get properly stress-tested.
Buried in finding three is a labeling-methodology improvement your own architecture already supports. You could not honestly label the hold-out candidate at −1.6% with 5x volume, and documented it as where the ambiguity lives, which is more useful than a forced label. But your evaluator signature already returns true | false | no_data, so the ground-truth format could admit an explicit abstain class: cases where the SME declines to answer, which then test whether tools decline too, instead of quietly rewarding a confident guess on an undefined boundary. Overconfidence on ambiguous input is exactly the failure compiled tools should inherit least. "Labels are a spec, and specs have bugs" is the line I am keeping, and the mid-gap timestamp story earns it: the agent reasoned about the moment while you reasoned about the gap, and the agent read the spec correctly. Looking forward to the label-free voting variant in part 2; that is where the paper's claims get properly stress-tested.