The KoBBQ result is the sharpest part of this for me: the jump from 95% to 0% accuracy on the exact same evidence, just from removing one option, shows abstention is a design choice, not a byproduct of confidence. What I'd want to know before trusting that 95% number in a new domain: is it doing that well because "unknown" sits in the schema, or because RLCD training actually saw enough examples where the correct label was "the evidence is insufficient" to learn that pattern specifically. If abstention has to be trained on ambiguity that looks like the ambiguity you'll actually see in production, adding an unknown option to a fresh schema in an unfamiliar domain doesn't automatically buy you that 95%, and you're back to measuring it, same as any other calibration claim in the piece.