Nice test of a failure mode that benchmarks often miss: instruction following can hide contradiction blindness. I would add adversarial pairs where the answer space stays simple but the premises conflict, then report whether models ask for clarification. That signal may be more useful than another aggregate accuracy score. Did any model explain why it trusted one premise over the other?
indiainfranotes
Nice test of a failure mode that benchmarks often miss: instruction following can hide contradiction blindness. I would add adversarial pairs where the answer space stays simple but the premises conflict, then report whether models ask for clarification. That signal may be more useful than another aggregate accuracy score. Did any model explain why it trusted one premise over the other?