Really neat findings Oleg. Modern models improving accuracy while regressing on score reliability is a classic trade off. The Platt correction results are impressive. Did you notice any major difference in performance between temperature scaling and Platt scaling on rarer classes?