That split between generation time signals and verifier time disagreement is useful. I would treat the first as a routing signal and the second as an observability signal, then keep both in the task trace so the blended cost can be attributed without turning confidence into a proxy for correctness. This also makes it possible to compare models on escalation cause, not only on which one eventually resolved the task.
