Good call waiting for independent tests. First thing I would check is long-output stability: does it keep a 2,000 line refactor consistent end to end, or does quality drop after the first chunk the way you describe with "continue"? Second is cost per finished task, not per token, since a bigger context only helps if it cuts retries. Which independent eval are you treating as the real signal?
iin1005h21