Deleting the "double-check your answer" lines is the tip I would act on first. Those prompts were written for models that under-verified; on a model that already self-verifies they read as an instruction to keep going, which is where a chunk of the over-engineering you describe comes from. Same experience on subagent sprawl: capping how many it may spawn changed our bill more than any effort-level setting did. One place I would soften the verdict: 55.2% vs 61.1% recall makes it sound unfit for code review, but a precision-heavy reviewer works well as the second voice in an ensemble where a cheaper high-recall pass goes first. As the only safety net, agreed, no.
Deleting the "double-check your answer" lines is the tip I would act on first. Those prompts were written for models that under-verified; on a model that already self-verifies they read as an instruction to keep going, which is where a chunk of the over-engineering you describe comes from. Same experience on subagent sprawl: capping how many it may spawn changed our bill more than any effort-level setting did. One place I would soften the verdict: 55.2% vs 61.1% recall makes it sound unfit for code review, but a precision-heavy reviewer works well as the second voice in an ensemble where a cheaper high-recall pass goes first. As the only safety net, agreed, no.