The point about training or evaluating with one harness and deploying with another deserves more attention. It's the agent version of train/serve skew: tool timeouts, context limits and approval rules differ between the eval run and production, so the scores describe a system nobody shipped. Reusing one harness core for both is the cleanest fix, and your point about skills carrying local expectations (how this team reviews migrations) is the part models won't absorb anytime soon.
The point about training or evaluating with one harness and deploying with another deserves more attention. It's the agent version of train/serve skew: tool timeouts, context limits and approval rules differ between the eval run and production, so the scores describe a system nobody shipped. Reusing one harness core for both is the cleanest fix, and your point about skills carrying local expectations (how this team reviews migrations) is the part models won't absorb anytime soon.