The harness is where most of the engineering lives, and the eval layer is what tells you whether the harness is actually working. I have seen agents where the runtime was clean but the offline evals scored the wrong axis, so production failures accumulated while the dashboard stayed green.