Agreed, and that's the bar I'm holding it to. From day 01 the loop only counts a step when a tool result comes back, and the map has to point at what it read: a file and a line, not a summary of a summary.
What I'd refuse to call done on day 30:
- a map whose claims don't cite a file and line the agent actually opened
- a "better" version that hasn't been scored against the eval set
- any version that can write to the repo it's reading. It writes docs, never code.
If it can't show its evidence, it's narration, and you're right that 30 days of that would teach the wrong thing. Thanks for pushing on it.