Step-level agent evals exist now. Most teams still grade the finish line.
The thesis: an agent that fails at step 2 of 7 and an agent that fails at step 7 of 7 get the same score from an outcome eval, and they should not. The first one picked the wrong tool while holding th
jamesoconnor.hashnode.dev8 min read