Your agent eval tests whether it succeeds. It should test whether it recovers.
Most agent evaluations I've read measure one thing: given a task, did the agent complete it. That's the happy path. It's necessary and it's not enough, because in production the interesting question i
jamesoconnor.hashnode.dev5 min read