You should score that path, and it's where most of the signal is. I'd add dedicated trajectory cases for exactly those scenarios: a mid-task change of request, a tool returning a 429 or timeout, a malformed tool response.
What gets scored there isn't the final answer alone. It's the decisions: did it back off and retry a bounded number of times, did it stop and report instead of looping, did it carry the new instruction forward, and did it avoid repeating side-effecting calls. Rule-based checks cover most of it (call counts, retry limits, no duplicate writes), with an LLM judge only for the fuzzier "did it handle the changed ask correctly" part.
Clean-path-only evals tend to look great right up until a flaky dependency shows up in production. How are you injecting failures in your setup, mocked tools or real chaos in staging?