The exchange in the responses is the most useful part for me: scoring the path, not just the final answer. A clean demo prompt tells you the model can answer. It says nothing about whether the agent backs off on a 429, stops instead of looping, or carries a changed instruction forward.
When we scope AI features for clients, we now treat the eval set as part of the deliverable and agree on it before building. Twenty or thirty cases, each with a written expected behaviour, including the awkward ones: the user changes the request halfway, a tool returns garbage, the right answer is "I cannot do that". The feature is done when those pass, and the same set runs on every model swap or prompt change.
It also settles arguments that otherwise drag on. "The assistant feels worse this week" becomes "these four cases regressed after Tuesday's prompt change", which is something a team can fix. The hardest part is usually getting the business owner to write the expected answers, because that is where the real definition of the feature lives.
Growing the eval set after every user-reported bug is useful, but I would version the dataset and keep the previous release’s cases as a stable comparison slice. Otherwise a score can fall because coverage improved, or rise because difficult cases disappeared, without the underlying system changing.
For trajectory evaluation, acceptable alternatives should be explicit too. Two agents may complete the same task using different valid tool sequences, so comparing against one canonical path can penalize a better strategy. Constraints on required effects, forbidden actions and resource use are a more durable contract than exact step-for-step matching.
indiainfranotes
Demo prompts hide the failure mode. The eval that matters is a case where the user changes the ask mid task, or the tool returns a 429 and the agent has to decide whether to retry or stop. Do you score that path, or only the first clean answer? iin1006h11