The exchange in the responses is the most useful part for me: scoring the path, not just the final answer. A clean demo prompt tells you the model can answer. It says nothing about whether the agent backs off on a 429, stops instead of looping, or carries a changed instruction forward.
When we scope AI features for clients, we now treat the eval set as part of the deliverable and agree on it before building. Twenty or thirty cases, each with a written expected behaviour, including the awkward ones: the user changes the request halfway, a tool returns garbage, the right answer is "I cannot do that". The feature is done when those pass, and the same set runs on every model swap or prompt change.
It also settles arguments that otherwise drag on. "The assistant feels worse this week" becomes "these four cases regressed after Tuesday's prompt change", which is something a team can fix. The hardest part is usually getting the business owner to write the expected answers, because that is where the real definition of the feature lives.