Growing the eval set after every user-reported bug is useful, but I would version the dataset and keep the previous release’s cases as a stable comparison slice. Otherwise a score can fall because coverage improved, or rise because difficult cases disappeared, without the underlying system changing.
For trajectory evaluation, acceptable alternatives should be explicit too. Two agents may complete the same task using different valid tool sequences, so comparing against one canonical path can penalize a better strategy. Constraints on required effects, forbidden actions and resource use are a more durable contract than exact step-for-step matching.