your agent passed the eval. it still failed the job.
Most AI projects add evals in roughly the same way.
Someone picks a benchmark, collects a set of prompts, runs a model, and turns the results into a percentage. The percentage becomes a line on a dash
blog.langersword.in12 min read