YLYaoshen Luoinluoyaoshen.hashnode.dev·2d ago · 9 min readBenchmarking Real Work Case 2: Why Good Agent Evaluation Wasn’t Enough for Production"They used it once and refused to touch it again." When that feedback landed two weeks in, I could barely believe it. An agent validated against a benchmark built on real-world data had still failed 00
YLYaoshen Luoinluoyaoshen.hashnode.dev·2d ago · 6 min readHow to Evaluate AI Agent with Benchmark in Practice?Is your agent worth evaluating? If so, how much should you invest—and how do you ensure the evaluation actually serves the business? 1. Is Your Agent's Work Worth Evaluating Yet? Many teams miss this00
YLYaoshen Luoinluoyaoshen.hashnode.dev·2d ago · 6 min readWhy AI Agent Benchmark MattersYou probably are not short of demos. What you lack is judgment. In 2026, competitive organizations have already put AI agents into real workflows. The question is no longer whether they can do it, bu00
YLYaoshen Luoinluoyaoshen.hashnode.dev·3d ago · 7 min readBenchmarking Real Work - Case 01: How I Built a Voice Agent Benchmark from Real Customer FailuresCustomer complaints are not the problem definition; they are the signal. This post captures a real-world case of agent benchmarking: how I built v1 of our benchmark from real customer failure logs—wi11L