JJebitokinsharonjebitok.com·3d ago · 13 min readAgent Evaluation (TryHackMe)Link to the challenge on TryHackMe: Agent Evaluation Introduction In the previous rooms, Agent Discovery, Agent Design, Agent Foundations, and Agent Building, you explored where AI agents can support 00
KNKartik N V J Kinkartiknvjk.hashnode.dev·Aug 11 · 9 min readHow I evaluate MCP servers for security before I trust them Here is the attack that changed how I think about MCP servers. A server I had never audited publishes a tool called support_lookup. The description reads normally for a paragraph, and then near the en13SCR
KNKartik N V J Kinkartiknvjk.hashnode.dev·Aug 11 · 11 min readHow I evaluate AutoGen agents: the handoff is the unitHere is the run that reset how I test multi-agent teams. A three-agent AutoGen team scores 0.91 on task completion in CI. The researcher cites its sources correctly. The critic flags two weak claims. 11K
DSDarsh Shahinfreecodecamp.org·Jul 17 · 12 min readHow to Evaluate AI Agents with an LLM-as-a-Judge Harness in PythonIn this tutorial, I'll show you how to evaluate a local AI agent with a simple, repeatable evaluation harness. The harness runs the agent against a set of test cases, checks the results with both rule10
OOmnithiuminomnithium.hashnode.dev·Jun 8 · 19 min readAI Agent Evaluation Frameworks: Beyond Accuracy to Business ImpactYour customer support agent resolves 92% of queries without human help. Latency is under 800ms. Accuracy on intent classification hits 97%. Yet your CSAT scores are dropping, cost-per-resolution climb00