© 2026 Hashnode
Walk into any talent acquisition retrospective after a failed engineering hire, and you will likely hear variations of the same corporate post-mortem: “They had a stellar pedigree, but they just didn’

If you’re building AI agents and need production-grade open-source observability and evaluation without vendor lock-in, Arize Phoenix is worth exploring. Phoenix is an open-source AI development platform built on OpenTelemetry and OpenInference instr...

Hallucinations are one of the biggest challenges in production AI agents. Here's how to detect them with Python: Five Types of Hallucinations in AI Agents Context Hallucinations: The agent invents facts not in the provided context. Tool Hallucinati...

Building AI agents is hard. Evaluating them is harder. Most teams I talk to are evaluating their agents the wrong way. They look at the final output and ask, "Is it correct?" But that's like grading a math test by only looking at the final answer, no...

When you start building with LLMs, it quickly becomes clear that not all models behave the same. One model may excel at creative writing but struggle with technical precision. Another might be thoughtful yet verbose. A third could be fast and efficie...

Written by Ian Unsworth at MindsDB AI evaluations ("evals") have emerged as the critical success factor distinguishing enterprises that successfully deploy AI solutions from those trapped in pilot purgatory. Only 10% of enterprises have Gen AI in pro...
