This is a great reminder that once AI agents start doing real work, “it ran successfully” isn’t really enough information anymore.
I particularly liked the focus on tracing the agent’s full journey rather than treating the LLM call as the only interesting part. Seeing the tool calls, decisions, errors, latency, and downstream effects in one trace makes debugging a completely different problem.
The governance side is important too. An agent can technically make a successful tool call while still doing something outside the intended scope, so having that context available in the trace feels much more useful than simply logging prompts and responses.
The OpenTelemetry approach also makes a lot of sense here. If agent observability can build on infrastructure teams already understand instead of creating another completely separate monitoring stack, adoption becomes much easier.
Definitely the kind of topic that becomes more important once agents move from demos into systems where people actually depend on them. Great practical guide. 👍