Kartik N V J K
AI Developer | Making AI reliable, trustworthy & accessible to everyone | Active community contributor
The answer-slip rule on the operational-error path is a useful invariant, but I would test it with one model response containing several tool calls. If the first operational failure raises immediately, adding its error ToolMessage still leaves the later calls in that same assistant message without answer slips.
A second run using the retained memory may then be rejected even though the failed call itself was paired correctly. One option is to append explicit not-executed results for the remaining calls before aborting; another is to discard that unfinished round before reuse. The State objects make that cleanup policy a concrete transition worth testing with a scripted model.
Your point that the trace has to be recording before you know there is a bug is the whole case for observability in agent loops, since a model's mistake may never reproduce and you get one shot to capture it. Wiring the Observer pattern so logging and trace files subscribe without the loop knowing who is listening keeps that always-on without coupling. Do your observers capture tool inputs and outputs, or just state transitions? The former is where the real post-mortem evidence usually lives.