Thanks Brian. The "replay why it flagged a cluster" part is the piece most agent stacks skip, and it is why I keep tool calls in the database instead of a log file.
In CerebrumKit each tool execution is persisted as its own row - the sender is the tool, the receiver is the calling agent, and the row carries the input and the serialised output. Reconstructing a run is a query rather than a log grep. And because a tool is itself a database row (an OpenAI function spec plus a Python body), the trail points at a specific version of the tool, not at whatever happened to be deployed that day.
Honest gap: that recording sits behind a show_tool_output flag today, so it is debug-mode only, and a multi-agent hand-off still lands as separate chats rather than one trace. Making the trail unconditional and giving it a run id is the next thing I want to fix.
Repo, if it is useful github islomkhon/CerebrumKit