The key insight is that adding a tool is not really a code-change event; it is an authority and behavior-change event. Even if the tool implementation is tiny, the agent's reachable state space has changed, so existing prompt tests no longer represent the same system.
In production agent work at IT Path Solutions, I’d make the previous failure cases part of a persistent regression suite, but also track the tool’s permission boundary and argument space separately. A model upgrade can preserve all existing tests while still changing tool-selection behavior, so testing whether the agent refuses something isn't enough test which tool it chooses, with which arguments, and what state ultimately changes.
That gives you a much stronger release signal than a generic agent smoke test: same failures must stay fixed, and newly reachable actions must have explicit authorization coverage.
The key insight is that adding a tool is not really a code-change event; it is an authority and behavior-change event. Even if the tool implementation is tiny, the agent's reachable state space has changed, so existing prompt tests no longer represent the same system.
In production agent work at IT Path Solutions, I’d make the previous failure cases part of a persistent regression suite, but also track the tool’s permission boundary and argument space separately. A model upgrade can preserve all existing tests while still changing tool-selection behavior, so testing whether the agent refuses something isn't enough test which tool it chooses, with which arguments, and what state ultimately changes.
That gives you a much stronger release signal than a generic agent smoke test: same failures must stay fixed, and newly reachable actions must have explicit authorization coverage.