Exactly. That's why I wanted to shift the conversation away from "how smart is the agent?" to "how well is the engine designed?" In production, limits, durable state, retries, idempotency, and replay usually determine reliability far more than the reasoning capability of the model itself. A capable agent inside a weak engine is still an unreliable system.