Well said. “The deployment decision preceded the cost model” is probably the line that stood out most to me.
We're seeing a familiar pattern from the early cloud era playing out with AI: teams optimize for capability first, then try to understand consumption after the system is already running at scale.
The interesting part is that the solution isn't necessarily using cheaper models. It's building the architecture to match the right level of intelligence to the right task.
Routing, caching, cost attribution, governance and observability need to become part of the platform—not decisions every engineer has to make independently.
AI consumption isn't the problem. Unmanaged consumption is
The organizations that build that orchestration layer early will have a very different cost curve as agentic workloads scale.