Separating "using an open model" from "operating an open model" is what makes this work. The per-hour versus per-token arithmetic misleads precisely because it prices the first and quietly ignores the second.
The number that decides it in practice is utilisation. A rented H100 costs the same at 5% as at 95%, so self-hosting only wins above a fairly high floor of steady traffic, and bursty workloads pay for idle silicon overnight. The other cost nobody quotes is the evaluation and rollback machinery you now own: when a serving change regresses quality, a provider's version pin is somebody else's problem and your own weights are not. Managed open weights sits in the middle for exactly that reason, keeping model access without buying the operations.
Separating "using an open model" from "operating an open model" is what makes this work. The per-hour versus per-token arithmetic misleads precisely because it prices the first and quietly ignores the second.
The number that decides it in practice is utilisation. A rented H100 costs the same at 5% as at 95%, so self-hosting only wins above a fairly high floor of steady traffic, and bursty workloads pay for idle silicon overnight. The other cost nobody quotes is the evaluation and rollback machinery you now own: when a serving change regresses quality, a provider's version pin is somebody else's problem and your own weights are not. Managed open weights sits in the middle for exactly that reason, keeping model access without buying the operations.