Separating "using an open model" from "operating an open model" is what makes this work. The per-hour versus per-token arithmetic misleads precisely because it prices the first and quietly ignores the second.
The number that decides it in practice is utilisation. A rented H100 costs the same at 5% as at 95%, so self-hosting only wins above a fairly high floor of steady traffic, and bursty workloads pay for idle silicon overnight. The other cost nobody quotes is the evaluation and rollback machinery you now own: when a serving change regresses quality, a provider's version pin is somebody else's problem and your own weights are not. Managed open weights sits in the middle for exactly that reason, keeping model access without buying the operations.