"Open is not the same as free" is the right thesis, and starting from the token is the correct move since most cost confusion is really unit confusion.
The number that decides it in practice is utilisation. A rented GPU costs the same at 5% as at 95%, so self-hosting only wins above a fairly high floor of steady traffic, and bursty workloads pay for idle silicon overnight.
Two costs that rarely make these comparisons: the evaluation and rollback machinery you now own, since a serving change that regresses quality becomes your problem rather than a provider's, and the engineer-hours spent keeping it running, which at small scale are usually larger than the GPU bill.