The prefill-vs-generation split is the part most of these comparisons skip, and agent loops make it worse than a single chat prompt would suggest. A loop that resends a growing transcript every turn is prefill-dominated by construction, so the local machine's disadvantage compounds turn over turn instead of staying fixed. One thing the break-even table doesn't account for: prompt caching on the hosted side. Anthropic (and most frontier providers now) discount cached input tokens heavily, so a long, mostly-repeated agent transcript costs a lot less on the API than the raw per-million-token rate implies -- which pushes the break-even point further out than the $5/million-input math suggests on its own. Local inference doesn't get that discount for free either, you'd need your own KV-cache reuse across turns to close the gap, and llama.cpp's prompt caching only helps when the prefix is byte-identical between calls, which an evolving agent transcript usually isn't.