The framing of TPUs as fixed-function versus the GPU's thousands of general cores is the right lens for inference cost, you're paying for flexibility you may not use. Tying it back to LLM inference economics is what makes this more than a hardware primer. When you're choosing for an inference workload, what tips you from GPU to TPU, batch size, model shape, or just the per-token math?