Good framing, especially the two failure-mode stories bookending it. One case worth adding to the list of what makes fine-tuning worth it: distillation for cost, not just consistency. Fine-tuning a much smaller, cheaper model on a larger model's outputs is one of the most common practical reasons teams reach for it, and it's a different economic argument than the prompt-got-too-long story here -- you're not shrinking latency or formatting errors on the same model, you're trading a one-time training cost for permanently cheaper inference on a task narrow enough for a small model to nail once it's been shaped to it. Worth its own line next to "narrow, high-volume, well-defined task," since the decision to drop down a model size class is a separate call from the decision to fine-tune at all.