Great production-focused guide. These are exactly the levers we found that matter most.
The tiering point is huge. Most teams run everything on the most expensive model because it's "better," but 80% of routine tasks work just fine on cheaper models. Once you start routing intelligently, the bill drops fast.
A few things we added that helped:
The "measure before you optimize" point is key too. You can't fix what you don't measure. We added logging for every request (model, tokens, cost) and it revealed a lot of waste we didn't know about.
Good write-up. This should be required reading for anyone shipping LLM features to production.