Good question, and you're right that above-zero temperature turns an eval into scoring a distribution. My take: use temperature 0 when you're comparing models, prompts, or changes, because you want differences you can attribute to the change and not to sampling noise. Even then it's not perfectly deterministic across hardware and batching, so I'd still repeat runs on anything close.
But I'd keep a matched run at your production settings too. Otherwise you're validating a configuration you never ship. If production runs at 0.7, the number that matters is the pass rate across several samples at 0.7, reported as a rate with a spread, not a single score.
So: zero for A/B comparisons, matched settings with multiple samples for the "does this feature actually work" call. How are you handling it in your evals?
The reframe that "creative versus precise" is a setting you control rather than a mood the model is in is the mental model most people are missing. The part I would add from testing is that temperature also quietly wrecks reproducibility in evals: benchmark at any temperature above zero and you are scoring a distribution, not the model, so the numbers wobble run to run. Do you set temperature to zero for eval runs and only raise it in production, or keep them matched?