The reframe that "creative versus precise" is a setting you control rather than a mood the model is in is the mental model most people are missing. The part I would add from testing is that temperature also quietly wrecks reproducibility in evals: benchmark at any temperature above zero and you are scoring a distribution, not the model, so the numbers wobble run to run. Do you set temperature to zero for eval runs and only raise it in production, or keep them matched?