the receipts here are one person's self-reported 43k-call log pulled off reddit, not something independently reproduced, and that matters more than the headline numbers -- a 65-day window is plenty of time for a team's own prompt and task mix to drift toward shorter, simpler calls without any vendor throttling involved, and that alone would move a median thinking-token count. worth treating this as a hypothesis rather than a confirmed pattern until someone runs a controlled replication: same fixed prompt, same params, same task difficulty, sampled at a fixed cadence, so the comparison is against a stable baseline instead of whatever mix of real traffic happened to land that day. a first-party canary like that is cheap to build and gives you a number you can actually trust when a model quietly changes under you, vendor-reported or not.