Great question, and I didn't cover it in the article. Across major versions, the Opus version was deliberate and dated. Every Opus bump is an explicit version in the code with a date: Opus 4.6 in March, then 4.7, 4.8 in May and 5 in July. Labels land in a per-annotator path (Hive-style S3 partitions: against=opus-4.8/, against=opus-5/), and mixing two partitions was forbidden by rule (defined and used by the ML agent, see the "Building It With AI" section). Also, an Opus version bump had to pass a Cohen's κ check on a reference set first (a measure of how much two labellers agree). For instance κ(Opus 4.8, Opus 5) = 0.813 for n=120, against a 0.80 floor (on my frozen data set). That pass is also why I kept the frozen recall benchmark rather than rebuilding it, so those 477 pairs are still Opus-4.8-labelled. Note that was a judgement call, not a measurement. But within a major version, you're right, and I had no protection. I used the short alias, not a dated snapshot, so a checkpoint moving under claude-opus-4-8 would have been invisible to me. How do you handle it on your side? Any unattended agent calling a model has the same problem: dated snapshots, or accept the drift and check the output?