This is one of the more honest measurement posts I've read on this topic - the selection-bias and reverse-causation objections you raise against your own result are exactly the ones that needed raising, and most people wouldn't have bothered. One thing I'm curious about in the tier-2 design: since the injected top-3 set is chosen dynamically by recent violation count, a newly-written clause for a brand-new rule has to compete for one of those three slots against whatever is currently violating the most. If it doesn't win a slot early, does it get less reinforcement than an established high-violator, and does that change how fast it accumulates the three violations needed to earn a code gate in the first place? That would be a form of selection bias in which rules ever make it into Group A - gated because they violated often enough, fast enough, to win a top-3 slot, not necessarily because they mattered most. Tracking how long a rule waits before first cracking the top 3, separate from the recurrence count you're already logging, might surface that.