The nightly precision audit is the part I'd want to stress-test hardest: it calls Opus fresh each night to re-judge delivered leads, while the recall benchmark runs off labels Opus produced once and froze. Was that nightly call pinned to a specific Opus checkpoint the whole seven months, or the latest alias? If Anthropic moved the model under it at any point, the audit's meaning shifts even though nothing in the pipeline changed -- and with 339k lines nobody read, that audit was the only signal anyone had that anything upstream was still doing its job.
Vivian Oliveres
Sharing thoughts as a Solo Developer