I thought I had already replied to this, but thanks for the comment by the way. The short answer, in my humble opinion, is that there will be some gates that may be tightly coupled to a spec, and some that are spec-agnostic - a gate associated with a linter, for example. I'm currently looking at a publicly available repo, getting my harness (Claude Code) to recreate each PR based on a spec derived from each PR's documentation, and then measuring both the original PR and the harness-generated one against CRAP tests, mutation tests, linter output, the metrics used by SlopCodeBench, etc. The art in doing this is to do it in such a way that the harness does not realise what the test is and game it by copying the original PRs verbatim.
Adam Lewis
Product Engineer/Architect navigating the AI revolution
A gate is only as good as what the spec says when it's wrong. That part stayed open. If the spec is the thing slop gets measured against, what happens when the agent satisfies every check and the feature is still the wrong one? Can the gate fail a spec, or only the code written against it?