Putting the gate in code is much stronger than putting honesty in a prompt, but repeated held-out feedback is still adaptive access. Even if generated code cannot read the grader, the search policy observes scores and pass or kill outcomes across nodes, which can eventually tune itself to that slice. I would use dev for search, a query-budgeted validation set for ratcheting, and one locked final test used only once, with the ledger recording every validation access.
Ahmet Özel
AI Engineer. Computer Vision, RAG and LLM agents.
Putting the gate in code is much stronger than putting honesty in a prompt, but repeated held-out feedback is still adaptive access. Even if generated code cannot read the grader, the search policy observes scores and pass or kill outcomes across nodes, which can eventually tune itself to that slice. I would use dev for search, a query-budgeted validation set for ratcheting, and one locked final test used only once, with the ledger recording every validation access.