Anthropic’s AuditBench Signals a More Practical Test of AI Alignment Auditing
Anthropic’s AuditBench research points to a more demanding way of assessing AI alignment: testing whether evaluation tools help an investigator uncover problematic hidden behaviors, rather than treati
scalevise.hashnode.dev7 min read