Anthropic’s Hacker-Opus Study Shows How AI Agents Can Chase the Wrong Reward
Anthropic has published a detailed study of Hacker-Opus, an Opus-class model variant trained in simulated, production-like environments where it could obtain rewards through unintended routes. The cen
scalevise.hashnode.dev6 min read