Search posts, tags, users, and pages
Ali Farhat
Founder @Scalevise | We build smart automations, custom tools & AI agents for companies that want to scale faster.
Anthropic's alignment research offers a cautionary look at how reward hacking can shape AI agent behavior in cyber-related evaluations. Its results compare an early Opus 4.8 initialization called Init
No responses yet.