Chain of Thought Reasoning Hides What AI Models Actually Do
TL;DR: Chain of Thought reasoning was supposed to make AI transparent. Anthropic’s own research shows it doesn’t. Claude mentioned planted hints in its reasoning only 25% of the time. When interpretab
toxsec.hashnode.dev5 min read