DECEPTIVE ALIGNMENT IN INTELLIGENT MODELS
Over the years, in our daily interactions, we must have come across people who seem to share our values or at least try to make us think they do. However, we may eventually realise that they never tru
josiahwrites.hashnode.dev10 min read
konyhea
Interesting, the 14% compliance rate in the Anthropic paper, any sense of whether that number scales up or down with model capability?