AKAyesha Khaninayeshakoder.hashnode.dev·2d ago · 8 min readDefending LLMs: What Actually Works, and Where It Still BreaksPart 1 of this series mapped how prompt injection and jailbreaking get malicious instructions into a model and past its guardrails. This post is the other half of the same coin: what defenses engineer00
AKAyesha Khaninayeshakoder.hashnode.dev·3d ago · 8 min readBreaking LLMs From the Outside In: A Field Guide to Prompt Injection and JailbreakingLarge language models don't have a firewall. They have an instruction hierarchy, a soft, learned sense of "the system prompt outranks the user, and the user outranks whatever text shows up inside a do00