Raphaël Millière @raphaelmilliere.com · Nov 7

Specifically, LLMs are vulnerable to prompt injection attacks that can "jailbreak" aligned models by transiently removing safety guardrails. These can be hand-crafted or automatically generated prompts; many of them are effective across LLMs, including state-of-the-art models. 8/

0 likes 1 replies

?

Replies

Raphaël Millière · Nov 7

Why do these attacks work so well, and why are they so hard to mitigate? In a nutshell, they exploit the very feature that makes LLMs so versatile and useful in the first place: their ability to learn "in context", that is, to induce tasks from the prompt in the forward pass. 9/