Raphaël Millière @raphaelmilliere.com · Nov 7

Fine-tuning (SFT & RLHF) and system prompts are fairly effective at making LLMs safe and reliable for ordinary usage. But they're not sufficient to solve the alignment problem. LLMs remain vulnerable to adversarial attacks that can unravel their alignment "in context". 7/

0 likes 1 replies

?

Replies

Raphaël Millière · Nov 7

Specifically, LLMs are vulnerable to prompt injection attacks that can "jailbreak" aligned models by transiently removing safety guardrails. These can be hand-crafted or automatically generated prompts; many of them are effective across LLMs, including state-of-the-art models. 8/