Specifically, LLMs are vulnerable to prompt injection attacks that can "jailbreak" aligned models by transiently removing safety guardrails. These can be hand-crafted or automatically generated prompts; many of them are effective across LLMs, including state-of-the-art models. 8/
0 likes 1 replies
?