Raphaël Millière @raphaelmilliere.com · Jun 10

The problem is that these norms often conflict. For example, a request for dangerous information (violating “harmlessness”) can be framed as an educational query (appealing to “helpfulness”). Many issues with LLM behavior can be framed through these normative conflicts. 3/13

2 likes 1 replies

?

Replies

Raphaël Millière · Jun 10

Adversarial prompt can exploit these normative conflicts to “jailbreak” models into producing harmful outputs. They often frame a malicious request within a context that makes one norm (e.g., helpfulness) more salient than another (e.g., harmlessness). 4/13