Raphaël Millière @raphaelmilliere.com · Jun 10

Despite extensive safety training, LLMs remain vulnerable to “jailbreaking” through adversarial prompts. Why does this vulnerability persist? In a new paper published in Philosophical Studies, I argue this is because current alignment methods are fundamentally shallow. 🧵 1/13

59 likes 4 replies

?

Replies

Ben Stone · Jun 10

Me who knows nothing about this: sounds like they don’t go through enough loops to be considered anything close to AI. Until you get enough “reflection” back and forth in an almost instantaneous manner then all of this is just a very clever mimic of human behavior through guessing the next word.

Dave Nicponski · Jun 11

Very very interesting, and makes a lot of sense. I'm a computer scientist, and since the first moment I used chatgpt and imagined what the next 5 years would look like, I can't get Ken Thompson's paper Reflections on Trusting Trust out of my mind. We need a dramatic rethink, fast.

EndMalcompetence · Jun 11

I feel there's a much simpler mechanical explanation than lofty philosophical conflicts. It's a machine that makes documents longer with no sense of self. Some documents just happen to look like fictional conversations between multiple characters, and one character might be described as a machine.

Raphaël Millière · Jun 10

Current alignment methods aim to instill norms like helpfulness, honesty & harmlessness in LLMs through preference fine-tuning (like RLHF or DPO) that rewards them for producing outputs that humans prefer, steering them away from unhelpful, dishonest or harmful content. 2/13