Raphaël Millière @raphaelmilliere.com · Nov 7

LLMs' vulnerability to in-context misalignment is far from a solved problem. While I wouldn't go so far as to claim it's unsolvable, current alignment strategies are insufficient and even exploitable to recover misaligned behavior in context. 15/

0 likes 1 replies

?

Replies

Raphaël Millière · Nov 7

This doesn't bode very well for the prospect of solving the alignment problem. Future architectures might not have the same vulnerabilities, but the failure to patch these vulnerabilities with current techniques highlights how unpredictable progress in this area can be. 16/