LLMs' vulnerability to in-context misalignment is far from a solved problem. While I wouldn't go so far as to claim it's unsolvable, current alignment strategies are insufficient and even exploitable to recover misaligned behavior in context. 15/
0 likes 1 replies
?