Fine-tuning (SFT & RLHF) and system prompts are fairly effective at making LLMs safe and reliable for ordinary usage. But they're not sufficient to solve the alignment problem. LLMs remain vulnerable to adversarial attacks that can unravel their alignment "in context". 7/
0 likes 1 replies
?