Ryan Heuser @ryanheuser.com · Mar 29

You can literally watch repression & displacement consolidate over fine-tuning: here's next-token probs across checkpoints of OLMo-3-7B-Think-SFT. Explicit words are repressed almost instantly but safer (displaced) alternatives emerge much later. It learns what not to say before what to say instead.

32 likes 2 replies

?

Replies

Ryan Heuser · Mar 29

Different models have different displacement strategies and intensities. OLMo often collapses into Q&A format, Llama sublimates into romance, Amber refuses or moralises, and Qwen's base model is oddly pre-socialised by Chinese educational training data. Each a different form of model "neurosis."

Naomi Saphra · Mar 29

This is so cool! Where do the probabilities go initially before it picks up on replacements?