Flashes Home Search Notifications Sign in Post
meowtase @sus.cat · Apr 10

TLDR: If you use a linear probe, it can only learn things that are easy to find in the model. If you use a non-linear probe, the probe might be learning things that the model wasn't thinking of I briefly walk through a fictional example, and a real OthelloGPT example

0 likes 1 replies

?

Replies

meowtase · Apr 10

https://open.substack.com/pub/suscat/p/linear-vs-non-linear-probes-for-interpretability?utm_campaign=post-expanded-share&utm_medium=web

Legal

Privacy Policy Terms and Conditions

Contact

FAQs Feedback

Follow

Bluesky Instagram Threads
Flashes for Bluesky Get it on the App Store

Feeds

  • Global
  • Trending
  • Blacksky Images
  • Photography
  • Discover
Browse all feeds
Feedback Ideas, bugs, and what ships next
Privacy Policy Terms and Conditions FAQs Bluesky Instagram Threads
Get it on the App Store