TLDR: If you use a linear probe, it can only learn things that are easy to find in the model. If you use a non-linear probe, the probe might be learning things that the model wasn't thinking of I briefly walk through a fictional example, and a real OthelloGPT example
0 likes 1 replies
?