When models know whether they’re being monitored, they can downplay their capabilities in order to avoid modification or ensure deployment. This is called sandbagging. We explore recent work by @apolloaisafety demonstrating sandbagging in LLMs.
0 likes 1 replies
?