Adam Binksmith @binksmith.com · Dec 20

When models know whether they’re being monitored, they can downplay their capabilities in order to avoid modification or ensure deployment. This is called sandbagging. We explore recent work by @apolloaisafety demonstrating sandbagging in LLMs.

0 likes 1 replies

?

Replies

Adam Binksmith · Dec 20

When models know whether they’re being monitored, they can pretend to be aligned with a goal in order to avoid modification. We explore work (released yesterday!) by @Anthropic and @redwood_ai showing this behavior in LLMs. It does come with caveats that we discuss in-depth.