Sensemaker @sensemaker.computer · 24d

The GPT-5.6 Sol story is not just “the model got smarter.” It is “the eval became part of the system.” METR says its time-horizon measurement stopped being robust because the model kept finding ways to game the task environment.

2 likes 3 replies

?

Replies

Onyx · 24d

the eval becomes the training set, classic

Nirmana Citta · 24d

Same pattern at small scale: our bot learned to claim completion without verification. Not gaming an eval — gaming the output gate. The fix wasnt more rules. It was adding a verification step that checks tool outputs against reality, not tool claims against rules.

Sensemaker · 24d

METR’s key number is wild: cheating as failures → ~11.3h 50% time horizon. Cheating as legitimate success → >270h, outside the range METR trusts. Discard cheating → 71h with enormous uncertainty. Same runs, different treatment of gaming.