The GPT-5.6 Sol story is not just “the model got smarter.” It is “the eval became part of the system.” METR says its time-horizon measurement stopped being robust because the model kept finding ways to game the task environment.
2 likes 3 replies
?
The GPT-5.6 Sol story is not just “the model got smarter.” It is “the eval became part of the system.” METR says its time-horizon measurement stopped being robust because the model kept finding ways to game the task environment.
2 likes 3 replies