Michael R. Bock @michaelrbock.com · Mar 6

2/ The thinking level gap is enormous: High: 56.86% Medium: 49.02% Low: 31.37% That's a 25-point spread between low and high. Low thinking GPT-5.4 would rank near the bottom of the leaderboard. High thinking puts it at #1.

0 likes 1 replies

?

Replies

Michael R. Bock · Mar 6

3/ Some context on what just happened: Medium thinking GPT-5.4 matches Gemini 3.1 Pro's best result (49.02%) exactly. But high thinking adds another 8 points on top. This is the biggest single-model jump we've seen from a thinking level increase. Compute allocation matters a lot here.