Michael R. Bock @michaelrbock.com · Mar 3

3/ An interesting wrinkle on thinking budgets: Ultrathink: 49.02% Medium: 49.02% High: 47.06% Lobotomized: 37.25% Low: 35.29% Medium thinking matches ultrathink exactly. More thinking doesn't always mean better, but some thinking is critical. The jump from low to medium is +14 points.

0 likes 1 replies

?

Replies

Michael R. Bock · Mar 3

4/ Updated rankings (strict: every line must be correct): Opus 4.6: 52.94% Gemini 3.1 Pro: 49.02% <-- new GPT-5 w/ Search: 41.67% GPT-5.2 Pro: 41.18% Sonnet 4.6: 37.25% Gemini 3 Pro: 36.27% Opus 4.5: 36.27% The gap at the top is closing fast. 8 months ago, 32% was SOTA.