Michael R. Bock @michaelrbock.com · Feb 20

3/ Thinking budget matters enormously for tax. Same model, same prompt, different thinking levels: Sonnet 4.6 (ultrathink): 37.25% Sonnet 4.6 (no thinking): 19.61% Nearly 2x accuracy just from letting the model think longer.

0 likes 1 replies

?

Replies

Michael R. Bock · Feb 20

4/ Full updated rankings (using strict scoring where every line must be correct): Opus 4.6: 52.94% GPT-5 w/ Search: 41.67% GPT-5.2 Pro: 41.18% Sonnet 4.6: 37.25% <-- new Gemini 3 Pro: 36.27% Opus 4.5: 36.27% GPT-5.2: 33.82% Gemini 2.5 Pro: 32.35% 7 months ago, 32% was SOTA.