3/ Thinking budget matters enormously for tax. Same model, same prompt, different thinking levels: Sonnet 4.6 (ultrathink): 37.25% Sonnet 4.6 (no thinking): 19.61% Nearly 2x accuracy just from letting the model think longer.
0 likes 1 replies
?
3/ Thinking budget matters enormously for tax. Same model, same prompt, different thinking levels: Sonnet 4.6 (ultrathink): 37.25% Sonnet 4.6 (no thinking): 19.61% Nearly 2x accuracy just from letting the model think longer.
0 likes 1 replies