Simon P. Couch @simonpcouch.com · May 27

New on my blog: the Claude 4 models are here! I evaluate the new releases of Sonnet and Opus against Claude 3.7 Sonnet and o4-mini on a dataset of challenging #rstats coding problems. www.simonpcouch.com/blog/2025-05...

27 likes 4 replies

?

Replies

Sharon Machlis · May 27

If I’m doing something complex, I tend to use the chatbot Web UIs so I don’t have to stress about tokens and pricing (I have flat-fee monthly subscriptions). Then I upload package docs as part of my initial prompts, even to Opus 😀 No data for this, but I feel like #RStats results are better.

A

Aymen · May 27

It's weird how there is no clear progress so far, thought Claude 4 would crush the competition.

Tyler Burch · May 27

This is awesome, thanks for putting it together. I’d be curious to see if it’s possible to add the Gemini models to the benchmark. I’ve had decent success using them to write R, but they also tend to over-engineer and burn a lot of tokens.

Platform Puff Watch · May 28

Great stuff. I feel Claude performs better with Python than R. Completely anecdotal. Have you investigated?