Simon P. Couch @simonpcouch.com · May 28

Re-ran this eval against Opus 4.8, Gemini 3.5 Flash, and GPT 5.5. Opus 4.8 is a modest improvement over the previously tested Opus models, but Gemini 3.5 Flash is the real stand-out! simonpcouch.github.io/bluffbench/

37 likes 5 replies

?

Replies

Jeremy Allen · May 28

I’ve been using Gemini 3.5 Flash in Antigravity CLI for a few days. It’s great at knowledge work and writes prose better than the Anthropic models IMHO but I haven’t tried Opus 4.8 yet. The Gemini model is so fast and relatively cheap, and it’s doing great at orchestrating multi-agent workflows.

Victor Geislinger · May 28

Totally reminds me of the 'gorilla hiding in the data' see LLMs, you're not always so sharp either ; )

Viktor Rognås · May 28

Are the ”high” thinking modes similar enough to compare; would it make sense to also compare the highest available thinking (eg max for opus 4.8), since that is what users probably will use for tricky data assessments?

Viktor Rognås · May 28

How well would a human perform on this test?