I've just read about the new math benchmark for AIs called FrontierMath and found this interesting graph on their homepage. Basically, it's saying that Sonnet 3.5 was much better compared to o1-preview, even though this area (solving hard math problems) is the specialty area of o1-preview.
1 likes 1 replies
?