Augment Code @augmentcode.com ยท Apr 15

๐†๐๐“-๐Ÿ’.๐Ÿ ๐š๐ฅ๐ฆ๐จ๐ฌ๐ญ ๐ญ๐จ๐ฉ๐ฌ ๐‚๐ฅ๐š๐ฎ๐๐ž ๐Ÿ‘.๐Ÿ• ๐จ๐ง ๐œ๐จ๐๐ข๐ง๐ ?! New eval dropping using our #1 SWE-bench coding agent! - GPT-4.1 beats Gemini 2.5 Pro and almost tops Claude 3.7 Sonnet! - Even GPT-4.1 mini matches Claude 3.5 Sonnet V2 performance. It was the top model just 2mo ago!

0 likes 1 replies

?

Replies

Augment Code ยท Apr 15

The evaluation is done through our proprietary codebase understanding benchmark AugmentQA. You can learn more at: www.augmentcode.com/blog/you-mak... Try our agent yourself at: augmentcode.com!