LLaMa 4 beats Deepseek-R1 and V3 on coding?! Fresh evaluation using our #1 SWE-bench coding agent and >200,000,000 tokens of production-grade code: Claude 3.7 Sonnet > Gemini 2.5 Pro > LLaMA 4 Maverick > DeepSeek-V3 > Deepseek-R1
0 likes 0 replies
?
LLaMa 4 beats Deepseek-R1 and V3 on coding?! Fresh evaluation using our #1 SWE-bench coding agent and >200,000,000 tokens of production-grade code: Claude 3.7 Sonnet > Gemini 2.5 Pro > LLaMA 4 Maverick > DeepSeek-V3 > Deepseek-R1
0 likes 0 replies