just to remind you what this is: this single box with two SSDs. is running deepseek v4.1 flash (a 552B model) in its native MXFP4, ~550GBs of weights plus engrams total. with prefill speeds of over 450tps and decode speeds of over 15tps. at full 1M context.
Best Open LLM
by @overby.me
Tracks discussion of the current best open-weight/open-source LLM: new releases, benchmarks, leaderboards, and head-to-head comparisons across DeepSeek, Qwen, Llama, Mistral, GLM, Kimi, MiniMax, GPT-OSS, Gemma and other open models.
Le Chonk: Mistral drops a 1T beast
congratulations to mistral for releasing the first pareto roadside-ditch model
Mistral Large 4 trails Chinese open models Independent tests rank Mistral’s new Large 4 the best model outside the US and China, but several Chinese open models still scored higher overall. 📉
A 27-billion-parameter model, running entirely in your browser tab. No install, no server, no API key. 3.8 GB one-time download, then it works offline, anywhere. Bonsai(Qwen) 27B (1-bit) on WebGPU running at ~28 tok/s on consumer hardware
You can now train your own Decision model like Jev locally! We increased Qwen3.5 0.8B’s aggregate accuracy from 20.7% to 74.3% across 3 decision benchmarks - on just 4GB VRAM. Turn any LLM, Qwen3.8, Gemma 4 into decision models. GitHub: github.com/unslothai/un... Guide: unsloth.ai/docs/basics/...
Reflection AI's Beam: a highly efficient agentic open model with 501B total parameters and 23B active. - Frontier reasoning efficiency - Advances the Western open frontier on coding & agentic tasks - Trained end-to-end from scratch Full weights release this month. reflection.ai/blog/introdu...
BEAM: a 501B highly agentic LLM from Reflection AI - open weights - MoE, A23B reflection.ai/blog/introdu...
whenever i feel sad and worthless i look at open model benchmarks to cheer myself up
The Last Translation Benchmark is still open (10 accepted submissions in any language/dialect for coauthorship 🤩). last-translation-benchmark.vilda.net Also submit your model & try to beat LTB. We added all the latest models: North Small Translate (@cohere.com), Mistral Large 4, Qwen 3.8, congrats!
yes please always assume 1 Eva for our estimates
this feels like a headline you'd wake up to if you were living in a @greatdismal.bsky.social novel
🇫🇷 ✨ French AI startup Mistral on Tuesday announced its biggest open-weight model, dubbed "Chonk", claiming enhanced performance and security capabilities for clients seeking to keep their data private instead of relying on closed models from the likes of Anthropic and OpenAI ➡️ u.afp.com/S9yE