iury souza @iurysouza.dev · Mar 2

Jumped from 20 TPS to 80 TPS ⚡ on my M4 Pro. The trick is that this version uses a MoE architecture with 3 billion active parameters. That makes it way faster. I was running the base model via GGUF (ollama), and now running on MLX (vlm - because this is a vision model too).

2 likes 2 replies

?

Replies

Sam Edwards · Mar 2

Is that possible with ollama? I'm still new to local models and ollama makes that easy for me still. I saw there is Qwen3.5 ollama.com/library/qwen... and tried the lowest one on a M4 but was still slow out of the box.