1 / 2
Jumped from 20 TPS to 80 TPS ⚡ on my M4 Pro. The trick is that this version uses a MoE architecture with 3 billion active parameters. That makes it way faster. I was running the base model via GGUF (ollama), and now running on MLX (vlm - because this is a vision model too).
2 likes 2 replies
?