famstack.dev @famstack.dev · Mar 13

Wow! My MLX vs llama.cpp benchmark hit #9 on r/LocalLLaMA today. Did not expect that. Takeaway: benchmark actual scenarios, do not rely on just the tok/s counter in your UI. Ran into a caching bug specific to Qwen 3.5 (35B-A3B) on MLX. Effective tokens/s is what we experience #MLX #LlamaCpp #Qwen

0 likes 1 replies

?

Replies

famstack.dev · Mar 13

#LocalAI #AppleSilicon #Mac #SelfHosted Reddit: www.reddit.com/r/LocalLLaMA... The initial article famstack.dev/guides/mlx-v... I am going to update the article with the insights from the community soon