John O'Reilly @johnoreilly.dev · Apr 9

Still trying local LLMs on M5 Max device here. Was comparing following and was interesting to see explanation from Gemini below about why Qwen model is faster (and certainly seeing that's the case here in testing). - gemma4:26b-a4b-it-q4_K_M - qwen3.5:35b-a3b-coding-nvfp4

1 likes 0 replies

?