Tried new qwen3.5:4b-nvfp4 Ollama model on M1 Max here (in project where it's used with Koog AI agent)...38% faster than qwen3.5:4b (averaged over 5 runs of the agent). I asked Gemini and it initially claimed that I shouldn't see increase in performance like this...asked again and got following 😀
3 likes 1 replies
?