John O'Reilly @johnoreilly.dev · Mar 31

Tried new qwen3.5:4b-nvfp4 Ollama model on M1 Max here (in project where it's used with Koog AI agent)...38% faster than qwen3.5:4b (averaged over 5 runs of the agent). I asked Gemini and it initially claimed that I shouldn't see increase in performance like this...asked again and got following 😀

3 likes 1 replies

?

Replies

Max Slinger · Mar 31

38% from a quant format swap on the same hardware is kinda nuts. curious what the accuracy tradeoff looks like on your specific agent tasks though