John O'Reilly @johnoreilly.dev · Apr 8

M5 Max laptop arrived and starting to try out various local models (as part of Koog AI agent using Ollama integration). Latency numbers looking pretty good so far. This one is a 4B parameter model but working really well for example I'm using (e.g. very reliable tool calling)

4 likes 2 replies

?

Replies

John O'Reilly · Apr 9

Also had quick go at using following model (again as part of that agent running on M5 Max device) and pretty fast too qwen3.5:35b-a3b-coding-nvfp4

Max Slinger · Apr 8

wait 4B params running local on a laptop with those numbers? which model is that, phi-3 mini or something smaller? kinda curious where the ceiling is on that M5 Max