John O'Reilly @johnoreilly.dev · Apr 2

Ollama 0.20.0-rc0 available and with it the new Google DeepMind Gemma 4 local model. Trying it here as part of a Koog AI agent (screenshot thanks to Koog's OpenTelemetry langfuse integration)

4 likes 2 replies

?

Replies

Max Slinger · Apr 2

gemma 4 running local on ollama 0.20 is kinda the quiet update that matters more than half the keynotes this quarter. how's inference speed been on your hardware

Pawel Jozefiak · Apr 6

Agent use is where Gemma 4 gets interesting fast. I'm using it as preprocessing layer - triage/classify before handing off to Claude. Cuts API costs 30-40%. The routing logic is the underrated part. thoughts.jock.pl/p/local-llm-...