rdwyer@fastmail.com · junk-mail
Ollama 0.31: Faster Gemma 4 on Apple Silicon with multi-token prediction (MTP)
Wed Jul 01, 2026 · 05:05 AM EDT
Gemma 4 is now significantly faster in Ollama 0.31 with MLX. On Apple Silicon, it generates tokens nearly 90% faster on a coding
agent benchmark, with no change to the model's output.
The speedup comes from improved multi-token prediction (MTP), now on by default. Ollama auto-tunes how many tokens to draft as it
runs, so it never slows generation down when speculation stops helping.
Gemma 4 is the first model to receive this, with more model support to follow.
Get started
Download Ollama [ollama.com/download]
Once downloaded, configure claude to use Gemma 4 and start working:
ollama launch claude --model gemma4:12b-mlx
ollama launch also works with Codex, Droid, OpenCode, Copilot, and other integrations [docs.ollama.com/integrations]. If
you previously downloaded Gemma 4, re-pull it to get the latest version with MTP using ollama pull gemma4:12b-mlx.
You can also chat with the model directly:
ollama run gemma4:12b-mlx
❤️ Ollama
You are receiving this email because you opted in to receive updates from Ollama
Ollama, 744 High Street, Palo Alto, CA 94301
Unsubscribe [app.loops.so/unsubscribe/cmr1unuy2cs23…237d9a5f6]