Sync a voice

rdwyer@fastmail.com · junk-mail

Ollama 0.31: Faster Gemma 4 on Apple Silicon with multi-token prediction (MTP)

Wed Jul 01, 2026 · 05:05 AM EDT

From
Ollama <hello@ollama.com>
To
rdwyer@fastmail.com

Gemma 4 is now significantly faster in Ollama 0.31 with MLX. On Apple Silicon, it generates tokens nearly 90% faster on a coding
agent benchmark, with no change to the model's output.

The speedup comes from improved multi-token prediction (MTP), now on by default. Ollama auto-tunes how many tokens to draft as it
runs, so it never slows generation down when speculation stops helping.

Gemma 4 is the first model to receive this, with more model support to follow.

Get started

Download Ollama [ollama.com/download]

Once downloaded, configure claude to use Gemma 4 and start working:

ollama launch claude --model gemma4:12b-mlx

ollama launch also works with Codex, Droid, OpenCode, Copilot, and other integrations [docs.ollama.com/integrations]. If
you previously downloaded Gemma 4, re-pull it to get the latest version with MTP using ollama pull gemma4:12b-mlx.

You can also chat with the model directly:

ollama run gemma4:12b-mlx

❤️ Ollama

You are receiving this email because you opted in to receive updates from Ollama
Ollama, 744 High Street, Palo Alto, CA 94301
Unsubscribe [app.loops.so/unsubscribe/cmr1unuy2cs23…237d9a5f6]