Sync a voice

rdwyer@fastmail.com · junk-mail

Qwen 3.8 27B is now available on Ollama

Sun Aug 16, 2026 · 07:57 AM EDT

From
Ollama <hello@ollama.com>
To
rdwyer@fastmail.com

Qwen 3.8 27B is now available to run with Ollama with substantial performance improvements across coding, professional work,
research, and long-horizon agentic tasks.

Download model [ollama.com/library/qwen3.8]

Qwen 3.8 27B is the most powerful Qwen model in its class:

* Built for agents: flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater
reliability.

* Multimodal: native multimodal dense model that includes support for image input.

* Long-context: 262K native context for long-running coding sessions.

* Widely accessible: available on all platforms supported by Ollama under the Apache 2.0 license.

Get started

Download Ollama [ollama.com/download] and run the model:

ollama run qwen3.8

To code with the model locally, run it using a coding agent such as Pi [docs.ollama.com/integrations/pi]:

ollama launch pi --model qwen3.8

See Ollama's documentation [docs.ollama.com/integrations] for more information on using Qwen 3.8 27B with coding agents
and personal assistants including Claude Code, Codex, Hermes, OpenClaw, and more.

Fast performance on Apple Silicon and NVIDIA hardware

We've collaborated with NVIDIA and Apple's MLX [opensource.apple.com/projects/mlx/] open-source project to optimize Qwen
3.8 27B's performance on NVIDIA and Apple Silicon hardware.

Apple Silicon

On M3 or newer devices, Ollama can leverage Apple's MLX framework for maximum performance, reaching over 70 output tokens/s when
running Qwen 3.8 27B on a MacBook with M5 Max chip:

ollama run qwen3.8:27b-mlx

Ollama also now supports image input when powered by MLX:

ollama run qwen3.8:27b-mlx "What's in this image? ./image.png"

NVIDIA

We've collaborated with NVIDIA [blogs.nvidia.com/blog/local-ai-open-so…nemotron/] for maximum
performance on Blackwell via llama.cpp, reaching over 131 output tokens/s when running Qwen 3.8 27B, optimized with multi-token
prediction (MTP).

If you have any feedback, reply directly to this email or join Ollama's Discord [discord.gg/ollama].

❤️ Ollama

You are receiving this email because you opted in to receive updates from Ollama
Ollama, 744 High Street, Palo Alto, CA 94301
Unsubscribe [app.loops.so/unsubscribe/cmsvr307be8js…5dc520d27]