RELEASE

Ollama v0.32.6

What's Changed Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model's MTP head for speculative decoding automatically /v1/chat/completions streaming now matches OpenAI's wire format: role only on the first chunk, finishreason on its own…

Source ollama/ollamaPublished 10d ago · Aug 5, 2026Posted on Bluesky

What we hold

Repo
ollama/ollama
Version
v0.32.6

Getting releases like this one by email is a separate problem from reading them here. GitHub’s custom watch, a repository’s releases.atom feed, Dependabot and NewReleases.io, compared with sources: how to get an email when a GitHub repo publishes a release.