Skip to content
AI.info

The Pulse

Ollama Adds DFlash and Image Input to Its Apple MLX Engine

Ollama’s August 10 announcement adds DFlash acceleration and image input for Muse Glimmer on Apple Silicon.

Ollama Adds DFlash and Image Input to Its Apple MLX Engine

AI.info Team ·

Ollama’s August 10, 2026, announcement added DFlash acceleration and image input support to its MLX engine on Apple Silicon, featuring Meta’s Muse Glimmer model. Ollama says DFlash makes Muse Glimmer run 1.5 to 1.8 times faster on Apple’s chips. The instructions specify the MLX-tagged model, muse-glimmer:30b-mlx; the post does not announce a new default for every model or request.

Muse Glimmer gets two MLX additions

Muse Glimmer is a 30-billion-parameter multimodal model released under the Apache 2.0 license, with a context length above 128,000 tokens, according to Ollama. The company presents it as a model for locally run agent workloads, including coding tools and personal assistants. It can be started from Ollama’s CLI or used with coding-agent applications including Claude Code, Codex and Pi, as well as personal-assistant frameworks such as OpenClaw and Hermes.

The MLX-specific option brings DFlash and image input. Muse Glimmer’s dedicated 1.8-billion-parameter perception encoder provides image understanding. Ollama lists tasks such as building a website or application from a drawing or mockup, using screenshots in computer-use applications, and reading documents, receipts and charts. These capabilities are described for Muse Glimmer through MLX; the post does not claim the image-input support or stated speed range applies to every model in Ollama’s catalog.

Ollama’s instructions show separate commands for the standard model and the MLX version: ollama run muse-glimmer and ollama run muse-glimmer:30b-mlx. The blog recommends the latter for performance on Apple Silicon. It also says Muse Glimmer supports controllable reasoning strength, with settings of low, medium, high and xhigh; the post recommends higher settings for complex coding and agentic tasks and lower settings when speed matters more.

The MLX rollout began in March

Ollama introduced its MLX-powered Apple Silicon experience as a preview on March 30, 2026. MLX is Apple’s machine-learning framework, and Ollama said the implementation uses Apple’s unified memory architecture. The preview highlighted coding agents and personal assistants as use cases. For its featured Qwen3.5 model, Ollama recommended a Mac with more than 32GB of unified memory.

On June 29, Ollama described a separate MLX performance improvement for Gemma 4: multi-token prediction, or MTP. In the Aider polyglot coding-agent benchmark, the company reported nearly 90% faster generation on average for Gemma 4 12B in NVFP4 on an M5 Max. That result applies to the specified model, quantization, hardware and benchmark; it is not a general speed claim for MLX or all Apple Silicon systems. Ollama described MTP as having the model draft several tokens and then verify them together, with the draft length adjusted while the model runs.

The latest listed release is a separate update

As of September 25, GitHub’s Ollama releases page listed v0.34.4 as the latest release; it was published September 23. Its notes include updates to llama.cpp, MLX and XGrammar, as well as a change to image resolution selection for Gemma 4 on Apple Silicon. Those release notes are separate from the August Muse Glimmer announcement and do not change the model-specific instructions in that post.

Ollama’s published materials document a sequence: the MLX preview in March, MTP performance work for Gemma 4 in June, and DFlash plus image input for Muse Glimmer in August. For the August feature set, the company’s specified command is ollama run muse-glimmer:30b-mlx.

Sources

Explore

More articles