Meta Releases Muse Glimmer 30B: Open Multimodal Model Optimized for Local Agentic Tasks
Meta's new 30B-parameter multimodal model, distilled from Muse and released under Apache 2.0, targets local deployment for privacy-focused coding, document analysis and…

- Muse Glimmer-30B outperforms Gemma4-31B and matches or exceeds Qwen3.6-27B on most agentic and coding benchmarks while showing competitive multimodal and reasoning scores.
- The model combines a 2B ViT-style perception encoder with a 28B text decoder using hybrid sliding-window and full attention plus gated grouped-query attention for efficiency.
- Day-0 integration in transformers, llama.cpp, vLLM and Inference Endpoints enables immediate local and cloud experimentation including multimodal tool calling and object detection.
Release and Intended Use
On August 10 2026 Meta introduced Muse Glimmer, a 30-billion-parameter multimodal model distilled from its larger Muse predecessor. The model is licensed under Apache 2.0 and published on the Hugging Face Hub under the meta-models organization.
According to the Hugging Face blog post the model was designed specifically for local agentic workloads. Target applications include privacy-sensitive coding assistants, document analysis tools and personal agents similar to Claw or Hermes setups. The release ships with immediate support across the open-source inference stack.
Benchmark Performance
The published benchmark table shows Muse Glimmer-30B achieving top scores among the three compared models on several agentic tasks. It records 75.5 on MCP Atlas, 74.6 on DeepSearch QA, 47.6 on WildClawBench and 65.9 on OSWorld-Verified.
On coding benchmarks it reaches 51.2 on SWE-Bench Pro and 76.0 on SWE-Bench Verified. Multimodal results include 78.8 on Charxiv Reasoning and 75.4 on ScreenSpot Pro. General reasoning scores are 77.0 on IFBench, 94.7 on AIME 2026 and 83.5 on GPQA Diamond.
Safety evaluations report a CI Memories violation rate of 26.4 with 64.8 coverage and an AgentDojo attack success rate of 28.4 with 94.2 utility. These numbers are presented as published by Meta and should be interpreted in the context of the specific test setups.
Architecture Details
Muse Glimmer consists of a 2-billion-parameter ViT-style perception encoder for images and video plus a 28-billion-parameter text decoder. The vision encoder processes 2-frame patches and applies 2D rotary embeddings with a repeating window-plus-full attention pattern.
The language model uses hybrid attention alternating three sliding-window layers of 2048 tokens with rotary embeddings and one full-attention layer without positional embeddings. It employs gated grouped-query attention reducing KV cache by 16x and applies query-key normalization with additional query scaling.
An optional DFlash speculative decoding drafter is included. It uses block diffusion to propose multiple future tokens and is reported to accelerate structured generation such as code while adding modest memory overhead.
Inference and Tooling Support
The model loads via AutoModelForMultimodalLM and AutoProcessor in the latest transformers library and runs unchanged on CUDA, ROCm and XPU hardware. Examples demonstrate text-only chat, image understanding, video question answering without audio, multimodal tool calling and open-ended object detection.
llama.cpp support arrives with pre-calibrated GGUF quants and DFlash integration. vLLM support uses the transformers backend for tensor-parallel serving. Managed deployment is available through Hugging Face Inference Endpoints exposing an OpenAI-compatible API.
Fine-tuning guidance using TRL covers LoRA supervised fine-tuning and GRPO on datasets such as MolmoWeb and OpenCode. Memory requirements range from a single 80 GB H100 for inference or LoRA to multiple GPUs for full fine-tuning.
Agentic Demos and Self-Deployment
The blog presents agent demonstrations in which Muse Glimmer quantizes itself to GGUF, launches a local llama-server, or deploys itself to a protected Inference Endpoint. These workflows rely on additional tooling such as OpenClaw, the Hugging Face MCP and specific prompts added to AGENTS.md.
Such capabilities illustrate the model's intended use as a local-scale personal assistant that can inspect hardware, interact with the Hub and manage its own deployment while preserving privacy or reducing costs.


