Muse Glimmer 30B is Meta Superintelligence Labs' open-weight agentic model: a 30B-parameter dense transformer with a 128K-token context, controllable reasoning strength, and reliable tool calling, released under Apache 2.0.
Muse Glimmer is a 52-layer dense causal transformer distilled from Muse Spark and purpose-built for autonomous agents on consumer hardware. Attention alternates three sliding-window layers of 2,048 tokens with one global layer, using grouped-query attention with 32 query and 2 KV heads and RoPE on local layers only. The model chains multi-step reasoning over long horizons, invokes tools with precise schemas, recovers from failed tool calls, supports low, medium, high, and xhigh reasoning strengths, and is trained on data from more than 100 languages with a 131,072-token context.
The model was pre-trained by logit distillation from Muse Spark, mid-trained on long-context agentic data, and post-trained with supervised fine-tuning, on-policy distillation, and reinforcement learning. All released artifacts are licensed under Apache 2.0. Full details are in Meta's announcement and the model card.
A Dedicated Endpoint is your own deployment of Muse Glimmer 30B BF16, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model meta-models/Muse-Glimmer-30B