Muse Glimmer 30B NVFP4 is RadixArk's (the SGLang team) mixed NVFP4/MXFP8 quantization of Meta's 30B-parameter open agentic model: text-only, 128K context, controllable reasoning, and tool calling, under Apache 2.0.
Muse Glimmer is Meta Superintelligence Labs' 52-layer dense causal transformer, distilled from Muse Spark and purpose-built for autonomous agents on consumer hardware. Attention alternates three sliding-window layers of 2,048 tokens with one global layer, using grouped-query attention with 32 query and 2 KV heads, a 202,048-token vocabulary, and a 131,072-token context. The model chains multi-step reasoning over long horizons, invokes tools with precise schemas throughout extended workflows, recovers from failed tool calls, supports low, medium, high, and xhigh reasoning strengths, and is trained on data from more than 100 languages.
This quantization was produced by RadixArk, the organization behind SGLang, with an fp_awq mixed-precision recipe: NVFP4 on most projections and MXFP8 on higher-sensitivity layers. It is text-only because the vision tower is not included. Weights are released under Apache 2.0. Full details are in Meta's announcement and the model card.
A Dedicated Endpoint is your own deployment of Muse Glimmer 30B NVFP4, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model RadixArk/Muse-Glimmer-NVFP4