Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

Qwen3.8 2.4T A95B NVFP4

Qwen3.8-2.4T-A95B is the first open-weight Qwen-Max-class model: a text-only 2.4-trillion-parameter Mixture-of-Experts with 95B active parameters per token, mandatory thinking mode, and a 262K native context extensible to 1M. This is Modal's NVFP4 quantization, released under the Qwen3.8-Max License.

Model

Qwen3.8-2.4T-A95B is the open-weight release of Qwen3.8-Max, the first time Qwen has opened a Max-class model. It is a 92-layer hybrid Mixture-of-Experts built on the Qwen3.5 architecture: each block of three Gated DeltaNet linear-attention layers is followed by one gated full-attention layer, and every layer routes tokens to 10 of 512 experts plus one shared expert, activating 95B of its 2.4T parameters. The model is text-only, always reasons before answering, and lets you tune depth with reasoning_effort (xhigh, medium, low) while preserve_thinking keeps prior reasoning across turns. Qwen positions it for coding, professional work, research, and long-horizon agentic tasks, with a 262,144-token native context extensible to about 1M tokens.

This is Modal's NVFP4 quantization of Qwen/Qwen3.8-2.4T-A95B. Routed experts are stored in NVFP4 with a group size of 16, while attention, linear attention, shared experts, router gates, embeddings, and the LM head keep their original precision, and the 262K native context is extended to 1M with RoPE scaling at serving time. The base weights are released under the Qwen3.8-Max License, an MIT-derived license with attribution requirements and a separate-licensing condition for large Model-as-a-Service and AI work-assistant businesses. Full details are in Qwen's announcement and the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of Qwen3.8 2.4T A95B NVFP4, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model modal-labs/Qwen3.8-2.4T-A95B-NVFP4

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute