Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

Qwen3.5 397B A17B FP8

Qwen3.5-397B-A17B is the first and largest model in Qwen's natively multimodal Qwen3.5 series: a Mixture-of-Experts model with 397B total and 17B active parameters, text and image input, a 256K-token context extensible to 1M, and thinking on by default. This is Qwen's official FP8 checkpoint, released under Apache 2.0.

Model

Qwen3.5-397B-A17B is a 60-layer hybrid model built on the Qwen3-Next architecture: each block of three Gated DeltaNet linear-attention layers is followed by one gated full-attention layer with 32 query heads and 2 key-value heads, and every layer routes through a 512-expert Mixture-of-Experts that activates 10 routed experts plus one shared expert, so only 17B of its 397B parameters are used per token. Trained with early vision-language fusion, it accepts images alongside text, thinks by default before answering (thinking can be disabled per request), and is tuned for tool calling and agentic use. It supports 201 languages and dialects, a 262,144-token native context extensible to 1,010,000 tokens with YaRN, and includes a multi-token-prediction head for speculative decoding.

This is Qwen's official FP8 quantization of the post-trained model, fine-grained with a block size of 128, which Qwen reports as nearly identical in quality to the original weights. It is the open-weight counterpart of the hosted Qwen3.5-Plus, which adds a 1M-token default context, built-in tools and adaptive tool use. Released in February 2026 as the first Qwen3.5 model, the weights are under the Apache 2.0 license. Full details are in Qwen's announcement and the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of Qwen3.5 397B A17B FP8, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model Qwen/Qwen3.5-397B-A17B-FP8

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute