Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

Qwen3.6 35B A3B FP8

Qwen3.6-35B-A3B-FP8 is Qwen's official FP8 quantization of its 35B-total, 3B-active MoE multimodal model: text and image input, 262K-token native context, thinking and non-thinking modes, and strong agentic coding at quality Qwen reports as nearly identical to BF16, released under Apache 2.0.

Model

Qwen3.6-35B-A3B is a 40-layer Mixture-of-Experts model on the Qwen3.5 hybrid architecture: 10 blocks that each stack three Gated DeltaNet linear-attention layers ahead of one gated full-attention layer, with 256 experts per MoE layer and 8 routed plus 1 shared expert active per token, so only about 3B of its 35B parameters fire on each step. A vision encoder handles image input. Context is 262,144 tokens natively and extends to 1,010,000 with YaRN. It thinks by default, can be switched to a non-thinking mode, and supports preserve_thinking to carry reasoning across turns in agentic workflows, plus a multi-token-prediction head for speculative decoding.

This is Qwen's own FP8 quantization of Qwen3.6-35B-A3B, fine-grained with a block size of 128 and dynamic activation scaling, which Qwen reports gives performance nearly identical to the original BF16 model. Architecture and context length are unchanged. Weights are released under the Apache 2.0 license. Full details are in Qwen's announcement and the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of Qwen3.6 35B A3B FP8, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model Qwen/Qwen3.6-35B-A3B-FP8

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute