Qwen3.6-27B-FP8 is Qwen's official FP8 quantization of its dense 27B vision-language model: text and image input, 262K-token native context, thinking and non-thinking modes, and flagship-level agentic coding at quality Qwen reports as nearly identical to BF16, released under Apache 2.0.
Qwen3.6-27B is a 64-layer dense model built on the Qwen3.5 hybrid architecture: 16 blocks that each stack three Gated DeltaNet linear-attention layers ahead of one gated full-attention layer, paired with a vision encoder for image input. It handles 262,144 tokens natively and extends to 1,010,000 tokens with YaRN. The model thinks by default, can be switched to a direct non-thinking mode, and adds a preserve_thinking option that carries reasoning from earlier turns into agentic sessions. Qwen reports it surpasses the 397B-parameter Qwen3.5-397B-A17B on SWE-bench Verified, SWE-bench Pro and Terminal-Bench 2.0, and it includes a multi-token-prediction head for speculative decoding.
This is Qwen's own FP8 quantization of Qwen3.6-27B, fine-grained with a block size of 128 and dynamic activation scaling, which Qwen reports gives performance nearly identical to the original BF16 model. Architecture and context length are unchanged. Weights are released under the Apache 2.0 license. Full details are in Qwen's announcement and the model card.
A Dedicated Endpoint is your own deployment of Qwen3.6 27B FP8, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model Qwen/Qwen3.6-27B-FP8