Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

Qwen3.6 27B

Qwen3.6-27B is Qwen's dense 27-billion-parameter vision-language model: text and image input, a 262K-token native context extensible to 1M, thinking and non-thinking modes, and agentic coding that Qwen reports beats its 397B Qwen3.5 flagship, released under Apache 2.0.

Model

Qwen3.6-27B is a 64-layer dense model built on the Qwen3.5 hybrid architecture: 16 blocks that each stack three Gated DeltaNet linear-attention layers ahead of one gated full-attention layer, paired with a vision encoder for image input. It handles 262,144 tokens natively and extends to 1,010,000 tokens with YaRN. The model thinks by default, can be switched to a direct non-thinking mode, and adds a preserve_thinking option that carries reasoning from earlier turns into agentic sessions. Qwen reports it surpasses the 397B-parameter Qwen3.5-397B-A17B on SWE-bench Verified, SWE-bench Pro and Terminal-Bench 2.0, and it includes a multi-token-prediction head for speculative decoding.

Qwen recommends keeping at least 128K tokens of context to preserve thinking and enabling YaRN only when inputs exceed the native window. Weights are released under the Apache 2.0 license. Full details are in Qwen's announcement and the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of Qwen3.6 27B, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model Qwen/Qwen3.6-27B

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute