Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

Qwen3.8 27B

Qwen3.8-27B brings the Qwen3.8 generation to a deployment-friendly dense model: 27B parameters, thinking on by default with tunable reasoning_effort, a 262K native context extensible to 1M with YaRN, and Apache 2.0 weights.

Model

Qwen3.8-27B is the dense member of the Qwen3.8 family, built on the hybrid architecture introduced with Qwen3.5. Its 64 layers are arranged as 16 blocks of three Gated DeltaNet linear-attention layers followed by one gated full-attention layer, with a 5,120-dimensional hidden state and a 17,408-dimensional feed-forward network, and it is trained with a multi-token-prediction head. Thinking is on by default and can be disabled per request, reasoning depth is tunable with reasoning_effort (xhigh, medium, low), and preserve_thinking keeps prior reasoning across turns. Qwen targets coding, professional work, research, and long-horizon agentic tasks.

The native 262,144-token context can be extended to 1,000,000 tokens with YaRN RoPE scaling, and Qwen recommends generous output budgets for agentic work. The weights are released under the Apache 2.0 license. Full details are in the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of Qwen3.8 27B, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model Qwen/Qwen3.8-27B

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute