Qwen3.5-27B is the largest dense model in Qwen's natively multimodal Qwen3.5 series: 27B parameters, text and image input, a 256K-token context extensible to 1M, thinking on by default, and tool calling. This is Qwen's official FP8 checkpoint, released under Apache 2.0.
Qwen3.5-27B is a 64-layer dense transformer that replaces most standard attention with linear attention: each block of three Gated DeltaNet layers is followed by one gated full-attention layer with 24 query heads and 4 key-value heads, and every layer uses a 17,408-wide feed-forward network over a 248,320-token vocabulary. Trained with early vision-language fusion, it accepts images alongside text, thinks by default before answering (thinking can be disabled per request), and is tuned for tool calling and agentic use. It supports 201 languages and dialects, a 262,144-token native context extensible to 1,010,000 tokens with YaRN, and includes a multi-token-prediction head for speculative decoding.
This is Qwen's official FP8 quantization of the post-trained model, fine-grained with a block size of 128, which Qwen reports as nearly identical in quality to the original weights. Qwen3.5-27B was released on February 24, 2026, alongside the 122B-A10B and 35B-A3B models, about a week after the flagship 397B-A17B, and its weights are released under the Apache 2.0 license. Full details are in Qwen's announcement and the model card.
A Dedicated Endpoint is your own deployment of Qwen3.5 27B FP8, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model Qwen/Qwen3.5-27B-FP8