Qwen3.5 9B is the largest of the four small dense models in Qwen's natively multimodal Qwen3.5 family: a 9B-parameter vision-language model with thinking mode on by default, a 262K-token context extensible to 1M with YaRN, 201-language coverage, and open weights under Apache 2.0.
Qwen3.5 9B is a 32-layer causal language model with a vision encoder, trained with early text-vision fusion so it handles text and images natively. Its hybrid layout repeats eight blocks of three Gated DeltaNet linear-attention layers followed by one gated full-attention layer, for high-throughput, low-latency inference across a native 262,144-token context that can be extended to 1,010,000 tokens with YaRN scaling. The model uses a 4096-dimensional hidden state, a 248K-token vocabulary, and multi-token prediction training. It thinks by default, emitting reasoning inside think tags before answering, and can be switched to direct instruct-style responses. It supports tool calling and 201 languages and dialects.
This is the post-trained model, fine-tuned from Qwen3.5-9B-Base. Qwen benchmarks it against GPT-OSS-120B, GPT-OSS-20B, Qwen3-Next-80B-A3B-Thinking, and Qwen3-VL-30B-A3B on language and vision tasks, and recommends a 32,768-token output budget for most queries. Weights are released under the Apache 2.0 license. Full details are in Qwen's announcement and the model card.
A Dedicated Endpoint is your own deployment of Qwen3.5 9B, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model Qwen/Qwen3.5-9B