Qwen3.5 2B is a compact member of Qwen's natively multimodal Qwen3.5 family: a dense 2B-parameter vision-language model with a 262K-token context window, 201-language coverage, optional thinking mode, and open weights under Apache 2.0.
Qwen3.5 2B is a 24-layer causal language model with a vision encoder, trained with early text-vision fusion so it handles text and images natively. Its hybrid layout repeats six blocks of three Gated DeltaNet linear-attention layers followed by one gated full-attention layer, for high-throughput, low-latency inference across its native 262,144-token context. The model uses a 2048-dimensional hidden state, a 248K-token vocabulary, tied input and output embeddings, and multi-token prediction training. It supports tool calling and 201 languages and dialects, and runs in non-thinking mode by default with thinking switchable through the chat template.
This is the post-trained model, fine-tuned from Qwen3.5-2B-Base. Qwen positions this size for prototyping, task-specific fine-tuning, and research or development use, and benchmarks it against Qwen3-1.7B, Qwen3-4B-2507, and the Qwen3-VL 2B and 4B models on language and vision tasks, recommending a 32,768-token output budget for most queries. Weights are released under the Apache 2.0 license. Full details are in Qwen's announcement and the model card.
A Dedicated Endpoint is your own deployment of Qwen3.5 2B, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model Qwen/Qwen3.5-2B