Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

Qwen3.5 4B

Qwen3.5 4B is a dense 4B-parameter vision-language model in Qwen's natively multimodal Qwen3.5 family, with thinking mode on by default, a 262K-token context extensible to 1M with YaRN, 201-language coverage, and open weights under Apache 2.0.

Model

Qwen3.5 4B is a 32-layer causal language model with a vision encoder, trained with early text-vision fusion so it handles text and images natively. Its hybrid layout repeats eight blocks of three Gated DeltaNet linear-attention layers followed by one gated full-attention layer, for high-throughput, low-latency inference across a native 262,144-token context that can be extended to 1,010,000 tokens with YaRN scaling. The model uses a 2560-dimensional hidden state, a 248K-token vocabulary, and multi-token prediction training. It thinks by default, emitting reasoning inside think tags before answering, and can be switched to direct instruct-style responses. It supports tool calling and 201 languages and dialects.

This is the post-trained model, fine-tuned from Qwen3.5-4B-Base. Qwen benchmarks it alongside Qwen3.5 9B against GPT-OSS-20B, Qwen3-30B-A3B-Thinking-2507, and Qwen3-VL-30B-A3B on language and vision tasks, and recommends a 32,768-token output budget for most queries. Weights are released under the Apache 2.0 license. Full details are in Qwen's announcement and the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of Qwen3.5 4B, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model Qwen/Qwen3.5-4B

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute