Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

Qwen3.5 0.8B

Qwen3.5 0.8B is the smallest member of Qwen's natively multimodal Qwen3.5 family: a dense 0.8B-parameter vision-language model with a 262K-token context window, 201-language coverage, optional thinking mode, and open weights under Apache 2.0.

Model

Qwen3.5 0.8B is a 24-layer causal language model with a vision encoder, trained with early text-vision fusion so it handles text and images natively. Its hybrid layout repeats six blocks of three Gated DeltaNet linear-attention layers followed by one gated full-attention layer, for high-throughput, low-latency inference across its native 262,144-token context. The model uses a 1024-dimensional hidden state, a 248K-token vocabulary, tied input and output embeddings, and multi-token prediction training. It supports tool calling and 201 languages and dialects, and runs in non-thinking mode by default with thinking switchable through the chat template.

This is the post-trained model, fine-tuned from Qwen3.5-0.8B-Base. Qwen positions this size for prototyping, task-specific fine-tuning, and research or development use, and notes that in thinking mode it is more prone to thinking loops than other Qwen3.5 models, recommending tuned sampling parameters and streaming so runaway generations can be cut off. Weights are released under the Apache 2.0 license. Full details are in Qwen's announcement and the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of Qwen3.5 0.8B, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model Qwen/Qwen3.5-0.8B

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute