Qwen3-Embedding-0.6B is the lightest model in Qwen's Qwen3 Embedding series: a 0.6-billion-parameter, 28-layer text embedding model built on Qwen3-0.6B-Base, with a 32K-token context, instruction-aware embeddings of up to 1024 dimensions, 100+ languages, and Apache 2.0 weights.
Qwen3-Embedding-0.6B is a 28-layer dense transformer fine-tuned with LoRA from Qwen3-0.6B-Base as a dual-encoder: it embeds a single text segment and takes the hidden state of the final [EOS] token as the vector. Output embeddings are up to 1024 dimensions and, with MRL support, can be set to any user-defined size from 32 to 1024. The model is instruction-aware: a one-sentence task description prepended to queries improves most downstream tasks by 1 to 5%, per Qwen. Training combines contrastive pre-training on weakly supervised pairs synthesized by Qwen3, supervised fine-tuning, and model merging. It handles 32K-token inputs and 100+ natural and programming languages, scoring 64.33 on MTEB (Multilingual) and 70.70 on MTEB (English v2).
It is the entry point to a series that also includes 4B and 8B embedding models and matching 0.6B, 4B and 8B rerankers, all released as open weights under the Apache 2.0 license. Full details are in Qwen's announcement and the model card.
A Dedicated Endpoint is your own deployment of Qwen3 Embedding 0.6B, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model Qwen/Qwen3-Embedding-0.6B