Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

Qwen3 Embedding 0.6B

Qwen3-Embedding-0.6B is the lightest model in Qwen's Qwen3 Embedding series: a 0.6-billion-parameter, 28-layer text embedding model built on Qwen3-0.6B-Base, with a 32K-token context, instruction-aware embeddings of up to 1024 dimensions, 100+ languages, and Apache 2.0 weights.

Model

Qwen3-Embedding-0.6B is a 28-layer dense transformer fine-tuned with LoRA from Qwen3-0.6B-Base as a dual-encoder: it embeds a single text segment and takes the hidden state of the final [EOS] token as the vector. Output embeddings are up to 1024 dimensions and, with MRL support, can be set to any user-defined size from 32 to 1024. The model is instruction-aware: a one-sentence task description prepended to queries improves most downstream tasks by 1 to 5%, per Qwen. Training combines contrastive pre-training on weakly supervised pairs synthesized by Qwen3, supervised fine-tuning, and model merging. It handles 32K-token inputs and 100+ natural and programming languages, scoring 64.33 on MTEB (Multilingual) and 70.70 on MTEB (English v2).

It is the entry point to a series that also includes 4B and 8B embedding models and matching 0.6B, 4B and 8B rerankers, all released as open weights under the Apache 2.0 license. Full details are in Qwen's announcement and the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of Qwen3 Embedding 0.6B, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model Qwen/Qwen3-Embedding-0.6B

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute