Qwen3-Embedding-8B is the flagship of Qwen's Qwen3 Embedding series: an 8-billion-parameter, 36-layer text embedding model built on Qwen3-8B-Base that topped the MTEB multilingual leaderboard at release, with 32K context, instruction-aware embeddings of up to 4096 dimensions, 100+ languages, and Apache 2.0 weights.
Qwen3-Embedding-8B is a 36-layer dense transformer fine-tuned with LoRA from Qwen3-8B-Base as a dual-encoder: it embeds one text segment and takes the hidden state of the final [EOS] token as the vector. Output embeddings are up to 4096 dimensions and, with MRL support, can be set to any user-defined size from 32 to 4096. The model is instruction-aware: a one-sentence task description prepended to queries improves most downstream tasks by 1 to 5%, per Qwen. Training combines contrastive pre-training on weakly supervised pairs synthesized by Qwen3, supervised fine-tuning, and model merging. It handles 32K-token inputs and 100+ natural and programming languages, and ranked No. 1 on the MTEB multilingual leaderboard at release (70.58; 75.22 on MTEB English v2).
It is the largest of a series that also includes 0.6B and 4B embedding models and matching 0.6B, 4B and 8B rerankers, all released as open weights under the Apache 2.0 license. Full details are in Qwen's announcement and the model card.
A Dedicated Endpoint is your own deployment of Qwen3 Embedding 8B, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model Qwen/Qwen3-Embedding-8B