Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

Qwen3 Embedding 8B

Qwen3-Embedding-8B is the flagship of Qwen's Qwen3 Embedding series: an 8-billion-parameter, 36-layer text embedding model built on Qwen3-8B-Base that topped the MTEB multilingual leaderboard at release, with 32K context, instruction-aware embeddings of up to 4096 dimensions, 100+ languages, and Apache 2.0 weights.

Model

Qwen3-Embedding-8B is a 36-layer dense transformer fine-tuned with LoRA from Qwen3-8B-Base as a dual-encoder: it embeds one text segment and takes the hidden state of the final [EOS] token as the vector. Output embeddings are up to 4096 dimensions and, with MRL support, can be set to any user-defined size from 32 to 4096. The model is instruction-aware: a one-sentence task description prepended to queries improves most downstream tasks by 1 to 5%, per Qwen. Training combines contrastive pre-training on weakly supervised pairs synthesized by Qwen3, supervised fine-tuning, and model merging. It handles 32K-token inputs and 100+ natural and programming languages, and ranked No. 1 on the MTEB multilingual leaderboard at release (70.58; 75.22 on MTEB English v2).

It is the largest of a series that also includes 0.6B and 4B embedding models and matching 0.6B, 4B and 8B rerankers, all released as open weights under the Apache 2.0 license. Full details are in Qwen's announcement and the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of Qwen3 Embedding 8B, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model Qwen/Qwen3-Embedding-8B

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute