Qwen3.8-Flash-Next is Qwen's open-weight preview of the Qwen4 architecture: a Mixture-of-Experts with a 125B main model, 6B active parameters per token, 51B of n-gram embeddings, and Qwen Sparse Attention for cheap long context. This FP8 variant is quantized by Qwen under the Qwen Community License 1.0.
Qwen3.8-Flash-Next is an experimental preview of the architecture that will underpin Qwen4. Its 48 layers pair Gated DeltaNet with Qwen Sparse Attention (QSA), which indexes context at micro-block granularity to cut long-sequence attention cost, and every layer routes to 10 of 512 experts plus one shared expert. Gated Residual widens the residual stream into four branches with data-dependent read and write gates, and a 20-million-entry n-gram embedding table adds 51B parameters of capacity with very little extra compute. The result is a 125B-parameter main model with 6B active per token, default-on thinking with tunable reasoning_effort, and a 262,144-token native context extensible to 1M with YaRN.
This is Qwen's official FP8 quantization of Qwen3.8-Flash-Next, fine-grained with a block size of 128; Qwen reports performance nearly identical to the bf16 original. The weights are released under the Qwen Community License 1.0, an MIT-derived license with attribution requirements that requires a separate license for Model-as-a-Service or AI work-assistant businesses. Full details are in Qwen's announcement and the model card.
A Dedicated Endpoint is your own deployment of Qwen3.8 Flash Next FP8, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model Qwen/Qwen3.8-Flash-Next-FP8