Qwen3.8-2.4T-A95B now available on Modal
Qwen3.8-2.4T-A95B just launched as an open weights model, and it’s now available on Modal.
Over Qwen 3.7, the latest model sees substantial improvement across coding, work, research, and long-horizon tasks.
We worked with Qwen ahead of the drop to bring day zero support to Modal Auto Endpoints, backed by SGLang and a custom DFlash speculator tuned to Qwen3.8’s shape.
Try it out now as a Shared Endpoint.
Speeding up inference with custom DFlash speculation
Just getting the model running is one thing, making it fast is another. For this, we once again turn to a DFlash speculator model because—say it with us now—Speculation is all you need.
A speculator only earns its keep when the target accepts the tokens it drafts, and acceptance comes down to whether the drafter has seen sequences like the ones it's predicting. Because of Max’s improvements in coding, research, and work (things that tend to use more tool calls) we leaned into that in our training data to increase accepts.
Try it now
Qwen3.8-2.4T-A95B (text only) is available for the next month as an OpenAI compatible Shared Endpoint with token-based pricing.