Qwen3.8-Max (text only; Qwen/Qwen3.8-2.4T-A95B) is the Max-class flagship of the Qwen3.8 family: a 2.4-trillion-parameter Mixture-of-Experts LLM with 95B parameters active per token and a 1M-token context window.
Qwen3.8-Max (text only) is built for coding and cowork: it was the first Max-class Qwen model released with open weights, on the architectural foundation of Qwen3.5. It runs long-horizon agentic work — multi-day autonomous coding runs, end-to-end research and task workflows — with adjustable reasoning effort (low, medium, xhigh). Full details are in Qwen's announcement.
On a Shared Endpoint, you pay per token. The endpoint is OpenAI-compatible and already live: point your existing SDK at it and start sending requests.