DeepSeek V4 Flash (0731) is the official release of DeepSeek's efficient V4 model: 284 billion Mixture-of-Experts parameters with 13B active per token, a one-million-token context, low/high/max reasoning effort, a built-in DSpark speculative decoding module, and MIT-licensed weights.
DeepSeek V4 Flash keeps the architecture and size of the April preview: a 43-layer Mixture-of-Experts transformer with 256 routed and one shared expert per layer, 13B of its 284B parameters active per token, and a hybrid of Compressed Sparse Attention and Heavily Compressed Attention that shrinks the KV cache across its one-million-token context. The 0731 checkpoint was re-post-trained for agentic work, with DeepSeek reporting Terminal Bench 2.1 rising from 61.8 to 82.7 and DeepSWE from 7.3 to 54.4 over the preview, ahead of the V4 Pro preview. A DSpark speculative decoding module is attached, and reasoning effort is selectable as low, high, or max.
DeepSeek recommends temperature 1.0 with top_p 0.95 for agentic use and a 384K maximum output length for the high and max effort levels. Weights are released under the MIT license. Full details are in DeepSeek's announcement and the model card.
A Dedicated Endpoint is your own deployment of DeepSeek V4 Flash, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model deepseek-ai/DeepSeek-V4-Flash-0731