Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

DeepSeek V4 Flash

DeepSeek V4 Flash (0731) is the official release of DeepSeek's efficient V4 model: 284 billion Mixture-of-Experts parameters with 13B active per token, a one-million-token context, low/high/max reasoning effort, a built-in DSpark speculative decoding module, and MIT-licensed weights.

Model

DeepSeek V4 Flash keeps the architecture and size of the April preview: a 43-layer Mixture-of-Experts transformer with 256 routed and one shared expert per layer, 13B of its 284B parameters active per token, and a hybrid of Compressed Sparse Attention and Heavily Compressed Attention that shrinks the KV cache across its one-million-token context. The 0731 checkpoint was re-post-trained for agentic work, with DeepSeek reporting Terminal Bench 2.1 rising from 61.8 to 82.7 and DeepSWE from 7.3 to 54.4 over the preview, ahead of the V4 Pro preview. A DSpark speculative decoding module is attached, and reasoning effort is selectable as low, high, or max.

DeepSeek recommends temperature 1.0 with top_p 0.95 for agentic use and a 384K maximum output length for the high and max effort levels. Weights are released under the MIT license. Full details are in DeepSeek's announcement and the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of DeepSeek V4 Flash, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model deepseek-ai/DeepSeek-V4-Flash-0731

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute