DeepSeek V4 Flash Preview is the smaller of DeepSeek's two V4 preview models: a 284-billion-parameter Mixture-of-Experts text model with 13B active parameters per token, a one-million-token context window, three reasoning effort modes, and MIT-licensed open weights.
DeepSeek V4 Flash Preview is a 43-layer Mixture-of-Experts transformer with 256 routed experts and one shared expert per layer, with six routed experts activated for each token. Its hybrid attention interleaves Compressed Sparse Attention, which compresses the KV cache four-to-one before applying DeepSeek Sparse Attention, with Heavily Compressed Attention at a 128-to-one ratio, cutting the compute and memory cost of its one-million-token context. Manifold-Constrained Hyper-Connections strengthen the conventional residual connections, and the model was pre-trained on 32 trillion tokens with the Muon optimizer. It supports tool calling and three reasoning modes: Non-think, Think High, and Think Max.
This is the preview checkpoint released in April 2026, superseded in July 2026 by DeepSeek V4 Flash (the 0731 release), which keeps the same architecture and size but was re-post-trained for stronger agentic performance. Weights are released under the MIT license. Full details are in DeepSeek's announcement and the model card.
A Dedicated Endpoint is your own deployment of DeepSeek V4 Flash Preview, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model deepseek-ai/DeepSeek-V4-Flash