DeepSeek V4 Pro Preview is the flagship of DeepSeek's V4 preview: a 1.6-trillion-parameter Mixture-of-Experts text model with 49B active parameters per token, a one-million-token context window, three reasoning effort modes, and MIT-licensed open weights.
DeepSeek V4 Pro Preview is a 61-layer Mixture-of-Experts transformer with 384 routed experts and one shared expert per layer, with six routed experts activated per token, for 49B active out of 1.6T total parameters. Its hybrid attention interleaves Compressed Sparse Attention, which compresses the KV cache four-to-one before applying DeepSeek Sparse Attention, with Heavily Compressed Attention at a 128-to-one ratio; at a one-million-token context it needs only 27 percent of the single-token inference FLOPs and 10 percent of the KV cache of DeepSeek V3.2. Pre-trained on more than 32 trillion tokens with the Muon optimizer, it supports tool calling and Non-think, Think High, and Think Max reasoning modes.
This is the preview checkpoint released in April 2026 and superseded in August 2026 by DeepSeek V4 Pro (the 0813 release), which builds on the same model structure with a DSpark speculative decoding module and stronger agentic performance. Weights are released under the MIT license. Full details are in DeepSeek's announcement and the model card.
A Dedicated Endpoint is your own deployment of DeepSeek V4 Pro Preview, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model deepseek-ai/DeepSeek-V4-Pro