DeepSeek V4 Pro (0813) is the GA release of DeepSeek's flagship V4 model: a 1.6-trillion-parameter Mixture-of-Experts text model with 49B active per token, a one-million-token context, low/high/max reasoning effort, a built-in DSpark speculative decoding module, and MIT-licensed weights.
DeepSeek V4 Pro is built on the V4 Pro preview's model structure: a 61-layer Mixture-of-Experts transformer with 384 routed and one shared expert per layer, 49B of its 1.6T parameters active per token, and hybrid Compressed Sparse and Heavily Compressed Attention that keeps the KV cache small across a one-million-token context. The 0813 checkpoint targets agentic workloads, with DeepSeek reporting Terminal Bench 2.1 rising from 72.1 to 87.9, DeepSWE from 12.8 to 62.7, and HLE with tools from 48.2 to 60.0 over the preview. A DSpark speculative decoding module is attached, and reasoning effort is selectable as low, high, or max.
DeepSeek recommends temperature 1.0 with top_p 0.95 for agentic use and a 384K maximum output length at the high and max effort levels. Weights are released under the MIT license. Full details are in DeepSeek's announcement and the model card.
A Dedicated Endpoint is your own deployment of DeepSeek V4 Pro, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model deepseek-ai/DeepSeek-V4-Pro-0813