DeepSeek V4.1 Flash is DeepSeek's multimodal Mixture-of-Experts model: 552 billion parameters with 8B active per token during prefill and 16B during decode, native image and text input, and a one-million-token context window built for input-heavy agentic workloads.
DeepSeek V4.1 Flash is the smallest model in DeepSeek's new architecture family and its first with native visual understanding. A 552-billion-parameter Mixture-of-Experts model with a Causal Encoder-Decoder design, it activates just 8B parameters per input token and 16B per output token, and its KV cache needs a quarter of the HBM of DeepSeek V4 Flash, improving cost efficiency. Larger-scale RL post-training puts it ahead of DeepSeek V4 Pro on agentic coding benchmarks such as Terminal-Bench and DeepSWE.
It supports tool calling, JSON mode, structured outputs and a controllable reasoning effort. Full details are in DeepSeek's announcement and on Hugging Face.
On a Shared Endpoint, you pay per token. The endpoint is OpenAI-compatible and already live: point your existing SDK at it and start sending requests.
A Dedicated Endpoint is your own deployment of DeepSeek V4.1 Flash, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model deepseek-ai/DeepSeek-V4.1-Flash