Model Library
Explore open models you can run and customize on Modal.
Language & vision models
46 models
Z.ai
GLM 5.3
A natively multimodal model for coding, agents, and long-context workloads.
Language
753B MoE 1M context
Z.ai
GLM 5.3 Flash
A natively multimodal model for coding, agents, and long-context workloads.
Vision + language
320B MoE 1M context
Qwen
Qwen3.8-Max
The open-weights, text-only Max-class model for coding and long-horizon agentic work.
Language
2.4T MoE Up to 1M context
Moonshot AI
Kimi K3
Open frontier intelligence for deep reasoning and knowledge work.
Vision + language
2.8T MoE 1M context
Thinking Machines
Inkling NVFP4
An open-weights multimodal model with controllable reasoning across text, images, and audio.
Multimodal
975B MoE 1M context
DeepSeek
DeepSeek V4.1 Flash
A strong MoE model with a 1M-token context, native visual understanding, and an 3.9× smaller KV cache than V4 Flash.
Vision + language
552B MoE 1M context
OpenAI
GPT-OSS 120B
OpenAI's open-weight reasoning model with configurable reasoning effort and native tool use, sized to fit on a single GPU.
Language
117B MoE 128K context
IFM
K2 Horizon 0.9B
IFM's smallest K2 Horizon model: a 0.9B dense reasoning model with 128K context, selectable reasoning effort, tool calling, open training checkpoints, and Apache 2.0 weights.
Language
0.9B dense 128K context
DeepSeek
DeepSeek V4 Pro
DeepSeek's official V4 Pro release: a 1.6T MoE model with 49B active parameters, 1M-token context, built-in DSpark speculative decoding, and MIT weights.
Language
1.6T MoE · 49B active 1M context
Z.ai
GLM 5.3
Z.ai's GLM-5.3 flagship coding model in Modal's NVFP4 quantization: a 753B sparse-attention MoE with 40B active parameters, 1M context, and three reasoning effort levels.
Language
753B MoE · 40B active 1M context
Meta
Muse Glimmer 30B NVFP4
RadixArk's NVFP4/MXFP8 quantization of Meta's Muse Glimmer 30B agentic model: text-only, 128K context, controllable reasoning, and tool calling.
Language
30B dense 128K context
Meta
Muse Glimmer 30B BF16
Meta's open agentic model for local hardware: a dense 30B model distilled from Muse Spark, with 128K context, controllable reasoning, and Apache 2.0 weights.
Language
30B dense 128K context
NVIDIA
Nemotron 3.5 Lightning 30B A3B NVFP4
NVIDIA's fast open agent workhorse: a hybrid Mamba-2 MoE with 30B total and 3B active parameters, 1M-token context, built-in speculative decoding, and OpenMDW-1.1 weights.
Language
30B MoE · 3B active 1M context
Pipecat
PhoneLLM Alpha 1
Pipecat's open-weights LLM for voice agents: a 30B Mamba-Transformer MoE (3.5B active) fine-tuned from Nemotron 3 Nano for low-latency tool calling, 256K context, BSD 2-Clause.
Language
30B MoE · 3.5B active 256K context
Qwen
Qwen3.8 27B
Qwen's compact dense Qwen3.8 model: a 27B model with switchable thinking, 256K context, and Apache 2.0 weights.
Language
27B dense 256K context
Qwen
Qwen3.8 Flash Next FP8
Qwen's open preview of the Qwen4 architecture: a 125B MoE with 6B active, Qwen Sparse Attention, 51B n-gram embeddings, and 256K context, in FP8.
Language
125B MoE · 6B active 256K context
Qwen
Qwen3.8 2.4T A95B NVFP4
Qwen's open Qwen-Max-class flagship: a text-only 2.4T-parameter MoE with 95B active per token, always-on reasoning, and 256K context, quantized to NVFP4 by Modal.
Language
2.4T MoE · 95B active 256K context
DeepSeek
DeepSeek V4 Flash
DeepSeek's official V4 Flash release: a 284B MoE model with 13B active parameters, 1M-token context, built-in DSpark speculative decoding, and MIT weights.
Language
284B MoE · 13B active 1M context
Tencent
Hy3 GPTQ Int4
Tencent's Hy3 in AngelSlim's GPTQ Int4 quantization: a 299B Mixture-of-Experts model with 21B active parameters, 256K context, selectable reasoning effort, and Apache 2.0 weights.
Language
299B MoE · 21B active 256K context
Z.ai
GLM 5.2 NVFP4
Z.ai's GLM-5.2 flagship in NVIDIA's NVFP4 quantization: a 753B sparse-attention MoE with 40B active parameters, 1M context, and MIT weights for Blackwell GPUs.
Language
753B MoE · 40B active 1M context
Z.ai
GLM 5.2 FP8
Z.ai's official FP8 checkpoint of GLM-5.2: a 753B sparse-attention MoE with 40B active parameters, a solid 1M context, configurable effort levels, and MIT weights.
Language
753B MoE · 40B active 1M context
NVIDIA
NVIDIA Nemotron 3 Ultra 550B A55B NVFP4
NVIDIA's frontier-scale open reasoning model: a 550B latent MoE with 55B active parameters, up to 1M-token context, NVFP4 weights, and OpenMDW-1.1 licensing.
Language
550B MoE · 55B active 1M context
Moonshot AI
Kimi K2.6 NVFP4
Moonshot AI's Kimi K2.6 in NVIDIA's NVFP4 quantization: a 1T-parameter MoE agentic model with 32B active per token, 256K context, and thinking mode.
Language
1T MoE · 32B active 256K context
DeepSeek
DeepSeek V4 Pro Preview
DeepSeek's flagship V4 preview: a 1.6T Mixture-of-Experts model with 49B active parameters, hybrid compressed sparse attention, 1M-token context, and MIT weights.
Language
1.6T MoE · 49B active 1M context
DeepSeek
DeepSeek V4 Flash Preview
DeepSeek's efficient V4 preview: a 284B Mixture-of-Experts model with 13B active parameters, hybrid compressed sparse attention, 1M-token context, and MIT weights.
Language
284B MoE · 13B active 1M context
Gemma 4 31B IT
Google DeepMind's largest open Gemma 4 model: a dense 31B vision-language model with configurable thinking, 256K context, and Apache 2.0 weights.
Vision + language
31B dense 256K context
Gemma 4 26B A4B IT
Google DeepMind's Mixture-of-Experts Gemma 4 model: 25.2B total parameters with 3.8B active per token, text and image input, 256K context, and Apache 2.0 weights.
Vision + language
25.2B MoE · 3.8B active 256K context
Gemma 4 E4B IT
Google DeepMind's on-device Gemma 4 model: an 8B (4.5B effective) vision-language model with text and image input, 128K context, and Apache 2.0 weights.
Vision + language
8B dense 128K context
Gemma 4 E2B IT
Google DeepMind's smallest Gemma 4 model: a 5.1B (2.3B effective) on-device vision-language model with text and image input, 128K context, and Apache 2.0 weights.
Vision + language
5.1B dense 128K context
Qwen
Qwen3.6 35B A3B
Qwen's first open-weight Qwen3.6 model: a 35B MoE vision-language model with 3B active parameters, default-on thinking, 262K context, and Apache 2.0 weights.
Vision + language
35B MoE · 3B active 262K context
Qwen
Qwen3.6 35B A3B FP8
Qwen's FP8 build of its 35B MoE vision-language model: 3B active parameters, default-on thinking, 262K context, quantized block-wise to FP8 under Apache 2.0.
Vision + language
35B MoE · 3B active 262K context
Qwen
Qwen3.6 27B
Qwen's flagship-level dense coder: a 27B vision-language model with default-on thinking, 262K native context, and Apache 2.0 weights.
Vision + language
27B dense 262K context
Qwen
Qwen3.6 27B FP8
Qwen's FP8 build of its 27B dense coder: same vision-language model with default-on thinking and 262K context, quantized block-wise to FP8 under Apache 2.0.
Vision + language
27B dense 262K context
NVIDIA
NVIDIA Nemotron 3 Super 120B A12B NVFP4
NVIDIA's hybrid Mamba-Transformer MoE for agentic reasoning: 120B total and 12B active parameters, up to 1M-token context, native NVFP4 pretraining, and multi-token prediction.
Language
120B MoE · 12B active 1M context
Qwen
Qwen3.5 9B
Qwen's 9B Qwen3.5 model: a dense vision-language model with a hybrid Gated DeltaNet architecture, thinking on by default, 256K context extensible to 1M, and Apache 2.0 weights.
Vision + language
9B dense 256K context
Qwen
Qwen3.5 4B
Qwen's 4B Qwen3.5 model: a dense vision-language model with a hybrid Gated DeltaNet architecture, thinking on by default, 256K context extensible to 1M, and Apache 2.0 weights.
Vision + language
4B dense 256K context
Qwen
Qwen3.5 2B
Qwen's compact Qwen3.5 model: a dense 2B vision-language model with a hybrid Gated DeltaNet architecture, 256K context, and Apache 2.0 weights.
Vision + language
2B dense 256K context
Qwen
Qwen3.5 0.8B
Qwen's smallest Qwen3.5 model: a dense 0.8B vision-language model with a hybrid Gated DeltaNet architecture, 256K context, and Apache 2.0 weights.
Vision + language
0.8B dense 256K context
Qwen
Qwen3.5 397B A17B FP8
Qwen's flagship Qwen3.5 model in FP8: a 397B Mixture-of-Experts vision-language model with 17B active parameters, hybrid linear attention, 256K context, and Apache 2.0 weights.
Vision + language
397B MoE · 17B active 256K context
Qwen
Qwen3.5 122B A10B FP8
Qwen's mid-size Mixture-of-Experts Qwen3.5 model in FP8: 122B total with 10B active parameters, native text and image input, 256K context, and Apache 2.0 weights.
Vision + language
122B MoE · 10B active 256K context
Qwen
Qwen3.5 35B A3B FP8
Qwen's compact Mixture-of-Experts Qwen3.5 model in FP8: 35B total with 3B active parameters, native text and image input, 256K context, and Apache 2.0 weights.
Vision + language
35B MoE · 3B active 256K context
Qwen
Qwen3.5 27B FP8
Qwen's largest dense Qwen3.5 model in FP8: a 27B vision-language model with hybrid Gated DeltaNet attention, thinking on by default, 256K context, and Apache 2.0 weights.
Vision + language
27B dense 256K context
Z.ai
GLM 4.7
Z.ai's GLM-4.7 agentic coding model: a 358B Mixture-of-Experts with 32B active parameters, interleaved and preserved thinking, 200K context, and MIT weights.
Language
358B MoE · 32B active 200K context
Qwen
Qwen3 Embedding 8B
Qwen's largest Qwen3 Embedding model: an 8B dense text embedder with 32K context, up to 4096-dim MRL vectors, 100+ languages, and Apache 2.0 weights.
Language
8B dense 32K context
Qwen
Qwen3 Embedding 0.6B
Qwen's smallest Qwen3 Embedding model: a 0.6B dense text embedder with 32K context, up to 1024-dim MRL vectors, 100+ languages, and Apache 2.0 weights.
Language
0.6B dense 32K context
Gemma 3 1B IT
Google DeepMind's compact Gemma 3 model: a 1B-parameter text-only instruction-tuned LLM with 32K context, 140+ language pretraining, and open weights under the Gemma Terms of Use.
Language
1B dense 32K context