Training
Infrastructure for training
Manage rollouts, evals, and clustered training from a single file.
Agent infrastructure that agents love.
Data pre-processing
vol:/data
s3 . hf . volume put
raw/*.json
512 containers
vol:/data
train/*.bin
Turn raw data into training data at any scale, with a single function.
Get startedParallelize with a single call
Fan out massive preprocessing jobs across containers with `.spawn_map()`
Match compute to your pipeline
Run processing on GPUs or CPUs for disaggregated, scalable pipelines.
Built-in distributed storage
Store raw and processed data in Volumes for reuse, or BYO cloud bucket.
Supervised Fine-Tuning
vol:/data
s3 . hf . volume put
train/*.bin
Containers on 4 nodes
vol:/data
train/*.bin
Bring your fine-tuning stack and scale it from a single GPU to a cluster.
Get startedBring your training stack
Run TRL, Unsloth, or custom PyTorch code in Modal containers.
Scale from one GPU to a cluster
Configure multi-GPU nodes and RDMA-connected clusters for distributed training.
Write checkpoints
Save checkpoints to Volumes to resume training, run evals, or serve the best one.
Reinforcement Learning
TRAINER
8 containers . 1 node
Inference
6/14
Sandboxes
2,048
Rollouts that scale fast without slowing down your workflow.
Get startedDisaggregated by default
Rollouts, inference, and training run as separate services with their own autoscalers.
Battle-tested with Open Source
Run the open RL frameworks trusted by researchers.
Native sandbox environments
Write to a Modal Volume for easy access for RL, evals, or serving.
vol:/data
s3 . hf . volume put
raw/*.json
512 containers
vol:/data
train/*.bin
Turn raw data into training data at any scale, with a single function.
Get startedParallelize with a single call
Fan out massive preprocessing jobs across containers with `.spawn_map()`
Match compute to your pipeline
Run processing on GPUs or CPUs for disaggregated, scalable pipelines.
Built-in distributed storage
Store raw and processed data in Volumes for reuse, or BYO cloud bucket.
Engineered for the speed and scale agents require
On-demand access to single and multi-node GPUs
Instant access to the compute shape you need, including RDMA-enabled GPUs, with a few lines of code.
Managed GPU and cluster health
Spend more time experimenting and less time managing the state of the GPUs you’re working with.
Battle-tested infra on Open Source frameworks
Tune everything, from post-training techniques to OSS frameworks and hyperparameters.
“By providing high-GPU-concurrency sweeps, persistent volumes, and observability tools, engineers could easily launch, monitor, and compare fine-tuning runs without worrying about cluster orchestration or storage plumbing.”
“Modal Sandboxes let us recreate complex, production-like systems - from multi-container infrastructure workloads to GPU-enabled deep learning environments - across the millions of simulations we run each month.”
“Modal lets us deploy new ML models in hours rather than weeks. We use it across spam detection, recommendations, audio transcription, and video pipelines, and it’s helped us move faster with far less complexity.”
Latest updates

Modal Clusters are generally available
Multi-node GPU clusters with RDMA, gang scheduled from Modal's shared capacity pool and billed by the second, behind a single decorator.

Quail: Speeding up AI-SQL by jointly optimizing query planner and inference engine
Maximizing perf on AI-SQL queries with the KV-optimal left-deep join

How to serve trillions of tokens for trillion-parameter coding agents
Learn how we optimized performance and efficiency serving the workload that is changing software engineering forever — and you can too.

