Skip to content
Runtime is here: Join TypeSafe AI, Cognition, DoorDash and more in SF. Last chance to register

Training

Infrastructure for training

Manage rollouts, evals, and clustered training from a single file.

Agent infrastructure that agents love.

Data pre-processing

Turn raw data into training data at any scale, with a single function.

Get started
  • Parallelize with a single call

    Fan out massive preprocessing jobs across containers with `.spawn_map()`

  • Match compute to your pipeline

    Run processing on GPUs or CPUs for disaggregated, scalable pipelines.

  • Built-in distributed storage

    Store raw and processed data in Volumes for reuse, or BYO cloud bucket.

Supervised Fine-Tuning

Bring your fine-tuning stack and scale it from a single GPU to a cluster.

Get started
  • Bring your training stack

    Run TRL, Unsloth, or custom PyTorch code in Modal containers.

  • Scale from one GPU to a cluster

    Configure multi-GPU nodes and RDMA-connected clusters for distributed training.

  • Write checkpoints

    Save checkpoints to Volumes to resume training, run evals, or serve the best one.

Reinforcement Learning

Rollouts that scale fast without slowing down your workflow.

Get started
  • Disaggregated by default

    Rollouts, inference, and training run as separate services with their own autoscalers.

  • Battle-tested with Open Source

    Run the open RL frameworks trusted by researchers.

  • Native sandbox environments

    Write to a Modal Volume for easy access for RL, evals, or serving.

Engineered for the speed and scale agents require

1

On-demand access to single and multi-node GPUs

Instant access to the compute shape you need, including RDMA-enabled GPUs, with a few lines of code.

Learn more
2

Managed GPU and cluster health

Spend more time experimenting and less time managing the state of the GPUs you’re working with.

Learn more
3

Battle-tested infra on Open Source frameworks

Tune everything, from post-training techniques to OSS frameworks and hyperparameters.

Learn more
Fine-Tuning
“By providing high-GPU-concurrency sweeps, persistent volumes, and observability tools, engineers could easily launch, monitor, and compare fine-tuning runs without worrying about cluster orchestration or storage plumbing.”
Cyrus Asgari Research Engineer, Decagon
Evals
“Modal Sandboxes let us recreate complex, production-like systems - from multi-container infrastructure workloads to GPU-enabled deep learning environments - across the millions of simulations we run each month.”
Rustem Feyzkhanov Sr. Engineering Manager — AI Platform, Snorkel AI
ML Pipelines
“Modal lets us deploy new ML models in hours rather than weeks. We use it across spam detection, recommendations, audio transcription, and video pipelines, and it’s helped us move faster with far less complexity.”
Mike Cohen Head of AI & ML Engineering, Substack

Ship your first app in minutes

$30 / month free compute