Low-latency, real-time serving
Global GPU fleet runs close to your users, wherever they are. Support for inference optimizations like prefill disaggregation and prefix-aware routing.
3 containers
p50 51 ms
Inference
Serve, scale and optimize inference globally, from low-latency LLM endpoints to async batch jobs.
Deploy fully managed endpoints or take advantage of low-level primitives to build, serve, and scale inference however your workload demands.
Deploy optimized models
Deploy a dedicated endpoint tuned to your traffic, workload, and performance needs.
Get startedOne-click deploy leading open models or bring your own weights and serving configurations.
Get engine and container-level metrics for better debugging and more granular control.
Mirror production traffic to a trial serving recipe, so you can compare quality, latency, throughput, and cost
Serve custom stacks
import modal MODEL_NAME = "black-forest-labs/FLUX.1-schnell"image = ( modal.Image.from_registry("nvidia/cuda:12.4.0-devel-ubuntu22.04") .pip_install("torch", "transformers", "diffusers", ...)) volume = modal.Volume.from_name("flux-lora-models") @app.cls(gpu="H100", image=image, volumes={"/loras": volume})class FluxWithLoRA: @modal.enter() def setup(self): self.pipeline = FluxPipeline.from_pretrained(MODEL_NAME).to("cuda") self.pipeline.load_and_fuse_lora() @modal.method() def generate_image(self, prompt: str): return self.pipeline(prompt).images[0] flux = FluxWithLoRA()flux.generate_image.remote("") Run your own custom code and model with the flexibility to configure everything your way.
Get startedDeploy SGLang, vLLM, or your own server on the GPUs you choose. Modal handles the infrastructure.
Run frontier or custom models from LLMs and audio to world models and biology.
Keep pre- and post-processing, pipelines, and agent tool execution in the same code as your model.
curl https://api.modal.com/v1/chat/completions \-H "Authorization: Bearer $MODAL_KEY" \-H "Content-Type: application/json" \-d '{ "model": "auto", "messages": [...]}'SHARED ENDPOINT
Live
AVAILABLE MODELS
Shared Endpoints gets you started with optimized OSS model serving built for speed and simplicity.
Get startedPay only for the tokens you use, with no infrastructure to manage or capacity to provision.
Built to handle the autoscaling and high concurrency that production requires, in a single OpenAI-compatible API.
Deployed with frontier techniques to meet your throughput and latency SLOs.
A custom runtime with GPU snapshotting and nearby weight caches, so capacity comes online the moment traffic spikes.
Autoscaling workloads let you match compute to demand in real time, scaling up instantly and back down to zero when idle.
Low-latency routing with configurable policies, keeping requests fast and performance consistent as traffic changes.
Global GPU fleet runs close to your users, wherever they are. Support for inference optimizations like prefill disaggregation and prefix-aware routing.
3 containers
p50 51 ms
Run large-scale batch processing, autoscale functions to thousands of parallel containers with zero additional configuration.
45,216 / 100,000 done
25 queued
“Modal served 1 trillion tokens for Jev within three days of launch. Their team was proactive and responsive as we scaled to meet unprecedented demand, and we could stay focused on shipping instead of managing infrastructure.”
“We use Modal to run edge inference with <10ms overhead and batch jobs at large scale. Our team loves the platform for the power and flexibility it gives us.”
“Modal's infrastructure gave us the performance and reliability we need to ship this in every global region, at production scale.”

Multi-node GPU clusters with RDMA, gang scheduled from Modal's shared capacity pool and billed by the second, behind a single decorator.

VM Sandboxes are built for those who need to give their agents the power of a full computer.

Everything announced live at Modal's inaugural conference, Runtime.
$30 / month free compute