Skip to content

Inference

Infra for Inference

Serve, scale and optimize inference globally, from low-latency LLM endpoints to async batch jobs.

Optimized inference from endpoint to engine.

Deploy fully managed endpoints or take advantage of low-level primitives to build, serve, and scale inference however your workload demands.

Access intelligence

Shared Endpoints gets you started with optimized OSS model serving built for speed and simplicity.

Get started
  • Token-based pricing

    Pay only for the tokens you use, with no infrastructure to manage or capacity to provision.

  • Production-ready

    Built to handle the autoscaling and high concurrency that production requires, in a single OpenAI-compatible API.

  • Performance optimized

    Deployed with frontier techniques to meet your throughput and latency SLOs.

Deploy optimized models

Deploy a dedicated endpoint tuned to your traffic, workload, and performance needs.

Get started
  • Deploy custom models

    One-click deploy leading open models or bring your own weights and serving configurations.

  • Fine-grained observability

    Get engine and container-level metrics for better debugging and more granular control.

  • Custom optimizations

    Mirror production traffic to a trial serving recipe, so you can compare quality, latency, throughput, and cost

Serve custom stacks

Generated abstract green flower illustration
import modal MODEL_NAME = "black-forest-labs/FLUX.1-schnell"image = (    modal.Image.from_registry("nvidia/cuda:12.4.0-devel-ubuntu22.04")    .pip_install("torch", "transformers", "diffusers", ...)) volume = modal.Volume.from_name("flux-lora-models") @app.cls(gpu="H100", image=image, volumes={"/loras": volume})class FluxWithLoRA:    @modal.enter()    def setup(self):        self.pipeline = FluxPipeline.from_pretrained(MODEL_NAME).to("cuda")        self.pipeline.load_and_fuse_lora()     @modal.method()    def generate_image(self, prompt: str):        return self.pipeline(prompt).images[0] flux = FluxWithLoRA()flux.generate_image.remote("")

Run your own custom code and model with the flexibility to configure everything your way.

Get started
  • Custom inference stack

    Deploy SGLang, vLLM, or your own server on the GPUs you choose. Modal handles the infrastructure.

  • Any modality

    Run frontier or custom models from LLMs and audio to world models and biology.

  • Full ownership

    Keep pre- and post-processing, pipelines, and agent tool execution in the same code as your model.

Engineered for speed at scale

1

Load models fast

A custom runtime with GPU snapshotting and nearby weight caches, so capacity comes online the moment traffic spikes.

Learn more
2

Elastic autoscaling on any GPU

Autoscaling workloads let you match compute to demand in real time, scaling up instantly and back down to zero when idle.

Learn more
3

Routing reimagined

Low-latency routing with configurable policies, keeping requests fast and performance consistent as traffic changes.

Learn more

Powerful primitives for production inference

Low-latency, real-time serving

Global GPU fleet runs close to your users, wherever they are. Support for inference optimizations like prefill disaggregation and prefix-aware routing.

3 containers

p50 51 ms

Batch and offline inference

Run large-scale batch processing, autoscale functions to thousands of parallel containers with zero additional configuration.

45,216 / 100,000 done

25 queued

Custom Models
“Modal served 1 trillion tokens for Jev within three days of launch. Their team was proactive and responsive as we scaled to meet unprecedented demand, and we could stay focused on shipping instead of managing infrastructure.”
Erik Gafni Co-founder and CTO, TypeSafe AI
Edge Inference
“We use Modal to run edge inference with <10ms overhead and batch jobs at large scale. Our team loves the platform for the power and flexibility it gives us.”
Brian Ichter Co-founder, Physical Intelligence
World Models
“Modal's infrastructure gave us the performance and reliability we need to ship this in every global region, at production scale.”
Kamil Sindi CTO, Runway

Ship your first app in minutes

$30 / month free compute