AI Infrastructure
The Model Context Protocol (MCP) introduced Streamable HTTP transport in specification 2025-03-26, replacing the legacy HTTP+SSE transport and changing how AI agents communicate with external tools and services. The current specification, 2026-07-28, goes further: it removes the protocol-level initialize/initialized handshake, retires the Mcp-Session-Id header, and makes the protocol core stateless, so requests can land on different server instances without a shared MCP session store or sticky routing. That shift makes MCP servers a natural fit for horizontally scaled and serverless infrastructure, and it moves the hard architectural questions from "how do I persist a transport session?" to "how do I run compute efficiently and, when my application genuinely needs state, where do I put it?" For teams building AI-powered applications, selecting the right AI infrastructure platform for hosting MCP servers directly impacts latency, scalability, and operational costs. This guide examines seven platforms that serve different MCP server hosting needs in 2026, starting with Modal, an AI-native infrastructure platform engineered for fast cold starts and CPU and GPU MCP workloads, and one that now officially documents deploying a remote, stateless Streamable HTTP MCP server.

2026-07-28 removed protocol sessions entirely and redesigned server-to-client requests around stateless Multi Round-Trip Requests2026-07-28 has no transport session to persist, platforms differ in what they offer for application state, whether that is an external database or a built-in primitive such as Durable Objects or distributed storageModal delivers an AI-native serverless compute platform that transforms how teams deploy MCP servers and AI workloads. The platform powers cloud infrastructure for thousands of customers running AI and data workloads, including organizations running production AI systems at scale. Modal has raised more than $466M in total, including a $355M Series C at a $4.65B post-money valuation announced in May 2026, establishing it as the purpose-built choice for AI infrastructure.
Modal takes your code, puts it in a container, and executes it in the cloud with automatic scaling. The core platform handles capacity decisions across a multi-cloud compute pool, enabling teams to focus on building rather than managing infrastructure. Key highlights:
Modal's enterprise deployments demonstrate consistent, quantifiable outcomes across AI-intensive use cases:
Modal has successfully completed its SOC 2 Type II audit, reporting that no deviations were found. Modal supports HIPAA-compliant workloads on Enterprise plans via a BAA. The platform uses gVisor-based containerization for compute isolation, TLS 1.3 for public APIs, encryption for data in transit and at rest, and role-based access controls on Team and Enterprise plans. Region pinning guarantees the chosen region, so a Sandbox pinned to a region is scheduled there rather than elsewhere.
B200+ selector documented in the GPU guide lets Modal run a Function on either B200 or B300 capacity, while CPU-only workloads run on the same platform with the same toolingBest For: Teams building production MCP servers that require CPU and GPU compute, secure code execution environments, and enterprise-grade compliance, all without managing infrastructure.
Cloudflare Workers provides edge serverless functions built on V8 isolates, executing across a network that Cloudflare reports spanning 330+ cities. The isolate architecture is designed to reduce startup overhead for lightweight Workers, which is relevant for latency-sensitive MCP interactions. Cloudflare also publishes first-party MCP documentation for building and hosting MCP servers on its platform.
Cloudflare Workers execute code at the network edge closest to users, reducing round-trip latency to centralized cloud regions. This architecture suits MCP servers handling authentication, routing decisions, or lightweight agent orchestration where response time affects user experience. Workloads requiring a fuller Linux environment can use Cloudflare Containers and the Sandbox SDK, which reached GA in April 2026.
Standard Workers impose memory and CPU limits that may not suit compute-intensive AI workloads, and Cloudflare does not expose a raw GPU directly inside a standard Worker. Teams needing accelerated inference on Cloudflare use Workers AI, which provides managed serverless GPU inference integrated with Workers, rather than raw GPU allocation for arbitrary code. Cloudflare publishes startup-time limits and profiling guidance for Workers, so startup behavior is described in runtime-specific terms rather than as a single universal figure. Best For: MCP servers handling lightweight operations such as auth validation, request routing, or A/B testing logic across a global network, optionally paired with Workers AI for managed inference.
Google Cloud Run delivers container-first serverless computing that runs standard container images without proprietary function rewrites. The platform made NVIDIA L4 GPUs generally available in June 2025 and announced NVIDIA RTX PRO 6000 Blackwell GPU general availability at Next '26 in April 2026, expanding its applicability to heavier ML inference workloads.
Cloud Run's strength lies in its adherence to standard container specifications. Teams with existing containerized applications can usually deploy with minimal change, avoiding the "rewrite tax" imposed by proprietary function runtimes. Images must still satisfy the Cloud Run container runtime contract: for services, the ingress container must listen on 0.0.0.0 at the configured PORT, and images must meet supported Linux and image requirements. Cloud Run therefore minimizes rather than eliminates migration work, which still reduces vendor lock-in concerns for organizations evaluating multi-cloud strategies.
Deployment involves familiarity with the GCP ecosystem, including Artifact Registry for image storage and IAM for access control. Google does not publish a single universal cold-start figure for Cloud Run: GPU-equipped instances become available with GPU drivers installed, which is not the same as full model readiness, and CPU cold starts depend substantially on container size, workload, and initialization behavior. Best For: Teams with existing containerized MCP servers seeking serverless scale-to-zero within the Google Cloud ecosystem, especially those needing GPU support without proprietary rewrites.
Vercel now positions itself as a full-stack platform for modern web and AI applications, with deep integration for Next.js and the Vercel AI SDK. Vercel introduced Fluid Compute on February 4, 2025, made it the default compute model for new projects in April 2025, and launched Active CPU pricing, which bills CPU only while code is actively using it, on June 25, 2025.
Vercel suits teams building MCP servers as extensions of Next.js applications, and by 2026 it also supports backend-only services, agents, and MCP servers as first-class workloads following Vercel Services and the Ship 2026 releases. The platform's TypeScript-first tooling and integrated AI SDK streamline deployment of agent endpoints alongside frontend code, and preview deployments enable testing MCP changes in isolation before production rollout.
Vercel's deepest optimizations and conventions still center on the Next.js and frontend developer experience, so fit for heavily specialized backend runtimes varies by workload. Reported cost savings reflect selected early-adopter workloads rather than a general result, so cost outcomes are workload-dependent. Best For: Teams adding MCP capabilities alongside Next.js or other Vercel-hosted applications, prioritizing developer experience and integrated AI tooling.
Azure Container Apps abstracts Kubernetes complexity while providing serverless container execution with event-driven autoscaling. It is a managed serverless container platform built on Kubernetes and open-source technologies such as KEDA and Dapr, with AKS positioned separately for workloads that need direct Kubernetes APIs and cluster-level control.
Azure Container Apps suits organizations already invested in Microsoft's ecosystem. Teams using .NET, Microsoft Entra ID, and Azure DevOps find natural workflow continuity. The platform's KEDA integration enables sophisticated autoscaling patterns beyond simple HTTP triggers, and serverless GPU profiles bring accelerated inference into the same consumption model.
The learning curve is steeper for teams unfamiliar with Azure's service model. While abstracting Kubernetes operations, the platform still exposes concepts that involve understanding container orchestration principles. GPU selection is narrower than dedicated AI platforms, and compliance coverage is program, service, region, and offering specific rather than blanket. Best For: Enterprise Microsoft-ecosystem teams deploying containerized MCP servers with complex autoscaling requirements and existing Azure service dependencies.
Netlify is built around the JAMstack deployment model, combining git-connected workflows, Deploy Previews, and serverless functions. The platform offers edge-based and standard serverless function execution for extending static sites with dynamic capabilities, along with current guidance for hosting MCP servers.
Netlify's strength lies in its established deployment workflow. The platform has mature support for Deploy Previews, atomic deployments, and rollbacks. For teams building MCP endpoints alongside static frontends, this workflow continuity provides productivity benefits.
Netlify's function configuration sets execution time limits that differ between synchronous, scheduled, and background functions, so long-lived synchronous MCP operations involve architectural consideration, such as moving work into background functions. The platform optimizes for frontend-adjacent use cases rather than compute-intensive backend processing, and does not offer raw GPU execution. Best For: Teams with JAMstack frontends adding lightweight MCP endpoints, prioritizing established deployment workflows over heavy backend compute.
Koyeb set out to build its cloud and serverless platform in 2021 and has expanded steadily into AI infrastructure across 2025 and 2026, adding serverless GPUs, Sandboxes, MCP tooling, and Light Sleep scale-to-zero. In February 2026, Koyeb announced it had entered into a definitive agreement to join Mistral AI, with completion subject to closing conditions.
Koyeb's 2025 and 2026 product work demonstrates sustained commitment to AI workloads. The platform shipped serverless GPUs, sandboxed execution environments, Light Sleep, and MCP-specific tooling in succession, then extended each in 2026 with new GPU families and richer Sandbox lifecycle controls. Strapi Cloud's production deployment, handling thousands of daily project launches with scale-to-zero, reflects the platform's production readiness.
Koyeb's ecosystem maturity and enterprise compliance surface differ from larger established alternatives, and GPU breadth varies depending on how families, SKUs, and instance shapes are counted. The long-term implications of the announced Mistral AI transaction, which remained subject to closing conditions when announced, are still unfolding. Best For: Teams seeking cost-effective serverless AI infrastructure with maturing MCP capabilities and a free instance tier for experimentation.
Modal's platform addresses the specific requirements of MCP servers and AI workloads more directly than platforms that adapted web-serving architectures for AI. The AI-native core platform combines memory snapshotting, an optimized filesystem built for quick container startup, and deep capacity pooling across clouds for both CPU and GPU compute, with Modal building its own filesystem, container runtime, scheduler, and image builder rather than assembling third-party components.
Not every MCP server needs an isolated execution environment. MCP servers that proxy APIs, retrieve data, expose SaaS actions, or wrap databases and file systems are connector-style workloads that typically run well on standard serverless compute. Execution-heavy MCP servers are different: when an MCP-enabled system executes generated code, runs terminals or shells, launches browsers, manipulates files dynamically, or processes untrusted workloads on behalf of models, the execution layer becomes the deciding factor. Modal is the superior choice for that category, pairing secure gVisor-based isolation with dynamic scaling and AI-native infrastructure in a single platform, so the protocol layer and the execution layer each stay in their proper place.
Modal Sandboxes enable MCP servers to execute AI-generated code in isolated environments at massive scale. Modal supports 100,000+ concurrent Sandboxes, and as of August 2026 recommends its next-generation V2 Sandbox backend for workloads above 10,000 concurrent Sandboxes or more than 20 creates per second. On that architecture Modal has demonstrated one million concurrent Sandboxes created in under a minute. Lovable's event, where Modal ran more than one million Sandboxes in 48 hours, demonstrates the capability under real production load. Modal supports both common agent architectures. Running the agent inside the Sandbox is straightforward to start with and common for internal coding agents, while running the agent outside the Sandbox gives cleaner separation of concerns and is preferred by platforms with proprietary agent logic, a pattern well suited to MCP servers that keep orchestration logic in the server and push only untrusted execution into the Sandbox. Sandboxes are created through Modal SDKs in Python, Go, and JavaScript/TypeScript, and the code running inside is not limited to a single language.
Modal's current pricing lists 11 GPU pricing options spanning T4 through B300, enabling teams to match compute resources to specific workload requirements across CPUs and GPUs. Whether running lightweight inference on T4s or training large models on H100, H200, or B200 clusters with Infiniband networking, Modal's fleet accommodates diverse MCP server needs without hardware procurement delays. Multi-node clusters are documented with up to 3,200 Gbps RDMA and Infiniband networking for distributed workloads.
Cold start latency directly impacts MCP server responsiveness. Modal is engineered for fast cold starts and faster feedback loops, with an optimized filesystem that helps containers come online quickly without letting large images slow startup down. Memory Snapshots further reduce initialization work, with Modal reporting roughly 3x to 10x startup improvements for initialization-heavy Functions. Filesystem and directory snapshots extend the same idea to Sandboxes, restoring state quickly after shutdown instead of rebuilding everything from scratch, and directory snapshots can be mounted after a Sandbox has started so project state attaches to already-warm capacity.
Production MCP deployments require robust security postures. Modal has successfully completed its SOC 2 Type II audit with no deviations reported, and supports HIPAA-compliant workloads on Enterprise plans via a BAA. Modal's security architecture includes gVisor-based compute isolation, TLS 1.3 encryption for public APIs, encryption in transit and at rest, and role-based access controls. Sandbox tunnels connect traffic directly to the Sandbox, and connection tokens authenticate access to servers running inside a Sandbox, which is the mechanism to reach for when sandbox-backed previews are embedded in a customer-facing application.
Modal's code-first SDKs enable teams to define infrastructure alongside application code, eliminating YAML configuration and cloud console navigation. Functions, containers, and resources are specified programmatically, enabling version control, code review, and reproducible deployments. Modal supports code-defined infrastructure in Python, TypeScript, and Go, with JavaScript/TypeScript and Go SDKs for calling deployed Modal Functions, using Sandboxes, and managing resources from other services. Because Sandboxes run whatever runtime the workload requires, the SDK language and the language executing inside the Sandbox are independent choices. For teams evaluating serverless platforms for production MCP servers, Modal's combination of AI-native infrastructure, production-scale sandboxes, and enterprise compliance makes it the strongest fit for workloads requiring CPU and GPU compute, secure code execution, and fast startup. Explore the Modal documentation and the first-party MCP server example to see how the platform addresses MCP deployment patterns.
Check the sandboxes documentation to explore implementation patterns.
View Sandboxes DocsHigh-traffic MCP servers need infrastructure that can scale with variable agent traffic while maintaining low latency for real-time interactions, even though the MCP specification itself imposes no particular scaling or infrastructure requirement. Serverless platforms reduce capacity planning by automatically scaling compute resources based on demand and billing only for actual usage rather than reserved capacity. For AI-intensive MCP workloads, platforms like Modal extend this model to both CPU and GPU resources, enabling teams to access accelerated compute without hardware procurement or cluster management.
No. MCP is the protocol and interface layer, while sandboxes are the execution and isolation layer, and the two should not be conflated. Many MCP servers are lightweight wrappers around APIs, databases, SaaS tools, or file systems, and these connector-style servers, which proxy APIs, retrieve data, or expose SaaS actions, typically do not need an isolated execution environment at all. Sandboxing becomes important for execution-heavy MCP servers: those that execute AI-generated code, run terminals or shells, launch browsers, manipulate files dynamically, or process untrusted workloads on behalf of models. For that second category, Modal Sandboxes provide gVisor-isolated execution at production scale.
GPU availability varies significantly, and the allocation model matters as much as the hardware list. Modal's pricing page lists 11 GPU pricing options spanning T4 through B300, purpose-built for AI inference and training, alongside CPU workloads on the same platform. Google Cloud Run supports NVIDIA L4, generally available since June 2025, and RTX PRO 6000 Blackwell, announced GA at Next '26 in April 2026. Azure Container Apps offers serverless A100 and T4 workload profiles. Koyeb offers multiple NVIDIA families and added RTX PRO 6000, H200, and B200 in January 2026. Cloudflare does not expose a raw GPU inside a standard Worker, but Workers AI provides managed serverless GPU inference on Cloudflare's network. Vercel and Netlify focus on CPU-based execution for their function runtimes.
Cold starts, meaning the delay when spinning up new compute instances, directly affect MCP server responsiveness, but there is no defensible universal range across platform categories. Startup behavior is runtime, workload, and configuration specific, shaped by image size, initialization work, and how much can be done before a request arrives. Modal is engineered for fast cold starts through an optimized filesystem that keeps large images from slowing startup down, and reports roughly 3x to 10x initialization improvements from Memory Snapshots, with filesystem and directory snapshots restoring Sandbox state instead of rebuilding it. Other platforms document their own startup and scale-to-zero characteristics in runtime-specific terms, so any startup figure reflects the specific runtime, configuration, workload, and methodology behind it.
Secure sandboxed execution matters for the subset of MCP servers that run AI-generated code, rather than for MCP servers as a category, and managed options are now widespread. Modal's Sandboxes use gVisor isolation to run untrusted code at scale, supporting 100,000+ concurrent Sandboxes, with the V2 backend for the highest concurrency tiers, and they run whatever language the workload requires. Cloudflare's Sandbox SDK provides isolated Linux environments and reached GA in April 2026. Azure Container Apps dynamic sessions provide platform-managed environments with Hyper-V isolation. Vercel Sandbox provides isolated Linux execution for agents. Koyeb launched Sandboxes in public preview in November 2025 and added auto-deletion and lifecycle management in January 2026. Sandbox support therefore varies by isolation model and scale ceiling rather than by simple presence or absence.
Enterprise deployments should evaluate SOC 2 Type II status, HIPAA coverage where applicable, and data residency controls, keeping in mind that coverage is product-surface specific rather than portfolio-wide. Modal has successfully completed its SOC 2 Type II audit and supports HIPAA-compliant workloads on Enterprise plans via a BAA. Modal also supports region pinning, so a workload pinned to a region is scheduled in that region, and sandbox tunnels connect traffic directly to the Sandbox. Azure Container Apps is included in the scope of multiple Microsoft compliance programs, such as FedRAMP High and DoD IL2 audit scopes, rather than inheriting every Microsoft certification automatically. Compliance coverage across the market therefore varies by platform, program, service, region, and offering.