Sandboxes

Best Sandboxes for Data-Analyst Agent RL in 2026

Training data-analyst agents through reinforcement learning demands infrastructure that can handle parallel rollouts, stateful environments, and compute resources across CPUs and GPUs for policy inference and gradient updates. The challenge extends beyond simple code execution: you need isolated environments that reset reliably, scale to the level of parallelism your rollout batch size and episode length actually require, and connect cleanly to the surrounding training stack. Selecting the right AI infrastructure can materially affect training throughput, reliability, cost, and scalability, alongside model quality, environment design, verifier quality, and reward design.

Modal TeamEngineering
August 202624 min read
Data-analyst agent reinforcement learning sandbox environments

One clarification before the list. The seven entries below do not all sit at the same layer of the stack, and treating them as interchangeable "sandboxes" is the most common source of confusion in this space. A sandbox execution substrate supplies isolated compute. An RL environment defines observations, actions, state, reset semantics, and often verifier behavior. An environment-generation system creates environments and tasks. A training recipe or orchestration framework trains or coordinates policies. A protocol layer standardizes communication between them. A reference implementation shows how the pieces fit together. Each entry below is labeled with the layer it actually occupies, so you can tell what you are choosing between. This guide examines seven sandboxes, environments, and frameworks suited for data-analyst agent reinforcement learning in 2026, starting with Modal, which combines CPU and GPU Sandbox execution with first-party training and inference infrastructure on the same stack.

Key Takeaways

GPU placement matters more than GPU presence: RL training generally requires GPUs for policy inference and gradient updates. Modal Sandboxes run across CPU and GPU compute, and GPU-enabled sandboxes are additionally useful when the agent's environment-side workload itself needs acceleration, such as GPU data processing, simulation, or in-environment model execution

Environment synthesis can alleviate the training data bottleneck, not solve it: Agent World Model generates 1,000 executable environments automatically, while its own error analysis reports that one evaluator judged 74% of generated environments to contain at least one bug, which converts scarcity into a quality-control problem

Standardized protocols reduce integration work: OpenEnv's Gymnasium-inspired API lets compliant trainers and compliant environments interoperate without bespoke glue code, though the project still labels itself experimental

Purpose-built training recipes report strong domain results: DataMind-14B reports a 71.16% average across the authors' selected data-analysis benchmarks, and its domain-specific approach is a plausible contributor even though the evaluation does not isolate it as the sole cause

Production infrastructure should scale with the training loop: Modal's production Sandboxes support 100k+ concurrency with fast cold starts, while its V2 backend targets million-scale workloads and has demonstrated one million concurrent Sandboxes

1. Modal (sandbox execution substrate)

Modal delivers cloud-native Sandboxes engineered for high-performance AI workloads, including data-analyst agent reinforcement learning. The platform provides dynamically defined containers with fast cold starts, GPU-enabled Sandbox execution, and 100k+ concurrent Sandboxes on the production configuration. Separately, Modal's V2 Sandbox backend, currently in beta, is built for extreme concurrency and has demonstrated one million concurrent Sandboxes created in under a minute.

Key Capabilities

CPU and GPU-enabled Sandboxes: Sandboxes run across CPU and GPU compute, and you can attach H100, H200, B200, B300, A100, and other accelerators to isolated Sandboxes. Sandbox resource documentation confirms GPU Sandbox support. Modal's homepage describes GPU Sandboxes scaling to thousands of concurrent runs, and RL rollout environments reaching hundreds of thousands concurrently in seconds

Fast cold starts: Engineered for fast cold starts and faster feedback loops, with an optimized filesystem that helps containers come online quickly without letting large images slow startup down. Filesystem, directory, and memory snapshots further reduce repeated setup work between rollouts, and directory snapshots let you capture only part of a Sandbox, such as separating project files from platform-owned dependencies

High concurrency: Scale from zero into the hundreds of thousands of concurrent rollout environments in seconds, without operating your own fixed container cluster. Concurrency scales with plan: current pricing lists 100 containers and 10 GPU concurrency on Starter, 5,000 containers and 50 GPU concurrency on Team, and custom capacity on Enterprise. The V2 backend is recommended for workloads above 10,000 concurrent Sandboxes

Unified platform: Training, inference, and Sandbox execution run on the same core platform, reducing the overhead of coordinating multiple vendors

Code-first SDK: Define environments programmatically, with configuration expressed as code and no YAML. Modal supports code-defined infrastructure in Python, TypeScript, and Go, and code running inside a Sandbox is not limited to one programming language: a Sandbox can run whatever runtime or language the workload requires

Usage-based billing: Sandboxes are charged by CPU and memory consumption, by the second, and you can tune requested CPU and memory to match the shape of your rollout workload

Infrastructure Architecture

Modal built its own container runtime, scheduler, filesystem, and container image builder, and describes its core platform as cloud infrastructure designed specifically for AI workloads. That investment is what makes the combination of fast cold starts, high Sandbox concurrency, and first-party training infrastructure possible on one stack. On security, Modal's security documentation states that compute jobs run under gVisor, that Modal has successfully completed a SOC 2 Type 2 audit, that public APIs use TLS 1.3, and that user data is encrypted in transit and at rest. Modal supports HIPAA-compliant workloads on Enterprise plans via a BAA.

Documented Results

Lovable's Founder and CEO Anton Osika is quoted on Modal's Sandboxes page saying Modal enabled running "tens of thousands of app creation sessions in an instant." For data-analyst agent RL, that maps onto collecting rollouts across large numbers of environments concurrently. Ramp is a production, agent-specific example: it runs background coding agents on Modal Sandboxes, and those agents generate code changes and write them back into commits or pull requests, a pattern documented in Modal's write-up on Ramp's background coding agent. The same loop of generate, execute, inspect, and iterate is what a data-analyst agent performs during RL rollouts. Best For: Teams that want GPU-enabled Sandbox execution and first-party training infrastructure on the same platform, particularly RL loops that need fast environment turnover at high concurrency.

2. DataMind (data-synthesis and agent-training recipe)

DataMind is an open-source data-synthesis and agent-training recipe focused on data-analytic agents, presented at ICLR 2026 as a scalable approach to post-training data scaling and multi-turn RL. Developed by researchers affiliated with Zhejiang University and Alibaba Group, it targets agents that work across heterogeneous data formats including CSV files, SQLite databases, and Excel spreadsheets. It is not a sandbox product, and it runs on top of whatever execution infrastructure you provide.

Core Recipe Components

Combined training pipeline: Supports both supervised fine-tuning (SFT) and reinforcement learning stages in a unified workflow

Multi-turn analysis via replayed code context: Rather than depending on a persistent live global-variable pool, DataMind preserves prior code chunks and concatenates them to reproduce the effect of global execution while maintaining isolation

Domain-specific dataset: The paper reports 11,707 filtered multi-turn trajectories across diverse data-analysis tasks and 18 task categories. The current Hugging Face dataset lists 12,187 rows

Related process-level reward work: DataPRM provides step-level verification for multi-step analysis tasks. It is a separate project in the same research line, not an intrinsic component of the ICLR 2026 DataMind recipe

Academic Validation

The project repository distinguishes three separate projects and venues: DataMind was accepted at ICLR 2026, DataMind-Analysis at AAAI 2026, and DataPRM at KDD 2026. DataMind-14B reports a 71.16% average across the authors' selected data-analysis benchmarks, outperforming the compared general-purpose and proprietary baselines in that evaluation. The domain-specific training approach is a plausible contributor to that result, though the evaluation does not isolate it experimentally as the sole cause.

Open Source Availability

Licensing differs by artifact. The DataMind-7B and DataMind-14B model repositories are Apache-2.0, while DataMind-12K is currently published under MIT. The code can be deployed on compatible compute infrastructure subject to its software dependencies and hardware requirements, and the reported training setup itself used substantial GPU resources. Modal provides training infrastructure spanning CPU and GPU compute, supports common frameworks including PyTorch and Hugging Face TRL, and lets you define the training environment as code in Python, TypeScript, or Go. Best For: Research teams and practitioners who want a published training methodology for data-analyst agents, with released models, data, and benchmarks intended to support reproduction and follow-on research.

3. Agent World Model (environment-generation system)

Agent World Model (AWM), developed by researchers associated with Snowflake AI Research and UNC-Chapel Hill, addresses a major bottleneck in agent RL: the scarcity of diverse training environments. AWM automatically synthesizes executable environments backed by SQLite databases, generating tasks at a scale that would be costly and labor-intensive to build manually. It generates environments; it does not supply the isolated compute they run on.

Environment Generation Capabilities

Automatic synthesis: Creates 1,000 executable SQL-backed environments from seed specifications

Broad tool coverage: Generates 35,062 tools across 10,000 tasks with MCP interface compatibility. MCP here is the protocol and interface layer that describes how tools are exposed, which is distinct from the execution and isolation layer that a sandbox provides

Database-backed state: Uses actual SQLite database instances whose schemas and sample data are synthetically generated, which provides explicit, persistent state transitions that can be checked programmatically rather than relying on purely LLM-simulated world state. Verification itself still uses code and/or LLM judging, and the paper discusses imperfections in environment and verifier signals

End-to-end pipeline: Automatically generates scenarios, tasks, database schemas, interfaces, and verifiers

Scale and Performance

The authors' RL experiments used 1,024 parallel environment instances per training step, an experimental configuration rather than a platform guarantee. The Arctic-AWM models (4B/8B/14B) were trained on a 526-environment, 3,315-task subset of the synthesized corpus, selected because of compute budget, and show improved out-of-distribution performance in the authors' evaluations, evidence consistent with generalization beyond the training tasks.

Research Impact

AWM reports an ICML 2026 acceptance, and its repository records that its infrastructure was merged into OpenEnv in May 2026. OpenEnv has since moved to huggingface/OpenEnv under broader multi-stakeholder governance, so as of August 2026 that integration sits inside a community-governed project rather than a Meta or PyTorch-only ecosystem. AWM is a recent large-scale approach to synthetic executable-environment generation for agentic RL, alongside other 2026 environment-generation research. Best For: Teams facing environment scarcity who need to generate large collections of executable, database-backed environments and tasks automatically, and who can account for quality assurance across generated environments.

4. OpenEnv (protocol and interoperability layer)

OpenEnv provides a standardized protocol for RL environments, functioning as the connective tissue between sandboxes, environments, and training systems. It was initially launched through collaboration between Meta/PyTorch and Hugging Face in 2025, and as of 2026 it is coordinated under a broader community governance model. It is an emerging open protocol and framework intended to standardize interfaces between agent environments and training systems.

Protocol Design

Gymnasium-inspired API: Familiar reset/step/state interface adapted for network-based environments

Multiple providers: The current README documents provider options including Docker, Docker Swarm, UV, Daytona, and Azure Container Apps Sandbox, with Kubernetes support described as planned. Some individual environments use their own backends: the Jupyter environment runs on E2B, which is environment-specific rather than a generic provider backend

Environment Hub: Centralized discovery for community-built environments

Experimental status: The project's own documentation describes it as experimental, so its API surface continues to evolve

Environment Catalog

OpenEnv's current catalog lists roughly 40 environments spanning coding (REPL, Jupyter, Git), games (Atari, Chess), finance (FinQA, FinRL), and web (BrowserGym) domains.

Ecosystem Integration

Current OpenEnv material documents integrations and examples for TRL, SkyRL, and TorchForge. The 2025 launch post also described integration work with verl, which is best treated as historical context rather than a documented current integration. OpenEnv can provide a common interface across compliant environments and trainers, subject to environment and runtime parity. One scope note: the current FinQA environment documentation designates the FinQA dataset as evaluation only and indicates it is not intended for training. Best For: Teams building interoperable agent training pipelines who want to reduce bespoke integration code between compatible trainers and environment backends, and who are comfortable with an evolving API surface.

5. FinQA RL with OpenEnv and Daytona (reference implementation)

This entry is a Daytona implementation guide, not a sandbox product. It documents end-to-end RL training for financial question-answering agents using OpenEnv's FinQA environment with Daytona sandboxes, and remains a detailed public reference implementation for data-analyst agent RL.

Training Implementation

500-sandbox GRPO run: Uses Qwen3-14B with LoRA adapters across 500 concurrent sandboxes, with four GPUs serving vLLM generation and the trainer while the rollout sandboxes execute the SQL environment separately

Multi-turn tool-calling: Agents interact with SQL databases to answer questions about SEC 10-K filing data

Documented accuracy improvement: The guide reports performance rising from roughly 21% to 52% over 10 iterations. This is a documented experimental result, and OpenEnv's current FinQA documentation designates the FinQA dataset as evaluation only rather than as training data

Technical Patterns

The guide covers practical production concerns often missing from research papers: sandbox pooling, stale WebSocket cleanup, staggered launches, lag-1 pipeline overlap for GPU efficiency, and batched vLLM generation. Warm pools of pre-started sandboxes are a common latency-optimization pattern generally, because they let upfront work such as starting the sandbox, launching a server, pulling a repo, or installing dependencies happen before the end user is waiting.

Learning Outcomes

Analysis of the trained model shows improvements in SQL quoting conventions, numeric formatting, and error recovery. The guide itself notes that the model did not learn a fundamentally new database-navigation strategy, and it documents dataset-specific conventions and failure modes, so it does not establish cross-task or cross-dataset transfer. Best For: Practitioners who want a detailed reference implementation and operational patterns for large-scale data-analyst agent RL, with the FinQA dataset designation in mind.

6. Harbor (task and environment orchestration framework)

Harbor is an independent agent-environment framework with an experimental integration in Hugging Face TRL. Its core concepts explicitly separate tasks and environments from container runtimes and cloud sandbox providers, so Harbor orchestrates sandboxes rather than being one.

Architecture Design

Task-suite abstraction: Each task bundles instruction, environment, and verifier into a self-contained unit

TRL integration: A HarborSpec wires a task suite into GRPOTrainer slots via train_dataset, environment_factory, and reward_funcs

External agent pattern: Supports in-process vLLM generation for efficient training loops

Multi-backend support: Works with Docker, E2B, Daytona, GKE, Modal, and Runloop. Modal publishes a direct example that runs Harbor evals on Modal Sandboxes, using the harbor[modal] dependency and Modal as the selected Harbor environment

Modularity and Current Scope

Harbor's design lets you swap sandbox backends, for example running the same task suite on Modal or on Docker, and experiment with different harnesses (bash, jupyter, terminal_notes) without rewriting training code. Individual task suites can be configured modularly, and the current TRL documentation states that a HarborSpec represents one task suite plus one harness, with multi-suite training outside the present scope. The integration sits under TRL's experimental surface, and sandbox provisioning is currently sequential.

Ecosystem

Harbor has an experimental integration with Hugging Face TRL, a widely used open-source post-training library, and example task suites include data agent environments designed for analyst workflows. Best For: Teams using TRL who want modular task definitions and the flexibility to experiment with different sandbox backends, within the current single-task-suite scope.

7. OpenEnv Jupyter (notebook environment, E2B-backed)

OpenEnv Jupyter is an environment implementation rather than a sandbox provider: its current v0.4.1 documentation says each episode creates a fresh E2B Code Interpreter sandbox. It provides a notebook-native surface for data-analyst agent RL, maintaining state across tool calls for multi-step data exploration.

Notebook-Specific Features

Stateful sessions: Variables, imports, files, and plots persist across tool calls within an episode, enabling iterative data analysis workflows

MCP tool interface: The documented tools are add_and_execute_code_cell, edit_and_execute_current_cell, execute_shell_command, a notebook-cell summarization tool, and final_answer. MCP defines how those tools are exposed to the model, while the isolated execution of the code they run is a separate infrastructure concern

Setup/verify hooks: Configure task-specific state and validation logic

Web UI included: Interactive notebook interface at the /web endpoint for debugging agent behavior

Data-Analyst Workflow Alignment

Unlike terminal-based environments, OpenEnv Jupyter matches notebook-oriented exploratory data-analysis workflows, where analysis is built up incrementally: loading data, exploring distributions, creating visualizations, and iterating on transformations. Analysts also work in IDEs, SQL clients, BI tools, terminals, and spreadsheets, so notebook alignment is a fit for a specific workflow rather than for the occupation as a whole.

Tool Coverage

The environment exposes five focused tools covering notebook execution, shell interaction, state and context handling, and answer submission. Best For: Teams training agents for exploratory data analysis tasks where notebook-style interaction matches the target workflow.

Why Modal Stands Out for Data-Analyst Agent RL

GPU-Enabled Sandboxes and Training on One Stack

Modal describes itself on its homepage as the only platform where sandboxes and training infrastructure are native to the same stack. That first-party positioning maps well onto data-analyst agent RL, where you need environment execution and model training coordinated tightly. Modal supports GPU-backed Sandboxes that scale to thousands of concurrent runs, which matters when the environment-side workload itself needs CUDA for accelerated data processing, simulation, or in-environment model execution, and the same Sandboxes run CPU-only analyst workloads just as readily. Modal's own RL architecture write-up separates GPU-intensive training, GPU-backed inference for rollout generation, and isolated environment containers, and notes that high Sandbox concurrency is what keeps the inference GPUs fed. Separately, Modal's V2 Sandbox architecture is designed for extreme concurrency and has demonstrated one million concurrent Sandboxes.

Infrastructure Built for AI Workloads

Modal's custom-built infrastructure includes a purpose-designed container runtime, scheduler, filesystem, and image builder. That investment shows up as fast cold starts and fast environment turnover, which compounds when training requires thousands of environment resets per hour. Filesystem and memory snapshots further reduce repeated setup work, and directory snapshots can be mounted after a Sandbox has started, which supports patterns such as attaching project-specific state to pre-warmed environments.

Unified Platform Reduces Vendor Sprawl

Data-analyst agent development requires training infrastructure, inference endpoints, and sandbox environments. Modal delivers all three on one platform, so you can train models, serve inference, and run sandboxed evaluation inside the same infrastructure ecosystem. Sandboxes, Functions, inference endpoints, and training jobs are distinct primitives with clearly defined roles, and keeping them in one ecosystem reduces cross-vendor integration while giving you consistent observability across the pipeline.

Code-First Developer Experience

Modal's SDK lets you define environments programmatically, with configuration expressed as code and no YAML. Modal supports code-defined infrastructure in Python, TypeScript, and Go, and Sandboxes are not tied to a single language: code inside a Sandbox can run whatever runtime or language the workload requires. This aligns with how ML engineers prefer to work: infrastructure as code that can be version-controlled, tested, and iterated alongside training scripts.

Enterprise-Ready Security

Modal has successfully completed a SOC 2 Type 2 audit, runs compute jobs under gVisor, and supports HIPAA-compliant workloads on Enterprise plans via a BAA. Sandboxes also support tunnels and connection tokens, which authenticate access to servers running inside a Sandbox, a useful control when sandbox-backed analysis surfaces are embedded in an application.

For teams building production data-analyst agents, Modal can serve as the scalable Sandbox execution backend beneath agent and evaluation frameworks such as Harbor. Explore the sandbox documentation to start building GPU-enabled environments for your agent training.

Explore the sandbox documentation to start building GPU-enabled environments for your agent training.

View Sandboxes Docs

Frequently asked questions

What defines a "sandbox" in data-analyst agent reinforcement learning?

A sandbox provides an isolated execution environment where agents can run code, query databases, and manipulate files without affecting the host system or other training instances. RL systems generally need resettable environments and sufficient rollout parallelism; reward computation can live in the environment, the verifier, or the training framework rather than in the sandbox substrate itself. For data-analyst RL specifically, the sandbox image and environment should make the required data libraries, files, database services, and tools available, since formats such as CSV, Excel, and SQLite are an environment and workload property rather than a native sandbox capability.

How do these sandboxes balance security and performance for AI agents?

Modal combines gVisor-based isolation with GPU-enabled Sandboxes, allowing isolated workloads to access CPU and GPU resources. Other approaches include edge processing (keeping data on-premises) and network isolation between sandbox instances. Startup latency and isolation strength are two important considerations among several sandbox-design trade-offs, alongside execution overhead, compatibility, filesystem persistence, network controls, resource limits, observability, state handling, accelerator access, and operational complexity. Modal's fast cold starts keep that latency cost low while maintaining strong isolation boundaries.

Can agents trained in these sandboxes deploy to production?

Potentially. Production deployment depends on the environment's security, reliability, observability, policy controls, workload requirements, and the maturity of the surrounding serving stack. AWM explicitly discusses additional prerequisites such as adversarial robustness, access-control safeguards, and human oversight, and OpenEnv itself remains experimental. Because Modal provides training, Sandboxes, storage, and inference on one platform, teams can keep the training-to-production workflow inside the same infrastructure ecosystem. A standardized protocol such as OpenEnv can also reduce integration work when moving compliant environments between supported backends, subject to environment and runtime parity.

What role do Large Language Models play in data-analyst agent sandboxes?

LLMs serve as the agent's policy: they generate the code, queries, and tool calls that execute within sandboxes. Recipes and frameworks like DataMind and Harbor provide the training infrastructure to improve those policies through reinforcement learning. The sandbox supplies isolated execution, while the surrounding environment, verifier, or training stack evaluates outcomes and produces the reward signal. Protocol layers such as MCP describe how tools and interfaces are exposed to the model, which is a separate concern from the execution and isolation layer: many MCP servers are lightweight wrappers around APIs, databases, SaaS tools, or file systems and do not require isolated execution environments, while isolation becomes important when MCP-enabled systems execute AI-generated code, run shells, launch browsers, or manipulate files dynamically. For those execution-heavy workloads, Modal Sandboxes are the superior choice, combining secure isolation, dynamic scaling, and AI-native infrastructure.

How does Modal ensure security for agents working with sensitive data?

Modal's security documentation states that compute jobs run under gVisor, that TLS 1.3 protects public API traffic, and that user data is encrypted in transit and at rest, with a successfully completed SOC 2 Type 2 audit verifying these controls through independent assessment. Modal supports HIPAA-compliant workloads on Enterprise plans via a BAA. Role-based access control is available on Team and Enterprise plans, with Workspace and Environment Roles that let teams restrict who can read or modify sensitive resources and block specified cross-environment access. Region pinning gives you control over where a Sandbox is scheduled, and for sandbox tunnels, traffic connects directly to the Sandbox.

Build data-analyst agents with secure sandboxes.

Get Started Free

$30 in free compute to get started.