Complete English translation
DeepSeek Elastic Compute (DSec): Sandbox Infrastructure for Large-Scale Agent Training
DeepSeek Elastic Compute (DSec) provides the sandboxes used throughout DeepSeek-V4 training, evaluation, and data preprocessing. To train a reliable agentic large language model, the agent needs to learn through repeated trial and error in real environments. It reads code, changes files, installs dependencies, runs tests, and starts services. Each action changes the environment, so the sandbox must preserve state across many rounds of interaction.
Sandbox creation requests arrive in sharp bursts. Once a sandbox starts, its CPU is idle most of the time, but its memory must stay resident. Agent tasks need many different environments, so base images are rarely reused. Tasks can run for a long time, and resource preemption may interrupt training. These workload characteristics shaped DSec's design.
Unified access for diverse tasks
Agent tasks differ in the isolation they need, the operating-system features they use, and the overhead they can tolerate. DSec supports four execution backends: FnCall, Container, MicroVM, and Full VM. A single Python SDK, libdsec, lets each task use the appropriate backend.
FnCall reuses pre-created containers for short tasks such as online evaluation. Container serves the most common software-engineering and tool-calling tasks with relatively fast startup and high deployment density. MicroVM provides a stronger isolation boundary for scenarios such as security tasks. Full VM provides a complete operating-system environment for applications such as graphical interfaces, graphics rendering, and Android.

Layered, composable environments
Agent training requires a large number of environments. The first challenge is how to build and update them in bulk. As an example from one week of production data in 2026, the container backend used 11,266 base images, 102,171 workspaces, and hundreds of toolkits. Under the traditional approach, updating or replacing any one of these components would require rebuilding a large number of images.
DSec splits each sandbox environment into three parts: a base image with the operating system and basic software, a workspace with the task's code repository and dependencies, and a toolkit such as DeepSeek Harness. They change at different rates, so DSec versions them separately and combines them at runtime.
DSec stores the base images, workspaces, and toolkits in EROFS format. EROFS also separates metadata from data and supports deduplication across images. When DSec creates a sandbox, OverlayFS combines the required EROFS layers on demand. An update to basic software, task code, or a toolkit requires rebuilding only the layer that changed.

On-demand image loading
We analyzed image data used in production and found that a running sandbox actually accesses only 4.2% to 13.3% of the image's total data. Pulling a complete image to local storage in advance would transfer mostly data that the sandbox never uses.
DSec therefore stores all image data in the 3FS distributed file system. It pulls only the metadata needed for an image to local storage (a property of EROFS), while the much larger data portion is read on demand when accessed. This makes I/O overhead scale with the amount of data actually used.
In an experiment that created 8,192 containers in a concentrated burst, on-demand loading reduced task completion time from more than 60 minutes to about 35 minutes compared with pulling full images, for a speedup of approximately 1.71, and reduced disk writes by about 57%. In another workspace provisioning experiment, replacing per-sandbox extraction of tar.gz files with direct mounting of EROFS layers reduced task completion time from 79 minutes to 45 minutes.

tar.gz files to directly mounting EROFS layers, task completion time fell from 79 minutes to 45 minutes, and total disk writes fell to about 1/5.5 of the original amount.
High-density resource management
During agent training, sandboxes spend most of their time waiting for the model's next action. As the figure shows, about 90% of sandboxes use no more than 5% of their requested CPU allocation on average. The CPU sits idle for long periods, while memory must stay resident to preserve files, processes, and other state. This allows a high degree of overcommitment. In our production environment, the overcommit ratio exceeds 50 times.

To deploy resources at high density, caches must be shared and idle memory reclaimed as much as possible. DSec uses virtio-pmem and DAX to let MicroVMs on the same host share a single copy of the host page cache. In an experiment, enabling this mechanism alone reduced peak host memory usage by 40.2% relative to the baseline. In an experiment, enabling memory reclamation alone (DAMON and balloon free-page reporting) reduced host memory consumption accumulated over time by 21.2% relative to the baseline. Overall memory consumption was lowest when the two mechanisms were used together.

Even at this deployment density, DSec gives priority to tasks with stricter response-time requirements, lets tasks with looser latency requirements use idle CPU capacity, and uses core scheduling to reduce interference between hyperthreads on the same physical core. Experiments show that, through CPU scheduling optimization, when other tasks on the same machine occupied 50% of the node's CPU capacity, the latency increase for latency-sensitive tasks relative to the no-interference baseline fell from 45.2% to 17.3%.


Decoupling rollout execution from GPU training
In reinforcement-learning (RL) training, an agent usually needs multiple rounds of interaction with a sandbox environment to complete a rollout. In the early training workflow, the agent execution loop and the GPU training task ran in the same Pod container. After a training task was interrupted by preemption, the environment sandbox was still there, but the agent execution loop responsible for advancing the interaction had stopped. On recovery, the system had to replay command logs to reconnect the progress saved by the training framework with the state actually executed in the sandbox. That recovery logic was very complex and increased the coordination cost among multiple components.
Starting with DeepSeek-V4.1, we moved this execution logic into DSec sandboxes and split it between an agent sandbox and a worker container. The agent sandbox runs the agent framework and toolkits; the worker container manages sandboxes and advances the interaction process. Both run outside the preemptible GPU resource pool and together preserve execution progress and environment state. As a result, when a GPU training task is preempted, the agent's execution state remains intact; once training resumes, the agent can continue from the point of interruption.
Using agents to build the environments in which agents run
Large-scale agent RL training and evaluation require highly diverse runtime environments, including, but not limited to, binary dependencies, code repositories, Harness toolkits, evaluation scripts, and a series of other components used to produce agent outputs and measure their correctness. Faced with these complex requirements, using agents to automate environment construction is an efficient and scalable way to make agent environments.
We noticed that the agent building an environment is already running inside the environment it is building. Rather than maintain one platform for agent training and another for environment construction, we use the same platform for both: DSec. This simplifies the system architecture and keeps the build environment consistent with the runtime environment.
For environment construction, DSec introduces a packaging mechanism called pack_diff. An agent can ask the platform to take an incremental snapshot of a sandbox and later restore it as a new sandbox. This lets us save sandbox state and even turn each round of agent interaction into a reusable environment.
These incremental snapshots can also be used to branch rollouts: save a snapshot at step k, restore multiple sandboxes from the same state, and let each branch continue exploring separately. The branches share read-only layers and record only changes, avoiding repeated execution of the steps before the branch. MicroVM snapshots support saving and restoring memory and process state, while the container side currently mainly supports disk-level snapshots.
Security boundaries for agents
As agents become more capable, the security boundaries of their training environments must continue to improve. In production, we have observed agents attempt to read leftover answers, forge RPC requests, overwrite /bin/bash to inject commands, and even use XFS_IOC_SWAPEXT to bypass access controls. These behaviors can affect training and evaluation results and can even damage the runtime environment.
If the environment offers a shortcut to obtaining a reward, a model may exploit it. DSec therefore makes fine-grained access control a basic feature: AppArmor constrains file reads and writes and socket access, and those restrictions remain effective even when an agent runs as an administrator; eBPF enforces a network-access allowlist for each sandbox, restricting the addresses, ports, and protocols it can connect to.
These measures can mitigate only some problems, however. There is still no general defense against destructive behavior such as triggering kernel vulnerabilities. We expect the contest between us and agents to continue as model capabilities improve, with the system's security mechanisms evolving and improving along with it.
Production data
DSec scales horizontally through shards. Each expansion shard contains about 160 servers, providing around 30,000 CPU cores and 250 TB of memory. A single shard serves about 3 million sandboxes per day, has peak concurrency above 380,000, and can create more than 5,000 sandboxes per second.
Multiple such shards are deployed in production and can support millions of concurrently running sandboxes.
From DeepSeek-V3.2 to DeepSeek-V4.1, DSec carried all sandbox workloads in agent training, evaluation, and data preprocessing. We have made the DSec technical report, DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale, available on arXiv to share the engineering practices supporting large-scale agent sandbox execution.
Conclusion
We believe the potential of agents is boundless. Next, we will increase the number and variety of agent runtime environments by hundreds or thousands of times and bring the capabilities developed through these tasks back to open models. That will require more diverse environments, more reliable infrastructure, and more development partners.