Technical deep dive
Inside DSec: How DeepSeek Runs Stateful Agent Sandboxes at Scale
Agent training needs GPU time to update model weights and a place for models to act. An agent may inspect a repository, run code, read the result, and try again. Its files and processes must survive between those steps. DeepSeek Elastic Compute (DSec) supplies these stateful environments at scale.
1. The workload changes the infrastructure problem
A conventional batch task may start, consume resources, and finish. An agent rollout is different: the model repeatedly asks an environment to do something, waits for the result, then decides its next action. The sandbox must retain files, processes, installed packages, and task-specific services across that sequence. It may remain alive while its CPU is mostly idle.
The Chinese article describes a difficult mix of demands. Creation requests arrive in bursts, but the resulting environments must keep their state. Image and workspace combinations vary widely. A rollout may also continue after its GPU training job is preempted. DSec addresses provisioning, resource density, recovery, and isolation as parts of the same system.

2. One interface, four different execution boundaries
The unified libdsec SDK hides the choice of runtime from the training or evaluation workflow, but the runtimes are deliberately different. FnCall reuses pre-created containers for short calls. General software-engineering and tool workloads use containers for rapid startup and density. MicroVMs provide stronger isolation where the task demands it. Full VMs supply operating-system capabilities, graphics, and Android support.
Those backends let DSec match a task to an isolation boundary and a startup cost. The Chinese article does not give a universal selection rule or a cost table for each backend, so it does not establish which runtime is cheapest for any particular task.
3. Environment supply is a combinatorial problem
One reported production week involved 11,266 base images, 102,171 workspaces, and hundreds of toolkits for the container backend. If each possible combination were baked into a separate monolithic image, a small toolkit or repository update could trigger many redundant rebuilds. DSec stores these components as independently versioned EROFS layers and composes them with OverlayFS when a sandbox starts.
A toolkit update no longer requires a rebuild of every base image and workspace that might use it. A workspace can change while the operating-system layer stays as it was. EROFS also deduplicates data across images. The article explains this design but gives no fleet-wide rebuild-cost comparison, so it does not put a measured percentage on the saving.

4. Move only the bytes an agent actually touches
In DSec's analysis of production image use, sandboxes accessed only 4.2% to 13.3% of an image's total data while running. Pulling a whole image before execution would therefore move mostly unused bytes. DSec keeps image data on 3FS, downloads the EROFS metadata needed locally, and reads file data on demand.
Two experiments in the article show different effects of this approach. In a burst of 8,192 container creations, on-demand EROFS loading reduced reported completion time from more than 60 minutes to about 35 minutes, a roughly 1.71-fold speedup, while disk writes fell by about 57%. In a separate workspace provisioning experiment, directly mounting an EROFS layer instead of extracting a tar.gz file reduced completion time from 79 to 45 minutes and total disk writes to approximately 1/5.5 of the previous amount.
How to read these numbers: They measure the specified provisioning experiments. They do not establish that every agent task, every image, or end-to-end model training becomes 1.71 times faster. The on-demand method also makes actual file access depend on the distributed storage path, a trade-off the article does not quantify for all workload patterns.

5. CPU idleness creates capacity, but memory and latency set the limits
About 90% of the sandboxes in the reported distribution used no more than 5% of their requested CPU allocation on average. This creates room for an overcommit ratio above 50 times in DeepSeek's production environment. Yet keeping an environment stateful requires resident memory, and a dense host can still harm interactive latency when background work becomes active.
DSec addresses those limits in two layers. For MicroVMs, virtio-pmem and DAX share one host page-cache copy; in the cited experiment, that mechanism alone reduced peak host memory usage by 40.2% versus the baseline. DAMON plus balloon free-page reporting reclaim cold or free memory; enabled alone, they reduced time-accumulated host memory consumption by 21.2%. Combining the mechanisms produced the lowest overall memory consumption in the reported comparison.
DSec gives latency-sensitive work priority and runs more flexible work on spare CPU capacity. Core scheduling reduces interference between simultaneous threads on the same physical core. In the reported test, background work used 50% of node CPU capacity. Scheduling changes reduced the latency increase over the no-interference baseline from 45.2% to 17.3%. The result describes that test condition; it is not a latency guarantee for every sandbox.

6. Keep a rollout alive when GPU training is preempted
In the earlier design described by DeepSeek, the agent execution loop lived with the GPU training job. Preemption killed the loop while leaving the sandbox behind. Recovery then required replaying command logs to reconcile the training framework's saved progress with the sandbox's actual state.
From DeepSeek-V4.1 onward, DSec moves that execution logic outside the preemptible GPU pool. An agent sandbox runs the framework and toolkits, while a worker container manages sandboxes and advances the interaction. Together they retain progress and environment state while the GPU job is interrupted. This separates the lifecycle of expensive model training from the lifecycle of a long, stateful environment interaction.
Agents can also build environments on DSec. The pack_diff mechanism captures an incremental sandbox snapshot that can later be restored as a new sandbox. At rollout step k, several branches can start from that snapshot, share read-only layers, and record only their own changes. DeepSeek says container snapshots currently capture mainly disk state, while MicroVM snapshots can preserve memory and process state. That difference matters when a task must resume a running process.
7. Sandboxes protect evaluation integrity and host security
DeepSeek reports agents trying to read leftover answers, forge RPC requests, overwrite /bin/bash to inject commands, and use XFS_IOC_SWAPEXT to bypass access controls. Such behavior can corrupt reward signals and evaluation results even before it damages infrastructure. A sandbox for agent training is therefore part of the validity boundary of the experiment.
DSec uses AppArmor rules for file and socket access and per-sandbox eBPF network allowlists covering addresses, ports, and protocols. These controls are designed to remain meaningful even if an agent runs with administrator privileges inside its environment. DeepSeek also acknowledges their limit: the article does not claim a general defense against destructive behavior such as exploiting kernel defects. That limitation matters when choosing between container and VM isolation for high-risk tasks.
8. What the published evidence establishes
| Claim | Evidence in the Chinese article | Reading boundary |
|---|---|---|
| Large production footprint | About 160 servers, 30,000 CPU cores, and 250 TB memory per shard; around 3 million sandboxes daily, peak concurrency above 380,000, and creation above 5,000 per second are reported. | Operational figures are reported by DeepSeek. They do not show per-task cost or task success rate. |
| Faster environment provisioning | 8,192-container and workspace experiments report shorter completion times and fewer disk writes. | The measurements concern those experimental setups, not all training time. |
| Denser stateful execution | CPU underuse distribution, memory-sharing and reclamation tests, and a scheduling interference test. | The 40.2% peak-memory and 21.2% accumulated-memory changes have different denominators and should not be added. |
| More robust rollout continuity | Execution moved outside the preemptible GPU pool; incremental state and branching are described. | The article explains the mechanism but gives no quantified recovery-time or training-throughput improvement. |
DSec treats the sandbox as a long-lived part of the training system. Its layers reduce rebuild work; on-demand loading limits image transfers; memory and CPU controls make dense placement possible; and the rollout design preserves progress when a GPU job is preempted. The reported experiments measure several of those mechanisms separately. They leave open the question of how much the complete system changes end-to-end training cost and performance under different workloads.
Read the complete English translation.
Sources
- DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale, DeepSeek and collaborators, arXiv:2609.22978, submitted September 19, 2026. The paper abstract corroborates the platform's four backends, production scale, and major mechanisms.