Agent infrastructure | September 29, 2026

DSec shows why agent infrastructure needs an execution plane, not a pile of sandboxes

DeepSeek Elastic Compute coordinates millions of short-lived, stateful tool environments across containers, microVMs, full VMs, and function calls. Its useful lesson is architectural: separate state, admission, images, placement, and preemption from the accelerator job that trains the model.

Burst: thousands of creates/sec State: local writes + shared layers Gate: final node admission
Elastic agent execution plane with placement, storage, and sandbox workers

The scarce resource is not the only lifecycle

Agentic post-training creates a mismatch. Expensive GPU workers advance model optimization, while tool calls fan out into CPU-heavy sandboxes whose duration, image, network access, memory, and failure modes vary wildly. If every sandbox belongs to the GPU job that requested it, a slow package install or blocked network call can idle the accelerator. If the sandbox is disposable and stateless, a preemption can erase a long agent trajectory. DSec addresses that mismatch by turning agent execution into a separate service.

DeepSeek published the DSec paper on September 19, 2026. The authors describe a production platform with one SDK across function calls, containers, Firecracker microVMs, and full VMs. They report roughly three million sandbox instances per day, a peak near 380,000 concurrent instances, more than 5,000 creations per second, and individual jobs requesting as many as 32,000 sandboxes. These are self-reported production figures, not independently audited benchmarks. The evaluation cluster is also distinct from production. Treat the numbers as evidence of the design target, not a promise that another installation will reproduce it.

The deeper contribution is the decomposition. DSec does not ask one scheduler to understand every container layer, state checkpoint, network rule, GPU step, and host race. It defines ownership boundaries: a placement service selects a candidate node, a per-node edge service makes final admission, a shared filesystem serves immutable data, local disks absorb irregular writes, and the training loop can pause or move without destroying the tool environment.

A sandbox is an isolation primitive. An execution plane is the system that keeps isolation, state, placement, images, and caller lifecycles coherent under burst.

What the paper does and does not establish

The paper supplies concrete architecture, operating observations, and controlled evaluations. It does not release a complete, ready-to-install DSec distribution. DeepSeek has open-sourced relevant storage work such as 3FS, and the paper connects to systems including DeltaBox and Aries, but a team still has to assemble or buy its own control plane. Hacker News discussion usefully probes density, orchestration, and storage tradeoffs; it does not prove production adoption outside DeepSeek.

Split desired placement from final admission

The DSec control path starts with identity and project quotas, moves through a placement engine and stateless API servers, and ends at an edge service on each worker. A watcher collects node health and load. The placement engine uses that view to choose a candidate. The edge owns the host's real-time truth and may reject a placement if resources changed after the global snapshot.

This is more than an optimization. Distributed schedulers always act on slightly stale information. If the central planner is the final authority, two concurrent decisions can overbook the same node. DSec keeps an in-memory overlay of placements in flight and still gives the edge the last word. Rejection is a normal scheduling result, not a corrupt state. The API server can route a sandbox request to the owning edge by sandbox ID, while watcher and placement services remain rebuildable rather than becoming a second durable database.

LayerOwnsMust not ownFailure response
Identity and quotaProject, caller, limits, allowed modeHost-level free memoryDeny before placement
Placement engineCandidate selection and in-flight overlayFinal host admissionRetry another candidate
API serverStateless request routingSandbox durable stateAny healthy instance resumes routing
Node edgeFinal admission, lifecycle, local resourcesGlobal scheduling policyReject or reconcile local instance
Shared storageImmutable layers and bulk readsHigh-churn writable overlayCache, backpressure, fail visibly
Sandbox runtimeExec, files, HTTP, streamsCaller authorization policyReturn bounded receipts

A practical API should expose desired properties, not runtime-specific commands. The caller requests isolation, resources, image layers, network policy, persistence, and preemption behavior. The platform selects container, microVM, or VM only if policy permits. Keeping that translation server-side makes it possible to change runtimes without rewriting every training or evaluation client.

apiVersion: execution.trae.example/v1
kind: AgentSandbox
metadata:
  project: code-rl
  run_id: run-2026-09-29-042
spec:
  isolation: microvm
  image:
    base: sha256:base-locked
    workspace: sha256:repo-snapshot
    toolkit: sha256:toolchain-17
  resources: {cpu: 2, memory_mib: 4096, scratch_gib: 12}
  network: {policy: eval-egress-v4, default: deny}
  lifecycle:
    ttl_seconds: 3600
    idle_timeout_seconds: 300
    checkpoint: incremental
    preemptible: true
  receipts: [image_digests, commands, network, exit, checkpoint]

The configuration deliberately binds receipts and policy to the lifecycle. Otherwise a resumed sandbox may keep state while losing the evidence needed to explain it. For high-impact agents, connect this plane to an execution-receipt contract rather than retaining only terminal output.

Make immutable data shared and irregular writes local

Agent images are a distribution problem before they are a startup problem. A monolithic image multiplied across thousands of concurrent sandboxes can saturate registries, network links, and host disks. DSec converts OCI image content to EROFS, serves shared read-only data through 3FS, and keeps writable overlays local. For microVMs it combines an EROFS base with OverlayBD and ublk-backed writable disks. Metadata is kept near the worker; file content can be read lazily or in bulk.

The composable image model separates a stable base, a workspace, and a toolkit. That prevents the Cartesian product of language runtime × repository × evaluator tools from becoming a unique rebuilt image for every job. The paper reports one observed week with 11,266 base images, 102,171 container workspaces, 53,590 microVM workspaces, and 103 toolkits; 67.8 percent of tasks required a workspace or toolkit. Again, those figures characterize DeepSeek's workload. The design implication is portable: independently addressable layers reduce rebuild and transfer amplification.

Lazy fetch is not free. It turns startup bandwidth into runtime latency, so prefetch policy should follow access traces. DSec reports that eager pulling increased completion time by 1.7× in one evaluation, while on-demand loading reduced cumulative disk writes by 57 percent. A team should reproduce that comparison with its own layer sizes and locality. Package-compilation jobs and browser agents have different access shapes.

StrategyStartupRuntime riskBest fit
Full image pullSlow and bandwidth-heavyPredictable after startSmall image, long job, good locality
Lazy immutable layersFastFirst-touch latency and shared-store loadLarge image, sparse access, short jobs
Warm local cacheFast on hitEviction and skewRepeated toolchains
Composable layersDepends on hit mixLayer compatibility and provenanceMany repositories and few toolkits

State ownership must be explicit. Immutable image layers belong to the content-addressed store. The writable filesystem belongs to the sandbox lifecycle. Agent trajectory, reward, and orchestration state belong to the training or evaluation service. Checkpoints bridge those owners; they should never blur them. A snapshot without its image digests, runtime version, policy, and parent run is not a reliable resume point.

Capacity-plan for concurrency, not average CPU

DSec reports that 90 percent of its container and microVM sandboxes average no more than 5 percent of requested CPU. That invites density, but an average can hide synchronized bursts. The platform uses high packing density, Linux scheduling controls, memory reclamation, and isolation to recover stranded capacity without pretending requests are meaningless.

The paper reports stable observations of at least 3,200 containers or 800 microVMs on a node. Those are operating points in its environment, not universal hard limits. File descriptors, process count, kernel memory, network connections, page cache, and storage IOPS can become the real constraint before CPU. Define a per-mode density envelope from load tests and keep a burst reserve.

admissible = min(
  cpu_burst_budget / observed_p95_cpu,
  memory_budget / observed_p99_rss,
  process_limit / p99_processes,
  fd_limit / p99_file_descriptors,
  storage_iops_budget / p95_iops,
  network_conn_budget / p95_connections,
  tested_runtime_density
) * safety_factor

For example, suppose a 128-core worker has a 96-core sandbox budget after system reserve. If a class requests two cores but p95 observed usage is 0.12 cores, CPU alone suggests 800 sandboxes. If the p99 resident set is 640 MiB and the node exposes 480 GiB to sandboxes, memory suggests 768. A tested microVM ceiling of 600 and a 0.8 safety factor produce an admission target of 480—not 800. The lowest credible resource bound wins.

Memory reclamation is also workload-specific. DSec evaluates DAMON monitoring with virtio-balloon free-page reporting and reports a 21.2 percent reduction in memory consumption in that test. Its QoS configuration uses idle scheduling and core scheduling to limit interference; one evaluation reduced simultaneous-multithreading latency inflation from 45.2 to 17.3 percent. These results show mechanisms worth testing, not defaults to copy blindly. Reclamation that helps idle shells can damage latency-sensitive browsers or compilers.

Burst outside the private fleet deliberately

The paper describes cloud bursting after on-premises utilization exceeds 80 percent. Two hundred cloud VMs absorbed roughly 30 percent of a reported peak overflow. Only eligible images are synchronized into a deduplicated cloud set: 30 TB reportedly covered 70 percent of container tasks. The important control is eligibility. If an image, dataset, network route, or jurisdiction cannot leave the private fleet, the scheduler must know before the queue spikes.

Decouple the agent trajectory from the GPU lease

In reinforcement-learning and evaluation loops, an agent may be waiting on a build, browser, or external service while the accelerator process is preempted. Killing the sandbox wastes the work already performed; keeping the GPU reservation wastes scarce capacity. An execution plane can pause the caller, preserve the sandbox, and later attach a replacement worker.

on_gpu_preemption(run):
    freeze_new_tool_requests(run)
    wait_for_inflight_calls_or_timeout(run, 20s)
    snapshot = sandbox.checkpoint(incremental=true)
    persist({run_id, trajectory_offset, sandbox_id, snapshot_digest,
             image_digests, policy_version, last_receipt_id})
    release_gpu(run)

on_resume(run, new_worker):
    state = load_run_manifest(run.id)
    sandbox.restore(state.snapshot_digest)
    verify_digests_and_policy(state)
    replay_receipts_after(state.last_committed_receipt)
    attach(new_worker, sandbox)
    permit_new_tool_requests(run)

The order matters. Stop new calls, settle or mark in-flight calls, checkpoint, persist a manifest, then release the scarce worker. On resume, verify the exact environment before accepting new work. An idempotency key should bind each action to the trajectory offset so replay does not repeat a purchase, message, or mutation. Even evaluation-only workloads need this protection when tools modify repositories or services.

Compare this with DeltaBox, which targets fast reset and fine-grained copy-on-write storage for agent sandboxes, and Aries, which explores elastic sandbox infrastructure for agentic AI. The systems emphasize different bottlenecks, but together they point to the same operational split: environment state has a lifecycle that cannot be reduced to a container start command.

Failure modes an execution plane must surface

FailureHidden causeObservable signalControl
Placement stormStale global view overbooks popular nodesRising edge rejection and retriesIn-flight overlay, jitter, final edge admission
Cold-layer cascadeMany jobs first-touch the same large layerShared-store tail latency and queue growthPopularity-aware prefetch, admission backpressure
Ghost sandboxAPI loses owner after edge or network failureCompute exists without live leaseLease expiry, reconciliation, idempotent cleanup
Resume driftImage, policy, or tool version changesCheckpoint restores into a different contractDigest-bound manifest; refuse mismatched resume
Noisy-neighbor burstLow averages hide synchronized workp99 latency rises while averages look idlePer-class QoS, burst reserve, core isolation
Writable-overlay exhaustionBuild or browser writes exceed assumptionsDisk pressure and abrupt sandbox failuresQuota, early warning, spill or clean stop
Cloud policy leakOverflow route ignores data or image boundaryIneligible workload enters cloud queuePlacement-time eligibility and deny-by-default
Duplicate side effectResume replays an uncommitted tool callTwo mutations share one intentIdempotency key and receipt reconciliation

Security still matters. DSec describes identity and quotas, AppArmor, and task-specific eBPF network rules. Those controls complement rather than replace runtime isolation. For the threat model and containment choices themselves, use the separate AI agent sandbox guide. The execution-plane question is how those controls remain attached as workloads place, burst, pause, resume, and terminate.

A six-gate build or buy checklist

Measure the workloadRecord concurrency, arrival bursts, lifetime, CPU, RSS, processes, file descriptors, IOPS, network calls, layer access, and writable growth by task class.
Choose an isolation ladderMap function, container, microVM, and VM modes to trust, kernel, device, and performance requirements.
Define state ownersSeparate immutable layers, local writes, agent trajectory, receipts, and checkpoint manifests.
Make admission race-safeUse global candidate selection with a local final authority, bounded retries, and visible rejection reasons.
Test interruptionKill callers, edges, stores, and GPU workers; verify lease cleanup, idempotent resume, and exact-policy restoration.
Prove economicsCompare completion time, accelerator idle time, transfer bytes, disk writes, tail latency, failed resumes, and cost per accepted trajectory.

Start with one repeatable class rather than a universal platform. A code-repair evaluator with pinned images and deterministic tests is easier to characterize than a browsing agent with unrestricted network access. Add isolation modes and cloud overflow only when the measurements show a bottleneck. The post-training side should use the same evidence discipline described in the agent-model scaling guide: accepted outcomes and stable trajectories matter more than raw rollout count.

FAQ

Is DSec open source?

The paper describes the production platform, but it is not a complete open-source DSec distribution. Related components and ideas, including 3FS and storage work referenced by the paper, are public. Evaluate what is actually released rather than treating a paper as installable software.

Why not use Kubernetes jobs?

Kubernetes can be part of the implementation. The missing layer is the workload contract: high-rate admission, heterogeneous isolation, content distribution, tool APIs, stateful pause/resume, receipts, and a lifecycle independent of the GPU caller. A team can build those services on Kubernetes, but the orchestration object alone does not supply them.

When should I choose a microVM over a container?

Choose from the threat and compatibility model. MicroVMs add a kernel boundary and usually more overhead. Containers can be denser for trusted or heavily constrained work. Measure startup, memory, density, syscall needs, and escape consequences for each task class.

Does low average CPU justify extreme overcommit?

No. Use percentile demand, synchronized-burst tests, secondary resource limits, and a tested density ceiling. Average CPU is one input, not the admission rule.

What is the minimum useful checkpoint?

It binds writable state to immutable image digests, runtime and policy versions, trajectory offset, committed receipt, and caller identity. If those fields cannot be verified, resume should stop rather than silently drift.

Sources and further reading

Current facts and public artifacts were accessed and verified on September 29, 2026. DSec scale and performance figures above are explicitly presented as the authors' reported measurements.

Related guides