AI Agent Sandbox: Restrict Code, Files and Network Access
Bind each execution to one task
Limit code, files and network in the worker
Original SANDBOX-6 contract and CRM example
Isolation failures and production tests

$ sandbox --contract SANDBOX-6
> bind: actor / task / input
> isolate: code / files / network
> limit: time / memory / output
> route: verified / deny / reconcileAn AI agent that executes code can reach processes, files and network endpoints. An AI agent sandbox bounds the effect of a mistaken or hostile model proposal through process isolation, explicit resources, expiry and a verifiable output. This is an engineering guide, not a universal security guarantee. For the broader decision about an AI specialist in Armenia, use the dedicated service page.
The original SANDBOX-6 design below uses a synthetic task: summarize an uploaded CSV and return only a JSON report. It is not a deployed customer claim or a measured benchmark. Related guides cover MCP server security and tool schemas.
The problem and requirements
A prompt saying “do not read secrets” cannot stop a process from reading an available file. The rule must be enforced outside the model. Before execution, bind the actor, task, authorized inputs, allowed output and expiry. First ask whether arbitrary code is necessary: a typed transformation tool may be enough for a narrow task.
Bound three paths of effect. Code should have no host privileges. Files should include only read-only inputs and a separate output directory. Network access should be denied by default or restricted to explicit destinations. CPU, memory, disk, process and time limits protect availability. Output and logs also need validation and retention rules because they may contain sensitive data.
SANDBOX-6 architecture
| Boundary | Contract | Evidence |
|---|---|---|
| 1. Task | request ID, actor and purpose | authorization before launch |
| 2. Image | pinned minimal image | digest and dependency review |
| 3. Code | unprivileged process, no host socket | process and syscall profile |
| 4. Files | read-only input, separate temporary output | mount manifest and path checks |
| 5. Network | deny by default, explicit gateway routes | egress policy and connection log |
| 6. Result | limits, artifact validation and cleanup | receipt, size, type and TTL |
An orchestrator accepts a typed request. Policy binds actor, tenant and source object, then produces an execution manifest. The worker sees a copy of the permitted CSV, a work directory and one output path. Access to a CRM or MCP tool remains a separately authorized call: process isolation does not replace target-system permissions.
Synthetic manifest
request_id: example-095
image_digest: sha256:reviewed-image-placeholder
user: nonroot
inputs:
- source: authorized-report.csv
mount: /input/report.csv
mode: read-only
output: /output/summary.json
network: deny
limits:
seconds: 30
memory_mb: 256
processes: 16
output_mb: 5This is a design example, not a ready configuration for a particular runtime. Tune limits from observed workloads. The model must not provide an arbitrary host path. The orchestrator resolves a permitted input ID, creates a temporary directory and records actual mounts.
Components and lifecycle
Policy checks document access and whether code execution is needed. A broker issues a short-lived job identifier without giving the worker a durable token. The worker uses a fixed image, unprivileged user and read-only root. Inputs are mounted read-only and output goes to a separate bounded directory. If the task needs network access, an egress gateway checks destinations, DNS resolution and redirects.
After completion, the orchestrator checks exit code, duration, artifact size and type. JSON must pass a schema; links and text inside it remain untrusted data. Only verified output is returned. On timeout, stop the worker, clean temporary files and reconcile any uncertain external effect before retrying. AI automation covers the surrounding workflow; prompt engineering helps define narrow tasks and tools.
// Illustrative orchestrator pseudocode, not a runtime implementation.
const job = policy.authorize(actor, tenant, inputId, "summarize_csv");
const manifest = broker.createManifest(job, { network: "deny", output: "summary.json" });
const run = await sandbox.executeOnce(manifest, requestId);
if (run.timedOut || run.externalEffectUnknown) return reconcile(requestId);
if (!artifactSchema.safeParse(run.output).success) return reject("invalid_output");
return publishVerifiedArtifact(run.output);Failure modes and recovery
| Scenario | Control | Route |
|---|---|---|
| Model names a host secret path | broker accepts permitted input ID, not path | deny |
| Symlink escapes the input directory | copy with resolved-path check and mount manifest | deny |
| Code attempts network access | egress deny and logging | stop and investigate |
| Dependency spawns child processes | pinned image, unprivileged user and process limit | stop |
| Output fills disk | quota and artifact size cap | reject output |
| Timeout follows an external call | request ID and destination receipt | reconcile before retry |
A container alone is not proof of isolation. A mounted Docker socket, privileged mode, shared writable volume or reachable metadata endpoint changes the real boundary. Sensitive tasks may need stronger isolation, validated on the actual infrastructure.
Tests and production checklist
Build fixtures for a normal CSV, unrelated file read, path traversal, symlink, network request, infinite loop, process fork, oversized output and timeout. For each, assert the expected denial or valid artifact and no unintended effect. Test mounts and egress in the production-like worker environment, beyond policy mocks.
Before production, assign a policy owner, pin the image digest, define dependency updates, resource limits, input and log retention, stop procedures and reconciliation. Test input and output permissions across tenants. Monitoring should show denial reasons and unknown outcomes without leaking secrets. Request an architecture review before adding tools or network privileges to the agent.
const job = policy.authorize(actor, tenant, inputId);
const run = await sandbox.executeOnce(broker.manifest(job), requestId);
if (run.timedOut || run.externalEffectUnknown) return reconcile(requestId);
if (!artifactSchema.safeParse(run.output).success) return reject;
return publishVerifiedArtifact(run.output);