# Sandboxing & devcontainers

Sandboxing decides where the agent runs and what it can reach if something goes wrong, which sets the blast radius of a mistake or a manipulated agent.

> Sandboxing limits where an agent runs and what it can read, write and reach on the network. A mistake or a manipulated agent can then only damage a disposable environment.

Source: https://ai-sw-factory.mellicci.dev/fundamentals/sandboxing

**Diagram:** Sandboxing nests boundaries around the agent's commands, so the home directory, keys and open internet stay outside.

- Host · laptop or CI (home dir, SSH keys, other repos):
  - Devcontainer (repo only, scoped token):
    - Agent:
      - Model + harness
      - → runs commands
      - Built-in sandbox:
        - Commands — test · git · build
- → allowlist only
- Allowed hosts — registry · git host · model API

Sandboxing answers one question: if this run goes wrong, what can it touch? A run goes wrong when the model makes a mistake, or when text it reads steers it into an action you did not intend. You cannot prevent both, so you limit where they can do harm.

There are two layers. The **built-in sandbox** is a feature of the agent: the harness starts the commands the agent runs under operating-system limits. The **environment** is what you put around the whole agent: a devcontainer, container, virtual machine or cloud sandbox. The built-in sandbox usually covers only the commands; the environment also covers file tools, MCP servers and hooks. Sandboxing is not strictly a coding-agent feature. It is the foundation for running agents unattended.

## Why it exists

By default, an agent started in your terminal runs as you. Its commands can read `~/.ssh`, your cloud credentials and every other repository on the disk, and they can reach any host your machine can. One bad command or one injected instruction is then a problem for your whole account, not for one repository.

## How it works

**Diagram:** Each control closes one path out of the environment, and rebuilding from a file makes any damage disposable.

- Definition file — describes the environment
- → rebuilds
- Sandbox:
  - Filesystem — writes to project and temp only
  - Network — host allowlist via proxy or firewall
  - Mounts — repo only, no home or Docker socket
  - Credentials — task-scoped token as env variable
- → after the run
- Container discarded — damage disappears with it

<warning>

Any path that stays open can leak what the agent can read. An allowed host can receive data, and a writable mount can be modified. Sandboxes reduce the blast radius; they do not remove it.

</warning>

## Use it when

- You want fewer approval prompts, because the boundary catches mistakes the prompts would have caught.
- The agent runs unattended, in CI or on a schedule.
- You work in a repository you have not fully read.

## Use something else when

- You need to decide which commands and tools the agent may use, and who approves them → [Security model](https://ai-sw-factory.mellicci.dev/fundamentals/security-model)
- A project rule must block a specific command every time → [Hooks](https://ai-sw-factory.mellicci.dev/fundamentals/hooks)
- You need to start the agent from a script and read its result → [Headless execution](https://ai-sw-factory.mellicci.dev/fundamentals/headless-execution)

## Key terms

- **Blast radius** — what a failed run could damage or leak.
- **Built-in sandbox** — OS-level limits on the agent's commands.
- **Devcontainer** — a repository-defined container environment.
- **Egress allowlist** — outbound traffic limited to named hosts.
- **Scoped credential** — a token limited to one task.
