# Security model

What an agent is allowed to do is set by permissions, sandboxing, approvals, secrets and network access, enforced by the harness rather than the model.

> An agent's security model is the set of limits the harness and operating system enforce on what it may do. They cover which tools it can call, what it can reach, and who must approve.

Source: https://ai-sw-factory.mellicci.dev/fundamentals/security-model

**Diagram:** The security model limits the actions; the sandbox around it limits the damage.

- Agent action — tool call or command
- → checked by
- Policy · enforced by harness (secrets, network rules):
  - Permission rules — allow · ask · deny
  - → ask
  - Approval — a person confirms
- → allowed
- Sandbox:
  - Action runs — files · shell · network

The model cannot be relied on to refuse a bad action, because its behavior is probabilistic and it can be steered by the text it reads. So limits must live where the harness and the operating system enforce them, and every action passes them on its way out, whatever untrusted text is in the context. Three controls set what the agent may do, and a fourth limits the damage when they fail.

**Permissions** decide which tools and commands the harness will run, ask about, or refuse. **Approvals** decide who confirms an action. **Secrets and network** rules decide which credentials and hosts are reachable.

The fourth is the sandbox, covered in [Sandboxing & devcontainers](https://ai-sw-factory.mellicci.dev/fundamentals/sandboxing): the sandbox limits the damage, the security model limits the actions.

## Why it exists

An agent reads issue text, web pages and tool output, and any of these can contain instructions aimed at the model. This is *prompt injection*. If the agent holds broad credentials and unrestricted network access, one crafted issue could make it read a token and send it elsewhere. Asking the model to "be careful" does not close that path.

## How it works

**Diagram:** Rules decide each action, a person or pre-set policy answers for "ask", and the sandbox, scoped credentials and logs cover what rules miss.

- Permission rules — match each tool call
- → allow · ask · deny
- Who answers "ask":
  - Interactive — a person decides
  - Headless — nobody; deny or pre-allow
- → allowed
- What rules miss:
  - Sandbox — files and network, regardless of rules
  - Scoped credentials — only what the task needs
  - Logs — what ran, for review

## Use it when

- Any agent can run commands or reach the network. This applies always; it matters most for unattended runs.

## Use something else when

- You want to enforce a project-specific rule, such as blocking a command → [Hooks](https://ai-sw-factory.mellicci.dev/fundamentals/hooks)
- You want to state a convention → [Instruction files](https://ai-sw-factory.mellicci.dev/fundamentals/instruction-files)
- You want to limit a role to reading → [Subagents](https://ai-sw-factory.mellicci.dev/fundamentals/subagents)
- You want to choose where the agent runs and what it can reach → [Sandboxing & devcontainers](https://ai-sw-factory.mellicci.dev/fundamentals/sandboxing)
- You are running without a person → [Headless execution](https://ai-sw-factory.mellicci.dev/fundamentals/headless-execution)

<note>

Hooks and permissions are complementary. A hook enforces your project rules, and a sandbox limits damage when a rule is missing.

</note>

## Key terms

- **Permission** — allow, ask or deny for a tool call.
- **Sandbox** — OS-level limits on files and network; see [Sandboxing & devcontainers](https://ai-sw-factory.mellicci.dev/fundamentals/sandboxing).
- **Approval** — confirmation before an action runs.
- **Prompt injection** — untrusted text steering the model.
- **Trust boundary** — where control ends.
