# Loops: example scenario

Loop an agent until ticket-service's flaky tests pass three runs in a row, then triage new failures every night.

Source: https://ai-sw-factory.mellicci.dev/fundamentals/loops/example-scenario

## Scenario

In `ticket-service`, `pnpm test` fails about one run in four: some tests share database rows, others depend on timing. Each fix needs a full run to confirm, so a person spends an afternoon re-running the suite and guessing. New flaky tests then appear every week, and nobody notices until a pull request goes red.

This page shows one way to fix that: a goal loop, "make the flaky test suite green", and a nightly recurring loop that triages new failures.

## Before → after

| | Before | After |
|---|---|---|
| Time | A person prompts, waits and re-runs for hours | The loop runs rounds unattended, up to a cap |
| Quality | One green run counts as fixed | Three green runs in a row count as fixed |
| Risk | Each attempt forgets the previous one | `.agent/loop-notes.md` and commits record what was tried |
| Consistency | New flaky tests surface in someone's pull request | A nightly issue labeled `flaky` lists them |

## Design

**Diagram:** Each round the agent works, the check runs three times, and loop control continues or stops.

- Agent works — reads notes, fixes one cause
- →
- Check runs 3× — pnpm test, all must pass
- →
- Loop control — rounds, budget, result
- → stop
- Stop — success, cap or human

**Diagram:** The nightly loop is independent each run, so its state lives in one issue.

- Nightly run — triage failures of the last day
- → opens or updates
- Issue "flaky" — one issue, all evidence

```text
# GENERIC loop definition; every agent expresses it differently
goal:        make the flaky test suite green
check:       pnpm test, run 3 times
success:     all 3 runs pass (exit code 0)
max_rounds:  10
budget:      a token or money cap, plus a time limit
notes:       .agent/loop-notes.md (tried, failed), plus one commit per round
stop:        success, max rounds, budget or time cap, or a human stop
```

## What happens at runtime

| Round | What failed | What the notes say | Decision |
|---|---|---|---|
| 1 | `tickets.test.ts` fails 1 of 3 runs (shared rows) | Empty | Continue; note "wrapped tests in transactions" |
| 2 | `sla.test.ts` times out (real clock) | Transactions done | Continue; note "fake timers in `sla.test.ts`" |
| 3 | Run 2 of 3 fails on a leftover row | Both fixes listed, one failure open | Continue; note "reset sequences" |
| … | … | Every attempt, with result | Continue |
| n | Nothing | Full history | Stop on success, or on round 10, a budget cap or a human stop |

The nightly loop has no rounds. Each run lists tests that failed in the last day and opens or updates the one `flaky` issue. Tomorrow's run starts fresh and reads that issue.

## What can go wrong

| Failure | How you notice | What to do |
|---|---|---|
| The check is gamed: the agent deletes or skips the flaky test | The suite is green but the test count drops, or `.skip` appears in the diff | Protect test files with a [hook](https://ai-sw-factory.mellicci.dev/fundamentals/hooks) that blocks edits removing tests; review the diff |
| No progress memory: the same fix is retried | Notes repeat an attempt, or rounds show identical failures | Require each round to write what was tried and the result, and to commit |
| Runaway cost: no caps | Spend grows per round with no end in sight | Set max rounds, a budget and a time limit before the first run |
| The nightly run has write access, unattended | It pushes or edits when it should only file an issue | Give it read-only code access plus issue write; see the [security model](https://ai-sw-factory.mellicci.dev/fundamentals/security-model) and [sandboxing](https://ai-sw-factory.mellicci.dev/fundamentals/sandboxing) |
| The notes file grows and costs context | Each round starts slower and cites old details | Keep entries to one line per attempt; summarize old rounds |
