Autonomous execution loop
The issue-driven delivery loop used to take Claire from the backlog to v1.
A loop driver picks the next GitHub ticket, builds it, tests it end to end in a headless browser, opens a pull request, and either auto-merges it (low risk) or parks it for review (high risk) — then repeats. This page is the source of truth; the /claire-loop command points back here.
Where the tickets live#
All work is tracked as GitHub Issues. Milestones are epics (M0–M7), and M0 — test harness and green CI — lands first because it unblocks every end-to-end ticket after it.
| Label group | Values |
|---|---|
area/* | server, client, ai, loops, notifications, platforms, testing, infra, auth, db |
p0–p3 | Priority; p0 is critical |
type/* | feature, bug, chore, test, docs |
risk/* | auto-merge — the loop may merge on green CI; human-gate — review required |
state | ready, blocked, needs-review |
Find the next ticket — lowest milestone, then lowest priority number:
gh issue list --state open --label ready --json number,title,labels,milestone \
--jq 'map(.pri = (([.labels[].name]|map(select(test("^p[0-9]$")))|first) // "p9"))
| sort_by(.milestone.title, .pri)
| .[] | "#\(.number) [\(.milestone.title|split(" ")[0])] \(.pri) \(.title)"'Skip anything labelled blocked or needs-review, and anything whose body lists a Depends on: issue that is still open.
One iteration#
Sync
git fetch origin. Branches are cut fromorigin/main, which works in both a fresh clone and a multi-worktree checkout.Pick the next ready issue
M0 first; within a milestone, p0 through p3. Respect declared dependencies.
Claim it
git worktree add ../wt-<num> -b feat/<area>-<slug> origin/mainAnd comment
🔁 loop: startingon the issue so a second driver does not pick it up.Build to the acceptance criteria
Reuse existing code. Add or extend mock Playwright end-to-end tests and unit tests.
Gate locally
(cd server && bun run lint && bun run typecheck && bun test) (cd mobile && bun run lint && bun run typecheck && bun test && MOCK_BRIDGE=true bunx playwright test)Red means fix in-loop, at most twice. Still red: label
blocked, comment the failure, dropready, and move on.Open the pull request
gh pr create --base mainwithCloses #<num>, the risk tier, and how it was verified.Decide the merge
risk/auto-mergelands withgh pr merge --squash --autoonce CI is green.risk/human-gatestays open, losesready, and gainsneeds-review.Record and clean up
Tick the milestone in the tracking issue and append a line to
.context/loop-state.md, then remove the worktree.
Guardrails#
- Ordering. M0 lands first. Tickets touching shared files — CI config, the Playwright config, the server message handler, migrations — run serially.
- Isolation. Each worker gets its own git worktree and rebases on main before opening a pull request.
- Human gate respected. The loop never merges auth, database, infrastructure, or bridge-core changes. It parks them.
- Idempotent restart. Issue labels and the ledger are the source of truth, so the loop resumes cleanly after any interruption, on any machine.
Stop conditions#
- The run bound is reached, or a milestone is exhausted.
- No ready, unblocked, auto-mergeable tickets remain.
- Definition of done: core-loop end-to-end green for WhatsApp, Telegram, and Instagram.
- A ticket hard-fails twice — mark it blocked and continue.
Merge policy#
| Risk | Examples | Action |
|---|---|---|
risk/auto-merge | Tests, docs, isolated UI, client screens | Squash-merge automatically on green CI |
risk/human-gate | Auth, schema and migrations, infra and CI, bridge core, anything that sends messages autonomously | Open the PR, label needs-review, wait for a human |
CI must be green either way: lint, typecheck, Jest, the web build, and the mock Playwright suite.
Testing model#
- Mock backend for CI and the fast loop. The server runs with
MOCK_BRIDGE=trueagainst deterministic fixtures. No Docker, no real device pairing. - Real stack nightly. A scheduled workflow boots Supabase and Matrix, seeds a session, and runs a subset across all three platforms.
- Core-loop end-to-end is the definition of done: sign in → inbox shows seeded messages → open chat → AI suggestion appears and is accepted → send → Loops lists the detected loop → mark complete → reminder scheduled.
Watching a run#
claude -p runs headless, so there is no live TUI. Visibility comes from three places:
- Live dashboard.
scripts/loop-watch.shshows the active ticket, runner heartbeat, open pull requests, latest CI, queue counts, and the ledger. - Live agent trace. The runner streams formatted JSON into
.context/loop-runner.log, with raw JSONL in.context/loop-trace.jsonl. - Durable trail. Issue comments, branches, pull requests, and
.context/loop-state.md.
tail -f .context/loop-trace.jsonl | jq -Rrf scripts/loop-fmt.jq