# OPERATOR-RUNBOOK — R720 agent-team coordinator incident handling The on-call runbook for the always-on `agent-team-coordinator` daemon on the `sh-secrev` R720 VM. Covers pipeline stalls, stuck/parked tasks, failed human-in-the-loop resumes, budget exhaustion mid-pipeline, transport outages, and the COMPLACENCY / COVERAGE alarms. Grounds every recovery in real code (design §5 escalation ladder + §6.6 contention/park policy; Phase-6 requirement). > **CLI used throughout: `run-team.py`** (the entry CLI), run from > `~/orchestrator/agent-team` with the venv active so it hits the default ledger > (`state/agent_team.sqlite`) and audit log (`state/audit.log.jsonl`). Do **not** > use `agent_team/operator_cli.py` — it has divergent verbs (no `show`, required > `--db`/`--audit-log`, and a `force-resume` that supersedes). See DEPLOY-AUDIT.md. ```bash ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120 cd ~/orchestrator/agent-team && . .venv/bin/activate ``` ## Verb reference (all confirmed in `run-team.py:build_parser`) | Verb | Effect | Destructive? | |---|---|---| | `list` / `list --all` / `list --parked` / `list --status ` | read pending questions (default `open`) | no | | `show ` | print one ledger row (JSON) | no | | `redeliver ` | clear `channel_ref` so the reconcile loop re-posts an `open` question | no (audit-logged) | | `expire --confirm` | force `open`→`expired` | yes | | `answer --answer

[--via ] --confirm` | answer-on-behalf (first-answer-wins CAS) | yes | | `force-resume --confirm` | reopen an `expired` (parked) question; for `answered` records resume intent | yes | | `supersede --confirm` | mark a stale `open`/`answered` row `superseded` | yes | | `start --task "..." [--transport ...] [--dry-run]` | start one task to the human gate | no | `--operator ` (global) sets the audit attribution; it defaults to the OS login. Destructive verbs require `--confirm` and write an attempt-then-outcome record to the audit log **before** mutating. First triage for any incident: ```bash systemctl status agent-team-coordinator.service journalctl -u agent-team-coordinator.service -e --since "-2h" | tail -100 python3 run-team.py list --all # full ledger snapshot python3 run-team.py list --parked # non-open rows (parked-task context) ``` --- ## Incident 1 — Pipeline stall (tasks not advancing) **Symptoms:** `list` shows `open` (or `answered`) rows that never progress; the journal shows no `tick` activity or repeated errors. **Diagnose:** ```bash systemctl is-active agent-team-coordinator.service # "active" expected journalctl -u agent-team-coordinator.service -e | tail -60 # Daemon dead/looping on restart? Check the unit + deps: .venv/bin/python -c "import langgraph, slack_sdk, slack_bolt, requests; print('deps ok')" ``` **Recover:** 1. If the daemon is dead and `Restart=on-failure` is flapping, read the journal for the import/auth error. A missing venv dep (D-4/D-5) or a missing `CLAUDE_CODE_OAUTH_TOKEN` is the usual cause — fix the env / venv, then: ```bash sudo systemctl restart agent-team-coordinator.service ``` 2. The startup `recover()` sweep re-drives `answered`-but-unresumed rows and re-posts `open` rows that lost their `channel_ref`, so a clean restart converges from the durable ledger. Confirm with `journalctl ... | tail` and `list --all`. 3. A single task stuck `open` with a stale/missing post: re-deliver it. ```bash python3 run-team.py show python3 run-team.py redeliver # clears channel_ref; reconcile re-posts ``` **Escalate** (see the ladder) if a restart does not clear it within one tick cadence and the journal shows a non-transient error. --- ## Incident 2 — Stuck / parked task A task parks (design §6.6) when its clarifier question **expires** with no answer, or its budget headroom drops below reserve, or a checkpoint is corrupt. Parked = the durable ledger row is no longer `open` (it is `expired`), and an ALARM was raised, not spun on. **Diagnose:** ```bash python3 run-team.py list --parked python3 run-team.py show # status, thread_id, deadline_at, answered_via ``` **Recover — depends on why it parked:** - **Expired with no answer (the common case)** — un-park by reopening the expired question; it is then re-delivered for an answer: ```bash python3 run-team.py force-resume --confirm # "force-resume: reopened expired question ; it will be re-delivered" python3 run-team.py show # status flips back to "open" ``` Then answer it (Slack or CLI) to drive it forward. - **Answered but not yet resumed** — the recovery sweep handles it; `force-resume` records intent and reports that (no mutation): ```bash python3 run-team.py force-resume --confirm # "force-resume: question is answered and pending resume; the recovery # sweep will resume it (intent recorded)" sudo systemctl restart agent-team-coordinator.service # forces the recover() sweep now ``` - **Stale / wrong question that should be abandoned** — supersede it so it stops surfacing as parked context: ```bash python3 run-team.py supersede --confirm ``` `MAX_PARK` FIFO-aging: a task that exceeds the park window escalates (ALARM + a Jira ticket per the ladder) rather than starving silently. --- ## Incident 3 — Failed human-in-the-loop resume **Symptoms:** an answer was submitted (Slack or CLI) but the graph did not advance. **Diagnose:** ```bash python3 run-team.py show # is status "answered"? what answered_via? journalctl -u agent-team-coordinator.service -e | grep -iE "resume|answer|" ``` **Recover:** - **Status is `answered` but no resume drained** — the resume queue is in-process; a daemon restart triggers the `recover()` sweep that re-drives `answered` rows via the turn-guarded `ResumeWorker` (idempotent — an already-advanced thread supersedes-and-skips): ```bash sudo systemctl restart agent-team-coordinator.service python3 run-team.py show # confirm it advanced ``` - **Live Slack answer never registered** — the listener fails closed. Check: ```bash journalctl -u agent-team-coordinator.service -e | grep -i "Slack answer" # "rejecting Slack answer: owner allowlist is unconfigured ..." -> set # AGENT_TEAM_SLACK_OWNER_IDS in ~/secrev.env and restart. # "rejecting Slack answer ... unauthorized sender" -> the answerer's user id is # not in the allowlist. Add it, or answer via the CLI on their behalf: python3 run-team.py answer --answer "" --via "cli:adam" --confirm ``` - **Answer lost the compare-and-set (`not open`)** — the row was already answered/expired/superseded. Inspect with `show`; if it parked, go to Incident 2. --- ## Incident 4 — Budget exhaustion mid-pipeline The shared Claude budget ledger (`budget_ledger` table) enforces per-call + total nightly caps (design §6.1, §6.6). A stage that would breach the reserve **parks** the task (deferred, not dropped) and ALARMs — it never loop-drains the pool. **Diagnose:** ```bash journalctl -u agent-team-coordinator.service -e | grep -iE "budget|reserve|park" # Inspect today's spend directly (no CLI verb for the budget ledger; read it). # Columns (agent_team/db/schema.py budget_ledger DDL): thread_id, stage, model, # billing_mode, input_tokens, output_tokens, usd_cost, recorded_at, day_bucket. sqlite3 state/agent_team.sqlite \ "SELECT day_bucket, thread_id, ROUND(SUM(usd_cost),4) AS usd, SUM(input_tokens) AS in_tok FROM budget_ledger GROUP BY day_bucket, thread_id ORDER BY day_bucket DESC LIMIT 20;" ``` **Recover:** - **Wait for the next budget window** — budget-exhausted roles/tasks are deferred via the rotation pointer and picked up next cycle; this is the intended behavior, not a failure. The parked task surfaces in `list --parked`. - **Force a specific parked task forward now** (e.g. it is urgent and headroom has since freed): un-park it and let the daemon re-run the stage within the remaining cap: ```bash python3 run-team.py force-resume --confirm ``` - **Confirm `ANTHROPIC_API_KEY` is absent** — its presence would silently meter to API rates and blow the budget model (the billing seam pops it defensively, but it must not be set): ```bash grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # both must print 0 ``` Do **not** raise the cap to push a task through without Adam's decision — budget realism is a deliberate guardrail. Escalate per the ladder if a task repeatedly parks on budget across nights. --- ## Incident 5 — Transport outage (Slack / GitHub down or misconfigured) **Symptoms:** clarifier posts fail; the journal shows transport errors or the inbound listener is not started. **Diagnose:** ```bash journalctl -u agent-team-coordinator.service -e | grep -iE "Slack|listener|post|transport" # Did the inbound listener start? journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener" # "started (Socket Mode ...)" -> inbound up # "not started (... SLACK_APP_TOKEN is unset)" -> Slack inbound intentionally off ``` **Recover:** - **Outbound post failed (Slack API / network)** — the ledger row stays `open` with no `channel_ref` (the responder leaves it for reconcile). After the transport recovers, the `recover()` sweep (on restart) or a manual `redeliver` re-posts: ```bash python3 run-team.py redeliver ``` - **Inbound listener down / never started** — the daemon still posts and expires; it just cannot hear Slack. **The CLI `answer` path is the outage fallback** — it runs the identical compare-and-set with no socket: ```bash python3 run-team.py answer --answer "" --via "cli:adam" --confirm ``` Fix `SLACK_APP_TOKEN` / `SLACK_BOT_TOKEN` in `~/secrev.env`, then `sudo systemctl restart agent-team-coordinator.service` to re-establish Socket Mode. (D-1: the listener starts only when the transport is live Slack AND `SLACK_APP_TOKEN` is set; it is optional by design.) - **Slack listener crash-looping** — the listener runs on an isolated daemon thread; a crash is logged (`inbound Slack listener thread exited with an error`) and takes down only the inbound socket, not the maintenance loop. Restart the service to relaunch the thread once the cause is fixed. --- ## Incident 5b — WS0–WS5 surfaces (HTTP API, /new-task, handbook context) These were added by the WS-rollout (see `agent-team/DEPLOY-R720.md` §4b). They layer onto the coordinator; none of them should take down the maintenance loop. - **HTTP API down / unreachable** — the WS1 FastAPI app (`agent_team/api.py`) is a **separate, opt-in process** (`api.serve()`, `127.0.0.1:8765`, bearer auth), **not** started by the coordinator daemon. If `/delegate` from Claude Code or `POST /tasks` over HTTP stops working, the coordinator itself is unaffected — check the API process separately: ```bash curl -sS -o /dev/null -w '%{http_code}\n' \ -H "Authorization: Bearer $AGENT_TEAM_API_TOKEN" http://127.0.0.1:8765/tasks # 405 = API up + authed (GET not allowed on /tasks); 000 = process down; # 401 = AGENT_TEAM_API_TOKEN mismatch (client vs ~/secrev.env). ``` The API refuses to start if `AGENT_TEAM_API_TOKEN` is unset/empty (logs a `RuntimeError`). Fix the token, restart the API process. Tasks already in the ledger are unaffected — the API is only an *intake/invoke* front door; answer via Slack or the CLI as usual. - **`/new-task` Slack command not responding** — the WS2 slash command is AUTHZ-01 owner-allowlist gated and routes through the same Socket Mode listener as answers. If it silently does nothing, it is almost always the owner allowlist (same failure mode as Incident 3's live-Slack path): ```bash journalctl -u agent-team-coordinator.service -e | grep -iE "new-task|owner|unauthorized" # unauthorized sender / unconfigured allowlist -> fix AGENT_TEAM_SLACK_OWNER_IDS ``` Fallback: start the task from the CLI (`run-team.py start --task "..."`) or the HTTP API. If the listener itself is down, see Incident 5 (inbound listener). - **Handbook dir missing → planner runs without handbook context** — the WS5 `context_provider` (`load_handbook_conventions`) is **fail-safe**: if `SEA_HAVEN_HANDBOOK_DIR` (or `~/.sea-haven/engineering-handbook`) is missing or unreadable it returns `""` and the planner runs normally, just without handbook conventions injected. This is **degraded, not broken** — no park, no alarm. Confirm and restore: ```bash grep '^SEA_HAVEN_HANDBOOK_DIR=' ~/secrev.env ls "$(grep '^SEA_HAVEN_HANDBOOK_DIR=' ~/secrev.env | cut -d= -f2)" # dir present + populated? ``` Re-sync the handbook (the deploy script does this) and restart the daemon so the planner picks it back up. The WS3 dispatch node is **inert** (gated) and should never appear in pipeline activity; if it does, treat as an unexpected state and escalate. --- ## Incident 6 — COMPLACENCY / COVERAGE alarms (Plane-1 checkers) These come from the nightly checker run, not the coordinator daemon (design §6.4, §6.6): - **COMPLACENCY ALARM** — a checker role missed a planted canary fault. That role is **skipped** for the night (it never runs silently degraded). Recover: inspect the canary corpus / the role's checker under `security-review/checkers/`, fix the regression, re-run that checker's dry-run, and confirm the canary passes before re-enabling. - **COVERAGE ALARM** — a role slipped its rotation slot (e.g. budget-deferred). It is **deferred via the rotation pointer, never dropped**, and picked up next cycle. Recover: confirm the rotation pointer advanced (it is rebuildable from report history) and that the role runs on the next cadence; investigate only if it slips repeatedly. Both alarms are **report + ALARM-only** (design D3): nothing posts on a clean state, no auto-Jira/Notion writes from the checker itself. Persistent alarms follow the escalation ladder below. --- ## Escalation ladder (design §5 / §6.6, resolves Q3) Anything that does not clear on the first ALARM escalates — but **ALARM-only in spirit** (nothing posts on a clean state): 1. **Re-alarm on a backoff.** A confirmed critical (or a COMPLACENCY / COVERAGE alarm) that persists re-alarms to Slack each night it is still unresolved, on a backoff so it does not spam. 2. **Open a Jira tracking ticket after `N` nights** (default **N = 3**). If the condition still has not cleared, the coordinator opens an **INFRA** Jira ticket so it cannot quietly linger. The same ladder applies to a parked task that exceeds `MAX_PARK`. 3. **Human (Adam) takes it from the Jira ticket.** For a parked task, recover via Incidents 2–4 above; for a checker alarm, via Incident 6. When you resolve an incident, record the action — destructive CLI verbs already write an attributable attempt+outcome record to `state/audit.log.jsonl`; for non-CLI recoveries note it on the Jira ticket. --- ## Post-incident - Confirm the daemon is `active` and the ledger has no unexpected parked rows (`list --parked`). - If you restored the VM snapshot or wiped the ledger, re-run the relevant PROVISIONING-RUNBOOK steps. - Update `project_r720_agent_team` memory if the incident revealed a durable fact (a new failure mode, a config that must change).