# OPERATOR-RUNBOOK — R720 agent-team coordinator incident handling The on-call runbook for the always-on `agent-team-coordinator` daemon on the `sh-secrev` R720 VM. Covers pipeline stalls, stuck/parked tasks, failed human-in-the-loop resumes, budget exhaustion mid-pipeline, transport outages, and the COMPLACENCY / COVERAGE alarms. Grounds every recovery in real code (design §5 escalation ladder + §6.6 contention/park policy; Phase-6 requirement). > **CLI used throughout: `run-team.py`** (the entry CLI), run from > `~/orchestrator/agent-team` with the venv active so it hits the default ledger > (`state/agent_team.sqlite`) and audit log (`state/audit.log.jsonl`). Do **not** > use `agent_team/operator_cli.py` — it has divergent verbs (no `show`, required > `--db`/`--audit-log`, and a `force-resume` that supersedes). See DEPLOY-AUDIT.md. ```bash ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120 cd ~/orchestrator/agent-team && . .venv/bin/activate ``` ## Verb reference (all confirmed in `run-team.py:build_parser`) | Verb | Effect | Destructive? | |---|---|---| | `list` / `list --all` / `list --parked` / `list --status ` | read pending questions (default `open`) | no | | `show ` | print one ledger row (JSON) | no | | `redeliver ` | clear `channel_ref` so the reconcile loop re-posts an `open` question | no (audit-logged) | | `expire --confirm` | force `open`→`expired` | yes | | `answer --answer

[--via ] --confirm` | answer-on-behalf (first-answer-wins CAS) | yes | | `force-resume --confirm` | reopen an `expired` (parked) question; for `answered` records resume intent | yes | | `supersede --confirm` | mark a stale `open`/`answered` row `superseded` | yes | | `start --task "..." [--transport ...] [--dry-run]` | start one task to the human gate | no | `--operator ` (global) sets the audit attribution; it defaults to the OS login. Destructive verbs require `--confirm` and write an attempt-then-outcome record to the audit log **before** mutating. First triage for any incident: ```bash systemctl status agent-team-coordinator.service journalctl -u agent-team-coordinator.service -e --since "-2h" | tail -100 python3 run-team.py list --all # full ledger snapshot python3 run-team.py list --parked # non-open rows (parked-task context) ``` --- ## Incident 1 — Pipeline stall (tasks not advancing) **Symptoms:** `list` shows `open` (or `answered`) rows that never progress; the journal shows no `tick` activity or repeated errors. **Diagnose:** ```bash systemctl is-active agent-team-coordinator.service # "active" expected journalctl -u agent-team-coordinator.service -e | tail -60 # Daemon dead/looping on restart? Check the unit + deps: .venv/bin/python -c "import langgraph, slack_sdk, slack_bolt, requests; print('deps ok')" ``` **Recover:** 1. If the daemon is dead and `Restart=on-failure` is flapping, read the journal for the import/auth error. A missing venv dep (D-4/D-5) or a missing `CLAUDE_CODE_OAUTH_TOKEN` is the usual cause — fix the env / venv, then: ```bash sudo systemctl restart agent-team-coordinator.service ``` 2. The startup `recover()` sweep re-drives `answered`-but-unresumed rows and re-posts `open` rows that lost their `channel_ref`, so a clean restart converges from the durable ledger. Confirm with `journalctl ... | tail` and `list --all`. 3. A single task stuck `open` with a stale/missing post: re-deliver it. ```bash python3 run-team.py show python3 run-team.py redeliver # clears channel_ref; reconcile re-posts ``` **Escalate** (see the ladder) if a restart does not clear it within one tick cadence and the journal shows a non-transient error. --- ## Incident 2 — Stuck / parked task A task parks (design §6.6) when its clarifier question **expires** with no answer, or its budget headroom drops below reserve, or a checkpoint is corrupt. Parked = the durable ledger row is no longer `open` (it is `expired`), and an ALARM was raised, not spun on. **Diagnose:** ```bash python3 run-team.py list --parked python3 run-team.py show # status, thread_id, deadline_at, answered_via ``` **Recover — depends on why it parked:** - **Expired with no answer (the common case)** — un-park by reopening the expired question; it is then re-delivered for an answer: ```bash python3 run-team.py force-resume --confirm # "force-resume: reopened expired question ; it will be re-delivered" python3 run-team.py show # status flips back to "open" ``` Then answer it (Slack or CLI) to drive it forward. - **Answered but not yet resumed** — the recovery sweep handles it; `force-resume` records intent and reports that (no mutation): ```bash python3 run-team.py force-resume --confirm # "force-resume: question is answered and pending resume; the recovery # sweep will resume it (intent recorded)" sudo systemctl restart agent-team-coordinator.service # forces the recover() sweep now ``` - **Stale / wrong question that should be abandoned** — supersede it so it stops surfacing as parked context: ```bash python3 run-team.py supersede --confirm ``` `MAX_PARK` FIFO-aging: a task that exceeds the park window escalates (ALARM + a Jira ticket per the ladder) rather than starving silently. --- ## Incident 3 — Failed human-in-the-loop resume **Symptoms:** an answer was submitted (Slack or CLI) but the graph did not advance. **Diagnose:** ```bash python3 run-team.py show # is status "answered"? what answered_via? journalctl -u agent-team-coordinator.service -e | grep -iE "resume|answer|" ``` **Recover:** - **Status is `answered` but no resume drained** — the resume queue is in-process; a daemon restart triggers the `recover()` sweep that re-drives `answered` rows via the turn-guarded `ResumeWorker` (idempotent — an already-advanced thread supersedes-and-skips): ```bash sudo systemctl restart agent-team-coordinator.service python3 run-team.py show # confirm it advanced ``` - **Live Slack answer never registered** — the listener fails closed. Check: ```bash journalctl -u agent-team-coordinator.service -e | grep -i "Slack answer" # "rejecting Slack answer: owner allowlist is unconfigured ..." -> set # AGENT_TEAM_SLACK_OWNER_IDS in ~/secrev.env and restart. # "rejecting Slack answer ... unauthorized sender" -> the answerer's user id is # not in the allowlist. Add it, or answer via the CLI on their behalf: python3 run-team.py answer --answer "" --via "cli:adam" --confirm ``` - **Answer lost the compare-and-set (`not open`)** — the row was already answered/expired/superseded. Inspect with `show`; if it parked, go to Incident 2. --- ## Incident 4 — Budget exhaustion mid-pipeline The shared Claude budget ledger (`budget_ledger` table) enforces per-call + total nightly caps (design §6.1, §6.6). A stage that would breach the reserve **parks** the task (deferred, not dropped) and ALARMs — it never loop-drains the pool. **Diagnose:** ```bash journalctl -u agent-team-coordinator.service -e | grep -iE "budget|reserve|park" # Inspect today's spend directly (no CLI verb for the budget ledger; read it). # Columns (agent_team/db/schema.py budget_ledger DDL): thread_id, stage, model, # billing_mode, input_tokens, output_tokens, usd_cost, recorded_at, day_bucket. sqlite3 state/agent_team.sqlite \ "SELECT day_bucket, thread_id, ROUND(SUM(usd_cost),4) AS usd, SUM(input_tokens) AS in_tok FROM budget_ledger GROUP BY day_bucket, thread_id ORDER BY day_bucket DESC LIMIT 20;" ``` **Recover:** - **Wait for the next budget window** — budget-exhausted roles/tasks are deferred via the rotation pointer and picked up next cycle; this is the intended behavior, not a failure. The parked task surfaces in `list --parked`. - **Force a specific parked task forward now** (e.g. it is urgent and headroom has since freed): un-park it and let the daemon re-run the stage within the remaining cap: ```bash python3 run-team.py force-resume --confirm ``` - **Confirm `ANTHROPIC_API_KEY` is absent** — its presence would silently meter to API rates and blow the budget model (the billing seam pops it defensively, but it must not be set): ```bash grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # both must print 0 ``` Do **not** raise the cap to push a task through without Adam's decision — budget realism is a deliberate guardrail. Escalate per the ladder if a task repeatedly parks on budget across nights. --- ## Incident 5 — Transport outage (Slack / GitHub down or misconfigured) **Symptoms:** clarifier posts fail; the journal shows transport errors or the inbound listener is not started. **Diagnose:** ```bash journalctl -u agent-team-coordinator.service -e | grep -iE "Slack|listener|post|transport" # Did the inbound listener start? journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener" # "started (Socket Mode ...)" -> inbound up # "not started (... SLACK_APP_TOKEN is unset)" -> Slack inbound intentionally off ``` **Recover:** - **Outbound post failed (Slack API / network)** — the ledger row stays `open` with no `channel_ref` (the responder leaves it for reconcile). After the transport recovers, the `recover()` sweep (on restart) or a manual `redeliver` re-posts: ```bash python3 run-team.py redeliver ``` - **Inbound listener down / never started** — the daemon still posts and expires; it just cannot hear Slack. **The CLI `answer` path is the outage fallback** — it runs the identical compare-and-set with no socket: ```bash python3 run-team.py answer --answer "" --via "cli:adam" --confirm ``` Fix `SLACK_APP_TOKEN` / `SLACK_BOT_TOKEN` in `~/secrev.env`, then `sudo systemctl restart agent-team-coordinator.service` to re-establish Socket Mode. (D-1: the listener starts only when the transport is live Slack AND `SLACK_APP_TOKEN` is set; it is optional by design.) - **Slack listener crash-looping** — the listener runs on an isolated daemon thread; a crash is logged (`inbound Slack listener thread exited with an error`) and takes down only the inbound socket, not the maintenance loop. Restart the service to relaunch the thread once the cause is fixed. --- ## Incident 6 — COMPLACENCY / COVERAGE alarms (Plane-1 checkers) These come from the nightly checker run, not the coordinator daemon (design §6.4, §6.6): - **COMPLACENCY ALARM** — a checker role missed a planted canary fault. That role is **skipped** for the night (it never runs silently degraded). Recover: inspect the canary corpus / the role's checker under `security-review/checkers/`, fix the regression, re-run that checker's dry-run, and confirm the canary passes before re-enabling. - **COVERAGE ALARM** — a role slipped its rotation slot (e.g. budget-deferred). It is **deferred via the rotation pointer, never dropped**, and picked up next cycle. Recover: confirm the rotation pointer advanced (it is rebuildable from report history) and that the role runs on the next cadence; investigate only if it slips repeatedly. Both alarms are **report + ALARM-only** (design D3): nothing posts on a clean state, no auto-Jira/Notion writes from the checker itself. Persistent alarms follow the escalation ladder below. --- ## Escalation ladder (design §5 / §6.6, resolves Q3) Anything that does not clear on the first ALARM escalates — but **ALARM-only in spirit** (nothing posts on a clean state): 1. **Re-alarm on a backoff.** A confirmed critical (or a COMPLACENCY / COVERAGE alarm) that persists re-alarms to Slack each night it is still unresolved, on a backoff so it does not spam. 2. **Open a Jira tracking ticket after `N` nights** (default **N = 3**). If the condition still has not cleared, the coordinator opens an **INFRA** Jira ticket so it cannot quietly linger. The same ladder applies to a parked task that exceeds `MAX_PARK`. 3. **Human (Adam) takes it from the Jira ticket.** For a parked task, recover via Incidents 2–4 above; for a checker alarm, via Incident 6. When you resolve an incident, record the action — destructive CLI verbs already write an attributable attempt+outcome record to `state/audit.log.jsonl`; for non-CLI recoveries note it on the Jira ticket. --- ## Post-incident - Confirm the daemon is `active` and the ledger has no unexpected parked rows (`list --parked`). - If you restored the VM snapshot or wiped the ledger, re-run the relevant PROVISIONING-RUNBOOK steps. - Update `project_r720_agent_team` memory if the incident revealed a durable fact (a new failure mode, a config that must change).