339 lines
15 KiB
Markdown
339 lines
15 KiB
Markdown
# OPERATOR-RUNBOOK — R720 agent-team coordinator incident handling
|
||
|
||
The on-call runbook for the always-on `agent-team-coordinator` daemon on the
|
||
`sh-secrev` R720 VM. Covers pipeline stalls, stuck/parked tasks, failed
|
||
human-in-the-loop resumes, budget exhaustion mid-pipeline, transport outages, and
|
||
the COMPLACENCY / COVERAGE alarms. Grounds every recovery in real code (design
|
||
§5 escalation ladder + §6.6 contention/park policy; Phase-6 requirement).
|
||
|
||
> **CLI used throughout: `run-team.py`** (the entry CLI), run from
|
||
> `~/orchestrator/agent-team` with the venv active so it hits the default ledger
|
||
> (`state/agent_team.sqlite`) and audit log (`state/audit.log.jsonl`). Do **not**
|
||
> use `agent_team/operator_cli.py` — it has divergent verbs (no `show`, required
|
||
> `--db`/`--audit-log`, and a `force-resume` that supersedes). See DEPLOY-AUDIT.md.
|
||
|
||
```bash
|
||
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
|
||
cd ~/orchestrator/agent-team && . .venv/bin/activate
|
||
```
|
||
|
||
## Verb reference (all confirmed in `run-team.py:build_parser`)
|
||
|
||
| Verb | Effect | Destructive? |
|
||
|---|---|---|
|
||
| `list` / `list --all` / `list --parked` / `list --status <state>` | read pending questions (default `open`) | no |
|
||
| `show <question_id>` | print one ledger row (JSON) | no |
|
||
| `redeliver <question_id>` | clear `channel_ref` so the reconcile loop re-posts an `open` question | no (audit-logged) |
|
||
| `expire <question_id> --confirm` | force `open`→`expired` | yes |
|
||
| `answer <question_id> --answer <p> [--via <id>] --confirm` | answer-on-behalf (first-answer-wins CAS) | yes |
|
||
| `force-resume <question_id> --confirm` | reopen an `expired` (parked) question; for `answered` records resume intent | yes |
|
||
| `supersede <question_id> --confirm` | mark a stale `open`/`answered` row `superseded` | yes |
|
||
| `start --task "..." [--transport ...] [--dry-run]` | start one task to the human gate | no |
|
||
|
||
`--operator <name>` (global) sets the audit attribution; it defaults to the OS
|
||
login. Destructive verbs require `--confirm` and write an attempt-then-outcome
|
||
record to the audit log **before** mutating.
|
||
|
||
First triage for any incident:
|
||
|
||
```bash
|
||
systemctl status agent-team-coordinator.service
|
||
journalctl -u agent-team-coordinator.service -e --since "-2h" | tail -100
|
||
python3 run-team.py list --all # full ledger snapshot
|
||
python3 run-team.py list --parked # non-open rows (parked-task context)
|
||
```
|
||
|
||
---
|
||
|
||
## Incident 1 — Pipeline stall (tasks not advancing)
|
||
|
||
**Symptoms:** `list` shows `open` (or `answered`) rows that never progress; the
|
||
journal shows no `tick` activity or repeated errors.
|
||
|
||
**Diagnose:**
|
||
```bash
|
||
systemctl is-active agent-team-coordinator.service # "active" expected
|
||
journalctl -u agent-team-coordinator.service -e | tail -60
|
||
# Daemon dead/looping on restart? Check the unit + deps:
|
||
.venv/bin/python -c "import langgraph, slack_sdk, slack_bolt, requests; print('deps ok')"
|
||
```
|
||
|
||
**Recover:**
|
||
1. If the daemon is dead and `Restart=on-failure` is flapping, read the journal
|
||
for the import/auth error. A missing venv dep (D-4/D-5) or a missing
|
||
`CLAUDE_CODE_OAUTH_TOKEN` is the usual cause — fix the env / venv, then:
|
||
```bash
|
||
sudo systemctl restart agent-team-coordinator.service
|
||
```
|
||
2. The startup `recover()` sweep re-drives `answered`-but-unresumed rows and
|
||
re-posts `open` rows that lost their `channel_ref`, so a clean restart
|
||
converges from the durable ledger. Confirm with `journalctl ... | tail` and
|
||
`list --all`.
|
||
3. A single task stuck `open` with a stale/missing post: re-deliver it.
|
||
```bash
|
||
python3 run-team.py show <qid>
|
||
python3 run-team.py redeliver <qid> # clears channel_ref; reconcile re-posts
|
||
```
|
||
|
||
**Escalate** (see the ladder) if a restart does not clear it within one tick
|
||
cadence and the journal shows a non-transient error.
|
||
|
||
---
|
||
|
||
## Incident 2 — Stuck / parked task
|
||
|
||
A task parks (design §6.6) when its clarifier question **expires** with no answer,
|
||
or its budget headroom drops below reserve, or a checkpoint is corrupt. Parked =
|
||
the durable ledger row is no longer `open` (it is `expired`), and an ALARM was
|
||
raised, not spun on.
|
||
|
||
**Diagnose:**
|
||
```bash
|
||
python3 run-team.py list --parked
|
||
python3 run-team.py show <qid> # status, thread_id, deadline_at, answered_via
|
||
```
|
||
|
||
**Recover — depends on why it parked:**
|
||
|
||
- **Expired with no answer (the common case)** — un-park by reopening the expired
|
||
question; it is then re-delivered for an answer:
|
||
```bash
|
||
python3 run-team.py force-resume <qid> --confirm
|
||
# "force-resume: reopened expired question <qid>; it will be re-delivered"
|
||
python3 run-team.py show <qid> # status flips back to "open"
|
||
```
|
||
Then answer it (Slack or CLI) to drive it forward.
|
||
|
||
- **Answered but not yet resumed** — the recovery sweep handles it; `force-resume`
|
||
records intent and reports that (no mutation):
|
||
```bash
|
||
python3 run-team.py force-resume <qid> --confirm
|
||
# "force-resume: question <qid> is answered and pending resume; the recovery
|
||
# sweep will resume it (intent recorded)"
|
||
sudo systemctl restart agent-team-coordinator.service # forces the recover() sweep now
|
||
```
|
||
|
||
- **Stale / wrong question that should be abandoned** — supersede it so it stops
|
||
surfacing as parked context:
|
||
```bash
|
||
python3 run-team.py supersede <qid> --confirm
|
||
```
|
||
|
||
`MAX_PARK` FIFO-aging: a task that exceeds the park window escalates (ALARM + a
|
||
Jira ticket per the ladder) rather than starving silently.
|
||
|
||
---
|
||
|
||
## Incident 3 — Failed human-in-the-loop resume
|
||
|
||
**Symptoms:** an answer was submitted (Slack or CLI) but the graph did not
|
||
advance.
|
||
|
||
**Diagnose:**
|
||
```bash
|
||
python3 run-team.py show <qid> # is status "answered"? what answered_via?
|
||
journalctl -u agent-team-coordinator.service -e | grep -iE "resume|answer|<qid>"
|
||
```
|
||
|
||
**Recover:**
|
||
- **Status is `answered` but no resume drained** — the resume queue is in-process;
|
||
a daemon restart triggers the `recover()` sweep that re-drives `answered` rows
|
||
via the turn-guarded `ResumeWorker` (idempotent — an already-advanced thread
|
||
supersedes-and-skips):
|
||
```bash
|
||
sudo systemctl restart agent-team-coordinator.service
|
||
python3 run-team.py show <qid> # confirm it advanced
|
||
```
|
||
- **Live Slack answer never registered** — the listener fails closed. Check:
|
||
```bash
|
||
journalctl -u agent-team-coordinator.service -e | grep -i "Slack answer"
|
||
# "rejecting Slack answer: owner allowlist is unconfigured ..." -> set
|
||
# AGENT_TEAM_SLACK_OWNER_IDS in ~/secrev.env and restart.
|
||
# "rejecting Slack answer ... unauthorized sender" -> the answerer's user id is
|
||
# not in the allowlist. Add it, or answer via the CLI on their behalf:
|
||
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
|
||
```
|
||
- **Answer lost the compare-and-set (`not open`)** — the row was already
|
||
answered/expired/superseded. Inspect with `show`; if it parked, go to Incident 2.
|
||
|
||
---
|
||
|
||
## Incident 4 — Budget exhaustion mid-pipeline
|
||
|
||
The shared Claude budget ledger (`budget_ledger` table) enforces per-call + total
|
||
nightly caps (design §6.1, §6.6). A stage that would breach the reserve **parks**
|
||
the task (deferred, not dropped) and ALARMs — it never loop-drains the pool.
|
||
|
||
**Diagnose:**
|
||
```bash
|
||
journalctl -u agent-team-coordinator.service -e | grep -iE "budget|reserve|park"
|
||
# Inspect today's spend directly (no CLI verb for the budget ledger; read it).
|
||
# Columns (agent_team/db/schema.py budget_ledger DDL): thread_id, stage, model,
|
||
# billing_mode, input_tokens, output_tokens, usd_cost, recorded_at, day_bucket.
|
||
sqlite3 state/agent_team.sqlite \
|
||
"SELECT day_bucket, thread_id, ROUND(SUM(usd_cost),4) AS usd, SUM(input_tokens) AS in_tok
|
||
FROM budget_ledger GROUP BY day_bucket, thread_id ORDER BY day_bucket DESC LIMIT 20;"
|
||
```
|
||
|
||
**Recover:**
|
||
- **Wait for the next budget window** — budget-exhausted roles/tasks are deferred
|
||
via the rotation pointer and picked up next cycle; this is the intended
|
||
behavior, not a failure. The parked task surfaces in `list --parked`.
|
||
- **Force a specific parked task forward now** (e.g. it is urgent and headroom has
|
||
since freed): un-park it and let the daemon re-run the stage within the
|
||
remaining cap:
|
||
```bash
|
||
python3 run-team.py force-resume <qid> --confirm
|
||
```
|
||
- **Confirm `ANTHROPIC_API_KEY` is absent** — its presence would silently meter to
|
||
API rates and blow the budget model (the billing seam pops it defensively, but
|
||
it must not be set):
|
||
```bash
|
||
grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # both must print 0
|
||
```
|
||
|
||
Do **not** raise the cap to push a task through without Adam's decision — budget
|
||
realism is a deliberate guardrail. Escalate per the ladder if a task repeatedly
|
||
parks on budget across nights.
|
||
|
||
---
|
||
|
||
## Incident 5 — Transport outage (Slack / GitHub down or misconfigured)
|
||
|
||
**Symptoms:** clarifier posts fail; the journal shows transport errors or the
|
||
inbound listener is not started.
|
||
|
||
**Diagnose:**
|
||
```bash
|
||
journalctl -u agent-team-coordinator.service -e | grep -iE "Slack|listener|post|transport"
|
||
# Did the inbound listener start?
|
||
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
|
||
# "started (Socket Mode ...)" -> inbound up
|
||
# "not started (... SLACK_APP_TOKEN is unset)" -> Slack inbound intentionally off
|
||
```
|
||
|
||
**Recover:**
|
||
- **Outbound post failed (Slack API / network)** — the ledger row stays `open`
|
||
with no `channel_ref` (the responder leaves it for reconcile). After the
|
||
transport recovers, the `recover()` sweep (on restart) or a manual `redeliver`
|
||
re-posts:
|
||
```bash
|
||
python3 run-team.py redeliver <qid>
|
||
```
|
||
- **Inbound listener down / never started** — the daemon still posts and expires;
|
||
it just cannot hear Slack. **The CLI `answer` path is the outage fallback** —
|
||
it runs the identical compare-and-set with no socket:
|
||
```bash
|
||
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
|
||
```
|
||
Fix `SLACK_APP_TOKEN` / `SLACK_BOT_TOKEN` in `~/secrev.env`, then
|
||
`sudo systemctl restart agent-team-coordinator.service` to re-establish Socket
|
||
Mode. (D-1: the listener starts only when the transport is live Slack AND
|
||
`SLACK_APP_TOKEN` is set; it is optional by design.)
|
||
- **Slack listener crash-looping** — the listener runs on an isolated daemon
|
||
thread; a crash is logged (`inbound Slack listener thread exited with an error`)
|
||
and takes down only the inbound socket, not the maintenance loop. Restart the
|
||
service to relaunch the thread once the cause is fixed.
|
||
|
||
---
|
||
|
||
## Incident 5b — WS0–WS5 surfaces (HTTP API, /new-task, handbook context)
|
||
|
||
These were added by the WS-rollout (see `agent-team/DEPLOY-R720.md` §4b). They
|
||
layer onto the coordinator; none of them should take down the maintenance loop.
|
||
|
||
- **HTTP API down / unreachable** — the WS1 FastAPI app (`agent_team/api.py`) is
|
||
a **separate, opt-in process** (`api.serve()`, `127.0.0.1:8765`, bearer auth),
|
||
**not** started by the coordinator daemon. If `/delegate` from Claude Code or
|
||
`POST /tasks` over HTTP stops working, the coordinator itself is unaffected —
|
||
check the API process separately:
|
||
```bash
|
||
curl -sS -o /dev/null -w '%{http_code}\n' \
|
||
-H "Authorization: Bearer $AGENT_TEAM_API_TOKEN" http://127.0.0.1:8765/tasks
|
||
# 405 = API up + authed (GET not allowed on /tasks); 000 = process down;
|
||
# 401 = AGENT_TEAM_API_TOKEN mismatch (client vs ~/secrev.env).
|
||
```
|
||
The API refuses to start if `AGENT_TEAM_API_TOKEN` is unset/empty (logs a
|
||
`RuntimeError`). Fix the token, restart the API process. Tasks already in the
|
||
ledger are unaffected — the API is only an *intake/invoke* front door; answer
|
||
via Slack or the CLI as usual.
|
||
|
||
- **`/new-task` Slack command not responding** — the WS2 slash command is
|
||
AUTHZ-01 owner-allowlist gated and routes through the same Socket Mode listener
|
||
as answers. If it silently does nothing, it is almost always the owner
|
||
allowlist (same failure mode as Incident 3's live-Slack path):
|
||
```bash
|
||
journalctl -u agent-team-coordinator.service -e | grep -iE "new-task|owner|unauthorized"
|
||
# unauthorized sender / unconfigured allowlist -> fix AGENT_TEAM_SLACK_OWNER_IDS
|
||
```
|
||
Fallback: start the task from the CLI (`run-team.py start --task "..."`) or the
|
||
HTTP API. If the listener itself is down, see Incident 5 (inbound listener).
|
||
|
||
- **Handbook dir missing → planner runs without handbook context** — the WS5
|
||
`context_provider` (`load_handbook_conventions`) is **fail-safe**: if
|
||
`SEA_HAVEN_HANDBOOK_DIR` (or `~/.sea-haven/engineering-handbook`) is missing or
|
||
unreadable it returns `""` and the planner runs normally, just without handbook
|
||
conventions injected. This is **degraded, not broken** — no park, no alarm.
|
||
Confirm and restore:
|
||
```bash
|
||
grep '^SEA_HAVEN_HANDBOOK_DIR=' ~/secrev.env
|
||
ls "$(grep '^SEA_HAVEN_HANDBOOK_DIR=' ~/secrev.env | cut -d= -f2)" # dir present + populated?
|
||
```
|
||
Re-sync the handbook (the deploy script does this) and restart the daemon so
|
||
the planner picks it back up. The WS3 dispatch node is **inert** (gated) and
|
||
should never appear in pipeline activity; if it does, treat as an unexpected
|
||
state and escalate.
|
||
|
||
---
|
||
|
||
## Incident 6 — COMPLACENCY / COVERAGE alarms (Plane-1 checkers)
|
||
|
||
These come from the nightly checker run, not the coordinator daemon (design §6.4,
|
||
§6.6):
|
||
|
||
- **COMPLACENCY ALARM** — a checker role missed a planted canary fault. That role
|
||
is **skipped** for the night (it never runs silently degraded). Recover: inspect
|
||
the canary corpus / the role's checker under `security-review/checkers/`, fix the
|
||
regression, re-run that checker's dry-run, and confirm the canary passes before
|
||
re-enabling.
|
||
- **COVERAGE ALARM** — a role slipped its rotation slot (e.g. budget-deferred). It
|
||
is **deferred via the rotation pointer, never dropped**, and picked up next
|
||
cycle. Recover: confirm the rotation pointer advanced (it is rebuildable from
|
||
report history) and that the role runs on the next cadence; investigate only if
|
||
it slips repeatedly.
|
||
|
||
Both alarms are **report + ALARM-only** (design D3): nothing posts on a clean
|
||
state, no auto-Jira/Notion writes from the checker itself. Persistent alarms
|
||
follow the escalation ladder below.
|
||
|
||
---
|
||
|
||
## Escalation ladder (design §5 / §6.6, resolves Q3)
|
||
|
||
Anything that does not clear on the first ALARM escalates — but **ALARM-only in
|
||
spirit** (nothing posts on a clean state):
|
||
|
||
1. **Re-alarm on a backoff.** A confirmed critical (or a COMPLACENCY / COVERAGE
|
||
alarm) that persists re-alarms to Slack each night it is still unresolved, on
|
||
a backoff so it does not spam.
|
||
2. **Open a Jira tracking ticket after `N` nights** (default **N = 3**). If the
|
||
condition still has not cleared, the coordinator opens an **INFRA** Jira ticket
|
||
so it cannot quietly linger. The same ladder applies to a parked task that
|
||
exceeds `MAX_PARK`.
|
||
3. **Human (Adam) takes it from the Jira ticket.** For a parked task, recover via
|
||
Incidents 2–4 above; for a checker alarm, via Incident 6.
|
||
|
||
When you resolve an incident, record the action — destructive CLI verbs already
|
||
write an attributable attempt+outcome record to `state/audit.log.jsonl`; for
|
||
non-CLI recoveries note it on the Jira ticket.
|
||
|
||
---
|
||
|
||
## Post-incident
|
||
|
||
- Confirm the daemon is `active` and the ledger has no unexpected parked rows
|
||
(`list --parked`).
|
||
- If you restored the VM snapshot or wiped the ledger, re-run the relevant
|
||
PROVISIONING-RUNBOOK steps.
|
||
- Update `project_r720_agent_team` memory if the incident revealed a durable
|
||
fact (a new failure mode, a config that must change).
|