This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/docs/provisioning/OPERATOR-RUNBOOK.md

291 lines
13 KiB
Markdown
Raw Normal View History

feat(agent-team): deploy-readiness — serve starts Slack listener + systemd + provisioning docs (#23) * fix(agent-team): serve() starts the inbound Slack listener (D-1) Coordinator.serve() now constructs and starts the SlackListener concurrently with the tick/drain loop on a background daemon thread, but ONLY when the live transport is a SlackTransport AND SLACK_APP_TOKEN is configured. When Slack is not the transport or the app token is absent, serve() behaves exactly as before (tick/recover only) — Slack is never made mandatory. - New injectable build_listener seam + default_slack_listener_factory sharing the coordinator's own transport, ledger db_path, and resume_queue put. - AUTHZ-01 owner-allowlist + open-status CAS untouched: serve() sources AGENT_TEAM_SLACK_OWNER_IDS in SlackListener.serve, which still fails closed. - SlackListener.close() added for clean Socket Mode teardown on shutdown; serve() stops the listener + joins the thread in a finally. - Tests: start-when-Slack+app-token, no-start otherwise, clean shutdown, idempotent start, serve start/stop around the loop, listener close(). * fix(agent-team): systemd unit loads ~/orchestrator/.env + uses venv python (D-2/D-7) D-2: add EnvironmentFile=-/home/adam/orchestrator/.env (optional '-') so the P2 GPT-4.1 review loop's cross_reviewer sub-process can read the non-Claude provider key once a task reaches REVIEW. Mirrors the sea-haven-secrev unit. D-7: point ExecStart at the agent-team venv interpreter (/home/adam/orchestrator/agent-team/.venv/bin/python) instead of /usr/bin/env python3, which resolved the system interpreter without the installed deps under systemd's PATH. All hardening (NoNewPrivileges / ProtectSystem=full / ProtectHome=read-only / ReadWritePaths) is retained unchanged (locked decision). * docs(agent-team): land provisioning + operator runbooks under docs/provisioning - PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5 intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked FIXED so the demo can use the live Slack answer path. - P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the Slack and operator-CLI answer paths documented for all four exit criteria. - DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the operator-CLI divergence kept as provisioning notes. - OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks, failed HITL resumes, budget exhaustion, transport outages, and COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6). * fix(agent-team): supervise the Slack listener thread — recurring ALARM + respawn sh-security-review (logic) MEDIUM: a crashed listener thread was logged once, then the daemon ran on 'deaf' — posting clarifier questions but receiving no answers, every gate silently parking, process never exiting so systemd Restart=on-failure never fired. serve() now calls _supervise_slack_listener() each pass: when the listener is enabled but its thread is dead, it emits a recurring ERROR ALARM and respawns via the idempotent starter (self-heal). No-op when alive or disabled. +3 tests. (authz detector: wiring clean — AUTHZ-01 fail-closed allowlist + open-status CAS intact, dead listener fails SAFE.)
2026-06-18 16:56:21 -04:00
# OPERATOR-RUNBOOK — R720 agent-team coordinator incident handling
The on-call runbook for the always-on `agent-team-coordinator` daemon on the
`sh-secrev` R720 VM. Covers pipeline stalls, stuck/parked tasks, failed
human-in-the-loop resumes, budget exhaustion mid-pipeline, transport outages, and
the COMPLACENCY / COVERAGE alarms. Grounds every recovery in real code (design
§5 escalation ladder + §6.6 contention/park policy; Phase-6 requirement).
> **CLI used throughout: `run-team.py`** (the entry CLI), run from
> `~/orchestrator/agent-team` with the venv active so it hits the default ledger
> (`state/agent_team.sqlite`) and audit log (`state/audit.log.jsonl`). Do **not**
> use `agent_team/operator_cli.py` — it has divergent verbs (no `show`, required
> `--db`/`--audit-log`, and a `force-resume` that supersedes). See DEPLOY-AUDIT.md.
```bash
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team && . .venv/bin/activate
```
## Verb reference (all confirmed in `run-team.py:build_parser`)
| Verb | Effect | Destructive? |
|---|---|---|
| `list` / `list --all` / `list --parked` / `list --status <state>` | read pending questions (default `open`) | no |
| `show <question_id>` | print one ledger row (JSON) | no |
| `redeliver <question_id>` | clear `channel_ref` so the reconcile loop re-posts an `open` question | no (audit-logged) |
| `expire <question_id> --confirm` | force `open`→`expired` | yes |
| `answer <question_id> --answer <p> [--via <id>] --confirm` | answer-on-behalf (first-answer-wins CAS) | yes |
| `force-resume <question_id> --confirm` | reopen an `expired` (parked) question; for `answered` records resume intent | yes |
| `supersede <question_id> --confirm` | mark a stale `open`/`answered` row `superseded` | yes |
| `start --task "..." [--transport ...] [--dry-run]` | start one task to the human gate | no |
`--operator <name>` (global) sets the audit attribution; it defaults to the OS
login. Destructive verbs require `--confirm` and write an attempt-then-outcome
record to the audit log **before** mutating.
First triage for any incident:
```bash
systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e --since "-2h" | tail -100
python3 run-team.py list --all # full ledger snapshot
python3 run-team.py list --parked # non-open rows (parked-task context)
```
---
## Incident 1 — Pipeline stall (tasks not advancing)
**Symptoms:** `list` shows `open` (or `answered`) rows that never progress; the
journal shows no `tick` activity or repeated errors.
**Diagnose:**
```bash
systemctl is-active agent-team-coordinator.service # "active" expected
journalctl -u agent-team-coordinator.service -e | tail -60
# Daemon dead/looping on restart? Check the unit + deps:
.venv/bin/python -c "import langgraph, slack_sdk, slack_bolt, requests; print('deps ok')"
```
**Recover:**
1. If the daemon is dead and `Restart=on-failure` is flapping, read the journal
for the import/auth error. A missing venv dep (D-4/D-5) or a missing
`CLAUDE_CODE_OAUTH_TOKEN` is the usual cause — fix the env / venv, then:
```bash
sudo systemctl restart agent-team-coordinator.service
```
2. The startup `recover()` sweep re-drives `answered`-but-unresumed rows and
re-posts `open` rows that lost their `channel_ref`, so a clean restart
converges from the durable ledger. Confirm with `journalctl ... | tail` and
`list --all`.
3. A single task stuck `open` with a stale/missing post: re-deliver it.
```bash
python3 run-team.py show <qid>
python3 run-team.py redeliver <qid> # clears channel_ref; reconcile re-posts
```
**Escalate** (see the ladder) if a restart does not clear it within one tick
cadence and the journal shows a non-transient error.
---
## Incident 2 — Stuck / parked task
A task parks (design §6.6) when its clarifier question **expires** with no answer,
or its budget headroom drops below reserve, or a checkpoint is corrupt. Parked =
the durable ledger row is no longer `open` (it is `expired`), and an ALARM was
raised, not spun on.
**Diagnose:**
```bash
python3 run-team.py list --parked
python3 run-team.py show <qid> # status, thread_id, deadline_at, answered_via
```
**Recover — depends on why it parked:**
- **Expired with no answer (the common case)** — un-park by reopening the expired
question; it is then re-delivered for an answer:
```bash
python3 run-team.py force-resume <qid> --confirm
# "force-resume: reopened expired question <qid>; it will be re-delivered"
python3 run-team.py show <qid> # status flips back to "open"
```
Then answer it (Slack or CLI) to drive it forward.
- **Answered but not yet resumed** — the recovery sweep handles it; `force-resume`
records intent and reports that (no mutation):
```bash
python3 run-team.py force-resume <qid> --confirm
# "force-resume: question <qid> is answered and pending resume; the recovery
# sweep will resume it (intent recorded)"
sudo systemctl restart agent-team-coordinator.service # forces the recover() sweep now
```
- **Stale / wrong question that should be abandoned** — supersede it so it stops
surfacing as parked context:
```bash
python3 run-team.py supersede <qid> --confirm
```
`MAX_PARK` FIFO-aging: a task that exceeds the park window escalates (ALARM + a
Jira ticket per the ladder) rather than starving silently.
---
## Incident 3 — Failed human-in-the-loop resume
**Symptoms:** an answer was submitted (Slack or CLI) but the graph did not
advance.
**Diagnose:**
```bash
python3 run-team.py show <qid> # is status "answered"? what answered_via?
journalctl -u agent-team-coordinator.service -e | grep -iE "resume|answer|<qid>"
```
**Recover:**
- **Status is `answered` but no resume drained** — the resume queue is in-process;
a daemon restart triggers the `recover()` sweep that re-drives `answered` rows
via the turn-guarded `ResumeWorker` (idempotent — an already-advanced thread
supersedes-and-skips):
```bash
sudo systemctl restart agent-team-coordinator.service
python3 run-team.py show <qid> # confirm it advanced
```
- **Live Slack answer never registered** — the listener fails closed. Check:
```bash
journalctl -u agent-team-coordinator.service -e | grep -i "Slack answer"
# "rejecting Slack answer: owner allowlist is unconfigured ..." -> set
# AGENT_TEAM_SLACK_OWNER_IDS in ~/secrev.env and restart.
# "rejecting Slack answer ... unauthorized sender" -> the answerer's user id is
# not in the allowlist. Add it, or answer via the CLI on their behalf:
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
```
- **Answer lost the compare-and-set (`not open`)** — the row was already
answered/expired/superseded. Inspect with `show`; if it parked, go to Incident 2.
---
## Incident 4 — Budget exhaustion mid-pipeline
The shared Claude budget ledger (`budget_ledger` table) enforces per-call + total
nightly caps (design §6.1, §6.6). A stage that would breach the reserve **parks**
the task (deferred, not dropped) and ALARMs — it never loop-drains the pool.
**Diagnose:**
```bash
journalctl -u agent-team-coordinator.service -e | grep -iE "budget|reserve|park"
# Inspect today's spend directly (no CLI verb for the budget ledger; read it).
# Columns (agent_team/db/schema.py budget_ledger DDL): thread_id, stage, model,
# billing_mode, input_tokens, output_tokens, usd_cost, recorded_at, day_bucket.
sqlite3 state/agent_team.sqlite \
"SELECT day_bucket, thread_id, ROUND(SUM(usd_cost),4) AS usd, SUM(input_tokens) AS in_tok
FROM budget_ledger GROUP BY day_bucket, thread_id ORDER BY day_bucket DESC LIMIT 20;"
```
**Recover:**
- **Wait for the next budget window** — budget-exhausted roles/tasks are deferred
via the rotation pointer and picked up next cycle; this is the intended
behavior, not a failure. The parked task surfaces in `list --parked`.
- **Force a specific parked task forward now** (e.g. it is urgent and headroom has
since freed): un-park it and let the daemon re-run the stage within the
remaining cap:
```bash
python3 run-team.py force-resume <qid> --confirm
```
- **Confirm `ANTHROPIC_API_KEY` is absent** — its presence would silently meter to
API rates and blow the budget model (the billing seam pops it defensively, but
it must not be set):
```bash
grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # both must print 0
```
Do **not** raise the cap to push a task through without Adam's decision — budget
realism is a deliberate guardrail. Escalate per the ladder if a task repeatedly
parks on budget across nights.
---
## Incident 5 — Transport outage (Slack / GitHub down or misconfigured)
**Symptoms:** clarifier posts fail; the journal shows transport errors or the
inbound listener is not started.
**Diagnose:**
```bash
journalctl -u agent-team-coordinator.service -e | grep -iE "Slack|listener|post|transport"
# Did the inbound listener start?
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
# "started (Socket Mode ...)" -> inbound up
# "not started (... SLACK_APP_TOKEN is unset)" -> Slack inbound intentionally off
```
**Recover:**
- **Outbound post failed (Slack API / network)** — the ledger row stays `open`
with no `channel_ref` (the responder leaves it for reconcile). After the
transport recovers, the `recover()` sweep (on restart) or a manual `redeliver`
re-posts:
```bash
python3 run-team.py redeliver <qid>
```
- **Inbound listener down / never started** — the daemon still posts and expires;
it just cannot hear Slack. **The CLI `answer` path is the outage fallback** —
it runs the identical compare-and-set with no socket:
```bash
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
```
Fix `SLACK_APP_TOKEN` / `SLACK_BOT_TOKEN` in `~/secrev.env`, then
`sudo systemctl restart agent-team-coordinator.service` to re-establish Socket
Mode. (D-1: the listener starts only when the transport is live Slack AND
`SLACK_APP_TOKEN` is set; it is optional by design.)
- **Slack listener crash-looping** — the listener runs on an isolated daemon
thread; a crash is logged (`inbound Slack listener thread exited with an error`)
and takes down only the inbound socket, not the maintenance loop. Restart the
service to relaunch the thread once the cause is fixed.
---
## Incident 6 — COMPLACENCY / COVERAGE alarms (Plane-1 checkers)
These come from the nightly checker run, not the coordinator daemon (design §6.4,
§6.6):
- **COMPLACENCY ALARM** — a checker role missed a planted canary fault. That role
is **skipped** for the night (it never runs silently degraded). Recover: inspect
the canary corpus / the role's checker under `security-review/checkers/`, fix the
regression, re-run that checker's dry-run, and confirm the canary passes before
re-enabling.
- **COVERAGE ALARM** — a role slipped its rotation slot (e.g. budget-deferred). It
is **deferred via the rotation pointer, never dropped**, and picked up next
cycle. Recover: confirm the rotation pointer advanced (it is rebuildable from
report history) and that the role runs on the next cadence; investigate only if
it slips repeatedly.
Both alarms are **report + ALARM-only** (design D3): nothing posts on a clean
state, no auto-Jira/Notion writes from the checker itself. Persistent alarms
follow the escalation ladder below.
---
## Escalation ladder (design §5 / §6.6, resolves Q3)
Anything that does not clear on the first ALARM escalates — but **ALARM-only in
spirit** (nothing posts on a clean state):
1. **Re-alarm on a backoff.** A confirmed critical (or a COMPLACENCY / COVERAGE
alarm) that persists re-alarms to Slack each night it is still unresolved, on
a backoff so it does not spam.
2. **Open a Jira tracking ticket after `N` nights** (default **N = 3**). If the
condition still has not cleared, the coordinator opens an **INFRA** Jira ticket
so it cannot quietly linger. The same ladder applies to a parked task that
exceeds `MAX_PARK`.
3. **Human (Adam) takes it from the Jira ticket.** For a parked task, recover via
Incidents 2–4 above; for a checker alarm, via Incident 6.
When you resolve an incident, record the action — destructive CLI verbs already
write an attributable attempt+outcome record to `state/audit.log.jsonl`; for
non-CLI recoveries note it on the Jira ticket.
---
## Post-incident
- Confirm the daemon is `active` and the ledger has no unexpected parked rows
(`list --parked`).
- If you restored the VM snapshot or wiped the ledger, re-run the relevant
PROVISIONING-RUNBOOK steps.
- Update `project_r720_agent_team` memory if the incident revealed a durable
fact (a new failure mode, a config that must change).