This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/docs/provisioning/OPERATOR-RUNBOOK.md
Adam Moussa 2d1dca0804
feat(agent-team): deploy-readiness — serve starts Slack listener + systemd + provisioning docs (#23)
* fix(agent-team): serve() starts the inbound Slack listener (D-1)

Coordinator.serve() now constructs and starts the SlackListener concurrently
with the tick/drain loop on a background daemon thread, but ONLY when the live
transport is a SlackTransport AND SLACK_APP_TOKEN is configured. When Slack is
not the transport or the app token is absent, serve() behaves exactly as before
(tick/recover only) — Slack is never made mandatory.

- New injectable build_listener seam + default_slack_listener_factory sharing
  the coordinator's own transport, ledger db_path, and resume_queue put.
- AUTHZ-01 owner-allowlist + open-status CAS untouched: serve() sources
  AGENT_TEAM_SLACK_OWNER_IDS in SlackListener.serve, which still fails closed.
- SlackListener.close() added for clean Socket Mode teardown on shutdown;
  serve() stops the listener + joins the thread in a finally.
- Tests: start-when-Slack+app-token, no-start otherwise, clean shutdown,
  idempotent start, serve start/stop around the loop, listener close().

* fix(agent-team): systemd unit loads ~/orchestrator/.env + uses venv python (D-2/D-7)

D-2: add EnvironmentFile=-/home/adam/orchestrator/.env (optional '-') so the P2
GPT-4.1 review loop's cross_reviewer sub-process can read the non-Claude provider
key once a task reaches REVIEW. Mirrors the sea-haven-secrev unit.

D-7: point ExecStart at the agent-team venv interpreter
(/home/adam/orchestrator/agent-team/.venv/bin/python) instead of
/usr/bin/env python3, which resolved the system interpreter without the
installed deps under systemd's PATH.

All hardening (NoNewPrivileges / ProtectSystem=full / ProtectHome=read-only /
ReadWritePaths) is retained unchanged (locked decision).

* docs(agent-team): land provisioning + operator runbooks under docs/provisioning

- PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5
  intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub
  App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked
  FIXED so the demo can use the live Slack answer path.
- P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the
  Slack and operator-CLI answer paths documented for all four exit criteria.
- DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the
  operator-CLI divergence kept as provisioning notes.
- OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks,
  failed HITL resumes, budget exhaustion, transport outages, and
  COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the
  re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6).

* fix(agent-team): supervise the Slack listener thread — recurring ALARM + respawn

sh-security-review (logic) MEDIUM: a crashed listener thread was logged once,
then the daemon ran on 'deaf' — posting clarifier questions but receiving no
answers, every gate silently parking, process never exiting so systemd
Restart=on-failure never fired. serve() now calls _supervise_slack_listener()
each pass: when the listener is enabled but its thread is dead, it emits a
recurring ERROR ALARM and respawns via the idempotent starter (self-heal).
No-op when alive or disabled. +3 tests. (authz detector: wiring clean — AUTHZ-01
fail-closed allowlist + open-status CAS intact, dead listener fails SAFE.)
2026-06-18 16:56:21 -04:00

290 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# OPERATOR-RUNBOOK — R720 agent-team coordinator incident handling
The on-call runbook for the always-on `agent-team-coordinator` daemon on the
`sh-secrev` R720 VM. Covers pipeline stalls, stuck/parked tasks, failed
human-in-the-loop resumes, budget exhaustion mid-pipeline, transport outages, and
the COMPLACENCY / COVERAGE alarms. Grounds every recovery in real code (design
§5 escalation ladder + §6.6 contention/park policy; Phase-6 requirement).
> **CLI used throughout: `run-team.py`** (the entry CLI), run from
> `~/orchestrator/agent-team` with the venv active so it hits the default ledger
> (`state/agent_team.sqlite`) and audit log (`state/audit.log.jsonl`). Do **not**
> use `agent_team/operator_cli.py` — it has divergent verbs (no `show`, required
> `--db`/`--audit-log`, and a `force-resume` that supersedes). See DEPLOY-AUDIT.md.
```bash
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team && . .venv/bin/activate
```
## Verb reference (all confirmed in `run-team.py:build_parser`)
| Verb | Effect | Destructive? |
|---|---|---|
| `list` / `list --all` / `list --parked` / `list --status <state>` | read pending questions (default `open`) | no |
| `show <question_id>` | print one ledger row (JSON) | no |
| `redeliver <question_id>` | clear `channel_ref` so the reconcile loop re-posts an `open` question | no (audit-logged) |
| `expire <question_id> --confirm` | force `open`→`expired` | yes |
| `answer <question_id> --answer <p> [--via <id>] --confirm` | answer-on-behalf (first-answer-wins CAS) | yes |
| `force-resume <question_id> --confirm` | reopen an `expired` (parked) question; for `answered` records resume intent | yes |
| `supersede <question_id> --confirm` | mark a stale `open`/`answered` row `superseded` | yes |
| `start --task "..." [--transport ...] [--dry-run]` | start one task to the human gate | no |
`--operator <name>` (global) sets the audit attribution; it defaults to the OS
login. Destructive verbs require `--confirm` and write an attempt-then-outcome
record to the audit log **before** mutating.
First triage for any incident:
```bash
systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e --since "-2h" | tail -100
python3 run-team.py list --all # full ledger snapshot
python3 run-team.py list --parked # non-open rows (parked-task context)
```
---
## Incident 1 — Pipeline stall (tasks not advancing)
**Symptoms:** `list` shows `open` (or `answered`) rows that never progress; the
journal shows no `tick` activity or repeated errors.
**Diagnose:**
```bash
systemctl is-active agent-team-coordinator.service # "active" expected
journalctl -u agent-team-coordinator.service -e | tail -60
# Daemon dead/looping on restart? Check the unit + deps:
.venv/bin/python -c "import langgraph, slack_sdk, slack_bolt, requests; print('deps ok')"
```
**Recover:**
1. If the daemon is dead and `Restart=on-failure` is flapping, read the journal
for the import/auth error. A missing venv dep (D-4/D-5) or a missing
`CLAUDE_CODE_OAUTH_TOKEN` is the usual cause — fix the env / venv, then:
```bash
sudo systemctl restart agent-team-coordinator.service
```
2. The startup `recover()` sweep re-drives `answered`-but-unresumed rows and
re-posts `open` rows that lost their `channel_ref`, so a clean restart
converges from the durable ledger. Confirm with `journalctl ... | tail` and
`list --all`.
3. A single task stuck `open` with a stale/missing post: re-deliver it.
```bash
python3 run-team.py show <qid>
python3 run-team.py redeliver <qid> # clears channel_ref; reconcile re-posts
```
**Escalate** (see the ladder) if a restart does not clear it within one tick
cadence and the journal shows a non-transient error.
---
## Incident 2 — Stuck / parked task
A task parks (design §6.6) when its clarifier question **expires** with no answer,
or its budget headroom drops below reserve, or a checkpoint is corrupt. Parked =
the durable ledger row is no longer `open` (it is `expired`), and an ALARM was
raised, not spun on.
**Diagnose:**
```bash
python3 run-team.py list --parked
python3 run-team.py show <qid> # status, thread_id, deadline_at, answered_via
```
**Recover — depends on why it parked:**
- **Expired with no answer (the common case)** — un-park by reopening the expired
question; it is then re-delivered for an answer:
```bash
python3 run-team.py force-resume <qid> --confirm
# "force-resume: reopened expired question <qid>; it will be re-delivered"
python3 run-team.py show <qid> # status flips back to "open"
```
Then answer it (Slack or CLI) to drive it forward.
- **Answered but not yet resumed** — the recovery sweep handles it; `force-resume`
records intent and reports that (no mutation):
```bash
python3 run-team.py force-resume <qid> --confirm
# "force-resume: question <qid> is answered and pending resume; the recovery
# sweep will resume it (intent recorded)"
sudo systemctl restart agent-team-coordinator.service # forces the recover() sweep now
```
- **Stale / wrong question that should be abandoned** — supersede it so it stops
surfacing as parked context:
```bash
python3 run-team.py supersede <qid> --confirm
```
`MAX_PARK` FIFO-aging: a task that exceeds the park window escalates (ALARM + a
Jira ticket per the ladder) rather than starving silently.
---
## Incident 3 — Failed human-in-the-loop resume
**Symptoms:** an answer was submitted (Slack or CLI) but the graph did not
advance.
**Diagnose:**
```bash
python3 run-team.py show <qid> # is status "answered"? what answered_via?
journalctl -u agent-team-coordinator.service -e | grep -iE "resume|answer|<qid>"
```
**Recover:**
- **Status is `answered` but no resume drained** — the resume queue is in-process;
a daemon restart triggers the `recover()` sweep that re-drives `answered` rows
via the turn-guarded `ResumeWorker` (idempotent — an already-advanced thread
supersedes-and-skips):
```bash
sudo systemctl restart agent-team-coordinator.service
python3 run-team.py show <qid> # confirm it advanced
```
- **Live Slack answer never registered** — the listener fails closed. Check:
```bash
journalctl -u agent-team-coordinator.service -e | grep -i "Slack answer"
# "rejecting Slack answer: owner allowlist is unconfigured ..." -> set
# AGENT_TEAM_SLACK_OWNER_IDS in ~/secrev.env and restart.
# "rejecting Slack answer ... unauthorized sender" -> the answerer's user id is
# not in the allowlist. Add it, or answer via the CLI on their behalf:
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
```
- **Answer lost the compare-and-set (`not open`)** — the row was already
answered/expired/superseded. Inspect with `show`; if it parked, go to Incident 2.
---
## Incident 4 — Budget exhaustion mid-pipeline
The shared Claude budget ledger (`budget_ledger` table) enforces per-call + total
nightly caps (design §6.1, §6.6). A stage that would breach the reserve **parks**
the task (deferred, not dropped) and ALARMs — it never loop-drains the pool.
**Diagnose:**
```bash
journalctl -u agent-team-coordinator.service -e | grep -iE "budget|reserve|park"
# Inspect today's spend directly (no CLI verb for the budget ledger; read it).
# Columns (agent_team/db/schema.py budget_ledger DDL): thread_id, stage, model,
# billing_mode, input_tokens, output_tokens, usd_cost, recorded_at, day_bucket.
sqlite3 state/agent_team.sqlite \
"SELECT day_bucket, thread_id, ROUND(SUM(usd_cost),4) AS usd, SUM(input_tokens) AS in_tok
FROM budget_ledger GROUP BY day_bucket, thread_id ORDER BY day_bucket DESC LIMIT 20;"
```
**Recover:**
- **Wait for the next budget window** — budget-exhausted roles/tasks are deferred
via the rotation pointer and picked up next cycle; this is the intended
behavior, not a failure. The parked task surfaces in `list --parked`.
- **Force a specific parked task forward now** (e.g. it is urgent and headroom has
since freed): un-park it and let the daemon re-run the stage within the
remaining cap:
```bash
python3 run-team.py force-resume <qid> --confirm
```
- **Confirm `ANTHROPIC_API_KEY` is absent** — its presence would silently meter to
API rates and blow the budget model (the billing seam pops it defensively, but
it must not be set):
```bash
grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # both must print 0
```
Do **not** raise the cap to push a task through without Adam's decision — budget
realism is a deliberate guardrail. Escalate per the ladder if a task repeatedly
parks on budget across nights.
---
## Incident 5 — Transport outage (Slack / GitHub down or misconfigured)
**Symptoms:** clarifier posts fail; the journal shows transport errors or the
inbound listener is not started.
**Diagnose:**
```bash
journalctl -u agent-team-coordinator.service -e | grep -iE "Slack|listener|post|transport"
# Did the inbound listener start?
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
# "started (Socket Mode ...)" -> inbound up
# "not started (... SLACK_APP_TOKEN is unset)" -> Slack inbound intentionally off
```
**Recover:**
- **Outbound post failed (Slack API / network)** — the ledger row stays `open`
with no `channel_ref` (the responder leaves it for reconcile). After the
transport recovers, the `recover()` sweep (on restart) or a manual `redeliver`
re-posts:
```bash
python3 run-team.py redeliver <qid>
```
- **Inbound listener down / never started** — the daemon still posts and expires;
it just cannot hear Slack. **The CLI `answer` path is the outage fallback** —
it runs the identical compare-and-set with no socket:
```bash
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
```
Fix `SLACK_APP_TOKEN` / `SLACK_BOT_TOKEN` in `~/secrev.env`, then
`sudo systemctl restart agent-team-coordinator.service` to re-establish Socket
Mode. (D-1: the listener starts only when the transport is live Slack AND
`SLACK_APP_TOKEN` is set; it is optional by design.)
- **Slack listener crash-looping** — the listener runs on an isolated daemon
thread; a crash is logged (`inbound Slack listener thread exited with an error`)
and takes down only the inbound socket, not the maintenance loop. Restart the
service to relaunch the thread once the cause is fixed.
---
## Incident 6 — COMPLACENCY / COVERAGE alarms (Plane-1 checkers)
These come from the nightly checker run, not the coordinator daemon (design §6.4,
§6.6):
- **COMPLACENCY ALARM** — a checker role missed a planted canary fault. That role
is **skipped** for the night (it never runs silently degraded). Recover: inspect
the canary corpus / the role's checker under `security-review/checkers/`, fix the
regression, re-run that checker's dry-run, and confirm the canary passes before
re-enabling.
- **COVERAGE ALARM** — a role slipped its rotation slot (e.g. budget-deferred). It
is **deferred via the rotation pointer, never dropped**, and picked up next
cycle. Recover: confirm the rotation pointer advanced (it is rebuildable from
report history) and that the role runs on the next cadence; investigate only if
it slips repeatedly.
Both alarms are **report + ALARM-only** (design D3): nothing posts on a clean
state, no auto-Jira/Notion writes from the checker itself. Persistent alarms
follow the escalation ladder below.
---
## Escalation ladder (design §5 / §6.6, resolves Q3)
Anything that does not clear on the first ALARM escalates — but **ALARM-only in
spirit** (nothing posts on a clean state):
1. **Re-alarm on a backoff.** A confirmed critical (or a COMPLACENCY / COVERAGE
alarm) that persists re-alarms to Slack each night it is still unresolved, on
a backoff so it does not spam.
2. **Open a Jira tracking ticket after `N` nights** (default **N = 3**). If the
condition still has not cleared, the coordinator opens an **INFRA** Jira ticket
so it cannot quietly linger. The same ladder applies to a parked task that
exceeds `MAX_PARK`.
3. **Human (Adam) takes it from the Jira ticket.** For a parked task, recover via
Incidents 2–4 above; for a checker alarm, via Incident 6.
When you resolve an incident, record the action — destructive CLI verbs already
write an attributable attempt+outcome record to `state/audit.log.jsonl`; for
non-CLI recoveries note it on the Jira ticket.
---
## Post-incident
- Confirm the daemon is `active` and the ledger has no unexpected parked rows
(`list --parked`).
- If you restored the VM snapshot or wiped the ledger, re-run the relevant
PROVISIONING-RUNBOOK steps.
- Update `project_r720_agent_team` memory if the incident revealed a durable
fact (a new failure mode, a config that must change).