diff --git a/docs/provisioning/DEPLOY-AUDIT.md b/docs/provisioning/DEPLOY-AUDIT.md new file mode 100644 index 0000000..91bdb39 --- /dev/null +++ b/docs/provisioning/DEPLOY-AUDIT.md @@ -0,0 +1,187 @@ +# DEPLOY-AUDIT β€” `agent-team/DEPLOY-R720.md` + systemd unit vs. actual code + +Cross-check of the drafted deploy doc (`agent-team/DEPLOY-R720.md`) and the +systemd unit (`agent-team/systemd/agent-team-coordinator.service`) against the +**actual current code** in `agent-team/agent_team/**` and `agent-team/run-team.py`. + +Each finding: **location β†’ claimed β†’ actual β†’ fix**. Severity: πŸ”΄ blocker / +🟠 should-fix / 🟑 nit. Items confirmed clean are stated explicitly. + +> **Status (deploy-readiness PR):** the three deploy-correctness bugs that +> blocked the live provisioning session β€” **D-1, D-2, D-7** β€” are **RESOLVED** in +> this PR (`feature/agent-team-deploy-readiness`). The remaining items +> (**D-4 / D-5** dependency pinning, the **operator-CLI divergence**) are kept +> below as **provisioning notes** β€” they do not block the coordinator deploy. + +--- + +## βœ… D-1 β€” `serve` now starts the Slack inbound listener β€” RESOLVED (this PR) + +- **Location:** `agent_team/coordinator.py` (`Coordinator.serve` + + `_maybe_start_slack_listener` / `_slack_listener_enabled` / `_stop_slack_listener` + / `default_slack_listener_factory`); `agent_team/transport/slack_listener.py` + (`SlackListener.serve` + new `close`). +- **Was:** `Coordinator.serve()` did only `bind_subscription_invoker()`, + `setup()`, `recover()`, then an infinite tick/sleep loop β€” it never constructed + or started `SlackListener`, so a deployed daemon posted clarifier questions and + expired them on deadline but could **not hear Slack answers**. +- **Now:** `serve()` starts the `SlackListener` on a background **daemon thread**, + concurrently with the tick/drain loop, **when** the live transport is a + `SlackTransport` AND `SLACK_APP_TOKEN` is set. It shares the coordinator's own + transport, ledger `db_path`, and `resume_queue` put; on shutdown it calls the + listener's new `close()` and joins the thread in a `finally`. When Slack is not + the transport or the app token is absent, no listener starts and `serve` behaves + exactly as before β€” **Slack is never made mandatory**. The AUTHZ-01 owner + allowlist + the open-status compare-and-set are untouched (still fail closed on + an empty `AGENT_TEAM_SLACK_OWNER_IDS`). +- **Tests:** `tests/test_coordinator.py` β€” start-when-Slack+app-token, no-start + without the token, no-start when the transport is not Slack, idempotent start, + clean shutdown, serve start/stop around the loop; `tests/test_slack_listener.py` + β€” `close()` no-op + handler teardown. + +--- + +## βœ… D-2 β€” coordinator unit loads `~/orchestrator/.env` β€” RESOLVED (this PR) + +- **Location:** `agent-team/systemd/agent-team-coordinator.service`; + `run-team.py:_build_coordinator` (always wires `default_review_wiring`); + `coordinator.py:default_review_wiring` β†’ `review_loop_llm.default_plan_reviewer` + (shells the local orchestrator `run.py` β†’ `cross_reviewer` GPT-4.1). +- **Was:** the unit loaded only `~/secrev.env`. `run-team.py serve` builds the + coordinator with the P2 review loop wired, and once a task reaches REVIEW the + reviewer shells the orchestrator `run.py`, whose GPT-4.1 call reads the + non-Claude provider key from `~/orchestrator/.env`. The review path would fail + to authenticate. +- **Now:** the unit adds `EnvironmentFile=-/home/adam/orchestrator/.env` + (optional `-`, mirroring the `sea-haven-secrev` unit). Does not affect the P1 + demo (P1 stops at PLAN before REVIEW); closes the latent P2 break. + +--- + +## 🟠 D-4 β€” pip install list omits `requests` (provisioning note) + +- **Location:** `agent_team/transport/github_live.py` / `github_intake.py` + (`import requests`). +- **Actual:** the GitHub transport + intake require `requests`; the Slack-first + path does not hit it, but any `--transport github` / `intake-github` use fails + with a clear RuntimeError without it. +- **Fix:** PROVISIONING-RUNBOOK Step 4 installs `requests` into the venv. + `requests` is also absent from `requirements.txt` (see D-5). + +--- + +## 🟠 D-5 β€” agent-team runtime deps are not pinned in `requirements.txt` (provisioning note) + +- **Location:** `requirements.txt` (repo root). +- **Actual:** `requirements.txt` pins `langgraph==1.1.10` and + `langgraph-checkpoint-sqlite==3.1.0`, but the agent-team runtime deps + `claude-agent-sdk`, `slack_sdk`, `slack_bolt`, `requests` (and `anthropic` for + api mode) are **not in `requirements.txt` at all** β€” they are installed ad-hoc + into the agent-team venv by the runbook. There is no pinned, reproducible source + of truth for the box's runtime set. +- **Fix (deferred):** add an `agent-team/requirements.txt` (or extras group) + pinning these, version-matched to the root `requirements.txt` langgraph pin. + Until then, PROVISIONING-RUNBOOK Step 4 pins `langgraph==1.1.10` / + `langgraph-checkpoint-sqlite==3.1.0` explicitly so the unpinned `pip install` + cannot pull a newer, untested major. **Do NOT modify `requirements.txt` or the + checkers in this PR** (out of scope). + +--- + +## βœ… D-6 β€” `slack_bolt` is now exercised by the daemon β€” RESOLVED (consequence of D-1) + +- **Location:** `slack_listener.py:serve` (the only `slack_bolt` import). +- **Now:** with D-1 fixed, `Coordinator.serve()` starts `SlackListener.serve()`, + which imports + uses `slack_bolt` for the Socket Mode handler. The dep is right + and now actually exercised on the live Slack path. + +--- + +## βœ… D-SLACKVAR (clean) β€” `SLACK_CHANNEL_ID` matches + +- `run-team.py:_build_transport` reads exactly `os.environ.get("SLACK_CHANNEL_ID")`. + The runbook, the unit comment, and the code all use `SLACK_CHANNEL_ID` (not + `SLACK_CHANNEL`). **CLEAN.** + +--- + +## βœ… D-ENV-SLACKBOT / OAUTH / OWNERS / APPTOKEN (clean) β€” names match + +- **`SLACK_BOT_TOKEN`** ↔ `slack_live.py` + `slack_listener` env read. **CLEAN.** +- **`CLAUDE_CODE_OAUTH_TOKEN`** ↔ `invoker.py`. **CLEAN.** +- **`AGENT_TEAM_SLACK_OWNER_IDS`** ↔ `slack_listener.py` (name + fail-closed + semantics). **CLEAN** β€” now read by the running daemon (D-1 fixed). +- **`SLACK_APP_TOKEN`** β€” now read in two places: `coordinator._slack_listener_enabled` + gates the listener on its presence, and `default_slack_listener_factory` / + `SlackListener.serve` source it to open the socket. **CLEAN** (read site exists + now that D-1 is fixed). + +--- + +## βœ… D-7 β€” `ExecStart` uses the venv interpreter β€” RESOLVED (this PR) + +- **Location:** `agent-team/systemd/agent-team-coordinator.service` ExecStart. +- **Was:** `ExecStart=/usr/bin/env python3 run-team.py serve` resolved the + **system** interpreter under systemd's PATH β€” not the venv where the deps were + installed, so the daemon would fail at import. +- **Now:** `ExecStart=/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve` + (matches the runbook venv path + `WorkingDirectory`). + +--- + +## βœ… D-SUBCMD (mostly clean) β€” run-team.py subcommands referenced exist + +Cross-checked every `run-team.py ` the deploy doc + demo name against +`run-team.py:build_parser`: `init-db`, `serve`, `list` (+ `--all` / `--parked`), +`show`, `expire`, `answer`, `redeliver`, `supersede`, `force-resume`, `start`, +`intake-github`, `intake-checker`, `fix` β€” all exist. No invented verbs. + +> ### Operator-CLI divergence (provisioning note) +> +> Two operator CLIs exist with **different verb names**: +> +> - `run-team.py` (the entry CLI): `init-db, list, show, redeliver, expire, +> answer, supersede, force-resume, start, serve, intake-github, intake-checker, +> fix`. Has `show` and `--parked`; `--db` / `--audit-log` default sensibly. +> - `agent_team/operator_cli.py`: `list, redeliver, force-expire, +> answer-on-behalf, force-resume` β€” **no `show`**, `--db` / `--audit-log` are +> **required**, and its `force-resume` **supersedes** (unlike `run-team.py`'s, +> which reopens an expired row and never supersedes). +> +> **Use `run-team.py` for provisioning + the demo + incident recovery.** The +> docs reference only `run-team.py`. Reconciling the two CLIs is a follow-up. + +--- + +## βœ… D-PYTHONPKG (clean) β€” package import bootstrap is correct + +`run-team.py` inserts its own dir into `sys.path` so the hyphenated script +imports the `agent_team` package without an editable install. **CLEAN.** + +--- + +## βœ… D-HARDENING (clean, and matches the locked decision) + +- Unit: `NoNewPrivileges=true`, `ProtectSystem=full`, `ProtectHome=read-only`, + `ReadWritePaths=/home/adam/orchestrator/agent-team/state`. **Retained unchanged** + in this PR (locked decision β€” do not revert to secrev parity). +- The `ReadWritePaths` carve-out matches the ledger + audit-log location + (`state/agent_team.sqlite`, `state/audit.log.jsonl`). **CLEAN.** + +--- + +## Summary table + +| ID | Sev | Status | One-line | +|---|---|---|---| +| D-1 | πŸ”΄ | βœ… RESOLVED (PR) | `serve` starts `SlackListener` (Slack + app-token gated; Slack stays optional) | +| D-2 | πŸ”΄ | βœ… RESOLVED (PR) | unit loads `~/orchestrator/.env` for the P2 GPT-4.1 review provider key | +| D-7 | 🟑 | βœ… RESOLVED (PR) | `ExecStart` points at the agent-team venv interpreter | +| D-6 | 🟑 | βœ… RESOLVED | `slack_bolt` now exercised by the daemon (consequence of D-1) | +| D-4 | 🟠 | NOTE | pip list omits `requests` β€” runbook Step 4 installs it | +| D-5 | 🟠 | NOTE | agent-team runtime deps not pinned in `requirements.txt` β€” runbook pins langgraph | +| operator-CLI | β€” | NOTE | `run-team.py` vs `operator_cli.py` divergent verbs β€” use `run-team.py` | +| D-SLACKVAR | βœ… | CLEAN | `SLACK_CHANNEL_ID` matches everywhere | +| D-ENV-* | βœ… | CLEAN | bot/oauth/owner/app-token env names match; all read sites now exist | +| D-SUBCMD | βœ… | CLEAN | every `run-team.py` verb/flag the docs cite exists | +| D-HARDENING | βœ… | CLEAN | unit hardening retained unchanged (locked decision) | diff --git a/docs/provisioning/OPERATOR-RUNBOOK.md b/docs/provisioning/OPERATOR-RUNBOOK.md new file mode 100644 index 0000000..806b583 --- /dev/null +++ b/docs/provisioning/OPERATOR-RUNBOOK.md @@ -0,0 +1,290 @@ +# OPERATOR-RUNBOOK β€” R720 agent-team coordinator incident handling + +The on-call runbook for the always-on `agent-team-coordinator` daemon on the +`sh-secrev` R720 VM. Covers pipeline stalls, stuck/parked tasks, failed +human-in-the-loop resumes, budget exhaustion mid-pipeline, transport outages, and +the COMPLACENCY / COVERAGE alarms. Grounds every recovery in real code (design +Β§5 escalation ladder + Β§6.6 contention/park policy; Phase-6 requirement). + +> **CLI used throughout: `run-team.py`** (the entry CLI), run from +> `~/orchestrator/agent-team` with the venv active so it hits the default ledger +> (`state/agent_team.sqlite`) and audit log (`state/audit.log.jsonl`). Do **not** +> use `agent_team/operator_cli.py` β€” it has divergent verbs (no `show`, required +> `--db`/`--audit-log`, and a `force-resume` that supersedes). See DEPLOY-AUDIT.md. + +```bash +ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120 +cd ~/orchestrator/agent-team && . .venv/bin/activate +``` + +## Verb reference (all confirmed in `run-team.py:build_parser`) + +| Verb | Effect | Destructive? | +|---|---|---| +| `list` / `list --all` / `list --parked` / `list --status ` | read pending questions (default `open`) | no | +| `show ` | print one ledger row (JSON) | no | +| `redeliver ` | clear `channel_ref` so the reconcile loop re-posts an `open` question | no (audit-logged) | +| `expire --confirm` | force `open`β†’`expired` | yes | +| `answer --answer

[--via ] --confirm` | answer-on-behalf (first-answer-wins CAS) | yes | +| `force-resume --confirm` | reopen an `expired` (parked) question; for `answered` records resume intent | yes | +| `supersede --confirm` | mark a stale `open`/`answered` row `superseded` | yes | +| `start --task "..." [--transport ...] [--dry-run]` | start one task to the human gate | no | + +`--operator ` (global) sets the audit attribution; it defaults to the OS +login. Destructive verbs require `--confirm` and write an attempt-then-outcome +record to the audit log **before** mutating. + +First triage for any incident: + +```bash +systemctl status agent-team-coordinator.service +journalctl -u agent-team-coordinator.service -e --since "-2h" | tail -100 +python3 run-team.py list --all # full ledger snapshot +python3 run-team.py list --parked # non-open rows (parked-task context) +``` + +--- + +## Incident 1 β€” Pipeline stall (tasks not advancing) + +**Symptoms:** `list` shows `open` (or `answered`) rows that never progress; the +journal shows no `tick` activity or repeated errors. + +**Diagnose:** +```bash +systemctl is-active agent-team-coordinator.service # "active" expected +journalctl -u agent-team-coordinator.service -e | tail -60 +# Daemon dead/looping on restart? Check the unit + deps: +.venv/bin/python -c "import langgraph, slack_sdk, slack_bolt, requests; print('deps ok')" +``` + +**Recover:** +1. If the daemon is dead and `Restart=on-failure` is flapping, read the journal + for the import/auth error. A missing venv dep (D-4/D-5) or a missing + `CLAUDE_CODE_OAUTH_TOKEN` is the usual cause β€” fix the env / venv, then: + ```bash + sudo systemctl restart agent-team-coordinator.service + ``` +2. The startup `recover()` sweep re-drives `answered`-but-unresumed rows and + re-posts `open` rows that lost their `channel_ref`, so a clean restart + converges from the durable ledger. Confirm with `journalctl ... | tail` and + `list --all`. +3. A single task stuck `open` with a stale/missing post: re-deliver it. + ```bash + python3 run-team.py show + python3 run-team.py redeliver # clears channel_ref; reconcile re-posts + ``` + +**Escalate** (see the ladder) if a restart does not clear it within one tick +cadence and the journal shows a non-transient error. + +--- + +## Incident 2 β€” Stuck / parked task + +A task parks (design Β§6.6) when its clarifier question **expires** with no answer, +or its budget headroom drops below reserve, or a checkpoint is corrupt. Parked = +the durable ledger row is no longer `open` (it is `expired`), and an ALARM was +raised, not spun on. + +**Diagnose:** +```bash +python3 run-team.py list --parked +python3 run-team.py show # status, thread_id, deadline_at, answered_via +``` + +**Recover β€” depends on why it parked:** + +- **Expired with no answer (the common case)** β€” un-park by reopening the expired + question; it is then re-delivered for an answer: + ```bash + python3 run-team.py force-resume --confirm + # "force-resume: reopened expired question ; it will be re-delivered" + python3 run-team.py show # status flips back to "open" + ``` + Then answer it (Slack or CLI) to drive it forward. + +- **Answered but not yet resumed** β€” the recovery sweep handles it; `force-resume` + records intent and reports that (no mutation): + ```bash + python3 run-team.py force-resume --confirm + # "force-resume: question is answered and pending resume; the recovery + # sweep will resume it (intent recorded)" + sudo systemctl restart agent-team-coordinator.service # forces the recover() sweep now + ``` + +- **Stale / wrong question that should be abandoned** β€” supersede it so it stops + surfacing as parked context: + ```bash + python3 run-team.py supersede --confirm + ``` + +`MAX_PARK` FIFO-aging: a task that exceeds the park window escalates (ALARM + a +Jira ticket per the ladder) rather than starving silently. + +--- + +## Incident 3 β€” Failed human-in-the-loop resume + +**Symptoms:** an answer was submitted (Slack or CLI) but the graph did not +advance. + +**Diagnose:** +```bash +python3 run-team.py show # is status "answered"? what answered_via? +journalctl -u agent-team-coordinator.service -e | grep -iE "resume|answer|" +``` + +**Recover:** +- **Status is `answered` but no resume drained** β€” the resume queue is in-process; + a daemon restart triggers the `recover()` sweep that re-drives `answered` rows + via the turn-guarded `ResumeWorker` (idempotent β€” an already-advanced thread + supersedes-and-skips): + ```bash + sudo systemctl restart agent-team-coordinator.service + python3 run-team.py show # confirm it advanced + ``` +- **Live Slack answer never registered** β€” the listener fails closed. Check: + ```bash + journalctl -u agent-team-coordinator.service -e | grep -i "Slack answer" + # "rejecting Slack answer: owner allowlist is unconfigured ..." -> set + # AGENT_TEAM_SLACK_OWNER_IDS in ~/secrev.env and restart. + # "rejecting Slack answer ... unauthorized sender" -> the answerer's user id is + # not in the allowlist. Add it, or answer via the CLI on their behalf: + python3 run-team.py answer --answer "" --via "cli:adam" --confirm + ``` +- **Answer lost the compare-and-set (`not open`)** β€” the row was already + answered/expired/superseded. Inspect with `show`; if it parked, go to Incident 2. + +--- + +## Incident 4 β€” Budget exhaustion mid-pipeline + +The shared Claude budget ledger (`budget_ledger` table) enforces per-call + total +nightly caps (design Β§6.1, Β§6.6). A stage that would breach the reserve **parks** +the task (deferred, not dropped) and ALARMs β€” it never loop-drains the pool. + +**Diagnose:** +```bash +journalctl -u agent-team-coordinator.service -e | grep -iE "budget|reserve|park" +# Inspect today's spend directly (no CLI verb for the budget ledger; read it). +# Columns (agent_team/db/schema.py budget_ledger DDL): thread_id, stage, model, +# billing_mode, input_tokens, output_tokens, usd_cost, recorded_at, day_bucket. +sqlite3 state/agent_team.sqlite \ + "SELECT day_bucket, thread_id, ROUND(SUM(usd_cost),4) AS usd, SUM(input_tokens) AS in_tok + FROM budget_ledger GROUP BY day_bucket, thread_id ORDER BY day_bucket DESC LIMIT 20;" +``` + +**Recover:** +- **Wait for the next budget window** β€” budget-exhausted roles/tasks are deferred + via the rotation pointer and picked up next cycle; this is the intended + behavior, not a failure. The parked task surfaces in `list --parked`. +- **Force a specific parked task forward now** (e.g. it is urgent and headroom has + since freed): un-park it and let the daemon re-run the stage within the + remaining cap: + ```bash + python3 run-team.py force-resume --confirm + ``` +- **Confirm `ANTHROPIC_API_KEY` is absent** β€” its presence would silently meter to + API rates and blow the budget model (the billing seam pops it defensively, but + it must not be set): + ```bash + grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # both must print 0 + ``` + +Do **not** raise the cap to push a task through without Adam's decision β€” budget +realism is a deliberate guardrail. Escalate per the ladder if a task repeatedly +parks on budget across nights. + +--- + +## Incident 5 β€” Transport outage (Slack / GitHub down or misconfigured) + +**Symptoms:** clarifier posts fail; the journal shows transport errors or the +inbound listener is not started. + +**Diagnose:** +```bash +journalctl -u agent-team-coordinator.service -e | grep -iE "Slack|listener|post|transport" +# Did the inbound listener start? +journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener" +# "started (Socket Mode ...)" -> inbound up +# "not started (... SLACK_APP_TOKEN is unset)" -> Slack inbound intentionally off +``` + +**Recover:** +- **Outbound post failed (Slack API / network)** β€” the ledger row stays `open` + with no `channel_ref` (the responder leaves it for reconcile). After the + transport recovers, the `recover()` sweep (on restart) or a manual `redeliver` + re-posts: + ```bash + python3 run-team.py redeliver + ``` +- **Inbound listener down / never started** β€” the daemon still posts and expires; + it just cannot hear Slack. **The CLI `answer` path is the outage fallback** β€” + it runs the identical compare-and-set with no socket: + ```bash + python3 run-team.py answer --answer "" --via "cli:adam" --confirm + ``` + Fix `SLACK_APP_TOKEN` / `SLACK_BOT_TOKEN` in `~/secrev.env`, then + `sudo systemctl restart agent-team-coordinator.service` to re-establish Socket + Mode. (D-1: the listener starts only when the transport is live Slack AND + `SLACK_APP_TOKEN` is set; it is optional by design.) +- **Slack listener crash-looping** β€” the listener runs on an isolated daemon + thread; a crash is logged (`inbound Slack listener thread exited with an error`) + and takes down only the inbound socket, not the maintenance loop. Restart the + service to relaunch the thread once the cause is fixed. + +--- + +## Incident 6 β€” COMPLACENCY / COVERAGE alarms (Plane-1 checkers) + +These come from the nightly checker run, not the coordinator daemon (design Β§6.4, +Β§6.6): + +- **COMPLACENCY ALARM** β€” a checker role missed a planted canary fault. That role + is **skipped** for the night (it never runs silently degraded). Recover: inspect + the canary corpus / the role's checker under `security-review/checkers/`, fix the + regression, re-run that checker's dry-run, and confirm the canary passes before + re-enabling. +- **COVERAGE ALARM** β€” a role slipped its rotation slot (e.g. budget-deferred). It + is **deferred via the rotation pointer, never dropped**, and picked up next + cycle. Recover: confirm the rotation pointer advanced (it is rebuildable from + report history) and that the role runs on the next cadence; investigate only if + it slips repeatedly. + +Both alarms are **report + ALARM-only** (design D3): nothing posts on a clean +state, no auto-Jira/Notion writes from the checker itself. Persistent alarms +follow the escalation ladder below. + +--- + +## Escalation ladder (design Β§5 / Β§6.6, resolves Q3) + +Anything that does not clear on the first ALARM escalates β€” but **ALARM-only in +spirit** (nothing posts on a clean state): + +1. **Re-alarm on a backoff.** A confirmed critical (or a COMPLACENCY / COVERAGE + alarm) that persists re-alarms to Slack each night it is still unresolved, on + a backoff so it does not spam. +2. **Open a Jira tracking ticket after `N` nights** (default **N = 3**). If the + condition still has not cleared, the coordinator opens an **INFRA** Jira ticket + so it cannot quietly linger. The same ladder applies to a parked task that + exceeds `MAX_PARK`. +3. **Human (Adam) takes it from the Jira ticket.** For a parked task, recover via + Incidents 2–4 above; for a checker alarm, via Incident 6. + +When you resolve an incident, record the action β€” destructive CLI verbs already +write an attributable attempt+outcome record to `state/audit.log.jsonl`; for +non-CLI recoveries note it on the Jira ticket. + +--- + +## Post-incident + +- Confirm the daemon is `active` and the ledger has no unexpected parked rows + (`list --parked`). +- If you restored the VM snapshot or wiped the ledger, re-run the relevant + PROVISIONING-RUNBOOK steps. +- Update `project_r720_agent_team` memory if the incident revealed a durable + fact (a new failure mode, a config that must change). diff --git a/docs/provisioning/P1-DEMO-SCRIPT.md b/docs/provisioning/P1-DEMO-SCRIPT.md new file mode 100644 index 0000000..aaafbd3 --- /dev/null +++ b/docs/provisioning/P1-DEMO-SCRIPT.md @@ -0,0 +1,286 @@ +# P1-DEMO-SCRIPT β€” live four-criteria acceptance demo (design Β§3.3.1 / Β§7.1 P1) + +The Β§7.1 P1 exit gate: demonstrate, on the live box, all four durable +human-in-the-loop criteria before P1 is accepted: + +- **(a)** kill the box mid-wait and have the task resume after restart; +- **(b)** submit a duplicate answer and confirm it no-ops; +- **(c)** submit an answer after the deadline expired and confirm it is rejected + and the task parks; +- **(d)** two tasks suspended concurrently resume independently to the correct + thread. + +Every command below is grounded in the **actual** code surface +(`run-team.py`, `coordinator.py`, `slack_listener.py`, `responder.py`, +`resume_worker.py`, `db/schema.py`, `graph.py`). No invented flags. Where the +code does not expose a needed knob (e.g. a short deadline), the script uses a +direct `sqlite3` write against the documented `pending_questions` schema and says +so. + +## Two answer paths β€” live Slack OR the operator CLI + +**D-1 is fixed:** `Coordinator.serve()` now starts the inbound `SlackListener` +when the live transport is Slack AND `SLACK_APP_TOKEN` is set, so the +**live Slack answer round-trip works**. You can run the demo either way: + +- **Live Slack** β€” Adam clicks the Block Kit button / replies in the channel; the + listener normalizes the event, runs the AUTHZ-01 owner check, and drives the + first-answer-wins compare-and-set (`responder.submit_answer` β†’ + `db.schema.answer_question`). +- **Operator CLI** β€” `run-team.py answer --answer ... --confirm` runs the + **identical** compare-and-set (audit-logged answer-on-behalf). Useful when the + Slack app is not yet provisioned, or to script the assertions. + +Both exercise the same durable mechanic; the human gate decision is always +Adam's. The assertions below assert on the **durable ledger status** (the Β§3.3.1 +source of truth) and are identical for either path. The examples use the CLI +`answer --confirm` form so they are copy-pasteable; substitute "Adam answers in +Slack" wherever you prefer the live path. + +> The automated proof of this mechanic is `tests/sim/test_p1_exit_criteria.py` +> (a `SimPipeline` harness) and `tests/test_coordinator.py` (the serve/listener +> wiring). This script is the live-box demonstration on top of that. + +## Verb map (use `run-team.py`, not `operator_cli.py`) + +| Need | `run-team.py` verb | Notes | +|---|---|---| +| start a task to the human gate | `start --task "..." [--transport slack] [--dry-run]` | mints a `thread_id`, posts the clarifier, writes the `open` ledger row | +| list waiting questions | `list` (default `open`) / `list --all` / `list --parked` | JSON rows | +| inspect one row | `show ` | JSON row | +| answer on the task's behalf | `answer --answer --confirm` | destructive, audit-logged; first-answer-wins CAS | +| force-expire an open question | `expire --confirm` | destructive; flips `open`β†’`expired` | +| un-park (reopen) an expired question | `force-resume --confirm` | only acts on `expired` rows (reopens them) | + +`operator_cli.py` has different verbs (`force-expire`, `answer-on-behalf`) and +**no `show`**, and requires `--db`/`--audit-log` β€” do not use it here. + +## Preconditions + +```bash +ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120 +cd ~/orchestrator/agent-team +. .venv/bin/activate # so run-team.py uses the venv deps +# Confirm the ledger exists (Step 5 of the runbook): +python3 run-team.py list --all # [] on a fresh DB is fine +# For the live-Slack path, confirm the listener started: +journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener" +# -> "inbound Slack listener started (Socket Mode, background thread)" +``` + +Run all `run-team.py` commands from `~/orchestrator/agent-team` so they hit the +default ledger (`state/agent_team.sqlite`) and default audit log +(`state/audit.log.jsonl`). The clarifier is posted to `SLACK_CHANNEL_ID`. + +> **Resume execution model.** A won answer (rowcount 1) enqueues a `ResumeJob` +> onto the coordinator's in-process `resume_queue`; the daemon drains it on the +> next `tick()`, or the startup `recover()` re-drives any `answered`-but-unresumed +> row after a restart. So a CLI `answer --confirm` (or a live Slack answer) flips +> the ledger row to `answered`; the **running daemon** then resumes the graph. +> The demo asserts on the durable ledger status. + +--- + +## (a) Crash-safe resume β€” kill mid-wait, restart, task resumes πŸ§‘ Adam answers + +**Setup β€” drive a task to the clarifier wait.** With the daemon already running: + +```bash +python3 run-team.py start --task "demo-a: trivial scoped task" +# prints a thread_id, e.g. 3f2a... (record it as $TID_A) +python3 run-team.py list +# expect one row: status="open", a thread_id, a question_id (record as $QID_A), +# channel_ref set (Slack ts) or null if posting is dry/unavailable. +``` + +**Kill the box mid-wait, then restart:** + +```bash +sudo systemctl stop agent-team-coordinator.service +# (optionally reboot the VM here for a stronger demonstration) +sudo systemctl start agent-team-coordinator.service +journalctl -u agent-team-coordinator.service -e | tail -40 # expect "starting" + recover() sweep +``` + +**ASSERT β€” durable state survived the kill (no answer yet):** + +```bash +python3 run-team.py show $QID_A +# PASS: status == "open" (the question was NOT lost across the restart) +``` + +**πŸ§‘ Human gate β€” Adam answers after the restart (Slack OR CLI):** + +```bash +# Live Slack: Adam replies/clicks in the channel. OR via the CLI: +python3 run-team.py answer $QID_A --answer "scope: just demo, no real change" --confirm +python3 run-team.py show $QID_A +# PASS: status == "answered" +# Wait one tick (~30s, DEFAULT_POLL_INTERVAL) for the daemon to drain the resume: +journalctl -u agent-team-coordinator.service -e | tail -20 +``` + +**PASS criterion (a):** the question stayed `open` across the kill, and an answer +submitted *after* the restart drove it forward. + +--- + +## (b) Duplicate answer is a no-op πŸ§‘ Adam answers + +```bash +python3 run-team.py start --task "demo-b: duplicate-answer test" +python3 run-team.py list # record the new question_id as $QID_B (status open) +``` + +**πŸ§‘ First answer (the winning one) β€” Slack OR CLI:** + +```bash +python3 run-team.py answer $QID_B --answer "first-answer" --via "slack:U1" --confirm +# expect: "answered question (via slack:U1)" (exit 0) +``` + +**Duplicate / second answer (must lose the compare-and-set):** + +```bash +python3 run-team.py answer $QID_B --answer "second-answer" --via "github:U2" --confirm +# expect: "answer no-op: question was not 'open' ..." on stderr, exit 1 +``` + +For the **live-Slack** variant, Adam (or a second authorized owner) answers the +same message twice; the second event loses the CAS and is logged +`ignored Slack answer for question_id=... (not open: duplicate ...)`. + +**ASSERT β€” the first answer is preserved:** + +```bash +python3 run-team.py show $QID_B +# PASS: status == "answered"; answered_via == "slack:U1" (the FIRST answer) +``` + +**PASS criterion (b):** first-answer-wins (`rowcount==1`); the duplicate hits the +`BEGIN IMMEDIATE` compare-and-set in `answer_question` and is ignored +(`rowcount==0`) β€” no second resume, no overwrite. + +--- + +## (c) Past-deadline answer rejected + task parks πŸ§‘ Adam observes + +> **Code reality:** `run-team.py start` always sets the clarifier deadline from +> `DEFAULT_CLARIFY_DEADLINE = 24h`. There is **no CLI flag for a short deadline**. +> Two faithful ways to demo expiry without waiting 24h: + +### Option C1 β€” force the deadline past, let the daemon's sweep expire it (most faithful) + +```bash +python3 run-team.py start --task "demo-c: deadline test" +python3 run-team.py list # record $QID_C (status open) + +# Set this question's deadline into the past directly in the ledger: +sqlite3 state/agent_team.sqlite \ + "UPDATE pending_questions SET deadline_at='2000-01-01T00:00:00+00:00' WHERE question_id='$QID_C';" + +# Wait one daemon tick (~30s) for the deadline sweep to flip it, OR observe: +journalctl -u agent-team-coordinator.service -e | tail -20 +# expect the park ALARM line: "task parked: clarifier question expired ..." +``` + +### Option C2 β€” operator force-expire (if you do not want to touch the DB) + +```bash +python3 run-team.py start --task "demo-c: deadline test" +python3 run-team.py list # record $QID_C +python3 run-team.py expire $QID_C --confirm # destructive, audit-logged; open->expired +``` + +**ASSERT β€” the question is expired and the late answer is rejected:** + +```bash +python3 run-team.py show $QID_C +# PASS: status == "expired" + +# πŸ§‘ Adam submits a LATE answer (Slack OR CLI) β€” it must lose the CAS: +python3 run-team.py answer $QID_C --answer "too-late" --confirm +# expect: "answer no-op: question was not 'open' ..." stderr, exit 1 +python3 run-team.py show $QID_C +# PASS: status still "expired"; answer_json still NULL + +# The parked task surfaces in the parked view: +python3 run-team.py list --parked # PASS: $QID_C appears here +``` + +**Deliberate un-park (proves the operator recovery path, Β§6.6):** + +```bash +python3 run-team.py force-resume $QID_C --confirm +# expect: "force-resume: reopened expired question ; it will be re-delivered" +python3 run-team.py show $QID_C +# status flips back to "open" (reopen_question), deadline_at cleared. +``` + +**PASS criterion (c):** the expired question rejects the late answer, the task +parks rather than spins, and the operator can deliberately un-park it via +`force-resume` (which reopens only an `expired` row). + +> **`force-resume` semantics (verified):** `run-team.py force-resume` reopens an +> `expired` question (the parked case). For an `answered` question it records +> intent and reports the recovery sweep will resume it (no mutation). For +> `open`/`superseded`/absent it is a no-op exit 1. It does **not** supersede. + +--- + +## (d) Two concurrent tasks resume independently πŸ§‘ Adam answers + +```bash +python3 run-team.py start --task "demo-d task A" # record $TID_A2, then: +python3 run-team.py start --task "demo-d task B" # record $TID_B2 +python3 run-team.py list +# expect TWO open rows with DISTINCT thread_id AND distinct question_id. +# Record $QID_A2 and $QID_B2 β€” match by thread_id. +``` + +**(Optional) restart the daemon first** to also show concurrent tasks survive a +restart, then answer. + +**πŸ§‘ Answer the SECOND task first, with a distinct answer, then the first +(Slack OR CLI):** + +```bash +python3 run-team.py answer $QID_B2 --answer "answer-for-B" --via "slack:U2" --confirm +python3 run-team.py answer $QID_A2 --answer "answer-for-A" --via "slack:U1" --confirm +# Wait one daemon tick (~30s) for both resumes to drain. +``` + +**ASSERT β€” each task carries its OWN answer; no cross-talk:** + +```bash +python3 run-team.py show $QID_A2 +# PASS: status "answered", answered_via "slack:U1", thread_id == $TID_A2 +python3 run-team.py show $QID_B2 +# PASS: status "answered", answered_via "slack:U2", thread_id == $TID_B2 +python3 run-team.py list --all +# PASS: the two rows resolved on their own thread_id; no cross-contamination. +``` + +**PASS criterion (d):** two concurrently-suspended tasks each resumed to their +own `thread_id` with their own answer β€” answering B before A did not misroute, +and the per-thread single-flight guard kept them independent. + +--- + +## Acceptance + +P1 is accepted only when **(a), (b), (c), and (d) all pass** on the live box. +Record the four `show` outputs (or `journalctl` excerpts) as evidence. Then +complete the PROVISIONING-RUNBOOK post-session definition-of-done (memory + +Confluence + the mandatory `/sh-security-review` on the Slack inbound listener). + +## What can only be verified on the live box + +- That the daemon actually drains the resume and the LangGraph checkpoint + advances (asserted via ledger status + journalctl; the graph-state advance is + observable only on the box). +- The live **Slack** post/answer round-trip end-to-end (the listener is wired β€” + D-1 fixed β€” but the live Socket Mode socket + a real `SLACK_BOT_TOKEN` / + `SLACK_APP_TOKEN` / channel membership are only present on the box). +- The exact `channel_ref` value (Slack message `ts`) β€” depends on a live Slack + post succeeding. diff --git a/docs/provisioning/PROVISIONING-RUNBOOK.md b/docs/provisioning/PROVISIONING-RUNBOOK.md new file mode 100644 index 0000000..d2e0328 --- /dev/null +++ b/docs/provisioning/PROVISIONING-RUNBOOK.md @@ -0,0 +1,402 @@ +# PROVISIONING RUNBOOK β€” R720 agent-team Plane-2 coordinator + +This is the ordered command sequence for the **operator-present** provisioning +session that stands up the always-on `agent-team-coordinator` daemon on the +`sh-secrev` R720 VM. Derived from `agent-team/DEPLOY-R720.md`, +`security-review/DEPLOY-R720.md`, and `docs/r720-agent-team-design.md` +(Β§3.3.1 / Β§7 / Β§7.1). + +> ## State of this runbook +> +> The build is **merged and final**: all six Plane-1 checkers +> (`aws-posture`, `compliance-drift`, `confluence-doc`, `dependency-cve`, +> `doc-drift`, `plan-groomer`), the Tier-3 dep-bump **fixer** +> (`run-team.py fix --dry-run`), and the Plane-1β†’Plane-2 **P5 cross-plane loop** +> (`run-team.py intake-checker`) exist in the tree. The three deploy-correctness +> bugs the provisioning-prep audit found are **FIXED** (see DEPLOY-AUDIT.md): +> +> - **D-1 RESOLVED** β€” `Coordinator.serve()` now starts the inbound Slack +> `SlackListener` concurrently with the tick/drain loop when Slack is the live +> transport AND `SLACK_APP_TOKEN` is set. **The demo can use the live Slack +> answer path** (P1-DEMO-SCRIPT.md), or the operator-CLI `answer` path. +> - **D-2 RESOLVED** β€” the systemd unit loads `~/orchestrator/.env` (for the P2 +> GPT-4.1 review loop's provider key) in addition to `~/secrev.env`. +> - **D-7 RESOLVED** β€” the unit's `ExecStart` points at the agent-team venv +> interpreter, not the system `python3`. +> +> Still open as provisioning notes (not blockers): **D-4/D-5** (agent-team +> runtime deps are installed ad-hoc into the venv and are not pinned in +> `requirements.txt`), and the **operator-CLI divergence** (`run-team.py` vs +> `agent_team/operator_cli.py` have different verb names β€” use `run-team.py`). + +--- + +## Scope and ground rules + +**Hard rules (carried from the design and global instructions):** + +- **Snapshot before any stateful change** (design Β§7; `feedback_ec2_replacement_snapshot`). +- **Every stateful step has an exercised rollback.** +- Secrets use **placeholder names only**; real values are entered by the operator + at the box and never echoed into shell history. +- The box is **read-only / subscription-OAuth only**. **No `ANTHROPIC_API_KEY`** + on this host (it would silently win over OAuth and meter to API rates β€” + `billing.claude_invoke` pops it defensively, but it must not be present). +- **No IAM / OIDC is involved in the coordinator deploy** (P1/P2). IAM enters + only at the **P3-live flip** (see the dedicated section), which is gated on the + mandatory GPT-4.1 cross-review + `/sh-security-review`. + +**Legend per step:** + +- πŸ§‘ **OPERATOR-REQUIRED** β€” needs the human (snapshot, secrets, the live human + gate, go/no-go). Cannot be automated. +- πŸ€– **MECHANICAL** β€” deterministic; an operator runs it but it needs no judgment. + +## Host facts (from both DEPLOY-R720.md files) + +- Hypervisor: R720 at `10.10.60.40` (Windows Server 2022, Hyper-V). +- VM: `sh-secrev`, Ubuntu 24.04, **4GB / 2 vCPU / 40GB** dynamic vhdx, `10.10.60.120`. +- Reach: `ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120` (key-only, NOPASSWD sudo). +- Repo on the box: `~/orchestrator/` (rsync from the Mac, **NOT** a git clone). + The package lives at `~/orchestrator/agent-team/`. +- Shares `~/secrev.env` (mode 600) with the secrev sweep, and `~/orchestrator/.env` + (mode 600) for non-Claude provider keys (same files the secrev unit loads). + +--- + +## STEP 1 β€” Snapshot the VM πŸ§‘ OPERATOR-REQUIRED + +**On the R720 host (Hyper-V), before anything else.** This is the one-command +undo for every change below. + +```powershell +# On the R720 Windows host (PowerShell, as admin): +Checkpoint-VM -Name sh-secrev -SnapshotName "pre-agent-team-coordinator-$(Get-Date -Format yyyyMMdd-HHmm)" +Get-VMSnapshot -VMName sh-secrev # confirm the checkpoint exists +``` + +**ROLLBACK (whole session):** +```powershell +Restore-VMSnapshot -VMName sh-secrev -Name "" -Confirm:$false +Start-VM -Name sh-secrev +``` + +> Do not proceed until the checkpoint is confirmed present. + +--- + +## STEP 2 β€” Rsync the repo to the box πŸ€– MECHANICAL (verify manifest πŸ§‘) + +**From the Mac.** Same pattern/excludes as the secrev deploy. The team **scans +the same `~/repo-mirrors` corpus secrev already maintains** β€” this rsync ships +*code*, not mirrors. + +```bash +# From the Mac (sync the canonical repo, not a worktree): +rsync -av --exclude .env --exclude .venv --exclude .git --exclude .claude \ + ~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/ +``` + +Surface that must be on the box: + +| Path (under `~/orchestrator/`) | Why it must ship | +|---|---| +| `agent-team/run-team.py` | the operator entry CLI | +| `agent-team/agent_team/**` | the package (coordinator, graph, ledger, transports, slack_listener, fixer) | +| `agent-team/systemd/agent-team-coordinator.service` | the daemon unit | +| `run.py` + the orchestrator package | the GPT-4.1 cross_reviewer the P2 review loop shells | +| `requirements.txt` | pin reference for langgraph / checkpoint-sqlite | +| `security-review/lib/**` | shared sweep substrate (Phase 0) | +| `security-review/checkers/**` | the six Plane-1 checkers + fixtures | + +**ROLLBACK:** rsync is additive; restore the Step-1 snapshot to revert code state. + +--- + +## STEP 3 β€” Write secrets πŸ§‘ OPERATOR-REQUIRED + +**On the VM.** The coordinator unit loads **two** EnvironmentFiles (both +optional via the leading `-`): `~/secrev.env` (agent-team runtime keys) and +`~/orchestrator/.env` (the non-Claude provider key for the P2 review loop). +**Placeholder names only β€” the operator pastes real values.** Do not echo real +tokens into shell history (use an editor or `read -s`). + +`~/secrev.env` (mode 600) β€” append the agent-team keys: + +```bash +ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120 +# Edit ~/secrev.env (mode 600) and add β€” values are placeholders: +# CLAUDE_CODE_OAUTH_TOKEN= # from `claude setup-token` +# SLACK_BOT_TOKEN= # xoxb-..., chat:write (posts questions) +# SLACK_APP_TOKEN= # xapp-..., connections:write (Socket Mode inbound) +# SLACK_CHANNEL_ID= # C0..., target clarifier channel +# AGENT_TEAM_SLACK_OWNER_IDS= # comma-separated U... ids of authorized answerers +chmod 600 ~/secrev.env +``` + +`~/orchestrator/.env` (mode 600) β€” the non-Claude provider key for the GPT-4.1 +review loop (same file secrev uses; if it already exists with the key, leave it): + +```bash +# OPENAI_API_KEY= # (or the provider key cross_reviewer/GPT-4.1 needs) +chmod 600 ~/orchestrator/.env +``` + +**Contract notes (verified against the code β€” see DEPLOY-AUDIT.md):** + +- `SLACK_CHANNEL_ID` is correct β€” `run-team.py _build_transport` reads exactly + `os.environ.get("SLACK_CHANNEL_ID")`. Do **not** use `SLACK_CHANNEL`. +- `SLACK_APP_TOKEN` is now **read by the daemon**: `Coordinator.serve()` starts + the inbound `SlackListener` when the transport is live Slack AND `SLACK_APP_TOKEN` + is set. Without it the daemon still runs (posts + expires) but never hears Slack + replies β€” Slack stays optional by design. +- `AGENT_TEAM_SLACK_OWNER_IDS` **fails closed** (AUTHZ-01): if unset/empty the + listener rejects **every** answer. It must be set for the live human gate. +- **CRITICAL:** confirm `ANTHROPIC_API_KEY` is NOT present: + ```bash + grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # must print 0 for both + ``` + +**ROLLBACK:** strip exactly the appended keys (keep secrev keys intact), or +restore the Step-1 snapshot. Do NOT blindly truncate β€” secrev keys live here too. + +--- + +## STEP 4 β€” Create the venv + install pip deps πŸ€– MECHANICAL + +**On the VM.** A dedicated venv under `agent-team/.venv` (excluded from rsync). +The systemd unit's `ExecStart` points at **this venv's interpreter** (D-7 fixed), +so the deps MUST land here. + +```bash +cd ~/orchestrator/agent-team +python3 -m venv .venv +. .venv/bin/activate +pip install langgraph==1.1.10 langgraph-checkpoint-sqlite==3.1.0 \ + claude-agent-sdk slack_sdk slack_bolt requests +``` + +> **Pin note (D-5):** `requirements.txt` pins `langgraph==1.1.10` / +> `langgraph-checkpoint-sqlite==3.1.0`; match those exactly here. The other +> runtime deps (`claude-agent-sdk`, `slack_sdk`, `slack_bolt`, `requests`) are +> not yet in `requirements.txt` (D-5 open) β€” installed ad-hoc here. `requests` +> is required by the GitHub transport/intake (D-4). `anthropic` is **not** +> installed (only the opt-in `api` billing mode needs it). +> `slack_bolt` is now actually exercised (D-1 fixed: the daemon starts the +> Socket Mode listener). + +**Verify the imports resolve (using the venv interpreter the unit will use):** +```bash +.venv/bin/python -c "import langgraph, langgraph.checkpoint.sqlite, slack_sdk, slack_bolt, requests; print('deps ok')" +.venv/bin/python -c "import claude_agent_sdk; print('agent-sdk ok')" +``` + +**ROLLBACK:** `deactivate 2>/dev/null; rm -rf ~/orchestrator/agent-team/.venv` + +--- + +## STEP 5 β€” Initialize the durable ledger DB πŸ€– MECHANICAL + +**On the VM, venv active.** Idempotent; creates `state/agent_team.sqlite` with +the `pending_questions` + `budget_ledger` + `schema_meta` tables (LangGraph +`SqliteSaver` creates its own tables in the same file on first run). + +```bash +cd ~/orchestrator/agent-team +. .venv/bin/activate +python3 run-team.py init-db +# Expect: "initialized ledger DB at .../state/agent_team.sqlite" +ls -l state/ # agent_team.sqlite present; state/ is gitignored +``` + +**ROLLBACK (reset the ledger only):** +```bash +cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak} +rm ~/orchestrator/agent-team/state/agent_team.sqlite* +# re-run `python3 run-team.py init-db` to recreate empty tables. +``` + +--- + +## STEP 6 β€” Install + start the systemd unit πŸ€– MECHANICAL (go/no-go πŸ§‘) + +**On the VM, as root.** Installs the long-running coordinator daemon. Keep the +**hardening as-shipped**: `NoNewPrivileges`, `ProtectSystem=full`, +`ProtectHome=read-only` + `ReadWritePaths=.../agent-team/state` (locked decision). + +```bash +sudo cp ~/orchestrator/agent-team/systemd/agent-team-coordinator.service /etc/systemd/system/ +sudo systemctl daemon-reload +sudo systemctl enable --now agent-team-coordinator.service +systemctl status agent-team-coordinator.service +journalctl -u agent-team-coordinator.service -e -f +# expect: "agent-team coordinator starting" +# and (if Slack + SLACK_APP_TOKEN provisioned): +# "inbound Slack listener started (Socket Mode, background thread)" +# or otherwise: +# "inbound Slack listener not started (... or SLACK_APP_TOKEN is unset) ..." +``` + +The unit (post-fix) runs +`/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve` from +`WorkingDirectory=/home/adam/orchestrator/agent-team` as `User=adam`, loading +**both** `EnvironmentFile=-/home/adam/secrev.env` and +`EnvironmentFile=-/home/adam/orchestrator/.env`, `Restart=on-failure`. + +> **D-7 fixed:** `ExecStart` now resolves the venv interpreter, so the Step-4 +> deps are on the path. **D-1 fixed:** `serve` starts the inbound `SlackListener` +> when Slack + app token are provisioned. **Go/no-go:** confirm the journal shows +> the listener line you expect for your transport choice before Step 8. + +**ROLLBACK (exercised β€” tear the unit down cleanly):** +```bash +sudo systemctl disable --now agent-team-coordinator.service +sudo rm /etc/systemd/system/agent-team-coordinator.service +sudo systemctl daemon-reload +systemctl status agent-team-coordinator.service # should report "could not be found" +``` +The ledger under `state/` is untouched. Step-1 snapshot is the whole-host fallback. + +--- + +## STEP 7 β€” Verify the Slack inbound listener πŸ§‘ OPERATOR-REQUIRED + +The clarifier gate is two halves: **outbound** (post the question β€” done by +`serve`/`start` via the Slack poster) and **inbound** (receive Adam's answer β€” +the `SlackListener` over Socket Mode). + +> **D-1 RESOLVED.** `Coordinator.serve()` now constructs and starts +> `SlackListener` on a background daemon thread, concurrently with the tick/drain +> loop, **when** the live transport is a `SlackTransport` AND `SLACK_APP_TOKEN` +> is set. It shares the coordinator's own transport, ledger, and resume queue, +> and stops cleanly on shutdown. The AUTHZ-01 owner allowlist + the open-status +> compare-and-set are unchanged β€” the listener still fails closed on an empty +> `AGENT_TEAM_SLACK_OWNER_IDS`. + +**Confirm the live inbound path is up:** + +```bash +# In the journal (Step 6) expect the "inbound Slack listener started" line. +# Verify the two gating vars are present in the unit's environment: +grep -c SLACK_APP_TOKEN ~/secrev.env # 1 +grep -c AGENT_TEAM_SLACK_OWNER_IDS ~/secrev.env # 1 (else the gate rejects all answers) +``` + +If you are **not** using Slack as the transport (or are deliberately running +without the app token), the daemon runs the maintenance loop only and the demo +uses the operator-CLI `answer` path β€” both are valid (P1-DEMO-SCRIPT.md). + +**ROLLBACK:** none needed β€” verification only, no host state change. + +--- + +## STEP 8 β€” Live P1 four-criteria acceptance demo πŸ§‘ OPERATOR-REQUIRED + +Run **P1-DEMO-SCRIPT.md** in full. All four Β§3.3.1 exit criteria must pass: +(a) crash-safe resume, (b) duplicate-answer no-op, (c) post-deadline rejection + +park, (d) two concurrent tasks resume independently. With D-1 fixed you may +exercise the **live Slack answer path** for (b)/(d); the operator-CLI `answer` +path remains available and exercises the identical compare-and-set. **Do not +accept P1 until all four pass.** + +**ROLLBACK:** the demo writes only ledger rows under `state/`; reset via the +Step-5 ledger rollback, or restore the Step-1 snapshot, then re-run. + +--- + +## STEP 9 β€” Plane-1 checker live dry-runs πŸ§‘ OPERATOR-REQUIRED + +The six checkers live under `security-review/checkers/`: +`aws-posture.sh`, `compliance-drift.sh`, `confluence-doc.sh`, +`dependency-cve.sh`, `doc-drift.sh`, `plan-groomer.sh`. All run **report + +ALARM-only** (design D3): a clean run posts nothing and lands a mode-600 report; +no auto-Jira/Notion writes. + +- Dry-run each checker against `~/repo-mirrors` in report-only mode; confirm a + clean run posts nothing and writes a mode-600 report. Confirm the exact + invocation against the checker scripts and the shared + `security-review/lib/` substrate on the box. +- Confirm the **canary suite** runs first (a planted-fault miss is a COMPLACENCY + ALARM and that role is skipped β€” design Β§6.4), and the **coverage rotation** + pointer advances (a slipped role is a COVERAGE ALARM, deferred-not-dropped). + +**ROLLBACK:** checkers are read-only over the mirror corpus; a dry-run produces +only a report file. Remove the report dir to revert; no host state change. + +--- + +## STEP 10 β€” P5 cross-plane loop dry-run πŸ§‘ OPERATOR-REQUIRED + +The Plane-1β†’Plane-2 loop turns confirmed checker findings into pipeline tasks: + +```bash +cd ~/orchestrator/agent-team && . .venv/bin/activate +# Read one or more checker report JSONs and start one task per confirmed +# at/above-threshold finding (default threshold: high). --dry-run posts nowhere. +python3 run-team.py intake-checker --report --threshold high --dry-run +``` + +The Tier-3 dep-bump **fixer** is dry-run only on the box (it holds no write +token, D2): + +```bash +python3 run-team.py fix --report security-review/<...>/dependency-cve.json \ + --finding-id --task-id --dry-run +# prints the fix spec + minimal bump patch + the CI workflow_dispatch inputs; +# dispatches NOTHING. Live dispatch is the P3-live flip below. +``` + +**ROLLBACK:** both are read-only / dry-run (no dispatch, no apply); `intake-checker` +de-dup is in-memory per process. No host state to revert beyond ledger rows from +a non-dry-run intake (Step-5 ledger rollback). + +--- + +## STEP 11 β€” Wire the schedule πŸ§‘ OPERATOR-REQUIRED + +The coordinator daemon (Step 6) is **always-on**, not timer-driven. The +**per-checker timers** (and the design's "shared timer with secrev", Β§8) attach +here. The existing `sea-haven-secrev.timer` (OnCalendar `02:00`, Persistent) is +untouched. Add a checker timer only after that checker is dry-run-validated +(Step 9). + +**ROLLBACK:** each timer gets its own `systemctl disable --now .timer` + `rm`. + +--- + +## The P3-live flip (deferred; IAM + GitHub App gated) πŸ§‘ OPERATOR-REQUIRED + +P3 (the buildβ†’verify apply-and-open-draft-PR loop) is **opt-in and inert** in +this deploy: `run-team.py` / `serve` pass `build_verify_wiring=None`, so no P3 +subgraph is assembled. Flipping it live is a **separate, gated** provisioning +session, not part of the coordinator deploy: + +1. **Mandatory reviews first.** The CI trust-boundary + OIDC IAM change is a + breaking IAM change β†’ **GPT-4.1 cross-review** (global instructions) AND + `/sh-security-review` on the apply/verify surface. Do not flip without both. +2. **Provision the GitHub App** for the trusted apply path (the App that opens + the draft PR), and the **`agent-apply` GitHub Actions environment** that holds + the apply path's scoped permissions. +3. **Bind the live buildβ†’verify wiring** via + `agent_team.coordinator.gated_build_verify_wiring(...)` (the read-only CI + result fetcher + the real diff builder) β€” the seam a leaf calls *after* the + gate clears. The CI fetcher is read-only and fails closed (missing token / + 404 / auth failure β†’ `None` β†’ the gate BLOCKs and the task parks). +4. **Set the apply env vars** the live path reads (the read-only CI-result token + and the dispatch target), then re-run the fixer **without** `--dry-run` only + once the dispatcher is bound. + +Until every step above is done, the box dispatches/applies nothing. + +--- + +## Post-session definition-of-done (design Β§7, global instructions) + +- [ ] All four P1 criteria demonstrated live (Step 8). +- [ ] `project_r720_agent_team` memory created/updated. +- [ ] Confluence "AWS Architecture Map" / IT host inventory updated to show + `sh-secrev` now also hosts the always-on agent-team coordinator daemon. +- [ ] `/sh-security-review` run on the Slack inbound listener surface + (auth + untrusted-input; mandatory) β€” flag outstanding if not run. +- [ ] OPERATOR-RUNBOOK.md reviewed by whoever holds the pager. +- [ ] Snapshot retained until the daemon runs clean for one full cycle, then pruned.