# PROVISIONING RUNBOOK — R720 agent-team Plane-2 coordinator This is the ordered command sequence for the **operator-present** provisioning session that stands up the always-on `agent-team-coordinator` daemon on the `sh-secrev` R720 VM. Derived from `agent-team/DEPLOY-R720.md`, `security-review/DEPLOY-R720.md`, and `docs/r720-agent-team-design.md` (§3.3.1 / §7 / §7.1). > ## State of this runbook > > The build is **merged and final**: all six Plane-1 checkers > (`aws-posture`, `compliance-drift`, `confluence-doc`, `dependency-cve`, > `doc-drift`, `plan-groomer`), the Tier-3 dep-bump **fixer** > (`run-team.py fix --dry-run`), and the Plane-1→Plane-2 **P5 cross-plane loop** > (`run-team.py intake-checker`) exist in the tree. The three deploy-correctness > bugs the provisioning-prep audit found are **FIXED** (see DEPLOY-AUDIT.md): > > - **D-1 RESOLVED** — `Coordinator.serve()` now starts the inbound Slack > `SlackListener` concurrently with the tick/drain loop when Slack is the live > transport AND `SLACK_APP_TOKEN` is set. **The demo can use the live Slack > answer path** (P1-DEMO-SCRIPT.md), or the operator-CLI `answer` path. > - **D-2 RESOLVED** — the systemd unit loads `~/orchestrator/.env` (for the P2 > GPT-4.1 review loop's provider key) in addition to `~/secrev.env`. > - **D-7 RESOLVED** — the unit's `ExecStart` points at the agent-team venv > interpreter, not the system `python3`. > > Still open as provisioning notes (not blockers): **D-4/D-5** (agent-team > runtime deps are installed ad-hoc into the venv and are not pinned in > `requirements.txt`), and the **operator-CLI divergence** (`run-team.py` vs > `agent_team/operator_cli.py` have different verb names — use `run-team.py`). --- ## Scope and ground rules **Hard rules (carried from the design and global instructions):** - **Snapshot before any stateful change** (design §7; `feedback_ec2_replacement_snapshot`). - **Every stateful step has an exercised rollback.** - Secrets use **placeholder names only**; real values are entered by the operator at the box and never echoed into shell history. - The box is **read-only / subscription-OAuth only**. **No `ANTHROPIC_API_KEY`** on this host (it would silently win over OAuth and meter to API rates — `billing.claude_invoke` pops it defensively, but it must not be present). - **No IAM / OIDC is involved in the coordinator deploy** (P1/P2). IAM enters only at the **P3-live flip** (see the dedicated section), which is gated on the mandatory GPT-4.1 cross-review + `/sh-security-review`. **Legend per step:** - 🧑 **OPERATOR-REQUIRED** — needs the human (snapshot, secrets, the live human gate, go/no-go). Cannot be automated. - 🤖 **MECHANICAL** — deterministic; an operator runs it but it needs no judgment. ## Host facts (from both DEPLOY-R720.md files) - Hypervisor: R720 at `10.10.60.40` (Windows Server 2022, Hyper-V). - VM: `sh-secrev`, Ubuntu 24.04, **4GB / 2 vCPU / 40GB** dynamic vhdx, `10.10.60.120`. - Reach: `ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120` (key-only, NOPASSWD sudo). - Repo on the box: `~/orchestrator/` (rsync from the Mac, **NOT** a git clone). The package lives at `~/orchestrator/agent-team/`. - Shares `~/secrev.env` (mode 600) with the secrev sweep, and `~/orchestrator/.env` (mode 600) for non-Claude provider keys (same files the secrev unit loads). --- ## STEP 1 — Snapshot the VM 🧑 OPERATOR-REQUIRED **On the R720 host (Hyper-V), before anything else.** This is the one-command undo for every change below. ```powershell # On the R720 Windows host (PowerShell, as admin): Checkpoint-VM -Name sh-secrev -SnapshotName "pre-agent-team-coordinator-$(Get-Date -Format yyyyMMdd-HHmm)" Get-VMSnapshot -VMName sh-secrev # confirm the checkpoint exists ``` **ROLLBACK (whole session):** ```powershell Restore-VMSnapshot -VMName sh-secrev -Name "" -Confirm:$false Start-VM -Name sh-secrev ``` > Do not proceed until the checkpoint is confirmed present. --- ## STEP 2 — Rsync the repo to the box 🤖 MECHANICAL (verify manifest 🧑) **From the Mac.** Same pattern/excludes as the secrev deploy. The team **scans the same `~/repo-mirrors` corpus secrev already maintains** — this rsync ships *code*, not mirrors. ```bash # From the Mac (sync the canonical repo, not a worktree): rsync -av --exclude .env --exclude .venv --exclude .git --exclude .claude \ ~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/ ``` Surface that must be on the box: | Path (under `~/orchestrator/`) | Why it must ship | |---|---| | `agent-team/run-team.py` | the operator entry CLI | | `agent-team/agent_team/**` | the package (coordinator, graph, ledger, transports, slack_listener, fixer) | | `agent-team/systemd/agent-team-coordinator.service` | the daemon unit | | `run.py` + the orchestrator package | the GPT-4.1 cross_reviewer the P2 review loop shells | | `requirements.txt` | pin reference for langgraph / checkpoint-sqlite | | `security-review/lib/**` | shared sweep substrate (Phase 0) | | `security-review/checkers/**` | the six Plane-1 checkers + fixtures | **ROLLBACK:** rsync is additive; restore the Step-1 snapshot to revert code state. --- ## STEP 3 — Write secrets 🧑 OPERATOR-REQUIRED **On the VM.** The coordinator unit loads **two** EnvironmentFiles (both optional via the leading `-`): `~/secrev.env` (agent-team runtime keys) and `~/orchestrator/.env` (the non-Claude provider key for the P2 review loop). **Placeholder names only — the operator pastes real values.** Do not echo real tokens into shell history (use an editor or `read -s`). `~/secrev.env` (mode 600) — append the agent-team keys: ```bash ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120 # Edit ~/secrev.env (mode 600) and add — values are placeholders: # CLAUDE_CODE_OAUTH_TOKEN= # from `claude setup-token` # SLACK_BOT_TOKEN= # xoxb-..., chat:write (posts questions) # SLACK_APP_TOKEN= # xapp-..., connections:write (Socket Mode inbound) # SLACK_CHANNEL_ID= # C0..., target clarifier channel # AGENT_TEAM_SLACK_OWNER_IDS= # comma-separated U... ids of authorized answerers chmod 600 ~/secrev.env ``` `~/orchestrator/.env` (mode 600) — the non-Claude provider key for the GPT-4.1 review loop (same file secrev uses; if it already exists with the key, leave it): ```bash # OPENAI_API_KEY= # (or the provider key cross_reviewer/GPT-4.1 needs) chmod 600 ~/orchestrator/.env ``` **Contract notes (verified against the code — see DEPLOY-AUDIT.md):** - `SLACK_CHANNEL_ID` is correct — `run-team.py _build_transport` reads exactly `os.environ.get("SLACK_CHANNEL_ID")`. Do **not** use `SLACK_CHANNEL`. - `SLACK_APP_TOKEN` is now **read by the daemon**: `Coordinator.serve()` starts the inbound `SlackListener` when the transport is live Slack AND `SLACK_APP_TOKEN` is set. Without it the daemon still runs (posts + expires) but never hears Slack replies — Slack stays optional by design. - `AGENT_TEAM_SLACK_OWNER_IDS` **fails closed** (AUTHZ-01): if unset/empty the listener rejects **every** answer. It must be set for the live human gate. - **CRITICAL:** confirm `ANTHROPIC_API_KEY` is NOT present: ```bash grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # must print 0 for both ``` **ROLLBACK:** strip exactly the appended keys (keep secrev keys intact), or restore the Step-1 snapshot. Do NOT blindly truncate — secrev keys live here too. --- ## STEP 4 — Create the venv + install pip deps 🤖 MECHANICAL **On the VM.** A dedicated venv under `agent-team/.venv` (excluded from rsync). The systemd unit's `ExecStart` points at **this venv's interpreter** (D-7 fixed), so the deps MUST land here. ```bash cd ~/orchestrator/agent-team python3 -m venv .venv . .venv/bin/activate pip install langgraph==1.2.5 langgraph-checkpoint-sqlite==3.1.0 \ claude-agent-sdk slack_sdk slack_bolt requests ``` > **Pin note (D-5):** `requirements.txt` pins `langgraph==1.2.5` / > `langgraph-checkpoint-sqlite==3.1.0`; match those exactly here. The other > runtime deps (`claude-agent-sdk`, `slack_sdk`, `slack_bolt`, `requests`) are > not yet in `requirements.txt` (D-5 open) — installed ad-hoc here. `requests` > is required by the GitHub transport/intake (D-4). `anthropic` is **not** > installed (only the opt-in `api` billing mode needs it). > `slack_bolt` is now actually exercised (D-1 fixed: the daemon starts the > Socket Mode listener). **Verify the imports resolve (using the venv interpreter the unit will use):** ```bash .venv/bin/python -c "import langgraph, langgraph.checkpoint.sqlite, slack_sdk, slack_bolt, requests; print('deps ok')" .venv/bin/python -c "import claude_agent_sdk; print('agent-sdk ok')" ``` **ROLLBACK:** `deactivate 2>/dev/null; rm -rf ~/orchestrator/agent-team/.venv` --- ## STEP 5 — Initialize the durable ledger DB 🤖 MECHANICAL **On the VM, venv active.** Idempotent; creates `state/agent_team.sqlite` with the `pending_questions` + `budget_ledger` + `schema_meta` tables (LangGraph `SqliteSaver` creates its own tables in the same file on first run). ```bash cd ~/orchestrator/agent-team . .venv/bin/activate python3 run-team.py init-db # Expect: "initialized ledger DB at .../state/agent_team.sqlite" ls -l state/ # agent_team.sqlite present; state/ is gitignored ``` **ROLLBACK (reset the ledger only):** ```bash cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak} rm ~/orchestrator/agent-team/state/agent_team.sqlite* # re-run `python3 run-team.py init-db` to recreate empty tables. ``` --- ## STEP 6 — Install + start the systemd unit 🤖 MECHANICAL (go/no-go 🧑) **On the VM, as root.** Installs the long-running coordinator daemon. Keep the **hardening as-shipped**: `NoNewPrivileges`, `ProtectSystem=full`, `ProtectHome=read-only` + `ReadWritePaths=.../agent-team/state` (locked decision). ```bash sudo cp ~/orchestrator/agent-team/systemd/agent-team-coordinator.service /etc/systemd/system/ sudo systemctl daemon-reload sudo systemctl enable --now agent-team-coordinator.service systemctl status agent-team-coordinator.service journalctl -u agent-team-coordinator.service -e -f # expect: "agent-team coordinator starting" # and (if Slack + SLACK_APP_TOKEN provisioned): # "inbound Slack listener started (Socket Mode, background thread)" # or otherwise: # "inbound Slack listener not started (... or SLACK_APP_TOKEN is unset) ..." ``` The unit (post-fix) runs `/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve` from `WorkingDirectory=/home/adam/orchestrator/agent-team` as `User=adam`, loading **both** `EnvironmentFile=-/home/adam/secrev.env` and `EnvironmentFile=-/home/adam/orchestrator/.env`, `Restart=on-failure`. > **D-7 fixed:** `ExecStart` now resolves the venv interpreter, so the Step-4 > deps are on the path. **D-1 fixed:** `serve` starts the inbound `SlackListener` > when Slack + app token are provisioned. **Go/no-go:** confirm the journal shows > the listener line you expect for your transport choice before Step 8. **ROLLBACK (exercised — tear the unit down cleanly):** ```bash sudo systemctl disable --now agent-team-coordinator.service sudo rm /etc/systemd/system/agent-team-coordinator.service sudo systemctl daemon-reload systemctl status agent-team-coordinator.service # should report "could not be found" ``` The ledger under `state/` is untouched. Step-1 snapshot is the whole-host fallback. --- ## STEP 7 — Verify the Slack inbound listener 🧑 OPERATOR-REQUIRED The clarifier gate is two halves: **outbound** (post the question — done by `serve`/`start` via the Slack poster) and **inbound** (receive Adam's answer — the `SlackListener` over Socket Mode). > **D-1 RESOLVED.** `Coordinator.serve()` now constructs and starts > `SlackListener` on a background daemon thread, concurrently with the tick/drain > loop, **when** the live transport is a `SlackTransport` AND `SLACK_APP_TOKEN` > is set. It shares the coordinator's own transport, ledger, and resume queue, > and stops cleanly on shutdown. The AUTHZ-01 owner allowlist + the open-status > compare-and-set are unchanged — the listener still fails closed on an empty > `AGENT_TEAM_SLACK_OWNER_IDS`. **Confirm the live inbound path is up:** ```bash # In the journal (Step 6) expect the "inbound Slack listener started" line. # Verify the two gating vars are present in the unit's environment: grep -c SLACK_APP_TOKEN ~/secrev.env # 1 grep -c AGENT_TEAM_SLACK_OWNER_IDS ~/secrev.env # 1 (else the gate rejects all answers) ``` If you are **not** using Slack as the transport (or are deliberately running without the app token), the daemon runs the maintenance loop only and the demo uses the operator-CLI `answer` path — both are valid (P1-DEMO-SCRIPT.md). **ROLLBACK:** none needed — verification only, no host state change. --- ## STEP 8 — Live P1 four-criteria acceptance demo 🧑 OPERATOR-REQUIRED Run **P1-DEMO-SCRIPT.md** in full. All four §3.3.1 exit criteria must pass: (a) crash-safe resume, (b) duplicate-answer no-op, (c) post-deadline rejection + park, (d) two concurrent tasks resume independently. With D-1 fixed you may exercise the **live Slack answer path** for (b)/(d); the operator-CLI `answer` path remains available and exercises the identical compare-and-set. **Do not accept P1 until all four pass.** **ROLLBACK:** the demo writes only ledger rows under `state/`; reset via the Step-5 ledger rollback, or restore the Step-1 snapshot, then re-run. --- ## STEP 9 — Plane-1 checker live dry-runs 🧑 OPERATOR-REQUIRED The six checkers live under `security-review/checkers/`: `aws-posture.sh`, `compliance-drift.sh`, `confluence-doc.sh`, `dependency-cve.sh`, `doc-drift.sh`, `plan-groomer.sh`. All run **report + ALARM-only** (design D3): a clean run posts nothing and lands a mode-600 report; no auto-Jira/Notion writes. - Dry-run each checker against `~/repo-mirrors` in report-only mode; confirm a clean run posts nothing and writes a mode-600 report. Confirm the exact invocation against the checker scripts and the shared `security-review/lib/` substrate on the box. - Confirm the **canary suite** runs first (a planted-fault miss is a COMPLACENCY ALARM and that role is skipped — design §6.4), and the **coverage rotation** pointer advances (a slipped role is a COVERAGE ALARM, deferred-not-dropped). **ROLLBACK:** checkers are read-only over the mirror corpus; a dry-run produces only a report file. Remove the report dir to revert; no host state change. --- ## STEP 10 — P5 cross-plane loop dry-run 🧑 OPERATOR-REQUIRED The Plane-1→Plane-2 loop turns confirmed checker findings into pipeline tasks: ```bash cd ~/orchestrator/agent-team && . .venv/bin/activate # Read one or more checker report JSONs and start one task per confirmed # at/above-threshold finding (default threshold: high). --dry-run posts nowhere. python3 run-team.py intake-checker --report --threshold high --dry-run ``` The Tier-3 dep-bump **fixer** is dry-run only on the box (it holds no write token, D2): ```bash python3 run-team.py fix --report security-review/<...>/dependency-cve.json \ --finding-id --task-id --dry-run # prints the fix spec + minimal bump patch + the CI workflow_dispatch inputs; # dispatches NOTHING. Live dispatch is the P3-live flip below. ``` **ROLLBACK:** both are read-only / dry-run (no dispatch, no apply); `intake-checker` de-dup is in-memory per process. No host state to revert beyond ledger rows from a non-dry-run intake (Step-5 ledger rollback). --- ## STEP 11 — Wire the schedule 🧑 OPERATOR-REQUIRED The coordinator daemon (Step 6) is **always-on**, not timer-driven. The **per-checker timers** (and the design's "shared timer with secrev", §8) attach here. The existing `sea-haven-secrev.timer` (OnCalendar `02:00`, Persistent) is untouched. Add a checker timer only after that checker is dry-run-validated (Step 9). **ROLLBACK:** each timer gets its own `systemctl disable --now .timer` + `rm`. --- ## The P3-live flip (deferred; IAM + GitHub App gated) 🧑 OPERATOR-REQUIRED P3 (the build→verify apply-and-open-draft-PR loop) is **opt-in and inert** in this deploy: `run-team.py` / `serve` pass `build_verify_wiring=None`, so no P3 subgraph is assembled. Flipping it live is a **separate, gated** provisioning session, not part of the coordinator deploy: 1. **Mandatory reviews first.** The CI trust-boundary + OIDC IAM change is a breaking IAM change → **GPT-4.1 cross-review** (global instructions) AND `/sh-security-review` on the apply/verify surface. Do not flip without both. 2. **Provision the GitHub App** for the trusted apply path (the App that opens the draft PR), and the **`agent-apply` GitHub Actions environment** that holds the apply path's scoped permissions. 3. **Bind the live build→verify wiring** via `agent_team.coordinator.gated_build_verify_wiring(...)` (the read-only CI result fetcher + the real diff builder) — the seam a leaf calls *after* the gate clears. The CI fetcher is read-only and fails closed (missing token / 404 / auth failure → `None` → the gate BLOCKs and the task parks). 4. **Set the apply env vars** the live path reads (the read-only CI-result token and the dispatch target), then re-run the fixer **without** `--dry-run` only once the dispatcher is bound. Until every step above is done, the box dispatches/applies nothing. --- ## Post-session definition-of-done (design §7, global instructions) - [ ] All four P1 criteria demonstrated live (Step 8). - [ ] `project_r720_agent_team` memory created/updated. - [ ] Confluence "AWS Architecture Map" / IT host inventory updated to show `sh-secrev` now also hosts the always-on agent-team coordinator daemon. - [ ] `/sh-security-review` run on the Slack inbound listener surface (auth + untrusted-input; mandatory) — flag outstanding if not run. - [ ] OPERATOR-RUNBOOK.md reviewed by whoever holds the pager. - [ ] Snapshot retained until the daemon runs clean for one full cycle, then pruned.