# P1 - agent-team Plane-2 coordinator deployment (R720 VM) Status: **P1 DEPLOY ARTIFACTS** - the systemd unit + this runbook. The coordinator daemon (`run-team.py serve`) is built by a separate agent; this file is the operator runbook for standing it up on the always-on R720 VM. The agent-team coordinator is the long-running Plane-2 brain: it drives the LangGraph pipeline, owns the durable `pending_questions` ledger, and runs the Slack Socket Mode inbound listener that receives clarifier answers. It shares the `sh-secrev` VM and the `~/secrev.env` secrets file with the Path B security sweep (see `../security-review/DEPLOY-R720.md`), but it is a **service** (always-on), not a timer-driven oneshot. ## Host - **Hypervisor:** R720 at `10.10.60.40` (Windows Server 2022, Hyper-V role). - **VM:** `sh-secrev`, always-on Ubuntu 24.04, Gen2, **4GB / 2 vCPU / 40GB** dynamic vhdx at `10.10.60.120`. - **Reach it:** `ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120` (key-only, NOPASSWD sudo). Operate on the VM, not from the Mac against the host by hand. ## 1. SNAPSHOT FIRST **Standing rule: snapshot the VM before any provisioning change.** This box is a 4GB / 2 vCPU / 40GB VM at `10.10.60.120`. Take a Hyper-V checkpoint on the R720 host **before** you install pip deps, the unit, or touch `~/secrev.env`, so the whole change is one-command reversible (see ROLLBACK). Do not skip this because "it is only a pip install" - a bad dep set or a wedged service is exactly what the snapshot exists to undo. ## 2. Prereqs on the VM Already present from the secrev deploy: - **Python 3** (3.12) and the `claude` CLI (Node) - the subscription-auth path. - **Repo:** `~/orchestrator/` (rsync from the Mac, NOT a git clone). The agent-team package lives at `~/orchestrator/agent-team/`. New for the coordinator - a dedicated venv under `agent-team/.venv` (excluded from rsync) with the coordinator/transport deps: | pip dep | Why | |---|---| | `langgraph` | the coordinator pipeline graph | | `langgraph-checkpoint-sqlite` | `SqliteSaver` checkpointer against the ledger DB | | `claude-agent-sdk` | subscription-auth Claude invocation seam | | `slack_sdk` | Slack Web API (post questions, `chat:write`) | | `slack_bolt` | Socket Mode inbound listener (receive answers) | The ledger DB defaults to `agent-team/state/agent_team.sqlite`; the audit log to `agent-team/state/audit.log.jsonl`. Both live under `state/` (gitignored, never committed). ## 3. Secrets - append to `~/secrev.env` (mode 600, never committed) The coordinator reads its secrets from the same `~/secrev.env` the secrev sweep uses. Append these (do not echo them into shell history files; lock the file down after): ``` echo 'CLAUDE_CODE_OAUTH_TOKEN=...' >> ~/secrev.env # from `claude setup-token` echo 'SLACK_BOT_TOKEN=xoxb-...' >> ~/secrev.env # bot token, chat:write echo 'SLACK_APP_TOKEN=xapp-...' >> ~/secrev.env # app-level, connections:write (Socket Mode) echo 'SLACK_CHANNEL_ID=C0XXXXXXX' >> ~/secrev.env # target clarifier channel echo 'AGENT_TEAM_SLACK_OWNER_IDS=U0XXXXXXX' >> ~/secrev.env # authorized answerer(s), comma-separated chmod 600 ~/secrev.env ``` - `CLAUDE_CODE_OAUTH_TOKEN` - subscription OAuth from `claude setup-token`. The same token type the secrev sweep uses. - `SLACK_BOT_TOKEN` (`xoxb-`) - bot token with `chat:write`; posts questions. - `SLACK_APP_TOKEN` (`xapp-`) - app-level token with `connections:write`; **required for Socket Mode** (opens the inbound WebSocket that hears answers). - `SLACK_CHANNEL_ID` - the channel id the coordinator posts clarifiers to. - `AGENT_TEAM_SLACK_OWNER_IDS` - comma-separated Slack **user ids** of the authorized answerers (e.g. Adam's `U…` id). The inbound listener enforces this as an owner allowlist (AUTHZ-01): only a sender in this set may answer/steer the pipeline. **The listener fails closed** - if this is unset/empty it rejects **every** answer (logs a warning naming `AGENT_TEAM_SLACK_OWNER_IDS`), so it must be set for the human gate to function. Look up your user id via Slack profile → "Copy member ID", or the `users.identity` / `auth.test` API. **CRITICAL:** `ANTHROPIC_API_KEY` must **NOT** be set on this host. A raw API key would silently win over the subscription OAuth and meter to API rates. The box runs on subscription OAuth only. ### 3a. P3 dispatch auth — GitHub App installation tokens (box-side) The P3 dispatcher (`agent_team/dispatcher.py`) carries an approved diff into org CI: it pushes the candidate head branch and triggers the `agent-team-apply-verify.yml` `workflow_dispatch`, then locates the resulting run id. By default those three seams shell out to `gh`/`git`. **`gh` is not installed on the box and the box's read-only PAT has no Actions permission**, so the default path cannot dispatch. Instead the box authenticates with a **GitHub App**: it holds the App private key and mints short-lived (~1h) **installation access tokens** on demand (`agent_team/github_app.py`). Set all three of these to turn the App path on (all-or-nothing — partial config logs one warning and falls back to the inert gh-default path, it never crashes serve): ``` echo 'AGENT_TEAM_GH_APP_ID=...' >> ~/secrev.env # the App's numeric app_id echo 'AGENT_TEAM_GH_APP_INSTALLATION_ID=...' >> ~/secrev.env # installation id on Sea-Haven-Industries/orchestrator echo 'AGENT_TEAM_GH_APP_PRIVATE_KEY=/home/adam/.sea-haven/agent-team-apply.pem' >> ~/secrev.env # PATH to the .pem chmod 600 ~/secrev.env ``` Place the App private key on the box and lock it down — it is a write-capable credential and must be owner-only: ``` install -m 600 /dev/stdin ~/.sea-haven/agent-team-apply.pem < the-downloaded-key.pem chmod 600 ~/.sea-haven/agent-team-apply.pem && ls -l ~/.sea-haven/agent-team-apply.pem # expect -rw------- ``` - `AGENT_TEAM_GH_APP_PRIVATE_KEY` is a **filesystem path** to the `.pem`, not the key material. The coordinator reads the file at graph-build time; an unreadable path logs one warning and falls back to gh-default (never crashes). - **App permission/scope audit (do before deploy).** The App (currently the reused **`agent-team-apply`** App) must be installed on **only** `Sea-Haven-Industries/orchestrator` with **Contents: Read/Write** + **Actions: Read/Write** and nothing broader. Minting an installation token grants exactly the App's scopes; over-broad scopes widen blast radius. (Follow-up: move to a dedicated least-privilege App to replace `agent-team-apply`.) - The minted token is short-lived (~1h), is **never** logged, never put in an exception message, and never written to the ledger/graph state; on the authenticated push it rides in a host-scoped `http.extraHeader` passed via `GIT_CONFIG_*` env (never in argv/`ps`). - **Env-file precedence (verify after editing).** The unit loads **both** `~/secrev.env` and `~/orchestrator/.env` (last-wins). Confirm only the intended values are present in both so a stale entry can't shadow the App config: ``` grep -nE 'AGENT_TEAM_GH_APP|AGENT_TEAM_REPO_(OWNER|NAME)' ~/secrev.env ~/orchestrator/.env 2>/dev/null ``` ## 4. Deploy steps ``` # 4a. From the Mac - rsync the repo (same pattern/excludes as secrev): rsync -av --exclude .env --exclude .venv \ ~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/ # 4b. On the VM - create + activate the agent-team venv and install deps: ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120 cd ~/orchestrator/agent-team python3 -m venv .venv . .venv/bin/activate pip install langgraph langgraph-checkpoint-sqlite claude-agent-sdk slack_sdk slack_bolt # 4c. Initialize the durable ledger DB (idempotent; creates state/agent_team.sqlite): python3 run-team.py init-db # 4d. Install + start the service: sudo cp systemd/agent-team-coordinator.service /etc/systemd/system/ sudo systemctl daemon-reload sudo systemctl enable --now agent-team-coordinator.service # 4e. Verify it is up: systemctl status agent-team-coordinator.service journalctl -u agent-team-coordinator.service -e -f ``` The unit runs `python3 run-team.py serve` from `WorkingDirectory=/home/adam/orchestrator/agent-team` as `User=adam`, loading secrets from `EnvironmentFile=/home/adam/secrev.env`. `Restart=on-failure` keeps it up across transient faults; `journalctl -u` is the live log. > **⚠️ This deploy carries a ledger migration (SCHEMA_VERSION → 4).** It adds a > `kind` column to `pending_questions` (values `clarify` | `plan_decision`, > existing rows default to `clarify`) for the plan-review decision gate. The > migration is an **additive, idempotent in-place `ALTER TABLE`** run on startup > (`init_db` / `migrate`, guarded so a second run is a no-op — it never recreates > the table), so it applies in place against the live box ledger. **Take the > ledger backup (§1 / the deploy script step) BEFORE restart** — it is the > migration's safety net (see Rollback, §6). After restart, confirm the column > landed and the daemon came up clean: > ```bash > sqlite3 ~/orchestrator/agent-team/state/agent_team.sqlite \ > "PRAGMA table_info(pending_questions);" | grep kind # expect a 'kind' row > sqlite3 ~/orchestrator/agent-team/state/agent_team.sqlite \ > "SELECT schema_version FROM schema_meta WHERE id=1;" # expect 4 > journalctl -u agent-team-coordinator.service -e | tail # no migration/import errors > ``` ## 4b. WS0–WS5 rollout — UPDATE an already-deployed box The steps above (§1–4) are the **first-time** P1 provision. To bring an already-deployed coordinator up to the WS0–WS5 rollout, use the attended update script rather than re-running the manual steps: ``` # From the Mac, after the WS branches have merged to main, snapshot first: agent-team/scripts/deploy-r720-ws-rollout.sh ``` It is an UPDATE (not a provision): it snapshots-reminds, rsyncs the new code, rsyncs the engineering handbook to the box, installs the new deps, appends the new secrets if absent, restarts the coordinator, and smoke-tests. It is idempotent and fails loudly. What goes **live** after it: WS1 in-process multi-model invokers (`bind_multi_invoker`, already wired), the WS5 handbook `context_provider` injected into the planner prompt, and the WS2 Slack `/new-task` command (AUTHZ-01 owner-allowlist gated). The P3 dispatch/build-verify path stays **inert** (gated behind the `agent-apply` GitHub Environment approval). **New venv deps** (the coordinator does not need them; only the optional HTTP API does) — now pinned in the root `requirements.txt`: | pip dep | Why | |---|---| | `fastapi==0.136.1` | the WS1 HTTP API app (`agent_team/api.py`) | | `uvicorn==0.46.0` | ASGI server for `api.serve()` | **New env vars** — append to `~/secrev.env` (mode 600, never committed): - `SEA_HAVEN_HANDBOOK_DIR` — where `load_handbook_conventions()` reads the engineering handbook (the script syncs it to `/home/adam/.sea-haven/engineering-handbook` by default; this var must match). Fail-safe: if the dir is missing the `context_provider` returns `""` and the planner runs without handbook context. - `AGENT_TEAM_API_TOKEN` — bearer token for the HTTP API / `/delegate` hook **only**. Not needed by the coordinator daemon itself. The HTTP API refuses to start if this is unset/empty. **The HTTP API is a separate, opt-in process** — it is **not** started by the coordinator daemon. Run it explicitly (`api.serve()`, binds `127.0.0.1:8765`, bearer auth) only if you want the `/delegate` Claude Code hook or the `POST /tasks` / `GET /tasks/{thread_id}` / `POST /orchestrator/invoke` endpoints. The `/docs` + `/openapi` routes are disabled and it binds loopback by design (do not change to `0.0.0.0`). See the deploy script's step 6 for how to start it. ### Status dashboard (optional, LAN/VPN-only, READ-ONLY) `agent_team/status_page.py` serves a **live visual pipeline map** of the coordinator. The top of the page is a hand-rolled inline-SVG diagram of the agent-team DAG (`INTAKE → CLARIFY ⇄ human gate → PLAN ⇄ REVIEW → [BUILD → VERIFY → DISPATCH] → DONE`); each stage node is labelled with its model role (Claude on subscription for CLARIFY/PLAN/VERIFY, GPT-4.1 cross_reviewer for REVIEW, Gemini for SCAN, DeepSeek fast_coder for BUILD, the Slack owner for the HUMAN GATE) and is colour-coded by live state — idle / active / awaiting-human (an **open** pending question) / parked — with a count badge of tasks in that stage. Hovering (or keyboard-focusing) a node shows what it is working on: the short `thread_id`, description, status, and waiting age of each task there. Below the map, the original detail tables remain: all tasks, the human-gate wait-list, and recent `budget_ledger` spend. The page **auto-updates without a full reload**: a `GET /api/state` JSON sidecar returns the same snapshot, and an inline vanilla-JS poller (`fetch()`, no libraries, no CDN) re-paints node states, counts, the cards, the tooltip data, and the "last updated" clock every ~4s in place, so hover/scroll/focus survive. A `