diff --git a/agent-team/DEPLOY-R720.md b/agent-team/DEPLOY-R720.md new file mode 100644 index 0000000..18cc9a7 --- /dev/null +++ b/agent-team/DEPLOY-R720.md @@ -0,0 +1,202 @@ +# P1 - agent-team Plane-2 coordinator deployment (R720 VM) + +Status: **P1 DEPLOY ARTIFACTS** - the systemd unit + this runbook. The coordinator +daemon (`run-team.py serve`) is built by a separate agent; this file is the +operator runbook for standing it up on the always-on R720 VM. + +The agent-team coordinator is the long-running Plane-2 brain: it drives the +LangGraph pipeline, owns the durable `pending_questions` ledger, and runs the +Slack Socket Mode inbound listener that receives clarifier answers. It shares the +`sh-secrev` VM and the `~/secrev.env` secrets file with the Path B security sweep +(see `../security-review/DEPLOY-R720.md`), but it is a **service** (always-on), +not a timer-driven oneshot. + +## Host + +- **Hypervisor:** R720 at `10.10.60.40` (Windows Server 2022, Hyper-V role). +- **VM:** `sh-secrev`, always-on Ubuntu 24.04, Gen2, **4GB / 2 vCPU / 40GB** + dynamic vhdx at `10.10.60.120`. +- **Reach it:** `ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120` (key-only, + NOPASSWD sudo). + +Operate on the VM, not from the Mac against the host by hand. + +## 1. SNAPSHOT FIRST + +**Standing rule: snapshot the VM before any provisioning change.** This box is a +4GB / 2 vCPU / 40GB VM at `10.10.60.120`. Take a Hyper-V checkpoint on the R720 +host **before** you install pip deps, the unit, or touch `~/secrev.env`, so the +whole change is one-command reversible (see ROLLBACK). Do not skip this because +"it is only a pip install" - a bad dep set or a wedged service is exactly what +the snapshot exists to undo. + +## 2. Prereqs on the VM + +Already present from the secrev deploy: + +- **Python 3** (3.12) and the `claude` CLI (Node) - the subscription-auth path. +- **Repo:** `~/orchestrator/` (rsync from the Mac, NOT a git clone). The + agent-team package lives at `~/orchestrator/agent-team/`. + +New for the coordinator - a dedicated venv under `agent-team/.venv` (excluded +from rsync) with the coordinator/transport deps: + +| pip dep | Why | +|---|---| +| `langgraph` | the coordinator pipeline graph | +| `langgraph-checkpoint-sqlite` | `SqliteSaver` checkpointer against the ledger DB | +| `claude-agent-sdk` | subscription-auth Claude invocation seam | +| `slack_sdk` | Slack Web API (post questions, `chat:write`) | +| `slack_bolt` | Socket Mode inbound listener (receive answers) | + +The ledger DB defaults to `agent-team/state/agent_team.sqlite`; the audit log to +`agent-team/state/audit.log.jsonl`. Both live under `state/` (gitignored, +never committed). + +## 3. Secrets - append to `~/secrev.env` (mode 600, never committed) + +The coordinator reads its secrets from the same `~/secrev.env` the secrev sweep +uses. Append these (do not echo them into shell history files; lock the file +down after): + +``` +echo 'CLAUDE_CODE_OAUTH_TOKEN=...' >> ~/secrev.env # from `claude setup-token` +echo 'SLACK_BOT_TOKEN=xoxb-...' >> ~/secrev.env # bot token, chat:write +echo 'SLACK_APP_TOKEN=xapp-...' >> ~/secrev.env # app-level, connections:write (Socket Mode) +echo 'SLACK_CHANNEL_ID=C0XXXXXXX' >> ~/secrev.env # target clarifier channel +echo 'AGENT_TEAM_SLACK_OWNER_IDS=U0XXXXXXX' >> ~/secrev.env # authorized answerer(s), comma-separated +chmod 600 ~/secrev.env +``` + +- `CLAUDE_CODE_OAUTH_TOKEN` - subscription OAuth from `claude setup-token`. The + same token type the secrev sweep uses. +- `SLACK_BOT_TOKEN` (`xoxb-`) - bot token with `chat:write`; posts questions. +- `SLACK_APP_TOKEN` (`xapp-`) - app-level token with `connections:write`; + **required for Socket Mode** (opens the inbound WebSocket that hears answers). +- `SLACK_CHANNEL_ID` - the channel id the coordinator posts clarifiers to. +- `AGENT_TEAM_SLACK_OWNER_IDS` - comma-separated Slack **user ids** of the + authorized answerers (e.g. Adam's `U…` id). The inbound listener enforces this + as an owner allowlist (AUTHZ-01): only a sender in this set may answer/steer + the pipeline. **The listener fails closed** - if this is unset/empty it rejects + **every** answer (logs a warning naming `AGENT_TEAM_SLACK_OWNER_IDS`), so it + must be set for the human gate to function. Look up your user id via Slack + profile → "Copy member ID", or the `users.identity` / `auth.test` API. + +**CRITICAL:** `ANTHROPIC_API_KEY` must **NOT** be set on this host. A raw API key +would silently win over the subscription OAuth and meter to API rates. The box +runs on subscription OAuth only. + +## 4. Deploy steps + +``` +# 4a. From the Mac - rsync the repo (same pattern/excludes as secrev): +rsync -av --exclude .env --exclude .venv \ + ~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/ + +# 4b. On the VM - create + activate the agent-team venv and install deps: +ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120 +cd ~/orchestrator/agent-team +python3 -m venv .venv +. .venv/bin/activate +pip install langgraph langgraph-checkpoint-sqlite claude-agent-sdk slack_sdk slack_bolt + +# 4c. Initialize the durable ledger DB (idempotent; creates state/agent_team.sqlite): +python3 run-team.py init-db + +# 4d. Install + start the service: +sudo cp systemd/agent-team-coordinator.service /etc/systemd/system/ +sudo systemctl daemon-reload +sudo systemctl enable --now agent-team-coordinator.service + +# 4e. Verify it is up: +systemctl status agent-team-coordinator.service +journalctl -u agent-team-coordinator.service -e -f +``` + +The unit runs `python3 run-team.py serve` from +`WorkingDirectory=/home/adam/orchestrator/agent-team` as `User=adam`, loading +secrets from `EnvironmentFile=/home/adam/secrev.env`. `Restart=on-failure` keeps +it up across transient faults; `journalctl -u` is the live log. + +## 5. P1 live exit-criteria demo (§3.3.1) + +Demonstrate all four once the service is live. Map each to the operator commands +(`run-team.py list / show / force-resume`, `systemctl`). Run the CLI from the +working dir so it hits the default ledger: `cd ~/orchestrator/agent-team`. + +**(a) Crash-safe resume - kill mid-wait, restart, task resumes.** +Start a task, get it to a clarifier wait (`run-team.py list` shows an `open` +question), then: +``` +sudo systemctl stop agent-team-coordinator.service +sudo systemctl start agent-team-coordinator.service +journalctl -u agent-team-coordinator.service -e # confirm the task resumes from the ledger/checkpoint +run-team.py show # the question is still open, not lost +``` +Pass: the task picks up the same waiting question after restart (the LangGraph +`SqliteSaver` checkpoint + the durable ledger survive the kill). + +**(b) Duplicate Slack answer is a no-op.** +Answer a question in Slack, then answer the **same** question again. +``` +run-team.py show # status flipped to answered exactly once; answered_via is the first answer +``` +Pass: the first answer wins (`rowcount == 1`); the duplicate hits the +`BEGIN IMMEDIATE` compare-and-set and is ignored (`rowcount == 0`) - no second +resume, no error. + +**(c) Past-deadline answer is rejected + the task parks.** +Let a question's `deadline_at` pass with no answer, then answer late. +``` +run-team.py show # status == expired (auto-expired at deadline) +run-team.py list --parked # the now-parked task surfaces here +``` +Pass: the expired question rejects the late answer and the task parks rather than +spins. To un-park it deliberately: +``` +run-team.py force-resume --confirm # reopens the expired question for re-delivery +``` + +**(d) Two concurrent tasks resume independently to the correct thread.** +Start two tasks concurrently, each reaching its own clarifier wait. +``` +run-team.py list # two distinct open questions, distinct thread_id values +``` +Restart the service (as in (a)); answer each in Slack. +Pass: each task resumes to its own `thread_id` / channel - no cross-talk, no +answer routed to the wrong task. + +## 6. Rollback + +``` +# Stop + disable the service and remove the unit: +sudo systemctl disable --now agent-team-coordinator.service +sudo rm /etc/systemd/system/agent-team-coordinator.service +sudo systemctl daemon-reload + +# Restore the VM from the pre-provision Hyper-V checkpoint (§1) to undo +# pip deps + any host changes in one step. +``` + +The ledger is **local state** under `agent-team/state/` (not in git). To reset +it without a full snapshot restore: back it up first, then wipe. +``` +cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak} # back up +rm ~/orchestrator/agent-team/state/agent_team.sqlite* # wipe (then re-run init-db) +``` +Note the secrets in `~/secrev.env` are NOT removed by rollback - leave them, or +strip the four agent-team keys if you are decommissioning entirely. + +## 7. Security + +- **Slack inbound listener (Socket Mode) is the auth + untrusted-input surface.** + It accepts inbound messages over a WebSocket and turns them into ledger + mutations (answering live clarifier questions). It **must pass + `/sh-security-review`** before this is enabled in production - that review is + mandatory for authentication/authorization and untrusted-input handling + changes, and this is both. +- **No IAM / OIDC is involved in P1.** The box runs on subscription OAuth + (`CLAUDE_CODE_OAUTH_TOKEN`) and Slack tokens only; there is no AWS role, no + OIDC trust relationship, no cloud permission surface in this deploy. +- `~/secrev.env` stays mode 600 and out of git; `state/` (ledger + audit log) is + gitignored and written 0600. diff --git a/agent-team/systemd/agent-team-coordinator.service b/agent-team/systemd/agent-team-coordinator.service new file mode 100644 index 0000000..20d25d2 --- /dev/null +++ b/agent-team/systemd/agent-team-coordinator.service @@ -0,0 +1,54 @@ +# agent-team-coordinator.service - R720 Plane-2 coordinator daemon (sh-secrev VM, user adam). +# +# Long-running coordinator for the agent-team SDLC pipeline. Unlike the +# sea-haven-secrev sweep (a oneshot driven by a timer), this is an always-on +# service: it serves the LangGraph coordinator, the durable pending_questions +# ledger, and the Slack Socket Mode inbound listener that answers clarifier +# questions. ExecStart runs the `serve` subcommand of the operator CLI. +# +# Install (on the VM, as root): +# sudo cp agent-team-coordinator.service /etc/systemd/system/ +# sudo systemctl daemon-reload +# sudo systemctl enable --now agent-team-coordinator.service +# systemctl status agent-team-coordinator.service +# journalctl -u agent-team-coordinator.service -e -f +# +# Secrets come from the EnvironmentFile (leading '-' = optional, no failure if +# absent), ~/secrev.env (mode 600, NOT in git): +# CLAUDE_CODE_OAUTH_TOKEN -> subscription OAuth (from `claude setup-token`). +# A raw ANTHROPIC_API_KEY must NOT be set on this +# box; it would silently win and meter to API +# rates. The billing seam pops it defensively. +# SLACK_BOT_TOKEN -> xoxb- bot token (chat:write) - posts questions. +# SLACK_APP_TOKEN -> xapp- app-level token (connections:write) - +# REQUIRED for Socket Mode; opens the inbound +# WebSocket that receives answers. Without it the +# coordinator can post but never hear replies. +# SLACK_CHANNEL_ID -> target channel for clarifier questions. + +[Unit] +Description=Sea Haven agent-team Plane-2 coordinator daemon +After=network-online.target +Wants=network-online.target + +[Service] +Type=simple +User=adam +WorkingDirectory=/home/adam/orchestrator/agent-team +EnvironmentFile=-/home/adam/secrev.env +ExecStart=/usr/bin/env python3 run-team.py serve +Restart=on-failure +RestartSec=5 +# Hardening - matches the level the sea-haven-secrev unit relies on, scoped for a +# long-running daemon that must READ ~/secrev.env and WRITE the local ledger. +NoNewPrivileges=true +ProtectSystem=full +# ProtectHome cannot be `true`: the daemon reads /home/adam/secrev.env and writes +# the ledger under the working dir. read-only home + an explicit RW carve-out for +# the state/ dir keeps the rest of $HOME unreadable/unwritable to the service. +ProtectHome=read-only +ReadWritePaths=/home/adam/orchestrator/agent-team/state +Nice=10 + +[Install] +WantedBy=multi-user.target