This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/agent-team/DEPLOY-R720.md

13 KiB
Raw Blame History

P1 - agent-team Plane-2 coordinator deployment (R720 VM)

Status: P1 DEPLOY ARTIFACTS - the systemd unit + this runbook. The coordinator daemon (run-team.py serve) is built by a separate agent; this file is the operator runbook for standing it up on the always-on R720 VM.

The agent-team coordinator is the long-running Plane-2 brain: it drives the LangGraph pipeline, owns the durable pending_questions ledger, and runs the Slack Socket Mode inbound listener that receives clarifier answers. It shares the sh-secrev VM and the ~/secrev.env secrets file with the Path B security sweep (see ../security-review/DEPLOY-R720.md), but it is a service (always-on), not a timer-driven oneshot.

Host

  • Hypervisor: R720 at 10.10.60.40 (Windows Server 2022, Hyper-V role).
  • VM: sh-secrev, always-on Ubuntu 24.04, Gen2, 4GB / 2 vCPU / 40GB dynamic vhdx at 10.10.60.120.
  • Reach it: ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120 (key-only, NOPASSWD sudo).

Operate on the VM, not from the Mac against the host by hand.

1. SNAPSHOT FIRST

Standing rule: snapshot the VM before any provisioning change. This box is a 4GB / 2 vCPU / 40GB VM at 10.10.60.120. Take a Hyper-V checkpoint on the R720 host before you install pip deps, the unit, or touch ~/secrev.env, so the whole change is one-command reversible (see ROLLBACK). Do not skip this because "it is only a pip install" - a bad dep set or a wedged service is exactly what the snapshot exists to undo.

2. Prereqs on the VM

Already present from the secrev deploy:

  • Python 3 (3.12) and the claude CLI (Node) - the subscription-auth path.
  • Repo: ~/orchestrator/ (rsync from the Mac, NOT a git clone). The agent-team package lives at ~/orchestrator/agent-team/.

New for the coordinator - a dedicated venv under agent-team/.venv (excluded from rsync) with the coordinator/transport deps:

pip dep Why
langgraph the coordinator pipeline graph
langgraph-checkpoint-sqlite SqliteSaver checkpointer against the ledger DB
claude-agent-sdk subscription-auth Claude invocation seam
slack_sdk Slack Web API (post questions, chat:write)
slack_bolt Socket Mode inbound listener (receive answers)

The ledger DB defaults to agent-team/state/agent_team.sqlite; the audit log to agent-team/state/audit.log.jsonl. Both live under state/ (gitignored, never committed).

3. Secrets - append to ~/secrev.env (mode 600, never committed)

The coordinator reads its secrets from the same ~/secrev.env the secrev sweep uses. Append these (do not echo them into shell history files; lock the file down after):

echo 'CLAUDE_CODE_OAUTH_TOKEN=...'   >> ~/secrev.env   # from `claude setup-token`
echo 'SLACK_BOT_TOKEN=xoxb-...'      >> ~/secrev.env   # bot token, chat:write
echo 'SLACK_APP_TOKEN=xapp-...'      >> ~/secrev.env   # app-level, connections:write (Socket Mode)
echo 'SLACK_CHANNEL_ID=C0XXXXXXX'    >> ~/secrev.env   # target clarifier channel
echo 'AGENT_TEAM_SLACK_OWNER_IDS=U0XXXXXXX'  >> ~/secrev.env   # authorized answerer(s), comma-separated
chmod 600 ~/secrev.env
  • CLAUDE_CODE_OAUTH_TOKEN - subscription OAuth from claude setup-token. The same token type the secrev sweep uses.
  • SLACK_BOT_TOKEN (xoxb-) - bot token with chat:write; posts questions.
  • SLACK_APP_TOKEN (xapp-) - app-level token with connections:write; required for Socket Mode (opens the inbound WebSocket that hears answers).
  • SLACK_CHANNEL_ID - the channel id the coordinator posts clarifiers to.
  • AGENT_TEAM_SLACK_OWNER_IDS - comma-separated Slack user ids of the authorized answerers (e.g. Adam's U… id). The inbound listener enforces this as an owner allowlist (AUTHZ-01): only a sender in this set may answer/steer the pipeline. The listener fails closed - if this is unset/empty it rejects every answer (logs a warning naming AGENT_TEAM_SLACK_OWNER_IDS), so it must be set for the human gate to function. Look up your user id via Slack profile → "Copy member ID", or the users.identity / auth.test API.

CRITICAL: ANTHROPIC_API_KEY must NOT be set on this host. A raw API key would silently win over the subscription OAuth and meter to API rates. The box runs on subscription OAuth only.

4. Deploy steps

# 4a. From the Mac - rsync the repo (same pattern/excludes as secrev):
rsync -av --exclude .env --exclude .venv \
  ~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/

# 4b. On the VM - create + activate the agent-team venv and install deps:
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team
python3 -m venv .venv
. .venv/bin/activate
pip install langgraph langgraph-checkpoint-sqlite claude-agent-sdk slack_sdk slack_bolt

# 4c. Initialize the durable ledger DB (idempotent; creates state/agent_team.sqlite):
python3 run-team.py init-db

# 4d. Install + start the service:
sudo cp systemd/agent-team-coordinator.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now agent-team-coordinator.service

# 4e. Verify it is up:
systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e -f

The unit runs python3 run-team.py serve from WorkingDirectory=/home/adam/orchestrator/agent-team as User=adam, loading secrets from EnvironmentFile=/home/adam/secrev.env. Restart=on-failure keeps it up across transient faults; journalctl -u is the live log.

4b. WS0–WS5 rollout — UPDATE an already-deployed box

The steps above (§1–4) are the first-time P1 provision. To bring an already-deployed coordinator up to the WS0–WS5 rollout, use the attended update script rather than re-running the manual steps:

# From the Mac, after the WS branches have merged to main, snapshot first:
agent-team/scripts/deploy-r720-ws-rollout.sh

It is an UPDATE (not a provision): it snapshots-reminds, rsyncs the new code, rsyncs the engineering handbook to the box, installs the new deps, appends the new secrets if absent, restarts the coordinator, and smoke-tests. It is idempotent and fails loudly. What goes live after it: WS1 in-process multi-model invokers (bind_multi_invoker, already wired), the WS5 handbook context_provider injected into the planner prompt, and the WS2 Slack /new-task command (AUTHZ-01 owner-allowlist gated). The P3 dispatch/build-verify path stays inert (gated behind the agent-apply GitHub Environment approval).

New venv deps (the coordinator does not need them; only the optional HTTP API does) — now pinned in the root requirements.txt:

pip dep Why
fastapi==0.136.1 the WS1 HTTP API app (agent_team/api.py)
uvicorn==0.46.0 ASGI server for api.serve()

New env vars — append to ~/secrev.env (mode 600, never committed):

  • SEA_HAVEN_HANDBOOK_DIR — where load_handbook_conventions() reads the engineering handbook (the script syncs it to /home/adam/.sea-haven/engineering-handbook by default; this var must match). Fail-safe: if the dir is missing the context_provider returns "" and the planner runs without handbook context.
  • AGENT_TEAM_API_TOKEN — bearer token for the HTTP API / /delegate hook only. Not needed by the coordinator daemon itself. The HTTP API refuses to start if this is unset/empty.

The HTTP API is a separate, opt-in process — it is not started by the coordinator daemon. Run it explicitly (api.serve(), binds 127.0.0.1:8765, bearer auth) only if you want the /delegate Claude Code hook or the POST /tasks / GET /tasks/{thread_id} / POST /orchestrator/invoke endpoints. The /docs + /openapi routes are disabled and it binds loopback by design (do not change to 0.0.0.0). See the deploy script's step 6 for how to start it.

Status dashboard (optional, LAN/VPN-only, READ-ONLY)

agent_team/status_page.py serves a tiny self-refreshing HTML page showing the coordinator queue: each task's short thread_id, description, current phase and status; which tasks have an open pending question (blocked on the human gate) vs. progressing; active/parked counts; and recent budget_ledger spend. It opens the SQLite ledger READ-ONLY (mode=ro) and has no mutating endpoints and no auth.

It is a separate, optional process — agent-team-status.service (mirrors the coordinator unit's hardening; User=adam, EnvironmentFile=-/home/adam/secrev.env, venv-python ExecStart, Restart=on-failure). Unlike the coordinator it needs no ReadWritePaths carve-out (it only reads). It can run side-by-side with the coordinator (RO SQLite opens coexist with the writer).

# install the unit
sudo cp ~/orchestrator/agent-team/systemd/agent-team-status.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now agent-team-status.service
systemctl status agent-team-status.service
journalctl -u agent-team-status.service -e -f

# or run it ad hoc from the venv
cd ~/orchestrator/agent-team && . .venv/bin/activate && \
  python3 -c "from agent_team.status_page import serve; serve()"

Then browse http://10.10.60.120:8770/ from the LAN/VPN.

Config (env): AGENT_TEAM_DB (default state/agent_team.sqlite), AGENT_TEAM_STATUS_HOST (default 0.0.0.0), AGENT_TEAM_STATUS_PORT (default 8770).

Posture: the sh-secrev VM (10.10.60.120, VLAN 60) has no public NIC and sits behind the UniFi firewall, so 0.0.0.0 reaches the LAN/VPN only. Task descriptions may be sensitive and the page is unauthenticated — keep it LAN/VPN-only, never expose it to the public internet. A missing/locked DB renders a friendly "no data" page rather than crashing.

5. P1 live exit-criteria demo (§3.3.1)

Demonstrate all four once the service is live. Map each to the operator commands (run-team.py list / show / force-resume, systemctl). Run the CLI from the working dir so it hits the default ledger: cd ~/orchestrator/agent-team.

(a) Crash-safe resume - kill mid-wait, restart, task resumes. Start a task, get it to a clarifier wait (run-team.py list shows an open question), then:

sudo systemctl stop agent-team-coordinator.service
sudo systemctl start agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e   # confirm the task resumes from the ledger/checkpoint
run-team.py show <question_id>                     # the question is still open, not lost

Pass: the task picks up the same waiting question after restart (the LangGraph SqliteSaver checkpoint + the durable ledger survive the kill).

(b) Duplicate Slack answer is a no-op. Answer a question in Slack, then answer the same question again.

run-team.py show <question_id>   # status flipped to answered exactly once; answered_via is the first answer

Pass: the first answer wins (rowcount == 1); the duplicate hits the BEGIN IMMEDIATE compare-and-set and is ignored (rowcount == 0) - no second resume, no error.

(c) Past-deadline answer is rejected + the task parks. Let a question's deadline_at pass with no answer, then answer late.

run-team.py show <question_id>           # status == expired (auto-expired at deadline)
run-team.py list --parked                # the now-parked task surfaces here

Pass: the expired question rejects the late answer and the task parks rather than spins. To un-park it deliberately:

run-team.py force-resume <question_id> --confirm   # reopens the expired question for re-delivery

(d) Two concurrent tasks resume independently to the correct thread. Start two tasks concurrently, each reaching its own clarifier wait.

run-team.py list   # two distinct open questions, distinct thread_id values

Restart the service (as in (a)); answer each in Slack. Pass: each task resumes to its own thread_id / channel - no cross-talk, no answer routed to the wrong task.

6. Rollback

# Stop + disable the service and remove the unit:
sudo systemctl disable --now agent-team-coordinator.service
sudo rm /etc/systemd/system/agent-team-coordinator.service
sudo systemctl daemon-reload

# Restore the VM from the pre-provision Hyper-V checkpoint (§1) to undo
# pip deps + any host changes in one step.

The ledger is local state under agent-team/state/ (not in git). To reset it without a full snapshot restore: back it up first, then wipe.

cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak}   # back up
rm ~/orchestrator/agent-team/state/agent_team.sqlite*         # wipe (then re-run init-db)

Note the secrets in ~/secrev.env are NOT removed by rollback - leave them, or strip the four agent-team keys if you are decommissioning entirely.

7. Security

  • Slack inbound listener (Socket Mode) is the auth + untrusted-input surface. It accepts inbound messages over a WebSocket and turns them into ledger mutations (answering live clarifier questions). It must pass /sh-security-review before this is enabled in production - that review is mandatory for authentication/authorization and untrusted-input handling changes, and this is both.
  • No IAM / OIDC is involved in P1. The box runs on subscription OAuth (CLAUDE_CODE_OAUTH_TOKEN) and Slack tokens only; there is no AWS role, no OIDC trust relationship, no cloud permission surface in this deploy.
  • ~/secrev.env stays mode 600 and out of git; state/ (ledger + audit log) is gitignored and written 0600.