13 KiB
P1 - agent-team Plane-2 coordinator deployment (R720 VM)
Status: P1 DEPLOY ARTIFACTS - the systemd unit + this runbook. The coordinator
daemon (run-team.py serve) is built by a separate agent; this file is the
operator runbook for standing it up on the always-on R720 VM.
The agent-team coordinator is the long-running Plane-2 brain: it drives the
LangGraph pipeline, owns the durable pending_questions ledger, and runs the
Slack Socket Mode inbound listener that receives clarifier answers. It shares the
sh-secrev VM and the ~/secrev.env secrets file with the Path B security sweep
(see ../security-review/DEPLOY-R720.md), but it is a service (always-on),
not a timer-driven oneshot.
Host
- Hypervisor: R720 at
10.10.60.40(Windows Server 2022, Hyper-V role). - VM:
sh-secrev, always-on Ubuntu 24.04, Gen2, 4GB / 2 vCPU / 40GB dynamic vhdx at10.10.60.120. - Reach it:
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120(key-only, NOPASSWD sudo).
Operate on the VM, not from the Mac against the host by hand.
1. SNAPSHOT FIRST
Standing rule: snapshot the VM before any provisioning change. This box is a
4GB / 2 vCPU / 40GB VM at 10.10.60.120. Take a Hyper-V checkpoint on the R720
host before you install pip deps, the unit, or touch ~/secrev.env, so the
whole change is one-command reversible (see ROLLBACK). Do not skip this because
"it is only a pip install" - a bad dep set or a wedged service is exactly what
the snapshot exists to undo.
2. Prereqs on the VM
Already present from the secrev deploy:
- Python 3 (3.12) and the
claudeCLI (Node) - the subscription-auth path. - Repo:
~/orchestrator/(rsync from the Mac, NOT a git clone). The agent-team package lives at~/orchestrator/agent-team/.
New for the coordinator - a dedicated venv under agent-team/.venv (excluded
from rsync) with the coordinator/transport deps:
| pip dep | Why |
|---|---|
langgraph |
the coordinator pipeline graph |
langgraph-checkpoint-sqlite |
SqliteSaver checkpointer against the ledger DB |
claude-agent-sdk |
subscription-auth Claude invocation seam |
slack_sdk |
Slack Web API (post questions, chat:write) |
slack_bolt |
Socket Mode inbound listener (receive answers) |
The ledger DB defaults to agent-team/state/agent_team.sqlite; the audit log to
agent-team/state/audit.log.jsonl. Both live under state/ (gitignored,
never committed).
3. Secrets - append to ~/secrev.env (mode 600, never committed)
The coordinator reads its secrets from the same ~/secrev.env the secrev sweep
uses. Append these (do not echo them into shell history files; lock the file
down after):
echo 'CLAUDE_CODE_OAUTH_TOKEN=...' >> ~/secrev.env # from `claude setup-token`
echo 'SLACK_BOT_TOKEN=xoxb-...' >> ~/secrev.env # bot token, chat:write
echo 'SLACK_APP_TOKEN=xapp-...' >> ~/secrev.env # app-level, connections:write (Socket Mode)
echo 'SLACK_CHANNEL_ID=C0XXXXXXX' >> ~/secrev.env # target clarifier channel
echo 'AGENT_TEAM_SLACK_OWNER_IDS=U0XXXXXXX' >> ~/secrev.env # authorized answerer(s), comma-separated
chmod 600 ~/secrev.env
CLAUDE_CODE_OAUTH_TOKEN- subscription OAuth fromclaude setup-token. The same token type the secrev sweep uses.SLACK_BOT_TOKEN(xoxb-) - bot token withchat:write; posts questions.SLACK_APP_TOKEN(xapp-) - app-level token withconnections:write; required for Socket Mode (opens the inbound WebSocket that hears answers).SLACK_CHANNEL_ID- the channel id the coordinator posts clarifiers to.AGENT_TEAM_SLACK_OWNER_IDS- comma-separated Slack user ids of the authorized answerers (e.g. Adam'sU…id). The inbound listener enforces this as an owner allowlist (AUTHZ-01): only a sender in this set may answer/steer the pipeline. The listener fails closed - if this is unset/empty it rejects every answer (logs a warning namingAGENT_TEAM_SLACK_OWNER_IDS), so it must be set for the human gate to function. Look up your user id via Slack profile → "Copy member ID", or theusers.identity/auth.testAPI.
CRITICAL: ANTHROPIC_API_KEY must NOT be set on this host. A raw API key
would silently win over the subscription OAuth and meter to API rates. The box
runs on subscription OAuth only.
4. Deploy steps
# 4a. From the Mac - rsync the repo (same pattern/excludes as secrev):
rsync -av --exclude .env --exclude .venv \
~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/
# 4b. On the VM - create + activate the agent-team venv and install deps:
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team
python3 -m venv .venv
. .venv/bin/activate
pip install langgraph langgraph-checkpoint-sqlite claude-agent-sdk slack_sdk slack_bolt
# 4c. Initialize the durable ledger DB (idempotent; creates state/agent_team.sqlite):
python3 run-team.py init-db
# 4d. Install + start the service:
sudo cp systemd/agent-team-coordinator.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now agent-team-coordinator.service
# 4e. Verify it is up:
systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e -f
The unit runs python3 run-team.py serve from
WorkingDirectory=/home/adam/orchestrator/agent-team as User=adam, loading
secrets from EnvironmentFile=/home/adam/secrev.env. Restart=on-failure keeps
it up across transient faults; journalctl -u is the live log.
4b. WS0–WS5 rollout — UPDATE an already-deployed box
The steps above (§1–4) are the first-time P1 provision. To bring an already-deployed coordinator up to the WS0–WS5 rollout, use the attended update script rather than re-running the manual steps:
# From the Mac, after the WS branches have merged to main, snapshot first:
agent-team/scripts/deploy-r720-ws-rollout.sh
It is an UPDATE (not a provision): it snapshots-reminds, rsyncs the new code,
rsyncs the engineering handbook to the box, installs the new deps, appends the
new secrets if absent, restarts the coordinator, and smoke-tests. It is
idempotent and fails loudly. What goes live after it: WS1 in-process
multi-model invokers (bind_multi_invoker, already wired), the WS5 handbook
context_provider injected into the planner prompt, and the WS2 Slack
/new-task command (AUTHZ-01 owner-allowlist gated). The P3 dispatch/build-verify
path stays inert (gated behind the agent-apply GitHub Environment approval).
New venv deps (the coordinator does not need them; only the optional HTTP
API does) — now pinned in the root requirements.txt:
| pip dep | Why |
|---|---|
fastapi==0.136.1 |
the WS1 HTTP API app (agent_team/api.py) |
uvicorn==0.46.0 |
ASGI server for api.serve() |
New env vars — append to ~/secrev.env (mode 600, never committed):
SEA_HAVEN_HANDBOOK_DIR— whereload_handbook_conventions()reads the engineering handbook (the script syncs it to/home/adam/.sea-haven/engineering-handbookby default; this var must match). Fail-safe: if the dir is missing thecontext_providerreturns""and the planner runs without handbook context.AGENT_TEAM_API_TOKEN— bearer token for the HTTP API //delegatehook only. Not needed by the coordinator daemon itself. The HTTP API refuses to start if this is unset/empty.
The HTTP API is a separate, opt-in process — it is not started by the
coordinator daemon. Run it explicitly (api.serve(), binds 127.0.0.1:8765,
bearer auth) only if you want the /delegate Claude Code hook or the
POST /tasks / GET /tasks/{thread_id} / POST /orchestrator/invoke endpoints.
The /docs + /openapi routes are disabled and it binds loopback by design (do
not change to 0.0.0.0). See the deploy script's step 6 for how to start it.
Status dashboard (optional, LAN/VPN-only, READ-ONLY)
agent_team/status_page.py serves a tiny self-refreshing HTML page showing the
coordinator queue: each task's short thread_id, description, current phase and
status; which tasks have an open pending question (blocked on the human gate)
vs. progressing; active/parked counts; and recent budget_ledger spend. It opens
the SQLite ledger READ-ONLY (mode=ro) and has no mutating endpoints and
no auth.
It is a separate, optional process — agent-team-status.service (mirrors the
coordinator unit's hardening; User=adam, EnvironmentFile=-/home/adam/secrev.env,
venv-python ExecStart, Restart=on-failure). Unlike the coordinator it needs no
ReadWritePaths carve-out (it only reads). It can run side-by-side with the
coordinator (RO SQLite opens coexist with the writer).
# install the unit
sudo cp ~/orchestrator/agent-team/systemd/agent-team-status.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now agent-team-status.service
systemctl status agent-team-status.service
journalctl -u agent-team-status.service -e -f
# or run it ad hoc from the venv
cd ~/orchestrator/agent-team && . .venv/bin/activate && \
python3 -c "from agent_team.status_page import serve; serve()"
Then browse http://10.10.60.120:8770/ from the LAN/VPN.
Config (env): AGENT_TEAM_DB (default state/agent_team.sqlite),
AGENT_TEAM_STATUS_HOST (default 0.0.0.0), AGENT_TEAM_STATUS_PORT (default
8770).
Posture: the sh-secrev VM (10.10.60.120, VLAN 60) has no public NIC and sits
behind the UniFi firewall, so 0.0.0.0 reaches the LAN/VPN only. Task
descriptions may be sensitive and the page is unauthenticated — keep it
LAN/VPN-only, never expose it to the public internet. A missing/locked DB renders
a friendly "no data" page rather than crashing.
5. P1 live exit-criteria demo (§3.3.1)
Demonstrate all four once the service is live. Map each to the operator commands
(run-team.py list / show / force-resume, systemctl). Run the CLI from the
working dir so it hits the default ledger: cd ~/orchestrator/agent-team.
(a) Crash-safe resume - kill mid-wait, restart, task resumes.
Start a task, get it to a clarifier wait (run-team.py list shows an open
question), then:
sudo systemctl stop agent-team-coordinator.service
sudo systemctl start agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e # confirm the task resumes from the ledger/checkpoint
run-team.py show <question_id> # the question is still open, not lost
Pass: the task picks up the same waiting question after restart (the LangGraph
SqliteSaver checkpoint + the durable ledger survive the kill).
(b) Duplicate Slack answer is a no-op. Answer a question in Slack, then answer the same question again.
run-team.py show <question_id> # status flipped to answered exactly once; answered_via is the first answer
Pass: the first answer wins (rowcount == 1); the duplicate hits the
BEGIN IMMEDIATE compare-and-set and is ignored (rowcount == 0) - no second
resume, no error.
(c) Past-deadline answer is rejected + the task parks.
Let a question's deadline_at pass with no answer, then answer late.
run-team.py show <question_id> # status == expired (auto-expired at deadline)
run-team.py list --parked # the now-parked task surfaces here
Pass: the expired question rejects the late answer and the task parks rather than spins. To un-park it deliberately:
run-team.py force-resume <question_id> --confirm # reopens the expired question for re-delivery
(d) Two concurrent tasks resume independently to the correct thread. Start two tasks concurrently, each reaching its own clarifier wait.
run-team.py list # two distinct open questions, distinct thread_id values
Restart the service (as in (a)); answer each in Slack.
Pass: each task resumes to its own thread_id / channel - no cross-talk, no
answer routed to the wrong task.
6. Rollback
# Stop + disable the service and remove the unit:
sudo systemctl disable --now agent-team-coordinator.service
sudo rm /etc/systemd/system/agent-team-coordinator.service
sudo systemctl daemon-reload
# Restore the VM from the pre-provision Hyper-V checkpoint (§1) to undo
# pip deps + any host changes in one step.
The ledger is local state under agent-team/state/ (not in git). To reset
it without a full snapshot restore: back it up first, then wipe.
cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak} # back up
rm ~/orchestrator/agent-team/state/agent_team.sqlite* # wipe (then re-run init-db)
Note the secrets in ~/secrev.env are NOT removed by rollback - leave them, or
strip the four agent-team keys if you are decommissioning entirely.
7. Security
- Slack inbound listener (Socket Mode) is the auth + untrusted-input surface.
It accepts inbound messages over a WebSocket and turns them into ledger
mutations (answering live clarifier questions). It must pass
/sh-security-reviewbefore this is enabled in production - that review is mandatory for authentication/authorization and untrusted-input handling changes, and this is both. - No IAM / OIDC is involved in P1. The box runs on subscription OAuth
(
CLAUDE_CODE_OAUTH_TOKEN) and Slack tokens only; there is no AWS role, no OIDC trust relationship, no cloud permission surface in this deploy. ~/secrev.envstays mode 600 and out of git;state/(ledger + audit log) is gitignored and written 0600.