P1c deploy artifacts: DEPLOY-R720.md (snapshot-first, rsync, venv deps, ~/secrev.env tokens incl. AGENT_TEAM_SLACK_OWNER_IDS, init-db, systemd, the 4-criteria live demo, rollback) and the long-running coordinator service unit.
9.1 KiB
P1 - agent-team Plane-2 coordinator deployment (R720 VM)
Status: P1 DEPLOY ARTIFACTS - the systemd unit + this runbook. The coordinator
daemon (run-team.py serve) is built by a separate agent; this file is the
operator runbook for standing it up on the always-on R720 VM.
The agent-team coordinator is the long-running Plane-2 brain: it drives the
LangGraph pipeline, owns the durable pending_questions ledger, and runs the
Slack Socket Mode inbound listener that receives clarifier answers. It shares the
sh-secrev VM and the ~/secrev.env secrets file with the Path B security sweep
(see ../security-review/DEPLOY-R720.md), but it is a service (always-on),
not a timer-driven oneshot.
Host
- Hypervisor: R720 at
10.10.60.40(Windows Server 2022, Hyper-V role). - VM:
sh-secrev, always-on Ubuntu 24.04, Gen2, 4GB / 2 vCPU / 40GB dynamic vhdx at10.10.60.120. - Reach it:
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120(key-only, NOPASSWD sudo).
Operate on the VM, not from the Mac against the host by hand.
1. SNAPSHOT FIRST
Standing rule: snapshot the VM before any provisioning change. This box is a
4GB / 2 vCPU / 40GB VM at 10.10.60.120. Take a Hyper-V checkpoint on the R720
host before you install pip deps, the unit, or touch ~/secrev.env, so the
whole change is one-command reversible (see ROLLBACK). Do not skip this because
"it is only a pip install" - a bad dep set or a wedged service is exactly what
the snapshot exists to undo.
2. Prereqs on the VM
Already present from the secrev deploy:
- Python 3 (3.12) and the
claudeCLI (Node) - the subscription-auth path. - Repo:
~/orchestrator/(rsync from the Mac, NOT a git clone). The agent-team package lives at~/orchestrator/agent-team/.
New for the coordinator - a dedicated venv under agent-team/.venv (excluded
from rsync) with the coordinator/transport deps:
| pip dep | Why |
|---|---|
langgraph |
the coordinator pipeline graph |
langgraph-checkpoint-sqlite |
SqliteSaver checkpointer against the ledger DB |
claude-agent-sdk |
subscription-auth Claude invocation seam |
slack_sdk |
Slack Web API (post questions, chat:write) |
slack_bolt |
Socket Mode inbound listener (receive answers) |
The ledger DB defaults to agent-team/state/agent_team.sqlite; the audit log to
agent-team/state/audit.log.jsonl. Both live under state/ (gitignored,
never committed).
3. Secrets - append to ~/secrev.env (mode 600, never committed)
The coordinator reads its secrets from the same ~/secrev.env the secrev sweep
uses. Append these (do not echo them into shell history files; lock the file
down after):
echo 'CLAUDE_CODE_OAUTH_TOKEN=...' >> ~/secrev.env # from `claude setup-token`
echo 'SLACK_BOT_TOKEN=xoxb-...' >> ~/secrev.env # bot token, chat:write
echo 'SLACK_APP_TOKEN=xapp-...' >> ~/secrev.env # app-level, connections:write (Socket Mode)
echo 'SLACK_CHANNEL_ID=C0XXXXXXX' >> ~/secrev.env # target clarifier channel
echo 'AGENT_TEAM_SLACK_OWNER_IDS=U0XXXXXXX' >> ~/secrev.env # authorized answerer(s), comma-separated
chmod 600 ~/secrev.env
CLAUDE_CODE_OAUTH_TOKEN- subscription OAuth fromclaude setup-token. The same token type the secrev sweep uses.SLACK_BOT_TOKEN(xoxb-) - bot token withchat:write; posts questions.SLACK_APP_TOKEN(xapp-) - app-level token withconnections:write; required for Socket Mode (opens the inbound WebSocket that hears answers).SLACK_CHANNEL_ID- the channel id the coordinator posts clarifiers to.AGENT_TEAM_SLACK_OWNER_IDS- comma-separated Slack user ids of the authorized answerers (e.g. Adam'sU…id). The inbound listener enforces this as an owner allowlist (AUTHZ-01): only a sender in this set may answer/steer the pipeline. The listener fails closed - if this is unset/empty it rejects every answer (logs a warning namingAGENT_TEAM_SLACK_OWNER_IDS), so it must be set for the human gate to function. Look up your user id via Slack profile → "Copy member ID", or theusers.identity/auth.testAPI.
CRITICAL: ANTHROPIC_API_KEY must NOT be set on this host. A raw API key
would silently win over the subscription OAuth and meter to API rates. The box
runs on subscription OAuth only.
4. Deploy steps
# 4a. From the Mac - rsync the repo (same pattern/excludes as secrev):
rsync -av --exclude .env --exclude .venv \
~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/
# 4b. On the VM - create + activate the agent-team venv and install deps:
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team
python3 -m venv .venv
. .venv/bin/activate
pip install langgraph langgraph-checkpoint-sqlite claude-agent-sdk slack_sdk slack_bolt
# 4c. Initialize the durable ledger DB (idempotent; creates state/agent_team.sqlite):
python3 run-team.py init-db
# 4d. Install + start the service:
sudo cp systemd/agent-team-coordinator.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now agent-team-coordinator.service
# 4e. Verify it is up:
systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e -f
The unit runs python3 run-team.py serve from
WorkingDirectory=/home/adam/orchestrator/agent-team as User=adam, loading
secrets from EnvironmentFile=/home/adam/secrev.env. Restart=on-failure keeps
it up across transient faults; journalctl -u is the live log.
5. P1 live exit-criteria demo (§3.3.1)
Demonstrate all four once the service is live. Map each to the operator commands
(run-team.py list / show / force-resume, systemctl). Run the CLI from the
working dir so it hits the default ledger: cd ~/orchestrator/agent-team.
(a) Crash-safe resume - kill mid-wait, restart, task resumes.
Start a task, get it to a clarifier wait (run-team.py list shows an open
question), then:
sudo systemctl stop agent-team-coordinator.service
sudo systemctl start agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e # confirm the task resumes from the ledger/checkpoint
run-team.py show <question_id> # the question is still open, not lost
Pass: the task picks up the same waiting question after restart (the LangGraph
SqliteSaver checkpoint + the durable ledger survive the kill).
(b) Duplicate Slack answer is a no-op. Answer a question in Slack, then answer the same question again.
run-team.py show <question_id> # status flipped to answered exactly once; answered_via is the first answer
Pass: the first answer wins (rowcount == 1); the duplicate hits the
BEGIN IMMEDIATE compare-and-set and is ignored (rowcount == 0) - no second
resume, no error.
(c) Past-deadline answer is rejected + the task parks.
Let a question's deadline_at pass with no answer, then answer late.
run-team.py show <question_id> # status == expired (auto-expired at deadline)
run-team.py list --parked # the now-parked task surfaces here
Pass: the expired question rejects the late answer and the task parks rather than spins. To un-park it deliberately:
run-team.py force-resume <question_id> --confirm # reopens the expired question for re-delivery
(d) Two concurrent tasks resume independently to the correct thread. Start two tasks concurrently, each reaching its own clarifier wait.
run-team.py list # two distinct open questions, distinct thread_id values
Restart the service (as in (a)); answer each in Slack.
Pass: each task resumes to its own thread_id / channel - no cross-talk, no
answer routed to the wrong task.
6. Rollback
# Stop + disable the service and remove the unit:
sudo systemctl disable --now agent-team-coordinator.service
sudo rm /etc/systemd/system/agent-team-coordinator.service
sudo systemctl daemon-reload
# Restore the VM from the pre-provision Hyper-V checkpoint (§1) to undo
# pip deps + any host changes in one step.
The ledger is local state under agent-team/state/ (not in git). To reset
it without a full snapshot restore: back it up first, then wipe.
cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak} # back up
rm ~/orchestrator/agent-team/state/agent_team.sqlite* # wipe (then re-run init-db)
Note the secrets in ~/secrev.env are NOT removed by rollback - leave them, or
strip the four agent-team keys if you are decommissioning entirely.
7. Security
- Slack inbound listener (Socket Mode) is the auth + untrusted-input surface.
It accepts inbound messages over a WebSocket and turns them into ledger
mutations (answering live clarifier questions). It must pass
/sh-security-reviewbefore this is enabled in production - that review is mandatory for authentication/authorization and untrusted-input handling changes, and this is both. - No IAM / OIDC is involved in P1. The box runs on subscription OAuth
(
CLAUDE_CODE_OAUTH_TOKEN) and Slack tokens only; there is no AWS role, no OIDC trust relationship, no cloud permission surface in this deploy. ~/secrev.envstays mode 600 and out of git;state/(ledger + audit log) is gitignored and written 0600.