This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/agent-team/DEPLOY-R720.md
Adam Moussa 0f33fa6f24 docs(agent-team): R720 deploy runbook + coordinator systemd unit
P1c deploy artifacts: DEPLOY-R720.md (snapshot-first, rsync, venv deps, ~/secrev.env tokens incl. AGENT_TEAM_SLACK_OWNER_IDS, init-db, systemd, the 4-criteria live demo, rollback) and the long-running coordinator service unit.
2026-06-18 12:56:42 -04:00

9.1 KiB

P1 - agent-team Plane-2 coordinator deployment (R720 VM)

Status: P1 DEPLOY ARTIFACTS - the systemd unit + this runbook. The coordinator daemon (run-team.py serve) is built by a separate agent; this file is the operator runbook for standing it up on the always-on R720 VM.

The agent-team coordinator is the long-running Plane-2 brain: it drives the LangGraph pipeline, owns the durable pending_questions ledger, and runs the Slack Socket Mode inbound listener that receives clarifier answers. It shares the sh-secrev VM and the ~/secrev.env secrets file with the Path B security sweep (see ../security-review/DEPLOY-R720.md), but it is a service (always-on), not a timer-driven oneshot.

Host

  • Hypervisor: R720 at 10.10.60.40 (Windows Server 2022, Hyper-V role).
  • VM: sh-secrev, always-on Ubuntu 24.04, Gen2, 4GB / 2 vCPU / 40GB dynamic vhdx at 10.10.60.120.
  • Reach it: ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120 (key-only, NOPASSWD sudo).

Operate on the VM, not from the Mac against the host by hand.

1. SNAPSHOT FIRST

Standing rule: snapshot the VM before any provisioning change. This box is a 4GB / 2 vCPU / 40GB VM at 10.10.60.120. Take a Hyper-V checkpoint on the R720 host before you install pip deps, the unit, or touch ~/secrev.env, so the whole change is one-command reversible (see ROLLBACK). Do not skip this because "it is only a pip install" - a bad dep set or a wedged service is exactly what the snapshot exists to undo.

2. Prereqs on the VM

Already present from the secrev deploy:

  • Python 3 (3.12) and the claude CLI (Node) - the subscription-auth path.
  • Repo: ~/orchestrator/ (rsync from the Mac, NOT a git clone). The agent-team package lives at ~/orchestrator/agent-team/.

New for the coordinator - a dedicated venv under agent-team/.venv (excluded from rsync) with the coordinator/transport deps:

pip dep Why
langgraph the coordinator pipeline graph
langgraph-checkpoint-sqlite SqliteSaver checkpointer against the ledger DB
claude-agent-sdk subscription-auth Claude invocation seam
slack_sdk Slack Web API (post questions, chat:write)
slack_bolt Socket Mode inbound listener (receive answers)

The ledger DB defaults to agent-team/state/agent_team.sqlite; the audit log to agent-team/state/audit.log.jsonl. Both live under state/ (gitignored, never committed).

3. Secrets - append to ~/secrev.env (mode 600, never committed)

The coordinator reads its secrets from the same ~/secrev.env the secrev sweep uses. Append these (do not echo them into shell history files; lock the file down after):

echo 'CLAUDE_CODE_OAUTH_TOKEN=...'   >> ~/secrev.env   # from `claude setup-token`
echo 'SLACK_BOT_TOKEN=xoxb-...'      >> ~/secrev.env   # bot token, chat:write
echo 'SLACK_APP_TOKEN=xapp-...'      >> ~/secrev.env   # app-level, connections:write (Socket Mode)
echo 'SLACK_CHANNEL_ID=C0XXXXXXX'    >> ~/secrev.env   # target clarifier channel
echo 'AGENT_TEAM_SLACK_OWNER_IDS=U0XXXXXXX'  >> ~/secrev.env   # authorized answerer(s), comma-separated
chmod 600 ~/secrev.env
  • CLAUDE_CODE_OAUTH_TOKEN - subscription OAuth from claude setup-token. The same token type the secrev sweep uses.
  • SLACK_BOT_TOKEN (xoxb-) - bot token with chat:write; posts questions.
  • SLACK_APP_TOKEN (xapp-) - app-level token with connections:write; required for Socket Mode (opens the inbound WebSocket that hears answers).
  • SLACK_CHANNEL_ID - the channel id the coordinator posts clarifiers to.
  • AGENT_TEAM_SLACK_OWNER_IDS - comma-separated Slack user ids of the authorized answerers (e.g. Adam's U… id). The inbound listener enforces this as an owner allowlist (AUTHZ-01): only a sender in this set may answer/steer the pipeline. The listener fails closed - if this is unset/empty it rejects every answer (logs a warning naming AGENT_TEAM_SLACK_OWNER_IDS), so it must be set for the human gate to function. Look up your user id via Slack profile → "Copy member ID", or the users.identity / auth.test API.

CRITICAL: ANTHROPIC_API_KEY must NOT be set on this host. A raw API key would silently win over the subscription OAuth and meter to API rates. The box runs on subscription OAuth only.

4. Deploy steps

# 4a. From the Mac - rsync the repo (same pattern/excludes as secrev):
rsync -av --exclude .env --exclude .venv \
  ~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/

# 4b. On the VM - create + activate the agent-team venv and install deps:
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team
python3 -m venv .venv
. .venv/bin/activate
pip install langgraph langgraph-checkpoint-sqlite claude-agent-sdk slack_sdk slack_bolt

# 4c. Initialize the durable ledger DB (idempotent; creates state/agent_team.sqlite):
python3 run-team.py init-db

# 4d. Install + start the service:
sudo cp systemd/agent-team-coordinator.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now agent-team-coordinator.service

# 4e. Verify it is up:
systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e -f

The unit runs python3 run-team.py serve from WorkingDirectory=/home/adam/orchestrator/agent-team as User=adam, loading secrets from EnvironmentFile=/home/adam/secrev.env. Restart=on-failure keeps it up across transient faults; journalctl -u is the live log.

5. P1 live exit-criteria demo (§3.3.1)

Demonstrate all four once the service is live. Map each to the operator commands (run-team.py list / show / force-resume, systemctl). Run the CLI from the working dir so it hits the default ledger: cd ~/orchestrator/agent-team.

(a) Crash-safe resume - kill mid-wait, restart, task resumes. Start a task, get it to a clarifier wait (run-team.py list shows an open question), then:

sudo systemctl stop agent-team-coordinator.service
sudo systemctl start agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e   # confirm the task resumes from the ledger/checkpoint
run-team.py show <question_id>                     # the question is still open, not lost

Pass: the task picks up the same waiting question after restart (the LangGraph SqliteSaver checkpoint + the durable ledger survive the kill).

(b) Duplicate Slack answer is a no-op. Answer a question in Slack, then answer the same question again.

run-team.py show <question_id>   # status flipped to answered exactly once; answered_via is the first answer

Pass: the first answer wins (rowcount == 1); the duplicate hits the BEGIN IMMEDIATE compare-and-set and is ignored (rowcount == 0) - no second resume, no error.

(c) Past-deadline answer is rejected + the task parks. Let a question's deadline_at pass with no answer, then answer late.

run-team.py show <question_id>           # status == expired (auto-expired at deadline)
run-team.py list --parked                # the now-parked task surfaces here

Pass: the expired question rejects the late answer and the task parks rather than spins. To un-park it deliberately:

run-team.py force-resume <question_id> --confirm   # reopens the expired question for re-delivery

(d) Two concurrent tasks resume independently to the correct thread. Start two tasks concurrently, each reaching its own clarifier wait.

run-team.py list   # two distinct open questions, distinct thread_id values

Restart the service (as in (a)); answer each in Slack. Pass: each task resumes to its own thread_id / channel - no cross-talk, no answer routed to the wrong task.

6. Rollback

# Stop + disable the service and remove the unit:
sudo systemctl disable --now agent-team-coordinator.service
sudo rm /etc/systemd/system/agent-team-coordinator.service
sudo systemctl daemon-reload

# Restore the VM from the pre-provision Hyper-V checkpoint (§1) to undo
# pip deps + any host changes in one step.

The ledger is local state under agent-team/state/ (not in git). To reset it without a full snapshot restore: back it up first, then wipe.

cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak}   # back up
rm ~/orchestrator/agent-team/state/agent_team.sqlite*         # wipe (then re-run init-db)

Note the secrets in ~/secrev.env are NOT removed by rollback - leave them, or strip the four agent-team keys if you are decommissioning entirely.

7. Security

  • Slack inbound listener (Socket Mode) is the auth + untrusted-input surface. It accepts inbound messages over a WebSocket and turns them into ledger mutations (answering live clarifier questions). It must pass /sh-security-review before this is enabled in production - that review is mandatory for authentication/authorization and untrusted-input handling changes, and this is both.
  • No IAM / OIDC is involved in P1. The box runs on subscription OAuth (CLAUDE_CODE_OAUTH_TOKEN) and Slack tokens only; there is no AWS role, no OIDC trust relationship, no cloud permission surface in this deploy.
  • ~/secrev.env stays mode 600 and out of git; state/ (ledger + audit log) is gitignored and written 0600.