This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/agent-team/DEPLOY-R720.md
Adam Moussa 4eb245e83b docs(agent-team): key lives at ~/.ssh/agent-team-apply.pem on the box
The App private key is placed under ~/.ssh (already mode 700) rather than
~/.sea-haven (which holds the synced engineering-handbook). Update the env path,
the secure-copy block, and the rotation/incident-response rm to ~/.ssh.
2026-06-24 17:37:07 -04:00

23 KiB
Raw Blame History

P1 - agent-team Plane-2 coordinator deployment (R720 VM)

Status: P1 DEPLOY ARTIFACTS - the systemd unit + this runbook. The coordinator daemon (run-team.py serve) is built by a separate agent; this file is the operator runbook for standing it up on the always-on R720 VM.

The agent-team coordinator is the long-running Plane-2 brain: it drives the LangGraph pipeline, owns the durable pending_questions ledger, and runs the Slack Socket Mode inbound listener that receives clarifier answers. It shares the sh-secrev VM and the ~/secrev.env secrets file with the Path B security sweep (see ../security-review/DEPLOY-R720.md), but it is a service (always-on), not a timer-driven oneshot.

Host

  • Hypervisor: R720 at 10.10.60.40 (Windows Server 2022, Hyper-V role).
  • VM: sh-secrev, always-on Ubuntu 24.04, Gen2, 4GB / 2 vCPU / 40GB dynamic vhdx at 10.10.60.120.
  • Reach it: ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120 (key-only, NOPASSWD sudo).

Operate on the VM, not from the Mac against the host by hand.

1. SNAPSHOT FIRST

Standing rule: snapshot the VM before any provisioning change. This box is a 4GB / 2 vCPU / 40GB VM at 10.10.60.120. Take a Hyper-V checkpoint on the R720 host before you install pip deps, the unit, or touch ~/secrev.env, so the whole change is one-command reversible (see ROLLBACK). Do not skip this because "it is only a pip install" - a bad dep set or a wedged service is exactly what the snapshot exists to undo.

2. Prereqs on the VM

Already present from the secrev deploy:

  • Python 3 (3.12) and the claude CLI (Node) - the subscription-auth path.
  • Repo: ~/orchestrator/ (rsync from the Mac, NOT a git clone). The agent-team package lives at ~/orchestrator/agent-team/.

New for the coordinator - a dedicated venv under agent-team/.venv (excluded from rsync) with the coordinator/transport deps:

pip dep Why
langgraph the coordinator pipeline graph
langgraph-checkpoint-sqlite SqliteSaver checkpointer against the ledger DB
claude-agent-sdk subscription-auth Claude invocation seam
slack_sdk Slack Web API (post questions, chat:write)
slack_bolt Socket Mode inbound listener (receive answers)

The ledger DB defaults to agent-team/state/agent_team.sqlite; the audit log to agent-team/state/audit.log.jsonl. Both live under state/ (gitignored, never committed).

3. Secrets - append to ~/secrev.env (mode 600, never committed)

The coordinator reads its secrets from the same ~/secrev.env the secrev sweep uses. Append these (do not echo them into shell history files; lock the file down after):

echo 'CLAUDE_CODE_OAUTH_TOKEN=...'   >> ~/secrev.env   # from `claude setup-token`
echo 'SLACK_BOT_TOKEN=xoxb-...'      >> ~/secrev.env   # bot token, chat:write
echo 'SLACK_APP_TOKEN=xapp-...'      >> ~/secrev.env   # app-level, connections:write (Socket Mode)
echo 'SLACK_CHANNEL_ID=C0XXXXXXX'    >> ~/secrev.env   # target clarifier channel
echo 'AGENT_TEAM_SLACK_OWNER_IDS=U0XXXXXXX'  >> ~/secrev.env   # authorized answerer(s), comma-separated
chmod 600 ~/secrev.env
  • CLAUDE_CODE_OAUTH_TOKEN - subscription OAuth from claude setup-token. The same token type the secrev sweep uses.
  • SLACK_BOT_TOKEN (xoxb-) - bot token with chat:write; posts questions.
  • SLACK_APP_TOKEN (xapp-) - app-level token with connections:write; required for Socket Mode (opens the inbound WebSocket that hears answers).
  • SLACK_CHANNEL_ID - the channel id the coordinator posts clarifiers to.
  • AGENT_TEAM_SLACK_OWNER_IDS - comma-separated Slack user ids of the authorized answerers (e.g. Adam's U… id). The inbound listener enforces this as an owner allowlist (AUTHZ-01): only a sender in this set may answer/steer the pipeline. The listener fails closed - if this is unset/empty it rejects every answer (logs a warning naming AGENT_TEAM_SLACK_OWNER_IDS), so it must be set for the human gate to function. Look up your user id via Slack profile → "Copy member ID", or the users.identity / auth.test API.

CRITICAL: ANTHROPIC_API_KEY must NOT be set on this host. A raw API key would silently win over the subscription OAuth and meter to API rates. The box runs on subscription OAuth only.

3a. P3 dispatch auth — GitHub App installation tokens (box-side)

The P3 dispatcher (agent_team/dispatcher.py) carries an approved diff into org CI: it pushes the candidate head branch and triggers the agent-team-apply-verify.yml workflow_dispatch, then locates the resulting run id. By default those three seams shell out to gh/git. gh is not installed on the box and the box's read-only PAT has no Actions permission, so the default path cannot dispatch. Instead the box authenticates with a GitHub App: it holds the App private key and mints short-lived (~1h) installation access tokens on demand (agent_team/github_app.py). Set all three of these to turn the App path on (all-or-nothing — partial config logs one warning and falls back to the inert gh-default path, it never crashes serve):

The App is the existing agent-team-apply App: app_id 4119505, installation 141992144 (on Sea-Haven-Industries/orchestrator). These two IDs are not secrets (only the .pem is). Set all three vars (all-or-nothing — partial config logs one warning and falls back to the inert gh-default path, it never crashes serve):

echo 'AGENT_TEAM_GH_APP_ID=4119505'              >> ~/secrev.env
echo 'AGENT_TEAM_GH_APP_INSTALLATION_ID=141992144' >> ~/secrev.env
echo 'AGENT_TEAM_GH_APP_PRIVATE_KEY=/home/adam/.ssh/agent-team-apply.pem' >> ~/secrev.env   # PATH to the .pem
chmod 600 ~/secrev.env

Place the App private key on the box and lock it down — it is a write-capable credential and must be owner-only. The CI-side Actions secret AGENT_APPLY_APP_PRIVATE_KEY is write-only and cannot be re-exported, so the box gets its OWN freshly-generated private key for the same App (generate one under the App settings — the App holds several keys; the CI key keeps working). Secure-copy it from the Mac (do NOT commit it; it is not in the repo):

# From the Mac (the key lives in ~/Downloads after generation):
scp ~/Downloads/agent-team-apply.*.private-key.pem secrev:.ssh/agent-team-apply.pem
# On the box (~/.ssh is already mode 700):
chmod 600 ~/.ssh/agent-team-apply.pem
ls -l ~/.ssh/agent-team-apply.pem   # expect -rw-------
# Then DELETE the Mac copy (the box now holds the only working copy):
#   rm ~/Downloads/agent-team-apply.*.private-key.pem
  • AGENT_TEAM_GH_APP_PRIVATE_KEY is a filesystem path to the .pem, not the key material. The coordinator reads the file at graph-build time; an unreadable path logs one warning and falls back to gh-default (never crashes).
  • App permission/scope (verified 2026-06-24). A test mint of an installation token for app 4119505 / install 141992144 returned {"actions":"write","contents":"write","metadata":"read","pull_requests":"write"} with repository_selection: selected — i.e. Contents R/W + Actions R/W are in place and the install is scoped to selected repos (not all). Minting grants exactly the App's scopes; keep it scoped to Sea-Haven-Industries/orchestrator only. (Follow-up: move to a dedicated least-privilege App to replace agent-team-apply, dropping pull_requests:write which the box path does not need.)
  • The minted token is short-lived (~1h), is never logged, never put in an exception message, and never written to the ledger/graph state; on the authenticated push it rides in a host-scoped http.extraHeader passed via GIT_CONFIG_* env (never in argv/ps).
  • Env-file precedence (verify after editing). The unit loads both ~/secrev.env and ~/orchestrator/.env (last-wins). Confirm only the intended values are present in both so a stale entry can't shadow the App config:
    grep -nE 'AGENT_TEAM_GH_APP|AGENT_TEAM_REPO_(OWNER|NAME)' ~/secrev.env ~/orchestrator/.env 2>/dev/null
    

4. Deploy steps

# 4a. From the Mac - rsync the repo (same pattern/excludes as secrev):
rsync -av --exclude .env --exclude .venv \
  ~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/

# 4b. On the VM - create + activate the agent-team venv and install deps:
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team
python3 -m venv .venv
. .venv/bin/activate
pip install langgraph langgraph-checkpoint-sqlite claude-agent-sdk slack_sdk slack_bolt

# 4c. Initialize the durable ledger DB (idempotent; creates state/agent_team.sqlite):
python3 run-team.py init-db

# 4d. Install + start the service:
sudo cp systemd/agent-team-coordinator.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now agent-team-coordinator.service

# 4e. Verify it is up:
systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e -f

The unit runs python3 run-team.py serve from WorkingDirectory=/home/adam/orchestrator/agent-team as User=adam, loading secrets from EnvironmentFile=/home/adam/secrev.env. Restart=on-failure keeps it up across transient faults; journalctl -u is the live log.

⚠️ This deploy carries a ledger migration (SCHEMA_VERSION → 4). It adds a kind column to pending_questions (values clarify | plan_decision, existing rows default to clarify) for the plan-review decision gate. The migration is an additive, idempotent in-place ALTER TABLE run on startup (init_db / migrate, guarded so a second run is a no-op — it never recreates the table), so it applies in place against the live box ledger. Take the ledger backup (§1 / the deploy script step) BEFORE restart — it is the migration's safety net (see Rollback, §6). After restart, confirm the column landed and the daemon came up clean:

sqlite3 ~/orchestrator/agent-team/state/agent_team.sqlite \
  "PRAGMA table_info(pending_questions);" | grep kind   # expect a 'kind' row
sqlite3 ~/orchestrator/agent-team/state/agent_team.sqlite \
  "SELECT schema_version FROM schema_meta WHERE id=1;"  # expect 4
journalctl -u agent-team-coordinator.service -e | tail  # no migration/import errors

4b. WS0–WS5 rollout — UPDATE an already-deployed box

The steps above (§1–4) are the first-time P1 provision. To bring an already-deployed coordinator up to the WS0–WS5 rollout, use the attended update script rather than re-running the manual steps:

# From the Mac, after the WS branches have merged to main, snapshot first:
agent-team/scripts/deploy-r720-ws-rollout.sh

It is an UPDATE (not a provision): it snapshots-reminds, rsyncs the new code, rsyncs the engineering handbook to the box, installs the new deps, appends the new secrets if absent, restarts the coordinator, and smoke-tests. It is idempotent and fails loudly. What goes live after it: WS1 in-process multi-model invokers (bind_multi_invoker, already wired), the WS5 handbook context_provider injected into the planner prompt, and the WS2 Slack /new-task command (AUTHZ-01 owner-allowlist gated). The P3 dispatch/build-verify path stays inert (gated behind the agent-apply GitHub Environment approval).

New venv deps (the coordinator does not need them; only the optional HTTP API does) — now pinned in the root requirements.txt:

pip dep Why
fastapi==0.136.1 the WS1 HTTP API app (agent_team/api.py)
uvicorn==0.46.0 ASGI server for api.serve()

New env vars — append to ~/secrev.env (mode 600, never committed):

  • SEA_HAVEN_HANDBOOK_DIR — where load_handbook_conventions() reads the engineering handbook (the script syncs it to /home/adam/.sea-haven/engineering-handbook by default; this var must match). Fail-safe: if the dir is missing the context_provider returns "" and the planner runs without handbook context.
  • AGENT_TEAM_API_TOKEN — bearer token for the HTTP API / /delegate hook only. Not needed by the coordinator daemon itself. The HTTP API refuses to start if this is unset/empty.

The HTTP API is a separate, opt-in process — it is not started by the coordinator daemon. Run it explicitly (api.serve(), binds 127.0.0.1:8765, bearer auth) only if you want the /delegate Claude Code hook or the POST /tasks / GET /tasks/{thread_id} / POST /orchestrator/invoke endpoints. The /docs + /openapi routes are disabled and it binds loopback by design (do not change to 0.0.0.0). See the deploy script's step 6 for how to start it.

Status dashboard (optional, LAN/VPN-only, READ-ONLY)

agent_team/status_page.py serves a live visual pipeline map of the coordinator. The top of the page is a hand-rolled inline-SVG diagram of the agent-team DAG (INTAKE → CLARIFY ⇄ human gate → PLAN ⇄ REVIEW → [BUILD → VERIFY → DISPATCH] → DONE); each stage node is labelled with its model role (Claude on subscription for CLARIFY/PLAN/VERIFY, GPT-4.1 cross_reviewer for REVIEW, Gemini for SCAN, DeepSeek fast_coder for BUILD, the Slack owner for the HUMAN GATE) and is colour-coded by live state — idle / active / awaiting-human (an open pending question) / parked — with a count badge of tasks in that stage. Hovering (or keyboard-focusing) a node shows what it is working on: the short thread_id, description, status, and waiting age of each task there. Below the map, the original detail tables remain: all tasks, the human-gate wait-list, and recent budget_ledger spend.

The page auto-updates without a full reload: a GET /api/state JSON sidecar returns the same snapshot, and an inline vanilla-JS poller (fetch(), no libraries, no CDN) re-paints node states, counts, the cards, the tooltip data, and the "last updated" clock every ~4s in place, so hover/scroll/focus survive. A <noscript> 10s meta-refresh is the JS-disabled fallback. Everything (SVG + CSS + JS) is inline in the served document — nothing is fetched from a CDN, because the VM is offline/LAN-only. It opens the SQLite ledger READ-ONLY (mode=ro), exposes only the two read GETs (/ and /api/state) and no mutating endpoints and no auth.

It is a separate, optional process — agent-team-status.service (mirrors the coordinator unit's hardening; User=adam, EnvironmentFile=-/home/adam/secrev.env, venv-python ExecStart, Restart=on-failure). Unlike the coordinator it needs no ReadWritePaths carve-out (it only reads). It can run side-by-side with the coordinator (RO SQLite opens coexist with the writer).

# install the unit
sudo cp ~/orchestrator/agent-team/systemd/agent-team-status.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now agent-team-status.service
systemctl status agent-team-status.service
journalctl -u agent-team-status.service -e -f

# or run it ad hoc from the venv
cd ~/orchestrator/agent-team && . .venv/bin/activate && \
  python3 -c "from agent_team.status_page import serve; serve()"

Then browse http://10.10.60.120:8770/ from the LAN/VPN (the live map polls http://10.10.60.120:8770/api/state itself).

Config (env): AGENT_TEAM_DB (default state/agent_team.sqlite), AGENT_TEAM_STATUS_HOST (default 0.0.0.0), AGENT_TEAM_STATUS_PORT (default 8770).

Posture: the sh-secrev VM (10.10.60.120, VLAN 60) has no public NIC and sits behind the UniFi firewall, so 0.0.0.0 reaches the LAN/VPN only. Task descriptions may be sensitive and the page is unauthenticated — keep it LAN/VPN-only, never expose it to the public internet. A missing/locked DB renders a friendly "no data" page rather than crashing.

5. P1 live exit-criteria demo (§3.3.1)

Demonstrate all four once the service is live. Map each to the operator commands (run-team.py list / show / force-resume, systemctl). Run the CLI from the working dir so it hits the default ledger: cd ~/orchestrator/agent-team.

(a) Crash-safe resume - kill mid-wait, restart, task resumes. Start a task, get it to a clarifier wait (run-team.py list shows an open question), then:

sudo systemctl stop agent-team-coordinator.service
sudo systemctl start agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e   # confirm the task resumes from the ledger/checkpoint
run-team.py show <question_id>                     # the question is still open, not lost

Pass: the task picks up the same waiting question after restart (the LangGraph SqliteSaver checkpoint + the durable ledger survive the kill).

(b) Duplicate Slack answer is a no-op. Answer a question in Slack, then answer the same question again.

run-team.py show <question_id>   # status flipped to answered exactly once; answered_via is the first answer

Pass: the first answer wins (rowcount == 1); the duplicate hits the BEGIN IMMEDIATE compare-and-set and is ignored (rowcount == 0) - no second resume, no error.

(c) Past-deadline answer is rejected + the task parks. Let a question's deadline_at pass with no answer, then answer late.

run-team.py show <question_id>           # status == expired (auto-expired at deadline)
run-team.py list --parked                # the now-parked task surfaces here

Pass: the expired question rejects the late answer and the task parks rather than spins. To un-park it deliberately:

run-team.py force-resume <question_id> --confirm   # reopens the expired question for re-delivery

(d) Two concurrent tasks resume independently to the correct thread. Start two tasks concurrently, each reaching its own clarifier wait.

run-team.py list   # two distinct open questions, distinct thread_id values

Restart the service (as in (a)); answer each in Slack. Pass: each task resumes to its own thread_id / channel - no cross-talk, no answer routed to the wrong task.

6. Rollback

# Stop + disable the service and remove the unit:
sudo systemctl disable --now agent-team-coordinator.service
sudo rm /etc/systemd/system/agent-team-coordinator.service
sudo systemctl daemon-reload

# Restore the VM from the pre-provision Hyper-V checkpoint (§1) to undo
# pip deps + any host changes in one step.

The ledger is local state under agent-team/state/ (not in git). To reset it without a full snapshot restore: back it up first, then wipe.

cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak}   # back up
rm ~/orchestrator/agent-team/state/agent_team.sqlite*         # wipe (then re-run init-db)

Ledger-migration rollback (the schema-v4 kind migration). The deploy takes a dated ledger backup (~/agent_team.sqlite.bak-<date>, written by the /sh-deploy-r720 flow / scripts/deploy-r720.sh) before restart — that backup is the migration's safety net. The kind migration is additive and idempotent, but if the init_db/migrate step fails, or the deploy is rolled back to pre-v4 code after the migration ran, restore the ledger from that backup BEFORE restarting the coordinator (old code does not expect the new column to matter, but restoring guarantees a clean, pre-migration ledger):

sudo systemctl stop agent-team-coordinator.service
cp ~/agent_team.sqlite.bak-<date> ~/orchestrator/agent-team/state/agent_team.sqlite
rm -f ~/orchestrator/agent-team/state/agent_team.sqlite-wal \
      ~/orchestrator/agent-team/state/agent_team.sqlite-shm   # drop stale WAL/SHM
sudo systemctl start agent-team-coordinator.service

Restore the ledger backup BEFORE the coordinator restarts — never start the daemon against a half-migrated or suspect ledger. Note the secrets in ~/secrev.env are NOT removed by rollback - leave them, or strip the four agent-team keys if you are decommissioning entirely.

GitHub App key — compromise / rotation / revocation

The box holds a write-capable App private key, so it needs its own incident path (separate from the VM snapshot rollback above):

# 1. Stop the daemon so no further token mints happen:
sudo systemctl stop agent-team-coordinator.service
# 2. Remove the key + the App env vars from the box (kills the App path -> inert):
rm -f ~/.ssh/agent-team-apply.pem
sed -i '/AGENT_TEAM_GH_APP_/d' ~/secrev.env
# 3. In GitHub: rotate (generate a new private key, delete the old one) under the
#    App's settings, or uninstall the App from the repo to revoke all access.
#    Existing installation tokens are short-lived (~1h) and expire on their own.
# 4. Restart inert (gh-default dispatch) and confirm no App env is loaded:
sudo systemctl start agent-team-coordinator.service
sudo systemctl show agent-team-coordinator.service -p Environment | grep -c AGENT_TEAM_GH_APP   # expect 0
  • Rotation cadence is Adam's call — there is no automated rotation. Rotate on any suspected box compromise, on operator turnover, and on a periodic cadence Adam sets. Document the rotation event in Confluence.
  • Removing the App env vars (or the key file) is a safe partial rollback on its own: the dispatcher falls back to the inert gh-default path (which is itself inert without gh), so no diff is ever pushed — never a fabricated pass.

7. Security

  • Slack inbound listener (Socket Mode) is the auth + untrusted-input surface. It accepts inbound messages over a WebSocket and turns them into ledger mutations (answering live clarifier questions). It must pass /sh-security-review before this is enabled in production - that review is mandatory for authentication/authorization and untrusted-input handling changes, and this is both.
  • No IAM / OIDC is involved in P1. The box runs on subscription OAuth (CLAUDE_CODE_OAUTH_TOKEN) and Slack tokens only; there is no AWS role, no OIDC trust relationship, no cloud permission surface in this deploy.
  • GitHub App key on the box (P3 dispatch) — conscious, mitigated trade-off. The dispatcher's original design comment said "never a box-held token / operator host only." Moving dispatch onto the box (so P3 reaches CI without gh) deliberately deviates from that. Mitigations: the App is scoped to one repo with Contents + Actions only; minted installation tokens are short-lived (~1h); the key file is mode 600; tokens are never logged / never in exception text / never in argv (push auth rides a host-scoped http.extraHeader via GIT_CONFIG_*); CI re-verifies the pushed content by hash and the draft PR is still gated by the agent-apply environment's required reviewer. Any mint/HTTP/git failure parks the task (no run_id → the verify gate fails closed) — never a fabricated pass. See §3a for the permission/scope audit, env-precedence check, and §6 for key rotation/ revocation. Credential-handling change → /sh-security-review is mandatory before merge (auth + secret handling surface).
  • ~/secrev.env stays mode 600 and out of git; state/ (ledger + audit log) is gitignored and written 0600.