This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/agent-team/DEPLOY-R720.md
Adam Moussa 4eb245e83b docs(agent-team): key lives at ~/.ssh/agent-team-apply.pem on the box
The App private key is placed under ~/.ssh (already mode 700) rather than
~/.sea-haven (which holds the synced engineering-handbook). Update the env path,
the secure-copy block, and the rotation/incident-response rm to ~/.ssh.
2026-06-24 17:37:07 -04:00

445 lines
23 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# P1 - agent-team Plane-2 coordinator deployment (R720 VM)
Status: **P1 DEPLOY ARTIFACTS** - the systemd unit + this runbook. The coordinator
daemon (`run-team.py serve`) is built by a separate agent; this file is the
operator runbook for standing it up on the always-on R720 VM.
The agent-team coordinator is the long-running Plane-2 brain: it drives the
LangGraph pipeline, owns the durable `pending_questions` ledger, and runs the
Slack Socket Mode inbound listener that receives clarifier answers. It shares the
`sh-secrev` VM and the `~/secrev.env` secrets file with the Path B security sweep
(see `../security-review/DEPLOY-R720.md`), but it is a **service** (always-on),
not a timer-driven oneshot.
## Host
- **Hypervisor:** R720 at `10.10.60.40` (Windows Server 2022, Hyper-V role).
- **VM:** `sh-secrev`, always-on Ubuntu 24.04, Gen2, **4GB / 2 vCPU / 40GB**
dynamic vhdx at `10.10.60.120`.
- **Reach it:** `ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120` (key-only,
NOPASSWD sudo).
Operate on the VM, not from the Mac against the host by hand.
## 1. SNAPSHOT FIRST
**Standing rule: snapshot the VM before any provisioning change.** This box is a
4GB / 2 vCPU / 40GB VM at `10.10.60.120`. Take a Hyper-V checkpoint on the R720
host **before** you install pip deps, the unit, or touch `~/secrev.env`, so the
whole change is one-command reversible (see ROLLBACK). Do not skip this because
"it is only a pip install" - a bad dep set or a wedged service is exactly what
the snapshot exists to undo.
## 2. Prereqs on the VM
Already present from the secrev deploy:
- **Python 3** (3.12) and the `claude` CLI (Node) - the subscription-auth path.
- **Repo:** `~/orchestrator/` (rsync from the Mac, NOT a git clone). The
agent-team package lives at `~/orchestrator/agent-team/`.
New for the coordinator - a dedicated venv under `agent-team/.venv` (excluded
from rsync) with the coordinator/transport deps:
| pip dep | Why |
|---|---|
| `langgraph` | the coordinator pipeline graph |
| `langgraph-checkpoint-sqlite` | `SqliteSaver` checkpointer against the ledger DB |
| `claude-agent-sdk` | subscription-auth Claude invocation seam |
| `slack_sdk` | Slack Web API (post questions, `chat:write`) |
| `slack_bolt` | Socket Mode inbound listener (receive answers) |
The ledger DB defaults to `agent-team/state/agent_team.sqlite`; the audit log to
`agent-team/state/audit.log.jsonl`. Both live under `state/` (gitignored,
never committed).
## 3. Secrets - append to `~/secrev.env` (mode 600, never committed)
The coordinator reads its secrets from the same `~/secrev.env` the secrev sweep
uses. Append these (do not echo them into shell history files; lock the file
down after):
```
echo 'CLAUDE_CODE_OAUTH_TOKEN=...' >> ~/secrev.env # from `claude setup-token`
echo 'SLACK_BOT_TOKEN=xoxb-...' >> ~/secrev.env # bot token, chat:write
echo 'SLACK_APP_TOKEN=xapp-...' >> ~/secrev.env # app-level, connections:write (Socket Mode)
echo 'SLACK_CHANNEL_ID=C0XXXXXXX' >> ~/secrev.env # target clarifier channel
echo 'AGENT_TEAM_SLACK_OWNER_IDS=U0XXXXXXX' >> ~/secrev.env # authorized answerer(s), comma-separated
chmod 600 ~/secrev.env
```
- `CLAUDE_CODE_OAUTH_TOKEN` - subscription OAuth from `claude setup-token`. The
same token type the secrev sweep uses.
- `SLACK_BOT_TOKEN` (`xoxb-`) - bot token with `chat:write`; posts questions.
- `SLACK_APP_TOKEN` (`xapp-`) - app-level token with `connections:write`;
**required for Socket Mode** (opens the inbound WebSocket that hears answers).
- `SLACK_CHANNEL_ID` - the channel id the coordinator posts clarifiers to.
- `AGENT_TEAM_SLACK_OWNER_IDS` - comma-separated Slack **user ids** of the
authorized answerers (e.g. Adam's `U…` id). The inbound listener enforces this
as an owner allowlist (AUTHZ-01): only a sender in this set may answer/steer
the pipeline. **The listener fails closed** - if this is unset/empty it rejects
**every** answer (logs a warning naming `AGENT_TEAM_SLACK_OWNER_IDS`), so it
must be set for the human gate to function. Look up your user id via Slack
profile → "Copy member ID", or the `users.identity` / `auth.test` API.
**CRITICAL:** `ANTHROPIC_API_KEY` must **NOT** be set on this host. A raw API key
would silently win over the subscription OAuth and meter to API rates. The box
runs on subscription OAuth only.
### 3a. P3 dispatch auth — GitHub App installation tokens (box-side)
The P3 dispatcher (`agent_team/dispatcher.py`) carries an approved diff into org
CI: it pushes the candidate head branch and triggers the
`agent-team-apply-verify.yml` `workflow_dispatch`, then locates the resulting
run id. By default those three seams shell out to `gh`/`git`. **`gh` is not
installed on the box and the box's read-only PAT has no Actions permission**, so
the default path cannot dispatch. Instead the box authenticates with a **GitHub
App**: it holds the App private key and mints short-lived (~1h) **installation
access tokens** on demand (`agent_team/github_app.py`). Set all three of these to
turn the App path on (all-or-nothing — partial config logs one warning and falls
back to the inert gh-default path, it never crashes serve):
The App is the existing **`agent-team-apply`** App: **app_id `4119505`**,
**installation `141992144`** (on `Sea-Haven-Industries/orchestrator`). These two
IDs are not secrets (only the `.pem` is). Set all three vars (all-or-nothing —
partial config logs one warning and falls back to the inert gh-default path, it
never crashes serve):
```
echo 'AGENT_TEAM_GH_APP_ID=4119505' >> ~/secrev.env
echo 'AGENT_TEAM_GH_APP_INSTALLATION_ID=141992144' >> ~/secrev.env
echo 'AGENT_TEAM_GH_APP_PRIVATE_KEY=/home/adam/.ssh/agent-team-apply.pem' >> ~/secrev.env # PATH to the .pem
chmod 600 ~/secrev.env
```
Place the App private key on the box and lock it down — it is a write-capable
credential and must be owner-only. The CI-side Actions secret
`AGENT_APPLY_APP_PRIVATE_KEY` is write-only and cannot be re-exported, so the
box gets its OWN freshly-generated private key for the same App (generate one
under the App settings — the App holds several keys; the CI key keeps working).
Secure-copy it from the Mac (do NOT commit it; it is not in the repo):
```
# From the Mac (the key lives in ~/Downloads after generation):
scp ~/Downloads/agent-team-apply.*.private-key.pem secrev:.ssh/agent-team-apply.pem
# On the box (~/.ssh is already mode 700):
chmod 600 ~/.ssh/agent-team-apply.pem
ls -l ~/.ssh/agent-team-apply.pem # expect -rw-------
# Then DELETE the Mac copy (the box now holds the only working copy):
# rm ~/Downloads/agent-team-apply.*.private-key.pem
```
- `AGENT_TEAM_GH_APP_PRIVATE_KEY` is a **filesystem path** to the `.pem`, not the
key material. The coordinator reads the file at graph-build time; an unreadable
path logs one warning and falls back to gh-default (never crashes).
- **App permission/scope (verified 2026-06-24).** A test mint of an installation
token for app `4119505` / install `141992144` returned
`{"actions":"write","contents":"write","metadata":"read","pull_requests":"write"}`
with `repository_selection: selected` — i.e. Contents R/W + Actions R/W are in
place and the install is scoped to selected repos (not all). Minting grants
exactly the App's scopes; keep it scoped to `Sea-Haven-Industries/orchestrator`
only. (Follow-up: move to a dedicated least-privilege App to replace
`agent-team-apply`, dropping `pull_requests:write` which the box path does not
need.)
- The minted token is short-lived (~1h), is **never** logged, never put in an
exception message, and never written to the ledger/graph state; on the
authenticated push it rides in a host-scoped `http.extraHeader` passed via
`GIT_CONFIG_*` env (never in argv/`ps`).
- **Env-file precedence (verify after editing).** The unit loads **both**
`~/secrev.env` and `~/orchestrator/.env` (last-wins). Confirm only the intended
values are present in both so a stale entry can't shadow the App config:
```
grep -nE 'AGENT_TEAM_GH_APP|AGENT_TEAM_REPO_(OWNER|NAME)' ~/secrev.env ~/orchestrator/.env 2>/dev/null
```
## 4. Deploy steps
```
# 4a. From the Mac - rsync the repo (same pattern/excludes as secrev):
rsync -av --exclude .env --exclude .venv \
~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/
# 4b. On the VM - create + activate the agent-team venv and install deps:
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team
python3 -m venv .venv
. .venv/bin/activate
pip install langgraph langgraph-checkpoint-sqlite claude-agent-sdk slack_sdk slack_bolt
# 4c. Initialize the durable ledger DB (idempotent; creates state/agent_team.sqlite):
python3 run-team.py init-db
# 4d. Install + start the service:
sudo cp systemd/agent-team-coordinator.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now agent-team-coordinator.service
# 4e. Verify it is up:
systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e -f
```
The unit runs `python3 run-team.py serve` from
`WorkingDirectory=/home/adam/orchestrator/agent-team` as `User=adam`, loading
secrets from `EnvironmentFile=/home/adam/secrev.env`. `Restart=on-failure` keeps
it up across transient faults; `journalctl -u` is the live log.
> **⚠️ This deploy carries a ledger migration (SCHEMA_VERSION → 4).** It adds a
> `kind` column to `pending_questions` (values `clarify` | `plan_decision`,
> existing rows default to `clarify`) for the plan-review decision gate. The
> migration is an **additive, idempotent in-place `ALTER TABLE`** run on startup
> (`init_db` / `migrate`, guarded so a second run is a no-op — it never recreates
> the table), so it applies in place against the live box ledger. **Take the
> ledger backup (§1 / the deploy script step) BEFORE restart** — it is the
> migration's safety net (see Rollback, §6). After restart, confirm the column
> landed and the daemon came up clean:
> ```bash
> sqlite3 ~/orchestrator/agent-team/state/agent_team.sqlite \
> "PRAGMA table_info(pending_questions);" | grep kind # expect a 'kind' row
> sqlite3 ~/orchestrator/agent-team/state/agent_team.sqlite \
> "SELECT schema_version FROM schema_meta WHERE id=1;" # expect 4
> journalctl -u agent-team-coordinator.service -e | tail # no migration/import errors
> ```
## 4b. WS0–WS5 rollout — UPDATE an already-deployed box
The steps above (§1–4) are the **first-time** P1 provision. To bring an
already-deployed coordinator up to the WS0–WS5 rollout, use the attended update
script rather than re-running the manual steps:
```
# From the Mac, after the WS branches have merged to main, snapshot first:
agent-team/scripts/deploy-r720-ws-rollout.sh
```
It is an UPDATE (not a provision): it snapshots-reminds, rsyncs the new code,
rsyncs the engineering handbook to the box, installs the new deps, appends the
new secrets if absent, restarts the coordinator, and smoke-tests. It is
idempotent and fails loudly. What goes **live** after it: WS1 in-process
multi-model invokers (`bind_multi_invoker`, already wired), the WS5 handbook
`context_provider` injected into the planner prompt, and the WS2 Slack
`/new-task` command (AUTHZ-01 owner-allowlist gated). The P3 dispatch/build-verify
path stays **inert** (gated behind the `agent-apply` GitHub Environment approval).
**New venv deps** (the coordinator does not need them; only the optional HTTP
API does) — now pinned in the root `requirements.txt`:
| pip dep | Why |
|---|---|
| `fastapi==0.136.1` | the WS1 HTTP API app (`agent_team/api.py`) |
| `uvicorn==0.46.0` | ASGI server for `api.serve()` |
**New env vars** — append to `~/secrev.env` (mode 600, never committed):
- `SEA_HAVEN_HANDBOOK_DIR` — where `load_handbook_conventions()` reads the
engineering handbook (the script syncs it to `/home/adam/.sea-haven/engineering-handbook`
by default; this var must match). Fail-safe: if the dir is missing the
`context_provider` returns `""` and the planner runs without handbook context.
- `AGENT_TEAM_API_TOKEN` — bearer token for the HTTP API / `/delegate` hook
**only**. Not needed by the coordinator daemon itself. The HTTP API refuses to
start if this is unset/empty.
**The HTTP API is a separate, opt-in process** — it is **not** started by the
coordinator daemon. Run it explicitly (`api.serve()`, binds `127.0.0.1:8765`,
bearer auth) only if you want the `/delegate` Claude Code hook or the
`POST /tasks` / `GET /tasks/{thread_id}` / `POST /orchestrator/invoke` endpoints.
The `/docs` + `/openapi` routes are disabled and it binds loopback by design (do
not change to `0.0.0.0`). See the deploy script's step 6 for how to start it.
### Status dashboard (optional, LAN/VPN-only, READ-ONLY)
`agent_team/status_page.py` serves a **live visual pipeline map** of the
coordinator. The top of the page is a hand-rolled inline-SVG diagram of the
agent-team DAG (`INTAKE → CLARIFY ⇄ human gate → PLAN ⇄ REVIEW →
[BUILD → VERIFY → DISPATCH] → DONE`); each stage node is labelled with its model
role (Claude on subscription for CLARIFY/PLAN/VERIFY, GPT-4.1 cross_reviewer for
REVIEW, Gemini for SCAN, DeepSeek fast_coder for BUILD, the Slack owner for the
HUMAN GATE) and is colour-coded by live state — idle / active / awaiting-human
(an **open** pending question) / parked — with a count badge of tasks in that
stage. Hovering (or keyboard-focusing) a node shows what it is working on: the
short `thread_id`, description, status, and waiting age of each task there.
Below the map, the original detail tables remain: all tasks, the human-gate
wait-list, and recent `budget_ledger` spend.
The page **auto-updates without a full reload**: a `GET /api/state` JSON sidecar
returns the same snapshot, and an inline vanilla-JS poller (`fetch()`, no
libraries, no CDN) re-paints node states, counts, the cards, the tooltip data,
and the "last updated" clock every ~4s in place, so hover/scroll/focus survive.
A `<noscript>` 10s meta-refresh is the JS-disabled fallback. Everything
(SVG + CSS + JS) is inline in the served document — nothing is fetched from a
CDN, because the VM is offline/LAN-only. It opens the SQLite ledger
**READ-ONLY** (`mode=ro`), exposes only the two read GETs (`/` and `/api/state`)
and **no mutating endpoints and no auth**.
It is a **separate, optional process** — `agent-team-status.service` (mirrors the
coordinator unit's hardening; `User=adam`, `EnvironmentFile=-/home/adam/secrev.env`,
venv-python ExecStart, `Restart=on-failure`). Unlike the coordinator it needs **no**
`ReadWritePaths` carve-out (it only reads). It can run side-by-side with the
coordinator (RO SQLite opens coexist with the writer).
```bash
# install the unit
sudo cp ~/orchestrator/agent-team/systemd/agent-team-status.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now agent-team-status.service
systemctl status agent-team-status.service
journalctl -u agent-team-status.service -e -f
# or run it ad hoc from the venv
cd ~/orchestrator/agent-team && . .venv/bin/activate && \
python3 -c "from agent_team.status_page import serve; serve()"
```
Then browse `http://10.10.60.120:8770/` from the LAN/VPN (the live map polls
`http://10.10.60.120:8770/api/state` itself).
**Config (env):** `AGENT_TEAM_DB` (default `state/agent_team.sqlite`),
`AGENT_TEAM_STATUS_HOST` (default `0.0.0.0`), `AGENT_TEAM_STATUS_PORT` (default
`8770`).
**Posture:** the sh-secrev VM (`10.10.60.120`, VLAN 60) has no public NIC and sits
behind the UniFi firewall, so `0.0.0.0` reaches the **LAN/VPN only**. Task
descriptions may be sensitive and the page is unauthenticated — **keep it
LAN/VPN-only, never expose it to the public internet.** A missing/locked DB renders
a friendly "no data" page rather than crashing.
## 5. P1 live exit-criteria demo (§3.3.1)
Demonstrate all four once the service is live. Map each to the operator commands
(`run-team.py list / show / force-resume`, `systemctl`). Run the CLI from the
working dir so it hits the default ledger: `cd ~/orchestrator/agent-team`.
**(a) Crash-safe resume - kill mid-wait, restart, task resumes.**
Start a task, get it to a clarifier wait (`run-team.py list` shows an `open`
question), then:
```
sudo systemctl stop agent-team-coordinator.service
sudo systemctl start agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e # confirm the task resumes from the ledger/checkpoint
run-team.py show <question_id> # the question is still open, not lost
```
Pass: the task picks up the same waiting question after restart (the LangGraph
`SqliteSaver` checkpoint + the durable ledger survive the kill).
**(b) Duplicate Slack answer is a no-op.**
Answer a question in Slack, then answer the **same** question again.
```
run-team.py show <question_id> # status flipped to answered exactly once; answered_via is the first answer
```
Pass: the first answer wins (`rowcount == 1`); the duplicate hits the
`BEGIN IMMEDIATE` compare-and-set and is ignored (`rowcount == 0`) - no second
resume, no error.
**(c) Past-deadline answer is rejected + the task parks.**
Let a question's `deadline_at` pass with no answer, then answer late.
```
run-team.py show <question_id> # status == expired (auto-expired at deadline)
run-team.py list --parked # the now-parked task surfaces here
```
Pass: the expired question rejects the late answer and the task parks rather than
spins. To un-park it deliberately:
```
run-team.py force-resume <question_id> --confirm # reopens the expired question for re-delivery
```
**(d) Two concurrent tasks resume independently to the correct thread.**
Start two tasks concurrently, each reaching its own clarifier wait.
```
run-team.py list # two distinct open questions, distinct thread_id values
```
Restart the service (as in (a)); answer each in Slack.
Pass: each task resumes to its own `thread_id` / channel - no cross-talk, no
answer routed to the wrong task.
## 6. Rollback
```
# Stop + disable the service and remove the unit:
sudo systemctl disable --now agent-team-coordinator.service
sudo rm /etc/systemd/system/agent-team-coordinator.service
sudo systemctl daemon-reload
# Restore the VM from the pre-provision Hyper-V checkpoint (§1) to undo
# pip deps + any host changes in one step.
```
The ledger is **local state** under `agent-team/state/` (not in git). To reset
it without a full snapshot restore: back it up first, then wipe.
```
cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak} # back up
rm ~/orchestrator/agent-team/state/agent_team.sqlite* # wipe (then re-run init-db)
```
**Ledger-migration rollback (the schema-v4 `kind` migration).** The deploy takes
a dated ledger backup (`~/agent_team.sqlite.bak-<date>`, written by the
`/sh-deploy-r720` flow / `scripts/deploy-r720.sh`) **before** restart — that
backup is the migration's safety net. The `kind` migration is additive and
idempotent, but **if the `init_db`/`migrate` step fails, or the deploy is rolled
back to pre-v4 code after the migration ran, restore the ledger from that backup
BEFORE restarting the coordinator** (old code does not expect the new column to
matter, but restoring guarantees a clean, pre-migration ledger):
```
sudo systemctl stop agent-team-coordinator.service
cp ~/agent_team.sqlite.bak-<date> ~/orchestrator/agent-team/state/agent_team.sqlite
rm -f ~/orchestrator/agent-team/state/agent_team.sqlite-wal \
~/orchestrator/agent-team/state/agent_team.sqlite-shm # drop stale WAL/SHM
sudo systemctl start agent-team-coordinator.service
```
Restore the ledger backup BEFORE the coordinator restarts — never start the
daemon against a half-migrated or suspect ledger.
Note the secrets in `~/secrev.env` are NOT removed by rollback - leave them, or
strip the four agent-team keys if you are decommissioning entirely.
### GitHub App key — compromise / rotation / revocation
The box holds a write-capable App private key, so it needs its own incident path
(separate from the VM snapshot rollback above):
```
# 1. Stop the daemon so no further token mints happen:
sudo systemctl stop agent-team-coordinator.service
# 2. Remove the key + the App env vars from the box (kills the App path -> inert):
rm -f ~/.ssh/agent-team-apply.pem
sed -i '/AGENT_TEAM_GH_APP_/d' ~/secrev.env
# 3. In GitHub: rotate (generate a new private key, delete the old one) under the
# App's settings, or uninstall the App from the repo to revoke all access.
# Existing installation tokens are short-lived (~1h) and expire on their own.
# 4. Restart inert (gh-default dispatch) and confirm no App env is loaded:
sudo systemctl start agent-team-coordinator.service
sudo systemctl show agent-team-coordinator.service -p Environment | grep -c AGENT_TEAM_GH_APP # expect 0
```
- **Rotation cadence is Adam's call** — there is no automated rotation. Rotate on
any suspected box compromise, on operator turnover, and on a periodic cadence
Adam sets. Document the rotation event in Confluence.
- Removing the App env vars (or the key file) is a safe partial rollback on its
own: the dispatcher falls back to the inert gh-default path (which is itself
inert without `gh`), so no diff is ever pushed — never a fabricated pass.
## 7. Security
- **Slack inbound listener (Socket Mode) is the auth + untrusted-input surface.**
It accepts inbound messages over a WebSocket and turns them into ledger
mutations (answering live clarifier questions). It **must pass
`/sh-security-review`** before this is enabled in production - that review is
mandatory for authentication/authorization and untrusted-input handling
changes, and this is both.
- **No IAM / OIDC is involved in P1.** The box runs on subscription OAuth
(`CLAUDE_CODE_OAUTH_TOKEN`) and Slack tokens only; there is no AWS role, no
OIDC trust relationship, no cloud permission surface in this deploy.
- **GitHub App key on the box (P3 dispatch) — conscious, mitigated trade-off.**
The dispatcher's original design comment said "never a box-held token / operator
host only." Moving dispatch onto the box (so P3 reaches CI without `gh`)
deliberately deviates from that. Mitigations: the App is scoped to **one repo**
with **Contents + Actions only**; minted installation tokens are **short-lived
(~1h)**; the key file is **mode 600**; tokens are **never logged / never in
exception text / never in argv** (push auth rides a host-scoped
`http.extraHeader` via `GIT_CONFIG_*`); CI **re-verifies** the pushed content by
hash and the draft PR is still gated by the `agent-apply` environment's required
reviewer. Any mint/HTTP/git failure **parks** the task (no `run_id` → the verify
gate fails closed) — never a fabricated pass. See §3a for the
permission/scope audit, env-precedence check, and §6 for key rotation/
revocation. **Credential-handling change → `/sh-security-review` is mandatory
before merge** (auth + secret handling surface).
- `~/secrev.env` stays mode 600 and out of git; `state/` (ledger + audit log) is
gitignored and written 0600.