This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/agent-team/DEPLOY-R720.md

247 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# P1 - agent-team Plane-2 coordinator deployment (R720 VM)
Status: **P1 DEPLOY ARTIFACTS** - the systemd unit + this runbook. The coordinator
daemon (`run-team.py serve`) is built by a separate agent; this file is the
operator runbook for standing it up on the always-on R720 VM.
The agent-team coordinator is the long-running Plane-2 brain: it drives the
LangGraph pipeline, owns the durable `pending_questions` ledger, and runs the
Slack Socket Mode inbound listener that receives clarifier answers. It shares the
`sh-secrev` VM and the `~/secrev.env` secrets file with the Path B security sweep
(see `../security-review/DEPLOY-R720.md`), but it is a **service** (always-on),
not a timer-driven oneshot.
## Host
- **Hypervisor:** R720 at `10.10.60.40` (Windows Server 2022, Hyper-V role).
- **VM:** `sh-secrev`, always-on Ubuntu 24.04, Gen2, **4GB / 2 vCPU / 40GB**
dynamic vhdx at `10.10.60.120`.
- **Reach it:** `ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120` (key-only,
NOPASSWD sudo).
Operate on the VM, not from the Mac against the host by hand.
## 1. SNAPSHOT FIRST
**Standing rule: snapshot the VM before any provisioning change.** This box is a
4GB / 2 vCPU / 40GB VM at `10.10.60.120`. Take a Hyper-V checkpoint on the R720
host **before** you install pip deps, the unit, or touch `~/secrev.env`, so the
whole change is one-command reversible (see ROLLBACK). Do not skip this because
"it is only a pip install" - a bad dep set or a wedged service is exactly what
the snapshot exists to undo.
## 2. Prereqs on the VM
Already present from the secrev deploy:
- **Python 3** (3.12) and the `claude` CLI (Node) - the subscription-auth path.
- **Repo:** `~/orchestrator/` (rsync from the Mac, NOT a git clone). The
agent-team package lives at `~/orchestrator/agent-team/`.
New for the coordinator - a dedicated venv under `agent-team/.venv` (excluded
from rsync) with the coordinator/transport deps:
| pip dep | Why |
|---|---|
| `langgraph` | the coordinator pipeline graph |
| `langgraph-checkpoint-sqlite` | `SqliteSaver` checkpointer against the ledger DB |
| `claude-agent-sdk` | subscription-auth Claude invocation seam |
| `slack_sdk` | Slack Web API (post questions, `chat:write`) |
| `slack_bolt` | Socket Mode inbound listener (receive answers) |
The ledger DB defaults to `agent-team/state/agent_team.sqlite`; the audit log to
`agent-team/state/audit.log.jsonl`. Both live under `state/` (gitignored,
never committed).
## 3. Secrets - append to `~/secrev.env` (mode 600, never committed)
The coordinator reads its secrets from the same `~/secrev.env` the secrev sweep
uses. Append these (do not echo them into shell history files; lock the file
down after):
```
echo 'CLAUDE_CODE_OAUTH_TOKEN=...' >> ~/secrev.env # from `claude setup-token`
echo 'SLACK_BOT_TOKEN=xoxb-...' >> ~/secrev.env # bot token, chat:write
echo 'SLACK_APP_TOKEN=xapp-...' >> ~/secrev.env # app-level, connections:write (Socket Mode)
echo 'SLACK_CHANNEL_ID=C0XXXXXXX' >> ~/secrev.env # target clarifier channel
echo 'AGENT_TEAM_SLACK_OWNER_IDS=U0XXXXXXX' >> ~/secrev.env # authorized answerer(s), comma-separated
chmod 600 ~/secrev.env
```
- `CLAUDE_CODE_OAUTH_TOKEN` - subscription OAuth from `claude setup-token`. The
same token type the secrev sweep uses.
- `SLACK_BOT_TOKEN` (`xoxb-`) - bot token with `chat:write`; posts questions.
- `SLACK_APP_TOKEN` (`xapp-`) - app-level token with `connections:write`;
**required for Socket Mode** (opens the inbound WebSocket that hears answers).
- `SLACK_CHANNEL_ID` - the channel id the coordinator posts clarifiers to.
- `AGENT_TEAM_SLACK_OWNER_IDS` - comma-separated Slack **user ids** of the
authorized answerers (e.g. Adam's `U…` id). The inbound listener enforces this
as an owner allowlist (AUTHZ-01): only a sender in this set may answer/steer
the pipeline. **The listener fails closed** - if this is unset/empty it rejects
**every** answer (logs a warning naming `AGENT_TEAM_SLACK_OWNER_IDS`), so it
must be set for the human gate to function. Look up your user id via Slack
profile → "Copy member ID", or the `users.identity` / `auth.test` API.
**CRITICAL:** `ANTHROPIC_API_KEY` must **NOT** be set on this host. A raw API key
would silently win over the subscription OAuth and meter to API rates. The box
runs on subscription OAuth only.
## 4. Deploy steps
```
# 4a. From the Mac - rsync the repo (same pattern/excludes as secrev):
rsync -av --exclude .env --exclude .venv \
~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/
# 4b. On the VM - create + activate the agent-team venv and install deps:
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team
python3 -m venv .venv
. .venv/bin/activate
pip install langgraph langgraph-checkpoint-sqlite claude-agent-sdk slack_sdk slack_bolt
# 4c. Initialize the durable ledger DB (idempotent; creates state/agent_team.sqlite):
python3 run-team.py init-db
# 4d. Install + start the service:
sudo cp systemd/agent-team-coordinator.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now agent-team-coordinator.service
# 4e. Verify it is up:
systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e -f
```
The unit runs `python3 run-team.py serve` from
`WorkingDirectory=/home/adam/orchestrator/agent-team` as `User=adam`, loading
secrets from `EnvironmentFile=/home/adam/secrev.env`. `Restart=on-failure` keeps
it up across transient faults; `journalctl -u` is the live log.
## 4b. WS0–WS5 rollout — UPDATE an already-deployed box
The steps above (§1–4) are the **first-time** P1 provision. To bring an
already-deployed coordinator up to the WS0–WS5 rollout, use the attended update
script rather than re-running the manual steps:
```
# From the Mac, after the WS branches have merged to main, snapshot first:
agent-team/scripts/deploy-r720-ws-rollout.sh
```
It is an UPDATE (not a provision): it snapshots-reminds, rsyncs the new code,
rsyncs the engineering handbook to the box, installs the new deps, appends the
new secrets if absent, restarts the coordinator, and smoke-tests. It is
idempotent and fails loudly. What goes **live** after it: WS1 in-process
multi-model invokers (`bind_multi_invoker`, already wired), the WS5 handbook
`context_provider` injected into the planner prompt, and the WS2 Slack
`/new-task` command (AUTHZ-01 owner-allowlist gated). The P3 dispatch/build-verify
path stays **inert** (gated behind the `agent-apply` GitHub Environment approval).
**New venv deps** (the coordinator does not need them; only the optional HTTP
API does) — now pinned in the root `requirements.txt`:
| pip dep | Why |
|---|---|
| `fastapi==0.136.1` | the WS1 HTTP API app (`agent_team/api.py`) |
| `uvicorn==0.46.0` | ASGI server for `api.serve()` |
**New env vars** — append to `~/secrev.env` (mode 600, never committed):
- `SEA_HAVEN_HANDBOOK_DIR` — where `load_handbook_conventions()` reads the
engineering handbook (the script syncs it to `/home/adam/.sea-haven/engineering-handbook`
by default; this var must match). Fail-safe: if the dir is missing the
`context_provider` returns `""` and the planner runs without handbook context.
- `AGENT_TEAM_API_TOKEN` — bearer token for the HTTP API / `/delegate` hook
**only**. Not needed by the coordinator daemon itself. The HTTP API refuses to
start if this is unset/empty.
**The HTTP API is a separate, opt-in process** — it is **not** started by the
coordinator daemon. Run it explicitly (`api.serve()`, binds `127.0.0.1:8765`,
bearer auth) only if you want the `/delegate` Claude Code hook or the
`POST /tasks` / `GET /tasks/{thread_id}` / `POST /orchestrator/invoke` endpoints.
The `/docs` + `/openapi` routes are disabled and it binds loopback by design (do
not change to `0.0.0.0`). See the deploy script's step 6 for how to start it.
## 5. P1 live exit-criteria demo (§3.3.1)
Demonstrate all four once the service is live. Map each to the operator commands
(`run-team.py list / show / force-resume`, `systemctl`). Run the CLI from the
working dir so it hits the default ledger: `cd ~/orchestrator/agent-team`.
**(a) Crash-safe resume - kill mid-wait, restart, task resumes.**
Start a task, get it to a clarifier wait (`run-team.py list` shows an `open`
question), then:
```
sudo systemctl stop agent-team-coordinator.service
sudo systemctl start agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e # confirm the task resumes from the ledger/checkpoint
run-team.py show <question_id> # the question is still open, not lost
```
Pass: the task picks up the same waiting question after restart (the LangGraph
`SqliteSaver` checkpoint + the durable ledger survive the kill).
**(b) Duplicate Slack answer is a no-op.**
Answer a question in Slack, then answer the **same** question again.
```
run-team.py show <question_id> # status flipped to answered exactly once; answered_via is the first answer
```
Pass: the first answer wins (`rowcount == 1`); the duplicate hits the
`BEGIN IMMEDIATE` compare-and-set and is ignored (`rowcount == 0`) - no second
resume, no error.
**(c) Past-deadline answer is rejected + the task parks.**
Let a question's `deadline_at` pass with no answer, then answer late.
```
run-team.py show <question_id> # status == expired (auto-expired at deadline)
run-team.py list --parked # the now-parked task surfaces here
```
Pass: the expired question rejects the late answer and the task parks rather than
spins. To un-park it deliberately:
```
run-team.py force-resume <question_id> --confirm # reopens the expired question for re-delivery
```
**(d) Two concurrent tasks resume independently to the correct thread.**
Start two tasks concurrently, each reaching its own clarifier wait.
```
run-team.py list # two distinct open questions, distinct thread_id values
```
Restart the service (as in (a)); answer each in Slack.
Pass: each task resumes to its own `thread_id` / channel - no cross-talk, no
answer routed to the wrong task.
## 6. Rollback
```
# Stop + disable the service and remove the unit:
sudo systemctl disable --now agent-team-coordinator.service
sudo rm /etc/systemd/system/agent-team-coordinator.service
sudo systemctl daemon-reload
# Restore the VM from the pre-provision Hyper-V checkpoint (§1) to undo
# pip deps + any host changes in one step.
```
The ledger is **local state** under `agent-team/state/` (not in git). To reset
it without a full snapshot restore: back it up first, then wipe.
```
cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak} # back up
rm ~/orchestrator/agent-team/state/agent_team.sqlite* # wipe (then re-run init-db)
```
Note the secrets in `~/secrev.env` are NOT removed by rollback - leave them, or
strip the four agent-team keys if you are decommissioning entirely.
## 7. Security
- **Slack inbound listener (Socket Mode) is the auth + untrusted-input surface.**
It accepts inbound messages over a WebSocket and turns them into ledger
mutations (answering live clarifier questions). It **must pass
`/sh-security-review`** before this is enabled in production - that review is
mandatory for authentication/authorization and untrusted-input handling
changes, and this is both.
- **No IAM / OIDC is involved in P1.** The box runs on subscription OAuth
(`CLAUDE_CODE_OAUTH_TOKEN`) and Slack tokens only; there is no AWS role, no
OIDC trust relationship, no cloud permission surface in this deploy.
- `~/secrev.env` stays mode 600 and out of git; `state/` (ledger + audit log) is
gitignored and written 0600.