docs(agent-team): R720 deploy runbook + coordinator systemd unit
P1c deploy artifacts: DEPLOY-R720.md (snapshot-first, rsync, venv deps, ~/secrev.env tokens incl. AGENT_TEAM_SLACK_OWNER_IDS, init-db, systemd, the 4-criteria live demo, rollback) and the long-running coordinator service unit.
This commit is contained in:
parent
22edd4143a
commit
0f33fa6f24
2 changed files with 256 additions and 0 deletions
202
agent-team/DEPLOY-R720.md
Normal file
202
agent-team/DEPLOY-R720.md
Normal file
|
|
@ -0,0 +1,202 @@
|
||||||
|
# P1 - agent-team Plane-2 coordinator deployment (R720 VM)
|
||||||
|
|
||||||
|
Status: **P1 DEPLOY ARTIFACTS** - the systemd unit + this runbook. The coordinator
|
||||||
|
daemon (`run-team.py serve`) is built by a separate agent; this file is the
|
||||||
|
operator runbook for standing it up on the always-on R720 VM.
|
||||||
|
|
||||||
|
The agent-team coordinator is the long-running Plane-2 brain: it drives the
|
||||||
|
LangGraph pipeline, owns the durable `pending_questions` ledger, and runs the
|
||||||
|
Slack Socket Mode inbound listener that receives clarifier answers. It shares the
|
||||||
|
`sh-secrev` VM and the `~/secrev.env` secrets file with the Path B security sweep
|
||||||
|
(see `../security-review/DEPLOY-R720.md`), but it is a **service** (always-on),
|
||||||
|
not a timer-driven oneshot.
|
||||||
|
|
||||||
|
## Host
|
||||||
|
|
||||||
|
- **Hypervisor:** R720 at `10.10.60.40` (Windows Server 2022, Hyper-V role).
|
||||||
|
- **VM:** `sh-secrev`, always-on Ubuntu 24.04, Gen2, **4GB / 2 vCPU / 40GB**
|
||||||
|
dynamic vhdx at `10.10.60.120`.
|
||||||
|
- **Reach it:** `ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120` (key-only,
|
||||||
|
NOPASSWD sudo).
|
||||||
|
|
||||||
|
Operate on the VM, not from the Mac against the host by hand.
|
||||||
|
|
||||||
|
## 1. SNAPSHOT FIRST
|
||||||
|
|
||||||
|
**Standing rule: snapshot the VM before any provisioning change.** This box is a
|
||||||
|
4GB / 2 vCPU / 40GB VM at `10.10.60.120`. Take a Hyper-V checkpoint on the R720
|
||||||
|
host **before** you install pip deps, the unit, or touch `~/secrev.env`, so the
|
||||||
|
whole change is one-command reversible (see ROLLBACK). Do not skip this because
|
||||||
|
"it is only a pip install" - a bad dep set or a wedged service is exactly what
|
||||||
|
the snapshot exists to undo.
|
||||||
|
|
||||||
|
## 2. Prereqs on the VM
|
||||||
|
|
||||||
|
Already present from the secrev deploy:
|
||||||
|
|
||||||
|
- **Python 3** (3.12) and the `claude` CLI (Node) - the subscription-auth path.
|
||||||
|
- **Repo:** `~/orchestrator/` (rsync from the Mac, NOT a git clone). The
|
||||||
|
agent-team package lives at `~/orchestrator/agent-team/`.
|
||||||
|
|
||||||
|
New for the coordinator - a dedicated venv under `agent-team/.venv` (excluded
|
||||||
|
from rsync) with the coordinator/transport deps:
|
||||||
|
|
||||||
|
| pip dep | Why |
|
||||||
|
|---|---|
|
||||||
|
| `langgraph` | the coordinator pipeline graph |
|
||||||
|
| `langgraph-checkpoint-sqlite` | `SqliteSaver` checkpointer against the ledger DB |
|
||||||
|
| `claude-agent-sdk` | subscription-auth Claude invocation seam |
|
||||||
|
| `slack_sdk` | Slack Web API (post questions, `chat:write`) |
|
||||||
|
| `slack_bolt` | Socket Mode inbound listener (receive answers) |
|
||||||
|
|
||||||
|
The ledger DB defaults to `agent-team/state/agent_team.sqlite`; the audit log to
|
||||||
|
`agent-team/state/audit.log.jsonl`. Both live under `state/` (gitignored,
|
||||||
|
never committed).
|
||||||
|
|
||||||
|
## 3. Secrets - append to `~/secrev.env` (mode 600, never committed)
|
||||||
|
|
||||||
|
The coordinator reads its secrets from the same `~/secrev.env` the secrev sweep
|
||||||
|
uses. Append these (do not echo them into shell history files; lock the file
|
||||||
|
down after):
|
||||||
|
|
||||||
|
```
|
||||||
|
echo 'CLAUDE_CODE_OAUTH_TOKEN=...' >> ~/secrev.env # from `claude setup-token`
|
||||||
|
echo 'SLACK_BOT_TOKEN=xoxb-...' >> ~/secrev.env # bot token, chat:write
|
||||||
|
echo 'SLACK_APP_TOKEN=xapp-...' >> ~/secrev.env # app-level, connections:write (Socket Mode)
|
||||||
|
echo 'SLACK_CHANNEL_ID=C0XXXXXXX' >> ~/secrev.env # target clarifier channel
|
||||||
|
echo 'AGENT_TEAM_SLACK_OWNER_IDS=U0XXXXXXX' >> ~/secrev.env # authorized answerer(s), comma-separated
|
||||||
|
chmod 600 ~/secrev.env
|
||||||
|
```
|
||||||
|
|
||||||
|
- `CLAUDE_CODE_OAUTH_TOKEN` - subscription OAuth from `claude setup-token`. The
|
||||||
|
same token type the secrev sweep uses.
|
||||||
|
- `SLACK_BOT_TOKEN` (`xoxb-`) - bot token with `chat:write`; posts questions.
|
||||||
|
- `SLACK_APP_TOKEN` (`xapp-`) - app-level token with `connections:write`;
|
||||||
|
**required for Socket Mode** (opens the inbound WebSocket that hears answers).
|
||||||
|
- `SLACK_CHANNEL_ID` - the channel id the coordinator posts clarifiers to.
|
||||||
|
- `AGENT_TEAM_SLACK_OWNER_IDS` - comma-separated Slack **user ids** of the
|
||||||
|
authorized answerers (e.g. Adam's `U…` id). The inbound listener enforces this
|
||||||
|
as an owner allowlist (AUTHZ-01): only a sender in this set may answer/steer
|
||||||
|
the pipeline. **The listener fails closed** - if this is unset/empty it rejects
|
||||||
|
**every** answer (logs a warning naming `AGENT_TEAM_SLACK_OWNER_IDS`), so it
|
||||||
|
must be set for the human gate to function. Look up your user id via Slack
|
||||||
|
profile → "Copy member ID", or the `users.identity` / `auth.test` API.
|
||||||
|
|
||||||
|
**CRITICAL:** `ANTHROPIC_API_KEY` must **NOT** be set on this host. A raw API key
|
||||||
|
would silently win over the subscription OAuth and meter to API rates. The box
|
||||||
|
runs on subscription OAuth only.
|
||||||
|
|
||||||
|
## 4. Deploy steps
|
||||||
|
|
||||||
|
```
|
||||||
|
# 4a. From the Mac - rsync the repo (same pattern/excludes as secrev):
|
||||||
|
rsync -av --exclude .env --exclude .venv \
|
||||||
|
~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/
|
||||||
|
|
||||||
|
# 4b. On the VM - create + activate the agent-team venv and install deps:
|
||||||
|
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
|
||||||
|
cd ~/orchestrator/agent-team
|
||||||
|
python3 -m venv .venv
|
||||||
|
. .venv/bin/activate
|
||||||
|
pip install langgraph langgraph-checkpoint-sqlite claude-agent-sdk slack_sdk slack_bolt
|
||||||
|
|
||||||
|
# 4c. Initialize the durable ledger DB (idempotent; creates state/agent_team.sqlite):
|
||||||
|
python3 run-team.py init-db
|
||||||
|
|
||||||
|
# 4d. Install + start the service:
|
||||||
|
sudo cp systemd/agent-team-coordinator.service /etc/systemd/system/
|
||||||
|
sudo systemctl daemon-reload
|
||||||
|
sudo systemctl enable --now agent-team-coordinator.service
|
||||||
|
|
||||||
|
# 4e. Verify it is up:
|
||||||
|
systemctl status agent-team-coordinator.service
|
||||||
|
journalctl -u agent-team-coordinator.service -e -f
|
||||||
|
```
|
||||||
|
|
||||||
|
The unit runs `python3 run-team.py serve` from
|
||||||
|
`WorkingDirectory=/home/adam/orchestrator/agent-team` as `User=adam`, loading
|
||||||
|
secrets from `EnvironmentFile=/home/adam/secrev.env`. `Restart=on-failure` keeps
|
||||||
|
it up across transient faults; `journalctl -u` is the live log.
|
||||||
|
|
||||||
|
## 5. P1 live exit-criteria demo (§3.3.1)
|
||||||
|
|
||||||
|
Demonstrate all four once the service is live. Map each to the operator commands
|
||||||
|
(`run-team.py list / show / force-resume`, `systemctl`). Run the CLI from the
|
||||||
|
working dir so it hits the default ledger: `cd ~/orchestrator/agent-team`.
|
||||||
|
|
||||||
|
**(a) Crash-safe resume - kill mid-wait, restart, task resumes.**
|
||||||
|
Start a task, get it to a clarifier wait (`run-team.py list` shows an `open`
|
||||||
|
question), then:
|
||||||
|
```
|
||||||
|
sudo systemctl stop agent-team-coordinator.service
|
||||||
|
sudo systemctl start agent-team-coordinator.service
|
||||||
|
journalctl -u agent-team-coordinator.service -e # confirm the task resumes from the ledger/checkpoint
|
||||||
|
run-team.py show <question_id> # the question is still open, not lost
|
||||||
|
```
|
||||||
|
Pass: the task picks up the same waiting question after restart (the LangGraph
|
||||||
|
`SqliteSaver` checkpoint + the durable ledger survive the kill).
|
||||||
|
|
||||||
|
**(b) Duplicate Slack answer is a no-op.**
|
||||||
|
Answer a question in Slack, then answer the **same** question again.
|
||||||
|
```
|
||||||
|
run-team.py show <question_id> # status flipped to answered exactly once; answered_via is the first answer
|
||||||
|
```
|
||||||
|
Pass: the first answer wins (`rowcount == 1`); the duplicate hits the
|
||||||
|
`BEGIN IMMEDIATE` compare-and-set and is ignored (`rowcount == 0`) - no second
|
||||||
|
resume, no error.
|
||||||
|
|
||||||
|
**(c) Past-deadline answer is rejected + the task parks.**
|
||||||
|
Let a question's `deadline_at` pass with no answer, then answer late.
|
||||||
|
```
|
||||||
|
run-team.py show <question_id> # status == expired (auto-expired at deadline)
|
||||||
|
run-team.py list --parked # the now-parked task surfaces here
|
||||||
|
```
|
||||||
|
Pass: the expired question rejects the late answer and the task parks rather than
|
||||||
|
spins. To un-park it deliberately:
|
||||||
|
```
|
||||||
|
run-team.py force-resume <question_id> --confirm # reopens the expired question for re-delivery
|
||||||
|
```
|
||||||
|
|
||||||
|
**(d) Two concurrent tasks resume independently to the correct thread.**
|
||||||
|
Start two tasks concurrently, each reaching its own clarifier wait.
|
||||||
|
```
|
||||||
|
run-team.py list # two distinct open questions, distinct thread_id values
|
||||||
|
```
|
||||||
|
Restart the service (as in (a)); answer each in Slack.
|
||||||
|
Pass: each task resumes to its own `thread_id` / channel - no cross-talk, no
|
||||||
|
answer routed to the wrong task.
|
||||||
|
|
||||||
|
## 6. Rollback
|
||||||
|
|
||||||
|
```
|
||||||
|
# Stop + disable the service and remove the unit:
|
||||||
|
sudo systemctl disable --now agent-team-coordinator.service
|
||||||
|
sudo rm /etc/systemd/system/agent-team-coordinator.service
|
||||||
|
sudo systemctl daemon-reload
|
||||||
|
|
||||||
|
# Restore the VM from the pre-provision Hyper-V checkpoint (§1) to undo
|
||||||
|
# pip deps + any host changes in one step.
|
||||||
|
```
|
||||||
|
|
||||||
|
The ledger is **local state** under `agent-team/state/` (not in git). To reset
|
||||||
|
it without a full snapshot restore: back it up first, then wipe.
|
||||||
|
```
|
||||||
|
cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak} # back up
|
||||||
|
rm ~/orchestrator/agent-team/state/agent_team.sqlite* # wipe (then re-run init-db)
|
||||||
|
```
|
||||||
|
Note the secrets in `~/secrev.env` are NOT removed by rollback - leave them, or
|
||||||
|
strip the four agent-team keys if you are decommissioning entirely.
|
||||||
|
|
||||||
|
## 7. Security
|
||||||
|
|
||||||
|
- **Slack inbound listener (Socket Mode) is the auth + untrusted-input surface.**
|
||||||
|
It accepts inbound messages over a WebSocket and turns them into ledger
|
||||||
|
mutations (answering live clarifier questions). It **must pass
|
||||||
|
`/sh-security-review`** before this is enabled in production - that review is
|
||||||
|
mandatory for authentication/authorization and untrusted-input handling
|
||||||
|
changes, and this is both.
|
||||||
|
- **No IAM / OIDC is involved in P1.** The box runs on subscription OAuth
|
||||||
|
(`CLAUDE_CODE_OAUTH_TOKEN`) and Slack tokens only; there is no AWS role, no
|
||||||
|
OIDC trust relationship, no cloud permission surface in this deploy.
|
||||||
|
- `~/secrev.env` stays mode 600 and out of git; `state/` (ledger + audit log) is
|
||||||
|
gitignored and written 0600.
|
||||||
54
agent-team/systemd/agent-team-coordinator.service
Normal file
54
agent-team/systemd/agent-team-coordinator.service
Normal file
|
|
@ -0,0 +1,54 @@
|
||||||
|
# agent-team-coordinator.service - R720 Plane-2 coordinator daemon (sh-secrev VM, user adam).
|
||||||
|
#
|
||||||
|
# Long-running coordinator for the agent-team SDLC pipeline. Unlike the
|
||||||
|
# sea-haven-secrev sweep (a oneshot driven by a timer), this is an always-on
|
||||||
|
# service: it serves the LangGraph coordinator, the durable pending_questions
|
||||||
|
# ledger, and the Slack Socket Mode inbound listener that answers clarifier
|
||||||
|
# questions. ExecStart runs the `serve` subcommand of the operator CLI.
|
||||||
|
#
|
||||||
|
# Install (on the VM, as root):
|
||||||
|
# sudo cp agent-team-coordinator.service /etc/systemd/system/
|
||||||
|
# sudo systemctl daemon-reload
|
||||||
|
# sudo systemctl enable --now agent-team-coordinator.service
|
||||||
|
# systemctl status agent-team-coordinator.service
|
||||||
|
# journalctl -u agent-team-coordinator.service -e -f
|
||||||
|
#
|
||||||
|
# Secrets come from the EnvironmentFile (leading '-' = optional, no failure if
|
||||||
|
# absent), ~/secrev.env (mode 600, NOT in git):
|
||||||
|
# CLAUDE_CODE_OAUTH_TOKEN -> subscription OAuth (from `claude setup-token`).
|
||||||
|
# A raw ANTHROPIC_API_KEY must NOT be set on this
|
||||||
|
# box; it would silently win and meter to API
|
||||||
|
# rates. The billing seam pops it defensively.
|
||||||
|
# SLACK_BOT_TOKEN -> xoxb- bot token (chat:write) - posts questions.
|
||||||
|
# SLACK_APP_TOKEN -> xapp- app-level token (connections:write) -
|
||||||
|
# REQUIRED for Socket Mode; opens the inbound
|
||||||
|
# WebSocket that receives answers. Without it the
|
||||||
|
# coordinator can post but never hear replies.
|
||||||
|
# SLACK_CHANNEL_ID -> target channel for clarifier questions.
|
||||||
|
|
||||||
|
[Unit]
|
||||||
|
Description=Sea Haven agent-team Plane-2 coordinator daemon
|
||||||
|
After=network-online.target
|
||||||
|
Wants=network-online.target
|
||||||
|
|
||||||
|
[Service]
|
||||||
|
Type=simple
|
||||||
|
User=adam
|
||||||
|
WorkingDirectory=/home/adam/orchestrator/agent-team
|
||||||
|
EnvironmentFile=-/home/adam/secrev.env
|
||||||
|
ExecStart=/usr/bin/env python3 run-team.py serve
|
||||||
|
Restart=on-failure
|
||||||
|
RestartSec=5
|
||||||
|
# Hardening - matches the level the sea-haven-secrev unit relies on, scoped for a
|
||||||
|
# long-running daemon that must READ ~/secrev.env and WRITE the local ledger.
|
||||||
|
NoNewPrivileges=true
|
||||||
|
ProtectSystem=full
|
||||||
|
# ProtectHome cannot be `true`: the daemon reads /home/adam/secrev.env and writes
|
||||||
|
# the ledger under the working dir. read-only home + an explicit RW carve-out for
|
||||||
|
# the state/ dir keeps the rest of $HOME unreadable/unwritable to the service.
|
||||||
|
ProtectHome=read-only
|
||||||
|
ReadWritePaths=/home/adam/orchestrator/agent-team/state
|
||||||
|
Nice=10
|
||||||
|
|
||||||
|
[Install]
|
||||||
|
WantedBy=multi-user.target
|
||||||
Reference in a new issue