- PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5 intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked FIXED so the demo can use the live Slack answer path. - P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the Slack and operator-CLI answer paths documented for all four exit criteria. - DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the operator-CLI divergence kept as provisioning notes. - OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks, failed HITL resumes, budget exhaustion, transport outages, and COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6).
18 KiB
PROVISIONING RUNBOOK — R720 agent-team Plane-2 coordinator
This is the ordered command sequence for the operator-present provisioning
session that stands up the always-on agent-team-coordinator daemon on the
sh-secrev R720 VM. Derived from agent-team/DEPLOY-R720.md,
security-review/DEPLOY-R720.md, and docs/r720-agent-team-design.md
(§3.3.1 / §7 / §7.1).
State of this runbook
The build is merged and final: all six Plane-1 checkers (
aws-posture,compliance-drift,confluence-doc,dependency-cve,doc-drift,plan-groomer), the Tier-3 dep-bump fixer (run-team.py fix --dry-run), and the Plane-1→Plane-2 P5 cross-plane loop (run-team.py intake-checker) exist in the tree. The three deploy-correctness bugs the provisioning-prep audit found are FIXED (see DEPLOY-AUDIT.md):
- D-1 RESOLVED —
Coordinator.serve()now starts the inbound SlackSlackListenerconcurrently with the tick/drain loop when Slack is the live transport ANDSLACK_APP_TOKENis set. The demo can use the live Slack answer path (P1-DEMO-SCRIPT.md), or the operator-CLIanswerpath.- D-2 RESOLVED — the systemd unit loads
~/orchestrator/.env(for the P2 GPT-4.1 review loop's provider key) in addition to~/secrev.env.- D-7 RESOLVED — the unit's
ExecStartpoints at the agent-team venv interpreter, not the systempython3.Still open as provisioning notes (not blockers): D-4/D-5 (agent-team runtime deps are installed ad-hoc into the venv and are not pinned in
requirements.txt), and the operator-CLI divergence (run-team.pyvsagent_team/operator_cli.pyhave different verb names — userun-team.py).
Scope and ground rules
Hard rules (carried from the design and global instructions):
- Snapshot before any stateful change (design §7;
feedback_ec2_replacement_snapshot). - Every stateful step has an exercised rollback.
- Secrets use placeholder names only; real values are entered by the operator at the box and never echoed into shell history.
- The box is read-only / subscription-OAuth only. No
ANTHROPIC_API_KEYon this host (it would silently win over OAuth and meter to API rates —billing.claude_invokepops it defensively, but it must not be present). - No IAM / OIDC is involved in the coordinator deploy (P1/P2). IAM enters
only at the P3-live flip (see the dedicated section), which is gated on the
mandatory GPT-4.1 cross-review +
/sh-security-review.
Legend per step:
- 🧑 OPERATOR-REQUIRED — needs the human (snapshot, secrets, the live human gate, go/no-go). Cannot be automated.
- 🤖 MECHANICAL — deterministic; an operator runs it but it needs no judgment.
Host facts (from both DEPLOY-R720.md files)
- Hypervisor: R720 at
10.10.60.40(Windows Server 2022, Hyper-V). - VM:
sh-secrev, Ubuntu 24.04, 4GB / 2 vCPU / 40GB dynamic vhdx,10.10.60.120. - Reach:
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120(key-only, NOPASSWD sudo). - Repo on the box:
~/orchestrator/(rsync from the Mac, NOT a git clone). The package lives at~/orchestrator/agent-team/. - Shares
~/secrev.env(mode 600) with the secrev sweep, and~/orchestrator/.env(mode 600) for non-Claude provider keys (same files the secrev unit loads).
STEP 1 — Snapshot the VM 🧑 OPERATOR-REQUIRED
On the R720 host (Hyper-V), before anything else. This is the one-command undo for every change below.
# On the R720 Windows host (PowerShell, as admin):
Checkpoint-VM -Name sh-secrev -SnapshotName "pre-agent-team-coordinator-$(Get-Date -Format yyyyMMdd-HHmm)"
Get-VMSnapshot -VMName sh-secrev # confirm the checkpoint exists
ROLLBACK (whole session):
Restore-VMSnapshot -VMName sh-secrev -Name "<the checkpoint name above>" -Confirm:$false
Start-VM -Name sh-secrev
Do not proceed until the checkpoint is confirmed present.
STEP 2 — Rsync the repo to the box 🤖 MECHANICAL (verify manifest 🧑)
From the Mac. Same pattern/excludes as the secrev deploy. The team scans
the same ~/repo-mirrors corpus secrev already maintains — this rsync ships
code, not mirrors.
# From the Mac (sync the canonical repo, not a worktree):
rsync -av --exclude .env --exclude .venv --exclude .git --exclude .claude \
~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/
Surface that must be on the box:
Path (under ~/orchestrator/) |
Why it must ship |
|---|---|
agent-team/run-team.py |
the operator entry CLI |
agent-team/agent_team/** |
the package (coordinator, graph, ledger, transports, slack_listener, fixer) |
agent-team/systemd/agent-team-coordinator.service |
the daemon unit |
run.py + the orchestrator package |
the GPT-4.1 cross_reviewer the P2 review loop shells |
requirements.txt |
pin reference for langgraph / checkpoint-sqlite |
security-review/lib/** |
shared sweep substrate (Phase 0) |
security-review/checkers/** |
the six Plane-1 checkers + fixtures |
ROLLBACK: rsync is additive; restore the Step-1 snapshot to revert code state.
STEP 3 — Write secrets 🧑 OPERATOR-REQUIRED
On the VM. The coordinator unit loads two EnvironmentFiles (both
optional via the leading -): ~/secrev.env (agent-team runtime keys) and
~/orchestrator/.env (the non-Claude provider key for the P2 review loop).
Placeholder names only — the operator pastes real values. Do not echo real
tokens into shell history (use an editor or read -s).
~/secrev.env (mode 600) — append the agent-team keys:
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
# Edit ~/secrev.env (mode 600) and add — values are placeholders:
# CLAUDE_CODE_OAUTH_TOKEN=<CLAUDE_CODE_OAUTH_TOKEN> # from `claude setup-token`
# SLACK_BOT_TOKEN=<SLACK_BOT_TOKEN> # xoxb-..., chat:write (posts questions)
# SLACK_APP_TOKEN=<SLACK_APP_TOKEN> # xapp-..., connections:write (Socket Mode inbound)
# SLACK_CHANNEL_ID=<SLACK_CHANNEL_ID> # C0..., target clarifier channel
# AGENT_TEAM_SLACK_OWNER_IDS=<SLACK_OWNER_USER_IDS> # comma-separated U... ids of authorized answerers
chmod 600 ~/secrev.env
~/orchestrator/.env (mode 600) — the non-Claude provider key for the GPT-4.1
review loop (same file secrev uses; if it already exists with the key, leave it):
# OPENAI_API_KEY=<OPENAI_API_KEY> # (or the provider key cross_reviewer/GPT-4.1 needs)
chmod 600 ~/orchestrator/.env
Contract notes (verified against the code — see DEPLOY-AUDIT.md):
SLACK_CHANNEL_IDis correct —run-team.py _build_transportreads exactlyos.environ.get("SLACK_CHANNEL_ID"). Do not useSLACK_CHANNEL.SLACK_APP_TOKENis now read by the daemon:Coordinator.serve()starts the inboundSlackListenerwhen the transport is live Slack ANDSLACK_APP_TOKENis set. Without it the daemon still runs (posts + expires) but never hears Slack replies — Slack stays optional by design.AGENT_TEAM_SLACK_OWNER_IDSfails closed (AUTHZ-01): if unset/empty the listener rejects every answer. It must be set for the live human gate.- CRITICAL: confirm
ANTHROPIC_API_KEYis NOT present:grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # must print 0 for both
ROLLBACK: strip exactly the appended keys (keep secrev keys intact), or restore the Step-1 snapshot. Do NOT blindly truncate — secrev keys live here too.
STEP 4 — Create the venv + install pip deps 🤖 MECHANICAL
On the VM. A dedicated venv under agent-team/.venv (excluded from rsync).
The systemd unit's ExecStart points at this venv's interpreter (D-7 fixed),
so the deps MUST land here.
cd ~/orchestrator/agent-team
python3 -m venv .venv
. .venv/bin/activate
pip install langgraph==1.1.10 langgraph-checkpoint-sqlite==3.1.0 \
claude-agent-sdk slack_sdk slack_bolt requests
Pin note (D-5):
requirements.txtpinslanggraph==1.1.10/langgraph-checkpoint-sqlite==3.1.0; match those exactly here. The other runtime deps (claude-agent-sdk,slack_sdk,slack_bolt,requests) are not yet inrequirements.txt(D-5 open) — installed ad-hoc here.requestsis required by the GitHub transport/intake (D-4).anthropicis not installed (only the opt-inapibilling mode needs it).slack_boltis now actually exercised (D-1 fixed: the daemon starts the Socket Mode listener).
Verify the imports resolve (using the venv interpreter the unit will use):
.venv/bin/python -c "import langgraph, langgraph.checkpoint.sqlite, slack_sdk, slack_bolt, requests; print('deps ok')"
.venv/bin/python -c "import claude_agent_sdk; print('agent-sdk ok')"
ROLLBACK: deactivate 2>/dev/null; rm -rf ~/orchestrator/agent-team/.venv
STEP 5 — Initialize the durable ledger DB 🤖 MECHANICAL
On the VM, venv active. Idempotent; creates state/agent_team.sqlite with
the pending_questions + budget_ledger + schema_meta tables (LangGraph
SqliteSaver creates its own tables in the same file on first run).
cd ~/orchestrator/agent-team
. .venv/bin/activate
python3 run-team.py init-db
# Expect: "initialized ledger DB at .../state/agent_team.sqlite"
ls -l state/ # agent_team.sqlite present; state/ is gitignored
ROLLBACK (reset the ledger only):
cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak}
rm ~/orchestrator/agent-team/state/agent_team.sqlite*
# re-run `python3 run-team.py init-db` to recreate empty tables.
STEP 6 — Install + start the systemd unit 🤖 MECHANICAL (go/no-go 🧑)
On the VM, as root. Installs the long-running coordinator daemon. Keep the
hardening as-shipped: NoNewPrivileges, ProtectSystem=full,
ProtectHome=read-only + ReadWritePaths=.../agent-team/state (locked decision).
sudo cp ~/orchestrator/agent-team/systemd/agent-team-coordinator.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now agent-team-coordinator.service
systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e -f
# expect: "agent-team coordinator starting"
# and (if Slack + SLACK_APP_TOKEN provisioned):
# "inbound Slack listener started (Socket Mode, background thread)"
# or otherwise:
# "inbound Slack listener not started (... or SLACK_APP_TOKEN is unset) ..."
The unit (post-fix) runs
/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve from
WorkingDirectory=/home/adam/orchestrator/agent-team as User=adam, loading
both EnvironmentFile=-/home/adam/secrev.env and
EnvironmentFile=-/home/adam/orchestrator/.env, Restart=on-failure.
D-7 fixed:
ExecStartnow resolves the venv interpreter, so the Step-4 deps are on the path. D-1 fixed:servestarts the inboundSlackListenerwhen Slack + app token are provisioned. Go/no-go: confirm the journal shows the listener line you expect for your transport choice before Step 8.
ROLLBACK (exercised — tear the unit down cleanly):
sudo systemctl disable --now agent-team-coordinator.service
sudo rm /etc/systemd/system/agent-team-coordinator.service
sudo systemctl daemon-reload
systemctl status agent-team-coordinator.service # should report "could not be found"
The ledger under state/ is untouched. Step-1 snapshot is the whole-host fallback.
STEP 7 — Verify the Slack inbound listener 🧑 OPERATOR-REQUIRED
The clarifier gate is two halves: outbound (post the question — done by
serve/start via the Slack poster) and inbound (receive Adam's answer —
the SlackListener over Socket Mode).
D-1 RESOLVED.
Coordinator.serve()now constructs and startsSlackListeneron a background daemon thread, concurrently with the tick/drain loop, when the live transport is aSlackTransportANDSLACK_APP_TOKENis set. It shares the coordinator's own transport, ledger, and resume queue, and stops cleanly on shutdown. The AUTHZ-01 owner allowlist + the open-status compare-and-set are unchanged — the listener still fails closed on an emptyAGENT_TEAM_SLACK_OWNER_IDS.
Confirm the live inbound path is up:
# In the journal (Step 6) expect the "inbound Slack listener started" line.
# Verify the two gating vars are present in the unit's environment:
grep -c SLACK_APP_TOKEN ~/secrev.env # 1
grep -c AGENT_TEAM_SLACK_OWNER_IDS ~/secrev.env # 1 (else the gate rejects all answers)
If you are not using Slack as the transport (or are deliberately running
without the app token), the daemon runs the maintenance loop only and the demo
uses the operator-CLI answer path — both are valid (P1-DEMO-SCRIPT.md).
ROLLBACK: none needed — verification only, no host state change.
STEP 8 — Live P1 four-criteria acceptance demo 🧑 OPERATOR-REQUIRED
Run P1-DEMO-SCRIPT.md in full. All four §3.3.1 exit criteria must pass:
(a) crash-safe resume, (b) duplicate-answer no-op, (c) post-deadline rejection +
park, (d) two concurrent tasks resume independently. With D-1 fixed you may
exercise the live Slack answer path for (b)/(d); the operator-CLI answer
path remains available and exercises the identical compare-and-set. Do not
accept P1 until all four pass.
ROLLBACK: the demo writes only ledger rows under state/; reset via the
Step-5 ledger rollback, or restore the Step-1 snapshot, then re-run.
STEP 9 — Plane-1 checker live dry-runs 🧑 OPERATOR-REQUIRED
The six checkers live under security-review/checkers/:
aws-posture.sh, compliance-drift.sh, confluence-doc.sh,
dependency-cve.sh, doc-drift.sh, plan-groomer.sh. All run report +
ALARM-only (design D3): a clean run posts nothing and lands a mode-600 report;
no auto-Jira/Notion writes.
- Dry-run each checker against
~/repo-mirrorsin report-only mode; confirm a clean run posts nothing and writes a mode-600 report. Confirm the exact invocation against the checker scripts and the sharedsecurity-review/lib/substrate on the box. - Confirm the canary suite runs first (a planted-fault miss is a COMPLACENCY ALARM and that role is skipped — design §6.4), and the coverage rotation pointer advances (a slipped role is a COVERAGE ALARM, deferred-not-dropped).
ROLLBACK: checkers are read-only over the mirror corpus; a dry-run produces only a report file. Remove the report dir to revert; no host state change.
STEP 10 — P5 cross-plane loop dry-run 🧑 OPERATOR-REQUIRED
The Plane-1→Plane-2 loop turns confirmed checker findings into pipeline tasks:
cd ~/orchestrator/agent-team && . .venv/bin/activate
# Read one or more checker report JSONs and start one task per confirmed
# at/above-threshold finding (default threshold: high). --dry-run posts nowhere.
python3 run-team.py intake-checker --report <path-to-report.json> --threshold high --dry-run
The Tier-3 dep-bump fixer is dry-run only on the box (it holds no write token, D2):
python3 run-team.py fix --report security-review/<...>/dependency-cve.json \
--finding-id <ID> --task-id <task> --dry-run
# prints the fix spec + minimal bump patch + the CI workflow_dispatch inputs;
# dispatches NOTHING. Live dispatch is the P3-live flip below.
ROLLBACK: both are read-only / dry-run (no dispatch, no apply); intake-checker
de-dup is in-memory per process. No host state to revert beyond ledger rows from
a non-dry-run intake (Step-5 ledger rollback).
STEP 11 — Wire the schedule 🧑 OPERATOR-REQUIRED
The coordinator daemon (Step 6) is always-on, not timer-driven. The
per-checker timers (and the design's "shared timer with secrev", §8) attach
here. The existing sea-haven-secrev.timer (OnCalendar 02:00, Persistent) is
untouched. Add a checker timer only after that checker is dry-run-validated
(Step 9).
ROLLBACK: each timer gets its own systemctl disable --now <unit>.timer + rm.
The P3-live flip (deferred; IAM + GitHub App gated) 🧑 OPERATOR-REQUIRED
P3 (the build→verify apply-and-open-draft-PR loop) is opt-in and inert in
this deploy: run-team.py / serve pass build_verify_wiring=None, so no P3
subgraph is assembled. Flipping it live is a separate, gated provisioning
session, not part of the coordinator deploy:
- Mandatory reviews first. The CI trust-boundary + OIDC IAM change is a
breaking IAM change → GPT-4.1 cross-review (global instructions) AND
/sh-security-reviewon the apply/verify surface. Do not flip without both. - Provision the GitHub App for the trusted apply path (the App that opens
the draft PR), and the
agent-applyGitHub Actions environment that holds the apply path's scoped permissions. - Bind the live build→verify wiring via
agent_team.coordinator.gated_build_verify_wiring(...)(the read-only CI result fetcher + the real diff builder) — the seam a leaf calls after the gate clears. The CI fetcher is read-only and fails closed (missing token / 404 / auth failure →None→ the gate BLOCKs and the task parks). - Set the apply env vars the live path reads (the read-only CI-result token
and the dispatch target), then re-run the fixer without
--dry-runonly once the dispatcher is bound.
Until every step above is done, the box dispatches/applies nothing.
Post-session definition-of-done (design §7, global instructions)
- All four P1 criteria demonstrated live (Step 8).
project_r720_agent_teammemory created/updated.- Confluence "AWS Architecture Map" / IT host inventory updated to show
sh-secrevnow also hosts the always-on agent-team coordinator daemon. /sh-security-reviewrun on the Slack inbound listener surface (auth + untrusted-input; mandatory) — flag outstanding if not run.- OPERATOR-RUNBOOK.md reviewed by whoever holds the pager.
- Snapshot retained until the daemon runs clean for one full cycle, then pruned.