This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/docs/provisioning/PROVISIONING-RUNBOOK.md
Adam Moussa 48bf4a8290 fix(agent-team): repair Slack listener block_actions matcher; add dedicated Slack app
The Socket Mode inbound listener crashed at registration time on first live
run: `@app.action({})` raised `BoltError: action ({}) must be any of str,
Pattern, and dict` under slack_bolt 1.28.0, killing the listener thread (the
whole inbound answer path — message/app_mention/block_actions — went down,
caught only by the coordinator's respawn watchdog). serve() is marked
`# pragma: no cover - live socket`, so this was never exercised until the R720
bring-up. Replace the unsupported empty-dict matcher with a catch-all
`re.compile(r".*")` action_id regex; handle_event still does the real filtering
+ AUTHZ-01 owner-allowlist gate, so over-matching is safe.

Verified on sh-secrev: listener connects (live Socket Mode WebSocket), outbound
chat.postMessage works, 0 errors. /sh-security-review PASS (no confirmed
critical/high; matcher change introduces no new findings).

Also adds the dedicated Slack app (manifest + README) backing the clarifier
gate — "Sea Haven agent-team" (A0BC7AT8NUD), workspace-scoped install to avoid
the Enterprise-Grid `scope_not_allowed_on_enterprise` org-install trap — and
patches the provisioning runbook's stale langgraph pin (1.1.10 -> 1.2.5).
2026-06-22 13:12:04 -04:00

18 KiB

PROVISIONING RUNBOOK — R720 agent-team Plane-2 coordinator

This is the ordered command sequence for the operator-present provisioning session that stands up the always-on agent-team-coordinator daemon on the sh-secrev R720 VM. Derived from agent-team/DEPLOY-R720.md, security-review/DEPLOY-R720.md, and docs/r720-agent-team-design.md (§3.3.1 / §7 / §7.1).

State of this runbook

The build is merged and final: all six Plane-1 checkers (aws-posture, compliance-drift, confluence-doc, dependency-cve, doc-drift, plan-groomer), the Tier-3 dep-bump fixer (run-team.py fix --dry-run), and the Plane-1→Plane-2 P5 cross-plane loop (run-team.py intake-checker) exist in the tree. The three deploy-correctness bugs the provisioning-prep audit found are FIXED (see DEPLOY-AUDIT.md):

  • D-1 RESOLVED — Coordinator.serve() now starts the inbound Slack SlackListener concurrently with the tick/drain loop when Slack is the live transport AND SLACK_APP_TOKEN is set. The demo can use the live Slack answer path (P1-DEMO-SCRIPT.md), or the operator-CLI answer path.
  • D-2 RESOLVED — the systemd unit loads ~/orchestrator/.env (for the P2 GPT-4.1 review loop's provider key) in addition to ~/secrev.env.
  • D-7 RESOLVED — the unit's ExecStart points at the agent-team venv interpreter, not the system python3.

Still open as provisioning notes (not blockers): D-4/D-5 (agent-team runtime deps are installed ad-hoc into the venv and are not pinned in requirements.txt), and the operator-CLI divergence (run-team.py vs agent_team/operator_cli.py have different verb names — use run-team.py).


Scope and ground rules

Hard rules (carried from the design and global instructions):

  • Snapshot before any stateful change (design §7; feedback_ec2_replacement_snapshot).
  • Every stateful step has an exercised rollback.
  • Secrets use placeholder names only; real values are entered by the operator at the box and never echoed into shell history.
  • The box is read-only / subscription-OAuth only. No ANTHROPIC_API_KEY on this host (it would silently win over OAuth and meter to API rates — billing.claude_invoke pops it defensively, but it must not be present).
  • No IAM / OIDC is involved in the coordinator deploy (P1/P2). IAM enters only at the P3-live flip (see the dedicated section), which is gated on the mandatory GPT-4.1 cross-review + /sh-security-review.

Legend per step:

  • 🧑 OPERATOR-REQUIRED — needs the human (snapshot, secrets, the live human gate, go/no-go). Cannot be automated.
  • 🤖 MECHANICAL — deterministic; an operator runs it but it needs no judgment.

Host facts (from both DEPLOY-R720.md files)

  • Hypervisor: R720 at 10.10.60.40 (Windows Server 2022, Hyper-V).
  • VM: sh-secrev, Ubuntu 24.04, 4GB / 2 vCPU / 40GB dynamic vhdx, 10.10.60.120.
  • Reach: ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120 (key-only, NOPASSWD sudo).
  • Repo on the box: ~/orchestrator/ (rsync from the Mac, NOT a git clone). The package lives at ~/orchestrator/agent-team/.
  • Shares ~/secrev.env (mode 600) with the secrev sweep, and ~/orchestrator/.env (mode 600) for non-Claude provider keys (same files the secrev unit loads).

STEP 1 — Snapshot the VM 🧑 OPERATOR-REQUIRED

On the R720 host (Hyper-V), before anything else. This is the one-command undo for every change below.

# On the R720 Windows host (PowerShell, as admin):
Checkpoint-VM -Name sh-secrev -SnapshotName "pre-agent-team-coordinator-$(Get-Date -Format yyyyMMdd-HHmm)"
Get-VMSnapshot -VMName sh-secrev    # confirm the checkpoint exists

ROLLBACK (whole session):

Restore-VMSnapshot -VMName sh-secrev -Name "<the checkpoint name above>" -Confirm:$false
Start-VM -Name sh-secrev

Do not proceed until the checkpoint is confirmed present.


STEP 2 — Rsync the repo to the box 🤖 MECHANICAL (verify manifest 🧑)

From the Mac. Same pattern/excludes as the secrev deploy. The team scans the same ~/repo-mirrors corpus secrev already maintains — this rsync ships code, not mirrors.

# From the Mac (sync the canonical repo, not a worktree):
rsync -av --exclude .env --exclude .venv --exclude .git --exclude .claude \
  ~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/

Surface that must be on the box:

Path (under ~/orchestrator/) Why it must ship
agent-team/run-team.py the operator entry CLI
agent-team/agent_team/** the package (coordinator, graph, ledger, transports, slack_listener, fixer)
agent-team/systemd/agent-team-coordinator.service the daemon unit
run.py + the orchestrator package the GPT-4.1 cross_reviewer the P2 review loop shells
requirements.txt pin reference for langgraph / checkpoint-sqlite
security-review/lib/** shared sweep substrate (Phase 0)
security-review/checkers/** the six Plane-1 checkers + fixtures

ROLLBACK: rsync is additive; restore the Step-1 snapshot to revert code state.


STEP 3 — Write secrets 🧑 OPERATOR-REQUIRED

On the VM. The coordinator unit loads two EnvironmentFiles (both optional via the leading -): ~/secrev.env (agent-team runtime keys) and ~/orchestrator/.env (the non-Claude provider key for the P2 review loop). Placeholder names only — the operator pastes real values. Do not echo real tokens into shell history (use an editor or read -s).

~/secrev.env (mode 600) — append the agent-team keys:

ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
# Edit ~/secrev.env (mode 600) and add — values are placeholders:
#   CLAUDE_CODE_OAUTH_TOKEN=<CLAUDE_CODE_OAUTH_TOKEN>   # from `claude setup-token`
#   SLACK_BOT_TOKEN=<SLACK_BOT_TOKEN>                   # xoxb-..., chat:write (posts questions)
#   SLACK_APP_TOKEN=<SLACK_APP_TOKEN>                   # xapp-..., connections:write (Socket Mode inbound)
#   SLACK_CHANNEL_ID=<SLACK_CHANNEL_ID>                # C0..., target clarifier channel
#   AGENT_TEAM_SLACK_OWNER_IDS=<SLACK_OWNER_USER_IDS>  # comma-separated U... ids of authorized answerers
chmod 600 ~/secrev.env

~/orchestrator/.env (mode 600) — the non-Claude provider key for the GPT-4.1 review loop (same file secrev uses; if it already exists with the key, leave it):

#   OPENAI_API_KEY=<OPENAI_API_KEY>   # (or the provider key cross_reviewer/GPT-4.1 needs)
chmod 600 ~/orchestrator/.env

Contract notes (verified against the code — see DEPLOY-AUDIT.md):

  • SLACK_CHANNEL_ID is correct — run-team.py _build_transport reads exactly os.environ.get("SLACK_CHANNEL_ID"). Do not use SLACK_CHANNEL.
  • SLACK_APP_TOKEN is now read by the daemon: Coordinator.serve() starts the inbound SlackListener when the transport is live Slack AND SLACK_APP_TOKEN is set. Without it the daemon still runs (posts + expires) but never hears Slack replies — Slack stays optional by design.
  • AGENT_TEAM_SLACK_OWNER_IDS fails closed (AUTHZ-01): if unset/empty the listener rejects every answer. It must be set for the live human gate.
  • CRITICAL: confirm ANTHROPIC_API_KEY is NOT present:
    grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env   # must print 0 for both
    

ROLLBACK: strip exactly the appended keys (keep secrev keys intact), or restore the Step-1 snapshot. Do NOT blindly truncate — secrev keys live here too.


STEP 4 — Create the venv + install pip deps 🤖 MECHANICAL

On the VM. A dedicated venv under agent-team/.venv (excluded from rsync). The systemd unit's ExecStart points at this venv's interpreter (D-7 fixed), so the deps MUST land here.

cd ~/orchestrator/agent-team
python3 -m venv .venv
. .venv/bin/activate
pip install langgraph==1.2.5 langgraph-checkpoint-sqlite==3.1.0 \
  claude-agent-sdk slack_sdk slack_bolt requests

Pin note (D-5): requirements.txt pins langgraph==1.2.5 / langgraph-checkpoint-sqlite==3.1.0; match those exactly here. The other runtime deps (claude-agent-sdk, slack_sdk, slack_bolt, requests) are not yet in requirements.txt (D-5 open) — installed ad-hoc here. requests is required by the GitHub transport/intake (D-4). anthropic is not installed (only the opt-in api billing mode needs it). slack_bolt is now actually exercised (D-1 fixed: the daemon starts the Socket Mode listener).

Verify the imports resolve (using the venv interpreter the unit will use):

.venv/bin/python -c "import langgraph, langgraph.checkpoint.sqlite, slack_sdk, slack_bolt, requests; print('deps ok')"
.venv/bin/python -c "import claude_agent_sdk; print('agent-sdk ok')"

ROLLBACK: deactivate 2>/dev/null; rm -rf ~/orchestrator/agent-team/.venv


STEP 5 — Initialize the durable ledger DB 🤖 MECHANICAL

On the VM, venv active. Idempotent; creates state/agent_team.sqlite with the pending_questions + budget_ledger + schema_meta tables (LangGraph SqliteSaver creates its own tables in the same file on first run).

cd ~/orchestrator/agent-team
. .venv/bin/activate
python3 run-team.py init-db
# Expect: "initialized ledger DB at .../state/agent_team.sqlite"
ls -l state/                      # agent_team.sqlite present; state/ is gitignored

ROLLBACK (reset the ledger only):

cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak}
rm ~/orchestrator/agent-team/state/agent_team.sqlite*
# re-run `python3 run-team.py init-db` to recreate empty tables.

STEP 6 — Install + start the systemd unit 🤖 MECHANICAL (go/no-go 🧑)

On the VM, as root. Installs the long-running coordinator daemon. Keep the hardening as-shipped: NoNewPrivileges, ProtectSystem=full, ProtectHome=read-only + ReadWritePaths=.../agent-team/state (locked decision).

sudo cp ~/orchestrator/agent-team/systemd/agent-team-coordinator.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now agent-team-coordinator.service
systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e -f
# expect: "agent-team coordinator starting"
# and (if Slack + SLACK_APP_TOKEN provisioned):
#   "inbound Slack listener started (Socket Mode, background thread)"
# or otherwise:
#   "inbound Slack listener not started (... or SLACK_APP_TOKEN is unset) ..."

The unit (post-fix) runs /home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve from WorkingDirectory=/home/adam/orchestrator/agent-team as User=adam, loading both EnvironmentFile=-/home/adam/secrev.env and EnvironmentFile=-/home/adam/orchestrator/.env, Restart=on-failure.

D-7 fixed: ExecStart now resolves the venv interpreter, so the Step-4 deps are on the path. D-1 fixed: serve starts the inbound SlackListener when Slack + app token are provisioned. Go/no-go: confirm the journal shows the listener line you expect for your transport choice before Step 8.

ROLLBACK (exercised — tear the unit down cleanly):

sudo systemctl disable --now agent-team-coordinator.service
sudo rm /etc/systemd/system/agent-team-coordinator.service
sudo systemctl daemon-reload
systemctl status agent-team-coordinator.service   # should report "could not be found"

The ledger under state/ is untouched. Step-1 snapshot is the whole-host fallback.


STEP 7 — Verify the Slack inbound listener 🧑 OPERATOR-REQUIRED

The clarifier gate is two halves: outbound (post the question — done by serve/start via the Slack poster) and inbound (receive Adam's answer — the SlackListener over Socket Mode).

D-1 RESOLVED. Coordinator.serve() now constructs and starts SlackListener on a background daemon thread, concurrently with the tick/drain loop, when the live transport is a SlackTransport AND SLACK_APP_TOKEN is set. It shares the coordinator's own transport, ledger, and resume queue, and stops cleanly on shutdown. The AUTHZ-01 owner allowlist + the open-status compare-and-set are unchanged — the listener still fails closed on an empty AGENT_TEAM_SLACK_OWNER_IDS.

Confirm the live inbound path is up:

# In the journal (Step 6) expect the "inbound Slack listener started" line.
# Verify the two gating vars are present in the unit's environment:
grep -c SLACK_APP_TOKEN ~/secrev.env            # 1
grep -c AGENT_TEAM_SLACK_OWNER_IDS ~/secrev.env # 1 (else the gate rejects all answers)

If you are not using Slack as the transport (or are deliberately running without the app token), the daemon runs the maintenance loop only and the demo uses the operator-CLI answer path — both are valid (P1-DEMO-SCRIPT.md).

ROLLBACK: none needed — verification only, no host state change.


STEP 8 — Live P1 four-criteria acceptance demo 🧑 OPERATOR-REQUIRED

Run P1-DEMO-SCRIPT.md in full. All four §3.3.1 exit criteria must pass: (a) crash-safe resume, (b) duplicate-answer no-op, (c) post-deadline rejection + park, (d) two concurrent tasks resume independently. With D-1 fixed you may exercise the live Slack answer path for (b)/(d); the operator-CLI answer path remains available and exercises the identical compare-and-set. Do not accept P1 until all four pass.

ROLLBACK: the demo writes only ledger rows under state/; reset via the Step-5 ledger rollback, or restore the Step-1 snapshot, then re-run.


STEP 9 — Plane-1 checker live dry-runs 🧑 OPERATOR-REQUIRED

The six checkers live under security-review/checkers/: aws-posture.sh, compliance-drift.sh, confluence-doc.sh, dependency-cve.sh, doc-drift.sh, plan-groomer.sh. All run report + ALARM-only (design D3): a clean run posts nothing and lands a mode-600 report; no auto-Jira/Notion writes.

  • Dry-run each checker against ~/repo-mirrors in report-only mode; confirm a clean run posts nothing and writes a mode-600 report. Confirm the exact invocation against the checker scripts and the shared security-review/lib/ substrate on the box.
  • Confirm the canary suite runs first (a planted-fault miss is a COMPLACENCY ALARM and that role is skipped — design §6.4), and the coverage rotation pointer advances (a slipped role is a COVERAGE ALARM, deferred-not-dropped).

ROLLBACK: checkers are read-only over the mirror corpus; a dry-run produces only a report file. Remove the report dir to revert; no host state change.


STEP 10 — P5 cross-plane loop dry-run 🧑 OPERATOR-REQUIRED

The Plane-1→Plane-2 loop turns confirmed checker findings into pipeline tasks:

cd ~/orchestrator/agent-team && . .venv/bin/activate
# Read one or more checker report JSONs and start one task per confirmed
# at/above-threshold finding (default threshold: high). --dry-run posts nowhere.
python3 run-team.py intake-checker --report <path-to-report.json> --threshold high --dry-run

The Tier-3 dep-bump fixer is dry-run only on the box (it holds no write token, D2):

python3 run-team.py fix --report security-review/<...>/dependency-cve.json \
  --finding-id <ID> --task-id <task> --dry-run
# prints the fix spec + minimal bump patch + the CI workflow_dispatch inputs;
# dispatches NOTHING. Live dispatch is the P3-live flip below.

ROLLBACK: both are read-only / dry-run (no dispatch, no apply); intake-checker de-dup is in-memory per process. No host state to revert beyond ledger rows from a non-dry-run intake (Step-5 ledger rollback).


STEP 11 — Wire the schedule 🧑 OPERATOR-REQUIRED

The coordinator daemon (Step 6) is always-on, not timer-driven. The per-checker timers (and the design's "shared timer with secrev", §8) attach here. The existing sea-haven-secrev.timer (OnCalendar 02:00, Persistent) is untouched. Add a checker timer only after that checker is dry-run-validated (Step 9).

ROLLBACK: each timer gets its own systemctl disable --now <unit>.timer + rm.


The P3-live flip (deferred; IAM + GitHub App gated) 🧑 OPERATOR-REQUIRED

P3 (the build→verify apply-and-open-draft-PR loop) is opt-in and inert in this deploy: run-team.py / serve pass build_verify_wiring=None, so no P3 subgraph is assembled. Flipping it live is a separate, gated provisioning session, not part of the coordinator deploy:

  1. Mandatory reviews first. The CI trust-boundary + OIDC IAM change is a breaking IAM change → GPT-4.1 cross-review (global instructions) AND /sh-security-review on the apply/verify surface. Do not flip without both.
  2. Provision the GitHub App for the trusted apply path (the App that opens the draft PR), and the agent-apply GitHub Actions environment that holds the apply path's scoped permissions.
  3. Bind the live build→verify wiring via agent_team.coordinator.gated_build_verify_wiring(...) (the read-only CI result fetcher + the real diff builder) — the seam a leaf calls after the gate clears. The CI fetcher is read-only and fails closed (missing token / 404 / auth failure → None → the gate BLOCKs and the task parks).
  4. Set the apply env vars the live path reads (the read-only CI-result token and the dispatch target), then re-run the fixer without --dry-run only once the dispatcher is bound.

Until every step above is done, the box dispatches/applies nothing.


Post-session definition-of-done (design §7, global instructions)

  • All four P1 criteria demonstrated live (Step 8).
  • project_r720_agent_team memory created/updated.
  • Confluence "AWS Architecture Map" / IT host inventory updated to show sh-secrev now also hosts the always-on agent-team coordinator daemon.
  • /sh-security-review run on the Slack inbound listener surface (auth + untrusted-input; mandatory) — flag outstanding if not run.
  • OPERATOR-RUNBOOK.md reviewed by whoever holds the pager.
  • Snapshot retained until the daemon runs clean for one full cycle, then pruned.