docs(agent-team): land provisioning + operator runbooks under docs/provisioning

- PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5
  intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub
  App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked
  FIXED so the demo can use the live Slack answer path.
- P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the
  Slack and operator-CLI answer paths documented for all four exit criteria.
- DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the
  operator-CLI divergence kept as provisioning notes.
- OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks,
  failed HITL resumes, budget exhaustion, transport outages, and
  COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the
  re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6).
This commit is contained in:
Adam Moussa 2026-06-18 16:45:31 -04:00
parent bec592bf5a
commit 2ed84389a9
4 changed files with 1165 additions and 0 deletions

View file

@ -0,0 +1,187 @@
# DEPLOY-AUDIT — `agent-team/DEPLOY-R720.md` + systemd unit vs. actual code
Cross-check of the drafted deploy doc (`agent-team/DEPLOY-R720.md`) and the
systemd unit (`agent-team/systemd/agent-team-coordinator.service`) against the
**actual current code** in `agent-team/agent_team/**` and `agent-team/run-team.py`.
Each finding: **location → claimed → actual → fix**. Severity: 🔴 blocker /
🟠 should-fix / 🟡 nit. Items confirmed clean are stated explicitly.
> **Status (deploy-readiness PR):** the three deploy-correctness bugs that
> blocked the live provisioning session — **D-1, D-2, D-7** — are **RESOLVED** in
> this PR (`feature/agent-team-deploy-readiness`). The remaining items
> (**D-4 / D-5** dependency pinning, the **operator-CLI divergence**) are kept
> below as **provisioning notes** — they do not block the coordinator deploy.
---
## ✅ D-1 — `serve` now starts the Slack inbound listener — RESOLVED (this PR)
- **Location:** `agent_team/coordinator.py` (`Coordinator.serve` +
`_maybe_start_slack_listener` / `_slack_listener_enabled` / `_stop_slack_listener`
/ `default_slack_listener_factory`); `agent_team/transport/slack_listener.py`
(`SlackListener.serve` + new `close`).
- **Was:** `Coordinator.serve()` did only `bind_subscription_invoker()`,
`setup()`, `recover()`, then an infinite tick/sleep loop — it never constructed
or started `SlackListener`, so a deployed daemon posted clarifier questions and
expired them on deadline but could **not hear Slack answers**.
- **Now:** `serve()` starts the `SlackListener` on a background **daemon thread**,
concurrently with the tick/drain loop, **when** the live transport is a
`SlackTransport` AND `SLACK_APP_TOKEN` is set. It shares the coordinator's own
transport, ledger `db_path`, and `resume_queue` put; on shutdown it calls the
listener's new `close()` and joins the thread in a `finally`. When Slack is not
the transport or the app token is absent, no listener starts and `serve` behaves
exactly as before — **Slack is never made mandatory**. The AUTHZ-01 owner
allowlist + the open-status compare-and-set are untouched (still fail closed on
an empty `AGENT_TEAM_SLACK_OWNER_IDS`).
- **Tests:** `tests/test_coordinator.py` — start-when-Slack+app-token, no-start
without the token, no-start when the transport is not Slack, idempotent start,
clean shutdown, serve start/stop around the loop; `tests/test_slack_listener.py`
— `close()` no-op + handler teardown.
---
## ✅ D-2 — coordinator unit loads `~/orchestrator/.env` — RESOLVED (this PR)
- **Location:** `agent-team/systemd/agent-team-coordinator.service`;
`run-team.py:_build_coordinator` (always wires `default_review_wiring`);
`coordinator.py:default_review_wiring` → `review_loop_llm.default_plan_reviewer`
(shells the local orchestrator `run.py` → `cross_reviewer` GPT-4.1).
- **Was:** the unit loaded only `~/secrev.env`. `run-team.py serve` builds the
coordinator with the P2 review loop wired, and once a task reaches REVIEW the
reviewer shells the orchestrator `run.py`, whose GPT-4.1 call reads the
non-Claude provider key from `~/orchestrator/.env`. The review path would fail
to authenticate.
- **Now:** the unit adds `EnvironmentFile=-/home/adam/orchestrator/.env`
(optional `-`, mirroring the `sea-haven-secrev` unit). Does not affect the P1
demo (P1 stops at PLAN before REVIEW); closes the latent P2 break.
---
## 🟠 D-4 — pip install list omits `requests` (provisioning note)
- **Location:** `agent_team/transport/github_live.py` / `github_intake.py`
(`import requests`).
- **Actual:** the GitHub transport + intake require `requests`; the Slack-first
path does not hit it, but any `--transport github` / `intake-github` use fails
with a clear RuntimeError without it.
- **Fix:** PROVISIONING-RUNBOOK Step 4 installs `requests` into the venv.
`requests` is also absent from `requirements.txt` (see D-5).
---
## 🟠 D-5 — agent-team runtime deps are not pinned in `requirements.txt` (provisioning note)
- **Location:** `requirements.txt` (repo root).
- **Actual:** `requirements.txt` pins `langgraph==1.1.10` and
`langgraph-checkpoint-sqlite==3.1.0`, but the agent-team runtime deps
`claude-agent-sdk`, `slack_sdk`, `slack_bolt`, `requests` (and `anthropic` for
api mode) are **not in `requirements.txt` at all** — they are installed ad-hoc
into the agent-team venv by the runbook. There is no pinned, reproducible source
of truth for the box's runtime set.
- **Fix (deferred):** add an `agent-team/requirements.txt` (or extras group)
pinning these, version-matched to the root `requirements.txt` langgraph pin.
Until then, PROVISIONING-RUNBOOK Step 4 pins `langgraph==1.1.10` /
`langgraph-checkpoint-sqlite==3.1.0` explicitly so the unpinned `pip install`
cannot pull a newer, untested major. **Do NOT modify `requirements.txt` or the
checkers in this PR** (out of scope).
---
## ✅ D-6 — `slack_bolt` is now exercised by the daemon — RESOLVED (consequence of D-1)
- **Location:** `slack_listener.py:serve` (the only `slack_bolt` import).
- **Now:** with D-1 fixed, `Coordinator.serve()` starts `SlackListener.serve()`,
which imports + uses `slack_bolt` for the Socket Mode handler. The dep is right
and now actually exercised on the live Slack path.
---
## ✅ D-SLACKVAR (clean) — `SLACK_CHANNEL_ID` matches
- `run-team.py:_build_transport` reads exactly `os.environ.get("SLACK_CHANNEL_ID")`.
The runbook, the unit comment, and the code all use `SLACK_CHANNEL_ID` (not
`SLACK_CHANNEL`). **CLEAN.**
---
## ✅ D-ENV-SLACKBOT / OAUTH / OWNERS / APPTOKEN (clean) — names match
- **`SLACK_BOT_TOKEN`** ↔ `slack_live.py` + `slack_listener` env read. **CLEAN.**
- **`CLAUDE_CODE_OAUTH_TOKEN`** ↔ `invoker.py`. **CLEAN.**
- **`AGENT_TEAM_SLACK_OWNER_IDS`** ↔ `slack_listener.py` (name + fail-closed
semantics). **CLEAN** — now read by the running daemon (D-1 fixed).
- **`SLACK_APP_TOKEN`** — now read in two places: `coordinator._slack_listener_enabled`
gates the listener on its presence, and `default_slack_listener_factory` /
`SlackListener.serve` source it to open the socket. **CLEAN** (read site exists
now that D-1 is fixed).
---
## ✅ D-7 — `ExecStart` uses the venv interpreter — RESOLVED (this PR)
- **Location:** `agent-team/systemd/agent-team-coordinator.service` ExecStart.
- **Was:** `ExecStart=/usr/bin/env python3 run-team.py serve` resolved the
**system** interpreter under systemd's PATH — not the venv where the deps were
installed, so the daemon would fail at import.
- **Now:** `ExecStart=/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve`
(matches the runbook venv path + `WorkingDirectory`).
---
## ✅ D-SUBCMD (mostly clean) — run-team.py subcommands referenced exist
Cross-checked every `run-team.py <sub>` the deploy doc + demo name against
`run-team.py:build_parser`: `init-db`, `serve`, `list` (+ `--all` / `--parked`),
`show`, `expire`, `answer`, `redeliver`, `supersede`, `force-resume`, `start`,
`intake-github`, `intake-checker`, `fix` — all exist. No invented verbs.
> ### Operator-CLI divergence (provisioning note)
>
> Two operator CLIs exist with **different verb names**:
>
> - `run-team.py` (the entry CLI): `init-db, list, show, redeliver, expire,
> answer, supersede, force-resume, start, serve, intake-github, intake-checker,
> fix`. Has `show` and `--parked`; `--db` / `--audit-log` default sensibly.
> - `agent_team/operator_cli.py`: `list, redeliver, force-expire,
> answer-on-behalf, force-resume` — **no `show`**, `--db` / `--audit-log` are
> **required**, and its `force-resume` **supersedes** (unlike `run-team.py`'s,
> which reopens an expired row and never supersedes).
>
> **Use `run-team.py` for provisioning + the demo + incident recovery.** The
> docs reference only `run-team.py`. Reconciling the two CLIs is a follow-up.
---
## ✅ D-PYTHONPKG (clean) — package import bootstrap is correct
`run-team.py` inserts its own dir into `sys.path` so the hyphenated script
imports the `agent_team` package without an editable install. **CLEAN.**
---
## ✅ D-HARDENING (clean, and matches the locked decision)
- Unit: `NoNewPrivileges=true`, `ProtectSystem=full`, `ProtectHome=read-only`,
`ReadWritePaths=/home/adam/orchestrator/agent-team/state`. **Retained unchanged**
in this PR (locked decision — do not revert to secrev parity).
- The `ReadWritePaths` carve-out matches the ledger + audit-log location
(`state/agent_team.sqlite`, `state/audit.log.jsonl`). **CLEAN.**
---
## Summary table
| ID | Sev | Status | One-line |
|---|---|---|---|
| D-1 | 🔴 | ✅ RESOLVED (PR) | `serve` starts `SlackListener` (Slack + app-token gated; Slack stays optional) |
| D-2 | 🔴 | ✅ RESOLVED (PR) | unit loads `~/orchestrator/.env` for the P2 GPT-4.1 review provider key |
| D-7 | 🟡 | ✅ RESOLVED (PR) | `ExecStart` points at the agent-team venv interpreter |
| D-6 | 🟡 | ✅ RESOLVED | `slack_bolt` now exercised by the daemon (consequence of D-1) |
| D-4 | 🟠 | NOTE | pip list omits `requests` — runbook Step 4 installs it |
| D-5 | 🟠 | NOTE | agent-team runtime deps not pinned in `requirements.txt` — runbook pins langgraph |
| operator-CLI | — | NOTE | `run-team.py` vs `operator_cli.py` divergent verbs — use `run-team.py` |
| D-SLACKVAR | ✅ | CLEAN | `SLACK_CHANNEL_ID` matches everywhere |
| D-ENV-* | ✅ | CLEAN | bot/oauth/owner/app-token env names match; all read sites now exist |
| D-SUBCMD | ✅ | CLEAN | every `run-team.py` verb/flag the docs cite exists |
| D-HARDENING | ✅ | CLEAN | unit hardening retained unchanged (locked decision) |

View file

@ -0,0 +1,290 @@
# OPERATOR-RUNBOOK — R720 agent-team coordinator incident handling
The on-call runbook for the always-on `agent-team-coordinator` daemon on the
`sh-secrev` R720 VM. Covers pipeline stalls, stuck/parked tasks, failed
human-in-the-loop resumes, budget exhaustion mid-pipeline, transport outages, and
the COMPLACENCY / COVERAGE alarms. Grounds every recovery in real code (design
§5 escalation ladder + §6.6 contention/park policy; Phase-6 requirement).
> **CLI used throughout: `run-team.py`** (the entry CLI), run from
> `~/orchestrator/agent-team` with the venv active so it hits the default ledger
> (`state/agent_team.sqlite`) and audit log (`state/audit.log.jsonl`). Do **not**
> use `agent_team/operator_cli.py` — it has divergent verbs (no `show`, required
> `--db`/`--audit-log`, and a `force-resume` that supersedes). See DEPLOY-AUDIT.md.
```bash
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team && . .venv/bin/activate
```
## Verb reference (all confirmed in `run-team.py:build_parser`)
| Verb | Effect | Destructive? |
|---|---|---|
| `list` / `list --all` / `list --parked` / `list --status <state>` | read pending questions (default `open`) | no |
| `show <question_id>` | print one ledger row (JSON) | no |
| `redeliver <question_id>` | clear `channel_ref` so the reconcile loop re-posts an `open` question | no (audit-logged) |
| `expire <question_id> --confirm` | force `open`→`expired` | yes |
| `answer <question_id> --answer <p> [--via <id>] --confirm` | answer-on-behalf (first-answer-wins CAS) | yes |
| `force-resume <question_id> --confirm` | reopen an `expired` (parked) question; for `answered` records resume intent | yes |
| `supersede <question_id> --confirm` | mark a stale `open`/`answered` row `superseded` | yes |
| `start --task "..." [--transport ...] [--dry-run]` | start one task to the human gate | no |
`--operator <name>` (global) sets the audit attribution; it defaults to the OS
login. Destructive verbs require `--confirm` and write an attempt-then-outcome
record to the audit log **before** mutating.
First triage for any incident:
```bash
systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e --since "-2h" | tail -100
python3 run-team.py list --all # full ledger snapshot
python3 run-team.py list --parked # non-open rows (parked-task context)
```
---
## Incident 1 — Pipeline stall (tasks not advancing)
**Symptoms:** `list` shows `open` (or `answered`) rows that never progress; the
journal shows no `tick` activity or repeated errors.
**Diagnose:**
```bash
systemctl is-active agent-team-coordinator.service # "active" expected
journalctl -u agent-team-coordinator.service -e | tail -60
# Daemon dead/looping on restart? Check the unit + deps:
.venv/bin/python -c "import langgraph, slack_sdk, slack_bolt, requests; print('deps ok')"
```
**Recover:**
1. If the daemon is dead and `Restart=on-failure` is flapping, read the journal
for the import/auth error. A missing venv dep (D-4/D-5) or a missing
`CLAUDE_CODE_OAUTH_TOKEN` is the usual cause — fix the env / venv, then:
```bash
sudo systemctl restart agent-team-coordinator.service
```
2. The startup `recover()` sweep re-drives `answered`-but-unresumed rows and
re-posts `open` rows that lost their `channel_ref`, so a clean restart
converges from the durable ledger. Confirm with `journalctl ... | tail` and
`list --all`.
3. A single task stuck `open` with a stale/missing post: re-deliver it.
```bash
python3 run-team.py show <qid>
python3 run-team.py redeliver <qid> # clears channel_ref; reconcile re-posts
```
**Escalate** (see the ladder) if a restart does not clear it within one tick
cadence and the journal shows a non-transient error.
---
## Incident 2 — Stuck / parked task
A task parks (design §6.6) when its clarifier question **expires** with no answer,
or its budget headroom drops below reserve, or a checkpoint is corrupt. Parked =
the durable ledger row is no longer `open` (it is `expired`), and an ALARM was
raised, not spun on.
**Diagnose:**
```bash
python3 run-team.py list --parked
python3 run-team.py show <qid> # status, thread_id, deadline_at, answered_via
```
**Recover — depends on why it parked:**
- **Expired with no answer (the common case)** — un-park by reopening the expired
question; it is then re-delivered for an answer:
```bash
python3 run-team.py force-resume <qid> --confirm
# "force-resume: reopened expired question <qid>; it will be re-delivered"
python3 run-team.py show <qid> # status flips back to "open"
```
Then answer it (Slack or CLI) to drive it forward.
- **Answered but not yet resumed** — the recovery sweep handles it; `force-resume`
records intent and reports that (no mutation):
```bash
python3 run-team.py force-resume <qid> --confirm
# "force-resume: question <qid> is answered and pending resume; the recovery
# sweep will resume it (intent recorded)"
sudo systemctl restart agent-team-coordinator.service # forces the recover() sweep now
```
- **Stale / wrong question that should be abandoned** — supersede it so it stops
surfacing as parked context:
```bash
python3 run-team.py supersede <qid> --confirm
```
`MAX_PARK` FIFO-aging: a task that exceeds the park window escalates (ALARM + a
Jira ticket per the ladder) rather than starving silently.
---
## Incident 3 — Failed human-in-the-loop resume
**Symptoms:** an answer was submitted (Slack or CLI) but the graph did not
advance.
**Diagnose:**
```bash
python3 run-team.py show <qid> # is status "answered"? what answered_via?
journalctl -u agent-team-coordinator.service -e | grep -iE "resume|answer|<qid>"
```
**Recover:**
- **Status is `answered` but no resume drained** — the resume queue is in-process;
a daemon restart triggers the `recover()` sweep that re-drives `answered` rows
via the turn-guarded `ResumeWorker` (idempotent — an already-advanced thread
supersedes-and-skips):
```bash
sudo systemctl restart agent-team-coordinator.service
python3 run-team.py show <qid> # confirm it advanced
```
- **Live Slack answer never registered** — the listener fails closed. Check:
```bash
journalctl -u agent-team-coordinator.service -e | grep -i "Slack answer"
# "rejecting Slack answer: owner allowlist is unconfigured ..." -> set
# AGENT_TEAM_SLACK_OWNER_IDS in ~/secrev.env and restart.
# "rejecting Slack answer ... unauthorized sender" -> the answerer's user id is
# not in the allowlist. Add it, or answer via the CLI on their behalf:
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
```
- **Answer lost the compare-and-set (`not open`)** — the row was already
answered/expired/superseded. Inspect with `show`; if it parked, go to Incident 2.
---
## Incident 4 — Budget exhaustion mid-pipeline
The shared Claude budget ledger (`budget_ledger` table) enforces per-call + total
nightly caps (design §6.1, §6.6). A stage that would breach the reserve **parks**
the task (deferred, not dropped) and ALARMs — it never loop-drains the pool.
**Diagnose:**
```bash
journalctl -u agent-team-coordinator.service -e | grep -iE "budget|reserve|park"
# Inspect today's spend directly (no CLI verb for the budget ledger; read it).
# Columns (agent_team/db/schema.py budget_ledger DDL): thread_id, stage, model,
# billing_mode, input_tokens, output_tokens, usd_cost, recorded_at, day_bucket.
sqlite3 state/agent_team.sqlite \
"SELECT day_bucket, thread_id, ROUND(SUM(usd_cost),4) AS usd, SUM(input_tokens) AS in_tok
FROM budget_ledger GROUP BY day_bucket, thread_id ORDER BY day_bucket DESC LIMIT 20;"
```
**Recover:**
- **Wait for the next budget window** — budget-exhausted roles/tasks are deferred
via the rotation pointer and picked up next cycle; this is the intended
behavior, not a failure. The parked task surfaces in `list --parked`.
- **Force a specific parked task forward now** (e.g. it is urgent and headroom has
since freed): un-park it and let the daemon re-run the stage within the
remaining cap:
```bash
python3 run-team.py force-resume <qid> --confirm
```
- **Confirm `ANTHROPIC_API_KEY` is absent** — its presence would silently meter to
API rates and blow the budget model (the billing seam pops it defensively, but
it must not be set):
```bash
grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # both must print 0
```
Do **not** raise the cap to push a task through without Adam's decision — budget
realism is a deliberate guardrail. Escalate per the ladder if a task repeatedly
parks on budget across nights.
---
## Incident 5 — Transport outage (Slack / GitHub down or misconfigured)
**Symptoms:** clarifier posts fail; the journal shows transport errors or the
inbound listener is not started.
**Diagnose:**
```bash
journalctl -u agent-team-coordinator.service -e | grep -iE "Slack|listener|post|transport"
# Did the inbound listener start?
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
# "started (Socket Mode ...)" -> inbound up
# "not started (... SLACK_APP_TOKEN is unset)" -> Slack inbound intentionally off
```
**Recover:**
- **Outbound post failed (Slack API / network)** — the ledger row stays `open`
with no `channel_ref` (the responder leaves it for reconcile). After the
transport recovers, the `recover()` sweep (on restart) or a manual `redeliver`
re-posts:
```bash
python3 run-team.py redeliver <qid>
```
- **Inbound listener down / never started** — the daemon still posts and expires;
it just cannot hear Slack. **The CLI `answer` path is the outage fallback** —
it runs the identical compare-and-set with no socket:
```bash
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
```
Fix `SLACK_APP_TOKEN` / `SLACK_BOT_TOKEN` in `~/secrev.env`, then
`sudo systemctl restart agent-team-coordinator.service` to re-establish Socket
Mode. (D-1: the listener starts only when the transport is live Slack AND
`SLACK_APP_TOKEN` is set; it is optional by design.)
- **Slack listener crash-looping** — the listener runs on an isolated daemon
thread; a crash is logged (`inbound Slack listener thread exited with an error`)
and takes down only the inbound socket, not the maintenance loop. Restart the
service to relaunch the thread once the cause is fixed.
---
## Incident 6 — COMPLACENCY / COVERAGE alarms (Plane-1 checkers)
These come from the nightly checker run, not the coordinator daemon (design §6.4,
§6.6):
- **COMPLACENCY ALARM** — a checker role missed a planted canary fault. That role
is **skipped** for the night (it never runs silently degraded). Recover: inspect
the canary corpus / the role's checker under `security-review/checkers/`, fix the
regression, re-run that checker's dry-run, and confirm the canary passes before
re-enabling.
- **COVERAGE ALARM** — a role slipped its rotation slot (e.g. budget-deferred). It
is **deferred via the rotation pointer, never dropped**, and picked up next
cycle. Recover: confirm the rotation pointer advanced (it is rebuildable from
report history) and that the role runs on the next cadence; investigate only if
it slips repeatedly.
Both alarms are **report + ALARM-only** (design D3): nothing posts on a clean
state, no auto-Jira/Notion writes from the checker itself. Persistent alarms
follow the escalation ladder below.
---
## Escalation ladder (design §5 / §6.6, resolves Q3)
Anything that does not clear on the first ALARM escalates — but **ALARM-only in
spirit** (nothing posts on a clean state):
1. **Re-alarm on a backoff.** A confirmed critical (or a COMPLACENCY / COVERAGE
alarm) that persists re-alarms to Slack each night it is still unresolved, on
a backoff so it does not spam.
2. **Open a Jira tracking ticket after `N` nights** (default **N = 3**). If the
condition still has not cleared, the coordinator opens an **INFRA** Jira ticket
so it cannot quietly linger. The same ladder applies to a parked task that
exceeds `MAX_PARK`.
3. **Human (Adam) takes it from the Jira ticket.** For a parked task, recover via
Incidents 2–4 above; for a checker alarm, via Incident 6.
When you resolve an incident, record the action — destructive CLI verbs already
write an attributable attempt+outcome record to `state/audit.log.jsonl`; for
non-CLI recoveries note it on the Jira ticket.
---
## Post-incident
- Confirm the daemon is `active` and the ledger has no unexpected parked rows
(`list --parked`).
- If you restored the VM snapshot or wiped the ledger, re-run the relevant
PROVISIONING-RUNBOOK steps.
- Update `project_r720_agent_team` memory if the incident revealed a durable
fact (a new failure mode, a config that must change).

View file

@ -0,0 +1,286 @@
# P1-DEMO-SCRIPT — live four-criteria acceptance demo (design §3.3.1 / §7.1 P1)
The §7.1 P1 exit gate: demonstrate, on the live box, all four durable
human-in-the-loop criteria before P1 is accepted:
- **(a)** kill the box mid-wait and have the task resume after restart;
- **(b)** submit a duplicate answer and confirm it no-ops;
- **(c)** submit an answer after the deadline expired and confirm it is rejected
and the task parks;
- **(d)** two tasks suspended concurrently resume independently to the correct
thread.
Every command below is grounded in the **actual** code surface
(`run-team.py`, `coordinator.py`, `slack_listener.py`, `responder.py`,
`resume_worker.py`, `db/schema.py`, `graph.py`). No invented flags. Where the
code does not expose a needed knob (e.g. a short deadline), the script uses a
direct `sqlite3` write against the documented `pending_questions` schema and says
so.
## Two answer paths — live Slack OR the operator CLI
**D-1 is fixed:** `Coordinator.serve()` now starts the inbound `SlackListener`
when the live transport is Slack AND `SLACK_APP_TOKEN` is set, so the
**live Slack answer round-trip works**. You can run the demo either way:
- **Live Slack** — Adam clicks the Block Kit button / replies in the channel; the
listener normalizes the event, runs the AUTHZ-01 owner check, and drives the
first-answer-wins compare-and-set (`responder.submit_answer` →
`db.schema.answer_question`).
- **Operator CLI** — `run-team.py answer <qid> --answer ... --confirm` runs the
**identical** compare-and-set (audit-logged answer-on-behalf). Useful when the
Slack app is not yet provisioned, or to script the assertions.
Both exercise the same durable mechanic; the human gate decision is always
Adam's. The assertions below assert on the **durable ledger status** (the §3.3.1
source of truth) and are identical for either path. The examples use the CLI
`answer --confirm` form so they are copy-pasteable; substitute "Adam answers in
Slack" wherever you prefer the live path.
> The automated proof of this mechanic is `tests/sim/test_p1_exit_criteria.py`
> (a `SimPipeline` harness) and `tests/test_coordinator.py` (the serve/listener
> wiring). This script is the live-box demonstration on top of that.
## Verb map (use `run-team.py`, not `operator_cli.py`)
| Need | `run-team.py` verb | Notes |
|---|---|---|
| start a task to the human gate | `start --task "..." [--transport slack] [--dry-run]` | mints a `thread_id`, posts the clarifier, writes the `open` ledger row |
| list waiting questions | `list` (default `open`) / `list --all` / `list --parked` | JSON rows |
| inspect one row | `show <question_id>` | JSON row |
| answer on the task's behalf | `answer <question_id> --answer <payload> --confirm` | destructive, audit-logged; first-answer-wins CAS |
| force-expire an open question | `expire <question_id> --confirm` | destructive; flips `open`→`expired` |
| un-park (reopen) an expired question | `force-resume <question_id> --confirm` | only acts on `expired` rows (reopens them) |
`operator_cli.py` has different verbs (`force-expire`, `answer-on-behalf`) and
**no `show`**, and requires `--db`/`--audit-log` — do not use it here.
## Preconditions
```bash
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team
. .venv/bin/activate # so run-team.py uses the venv deps
# Confirm the ledger exists (Step 5 of the runbook):
python3 run-team.py list --all # [] on a fresh DB is fine
# For the live-Slack path, confirm the listener started:
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
# -> "inbound Slack listener started (Socket Mode, background thread)"
```
Run all `run-team.py` commands from `~/orchestrator/agent-team` so they hit the
default ledger (`state/agent_team.sqlite`) and default audit log
(`state/audit.log.jsonl`). The clarifier is posted to `SLACK_CHANNEL_ID`.
> **Resume execution model.** A won answer (rowcount 1) enqueues a `ResumeJob`
> onto the coordinator's in-process `resume_queue`; the daemon drains it on the
> next `tick()`, or the startup `recover()` re-drives any `answered`-but-unresumed
> row after a restart. So a CLI `answer --confirm` (or a live Slack answer) flips
> the ledger row to `answered`; the **running daemon** then resumes the graph.
> The demo asserts on the durable ledger status.
---
## (a) Crash-safe resume — kill mid-wait, restart, task resumes 🧑 Adam answers
**Setup — drive a task to the clarifier wait.** With the daemon already running:
```bash
python3 run-team.py start --task "demo-a: trivial scoped task"
# prints a thread_id, e.g. 3f2a... (record it as $TID_A)
python3 run-team.py list
# expect one row: status="open", a thread_id, a question_id (record as $QID_A),
# channel_ref set (Slack ts) or null if posting is dry/unavailable.
```
**Kill the box mid-wait, then restart:**
```bash
sudo systemctl stop agent-team-coordinator.service
# (optionally reboot the VM here for a stronger demonstration)
sudo systemctl start agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e | tail -40 # expect "starting" + recover() sweep
```
**ASSERT — durable state survived the kill (no answer yet):**
```bash
python3 run-team.py show $QID_A
# PASS: status == "open" (the question was NOT lost across the restart)
```
**🧑 Human gate — Adam answers after the restart (Slack OR CLI):**
```bash
# Live Slack: Adam replies/clicks in the channel. OR via the CLI:
python3 run-team.py answer $QID_A --answer "scope: just demo, no real change" --confirm
python3 run-team.py show $QID_A
# PASS: status == "answered"
# Wait one tick (~30s, DEFAULT_POLL_INTERVAL) for the daemon to drain the resume:
journalctl -u agent-team-coordinator.service -e | tail -20
```
**PASS criterion (a):** the question stayed `open` across the kill, and an answer
submitted *after* the restart drove it forward.
---
## (b) Duplicate answer is a no-op 🧑 Adam answers
```bash
python3 run-team.py start --task "demo-b: duplicate-answer test"
python3 run-team.py list # record the new question_id as $QID_B (status open)
```
**🧑 First answer (the winning one) — Slack OR CLI:**
```bash
python3 run-team.py answer $QID_B --answer "first-answer" --via "slack:U1" --confirm
# expect: "answered question <QID_B> (via slack:U1)" (exit 0)
```
**Duplicate / second answer (must lose the compare-and-set):**
```bash
python3 run-team.py answer $QID_B --answer "second-answer" --via "github:U2" --confirm
# expect: "answer no-op: question <QID_B> was not 'open' ..." on stderr, exit 1
```
For the **live-Slack** variant, Adam (or a second authorized owner) answers the
same message twice; the second event loses the CAS and is logged
`ignored Slack answer for question_id=... (not open: duplicate ...)`.
**ASSERT — the first answer is preserved:**
```bash
python3 run-team.py show $QID_B
# PASS: status == "answered"; answered_via == "slack:U1" (the FIRST answer)
```
**PASS criterion (b):** first-answer-wins (`rowcount==1`); the duplicate hits the
`BEGIN IMMEDIATE` compare-and-set in `answer_question` and is ignored
(`rowcount==0`) — no second resume, no overwrite.
---
## (c) Past-deadline answer rejected + task parks 🧑 Adam observes
> **Code reality:** `run-team.py start` always sets the clarifier deadline from
> `DEFAULT_CLARIFY_DEADLINE = 24h`. There is **no CLI flag for a short deadline**.
> Two faithful ways to demo expiry without waiting 24h:
### Option C1 — force the deadline past, let the daemon's sweep expire it (most faithful)
```bash
python3 run-team.py start --task "demo-c: deadline test"
python3 run-team.py list # record $QID_C (status open)
# Set this question's deadline into the past directly in the ledger:
sqlite3 state/agent_team.sqlite \
"UPDATE pending_questions SET deadline_at='2000-01-01T00:00:00+00:00' WHERE question_id='$QID_C';"
# Wait one daemon tick (~30s) for the deadline sweep to flip it, OR observe:
journalctl -u agent-team-coordinator.service -e | tail -20
# expect the park ALARM line: "task parked: clarifier question <QID_C> expired ..."
```
### Option C2 — operator force-expire (if you do not want to touch the DB)
```bash
python3 run-team.py start --task "demo-c: deadline test"
python3 run-team.py list # record $QID_C
python3 run-team.py expire $QID_C --confirm # destructive, audit-logged; open->expired
```
**ASSERT — the question is expired and the late answer is rejected:**
```bash
python3 run-team.py show $QID_C
# PASS: status == "expired"
# 🧑 Adam submits a LATE answer (Slack OR CLI) — it must lose the CAS:
python3 run-team.py answer $QID_C --answer "too-late" --confirm
# expect: "answer no-op: question <QID_C> was not 'open' ..." stderr, exit 1
python3 run-team.py show $QID_C
# PASS: status still "expired"; answer_json still NULL
# The parked task surfaces in the parked view:
python3 run-team.py list --parked # PASS: $QID_C appears here
```
**Deliberate un-park (proves the operator recovery path, §6.6):**
```bash
python3 run-team.py force-resume $QID_C --confirm
# expect: "force-resume: reopened expired question <QID_C>; it will be re-delivered"
python3 run-team.py show $QID_C
# status flips back to "open" (reopen_question), deadline_at cleared.
```
**PASS criterion (c):** the expired question rejects the late answer, the task
parks rather than spins, and the operator can deliberately un-park it via
`force-resume` (which reopens only an `expired` row).
> **`force-resume` semantics (verified):** `run-team.py force-resume` reopens an
> `expired` question (the parked case). For an `answered` question it records
> intent and reports the recovery sweep will resume it (no mutation). For
> `open`/`superseded`/absent it is a no-op exit 1. It does **not** supersede.
---
## (d) Two concurrent tasks resume independently 🧑 Adam answers
```bash
python3 run-team.py start --task "demo-d task A" # record $TID_A2, then:
python3 run-team.py start --task "demo-d task B" # record $TID_B2
python3 run-team.py list
# expect TWO open rows with DISTINCT thread_id AND distinct question_id.
# Record $QID_A2 and $QID_B2 — match by thread_id.
```
**(Optional) restart the daemon first** to also show concurrent tasks survive a
restart, then answer.
**🧑 Answer the SECOND task first, with a distinct answer, then the first
(Slack OR CLI):**
```bash
python3 run-team.py answer $QID_B2 --answer "answer-for-B" --via "slack:U2" --confirm
python3 run-team.py answer $QID_A2 --answer "answer-for-A" --via "slack:U1" --confirm
# Wait one daemon tick (~30s) for both resumes to drain.
```
**ASSERT — each task carries its OWN answer; no cross-talk:**
```bash
python3 run-team.py show $QID_A2
# PASS: status "answered", answered_via "slack:U1", thread_id == $TID_A2
python3 run-team.py show $QID_B2
# PASS: status "answered", answered_via "slack:U2", thread_id == $TID_B2
python3 run-team.py list --all
# PASS: the two rows resolved on their own thread_id; no cross-contamination.
```
**PASS criterion (d):** two concurrently-suspended tasks each resumed to their
own `thread_id` with their own answer — answering B before A did not misroute,
and the per-thread single-flight guard kept them independent.
---
## Acceptance
P1 is accepted only when **(a), (b), (c), and (d) all pass** on the live box.
Record the four `show` outputs (or `journalctl` excerpts) as evidence. Then
complete the PROVISIONING-RUNBOOK post-session definition-of-done (memory +
Confluence + the mandatory `/sh-security-review` on the Slack inbound listener).
## What can only be verified on the live box
- That the daemon actually drains the resume and the LangGraph checkpoint
advances (asserted via ledger status + journalctl; the graph-state advance is
observable only on the box).
- The live **Slack** post/answer round-trip end-to-end (the listener is wired —
D-1 fixed — but the live Socket Mode socket + a real `SLACK_BOT_TOKEN` /
`SLACK_APP_TOKEN` / channel membership are only present on the box).
- The exact `channel_ref` value (Slack message `ts`) — depends on a live Slack
post succeeding.

View file

@ -0,0 +1,402 @@
# PROVISIONING RUNBOOK — R720 agent-team Plane-2 coordinator
This is the ordered command sequence for the **operator-present** provisioning
session that stands up the always-on `agent-team-coordinator` daemon on the
`sh-secrev` R720 VM. Derived from `agent-team/DEPLOY-R720.md`,
`security-review/DEPLOY-R720.md`, and `docs/r720-agent-team-design.md`
(§3.3.1 / §7 / §7.1).
> ## State of this runbook
>
> The build is **merged and final**: all six Plane-1 checkers
> (`aws-posture`, `compliance-drift`, `confluence-doc`, `dependency-cve`,
> `doc-drift`, `plan-groomer`), the Tier-3 dep-bump **fixer**
> (`run-team.py fix --dry-run`), and the Plane-1→Plane-2 **P5 cross-plane loop**
> (`run-team.py intake-checker`) exist in the tree. The three deploy-correctness
> bugs the provisioning-prep audit found are **FIXED** (see DEPLOY-AUDIT.md):
>
> - **D-1 RESOLVED** — `Coordinator.serve()` now starts the inbound Slack
> `SlackListener` concurrently with the tick/drain loop when Slack is the live
> transport AND `SLACK_APP_TOKEN` is set. **The demo can use the live Slack
> answer path** (P1-DEMO-SCRIPT.md), or the operator-CLI `answer` path.
> - **D-2 RESOLVED** — the systemd unit loads `~/orchestrator/.env` (for the P2
> GPT-4.1 review loop's provider key) in addition to `~/secrev.env`.
> - **D-7 RESOLVED** — the unit's `ExecStart` points at the agent-team venv
> interpreter, not the system `python3`.
>
> Still open as provisioning notes (not blockers): **D-4/D-5** (agent-team
> runtime deps are installed ad-hoc into the venv and are not pinned in
> `requirements.txt`), and the **operator-CLI divergence** (`run-team.py` vs
> `agent_team/operator_cli.py` have different verb names — use `run-team.py`).
---
## Scope and ground rules
**Hard rules (carried from the design and global instructions):**
- **Snapshot before any stateful change** (design §7; `feedback_ec2_replacement_snapshot`).
- **Every stateful step has an exercised rollback.**
- Secrets use **placeholder names only**; real values are entered by the operator
at the box and never echoed into shell history.
- The box is **read-only / subscription-OAuth only**. **No `ANTHROPIC_API_KEY`**
on this host (it would silently win over OAuth and meter to API rates —
`billing.claude_invoke` pops it defensively, but it must not be present).
- **No IAM / OIDC is involved in the coordinator deploy** (P1/P2). IAM enters
only at the **P3-live flip** (see the dedicated section), which is gated on the
mandatory GPT-4.1 cross-review + `/sh-security-review`.
**Legend per step:**
- 🧑 **OPERATOR-REQUIRED** — needs the human (snapshot, secrets, the live human
gate, go/no-go). Cannot be automated.
- 🤖 **MECHANICAL** — deterministic; an operator runs it but it needs no judgment.
## Host facts (from both DEPLOY-R720.md files)
- Hypervisor: R720 at `10.10.60.40` (Windows Server 2022, Hyper-V).
- VM: `sh-secrev`, Ubuntu 24.04, **4GB / 2 vCPU / 40GB** dynamic vhdx, `10.10.60.120`.
- Reach: `ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120` (key-only, NOPASSWD sudo).
- Repo on the box: `~/orchestrator/` (rsync from the Mac, **NOT** a git clone).
The package lives at `~/orchestrator/agent-team/`.
- Shares `~/secrev.env` (mode 600) with the secrev sweep, and `~/orchestrator/.env`
(mode 600) for non-Claude provider keys (same files the secrev unit loads).
---
## STEP 1 — Snapshot the VM 🧑 OPERATOR-REQUIRED
**On the R720 host (Hyper-V), before anything else.** This is the one-command
undo for every change below.
```powershell
# On the R720 Windows host (PowerShell, as admin):
Checkpoint-VM -Name sh-secrev -SnapshotName "pre-agent-team-coordinator-$(Get-Date -Format yyyyMMdd-HHmm)"
Get-VMSnapshot -VMName sh-secrev # confirm the checkpoint exists
```
**ROLLBACK (whole session):**
```powershell
Restore-VMSnapshot -VMName sh-secrev -Name "<the checkpoint name above>" -Confirm:$false
Start-VM -Name sh-secrev
```
> Do not proceed until the checkpoint is confirmed present.
---
## STEP 2 — Rsync the repo to the box 🤖 MECHANICAL (verify manifest 🧑)
**From the Mac.** Same pattern/excludes as the secrev deploy. The team **scans
the same `~/repo-mirrors` corpus secrev already maintains** — this rsync ships
*code*, not mirrors.
```bash
# From the Mac (sync the canonical repo, not a worktree):
rsync -av --exclude .env --exclude .venv --exclude .git --exclude .claude \
~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/
```
Surface that must be on the box:
| Path (under `~/orchestrator/`) | Why it must ship |
|---|---|
| `agent-team/run-team.py` | the operator entry CLI |
| `agent-team/agent_team/**` | the package (coordinator, graph, ledger, transports, slack_listener, fixer) |
| `agent-team/systemd/agent-team-coordinator.service` | the daemon unit |
| `run.py` + the orchestrator package | the GPT-4.1 cross_reviewer the P2 review loop shells |
| `requirements.txt` | pin reference for langgraph / checkpoint-sqlite |
| `security-review/lib/**` | shared sweep substrate (Phase 0) |
| `security-review/checkers/**` | the six Plane-1 checkers + fixtures |
**ROLLBACK:** rsync is additive; restore the Step-1 snapshot to revert code state.
---
## STEP 3 — Write secrets 🧑 OPERATOR-REQUIRED
**On the VM.** The coordinator unit loads **two** EnvironmentFiles (both
optional via the leading `-`): `~/secrev.env` (agent-team runtime keys) and
`~/orchestrator/.env` (the non-Claude provider key for the P2 review loop).
**Placeholder names only — the operator pastes real values.** Do not echo real
tokens into shell history (use an editor or `read -s`).
`~/secrev.env` (mode 600) — append the agent-team keys:
```bash
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
# Edit ~/secrev.env (mode 600) and add — values are placeholders:
# CLAUDE_CODE_OAUTH_TOKEN=<CLAUDE_CODE_OAUTH_TOKEN> # from `claude setup-token`
# SLACK_BOT_TOKEN=<SLACK_BOT_TOKEN> # xoxb-..., chat:write (posts questions)
# SLACK_APP_TOKEN=<SLACK_APP_TOKEN> # xapp-..., connections:write (Socket Mode inbound)
# SLACK_CHANNEL_ID=<SLACK_CHANNEL_ID> # C0..., target clarifier channel
# AGENT_TEAM_SLACK_OWNER_IDS=<SLACK_OWNER_USER_IDS> # comma-separated U... ids of authorized answerers
chmod 600 ~/secrev.env
```
`~/orchestrator/.env` (mode 600) — the non-Claude provider key for the GPT-4.1
review loop (same file secrev uses; if it already exists with the key, leave it):
```bash
# OPENAI_API_KEY=<OPENAI_API_KEY> # (or the provider key cross_reviewer/GPT-4.1 needs)
chmod 600 ~/orchestrator/.env
```
**Contract notes (verified against the code — see DEPLOY-AUDIT.md):**
- `SLACK_CHANNEL_ID` is correct — `run-team.py _build_transport` reads exactly
`os.environ.get("SLACK_CHANNEL_ID")`. Do **not** use `SLACK_CHANNEL`.
- `SLACK_APP_TOKEN` is now **read by the daemon**: `Coordinator.serve()` starts
the inbound `SlackListener` when the transport is live Slack AND `SLACK_APP_TOKEN`
is set. Without it the daemon still runs (posts + expires) but never hears Slack
replies — Slack stays optional by design.
- `AGENT_TEAM_SLACK_OWNER_IDS` **fails closed** (AUTHZ-01): if unset/empty the
listener rejects **every** answer. It must be set for the live human gate.
- **CRITICAL:** confirm `ANTHROPIC_API_KEY` is NOT present:
```bash
grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # must print 0 for both
```
**ROLLBACK:** strip exactly the appended keys (keep secrev keys intact), or
restore the Step-1 snapshot. Do NOT blindly truncate — secrev keys live here too.
---
## STEP 4 — Create the venv + install pip deps 🤖 MECHANICAL
**On the VM.** A dedicated venv under `agent-team/.venv` (excluded from rsync).
The systemd unit's `ExecStart` points at **this venv's interpreter** (D-7 fixed),
so the deps MUST land here.
```bash
cd ~/orchestrator/agent-team
python3 -m venv .venv
. .venv/bin/activate
pip install langgraph==1.1.10 langgraph-checkpoint-sqlite==3.1.0 \
claude-agent-sdk slack_sdk slack_bolt requests
```
> **Pin note (D-5):** `requirements.txt` pins `langgraph==1.1.10` /
> `langgraph-checkpoint-sqlite==3.1.0`; match those exactly here. The other
> runtime deps (`claude-agent-sdk`, `slack_sdk`, `slack_bolt`, `requests`) are
> not yet in `requirements.txt` (D-5 open) — installed ad-hoc here. `requests`
> is required by the GitHub transport/intake (D-4). `anthropic` is **not**
> installed (only the opt-in `api` billing mode needs it).
> `slack_bolt` is now actually exercised (D-1 fixed: the daemon starts the
> Socket Mode listener).
**Verify the imports resolve (using the venv interpreter the unit will use):**
```bash
.venv/bin/python -c "import langgraph, langgraph.checkpoint.sqlite, slack_sdk, slack_bolt, requests; print('deps ok')"
.venv/bin/python -c "import claude_agent_sdk; print('agent-sdk ok')"
```
**ROLLBACK:** `deactivate 2>/dev/null; rm -rf ~/orchestrator/agent-team/.venv`
---
## STEP 5 — Initialize the durable ledger DB 🤖 MECHANICAL
**On the VM, venv active.** Idempotent; creates `state/agent_team.sqlite` with
the `pending_questions` + `budget_ledger` + `schema_meta` tables (LangGraph
`SqliteSaver` creates its own tables in the same file on first run).
```bash
cd ~/orchestrator/agent-team
. .venv/bin/activate
python3 run-team.py init-db
# Expect: "initialized ledger DB at .../state/agent_team.sqlite"
ls -l state/ # agent_team.sqlite present; state/ is gitignored
```
**ROLLBACK (reset the ledger only):**
```bash
cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak}
rm ~/orchestrator/agent-team/state/agent_team.sqlite*
# re-run `python3 run-team.py init-db` to recreate empty tables.
```
---
## STEP 6 — Install + start the systemd unit 🤖 MECHANICAL (go/no-go 🧑)
**On the VM, as root.** Installs the long-running coordinator daemon. Keep the
**hardening as-shipped**: `NoNewPrivileges`, `ProtectSystem=full`,
`ProtectHome=read-only` + `ReadWritePaths=.../agent-team/state` (locked decision).
```bash
sudo cp ~/orchestrator/agent-team/systemd/agent-team-coordinator.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now agent-team-coordinator.service
systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e -f
# expect: "agent-team coordinator starting"
# and (if Slack + SLACK_APP_TOKEN provisioned):
# "inbound Slack listener started (Socket Mode, background thread)"
# or otherwise:
# "inbound Slack listener not started (... or SLACK_APP_TOKEN is unset) ..."
```
The unit (post-fix) runs
`/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve` from
`WorkingDirectory=/home/adam/orchestrator/agent-team` as `User=adam`, loading
**both** `EnvironmentFile=-/home/adam/secrev.env` and
`EnvironmentFile=-/home/adam/orchestrator/.env`, `Restart=on-failure`.
> **D-7 fixed:** `ExecStart` now resolves the venv interpreter, so the Step-4
> deps are on the path. **D-1 fixed:** `serve` starts the inbound `SlackListener`
> when Slack + app token are provisioned. **Go/no-go:** confirm the journal shows
> the listener line you expect for your transport choice before Step 8.
**ROLLBACK (exercised — tear the unit down cleanly):**
```bash
sudo systemctl disable --now agent-team-coordinator.service
sudo rm /etc/systemd/system/agent-team-coordinator.service
sudo systemctl daemon-reload
systemctl status agent-team-coordinator.service # should report "could not be found"
```
The ledger under `state/` is untouched. Step-1 snapshot is the whole-host fallback.
---
## STEP 7 — Verify the Slack inbound listener 🧑 OPERATOR-REQUIRED
The clarifier gate is two halves: **outbound** (post the question — done by
`serve`/`start` via the Slack poster) and **inbound** (receive Adam's answer —
the `SlackListener` over Socket Mode).
> **D-1 RESOLVED.** `Coordinator.serve()` now constructs and starts
> `SlackListener` on a background daemon thread, concurrently with the tick/drain
> loop, **when** the live transport is a `SlackTransport` AND `SLACK_APP_TOKEN`
> is set. It shares the coordinator's own transport, ledger, and resume queue,
> and stops cleanly on shutdown. The AUTHZ-01 owner allowlist + the open-status
> compare-and-set are unchanged — the listener still fails closed on an empty
> `AGENT_TEAM_SLACK_OWNER_IDS`.
**Confirm the live inbound path is up:**
```bash
# In the journal (Step 6) expect the "inbound Slack listener started" line.
# Verify the two gating vars are present in the unit's environment:
grep -c SLACK_APP_TOKEN ~/secrev.env # 1
grep -c AGENT_TEAM_SLACK_OWNER_IDS ~/secrev.env # 1 (else the gate rejects all answers)
```
If you are **not** using Slack as the transport (or are deliberately running
without the app token), the daemon runs the maintenance loop only and the demo
uses the operator-CLI `answer` path — both are valid (P1-DEMO-SCRIPT.md).
**ROLLBACK:** none needed — verification only, no host state change.
---
## STEP 8 — Live P1 four-criteria acceptance demo 🧑 OPERATOR-REQUIRED
Run **P1-DEMO-SCRIPT.md** in full. All four §3.3.1 exit criteria must pass:
(a) crash-safe resume, (b) duplicate-answer no-op, (c) post-deadline rejection +
park, (d) two concurrent tasks resume independently. With D-1 fixed you may
exercise the **live Slack answer path** for (b)/(d); the operator-CLI `answer`
path remains available and exercises the identical compare-and-set. **Do not
accept P1 until all four pass.**
**ROLLBACK:** the demo writes only ledger rows under `state/`; reset via the
Step-5 ledger rollback, or restore the Step-1 snapshot, then re-run.
---
## STEP 9 — Plane-1 checker live dry-runs 🧑 OPERATOR-REQUIRED
The six checkers live under `security-review/checkers/`:
`aws-posture.sh`, `compliance-drift.sh`, `confluence-doc.sh`,
`dependency-cve.sh`, `doc-drift.sh`, `plan-groomer.sh`. All run **report +
ALARM-only** (design D3): a clean run posts nothing and lands a mode-600 report;
no auto-Jira/Notion writes.
- Dry-run each checker against `~/repo-mirrors` in report-only mode; confirm a
clean run posts nothing and writes a mode-600 report. Confirm the exact
invocation against the checker scripts and the shared
`security-review/lib/` substrate on the box.
- Confirm the **canary suite** runs first (a planted-fault miss is a COMPLACENCY
ALARM and that role is skipped — design §6.4), and the **coverage rotation**
pointer advances (a slipped role is a COVERAGE ALARM, deferred-not-dropped).
**ROLLBACK:** checkers are read-only over the mirror corpus; a dry-run produces
only a report file. Remove the report dir to revert; no host state change.
---
## STEP 10 — P5 cross-plane loop dry-run 🧑 OPERATOR-REQUIRED
The Plane-1→Plane-2 loop turns confirmed checker findings into pipeline tasks:
```bash
cd ~/orchestrator/agent-team && . .venv/bin/activate
# Read one or more checker report JSONs and start one task per confirmed
# at/above-threshold finding (default threshold: high). --dry-run posts nowhere.
python3 run-team.py intake-checker --report <path-to-report.json> --threshold high --dry-run
```
The Tier-3 dep-bump **fixer** is dry-run only on the box (it holds no write
token, D2):
```bash
python3 run-team.py fix --report security-review/<...>/dependency-cve.json \
--finding-id <ID> --task-id <task> --dry-run
# prints the fix spec + minimal bump patch + the CI workflow_dispatch inputs;
# dispatches NOTHING. Live dispatch is the P3-live flip below.
```
**ROLLBACK:** both are read-only / dry-run (no dispatch, no apply); `intake-checker`
de-dup is in-memory per process. No host state to revert beyond ledger rows from
a non-dry-run intake (Step-5 ledger rollback).
---
## STEP 11 — Wire the schedule 🧑 OPERATOR-REQUIRED
The coordinator daemon (Step 6) is **always-on**, not timer-driven. The
**per-checker timers** (and the design's "shared timer with secrev", §8) attach
here. The existing `sea-haven-secrev.timer` (OnCalendar `02:00`, Persistent) is
untouched. Add a checker timer only after that checker is dry-run-validated
(Step 9).
**ROLLBACK:** each timer gets its own `systemctl disable --now <unit>.timer` + `rm`.
---
## The P3-live flip (deferred; IAM + GitHub App gated) 🧑 OPERATOR-REQUIRED
P3 (the build→verify apply-and-open-draft-PR loop) is **opt-in and inert** in
this deploy: `run-team.py` / `serve` pass `build_verify_wiring=None`, so no P3
subgraph is assembled. Flipping it live is a **separate, gated** provisioning
session, not part of the coordinator deploy:
1. **Mandatory reviews first.** The CI trust-boundary + OIDC IAM change is a
breaking IAM change → **GPT-4.1 cross-review** (global instructions) AND
`/sh-security-review` on the apply/verify surface. Do not flip without both.
2. **Provision the GitHub App** for the trusted apply path (the App that opens
the draft PR), and the **`agent-apply` GitHub Actions environment** that holds
the apply path's scoped permissions.
3. **Bind the live build→verify wiring** via
`agent_team.coordinator.gated_build_verify_wiring(...)` (the read-only CI
result fetcher + the real diff builder) — the seam a leaf calls *after* the
gate clears. The CI fetcher is read-only and fails closed (missing token /
404 / auth failure → `None` → the gate BLOCKs and the task parks).
4. **Set the apply env vars** the live path reads (the read-only CI-result token
and the dispatch target), then re-run the fixer **without** `--dry-run` only
once the dispatcher is bound.
Until every step above is done, the box dispatches/applies nothing.
---
## Post-session definition-of-done (design §7, global instructions)
- [ ] All four P1 criteria demonstrated live (Step 8).
- [ ] `project_r720_agent_team` memory created/updated.
- [ ] Confluence "AWS Architecture Map" / IT host inventory updated to show
`sh-secrev` now also hosts the always-on agent-team coordinator daemon.
- [ ] `/sh-security-review` run on the Slack inbound listener surface
(auth + untrusted-input; mandatory) — flag outstanding if not run.
- [ ] OPERATOR-RUNBOOK.md reviewed by whoever holds the pager.
- [ ] Snapshot retained until the daemon runs clean for one full cycle, then pruned.