docs(agent-team): land provisioning + operator runbooks under docs/provisioning
- PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5 intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked FIXED so the demo can use the live Slack answer path. - P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the Slack and operator-CLI answer paths documented for all four exit criteria. - DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the operator-CLI divergence kept as provisioning notes. - OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks, failed HITL resumes, budget exhaustion, transport outages, and COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6).
This commit is contained in:
parent
bec592bf5a
commit
2ed84389a9
4 changed files with 1165 additions and 0 deletions
187
docs/provisioning/DEPLOY-AUDIT.md
Normal file
187
docs/provisioning/DEPLOY-AUDIT.md
Normal file
|
|
@ -0,0 +1,187 @@
|
|||
# DEPLOY-AUDIT — `agent-team/DEPLOY-R720.md` + systemd unit vs. actual code
|
||||
|
||||
Cross-check of the drafted deploy doc (`agent-team/DEPLOY-R720.md`) and the
|
||||
systemd unit (`agent-team/systemd/agent-team-coordinator.service`) against the
|
||||
**actual current code** in `agent-team/agent_team/**` and `agent-team/run-team.py`.
|
||||
|
||||
Each finding: **location → claimed → actual → fix**. Severity: 🔴 blocker /
|
||||
🟠 should-fix / 🟡 nit. Items confirmed clean are stated explicitly.
|
||||
|
||||
> **Status (deploy-readiness PR):** the three deploy-correctness bugs that
|
||||
> blocked the live provisioning session — **D-1, D-2, D-7** — are **RESOLVED** in
|
||||
> this PR (`feature/agent-team-deploy-readiness`). The remaining items
|
||||
> (**D-4 / D-5** dependency pinning, the **operator-CLI divergence**) are kept
|
||||
> below as **provisioning notes** — they do not block the coordinator deploy.
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-1 — `serve` now starts the Slack inbound listener — RESOLVED (this PR)
|
||||
|
||||
- **Location:** `agent_team/coordinator.py` (`Coordinator.serve` +
|
||||
`_maybe_start_slack_listener` / `_slack_listener_enabled` / `_stop_slack_listener`
|
||||
/ `default_slack_listener_factory`); `agent_team/transport/slack_listener.py`
|
||||
(`SlackListener.serve` + new `close`).
|
||||
- **Was:** `Coordinator.serve()` did only `bind_subscription_invoker()`,
|
||||
`setup()`, `recover()`, then an infinite tick/sleep loop — it never constructed
|
||||
or started `SlackListener`, so a deployed daemon posted clarifier questions and
|
||||
expired them on deadline but could **not hear Slack answers**.
|
||||
- **Now:** `serve()` starts the `SlackListener` on a background **daemon thread**,
|
||||
concurrently with the tick/drain loop, **when** the live transport is a
|
||||
`SlackTransport` AND `SLACK_APP_TOKEN` is set. It shares the coordinator's own
|
||||
transport, ledger `db_path`, and `resume_queue` put; on shutdown it calls the
|
||||
listener's new `close()` and joins the thread in a `finally`. When Slack is not
|
||||
the transport or the app token is absent, no listener starts and `serve` behaves
|
||||
exactly as before — **Slack is never made mandatory**. The AUTHZ-01 owner
|
||||
allowlist + the open-status compare-and-set are untouched (still fail closed on
|
||||
an empty `AGENT_TEAM_SLACK_OWNER_IDS`).
|
||||
- **Tests:** `tests/test_coordinator.py` — start-when-Slack+app-token, no-start
|
||||
without the token, no-start when the transport is not Slack, idempotent start,
|
||||
clean shutdown, serve start/stop around the loop; `tests/test_slack_listener.py`
|
||||
— `close()` no-op + handler teardown.
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-2 — coordinator unit loads `~/orchestrator/.env` — RESOLVED (this PR)
|
||||
|
||||
- **Location:** `agent-team/systemd/agent-team-coordinator.service`;
|
||||
`run-team.py:_build_coordinator` (always wires `default_review_wiring`);
|
||||
`coordinator.py:default_review_wiring` → `review_loop_llm.default_plan_reviewer`
|
||||
(shells the local orchestrator `run.py` → `cross_reviewer` GPT-4.1).
|
||||
- **Was:** the unit loaded only `~/secrev.env`. `run-team.py serve` builds the
|
||||
coordinator with the P2 review loop wired, and once a task reaches REVIEW the
|
||||
reviewer shells the orchestrator `run.py`, whose GPT-4.1 call reads the
|
||||
non-Claude provider key from `~/orchestrator/.env`. The review path would fail
|
||||
to authenticate.
|
||||
- **Now:** the unit adds `EnvironmentFile=-/home/adam/orchestrator/.env`
|
||||
(optional `-`, mirroring the `sea-haven-secrev` unit). Does not affect the P1
|
||||
demo (P1 stops at PLAN before REVIEW); closes the latent P2 break.
|
||||
|
||||
---
|
||||
|
||||
## 🟠 D-4 — pip install list omits `requests` (provisioning note)
|
||||
|
||||
- **Location:** `agent_team/transport/github_live.py` / `github_intake.py`
|
||||
(`import requests`).
|
||||
- **Actual:** the GitHub transport + intake require `requests`; the Slack-first
|
||||
path does not hit it, but any `--transport github` / `intake-github` use fails
|
||||
with a clear RuntimeError without it.
|
||||
- **Fix:** PROVISIONING-RUNBOOK Step 4 installs `requests` into the venv.
|
||||
`requests` is also absent from `requirements.txt` (see D-5).
|
||||
|
||||
---
|
||||
|
||||
## 🟠 D-5 — agent-team runtime deps are not pinned in `requirements.txt` (provisioning note)
|
||||
|
||||
- **Location:** `requirements.txt` (repo root).
|
||||
- **Actual:** `requirements.txt` pins `langgraph==1.1.10` and
|
||||
`langgraph-checkpoint-sqlite==3.1.0`, but the agent-team runtime deps
|
||||
`claude-agent-sdk`, `slack_sdk`, `slack_bolt`, `requests` (and `anthropic` for
|
||||
api mode) are **not in `requirements.txt` at all** — they are installed ad-hoc
|
||||
into the agent-team venv by the runbook. There is no pinned, reproducible source
|
||||
of truth for the box's runtime set.
|
||||
- **Fix (deferred):** add an `agent-team/requirements.txt` (or extras group)
|
||||
pinning these, version-matched to the root `requirements.txt` langgraph pin.
|
||||
Until then, PROVISIONING-RUNBOOK Step 4 pins `langgraph==1.1.10` /
|
||||
`langgraph-checkpoint-sqlite==3.1.0` explicitly so the unpinned `pip install`
|
||||
cannot pull a newer, untested major. **Do NOT modify `requirements.txt` or the
|
||||
checkers in this PR** (out of scope).
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-6 — `slack_bolt` is now exercised by the daemon — RESOLVED (consequence of D-1)
|
||||
|
||||
- **Location:** `slack_listener.py:serve` (the only `slack_bolt` import).
|
||||
- **Now:** with D-1 fixed, `Coordinator.serve()` starts `SlackListener.serve()`,
|
||||
which imports + uses `slack_bolt` for the Socket Mode handler. The dep is right
|
||||
and now actually exercised on the live Slack path.
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-SLACKVAR (clean) — `SLACK_CHANNEL_ID` matches
|
||||
|
||||
- `run-team.py:_build_transport` reads exactly `os.environ.get("SLACK_CHANNEL_ID")`.
|
||||
The runbook, the unit comment, and the code all use `SLACK_CHANNEL_ID` (not
|
||||
`SLACK_CHANNEL`). **CLEAN.**
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-ENV-SLACKBOT / OAUTH / OWNERS / APPTOKEN (clean) — names match
|
||||
|
||||
- **`SLACK_BOT_TOKEN`** ↔ `slack_live.py` + `slack_listener` env read. **CLEAN.**
|
||||
- **`CLAUDE_CODE_OAUTH_TOKEN`** ↔ `invoker.py`. **CLEAN.**
|
||||
- **`AGENT_TEAM_SLACK_OWNER_IDS`** ↔ `slack_listener.py` (name + fail-closed
|
||||
semantics). **CLEAN** — now read by the running daemon (D-1 fixed).
|
||||
- **`SLACK_APP_TOKEN`** — now read in two places: `coordinator._slack_listener_enabled`
|
||||
gates the listener on its presence, and `default_slack_listener_factory` /
|
||||
`SlackListener.serve` source it to open the socket. **CLEAN** (read site exists
|
||||
now that D-1 is fixed).
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-7 — `ExecStart` uses the venv interpreter — RESOLVED (this PR)
|
||||
|
||||
- **Location:** `agent-team/systemd/agent-team-coordinator.service` ExecStart.
|
||||
- **Was:** `ExecStart=/usr/bin/env python3 run-team.py serve` resolved the
|
||||
**system** interpreter under systemd's PATH — not the venv where the deps were
|
||||
installed, so the daemon would fail at import.
|
||||
- **Now:** `ExecStart=/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve`
|
||||
(matches the runbook venv path + `WorkingDirectory`).
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-SUBCMD (mostly clean) — run-team.py subcommands referenced exist
|
||||
|
||||
Cross-checked every `run-team.py <sub>` the deploy doc + demo name against
|
||||
`run-team.py:build_parser`: `init-db`, `serve`, `list` (+ `--all` / `--parked`),
|
||||
`show`, `expire`, `answer`, `redeliver`, `supersede`, `force-resume`, `start`,
|
||||
`intake-github`, `intake-checker`, `fix` — all exist. No invented verbs.
|
||||
|
||||
> ### Operator-CLI divergence (provisioning note)
|
||||
>
|
||||
> Two operator CLIs exist with **different verb names**:
|
||||
>
|
||||
> - `run-team.py` (the entry CLI): `init-db, list, show, redeliver, expire,
|
||||
> answer, supersede, force-resume, start, serve, intake-github, intake-checker,
|
||||
> fix`. Has `show` and `--parked`; `--db` / `--audit-log` default sensibly.
|
||||
> - `agent_team/operator_cli.py`: `list, redeliver, force-expire,
|
||||
> answer-on-behalf, force-resume` — **no `show`**, `--db` / `--audit-log` are
|
||||
> **required**, and its `force-resume` **supersedes** (unlike `run-team.py`'s,
|
||||
> which reopens an expired row and never supersedes).
|
||||
>
|
||||
> **Use `run-team.py` for provisioning + the demo + incident recovery.** The
|
||||
> docs reference only `run-team.py`. Reconciling the two CLIs is a follow-up.
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-PYTHONPKG (clean) — package import bootstrap is correct
|
||||
|
||||
`run-team.py` inserts its own dir into `sys.path` so the hyphenated script
|
||||
imports the `agent_team` package without an editable install. **CLEAN.**
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-HARDENING (clean, and matches the locked decision)
|
||||
|
||||
- Unit: `NoNewPrivileges=true`, `ProtectSystem=full`, `ProtectHome=read-only`,
|
||||
`ReadWritePaths=/home/adam/orchestrator/agent-team/state`. **Retained unchanged**
|
||||
in this PR (locked decision — do not revert to secrev parity).
|
||||
- The `ReadWritePaths` carve-out matches the ledger + audit-log location
|
||||
(`state/agent_team.sqlite`, `state/audit.log.jsonl`). **CLEAN.**
|
||||
|
||||
---
|
||||
|
||||
## Summary table
|
||||
|
||||
| ID | Sev | Status | One-line |
|
||||
|---|---|---|---|
|
||||
| D-1 | 🔴 | ✅ RESOLVED (PR) | `serve` starts `SlackListener` (Slack + app-token gated; Slack stays optional) |
|
||||
| D-2 | 🔴 | ✅ RESOLVED (PR) | unit loads `~/orchestrator/.env` for the P2 GPT-4.1 review provider key |
|
||||
| D-7 | 🟡 | ✅ RESOLVED (PR) | `ExecStart` points at the agent-team venv interpreter |
|
||||
| D-6 | 🟡 | ✅ RESOLVED | `slack_bolt` now exercised by the daemon (consequence of D-1) |
|
||||
| D-4 | 🟠 | NOTE | pip list omits `requests` — runbook Step 4 installs it |
|
||||
| D-5 | 🟠 | NOTE | agent-team runtime deps not pinned in `requirements.txt` — runbook pins langgraph |
|
||||
| operator-CLI | — | NOTE | `run-team.py` vs `operator_cli.py` divergent verbs — use `run-team.py` |
|
||||
| D-SLACKVAR | ✅ | CLEAN | `SLACK_CHANNEL_ID` matches everywhere |
|
||||
| D-ENV-* | ✅ | CLEAN | bot/oauth/owner/app-token env names match; all read sites now exist |
|
||||
| D-SUBCMD | ✅ | CLEAN | every `run-team.py` verb/flag the docs cite exists |
|
||||
| D-HARDENING | ✅ | CLEAN | unit hardening retained unchanged (locked decision) |
|
||||
290
docs/provisioning/OPERATOR-RUNBOOK.md
Normal file
290
docs/provisioning/OPERATOR-RUNBOOK.md
Normal file
|
|
@ -0,0 +1,290 @@
|
|||
# OPERATOR-RUNBOOK — R720 agent-team coordinator incident handling
|
||||
|
||||
The on-call runbook for the always-on `agent-team-coordinator` daemon on the
|
||||
`sh-secrev` R720 VM. Covers pipeline stalls, stuck/parked tasks, failed
|
||||
human-in-the-loop resumes, budget exhaustion mid-pipeline, transport outages, and
|
||||
the COMPLACENCY / COVERAGE alarms. Grounds every recovery in real code (design
|
||||
§5 escalation ladder + §6.6 contention/park policy; Phase-6 requirement).
|
||||
|
||||
> **CLI used throughout: `run-team.py`** (the entry CLI), run from
|
||||
> `~/orchestrator/agent-team` with the venv active so it hits the default ledger
|
||||
> (`state/agent_team.sqlite`) and audit log (`state/audit.log.jsonl`). Do **not**
|
||||
> use `agent_team/operator_cli.py` — it has divergent verbs (no `show`, required
|
||||
> `--db`/`--audit-log`, and a `force-resume` that supersedes). See DEPLOY-AUDIT.md.
|
||||
|
||||
```bash
|
||||
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
|
||||
cd ~/orchestrator/agent-team && . .venv/bin/activate
|
||||
```
|
||||
|
||||
## Verb reference (all confirmed in `run-team.py:build_parser`)
|
||||
|
||||
| Verb | Effect | Destructive? |
|
||||
|---|---|---|
|
||||
| `list` / `list --all` / `list --parked` / `list --status <state>` | read pending questions (default `open`) | no |
|
||||
| `show <question_id>` | print one ledger row (JSON) | no |
|
||||
| `redeliver <question_id>` | clear `channel_ref` so the reconcile loop re-posts an `open` question | no (audit-logged) |
|
||||
| `expire <question_id> --confirm` | force `open`→`expired` | yes |
|
||||
| `answer <question_id> --answer <p> [--via <id>] --confirm` | answer-on-behalf (first-answer-wins CAS) | yes |
|
||||
| `force-resume <question_id> --confirm` | reopen an `expired` (parked) question; for `answered` records resume intent | yes |
|
||||
| `supersede <question_id> --confirm` | mark a stale `open`/`answered` row `superseded` | yes |
|
||||
| `start --task "..." [--transport ...] [--dry-run]` | start one task to the human gate | no |
|
||||
|
||||
`--operator <name>` (global) sets the audit attribution; it defaults to the OS
|
||||
login. Destructive verbs require `--confirm` and write an attempt-then-outcome
|
||||
record to the audit log **before** mutating.
|
||||
|
||||
First triage for any incident:
|
||||
|
||||
```bash
|
||||
systemctl status agent-team-coordinator.service
|
||||
journalctl -u agent-team-coordinator.service -e --since "-2h" | tail -100
|
||||
python3 run-team.py list --all # full ledger snapshot
|
||||
python3 run-team.py list --parked # non-open rows (parked-task context)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Incident 1 — Pipeline stall (tasks not advancing)
|
||||
|
||||
**Symptoms:** `list` shows `open` (or `answered`) rows that never progress; the
|
||||
journal shows no `tick` activity or repeated errors.
|
||||
|
||||
**Diagnose:**
|
||||
```bash
|
||||
systemctl is-active agent-team-coordinator.service # "active" expected
|
||||
journalctl -u agent-team-coordinator.service -e | tail -60
|
||||
# Daemon dead/looping on restart? Check the unit + deps:
|
||||
.venv/bin/python -c "import langgraph, slack_sdk, slack_bolt, requests; print('deps ok')"
|
||||
```
|
||||
|
||||
**Recover:**
|
||||
1. If the daemon is dead and `Restart=on-failure` is flapping, read the journal
|
||||
for the import/auth error. A missing venv dep (D-4/D-5) or a missing
|
||||
`CLAUDE_CODE_OAUTH_TOKEN` is the usual cause — fix the env / venv, then:
|
||||
```bash
|
||||
sudo systemctl restart agent-team-coordinator.service
|
||||
```
|
||||
2. The startup `recover()` sweep re-drives `answered`-but-unresumed rows and
|
||||
re-posts `open` rows that lost their `channel_ref`, so a clean restart
|
||||
converges from the durable ledger. Confirm with `journalctl ... | tail` and
|
||||
`list --all`.
|
||||
3. A single task stuck `open` with a stale/missing post: re-deliver it.
|
||||
```bash
|
||||
python3 run-team.py show <qid>
|
||||
python3 run-team.py redeliver <qid> # clears channel_ref; reconcile re-posts
|
||||
```
|
||||
|
||||
**Escalate** (see the ladder) if a restart does not clear it within one tick
|
||||
cadence and the journal shows a non-transient error.
|
||||
|
||||
---
|
||||
|
||||
## Incident 2 — Stuck / parked task
|
||||
|
||||
A task parks (design §6.6) when its clarifier question **expires** with no answer,
|
||||
or its budget headroom drops below reserve, or a checkpoint is corrupt. Parked =
|
||||
the durable ledger row is no longer `open` (it is `expired`), and an ALARM was
|
||||
raised, not spun on.
|
||||
|
||||
**Diagnose:**
|
||||
```bash
|
||||
python3 run-team.py list --parked
|
||||
python3 run-team.py show <qid> # status, thread_id, deadline_at, answered_via
|
||||
```
|
||||
|
||||
**Recover — depends on why it parked:**
|
||||
|
||||
- **Expired with no answer (the common case)** — un-park by reopening the expired
|
||||
question; it is then re-delivered for an answer:
|
||||
```bash
|
||||
python3 run-team.py force-resume <qid> --confirm
|
||||
# "force-resume: reopened expired question <qid>; it will be re-delivered"
|
||||
python3 run-team.py show <qid> # status flips back to "open"
|
||||
```
|
||||
Then answer it (Slack or CLI) to drive it forward.
|
||||
|
||||
- **Answered but not yet resumed** — the recovery sweep handles it; `force-resume`
|
||||
records intent and reports that (no mutation):
|
||||
```bash
|
||||
python3 run-team.py force-resume <qid> --confirm
|
||||
# "force-resume: question <qid> is answered and pending resume; the recovery
|
||||
# sweep will resume it (intent recorded)"
|
||||
sudo systemctl restart agent-team-coordinator.service # forces the recover() sweep now
|
||||
```
|
||||
|
||||
- **Stale / wrong question that should be abandoned** — supersede it so it stops
|
||||
surfacing as parked context:
|
||||
```bash
|
||||
python3 run-team.py supersede <qid> --confirm
|
||||
```
|
||||
|
||||
`MAX_PARK` FIFO-aging: a task that exceeds the park window escalates (ALARM + a
|
||||
Jira ticket per the ladder) rather than starving silently.
|
||||
|
||||
---
|
||||
|
||||
## Incident 3 — Failed human-in-the-loop resume
|
||||
|
||||
**Symptoms:** an answer was submitted (Slack or CLI) but the graph did not
|
||||
advance.
|
||||
|
||||
**Diagnose:**
|
||||
```bash
|
||||
python3 run-team.py show <qid> # is status "answered"? what answered_via?
|
||||
journalctl -u agent-team-coordinator.service -e | grep -iE "resume|answer|<qid>"
|
||||
```
|
||||
|
||||
**Recover:**
|
||||
- **Status is `answered` but no resume drained** — the resume queue is in-process;
|
||||
a daemon restart triggers the `recover()` sweep that re-drives `answered` rows
|
||||
via the turn-guarded `ResumeWorker` (idempotent — an already-advanced thread
|
||||
supersedes-and-skips):
|
||||
```bash
|
||||
sudo systemctl restart agent-team-coordinator.service
|
||||
python3 run-team.py show <qid> # confirm it advanced
|
||||
```
|
||||
- **Live Slack answer never registered** — the listener fails closed. Check:
|
||||
```bash
|
||||
journalctl -u agent-team-coordinator.service -e | grep -i "Slack answer"
|
||||
# "rejecting Slack answer: owner allowlist is unconfigured ..." -> set
|
||||
# AGENT_TEAM_SLACK_OWNER_IDS in ~/secrev.env and restart.
|
||||
# "rejecting Slack answer ... unauthorized sender" -> the answerer's user id is
|
||||
# not in the allowlist. Add it, or answer via the CLI on their behalf:
|
||||
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
|
||||
```
|
||||
- **Answer lost the compare-and-set (`not open`)** — the row was already
|
||||
answered/expired/superseded. Inspect with `show`; if it parked, go to Incident 2.
|
||||
|
||||
---
|
||||
|
||||
## Incident 4 — Budget exhaustion mid-pipeline
|
||||
|
||||
The shared Claude budget ledger (`budget_ledger` table) enforces per-call + total
|
||||
nightly caps (design §6.1, §6.6). A stage that would breach the reserve **parks**
|
||||
the task (deferred, not dropped) and ALARMs — it never loop-drains the pool.
|
||||
|
||||
**Diagnose:**
|
||||
```bash
|
||||
journalctl -u agent-team-coordinator.service -e | grep -iE "budget|reserve|park"
|
||||
# Inspect today's spend directly (no CLI verb for the budget ledger; read it).
|
||||
# Columns (agent_team/db/schema.py budget_ledger DDL): thread_id, stage, model,
|
||||
# billing_mode, input_tokens, output_tokens, usd_cost, recorded_at, day_bucket.
|
||||
sqlite3 state/agent_team.sqlite \
|
||||
"SELECT day_bucket, thread_id, ROUND(SUM(usd_cost),4) AS usd, SUM(input_tokens) AS in_tok
|
||||
FROM budget_ledger GROUP BY day_bucket, thread_id ORDER BY day_bucket DESC LIMIT 20;"
|
||||
```
|
||||
|
||||
**Recover:**
|
||||
- **Wait for the next budget window** — budget-exhausted roles/tasks are deferred
|
||||
via the rotation pointer and picked up next cycle; this is the intended
|
||||
behavior, not a failure. The parked task surfaces in `list --parked`.
|
||||
- **Force a specific parked task forward now** (e.g. it is urgent and headroom has
|
||||
since freed): un-park it and let the daemon re-run the stage within the
|
||||
remaining cap:
|
||||
```bash
|
||||
python3 run-team.py force-resume <qid> --confirm
|
||||
```
|
||||
- **Confirm `ANTHROPIC_API_KEY` is absent** — its presence would silently meter to
|
||||
API rates and blow the budget model (the billing seam pops it defensively, but
|
||||
it must not be set):
|
||||
```bash
|
||||
grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # both must print 0
|
||||
```
|
||||
|
||||
Do **not** raise the cap to push a task through without Adam's decision — budget
|
||||
realism is a deliberate guardrail. Escalate per the ladder if a task repeatedly
|
||||
parks on budget across nights.
|
||||
|
||||
---
|
||||
|
||||
## Incident 5 — Transport outage (Slack / GitHub down or misconfigured)
|
||||
|
||||
**Symptoms:** clarifier posts fail; the journal shows transport errors or the
|
||||
inbound listener is not started.
|
||||
|
||||
**Diagnose:**
|
||||
```bash
|
||||
journalctl -u agent-team-coordinator.service -e | grep -iE "Slack|listener|post|transport"
|
||||
# Did the inbound listener start?
|
||||
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
|
||||
# "started (Socket Mode ...)" -> inbound up
|
||||
# "not started (... SLACK_APP_TOKEN is unset)" -> Slack inbound intentionally off
|
||||
```
|
||||
|
||||
**Recover:**
|
||||
- **Outbound post failed (Slack API / network)** — the ledger row stays `open`
|
||||
with no `channel_ref` (the responder leaves it for reconcile). After the
|
||||
transport recovers, the `recover()` sweep (on restart) or a manual `redeliver`
|
||||
re-posts:
|
||||
```bash
|
||||
python3 run-team.py redeliver <qid>
|
||||
```
|
||||
- **Inbound listener down / never started** — the daemon still posts and expires;
|
||||
it just cannot hear Slack. **The CLI `answer` path is the outage fallback** —
|
||||
it runs the identical compare-and-set with no socket:
|
||||
```bash
|
||||
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
|
||||
```
|
||||
Fix `SLACK_APP_TOKEN` / `SLACK_BOT_TOKEN` in `~/secrev.env`, then
|
||||
`sudo systemctl restart agent-team-coordinator.service` to re-establish Socket
|
||||
Mode. (D-1: the listener starts only when the transport is live Slack AND
|
||||
`SLACK_APP_TOKEN` is set; it is optional by design.)
|
||||
- **Slack listener crash-looping** — the listener runs on an isolated daemon
|
||||
thread; a crash is logged (`inbound Slack listener thread exited with an error`)
|
||||
and takes down only the inbound socket, not the maintenance loop. Restart the
|
||||
service to relaunch the thread once the cause is fixed.
|
||||
|
||||
---
|
||||
|
||||
## Incident 6 — COMPLACENCY / COVERAGE alarms (Plane-1 checkers)
|
||||
|
||||
These come from the nightly checker run, not the coordinator daemon (design §6.4,
|
||||
§6.6):
|
||||
|
||||
- **COMPLACENCY ALARM** — a checker role missed a planted canary fault. That role
|
||||
is **skipped** for the night (it never runs silently degraded). Recover: inspect
|
||||
the canary corpus / the role's checker under `security-review/checkers/`, fix the
|
||||
regression, re-run that checker's dry-run, and confirm the canary passes before
|
||||
re-enabling.
|
||||
- **COVERAGE ALARM** — a role slipped its rotation slot (e.g. budget-deferred). It
|
||||
is **deferred via the rotation pointer, never dropped**, and picked up next
|
||||
cycle. Recover: confirm the rotation pointer advanced (it is rebuildable from
|
||||
report history) and that the role runs on the next cadence; investigate only if
|
||||
it slips repeatedly.
|
||||
|
||||
Both alarms are **report + ALARM-only** (design D3): nothing posts on a clean
|
||||
state, no auto-Jira/Notion writes from the checker itself. Persistent alarms
|
||||
follow the escalation ladder below.
|
||||
|
||||
---
|
||||
|
||||
## Escalation ladder (design §5 / §6.6, resolves Q3)
|
||||
|
||||
Anything that does not clear on the first ALARM escalates — but **ALARM-only in
|
||||
spirit** (nothing posts on a clean state):
|
||||
|
||||
1. **Re-alarm on a backoff.** A confirmed critical (or a COMPLACENCY / COVERAGE
|
||||
alarm) that persists re-alarms to Slack each night it is still unresolved, on
|
||||
a backoff so it does not spam.
|
||||
2. **Open a Jira tracking ticket after `N` nights** (default **N = 3**). If the
|
||||
condition still has not cleared, the coordinator opens an **INFRA** Jira ticket
|
||||
so it cannot quietly linger. The same ladder applies to a parked task that
|
||||
exceeds `MAX_PARK`.
|
||||
3. **Human (Adam) takes it from the Jira ticket.** For a parked task, recover via
|
||||
Incidents 2–4 above; for a checker alarm, via Incident 6.
|
||||
|
||||
When you resolve an incident, record the action — destructive CLI verbs already
|
||||
write an attributable attempt+outcome record to `state/audit.log.jsonl`; for
|
||||
non-CLI recoveries note it on the Jira ticket.
|
||||
|
||||
---
|
||||
|
||||
## Post-incident
|
||||
|
||||
- Confirm the daemon is `active` and the ledger has no unexpected parked rows
|
||||
(`list --parked`).
|
||||
- If you restored the VM snapshot or wiped the ledger, re-run the relevant
|
||||
PROVISIONING-RUNBOOK steps.
|
||||
- Update `project_r720_agent_team` memory if the incident revealed a durable
|
||||
fact (a new failure mode, a config that must change).
|
||||
286
docs/provisioning/P1-DEMO-SCRIPT.md
Normal file
286
docs/provisioning/P1-DEMO-SCRIPT.md
Normal file
|
|
@ -0,0 +1,286 @@
|
|||
# P1-DEMO-SCRIPT — live four-criteria acceptance demo (design §3.3.1 / §7.1 P1)
|
||||
|
||||
The §7.1 P1 exit gate: demonstrate, on the live box, all four durable
|
||||
human-in-the-loop criteria before P1 is accepted:
|
||||
|
||||
- **(a)** kill the box mid-wait and have the task resume after restart;
|
||||
- **(b)** submit a duplicate answer and confirm it no-ops;
|
||||
- **(c)** submit an answer after the deadline expired and confirm it is rejected
|
||||
and the task parks;
|
||||
- **(d)** two tasks suspended concurrently resume independently to the correct
|
||||
thread.
|
||||
|
||||
Every command below is grounded in the **actual** code surface
|
||||
(`run-team.py`, `coordinator.py`, `slack_listener.py`, `responder.py`,
|
||||
`resume_worker.py`, `db/schema.py`, `graph.py`). No invented flags. Where the
|
||||
code does not expose a needed knob (e.g. a short deadline), the script uses a
|
||||
direct `sqlite3` write against the documented `pending_questions` schema and says
|
||||
so.
|
||||
|
||||
## Two answer paths — live Slack OR the operator CLI
|
||||
|
||||
**D-1 is fixed:** `Coordinator.serve()` now starts the inbound `SlackListener`
|
||||
when the live transport is Slack AND `SLACK_APP_TOKEN` is set, so the
|
||||
**live Slack answer round-trip works**. You can run the demo either way:
|
||||
|
||||
- **Live Slack** — Adam clicks the Block Kit button / replies in the channel; the
|
||||
listener normalizes the event, runs the AUTHZ-01 owner check, and drives the
|
||||
first-answer-wins compare-and-set (`responder.submit_answer` →
|
||||
`db.schema.answer_question`).
|
||||
- **Operator CLI** — `run-team.py answer <qid> --answer ... --confirm` runs the
|
||||
**identical** compare-and-set (audit-logged answer-on-behalf). Useful when the
|
||||
Slack app is not yet provisioned, or to script the assertions.
|
||||
|
||||
Both exercise the same durable mechanic; the human gate decision is always
|
||||
Adam's. The assertions below assert on the **durable ledger status** (the §3.3.1
|
||||
source of truth) and are identical for either path. The examples use the CLI
|
||||
`answer --confirm` form so they are copy-pasteable; substitute "Adam answers in
|
||||
Slack" wherever you prefer the live path.
|
||||
|
||||
> The automated proof of this mechanic is `tests/sim/test_p1_exit_criteria.py`
|
||||
> (a `SimPipeline` harness) and `tests/test_coordinator.py` (the serve/listener
|
||||
> wiring). This script is the live-box demonstration on top of that.
|
||||
|
||||
## Verb map (use `run-team.py`, not `operator_cli.py`)
|
||||
|
||||
| Need | `run-team.py` verb | Notes |
|
||||
|---|---|---|
|
||||
| start a task to the human gate | `start --task "..." [--transport slack] [--dry-run]` | mints a `thread_id`, posts the clarifier, writes the `open` ledger row |
|
||||
| list waiting questions | `list` (default `open`) / `list --all` / `list --parked` | JSON rows |
|
||||
| inspect one row | `show <question_id>` | JSON row |
|
||||
| answer on the task's behalf | `answer <question_id> --answer <payload> --confirm` | destructive, audit-logged; first-answer-wins CAS |
|
||||
| force-expire an open question | `expire <question_id> --confirm` | destructive; flips `open`→`expired` |
|
||||
| un-park (reopen) an expired question | `force-resume <question_id> --confirm` | only acts on `expired` rows (reopens them) |
|
||||
|
||||
`operator_cli.py` has different verbs (`force-expire`, `answer-on-behalf`) and
|
||||
**no `show`**, and requires `--db`/`--audit-log` — do not use it here.
|
||||
|
||||
## Preconditions
|
||||
|
||||
```bash
|
||||
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
|
||||
cd ~/orchestrator/agent-team
|
||||
. .venv/bin/activate # so run-team.py uses the venv deps
|
||||
# Confirm the ledger exists (Step 5 of the runbook):
|
||||
python3 run-team.py list --all # [] on a fresh DB is fine
|
||||
# For the live-Slack path, confirm the listener started:
|
||||
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
|
||||
# -> "inbound Slack listener started (Socket Mode, background thread)"
|
||||
```
|
||||
|
||||
Run all `run-team.py` commands from `~/orchestrator/agent-team` so they hit the
|
||||
default ledger (`state/agent_team.sqlite`) and default audit log
|
||||
(`state/audit.log.jsonl`). The clarifier is posted to `SLACK_CHANNEL_ID`.
|
||||
|
||||
> **Resume execution model.** A won answer (rowcount 1) enqueues a `ResumeJob`
|
||||
> onto the coordinator's in-process `resume_queue`; the daemon drains it on the
|
||||
> next `tick()`, or the startup `recover()` re-drives any `answered`-but-unresumed
|
||||
> row after a restart. So a CLI `answer --confirm` (or a live Slack answer) flips
|
||||
> the ledger row to `answered`; the **running daemon** then resumes the graph.
|
||||
> The demo asserts on the durable ledger status.
|
||||
|
||||
---
|
||||
|
||||
## (a) Crash-safe resume — kill mid-wait, restart, task resumes 🧑 Adam answers
|
||||
|
||||
**Setup — drive a task to the clarifier wait.** With the daemon already running:
|
||||
|
||||
```bash
|
||||
python3 run-team.py start --task "demo-a: trivial scoped task"
|
||||
# prints a thread_id, e.g. 3f2a... (record it as $TID_A)
|
||||
python3 run-team.py list
|
||||
# expect one row: status="open", a thread_id, a question_id (record as $QID_A),
|
||||
# channel_ref set (Slack ts) or null if posting is dry/unavailable.
|
||||
```
|
||||
|
||||
**Kill the box mid-wait, then restart:**
|
||||
|
||||
```bash
|
||||
sudo systemctl stop agent-team-coordinator.service
|
||||
# (optionally reboot the VM here for a stronger demonstration)
|
||||
sudo systemctl start agent-team-coordinator.service
|
||||
journalctl -u agent-team-coordinator.service -e | tail -40 # expect "starting" + recover() sweep
|
||||
```
|
||||
|
||||
**ASSERT — durable state survived the kill (no answer yet):**
|
||||
|
||||
```bash
|
||||
python3 run-team.py show $QID_A
|
||||
# PASS: status == "open" (the question was NOT lost across the restart)
|
||||
```
|
||||
|
||||
**🧑 Human gate — Adam answers after the restart (Slack OR CLI):**
|
||||
|
||||
```bash
|
||||
# Live Slack: Adam replies/clicks in the channel. OR via the CLI:
|
||||
python3 run-team.py answer $QID_A --answer "scope: just demo, no real change" --confirm
|
||||
python3 run-team.py show $QID_A
|
||||
# PASS: status == "answered"
|
||||
# Wait one tick (~30s, DEFAULT_POLL_INTERVAL) for the daemon to drain the resume:
|
||||
journalctl -u agent-team-coordinator.service -e | tail -20
|
||||
```
|
||||
|
||||
**PASS criterion (a):** the question stayed `open` across the kill, and an answer
|
||||
submitted *after* the restart drove it forward.
|
||||
|
||||
---
|
||||
|
||||
## (b) Duplicate answer is a no-op 🧑 Adam answers
|
||||
|
||||
```bash
|
||||
python3 run-team.py start --task "demo-b: duplicate-answer test"
|
||||
python3 run-team.py list # record the new question_id as $QID_B (status open)
|
||||
```
|
||||
|
||||
**🧑 First answer (the winning one) — Slack OR CLI:**
|
||||
|
||||
```bash
|
||||
python3 run-team.py answer $QID_B --answer "first-answer" --via "slack:U1" --confirm
|
||||
# expect: "answered question <QID_B> (via slack:U1)" (exit 0)
|
||||
```
|
||||
|
||||
**Duplicate / second answer (must lose the compare-and-set):**
|
||||
|
||||
```bash
|
||||
python3 run-team.py answer $QID_B --answer "second-answer" --via "github:U2" --confirm
|
||||
# expect: "answer no-op: question <QID_B> was not 'open' ..." on stderr, exit 1
|
||||
```
|
||||
|
||||
For the **live-Slack** variant, Adam (or a second authorized owner) answers the
|
||||
same message twice; the second event loses the CAS and is logged
|
||||
`ignored Slack answer for question_id=... (not open: duplicate ...)`.
|
||||
|
||||
**ASSERT — the first answer is preserved:**
|
||||
|
||||
```bash
|
||||
python3 run-team.py show $QID_B
|
||||
# PASS: status == "answered"; answered_via == "slack:U1" (the FIRST answer)
|
||||
```
|
||||
|
||||
**PASS criterion (b):** first-answer-wins (`rowcount==1`); the duplicate hits the
|
||||
`BEGIN IMMEDIATE` compare-and-set in `answer_question` and is ignored
|
||||
(`rowcount==0`) — no second resume, no overwrite.
|
||||
|
||||
---
|
||||
|
||||
## (c) Past-deadline answer rejected + task parks 🧑 Adam observes
|
||||
|
||||
> **Code reality:** `run-team.py start` always sets the clarifier deadline from
|
||||
> `DEFAULT_CLARIFY_DEADLINE = 24h`. There is **no CLI flag for a short deadline**.
|
||||
> Two faithful ways to demo expiry without waiting 24h:
|
||||
|
||||
### Option C1 — force the deadline past, let the daemon's sweep expire it (most faithful)
|
||||
|
||||
```bash
|
||||
python3 run-team.py start --task "demo-c: deadline test"
|
||||
python3 run-team.py list # record $QID_C (status open)
|
||||
|
||||
# Set this question's deadline into the past directly in the ledger:
|
||||
sqlite3 state/agent_team.sqlite \
|
||||
"UPDATE pending_questions SET deadline_at='2000-01-01T00:00:00+00:00' WHERE question_id='$QID_C';"
|
||||
|
||||
# Wait one daemon tick (~30s) for the deadline sweep to flip it, OR observe:
|
||||
journalctl -u agent-team-coordinator.service -e | tail -20
|
||||
# expect the park ALARM line: "task parked: clarifier question <QID_C> expired ..."
|
||||
```
|
||||
|
||||
### Option C2 — operator force-expire (if you do not want to touch the DB)
|
||||
|
||||
```bash
|
||||
python3 run-team.py start --task "demo-c: deadline test"
|
||||
python3 run-team.py list # record $QID_C
|
||||
python3 run-team.py expire $QID_C --confirm # destructive, audit-logged; open->expired
|
||||
```
|
||||
|
||||
**ASSERT — the question is expired and the late answer is rejected:**
|
||||
|
||||
```bash
|
||||
python3 run-team.py show $QID_C
|
||||
# PASS: status == "expired"
|
||||
|
||||
# 🧑 Adam submits a LATE answer (Slack OR CLI) — it must lose the CAS:
|
||||
python3 run-team.py answer $QID_C --answer "too-late" --confirm
|
||||
# expect: "answer no-op: question <QID_C> was not 'open' ..." stderr, exit 1
|
||||
python3 run-team.py show $QID_C
|
||||
# PASS: status still "expired"; answer_json still NULL
|
||||
|
||||
# The parked task surfaces in the parked view:
|
||||
python3 run-team.py list --parked # PASS: $QID_C appears here
|
||||
```
|
||||
|
||||
**Deliberate un-park (proves the operator recovery path, §6.6):**
|
||||
|
||||
```bash
|
||||
python3 run-team.py force-resume $QID_C --confirm
|
||||
# expect: "force-resume: reopened expired question <QID_C>; it will be re-delivered"
|
||||
python3 run-team.py show $QID_C
|
||||
# status flips back to "open" (reopen_question), deadline_at cleared.
|
||||
```
|
||||
|
||||
**PASS criterion (c):** the expired question rejects the late answer, the task
|
||||
parks rather than spins, and the operator can deliberately un-park it via
|
||||
`force-resume` (which reopens only an `expired` row).
|
||||
|
||||
> **`force-resume` semantics (verified):** `run-team.py force-resume` reopens an
|
||||
> `expired` question (the parked case). For an `answered` question it records
|
||||
> intent and reports the recovery sweep will resume it (no mutation). For
|
||||
> `open`/`superseded`/absent it is a no-op exit 1. It does **not** supersede.
|
||||
|
||||
---
|
||||
|
||||
## (d) Two concurrent tasks resume independently 🧑 Adam answers
|
||||
|
||||
```bash
|
||||
python3 run-team.py start --task "demo-d task A" # record $TID_A2, then:
|
||||
python3 run-team.py start --task "demo-d task B" # record $TID_B2
|
||||
python3 run-team.py list
|
||||
# expect TWO open rows with DISTINCT thread_id AND distinct question_id.
|
||||
# Record $QID_A2 and $QID_B2 — match by thread_id.
|
||||
```
|
||||
|
||||
**(Optional) restart the daemon first** to also show concurrent tasks survive a
|
||||
restart, then answer.
|
||||
|
||||
**🧑 Answer the SECOND task first, with a distinct answer, then the first
|
||||
(Slack OR CLI):**
|
||||
|
||||
```bash
|
||||
python3 run-team.py answer $QID_B2 --answer "answer-for-B" --via "slack:U2" --confirm
|
||||
python3 run-team.py answer $QID_A2 --answer "answer-for-A" --via "slack:U1" --confirm
|
||||
# Wait one daemon tick (~30s) for both resumes to drain.
|
||||
```
|
||||
|
||||
**ASSERT — each task carries its OWN answer; no cross-talk:**
|
||||
|
||||
```bash
|
||||
python3 run-team.py show $QID_A2
|
||||
# PASS: status "answered", answered_via "slack:U1", thread_id == $TID_A2
|
||||
python3 run-team.py show $QID_B2
|
||||
# PASS: status "answered", answered_via "slack:U2", thread_id == $TID_B2
|
||||
python3 run-team.py list --all
|
||||
# PASS: the two rows resolved on their own thread_id; no cross-contamination.
|
||||
```
|
||||
|
||||
**PASS criterion (d):** two concurrently-suspended tasks each resumed to their
|
||||
own `thread_id` with their own answer — answering B before A did not misroute,
|
||||
and the per-thread single-flight guard kept them independent.
|
||||
|
||||
---
|
||||
|
||||
## Acceptance
|
||||
|
||||
P1 is accepted only when **(a), (b), (c), and (d) all pass** on the live box.
|
||||
Record the four `show` outputs (or `journalctl` excerpts) as evidence. Then
|
||||
complete the PROVISIONING-RUNBOOK post-session definition-of-done (memory +
|
||||
Confluence + the mandatory `/sh-security-review` on the Slack inbound listener).
|
||||
|
||||
## What can only be verified on the live box
|
||||
|
||||
- That the daemon actually drains the resume and the LangGraph checkpoint
|
||||
advances (asserted via ledger status + journalctl; the graph-state advance is
|
||||
observable only on the box).
|
||||
- The live **Slack** post/answer round-trip end-to-end (the listener is wired —
|
||||
D-1 fixed — but the live Socket Mode socket + a real `SLACK_BOT_TOKEN` /
|
||||
`SLACK_APP_TOKEN` / channel membership are only present on the box).
|
||||
- The exact `channel_ref` value (Slack message `ts`) — depends on a live Slack
|
||||
post succeeding.
|
||||
402
docs/provisioning/PROVISIONING-RUNBOOK.md
Normal file
402
docs/provisioning/PROVISIONING-RUNBOOK.md
Normal file
|
|
@ -0,0 +1,402 @@
|
|||
# PROVISIONING RUNBOOK — R720 agent-team Plane-2 coordinator
|
||||
|
||||
This is the ordered command sequence for the **operator-present** provisioning
|
||||
session that stands up the always-on `agent-team-coordinator` daemon on the
|
||||
`sh-secrev` R720 VM. Derived from `agent-team/DEPLOY-R720.md`,
|
||||
`security-review/DEPLOY-R720.md`, and `docs/r720-agent-team-design.md`
|
||||
(§3.3.1 / §7 / §7.1).
|
||||
|
||||
> ## State of this runbook
|
||||
>
|
||||
> The build is **merged and final**: all six Plane-1 checkers
|
||||
> (`aws-posture`, `compliance-drift`, `confluence-doc`, `dependency-cve`,
|
||||
> `doc-drift`, `plan-groomer`), the Tier-3 dep-bump **fixer**
|
||||
> (`run-team.py fix --dry-run`), and the Plane-1→Plane-2 **P5 cross-plane loop**
|
||||
> (`run-team.py intake-checker`) exist in the tree. The three deploy-correctness
|
||||
> bugs the provisioning-prep audit found are **FIXED** (see DEPLOY-AUDIT.md):
|
||||
>
|
||||
> - **D-1 RESOLVED** — `Coordinator.serve()` now starts the inbound Slack
|
||||
> `SlackListener` concurrently with the tick/drain loop when Slack is the live
|
||||
> transport AND `SLACK_APP_TOKEN` is set. **The demo can use the live Slack
|
||||
> answer path** (P1-DEMO-SCRIPT.md), or the operator-CLI `answer` path.
|
||||
> - **D-2 RESOLVED** — the systemd unit loads `~/orchestrator/.env` (for the P2
|
||||
> GPT-4.1 review loop's provider key) in addition to `~/secrev.env`.
|
||||
> - **D-7 RESOLVED** — the unit's `ExecStart` points at the agent-team venv
|
||||
> interpreter, not the system `python3`.
|
||||
>
|
||||
> Still open as provisioning notes (not blockers): **D-4/D-5** (agent-team
|
||||
> runtime deps are installed ad-hoc into the venv and are not pinned in
|
||||
> `requirements.txt`), and the **operator-CLI divergence** (`run-team.py` vs
|
||||
> `agent_team/operator_cli.py` have different verb names — use `run-team.py`).
|
||||
|
||||
---
|
||||
|
||||
## Scope and ground rules
|
||||
|
||||
**Hard rules (carried from the design and global instructions):**
|
||||
|
||||
- **Snapshot before any stateful change** (design §7; `feedback_ec2_replacement_snapshot`).
|
||||
- **Every stateful step has an exercised rollback.**
|
||||
- Secrets use **placeholder names only**; real values are entered by the operator
|
||||
at the box and never echoed into shell history.
|
||||
- The box is **read-only / subscription-OAuth only**. **No `ANTHROPIC_API_KEY`**
|
||||
on this host (it would silently win over OAuth and meter to API rates —
|
||||
`billing.claude_invoke` pops it defensively, but it must not be present).
|
||||
- **No IAM / OIDC is involved in the coordinator deploy** (P1/P2). IAM enters
|
||||
only at the **P3-live flip** (see the dedicated section), which is gated on the
|
||||
mandatory GPT-4.1 cross-review + `/sh-security-review`.
|
||||
|
||||
**Legend per step:**
|
||||
|
||||
- 🧑 **OPERATOR-REQUIRED** — needs the human (snapshot, secrets, the live human
|
||||
gate, go/no-go). Cannot be automated.
|
||||
- 🤖 **MECHANICAL** — deterministic; an operator runs it but it needs no judgment.
|
||||
|
||||
## Host facts (from both DEPLOY-R720.md files)
|
||||
|
||||
- Hypervisor: R720 at `10.10.60.40` (Windows Server 2022, Hyper-V).
|
||||
- VM: `sh-secrev`, Ubuntu 24.04, **4GB / 2 vCPU / 40GB** dynamic vhdx, `10.10.60.120`.
|
||||
- Reach: `ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120` (key-only, NOPASSWD sudo).
|
||||
- Repo on the box: `~/orchestrator/` (rsync from the Mac, **NOT** a git clone).
|
||||
The package lives at `~/orchestrator/agent-team/`.
|
||||
- Shares `~/secrev.env` (mode 600) with the secrev sweep, and `~/orchestrator/.env`
|
||||
(mode 600) for non-Claude provider keys (same files the secrev unit loads).
|
||||
|
||||
---
|
||||
|
||||
## STEP 1 — Snapshot the VM 🧑 OPERATOR-REQUIRED
|
||||
|
||||
**On the R720 host (Hyper-V), before anything else.** This is the one-command
|
||||
undo for every change below.
|
||||
|
||||
```powershell
|
||||
# On the R720 Windows host (PowerShell, as admin):
|
||||
Checkpoint-VM -Name sh-secrev -SnapshotName "pre-agent-team-coordinator-$(Get-Date -Format yyyyMMdd-HHmm)"
|
||||
Get-VMSnapshot -VMName sh-secrev # confirm the checkpoint exists
|
||||
```
|
||||
|
||||
**ROLLBACK (whole session):**
|
||||
```powershell
|
||||
Restore-VMSnapshot -VMName sh-secrev -Name "<the checkpoint name above>" -Confirm:$false
|
||||
Start-VM -Name sh-secrev
|
||||
```
|
||||
|
||||
> Do not proceed until the checkpoint is confirmed present.
|
||||
|
||||
---
|
||||
|
||||
## STEP 2 — Rsync the repo to the box 🤖 MECHANICAL (verify manifest 🧑)
|
||||
|
||||
**From the Mac.** Same pattern/excludes as the secrev deploy. The team **scans
|
||||
the same `~/repo-mirrors` corpus secrev already maintains** — this rsync ships
|
||||
*code*, not mirrors.
|
||||
|
||||
```bash
|
||||
# From the Mac (sync the canonical repo, not a worktree):
|
||||
rsync -av --exclude .env --exclude .venv --exclude .git --exclude .claude \
|
||||
~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/
|
||||
```
|
||||
|
||||
Surface that must be on the box:
|
||||
|
||||
| Path (under `~/orchestrator/`) | Why it must ship |
|
||||
|---|---|
|
||||
| `agent-team/run-team.py` | the operator entry CLI |
|
||||
| `agent-team/agent_team/**` | the package (coordinator, graph, ledger, transports, slack_listener, fixer) |
|
||||
| `agent-team/systemd/agent-team-coordinator.service` | the daemon unit |
|
||||
| `run.py` + the orchestrator package | the GPT-4.1 cross_reviewer the P2 review loop shells |
|
||||
| `requirements.txt` | pin reference for langgraph / checkpoint-sqlite |
|
||||
| `security-review/lib/**` | shared sweep substrate (Phase 0) |
|
||||
| `security-review/checkers/**` | the six Plane-1 checkers + fixtures |
|
||||
|
||||
**ROLLBACK:** rsync is additive; restore the Step-1 snapshot to revert code state.
|
||||
|
||||
---
|
||||
|
||||
## STEP 3 — Write secrets 🧑 OPERATOR-REQUIRED
|
||||
|
||||
**On the VM.** The coordinator unit loads **two** EnvironmentFiles (both
|
||||
optional via the leading `-`): `~/secrev.env` (agent-team runtime keys) and
|
||||
`~/orchestrator/.env` (the non-Claude provider key for the P2 review loop).
|
||||
**Placeholder names only — the operator pastes real values.** Do not echo real
|
||||
tokens into shell history (use an editor or `read -s`).
|
||||
|
||||
`~/secrev.env` (mode 600) — append the agent-team keys:
|
||||
|
||||
```bash
|
||||
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
|
||||
# Edit ~/secrev.env (mode 600) and add — values are placeholders:
|
||||
# CLAUDE_CODE_OAUTH_TOKEN=<CLAUDE_CODE_OAUTH_TOKEN> # from `claude setup-token`
|
||||
# SLACK_BOT_TOKEN=<SLACK_BOT_TOKEN> # xoxb-..., chat:write (posts questions)
|
||||
# SLACK_APP_TOKEN=<SLACK_APP_TOKEN> # xapp-..., connections:write (Socket Mode inbound)
|
||||
# SLACK_CHANNEL_ID=<SLACK_CHANNEL_ID> # C0..., target clarifier channel
|
||||
# AGENT_TEAM_SLACK_OWNER_IDS=<SLACK_OWNER_USER_IDS> # comma-separated U... ids of authorized answerers
|
||||
chmod 600 ~/secrev.env
|
||||
```
|
||||
|
||||
`~/orchestrator/.env` (mode 600) — the non-Claude provider key for the GPT-4.1
|
||||
review loop (same file secrev uses; if it already exists with the key, leave it):
|
||||
|
||||
```bash
|
||||
# OPENAI_API_KEY=<OPENAI_API_KEY> # (or the provider key cross_reviewer/GPT-4.1 needs)
|
||||
chmod 600 ~/orchestrator/.env
|
||||
```
|
||||
|
||||
**Contract notes (verified against the code — see DEPLOY-AUDIT.md):**
|
||||
|
||||
- `SLACK_CHANNEL_ID` is correct — `run-team.py _build_transport` reads exactly
|
||||
`os.environ.get("SLACK_CHANNEL_ID")`. Do **not** use `SLACK_CHANNEL`.
|
||||
- `SLACK_APP_TOKEN` is now **read by the daemon**: `Coordinator.serve()` starts
|
||||
the inbound `SlackListener` when the transport is live Slack AND `SLACK_APP_TOKEN`
|
||||
is set. Without it the daemon still runs (posts + expires) but never hears Slack
|
||||
replies — Slack stays optional by design.
|
||||
- `AGENT_TEAM_SLACK_OWNER_IDS` **fails closed** (AUTHZ-01): if unset/empty the
|
||||
listener rejects **every** answer. It must be set for the live human gate.
|
||||
- **CRITICAL:** confirm `ANTHROPIC_API_KEY` is NOT present:
|
||||
```bash
|
||||
grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # must print 0 for both
|
||||
```
|
||||
|
||||
**ROLLBACK:** strip exactly the appended keys (keep secrev keys intact), or
|
||||
restore the Step-1 snapshot. Do NOT blindly truncate — secrev keys live here too.
|
||||
|
||||
---
|
||||
|
||||
## STEP 4 — Create the venv + install pip deps 🤖 MECHANICAL
|
||||
|
||||
**On the VM.** A dedicated venv under `agent-team/.venv` (excluded from rsync).
|
||||
The systemd unit's `ExecStart` points at **this venv's interpreter** (D-7 fixed),
|
||||
so the deps MUST land here.
|
||||
|
||||
```bash
|
||||
cd ~/orchestrator/agent-team
|
||||
python3 -m venv .venv
|
||||
. .venv/bin/activate
|
||||
pip install langgraph==1.1.10 langgraph-checkpoint-sqlite==3.1.0 \
|
||||
claude-agent-sdk slack_sdk slack_bolt requests
|
||||
```
|
||||
|
||||
> **Pin note (D-5):** `requirements.txt` pins `langgraph==1.1.10` /
|
||||
> `langgraph-checkpoint-sqlite==3.1.0`; match those exactly here. The other
|
||||
> runtime deps (`claude-agent-sdk`, `slack_sdk`, `slack_bolt`, `requests`) are
|
||||
> not yet in `requirements.txt` (D-5 open) — installed ad-hoc here. `requests`
|
||||
> is required by the GitHub transport/intake (D-4). `anthropic` is **not**
|
||||
> installed (only the opt-in `api` billing mode needs it).
|
||||
> `slack_bolt` is now actually exercised (D-1 fixed: the daemon starts the
|
||||
> Socket Mode listener).
|
||||
|
||||
**Verify the imports resolve (using the venv interpreter the unit will use):**
|
||||
```bash
|
||||
.venv/bin/python -c "import langgraph, langgraph.checkpoint.sqlite, slack_sdk, slack_bolt, requests; print('deps ok')"
|
||||
.venv/bin/python -c "import claude_agent_sdk; print('agent-sdk ok')"
|
||||
```
|
||||
|
||||
**ROLLBACK:** `deactivate 2>/dev/null; rm -rf ~/orchestrator/agent-team/.venv`
|
||||
|
||||
---
|
||||
|
||||
## STEP 5 — Initialize the durable ledger DB 🤖 MECHANICAL
|
||||
|
||||
**On the VM, venv active.** Idempotent; creates `state/agent_team.sqlite` with
|
||||
the `pending_questions` + `budget_ledger` + `schema_meta` tables (LangGraph
|
||||
`SqliteSaver` creates its own tables in the same file on first run).
|
||||
|
||||
```bash
|
||||
cd ~/orchestrator/agent-team
|
||||
. .venv/bin/activate
|
||||
python3 run-team.py init-db
|
||||
# Expect: "initialized ledger DB at .../state/agent_team.sqlite"
|
||||
ls -l state/ # agent_team.sqlite present; state/ is gitignored
|
||||
```
|
||||
|
||||
**ROLLBACK (reset the ledger only):**
|
||||
```bash
|
||||
cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak}
|
||||
rm ~/orchestrator/agent-team/state/agent_team.sqlite*
|
||||
# re-run `python3 run-team.py init-db` to recreate empty tables.
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## STEP 6 — Install + start the systemd unit 🤖 MECHANICAL (go/no-go 🧑)
|
||||
|
||||
**On the VM, as root.** Installs the long-running coordinator daemon. Keep the
|
||||
**hardening as-shipped**: `NoNewPrivileges`, `ProtectSystem=full`,
|
||||
`ProtectHome=read-only` + `ReadWritePaths=.../agent-team/state` (locked decision).
|
||||
|
||||
```bash
|
||||
sudo cp ~/orchestrator/agent-team/systemd/agent-team-coordinator.service /etc/systemd/system/
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl enable --now agent-team-coordinator.service
|
||||
systemctl status agent-team-coordinator.service
|
||||
journalctl -u agent-team-coordinator.service -e -f
|
||||
# expect: "agent-team coordinator starting"
|
||||
# and (if Slack + SLACK_APP_TOKEN provisioned):
|
||||
# "inbound Slack listener started (Socket Mode, background thread)"
|
||||
# or otherwise:
|
||||
# "inbound Slack listener not started (... or SLACK_APP_TOKEN is unset) ..."
|
||||
```
|
||||
|
||||
The unit (post-fix) runs
|
||||
`/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve` from
|
||||
`WorkingDirectory=/home/adam/orchestrator/agent-team` as `User=adam`, loading
|
||||
**both** `EnvironmentFile=-/home/adam/secrev.env` and
|
||||
`EnvironmentFile=-/home/adam/orchestrator/.env`, `Restart=on-failure`.
|
||||
|
||||
> **D-7 fixed:** `ExecStart` now resolves the venv interpreter, so the Step-4
|
||||
> deps are on the path. **D-1 fixed:** `serve` starts the inbound `SlackListener`
|
||||
> when Slack + app token are provisioned. **Go/no-go:** confirm the journal shows
|
||||
> the listener line you expect for your transport choice before Step 8.
|
||||
|
||||
**ROLLBACK (exercised — tear the unit down cleanly):**
|
||||
```bash
|
||||
sudo systemctl disable --now agent-team-coordinator.service
|
||||
sudo rm /etc/systemd/system/agent-team-coordinator.service
|
||||
sudo systemctl daemon-reload
|
||||
systemctl status agent-team-coordinator.service # should report "could not be found"
|
||||
```
|
||||
The ledger under `state/` is untouched. Step-1 snapshot is the whole-host fallback.
|
||||
|
||||
---
|
||||
|
||||
## STEP 7 — Verify the Slack inbound listener 🧑 OPERATOR-REQUIRED
|
||||
|
||||
The clarifier gate is two halves: **outbound** (post the question — done by
|
||||
`serve`/`start` via the Slack poster) and **inbound** (receive Adam's answer —
|
||||
the `SlackListener` over Socket Mode).
|
||||
|
||||
> **D-1 RESOLVED.** `Coordinator.serve()` now constructs and starts
|
||||
> `SlackListener` on a background daemon thread, concurrently with the tick/drain
|
||||
> loop, **when** the live transport is a `SlackTransport` AND `SLACK_APP_TOKEN`
|
||||
> is set. It shares the coordinator's own transport, ledger, and resume queue,
|
||||
> and stops cleanly on shutdown. The AUTHZ-01 owner allowlist + the open-status
|
||||
> compare-and-set are unchanged — the listener still fails closed on an empty
|
||||
> `AGENT_TEAM_SLACK_OWNER_IDS`.
|
||||
|
||||
**Confirm the live inbound path is up:**
|
||||
|
||||
```bash
|
||||
# In the journal (Step 6) expect the "inbound Slack listener started" line.
|
||||
# Verify the two gating vars are present in the unit's environment:
|
||||
grep -c SLACK_APP_TOKEN ~/secrev.env # 1
|
||||
grep -c AGENT_TEAM_SLACK_OWNER_IDS ~/secrev.env # 1 (else the gate rejects all answers)
|
||||
```
|
||||
|
||||
If you are **not** using Slack as the transport (or are deliberately running
|
||||
without the app token), the daemon runs the maintenance loop only and the demo
|
||||
uses the operator-CLI `answer` path — both are valid (P1-DEMO-SCRIPT.md).
|
||||
|
||||
**ROLLBACK:** none needed — verification only, no host state change.
|
||||
|
||||
---
|
||||
|
||||
## STEP 8 — Live P1 four-criteria acceptance demo 🧑 OPERATOR-REQUIRED
|
||||
|
||||
Run **P1-DEMO-SCRIPT.md** in full. All four §3.3.1 exit criteria must pass:
|
||||
(a) crash-safe resume, (b) duplicate-answer no-op, (c) post-deadline rejection +
|
||||
park, (d) two concurrent tasks resume independently. With D-1 fixed you may
|
||||
exercise the **live Slack answer path** for (b)/(d); the operator-CLI `answer`
|
||||
path remains available and exercises the identical compare-and-set. **Do not
|
||||
accept P1 until all four pass.**
|
||||
|
||||
**ROLLBACK:** the demo writes only ledger rows under `state/`; reset via the
|
||||
Step-5 ledger rollback, or restore the Step-1 snapshot, then re-run.
|
||||
|
||||
---
|
||||
|
||||
## STEP 9 — Plane-1 checker live dry-runs 🧑 OPERATOR-REQUIRED
|
||||
|
||||
The six checkers live under `security-review/checkers/`:
|
||||
`aws-posture.sh`, `compliance-drift.sh`, `confluence-doc.sh`,
|
||||
`dependency-cve.sh`, `doc-drift.sh`, `plan-groomer.sh`. All run **report +
|
||||
ALARM-only** (design D3): a clean run posts nothing and lands a mode-600 report;
|
||||
no auto-Jira/Notion writes.
|
||||
|
||||
- Dry-run each checker against `~/repo-mirrors` in report-only mode; confirm a
|
||||
clean run posts nothing and writes a mode-600 report. Confirm the exact
|
||||
invocation against the checker scripts and the shared
|
||||
`security-review/lib/` substrate on the box.
|
||||
- Confirm the **canary suite** runs first (a planted-fault miss is a COMPLACENCY
|
||||
ALARM and that role is skipped — design §6.4), and the **coverage rotation**
|
||||
pointer advances (a slipped role is a COVERAGE ALARM, deferred-not-dropped).
|
||||
|
||||
**ROLLBACK:** checkers are read-only over the mirror corpus; a dry-run produces
|
||||
only a report file. Remove the report dir to revert; no host state change.
|
||||
|
||||
---
|
||||
|
||||
## STEP 10 — P5 cross-plane loop dry-run 🧑 OPERATOR-REQUIRED
|
||||
|
||||
The Plane-1→Plane-2 loop turns confirmed checker findings into pipeline tasks:
|
||||
|
||||
```bash
|
||||
cd ~/orchestrator/agent-team && . .venv/bin/activate
|
||||
# Read one or more checker report JSONs and start one task per confirmed
|
||||
# at/above-threshold finding (default threshold: high). --dry-run posts nowhere.
|
||||
python3 run-team.py intake-checker --report <path-to-report.json> --threshold high --dry-run
|
||||
```
|
||||
|
||||
The Tier-3 dep-bump **fixer** is dry-run only on the box (it holds no write
|
||||
token, D2):
|
||||
|
||||
```bash
|
||||
python3 run-team.py fix --report security-review/<...>/dependency-cve.json \
|
||||
--finding-id <ID> --task-id <task> --dry-run
|
||||
# prints the fix spec + minimal bump patch + the CI workflow_dispatch inputs;
|
||||
# dispatches NOTHING. Live dispatch is the P3-live flip below.
|
||||
```
|
||||
|
||||
**ROLLBACK:** both are read-only / dry-run (no dispatch, no apply); `intake-checker`
|
||||
de-dup is in-memory per process. No host state to revert beyond ledger rows from
|
||||
a non-dry-run intake (Step-5 ledger rollback).
|
||||
|
||||
---
|
||||
|
||||
## STEP 11 — Wire the schedule 🧑 OPERATOR-REQUIRED
|
||||
|
||||
The coordinator daemon (Step 6) is **always-on**, not timer-driven. The
|
||||
**per-checker timers** (and the design's "shared timer with secrev", §8) attach
|
||||
here. The existing `sea-haven-secrev.timer` (OnCalendar `02:00`, Persistent) is
|
||||
untouched. Add a checker timer only after that checker is dry-run-validated
|
||||
(Step 9).
|
||||
|
||||
**ROLLBACK:** each timer gets its own `systemctl disable --now <unit>.timer` + `rm`.
|
||||
|
||||
---
|
||||
|
||||
## The P3-live flip (deferred; IAM + GitHub App gated) 🧑 OPERATOR-REQUIRED
|
||||
|
||||
P3 (the build→verify apply-and-open-draft-PR loop) is **opt-in and inert** in
|
||||
this deploy: `run-team.py` / `serve` pass `build_verify_wiring=None`, so no P3
|
||||
subgraph is assembled. Flipping it live is a **separate, gated** provisioning
|
||||
session, not part of the coordinator deploy:
|
||||
|
||||
1. **Mandatory reviews first.** The CI trust-boundary + OIDC IAM change is a
|
||||
breaking IAM change → **GPT-4.1 cross-review** (global instructions) AND
|
||||
`/sh-security-review` on the apply/verify surface. Do not flip without both.
|
||||
2. **Provision the GitHub App** for the trusted apply path (the App that opens
|
||||
the draft PR), and the **`agent-apply` GitHub Actions environment** that holds
|
||||
the apply path's scoped permissions.
|
||||
3. **Bind the live build→verify wiring** via
|
||||
`agent_team.coordinator.gated_build_verify_wiring(...)` (the read-only CI
|
||||
result fetcher + the real diff builder) — the seam a leaf calls *after* the
|
||||
gate clears. The CI fetcher is read-only and fails closed (missing token /
|
||||
404 / auth failure → `None` → the gate BLOCKs and the task parks).
|
||||
4. **Set the apply env vars** the live path reads (the read-only CI-result token
|
||||
and the dispatch target), then re-run the fixer **without** `--dry-run` only
|
||||
once the dispatcher is bound.
|
||||
|
||||
Until every step above is done, the box dispatches/applies nothing.
|
||||
|
||||
---
|
||||
|
||||
## Post-session definition-of-done (design §7, global instructions)
|
||||
|
||||
- [ ] All four P1 criteria demonstrated live (Step 8).
|
||||
- [ ] `project_r720_agent_team` memory created/updated.
|
||||
- [ ] Confluence "AWS Architecture Map" / IT host inventory updated to show
|
||||
`sh-secrev` now also hosts the always-on agent-team coordinator daemon.
|
||||
- [ ] `/sh-security-review` run on the Slack inbound listener surface
|
||||
(auth + untrusted-input; mandatory) — flag outstanding if not run.
|
||||
- [ ] OPERATOR-RUNBOOK.md reviewed by whoever holds the pager.
|
||||
- [ ] Snapshot retained until the daemon runs clean for one full cycle, then pruned.
|
||||
Reference in a new issue