This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/docs/provisioning/P1-DEMO-SCRIPT.md

287 lines
12 KiB
Markdown
Raw Normal View History

feat(agent-team): deploy-readiness — serve starts Slack listener + systemd + provisioning docs (#23) * fix(agent-team): serve() starts the inbound Slack listener (D-1) Coordinator.serve() now constructs and starts the SlackListener concurrently with the tick/drain loop on a background daemon thread, but ONLY when the live transport is a SlackTransport AND SLACK_APP_TOKEN is configured. When Slack is not the transport or the app token is absent, serve() behaves exactly as before (tick/recover only) — Slack is never made mandatory. - New injectable build_listener seam + default_slack_listener_factory sharing the coordinator's own transport, ledger db_path, and resume_queue put. - AUTHZ-01 owner-allowlist + open-status CAS untouched: serve() sources AGENT_TEAM_SLACK_OWNER_IDS in SlackListener.serve, which still fails closed. - SlackListener.close() added for clean Socket Mode teardown on shutdown; serve() stops the listener + joins the thread in a finally. - Tests: start-when-Slack+app-token, no-start otherwise, clean shutdown, idempotent start, serve start/stop around the loop, listener close(). * fix(agent-team): systemd unit loads ~/orchestrator/.env + uses venv python (D-2/D-7) D-2: add EnvironmentFile=-/home/adam/orchestrator/.env (optional '-') so the P2 GPT-4.1 review loop's cross_reviewer sub-process can read the non-Claude provider key once a task reaches REVIEW. Mirrors the sea-haven-secrev unit. D-7: point ExecStart at the agent-team venv interpreter (/home/adam/orchestrator/agent-team/.venv/bin/python) instead of /usr/bin/env python3, which resolved the system interpreter without the installed deps under systemd's PATH. All hardening (NoNewPrivileges / ProtectSystem=full / ProtectHome=read-only / ReadWritePaths) is retained unchanged (locked decision). * docs(agent-team): land provisioning + operator runbooks under docs/provisioning - PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5 intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked FIXED so the demo can use the live Slack answer path. - P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the Slack and operator-CLI answer paths documented for all four exit criteria. - DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the operator-CLI divergence kept as provisioning notes. - OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks, failed HITL resumes, budget exhaustion, transport outages, and COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6). * fix(agent-team): supervise the Slack listener thread — recurring ALARM + respawn sh-security-review (logic) MEDIUM: a crashed listener thread was logged once, then the daemon ran on 'deaf' — posting clarifier questions but receiving no answers, every gate silently parking, process never exiting so systemd Restart=on-failure never fired. serve() now calls _supervise_slack_listener() each pass: when the listener is enabled but its thread is dead, it emits a recurring ERROR ALARM and respawns via the idempotent starter (self-heal). No-op when alive or disabled. +3 tests. (authz detector: wiring clean — AUTHZ-01 fail-closed allowlist + open-status CAS intact, dead listener fails SAFE.)
2026-06-18 16:56:21 -04:00
# P1-DEMO-SCRIPT — live four-criteria acceptance demo (design §3.3.1 / §7.1 P1)
The §7.1 P1 exit gate: demonstrate, on the live box, all four durable
human-in-the-loop criteria before P1 is accepted:
- **(a)** kill the box mid-wait and have the task resume after restart;
- **(b)** submit a duplicate answer and confirm it no-ops;
- **(c)** submit an answer after the deadline expired and confirm it is rejected
and the task parks;
- **(d)** two tasks suspended concurrently resume independently to the correct
thread.
Every command below is grounded in the **actual** code surface
(`run-team.py`, `coordinator.py`, `slack_listener.py`, `responder.py`,
`resume_worker.py`, `db/schema.py`, `graph.py`). No invented flags. Where the
code does not expose a needed knob (e.g. a short deadline), the script uses a
direct `sqlite3` write against the documented `pending_questions` schema and says
so.
## Two answer paths — live Slack OR the operator CLI
**D-1 is fixed:** `Coordinator.serve()` now starts the inbound `SlackListener`
when the live transport is Slack AND `SLACK_APP_TOKEN` is set, so the
**live Slack answer round-trip works**. You can run the demo either way:
- **Live Slack** — Adam clicks the Block Kit button / replies in the channel; the
listener normalizes the event, runs the AUTHZ-01 owner check, and drives the
first-answer-wins compare-and-set (`responder.submit_answer` →
`db.schema.answer_question`).
- **Operator CLI** — `run-team.py answer <qid> --answer ... --confirm` runs the
**identical** compare-and-set (audit-logged answer-on-behalf). Useful when the
Slack app is not yet provisioned, or to script the assertions.
Both exercise the same durable mechanic; the human gate decision is always
Adam's. The assertions below assert on the **durable ledger status** (the §3.3.1
source of truth) and are identical for either path. The examples use the CLI
`answer --confirm` form so they are copy-pasteable; substitute "Adam answers in
Slack" wherever you prefer the live path.
> The automated proof of this mechanic is `tests/sim/test_p1_exit_criteria.py`
> (a `SimPipeline` harness) and `tests/test_coordinator.py` (the serve/listener
> wiring). This script is the live-box demonstration on top of that.
## Verb map (use `run-team.py`, not `operator_cli.py`)
| Need | `run-team.py` verb | Notes |
|---|---|---|
| start a task to the human gate | `start --task "..." [--transport slack] [--dry-run]` | mints a `thread_id`, posts the clarifier, writes the `open` ledger row |
| list waiting questions | `list` (default `open`) / `list --all` / `list --parked` | JSON rows |
| inspect one row | `show <question_id>` | JSON row |
| answer on the task's behalf | `answer <question_id> --answer <payload> --confirm` | destructive, audit-logged; first-answer-wins CAS |
| force-expire an open question | `expire <question_id> --confirm` | destructive; flips `open`→`expired` |
| un-park (reopen) an expired question | `force-resume <question_id> --confirm` | only acts on `expired` rows (reopens them) |
`operator_cli.py` has different verbs (`force-expire`, `answer-on-behalf`) and
**no `show`**, and requires `--db`/`--audit-log` — do not use it here.
## Preconditions
```bash
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team
. .venv/bin/activate # so run-team.py uses the venv deps
# Confirm the ledger exists (Step 5 of the runbook):
python3 run-team.py list --all # [] on a fresh DB is fine
# For the live-Slack path, confirm the listener started:
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
# -> "inbound Slack listener started (Socket Mode, background thread)"
```
Run all `run-team.py` commands from `~/orchestrator/agent-team` so they hit the
default ledger (`state/agent_team.sqlite`) and default audit log
(`state/audit.log.jsonl`). The clarifier is posted to `SLACK_CHANNEL_ID`.
> **Resume execution model.** A won answer (rowcount 1) enqueues a `ResumeJob`
> onto the coordinator's in-process `resume_queue`; the daemon drains it on the
> next `tick()`, or the startup `recover()` re-drives any `answered`-but-unresumed
> row after a restart. So a CLI `answer --confirm` (or a live Slack answer) flips
> the ledger row to `answered`; the **running daemon** then resumes the graph.
> The demo asserts on the durable ledger status.
---
## (a) Crash-safe resume — kill mid-wait, restart, task resumes 🧑 Adam answers
**Setup — drive a task to the clarifier wait.** With the daemon already running:
```bash
python3 run-team.py start --task "demo-a: trivial scoped task"
# prints a thread_id, e.g. 3f2a... (record it as $TID_A)
python3 run-team.py list
# expect one row: status="open", a thread_id, a question_id (record as $QID_A),
# channel_ref set (Slack ts) or null if posting is dry/unavailable.
```
**Kill the box mid-wait, then restart:**
```bash
sudo systemctl stop agent-team-coordinator.service
# (optionally reboot the VM here for a stronger demonstration)
sudo systemctl start agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e | tail -40 # expect "starting" + recover() sweep
```
**ASSERT — durable state survived the kill (no answer yet):**
```bash
python3 run-team.py show $QID_A
# PASS: status == "open" (the question was NOT lost across the restart)
```
**🧑 Human gate — Adam answers after the restart (Slack OR CLI):**
```bash
# Live Slack: Adam replies/clicks in the channel. OR via the CLI:
python3 run-team.py answer $QID_A --answer "scope: just demo, no real change" --confirm
python3 run-team.py show $QID_A
# PASS: status == "answered"
# Wait one tick (~30s, DEFAULT_POLL_INTERVAL) for the daemon to drain the resume:
journalctl -u agent-team-coordinator.service -e | tail -20
```
**PASS criterion (a):** the question stayed `open` across the kill, and an answer
submitted *after* the restart drove it forward.
---
## (b) Duplicate answer is a no-op 🧑 Adam answers
```bash
python3 run-team.py start --task "demo-b: duplicate-answer test"
python3 run-team.py list # record the new question_id as $QID_B (status open)
```
**🧑 First answer (the winning one) — Slack OR CLI:**
```bash
python3 run-team.py answer $QID_B --answer "first-answer" --via "slack:U1" --confirm
# expect: "answered question <QID_B> (via slack:U1)" (exit 0)
```
**Duplicate / second answer (must lose the compare-and-set):**
```bash
python3 run-team.py answer $QID_B --answer "second-answer" --via "github:U2" --confirm
# expect: "answer no-op: question <QID_B> was not 'open' ..." on stderr, exit 1
```
For the **live-Slack** variant, Adam (or a second authorized owner) answers the
same message twice; the second event loses the CAS and is logged
`ignored Slack answer for question_id=... (not open: duplicate ...)`.
**ASSERT — the first answer is preserved:**
```bash
python3 run-team.py show $QID_B
# PASS: status == "answered"; answered_via == "slack:U1" (the FIRST answer)
```
**PASS criterion (b):** first-answer-wins (`rowcount==1`); the duplicate hits the
`BEGIN IMMEDIATE` compare-and-set in `answer_question` and is ignored
(`rowcount==0`) — no second resume, no overwrite.
---
## (c) Past-deadline answer rejected + task parks 🧑 Adam observes
> **Code reality:** `run-team.py start` always sets the clarifier deadline from
> `DEFAULT_CLARIFY_DEADLINE = 24h`. There is **no CLI flag for a short deadline**.
> Two faithful ways to demo expiry without waiting 24h:
### Option C1 — force the deadline past, let the daemon's sweep expire it (most faithful)
```bash
python3 run-team.py start --task "demo-c: deadline test"
python3 run-team.py list # record $QID_C (status open)
# Set this question's deadline into the past directly in the ledger:
sqlite3 state/agent_team.sqlite \
"UPDATE pending_questions SET deadline_at='2000-01-01T00:00:00+00:00' WHERE question_id='$QID_C';"
# Wait one daemon tick (~30s) for the deadline sweep to flip it, OR observe:
journalctl -u agent-team-coordinator.service -e | tail -20
# expect the park ALARM line: "task parked: clarifier question <QID_C> expired ..."
```
### Option C2 — operator force-expire (if you do not want to touch the DB)
```bash
python3 run-team.py start --task "demo-c: deadline test"
python3 run-team.py list # record $QID_C
python3 run-team.py expire $QID_C --confirm # destructive, audit-logged; open->expired
```
**ASSERT — the question is expired and the late answer is rejected:**
```bash
python3 run-team.py show $QID_C
# PASS: status == "expired"
# 🧑 Adam submits a LATE answer (Slack OR CLI) — it must lose the CAS:
python3 run-team.py answer $QID_C --answer "too-late" --confirm
# expect: "answer no-op: question <QID_C> was not 'open' ..." stderr, exit 1
python3 run-team.py show $QID_C
# PASS: status still "expired"; answer_json still NULL
# The parked task surfaces in the parked view:
python3 run-team.py list --parked # PASS: $QID_C appears here
```
**Deliberate un-park (proves the operator recovery path, §6.6):**
```bash
python3 run-team.py force-resume $QID_C --confirm
# expect: "force-resume: reopened expired question <QID_C>; it will be re-delivered"
python3 run-team.py show $QID_C
# status flips back to "open" (reopen_question), deadline_at cleared.
```
**PASS criterion (c):** the expired question rejects the late answer, the task
parks rather than spins, and the operator can deliberately un-park it via
`force-resume` (which reopens only an `expired` row).
> **`force-resume` semantics (verified):** `run-team.py force-resume` reopens an
> `expired` question (the parked case). For an `answered` question it records
> intent and reports the recovery sweep will resume it (no mutation). For
> `open`/`superseded`/absent it is a no-op exit 1. It does **not** supersede.
---
## (d) Two concurrent tasks resume independently 🧑 Adam answers
```bash
python3 run-team.py start --task "demo-d task A" # record $TID_A2, then:
python3 run-team.py start --task "demo-d task B" # record $TID_B2
python3 run-team.py list
# expect TWO open rows with DISTINCT thread_id AND distinct question_id.
# Record $QID_A2 and $QID_B2 — match by thread_id.
```
**(Optional) restart the daemon first** to also show concurrent tasks survive a
restart, then answer.
**🧑 Answer the SECOND task first, with a distinct answer, then the first
(Slack OR CLI):**
```bash
python3 run-team.py answer $QID_B2 --answer "answer-for-B" --via "slack:U2" --confirm
python3 run-team.py answer $QID_A2 --answer "answer-for-A" --via "slack:U1" --confirm
# Wait one daemon tick (~30s) for both resumes to drain.
```
**ASSERT — each task carries its OWN answer; no cross-talk:**
```bash
python3 run-team.py show $QID_A2
# PASS: status "answered", answered_via "slack:U1", thread_id == $TID_A2
python3 run-team.py show $QID_B2
# PASS: status "answered", answered_via "slack:U2", thread_id == $TID_B2
python3 run-team.py list --all
# PASS: the two rows resolved on their own thread_id; no cross-contamination.
```
**PASS criterion (d):** two concurrently-suspended tasks each resumed to their
own `thread_id` with their own answer — answering B before A did not misroute,
and the per-thread single-flight guard kept them independent.
---
## Acceptance
P1 is accepted only when **(a), (b), (c), and (d) all pass** on the live box.
Record the four `show` outputs (or `journalctl` excerpts) as evidence. Then
complete the PROVISIONING-RUNBOOK post-session definition-of-done (memory +
Confluence + the mandatory `/sh-security-review` on the Slack inbound listener).
## What can only be verified on the live box
- That the daemon actually drains the resume and the LangGraph checkpoint
advances (asserted via ledger status + journalctl; the graph-state advance is
observable only on the box).
- The live **Slack** post/answer round-trip end-to-end (the listener is wired —
D-1 fixed — but the live Socket Mode socket + a real `SLACK_BOT_TOKEN` /
`SLACK_APP_TOKEN` / channel membership are only present on the box).
- The exact `channel_ref` value (Slack message `ts`) — depends on a live Slack
post succeeding.