287 lines
12 KiB
Markdown
287 lines
12 KiB
Markdown
|
|
# P1-DEMO-SCRIPT — live four-criteria acceptance demo (design §3.3.1 / §7.1 P1)
|
||
|
|
|
||
|
|
The §7.1 P1 exit gate: demonstrate, on the live box, all four durable
|
||
|
|
human-in-the-loop criteria before P1 is accepted:
|
||
|
|
|
||
|
|
- **(a)** kill the box mid-wait and have the task resume after restart;
|
||
|
|
- **(b)** submit a duplicate answer and confirm it no-ops;
|
||
|
|
- **(c)** submit an answer after the deadline expired and confirm it is rejected
|
||
|
|
and the task parks;
|
||
|
|
- **(d)** two tasks suspended concurrently resume independently to the correct
|
||
|
|
thread.
|
||
|
|
|
||
|
|
Every command below is grounded in the **actual** code surface
|
||
|
|
(`run-team.py`, `coordinator.py`, `slack_listener.py`, `responder.py`,
|
||
|
|
`resume_worker.py`, `db/schema.py`, `graph.py`). No invented flags. Where the
|
||
|
|
code does not expose a needed knob (e.g. a short deadline), the script uses a
|
||
|
|
direct `sqlite3` write against the documented `pending_questions` schema and says
|
||
|
|
so.
|
||
|
|
|
||
|
|
## Two answer paths — live Slack OR the operator CLI
|
||
|
|
|
||
|
|
**D-1 is fixed:** `Coordinator.serve()` now starts the inbound `SlackListener`
|
||
|
|
when the live transport is Slack AND `SLACK_APP_TOKEN` is set, so the
|
||
|
|
**live Slack answer round-trip works**. You can run the demo either way:
|
||
|
|
|
||
|
|
- **Live Slack** — Adam clicks the Block Kit button / replies in the channel; the
|
||
|
|
listener normalizes the event, runs the AUTHZ-01 owner check, and drives the
|
||
|
|
first-answer-wins compare-and-set (`responder.submit_answer` →
|
||
|
|
`db.schema.answer_question`).
|
||
|
|
- **Operator CLI** — `run-team.py answer <qid> --answer ... --confirm` runs the
|
||
|
|
**identical** compare-and-set (audit-logged answer-on-behalf). Useful when the
|
||
|
|
Slack app is not yet provisioned, or to script the assertions.
|
||
|
|
|
||
|
|
Both exercise the same durable mechanic; the human gate decision is always
|
||
|
|
Adam's. The assertions below assert on the **durable ledger status** (the §3.3.1
|
||
|
|
source of truth) and are identical for either path. The examples use the CLI
|
||
|
|
`answer --confirm` form so they are copy-pasteable; substitute "Adam answers in
|
||
|
|
Slack" wherever you prefer the live path.
|
||
|
|
|
||
|
|
> The automated proof of this mechanic is `tests/sim/test_p1_exit_criteria.py`
|
||
|
|
> (a `SimPipeline` harness) and `tests/test_coordinator.py` (the serve/listener
|
||
|
|
> wiring). This script is the live-box demonstration on top of that.
|
||
|
|
|
||
|
|
## Verb map (use `run-team.py`, not `operator_cli.py`)
|
||
|
|
|
||
|
|
| Need | `run-team.py` verb | Notes |
|
||
|
|
|---|---|---|
|
||
|
|
| start a task to the human gate | `start --task "..." [--transport slack] [--dry-run]` | mints a `thread_id`, posts the clarifier, writes the `open` ledger row |
|
||
|
|
| list waiting questions | `list` (default `open`) / `list --all` / `list --parked` | JSON rows |
|
||
|
|
| inspect one row | `show <question_id>` | JSON row |
|
||
|
|
| answer on the task's behalf | `answer <question_id> --answer <payload> --confirm` | destructive, audit-logged; first-answer-wins CAS |
|
||
|
|
| force-expire an open question | `expire <question_id> --confirm` | destructive; flips `open`→`expired` |
|
||
|
|
| un-park (reopen) an expired question | `force-resume <question_id> --confirm` | only acts on `expired` rows (reopens them) |
|
||
|
|
|
||
|
|
`operator_cli.py` has different verbs (`force-expire`, `answer-on-behalf`) and
|
||
|
|
**no `show`**, and requires `--db`/`--audit-log` — do not use it here.
|
||
|
|
|
||
|
|
## Preconditions
|
||
|
|
|
||
|
|
```bash
|
||
|
|
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
|
||
|
|
cd ~/orchestrator/agent-team
|
||
|
|
. .venv/bin/activate # so run-team.py uses the venv deps
|
||
|
|
# Confirm the ledger exists (Step 5 of the runbook):
|
||
|
|
python3 run-team.py list --all # [] on a fresh DB is fine
|
||
|
|
# For the live-Slack path, confirm the listener started:
|
||
|
|
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
|
||
|
|
# -> "inbound Slack listener started (Socket Mode, background thread)"
|
||
|
|
```
|
||
|
|
|
||
|
|
Run all `run-team.py` commands from `~/orchestrator/agent-team` so they hit the
|
||
|
|
default ledger (`state/agent_team.sqlite`) and default audit log
|
||
|
|
(`state/audit.log.jsonl`). The clarifier is posted to `SLACK_CHANNEL_ID`.
|
||
|
|
|
||
|
|
> **Resume execution model.** A won answer (rowcount 1) enqueues a `ResumeJob`
|
||
|
|
> onto the coordinator's in-process `resume_queue`; the daemon drains it on the
|
||
|
|
> next `tick()`, or the startup `recover()` re-drives any `answered`-but-unresumed
|
||
|
|
> row after a restart. So a CLI `answer --confirm` (or a live Slack answer) flips
|
||
|
|
> the ledger row to `answered`; the **running daemon** then resumes the graph.
|
||
|
|
> The demo asserts on the durable ledger status.
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## (a) Crash-safe resume — kill mid-wait, restart, task resumes 🧑 Adam answers
|
||
|
|
|
||
|
|
**Setup — drive a task to the clarifier wait.** With the daemon already running:
|
||
|
|
|
||
|
|
```bash
|
||
|
|
python3 run-team.py start --task "demo-a: trivial scoped task"
|
||
|
|
# prints a thread_id, e.g. 3f2a... (record it as $TID_A)
|
||
|
|
python3 run-team.py list
|
||
|
|
# expect one row: status="open", a thread_id, a question_id (record as $QID_A),
|
||
|
|
# channel_ref set (Slack ts) or null if posting is dry/unavailable.
|
||
|
|
```
|
||
|
|
|
||
|
|
**Kill the box mid-wait, then restart:**
|
||
|
|
|
||
|
|
```bash
|
||
|
|
sudo systemctl stop agent-team-coordinator.service
|
||
|
|
# (optionally reboot the VM here for a stronger demonstration)
|
||
|
|
sudo systemctl start agent-team-coordinator.service
|
||
|
|
journalctl -u agent-team-coordinator.service -e | tail -40 # expect "starting" + recover() sweep
|
||
|
|
```
|
||
|
|
|
||
|
|
**ASSERT — durable state survived the kill (no answer yet):**
|
||
|
|
|
||
|
|
```bash
|
||
|
|
python3 run-team.py show $QID_A
|
||
|
|
# PASS: status == "open" (the question was NOT lost across the restart)
|
||
|
|
```
|
||
|
|
|
||
|
|
**🧑 Human gate — Adam answers after the restart (Slack OR CLI):**
|
||
|
|
|
||
|
|
```bash
|
||
|
|
# Live Slack: Adam replies/clicks in the channel. OR via the CLI:
|
||
|
|
python3 run-team.py answer $QID_A --answer "scope: just demo, no real change" --confirm
|
||
|
|
python3 run-team.py show $QID_A
|
||
|
|
# PASS: status == "answered"
|
||
|
|
# Wait one tick (~30s, DEFAULT_POLL_INTERVAL) for the daemon to drain the resume:
|
||
|
|
journalctl -u agent-team-coordinator.service -e | tail -20
|
||
|
|
```
|
||
|
|
|
||
|
|
**PASS criterion (a):** the question stayed `open` across the kill, and an answer
|
||
|
|
submitted *after* the restart drove it forward.
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## (b) Duplicate answer is a no-op 🧑 Adam answers
|
||
|
|
|
||
|
|
```bash
|
||
|
|
python3 run-team.py start --task "demo-b: duplicate-answer test"
|
||
|
|
python3 run-team.py list # record the new question_id as $QID_B (status open)
|
||
|
|
```
|
||
|
|
|
||
|
|
**🧑 First answer (the winning one) — Slack OR CLI:**
|
||
|
|
|
||
|
|
```bash
|
||
|
|
python3 run-team.py answer $QID_B --answer "first-answer" --via "slack:U1" --confirm
|
||
|
|
# expect: "answered question <QID_B> (via slack:U1)" (exit 0)
|
||
|
|
```
|
||
|
|
|
||
|
|
**Duplicate / second answer (must lose the compare-and-set):**
|
||
|
|
|
||
|
|
```bash
|
||
|
|
python3 run-team.py answer $QID_B --answer "second-answer" --via "github:U2" --confirm
|
||
|
|
# expect: "answer no-op: question <QID_B> was not 'open' ..." on stderr, exit 1
|
||
|
|
```
|
||
|
|
|
||
|
|
For the **live-Slack** variant, Adam (or a second authorized owner) answers the
|
||
|
|
same message twice; the second event loses the CAS and is logged
|
||
|
|
`ignored Slack answer for question_id=... (not open: duplicate ...)`.
|
||
|
|
|
||
|
|
**ASSERT — the first answer is preserved:**
|
||
|
|
|
||
|
|
```bash
|
||
|
|
python3 run-team.py show $QID_B
|
||
|
|
# PASS: status == "answered"; answered_via == "slack:U1" (the FIRST answer)
|
||
|
|
```
|
||
|
|
|
||
|
|
**PASS criterion (b):** first-answer-wins (`rowcount==1`); the duplicate hits the
|
||
|
|
`BEGIN IMMEDIATE` compare-and-set in `answer_question` and is ignored
|
||
|
|
(`rowcount==0`) — no second resume, no overwrite.
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## (c) Past-deadline answer rejected + task parks 🧑 Adam observes
|
||
|
|
|
||
|
|
> **Code reality:** `run-team.py start` always sets the clarifier deadline from
|
||
|
|
> `DEFAULT_CLARIFY_DEADLINE = 24h`. There is **no CLI flag for a short deadline**.
|
||
|
|
> Two faithful ways to demo expiry without waiting 24h:
|
||
|
|
|
||
|
|
### Option C1 — force the deadline past, let the daemon's sweep expire it (most faithful)
|
||
|
|
|
||
|
|
```bash
|
||
|
|
python3 run-team.py start --task "demo-c: deadline test"
|
||
|
|
python3 run-team.py list # record $QID_C (status open)
|
||
|
|
|
||
|
|
# Set this question's deadline into the past directly in the ledger:
|
||
|
|
sqlite3 state/agent_team.sqlite \
|
||
|
|
"UPDATE pending_questions SET deadline_at='2000-01-01T00:00:00+00:00' WHERE question_id='$QID_C';"
|
||
|
|
|
||
|
|
# Wait one daemon tick (~30s) for the deadline sweep to flip it, OR observe:
|
||
|
|
journalctl -u agent-team-coordinator.service -e | tail -20
|
||
|
|
# expect the park ALARM line: "task parked: clarifier question <QID_C> expired ..."
|
||
|
|
```
|
||
|
|
|
||
|
|
### Option C2 — operator force-expire (if you do not want to touch the DB)
|
||
|
|
|
||
|
|
```bash
|
||
|
|
python3 run-team.py start --task "demo-c: deadline test"
|
||
|
|
python3 run-team.py list # record $QID_C
|
||
|
|
python3 run-team.py expire $QID_C --confirm # destructive, audit-logged; open->expired
|
||
|
|
```
|
||
|
|
|
||
|
|
**ASSERT — the question is expired and the late answer is rejected:**
|
||
|
|
|
||
|
|
```bash
|
||
|
|
python3 run-team.py show $QID_C
|
||
|
|
# PASS: status == "expired"
|
||
|
|
|
||
|
|
# 🧑 Adam submits a LATE answer (Slack OR CLI) — it must lose the CAS:
|
||
|
|
python3 run-team.py answer $QID_C --answer "too-late" --confirm
|
||
|
|
# expect: "answer no-op: question <QID_C> was not 'open' ..." stderr, exit 1
|
||
|
|
python3 run-team.py show $QID_C
|
||
|
|
# PASS: status still "expired"; answer_json still NULL
|
||
|
|
|
||
|
|
# The parked task surfaces in the parked view:
|
||
|
|
python3 run-team.py list --parked # PASS: $QID_C appears here
|
||
|
|
```
|
||
|
|
|
||
|
|
**Deliberate un-park (proves the operator recovery path, §6.6):**
|
||
|
|
|
||
|
|
```bash
|
||
|
|
python3 run-team.py force-resume $QID_C --confirm
|
||
|
|
# expect: "force-resume: reopened expired question <QID_C>; it will be re-delivered"
|
||
|
|
python3 run-team.py show $QID_C
|
||
|
|
# status flips back to "open" (reopen_question), deadline_at cleared.
|
||
|
|
```
|
||
|
|
|
||
|
|
**PASS criterion (c):** the expired question rejects the late answer, the task
|
||
|
|
parks rather than spins, and the operator can deliberately un-park it via
|
||
|
|
`force-resume` (which reopens only an `expired` row).
|
||
|
|
|
||
|
|
> **`force-resume` semantics (verified):** `run-team.py force-resume` reopens an
|
||
|
|
> `expired` question (the parked case). For an `answered` question it records
|
||
|
|
> intent and reports the recovery sweep will resume it (no mutation). For
|
||
|
|
> `open`/`superseded`/absent it is a no-op exit 1. It does **not** supersede.
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## (d) Two concurrent tasks resume independently 🧑 Adam answers
|
||
|
|
|
||
|
|
```bash
|
||
|
|
python3 run-team.py start --task "demo-d task A" # record $TID_A2, then:
|
||
|
|
python3 run-team.py start --task "demo-d task B" # record $TID_B2
|
||
|
|
python3 run-team.py list
|
||
|
|
# expect TWO open rows with DISTINCT thread_id AND distinct question_id.
|
||
|
|
# Record $QID_A2 and $QID_B2 — match by thread_id.
|
||
|
|
```
|
||
|
|
|
||
|
|
**(Optional) restart the daemon first** to also show concurrent tasks survive a
|
||
|
|
restart, then answer.
|
||
|
|
|
||
|
|
**🧑 Answer the SECOND task first, with a distinct answer, then the first
|
||
|
|
(Slack OR CLI):**
|
||
|
|
|
||
|
|
```bash
|
||
|
|
python3 run-team.py answer $QID_B2 --answer "answer-for-B" --via "slack:U2" --confirm
|
||
|
|
python3 run-team.py answer $QID_A2 --answer "answer-for-A" --via "slack:U1" --confirm
|
||
|
|
# Wait one daemon tick (~30s) for both resumes to drain.
|
||
|
|
```
|
||
|
|
|
||
|
|
**ASSERT — each task carries its OWN answer; no cross-talk:**
|
||
|
|
|
||
|
|
```bash
|
||
|
|
python3 run-team.py show $QID_A2
|
||
|
|
# PASS: status "answered", answered_via "slack:U1", thread_id == $TID_A2
|
||
|
|
python3 run-team.py show $QID_B2
|
||
|
|
# PASS: status "answered", answered_via "slack:U2", thread_id == $TID_B2
|
||
|
|
python3 run-team.py list --all
|
||
|
|
# PASS: the two rows resolved on their own thread_id; no cross-contamination.
|
||
|
|
```
|
||
|
|
|
||
|
|
**PASS criterion (d):** two concurrently-suspended tasks each resumed to their
|
||
|
|
own `thread_id` with their own answer — answering B before A did not misroute,
|
||
|
|
and the per-thread single-flight guard kept them independent.
|
||
|
|
|
||
|
|
---
|
||
|
|
|
||
|
|
## Acceptance
|
||
|
|
|
||
|
|
P1 is accepted only when **(a), (b), (c), and (d) all pass** on the live box.
|
||
|
|
Record the four `show` outputs (or `journalctl` excerpts) as evidence. Then
|
||
|
|
complete the PROVISIONING-RUNBOOK post-session definition-of-done (memory +
|
||
|
|
Confluence + the mandatory `/sh-security-review` on the Slack inbound listener).
|
||
|
|
|
||
|
|
## What can only be verified on the live box
|
||
|
|
|
||
|
|
- That the daemon actually drains the resume and the LangGraph checkpoint
|
||
|
|
advances (asserted via ledger status + journalctl; the graph-state advance is
|
||
|
|
observable only on the box).
|
||
|
|
- The live **Slack** post/answer round-trip end-to-end (the listener is wired —
|
||
|
|
D-1 fixed — but the live Socket Mode socket + a real `SLACK_BOT_TOKEN` /
|
||
|
|
`SLACK_APP_TOKEN` / channel membership are only present on the box).
|
||
|
|
- The exact `channel_ref` value (Slack message `ts`) — depends on a live Slack
|
||
|
|
post succeeding.
|