- PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5 intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked FIXED so the demo can use the live Slack answer path. - P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the Slack and operator-CLI answer paths documented for all four exit criteria. - DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the operator-CLI divergence kept as provisioning notes. - OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks, failed HITL resumes, budget exhaustion, transport outages, and COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6).
12 KiB
P1-DEMO-SCRIPT — live four-criteria acceptance demo (design §3.3.1 / §7.1 P1)
The §7.1 P1 exit gate: demonstrate, on the live box, all four durable human-in-the-loop criteria before P1 is accepted:
- (a) kill the box mid-wait and have the task resume after restart;
- (b) submit a duplicate answer and confirm it no-ops;
- (c) submit an answer after the deadline expired and confirm it is rejected and the task parks;
- (d) two tasks suspended concurrently resume independently to the correct thread.
Every command below is grounded in the actual code surface
(run-team.py, coordinator.py, slack_listener.py, responder.py,
resume_worker.py, db/schema.py, graph.py). No invented flags. Where the
code does not expose a needed knob (e.g. a short deadline), the script uses a
direct sqlite3 write against the documented pending_questions schema and says
so.
Two answer paths — live Slack OR the operator CLI
D-1 is fixed: Coordinator.serve() now starts the inbound SlackListener
when the live transport is Slack AND SLACK_APP_TOKEN is set, so the
live Slack answer round-trip works. You can run the demo either way:
- Live Slack — Adam clicks the Block Kit button / replies in the channel; the
listener normalizes the event, runs the AUTHZ-01 owner check, and drives the
first-answer-wins compare-and-set (
responder.submit_answer→db.schema.answer_question). - Operator CLI —
run-team.py answer <qid> --answer ... --confirmruns the identical compare-and-set (audit-logged answer-on-behalf). Useful when the Slack app is not yet provisioned, or to script the assertions.
Both exercise the same durable mechanic; the human gate decision is always
Adam's. The assertions below assert on the durable ledger status (the §3.3.1
source of truth) and are identical for either path. The examples use the CLI
answer --confirm form so they are copy-pasteable; substitute "Adam answers in
Slack" wherever you prefer the live path.
The automated proof of this mechanic is
tests/sim/test_p1_exit_criteria.py(aSimPipelineharness) andtests/test_coordinator.py(the serve/listener wiring). This script is the live-box demonstration on top of that.
Verb map (use run-team.py, not operator_cli.py)
| Need | run-team.py verb |
Notes |
|---|---|---|
| start a task to the human gate | start --task "..." [--transport slack] [--dry-run] |
mints a thread_id, posts the clarifier, writes the open ledger row |
| list waiting questions | list (default open) / list --all / list --parked |
JSON rows |
| inspect one row | show <question_id> |
JSON row |
| answer on the task's behalf | answer <question_id> --answer <payload> --confirm |
destructive, audit-logged; first-answer-wins CAS |
| force-expire an open question | expire <question_id> --confirm |
destructive; flips open→expired |
| un-park (reopen) an expired question | force-resume <question_id> --confirm |
only acts on expired rows (reopens them) |
operator_cli.py has different verbs (force-expire, answer-on-behalf) and
no show, and requires --db/--audit-log — do not use it here.
Preconditions
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team
. .venv/bin/activate # so run-team.py uses the venv deps
# Confirm the ledger exists (Step 5 of the runbook):
python3 run-team.py list --all # [] on a fresh DB is fine
# For the live-Slack path, confirm the listener started:
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
# -> "inbound Slack listener started (Socket Mode, background thread)"
Run all run-team.py commands from ~/orchestrator/agent-team so they hit the
default ledger (state/agent_team.sqlite) and default audit log
(state/audit.log.jsonl). The clarifier is posted to SLACK_CHANNEL_ID.
Resume execution model. A won answer (rowcount 1) enqueues a
ResumeJobonto the coordinator's in-processresume_queue; the daemon drains it on the nexttick(), or the startuprecover()re-drives anyanswered-but-unresumed row after a restart. So a CLIanswer --confirm(or a live Slack answer) flips the ledger row toanswered; the running daemon then resumes the graph. The demo asserts on the durable ledger status.
(a) Crash-safe resume — kill mid-wait, restart, task resumes 🧑 Adam answers
Setup — drive a task to the clarifier wait. With the daemon already running:
python3 run-team.py start --task "demo-a: trivial scoped task"
# prints a thread_id, e.g. 3f2a... (record it as $TID_A)
python3 run-team.py list
# expect one row: status="open", a thread_id, a question_id (record as $QID_A),
# channel_ref set (Slack ts) or null if posting is dry/unavailable.
Kill the box mid-wait, then restart:
sudo systemctl stop agent-team-coordinator.service
# (optionally reboot the VM here for a stronger demonstration)
sudo systemctl start agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e | tail -40 # expect "starting" + recover() sweep
ASSERT — durable state survived the kill (no answer yet):
python3 run-team.py show $QID_A
# PASS: status == "open" (the question was NOT lost across the restart)
🧑 Human gate — Adam answers after the restart (Slack OR CLI):
# Live Slack: Adam replies/clicks in the channel. OR via the CLI:
python3 run-team.py answer $QID_A --answer "scope: just demo, no real change" --confirm
python3 run-team.py show $QID_A
# PASS: status == "answered"
# Wait one tick (~30s, DEFAULT_POLL_INTERVAL) for the daemon to drain the resume:
journalctl -u agent-team-coordinator.service -e | tail -20
PASS criterion (a): the question stayed open across the kill, and an answer
submitted after the restart drove it forward.
(b) Duplicate answer is a no-op 🧑 Adam answers
python3 run-team.py start --task "demo-b: duplicate-answer test"
python3 run-team.py list # record the new question_id as $QID_B (status open)
🧑 First answer (the winning one) — Slack OR CLI:
python3 run-team.py answer $QID_B --answer "first-answer" --via "slack:U1" --confirm
# expect: "answered question <QID_B> (via slack:U1)" (exit 0)
Duplicate / second answer (must lose the compare-and-set):
python3 run-team.py answer $QID_B --answer "second-answer" --via "github:U2" --confirm
# expect: "answer no-op: question <QID_B> was not 'open' ..." on stderr, exit 1
For the live-Slack variant, Adam (or a second authorized owner) answers the
same message twice; the second event loses the CAS and is logged
ignored Slack answer for question_id=... (not open: duplicate ...).
ASSERT — the first answer is preserved:
python3 run-team.py show $QID_B
# PASS: status == "answered"; answered_via == "slack:U1" (the FIRST answer)
PASS criterion (b): first-answer-wins (rowcount==1); the duplicate hits the
BEGIN IMMEDIATE compare-and-set in answer_question and is ignored
(rowcount==0) — no second resume, no overwrite.
(c) Past-deadline answer rejected + task parks 🧑 Adam observes
Code reality:
run-team.py startalways sets the clarifier deadline fromDEFAULT_CLARIFY_DEADLINE = 24h. There is no CLI flag for a short deadline. Two faithful ways to demo expiry without waiting 24h:
Option C1 — force the deadline past, let the daemon's sweep expire it (most faithful)
python3 run-team.py start --task "demo-c: deadline test"
python3 run-team.py list # record $QID_C (status open)
# Set this question's deadline into the past directly in the ledger:
sqlite3 state/agent_team.sqlite \
"UPDATE pending_questions SET deadline_at='2000-01-01T00:00:00+00:00' WHERE question_id='$QID_C';"
# Wait one daemon tick (~30s) for the deadline sweep to flip it, OR observe:
journalctl -u agent-team-coordinator.service -e | tail -20
# expect the park ALARM line: "task parked: clarifier question <QID_C> expired ..."
Option C2 — operator force-expire (if you do not want to touch the DB)
python3 run-team.py start --task "demo-c: deadline test"
python3 run-team.py list # record $QID_C
python3 run-team.py expire $QID_C --confirm # destructive, audit-logged; open->expired
ASSERT — the question is expired and the late answer is rejected:
python3 run-team.py show $QID_C
# PASS: status == "expired"
# 🧑 Adam submits a LATE answer (Slack OR CLI) — it must lose the CAS:
python3 run-team.py answer $QID_C --answer "too-late" --confirm
# expect: "answer no-op: question <QID_C> was not 'open' ..." stderr, exit 1
python3 run-team.py show $QID_C
# PASS: status still "expired"; answer_json still NULL
# The parked task surfaces in the parked view:
python3 run-team.py list --parked # PASS: $QID_C appears here
Deliberate un-park (proves the operator recovery path, §6.6):
python3 run-team.py force-resume $QID_C --confirm
# expect: "force-resume: reopened expired question <QID_C>; it will be re-delivered"
python3 run-team.py show $QID_C
# status flips back to "open" (reopen_question), deadline_at cleared.
PASS criterion (c): the expired question rejects the late answer, the task
parks rather than spins, and the operator can deliberately un-park it via
force-resume (which reopens only an expired row).
force-resumesemantics (verified):run-team.py force-resumereopens anexpiredquestion (the parked case). For anansweredquestion it records intent and reports the recovery sweep will resume it (no mutation). Foropen/superseded/absent it is a no-op exit 1. It does not supersede.
(d) Two concurrent tasks resume independently 🧑 Adam answers
python3 run-team.py start --task "demo-d task A" # record $TID_A2, then:
python3 run-team.py start --task "demo-d task B" # record $TID_B2
python3 run-team.py list
# expect TWO open rows with DISTINCT thread_id AND distinct question_id.
# Record $QID_A2 and $QID_B2 — match by thread_id.
(Optional) restart the daemon first to also show concurrent tasks survive a restart, then answer.
🧑 Answer the SECOND task first, with a distinct answer, then the first (Slack OR CLI):
python3 run-team.py answer $QID_B2 --answer "answer-for-B" --via "slack:U2" --confirm
python3 run-team.py answer $QID_A2 --answer "answer-for-A" --via "slack:U1" --confirm
# Wait one daemon tick (~30s) for both resumes to drain.
ASSERT — each task carries its OWN answer; no cross-talk:
python3 run-team.py show $QID_A2
# PASS: status "answered", answered_via "slack:U1", thread_id == $TID_A2
python3 run-team.py show $QID_B2
# PASS: status "answered", answered_via "slack:U2", thread_id == $TID_B2
python3 run-team.py list --all
# PASS: the two rows resolved on their own thread_id; no cross-contamination.
PASS criterion (d): two concurrently-suspended tasks each resumed to their
own thread_id with their own answer — answering B before A did not misroute,
and the per-thread single-flight guard kept them independent.
Acceptance
P1 is accepted only when (a), (b), (c), and (d) all pass on the live box.
Record the four show outputs (or journalctl excerpts) as evidence. Then
complete the PROVISIONING-RUNBOOK post-session definition-of-done (memory +
Confluence + the mandatory /sh-security-review on the Slack inbound listener).
What can only be verified on the live box
- That the daemon actually drains the resume and the LangGraph checkpoint advances (asserted via ledger status + journalctl; the graph-state advance is observable only on the box).
- The live Slack post/answer round-trip end-to-end (the listener is wired —
D-1 fixed — but the live Socket Mode socket + a real
SLACK_BOT_TOKEN/SLACK_APP_TOKEN/ channel membership are only present on the box). - The exact
channel_refvalue (Slack messagets) — depends on a live Slack post succeeding.