This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/docs/provisioning/P1-DEMO-SCRIPT.md
Adam Moussa 2ed84389a9 docs(agent-team): land provisioning + operator runbooks under docs/provisioning
- PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5
  intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub
  App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked
  FIXED so the demo can use the live Slack answer path.
- P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the
  Slack and operator-CLI answer paths documented for all four exit criteria.
- DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the
  operator-CLI divergence kept as provisioning notes.
- OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks,
  failed HITL resumes, budget exhaustion, transport outages, and
  COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the
  re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6).
2026-06-18 16:45:31 -04:00

12 KiB

P1-DEMO-SCRIPT — live four-criteria acceptance demo (design §3.3.1 / §7.1 P1)

The §7.1 P1 exit gate: demonstrate, on the live box, all four durable human-in-the-loop criteria before P1 is accepted:

  • (a) kill the box mid-wait and have the task resume after restart;
  • (b) submit a duplicate answer and confirm it no-ops;
  • (c) submit an answer after the deadline expired and confirm it is rejected and the task parks;
  • (d) two tasks suspended concurrently resume independently to the correct thread.

Every command below is grounded in the actual code surface (run-team.py, coordinator.py, slack_listener.py, responder.py, resume_worker.py, db/schema.py, graph.py). No invented flags. Where the code does not expose a needed knob (e.g. a short deadline), the script uses a direct sqlite3 write against the documented pending_questions schema and says so.

Two answer paths — live Slack OR the operator CLI

D-1 is fixed: Coordinator.serve() now starts the inbound SlackListener when the live transport is Slack AND SLACK_APP_TOKEN is set, so the live Slack answer round-trip works. You can run the demo either way:

  • Live Slack — Adam clicks the Block Kit button / replies in the channel; the listener normalizes the event, runs the AUTHZ-01 owner check, and drives the first-answer-wins compare-and-set (responder.submit_answer → db.schema.answer_question).
  • Operator CLI — run-team.py answer <qid> --answer ... --confirm runs the identical compare-and-set (audit-logged answer-on-behalf). Useful when the Slack app is not yet provisioned, or to script the assertions.

Both exercise the same durable mechanic; the human gate decision is always Adam's. The assertions below assert on the durable ledger status (the §3.3.1 source of truth) and are identical for either path. The examples use the CLI answer --confirm form so they are copy-pasteable; substitute "Adam answers in Slack" wherever you prefer the live path.

The automated proof of this mechanic is tests/sim/test_p1_exit_criteria.py (a SimPipeline harness) and tests/test_coordinator.py (the serve/listener wiring). This script is the live-box demonstration on top of that.

Verb map (use run-team.py, not operator_cli.py)

Need run-team.py verb Notes
start a task to the human gate start --task "..." [--transport slack] [--dry-run] mints a thread_id, posts the clarifier, writes the open ledger row
list waiting questions list (default open) / list --all / list --parked JSON rows
inspect one row show <question_id> JSON row
answer on the task's behalf answer <question_id> --answer <payload> --confirm destructive, audit-logged; first-answer-wins CAS
force-expire an open question expire <question_id> --confirm destructive; flips open→expired
un-park (reopen) an expired question force-resume <question_id> --confirm only acts on expired rows (reopens them)

operator_cli.py has different verbs (force-expire, answer-on-behalf) and no show, and requires --db/--audit-log — do not use it here.

Preconditions

ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team
. .venv/bin/activate                 # so run-team.py uses the venv deps
# Confirm the ledger exists (Step 5 of the runbook):
python3 run-team.py list --all       # [] on a fresh DB is fine
# For the live-Slack path, confirm the listener started:
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
#   -> "inbound Slack listener started (Socket Mode, background thread)"

Run all run-team.py commands from ~/orchestrator/agent-team so they hit the default ledger (state/agent_team.sqlite) and default audit log (state/audit.log.jsonl). The clarifier is posted to SLACK_CHANNEL_ID.

Resume execution model. A won answer (rowcount 1) enqueues a ResumeJob onto the coordinator's in-process resume_queue; the daemon drains it on the next tick(), or the startup recover() re-drives any answered-but-unresumed row after a restart. So a CLI answer --confirm (or a live Slack answer) flips the ledger row to answered; the running daemon then resumes the graph. The demo asserts on the durable ledger status.


(a) Crash-safe resume — kill mid-wait, restart, task resumes 🧑 Adam answers

Setup — drive a task to the clarifier wait. With the daemon already running:

python3 run-team.py start --task "demo-a: trivial scoped task"
# prints a thread_id, e.g.  3f2a...   (record it as $TID_A)
python3 run-team.py list
# expect one row: status="open", a thread_id, a question_id (record as $QID_A),
# channel_ref set (Slack ts) or null if posting is dry/unavailable.

Kill the box mid-wait, then restart:

sudo systemctl stop agent-team-coordinator.service
# (optionally reboot the VM here for a stronger demonstration)
sudo systemctl start agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e | tail -40   # expect "starting" + recover() sweep

ASSERT — durable state survived the kill (no answer yet):

python3 run-team.py show $QID_A
# PASS: status == "open"  (the question was NOT lost across the restart)

🧑 Human gate — Adam answers after the restart (Slack OR CLI):

# Live Slack: Adam replies/clicks in the channel. OR via the CLI:
python3 run-team.py answer $QID_A --answer "scope: just demo, no real change" --confirm
python3 run-team.py show $QID_A
# PASS: status == "answered"
# Wait one tick (~30s, DEFAULT_POLL_INTERVAL) for the daemon to drain the resume:
journalctl -u agent-team-coordinator.service -e | tail -20

PASS criterion (a): the question stayed open across the kill, and an answer submitted after the restart drove it forward.


(b) Duplicate answer is a no-op 🧑 Adam answers

python3 run-team.py start --task "demo-b: duplicate-answer test"
python3 run-team.py list           # record the new question_id as $QID_B (status open)

🧑 First answer (the winning one) — Slack OR CLI:

python3 run-team.py answer $QID_B --answer "first-answer" --via "slack:U1" --confirm
# expect: "answered question <QID_B> (via slack:U1)"  (exit 0)

Duplicate / second answer (must lose the compare-and-set):

python3 run-team.py answer $QID_B --answer "second-answer" --via "github:U2" --confirm
# expect: "answer no-op: question <QID_B> was not 'open' ..." on stderr, exit 1

For the live-Slack variant, Adam (or a second authorized owner) answers the same message twice; the second event loses the CAS and is logged ignored Slack answer for question_id=... (not open: duplicate ...).

ASSERT — the first answer is preserved:

python3 run-team.py show $QID_B
# PASS: status == "answered"; answered_via == "slack:U1" (the FIRST answer)

PASS criterion (b): first-answer-wins (rowcount==1); the duplicate hits the BEGIN IMMEDIATE compare-and-set in answer_question and is ignored (rowcount==0) — no second resume, no overwrite.


(c) Past-deadline answer rejected + task parks 🧑 Adam observes

Code reality: run-team.py start always sets the clarifier deadline from DEFAULT_CLARIFY_DEADLINE = 24h. There is no CLI flag for a short deadline. Two faithful ways to demo expiry without waiting 24h:

Option C1 — force the deadline past, let the daemon's sweep expire it (most faithful)

python3 run-team.py start --task "demo-c: deadline test"
python3 run-team.py list                       # record $QID_C (status open)

# Set this question's deadline into the past directly in the ledger:
sqlite3 state/agent_team.sqlite \
  "UPDATE pending_questions SET deadline_at='2000-01-01T00:00:00+00:00' WHERE question_id='$QID_C';"

# Wait one daemon tick (~30s) for the deadline sweep to flip it, OR observe:
journalctl -u agent-team-coordinator.service -e | tail -20
# expect the park ALARM line: "task parked: clarifier question <QID_C> expired ..."

Option C2 — operator force-expire (if you do not want to touch the DB)

python3 run-team.py start --task "demo-c: deadline test"
python3 run-team.py list                       # record $QID_C
python3 run-team.py expire $QID_C --confirm    # destructive, audit-logged; open->expired

ASSERT — the question is expired and the late answer is rejected:

python3 run-team.py show $QID_C
# PASS: status == "expired"

# 🧑 Adam submits a LATE answer (Slack OR CLI) — it must lose the CAS:
python3 run-team.py answer $QID_C --answer "too-late" --confirm
# expect: "answer no-op: question <QID_C> was not 'open' ..." stderr, exit 1
python3 run-team.py show $QID_C
# PASS: status still "expired"; answer_json still NULL

# The parked task surfaces in the parked view:
python3 run-team.py list --parked              # PASS: $QID_C appears here

Deliberate un-park (proves the operator recovery path, §6.6):

python3 run-team.py force-resume $QID_C --confirm
# expect: "force-resume: reopened expired question <QID_C>; it will be re-delivered"
python3 run-team.py show $QID_C
# status flips back to "open" (reopen_question), deadline_at cleared.

PASS criterion (c): the expired question rejects the late answer, the task parks rather than spins, and the operator can deliberately un-park it via force-resume (which reopens only an expired row).

force-resume semantics (verified): run-team.py force-resume reopens an expired question (the parked case). For an answered question it records intent and reports the recovery sweep will resume it (no mutation). For open/superseded/absent it is a no-op exit 1. It does not supersede.


(d) Two concurrent tasks resume independently 🧑 Adam answers

python3 run-team.py start --task "demo-d task A"     # record $TID_A2, then:
python3 run-team.py start --task "demo-d task B"     # record $TID_B2
python3 run-team.py list
# expect TWO open rows with DISTINCT thread_id AND distinct question_id.
# Record $QID_A2 and $QID_B2 — match by thread_id.

(Optional) restart the daemon first to also show concurrent tasks survive a restart, then answer.

🧑 Answer the SECOND task first, with a distinct answer, then the first (Slack OR CLI):

python3 run-team.py answer $QID_B2 --answer "answer-for-B" --via "slack:U2" --confirm
python3 run-team.py answer $QID_A2 --answer "answer-for-A" --via "slack:U1" --confirm
# Wait one daemon tick (~30s) for both resumes to drain.

ASSERT — each task carries its OWN answer; no cross-talk:

python3 run-team.py show $QID_A2
# PASS: status "answered", answered_via "slack:U1", thread_id == $TID_A2
python3 run-team.py show $QID_B2
# PASS: status "answered", answered_via "slack:U2", thread_id == $TID_B2
python3 run-team.py list --all
# PASS: the two rows resolved on their own thread_id; no cross-contamination.

PASS criterion (d): two concurrently-suspended tasks each resumed to their own thread_id with their own answer — answering B before A did not misroute, and the per-thread single-flight guard kept them independent.


Acceptance

P1 is accepted only when (a), (b), (c), and (d) all pass on the live box. Record the four show outputs (or journalctl excerpts) as evidence. Then complete the PROVISIONING-RUNBOOK post-session definition-of-done (memory + Confluence + the mandatory /sh-security-review on the Slack inbound listener).

What can only be verified on the live box

  • That the daemon actually drains the resume and the LangGraph checkpoint advances (asserted via ledger status + journalctl; the graph-state advance is observable only on the box).
  • The live Slack post/answer round-trip end-to-end (the listener is wired — D-1 fixed — but the live Socket Mode socket + a real SLACK_BOT_TOKEN / SLACK_APP_TOKEN / channel membership are only present on the box).
  • The exact channel_ref value (Slack message ts) — depends on a live Slack post succeeding.