From 2d1dca0804c44736f37abf35b8fd042d30c9543f Mon Sep 17 00:00:00 2001 From: Adam Moussa <166072409+amoussa1229@users.noreply.github.com> Date: Thu, 18 Jun 2026 16:56:21 -0400 Subject: [PATCH] =?UTF-8?q?feat(agent-team):=20deploy-readiness=20?= =?UTF-8?q?=E2=80=94=20serve=20starts=20Slack=20listener=20+=20systemd=20+?= =?UTF-8?q?=20provisioning=20docs=20(#23)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(agent-team): serve() starts the inbound Slack listener (D-1) Coordinator.serve() now constructs and starts the SlackListener concurrently with the tick/drain loop on a background daemon thread, but ONLY when the live transport is a SlackTransport AND SLACK_APP_TOKEN is configured. When Slack is not the transport or the app token is absent, serve() behaves exactly as before (tick/recover only) — Slack is never made mandatory. - New injectable build_listener seam + default_slack_listener_factory sharing the coordinator's own transport, ledger db_path, and resume_queue put. - AUTHZ-01 owner-allowlist + open-status CAS untouched: serve() sources AGENT_TEAM_SLACK_OWNER_IDS in SlackListener.serve, which still fails closed. - SlackListener.close() added for clean Socket Mode teardown on shutdown; serve() stops the listener + joins the thread in a finally. - Tests: start-when-Slack+app-token, no-start otherwise, clean shutdown, idempotent start, serve start/stop around the loop, listener close(). * fix(agent-team): systemd unit loads ~/orchestrator/.env + uses venv python (D-2/D-7) D-2: add EnvironmentFile=-/home/adam/orchestrator/.env (optional '-') so the P2 GPT-4.1 review loop's cross_reviewer sub-process can read the non-Claude provider key once a task reaches REVIEW. Mirrors the sea-haven-secrev unit. D-7: point ExecStart at the agent-team venv interpreter (/home/adam/orchestrator/agent-team/.venv/bin/python) instead of /usr/bin/env python3, which resolved the system interpreter without the installed deps under systemd's PATH. All hardening (NoNewPrivileges / ProtectSystem=full / ProtectHome=read-only / ReadWritePaths) is retained unchanged (locked decision). * docs(agent-team): land provisioning + operator runbooks under docs/provisioning - PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5 intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked FIXED so the demo can use the live Slack answer path. - P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the Slack and operator-CLI answer paths documented for all four exit criteria. - DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the operator-CLI divergence kept as provisioning notes. - OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks, failed HITL resumes, budget exhaustion, transport outages, and COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6). * fix(agent-team): supervise the Slack listener thread — recurring ALARM + respawn sh-security-review (logic) MEDIUM: a crashed listener thread was logged once, then the daemon ran on 'deaf' — posting clarifier questions but receiving no answers, every gate silently parking, process never exiting so systemd Restart=on-failure never fired. serve() now calls _supervise_slack_listener() each pass: when the listener is enabled but its thread is dead, it emits a recurring ERROR ALARM and respawns via the idempotent starter (self-heal). No-op when alive or disabled. +3 tests. (authz detector: wiring clean — AUTHZ-01 fail-closed allowlist + open-status CAS intact, dead listener fails SAFE.) --- agent-team/agent_team/coordinator.py | 215 +++++++++- .../agent_team/transport/slack_listener.py | 34 +- .../systemd/agent-team-coordinator.service | 23 +- agent-team/tests/test_coordinator.py | 232 ++++++++++ agent-team/tests/test_slack_listener.py | 30 ++ docs/provisioning/DEPLOY-AUDIT.md | 187 ++++++++ docs/provisioning/OPERATOR-RUNBOOK.md | 290 +++++++++++++ docs/provisioning/P1-DEMO-SCRIPT.md | 286 +++++++++++++ docs/provisioning/PROVISIONING-RUNBOOK.md | 402 ++++++++++++++++++ 9 files changed, 1686 insertions(+), 13 deletions(-) create mode 100644 docs/provisioning/DEPLOY-AUDIT.md create mode 100644 docs/provisioning/OPERATOR-RUNBOOK.md create mode 100644 docs/provisioning/P1-DEMO-SCRIPT.md create mode 100644 docs/provisioning/PROVISIONING-RUNBOOK.md diff --git a/agent-team/agent_team/coordinator.py b/agent-team/agent_team/coordinator.py index eee73b0..d4f31e0 100644 --- a/agent-team/agent_team/coordinator.py +++ b/agent-team/agent_team/coordinator.py @@ -50,7 +50,9 @@ committed leaves (responder, resume_worker, db.schema, graph, transport). from __future__ import annotations import logging +import os import queue +import threading from datetime import timedelta from pathlib import Path from typing import TYPE_CHECKING, Any, Callable @@ -60,6 +62,7 @@ from agent_team import responder as responder_mod from agent_team.db.schema import connect, init_db from agent_team.resume_worker import ResumeResult, ResumeWorker from agent_team.transport.base import Transport +from agent_team.transport.slack_adapter import SlackTransport if TYPE_CHECKING: # pragma: no cover - typing only from agent_team.task_model import PipelineState @@ -68,6 +71,7 @@ __all__ = [ "Coordinator", "build_verify_wiring", "default_clarify_node_factory", + "default_slack_listener_factory", "gated_build_verify_wiring", ] @@ -124,6 +128,48 @@ CheckpointerFactory = Callable[[Path], Any] # so the deadline policy stays I/O-free in tests (§6.6). AlarmHook = Callable[[str], None] +# A listener factory: builds the inbound Slack Socket Mode listener +# (:class:`~agent_team.transport.slack_listener.SlackListener`) the serve loop +# starts on a background thread when Slack is the live transport and the app +# token is present. Injected so :meth:`Coordinator.serve` is testable with a fake +# listener and no live socket. The default +# (:func:`default_slack_listener_factory`) builds the real listener from the +# coordinator's shared transport / db / resume-queue and the Slack env tokens. +ListenerFactory = Callable[[], Any] + + +def default_slack_listener_factory( + *, + transport: SlackTransport, + db_path: Path, + enqueue_resume: Callable[[Any], None], +) -> Any: + """Build the live :class:`SlackListener` from the coordinator's seams (D-1). + + The production listener shares the coordinator's *own* live ``SlackTransport`` + (so ``parse_answer`` matches the outbound ``post_question`` wiring), the same + durable ledger ``db_path``, and the same ``resume_queue`` ``put`` callable + (the documented slack_listener ⇄ drainer handoff seam — ``coordinator.py`` + module docstring §3.3). The app/bot tokens and the owner allowlist are sourced + from the environment by :meth:`SlackListener.serve` itself + (``SLACK_APP_TOKEN`` / ``SLACK_BOT_TOKEN`` / ``AGENT_TEAM_SLACK_OWNER_IDS``); + the allowlist still FAILS CLOSED if ``AGENT_TEAM_SLACK_OWNER_IDS`` is unset + (AUTHZ-01), so this factory deliberately does not weaken that — it injects no + ``owner_ids`` and lets ``serve`` read + enforce them. + + Imported lazily for the same import-hygiene reason as the clarifier / planner + factories (the listener pulls the transport + responder leaves). + """ + from agent_team.transport.slack_listener import SlackListener + + return SlackListener( + transport, + db_path, + enqueue_resume, + app_token=os.environ.get("SLACK_APP_TOKEN") or None, + bot_token=os.environ.get("SLACK_BOT_TOKEN") or None, + ) + def default_clarify_node_factory() -> Callable[[PipelineState], PipelineState]: """Build the live Claude-backed clarifier node (§3.3, §7.1 P1). @@ -338,6 +384,7 @@ class Coordinator: resume_queue: "queue.Queue[Any] | None" = None, deadline_window: timedelta | None = None, alarm_hook: AlarmHook | None = None, + build_listener: ListenerFactory | None = None, ) -> None: self._db_path = Path(db_path) self._transport = transport @@ -357,6 +404,11 @@ class Coordinator: self._resume_queue: "queue.Queue[Any]" = resume_queue or queue.Queue() self._deadline_window = deadline_window or graph_mod.DEFAULT_CLARIFY_DEADLINE self._alarm_hook = alarm_hook or self._default_alarm_hook + # The inbound Slack listener factory (D-1). Left None, serve() uses the + # default that builds the real SlackListener; tests inject a fake to prove + # the start/no-start/shutdown wiring with no live socket. Only consulted + # by serve() when Slack is the live transport AND the app token is set. + self._build_listener = build_listener # Built by setup(). self._graph: Any = None @@ -364,6 +416,10 @@ class Coordinator: # Retains a context-manager checkpointer (production SQLite saver) so its # __exit__ is not run early; held open for the daemon's lifetime. self._checkpointer_cm: Any = None + # The running inbound listener + its daemon thread (None until serve() + # starts one). Held so serve()'s shutdown can stop it cleanly. + self._listener: Any = None + self._listener_thread: threading.Thread | None = None # ------------------------------------------------------------------ # # Accessors (the shared queue is the slack_listener handoff seam). @@ -814,32 +870,173 @@ class Coordinator: # ------------------------------------------------------------------ # def serve(self, *, poll_interval: timedelta | None = None) -> None: - """Run the live daemon: bind invoker, setup, recover, then tick forever. + """Run the live daemon: bind invoker, setup, recover, start listener, tick. The production entry. Binds the real Claude invoker (:func:`agent_team.invoker.bind_subscription_invoker`) BEFORE :meth:`setup` builds the clarifier node (so the node's ``billing.claude_invoke`` calls hit the live subscription path), runs the - startup :meth:`recover` sweep, then loops calling :meth:`tick` on the + startup :meth:`recover` sweep, **starts the inbound Slack listener** + (:meth:`_maybe_start_slack_listener`) when Slack is the live transport + and the app token is configured, then loops calling :meth:`tick` on the deadline cadence. - The actual Slack inbound feed is the slack_listener's job; the - coordinator exposes :meth:`submit_answer` and the shared - :attr:`resume_queue` for it. This loop owns only the deadline/recovery - maintenance cadence. + The Slack inbound feed (Socket Mode) is the slack_listener's job and runs + concurrently on a background daemon thread; this loop owns the + deadline/recovery maintenance cadence. On shutdown (Ctrl-C / + SIGTERM-driven ``KeyboardInterrupt``) the listener is stopped cleanly via + :meth:`_stop_slack_listener` in a ``finally``. + + Slack is NOT mandatory: when the transport is not the live ``SlackTransport`` + or the app token is absent, no listener is started and ``serve`` behaves + exactly as before (tick/recover only). """ from agent_team.invoker import bind_subscription_invoker bind_subscription_invoker() self.setup() self.recover() + self._maybe_start_slack_listener() interval = (poll_interval or DEFAULT_POLL_INTERVAL).total_seconds() import time - while True: # pragma: no cover - the infinite daemon loop - self.tick() - time.sleep(interval) + try: + while True: # pragma: no cover - the infinite daemon loop + self.tick() + self._supervise_slack_listener() + time.sleep(interval) + finally: + self._stop_slack_listener() + + # ------------------------------------------------------------------ # + # Inbound Slack listener wiring (D-1): start concurrently with the + # tick/drain loop ONLY when Slack is the live transport and the app + # token is configured; stop cleanly on shutdown. + # ------------------------------------------------------------------ # + + def _slack_listener_enabled(self) -> bool: + """Return ``True`` iff the inbound Slack listener should run (D-1 gate). + + Two conditions, both required (and Slack must stay OPTIONAL): + + * the live transport is a :class:`SlackTransport` (the inbound + ``parse_answer`` must match the outbound poster); and + * ``SLACK_APP_TOKEN`` is present in the environment (Socket Mode needs the + app-level token to open the outbound WebSocket). + + If either is false the daemon runs WITHOUT an inbound listener exactly as + before — Slack is never made mandatory. The owner allowlist + (``AGENT_TEAM_SLACK_OWNER_IDS``) is intentionally NOT part of this gate: + the listener fails closed on an empty allowlist (AUTHZ-01), so starting it + unprovisioned safely rejects every answer rather than silently never + hearing Slack. + """ + if not isinstance(self._transport, SlackTransport): + return False + return bool(os.environ.get("SLACK_APP_TOKEN")) + + def _maybe_start_slack_listener(self) -> bool: + """Start the inbound Slack listener on a daemon thread if enabled (D-1). + + Builds the listener via the injected ``build_listener`` factory (or the + :func:`default_slack_listener_factory`, sharing the coordinator's own live + transport, durable ``db_path`` and ``resume_queue`` put), then runs its + blocking :meth:`~agent_team.transport.slack_listener.SlackListener.serve` + on a background **daemon** thread so the Socket Mode socket and the + tick/drain loop run concurrently. Returns ``True`` iff a listener was + started; ``False`` (and no thread) when :meth:`_slack_listener_enabled` + is false — the Slack-optional contract. + + Idempotent: a second call while a listener thread is already alive is a + no-op. + """ + if self._listener_thread is not None and self._listener_thread.is_alive(): + return True + if not self._slack_listener_enabled(): + _LOG.info( + "inbound Slack listener not started (transport is not live Slack " + "or SLACK_APP_TOKEN is unset); daemon runs tick/recover only" + ) + return False + + if self._build_listener is not None: + listener = self._build_listener() + else: + listener = default_slack_listener_factory( + transport=self._transport, + db_path=self._db_path, + enqueue_resume=self._resume_queue.put, + ) + self._listener = listener + thread = threading.Thread( + target=self._run_listener, + name="agent-team-slack-listener", + daemon=True, + ) + self._listener_thread = thread + thread.start() + _LOG.info("inbound Slack listener started (Socket Mode, background thread)") + return True + + def _supervise_slack_listener(self) -> None: + """Detect a dead inbound listener thread, ALARM (recurring), and respawn. + + Called every :meth:`serve` pass. Without this, a crashed listener thread + (:meth:`_run_listener` swallows its exception) leaves the daemon "deaf": + it keeps posting clarifier questions but receives no answers, every gate + silently times out to PARK, and the process stays up so systemd + ``Restart=on-failure`` never fires. This makes that failure LOUD on every + pass (not a one-shot crash log) and self-heals via the idempotent + :meth:`_maybe_start_slack_listener`. No-op when the listener is not + enabled or the thread is alive. + """ + if not self._slack_listener_enabled(): + return + thread = self._listener_thread + if thread is not None and thread.is_alive(): + return + _LOG.error( + "ALARM: inbound Slack listener is DOWN — answers are NOT being " + "received; clarifier gates will park. Respawning the listener." + ) + # Drop the dead handle so the idempotent starter actually respawns. + self._listener_thread = None + self._maybe_start_slack_listener() + + def _run_listener(self) -> None: + """Thread target: run the listener's blocking serve, log on exit. + + Swallows any listener exception into a log line so a listener crash takes + down only the inbound socket (which Restart=on-failure / a manual restart + recovers), never the maintenance loop or the whole process from a + background thread. + """ + try: + self._listener.serve() + except Exception: # noqa: BLE001 - isolate the listener thread + _LOG.exception("inbound Slack listener thread exited with an error") + + def _stop_slack_listener(self) -> None: + """Stop the inbound listener cleanly on daemon shutdown (D-1). + + Best-effort: closes the Socket Mode connection via the listener's + ``close`` (if it exposes one), then joins the daemon thread briefly. A + teardown error never propagates — shutdown must complete regardless. + """ + listener = self._listener + thread = self._listener_thread + self._listener = None + self._listener_thread = None + if listener is not None: + close = getattr(listener, "close", None) + if close is not None: + try: + close() + except Exception: # noqa: BLE001 - shutdown must not raise + _LOG.debug("listener close raised during shutdown; ignoring") + if thread is not None and thread.is_alive(): + thread.join(timeout=5.0) # ------------------------------------------------------------------ # # Internals. diff --git a/agent-team/agent_team/transport/slack_listener.py b/agent-team/agent_team/transport/slack_listener.py index b592a56..d0472dd 100644 --- a/agent-team/agent_team/transport/slack_listener.py +++ b/agent-team/agent_team/transport/slack_listener.py @@ -147,6 +147,10 @@ class SlackListener: # The owner allowlist (AUTHZ-01). An empty set is the fail-closed default: # an unconfigured deploy rejects every answer. self._owner_ids: set[str] = set(owner_ids) if owner_ids else set() + # The live Socket Mode handler, retained by :meth:`serve` so :meth:`close` + # can stop it cleanly on daemon shutdown. ``None`` until ``serve`` opens + # the socket. + self._socket_handler: Any = None def handle_event(self, raw_payload: Any) -> AnswerOutcome | None: """Normalize + submit one inbound event; return its outcome or ``None``. @@ -331,7 +335,35 @@ class SlackListener: def _on_mention(body: Mapping[str, Any]) -> None: _forward(body) - SocketModeHandler(app, self._app_token).start() + handler = SocketModeHandler(app, self._app_token) + self._socket_handler = handler + handler.start() + + def close(self) -> None: + """Best-effort clean stop of the Socket Mode connection (daemon shutdown). + + The coordinator's :meth:`~agent_team.coordinator.Coordinator.serve` calls + this when it shuts the daemon down so the outbound WebSocket is closed + cleanly rather than only dying with the process. Tolerant: if the handler + was never opened, or the SDK's ``close``/``disconnect`` raises, the error + is swallowed — shutdown must never be blocked by a transport teardown + failure. The handler is dropped afterward so a second ``close`` is a + no-op. + """ + handler = self._socket_handler + if handler is None: + return + self._socket_handler = None + # slack_bolt's SocketModeHandler exposes ``close`` (and the underlying + # client a ``disconnect``); try the most specific available, swallow any + # teardown error. + closer = getattr(handler, "close", None) or getattr(handler, "disconnect", None) + if closer is None: + return + try: + closer() + except Exception: # noqa: BLE001 - shutdown must not raise + _LOG.debug("SlackListener.close: handler teardown raised; ignoring") def _extract_sender_id(raw_payload: Mapping[str, Any]) -> str | None: diff --git a/agent-team/systemd/agent-team-coordinator.service b/agent-team/systemd/agent-team-coordinator.service index 20d25d2..3261b3d 100644 --- a/agent-team/systemd/agent-team-coordinator.service +++ b/agent-team/systemd/agent-team-coordinator.service @@ -13,8 +13,10 @@ # systemctl status agent-team-coordinator.service # journalctl -u agent-team-coordinator.service -e -f # -# Secrets come from the EnvironmentFile (leading '-' = optional, no failure if -# absent), ~/secrev.env (mode 600, NOT in git): +# Secrets come from the EnvironmentFile(s) (leading '-' = optional, no failure if +# absent). Two files are loaded, mirroring the sea-haven-secrev unit: +# +# ~/secrev.env (mode 600, NOT in git) - the agent-team runtime keys: # CLAUDE_CODE_OAUTH_TOKEN -> subscription OAuth (from `claude setup-token`). # A raw ANTHROPIC_API_KEY must NOT be set on this # box; it would silently win and meter to API @@ -24,7 +26,18 @@ # REQUIRED for Socket Mode; opens the inbound # WebSocket that receives answers. Without it the # coordinator can post but never hear replies. +# serve() starts the inbound SlackListener only when +# the transport is live Slack AND this token is set. # SLACK_CHANNEL_ID -> target channel for clarifier questions. +# AGENT_TEAM_SLACK_OWNER_IDS -> comma-separated authorized answerer ids +# (AUTHZ-01). The listener FAILS CLOSED if unset. +# +# ~/orchestrator/.env (mode 600, NOT in git) - the P2 review-loop provider key: +# The production serve() wires the GPT-4.1 cross-review loop, which shells the +# local orchestrator run.py -> cross_reviewer once a task reaches REVIEW. That +# sub-process needs the non-Claude provider key from ~/orchestrator/.env (same +# file the sea-haven-secrev unit loads). Optional ('-') so the daemon still +# starts if it is absent; the review path then fails loudly only at REVIEW. [Unit] Description=Sea Haven agent-team Plane-2 coordinator daemon @@ -36,7 +49,11 @@ Type=simple User=adam WorkingDirectory=/home/adam/orchestrator/agent-team EnvironmentFile=-/home/adam/secrev.env -ExecStart=/usr/bin/env python3 run-team.py serve +EnvironmentFile=-/home/adam/orchestrator/.env +# Use the agent-team venv interpreter (where the runtime deps are installed by +# the DEPLOY-R720 / PROVISIONING-RUNBOOK step), NOT the bare system python3 that +# `/usr/bin/env python3` would resolve under systemd's PATH (D-7). +ExecStart=/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve Restart=on-failure RestartSec=5 # Hardening - matches the level the sea-haven-secrev unit relies on, scoped for a diff --git a/agent-team/tests/test_coordinator.py b/agent-team/tests/test_coordinator.py index d0d67c5..0b9439c 100644 --- a/agent-team/tests/test_coordinator.py +++ b/agent-team/tests/test_coordinator.py @@ -602,3 +602,235 @@ def test_setup_with_p2_factories_builds_a_review_node(db_path: Path) -> None: assert graph_mod.REVIEW in coord.graph.get_graph().nodes finally: review_loop._review_invoker = saved + + +# --------------------------------------------------------------------------- # +# Inbound Slack listener wiring in serve() (D-1) +# --------------------------------------------------------------------------- # + + +class _FakeListener: + """A record-only stand-in for SlackListener (no socket, no SDK). + + ``serve`` blocks on an Event until ``close`` is called, mirroring the live + listener whose ``serve`` blocks on the Socket Mode handler until shut down. + Records that ``serve`` ran and that ``close`` was called so the coordinator + start/shutdown wiring can be asserted. + """ + + def __init__(self) -> None: + import threading as _t + + self.served = False + self.closed = False + self._stop = _t.Event() + + def serve(self) -> None: + self.served = True + self._stop.wait(timeout=5.0) + + def close(self) -> None: + self.closed = True + self._stop.set() + + +def _slack_coordinator( + db_path: Path, + *, + build_listener: Any = None, +) -> Coordinator: + """A Coordinator whose live transport IS a SlackTransport (listener gate).""" + from agent_team.transport.slack_adapter import SlackTransport + + saver = _Saver() + return Coordinator( + db_path=db_path, + transport=SlackTransport("C123"), + build_clarify_node=lambda: graph_mod.clarify_node, + build_checkpointer=lambda _path: saver, + build_listener=build_listener, + ) + + +def test_serve_starts_listener_when_slack_and_app_token( + db_path: Path, monkeypatch: pytest.MonkeyPatch +) -> None: + """serve() starts the inbound listener when Slack is the transport AND the + app token is configured, then stops it cleanly on shutdown.""" + monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test") + listener = _FakeListener() + coord = _slack_coordinator(db_path, build_listener=lambda: listener) + coord.setup() + + assert coord._maybe_start_slack_listener() is True + # The background daemon thread actually ran the listener's serve(). + import time + + for _ in range(50): + if listener.served: + break + time.sleep(0.01) + assert listener.served is True + + # Clean shutdown stops the listener and joins the thread. + coord._stop_slack_listener() + assert listener.closed is True + assert coord._listener is None + assert coord._listener_thread is None + + +def test_serve_does_not_start_listener_without_app_token( + db_path: Path, monkeypatch: pytest.MonkeyPatch +) -> None: + """No app token => no inbound listener (Slack stays optional).""" + monkeypatch.delenv("SLACK_APP_TOKEN", raising=False) + listener = _FakeListener() + coord = _slack_coordinator(db_path, build_listener=lambda: listener) + coord.setup() + + assert coord._slack_listener_enabled() is False + assert coord._maybe_start_slack_listener() is False + assert listener.served is False + assert coord._listener is None + assert coord._listener_thread is None + + +def test_serve_does_not_start_listener_when_transport_not_slack( + db_path: Path, monkeypatch: pytest.MonkeyPatch +) -> None: + """A non-Slack transport never starts the listener even with the app token.""" + monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test") + listener = _FakeListener() + saver = _Saver() + coord = Coordinator( + db_path=db_path, + transport=FakeTransport(), # not a SlackTransport + build_clarify_node=lambda: graph_mod.clarify_node, + build_checkpointer=lambda _path: saver, + build_listener=lambda: listener, + ) + coord.setup() + + assert coord._slack_listener_enabled() is False + assert coord._maybe_start_slack_listener() is False + assert listener.served is False + + +def test_maybe_start_listener_is_idempotent( + db_path: Path, monkeypatch: pytest.MonkeyPatch +) -> None: + """A second start while the thread is alive is a no-op (one listener).""" + monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test") + built: list[_FakeListener] = [] + + def _factory() -> _FakeListener: + lst = _FakeListener() + built.append(lst) + return lst + + coord = _slack_coordinator(db_path, build_listener=_factory) + coord.setup() + try: + assert coord._maybe_start_slack_listener() is True + assert coord._maybe_start_slack_listener() is True + assert len(built) == 1 # not rebuilt/restarted + finally: + coord._stop_slack_listener() + + +def test_supervise_respawns_dead_listener_and_alarms( + db_path: Path, monkeypatch: pytest.MonkeyPatch, caplog: pytest.LogCaptureFixture +) -> None: + """A dead listener thread is respawned (self-heal) with a recurring ALARM, so + the daemon never silently goes 'deaf' to Slack answers (D1-silent-stall).""" + import logging + + monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test") + listeners: list[_FakeListener] = [] + + def _factory() -> _FakeListener: + lst = _FakeListener() + listeners.append(lst) + return lst + + coord = _slack_coordinator(db_path, build_listener=_factory) + coord.setup() + try: + assert coord._maybe_start_slack_listener() is True + first = coord._listener_thread + assert first is not None and first.is_alive() + # Kill the listener thread: close() fires the stop event so serve() returns. + listeners[0].close() + first.join(timeout=5.0) + assert not first.is_alive() + # The supervisor detects the death, ALARMs, and respawns a fresh listener. + with caplog.at_level(logging.ERROR): + coord._supervise_slack_listener() + assert any("listener is DOWN" in r.message for r in caplog.records) + second = coord._listener_thread + assert second is not None and second is not first and second.is_alive() + assert len(listeners) == 2 # a new listener was built on respawn + finally: + coord._stop_slack_listener() + + +def test_supervise_is_noop_when_listener_alive( + db_path: Path, monkeypatch: pytest.MonkeyPatch +) -> None: + """The supervisor leaves a healthy listener thread untouched (no churn).""" + monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test") + listener = _FakeListener() + coord = _slack_coordinator(db_path, build_listener=lambda: listener) + coord.setup() + try: + coord._maybe_start_slack_listener() + alive = coord._listener_thread + coord._supervise_slack_listener() + assert coord._listener_thread is alive + finally: + coord._stop_slack_listener() + + +def test_supervise_is_noop_when_listener_disabled( + db_path: Path, monkeypatch: pytest.MonkeyPatch +) -> None: + """No app token -> listener disabled -> supervisor never spawns a thread.""" + monkeypatch.delenv("SLACK_APP_TOKEN", raising=False) + coord = _slack_coordinator(db_path, build_listener=_FakeListener) + coord.setup() + coord._supervise_slack_listener() + assert coord._listener_thread is None + + +def test_stop_listener_is_safe_when_none(db_path: Path) -> None: + """Stopping with no listener running never raises (clean-shutdown safety).""" + coord = _slack_coordinator(db_path) + coord._stop_slack_listener() # must be a quiet no-op + assert coord._listener is None + + +def test_serve_wires_start_and_stop_around_the_loop( + db_path: Path, monkeypatch: pytest.MonkeyPatch +) -> None: + """serve() starts the listener before the loop and stops it in finally even + when the loop is interrupted (KeyboardInterrupt).""" + monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test") + listener = _FakeListener() + coord = _slack_coordinator(db_path, build_listener=lambda: listener) + + # Neuter the live invoker bind so serve() needs no Claude SDK / token. + # serve() does `from agent_team.invoker import bind_subscription_invoker`, + # so patch the name on the invoker module it imports from. + import agent_team.invoker as invoker_mod + + monkeypatch.setattr(invoker_mod, "bind_subscription_invoker", lambda: None) + # Break the infinite loop on the first tick. + monkeypatch.setattr( + coord, "tick", lambda: (_ for _ in ()).throw(KeyboardInterrupt()) + ) + + with pytest.raises(KeyboardInterrupt): + coord.serve(poll_interval=timedelta(seconds=0)) + + assert listener.served is True # started before the loop + assert listener.closed is True # stopped in finally on interrupt diff --git a/agent-team/tests/test_slack_listener.py b/agent-team/tests/test_slack_listener.py index 10eaf85..f1d04ff 100644 --- a/agent-team/tests/test_slack_listener.py +++ b/agent-team/tests/test_slack_listener.py @@ -449,6 +449,36 @@ def test_serve_requires_tokens(db_path: Path) -> None: listener.serve() +def test_close_without_open_handler_is_noop(db_path: Path) -> None: + """close() before serve() ever opened a socket is a quiet no-op.""" + listener = _listener(db_path, RecordingQueue()) + listener.close() # must not raise + listener.close() # idempotent + + +def test_close_stops_socket_handler(db_path: Path) -> None: + """close() invokes the retained Socket Mode handler's close and drops it.""" + + class _FakeHandler: + def __init__(self) -> None: + self.closed = False + + def close(self) -> None: + self.closed = True + + listener = _listener(db_path, RecordingQueue()) + handler = _FakeHandler() + listener._socket_handler = handler + listener.close() + assert handler.closed is True + assert listener._socket_handler is None + # A teardown that raises is swallowed (shutdown must never raise). + listener._socket_handler = type( + "_Boom", (), {"close": lambda self: (_ for _ in ()).throw(RuntimeError("x"))} + )() + listener.close() # no exception + + def test_real_slack_transport_is_a_transport() -> None: """Sanity: the injected SlackTransport is the contract the listener expects.""" assert isinstance(SlackTransport(channel="C123"), Transport) diff --git a/docs/provisioning/DEPLOY-AUDIT.md b/docs/provisioning/DEPLOY-AUDIT.md new file mode 100644 index 0000000..91bdb39 --- /dev/null +++ b/docs/provisioning/DEPLOY-AUDIT.md @@ -0,0 +1,187 @@ +# DEPLOY-AUDIT — `agent-team/DEPLOY-R720.md` + systemd unit vs. actual code + +Cross-check of the drafted deploy doc (`agent-team/DEPLOY-R720.md`) and the +systemd unit (`agent-team/systemd/agent-team-coordinator.service`) against the +**actual current code** in `agent-team/agent_team/**` and `agent-team/run-team.py`. + +Each finding: **location → claimed → actual → fix**. Severity: 🔴 blocker / +🟠 should-fix / 🟡 nit. Items confirmed clean are stated explicitly. + +> **Status (deploy-readiness PR):** the three deploy-correctness bugs that +> blocked the live provisioning session — **D-1, D-2, D-7** — are **RESOLVED** in +> this PR (`feature/agent-team-deploy-readiness`). The remaining items +> (**D-4 / D-5** dependency pinning, the **operator-CLI divergence**) are kept +> below as **provisioning notes** — they do not block the coordinator deploy. + +--- + +## ✅ D-1 — `serve` now starts the Slack inbound listener — RESOLVED (this PR) + +- **Location:** `agent_team/coordinator.py` (`Coordinator.serve` + + `_maybe_start_slack_listener` / `_slack_listener_enabled` / `_stop_slack_listener` + / `default_slack_listener_factory`); `agent_team/transport/slack_listener.py` + (`SlackListener.serve` + new `close`). +- **Was:** `Coordinator.serve()` did only `bind_subscription_invoker()`, + `setup()`, `recover()`, then an infinite tick/sleep loop — it never constructed + or started `SlackListener`, so a deployed daemon posted clarifier questions and + expired them on deadline but could **not hear Slack answers**. +- **Now:** `serve()` starts the `SlackListener` on a background **daemon thread**, + concurrently with the tick/drain loop, **when** the live transport is a + `SlackTransport` AND `SLACK_APP_TOKEN` is set. It shares the coordinator's own + transport, ledger `db_path`, and `resume_queue` put; on shutdown it calls the + listener's new `close()` and joins the thread in a `finally`. When Slack is not + the transport or the app token is absent, no listener starts and `serve` behaves + exactly as before — **Slack is never made mandatory**. The AUTHZ-01 owner + allowlist + the open-status compare-and-set are untouched (still fail closed on + an empty `AGENT_TEAM_SLACK_OWNER_IDS`). +- **Tests:** `tests/test_coordinator.py` — start-when-Slack+app-token, no-start + without the token, no-start when the transport is not Slack, idempotent start, + clean shutdown, serve start/stop around the loop; `tests/test_slack_listener.py` + — `close()` no-op + handler teardown. + +--- + +## ✅ D-2 — coordinator unit loads `~/orchestrator/.env` — RESOLVED (this PR) + +- **Location:** `agent-team/systemd/agent-team-coordinator.service`; + `run-team.py:_build_coordinator` (always wires `default_review_wiring`); + `coordinator.py:default_review_wiring` → `review_loop_llm.default_plan_reviewer` + (shells the local orchestrator `run.py` → `cross_reviewer` GPT-4.1). +- **Was:** the unit loaded only `~/secrev.env`. `run-team.py serve` builds the + coordinator with the P2 review loop wired, and once a task reaches REVIEW the + reviewer shells the orchestrator `run.py`, whose GPT-4.1 call reads the + non-Claude provider key from `~/orchestrator/.env`. The review path would fail + to authenticate. +- **Now:** the unit adds `EnvironmentFile=-/home/adam/orchestrator/.env` + (optional `-`, mirroring the `sea-haven-secrev` unit). Does not affect the P1 + demo (P1 stops at PLAN before REVIEW); closes the latent P2 break. + +--- + +## 🟠 D-4 — pip install list omits `requests` (provisioning note) + +- **Location:** `agent_team/transport/github_live.py` / `github_intake.py` + (`import requests`). +- **Actual:** the GitHub transport + intake require `requests`; the Slack-first + path does not hit it, but any `--transport github` / `intake-github` use fails + with a clear RuntimeError without it. +- **Fix:** PROVISIONING-RUNBOOK Step 4 installs `requests` into the venv. + `requests` is also absent from `requirements.txt` (see D-5). + +--- + +## 🟠 D-5 — agent-team runtime deps are not pinned in `requirements.txt` (provisioning note) + +- **Location:** `requirements.txt` (repo root). +- **Actual:** `requirements.txt` pins `langgraph==1.1.10` and + `langgraph-checkpoint-sqlite==3.1.0`, but the agent-team runtime deps + `claude-agent-sdk`, `slack_sdk`, `slack_bolt`, `requests` (and `anthropic` for + api mode) are **not in `requirements.txt` at all** — they are installed ad-hoc + into the agent-team venv by the runbook. There is no pinned, reproducible source + of truth for the box's runtime set. +- **Fix (deferred):** add an `agent-team/requirements.txt` (or extras group) + pinning these, version-matched to the root `requirements.txt` langgraph pin. + Until then, PROVISIONING-RUNBOOK Step 4 pins `langgraph==1.1.10` / + `langgraph-checkpoint-sqlite==3.1.0` explicitly so the unpinned `pip install` + cannot pull a newer, untested major. **Do NOT modify `requirements.txt` or the + checkers in this PR** (out of scope). + +--- + +## ✅ D-6 — `slack_bolt` is now exercised by the daemon — RESOLVED (consequence of D-1) + +- **Location:** `slack_listener.py:serve` (the only `slack_bolt` import). +- **Now:** with D-1 fixed, `Coordinator.serve()` starts `SlackListener.serve()`, + which imports + uses `slack_bolt` for the Socket Mode handler. The dep is right + and now actually exercised on the live Slack path. + +--- + +## ✅ D-SLACKVAR (clean) — `SLACK_CHANNEL_ID` matches + +- `run-team.py:_build_transport` reads exactly `os.environ.get("SLACK_CHANNEL_ID")`. + The runbook, the unit comment, and the code all use `SLACK_CHANNEL_ID` (not + `SLACK_CHANNEL`). **CLEAN.** + +--- + +## ✅ D-ENV-SLACKBOT / OAUTH / OWNERS / APPTOKEN (clean) — names match + +- **`SLACK_BOT_TOKEN`** ↔ `slack_live.py` + `slack_listener` env read. **CLEAN.** +- **`CLAUDE_CODE_OAUTH_TOKEN`** ↔ `invoker.py`. **CLEAN.** +- **`AGENT_TEAM_SLACK_OWNER_IDS`** ↔ `slack_listener.py` (name + fail-closed + semantics). **CLEAN** — now read by the running daemon (D-1 fixed). +- **`SLACK_APP_TOKEN`** — now read in two places: `coordinator._slack_listener_enabled` + gates the listener on its presence, and `default_slack_listener_factory` / + `SlackListener.serve` source it to open the socket. **CLEAN** (read site exists + now that D-1 is fixed). + +--- + +## ✅ D-7 — `ExecStart` uses the venv interpreter — RESOLVED (this PR) + +- **Location:** `agent-team/systemd/agent-team-coordinator.service` ExecStart. +- **Was:** `ExecStart=/usr/bin/env python3 run-team.py serve` resolved the + **system** interpreter under systemd's PATH — not the venv where the deps were + installed, so the daemon would fail at import. +- **Now:** `ExecStart=/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve` + (matches the runbook venv path + `WorkingDirectory`). + +--- + +## ✅ D-SUBCMD (mostly clean) — run-team.py subcommands referenced exist + +Cross-checked every `run-team.py ` the deploy doc + demo name against +`run-team.py:build_parser`: `init-db`, `serve`, `list` (+ `--all` / `--parked`), +`show`, `expire`, `answer`, `redeliver`, `supersede`, `force-resume`, `start`, +`intake-github`, `intake-checker`, `fix` — all exist. No invented verbs. + +> ### Operator-CLI divergence (provisioning note) +> +> Two operator CLIs exist with **different verb names**: +> +> - `run-team.py` (the entry CLI): `init-db, list, show, redeliver, expire, +> answer, supersede, force-resume, start, serve, intake-github, intake-checker, +> fix`. Has `show` and `--parked`; `--db` / `--audit-log` default sensibly. +> - `agent_team/operator_cli.py`: `list, redeliver, force-expire, +> answer-on-behalf, force-resume` — **no `show`**, `--db` / `--audit-log` are +> **required**, and its `force-resume` **supersedes** (unlike `run-team.py`'s, +> which reopens an expired row and never supersedes). +> +> **Use `run-team.py` for provisioning + the demo + incident recovery.** The +> docs reference only `run-team.py`. Reconciling the two CLIs is a follow-up. + +--- + +## ✅ D-PYTHONPKG (clean) — package import bootstrap is correct + +`run-team.py` inserts its own dir into `sys.path` so the hyphenated script +imports the `agent_team` package without an editable install. **CLEAN.** + +--- + +## ✅ D-HARDENING (clean, and matches the locked decision) + +- Unit: `NoNewPrivileges=true`, `ProtectSystem=full`, `ProtectHome=read-only`, + `ReadWritePaths=/home/adam/orchestrator/agent-team/state`. **Retained unchanged** + in this PR (locked decision — do not revert to secrev parity). +- The `ReadWritePaths` carve-out matches the ledger + audit-log location + (`state/agent_team.sqlite`, `state/audit.log.jsonl`). **CLEAN.** + +--- + +## Summary table + +| ID | Sev | Status | One-line | +|---|---|---|---| +| D-1 | 🔴 | ✅ RESOLVED (PR) | `serve` starts `SlackListener` (Slack + app-token gated; Slack stays optional) | +| D-2 | 🔴 | ✅ RESOLVED (PR) | unit loads `~/orchestrator/.env` for the P2 GPT-4.1 review provider key | +| D-7 | 🟡 | ✅ RESOLVED (PR) | `ExecStart` points at the agent-team venv interpreter | +| D-6 | 🟡 | ✅ RESOLVED | `slack_bolt` now exercised by the daemon (consequence of D-1) | +| D-4 | 🟠 | NOTE | pip list omits `requests` — runbook Step 4 installs it | +| D-5 | 🟠 | NOTE | agent-team runtime deps not pinned in `requirements.txt` — runbook pins langgraph | +| operator-CLI | — | NOTE | `run-team.py` vs `operator_cli.py` divergent verbs — use `run-team.py` | +| D-SLACKVAR | ✅ | CLEAN | `SLACK_CHANNEL_ID` matches everywhere | +| D-ENV-* | ✅ | CLEAN | bot/oauth/owner/app-token env names match; all read sites now exist | +| D-SUBCMD | ✅ | CLEAN | every `run-team.py` verb/flag the docs cite exists | +| D-HARDENING | ✅ | CLEAN | unit hardening retained unchanged (locked decision) | diff --git a/docs/provisioning/OPERATOR-RUNBOOK.md b/docs/provisioning/OPERATOR-RUNBOOK.md new file mode 100644 index 0000000..806b583 --- /dev/null +++ b/docs/provisioning/OPERATOR-RUNBOOK.md @@ -0,0 +1,290 @@ +# OPERATOR-RUNBOOK — R720 agent-team coordinator incident handling + +The on-call runbook for the always-on `agent-team-coordinator` daemon on the +`sh-secrev` R720 VM. Covers pipeline stalls, stuck/parked tasks, failed +human-in-the-loop resumes, budget exhaustion mid-pipeline, transport outages, and +the COMPLACENCY / COVERAGE alarms. Grounds every recovery in real code (design +§5 escalation ladder + §6.6 contention/park policy; Phase-6 requirement). + +> **CLI used throughout: `run-team.py`** (the entry CLI), run from +> `~/orchestrator/agent-team` with the venv active so it hits the default ledger +> (`state/agent_team.sqlite`) and audit log (`state/audit.log.jsonl`). Do **not** +> use `agent_team/operator_cli.py` — it has divergent verbs (no `show`, required +> `--db`/`--audit-log`, and a `force-resume` that supersedes). See DEPLOY-AUDIT.md. + +```bash +ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120 +cd ~/orchestrator/agent-team && . .venv/bin/activate +``` + +## Verb reference (all confirmed in `run-team.py:build_parser`) + +| Verb | Effect | Destructive? | +|---|---|---| +| `list` / `list --all` / `list --parked` / `list --status ` | read pending questions (default `open`) | no | +| `show ` | print one ledger row (JSON) | no | +| `redeliver ` | clear `channel_ref` so the reconcile loop re-posts an `open` question | no (audit-logged) | +| `expire --confirm` | force `open`→`expired` | yes | +| `answer --answer

[--via ] --confirm` | answer-on-behalf (first-answer-wins CAS) | yes | +| `force-resume --confirm` | reopen an `expired` (parked) question; for `answered` records resume intent | yes | +| `supersede --confirm` | mark a stale `open`/`answered` row `superseded` | yes | +| `start --task "..." [--transport ...] [--dry-run]` | start one task to the human gate | no | + +`--operator ` (global) sets the audit attribution; it defaults to the OS +login. Destructive verbs require `--confirm` and write an attempt-then-outcome +record to the audit log **before** mutating. + +First triage for any incident: + +```bash +systemctl status agent-team-coordinator.service +journalctl -u agent-team-coordinator.service -e --since "-2h" | tail -100 +python3 run-team.py list --all # full ledger snapshot +python3 run-team.py list --parked # non-open rows (parked-task context) +``` + +--- + +## Incident 1 — Pipeline stall (tasks not advancing) + +**Symptoms:** `list` shows `open` (or `answered`) rows that never progress; the +journal shows no `tick` activity or repeated errors. + +**Diagnose:** +```bash +systemctl is-active agent-team-coordinator.service # "active" expected +journalctl -u agent-team-coordinator.service -e | tail -60 +# Daemon dead/looping on restart? Check the unit + deps: +.venv/bin/python -c "import langgraph, slack_sdk, slack_bolt, requests; print('deps ok')" +``` + +**Recover:** +1. If the daemon is dead and `Restart=on-failure` is flapping, read the journal + for the import/auth error. A missing venv dep (D-4/D-5) or a missing + `CLAUDE_CODE_OAUTH_TOKEN` is the usual cause — fix the env / venv, then: + ```bash + sudo systemctl restart agent-team-coordinator.service + ``` +2. The startup `recover()` sweep re-drives `answered`-but-unresumed rows and + re-posts `open` rows that lost their `channel_ref`, so a clean restart + converges from the durable ledger. Confirm with `journalctl ... | tail` and + `list --all`. +3. A single task stuck `open` with a stale/missing post: re-deliver it. + ```bash + python3 run-team.py show + python3 run-team.py redeliver # clears channel_ref; reconcile re-posts + ``` + +**Escalate** (see the ladder) if a restart does not clear it within one tick +cadence and the journal shows a non-transient error. + +--- + +## Incident 2 — Stuck / parked task + +A task parks (design §6.6) when its clarifier question **expires** with no answer, +or its budget headroom drops below reserve, or a checkpoint is corrupt. Parked = +the durable ledger row is no longer `open` (it is `expired`), and an ALARM was +raised, not spun on. + +**Diagnose:** +```bash +python3 run-team.py list --parked +python3 run-team.py show # status, thread_id, deadline_at, answered_via +``` + +**Recover — depends on why it parked:** + +- **Expired with no answer (the common case)** — un-park by reopening the expired + question; it is then re-delivered for an answer: + ```bash + python3 run-team.py force-resume --confirm + # "force-resume: reopened expired question ; it will be re-delivered" + python3 run-team.py show # status flips back to "open" + ``` + Then answer it (Slack or CLI) to drive it forward. + +- **Answered but not yet resumed** — the recovery sweep handles it; `force-resume` + records intent and reports that (no mutation): + ```bash + python3 run-team.py force-resume --confirm + # "force-resume: question is answered and pending resume; the recovery + # sweep will resume it (intent recorded)" + sudo systemctl restart agent-team-coordinator.service # forces the recover() sweep now + ``` + +- **Stale / wrong question that should be abandoned** — supersede it so it stops + surfacing as parked context: + ```bash + python3 run-team.py supersede --confirm + ``` + +`MAX_PARK` FIFO-aging: a task that exceeds the park window escalates (ALARM + a +Jira ticket per the ladder) rather than starving silently. + +--- + +## Incident 3 — Failed human-in-the-loop resume + +**Symptoms:** an answer was submitted (Slack or CLI) but the graph did not +advance. + +**Diagnose:** +```bash +python3 run-team.py show # is status "answered"? what answered_via? +journalctl -u agent-team-coordinator.service -e | grep -iE "resume|answer|" +``` + +**Recover:** +- **Status is `answered` but no resume drained** — the resume queue is in-process; + a daemon restart triggers the `recover()` sweep that re-drives `answered` rows + via the turn-guarded `ResumeWorker` (idempotent — an already-advanced thread + supersedes-and-skips): + ```bash + sudo systemctl restart agent-team-coordinator.service + python3 run-team.py show # confirm it advanced + ``` +- **Live Slack answer never registered** — the listener fails closed. Check: + ```bash + journalctl -u agent-team-coordinator.service -e | grep -i "Slack answer" + # "rejecting Slack answer: owner allowlist is unconfigured ..." -> set + # AGENT_TEAM_SLACK_OWNER_IDS in ~/secrev.env and restart. + # "rejecting Slack answer ... unauthorized sender" -> the answerer's user id is + # not in the allowlist. Add it, or answer via the CLI on their behalf: + python3 run-team.py answer --answer "" --via "cli:adam" --confirm + ``` +- **Answer lost the compare-and-set (`not open`)** — the row was already + answered/expired/superseded. Inspect with `show`; if it parked, go to Incident 2. + +--- + +## Incident 4 — Budget exhaustion mid-pipeline + +The shared Claude budget ledger (`budget_ledger` table) enforces per-call + total +nightly caps (design §6.1, §6.6). A stage that would breach the reserve **parks** +the task (deferred, not dropped) and ALARMs — it never loop-drains the pool. + +**Diagnose:** +```bash +journalctl -u agent-team-coordinator.service -e | grep -iE "budget|reserve|park" +# Inspect today's spend directly (no CLI verb for the budget ledger; read it). +# Columns (agent_team/db/schema.py budget_ledger DDL): thread_id, stage, model, +# billing_mode, input_tokens, output_tokens, usd_cost, recorded_at, day_bucket. +sqlite3 state/agent_team.sqlite \ + "SELECT day_bucket, thread_id, ROUND(SUM(usd_cost),4) AS usd, SUM(input_tokens) AS in_tok + FROM budget_ledger GROUP BY day_bucket, thread_id ORDER BY day_bucket DESC LIMIT 20;" +``` + +**Recover:** +- **Wait for the next budget window** — budget-exhausted roles/tasks are deferred + via the rotation pointer and picked up next cycle; this is the intended + behavior, not a failure. The parked task surfaces in `list --parked`. +- **Force a specific parked task forward now** (e.g. it is urgent and headroom has + since freed): un-park it and let the daemon re-run the stage within the + remaining cap: + ```bash + python3 run-team.py force-resume --confirm + ``` +- **Confirm `ANTHROPIC_API_KEY` is absent** — its presence would silently meter to + API rates and blow the budget model (the billing seam pops it defensively, but + it must not be set): + ```bash + grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # both must print 0 + ``` + +Do **not** raise the cap to push a task through without Adam's decision — budget +realism is a deliberate guardrail. Escalate per the ladder if a task repeatedly +parks on budget across nights. + +--- + +## Incident 5 — Transport outage (Slack / GitHub down or misconfigured) + +**Symptoms:** clarifier posts fail; the journal shows transport errors or the +inbound listener is not started. + +**Diagnose:** +```bash +journalctl -u agent-team-coordinator.service -e | grep -iE "Slack|listener|post|transport" +# Did the inbound listener start? +journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener" +# "started (Socket Mode ...)" -> inbound up +# "not started (... SLACK_APP_TOKEN is unset)" -> Slack inbound intentionally off +``` + +**Recover:** +- **Outbound post failed (Slack API / network)** — the ledger row stays `open` + with no `channel_ref` (the responder leaves it for reconcile). After the + transport recovers, the `recover()` sweep (on restart) or a manual `redeliver` + re-posts: + ```bash + python3 run-team.py redeliver + ``` +- **Inbound listener down / never started** — the daemon still posts and expires; + it just cannot hear Slack. **The CLI `answer` path is the outage fallback** — + it runs the identical compare-and-set with no socket: + ```bash + python3 run-team.py answer --answer "" --via "cli:adam" --confirm + ``` + Fix `SLACK_APP_TOKEN` / `SLACK_BOT_TOKEN` in `~/secrev.env`, then + `sudo systemctl restart agent-team-coordinator.service` to re-establish Socket + Mode. (D-1: the listener starts only when the transport is live Slack AND + `SLACK_APP_TOKEN` is set; it is optional by design.) +- **Slack listener crash-looping** — the listener runs on an isolated daemon + thread; a crash is logged (`inbound Slack listener thread exited with an error`) + and takes down only the inbound socket, not the maintenance loop. Restart the + service to relaunch the thread once the cause is fixed. + +--- + +## Incident 6 — COMPLACENCY / COVERAGE alarms (Plane-1 checkers) + +These come from the nightly checker run, not the coordinator daemon (design §6.4, +§6.6): + +- **COMPLACENCY ALARM** — a checker role missed a planted canary fault. That role + is **skipped** for the night (it never runs silently degraded). Recover: inspect + the canary corpus / the role's checker under `security-review/checkers/`, fix the + regression, re-run that checker's dry-run, and confirm the canary passes before + re-enabling. +- **COVERAGE ALARM** — a role slipped its rotation slot (e.g. budget-deferred). It + is **deferred via the rotation pointer, never dropped**, and picked up next + cycle. Recover: confirm the rotation pointer advanced (it is rebuildable from + report history) and that the role runs on the next cadence; investigate only if + it slips repeatedly. + +Both alarms are **report + ALARM-only** (design D3): nothing posts on a clean +state, no auto-Jira/Notion writes from the checker itself. Persistent alarms +follow the escalation ladder below. + +--- + +## Escalation ladder (design §5 / §6.6, resolves Q3) + +Anything that does not clear on the first ALARM escalates — but **ALARM-only in +spirit** (nothing posts on a clean state): + +1. **Re-alarm on a backoff.** A confirmed critical (or a COMPLACENCY / COVERAGE + alarm) that persists re-alarms to Slack each night it is still unresolved, on + a backoff so it does not spam. +2. **Open a Jira tracking ticket after `N` nights** (default **N = 3**). If the + condition still has not cleared, the coordinator opens an **INFRA** Jira ticket + so it cannot quietly linger. The same ladder applies to a parked task that + exceeds `MAX_PARK`. +3. **Human (Adam) takes it from the Jira ticket.** For a parked task, recover via + Incidents 2–4 above; for a checker alarm, via Incident 6. + +When you resolve an incident, record the action — destructive CLI verbs already +write an attributable attempt+outcome record to `state/audit.log.jsonl`; for +non-CLI recoveries note it on the Jira ticket. + +--- + +## Post-incident + +- Confirm the daemon is `active` and the ledger has no unexpected parked rows + (`list --parked`). +- If you restored the VM snapshot or wiped the ledger, re-run the relevant + PROVISIONING-RUNBOOK steps. +- Update `project_r720_agent_team` memory if the incident revealed a durable + fact (a new failure mode, a config that must change). diff --git a/docs/provisioning/P1-DEMO-SCRIPT.md b/docs/provisioning/P1-DEMO-SCRIPT.md new file mode 100644 index 0000000..aaafbd3 --- /dev/null +++ b/docs/provisioning/P1-DEMO-SCRIPT.md @@ -0,0 +1,286 @@ +# P1-DEMO-SCRIPT — live four-criteria acceptance demo (design §3.3.1 / §7.1 P1) + +The §7.1 P1 exit gate: demonstrate, on the live box, all four durable +human-in-the-loop criteria before P1 is accepted: + +- **(a)** kill the box mid-wait and have the task resume after restart; +- **(b)** submit a duplicate answer and confirm it no-ops; +- **(c)** submit an answer after the deadline expired and confirm it is rejected + and the task parks; +- **(d)** two tasks suspended concurrently resume independently to the correct + thread. + +Every command below is grounded in the **actual** code surface +(`run-team.py`, `coordinator.py`, `slack_listener.py`, `responder.py`, +`resume_worker.py`, `db/schema.py`, `graph.py`). No invented flags. Where the +code does not expose a needed knob (e.g. a short deadline), the script uses a +direct `sqlite3` write against the documented `pending_questions` schema and says +so. + +## Two answer paths — live Slack OR the operator CLI + +**D-1 is fixed:** `Coordinator.serve()` now starts the inbound `SlackListener` +when the live transport is Slack AND `SLACK_APP_TOKEN` is set, so the +**live Slack answer round-trip works**. You can run the demo either way: + +- **Live Slack** — Adam clicks the Block Kit button / replies in the channel; the + listener normalizes the event, runs the AUTHZ-01 owner check, and drives the + first-answer-wins compare-and-set (`responder.submit_answer` → + `db.schema.answer_question`). +- **Operator CLI** — `run-team.py answer --answer ... --confirm` runs the + **identical** compare-and-set (audit-logged answer-on-behalf). Useful when the + Slack app is not yet provisioned, or to script the assertions. + +Both exercise the same durable mechanic; the human gate decision is always +Adam's. The assertions below assert on the **durable ledger status** (the §3.3.1 +source of truth) and are identical for either path. The examples use the CLI +`answer --confirm` form so they are copy-pasteable; substitute "Adam answers in +Slack" wherever you prefer the live path. + +> The automated proof of this mechanic is `tests/sim/test_p1_exit_criteria.py` +> (a `SimPipeline` harness) and `tests/test_coordinator.py` (the serve/listener +> wiring). This script is the live-box demonstration on top of that. + +## Verb map (use `run-team.py`, not `operator_cli.py`) + +| Need | `run-team.py` verb | Notes | +|---|---|---| +| start a task to the human gate | `start --task "..." [--transport slack] [--dry-run]` | mints a `thread_id`, posts the clarifier, writes the `open` ledger row | +| list waiting questions | `list` (default `open`) / `list --all` / `list --parked` | JSON rows | +| inspect one row | `show ` | JSON row | +| answer on the task's behalf | `answer --answer --confirm` | destructive, audit-logged; first-answer-wins CAS | +| force-expire an open question | `expire --confirm` | destructive; flips `open`→`expired` | +| un-park (reopen) an expired question | `force-resume --confirm` | only acts on `expired` rows (reopens them) | + +`operator_cli.py` has different verbs (`force-expire`, `answer-on-behalf`) and +**no `show`**, and requires `--db`/`--audit-log` — do not use it here. + +## Preconditions + +```bash +ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120 +cd ~/orchestrator/agent-team +. .venv/bin/activate # so run-team.py uses the venv deps +# Confirm the ledger exists (Step 5 of the runbook): +python3 run-team.py list --all # [] on a fresh DB is fine +# For the live-Slack path, confirm the listener started: +journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener" +# -> "inbound Slack listener started (Socket Mode, background thread)" +``` + +Run all `run-team.py` commands from `~/orchestrator/agent-team` so they hit the +default ledger (`state/agent_team.sqlite`) and default audit log +(`state/audit.log.jsonl`). The clarifier is posted to `SLACK_CHANNEL_ID`. + +> **Resume execution model.** A won answer (rowcount 1) enqueues a `ResumeJob` +> onto the coordinator's in-process `resume_queue`; the daemon drains it on the +> next `tick()`, or the startup `recover()` re-drives any `answered`-but-unresumed +> row after a restart. So a CLI `answer --confirm` (or a live Slack answer) flips +> the ledger row to `answered`; the **running daemon** then resumes the graph. +> The demo asserts on the durable ledger status. + +--- + +## (a) Crash-safe resume — kill mid-wait, restart, task resumes 🧑 Adam answers + +**Setup — drive a task to the clarifier wait.** With the daemon already running: + +```bash +python3 run-team.py start --task "demo-a: trivial scoped task" +# prints a thread_id, e.g. 3f2a... (record it as $TID_A) +python3 run-team.py list +# expect one row: status="open", a thread_id, a question_id (record as $QID_A), +# channel_ref set (Slack ts) or null if posting is dry/unavailable. +``` + +**Kill the box mid-wait, then restart:** + +```bash +sudo systemctl stop agent-team-coordinator.service +# (optionally reboot the VM here for a stronger demonstration) +sudo systemctl start agent-team-coordinator.service +journalctl -u agent-team-coordinator.service -e | tail -40 # expect "starting" + recover() sweep +``` + +**ASSERT — durable state survived the kill (no answer yet):** + +```bash +python3 run-team.py show $QID_A +# PASS: status == "open" (the question was NOT lost across the restart) +``` + +**🧑 Human gate — Adam answers after the restart (Slack OR CLI):** + +```bash +# Live Slack: Adam replies/clicks in the channel. OR via the CLI: +python3 run-team.py answer $QID_A --answer "scope: just demo, no real change" --confirm +python3 run-team.py show $QID_A +# PASS: status == "answered" +# Wait one tick (~30s, DEFAULT_POLL_INTERVAL) for the daemon to drain the resume: +journalctl -u agent-team-coordinator.service -e | tail -20 +``` + +**PASS criterion (a):** the question stayed `open` across the kill, and an answer +submitted *after* the restart drove it forward. + +--- + +## (b) Duplicate answer is a no-op 🧑 Adam answers + +```bash +python3 run-team.py start --task "demo-b: duplicate-answer test" +python3 run-team.py list # record the new question_id as $QID_B (status open) +``` + +**🧑 First answer (the winning one) — Slack OR CLI:** + +```bash +python3 run-team.py answer $QID_B --answer "first-answer" --via "slack:U1" --confirm +# expect: "answered question (via slack:U1)" (exit 0) +``` + +**Duplicate / second answer (must lose the compare-and-set):** + +```bash +python3 run-team.py answer $QID_B --answer "second-answer" --via "github:U2" --confirm +# expect: "answer no-op: question was not 'open' ..." on stderr, exit 1 +``` + +For the **live-Slack** variant, Adam (or a second authorized owner) answers the +same message twice; the second event loses the CAS and is logged +`ignored Slack answer for question_id=... (not open: duplicate ...)`. + +**ASSERT — the first answer is preserved:** + +```bash +python3 run-team.py show $QID_B +# PASS: status == "answered"; answered_via == "slack:U1" (the FIRST answer) +``` + +**PASS criterion (b):** first-answer-wins (`rowcount==1`); the duplicate hits the +`BEGIN IMMEDIATE` compare-and-set in `answer_question` and is ignored +(`rowcount==0`) — no second resume, no overwrite. + +--- + +## (c) Past-deadline answer rejected + task parks 🧑 Adam observes + +> **Code reality:** `run-team.py start` always sets the clarifier deadline from +> `DEFAULT_CLARIFY_DEADLINE = 24h`. There is **no CLI flag for a short deadline**. +> Two faithful ways to demo expiry without waiting 24h: + +### Option C1 — force the deadline past, let the daemon's sweep expire it (most faithful) + +```bash +python3 run-team.py start --task "demo-c: deadline test" +python3 run-team.py list # record $QID_C (status open) + +# Set this question's deadline into the past directly in the ledger: +sqlite3 state/agent_team.sqlite \ + "UPDATE pending_questions SET deadline_at='2000-01-01T00:00:00+00:00' WHERE question_id='$QID_C';" + +# Wait one daemon tick (~30s) for the deadline sweep to flip it, OR observe: +journalctl -u agent-team-coordinator.service -e | tail -20 +# expect the park ALARM line: "task parked: clarifier question expired ..." +``` + +### Option C2 — operator force-expire (if you do not want to touch the DB) + +```bash +python3 run-team.py start --task "demo-c: deadline test" +python3 run-team.py list # record $QID_C +python3 run-team.py expire $QID_C --confirm # destructive, audit-logged; open->expired +``` + +**ASSERT — the question is expired and the late answer is rejected:** + +```bash +python3 run-team.py show $QID_C +# PASS: status == "expired" + +# 🧑 Adam submits a LATE answer (Slack OR CLI) — it must lose the CAS: +python3 run-team.py answer $QID_C --answer "too-late" --confirm +# expect: "answer no-op: question was not 'open' ..." stderr, exit 1 +python3 run-team.py show $QID_C +# PASS: status still "expired"; answer_json still NULL + +# The parked task surfaces in the parked view: +python3 run-team.py list --parked # PASS: $QID_C appears here +``` + +**Deliberate un-park (proves the operator recovery path, §6.6):** + +```bash +python3 run-team.py force-resume $QID_C --confirm +# expect: "force-resume: reopened expired question ; it will be re-delivered" +python3 run-team.py show $QID_C +# status flips back to "open" (reopen_question), deadline_at cleared. +``` + +**PASS criterion (c):** the expired question rejects the late answer, the task +parks rather than spins, and the operator can deliberately un-park it via +`force-resume` (which reopens only an `expired` row). + +> **`force-resume` semantics (verified):** `run-team.py force-resume` reopens an +> `expired` question (the parked case). For an `answered` question it records +> intent and reports the recovery sweep will resume it (no mutation). For +> `open`/`superseded`/absent it is a no-op exit 1. It does **not** supersede. + +--- + +## (d) Two concurrent tasks resume independently 🧑 Adam answers + +```bash +python3 run-team.py start --task "demo-d task A" # record $TID_A2, then: +python3 run-team.py start --task "demo-d task B" # record $TID_B2 +python3 run-team.py list +# expect TWO open rows with DISTINCT thread_id AND distinct question_id. +# Record $QID_A2 and $QID_B2 — match by thread_id. +``` + +**(Optional) restart the daemon first** to also show concurrent tasks survive a +restart, then answer. + +**🧑 Answer the SECOND task first, with a distinct answer, then the first +(Slack OR CLI):** + +```bash +python3 run-team.py answer $QID_B2 --answer "answer-for-B" --via "slack:U2" --confirm +python3 run-team.py answer $QID_A2 --answer "answer-for-A" --via "slack:U1" --confirm +# Wait one daemon tick (~30s) for both resumes to drain. +``` + +**ASSERT — each task carries its OWN answer; no cross-talk:** + +```bash +python3 run-team.py show $QID_A2 +# PASS: status "answered", answered_via "slack:U1", thread_id == $TID_A2 +python3 run-team.py show $QID_B2 +# PASS: status "answered", answered_via "slack:U2", thread_id == $TID_B2 +python3 run-team.py list --all +# PASS: the two rows resolved on their own thread_id; no cross-contamination. +``` + +**PASS criterion (d):** two concurrently-suspended tasks each resumed to their +own `thread_id` with their own answer — answering B before A did not misroute, +and the per-thread single-flight guard kept them independent. + +--- + +## Acceptance + +P1 is accepted only when **(a), (b), (c), and (d) all pass** on the live box. +Record the four `show` outputs (or `journalctl` excerpts) as evidence. Then +complete the PROVISIONING-RUNBOOK post-session definition-of-done (memory + +Confluence + the mandatory `/sh-security-review` on the Slack inbound listener). + +## What can only be verified on the live box + +- That the daemon actually drains the resume and the LangGraph checkpoint + advances (asserted via ledger status + journalctl; the graph-state advance is + observable only on the box). +- The live **Slack** post/answer round-trip end-to-end (the listener is wired — + D-1 fixed — but the live Socket Mode socket + a real `SLACK_BOT_TOKEN` / + `SLACK_APP_TOKEN` / channel membership are only present on the box). +- The exact `channel_ref` value (Slack message `ts`) — depends on a live Slack + post succeeding. diff --git a/docs/provisioning/PROVISIONING-RUNBOOK.md b/docs/provisioning/PROVISIONING-RUNBOOK.md new file mode 100644 index 0000000..d2e0328 --- /dev/null +++ b/docs/provisioning/PROVISIONING-RUNBOOK.md @@ -0,0 +1,402 @@ +# PROVISIONING RUNBOOK — R720 agent-team Plane-2 coordinator + +This is the ordered command sequence for the **operator-present** provisioning +session that stands up the always-on `agent-team-coordinator` daemon on the +`sh-secrev` R720 VM. Derived from `agent-team/DEPLOY-R720.md`, +`security-review/DEPLOY-R720.md`, and `docs/r720-agent-team-design.md` +(§3.3.1 / §7 / §7.1). + +> ## State of this runbook +> +> The build is **merged and final**: all six Plane-1 checkers +> (`aws-posture`, `compliance-drift`, `confluence-doc`, `dependency-cve`, +> `doc-drift`, `plan-groomer`), the Tier-3 dep-bump **fixer** +> (`run-team.py fix --dry-run`), and the Plane-1→Plane-2 **P5 cross-plane loop** +> (`run-team.py intake-checker`) exist in the tree. The three deploy-correctness +> bugs the provisioning-prep audit found are **FIXED** (see DEPLOY-AUDIT.md): +> +> - **D-1 RESOLVED** — `Coordinator.serve()` now starts the inbound Slack +> `SlackListener` concurrently with the tick/drain loop when Slack is the live +> transport AND `SLACK_APP_TOKEN` is set. **The demo can use the live Slack +> answer path** (P1-DEMO-SCRIPT.md), or the operator-CLI `answer` path. +> - **D-2 RESOLVED** — the systemd unit loads `~/orchestrator/.env` (for the P2 +> GPT-4.1 review loop's provider key) in addition to `~/secrev.env`. +> - **D-7 RESOLVED** — the unit's `ExecStart` points at the agent-team venv +> interpreter, not the system `python3`. +> +> Still open as provisioning notes (not blockers): **D-4/D-5** (agent-team +> runtime deps are installed ad-hoc into the venv and are not pinned in +> `requirements.txt`), and the **operator-CLI divergence** (`run-team.py` vs +> `agent_team/operator_cli.py` have different verb names — use `run-team.py`). + +--- + +## Scope and ground rules + +**Hard rules (carried from the design and global instructions):** + +- **Snapshot before any stateful change** (design §7; `feedback_ec2_replacement_snapshot`). +- **Every stateful step has an exercised rollback.** +- Secrets use **placeholder names only**; real values are entered by the operator + at the box and never echoed into shell history. +- The box is **read-only / subscription-OAuth only**. **No `ANTHROPIC_API_KEY`** + on this host (it would silently win over OAuth and meter to API rates — + `billing.claude_invoke` pops it defensively, but it must not be present). +- **No IAM / OIDC is involved in the coordinator deploy** (P1/P2). IAM enters + only at the **P3-live flip** (see the dedicated section), which is gated on the + mandatory GPT-4.1 cross-review + `/sh-security-review`. + +**Legend per step:** + +- 🧑 **OPERATOR-REQUIRED** — needs the human (snapshot, secrets, the live human + gate, go/no-go). Cannot be automated. +- 🤖 **MECHANICAL** — deterministic; an operator runs it but it needs no judgment. + +## Host facts (from both DEPLOY-R720.md files) + +- Hypervisor: R720 at `10.10.60.40` (Windows Server 2022, Hyper-V). +- VM: `sh-secrev`, Ubuntu 24.04, **4GB / 2 vCPU / 40GB** dynamic vhdx, `10.10.60.120`. +- Reach: `ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120` (key-only, NOPASSWD sudo). +- Repo on the box: `~/orchestrator/` (rsync from the Mac, **NOT** a git clone). + The package lives at `~/orchestrator/agent-team/`. +- Shares `~/secrev.env` (mode 600) with the secrev sweep, and `~/orchestrator/.env` + (mode 600) for non-Claude provider keys (same files the secrev unit loads). + +--- + +## STEP 1 — Snapshot the VM 🧑 OPERATOR-REQUIRED + +**On the R720 host (Hyper-V), before anything else.** This is the one-command +undo for every change below. + +```powershell +# On the R720 Windows host (PowerShell, as admin): +Checkpoint-VM -Name sh-secrev -SnapshotName "pre-agent-team-coordinator-$(Get-Date -Format yyyyMMdd-HHmm)" +Get-VMSnapshot -VMName sh-secrev # confirm the checkpoint exists +``` + +**ROLLBACK (whole session):** +```powershell +Restore-VMSnapshot -VMName sh-secrev -Name "" -Confirm:$false +Start-VM -Name sh-secrev +``` + +> Do not proceed until the checkpoint is confirmed present. + +--- + +## STEP 2 — Rsync the repo to the box 🤖 MECHANICAL (verify manifest 🧑) + +**From the Mac.** Same pattern/excludes as the secrev deploy. The team **scans +the same `~/repo-mirrors` corpus secrev already maintains** — this rsync ships +*code*, not mirrors. + +```bash +# From the Mac (sync the canonical repo, not a worktree): +rsync -av --exclude .env --exclude .venv --exclude .git --exclude .claude \ + ~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/ +``` + +Surface that must be on the box: + +| Path (under `~/orchestrator/`) | Why it must ship | +|---|---| +| `agent-team/run-team.py` | the operator entry CLI | +| `agent-team/agent_team/**` | the package (coordinator, graph, ledger, transports, slack_listener, fixer) | +| `agent-team/systemd/agent-team-coordinator.service` | the daemon unit | +| `run.py` + the orchestrator package | the GPT-4.1 cross_reviewer the P2 review loop shells | +| `requirements.txt` | pin reference for langgraph / checkpoint-sqlite | +| `security-review/lib/**` | shared sweep substrate (Phase 0) | +| `security-review/checkers/**` | the six Plane-1 checkers + fixtures | + +**ROLLBACK:** rsync is additive; restore the Step-1 snapshot to revert code state. + +--- + +## STEP 3 — Write secrets 🧑 OPERATOR-REQUIRED + +**On the VM.** The coordinator unit loads **two** EnvironmentFiles (both +optional via the leading `-`): `~/secrev.env` (agent-team runtime keys) and +`~/orchestrator/.env` (the non-Claude provider key for the P2 review loop). +**Placeholder names only — the operator pastes real values.** Do not echo real +tokens into shell history (use an editor or `read -s`). + +`~/secrev.env` (mode 600) — append the agent-team keys: + +```bash +ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120 +# Edit ~/secrev.env (mode 600) and add — values are placeholders: +# CLAUDE_CODE_OAUTH_TOKEN= # from `claude setup-token` +# SLACK_BOT_TOKEN= # xoxb-..., chat:write (posts questions) +# SLACK_APP_TOKEN= # xapp-..., connections:write (Socket Mode inbound) +# SLACK_CHANNEL_ID= # C0..., target clarifier channel +# AGENT_TEAM_SLACK_OWNER_IDS= # comma-separated U... ids of authorized answerers +chmod 600 ~/secrev.env +``` + +`~/orchestrator/.env` (mode 600) — the non-Claude provider key for the GPT-4.1 +review loop (same file secrev uses; if it already exists with the key, leave it): + +```bash +# OPENAI_API_KEY= # (or the provider key cross_reviewer/GPT-4.1 needs) +chmod 600 ~/orchestrator/.env +``` + +**Contract notes (verified against the code — see DEPLOY-AUDIT.md):** + +- `SLACK_CHANNEL_ID` is correct — `run-team.py _build_transport` reads exactly + `os.environ.get("SLACK_CHANNEL_ID")`. Do **not** use `SLACK_CHANNEL`. +- `SLACK_APP_TOKEN` is now **read by the daemon**: `Coordinator.serve()` starts + the inbound `SlackListener` when the transport is live Slack AND `SLACK_APP_TOKEN` + is set. Without it the daemon still runs (posts + expires) but never hears Slack + replies — Slack stays optional by design. +- `AGENT_TEAM_SLACK_OWNER_IDS` **fails closed** (AUTHZ-01): if unset/empty the + listener rejects **every** answer. It must be set for the live human gate. +- **CRITICAL:** confirm `ANTHROPIC_API_KEY` is NOT present: + ```bash + grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # must print 0 for both + ``` + +**ROLLBACK:** strip exactly the appended keys (keep secrev keys intact), or +restore the Step-1 snapshot. Do NOT blindly truncate — secrev keys live here too. + +--- + +## STEP 4 — Create the venv + install pip deps 🤖 MECHANICAL + +**On the VM.** A dedicated venv under `agent-team/.venv` (excluded from rsync). +The systemd unit's `ExecStart` points at **this venv's interpreter** (D-7 fixed), +so the deps MUST land here. + +```bash +cd ~/orchestrator/agent-team +python3 -m venv .venv +. .venv/bin/activate +pip install langgraph==1.1.10 langgraph-checkpoint-sqlite==3.1.0 \ + claude-agent-sdk slack_sdk slack_bolt requests +``` + +> **Pin note (D-5):** `requirements.txt` pins `langgraph==1.1.10` / +> `langgraph-checkpoint-sqlite==3.1.0`; match those exactly here. The other +> runtime deps (`claude-agent-sdk`, `slack_sdk`, `slack_bolt`, `requests`) are +> not yet in `requirements.txt` (D-5 open) — installed ad-hoc here. `requests` +> is required by the GitHub transport/intake (D-4). `anthropic` is **not** +> installed (only the opt-in `api` billing mode needs it). +> `slack_bolt` is now actually exercised (D-1 fixed: the daemon starts the +> Socket Mode listener). + +**Verify the imports resolve (using the venv interpreter the unit will use):** +```bash +.venv/bin/python -c "import langgraph, langgraph.checkpoint.sqlite, slack_sdk, slack_bolt, requests; print('deps ok')" +.venv/bin/python -c "import claude_agent_sdk; print('agent-sdk ok')" +``` + +**ROLLBACK:** `deactivate 2>/dev/null; rm -rf ~/orchestrator/agent-team/.venv` + +--- + +## STEP 5 — Initialize the durable ledger DB 🤖 MECHANICAL + +**On the VM, venv active.** Idempotent; creates `state/agent_team.sqlite` with +the `pending_questions` + `budget_ledger` + `schema_meta` tables (LangGraph +`SqliteSaver` creates its own tables in the same file on first run). + +```bash +cd ~/orchestrator/agent-team +. .venv/bin/activate +python3 run-team.py init-db +# Expect: "initialized ledger DB at .../state/agent_team.sqlite" +ls -l state/ # agent_team.sqlite present; state/ is gitignored +``` + +**ROLLBACK (reset the ledger only):** +```bash +cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak} +rm ~/orchestrator/agent-team/state/agent_team.sqlite* +# re-run `python3 run-team.py init-db` to recreate empty tables. +``` + +--- + +## STEP 6 — Install + start the systemd unit 🤖 MECHANICAL (go/no-go 🧑) + +**On the VM, as root.** Installs the long-running coordinator daemon. Keep the +**hardening as-shipped**: `NoNewPrivileges`, `ProtectSystem=full`, +`ProtectHome=read-only` + `ReadWritePaths=.../agent-team/state` (locked decision). + +```bash +sudo cp ~/orchestrator/agent-team/systemd/agent-team-coordinator.service /etc/systemd/system/ +sudo systemctl daemon-reload +sudo systemctl enable --now agent-team-coordinator.service +systemctl status agent-team-coordinator.service +journalctl -u agent-team-coordinator.service -e -f +# expect: "agent-team coordinator starting" +# and (if Slack + SLACK_APP_TOKEN provisioned): +# "inbound Slack listener started (Socket Mode, background thread)" +# or otherwise: +# "inbound Slack listener not started (... or SLACK_APP_TOKEN is unset) ..." +``` + +The unit (post-fix) runs +`/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve` from +`WorkingDirectory=/home/adam/orchestrator/agent-team` as `User=adam`, loading +**both** `EnvironmentFile=-/home/adam/secrev.env` and +`EnvironmentFile=-/home/adam/orchestrator/.env`, `Restart=on-failure`. + +> **D-7 fixed:** `ExecStart` now resolves the venv interpreter, so the Step-4 +> deps are on the path. **D-1 fixed:** `serve` starts the inbound `SlackListener` +> when Slack + app token are provisioned. **Go/no-go:** confirm the journal shows +> the listener line you expect for your transport choice before Step 8. + +**ROLLBACK (exercised — tear the unit down cleanly):** +```bash +sudo systemctl disable --now agent-team-coordinator.service +sudo rm /etc/systemd/system/agent-team-coordinator.service +sudo systemctl daemon-reload +systemctl status agent-team-coordinator.service # should report "could not be found" +``` +The ledger under `state/` is untouched. Step-1 snapshot is the whole-host fallback. + +--- + +## STEP 7 — Verify the Slack inbound listener 🧑 OPERATOR-REQUIRED + +The clarifier gate is two halves: **outbound** (post the question — done by +`serve`/`start` via the Slack poster) and **inbound** (receive Adam's answer — +the `SlackListener` over Socket Mode). + +> **D-1 RESOLVED.** `Coordinator.serve()` now constructs and starts +> `SlackListener` on a background daemon thread, concurrently with the tick/drain +> loop, **when** the live transport is a `SlackTransport` AND `SLACK_APP_TOKEN` +> is set. It shares the coordinator's own transport, ledger, and resume queue, +> and stops cleanly on shutdown. The AUTHZ-01 owner allowlist + the open-status +> compare-and-set are unchanged — the listener still fails closed on an empty +> `AGENT_TEAM_SLACK_OWNER_IDS`. + +**Confirm the live inbound path is up:** + +```bash +# In the journal (Step 6) expect the "inbound Slack listener started" line. +# Verify the two gating vars are present in the unit's environment: +grep -c SLACK_APP_TOKEN ~/secrev.env # 1 +grep -c AGENT_TEAM_SLACK_OWNER_IDS ~/secrev.env # 1 (else the gate rejects all answers) +``` + +If you are **not** using Slack as the transport (or are deliberately running +without the app token), the daemon runs the maintenance loop only and the demo +uses the operator-CLI `answer` path — both are valid (P1-DEMO-SCRIPT.md). + +**ROLLBACK:** none needed — verification only, no host state change. + +--- + +## STEP 8 — Live P1 four-criteria acceptance demo 🧑 OPERATOR-REQUIRED + +Run **P1-DEMO-SCRIPT.md** in full. All four §3.3.1 exit criteria must pass: +(a) crash-safe resume, (b) duplicate-answer no-op, (c) post-deadline rejection + +park, (d) two concurrent tasks resume independently. With D-1 fixed you may +exercise the **live Slack answer path** for (b)/(d); the operator-CLI `answer` +path remains available and exercises the identical compare-and-set. **Do not +accept P1 until all four pass.** + +**ROLLBACK:** the demo writes only ledger rows under `state/`; reset via the +Step-5 ledger rollback, or restore the Step-1 snapshot, then re-run. + +--- + +## STEP 9 — Plane-1 checker live dry-runs 🧑 OPERATOR-REQUIRED + +The six checkers live under `security-review/checkers/`: +`aws-posture.sh`, `compliance-drift.sh`, `confluence-doc.sh`, +`dependency-cve.sh`, `doc-drift.sh`, `plan-groomer.sh`. All run **report + +ALARM-only** (design D3): a clean run posts nothing and lands a mode-600 report; +no auto-Jira/Notion writes. + +- Dry-run each checker against `~/repo-mirrors` in report-only mode; confirm a + clean run posts nothing and writes a mode-600 report. Confirm the exact + invocation against the checker scripts and the shared + `security-review/lib/` substrate on the box. +- Confirm the **canary suite** runs first (a planted-fault miss is a COMPLACENCY + ALARM and that role is skipped — design §6.4), and the **coverage rotation** + pointer advances (a slipped role is a COVERAGE ALARM, deferred-not-dropped). + +**ROLLBACK:** checkers are read-only over the mirror corpus; a dry-run produces +only a report file. Remove the report dir to revert; no host state change. + +--- + +## STEP 10 — P5 cross-plane loop dry-run 🧑 OPERATOR-REQUIRED + +The Plane-1→Plane-2 loop turns confirmed checker findings into pipeline tasks: + +```bash +cd ~/orchestrator/agent-team && . .venv/bin/activate +# Read one or more checker report JSONs and start one task per confirmed +# at/above-threshold finding (default threshold: high). --dry-run posts nowhere. +python3 run-team.py intake-checker --report --threshold high --dry-run +``` + +The Tier-3 dep-bump **fixer** is dry-run only on the box (it holds no write +token, D2): + +```bash +python3 run-team.py fix --report security-review/<...>/dependency-cve.json \ + --finding-id --task-id --dry-run +# prints the fix spec + minimal bump patch + the CI workflow_dispatch inputs; +# dispatches NOTHING. Live dispatch is the P3-live flip below. +``` + +**ROLLBACK:** both are read-only / dry-run (no dispatch, no apply); `intake-checker` +de-dup is in-memory per process. No host state to revert beyond ledger rows from +a non-dry-run intake (Step-5 ledger rollback). + +--- + +## STEP 11 — Wire the schedule 🧑 OPERATOR-REQUIRED + +The coordinator daemon (Step 6) is **always-on**, not timer-driven. The +**per-checker timers** (and the design's "shared timer with secrev", §8) attach +here. The existing `sea-haven-secrev.timer` (OnCalendar `02:00`, Persistent) is +untouched. Add a checker timer only after that checker is dry-run-validated +(Step 9). + +**ROLLBACK:** each timer gets its own `systemctl disable --now .timer` + `rm`. + +--- + +## The P3-live flip (deferred; IAM + GitHub App gated) 🧑 OPERATOR-REQUIRED + +P3 (the build→verify apply-and-open-draft-PR loop) is **opt-in and inert** in +this deploy: `run-team.py` / `serve` pass `build_verify_wiring=None`, so no P3 +subgraph is assembled. Flipping it live is a **separate, gated** provisioning +session, not part of the coordinator deploy: + +1. **Mandatory reviews first.** The CI trust-boundary + OIDC IAM change is a + breaking IAM change → **GPT-4.1 cross-review** (global instructions) AND + `/sh-security-review` on the apply/verify surface. Do not flip without both. +2. **Provision the GitHub App** for the trusted apply path (the App that opens + the draft PR), and the **`agent-apply` GitHub Actions environment** that holds + the apply path's scoped permissions. +3. **Bind the live build→verify wiring** via + `agent_team.coordinator.gated_build_verify_wiring(...)` (the read-only CI + result fetcher + the real diff builder) — the seam a leaf calls *after* the + gate clears. The CI fetcher is read-only and fails closed (missing token / + 404 / auth failure → `None` → the gate BLOCKs and the task parks). +4. **Set the apply env vars** the live path reads (the read-only CI-result token + and the dispatch target), then re-run the fixer **without** `--dry-run` only + once the dispatcher is bound. + +Until every step above is done, the box dispatches/applies nothing. + +--- + +## Post-session definition-of-done (design §7, global instructions) + +- [ ] All four P1 criteria demonstrated live (Step 8). +- [ ] `project_r720_agent_team` memory created/updated. +- [ ] Confluence "AWS Architecture Map" / IT host inventory updated to show + `sh-secrev` now also hosts the always-on agent-team coordinator daemon. +- [ ] `/sh-security-review` run on the Slack inbound listener surface + (auth + untrusted-input; mandatory) — flag outstanding if not run. +- [ ] OPERATOR-RUNBOOK.md reviewed by whoever holds the pager. +- [ ] Snapshot retained until the daemon runs clean for one full cycle, then pruned.