feat(agent-team): deploy-readiness — serve starts Slack listener + systemd + provisioning docs (#23)

* fix(agent-team): serve() starts the inbound Slack listener (D-1)

Coordinator.serve() now constructs and starts the SlackListener concurrently
with the tick/drain loop on a background daemon thread, but ONLY when the live
transport is a SlackTransport AND SLACK_APP_TOKEN is configured. When Slack is
not the transport or the app token is absent, serve() behaves exactly as before
(tick/recover only) — Slack is never made mandatory.

- New injectable build_listener seam + default_slack_listener_factory sharing
  the coordinator's own transport, ledger db_path, and resume_queue put.
- AUTHZ-01 owner-allowlist + open-status CAS untouched: serve() sources
  AGENT_TEAM_SLACK_OWNER_IDS in SlackListener.serve, which still fails closed.
- SlackListener.close() added for clean Socket Mode teardown on shutdown;
  serve() stops the listener + joins the thread in a finally.
- Tests: start-when-Slack+app-token, no-start otherwise, clean shutdown,
  idempotent start, serve start/stop around the loop, listener close().

* fix(agent-team): systemd unit loads ~/orchestrator/.env + uses venv python (D-2/D-7)

D-2: add EnvironmentFile=-/home/adam/orchestrator/.env (optional '-') so the P2
GPT-4.1 review loop's cross_reviewer sub-process can read the non-Claude provider
key once a task reaches REVIEW. Mirrors the sea-haven-secrev unit.

D-7: point ExecStart at the agent-team venv interpreter
(/home/adam/orchestrator/agent-team/.venv/bin/python) instead of
/usr/bin/env python3, which resolved the system interpreter without the
installed deps under systemd's PATH.

All hardening (NoNewPrivileges / ProtectSystem=full / ProtectHome=read-only /
ReadWritePaths) is retained unchanged (locked decision).

* docs(agent-team): land provisioning + operator runbooks under docs/provisioning

- PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5
  intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub
  App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked
  FIXED so the demo can use the live Slack answer path.
- P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the
  Slack and operator-CLI answer paths documented for all four exit criteria.
- DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the
  operator-CLI divergence kept as provisioning notes.
- OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks,
  failed HITL resumes, budget exhaustion, transport outages, and
  COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the
  re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6).

* fix(agent-team): supervise the Slack listener thread — recurring ALARM + respawn

sh-security-review (logic) MEDIUM: a crashed listener thread was logged once,
then the daemon ran on 'deaf' — posting clarifier questions but receiving no
answers, every gate silently parking, process never exiting so systemd
Restart=on-failure never fired. serve() now calls _supervise_slack_listener()
each pass: when the listener is enabled but its thread is dead, it emits a
recurring ERROR ALARM and respawns via the idempotent starter (self-heal).
No-op when alive or disabled. +3 tests. (authz detector: wiring clean — AUTHZ-01
fail-closed allowlist + open-status CAS intact, dead listener fails SAFE.)
This commit is contained in:
Adam Moussa 2026-06-18 16:56:21 -04:00 • committed by GitHub
parent 95e481890a
commit 2d1dca0804
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
9 changed files with 1686 additions and 13 deletions

View file

@ -50,7 +50,9 @@ committed leaves (responder, resume_worker, db.schema, graph, transport).
from __future__ import annotations from __future__ import annotations
import logging import logging
import os
import queue import queue
import threading
from datetime import timedelta from datetime import timedelta
from pathlib import Path from pathlib import Path
from typing import TYPE_CHECKING, Any, Callable from typing import TYPE_CHECKING, Any, Callable
@ -60,6 +62,7 @@ from agent_team import responder as responder_mod
from agent_team.db.schema import connect, init_db from agent_team.db.schema import connect, init_db
from agent_team.resume_worker import ResumeResult, ResumeWorker from agent_team.resume_worker import ResumeResult, ResumeWorker
from agent_team.transport.base import Transport from agent_team.transport.base import Transport
from agent_team.transport.slack_adapter import SlackTransport
if TYPE_CHECKING: # pragma: no cover - typing only if TYPE_CHECKING: # pragma: no cover - typing only
from agent_team.task_model import PipelineState from agent_team.task_model import PipelineState
@ -68,6 +71,7 @@ __all__ = [
"Coordinator", "Coordinator",
"build_verify_wiring", "build_verify_wiring",
"default_clarify_node_factory", "default_clarify_node_factory",
"default_slack_listener_factory",
"gated_build_verify_wiring", "gated_build_verify_wiring",
] ]
@ -124,6 +128,48 @@ CheckpointerFactory = Callable[[Path], Any]
# so the deadline policy stays I/O-free in tests (§6.6). # so the deadline policy stays I/O-free in tests (§6.6).
AlarmHook = Callable[[str], None] AlarmHook = Callable[[str], None]
# A listener factory: builds the inbound Slack Socket Mode listener
# (:class:`~agent_team.transport.slack_listener.SlackListener`) the serve loop
# starts on a background thread when Slack is the live transport and the app
# token is present. Injected so :meth:`Coordinator.serve` is testable with a fake
# listener and no live socket. The default
# (:func:`default_slack_listener_factory`) builds the real listener from the
# coordinator's shared transport / db / resume-queue and the Slack env tokens.
ListenerFactory = Callable[[], Any]
def default_slack_listener_factory(
*,
transport: SlackTransport,
db_path: Path,
enqueue_resume: Callable[[Any], None],
) -> Any:
"""Build the live :class:`SlackListener` from the coordinator's seams (D-1).
The production listener shares the coordinator's *own* live ``SlackTransport``
(so ``parse_answer`` matches the outbound ``post_question`` wiring), the same
durable ledger ``db_path``, and the same ``resume_queue`` ``put`` callable
(the documented slack_listener ⇄ drainer handoff seam — ``coordinator.py``
module docstring §3.3). The app/bot tokens and the owner allowlist are sourced
from the environment by :meth:`SlackListener.serve` itself
(``SLACK_APP_TOKEN`` / ``SLACK_BOT_TOKEN`` / ``AGENT_TEAM_SLACK_OWNER_IDS``);
the allowlist still FAILS CLOSED if ``AGENT_TEAM_SLACK_OWNER_IDS`` is unset
(AUTHZ-01), so this factory deliberately does not weaken that — it injects no
``owner_ids`` and lets ``serve`` read + enforce them.
Imported lazily for the same import-hygiene reason as the clarifier / planner
factories (the listener pulls the transport + responder leaves).
"""
from agent_team.transport.slack_listener import SlackListener
return SlackListener(
transport,
db_path,
enqueue_resume,
app_token=os.environ.get("SLACK_APP_TOKEN") or None,
bot_token=os.environ.get("SLACK_BOT_TOKEN") or None,
)
def default_clarify_node_factory() -> Callable[[PipelineState], PipelineState]: def default_clarify_node_factory() -> Callable[[PipelineState], PipelineState]:
"""Build the live Claude-backed clarifier node (§3.3, §7.1 P1). """Build the live Claude-backed clarifier node (§3.3, §7.1 P1).
@ -338,6 +384,7 @@ class Coordinator:
resume_queue: "queue.Queue[Any] | None" = None, resume_queue: "queue.Queue[Any] | None" = None,
deadline_window: timedelta | None = None, deadline_window: timedelta | None = None,
alarm_hook: AlarmHook | None = None, alarm_hook: AlarmHook | None = None,
build_listener: ListenerFactory | None = None,
) -> None: ) -> None:
self._db_path = Path(db_path) self._db_path = Path(db_path)
self._transport = transport self._transport = transport
@ -357,6 +404,11 @@ class Coordinator:
self._resume_queue: "queue.Queue[Any]" = resume_queue or queue.Queue() self._resume_queue: "queue.Queue[Any]" = resume_queue or queue.Queue()
self._deadline_window = deadline_window or graph_mod.DEFAULT_CLARIFY_DEADLINE self._deadline_window = deadline_window or graph_mod.DEFAULT_CLARIFY_DEADLINE
self._alarm_hook = alarm_hook or self._default_alarm_hook self._alarm_hook = alarm_hook or self._default_alarm_hook
# The inbound Slack listener factory (D-1). Left None, serve() uses the
# default that builds the real SlackListener; tests inject a fake to prove
# the start/no-start/shutdown wiring with no live socket. Only consulted
# by serve() when Slack is the live transport AND the app token is set.
self._build_listener = build_listener
# Built by setup(). # Built by setup().
self._graph: Any = None self._graph: Any = None
@ -364,6 +416,10 @@ class Coordinator:
# Retains a context-manager checkpointer (production SQLite saver) so its # Retains a context-manager checkpointer (production SQLite saver) so its
# __exit__ is not run early; held open for the daemon's lifetime. # __exit__ is not run early; held open for the daemon's lifetime.
self._checkpointer_cm: Any = None self._checkpointer_cm: Any = None
# The running inbound listener + its daemon thread (None until serve()
# starts one). Held so serve()'s shutdown can stop it cleanly.
self._listener: Any = None
self._listener_thread: threading.Thread | None = None
# ------------------------------------------------------------------ # # ------------------------------------------------------------------ #
# Accessors (the shared queue is the slack_listener handoff seam). # Accessors (the shared queue is the slack_listener handoff seam).
@ -814,32 +870,173 @@ class Coordinator:
# ------------------------------------------------------------------ # # ------------------------------------------------------------------ #
def serve(self, *, poll_interval: timedelta | None = None) -> None: def serve(self, *, poll_interval: timedelta | None = None) -> None:
"""Run the live daemon: bind invoker, setup, recover, then tick forever. """Run the live daemon: bind invoker, setup, recover, start listener, tick.
The production entry. Binds the real Claude invoker The production entry. Binds the real Claude invoker
(:func:`agent_team.invoker.bind_subscription_invoker`) BEFORE (:func:`agent_team.invoker.bind_subscription_invoker`) BEFORE
:meth:`setup` builds the clarifier node (so the node's :meth:`setup` builds the clarifier node (so the node's
``billing.claude_invoke`` calls hit the live subscription path), runs the ``billing.claude_invoke`` calls hit the live subscription path), runs the
startup :meth:`recover` sweep, then loops calling :meth:`tick` on the startup :meth:`recover` sweep, **starts the inbound Slack listener**
(:meth:`_maybe_start_slack_listener`) when Slack is the live transport
and the app token is configured, then loops calling :meth:`tick` on the
deadline cadence. deadline cadence.
The actual Slack inbound feed is the slack_listener's job; the The Slack inbound feed (Socket Mode) is the slack_listener's job and runs
coordinator exposes :meth:`submit_answer` and the shared concurrently on a background daemon thread; this loop owns the
:attr:`resume_queue` for it. This loop owns only the deadline/recovery deadline/recovery maintenance cadence. On shutdown (Ctrl-C /
maintenance cadence. SIGTERM-driven ``KeyboardInterrupt``) the listener is stopped cleanly via
:meth:`_stop_slack_listener` in a ``finally``.
Slack is NOT mandatory: when the transport is not the live ``SlackTransport``
or the app token is absent, no listener is started and ``serve`` behaves
exactly as before (tick/recover only).
""" """
from agent_team.invoker import bind_subscription_invoker from agent_team.invoker import bind_subscription_invoker
bind_subscription_invoker() bind_subscription_invoker()
self.setup() self.setup()
self.recover() self.recover()
self._maybe_start_slack_listener()
interval = (poll_interval or DEFAULT_POLL_INTERVAL).total_seconds() interval = (poll_interval or DEFAULT_POLL_INTERVAL).total_seconds()
import time import time
while True: # pragma: no cover - the infinite daemon loop try:
self.tick() while True: # pragma: no cover - the infinite daemon loop
time.sleep(interval) self.tick()
self._supervise_slack_listener()
time.sleep(interval)
finally:
self._stop_slack_listener()
# ------------------------------------------------------------------ #
# Inbound Slack listener wiring (D-1): start concurrently with the
# tick/drain loop ONLY when Slack is the live transport and the app
# token is configured; stop cleanly on shutdown.
# ------------------------------------------------------------------ #
def _slack_listener_enabled(self) -> bool:
"""Return ``True`` iff the inbound Slack listener should run (D-1 gate).
Two conditions, both required (and Slack must stay OPTIONAL):
* the live transport is a :class:`SlackTransport` (the inbound
``parse_answer`` must match the outbound poster); and
* ``SLACK_APP_TOKEN`` is present in the environment (Socket Mode needs the
app-level token to open the outbound WebSocket).
If either is false the daemon runs WITHOUT an inbound listener exactly as
before — Slack is never made mandatory. The owner allowlist
(``AGENT_TEAM_SLACK_OWNER_IDS``) is intentionally NOT part of this gate:
the listener fails closed on an empty allowlist (AUTHZ-01), so starting it
unprovisioned safely rejects every answer rather than silently never
hearing Slack.
"""
if not isinstance(self._transport, SlackTransport):
return False
return bool(os.environ.get("SLACK_APP_TOKEN"))
def _maybe_start_slack_listener(self) -> bool:
"""Start the inbound Slack listener on a daemon thread if enabled (D-1).
Builds the listener via the injected ``build_listener`` factory (or the
:func:`default_slack_listener_factory`, sharing the coordinator's own live
transport, durable ``db_path`` and ``resume_queue`` put), then runs its
blocking :meth:`~agent_team.transport.slack_listener.SlackListener.serve`
on a background **daemon** thread so the Socket Mode socket and the
tick/drain loop run concurrently. Returns ``True`` iff a listener was
started; ``False`` (and no thread) when :meth:`_slack_listener_enabled`
is false — the Slack-optional contract.
Idempotent: a second call while a listener thread is already alive is a
no-op.
"""
if self._listener_thread is not None and self._listener_thread.is_alive():
return True
if not self._slack_listener_enabled():
_LOG.info(
"inbound Slack listener not started (transport is not live Slack "
"or SLACK_APP_TOKEN is unset); daemon runs tick/recover only"
)
return False
if self._build_listener is not None:
listener = self._build_listener()
else:
listener = default_slack_listener_factory(
transport=self._transport,
db_path=self._db_path,
enqueue_resume=self._resume_queue.put,
)
self._listener = listener
thread = threading.Thread(
target=self._run_listener,
name="agent-team-slack-listener",
daemon=True,
)
self._listener_thread = thread
thread.start()
_LOG.info("inbound Slack listener started (Socket Mode, background thread)")
return True
def _supervise_slack_listener(self) -> None:
"""Detect a dead inbound listener thread, ALARM (recurring), and respawn.
Called every :meth:`serve` pass. Without this, a crashed listener thread
(:meth:`_run_listener` swallows its exception) leaves the daemon "deaf":
it keeps posting clarifier questions but receives no answers, every gate
silently times out to PARK, and the process stays up so systemd
``Restart=on-failure`` never fires. This makes that failure LOUD on every
pass (not a one-shot crash log) and self-heals via the idempotent
:meth:`_maybe_start_slack_listener`. No-op when the listener is not
enabled or the thread is alive.
"""
if not self._slack_listener_enabled():
return
thread = self._listener_thread
if thread is not None and thread.is_alive():
return
_LOG.error(
"ALARM: inbound Slack listener is DOWN — answers are NOT being "
"received; clarifier gates will park. Respawning the listener."
)
# Drop the dead handle so the idempotent starter actually respawns.
self._listener_thread = None
self._maybe_start_slack_listener()
def _run_listener(self) -> None:
"""Thread target: run the listener's blocking serve, log on exit.
Swallows any listener exception into a log line so a listener crash takes
down only the inbound socket (which Restart=on-failure / a manual restart
recovers), never the maintenance loop or the whole process from a
background thread.
"""
try:
self._listener.serve()
except Exception: # noqa: BLE001 - isolate the listener thread
_LOG.exception("inbound Slack listener thread exited with an error")
def _stop_slack_listener(self) -> None:
"""Stop the inbound listener cleanly on daemon shutdown (D-1).
Best-effort: closes the Socket Mode connection via the listener's
``close`` (if it exposes one), then joins the daemon thread briefly. A
teardown error never propagates — shutdown must complete regardless.
"""
listener = self._listener
thread = self._listener_thread
self._listener = None
self._listener_thread = None
if listener is not None:
close = getattr(listener, "close", None)
if close is not None:
try:
close()
except Exception: # noqa: BLE001 - shutdown must not raise
_LOG.debug("listener close raised during shutdown; ignoring")
if thread is not None and thread.is_alive():
thread.join(timeout=5.0)
# ------------------------------------------------------------------ # # ------------------------------------------------------------------ #
# Internals. # Internals.

View file

@ -147,6 +147,10 @@ class SlackListener:
# The owner allowlist (AUTHZ-01). An empty set is the fail-closed default: # The owner allowlist (AUTHZ-01). An empty set is the fail-closed default:
# an unconfigured deploy rejects every answer. # an unconfigured deploy rejects every answer.
self._owner_ids: set[str] = set(owner_ids) if owner_ids else set() self._owner_ids: set[str] = set(owner_ids) if owner_ids else set()
# The live Socket Mode handler, retained by :meth:`serve` so :meth:`close`
# can stop it cleanly on daemon shutdown. ``None`` until ``serve`` opens
# the socket.
self._socket_handler: Any = None
def handle_event(self, raw_payload: Any) -> AnswerOutcome | None: def handle_event(self, raw_payload: Any) -> AnswerOutcome | None:
"""Normalize + submit one inbound event; return its outcome or ``None``. """Normalize + submit one inbound event; return its outcome or ``None``.
@ -331,7 +335,35 @@ class SlackListener:
def _on_mention(body: Mapping[str, Any]) -> None: def _on_mention(body: Mapping[str, Any]) -> None:
_forward(body) _forward(body)
SocketModeHandler(app, self._app_token).start() handler = SocketModeHandler(app, self._app_token)
self._socket_handler = handler
handler.start()
def close(self) -> None:
"""Best-effort clean stop of the Socket Mode connection (daemon shutdown).
The coordinator's :meth:`~agent_team.coordinator.Coordinator.serve` calls
this when it shuts the daemon down so the outbound WebSocket is closed
cleanly rather than only dying with the process. Tolerant: if the handler
was never opened, or the SDK's ``close``/``disconnect`` raises, the error
is swallowed — shutdown must never be blocked by a transport teardown
failure. The handler is dropped afterward so a second ``close`` is a
no-op.
"""
handler = self._socket_handler
if handler is None:
return
self._socket_handler = None
# slack_bolt's SocketModeHandler exposes ``close`` (and the underlying
# client a ``disconnect``); try the most specific available, swallow any
# teardown error.
closer = getattr(handler, "close", None) or getattr(handler, "disconnect", None)
if closer is None:
return
try:
closer()
except Exception: # noqa: BLE001 - shutdown must not raise
_LOG.debug("SlackListener.close: handler teardown raised; ignoring")
def _extract_sender_id(raw_payload: Mapping[str, Any]) -> str | None: def _extract_sender_id(raw_payload: Mapping[str, Any]) -> str | None:

View file

@ -13,8 +13,10 @@
# systemctl status agent-team-coordinator.service # systemctl status agent-team-coordinator.service
# journalctl -u agent-team-coordinator.service -e -f # journalctl -u agent-team-coordinator.service -e -f
# #
# Secrets come from the EnvironmentFile (leading '-' = optional, no failure if # Secrets come from the EnvironmentFile(s) (leading '-' = optional, no failure if
# absent), ~/secrev.env (mode 600, NOT in git): # absent). Two files are loaded, mirroring the sea-haven-secrev unit:
#
# ~/secrev.env (mode 600, NOT in git) - the agent-team runtime keys:
# CLAUDE_CODE_OAUTH_TOKEN -> subscription OAuth (from `claude setup-token`). # CLAUDE_CODE_OAUTH_TOKEN -> subscription OAuth (from `claude setup-token`).
# A raw ANTHROPIC_API_KEY must NOT be set on this # A raw ANTHROPIC_API_KEY must NOT be set on this
# box; it would silently win and meter to API # box; it would silently win and meter to API
@ -24,7 +26,18 @@
# REQUIRED for Socket Mode; opens the inbound # REQUIRED for Socket Mode; opens the inbound
# WebSocket that receives answers. Without it the # WebSocket that receives answers. Without it the
# coordinator can post but never hear replies. # coordinator can post but never hear replies.
# serve() starts the inbound SlackListener only when
# the transport is live Slack AND this token is set.
# SLACK_CHANNEL_ID -> target channel for clarifier questions. # SLACK_CHANNEL_ID -> target channel for clarifier questions.
# AGENT_TEAM_SLACK_OWNER_IDS -> comma-separated authorized answerer ids
# (AUTHZ-01). The listener FAILS CLOSED if unset.
#
# ~/orchestrator/.env (mode 600, NOT in git) - the P2 review-loop provider key:
# The production serve() wires the GPT-4.1 cross-review loop, which shells the
# local orchestrator run.py -> cross_reviewer once a task reaches REVIEW. That
# sub-process needs the non-Claude provider key from ~/orchestrator/.env (same
# file the sea-haven-secrev unit loads). Optional ('-') so the daemon still
# starts if it is absent; the review path then fails loudly only at REVIEW.
[Unit] [Unit]
Description=Sea Haven agent-team Plane-2 coordinator daemon Description=Sea Haven agent-team Plane-2 coordinator daemon
@ -36,7 +49,11 @@ Type=simple
User=adam User=adam
WorkingDirectory=/home/adam/orchestrator/agent-team WorkingDirectory=/home/adam/orchestrator/agent-team
EnvironmentFile=-/home/adam/secrev.env EnvironmentFile=-/home/adam/secrev.env
ExecStart=/usr/bin/env python3 run-team.py serve EnvironmentFile=-/home/adam/orchestrator/.env
# Use the agent-team venv interpreter (where the runtime deps are installed by
# the DEPLOY-R720 / PROVISIONING-RUNBOOK step), NOT the bare system python3 that
# `/usr/bin/env python3` would resolve under systemd's PATH (D-7).
ExecStart=/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve
Restart=on-failure Restart=on-failure
RestartSec=5 RestartSec=5
# Hardening - matches the level the sea-haven-secrev unit relies on, scoped for a # Hardening - matches the level the sea-haven-secrev unit relies on, scoped for a

View file

@ -602,3 +602,235 @@ def test_setup_with_p2_factories_builds_a_review_node(db_path: Path) -> None:
assert graph_mod.REVIEW in coord.graph.get_graph().nodes assert graph_mod.REVIEW in coord.graph.get_graph().nodes
finally: finally:
review_loop._review_invoker = saved review_loop._review_invoker = saved
# --------------------------------------------------------------------------- #
# Inbound Slack listener wiring in serve() (D-1)
# --------------------------------------------------------------------------- #
class _FakeListener:
"""A record-only stand-in for SlackListener (no socket, no SDK).
``serve`` blocks on an Event until ``close`` is called, mirroring the live
listener whose ``serve`` blocks on the Socket Mode handler until shut down.
Records that ``serve`` ran and that ``close`` was called so the coordinator
start/shutdown wiring can be asserted.
"""
def __init__(self) -> None:
import threading as _t
self.served = False
self.closed = False
self._stop = _t.Event()
def serve(self) -> None:
self.served = True
self._stop.wait(timeout=5.0)
def close(self) -> None:
self.closed = True
self._stop.set()
def _slack_coordinator(
db_path: Path,
*,
build_listener: Any = None,
) -> Coordinator:
"""A Coordinator whose live transport IS a SlackTransport (listener gate)."""
from agent_team.transport.slack_adapter import SlackTransport
saver = _Saver()
return Coordinator(
db_path=db_path,
transport=SlackTransport("C123"),
build_clarify_node=lambda: graph_mod.clarify_node,
build_checkpointer=lambda _path: saver,
build_listener=build_listener,
)
def test_serve_starts_listener_when_slack_and_app_token(
db_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
"""serve() starts the inbound listener when Slack is the transport AND the
app token is configured, then stops it cleanly on shutdown."""
monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test")
listener = _FakeListener()
coord = _slack_coordinator(db_path, build_listener=lambda: listener)
coord.setup()
assert coord._maybe_start_slack_listener() is True
# The background daemon thread actually ran the listener's serve().
import time
for _ in range(50):
if listener.served:
break
time.sleep(0.01)
assert listener.served is True
# Clean shutdown stops the listener and joins the thread.
coord._stop_slack_listener()
assert listener.closed is True
assert coord._listener is None
assert coord._listener_thread is None
def test_serve_does_not_start_listener_without_app_token(
db_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
"""No app token => no inbound listener (Slack stays optional)."""
monkeypatch.delenv("SLACK_APP_TOKEN", raising=False)
listener = _FakeListener()
coord = _slack_coordinator(db_path, build_listener=lambda: listener)
coord.setup()
assert coord._slack_listener_enabled() is False
assert coord._maybe_start_slack_listener() is False
assert listener.served is False
assert coord._listener is None
assert coord._listener_thread is None
def test_serve_does_not_start_listener_when_transport_not_slack(
db_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
"""A non-Slack transport never starts the listener even with the app token."""
monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test")
listener = _FakeListener()
saver = _Saver()
coord = Coordinator(
db_path=db_path,
transport=FakeTransport(), # not a SlackTransport
build_clarify_node=lambda: graph_mod.clarify_node,
build_checkpointer=lambda _path: saver,
build_listener=lambda: listener,
)
coord.setup()
assert coord._slack_listener_enabled() is False
assert coord._maybe_start_slack_listener() is False
assert listener.served is False
def test_maybe_start_listener_is_idempotent(
db_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
"""A second start while the thread is alive is a no-op (one listener)."""
monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test")
built: list[_FakeListener] = []
def _factory() -> _FakeListener:
lst = _FakeListener()
built.append(lst)
return lst
coord = _slack_coordinator(db_path, build_listener=_factory)
coord.setup()
try:
assert coord._maybe_start_slack_listener() is True
assert coord._maybe_start_slack_listener() is True
assert len(built) == 1 # not rebuilt/restarted
finally:
coord._stop_slack_listener()
def test_supervise_respawns_dead_listener_and_alarms(
db_path: Path, monkeypatch: pytest.MonkeyPatch, caplog: pytest.LogCaptureFixture
) -> None:
"""A dead listener thread is respawned (self-heal) with a recurring ALARM, so
the daemon never silently goes 'deaf' to Slack answers (D1-silent-stall)."""
import logging
monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test")
listeners: list[_FakeListener] = []
def _factory() -> _FakeListener:
lst = _FakeListener()
listeners.append(lst)
return lst
coord = _slack_coordinator(db_path, build_listener=_factory)
coord.setup()
try:
assert coord._maybe_start_slack_listener() is True
first = coord._listener_thread
assert first is not None and first.is_alive()
# Kill the listener thread: close() fires the stop event so serve() returns.
listeners[0].close()
first.join(timeout=5.0)
assert not first.is_alive()
# The supervisor detects the death, ALARMs, and respawns a fresh listener.
with caplog.at_level(logging.ERROR):
coord._supervise_slack_listener()
assert any("listener is DOWN" in r.message for r in caplog.records)
second = coord._listener_thread
assert second is not None and second is not first and second.is_alive()
assert len(listeners) == 2 # a new listener was built on respawn
finally:
coord._stop_slack_listener()
def test_supervise_is_noop_when_listener_alive(
db_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
"""The supervisor leaves a healthy listener thread untouched (no churn)."""
monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test")
listener = _FakeListener()
coord = _slack_coordinator(db_path, build_listener=lambda: listener)
coord.setup()
try:
coord._maybe_start_slack_listener()
alive = coord._listener_thread
coord._supervise_slack_listener()
assert coord._listener_thread is alive
finally:
coord._stop_slack_listener()
def test_supervise_is_noop_when_listener_disabled(
db_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
"""No app token -> listener disabled -> supervisor never spawns a thread."""
monkeypatch.delenv("SLACK_APP_TOKEN", raising=False)
coord = _slack_coordinator(db_path, build_listener=_FakeListener)
coord.setup()
coord._supervise_slack_listener()
assert coord._listener_thread is None
def test_stop_listener_is_safe_when_none(db_path: Path) -> None:
"""Stopping with no listener running never raises (clean-shutdown safety)."""
coord = _slack_coordinator(db_path)
coord._stop_slack_listener() # must be a quiet no-op
assert coord._listener is None
def test_serve_wires_start_and_stop_around_the_loop(
db_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
"""serve() starts the listener before the loop and stops it in finally even
when the loop is interrupted (KeyboardInterrupt)."""
monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test")
listener = _FakeListener()
coord = _slack_coordinator(db_path, build_listener=lambda: listener)
# Neuter the live invoker bind so serve() needs no Claude SDK / token.
# serve() does `from agent_team.invoker import bind_subscription_invoker`,
# so patch the name on the invoker module it imports from.
import agent_team.invoker as invoker_mod
monkeypatch.setattr(invoker_mod, "bind_subscription_invoker", lambda: None)
# Break the infinite loop on the first tick.
monkeypatch.setattr(
coord, "tick", lambda: (_ for _ in ()).throw(KeyboardInterrupt())
)
with pytest.raises(KeyboardInterrupt):
coord.serve(poll_interval=timedelta(seconds=0))
assert listener.served is True # started before the loop
assert listener.closed is True # stopped in finally on interrupt

View file

@ -449,6 +449,36 @@ def test_serve_requires_tokens(db_path: Path) -> None:
listener.serve() listener.serve()
def test_close_without_open_handler_is_noop(db_path: Path) -> None:
"""close() before serve() ever opened a socket is a quiet no-op."""
listener = _listener(db_path, RecordingQueue())
listener.close() # must not raise
listener.close() # idempotent
def test_close_stops_socket_handler(db_path: Path) -> None:
"""close() invokes the retained Socket Mode handler's close and drops it."""
class _FakeHandler:
def __init__(self) -> None:
self.closed = False
def close(self) -> None:
self.closed = True
listener = _listener(db_path, RecordingQueue())
handler = _FakeHandler()
listener._socket_handler = handler
listener.close()
assert handler.closed is True
assert listener._socket_handler is None
# A teardown that raises is swallowed (shutdown must never raise).
listener._socket_handler = type(
"_Boom", (), {"close": lambda self: (_ for _ in ()).throw(RuntimeError("x"))}
)()
listener.close() # no exception
def test_real_slack_transport_is_a_transport() -> None: def test_real_slack_transport_is_a_transport() -> None:
"""Sanity: the injected SlackTransport is the contract the listener expects.""" """Sanity: the injected SlackTransport is the contract the listener expects."""
assert isinstance(SlackTransport(channel="C123"), Transport) assert isinstance(SlackTransport(channel="C123"), Transport)

View file

@ -0,0 +1,187 @@
# DEPLOY-AUDIT — `agent-team/DEPLOY-R720.md` + systemd unit vs. actual code
Cross-check of the drafted deploy doc (`agent-team/DEPLOY-R720.md`) and the
systemd unit (`agent-team/systemd/agent-team-coordinator.service`) against the
**actual current code** in `agent-team/agent_team/**` and `agent-team/run-team.py`.
Each finding: **location → claimed → actual → fix**. Severity: 🔴 blocker /
🟠 should-fix / 🟡 nit. Items confirmed clean are stated explicitly.
> **Status (deploy-readiness PR):** the three deploy-correctness bugs that
> blocked the live provisioning session — **D-1, D-2, D-7** — are **RESOLVED** in
> this PR (`feature/agent-team-deploy-readiness`). The remaining items
> (**D-4 / D-5** dependency pinning, the **operator-CLI divergence**) are kept
> below as **provisioning notes** — they do not block the coordinator deploy.
---
## ✅ D-1 — `serve` now starts the Slack inbound listener — RESOLVED (this PR)
- **Location:** `agent_team/coordinator.py` (`Coordinator.serve` +
`_maybe_start_slack_listener` / `_slack_listener_enabled` / `_stop_slack_listener`
/ `default_slack_listener_factory`); `agent_team/transport/slack_listener.py`
(`SlackListener.serve` + new `close`).
- **Was:** `Coordinator.serve()` did only `bind_subscription_invoker()`,
`setup()`, `recover()`, then an infinite tick/sleep loop — it never constructed
or started `SlackListener`, so a deployed daemon posted clarifier questions and
expired them on deadline but could **not hear Slack answers**.
- **Now:** `serve()` starts the `SlackListener` on a background **daemon thread**,
concurrently with the tick/drain loop, **when** the live transport is a
`SlackTransport` AND `SLACK_APP_TOKEN` is set. It shares the coordinator's own
transport, ledger `db_path`, and `resume_queue` put; on shutdown it calls the
listener's new `close()` and joins the thread in a `finally`. When Slack is not
the transport or the app token is absent, no listener starts and `serve` behaves
exactly as before — **Slack is never made mandatory**. The AUTHZ-01 owner
allowlist + the open-status compare-and-set are untouched (still fail closed on
an empty `AGENT_TEAM_SLACK_OWNER_IDS`).
- **Tests:** `tests/test_coordinator.py` — start-when-Slack+app-token, no-start
without the token, no-start when the transport is not Slack, idempotent start,
clean shutdown, serve start/stop around the loop; `tests/test_slack_listener.py`
— `close()` no-op + handler teardown.
---
## ✅ D-2 — coordinator unit loads `~/orchestrator/.env` — RESOLVED (this PR)
- **Location:** `agent-team/systemd/agent-team-coordinator.service`;
`run-team.py:_build_coordinator` (always wires `default_review_wiring`);
`coordinator.py:default_review_wiring` → `review_loop_llm.default_plan_reviewer`
(shells the local orchestrator `run.py` → `cross_reviewer` GPT-4.1).
- **Was:** the unit loaded only `~/secrev.env`. `run-team.py serve` builds the
coordinator with the P2 review loop wired, and once a task reaches REVIEW the
reviewer shells the orchestrator `run.py`, whose GPT-4.1 call reads the
non-Claude provider key from `~/orchestrator/.env`. The review path would fail
to authenticate.
- **Now:** the unit adds `EnvironmentFile=-/home/adam/orchestrator/.env`
(optional `-`, mirroring the `sea-haven-secrev` unit). Does not affect the P1
demo (P1 stops at PLAN before REVIEW); closes the latent P2 break.
---
## 🟠 D-4 — pip install list omits `requests` (provisioning note)
- **Location:** `agent_team/transport/github_live.py` / `github_intake.py`
(`import requests`).
- **Actual:** the GitHub transport + intake require `requests`; the Slack-first
path does not hit it, but any `--transport github` / `intake-github` use fails
with a clear RuntimeError without it.
- **Fix:** PROVISIONING-RUNBOOK Step 4 installs `requests` into the venv.
`requests` is also absent from `requirements.txt` (see D-5).
---
## 🟠 D-5 — agent-team runtime deps are not pinned in `requirements.txt` (provisioning note)
- **Location:** `requirements.txt` (repo root).
- **Actual:** `requirements.txt` pins `langgraph==1.1.10` and
`langgraph-checkpoint-sqlite==3.1.0`, but the agent-team runtime deps
`claude-agent-sdk`, `slack_sdk`, `slack_bolt`, `requests` (and `anthropic` for
api mode) are **not in `requirements.txt` at all** — they are installed ad-hoc
into the agent-team venv by the runbook. There is no pinned, reproducible source
of truth for the box's runtime set.
- **Fix (deferred):** add an `agent-team/requirements.txt` (or extras group)
pinning these, version-matched to the root `requirements.txt` langgraph pin.
Until then, PROVISIONING-RUNBOOK Step 4 pins `langgraph==1.1.10` /
`langgraph-checkpoint-sqlite==3.1.0` explicitly so the unpinned `pip install`
cannot pull a newer, untested major. **Do NOT modify `requirements.txt` or the
checkers in this PR** (out of scope).
---
## ✅ D-6 — `slack_bolt` is now exercised by the daemon — RESOLVED (consequence of D-1)
- **Location:** `slack_listener.py:serve` (the only `slack_bolt` import).
- **Now:** with D-1 fixed, `Coordinator.serve()` starts `SlackListener.serve()`,
which imports + uses `slack_bolt` for the Socket Mode handler. The dep is right
and now actually exercised on the live Slack path.
---
## ✅ D-SLACKVAR (clean) — `SLACK_CHANNEL_ID` matches
- `run-team.py:_build_transport` reads exactly `os.environ.get("SLACK_CHANNEL_ID")`.
The runbook, the unit comment, and the code all use `SLACK_CHANNEL_ID` (not
`SLACK_CHANNEL`). **CLEAN.**
---
## ✅ D-ENV-SLACKBOT / OAUTH / OWNERS / APPTOKEN (clean) — names match
- **`SLACK_BOT_TOKEN`** ↔ `slack_live.py` + `slack_listener` env read. **CLEAN.**
- **`CLAUDE_CODE_OAUTH_TOKEN`** ↔ `invoker.py`. **CLEAN.**
- **`AGENT_TEAM_SLACK_OWNER_IDS`** ↔ `slack_listener.py` (name + fail-closed
semantics). **CLEAN** — now read by the running daemon (D-1 fixed).
- **`SLACK_APP_TOKEN`** — now read in two places: `coordinator._slack_listener_enabled`
gates the listener on its presence, and `default_slack_listener_factory` /
`SlackListener.serve` source it to open the socket. **CLEAN** (read site exists
now that D-1 is fixed).
---
## ✅ D-7 — `ExecStart` uses the venv interpreter — RESOLVED (this PR)
- **Location:** `agent-team/systemd/agent-team-coordinator.service` ExecStart.
- **Was:** `ExecStart=/usr/bin/env python3 run-team.py serve` resolved the
**system** interpreter under systemd's PATH — not the venv where the deps were
installed, so the daemon would fail at import.
- **Now:** `ExecStart=/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve`
(matches the runbook venv path + `WorkingDirectory`).
---
## ✅ D-SUBCMD (mostly clean) — run-team.py subcommands referenced exist
Cross-checked every `run-team.py <sub>` the deploy doc + demo name against
`run-team.py:build_parser`: `init-db`, `serve`, `list` (+ `--all` / `--parked`),
`show`, `expire`, `answer`, `redeliver`, `supersede`, `force-resume`, `start`,
`intake-github`, `intake-checker`, `fix` — all exist. No invented verbs.
> ### Operator-CLI divergence (provisioning note)
>
> Two operator CLIs exist with **different verb names**:
>
> - `run-team.py` (the entry CLI): `init-db, list, show, redeliver, expire,
> answer, supersede, force-resume, start, serve, intake-github, intake-checker,
> fix`. Has `show` and `--parked`; `--db` / `--audit-log` default sensibly.
> - `agent_team/operator_cli.py`: `list, redeliver, force-expire,
> answer-on-behalf, force-resume` — **no `show`**, `--db` / `--audit-log` are
> **required**, and its `force-resume` **supersedes** (unlike `run-team.py`'s,
> which reopens an expired row and never supersedes).
>
> **Use `run-team.py` for provisioning + the demo + incident recovery.** The
> docs reference only `run-team.py`. Reconciling the two CLIs is a follow-up.
---
## ✅ D-PYTHONPKG (clean) — package import bootstrap is correct
`run-team.py` inserts its own dir into `sys.path` so the hyphenated script
imports the `agent_team` package without an editable install. **CLEAN.**
---
## ✅ D-HARDENING (clean, and matches the locked decision)
- Unit: `NoNewPrivileges=true`, `ProtectSystem=full`, `ProtectHome=read-only`,
`ReadWritePaths=/home/adam/orchestrator/agent-team/state`. **Retained unchanged**
in this PR (locked decision — do not revert to secrev parity).
- The `ReadWritePaths` carve-out matches the ledger + audit-log location
(`state/agent_team.sqlite`, `state/audit.log.jsonl`). **CLEAN.**
---
## Summary table
| ID | Sev | Status | One-line |
|---|---|---|---|
| D-1 | 🔴 | ✅ RESOLVED (PR) | `serve` starts `SlackListener` (Slack + app-token gated; Slack stays optional) |
| D-2 | 🔴 | ✅ RESOLVED (PR) | unit loads `~/orchestrator/.env` for the P2 GPT-4.1 review provider key |
| D-7 | 🟡 | ✅ RESOLVED (PR) | `ExecStart` points at the agent-team venv interpreter |
| D-6 | 🟡 | ✅ RESOLVED | `slack_bolt` now exercised by the daemon (consequence of D-1) |
| D-4 | 🟠 | NOTE | pip list omits `requests` — runbook Step 4 installs it |
| D-5 | 🟠 | NOTE | agent-team runtime deps not pinned in `requirements.txt` — runbook pins langgraph |
| operator-CLI | — | NOTE | `run-team.py` vs `operator_cli.py` divergent verbs — use `run-team.py` |
| D-SLACKVAR | ✅ | CLEAN | `SLACK_CHANNEL_ID` matches everywhere |
| D-ENV-* | ✅ | CLEAN | bot/oauth/owner/app-token env names match; all read sites now exist |
| D-SUBCMD | ✅ | CLEAN | every `run-team.py` verb/flag the docs cite exists |
| D-HARDENING | ✅ | CLEAN | unit hardening retained unchanged (locked decision) |

View file

@ -0,0 +1,290 @@
# OPERATOR-RUNBOOK — R720 agent-team coordinator incident handling
The on-call runbook for the always-on `agent-team-coordinator` daemon on the
`sh-secrev` R720 VM. Covers pipeline stalls, stuck/parked tasks, failed
human-in-the-loop resumes, budget exhaustion mid-pipeline, transport outages, and
the COMPLACENCY / COVERAGE alarms. Grounds every recovery in real code (design
§5 escalation ladder + §6.6 contention/park policy; Phase-6 requirement).
> **CLI used throughout: `run-team.py`** (the entry CLI), run from
> `~/orchestrator/agent-team` with the venv active so it hits the default ledger
> (`state/agent_team.sqlite`) and audit log (`state/audit.log.jsonl`). Do **not**
> use `agent_team/operator_cli.py` — it has divergent verbs (no `show`, required
> `--db`/`--audit-log`, and a `force-resume` that supersedes). See DEPLOY-AUDIT.md.
```bash
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team && . .venv/bin/activate
```
## Verb reference (all confirmed in `run-team.py:build_parser`)
| Verb | Effect | Destructive? |
|---|---|---|
| `list` / `list --all` / `list --parked` / `list --status <state>` | read pending questions (default `open`) | no |
| `show <question_id>` | print one ledger row (JSON) | no |
| `redeliver <question_id>` | clear `channel_ref` so the reconcile loop re-posts an `open` question | no (audit-logged) |
| `expire <question_id> --confirm` | force `open`→`expired` | yes |
| `answer <question_id> --answer <p> [--via <id>] --confirm` | answer-on-behalf (first-answer-wins CAS) | yes |
| `force-resume <question_id> --confirm` | reopen an `expired` (parked) question; for `answered` records resume intent | yes |
| `supersede <question_id> --confirm` | mark a stale `open`/`answered` row `superseded` | yes |
| `start --task "..." [--transport ...] [--dry-run]` | start one task to the human gate | no |
`--operator <name>` (global) sets the audit attribution; it defaults to the OS
login. Destructive verbs require `--confirm` and write an attempt-then-outcome
record to the audit log **before** mutating.
First triage for any incident:
```bash
systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e --since "-2h" | tail -100
python3 run-team.py list --all # full ledger snapshot
python3 run-team.py list --parked # non-open rows (parked-task context)
```
---
## Incident 1 — Pipeline stall (tasks not advancing)
**Symptoms:** `list` shows `open` (or `answered`) rows that never progress; the
journal shows no `tick` activity or repeated errors.
**Diagnose:**
```bash
systemctl is-active agent-team-coordinator.service # "active" expected
journalctl -u agent-team-coordinator.service -e | tail -60
# Daemon dead/looping on restart? Check the unit + deps:
.venv/bin/python -c "import langgraph, slack_sdk, slack_bolt, requests; print('deps ok')"
```
**Recover:**
1. If the daemon is dead and `Restart=on-failure` is flapping, read the journal
for the import/auth error. A missing venv dep (D-4/D-5) or a missing
`CLAUDE_CODE_OAUTH_TOKEN` is the usual cause — fix the env / venv, then:
```bash
sudo systemctl restart agent-team-coordinator.service
```
2. The startup `recover()` sweep re-drives `answered`-but-unresumed rows and
re-posts `open` rows that lost their `channel_ref`, so a clean restart
converges from the durable ledger. Confirm with `journalctl ... | tail` and
`list --all`.
3. A single task stuck `open` with a stale/missing post: re-deliver it.
```bash
python3 run-team.py show <qid>
python3 run-team.py redeliver <qid> # clears channel_ref; reconcile re-posts
```
**Escalate** (see the ladder) if a restart does not clear it within one tick
cadence and the journal shows a non-transient error.
---
## Incident 2 — Stuck / parked task
A task parks (design §6.6) when its clarifier question **expires** with no answer,
or its budget headroom drops below reserve, or a checkpoint is corrupt. Parked =
the durable ledger row is no longer `open` (it is `expired`), and an ALARM was
raised, not spun on.
**Diagnose:**
```bash
python3 run-team.py list --parked
python3 run-team.py show <qid> # status, thread_id, deadline_at, answered_via
```
**Recover — depends on why it parked:**
- **Expired with no answer (the common case)** — un-park by reopening the expired
question; it is then re-delivered for an answer:
```bash
python3 run-team.py force-resume <qid> --confirm
# "force-resume: reopened expired question <qid>; it will be re-delivered"
python3 run-team.py show <qid> # status flips back to "open"
```
Then answer it (Slack or CLI) to drive it forward.
- **Answered but not yet resumed** — the recovery sweep handles it; `force-resume`
records intent and reports that (no mutation):
```bash
python3 run-team.py force-resume <qid> --confirm
# "force-resume: question <qid> is answered and pending resume; the recovery
# sweep will resume it (intent recorded)"
sudo systemctl restart agent-team-coordinator.service # forces the recover() sweep now
```
- **Stale / wrong question that should be abandoned** — supersede it so it stops
surfacing as parked context:
```bash
python3 run-team.py supersede <qid> --confirm
```
`MAX_PARK` FIFO-aging: a task that exceeds the park window escalates (ALARM + a
Jira ticket per the ladder) rather than starving silently.
---
## Incident 3 — Failed human-in-the-loop resume
**Symptoms:** an answer was submitted (Slack or CLI) but the graph did not
advance.
**Diagnose:**
```bash
python3 run-team.py show <qid> # is status "answered"? what answered_via?
journalctl -u agent-team-coordinator.service -e | grep -iE "resume|answer|<qid>"
```
**Recover:**
- **Status is `answered` but no resume drained** — the resume queue is in-process;
a daemon restart triggers the `recover()` sweep that re-drives `answered` rows
via the turn-guarded `ResumeWorker` (idempotent — an already-advanced thread
supersedes-and-skips):
```bash
sudo systemctl restart agent-team-coordinator.service
python3 run-team.py show <qid> # confirm it advanced
```
- **Live Slack answer never registered** — the listener fails closed. Check:
```bash
journalctl -u agent-team-coordinator.service -e | grep -i "Slack answer"
# "rejecting Slack answer: owner allowlist is unconfigured ..." -> set
# AGENT_TEAM_SLACK_OWNER_IDS in ~/secrev.env and restart.
# "rejecting Slack answer ... unauthorized sender" -> the answerer's user id is
# not in the allowlist. Add it, or answer via the CLI on their behalf:
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
```
- **Answer lost the compare-and-set (`not open`)** — the row was already
answered/expired/superseded. Inspect with `show`; if it parked, go to Incident 2.
---
## Incident 4 — Budget exhaustion mid-pipeline
The shared Claude budget ledger (`budget_ledger` table) enforces per-call + total
nightly caps (design §6.1, §6.6). A stage that would breach the reserve **parks**
the task (deferred, not dropped) and ALARMs — it never loop-drains the pool.
**Diagnose:**
```bash
journalctl -u agent-team-coordinator.service -e | grep -iE "budget|reserve|park"
# Inspect today's spend directly (no CLI verb for the budget ledger; read it).
# Columns (agent_team/db/schema.py budget_ledger DDL): thread_id, stage, model,
# billing_mode, input_tokens, output_tokens, usd_cost, recorded_at, day_bucket.
sqlite3 state/agent_team.sqlite \
"SELECT day_bucket, thread_id, ROUND(SUM(usd_cost),4) AS usd, SUM(input_tokens) AS in_tok
FROM budget_ledger GROUP BY day_bucket, thread_id ORDER BY day_bucket DESC LIMIT 20;"
```
**Recover:**
- **Wait for the next budget window** — budget-exhausted roles/tasks are deferred
via the rotation pointer and picked up next cycle; this is the intended
behavior, not a failure. The parked task surfaces in `list --parked`.
- **Force a specific parked task forward now** (e.g. it is urgent and headroom has
since freed): un-park it and let the daemon re-run the stage within the
remaining cap:
```bash
python3 run-team.py force-resume <qid> --confirm
```
- **Confirm `ANTHROPIC_API_KEY` is absent** — its presence would silently meter to
API rates and blow the budget model (the billing seam pops it defensively, but
it must not be set):
```bash
grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # both must print 0
```
Do **not** raise the cap to push a task through without Adam's decision — budget
realism is a deliberate guardrail. Escalate per the ladder if a task repeatedly
parks on budget across nights.
---
## Incident 5 — Transport outage (Slack / GitHub down or misconfigured)
**Symptoms:** clarifier posts fail; the journal shows transport errors or the
inbound listener is not started.
**Diagnose:**
```bash
journalctl -u agent-team-coordinator.service -e | grep -iE "Slack|listener|post|transport"
# Did the inbound listener start?
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
# "started (Socket Mode ...)" -> inbound up
# "not started (... SLACK_APP_TOKEN is unset)" -> Slack inbound intentionally off
```
**Recover:**
- **Outbound post failed (Slack API / network)** — the ledger row stays `open`
with no `channel_ref` (the responder leaves it for reconcile). After the
transport recovers, the `recover()` sweep (on restart) or a manual `redeliver`
re-posts:
```bash
python3 run-team.py redeliver <qid>
```
- **Inbound listener down / never started** — the daemon still posts and expires;
it just cannot hear Slack. **The CLI `answer` path is the outage fallback** —
it runs the identical compare-and-set with no socket:
```bash
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
```
Fix `SLACK_APP_TOKEN` / `SLACK_BOT_TOKEN` in `~/secrev.env`, then
`sudo systemctl restart agent-team-coordinator.service` to re-establish Socket
Mode. (D-1: the listener starts only when the transport is live Slack AND
`SLACK_APP_TOKEN` is set; it is optional by design.)
- **Slack listener crash-looping** — the listener runs on an isolated daemon
thread; a crash is logged (`inbound Slack listener thread exited with an error`)
and takes down only the inbound socket, not the maintenance loop. Restart the
service to relaunch the thread once the cause is fixed.
---
## Incident 6 — COMPLACENCY / COVERAGE alarms (Plane-1 checkers)
These come from the nightly checker run, not the coordinator daemon (design §6.4,
§6.6):
- **COMPLACENCY ALARM** — a checker role missed a planted canary fault. That role
is **skipped** for the night (it never runs silently degraded). Recover: inspect
the canary corpus / the role's checker under `security-review/checkers/`, fix the
regression, re-run that checker's dry-run, and confirm the canary passes before
re-enabling.
- **COVERAGE ALARM** — a role slipped its rotation slot (e.g. budget-deferred). It
is **deferred via the rotation pointer, never dropped**, and picked up next
cycle. Recover: confirm the rotation pointer advanced (it is rebuildable from
report history) and that the role runs on the next cadence; investigate only if
it slips repeatedly.
Both alarms are **report + ALARM-only** (design D3): nothing posts on a clean
state, no auto-Jira/Notion writes from the checker itself. Persistent alarms
follow the escalation ladder below.
---
## Escalation ladder (design §5 / §6.6, resolves Q3)
Anything that does not clear on the first ALARM escalates — but **ALARM-only in
spirit** (nothing posts on a clean state):
1. **Re-alarm on a backoff.** A confirmed critical (or a COMPLACENCY / COVERAGE
alarm) that persists re-alarms to Slack each night it is still unresolved, on
a backoff so it does not spam.
2. **Open a Jira tracking ticket after `N` nights** (default **N = 3**). If the
condition still has not cleared, the coordinator opens an **INFRA** Jira ticket
so it cannot quietly linger. The same ladder applies to a parked task that
exceeds `MAX_PARK`.
3. **Human (Adam) takes it from the Jira ticket.** For a parked task, recover via
Incidents 2–4 above; for a checker alarm, via Incident 6.
When you resolve an incident, record the action — destructive CLI verbs already
write an attributable attempt+outcome record to `state/audit.log.jsonl`; for
non-CLI recoveries note it on the Jira ticket.
---
## Post-incident
- Confirm the daemon is `active` and the ledger has no unexpected parked rows
(`list --parked`).
- If you restored the VM snapshot or wiped the ledger, re-run the relevant
PROVISIONING-RUNBOOK steps.
- Update `project_r720_agent_team` memory if the incident revealed a durable
fact (a new failure mode, a config that must change).

View file

@ -0,0 +1,286 @@
# P1-DEMO-SCRIPT — live four-criteria acceptance demo (design §3.3.1 / §7.1 P1)
The §7.1 P1 exit gate: demonstrate, on the live box, all four durable
human-in-the-loop criteria before P1 is accepted:
- **(a)** kill the box mid-wait and have the task resume after restart;
- **(b)** submit a duplicate answer and confirm it no-ops;
- **(c)** submit an answer after the deadline expired and confirm it is rejected
and the task parks;
- **(d)** two tasks suspended concurrently resume independently to the correct
thread.
Every command below is grounded in the **actual** code surface
(`run-team.py`, `coordinator.py`, `slack_listener.py`, `responder.py`,
`resume_worker.py`, `db/schema.py`, `graph.py`). No invented flags. Where the
code does not expose a needed knob (e.g. a short deadline), the script uses a
direct `sqlite3` write against the documented `pending_questions` schema and says
so.
## Two answer paths — live Slack OR the operator CLI
**D-1 is fixed:** `Coordinator.serve()` now starts the inbound `SlackListener`
when the live transport is Slack AND `SLACK_APP_TOKEN` is set, so the
**live Slack answer round-trip works**. You can run the demo either way:
- **Live Slack** — Adam clicks the Block Kit button / replies in the channel; the
listener normalizes the event, runs the AUTHZ-01 owner check, and drives the
first-answer-wins compare-and-set (`responder.submit_answer` →
`db.schema.answer_question`).
- **Operator CLI** — `run-team.py answer <qid> --answer ... --confirm` runs the
**identical** compare-and-set (audit-logged answer-on-behalf). Useful when the
Slack app is not yet provisioned, or to script the assertions.
Both exercise the same durable mechanic; the human gate decision is always
Adam's. The assertions below assert on the **durable ledger status** (the §3.3.1
source of truth) and are identical for either path. The examples use the CLI
`answer --confirm` form so they are copy-pasteable; substitute "Adam answers in
Slack" wherever you prefer the live path.
> The automated proof of this mechanic is `tests/sim/test_p1_exit_criteria.py`
> (a `SimPipeline` harness) and `tests/test_coordinator.py` (the serve/listener
> wiring). This script is the live-box demonstration on top of that.
## Verb map (use `run-team.py`, not `operator_cli.py`)
| Need | `run-team.py` verb | Notes |
|---|---|---|
| start a task to the human gate | `start --task "..." [--transport slack] [--dry-run]` | mints a `thread_id`, posts the clarifier, writes the `open` ledger row |
| list waiting questions | `list` (default `open`) / `list --all` / `list --parked` | JSON rows |
| inspect one row | `show <question_id>` | JSON row |
| answer on the task's behalf | `answer <question_id> --answer <payload> --confirm` | destructive, audit-logged; first-answer-wins CAS |
| force-expire an open question | `expire <question_id> --confirm` | destructive; flips `open`→`expired` |
| un-park (reopen) an expired question | `force-resume <question_id> --confirm` | only acts on `expired` rows (reopens them) |
`operator_cli.py` has different verbs (`force-expire`, `answer-on-behalf`) and
**no `show`**, and requires `--db`/`--audit-log` — do not use it here.
## Preconditions
```bash
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
cd ~/orchestrator/agent-team
. .venv/bin/activate # so run-team.py uses the venv deps
# Confirm the ledger exists (Step 5 of the runbook):
python3 run-team.py list --all # [] on a fresh DB is fine
# For the live-Slack path, confirm the listener started:
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
# -> "inbound Slack listener started (Socket Mode, background thread)"
```
Run all `run-team.py` commands from `~/orchestrator/agent-team` so they hit the
default ledger (`state/agent_team.sqlite`) and default audit log
(`state/audit.log.jsonl`). The clarifier is posted to `SLACK_CHANNEL_ID`.
> **Resume execution model.** A won answer (rowcount 1) enqueues a `ResumeJob`
> onto the coordinator's in-process `resume_queue`; the daemon drains it on the
> next `tick()`, or the startup `recover()` re-drives any `answered`-but-unresumed
> row after a restart. So a CLI `answer --confirm` (or a live Slack answer) flips
> the ledger row to `answered`; the **running daemon** then resumes the graph.
> The demo asserts on the durable ledger status.
---
## (a) Crash-safe resume — kill mid-wait, restart, task resumes 🧑 Adam answers
**Setup — drive a task to the clarifier wait.** With the daemon already running:
```bash
python3 run-team.py start --task "demo-a: trivial scoped task"
# prints a thread_id, e.g. 3f2a... (record it as $TID_A)
python3 run-team.py list
# expect one row: status="open", a thread_id, a question_id (record as $QID_A),
# channel_ref set (Slack ts) or null if posting is dry/unavailable.
```
**Kill the box mid-wait, then restart:**
```bash
sudo systemctl stop agent-team-coordinator.service
# (optionally reboot the VM here for a stronger demonstration)
sudo systemctl start agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e | tail -40 # expect "starting" + recover() sweep
```
**ASSERT — durable state survived the kill (no answer yet):**
```bash
python3 run-team.py show $QID_A
# PASS: status == "open" (the question was NOT lost across the restart)
```
**🧑 Human gate — Adam answers after the restart (Slack OR CLI):**
```bash
# Live Slack: Adam replies/clicks in the channel. OR via the CLI:
python3 run-team.py answer $QID_A --answer "scope: just demo, no real change" --confirm
python3 run-team.py show $QID_A
# PASS: status == "answered"
# Wait one tick (~30s, DEFAULT_POLL_INTERVAL) for the daemon to drain the resume:
journalctl -u agent-team-coordinator.service -e | tail -20
```
**PASS criterion (a):** the question stayed `open` across the kill, and an answer
submitted *after* the restart drove it forward.
---
## (b) Duplicate answer is a no-op 🧑 Adam answers
```bash
python3 run-team.py start --task "demo-b: duplicate-answer test"
python3 run-team.py list # record the new question_id as $QID_B (status open)
```
**🧑 First answer (the winning one) — Slack OR CLI:**
```bash
python3 run-team.py answer $QID_B --answer "first-answer" --via "slack:U1" --confirm
# expect: "answered question <QID_B> (via slack:U1)" (exit 0)
```
**Duplicate / second answer (must lose the compare-and-set):**
```bash
python3 run-team.py answer $QID_B --answer "second-answer" --via "github:U2" --confirm
# expect: "answer no-op: question <QID_B> was not 'open' ..." on stderr, exit 1
```
For the **live-Slack** variant, Adam (or a second authorized owner) answers the
same message twice; the second event loses the CAS and is logged
`ignored Slack answer for question_id=... (not open: duplicate ...)`.
**ASSERT — the first answer is preserved:**
```bash
python3 run-team.py show $QID_B
# PASS: status == "answered"; answered_via == "slack:U1" (the FIRST answer)
```
**PASS criterion (b):** first-answer-wins (`rowcount==1`); the duplicate hits the
`BEGIN IMMEDIATE` compare-and-set in `answer_question` and is ignored
(`rowcount==0`) — no second resume, no overwrite.
---
## (c) Past-deadline answer rejected + task parks 🧑 Adam observes
> **Code reality:** `run-team.py start` always sets the clarifier deadline from
> `DEFAULT_CLARIFY_DEADLINE = 24h`. There is **no CLI flag for a short deadline**.
> Two faithful ways to demo expiry without waiting 24h:
### Option C1 — force the deadline past, let the daemon's sweep expire it (most faithful)
```bash
python3 run-team.py start --task "demo-c: deadline test"
python3 run-team.py list # record $QID_C (status open)
# Set this question's deadline into the past directly in the ledger:
sqlite3 state/agent_team.sqlite \
"UPDATE pending_questions SET deadline_at='2000-01-01T00:00:00+00:00' WHERE question_id='$QID_C';"
# Wait one daemon tick (~30s) for the deadline sweep to flip it, OR observe:
journalctl -u agent-team-coordinator.service -e | tail -20
# expect the park ALARM line: "task parked: clarifier question <QID_C> expired ..."
```
### Option C2 — operator force-expire (if you do not want to touch the DB)
```bash
python3 run-team.py start --task "demo-c: deadline test"
python3 run-team.py list # record $QID_C
python3 run-team.py expire $QID_C --confirm # destructive, audit-logged; open->expired
```
**ASSERT — the question is expired and the late answer is rejected:**
```bash
python3 run-team.py show $QID_C
# PASS: status == "expired"
# 🧑 Adam submits a LATE answer (Slack OR CLI) — it must lose the CAS:
python3 run-team.py answer $QID_C --answer "too-late" --confirm
# expect: "answer no-op: question <QID_C> was not 'open' ..." stderr, exit 1
python3 run-team.py show $QID_C
# PASS: status still "expired"; answer_json still NULL
# The parked task surfaces in the parked view:
python3 run-team.py list --parked # PASS: $QID_C appears here
```
**Deliberate un-park (proves the operator recovery path, §6.6):**
```bash
python3 run-team.py force-resume $QID_C --confirm
# expect: "force-resume: reopened expired question <QID_C>; it will be re-delivered"
python3 run-team.py show $QID_C
# status flips back to "open" (reopen_question), deadline_at cleared.
```
**PASS criterion (c):** the expired question rejects the late answer, the task
parks rather than spins, and the operator can deliberately un-park it via
`force-resume` (which reopens only an `expired` row).
> **`force-resume` semantics (verified):** `run-team.py force-resume` reopens an
> `expired` question (the parked case). For an `answered` question it records
> intent and reports the recovery sweep will resume it (no mutation). For
> `open`/`superseded`/absent it is a no-op exit 1. It does **not** supersede.
---
## (d) Two concurrent tasks resume independently 🧑 Adam answers
```bash
python3 run-team.py start --task "demo-d task A" # record $TID_A2, then:
python3 run-team.py start --task "demo-d task B" # record $TID_B2
python3 run-team.py list
# expect TWO open rows with DISTINCT thread_id AND distinct question_id.
# Record $QID_A2 and $QID_B2 — match by thread_id.
```
**(Optional) restart the daemon first** to also show concurrent tasks survive a
restart, then answer.
**🧑 Answer the SECOND task first, with a distinct answer, then the first
(Slack OR CLI):**
```bash
python3 run-team.py answer $QID_B2 --answer "answer-for-B" --via "slack:U2" --confirm
python3 run-team.py answer $QID_A2 --answer "answer-for-A" --via "slack:U1" --confirm
# Wait one daemon tick (~30s) for both resumes to drain.
```
**ASSERT — each task carries its OWN answer; no cross-talk:**
```bash
python3 run-team.py show $QID_A2
# PASS: status "answered", answered_via "slack:U1", thread_id == $TID_A2
python3 run-team.py show $QID_B2
# PASS: status "answered", answered_via "slack:U2", thread_id == $TID_B2
python3 run-team.py list --all
# PASS: the two rows resolved on their own thread_id; no cross-contamination.
```
**PASS criterion (d):** two concurrently-suspended tasks each resumed to their
own `thread_id` with their own answer — answering B before A did not misroute,
and the per-thread single-flight guard kept them independent.
---
## Acceptance
P1 is accepted only when **(a), (b), (c), and (d) all pass** on the live box.
Record the four `show` outputs (or `journalctl` excerpts) as evidence. Then
complete the PROVISIONING-RUNBOOK post-session definition-of-done (memory +
Confluence + the mandatory `/sh-security-review` on the Slack inbound listener).
## What can only be verified on the live box
- That the daemon actually drains the resume and the LangGraph checkpoint
advances (asserted via ledger status + journalctl; the graph-state advance is
observable only on the box).
- The live **Slack** post/answer round-trip end-to-end (the listener is wired —
D-1 fixed — but the live Socket Mode socket + a real `SLACK_BOT_TOKEN` /
`SLACK_APP_TOKEN` / channel membership are only present on the box).
- The exact `channel_ref` value (Slack message `ts`) — depends on a live Slack
post succeeding.

View file

@ -0,0 +1,402 @@
# PROVISIONING RUNBOOK — R720 agent-team Plane-2 coordinator
This is the ordered command sequence for the **operator-present** provisioning
session that stands up the always-on `agent-team-coordinator` daemon on the
`sh-secrev` R720 VM. Derived from `agent-team/DEPLOY-R720.md`,
`security-review/DEPLOY-R720.md`, and `docs/r720-agent-team-design.md`
(§3.3.1 / §7 / §7.1).
> ## State of this runbook
>
> The build is **merged and final**: all six Plane-1 checkers
> (`aws-posture`, `compliance-drift`, `confluence-doc`, `dependency-cve`,
> `doc-drift`, `plan-groomer`), the Tier-3 dep-bump **fixer**
> (`run-team.py fix --dry-run`), and the Plane-1→Plane-2 **P5 cross-plane loop**
> (`run-team.py intake-checker`) exist in the tree. The three deploy-correctness
> bugs the provisioning-prep audit found are **FIXED** (see DEPLOY-AUDIT.md):
>
> - **D-1 RESOLVED** — `Coordinator.serve()` now starts the inbound Slack
> `SlackListener` concurrently with the tick/drain loop when Slack is the live
> transport AND `SLACK_APP_TOKEN` is set. **The demo can use the live Slack
> answer path** (P1-DEMO-SCRIPT.md), or the operator-CLI `answer` path.
> - **D-2 RESOLVED** — the systemd unit loads `~/orchestrator/.env` (for the P2
> GPT-4.1 review loop's provider key) in addition to `~/secrev.env`.
> - **D-7 RESOLVED** — the unit's `ExecStart` points at the agent-team venv
> interpreter, not the system `python3`.
>
> Still open as provisioning notes (not blockers): **D-4/D-5** (agent-team
> runtime deps are installed ad-hoc into the venv and are not pinned in
> `requirements.txt`), and the **operator-CLI divergence** (`run-team.py` vs
> `agent_team/operator_cli.py` have different verb names — use `run-team.py`).
---
## Scope and ground rules
**Hard rules (carried from the design and global instructions):**
- **Snapshot before any stateful change** (design §7; `feedback_ec2_replacement_snapshot`).
- **Every stateful step has an exercised rollback.**
- Secrets use **placeholder names only**; real values are entered by the operator
at the box and never echoed into shell history.
- The box is **read-only / subscription-OAuth only**. **No `ANTHROPIC_API_KEY`**
on this host (it would silently win over OAuth and meter to API rates —
`billing.claude_invoke` pops it defensively, but it must not be present).
- **No IAM / OIDC is involved in the coordinator deploy** (P1/P2). IAM enters
only at the **P3-live flip** (see the dedicated section), which is gated on the
mandatory GPT-4.1 cross-review + `/sh-security-review`.
**Legend per step:**
- 🧑 **OPERATOR-REQUIRED** — needs the human (snapshot, secrets, the live human
gate, go/no-go). Cannot be automated.
- 🤖 **MECHANICAL** — deterministic; an operator runs it but it needs no judgment.
## Host facts (from both DEPLOY-R720.md files)
- Hypervisor: R720 at `10.10.60.40` (Windows Server 2022, Hyper-V).
- VM: `sh-secrev`, Ubuntu 24.04, **4GB / 2 vCPU / 40GB** dynamic vhdx, `10.10.60.120`.
- Reach: `ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120` (key-only, NOPASSWD sudo).
- Repo on the box: `~/orchestrator/` (rsync from the Mac, **NOT** a git clone).
The package lives at `~/orchestrator/agent-team/`.
- Shares `~/secrev.env` (mode 600) with the secrev sweep, and `~/orchestrator/.env`
(mode 600) for non-Claude provider keys (same files the secrev unit loads).
---
## STEP 1 — Snapshot the VM 🧑 OPERATOR-REQUIRED
**On the R720 host (Hyper-V), before anything else.** This is the one-command
undo for every change below.
```powershell
# On the R720 Windows host (PowerShell, as admin):
Checkpoint-VM -Name sh-secrev -SnapshotName "pre-agent-team-coordinator-$(Get-Date -Format yyyyMMdd-HHmm)"
Get-VMSnapshot -VMName sh-secrev # confirm the checkpoint exists
```
**ROLLBACK (whole session):**
```powershell
Restore-VMSnapshot -VMName sh-secrev -Name "<the checkpoint name above>" -Confirm:$false
Start-VM -Name sh-secrev
```
> Do not proceed until the checkpoint is confirmed present.
---
## STEP 2 — Rsync the repo to the box 🤖 MECHANICAL (verify manifest 🧑)
**From the Mac.** Same pattern/excludes as the secrev deploy. The team **scans
the same `~/repo-mirrors` corpus secrev already maintains** — this rsync ships
*code*, not mirrors.
```bash
# From the Mac (sync the canonical repo, not a worktree):
rsync -av --exclude .env --exclude .venv --exclude .git --exclude .claude \
~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/
```
Surface that must be on the box:
| Path (under `~/orchestrator/`) | Why it must ship |
|---|---|
| `agent-team/run-team.py` | the operator entry CLI |
| `agent-team/agent_team/**` | the package (coordinator, graph, ledger, transports, slack_listener, fixer) |
| `agent-team/systemd/agent-team-coordinator.service` | the daemon unit |
| `run.py` + the orchestrator package | the GPT-4.1 cross_reviewer the P2 review loop shells |
| `requirements.txt` | pin reference for langgraph / checkpoint-sqlite |
| `security-review/lib/**` | shared sweep substrate (Phase 0) |
| `security-review/checkers/**` | the six Plane-1 checkers + fixtures |
**ROLLBACK:** rsync is additive; restore the Step-1 snapshot to revert code state.
---
## STEP 3 — Write secrets 🧑 OPERATOR-REQUIRED
**On the VM.** The coordinator unit loads **two** EnvironmentFiles (both
optional via the leading `-`): `~/secrev.env` (agent-team runtime keys) and
`~/orchestrator/.env` (the non-Claude provider key for the P2 review loop).
**Placeholder names only — the operator pastes real values.** Do not echo real
tokens into shell history (use an editor or `read -s`).
`~/secrev.env` (mode 600) — append the agent-team keys:
```bash
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
# Edit ~/secrev.env (mode 600) and add — values are placeholders:
# CLAUDE_CODE_OAUTH_TOKEN=<CLAUDE_CODE_OAUTH_TOKEN> # from `claude setup-token`
# SLACK_BOT_TOKEN=<SLACK_BOT_TOKEN> # xoxb-..., chat:write (posts questions)
# SLACK_APP_TOKEN=<SLACK_APP_TOKEN> # xapp-..., connections:write (Socket Mode inbound)
# SLACK_CHANNEL_ID=<SLACK_CHANNEL_ID> # C0..., target clarifier channel
# AGENT_TEAM_SLACK_OWNER_IDS=<SLACK_OWNER_USER_IDS> # comma-separated U... ids of authorized answerers
chmod 600 ~/secrev.env
```
`~/orchestrator/.env` (mode 600) — the non-Claude provider key for the GPT-4.1
review loop (same file secrev uses; if it already exists with the key, leave it):
```bash
# OPENAI_API_KEY=<OPENAI_API_KEY> # (or the provider key cross_reviewer/GPT-4.1 needs)
chmod 600 ~/orchestrator/.env
```
**Contract notes (verified against the code — see DEPLOY-AUDIT.md):**
- `SLACK_CHANNEL_ID` is correct — `run-team.py _build_transport` reads exactly
`os.environ.get("SLACK_CHANNEL_ID")`. Do **not** use `SLACK_CHANNEL`.
- `SLACK_APP_TOKEN` is now **read by the daemon**: `Coordinator.serve()` starts
the inbound `SlackListener` when the transport is live Slack AND `SLACK_APP_TOKEN`
is set. Without it the daemon still runs (posts + expires) but never hears Slack
replies — Slack stays optional by design.
- `AGENT_TEAM_SLACK_OWNER_IDS` **fails closed** (AUTHZ-01): if unset/empty the
listener rejects **every** answer. It must be set for the live human gate.
- **CRITICAL:** confirm `ANTHROPIC_API_KEY` is NOT present:
```bash
grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # must print 0 for both
```
**ROLLBACK:** strip exactly the appended keys (keep secrev keys intact), or
restore the Step-1 snapshot. Do NOT blindly truncate — secrev keys live here too.
---
## STEP 4 — Create the venv + install pip deps 🤖 MECHANICAL
**On the VM.** A dedicated venv under `agent-team/.venv` (excluded from rsync).
The systemd unit's `ExecStart` points at **this venv's interpreter** (D-7 fixed),
so the deps MUST land here.
```bash
cd ~/orchestrator/agent-team
python3 -m venv .venv
. .venv/bin/activate
pip install langgraph==1.1.10 langgraph-checkpoint-sqlite==3.1.0 \
claude-agent-sdk slack_sdk slack_bolt requests
```
> **Pin note (D-5):** `requirements.txt` pins `langgraph==1.1.10` /
> `langgraph-checkpoint-sqlite==3.1.0`; match those exactly here. The other
> runtime deps (`claude-agent-sdk`, `slack_sdk`, `slack_bolt`, `requests`) are
> not yet in `requirements.txt` (D-5 open) — installed ad-hoc here. `requests`
> is required by the GitHub transport/intake (D-4). `anthropic` is **not**
> installed (only the opt-in `api` billing mode needs it).
> `slack_bolt` is now actually exercised (D-1 fixed: the daemon starts the
> Socket Mode listener).
**Verify the imports resolve (using the venv interpreter the unit will use):**
```bash
.venv/bin/python -c "import langgraph, langgraph.checkpoint.sqlite, slack_sdk, slack_bolt, requests; print('deps ok')"
.venv/bin/python -c "import claude_agent_sdk; print('agent-sdk ok')"
```
**ROLLBACK:** `deactivate 2>/dev/null; rm -rf ~/orchestrator/agent-team/.venv`
---
## STEP 5 — Initialize the durable ledger DB 🤖 MECHANICAL
**On the VM, venv active.** Idempotent; creates `state/agent_team.sqlite` with
the `pending_questions` + `budget_ledger` + `schema_meta` tables (LangGraph
`SqliteSaver` creates its own tables in the same file on first run).
```bash
cd ~/orchestrator/agent-team
. .venv/bin/activate
python3 run-team.py init-db
# Expect: "initialized ledger DB at .../state/agent_team.sqlite"
ls -l state/ # agent_team.sqlite present; state/ is gitignored
```
**ROLLBACK (reset the ledger only):**
```bash
cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak}
rm ~/orchestrator/agent-team/state/agent_team.sqlite*
# re-run `python3 run-team.py init-db` to recreate empty tables.
```
---
## STEP 6 — Install + start the systemd unit 🤖 MECHANICAL (go/no-go 🧑)
**On the VM, as root.** Installs the long-running coordinator daemon. Keep the
**hardening as-shipped**: `NoNewPrivileges`, `ProtectSystem=full`,
`ProtectHome=read-only` + `ReadWritePaths=.../agent-team/state` (locked decision).
```bash
sudo cp ~/orchestrator/agent-team/systemd/agent-team-coordinator.service /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now agent-team-coordinator.service
systemctl status agent-team-coordinator.service
journalctl -u agent-team-coordinator.service -e -f
# expect: "agent-team coordinator starting"
# and (if Slack + SLACK_APP_TOKEN provisioned):
# "inbound Slack listener started (Socket Mode, background thread)"
# or otherwise:
# "inbound Slack listener not started (... or SLACK_APP_TOKEN is unset) ..."
```
The unit (post-fix) runs
`/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve` from
`WorkingDirectory=/home/adam/orchestrator/agent-team` as `User=adam`, loading
**both** `EnvironmentFile=-/home/adam/secrev.env` and
`EnvironmentFile=-/home/adam/orchestrator/.env`, `Restart=on-failure`.
> **D-7 fixed:** `ExecStart` now resolves the venv interpreter, so the Step-4
> deps are on the path. **D-1 fixed:** `serve` starts the inbound `SlackListener`
> when Slack + app token are provisioned. **Go/no-go:** confirm the journal shows
> the listener line you expect for your transport choice before Step 8.
**ROLLBACK (exercised — tear the unit down cleanly):**
```bash
sudo systemctl disable --now agent-team-coordinator.service
sudo rm /etc/systemd/system/agent-team-coordinator.service
sudo systemctl daemon-reload
systemctl status agent-team-coordinator.service # should report "could not be found"
```
The ledger under `state/` is untouched. Step-1 snapshot is the whole-host fallback.
---
## STEP 7 — Verify the Slack inbound listener 🧑 OPERATOR-REQUIRED
The clarifier gate is two halves: **outbound** (post the question — done by
`serve`/`start` via the Slack poster) and **inbound** (receive Adam's answer —
the `SlackListener` over Socket Mode).
> **D-1 RESOLVED.** `Coordinator.serve()` now constructs and starts
> `SlackListener` on a background daemon thread, concurrently with the tick/drain
> loop, **when** the live transport is a `SlackTransport` AND `SLACK_APP_TOKEN`
> is set. It shares the coordinator's own transport, ledger, and resume queue,
> and stops cleanly on shutdown. The AUTHZ-01 owner allowlist + the open-status
> compare-and-set are unchanged — the listener still fails closed on an empty
> `AGENT_TEAM_SLACK_OWNER_IDS`.
**Confirm the live inbound path is up:**
```bash
# In the journal (Step 6) expect the "inbound Slack listener started" line.
# Verify the two gating vars are present in the unit's environment:
grep -c SLACK_APP_TOKEN ~/secrev.env # 1
grep -c AGENT_TEAM_SLACK_OWNER_IDS ~/secrev.env # 1 (else the gate rejects all answers)
```
If you are **not** using Slack as the transport (or are deliberately running
without the app token), the daemon runs the maintenance loop only and the demo
uses the operator-CLI `answer` path — both are valid (P1-DEMO-SCRIPT.md).
**ROLLBACK:** none needed — verification only, no host state change.
---
## STEP 8 — Live P1 four-criteria acceptance demo 🧑 OPERATOR-REQUIRED
Run **P1-DEMO-SCRIPT.md** in full. All four §3.3.1 exit criteria must pass:
(a) crash-safe resume, (b) duplicate-answer no-op, (c) post-deadline rejection +
park, (d) two concurrent tasks resume independently. With D-1 fixed you may
exercise the **live Slack answer path** for (b)/(d); the operator-CLI `answer`
path remains available and exercises the identical compare-and-set. **Do not
accept P1 until all four pass.**
**ROLLBACK:** the demo writes only ledger rows under `state/`; reset via the
Step-5 ledger rollback, or restore the Step-1 snapshot, then re-run.
---
## STEP 9 — Plane-1 checker live dry-runs 🧑 OPERATOR-REQUIRED
The six checkers live under `security-review/checkers/`:
`aws-posture.sh`, `compliance-drift.sh`, `confluence-doc.sh`,
`dependency-cve.sh`, `doc-drift.sh`, `plan-groomer.sh`. All run **report +
ALARM-only** (design D3): a clean run posts nothing and lands a mode-600 report;
no auto-Jira/Notion writes.
- Dry-run each checker against `~/repo-mirrors` in report-only mode; confirm a
clean run posts nothing and writes a mode-600 report. Confirm the exact
invocation against the checker scripts and the shared
`security-review/lib/` substrate on the box.
- Confirm the **canary suite** runs first (a planted-fault miss is a COMPLACENCY
ALARM and that role is skipped — design §6.4), and the **coverage rotation**
pointer advances (a slipped role is a COVERAGE ALARM, deferred-not-dropped).
**ROLLBACK:** checkers are read-only over the mirror corpus; a dry-run produces
only a report file. Remove the report dir to revert; no host state change.
---
## STEP 10 — P5 cross-plane loop dry-run 🧑 OPERATOR-REQUIRED
The Plane-1→Plane-2 loop turns confirmed checker findings into pipeline tasks:
```bash
cd ~/orchestrator/agent-team && . .venv/bin/activate
# Read one or more checker report JSONs and start one task per confirmed
# at/above-threshold finding (default threshold: high). --dry-run posts nowhere.
python3 run-team.py intake-checker --report <path-to-report.json> --threshold high --dry-run
```
The Tier-3 dep-bump **fixer** is dry-run only on the box (it holds no write
token, D2):
```bash
python3 run-team.py fix --report security-review/<...>/dependency-cve.json \
--finding-id <ID> --task-id <task> --dry-run
# prints the fix spec + minimal bump patch + the CI workflow_dispatch inputs;
# dispatches NOTHING. Live dispatch is the P3-live flip below.
```
**ROLLBACK:** both are read-only / dry-run (no dispatch, no apply); `intake-checker`
de-dup is in-memory per process. No host state to revert beyond ledger rows from
a non-dry-run intake (Step-5 ledger rollback).
---
## STEP 11 — Wire the schedule 🧑 OPERATOR-REQUIRED
The coordinator daemon (Step 6) is **always-on**, not timer-driven. The
**per-checker timers** (and the design's "shared timer with secrev", §8) attach
here. The existing `sea-haven-secrev.timer` (OnCalendar `02:00`, Persistent) is
untouched. Add a checker timer only after that checker is dry-run-validated
(Step 9).
**ROLLBACK:** each timer gets its own `systemctl disable --now <unit>.timer` + `rm`.
---
## The P3-live flip (deferred; IAM + GitHub App gated) 🧑 OPERATOR-REQUIRED
P3 (the build→verify apply-and-open-draft-PR loop) is **opt-in and inert** in
this deploy: `run-team.py` / `serve` pass `build_verify_wiring=None`, so no P3
subgraph is assembled. Flipping it live is a **separate, gated** provisioning
session, not part of the coordinator deploy:
1. **Mandatory reviews first.** The CI trust-boundary + OIDC IAM change is a
breaking IAM change → **GPT-4.1 cross-review** (global instructions) AND
`/sh-security-review` on the apply/verify surface. Do not flip without both.
2. **Provision the GitHub App** for the trusted apply path (the App that opens
the draft PR), and the **`agent-apply` GitHub Actions environment** that holds
the apply path's scoped permissions.
3. **Bind the live build→verify wiring** via
`agent_team.coordinator.gated_build_verify_wiring(...)` (the read-only CI
result fetcher + the real diff builder) — the seam a leaf calls *after* the
gate clears. The CI fetcher is read-only and fails closed (missing token /
404 / auth failure → `None` → the gate BLOCKs and the task parks).
4. **Set the apply env vars** the live path reads (the read-only CI-result token
and the dispatch target), then re-run the fixer **without** `--dry-run` only
once the dispatcher is bound.
Until every step above is done, the box dispatches/applies nothing.
---
## Post-session definition-of-done (design §7, global instructions)
- [ ] All four P1 criteria demonstrated live (Step 8).
- [ ] `project_r720_agent_team` memory created/updated.
- [ ] Confluence "AWS Architecture Map" / IT host inventory updated to show
`sh-secrev` now also hosts the always-on agent-team coordinator daemon.
- [ ] `/sh-security-review` run on the Slack inbound listener surface
(auth + untrusted-input; mandatory) — flag outstanding if not run.
- [ ] OPERATOR-RUNBOOK.md reviewed by whoever holds the pager.
- [ ] Snapshot retained until the daemon runs clean for one full cycle, then pruned.