feat(agent-team): deploy-readiness — serve starts Slack listener + systemd + provisioning docs (#23)
* fix(agent-team): serve() starts the inbound Slack listener (D-1) Coordinator.serve() now constructs and starts the SlackListener concurrently with the tick/drain loop on a background daemon thread, but ONLY when the live transport is a SlackTransport AND SLACK_APP_TOKEN is configured. When Slack is not the transport or the app token is absent, serve() behaves exactly as before (tick/recover only) — Slack is never made mandatory. - New injectable build_listener seam + default_slack_listener_factory sharing the coordinator's own transport, ledger db_path, and resume_queue put. - AUTHZ-01 owner-allowlist + open-status CAS untouched: serve() sources AGENT_TEAM_SLACK_OWNER_IDS in SlackListener.serve, which still fails closed. - SlackListener.close() added for clean Socket Mode teardown on shutdown; serve() stops the listener + joins the thread in a finally. - Tests: start-when-Slack+app-token, no-start otherwise, clean shutdown, idempotent start, serve start/stop around the loop, listener close(). * fix(agent-team): systemd unit loads ~/orchestrator/.env + uses venv python (D-2/D-7) D-2: add EnvironmentFile=-/home/adam/orchestrator/.env (optional '-') so the P2 GPT-4.1 review loop's cross_reviewer sub-process can read the non-Claude provider key once a task reaches REVIEW. Mirrors the sea-haven-secrev unit. D-7: point ExecStart at the agent-team venv interpreter (/home/adam/orchestrator/agent-team/.venv/bin/python) instead of /usr/bin/env python3, which resolved the system interpreter without the installed deps under systemd's PATH. All hardening (NoNewPrivileges / ProtectSystem=full / ProtectHome=read-only / ReadWritePaths) is retained unchanged (locked decision). * docs(agent-team): land provisioning + operator runbooks under docs/provisioning - PROVISIONING-RUNBOOK.md: merged final state (6 checkers, dep-bump fixer, P5 intake-checker loop), SLACK_CHANNEL_ID, the gated P3-live flip steps (GitHub App + agent-apply env + gated_build_verify_wiring), and D-1/D-2/D-7 marked FIXED so the demo can use the live Slack answer path. - P1-DEMO-SCRIPT.md: live Slack answer path now available (D-1 fixed); both the Slack and operator-CLI answer paths documented for all four exit criteria. - DEPLOY-AUDIT.md: D-1/D-2/D-7 RESOLVED (this PR); D-4/D-5 dep pinning and the operator-CLI divergence kept as provisioning notes. - OPERATOR-RUNBOOK.md (new): incident handling for pipeline stalls, parked tasks, failed HITL resumes, budget exhaustion, transport outages, and COMPLACENCY/COVERAGE alarms — each grounded in real run-team.py verbs, plus the re-alarm-backoff -> Jira-after-N-nights escalation ladder (design §5/§6.6). * fix(agent-team): supervise the Slack listener thread — recurring ALARM + respawn sh-security-review (logic) MEDIUM: a crashed listener thread was logged once, then the daemon ran on 'deaf' — posting clarifier questions but receiving no answers, every gate silently parking, process never exiting so systemd Restart=on-failure never fired. serve() now calls _supervise_slack_listener() each pass: when the listener is enabled but its thread is dead, it emits a recurring ERROR ALARM and respawns via the idempotent starter (self-heal). No-op when alive or disabled. +3 tests. (authz detector: wiring clean — AUTHZ-01 fail-closed allowlist + open-status CAS intact, dead listener fails SAFE.)
This commit is contained in:
parent
95e481890a
commit
2d1dca0804
9 changed files with 1686 additions and 13 deletions
|
|
@ -50,7 +50,9 @@ committed leaves (responder, resume_worker, db.schema, graph, transport).
|
|||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
import os
|
||||
import queue
|
||||
import threading
|
||||
from datetime import timedelta
|
||||
from pathlib import Path
|
||||
from typing import TYPE_CHECKING, Any, Callable
|
||||
|
|
@ -60,6 +62,7 @@ from agent_team import responder as responder_mod
|
|||
from agent_team.db.schema import connect, init_db
|
||||
from agent_team.resume_worker import ResumeResult, ResumeWorker
|
||||
from agent_team.transport.base import Transport
|
||||
from agent_team.transport.slack_adapter import SlackTransport
|
||||
|
||||
if TYPE_CHECKING: # pragma: no cover - typing only
|
||||
from agent_team.task_model import PipelineState
|
||||
|
|
@ -68,6 +71,7 @@ __all__ = [
|
|||
"Coordinator",
|
||||
"build_verify_wiring",
|
||||
"default_clarify_node_factory",
|
||||
"default_slack_listener_factory",
|
||||
"gated_build_verify_wiring",
|
||||
]
|
||||
|
||||
|
|
@ -124,6 +128,48 @@ CheckpointerFactory = Callable[[Path], Any]
|
|||
# so the deadline policy stays I/O-free in tests (§6.6).
|
||||
AlarmHook = Callable[[str], None]
|
||||
|
||||
# A listener factory: builds the inbound Slack Socket Mode listener
|
||||
# (:class:`~agent_team.transport.slack_listener.SlackListener`) the serve loop
|
||||
# starts on a background thread when Slack is the live transport and the app
|
||||
# token is present. Injected so :meth:`Coordinator.serve` is testable with a fake
|
||||
# listener and no live socket. The default
|
||||
# (:func:`default_slack_listener_factory`) builds the real listener from the
|
||||
# coordinator's shared transport / db / resume-queue and the Slack env tokens.
|
||||
ListenerFactory = Callable[[], Any]
|
||||
|
||||
|
||||
def default_slack_listener_factory(
|
||||
*,
|
||||
transport: SlackTransport,
|
||||
db_path: Path,
|
||||
enqueue_resume: Callable[[Any], None],
|
||||
) -> Any:
|
||||
"""Build the live :class:`SlackListener` from the coordinator's seams (D-1).
|
||||
|
||||
The production listener shares the coordinator's *own* live ``SlackTransport``
|
||||
(so ``parse_answer`` matches the outbound ``post_question`` wiring), the same
|
||||
durable ledger ``db_path``, and the same ``resume_queue`` ``put`` callable
|
||||
(the documented slack_listener ⇄ drainer handoff seam — ``coordinator.py``
|
||||
module docstring §3.3). The app/bot tokens and the owner allowlist are sourced
|
||||
from the environment by :meth:`SlackListener.serve` itself
|
||||
(``SLACK_APP_TOKEN`` / ``SLACK_BOT_TOKEN`` / ``AGENT_TEAM_SLACK_OWNER_IDS``);
|
||||
the allowlist still FAILS CLOSED if ``AGENT_TEAM_SLACK_OWNER_IDS`` is unset
|
||||
(AUTHZ-01), so this factory deliberately does not weaken that — it injects no
|
||||
``owner_ids`` and lets ``serve`` read + enforce them.
|
||||
|
||||
Imported lazily for the same import-hygiene reason as the clarifier / planner
|
||||
factories (the listener pulls the transport + responder leaves).
|
||||
"""
|
||||
from agent_team.transport.slack_listener import SlackListener
|
||||
|
||||
return SlackListener(
|
||||
transport,
|
||||
db_path,
|
||||
enqueue_resume,
|
||||
app_token=os.environ.get("SLACK_APP_TOKEN") or None,
|
||||
bot_token=os.environ.get("SLACK_BOT_TOKEN") or None,
|
||||
)
|
||||
|
||||
|
||||
def default_clarify_node_factory() -> Callable[[PipelineState], PipelineState]:
|
||||
"""Build the live Claude-backed clarifier node (§3.3, §7.1 P1).
|
||||
|
|
@ -338,6 +384,7 @@ class Coordinator:
|
|||
resume_queue: "queue.Queue[Any] | None" = None,
|
||||
deadline_window: timedelta | None = None,
|
||||
alarm_hook: AlarmHook | None = None,
|
||||
build_listener: ListenerFactory | None = None,
|
||||
) -> None:
|
||||
self._db_path = Path(db_path)
|
||||
self._transport = transport
|
||||
|
|
@ -357,6 +404,11 @@ class Coordinator:
|
|||
self._resume_queue: "queue.Queue[Any]" = resume_queue or queue.Queue()
|
||||
self._deadline_window = deadline_window or graph_mod.DEFAULT_CLARIFY_DEADLINE
|
||||
self._alarm_hook = alarm_hook or self._default_alarm_hook
|
||||
# The inbound Slack listener factory (D-1). Left None, serve() uses the
|
||||
# default that builds the real SlackListener; tests inject a fake to prove
|
||||
# the start/no-start/shutdown wiring with no live socket. Only consulted
|
||||
# by serve() when Slack is the live transport AND the app token is set.
|
||||
self._build_listener = build_listener
|
||||
|
||||
# Built by setup().
|
||||
self._graph: Any = None
|
||||
|
|
@ -364,6 +416,10 @@ class Coordinator:
|
|||
# Retains a context-manager checkpointer (production SQLite saver) so its
|
||||
# __exit__ is not run early; held open for the daemon's lifetime.
|
||||
self._checkpointer_cm: Any = None
|
||||
# The running inbound listener + its daemon thread (None until serve()
|
||||
# starts one). Held so serve()'s shutdown can stop it cleanly.
|
||||
self._listener: Any = None
|
||||
self._listener_thread: threading.Thread | None = None
|
||||
|
||||
# ------------------------------------------------------------------ #
|
||||
# Accessors (the shared queue is the slack_listener handoff seam).
|
||||
|
|
@ -814,32 +870,173 @@ class Coordinator:
|
|||
# ------------------------------------------------------------------ #
|
||||
|
||||
def serve(self, *, poll_interval: timedelta | None = None) -> None:
|
||||
"""Run the live daemon: bind invoker, setup, recover, then tick forever.
|
||||
"""Run the live daemon: bind invoker, setup, recover, start listener, tick.
|
||||
|
||||
The production entry. Binds the real Claude invoker
|
||||
(:func:`agent_team.invoker.bind_subscription_invoker`) BEFORE
|
||||
:meth:`setup` builds the clarifier node (so the node's
|
||||
``billing.claude_invoke`` calls hit the live subscription path), runs the
|
||||
startup :meth:`recover` sweep, then loops calling :meth:`tick` on the
|
||||
startup :meth:`recover` sweep, **starts the inbound Slack listener**
|
||||
(:meth:`_maybe_start_slack_listener`) when Slack is the live transport
|
||||
and the app token is configured, then loops calling :meth:`tick` on the
|
||||
deadline cadence.
|
||||
|
||||
The actual Slack inbound feed is the slack_listener's job; the
|
||||
coordinator exposes :meth:`submit_answer` and the shared
|
||||
:attr:`resume_queue` for it. This loop owns only the deadline/recovery
|
||||
maintenance cadence.
|
||||
The Slack inbound feed (Socket Mode) is the slack_listener's job and runs
|
||||
concurrently on a background daemon thread; this loop owns the
|
||||
deadline/recovery maintenance cadence. On shutdown (Ctrl-C /
|
||||
SIGTERM-driven ``KeyboardInterrupt``) the listener is stopped cleanly via
|
||||
:meth:`_stop_slack_listener` in a ``finally``.
|
||||
|
||||
Slack is NOT mandatory: when the transport is not the live ``SlackTransport``
|
||||
or the app token is absent, no listener is started and ``serve`` behaves
|
||||
exactly as before (tick/recover only).
|
||||
"""
|
||||
from agent_team.invoker import bind_subscription_invoker
|
||||
|
||||
bind_subscription_invoker()
|
||||
self.setup()
|
||||
self.recover()
|
||||
self._maybe_start_slack_listener()
|
||||
|
||||
interval = (poll_interval or DEFAULT_POLL_INTERVAL).total_seconds()
|
||||
import time
|
||||
|
||||
while True: # pragma: no cover - the infinite daemon loop
|
||||
self.tick()
|
||||
time.sleep(interval)
|
||||
try:
|
||||
while True: # pragma: no cover - the infinite daemon loop
|
||||
self.tick()
|
||||
self._supervise_slack_listener()
|
||||
time.sleep(interval)
|
||||
finally:
|
||||
self._stop_slack_listener()
|
||||
|
||||
# ------------------------------------------------------------------ #
|
||||
# Inbound Slack listener wiring (D-1): start concurrently with the
|
||||
# tick/drain loop ONLY when Slack is the live transport and the app
|
||||
# token is configured; stop cleanly on shutdown.
|
||||
# ------------------------------------------------------------------ #
|
||||
|
||||
def _slack_listener_enabled(self) -> bool:
|
||||
"""Return ``True`` iff the inbound Slack listener should run (D-1 gate).
|
||||
|
||||
Two conditions, both required (and Slack must stay OPTIONAL):
|
||||
|
||||
* the live transport is a :class:`SlackTransport` (the inbound
|
||||
``parse_answer`` must match the outbound poster); and
|
||||
* ``SLACK_APP_TOKEN`` is present in the environment (Socket Mode needs the
|
||||
app-level token to open the outbound WebSocket).
|
||||
|
||||
If either is false the daemon runs WITHOUT an inbound listener exactly as
|
||||
before — Slack is never made mandatory. The owner allowlist
|
||||
(``AGENT_TEAM_SLACK_OWNER_IDS``) is intentionally NOT part of this gate:
|
||||
the listener fails closed on an empty allowlist (AUTHZ-01), so starting it
|
||||
unprovisioned safely rejects every answer rather than silently never
|
||||
hearing Slack.
|
||||
"""
|
||||
if not isinstance(self._transport, SlackTransport):
|
||||
return False
|
||||
return bool(os.environ.get("SLACK_APP_TOKEN"))
|
||||
|
||||
def _maybe_start_slack_listener(self) -> bool:
|
||||
"""Start the inbound Slack listener on a daemon thread if enabled (D-1).
|
||||
|
||||
Builds the listener via the injected ``build_listener`` factory (or the
|
||||
:func:`default_slack_listener_factory`, sharing the coordinator's own live
|
||||
transport, durable ``db_path`` and ``resume_queue`` put), then runs its
|
||||
blocking :meth:`~agent_team.transport.slack_listener.SlackListener.serve`
|
||||
on a background **daemon** thread so the Socket Mode socket and the
|
||||
tick/drain loop run concurrently. Returns ``True`` iff a listener was
|
||||
started; ``False`` (and no thread) when :meth:`_slack_listener_enabled`
|
||||
is false — the Slack-optional contract.
|
||||
|
||||
Idempotent: a second call while a listener thread is already alive is a
|
||||
no-op.
|
||||
"""
|
||||
if self._listener_thread is not None and self._listener_thread.is_alive():
|
||||
return True
|
||||
if not self._slack_listener_enabled():
|
||||
_LOG.info(
|
||||
"inbound Slack listener not started (transport is not live Slack "
|
||||
"or SLACK_APP_TOKEN is unset); daemon runs tick/recover only"
|
||||
)
|
||||
return False
|
||||
|
||||
if self._build_listener is not None:
|
||||
listener = self._build_listener()
|
||||
else:
|
||||
listener = default_slack_listener_factory(
|
||||
transport=self._transport,
|
||||
db_path=self._db_path,
|
||||
enqueue_resume=self._resume_queue.put,
|
||||
)
|
||||
self._listener = listener
|
||||
thread = threading.Thread(
|
||||
target=self._run_listener,
|
||||
name="agent-team-slack-listener",
|
||||
daemon=True,
|
||||
)
|
||||
self._listener_thread = thread
|
||||
thread.start()
|
||||
_LOG.info("inbound Slack listener started (Socket Mode, background thread)")
|
||||
return True
|
||||
|
||||
def _supervise_slack_listener(self) -> None:
|
||||
"""Detect a dead inbound listener thread, ALARM (recurring), and respawn.
|
||||
|
||||
Called every :meth:`serve` pass. Without this, a crashed listener thread
|
||||
(:meth:`_run_listener` swallows its exception) leaves the daemon "deaf":
|
||||
it keeps posting clarifier questions but receives no answers, every gate
|
||||
silently times out to PARK, and the process stays up so systemd
|
||||
``Restart=on-failure`` never fires. This makes that failure LOUD on every
|
||||
pass (not a one-shot crash log) and self-heals via the idempotent
|
||||
:meth:`_maybe_start_slack_listener`. No-op when the listener is not
|
||||
enabled or the thread is alive.
|
||||
"""
|
||||
if not self._slack_listener_enabled():
|
||||
return
|
||||
thread = self._listener_thread
|
||||
if thread is not None and thread.is_alive():
|
||||
return
|
||||
_LOG.error(
|
||||
"ALARM: inbound Slack listener is DOWN — answers are NOT being "
|
||||
"received; clarifier gates will park. Respawning the listener."
|
||||
)
|
||||
# Drop the dead handle so the idempotent starter actually respawns.
|
||||
self._listener_thread = None
|
||||
self._maybe_start_slack_listener()
|
||||
|
||||
def _run_listener(self) -> None:
|
||||
"""Thread target: run the listener's blocking serve, log on exit.
|
||||
|
||||
Swallows any listener exception into a log line so a listener crash takes
|
||||
down only the inbound socket (which Restart=on-failure / a manual restart
|
||||
recovers), never the maintenance loop or the whole process from a
|
||||
background thread.
|
||||
"""
|
||||
try:
|
||||
self._listener.serve()
|
||||
except Exception: # noqa: BLE001 - isolate the listener thread
|
||||
_LOG.exception("inbound Slack listener thread exited with an error")
|
||||
|
||||
def _stop_slack_listener(self) -> None:
|
||||
"""Stop the inbound listener cleanly on daemon shutdown (D-1).
|
||||
|
||||
Best-effort: closes the Socket Mode connection via the listener's
|
||||
``close`` (if it exposes one), then joins the daemon thread briefly. A
|
||||
teardown error never propagates — shutdown must complete regardless.
|
||||
"""
|
||||
listener = self._listener
|
||||
thread = self._listener_thread
|
||||
self._listener = None
|
||||
self._listener_thread = None
|
||||
if listener is not None:
|
||||
close = getattr(listener, "close", None)
|
||||
if close is not None:
|
||||
try:
|
||||
close()
|
||||
except Exception: # noqa: BLE001 - shutdown must not raise
|
||||
_LOG.debug("listener close raised during shutdown; ignoring")
|
||||
if thread is not None and thread.is_alive():
|
||||
thread.join(timeout=5.0)
|
||||
|
||||
# ------------------------------------------------------------------ #
|
||||
# Internals.
|
||||
|
|
|
|||
|
|
@ -147,6 +147,10 @@ class SlackListener:
|
|||
# The owner allowlist (AUTHZ-01). An empty set is the fail-closed default:
|
||||
# an unconfigured deploy rejects every answer.
|
||||
self._owner_ids: set[str] = set(owner_ids) if owner_ids else set()
|
||||
# The live Socket Mode handler, retained by :meth:`serve` so :meth:`close`
|
||||
# can stop it cleanly on daemon shutdown. ``None`` until ``serve`` opens
|
||||
# the socket.
|
||||
self._socket_handler: Any = None
|
||||
|
||||
def handle_event(self, raw_payload: Any) -> AnswerOutcome | None:
|
||||
"""Normalize + submit one inbound event; return its outcome or ``None``.
|
||||
|
|
@ -331,7 +335,35 @@ class SlackListener:
|
|||
def _on_mention(body: Mapping[str, Any]) -> None:
|
||||
_forward(body)
|
||||
|
||||
SocketModeHandler(app, self._app_token).start()
|
||||
handler = SocketModeHandler(app, self._app_token)
|
||||
self._socket_handler = handler
|
||||
handler.start()
|
||||
|
||||
def close(self) -> None:
|
||||
"""Best-effort clean stop of the Socket Mode connection (daemon shutdown).
|
||||
|
||||
The coordinator's :meth:`~agent_team.coordinator.Coordinator.serve` calls
|
||||
this when it shuts the daemon down so the outbound WebSocket is closed
|
||||
cleanly rather than only dying with the process. Tolerant: if the handler
|
||||
was never opened, or the SDK's ``close``/``disconnect`` raises, the error
|
||||
is swallowed — shutdown must never be blocked by a transport teardown
|
||||
failure. The handler is dropped afterward so a second ``close`` is a
|
||||
no-op.
|
||||
"""
|
||||
handler = self._socket_handler
|
||||
if handler is None:
|
||||
return
|
||||
self._socket_handler = None
|
||||
# slack_bolt's SocketModeHandler exposes ``close`` (and the underlying
|
||||
# client a ``disconnect``); try the most specific available, swallow any
|
||||
# teardown error.
|
||||
closer = getattr(handler, "close", None) or getattr(handler, "disconnect", None)
|
||||
if closer is None:
|
||||
return
|
||||
try:
|
||||
closer()
|
||||
except Exception: # noqa: BLE001 - shutdown must not raise
|
||||
_LOG.debug("SlackListener.close: handler teardown raised; ignoring")
|
||||
|
||||
|
||||
def _extract_sender_id(raw_payload: Mapping[str, Any]) -> str | None:
|
||||
|
|
|
|||
|
|
@ -13,8 +13,10 @@
|
|||
# systemctl status agent-team-coordinator.service
|
||||
# journalctl -u agent-team-coordinator.service -e -f
|
||||
#
|
||||
# Secrets come from the EnvironmentFile (leading '-' = optional, no failure if
|
||||
# absent), ~/secrev.env (mode 600, NOT in git):
|
||||
# Secrets come from the EnvironmentFile(s) (leading '-' = optional, no failure if
|
||||
# absent). Two files are loaded, mirroring the sea-haven-secrev unit:
|
||||
#
|
||||
# ~/secrev.env (mode 600, NOT in git) - the agent-team runtime keys:
|
||||
# CLAUDE_CODE_OAUTH_TOKEN -> subscription OAuth (from `claude setup-token`).
|
||||
# A raw ANTHROPIC_API_KEY must NOT be set on this
|
||||
# box; it would silently win and meter to API
|
||||
|
|
@ -24,7 +26,18 @@
|
|||
# REQUIRED for Socket Mode; opens the inbound
|
||||
# WebSocket that receives answers. Without it the
|
||||
# coordinator can post but never hear replies.
|
||||
# serve() starts the inbound SlackListener only when
|
||||
# the transport is live Slack AND this token is set.
|
||||
# SLACK_CHANNEL_ID -> target channel for clarifier questions.
|
||||
# AGENT_TEAM_SLACK_OWNER_IDS -> comma-separated authorized answerer ids
|
||||
# (AUTHZ-01). The listener FAILS CLOSED if unset.
|
||||
#
|
||||
# ~/orchestrator/.env (mode 600, NOT in git) - the P2 review-loop provider key:
|
||||
# The production serve() wires the GPT-4.1 cross-review loop, which shells the
|
||||
# local orchestrator run.py -> cross_reviewer once a task reaches REVIEW. That
|
||||
# sub-process needs the non-Claude provider key from ~/orchestrator/.env (same
|
||||
# file the sea-haven-secrev unit loads). Optional ('-') so the daemon still
|
||||
# starts if it is absent; the review path then fails loudly only at REVIEW.
|
||||
|
||||
[Unit]
|
||||
Description=Sea Haven agent-team Plane-2 coordinator daemon
|
||||
|
|
@ -36,7 +49,11 @@ Type=simple
|
|||
User=adam
|
||||
WorkingDirectory=/home/adam/orchestrator/agent-team
|
||||
EnvironmentFile=-/home/adam/secrev.env
|
||||
ExecStart=/usr/bin/env python3 run-team.py serve
|
||||
EnvironmentFile=-/home/adam/orchestrator/.env
|
||||
# Use the agent-team venv interpreter (where the runtime deps are installed by
|
||||
# the DEPLOY-R720 / PROVISIONING-RUNBOOK step), NOT the bare system python3 that
|
||||
# `/usr/bin/env python3` would resolve under systemd's PATH (D-7).
|
||||
ExecStart=/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve
|
||||
Restart=on-failure
|
||||
RestartSec=5
|
||||
# Hardening - matches the level the sea-haven-secrev unit relies on, scoped for a
|
||||
|
|
|
|||
|
|
@ -602,3 +602,235 @@ def test_setup_with_p2_factories_builds_a_review_node(db_path: Path) -> None:
|
|||
assert graph_mod.REVIEW in coord.graph.get_graph().nodes
|
||||
finally:
|
||||
review_loop._review_invoker = saved
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------- #
|
||||
# Inbound Slack listener wiring in serve() (D-1)
|
||||
# --------------------------------------------------------------------------- #
|
||||
|
||||
|
||||
class _FakeListener:
|
||||
"""A record-only stand-in for SlackListener (no socket, no SDK).
|
||||
|
||||
``serve`` blocks on an Event until ``close`` is called, mirroring the live
|
||||
listener whose ``serve`` blocks on the Socket Mode handler until shut down.
|
||||
Records that ``serve`` ran and that ``close`` was called so the coordinator
|
||||
start/shutdown wiring can be asserted.
|
||||
"""
|
||||
|
||||
def __init__(self) -> None:
|
||||
import threading as _t
|
||||
|
||||
self.served = False
|
||||
self.closed = False
|
||||
self._stop = _t.Event()
|
||||
|
||||
def serve(self) -> None:
|
||||
self.served = True
|
||||
self._stop.wait(timeout=5.0)
|
||||
|
||||
def close(self) -> None:
|
||||
self.closed = True
|
||||
self._stop.set()
|
||||
|
||||
|
||||
def _slack_coordinator(
|
||||
db_path: Path,
|
||||
*,
|
||||
build_listener: Any = None,
|
||||
) -> Coordinator:
|
||||
"""A Coordinator whose live transport IS a SlackTransport (listener gate)."""
|
||||
from agent_team.transport.slack_adapter import SlackTransport
|
||||
|
||||
saver = _Saver()
|
||||
return Coordinator(
|
||||
db_path=db_path,
|
||||
transport=SlackTransport("C123"),
|
||||
build_clarify_node=lambda: graph_mod.clarify_node,
|
||||
build_checkpointer=lambda _path: saver,
|
||||
build_listener=build_listener,
|
||||
)
|
||||
|
||||
|
||||
def test_serve_starts_listener_when_slack_and_app_token(
|
||||
db_path: Path, monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""serve() starts the inbound listener when Slack is the transport AND the
|
||||
app token is configured, then stops it cleanly on shutdown."""
|
||||
monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test")
|
||||
listener = _FakeListener()
|
||||
coord = _slack_coordinator(db_path, build_listener=lambda: listener)
|
||||
coord.setup()
|
||||
|
||||
assert coord._maybe_start_slack_listener() is True
|
||||
# The background daemon thread actually ran the listener's serve().
|
||||
import time
|
||||
|
||||
for _ in range(50):
|
||||
if listener.served:
|
||||
break
|
||||
time.sleep(0.01)
|
||||
assert listener.served is True
|
||||
|
||||
# Clean shutdown stops the listener and joins the thread.
|
||||
coord._stop_slack_listener()
|
||||
assert listener.closed is True
|
||||
assert coord._listener is None
|
||||
assert coord._listener_thread is None
|
||||
|
||||
|
||||
def test_serve_does_not_start_listener_without_app_token(
|
||||
db_path: Path, monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""No app token => no inbound listener (Slack stays optional)."""
|
||||
monkeypatch.delenv("SLACK_APP_TOKEN", raising=False)
|
||||
listener = _FakeListener()
|
||||
coord = _slack_coordinator(db_path, build_listener=lambda: listener)
|
||||
coord.setup()
|
||||
|
||||
assert coord._slack_listener_enabled() is False
|
||||
assert coord._maybe_start_slack_listener() is False
|
||||
assert listener.served is False
|
||||
assert coord._listener is None
|
||||
assert coord._listener_thread is None
|
||||
|
||||
|
||||
def test_serve_does_not_start_listener_when_transport_not_slack(
|
||||
db_path: Path, monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""A non-Slack transport never starts the listener even with the app token."""
|
||||
monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test")
|
||||
listener = _FakeListener()
|
||||
saver = _Saver()
|
||||
coord = Coordinator(
|
||||
db_path=db_path,
|
||||
transport=FakeTransport(), # not a SlackTransport
|
||||
build_clarify_node=lambda: graph_mod.clarify_node,
|
||||
build_checkpointer=lambda _path: saver,
|
||||
build_listener=lambda: listener,
|
||||
)
|
||||
coord.setup()
|
||||
|
||||
assert coord._slack_listener_enabled() is False
|
||||
assert coord._maybe_start_slack_listener() is False
|
||||
assert listener.served is False
|
||||
|
||||
|
||||
def test_maybe_start_listener_is_idempotent(
|
||||
db_path: Path, monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""A second start while the thread is alive is a no-op (one listener)."""
|
||||
monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test")
|
||||
built: list[_FakeListener] = []
|
||||
|
||||
def _factory() -> _FakeListener:
|
||||
lst = _FakeListener()
|
||||
built.append(lst)
|
||||
return lst
|
||||
|
||||
coord = _slack_coordinator(db_path, build_listener=_factory)
|
||||
coord.setup()
|
||||
try:
|
||||
assert coord._maybe_start_slack_listener() is True
|
||||
assert coord._maybe_start_slack_listener() is True
|
||||
assert len(built) == 1 # not rebuilt/restarted
|
||||
finally:
|
||||
coord._stop_slack_listener()
|
||||
|
||||
|
||||
def test_supervise_respawns_dead_listener_and_alarms(
|
||||
db_path: Path, monkeypatch: pytest.MonkeyPatch, caplog: pytest.LogCaptureFixture
|
||||
) -> None:
|
||||
"""A dead listener thread is respawned (self-heal) with a recurring ALARM, so
|
||||
the daemon never silently goes 'deaf' to Slack answers (D1-silent-stall)."""
|
||||
import logging
|
||||
|
||||
monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test")
|
||||
listeners: list[_FakeListener] = []
|
||||
|
||||
def _factory() -> _FakeListener:
|
||||
lst = _FakeListener()
|
||||
listeners.append(lst)
|
||||
return lst
|
||||
|
||||
coord = _slack_coordinator(db_path, build_listener=_factory)
|
||||
coord.setup()
|
||||
try:
|
||||
assert coord._maybe_start_slack_listener() is True
|
||||
first = coord._listener_thread
|
||||
assert first is not None and first.is_alive()
|
||||
# Kill the listener thread: close() fires the stop event so serve() returns.
|
||||
listeners[0].close()
|
||||
first.join(timeout=5.0)
|
||||
assert not first.is_alive()
|
||||
# The supervisor detects the death, ALARMs, and respawns a fresh listener.
|
||||
with caplog.at_level(logging.ERROR):
|
||||
coord._supervise_slack_listener()
|
||||
assert any("listener is DOWN" in r.message for r in caplog.records)
|
||||
second = coord._listener_thread
|
||||
assert second is not None and second is not first and second.is_alive()
|
||||
assert len(listeners) == 2 # a new listener was built on respawn
|
||||
finally:
|
||||
coord._stop_slack_listener()
|
||||
|
||||
|
||||
def test_supervise_is_noop_when_listener_alive(
|
||||
db_path: Path, monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""The supervisor leaves a healthy listener thread untouched (no churn)."""
|
||||
monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test")
|
||||
listener = _FakeListener()
|
||||
coord = _slack_coordinator(db_path, build_listener=lambda: listener)
|
||||
coord.setup()
|
||||
try:
|
||||
coord._maybe_start_slack_listener()
|
||||
alive = coord._listener_thread
|
||||
coord._supervise_slack_listener()
|
||||
assert coord._listener_thread is alive
|
||||
finally:
|
||||
coord._stop_slack_listener()
|
||||
|
||||
|
||||
def test_supervise_is_noop_when_listener_disabled(
|
||||
db_path: Path, monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""No app token -> listener disabled -> supervisor never spawns a thread."""
|
||||
monkeypatch.delenv("SLACK_APP_TOKEN", raising=False)
|
||||
coord = _slack_coordinator(db_path, build_listener=_FakeListener)
|
||||
coord.setup()
|
||||
coord._supervise_slack_listener()
|
||||
assert coord._listener_thread is None
|
||||
|
||||
|
||||
def test_stop_listener_is_safe_when_none(db_path: Path) -> None:
|
||||
"""Stopping with no listener running never raises (clean-shutdown safety)."""
|
||||
coord = _slack_coordinator(db_path)
|
||||
coord._stop_slack_listener() # must be a quiet no-op
|
||||
assert coord._listener is None
|
||||
|
||||
|
||||
def test_serve_wires_start_and_stop_around_the_loop(
|
||||
db_path: Path, monkeypatch: pytest.MonkeyPatch
|
||||
) -> None:
|
||||
"""serve() starts the listener before the loop and stops it in finally even
|
||||
when the loop is interrupted (KeyboardInterrupt)."""
|
||||
monkeypatch.setenv("SLACK_APP_TOKEN", "xapp-test")
|
||||
listener = _FakeListener()
|
||||
coord = _slack_coordinator(db_path, build_listener=lambda: listener)
|
||||
|
||||
# Neuter the live invoker bind so serve() needs no Claude SDK / token.
|
||||
# serve() does `from agent_team.invoker import bind_subscription_invoker`,
|
||||
# so patch the name on the invoker module it imports from.
|
||||
import agent_team.invoker as invoker_mod
|
||||
|
||||
monkeypatch.setattr(invoker_mod, "bind_subscription_invoker", lambda: None)
|
||||
# Break the infinite loop on the first tick.
|
||||
monkeypatch.setattr(
|
||||
coord, "tick", lambda: (_ for _ in ()).throw(KeyboardInterrupt())
|
||||
)
|
||||
|
||||
with pytest.raises(KeyboardInterrupt):
|
||||
coord.serve(poll_interval=timedelta(seconds=0))
|
||||
|
||||
assert listener.served is True # started before the loop
|
||||
assert listener.closed is True # stopped in finally on interrupt
|
||||
|
|
|
|||
|
|
@ -449,6 +449,36 @@ def test_serve_requires_tokens(db_path: Path) -> None:
|
|||
listener.serve()
|
||||
|
||||
|
||||
def test_close_without_open_handler_is_noop(db_path: Path) -> None:
|
||||
"""close() before serve() ever opened a socket is a quiet no-op."""
|
||||
listener = _listener(db_path, RecordingQueue())
|
||||
listener.close() # must not raise
|
||||
listener.close() # idempotent
|
||||
|
||||
|
||||
def test_close_stops_socket_handler(db_path: Path) -> None:
|
||||
"""close() invokes the retained Socket Mode handler's close and drops it."""
|
||||
|
||||
class _FakeHandler:
|
||||
def __init__(self) -> None:
|
||||
self.closed = False
|
||||
|
||||
def close(self) -> None:
|
||||
self.closed = True
|
||||
|
||||
listener = _listener(db_path, RecordingQueue())
|
||||
handler = _FakeHandler()
|
||||
listener._socket_handler = handler
|
||||
listener.close()
|
||||
assert handler.closed is True
|
||||
assert listener._socket_handler is None
|
||||
# A teardown that raises is swallowed (shutdown must never raise).
|
||||
listener._socket_handler = type(
|
||||
"_Boom", (), {"close": lambda self: (_ for _ in ()).throw(RuntimeError("x"))}
|
||||
)()
|
||||
listener.close() # no exception
|
||||
|
||||
|
||||
def test_real_slack_transport_is_a_transport() -> None:
|
||||
"""Sanity: the injected SlackTransport is the contract the listener expects."""
|
||||
assert isinstance(SlackTransport(channel="C123"), Transport)
|
||||
|
|
|
|||
187
docs/provisioning/DEPLOY-AUDIT.md
Normal file
187
docs/provisioning/DEPLOY-AUDIT.md
Normal file
|
|
@ -0,0 +1,187 @@
|
|||
# DEPLOY-AUDIT — `agent-team/DEPLOY-R720.md` + systemd unit vs. actual code
|
||||
|
||||
Cross-check of the drafted deploy doc (`agent-team/DEPLOY-R720.md`) and the
|
||||
systemd unit (`agent-team/systemd/agent-team-coordinator.service`) against the
|
||||
**actual current code** in `agent-team/agent_team/**` and `agent-team/run-team.py`.
|
||||
|
||||
Each finding: **location → claimed → actual → fix**. Severity: 🔴 blocker /
|
||||
🟠 should-fix / 🟡 nit. Items confirmed clean are stated explicitly.
|
||||
|
||||
> **Status (deploy-readiness PR):** the three deploy-correctness bugs that
|
||||
> blocked the live provisioning session — **D-1, D-2, D-7** — are **RESOLVED** in
|
||||
> this PR (`feature/agent-team-deploy-readiness`). The remaining items
|
||||
> (**D-4 / D-5** dependency pinning, the **operator-CLI divergence**) are kept
|
||||
> below as **provisioning notes** — they do not block the coordinator deploy.
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-1 — `serve` now starts the Slack inbound listener — RESOLVED (this PR)
|
||||
|
||||
- **Location:** `agent_team/coordinator.py` (`Coordinator.serve` +
|
||||
`_maybe_start_slack_listener` / `_slack_listener_enabled` / `_stop_slack_listener`
|
||||
/ `default_slack_listener_factory`); `agent_team/transport/slack_listener.py`
|
||||
(`SlackListener.serve` + new `close`).
|
||||
- **Was:** `Coordinator.serve()` did only `bind_subscription_invoker()`,
|
||||
`setup()`, `recover()`, then an infinite tick/sleep loop — it never constructed
|
||||
or started `SlackListener`, so a deployed daemon posted clarifier questions and
|
||||
expired them on deadline but could **not hear Slack answers**.
|
||||
- **Now:** `serve()` starts the `SlackListener` on a background **daemon thread**,
|
||||
concurrently with the tick/drain loop, **when** the live transport is a
|
||||
`SlackTransport` AND `SLACK_APP_TOKEN` is set. It shares the coordinator's own
|
||||
transport, ledger `db_path`, and `resume_queue` put; on shutdown it calls the
|
||||
listener's new `close()` and joins the thread in a `finally`. When Slack is not
|
||||
the transport or the app token is absent, no listener starts and `serve` behaves
|
||||
exactly as before — **Slack is never made mandatory**. The AUTHZ-01 owner
|
||||
allowlist + the open-status compare-and-set are untouched (still fail closed on
|
||||
an empty `AGENT_TEAM_SLACK_OWNER_IDS`).
|
||||
- **Tests:** `tests/test_coordinator.py` — start-when-Slack+app-token, no-start
|
||||
without the token, no-start when the transport is not Slack, idempotent start,
|
||||
clean shutdown, serve start/stop around the loop; `tests/test_slack_listener.py`
|
||||
— `close()` no-op + handler teardown.
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-2 — coordinator unit loads `~/orchestrator/.env` — RESOLVED (this PR)
|
||||
|
||||
- **Location:** `agent-team/systemd/agent-team-coordinator.service`;
|
||||
`run-team.py:_build_coordinator` (always wires `default_review_wiring`);
|
||||
`coordinator.py:default_review_wiring` → `review_loop_llm.default_plan_reviewer`
|
||||
(shells the local orchestrator `run.py` → `cross_reviewer` GPT-4.1).
|
||||
- **Was:** the unit loaded only `~/secrev.env`. `run-team.py serve` builds the
|
||||
coordinator with the P2 review loop wired, and once a task reaches REVIEW the
|
||||
reviewer shells the orchestrator `run.py`, whose GPT-4.1 call reads the
|
||||
non-Claude provider key from `~/orchestrator/.env`. The review path would fail
|
||||
to authenticate.
|
||||
- **Now:** the unit adds `EnvironmentFile=-/home/adam/orchestrator/.env`
|
||||
(optional `-`, mirroring the `sea-haven-secrev` unit). Does not affect the P1
|
||||
demo (P1 stops at PLAN before REVIEW); closes the latent P2 break.
|
||||
|
||||
---
|
||||
|
||||
## 🟠 D-4 — pip install list omits `requests` (provisioning note)
|
||||
|
||||
- **Location:** `agent_team/transport/github_live.py` / `github_intake.py`
|
||||
(`import requests`).
|
||||
- **Actual:** the GitHub transport + intake require `requests`; the Slack-first
|
||||
path does not hit it, but any `--transport github` / `intake-github` use fails
|
||||
with a clear RuntimeError without it.
|
||||
- **Fix:** PROVISIONING-RUNBOOK Step 4 installs `requests` into the venv.
|
||||
`requests` is also absent from `requirements.txt` (see D-5).
|
||||
|
||||
---
|
||||
|
||||
## 🟠 D-5 — agent-team runtime deps are not pinned in `requirements.txt` (provisioning note)
|
||||
|
||||
- **Location:** `requirements.txt` (repo root).
|
||||
- **Actual:** `requirements.txt` pins `langgraph==1.1.10` and
|
||||
`langgraph-checkpoint-sqlite==3.1.0`, but the agent-team runtime deps
|
||||
`claude-agent-sdk`, `slack_sdk`, `slack_bolt`, `requests` (and `anthropic` for
|
||||
api mode) are **not in `requirements.txt` at all** — they are installed ad-hoc
|
||||
into the agent-team venv by the runbook. There is no pinned, reproducible source
|
||||
of truth for the box's runtime set.
|
||||
- **Fix (deferred):** add an `agent-team/requirements.txt` (or extras group)
|
||||
pinning these, version-matched to the root `requirements.txt` langgraph pin.
|
||||
Until then, PROVISIONING-RUNBOOK Step 4 pins `langgraph==1.1.10` /
|
||||
`langgraph-checkpoint-sqlite==3.1.0` explicitly so the unpinned `pip install`
|
||||
cannot pull a newer, untested major. **Do NOT modify `requirements.txt` or the
|
||||
checkers in this PR** (out of scope).
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-6 — `slack_bolt` is now exercised by the daemon — RESOLVED (consequence of D-1)
|
||||
|
||||
- **Location:** `slack_listener.py:serve` (the only `slack_bolt` import).
|
||||
- **Now:** with D-1 fixed, `Coordinator.serve()` starts `SlackListener.serve()`,
|
||||
which imports + uses `slack_bolt` for the Socket Mode handler. The dep is right
|
||||
and now actually exercised on the live Slack path.
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-SLACKVAR (clean) — `SLACK_CHANNEL_ID` matches
|
||||
|
||||
- `run-team.py:_build_transport` reads exactly `os.environ.get("SLACK_CHANNEL_ID")`.
|
||||
The runbook, the unit comment, and the code all use `SLACK_CHANNEL_ID` (not
|
||||
`SLACK_CHANNEL`). **CLEAN.**
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-ENV-SLACKBOT / OAUTH / OWNERS / APPTOKEN (clean) — names match
|
||||
|
||||
- **`SLACK_BOT_TOKEN`** ↔ `slack_live.py` + `slack_listener` env read. **CLEAN.**
|
||||
- **`CLAUDE_CODE_OAUTH_TOKEN`** ↔ `invoker.py`. **CLEAN.**
|
||||
- **`AGENT_TEAM_SLACK_OWNER_IDS`** ↔ `slack_listener.py` (name + fail-closed
|
||||
semantics). **CLEAN** — now read by the running daemon (D-1 fixed).
|
||||
- **`SLACK_APP_TOKEN`** — now read in two places: `coordinator._slack_listener_enabled`
|
||||
gates the listener on its presence, and `default_slack_listener_factory` /
|
||||
`SlackListener.serve` source it to open the socket. **CLEAN** (read site exists
|
||||
now that D-1 is fixed).
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-7 — `ExecStart` uses the venv interpreter — RESOLVED (this PR)
|
||||
|
||||
- **Location:** `agent-team/systemd/agent-team-coordinator.service` ExecStart.
|
||||
- **Was:** `ExecStart=/usr/bin/env python3 run-team.py serve` resolved the
|
||||
**system** interpreter under systemd's PATH — not the venv where the deps were
|
||||
installed, so the daemon would fail at import.
|
||||
- **Now:** `ExecStart=/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve`
|
||||
(matches the runbook venv path + `WorkingDirectory`).
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-SUBCMD (mostly clean) — run-team.py subcommands referenced exist
|
||||
|
||||
Cross-checked every `run-team.py <sub>` the deploy doc + demo name against
|
||||
`run-team.py:build_parser`: `init-db`, `serve`, `list` (+ `--all` / `--parked`),
|
||||
`show`, `expire`, `answer`, `redeliver`, `supersede`, `force-resume`, `start`,
|
||||
`intake-github`, `intake-checker`, `fix` — all exist. No invented verbs.
|
||||
|
||||
> ### Operator-CLI divergence (provisioning note)
|
||||
>
|
||||
> Two operator CLIs exist with **different verb names**:
|
||||
>
|
||||
> - `run-team.py` (the entry CLI): `init-db, list, show, redeliver, expire,
|
||||
> answer, supersede, force-resume, start, serve, intake-github, intake-checker,
|
||||
> fix`. Has `show` and `--parked`; `--db` / `--audit-log` default sensibly.
|
||||
> - `agent_team/operator_cli.py`: `list, redeliver, force-expire,
|
||||
> answer-on-behalf, force-resume` — **no `show`**, `--db` / `--audit-log` are
|
||||
> **required**, and its `force-resume` **supersedes** (unlike `run-team.py`'s,
|
||||
> which reopens an expired row and never supersedes).
|
||||
>
|
||||
> **Use `run-team.py` for provisioning + the demo + incident recovery.** The
|
||||
> docs reference only `run-team.py`. Reconciling the two CLIs is a follow-up.
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-PYTHONPKG (clean) — package import bootstrap is correct
|
||||
|
||||
`run-team.py` inserts its own dir into `sys.path` so the hyphenated script
|
||||
imports the `agent_team` package without an editable install. **CLEAN.**
|
||||
|
||||
---
|
||||
|
||||
## ✅ D-HARDENING (clean, and matches the locked decision)
|
||||
|
||||
- Unit: `NoNewPrivileges=true`, `ProtectSystem=full`, `ProtectHome=read-only`,
|
||||
`ReadWritePaths=/home/adam/orchestrator/agent-team/state`. **Retained unchanged**
|
||||
in this PR (locked decision — do not revert to secrev parity).
|
||||
- The `ReadWritePaths` carve-out matches the ledger + audit-log location
|
||||
(`state/agent_team.sqlite`, `state/audit.log.jsonl`). **CLEAN.**
|
||||
|
||||
---
|
||||
|
||||
## Summary table
|
||||
|
||||
| ID | Sev | Status | One-line |
|
||||
|---|---|---|---|
|
||||
| D-1 | 🔴 | ✅ RESOLVED (PR) | `serve` starts `SlackListener` (Slack + app-token gated; Slack stays optional) |
|
||||
| D-2 | 🔴 | ✅ RESOLVED (PR) | unit loads `~/orchestrator/.env` for the P2 GPT-4.1 review provider key |
|
||||
| D-7 | 🟡 | ✅ RESOLVED (PR) | `ExecStart` points at the agent-team venv interpreter |
|
||||
| D-6 | 🟡 | ✅ RESOLVED | `slack_bolt` now exercised by the daemon (consequence of D-1) |
|
||||
| D-4 | 🟠 | NOTE | pip list omits `requests` — runbook Step 4 installs it |
|
||||
| D-5 | 🟠 | NOTE | agent-team runtime deps not pinned in `requirements.txt` — runbook pins langgraph |
|
||||
| operator-CLI | — | NOTE | `run-team.py` vs `operator_cli.py` divergent verbs — use `run-team.py` |
|
||||
| D-SLACKVAR | ✅ | CLEAN | `SLACK_CHANNEL_ID` matches everywhere |
|
||||
| D-ENV-* | ✅ | CLEAN | bot/oauth/owner/app-token env names match; all read sites now exist |
|
||||
| D-SUBCMD | ✅ | CLEAN | every `run-team.py` verb/flag the docs cite exists |
|
||||
| D-HARDENING | ✅ | CLEAN | unit hardening retained unchanged (locked decision) |
|
||||
290
docs/provisioning/OPERATOR-RUNBOOK.md
Normal file
290
docs/provisioning/OPERATOR-RUNBOOK.md
Normal file
|
|
@ -0,0 +1,290 @@
|
|||
# OPERATOR-RUNBOOK — R720 agent-team coordinator incident handling
|
||||
|
||||
The on-call runbook for the always-on `agent-team-coordinator` daemon on the
|
||||
`sh-secrev` R720 VM. Covers pipeline stalls, stuck/parked tasks, failed
|
||||
human-in-the-loop resumes, budget exhaustion mid-pipeline, transport outages, and
|
||||
the COMPLACENCY / COVERAGE alarms. Grounds every recovery in real code (design
|
||||
§5 escalation ladder + §6.6 contention/park policy; Phase-6 requirement).
|
||||
|
||||
> **CLI used throughout: `run-team.py`** (the entry CLI), run from
|
||||
> `~/orchestrator/agent-team` with the venv active so it hits the default ledger
|
||||
> (`state/agent_team.sqlite`) and audit log (`state/audit.log.jsonl`). Do **not**
|
||||
> use `agent_team/operator_cli.py` — it has divergent verbs (no `show`, required
|
||||
> `--db`/`--audit-log`, and a `force-resume` that supersedes). See DEPLOY-AUDIT.md.
|
||||
|
||||
```bash
|
||||
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
|
||||
cd ~/orchestrator/agent-team && . .venv/bin/activate
|
||||
```
|
||||
|
||||
## Verb reference (all confirmed in `run-team.py:build_parser`)
|
||||
|
||||
| Verb | Effect | Destructive? |
|
||||
|---|---|---|
|
||||
| `list` / `list --all` / `list --parked` / `list --status <state>` | read pending questions (default `open`) | no |
|
||||
| `show <question_id>` | print one ledger row (JSON) | no |
|
||||
| `redeliver <question_id>` | clear `channel_ref` so the reconcile loop re-posts an `open` question | no (audit-logged) |
|
||||
| `expire <question_id> --confirm` | force `open`→`expired` | yes |
|
||||
| `answer <question_id> --answer <p> [--via <id>] --confirm` | answer-on-behalf (first-answer-wins CAS) | yes |
|
||||
| `force-resume <question_id> --confirm` | reopen an `expired` (parked) question; for `answered` records resume intent | yes |
|
||||
| `supersede <question_id> --confirm` | mark a stale `open`/`answered` row `superseded` | yes |
|
||||
| `start --task "..." [--transport ...] [--dry-run]` | start one task to the human gate | no |
|
||||
|
||||
`--operator <name>` (global) sets the audit attribution; it defaults to the OS
|
||||
login. Destructive verbs require `--confirm` and write an attempt-then-outcome
|
||||
record to the audit log **before** mutating.
|
||||
|
||||
First triage for any incident:
|
||||
|
||||
```bash
|
||||
systemctl status agent-team-coordinator.service
|
||||
journalctl -u agent-team-coordinator.service -e --since "-2h" | tail -100
|
||||
python3 run-team.py list --all # full ledger snapshot
|
||||
python3 run-team.py list --parked # non-open rows (parked-task context)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Incident 1 — Pipeline stall (tasks not advancing)
|
||||
|
||||
**Symptoms:** `list` shows `open` (or `answered`) rows that never progress; the
|
||||
journal shows no `tick` activity or repeated errors.
|
||||
|
||||
**Diagnose:**
|
||||
```bash
|
||||
systemctl is-active agent-team-coordinator.service # "active" expected
|
||||
journalctl -u agent-team-coordinator.service -e | tail -60
|
||||
# Daemon dead/looping on restart? Check the unit + deps:
|
||||
.venv/bin/python -c "import langgraph, slack_sdk, slack_bolt, requests; print('deps ok')"
|
||||
```
|
||||
|
||||
**Recover:**
|
||||
1. If the daemon is dead and `Restart=on-failure` is flapping, read the journal
|
||||
for the import/auth error. A missing venv dep (D-4/D-5) or a missing
|
||||
`CLAUDE_CODE_OAUTH_TOKEN` is the usual cause — fix the env / venv, then:
|
||||
```bash
|
||||
sudo systemctl restart agent-team-coordinator.service
|
||||
```
|
||||
2. The startup `recover()` sweep re-drives `answered`-but-unresumed rows and
|
||||
re-posts `open` rows that lost their `channel_ref`, so a clean restart
|
||||
converges from the durable ledger. Confirm with `journalctl ... | tail` and
|
||||
`list --all`.
|
||||
3. A single task stuck `open` with a stale/missing post: re-deliver it.
|
||||
```bash
|
||||
python3 run-team.py show <qid>
|
||||
python3 run-team.py redeliver <qid> # clears channel_ref; reconcile re-posts
|
||||
```
|
||||
|
||||
**Escalate** (see the ladder) if a restart does not clear it within one tick
|
||||
cadence and the journal shows a non-transient error.
|
||||
|
||||
---
|
||||
|
||||
## Incident 2 — Stuck / parked task
|
||||
|
||||
A task parks (design §6.6) when its clarifier question **expires** with no answer,
|
||||
or its budget headroom drops below reserve, or a checkpoint is corrupt. Parked =
|
||||
the durable ledger row is no longer `open` (it is `expired`), and an ALARM was
|
||||
raised, not spun on.
|
||||
|
||||
**Diagnose:**
|
||||
```bash
|
||||
python3 run-team.py list --parked
|
||||
python3 run-team.py show <qid> # status, thread_id, deadline_at, answered_via
|
||||
```
|
||||
|
||||
**Recover — depends on why it parked:**
|
||||
|
||||
- **Expired with no answer (the common case)** — un-park by reopening the expired
|
||||
question; it is then re-delivered for an answer:
|
||||
```bash
|
||||
python3 run-team.py force-resume <qid> --confirm
|
||||
# "force-resume: reopened expired question <qid>; it will be re-delivered"
|
||||
python3 run-team.py show <qid> # status flips back to "open"
|
||||
```
|
||||
Then answer it (Slack or CLI) to drive it forward.
|
||||
|
||||
- **Answered but not yet resumed** — the recovery sweep handles it; `force-resume`
|
||||
records intent and reports that (no mutation):
|
||||
```bash
|
||||
python3 run-team.py force-resume <qid> --confirm
|
||||
# "force-resume: question <qid> is answered and pending resume; the recovery
|
||||
# sweep will resume it (intent recorded)"
|
||||
sudo systemctl restart agent-team-coordinator.service # forces the recover() sweep now
|
||||
```
|
||||
|
||||
- **Stale / wrong question that should be abandoned** — supersede it so it stops
|
||||
surfacing as parked context:
|
||||
```bash
|
||||
python3 run-team.py supersede <qid> --confirm
|
||||
```
|
||||
|
||||
`MAX_PARK` FIFO-aging: a task that exceeds the park window escalates (ALARM + a
|
||||
Jira ticket per the ladder) rather than starving silently.
|
||||
|
||||
---
|
||||
|
||||
## Incident 3 — Failed human-in-the-loop resume
|
||||
|
||||
**Symptoms:** an answer was submitted (Slack or CLI) but the graph did not
|
||||
advance.
|
||||
|
||||
**Diagnose:**
|
||||
```bash
|
||||
python3 run-team.py show <qid> # is status "answered"? what answered_via?
|
||||
journalctl -u agent-team-coordinator.service -e | grep -iE "resume|answer|<qid>"
|
||||
```
|
||||
|
||||
**Recover:**
|
||||
- **Status is `answered` but no resume drained** — the resume queue is in-process;
|
||||
a daemon restart triggers the `recover()` sweep that re-drives `answered` rows
|
||||
via the turn-guarded `ResumeWorker` (idempotent — an already-advanced thread
|
||||
supersedes-and-skips):
|
||||
```bash
|
||||
sudo systemctl restart agent-team-coordinator.service
|
||||
python3 run-team.py show <qid> # confirm it advanced
|
||||
```
|
||||
- **Live Slack answer never registered** — the listener fails closed. Check:
|
||||
```bash
|
||||
journalctl -u agent-team-coordinator.service -e | grep -i "Slack answer"
|
||||
# "rejecting Slack answer: owner allowlist is unconfigured ..." -> set
|
||||
# AGENT_TEAM_SLACK_OWNER_IDS in ~/secrev.env and restart.
|
||||
# "rejecting Slack answer ... unauthorized sender" -> the answerer's user id is
|
||||
# not in the allowlist. Add it, or answer via the CLI on their behalf:
|
||||
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
|
||||
```
|
||||
- **Answer lost the compare-and-set (`not open`)** — the row was already
|
||||
answered/expired/superseded. Inspect with `show`; if it parked, go to Incident 2.
|
||||
|
||||
---
|
||||
|
||||
## Incident 4 — Budget exhaustion mid-pipeline
|
||||
|
||||
The shared Claude budget ledger (`budget_ledger` table) enforces per-call + total
|
||||
nightly caps (design §6.1, §6.6). A stage that would breach the reserve **parks**
|
||||
the task (deferred, not dropped) and ALARMs — it never loop-drains the pool.
|
||||
|
||||
**Diagnose:**
|
||||
```bash
|
||||
journalctl -u agent-team-coordinator.service -e | grep -iE "budget|reserve|park"
|
||||
# Inspect today's spend directly (no CLI verb for the budget ledger; read it).
|
||||
# Columns (agent_team/db/schema.py budget_ledger DDL): thread_id, stage, model,
|
||||
# billing_mode, input_tokens, output_tokens, usd_cost, recorded_at, day_bucket.
|
||||
sqlite3 state/agent_team.sqlite \
|
||||
"SELECT day_bucket, thread_id, ROUND(SUM(usd_cost),4) AS usd, SUM(input_tokens) AS in_tok
|
||||
FROM budget_ledger GROUP BY day_bucket, thread_id ORDER BY day_bucket DESC LIMIT 20;"
|
||||
```
|
||||
|
||||
**Recover:**
|
||||
- **Wait for the next budget window** — budget-exhausted roles/tasks are deferred
|
||||
via the rotation pointer and picked up next cycle; this is the intended
|
||||
behavior, not a failure. The parked task surfaces in `list --parked`.
|
||||
- **Force a specific parked task forward now** (e.g. it is urgent and headroom has
|
||||
since freed): un-park it and let the daemon re-run the stage within the
|
||||
remaining cap:
|
||||
```bash
|
||||
python3 run-team.py force-resume <qid> --confirm
|
||||
```
|
||||
- **Confirm `ANTHROPIC_API_KEY` is absent** — its presence would silently meter to
|
||||
API rates and blow the budget model (the billing seam pops it defensively, but
|
||||
it must not be set):
|
||||
```bash
|
||||
grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # both must print 0
|
||||
```
|
||||
|
||||
Do **not** raise the cap to push a task through without Adam's decision — budget
|
||||
realism is a deliberate guardrail. Escalate per the ladder if a task repeatedly
|
||||
parks on budget across nights.
|
||||
|
||||
---
|
||||
|
||||
## Incident 5 — Transport outage (Slack / GitHub down or misconfigured)
|
||||
|
||||
**Symptoms:** clarifier posts fail; the journal shows transport errors or the
|
||||
inbound listener is not started.
|
||||
|
||||
**Diagnose:**
|
||||
```bash
|
||||
journalctl -u agent-team-coordinator.service -e | grep -iE "Slack|listener|post|transport"
|
||||
# Did the inbound listener start?
|
||||
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
|
||||
# "started (Socket Mode ...)" -> inbound up
|
||||
# "not started (... SLACK_APP_TOKEN is unset)" -> Slack inbound intentionally off
|
||||
```
|
||||
|
||||
**Recover:**
|
||||
- **Outbound post failed (Slack API / network)** — the ledger row stays `open`
|
||||
with no `channel_ref` (the responder leaves it for reconcile). After the
|
||||
transport recovers, the `recover()` sweep (on restart) or a manual `redeliver`
|
||||
re-posts:
|
||||
```bash
|
||||
python3 run-team.py redeliver <qid>
|
||||
```
|
||||
- **Inbound listener down / never started** — the daemon still posts and expires;
|
||||
it just cannot hear Slack. **The CLI `answer` path is the outage fallback** —
|
||||
it runs the identical compare-and-set with no socket:
|
||||
```bash
|
||||
python3 run-team.py answer <qid> --answer "<payload>" --via "cli:adam" --confirm
|
||||
```
|
||||
Fix `SLACK_APP_TOKEN` / `SLACK_BOT_TOKEN` in `~/secrev.env`, then
|
||||
`sudo systemctl restart agent-team-coordinator.service` to re-establish Socket
|
||||
Mode. (D-1: the listener starts only when the transport is live Slack AND
|
||||
`SLACK_APP_TOKEN` is set; it is optional by design.)
|
||||
- **Slack listener crash-looping** — the listener runs on an isolated daemon
|
||||
thread; a crash is logged (`inbound Slack listener thread exited with an error`)
|
||||
and takes down only the inbound socket, not the maintenance loop. Restart the
|
||||
service to relaunch the thread once the cause is fixed.
|
||||
|
||||
---
|
||||
|
||||
## Incident 6 — COMPLACENCY / COVERAGE alarms (Plane-1 checkers)
|
||||
|
||||
These come from the nightly checker run, not the coordinator daemon (design §6.4,
|
||||
§6.6):
|
||||
|
||||
- **COMPLACENCY ALARM** — a checker role missed a planted canary fault. That role
|
||||
is **skipped** for the night (it never runs silently degraded). Recover: inspect
|
||||
the canary corpus / the role's checker under `security-review/checkers/`, fix the
|
||||
regression, re-run that checker's dry-run, and confirm the canary passes before
|
||||
re-enabling.
|
||||
- **COVERAGE ALARM** — a role slipped its rotation slot (e.g. budget-deferred). It
|
||||
is **deferred via the rotation pointer, never dropped**, and picked up next
|
||||
cycle. Recover: confirm the rotation pointer advanced (it is rebuildable from
|
||||
report history) and that the role runs on the next cadence; investigate only if
|
||||
it slips repeatedly.
|
||||
|
||||
Both alarms are **report + ALARM-only** (design D3): nothing posts on a clean
|
||||
state, no auto-Jira/Notion writes from the checker itself. Persistent alarms
|
||||
follow the escalation ladder below.
|
||||
|
||||
---
|
||||
|
||||
## Escalation ladder (design §5 / §6.6, resolves Q3)
|
||||
|
||||
Anything that does not clear on the first ALARM escalates — but **ALARM-only in
|
||||
spirit** (nothing posts on a clean state):
|
||||
|
||||
1. **Re-alarm on a backoff.** A confirmed critical (or a COMPLACENCY / COVERAGE
|
||||
alarm) that persists re-alarms to Slack each night it is still unresolved, on
|
||||
a backoff so it does not spam.
|
||||
2. **Open a Jira tracking ticket after `N` nights** (default **N = 3**). If the
|
||||
condition still has not cleared, the coordinator opens an **INFRA** Jira ticket
|
||||
so it cannot quietly linger. The same ladder applies to a parked task that
|
||||
exceeds `MAX_PARK`.
|
||||
3. **Human (Adam) takes it from the Jira ticket.** For a parked task, recover via
|
||||
Incidents 2–4 above; for a checker alarm, via Incident 6.
|
||||
|
||||
When you resolve an incident, record the action — destructive CLI verbs already
|
||||
write an attributable attempt+outcome record to `state/audit.log.jsonl`; for
|
||||
non-CLI recoveries note it on the Jira ticket.
|
||||
|
||||
---
|
||||
|
||||
## Post-incident
|
||||
|
||||
- Confirm the daemon is `active` and the ledger has no unexpected parked rows
|
||||
(`list --parked`).
|
||||
- If you restored the VM snapshot or wiped the ledger, re-run the relevant
|
||||
PROVISIONING-RUNBOOK steps.
|
||||
- Update `project_r720_agent_team` memory if the incident revealed a durable
|
||||
fact (a new failure mode, a config that must change).
|
||||
286
docs/provisioning/P1-DEMO-SCRIPT.md
Normal file
286
docs/provisioning/P1-DEMO-SCRIPT.md
Normal file
|
|
@ -0,0 +1,286 @@
|
|||
# P1-DEMO-SCRIPT — live four-criteria acceptance demo (design §3.3.1 / §7.1 P1)
|
||||
|
||||
The §7.1 P1 exit gate: demonstrate, on the live box, all four durable
|
||||
human-in-the-loop criteria before P1 is accepted:
|
||||
|
||||
- **(a)** kill the box mid-wait and have the task resume after restart;
|
||||
- **(b)** submit a duplicate answer and confirm it no-ops;
|
||||
- **(c)** submit an answer after the deadline expired and confirm it is rejected
|
||||
and the task parks;
|
||||
- **(d)** two tasks suspended concurrently resume independently to the correct
|
||||
thread.
|
||||
|
||||
Every command below is grounded in the **actual** code surface
|
||||
(`run-team.py`, `coordinator.py`, `slack_listener.py`, `responder.py`,
|
||||
`resume_worker.py`, `db/schema.py`, `graph.py`). No invented flags. Where the
|
||||
code does not expose a needed knob (e.g. a short deadline), the script uses a
|
||||
direct `sqlite3` write against the documented `pending_questions` schema and says
|
||||
so.
|
||||
|
||||
## Two answer paths — live Slack OR the operator CLI
|
||||
|
||||
**D-1 is fixed:** `Coordinator.serve()` now starts the inbound `SlackListener`
|
||||
when the live transport is Slack AND `SLACK_APP_TOKEN` is set, so the
|
||||
**live Slack answer round-trip works**. You can run the demo either way:
|
||||
|
||||
- **Live Slack** — Adam clicks the Block Kit button / replies in the channel; the
|
||||
listener normalizes the event, runs the AUTHZ-01 owner check, and drives the
|
||||
first-answer-wins compare-and-set (`responder.submit_answer` →
|
||||
`db.schema.answer_question`).
|
||||
- **Operator CLI** — `run-team.py answer <qid> --answer ... --confirm` runs the
|
||||
**identical** compare-and-set (audit-logged answer-on-behalf). Useful when the
|
||||
Slack app is not yet provisioned, or to script the assertions.
|
||||
|
||||
Both exercise the same durable mechanic; the human gate decision is always
|
||||
Adam's. The assertions below assert on the **durable ledger status** (the §3.3.1
|
||||
source of truth) and are identical for either path. The examples use the CLI
|
||||
`answer --confirm` form so they are copy-pasteable; substitute "Adam answers in
|
||||
Slack" wherever you prefer the live path.
|
||||
|
||||
> The automated proof of this mechanic is `tests/sim/test_p1_exit_criteria.py`
|
||||
> (a `SimPipeline` harness) and `tests/test_coordinator.py` (the serve/listener
|
||||
> wiring). This script is the live-box demonstration on top of that.
|
||||
|
||||
## Verb map (use `run-team.py`, not `operator_cli.py`)
|
||||
|
||||
| Need | `run-team.py` verb | Notes |
|
||||
|---|---|---|
|
||||
| start a task to the human gate | `start --task "..." [--transport slack] [--dry-run]` | mints a `thread_id`, posts the clarifier, writes the `open` ledger row |
|
||||
| list waiting questions | `list` (default `open`) / `list --all` / `list --parked` | JSON rows |
|
||||
| inspect one row | `show <question_id>` | JSON row |
|
||||
| answer on the task's behalf | `answer <question_id> --answer <payload> --confirm` | destructive, audit-logged; first-answer-wins CAS |
|
||||
| force-expire an open question | `expire <question_id> --confirm` | destructive; flips `open`→`expired` |
|
||||
| un-park (reopen) an expired question | `force-resume <question_id> --confirm` | only acts on `expired` rows (reopens them) |
|
||||
|
||||
`operator_cli.py` has different verbs (`force-expire`, `answer-on-behalf`) and
|
||||
**no `show`**, and requires `--db`/`--audit-log` — do not use it here.
|
||||
|
||||
## Preconditions
|
||||
|
||||
```bash
|
||||
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
|
||||
cd ~/orchestrator/agent-team
|
||||
. .venv/bin/activate # so run-team.py uses the venv deps
|
||||
# Confirm the ledger exists (Step 5 of the runbook):
|
||||
python3 run-team.py list --all # [] on a fresh DB is fine
|
||||
# For the live-Slack path, confirm the listener started:
|
||||
journalctl -u agent-team-coordinator.service -e | grep -i "inbound Slack listener"
|
||||
# -> "inbound Slack listener started (Socket Mode, background thread)"
|
||||
```
|
||||
|
||||
Run all `run-team.py` commands from `~/orchestrator/agent-team` so they hit the
|
||||
default ledger (`state/agent_team.sqlite`) and default audit log
|
||||
(`state/audit.log.jsonl`). The clarifier is posted to `SLACK_CHANNEL_ID`.
|
||||
|
||||
> **Resume execution model.** A won answer (rowcount 1) enqueues a `ResumeJob`
|
||||
> onto the coordinator's in-process `resume_queue`; the daemon drains it on the
|
||||
> next `tick()`, or the startup `recover()` re-drives any `answered`-but-unresumed
|
||||
> row after a restart. So a CLI `answer --confirm` (or a live Slack answer) flips
|
||||
> the ledger row to `answered`; the **running daemon** then resumes the graph.
|
||||
> The demo asserts on the durable ledger status.
|
||||
|
||||
---
|
||||
|
||||
## (a) Crash-safe resume — kill mid-wait, restart, task resumes 🧑 Adam answers
|
||||
|
||||
**Setup — drive a task to the clarifier wait.** With the daemon already running:
|
||||
|
||||
```bash
|
||||
python3 run-team.py start --task "demo-a: trivial scoped task"
|
||||
# prints a thread_id, e.g. 3f2a... (record it as $TID_A)
|
||||
python3 run-team.py list
|
||||
# expect one row: status="open", a thread_id, a question_id (record as $QID_A),
|
||||
# channel_ref set (Slack ts) or null if posting is dry/unavailable.
|
||||
```
|
||||
|
||||
**Kill the box mid-wait, then restart:**
|
||||
|
||||
```bash
|
||||
sudo systemctl stop agent-team-coordinator.service
|
||||
# (optionally reboot the VM here for a stronger demonstration)
|
||||
sudo systemctl start agent-team-coordinator.service
|
||||
journalctl -u agent-team-coordinator.service -e | tail -40 # expect "starting" + recover() sweep
|
||||
```
|
||||
|
||||
**ASSERT — durable state survived the kill (no answer yet):**
|
||||
|
||||
```bash
|
||||
python3 run-team.py show $QID_A
|
||||
# PASS: status == "open" (the question was NOT lost across the restart)
|
||||
```
|
||||
|
||||
**🧑 Human gate — Adam answers after the restart (Slack OR CLI):**
|
||||
|
||||
```bash
|
||||
# Live Slack: Adam replies/clicks in the channel. OR via the CLI:
|
||||
python3 run-team.py answer $QID_A --answer "scope: just demo, no real change" --confirm
|
||||
python3 run-team.py show $QID_A
|
||||
# PASS: status == "answered"
|
||||
# Wait one tick (~30s, DEFAULT_POLL_INTERVAL) for the daemon to drain the resume:
|
||||
journalctl -u agent-team-coordinator.service -e | tail -20
|
||||
```
|
||||
|
||||
**PASS criterion (a):** the question stayed `open` across the kill, and an answer
|
||||
submitted *after* the restart drove it forward.
|
||||
|
||||
---
|
||||
|
||||
## (b) Duplicate answer is a no-op 🧑 Adam answers
|
||||
|
||||
```bash
|
||||
python3 run-team.py start --task "demo-b: duplicate-answer test"
|
||||
python3 run-team.py list # record the new question_id as $QID_B (status open)
|
||||
```
|
||||
|
||||
**🧑 First answer (the winning one) — Slack OR CLI:**
|
||||
|
||||
```bash
|
||||
python3 run-team.py answer $QID_B --answer "first-answer" --via "slack:U1" --confirm
|
||||
# expect: "answered question <QID_B> (via slack:U1)" (exit 0)
|
||||
```
|
||||
|
||||
**Duplicate / second answer (must lose the compare-and-set):**
|
||||
|
||||
```bash
|
||||
python3 run-team.py answer $QID_B --answer "second-answer" --via "github:U2" --confirm
|
||||
# expect: "answer no-op: question <QID_B> was not 'open' ..." on stderr, exit 1
|
||||
```
|
||||
|
||||
For the **live-Slack** variant, Adam (or a second authorized owner) answers the
|
||||
same message twice; the second event loses the CAS and is logged
|
||||
`ignored Slack answer for question_id=... (not open: duplicate ...)`.
|
||||
|
||||
**ASSERT — the first answer is preserved:**
|
||||
|
||||
```bash
|
||||
python3 run-team.py show $QID_B
|
||||
# PASS: status == "answered"; answered_via == "slack:U1" (the FIRST answer)
|
||||
```
|
||||
|
||||
**PASS criterion (b):** first-answer-wins (`rowcount==1`); the duplicate hits the
|
||||
`BEGIN IMMEDIATE` compare-and-set in `answer_question` and is ignored
|
||||
(`rowcount==0`) — no second resume, no overwrite.
|
||||
|
||||
---
|
||||
|
||||
## (c) Past-deadline answer rejected + task parks 🧑 Adam observes
|
||||
|
||||
> **Code reality:** `run-team.py start` always sets the clarifier deadline from
|
||||
> `DEFAULT_CLARIFY_DEADLINE = 24h`. There is **no CLI flag for a short deadline**.
|
||||
> Two faithful ways to demo expiry without waiting 24h:
|
||||
|
||||
### Option C1 — force the deadline past, let the daemon's sweep expire it (most faithful)
|
||||
|
||||
```bash
|
||||
python3 run-team.py start --task "demo-c: deadline test"
|
||||
python3 run-team.py list # record $QID_C (status open)
|
||||
|
||||
# Set this question's deadline into the past directly in the ledger:
|
||||
sqlite3 state/agent_team.sqlite \
|
||||
"UPDATE pending_questions SET deadline_at='2000-01-01T00:00:00+00:00' WHERE question_id='$QID_C';"
|
||||
|
||||
# Wait one daemon tick (~30s) for the deadline sweep to flip it, OR observe:
|
||||
journalctl -u agent-team-coordinator.service -e | tail -20
|
||||
# expect the park ALARM line: "task parked: clarifier question <QID_C> expired ..."
|
||||
```
|
||||
|
||||
### Option C2 — operator force-expire (if you do not want to touch the DB)
|
||||
|
||||
```bash
|
||||
python3 run-team.py start --task "demo-c: deadline test"
|
||||
python3 run-team.py list # record $QID_C
|
||||
python3 run-team.py expire $QID_C --confirm # destructive, audit-logged; open->expired
|
||||
```
|
||||
|
||||
**ASSERT — the question is expired and the late answer is rejected:**
|
||||
|
||||
```bash
|
||||
python3 run-team.py show $QID_C
|
||||
# PASS: status == "expired"
|
||||
|
||||
# 🧑 Adam submits a LATE answer (Slack OR CLI) — it must lose the CAS:
|
||||
python3 run-team.py answer $QID_C --answer "too-late" --confirm
|
||||
# expect: "answer no-op: question <QID_C> was not 'open' ..." stderr, exit 1
|
||||
python3 run-team.py show $QID_C
|
||||
# PASS: status still "expired"; answer_json still NULL
|
||||
|
||||
# The parked task surfaces in the parked view:
|
||||
python3 run-team.py list --parked # PASS: $QID_C appears here
|
||||
```
|
||||
|
||||
**Deliberate un-park (proves the operator recovery path, §6.6):**
|
||||
|
||||
```bash
|
||||
python3 run-team.py force-resume $QID_C --confirm
|
||||
# expect: "force-resume: reopened expired question <QID_C>; it will be re-delivered"
|
||||
python3 run-team.py show $QID_C
|
||||
# status flips back to "open" (reopen_question), deadline_at cleared.
|
||||
```
|
||||
|
||||
**PASS criterion (c):** the expired question rejects the late answer, the task
|
||||
parks rather than spins, and the operator can deliberately un-park it via
|
||||
`force-resume` (which reopens only an `expired` row).
|
||||
|
||||
> **`force-resume` semantics (verified):** `run-team.py force-resume` reopens an
|
||||
> `expired` question (the parked case). For an `answered` question it records
|
||||
> intent and reports the recovery sweep will resume it (no mutation). For
|
||||
> `open`/`superseded`/absent it is a no-op exit 1. It does **not** supersede.
|
||||
|
||||
---
|
||||
|
||||
## (d) Two concurrent tasks resume independently 🧑 Adam answers
|
||||
|
||||
```bash
|
||||
python3 run-team.py start --task "demo-d task A" # record $TID_A2, then:
|
||||
python3 run-team.py start --task "demo-d task B" # record $TID_B2
|
||||
python3 run-team.py list
|
||||
# expect TWO open rows with DISTINCT thread_id AND distinct question_id.
|
||||
# Record $QID_A2 and $QID_B2 — match by thread_id.
|
||||
```
|
||||
|
||||
**(Optional) restart the daemon first** to also show concurrent tasks survive a
|
||||
restart, then answer.
|
||||
|
||||
**🧑 Answer the SECOND task first, with a distinct answer, then the first
|
||||
(Slack OR CLI):**
|
||||
|
||||
```bash
|
||||
python3 run-team.py answer $QID_B2 --answer "answer-for-B" --via "slack:U2" --confirm
|
||||
python3 run-team.py answer $QID_A2 --answer "answer-for-A" --via "slack:U1" --confirm
|
||||
# Wait one daemon tick (~30s) for both resumes to drain.
|
||||
```
|
||||
|
||||
**ASSERT — each task carries its OWN answer; no cross-talk:**
|
||||
|
||||
```bash
|
||||
python3 run-team.py show $QID_A2
|
||||
# PASS: status "answered", answered_via "slack:U1", thread_id == $TID_A2
|
||||
python3 run-team.py show $QID_B2
|
||||
# PASS: status "answered", answered_via "slack:U2", thread_id == $TID_B2
|
||||
python3 run-team.py list --all
|
||||
# PASS: the two rows resolved on their own thread_id; no cross-contamination.
|
||||
```
|
||||
|
||||
**PASS criterion (d):** two concurrently-suspended tasks each resumed to their
|
||||
own `thread_id` with their own answer — answering B before A did not misroute,
|
||||
and the per-thread single-flight guard kept them independent.
|
||||
|
||||
---
|
||||
|
||||
## Acceptance
|
||||
|
||||
P1 is accepted only when **(a), (b), (c), and (d) all pass** on the live box.
|
||||
Record the four `show` outputs (or `journalctl` excerpts) as evidence. Then
|
||||
complete the PROVISIONING-RUNBOOK post-session definition-of-done (memory +
|
||||
Confluence + the mandatory `/sh-security-review` on the Slack inbound listener).
|
||||
|
||||
## What can only be verified on the live box
|
||||
|
||||
- That the daemon actually drains the resume and the LangGraph checkpoint
|
||||
advances (asserted via ledger status + journalctl; the graph-state advance is
|
||||
observable only on the box).
|
||||
- The live **Slack** post/answer round-trip end-to-end (the listener is wired —
|
||||
D-1 fixed — but the live Socket Mode socket + a real `SLACK_BOT_TOKEN` /
|
||||
`SLACK_APP_TOKEN` / channel membership are only present on the box).
|
||||
- The exact `channel_ref` value (Slack message `ts`) — depends on a live Slack
|
||||
post succeeding.
|
||||
402
docs/provisioning/PROVISIONING-RUNBOOK.md
Normal file
402
docs/provisioning/PROVISIONING-RUNBOOK.md
Normal file
|
|
@ -0,0 +1,402 @@
|
|||
# PROVISIONING RUNBOOK — R720 agent-team Plane-2 coordinator
|
||||
|
||||
This is the ordered command sequence for the **operator-present** provisioning
|
||||
session that stands up the always-on `agent-team-coordinator` daemon on the
|
||||
`sh-secrev` R720 VM. Derived from `agent-team/DEPLOY-R720.md`,
|
||||
`security-review/DEPLOY-R720.md`, and `docs/r720-agent-team-design.md`
|
||||
(§3.3.1 / §7 / §7.1).
|
||||
|
||||
> ## State of this runbook
|
||||
>
|
||||
> The build is **merged and final**: all six Plane-1 checkers
|
||||
> (`aws-posture`, `compliance-drift`, `confluence-doc`, `dependency-cve`,
|
||||
> `doc-drift`, `plan-groomer`), the Tier-3 dep-bump **fixer**
|
||||
> (`run-team.py fix --dry-run`), and the Plane-1→Plane-2 **P5 cross-plane loop**
|
||||
> (`run-team.py intake-checker`) exist in the tree. The three deploy-correctness
|
||||
> bugs the provisioning-prep audit found are **FIXED** (see DEPLOY-AUDIT.md):
|
||||
>
|
||||
> - **D-1 RESOLVED** — `Coordinator.serve()` now starts the inbound Slack
|
||||
> `SlackListener` concurrently with the tick/drain loop when Slack is the live
|
||||
> transport AND `SLACK_APP_TOKEN` is set. **The demo can use the live Slack
|
||||
> answer path** (P1-DEMO-SCRIPT.md), or the operator-CLI `answer` path.
|
||||
> - **D-2 RESOLVED** — the systemd unit loads `~/orchestrator/.env` (for the P2
|
||||
> GPT-4.1 review loop's provider key) in addition to `~/secrev.env`.
|
||||
> - **D-7 RESOLVED** — the unit's `ExecStart` points at the agent-team venv
|
||||
> interpreter, not the system `python3`.
|
||||
>
|
||||
> Still open as provisioning notes (not blockers): **D-4/D-5** (agent-team
|
||||
> runtime deps are installed ad-hoc into the venv and are not pinned in
|
||||
> `requirements.txt`), and the **operator-CLI divergence** (`run-team.py` vs
|
||||
> `agent_team/operator_cli.py` have different verb names — use `run-team.py`).
|
||||
|
||||
---
|
||||
|
||||
## Scope and ground rules
|
||||
|
||||
**Hard rules (carried from the design and global instructions):**
|
||||
|
||||
- **Snapshot before any stateful change** (design §7; `feedback_ec2_replacement_snapshot`).
|
||||
- **Every stateful step has an exercised rollback.**
|
||||
- Secrets use **placeholder names only**; real values are entered by the operator
|
||||
at the box and never echoed into shell history.
|
||||
- The box is **read-only / subscription-OAuth only**. **No `ANTHROPIC_API_KEY`**
|
||||
on this host (it would silently win over OAuth and meter to API rates —
|
||||
`billing.claude_invoke` pops it defensively, but it must not be present).
|
||||
- **No IAM / OIDC is involved in the coordinator deploy** (P1/P2). IAM enters
|
||||
only at the **P3-live flip** (see the dedicated section), which is gated on the
|
||||
mandatory GPT-4.1 cross-review + `/sh-security-review`.
|
||||
|
||||
**Legend per step:**
|
||||
|
||||
- 🧑 **OPERATOR-REQUIRED** — needs the human (snapshot, secrets, the live human
|
||||
gate, go/no-go). Cannot be automated.
|
||||
- 🤖 **MECHANICAL** — deterministic; an operator runs it but it needs no judgment.
|
||||
|
||||
## Host facts (from both DEPLOY-R720.md files)
|
||||
|
||||
- Hypervisor: R720 at `10.10.60.40` (Windows Server 2022, Hyper-V).
|
||||
- VM: `sh-secrev`, Ubuntu 24.04, **4GB / 2 vCPU / 40GB** dynamic vhdx, `10.10.60.120`.
|
||||
- Reach: `ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120` (key-only, NOPASSWD sudo).
|
||||
- Repo on the box: `~/orchestrator/` (rsync from the Mac, **NOT** a git clone).
|
||||
The package lives at `~/orchestrator/agent-team/`.
|
||||
- Shares `~/secrev.env` (mode 600) with the secrev sweep, and `~/orchestrator/.env`
|
||||
(mode 600) for non-Claude provider keys (same files the secrev unit loads).
|
||||
|
||||
---
|
||||
|
||||
## STEP 1 — Snapshot the VM 🧑 OPERATOR-REQUIRED
|
||||
|
||||
**On the R720 host (Hyper-V), before anything else.** This is the one-command
|
||||
undo for every change below.
|
||||
|
||||
```powershell
|
||||
# On the R720 Windows host (PowerShell, as admin):
|
||||
Checkpoint-VM -Name sh-secrev -SnapshotName "pre-agent-team-coordinator-$(Get-Date -Format yyyyMMdd-HHmm)"
|
||||
Get-VMSnapshot -VMName sh-secrev # confirm the checkpoint exists
|
||||
```
|
||||
|
||||
**ROLLBACK (whole session):**
|
||||
```powershell
|
||||
Restore-VMSnapshot -VMName sh-secrev -Name "<the checkpoint name above>" -Confirm:$false
|
||||
Start-VM -Name sh-secrev
|
||||
```
|
||||
|
||||
> Do not proceed until the checkpoint is confirmed present.
|
||||
|
||||
---
|
||||
|
||||
## STEP 2 — Rsync the repo to the box 🤖 MECHANICAL (verify manifest 🧑)
|
||||
|
||||
**From the Mac.** Same pattern/excludes as the secrev deploy. The team **scans
|
||||
the same `~/repo-mirrors` corpus secrev already maintains** — this rsync ships
|
||||
*code*, not mirrors.
|
||||
|
||||
```bash
|
||||
# From the Mac (sync the canonical repo, not a worktree):
|
||||
rsync -av --exclude .env --exclude .venv --exclude .git --exclude .claude \
|
||||
~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/
|
||||
```
|
||||
|
||||
Surface that must be on the box:
|
||||
|
||||
| Path (under `~/orchestrator/`) | Why it must ship |
|
||||
|---|---|
|
||||
| `agent-team/run-team.py` | the operator entry CLI |
|
||||
| `agent-team/agent_team/**` | the package (coordinator, graph, ledger, transports, slack_listener, fixer) |
|
||||
| `agent-team/systemd/agent-team-coordinator.service` | the daemon unit |
|
||||
| `run.py` + the orchestrator package | the GPT-4.1 cross_reviewer the P2 review loop shells |
|
||||
| `requirements.txt` | pin reference for langgraph / checkpoint-sqlite |
|
||||
| `security-review/lib/**` | shared sweep substrate (Phase 0) |
|
||||
| `security-review/checkers/**` | the six Plane-1 checkers + fixtures |
|
||||
|
||||
**ROLLBACK:** rsync is additive; restore the Step-1 snapshot to revert code state.
|
||||
|
||||
---
|
||||
|
||||
## STEP 3 — Write secrets 🧑 OPERATOR-REQUIRED
|
||||
|
||||
**On the VM.** The coordinator unit loads **two** EnvironmentFiles (both
|
||||
optional via the leading `-`): `~/secrev.env` (agent-team runtime keys) and
|
||||
`~/orchestrator/.env` (the non-Claude provider key for the P2 review loop).
|
||||
**Placeholder names only — the operator pastes real values.** Do not echo real
|
||||
tokens into shell history (use an editor or `read -s`).
|
||||
|
||||
`~/secrev.env` (mode 600) — append the agent-team keys:
|
||||
|
||||
```bash
|
||||
ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120
|
||||
# Edit ~/secrev.env (mode 600) and add — values are placeholders:
|
||||
# CLAUDE_CODE_OAUTH_TOKEN=<CLAUDE_CODE_OAUTH_TOKEN> # from `claude setup-token`
|
||||
# SLACK_BOT_TOKEN=<SLACK_BOT_TOKEN> # xoxb-..., chat:write (posts questions)
|
||||
# SLACK_APP_TOKEN=<SLACK_APP_TOKEN> # xapp-..., connections:write (Socket Mode inbound)
|
||||
# SLACK_CHANNEL_ID=<SLACK_CHANNEL_ID> # C0..., target clarifier channel
|
||||
# AGENT_TEAM_SLACK_OWNER_IDS=<SLACK_OWNER_USER_IDS> # comma-separated U... ids of authorized answerers
|
||||
chmod 600 ~/secrev.env
|
||||
```
|
||||
|
||||
`~/orchestrator/.env` (mode 600) — the non-Claude provider key for the GPT-4.1
|
||||
review loop (same file secrev uses; if it already exists with the key, leave it):
|
||||
|
||||
```bash
|
||||
# OPENAI_API_KEY=<OPENAI_API_KEY> # (or the provider key cross_reviewer/GPT-4.1 needs)
|
||||
chmod 600 ~/orchestrator/.env
|
||||
```
|
||||
|
||||
**Contract notes (verified against the code — see DEPLOY-AUDIT.md):**
|
||||
|
||||
- `SLACK_CHANNEL_ID` is correct — `run-team.py _build_transport` reads exactly
|
||||
`os.environ.get("SLACK_CHANNEL_ID")`. Do **not** use `SLACK_CHANNEL`.
|
||||
- `SLACK_APP_TOKEN` is now **read by the daemon**: `Coordinator.serve()` starts
|
||||
the inbound `SlackListener` when the transport is live Slack AND `SLACK_APP_TOKEN`
|
||||
is set. Without it the daemon still runs (posts + expires) but never hears Slack
|
||||
replies — Slack stays optional by design.
|
||||
- `AGENT_TEAM_SLACK_OWNER_IDS` **fails closed** (AUTHZ-01): if unset/empty the
|
||||
listener rejects **every** answer. It must be set for the live human gate.
|
||||
- **CRITICAL:** confirm `ANTHROPIC_API_KEY` is NOT present:
|
||||
```bash
|
||||
grep -c ANTHROPIC_API_KEY ~/secrev.env ~/orchestrator/.env # must print 0 for both
|
||||
```
|
||||
|
||||
**ROLLBACK:** strip exactly the appended keys (keep secrev keys intact), or
|
||||
restore the Step-1 snapshot. Do NOT blindly truncate — secrev keys live here too.
|
||||
|
||||
---
|
||||
|
||||
## STEP 4 — Create the venv + install pip deps 🤖 MECHANICAL
|
||||
|
||||
**On the VM.** A dedicated venv under `agent-team/.venv` (excluded from rsync).
|
||||
The systemd unit's `ExecStart` points at **this venv's interpreter** (D-7 fixed),
|
||||
so the deps MUST land here.
|
||||
|
||||
```bash
|
||||
cd ~/orchestrator/agent-team
|
||||
python3 -m venv .venv
|
||||
. .venv/bin/activate
|
||||
pip install langgraph==1.1.10 langgraph-checkpoint-sqlite==3.1.0 \
|
||||
claude-agent-sdk slack_sdk slack_bolt requests
|
||||
```
|
||||
|
||||
> **Pin note (D-5):** `requirements.txt` pins `langgraph==1.1.10` /
|
||||
> `langgraph-checkpoint-sqlite==3.1.0`; match those exactly here. The other
|
||||
> runtime deps (`claude-agent-sdk`, `slack_sdk`, `slack_bolt`, `requests`) are
|
||||
> not yet in `requirements.txt` (D-5 open) — installed ad-hoc here. `requests`
|
||||
> is required by the GitHub transport/intake (D-4). `anthropic` is **not**
|
||||
> installed (only the opt-in `api` billing mode needs it).
|
||||
> `slack_bolt` is now actually exercised (D-1 fixed: the daemon starts the
|
||||
> Socket Mode listener).
|
||||
|
||||
**Verify the imports resolve (using the venv interpreter the unit will use):**
|
||||
```bash
|
||||
.venv/bin/python -c "import langgraph, langgraph.checkpoint.sqlite, slack_sdk, slack_bolt, requests; print('deps ok')"
|
||||
.venv/bin/python -c "import claude_agent_sdk; print('agent-sdk ok')"
|
||||
```
|
||||
|
||||
**ROLLBACK:** `deactivate 2>/dev/null; rm -rf ~/orchestrator/agent-team/.venv`
|
||||
|
||||
---
|
||||
|
||||
## STEP 5 — Initialize the durable ledger DB 🤖 MECHANICAL
|
||||
|
||||
**On the VM, venv active.** Idempotent; creates `state/agent_team.sqlite` with
|
||||
the `pending_questions` + `budget_ledger` + `schema_meta` tables (LangGraph
|
||||
`SqliteSaver` creates its own tables in the same file on first run).
|
||||
|
||||
```bash
|
||||
cd ~/orchestrator/agent-team
|
||||
. .venv/bin/activate
|
||||
python3 run-team.py init-db
|
||||
# Expect: "initialized ledger DB at .../state/agent_team.sqlite"
|
||||
ls -l state/ # agent_team.sqlite present; state/ is gitignored
|
||||
```
|
||||
|
||||
**ROLLBACK (reset the ledger only):**
|
||||
```bash
|
||||
cp ~/orchestrator/agent-team/state/agent_team.sqlite{,.bak}
|
||||
rm ~/orchestrator/agent-team/state/agent_team.sqlite*
|
||||
# re-run `python3 run-team.py init-db` to recreate empty tables.
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## STEP 6 — Install + start the systemd unit 🤖 MECHANICAL (go/no-go 🧑)
|
||||
|
||||
**On the VM, as root.** Installs the long-running coordinator daemon. Keep the
|
||||
**hardening as-shipped**: `NoNewPrivileges`, `ProtectSystem=full`,
|
||||
`ProtectHome=read-only` + `ReadWritePaths=.../agent-team/state` (locked decision).
|
||||
|
||||
```bash
|
||||
sudo cp ~/orchestrator/agent-team/systemd/agent-team-coordinator.service /etc/systemd/system/
|
||||
sudo systemctl daemon-reload
|
||||
sudo systemctl enable --now agent-team-coordinator.service
|
||||
systemctl status agent-team-coordinator.service
|
||||
journalctl -u agent-team-coordinator.service -e -f
|
||||
# expect: "agent-team coordinator starting"
|
||||
# and (if Slack + SLACK_APP_TOKEN provisioned):
|
||||
# "inbound Slack listener started (Socket Mode, background thread)"
|
||||
# or otherwise:
|
||||
# "inbound Slack listener not started (... or SLACK_APP_TOKEN is unset) ..."
|
||||
```
|
||||
|
||||
The unit (post-fix) runs
|
||||
`/home/adam/orchestrator/agent-team/.venv/bin/python run-team.py serve` from
|
||||
`WorkingDirectory=/home/adam/orchestrator/agent-team` as `User=adam`, loading
|
||||
**both** `EnvironmentFile=-/home/adam/secrev.env` and
|
||||
`EnvironmentFile=-/home/adam/orchestrator/.env`, `Restart=on-failure`.
|
||||
|
||||
> **D-7 fixed:** `ExecStart` now resolves the venv interpreter, so the Step-4
|
||||
> deps are on the path. **D-1 fixed:** `serve` starts the inbound `SlackListener`
|
||||
> when Slack + app token are provisioned. **Go/no-go:** confirm the journal shows
|
||||
> the listener line you expect for your transport choice before Step 8.
|
||||
|
||||
**ROLLBACK (exercised — tear the unit down cleanly):**
|
||||
```bash
|
||||
sudo systemctl disable --now agent-team-coordinator.service
|
||||
sudo rm /etc/systemd/system/agent-team-coordinator.service
|
||||
sudo systemctl daemon-reload
|
||||
systemctl status agent-team-coordinator.service # should report "could not be found"
|
||||
```
|
||||
The ledger under `state/` is untouched. Step-1 snapshot is the whole-host fallback.
|
||||
|
||||
---
|
||||
|
||||
## STEP 7 — Verify the Slack inbound listener 🧑 OPERATOR-REQUIRED
|
||||
|
||||
The clarifier gate is two halves: **outbound** (post the question — done by
|
||||
`serve`/`start` via the Slack poster) and **inbound** (receive Adam's answer —
|
||||
the `SlackListener` over Socket Mode).
|
||||
|
||||
> **D-1 RESOLVED.** `Coordinator.serve()` now constructs and starts
|
||||
> `SlackListener` on a background daemon thread, concurrently with the tick/drain
|
||||
> loop, **when** the live transport is a `SlackTransport` AND `SLACK_APP_TOKEN`
|
||||
> is set. It shares the coordinator's own transport, ledger, and resume queue,
|
||||
> and stops cleanly on shutdown. The AUTHZ-01 owner allowlist + the open-status
|
||||
> compare-and-set are unchanged — the listener still fails closed on an empty
|
||||
> `AGENT_TEAM_SLACK_OWNER_IDS`.
|
||||
|
||||
**Confirm the live inbound path is up:**
|
||||
|
||||
```bash
|
||||
# In the journal (Step 6) expect the "inbound Slack listener started" line.
|
||||
# Verify the two gating vars are present in the unit's environment:
|
||||
grep -c SLACK_APP_TOKEN ~/secrev.env # 1
|
||||
grep -c AGENT_TEAM_SLACK_OWNER_IDS ~/secrev.env # 1 (else the gate rejects all answers)
|
||||
```
|
||||
|
||||
If you are **not** using Slack as the transport (or are deliberately running
|
||||
without the app token), the daemon runs the maintenance loop only and the demo
|
||||
uses the operator-CLI `answer` path — both are valid (P1-DEMO-SCRIPT.md).
|
||||
|
||||
**ROLLBACK:** none needed — verification only, no host state change.
|
||||
|
||||
---
|
||||
|
||||
## STEP 8 — Live P1 four-criteria acceptance demo 🧑 OPERATOR-REQUIRED
|
||||
|
||||
Run **P1-DEMO-SCRIPT.md** in full. All four §3.3.1 exit criteria must pass:
|
||||
(a) crash-safe resume, (b) duplicate-answer no-op, (c) post-deadline rejection +
|
||||
park, (d) two concurrent tasks resume independently. With D-1 fixed you may
|
||||
exercise the **live Slack answer path** for (b)/(d); the operator-CLI `answer`
|
||||
path remains available and exercises the identical compare-and-set. **Do not
|
||||
accept P1 until all four pass.**
|
||||
|
||||
**ROLLBACK:** the demo writes only ledger rows under `state/`; reset via the
|
||||
Step-5 ledger rollback, or restore the Step-1 snapshot, then re-run.
|
||||
|
||||
---
|
||||
|
||||
## STEP 9 — Plane-1 checker live dry-runs 🧑 OPERATOR-REQUIRED
|
||||
|
||||
The six checkers live under `security-review/checkers/`:
|
||||
`aws-posture.sh`, `compliance-drift.sh`, `confluence-doc.sh`,
|
||||
`dependency-cve.sh`, `doc-drift.sh`, `plan-groomer.sh`. All run **report +
|
||||
ALARM-only** (design D3): a clean run posts nothing and lands a mode-600 report;
|
||||
no auto-Jira/Notion writes.
|
||||
|
||||
- Dry-run each checker against `~/repo-mirrors` in report-only mode; confirm a
|
||||
clean run posts nothing and writes a mode-600 report. Confirm the exact
|
||||
invocation against the checker scripts and the shared
|
||||
`security-review/lib/` substrate on the box.
|
||||
- Confirm the **canary suite** runs first (a planted-fault miss is a COMPLACENCY
|
||||
ALARM and that role is skipped — design §6.4), and the **coverage rotation**
|
||||
pointer advances (a slipped role is a COVERAGE ALARM, deferred-not-dropped).
|
||||
|
||||
**ROLLBACK:** checkers are read-only over the mirror corpus; a dry-run produces
|
||||
only a report file. Remove the report dir to revert; no host state change.
|
||||
|
||||
---
|
||||
|
||||
## STEP 10 — P5 cross-plane loop dry-run 🧑 OPERATOR-REQUIRED
|
||||
|
||||
The Plane-1→Plane-2 loop turns confirmed checker findings into pipeline tasks:
|
||||
|
||||
```bash
|
||||
cd ~/orchestrator/agent-team && . .venv/bin/activate
|
||||
# Read one or more checker report JSONs and start one task per confirmed
|
||||
# at/above-threshold finding (default threshold: high). --dry-run posts nowhere.
|
||||
python3 run-team.py intake-checker --report <path-to-report.json> --threshold high --dry-run
|
||||
```
|
||||
|
||||
The Tier-3 dep-bump **fixer** is dry-run only on the box (it holds no write
|
||||
token, D2):
|
||||
|
||||
```bash
|
||||
python3 run-team.py fix --report security-review/<...>/dependency-cve.json \
|
||||
--finding-id <ID> --task-id <task> --dry-run
|
||||
# prints the fix spec + minimal bump patch + the CI workflow_dispatch inputs;
|
||||
# dispatches NOTHING. Live dispatch is the P3-live flip below.
|
||||
```
|
||||
|
||||
**ROLLBACK:** both are read-only / dry-run (no dispatch, no apply); `intake-checker`
|
||||
de-dup is in-memory per process. No host state to revert beyond ledger rows from
|
||||
a non-dry-run intake (Step-5 ledger rollback).
|
||||
|
||||
---
|
||||
|
||||
## STEP 11 — Wire the schedule 🧑 OPERATOR-REQUIRED
|
||||
|
||||
The coordinator daemon (Step 6) is **always-on**, not timer-driven. The
|
||||
**per-checker timers** (and the design's "shared timer with secrev", §8) attach
|
||||
here. The existing `sea-haven-secrev.timer` (OnCalendar `02:00`, Persistent) is
|
||||
untouched. Add a checker timer only after that checker is dry-run-validated
|
||||
(Step 9).
|
||||
|
||||
**ROLLBACK:** each timer gets its own `systemctl disable --now <unit>.timer` + `rm`.
|
||||
|
||||
---
|
||||
|
||||
## The P3-live flip (deferred; IAM + GitHub App gated) 🧑 OPERATOR-REQUIRED
|
||||
|
||||
P3 (the build→verify apply-and-open-draft-PR loop) is **opt-in and inert** in
|
||||
this deploy: `run-team.py` / `serve` pass `build_verify_wiring=None`, so no P3
|
||||
subgraph is assembled. Flipping it live is a **separate, gated** provisioning
|
||||
session, not part of the coordinator deploy:
|
||||
|
||||
1. **Mandatory reviews first.** The CI trust-boundary + OIDC IAM change is a
|
||||
breaking IAM change → **GPT-4.1 cross-review** (global instructions) AND
|
||||
`/sh-security-review` on the apply/verify surface. Do not flip without both.
|
||||
2. **Provision the GitHub App** for the trusted apply path (the App that opens
|
||||
the draft PR), and the **`agent-apply` GitHub Actions environment** that holds
|
||||
the apply path's scoped permissions.
|
||||
3. **Bind the live build→verify wiring** via
|
||||
`agent_team.coordinator.gated_build_verify_wiring(...)` (the read-only CI
|
||||
result fetcher + the real diff builder) — the seam a leaf calls *after* the
|
||||
gate clears. The CI fetcher is read-only and fails closed (missing token /
|
||||
404 / auth failure → `None` → the gate BLOCKs and the task parks).
|
||||
4. **Set the apply env vars** the live path reads (the read-only CI-result token
|
||||
and the dispatch target), then re-run the fixer **without** `--dry-run` only
|
||||
once the dispatcher is bound.
|
||||
|
||||
Until every step above is done, the box dispatches/applies nothing.
|
||||
|
||||
---
|
||||
|
||||
## Post-session definition-of-done (design §7, global instructions)
|
||||
|
||||
- [ ] All four P1 criteria demonstrated live (Step 8).
|
||||
- [ ] `project_r720_agent_team` memory created/updated.
|
||||
- [ ] Confluence "AWS Architecture Map" / IT host inventory updated to show
|
||||
`sh-secrev` now also hosts the always-on agent-team coordinator daemon.
|
||||
- [ ] `/sh-security-review` run on the Slack inbound listener surface
|
||||
(auth + untrusted-input; mandatory) — flag outstanding if not run.
|
||||
- [ ] OPERATOR-RUNBOOK.md reviewed by whoever holds the pager.
|
||||
- [ ] Snapshot retained until the daemon runs clean for one full cycle, then pruned.
|
||||
Reference in a new issue