This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/agent-team/run-team.py
Adam Moussa f0a2dc27a5 feat(notify): Slack lifecycle notifications + deliver multi-turn questions
The bot only ever posted clarifier questions; parks/completions were silent and
multi-turn follow-up questions were never posted during normal operation (only
the startup recover sweep posted them). So an answered task was a black box.

- Coordinator gains a 'notify' sink + _emit() (guarded, never breaks the loop).
- tick() now runs _post_resume_followups(results) after the drain: for each
  resumed thread it (a) POSTS a newly-pending clarifier question (fixes silent
  multi-turn — the drain path left it unposted) and emits 'needs more input';
  else emits 'parked — needs attention' or 'plan ready for review' from the
  settled state.
- run-team serve wires notify -> Slack channel (build_slack_poster) and an
  alarm_hook that logs the deadline-park WARNING AND posts a parked notice.
  Live-Slack only; dry-run/non-Slack/no-channel = silent (None), no token needed.

Tests: +4 (needs-input/parked/plan-ready emits + notify-failure swallow);
_FakeCoordinator gains notify/alarm_hook. 1147 passed, ruff clean.
2026-06-23 15:49:40 -04:00

1336 lines
50 KiB
Python

#!/usr/bin/env python3
"""``run-team.py`` — R720 agent-team operator entry CLI (design §3.3.1, §7.1 P1).
This is the **entry CLI** named in the design (§2 "entry CLI ``run-team.py``";
§9 "operator CLI"). It is the small manual path over the durable
``pending_questions`` ledger that §3.3.1 ("Manual path") requires::
A small CLI over the ledger lets an operator list ``open``/``parked``
questions, re-deliver, force-expire, or answer on a task's behalf; a stuck
task parks rather than spins. Destructive CLI actions (force-expire,
answer-on-behalf, force-resume) are audit-logged and require an explicit
confirmation flag.
It imports the committed FOUNDATION contracts verbatim — it does not redefine
them:
* :mod:`agent_team.db.schema` — :func:`connect`, :func:`init_db`,
:func:`answer_question`, :func:`expire_question`, :func:`supersede_question`,
:data:`QUESTION_STATES`.
* :mod:`agent_team.state_store` — :func:`atomic_write` for the append-only,
crash-safe audit log of destructive actions (§6.7 discipline).
Per the build constraints this is **pre-deployment scaffolding**: it provisions
nothing, enables no live CI, and performs no network or rsync. It only reads and
mutates the local SQLite ledger and writes a local audit log.
Subcommands (P1 surface):
* ``init-db`` — create/upgrade the agent-team tables in the ledger DB
(idempotent; wraps :func:`init_db`).
* ``list`` — list ``open`` (default) or any-status pending questions; with
``--parked`` it lists questions whose status is read as parked context. Pure
read; no confirmation needed.
* ``show`` — print one question row by ``question_id``. Pure read.
* ``expire`` — force-expire an ``open`` question (DESTRUCTIVE: requires
``--confirm``; audit-logged). Maps to :func:`expire_question`.
* ``answer`` — answer a question on a task's behalf (DESTRUCTIVE: requires
``--confirm``; audit-logged). Maps to :func:`answer_question`.
* ``supersede`` — mark a stale question ``superseded`` (DESTRUCTIVE: requires
``--confirm``; audit-logged). Maps to :func:`supersede_question`.
* ``intake-github`` — poll a repo for labeled open issues and start one pipeline
task per not-yet-ingested issue (one pass). Reads issues + calls the committed
coordinator intake entry only; no CI, OIDC, or git/patch apply. Opt-in/inert.
* ``intake-checker`` — the Plane-1 -> Plane-2 cross-plane loop (P5): read one or
more checker report JSON files / a dir, select CONFIRMED findings at/above a
severity threshold (default ``high``), de-dup, and start one remediation task
per unique finding via the committed coordinator intake entry. Reads local
report JSON only; no CI, OIDC, network, or git/patch apply. Opt-in/inert.
* ``fix`` — plan a Plane-1 Tier-3 dep-bump fix for one confirmed dependency-cve
finding (spec via Claude + minimal bump patch via DeepSeek) and, in
``--dry-run``, print the spec + patch + the org-CI ``workflow_dispatch`` inputs
WITHOUT dispatching. Opt-in/inert: it binds no dispatcher and holds no write
token; live dispatch is provisioning-gated (the P3-live apply/verify surface).
Exit codes: ``0`` success, ``1`` operational failure (e.g. row not found, the
compare-and-set lost the race), ``2`` usage error (argparse).
"""
from __future__ import annotations
import argparse
import getpass
import json
import logging
import os
import sqlite3
import sys
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Callable, Sequence
# ``run-team.py`` lives in ``agent-team/`` next to the importable ``agent_team``
# package. The hyphenated filename cannot itself be imported, so when run as a
# script we make the sibling package importable without an editable install
# (mirrors tests/conftest.py).
_CLI_DIR = Path(__file__).resolve().parent
if str(_CLI_DIR) not in sys.path:
sys.path.insert(0, str(_CLI_DIR))
from agent_team.db.schema import ( # noqa: E402 (path bootstrap must precede)
QUESTION_STATES,
answer_question,
connect,
expire_question,
init_db,
reopen_question,
supersede_question,
)
from agent_team.transport.base import Transport # noqa: E402 (path bootstrap)
# Transport choices the start/serve/intake commands accept (§3.3.1 D10). All
# three now have a live human-gate adapter wired in _build_transport: ``slack``
# (slack_sdk), ``github`` (requests issue/PR comment), and ``claude_code`` (local
# file drop). Each builder defers its optional SDK / token resolution to call
# time, so a missing dep/credential fails loudly only when that transport is
# actually selected.
_TRANSPORT_CHOICES: tuple[str, ...] = ("slack", "github", "claude_code")
__all__ = [
"build_parser",
"main",
]
# Default ledger DB location. Kept out of the repo (the package .gitignore
# excludes ``state/`` and ``*.sqlite``) so durable state is never committed.
_DEFAULT_DB = _CLI_DIR / "state" / "agent_team.sqlite"
# Default audit log for destructive actions, alongside the ledger DB.
_DEFAULT_AUDIT_LOG = _CLI_DIR / "state" / "audit.log.jsonl"
_LOG = logging.getLogger("agent_team.run_team")
# Columns selected for list/show rendering, in display order.
_QUESTION_COLUMNS: tuple[str, ...] = (
"question_id",
"thread_id",
"turn",
"status",
"transport",
"channel_ref",
"posted_at",
"deadline_at",
"answered_at",
"answered_via",
)
# Destructive subcommands that require ``--confirm`` and are audit-logged.
# ``force-resume`` is the design-named operator verb (§3.3.1/§6.6); ``supersede``
# is kept as its lower-level alias. ``redeliver`` is NOT here — it is idempotent
# and non-destructive (it only clears a delivery ref), though it is still
# audit-logged for provenance.
_DESTRUCTIVE_ACTIONS: frozenset[str] = frozenset(
{"expire", "answer", "supersede", "force-resume"}
)
def _utc_now_iso() -> str:
"""Return the current UTC time as an ISO-8601 string (audit timestamps)."""
return datetime.now(timezone.utc).isoformat()
def _default_operator() -> str:
"""Best-effort OS login for audit attribution (never an empty string).
A previous empty default left destructive actions non-attributable (the
audit record named no one). Defaulting to the OS login keeps the §3.3.1
"audit-logged AND attributable" guarantee even when --operator is omitted;
falls back to "unknown" only if the login cannot be resolved.
"""
try:
user = getpass.getuser()
except Exception: # noqa: BLE001 - getuser can raise on odd environments
return "unknown"
return user or "unknown"
def _row_to_dict(row: sqlite3.Row) -> dict[str, Any]:
"""Project a ``pending_questions`` row to a plain dict for display."""
return {col: row[col] for col in _QUESTION_COLUMNS if col in row.keys()}
def _append_audit(audit_log: Path, entry: dict[str, Any]) -> None:
"""Append one JSON audit record with an atomic ``O_APPEND`` single write.
Destructive actions (§3.3.1) must leave an attributable trail. A previous
read-modify-rewrite design lost records under concurrent operators (two
processes each read the same bytes and the last rewrite wins). Instead each
record is one line written with ``O_APPEND``: the kernel serializes the
append and a write below ``PIPE_BUF`` is atomic on POSIX, so concurrent
appends never clobber each other. The file is created mode ``0600``
(operator identity / action content is sensitive) and re-chmod'd in case it
pre-existed wider. A failure here raises ``OSError`` BEFORE any ledger
mutation, preserving the audit-before-mutate guarantee.
"""
audit_log = Path(audit_log)
audit_log.parent.mkdir(parents=True, exist_ok=True)
line = (json.dumps(entry, sort_keys=True) + "\n").encode("utf-8")
fd = os.open(str(audit_log), os.O_WRONLY | os.O_CREAT | os.O_APPEND, 0o600)
try:
os.write(fd, line)
finally:
os.close(fd)
os.chmod(audit_log, 0o600)
def _audit_attempt(
audit_log: Path,
action: str,
*,
question_id: str,
operator: str,
detail: dict[str, Any] | None = None,
) -> None:
"""Record the *intent* to perform a destructive action BEFORE it mutates.
§3.3.1 requires every destructive action to be audit-logged. Writing the
attempt before the ledger mutation closes the "mutation applied with no
audit record" gap: if this append fails (e.g. an unwritable audit path) it
raises before any ledger row is touched, so the action aborts cleanly with
nothing changed. The matching :func:`_audit_outcome` records what happened.
"""
_append_audit(
audit_log,
{
"ts": _utc_now_iso(),
"action": action,
"phase": "attempt",
"question_id": question_id,
"operator": operator,
**(detail or {}),
},
)
def _audit_outcome(
audit_log: Path,
action: str,
*,
question_id: str,
operator: str,
applied: bool,
detail: dict[str, Any] | None = None,
) -> None:
"""Record the *result* of a destructive action AFTER it ran.
Carries ``applied`` (did the compare-and-set change a row). Pairs with the
:func:`_audit_attempt` record written before the mutation, so even if this
outcome append fails the attempt already proves the action was made.
"""
_append_audit(
audit_log,
{
"ts": _utc_now_iso(),
"action": action,
"phase": "outcome",
"question_id": question_id,
"operator": operator,
"applied": applied,
**(detail or {}),
},
)
def _require_confirm(action: str, *, confirm: bool) -> None:
"""Raise unless a destructive ``action`` was explicitly confirmed.
Mirrors the §3.3.1 rule: force-expire, answer-on-behalf, and force-resume
are audit-logged AND require an explicit confirmation flag. Failing closed
here means a typo can never silently mutate a live task's ledger row.
"""
if action in _DESTRUCTIVE_ACTIONS and not confirm:
raise PermissionError(
f"refusing destructive action '{action}' without --confirm "
f"(force-expire / answer-on-behalf / supersede are gated, §3.3.1)"
)
def _fetch_question(conn: sqlite3.Connection, question_id: str) -> sqlite3.Row | None:
"""Return the ledger row for ``question_id`` or ``None`` if absent."""
return conn.execute(
"SELECT * FROM pending_questions WHERE question_id = ?",
(question_id,),
).fetchone()
# --------------------------------------------------------------------------- #
# Subcommand handlers. Each returns a process exit code (0 ok, 1 op failure).
# --------------------------------------------------------------------------- #
def _cmd_init_db(args: argparse.Namespace, *, out: Any) -> int:
"""Create/upgrade the agent-team tables (idempotent)."""
init_db(args.db)
print(f"initialized ledger DB at {args.db}", file=out)
return 0
def _cmd_list(args: argparse.Namespace, *, out: Any) -> int:
"""List pending questions, optionally filtered by status.
Default lists ``open`` questions (the operator's "what is waiting" view).
``--status STATE`` narrows to one lifecycle state; ``--all`` lists every
state. ``--parked`` is a convenience alias that surfaces the parked-task
context an operator chases: questions that are no longer ``open`` (answered
but never resumed, expired, or superseded) and so may back a parked task.
"""
conn = connect(args.db)
try:
if args.all:
rows = conn.execute(
"SELECT * FROM pending_questions ORDER BY thread_id, turn"
).fetchall()
elif args.parked:
placeholders = ",".join("?" for _ in _PARKED_STATES)
rows = conn.execute(
f"SELECT * FROM pending_questions WHERE status IN ({placeholders}) "
"ORDER BY thread_id, turn",
tuple(_PARKED_STATES),
).fetchall()
else:
rows = conn.execute(
"SELECT * FROM pending_questions WHERE status = ? "
"ORDER BY thread_id, turn",
(args.status,),
).fetchall()
finally:
conn.close()
payload = [_row_to_dict(row) for row in rows]
print(json.dumps(payload, indent=2, sort_keys=True), file=out)
return 0
def _cmd_show(args: argparse.Namespace, *, out: Any) -> int:
"""Print one question row by ``question_id`` (pure read)."""
conn = connect(args.db)
try:
row = _fetch_question(conn, args.question_id)
finally:
conn.close()
if row is None:
print(f"no such question: {args.question_id}", file=sys.stderr)
return 1
print(json.dumps(_row_to_dict(row), indent=2, sort_keys=True), file=out)
return 0
def _cmd_expire(args: argparse.Namespace, *, out: Any) -> int:
"""Force-expire an ``open`` question (destructive; audit-logged).
Audits the attempt BEFORE mutating so a mutation can never land without a
trail (§3.3.1); records the outcome after.
"""
_require_confirm("expire", confirm=args.confirm)
_audit_attempt(
args.audit_log, "expire", question_id=args.question_id, operator=args.operator
)
conn = connect(args.db)
try:
changed = expire_question(conn, question_id=args.question_id)
finally:
conn.close()
_audit_outcome(
args.audit_log,
"expire",
question_id=args.question_id,
operator=args.operator,
applied=changed,
)
if not changed:
print(
f"expire no-op: question {args.question_id} was not 'open' "
"(already answered/expired/superseded or absent)",
file=sys.stderr,
)
return 1
print(f"expired question {args.question_id}", file=out)
return 0
def _cmd_redeliver(args: argparse.Namespace, *, out: Any) -> int:
"""Clear an ``open`` question's ``channel_ref`` so it is re-posted (§3.3.1).
The design's "re-deliver" operator action. Re-delivery itself is performed
by the transport reconcile loop; clearing ``channel_ref`` makes that loop
re-post and record a fresh ref. Idempotent and non-destructive (the question
stays ``open``), so it needs no ``--confirm`` — but it is audit-logged for
provenance. Returns ``1`` if the question is absent or not ``open``.
"""
_audit_attempt(
args.audit_log,
"redeliver",
question_id=args.question_id,
operator=args.operator,
)
conn = connect(args.db)
try:
row = _fetch_question(conn, args.question_id)
if row is None:
applied = False
prior_ref = None
status = None
elif row["status"] != "open":
applied = False
prior_ref = row["channel_ref"]
status = row["status"]
else:
prior_ref = row["channel_ref"]
status = "open"
conn.execute(
"UPDATE pending_questions SET channel_ref=NULL WHERE question_id=?",
(args.question_id,),
)
applied = True
finally:
conn.close()
_audit_outcome(
args.audit_log,
"redeliver",
question_id=args.question_id,
operator=args.operator,
applied=applied,
detail={"prior_channel_ref": prior_ref},
)
if not applied:
reason = "absent" if status is None else f"status={status}, not open"
print(
f"redeliver no-op: question {args.question_id} ({reason}); "
"nothing to re-post",
file=sys.stderr,
)
return 1
print(
f"cleared channel_ref for {args.question_id}; reconcile loop will re-post",
file=out,
)
return 0
def _cmd_answer(args: argparse.Namespace, *, out: Any) -> int:
"""Answer a question on a task's behalf (destructive; audit-logged).
Uses the foundation first-answer-wins compare-and-set: succeeds only if the
question is still ``open``. ``--answer`` is stored verbatim as the answer
payload string; ``--via`` records the answering identity for the audit
trail. The audit log records the operator regardless of outcome.
"""
_require_confirm("answer", confirm=args.confirm)
via = args.via or f"cli:{args.operator}"
_audit_attempt(
args.audit_log,
"answer",
question_id=args.question_id,
operator=args.operator,
detail={"answered_via": via},
)
conn = connect(args.db)
try:
changed = answer_question(
conn,
question_id=args.question_id,
answer_json=args.answer,
answered_via=via,
)
finally:
conn.close()
_audit_outcome(
args.audit_log,
"answer",
question_id=args.question_id,
operator=args.operator,
applied=changed,
detail={"answered_via": via},
)
if not changed:
print(
f"answer no-op: question {args.question_id} was not 'open' "
"(already answered/expired/superseded or absent)",
file=sys.stderr,
)
return 1
print(f"answered question {args.question_id} (via {via})", file=out)
return 0
def _cmd_supersede(args: argparse.Namespace, *, out: Any) -> int:
"""Mark a stale ``open``/``answered`` question ``superseded`` (destructive)."""
_require_confirm("supersede", confirm=args.confirm)
_audit_attempt(
args.audit_log,
"supersede",
question_id=args.question_id,
operator=args.operator,
)
conn = connect(args.db)
try:
changed = supersede_question(conn, question_id=args.question_id)
finally:
conn.close()
_audit_outcome(
args.audit_log,
"supersede",
question_id=args.question_id,
operator=args.operator,
applied=changed,
)
if not changed:
print(
f"supersede no-op: question {args.question_id} was not "
"'open'/'answered' (already expired/superseded or absent)",
file=sys.stderr,
)
return 1
print(f"superseded question {args.question_id}", file=out)
return 0
def _build_context_provider() -> "Callable[[], str]":
"""Return the WS5 (D10) context provider: the Sea Haven handbook conventions.
The ``context_provider`` seam is zero-arg (``Callable[[], str]``), so it
supplies STATIC context — the handbook conventions loaded from
``SEA_HAVEN_HANDBOOK_DIR`` (or the default dir) via
:func:`agent_team.nodes.handbook.load_handbook_conventions`. That loader is
itself fail-safe (caps file count/size; returns ``""`` and never raises when
the dir is absent/empty/unreadable), so wiring it is safe even on a box where
the handbook has not been synced yet.
(Task-keyed memory retrieval is NOT wired here: the seam takes no task text,
so it cannot form a retrieval query — that would need a task-aware seam.)
Imported lazily for the same import-hygiene reason as the coordinator
factories (keeps ``--help`` / ledger commands import-clean).
"""
from agent_team.nodes.handbook import load_handbook_conventions
return load_handbook_conventions
def _build_notifiers(
args: argparse.Namespace,
) -> "tuple[Callable[[str], None] | None, Callable[[str], None] | None]":
"""Build the (notify, alarm_hook) Slack notifiers for the coordinator.
Returns ``(None, None)`` for dry-run / non-Slack / no-channel so import,
``--help``, ledger commands, and token-less dry runs stay silent and need no
Slack credentials. For live Slack with ``SLACK_CHANNEL_ID`` set, ``notify``
posts a plain status line to the channel (via the same ``build_slack_poster``
the transport uses), and ``alarm_hook`` logs the deadline-park WARNING AND
posts a parked-task notice. Both are best-effort — the coordinator wraps the
notify sink so a Slack failure never disturbs the pipeline.
"""
if getattr(args, "dry_run", False) or args.transport != "slack":
return None, None
channel = os.environ.get("SLACK_CHANNEL_ID", "")
if not channel:
return None, None
try:
from agent_team.transport.slack_live import build_slack_poster
poster = build_slack_poster()
except Exception: # noqa: BLE001 - no token / SDK -> run without notifications
_LOG.warning(
"Slack notifier unavailable; coordinator runs without notifications"
)
return None, None
def notify(message: str) -> None:
poster({"channel": channel, "text": message})
def alarm_hook(question_id: str) -> None:
_LOG.warning("park ALARM: clarifier question %s expired", question_id)
try:
notify(
f"⚠️ Task parked: clarifier question {question_id[:8]} expired with "
"no answer in the window. Re-assign or answer to resume."
)
except Exception: # noqa: BLE001 - notify failure must not break the park path
pass
return notify, alarm_hook
def _build_coordinator(args: argparse.Namespace) -> Any:
"""Construct a :class:`Coordinator` for the ``start`` / ``serve`` commands.
The live transport is built LAZILY here (never at import) so ``run-team.py``
imports, ``--help``, and the ledger subcommands all work with no Slack token
present. ``--dry-run`` substitutes a non-posting transport so an operator can
drive intake without a token (the ledger row is still written; only the
transport post is a no-op).
The coordinator is imported inside this function for the same reason: pulling
in the live runtime (and its optional SDK-adjacent deps) must not happen just
to render ``--help`` or run a read-only ledger command.
"""
from agent_team.coordinator import (
Coordinator,
default_clarify_node_factory,
default_plan_node_factory,
default_review_wiring,
)
transport = _build_transport(args)
# WS5 (D10): inject the Sea Haven handbook conventions into BOTH the clarifier
# and the planner prompts via the context_provider seam. The provider is the
# zero-arg handbook loader, which is itself fail-safe (returns "" when the
# handbook dir is absent), and the clarifier/planner additionally swallow
# provider errors — so this never affects a run where the handbook is
# unavailable.
context_provider = _build_context_provider()
# Lifecycle notifications: post plain status lines to the Slack channel so a
# task is never a black box (parked / needs-more-input / plan-ready, and
# deadline-park ALARMs). Live-Slack only; dry-run / non-Slack / no-channel =
# silent (notify None) so import + ledger commands need no token.
notify, alarm_hook = _build_notifiers(args)
# Production runs the full P2 graph: the wrapped real planner + the bound
# GPT-4.1 review loop (Plane-2 depth-first). These factories are lazy and
# only build/bind the model seams when a task actually runs.
coordinator = Coordinator(
db_path=args.db,
transport=transport,
build_clarify_node=lambda: default_clarify_node_factory(
context_provider=context_provider
),
build_plan_node=lambda: default_plan_node_factory(
context_provider=context_provider
),
review_wiring=default_review_wiring,
notify=notify,
alarm_hook=alarm_hook,
)
# WS2: an allowlisted Slack /new-task starts a task on THIS coordinator. Set
# post-construction (the adapter closes over the just-built coordinator), and
# before serve() builds the listener. AUTHZ-01 (owner allowlist) gates this
# upstream in the listener; the source label is always "slack".
coordinator.set_new_task_callback(
lambda task_text, _source: coordinator.start_task(
task_text=task_text, transport_name="slack"
)
)
return coordinator
def _build_transport(args: argparse.Namespace) -> Any:
"""Build the transport for a coordinator command (lazy; token-tolerant).
``--dry-run`` (or any transport in dry-run) yields a non-posting transport so
intake works without credentials. Otherwise the live transport is built
lazily from the environment for the chosen ``--transport`` (so import,
``--help``, and ledger commands never need a token):
* ``slack`` -> :func:`build_live_slack_transport` over ``SLACK_BOT_TOKEN`` /
``SLACK_CHANNEL_ID``.
* ``github`` -> :func:`build_live_github_transport` over ``GITHUB_TOKEN`` and
the issue thread ``GITHUB_OWNER`` / ``GITHUB_REPO`` /
``GITHUB_ISSUE_NUMBER``. This is the §3.3.1 human-gate I/O only (post an
issue/PR comment); it carries NO CI, OIDC, git/patch apply, or GitHub
Actions network (the P3 apply/verify workflow is held for the security
gate).
* ``claude_code`` -> :func:`build_live_claude_code_transport` over the local
file-drop directory ``CLAUDE_CODE_DROP_DIR`` the Mac harness polls (D10).
Each live builder defers its optional SDK / token resolution to call time, so
a missing dependency or credential fails loudly here rather than at import.
"""
if getattr(args, "dry_run", False):
return _DryRunTransport()
if args.transport == "slack":
from agent_team.transport.slack_live import build_live_slack_transport
channel = os.environ.get("SLACK_CHANNEL_ID", "")
return build_live_slack_transport(channel)
if args.transport == "github":
from agent_team.transport.github_live import build_live_github_transport
owner, repo, issue_number = _github_thread_from_env()
# Token resolves from GITHUB_TOKEN inside the builder (call-time read);
# a missing token fails loudly there rather than being captured here.
return build_live_github_transport(
owner=owner,
repo=repo,
issue_number=issue_number,
token=os.environ.get("GITHUB_TOKEN") or None,
)
if args.transport == "claude_code":
from agent_team.transport.claude_code_live import (
build_live_claude_code_transport,
)
drop_dir = os.environ.get("CLAUDE_CODE_DROP_DIR", "")
if not drop_dir:
raise SystemExit(
"live transport 'claude_code' requires CLAUDE_CODE_DROP_DIR (the "
"Mac file-drop directory the Claude-Code harness polls); set it, "
"or use --dry-run for a no-token dry run"
)
return build_live_claude_code_transport(drop_dir)
raise SystemExit(
f"live transport '{args.transport}' is not wired for the run-team CLI; "
"use --transport slack/github/claude_code, or --dry-run for a no-token "
"dry run"
)
def _github_thread_from_env() -> tuple[str, str, int]:
"""Resolve the GitHub issue thread (owner/repo/issue) from the environment.
The live GitHub transport posts the clarifier question-set as a comment on a
fixed ``owner/repo#issue_number`` thread, so the thread is configured via
``GITHUB_OWNER`` / ``GITHUB_REPO`` / ``GITHUB_ISSUE_NUMBER`` (read lazily so
the value is never captured at import). A missing or non-integer value raises
a clear :class:`SystemExit` rather than building a half-configured transport.
"""
owner = os.environ.get("GITHUB_OWNER", "")
repo = os.environ.get("GITHUB_REPO", "")
raw_issue = os.environ.get("GITHUB_ISSUE_NUMBER", "")
missing = [
name
for name, value in (
("GITHUB_OWNER", owner),
("GITHUB_REPO", repo),
("GITHUB_ISSUE_NUMBER", raw_issue),
)
if not value
]
if missing:
raise SystemExit(
"live transport 'github' requires "
f"{', '.join(missing)}; set the issue thread (owner/repo/issue), or "
"use --dry-run for a no-token dry run"
)
try:
issue_number = int(raw_issue)
except ValueError:
raise SystemExit(
f"GITHUB_ISSUE_NUMBER must be an integer, got {raw_issue!r}"
) from None
return owner, repo, issue_number
class _DryRunTransport(Transport):
"""A non-posting transport for ``--dry-run`` intake (no token, no Slack).
``post_question`` records nothing on a real channel — it returns a synthetic
``channel_ref`` so :func:`agent_team.responder.notify_question` still writes
and stamps the durable ledger row (the durable seam is exercised; only the
side-effecting post is skipped). ``parse_answer`` is unused by the CLI path
but implemented so the ABC is concrete.
"""
def post_question(
self,
*,
thread_id: str,
question_id: str,
turn: int,
question_set: Any,
deadline: str,
) -> str:
print(
f"[dry-run] would post question {question_id} (turn {turn}) "
f"for thread {thread_id}",
file=sys.stderr,
)
return f"dry-run:{question_id}"
def parse_answer(self, raw: Any) -> tuple[str, Any, str]:
raise NotImplementedError("dry-run transport does not parse answers")
def _cmd_start(args: argparse.Namespace, *, out: Any) -> int:
"""Intake: start one task and run it to the first human gate (§3.3, §3.3.1).
Builds a :class:`Coordinator` (transport from the lazy factory; ``--dry-run``
posts nowhere), runs ``setup`` + ``start_task``, and prints the minted
``thread_id``. The clarifier question-set is delivered over the chosen
transport (or no-op under ``--dry-run``); the durable ledger row is written
either way.
"""
coordinator = _build_coordinator(args)
# start_task runs the clarifier graph to the first human gate IN THIS PROCESS,
# and the clarifier calls Claude (assess_confidence). The invoker is a
# process-local binding, so a running daemon does not help this CLI process —
# bind the real subscription invoker here, mirroring Coordinator.serve(), or
# start_task fails with "claude_invoke has no invoker bound".
from agent_team.invoker import bind_subscription_invoker
bind_subscription_invoker()
coordinator.setup()
thread_id = coordinator.start_task(
task_text=args.task, transport_name=args.transport
)
print(thread_id, file=out)
return 0
def _cmd_serve(args: argparse.Namespace, *, out: Any) -> int:
"""Run the coordinator daemon loop (binds the live invoker; §7.1 P1).
Delegates to :meth:`agent_team.coordinator.Coordinator.serve`, which binds
the real Claude subscription invoker, runs the startup recovery sweep, then
loops on the deadline cadence. The Slack inbound feed is the slack_listener's
job; this command owns the maintenance loop. Runs until interrupted.
"""
coordinator = _build_coordinator(args)
# Bind the in-process non-Claude invokers (GPT-4.1 review, DeepSeek build)
# alongside the Claude subscription invoker (WS1).
from agent_team.invoker_multi import bind_multi_invoker # noqa: PLC0415
bind_multi_invoker()
print("agent-team coordinator starting (Ctrl-C to stop)", file=out)
coordinator.serve()
return 0 # pragma: no cover - serve() loops until interrupted
def _cmd_intake_github(args: argparse.Namespace, *, out: Any) -> int:
"""Poll a repo for labeled issues and start one task per new issue (§3.3.1).
The GitHub-issue intake front door: builds a :class:`Coordinator` (transport
from the lazy factory; ``--dry-run`` posts nowhere), runs ``setup``, then
constructs a :class:`agent_team.transport.github_intake.GithubIntake` over a
read-only REST issue client
(:func:`~agent_team.transport.github_intake.build_default_issue_client`,
deferred-import, reads ``GITHUB_TOKEN`` at call time) and runs ONE
:meth:`~agent_team.transport.github_intake.GithubIntake.poll_once`. An
operator (or a cron) re-runs the command on a cadence; de-dup is DURABLE
(the ``ingested_issues`` ledger table), so re-running over a still-open
labeled issue never starts a second task — safe for a scheduled timer.
OPT-IN and INERT: this only reads labeled issues and calls the committed
coordinator intake entry — no CI, OIDC, git/patch apply, or GitHub-Actions
network. Owner/repo/label come from CLI flags; the token comes from
``GITHUB_TOKEN``. Prints the issue ids ingested on this pass (one per line).
"""
from agent_team.transport.github_intake import (
GithubIntake,
build_default_issue_client,
build_ledger_ingest_store,
)
coordinator = _build_coordinator(args)
coordinator.setup()
client = build_default_issue_client(owner=args.owner, repo=args.repo)
# Durable de-dup: a scheduled/cron intake runs as a fresh process each time,
# so the in-memory default would re-ingest every still-open labeled issue and
# spawn duplicate tasks. The ledger store keys on (source, issue_id) in the
# same SQLite DB the coordinator uses, making intake idempotent across runs.
store = build_ledger_ingest_store(
db_path=args.db, source=f"github:{args.owner}/{args.repo}"
)
intake = GithubIntake(
client=client,
coordinator=coordinator,
label=args.label,
store=store,
)
ingested = intake.poll_once()
for issue_id in ingested:
print(issue_id, file=out)
if not ingested:
print("github-intake: no new labeled issues to ingest", file=sys.stderr)
return 0
def _cmd_intake_checker(args: argparse.Namespace, *, out: Any) -> int:
"""Read checker reports and start one task per confirmed eligible finding.
The Plane-1 -> Plane-2 cross-plane front door (P5): builds a
:class:`Coordinator` (transport from the lazy factory; ``--dry-run`` posts
nowhere), runs ``setup``, then constructs a
:class:`agent_team.transport.checker_intake.CheckerFindingIntake` and runs
ONE :meth:`~agent_team.transport.checker_intake.CheckerFindingIntake.ingest_reports`
over the report paths. Each ``--report`` may be a checker report JSON file or
a directory of ``*.json`` reports.
OPT-IN and INERT: this only reads local report JSON and calls the committed
coordinator intake entry — no CI, OIDC, git/patch apply, or network. The
severity threshold (default ``high``) and transport come from CLI flags.
De-dup is in-memory per process, so each run is a single pass. Prints the
finding identities ingested on this pass (one per line).
"""
from agent_team.transport.checker_intake import CheckerFindingIntake
coordinator = _build_coordinator(args)
coordinator.setup()
intake = CheckerFindingIntake(
coordinator=coordinator,
threshold=args.threshold,
transport_name=args.transport,
)
ingested = intake.ingest_reports([Path(p) for p in args.report])
for identity in ingested:
print(identity, file=out)
if not ingested:
print(
"checker-intake: no new confirmed at/above-threshold findings to ingest",
file=sys.stderr,
)
return 0
def _load_finding_for_fix(args: argparse.Namespace) -> dict[str, Any]:
"""Load the single confirmed dependency-cve finding to fix from a report file.
Reads the ``dependency-cve.json`` report (the checker's output shape) at
``--report`` and selects the finding by ``--finding-id``. Returns the finding
mapping. Raises :class:`SystemExit` on any load/selection error so the CLI
fails loudly rather than dispatching against a half-resolved finding. Pure
read; no network, no CI, no git.
"""
report_path = Path(args.report)
try:
raw = report_path.read_text(encoding="utf-8")
except OSError as exc:
raise SystemExit(f"cannot read finding report {report_path}: {exc}") from exc
try:
report = json.loads(raw)
except ValueError as exc:
raise SystemExit(
f"finding report {report_path} is not valid JSON: {exc}"
) from exc
findings = report.get("findings") if isinstance(report, dict) else None
if not isinstance(findings, list):
raise SystemExit(
f"finding report {report_path} has no 'findings' array (got "
f"{type(report).__name__})"
)
matches = [
f for f in findings if isinstance(f, dict) and f.get("id") == args.finding_id
]
if not matches:
raise SystemExit(f"no finding with id {args.finding_id!r} in {report_path}")
if len(matches) > 1:
raise SystemExit(
f"ambiguous: {len(matches)} findings share id {args.finding_id!r} in "
f"{report_path}"
)
return matches[0]
def _cmd_fix(args: argparse.Namespace, *, out: Any) -> int:
"""Plan a Plane-1 dep-bump fix and show what it WOULD dispatch (§7 Phase 5).
The Tier-3 fixer front door (design §4 fixer row, §3.3.2). Loads ONE confirmed
``dependency-cve`` finding from the report, asks the fixer to produce a fix
spec (Claude) + a minimal bump patch (DeepSeek via the orchestrator) + the CI
``workflow_dispatch`` inputs, and — in ``--dry-run`` (the only mode wired
here) — PRINTS the spec, patch, and dispatch inputs without dispatching
anything.
OPT-IN / INERT: this command never dispatches. It binds NO workflow dispatcher
(the box holds no write token, D2), so even an ``ok`` plan only prints. Live
dispatch is a provisioning-time wiring of the trusted apply path's dispatcher,
deliberately not reachable from this CLI. A non-fixable finding prints the
fail-safe reason and exits non-zero.
Returns ``0`` when a fix plan was produced (dry-run printed), ``1`` when the
finding is not fixable (fail-safe; nothing planned).
"""
from agent_team.nodes.fixer import describe_plan, plan_fix
if not args.dry_run:
# Live dispatch is provisioning-gated and not wired into the CLI; refuse
# to run without --dry-run rather than silently doing nothing.
raise SystemExit(
"fix supports only --dry-run in this build (live dispatch is "
"provisioning-gated; the box holds no write token, D2). Re-run with "
"--dry-run to see what it WOULD dispatch."
)
finding = _load_finding_for_fix(args)
plan = plan_fix(finding, task_id=args.task_id)
print(describe_plan(plan), file=out)
return 0 if plan.ok else 1
def _cmd_force_resume(args: argparse.Namespace, *, out: Any) -> int:
"""Force-resume a parked task's question (destructive; audit-logged).
The design-named operator verb (§3.3.1 / §6.6 "an operator can force-resume
... a parked task via the CLI"). A task parks when its clarifier question
EXPIRES with no answer, so the un-park action is to RE-OPEN that expired
question (:func:`agent_team.db.schema.reopen_question`) so the normal
delivery → answer → resume flow can proceed.
Crucially this does NOT ``supersede`` the row: superseding an ``answered``
row would flip it out of the state the recovery sweep resumes from, making a
stuck-but-answered task permanently un-resumable — the opposite of
force-resume. So:
* ``expired`` (the parked case) → reopened; returns 0.
* ``answered`` (answered but not yet resumed) → already eligible for the
recovery resume sweep; intent is recorded and we report that, no mutation.
* ``open`` / ``superseded`` / absent → nothing to force; reported as a no-op.
"""
_require_confirm("force-resume", confirm=args.confirm)
_audit_attempt(
args.audit_log,
"force-resume",
question_id=args.question_id,
operator=args.operator,
detail={"resume_requested": True},
)
conn = connect(args.db)
try:
row = _fetch_question(conn, args.question_id)
status = None if row is None else row["status"]
reopened = False
if status == "expired":
reopened = reopen_question(conn, question_id=args.question_id)
finally:
conn.close()
_audit_outcome(
args.audit_log,
"force-resume",
question_id=args.question_id,
operator=args.operator,
applied=reopened,
detail={"resume_requested": True, "prior_status": status},
)
if reopened:
print(
f"force-resume: reopened expired question {args.question_id}; "
"it will be re-delivered for an answer",
file=out,
)
return 0
if status == "answered":
print(
f"force-resume: question {args.question_id} is answered and pending "
"resume; the recovery sweep will resume it (intent recorded)",
file=out,
)
return 0
print(
f"force-resume no-op: question {args.question_id} "
f"({'absent' if status is None else f'status={status}'}) is not parked",
file=sys.stderr,
)
return 1
# Statuses an operator treats as "parked context": a task whose only pending
# question is no longer open may be parked (answered-but-unresumed, expired, or
# superseded). ``open`` is excluded — that is the live-waiting view (default
# ``list``). Derived from the foundation QUESTION_STATES so it stays in sync.
_PARKED_STATES: tuple[str, ...] = tuple(s for s in QUESTION_STATES if s != "open")
def build_parser() -> argparse.ArgumentParser:
"""Construct the argparse parser for ``run-team.py`` (no side effects)."""
parser = argparse.ArgumentParser(
prog="run-team.py",
description=(
"R720 agent-team operator CLI — manual path over the durable "
"pending_questions ledger (design §3.3.1)."
),
)
parser.add_argument(
"--db",
type=Path,
default=_DEFAULT_DB,
help=f"path to the agent-team SQLite ledger (default: {_DEFAULT_DB})",
)
parser.add_argument(
"--audit-log",
type=Path,
default=_DEFAULT_AUDIT_LOG,
dest="audit_log",
help=(
"append-only JSONL audit log for destructive actions "
f"(default: {_DEFAULT_AUDIT_LOG})"
),
)
parser.add_argument(
"--operator",
default=_default_operator(),
help="operator identity recorded in the audit log for destructive actions "
"(defaults to the OS login so the trail is always attributable)",
)
sub = parser.add_subparsers(dest="command", required=True)
p_init = sub.add_parser("init-db", help="create/upgrade the ledger tables")
p_init.set_defaults(func=_cmd_init_db)
p_list = sub.add_parser("list", help="list pending questions (read-only)")
list_filter = p_list.add_mutually_exclusive_group()
list_filter.add_argument(
"--status",
choices=QUESTION_STATES,
default="open",
help="lifecycle status to list (default: open)",
)
list_filter.add_argument(
"--all",
action="store_true",
help="list questions in every lifecycle status",
)
list_filter.add_argument(
"--parked",
action="store_true",
help="list non-open questions (parked-task context)",
)
p_list.set_defaults(func=_cmd_list)
p_show = sub.add_parser("show", help="print one question row (read-only)")
p_show.add_argument("question_id", help="the question_id to show")
p_show.set_defaults(func=_cmd_show)
p_redeliver = sub.add_parser(
"redeliver",
help="clear an open question's channel_ref so it is re-posted",
)
p_redeliver.add_argument("question_id", help="the question_id to re-deliver")
p_redeliver.set_defaults(func=_cmd_redeliver)
p_expire = sub.add_parser(
"expire", help="force-expire an open question (destructive)"
)
p_expire.add_argument("question_id", help="the question_id to expire")
p_expire.add_argument(
"--confirm",
action="store_true",
help="required: confirm this destructive, audit-logged action",
)
p_expire.set_defaults(func=_cmd_expire)
p_answer = sub.add_parser(
"answer", help="answer a question on a task's behalf (destructive)"
)
p_answer.add_argument("question_id", help="the question_id to answer")
p_answer.add_argument(
"--answer",
required=True,
help="the answer payload (stored verbatim as answer_json)",
)
p_answer.add_argument(
"--via",
default="",
help="answering identity for answered_via (default: cli:<operator>)",
)
p_answer.add_argument(
"--confirm",
action="store_true",
help="required: confirm this destructive, audit-logged action",
)
p_answer.set_defaults(func=_cmd_answer)
p_supersede = sub.add_parser(
"supersede", help="mark a stale question superseded (destructive)"
)
p_supersede.add_argument("question_id", help="the question_id to supersede")
p_supersede.add_argument(
"--confirm",
action="store_true",
help="required: confirm this destructive, audit-logged action",
)
p_supersede.set_defaults(func=_cmd_supersede)
p_resume = sub.add_parser(
"force-resume",
help="force-resume a parked task's question (destructive)",
)
p_resume.add_argument("question_id", help="the question_id to force-resume")
p_resume.add_argument(
"--confirm",
action="store_true",
help="required: confirm this destructive, audit-logged action",
)
p_resume.set_defaults(func=_cmd_force_resume)
p_start = sub.add_parser(
"start",
help="intake: start one task and run it to the first human gate",
)
p_start.add_argument(
"--task",
required=True,
help="the task description (intake text) to run through the pipeline",
)
p_start.add_argument(
"--transport",
choices=_TRANSPORT_CHOICES,
default="slack",
help="channel for delivering clarifier questions (default: slack)",
)
p_start.add_argument(
"--dry-run",
action="store_true",
dest="dry_run",
help="use a non-posting transport (no token needed; ledger still written)",
)
p_start.set_defaults(func=_cmd_start)
p_serve = sub.add_parser(
"serve",
help="run the coordinator daemon (binds invoker, runs the loop)",
)
p_serve.add_argument(
"--transport",
choices=_TRANSPORT_CHOICES,
default="slack",
help="channel for delivering clarifier questions (default: slack)",
)
p_serve.add_argument(
"--dry-run",
action="store_true",
dest="dry_run",
help="use a non-posting transport (no token needed)",
)
p_serve.set_defaults(func=_cmd_serve)
p_intake = sub.add_parser(
"intake-github",
help="poll a repo for labeled issues and start one task per new issue",
)
p_intake.add_argument(
"--owner",
required=True,
help="GitHub repository owner / org login to poll for intake issues",
)
p_intake.add_argument(
"--repo",
required=True,
help="GitHub repository name to poll for intake issues",
)
p_intake.add_argument(
"--label",
required=True,
help="issue label that flags an issue as pipeline intake (non-empty)",
)
p_intake.add_argument(
"--transport",
choices=_TRANSPORT_CHOICES,
default="github",
help="channel for delivering clarifier questions (default: github)",
)
p_intake.add_argument(
"--dry-run",
action="store_true",
dest="dry_run",
help="use a non-posting transport (no token needed; ingest still runs)",
)
p_intake.set_defaults(func=_cmd_intake_github)
p_intake_checker = sub.add_parser(
"intake-checker",
help=(
"read checker reports and start one remediation task per confirmed "
"at/above-threshold finding (P5 cross-plane loop)"
),
)
p_intake_checker.add_argument(
"--report",
required=True,
action="append",
metavar="PATH",
help=(
"checker report JSON file or directory of *.json reports; repeatable "
"to ingest several reports in one pass"
),
)
p_intake_checker.add_argument(
"--threshold",
default="high",
choices=("low", "medium", "high", "critical"),
help=(
"minimum severity a confirmed finding must reach to spawn a task "
"(default: high)"
),
)
p_intake_checker.add_argument(
"--transport",
choices=_TRANSPORT_CHOICES,
default="github",
help="channel for delivering clarifier questions (default: github)",
)
p_intake_checker.add_argument(
"--dry-run",
action="store_true",
dest="dry_run",
help="use a non-posting transport (no token needed; ingest still runs)",
)
p_intake_checker.set_defaults(func=_cmd_intake_checker)
p_fix = sub.add_parser(
"fix",
help=(
"plan a Plane-1 dep-bump fix for a confirmed dependency-cve finding "
"and (dry-run) show what it WOULD dispatch to org CI"
),
)
p_fix.add_argument(
"--report",
required=True,
help="path to the dependency-cve.json finding report to read the finding from",
)
p_fix.add_argument(
"--finding-id",
required=True,
dest="finding_id",
help="the finding id (report findings[].id) to fix",
)
p_fix.add_argument(
"--task-id",
required=True,
dest="task_id",
help="pipeline task id (provenance; becomes the CI dispatch task_id)",
)
p_fix.add_argument(
"--dry-run",
action="store_true",
dest="dry_run",
help=(
"show the spec + patch + dispatch inputs WITHOUT dispatching "
"(the only supported mode; live dispatch is provisioning-gated)"
),
)
p_fix.set_defaults(func=_cmd_fix)
return parser
def main(argv: Sequence[str] | None = None, *, out: Any = None) -> int:
"""CLI entry point. Returns a process exit code.
``argv`` defaults to ``sys.argv[1:]``; ``out`` defaults to ``sys.stdout``
(injectable for tests). Operational failures return ``1``; a missing
``--confirm`` on a destructive action raises :class:`PermissionError`,
surfaced as exit code ``1`` with a stderr message.
"""
out = out if out is not None else sys.stdout
parser = build_parser()
args = parser.parse_args(argv)
try:
return int(args.func(args, out=out))
except PermissionError as exc:
# A refused destructive action (no --confirm) or an unwritable audit
# path. The attempt-before-mutate ordering means nothing was mutated.
print(f"error: {exc}", file=sys.stderr)
return 1
except OSError as exc:
# Any other audit-log / filesystem failure (e.g. the audit append could
# not be written). Surfaced cleanly instead of as an uncaught traceback;
# if the attempt record was written, the action is on the trail.
print(f"error: audit/IO failure: {exc}", file=sys.stderr)
return 1
if __name__ == "__main__": # pragma: no cover
raise SystemExit(main())