diff --git a/.github/workflows/agent-team-apply-verify.yml b/.github/workflows/agent-team-apply-verify.yml index 077b9d9..9caaea0 100644 --- a/.github/workflows/agent-team-apply-verify.yml +++ b/.github/workflows/agent-team-apply-verify.yml @@ -55,6 +55,18 @@ name: agent-team-apply-verify +# Run name surfaces the dispatching task's thread_id so the box-side dispatcher +# can correlate the triggered run back to its task via `gh run list --json name` +# (workflow inputs are NOT queryable; a workflow_dispatch run reports against the +# `main` ref, not the head branch — so the task_id in the run name is the +# correlation key). The `concurrency` group below already guarantees ONE in-flight +# run per task_id, so this name + the dispatched-at watermark match the run +# unambiguously even under many simultaneous task dispatches. +# P3-BOX-INTEGRATION (feat/agent-team-p3-box-integration): additive run-name only; +# no privilege/permission/trigger change. Flagged for the C1 re-run of +# /sh-security-review + GPT-4.1 cross-review on this trust-boundary workflow. +run-name: "agent-team-apply ${{ inputs.task_id }}" + # Manual / API trigger only. The trusted, separate apply path (which owns the # GitHub App write token) invokes this with the candidate-diff # artifact + the ledger-recorded hash + the declared scope. There is NO diff --git a/agent-team/README.md b/agent-team/README.md index 59fac93..95729af 100644 --- a/agent-team/README.md +++ b/agent-team/README.md @@ -2,7 +2,7 @@ The durable, human-gated agentic SDLC pipeline for the R720 (`sh-secrev` VM), design: `../docs/r720-agent-team-design.md`. A task flows INTAKE → CLARIFY (the -human gate) → PLAN → REVIEW, and (opt-in, deploy-gated) → BUILD → VERIFY → draft +human gate) → PLAN → REVIEW, and (P3) → BUILD → DISPATCH → VERIFY → draft PR. Every stage is durable and resumable (LangGraph + a SQLite checkpointer); the human gate suspends on `interrupt()` and resumes on a real answer. @@ -11,20 +11,25 @@ INTAKE → CLARIFY (Claude, human gate) → PLAN (Claude) → REVIEW (GPT-4.1) ▲ │ └── loop-back ───┤ approve/escalate → END - (P3, opt-in + INERT until the CI gate clears): - approve → BUILD (DeepSeek) → VERIFY (ci_gate) → draft PR + (P3 box path — gated behind the C1 re-review): + approve → BUILD (DeepSeek) → DISPATCH (push branch, trigger CI, + capture run_id, suspend) → [CI-watcher resumes on terminal + conclusion] → VERIFY (ci_gate) → draft PR ``` -> **Status (2026-06-18).** P1 (human gate) + P2 (planner + adversarial review +> **Status (2026-06-23).** P1 (human gate) + P2 (planner + adversarial review > loop) + the live runtime (coordinator, Slack/GitHub/Claude-Code transports, -> intake) + P3-inert (build/verify subgraph, opt-in) + P4 (more transports + -> GitHub-issue intake) are **built, reviewed, and merged to main** (~795 tests). -> **Nothing is provisioned**: not rsync'd to the box, no live tokens, no -> systemd, no live CI. Production default runs **P2** (no builders). -> **Deploy-gated / not yet built:** the P3 *live* CI apply/verify + OIDC role -> (held behind `/sh-security-review` + the mandatory GPT-4.1 cross-review), and -> all provisioning. See the project memory `project_r720_agent_team` and §7 of -> the design for the phased rollout. +> intake) + P3 (build/dispatch/verify subgraph) + P4 (more transports + +> GitHub-issue intake) are **built, reviewed, and merged to main**. +> **The CI apply/verify workflow is LIVE + provisioned** (the `agent-apply` +> environment, the `AGENT_APPLY_APP_*` secrets, and dispatched runs all exist as +> of 2026-06-22). **The box-side integration that drives it** (run_id capture, +> the BUILD → DISPATCH → VERIFY reorder, async CI-watch, and the fail-safe bound +> `serve` default) is built on this branch but **gated behind the C1 re-review** +> — the `/sh-security-review` + mandatory GPT-4.1 cross-review of the new +> CI-trigger boundary — before it ships to the box. See the project memory +> `project_r720_agent_team` and §7 of the design (and `docs/P3-PHASE0-DESIGN.md`) +> for the phased rollout. ## Layout @@ -62,11 +67,17 @@ agent-team/ planner.py # plan (Claude) review_loop.py + review_loop_llm.py # adversarial review (GPT-4.1 via orchestrator) builders.py + builders_llm.py # candidate diff (DeepSeek) — INERT, proposes only - verifier.py + verifier_llm.py # ci_gate sole PASS authority; LLM = fix-proposer - build_verify_subgraph.py # P3 BUILD→VERIFY topology (opt-in) + verifier.py + verifier_llm.py # ci_gate sole PASS authority; LLM = fix-proposer; + # binds expected_run_id per-task from state["run_id"] + build_verify_subgraph.py # P3 BUILD→DISPATCH→VERIFY topology + the tick()-driven + # CI-watcher that resumes a suspended task on the + # dispatched run's terminal conclusion (or parks on timeout) handbook.py # WS5 load_handbook_conventions (handbook seam, # fail-safe → "" if dir missing); planner context - dispatch_invoker.py # WS3 auto-dispatch node — INERT (NOT wired live) + dispatch_invoker.py # P3 DISPATCH node — pushes the per-dispatch head branch, + # triggers CI (gh workflow run), and captures run_id + + # dispatched_at into state via a correlation-tagged poll + # of gh run list (fails closed on an unfound run) transport/ # one adapter contract + a live impl per channel base.py # Transport ABC + QuestionSet / NormalizedAnswer slack_adapter.py + slack_live.py + slack_listener.py # Block Kit + Socket Mode + /new-task @@ -75,7 +86,7 @@ agent-team/ web/ # React/Vite/TypeScript status-dashboard SPA (React Flow map, # task list, click-through task history); built to web/dist scripts/ # deploy-r720-ws-rollout.sh — attended WS0–WS5 UPDATE of the box - ci/ # §3.3.2 split-job CI apply/verify workflow (DEPLOY-GATED) + ci/ # §3.3.2 split-job CI apply/verify workflow (LIVE since 2026-06-22) systemd/ # agent-team-coordinator.service + agent-team-status.service DEPLOY-R720.md # provisioning runbook (snapshot-first, rsync, tokens, demo) tests/ # pytest, one module per source module + sim harness @@ -149,7 +160,7 @@ P1–P4 pipeline. What is **live** vs **inert** after the rollout: | WS1 | `invoker_multi.py` (in-process GPT-4.1 / DeepSeek / Gemini via the orchestrator's `models.py`) + `api.py` (FastAPI HTTP API, bearer auth via `AGENT_TEAM_API_TOKEN`, binds `127.0.0.1:8765`, `/docs`+`/openapi` disabled, concurrency-capped) | `bind_multi_invoker()` wired in `run-team.py` `_cmd_serve` (LIVE); the **HTTP API is a separate opt-in process** (`api.serve()`), NOT started by the coordinator | | WS5 | `nodes/handbook.py` `load_handbook_conventions` (reads `SEA_HAVEN_HANDBOOK_DIR` or `~/.sea-haven/engineering-handbook`, fail-safe → `""`); `retriever.py` `save_memory` writes to a `_box-drafts/` review queue | LIVE — the planner prompt receives the handbook via the `context_provider` seam in `run-team.py` `_build_coordinator` | | WS2/WS0/WS4 | Slack `/new-task` slash command (AUTHZ-01 owner-allowlist gated) → `Coordinator.set_new_task_callback`; the `sea-haven-claude-plugin/` (CLAUDE.md, settings, `/delegate` `UserPromptSubmit` hook) | LIVE (`/new-task` wired in `serve`); the plugin/HTTP-API path is opt-in | -| WS3 | `nodes/dispatch_invoker.py` (auto-dispatch LangGraph node) + graph/coordinator wiring | **INERT — NOT wired live.** The `agent-apply` GitHub Environment human-approval gate is KEPT; the P3 dispatch/build-verify path stays inert pending per-task `run_id` plumbing + a CI-boundary security re-review | +| WS3 / P3 | `nodes/dispatch_invoker.py` (DISPATCH LangGraph node) + the BUILD → DISPATCH → VERIFY reorder, the `tick()`-driven CI-watcher, per-task `run_id` plumbing, and the bound `serve` default | **Built on `feat/agent-team-p3-box-integration`, gated behind the C1 re-review** before it ships to the box. The `agent-apply` GitHub Environment human-approval gate is KEPT; dispatch is **operator-initiated** (the box holds no standing write token — the branch push + `gh workflow run` use operator-host credentials). The bound P3 wiring is the new fail-safe `serve` default: a missing `AGENT_TEAM_REPO_OWNER`/`_NAME` or CI-read token degrades to the INERT P3 path (task parks + a `#agent-team` notice), never a serve-start crash | The HTTP API endpoints: `POST /tasks` (start a task), `GET /tasks/{thread_id}` (status), `POST /orchestrator/invoke` (one-shot model invoke). See diff --git a/agent-team/agent_team/ci_fetcher.py b/agent-team/agent_team/ci_fetcher.py index b2bbdc7..049a0c5 100644 --- a/agent-team/agent_team/ci_fetcher.py +++ b/agent-team/agent_team/ci_fetcher.py @@ -101,10 +101,23 @@ _ALLOWED_CONCLUSIONS: frozenset[str] = frozenset( } ) -# The dedicated read-only token env var (preferred), falling back to the generic -# GITHUB_TOKEN (the idiom github_live uses). The token MUST be read-only — the -# fetcher only ever GETs; a write-scoped token here would be unnecessary blast -# radius (provisioning issues a read-only fine-grained PAT / read-only var). +# The dedicated read-only token env var (PREFERRED on the live box), falling back +# to the generic GITHUB_TOKEN (the idiom github_live uses). The token MUST be +# read-only — the fetcher only ever GETs; a write-scoped token here would be +# unnecessary blast radius (provisioning issues a read-only fine-grained PAT / +# read-only var). +# +# SEC-04 (CWE-269): the GITHUB_TOKEN fallback is a *documented, explicitly- +# narrowed convenience* for the github_live idiom — on the live box GITHUB_TOKEN +# MUST be read-only (contents:read). Provisioning issues the dedicated +# AGENT_TEAM_CI_READ_TOKEN for this read path and that name is the preferred +# source; GITHUB_TOKEN is only the fallback. Because the no-standing-write-token +# audit (scripts/assert_no_write_token.py) name-exempts GITHUB_TOKEN from the +# write-token *name* heuristic, it deliberately does NOT exempt GITHUB_TOKEN from +# the write-token *value* scan: a write-scoped PAT/installation token +# (ghp_/ghs_/...) parked in GITHUB_TOKEN on the box is still flagged there. This +# fetcher itself is unconditionally read-only (single GET, no write verbs), so +# even a mistakenly write-scoped token is never exercised for a write here. CI_READ_TOKEN_ENV = "AGENT_TEAM_CI_READ_TOKEN" _FALLBACK_TOKEN_ENV = "GITHUB_TOKEN" @@ -118,9 +131,13 @@ _DEFAULT_TIMEOUT_S = 15.0 def _resolve_read_token() -> str | None: """Resolve the read-only GitHub token at call time, or ``None``. - Prefers :data:`CI_READ_TOKEN_ENV`, falls back to ``GITHUB_TOKEN``. Returns - ``None`` when neither is set so the fetcher fails closed (the caller maps a - missing token to a ``None`` result -> gate BLOCK), rather than raising. + PREFERS the dedicated read-only :data:`CI_READ_TOKEN_ENV` + (``AGENT_TEAM_CI_READ_TOKEN``); falls back to ``GITHUB_TOKEN`` only as the + documented, explicitly-narrowed convenience described above — on the live box + ``GITHUB_TOKEN`` MUST be read-only (the no-standing-write-token audit + value-scans it, and this fetcher only ever issues a single read-only GET). + Returns ``None`` when neither is set so the fetcher fails closed (the caller + maps a missing token to a ``None`` result -> gate BLOCK), rather than raising. """ return os.environ.get(CI_READ_TOKEN_ENV) or os.environ.get(_FALLBACK_TOKEN_ENV) diff --git a/agent-team/agent_team/ci_gate.py b/agent-team/agent_team/ci_gate.py index d789be7..7774807 100644 --- a/agent-team/agent_team/ci_gate.py +++ b/agent-team/agent_team/ci_gate.py @@ -473,7 +473,7 @@ def evaluate_ci_gate( candidate_diff: str, ledger_hash: str | None, ci_result: Mapping[str, Any] | None, - expected_run_id: str, + expected_run_id: str | None, allowed_scope: Sequence[str] | None = None, ) -> GateResult: """Make the deterministic pass/fail/block decision (§3.3.2 boundary #4). @@ -486,26 +486,47 @@ def evaluate_ci_gate( read-only PAT (GitHub Checks/Actions API). The gate reads only ``run_id``, ``conclusion``, and (optionally) ``diff_hash`` from it; it **never** reads a patch-written success file/artifact. - * ``expected_run_id`` — the run id the verifier dispatched for this exact - diff; the conclusion must be keyed to it (a stale/substituted run id is a - BLOCK). + * ``expected_run_id`` — the run id THIS task dispatched the apply/verify + workflow under (per-task, sourced from ``state["run_id"]`` at node-run + time, NOT a static wiring-time constant). The conclusion must be keyed to + it (a stale/substituted run id is a BLOCK). A ``None``/empty value means + the dispatcher captured no run id to bind to (e.g. the run-id poll fell + through) — there is nothing to anchor the verdict against, so the gate + BLOCKs (refuse-to-proceed, never a vacuous pass). * ``allowed_scope`` — optional declared-scope prefixes for the task. Decision order (a trust violation always wins over a CI verdict): - 1. **Denylist / scope** — any violation -> :data:`GateDecision.BLOCK`. - 2. **Diff-hash integrity** — hash mismatch (ledger or CI-verified) -> + 1. **Missing binding** — a ``None``/empty ``expected_run_id`` -> + :data:`GateDecision.BLOCK` (no per-task run to bind the verdict to). + 2. **Denylist / scope** — any violation -> :data:`GateDecision.BLOCK`. + 3. **Diff-hash integrity** — hash mismatch (ledger or CI-verified) -> ``BLOCK``. - 3. **Authenticated conclusion** — missing result, a ``run_id`` that does not + 4. **Authenticated conclusion** — missing result, a ``run_id`` that does not match ``expected_run_id``, or an unrecognised/ambiguous conclusion -> ``BLOCK``; a recognised failure -> :data:`GateDecision.FAIL`; ``success`` -> :data:`GateDecision.PASS`. Returns a :class:`GateResult` with the decision and the reasons behind it. - Raises :class:`CiGateError` on structurally invalid inputs. + Raises :class:`CiGateError` on structurally invalid inputs (a non-string + ``expected_run_id`` is structurally invalid; ``None``/empty is a BLOCK, not + an exception, because it is the legitimate "no run captured" runtime state). """ - if not isinstance(expected_run_id, str) or not expected_run_id: - raise CiGateError("expected_run_id must be a non-empty string") + if expected_run_id is not None and not isinstance(expected_run_id, str): + raise CiGateError("expected_run_id must be a string or None") + + # (0) Missing per-task binding: the dispatcher captured no run id for this + # task (None/empty). There is nothing to anchor the verdict to, so refuse to + # proceed rather than gating against an empty string (which a substituted + # ``ci_result`` with no/empty run_id could otherwise vacuously satisfy). + if not expected_run_id: + return GateResult( + decision=GateDecision.BLOCK, + reasons=["no per-task run_id to bind the verdict to (dispatch unresolved)"], + run_id=None, + diff_hash=ledger_hash, + ci_conclusion=None, + ) reasons: list[str] = [] diff --git a/agent-team/agent_team/ci_watcher.py b/agent-team/agent_team/ci_watcher.py new file mode 100644 index 0000000..b6f809c --- /dev/null +++ b/agent-team/agent_team/ci_watcher.py @@ -0,0 +1,550 @@ +"""CI-watcher sweep — async resume-on-CI-complete for suspended VERIFY (design §3.3.2, P3 Decision 2). + +A CI apply/verify run takes ~7 minutes. The coordinator is a single durable +daemon: a multi-minute *blocking* VERIFY node would stall the ``tick()`` loop and +every other task. So the §4 Decision-2 shape is **async resume-on-CI-complete**, +NOT a blocking poll: + + BUILD → DISPATCH (push branch, trigger CI, capture run_id, suspend at VERIFY) + → [CI-watcher resumes on terminal conclusion] → VERIFY (read result, gate) + +This module is that CI-watcher: a ``tick()``-driven sweep that **mirrors +:mod:`agent_team.deadline_timer`'s shape** (a pure, restart-safe maintenance pass +whose side effects are injected callables, so it is unit-testable with no network +and no live graph). For each task suspended awaiting CI it: + +1. reads the task's ``run_id`` / ``dispatched_at`` (written ONLY by the trusted + dispatch node — never by an LLM/builder/verifier node, mirroring the + ``ci_fetcher`` TRUST SOURCE note); +2. polls that run **read-only** via an injected poller (the production default + reuses :func:`agent_team.ci_fetcher.fetch_ci_result` — a single read-only GET; + it NEVER writes, dispatches, or applies anything); and +3. acts on the poll outcome: + + * **TERMINAL** (the run concluded with a recognised conclusion) → RESUME the + suspended task. VERIFY then re-reads the now-terminal authenticated result + and the pure-code gate (:mod:`agent_team.ci_gate`) decides PASS/FAIL. The + watcher itself NEVER decides the verdict — it only un-suspends the task. + * **PENDING** (the run has no conclusion yet) → if ``dispatched_at + timeout`` + has elapsed, PARK the task (a CI run that never terminates must not wait + forever); otherwise leave it suspended for a later pass. + * **ERROR** (the read-only poll raised / could not produce a result) → PARK + the task. **Fail-closed:** a broken poll parks for a human rather than + spinning or fabricating progress. + +Fail-closed discipline (mirroring ``ci_fetcher`` / ``deadline_timer``): + +* A task whose state carries no usable ``run_id`` or ``dispatched_at`` is + **parked immediately** (a suspended-awaiting-CI task with no run to watch is + unrecoverable by polling — it can never resume on a conclusion that has no run + id to read). This is the "None result" fail-closed case. +* Any poll exception is isolated per task (caught → PARK that one task) so one + task's poll failure never aborts the whole sweep or crashes the daemon. +* All inputs are read fresh from the durable task state each pass; the watcher + holds no in-memory-only state, so a reboot mid-sweep simply re-runs the + remaining suspended tasks on the next pass. RESUME is idempotent via the + resume worker's turn guard; PARK is idempotent (a parked task is no longer + surfaced as awaiting-CI). + +Side effects (RESUME, PARK, ALARM) are injected as callables — exactly what makes +the terminal/timeout/error branches testable without a live graph, a network, or +a real coordinator. +""" + +from __future__ import annotations + +import logging +from collections.abc import Callable, Mapping +from dataclasses import dataclass, field +from datetime import datetime, timedelta, timezone +from enum import Enum +from typing import Any + +__all__ = [ + "AWAITING_CI_KEY", + "AlarmFn", + "CiPollOutcome", + "CiPollResult", + "CiWatchAction", + "CiWatchOutcome", + "CiWatchReport", + "ParkFn", + "PendingCiTask", + "PollFn", + "ResumeFn", + "awaiting_ci_run_id", + "default_ci_poller", + "run_ci_watcher", + "snapshot_awaiting_ci_run_id", +] + +logger = logging.getLogger(__name__) + +# The marker key the VERIFY node stamps on its ``interrupt()`` payload when it +# suspends awaiting a CI conclusion (see +# :func:`agent_team.nodes.build_verify_subgraph._await_ci`, whose payload is +# ``{"awaiting_ci": True, "run_id": }``). Both the durable pending-task +# enumerator (which threads are suspended-at-VERIFY?) and the resume turn-guard +# (is this thread STILL suspended awaiting THIS run?) key off this single marker +# so they recognise a suspended-at-VERIFY thread the exact same way. +AWAITING_CI_KEY = "awaiting_ci" + +# Conservative default: a CI apply/verify run that has not reached a terminal +# conclusion within this window is treated as stuck and the task parks (§4 +# Decision 2 "dispatched_at + timeout elapses -> park"). A normal run is ~7 min; +# 30 min leaves generous headroom for queueing / retries before we ALARM. +DEFAULT_CI_TIMEOUT = timedelta(minutes=30) + + +class CiPollOutcome(Enum): + """The classification of one read-only CI poll (the poller's verdict). + + ``TERMINAL`` — the run reached a recognised, authenticated conclusion (the + poll ``result`` mapping is populated); the task should RESUME so VERIFY can + gate it. ``PENDING`` — the run is still in progress (no conclusion yet); keep + waiting unless the timeout has elapsed. ``ERROR`` — the poll could not produce + a result (a fetch error / fail-closed ``None``); the task PARKS. + """ + + TERMINAL = "terminal" + PENDING = "pending" + ERROR = "error" + + +@dataclass(frozen=True) +class CiPollResult: + """The outcome of one read-only poll of a task's CI run. + + ``outcome`` is the :class:`CiPollOutcome`; ``result`` carries the + authenticated CI conclusion mapping (``run_id`` / ``conclusion`` / + ``diff_hash``) when (and only when) ``outcome`` is ``TERMINAL`` — the same + DATA shape :func:`agent_team.ci_gate.evaluate_ci_gate` consumes. The watcher + never derives a verdict from it; it only decides resume-vs-wait-vs-park. + """ + + outcome: CiPollOutcome + result: Mapping[str, Any] | None = None + + @classmethod + def terminal(cls, result: Mapping[str, Any]) -> CiPollResult: + """A run that reached a recognised terminal conclusion.""" + return cls(outcome=CiPollOutcome.TERMINAL, result=result) + + @classmethod + def pending(cls) -> CiPollResult: + """A run still in progress (no authenticated conclusion yet).""" + return cls(outcome=CiPollOutcome.PENDING) + + @classmethod + def error(cls) -> CiPollResult: + """A poll that could not produce a result (fail-closed).""" + return cls(outcome=CiPollOutcome.ERROR) + + +class CiWatchAction(Enum): + """The action the sweep actually took for one suspended-awaiting-CI task. + + ``RESUMED`` — the run terminated, the task was resumed so VERIFY can gate it. + ``PARKED_TIMEOUT`` — the run never terminated within the timeout; the task + parked. ``PARKED_ERROR`` — the poll failed / the state was unusable; the task + parked (fail-closed). ``WAITING`` — the run is still in progress and within + the timeout; the task stays suspended for a later pass. + """ + + RESUMED = "resumed" + PARKED_TIMEOUT = "parked_timeout" + PARKED_ERROR = "parked_error" + WAITING = "waiting" + + +@dataclass(frozen=True) +class PendingCiTask: + """A minimal read-snapshot of one task suspended awaiting CI (§3.3.2). + + Only the fields the watcher needs: identity (``thread_id``) and the trusted + dispatch watermarks (``run_id`` / ``dispatched_at``) the poll + timeout key + off. Frozen because it is a snapshot — the watcher never mutates a task + in-place; it acts through the injected resume/park callables. + """ + + thread_id: str + run_id: str | None = None + dispatched_at: str | None = None + + @classmethod + def from_state(cls, thread_id: str, state: Mapping[str, Any]) -> PendingCiTask: + """Build a :class:`PendingCiTask` from a durable task ``state`` mapping.""" + run_id = state.get("run_id") + dispatched_at = state.get("dispatched_at") + return cls( + thread_id=thread_id, + run_id=str(run_id) if run_id is not None and str(run_id) != "" else None, + dispatched_at=( + str(dispatched_at) + if dispatched_at is not None and str(dispatched_at) != "" + else None + ), + ) + + +@dataclass(frozen=True) +class CiWatchOutcome: + """The result of processing one suspended-awaiting-CI task this pass.""" + + thread_id: str + action: CiWatchAction + run_id: str | None = None + error: str | None = None + + +@dataclass +class CiWatchReport: + """Aggregate result of one CI-watcher pass (mirrors :class:`~agent_team.deadline_timer.TimerLoopReport`). + + ``outcomes`` is one entry per suspended task examined. The summary counters + let the coordinator decide whether to ALARM (any ``parked_error``) without + re-walking the list. + """ + + outcomes: list[CiWatchOutcome] = field(default_factory=list) + + @property + def examined(self) -> int: + """Number of suspended-awaiting-CI tasks examined this pass.""" + return len(self.outcomes) + + @property + def resumed(self) -> int: + """Tasks whose run terminated and were resumed.""" + return sum(1 for o in self.outcomes if o.action is CiWatchAction.RESUMED) + + @property + def parked_timeout(self) -> int: + """Tasks parked because their run never terminated within the timeout.""" + return sum(1 for o in self.outcomes if o.action is CiWatchAction.PARKED_TIMEOUT) + + @property + def parked_error(self) -> int: + """Tasks parked because the poll failed / the state was unusable.""" + return sum(1 for o in self.outcomes if o.action is CiWatchAction.PARKED_ERROR) + + @property + def waiting(self) -> int: + """Tasks still in progress and left suspended for a later pass.""" + return sum(1 for o in self.outcomes if o.action is CiWatchAction.WAITING) + + @property + def parked(self) -> int: + """Total tasks parked this pass (timeout + error).""" + return self.parked_timeout + self.parked_error + + +# Injected seams. Keeping these as callables means the watcher performs no +# transport / graph / network I/O of its own (testable, and faithful to the +# deadline_timer shape). +PollFn = Callable[[PendingCiTask], CiPollResult] +"""Poll one task's CI run read-only and classify it. The production default +(:func:`default_ci_poller`) reuses :func:`agent_team.ci_fetcher.fetch_ci_result` +(a single read-only GET; never a write).""" + +ResumeFn = Callable[[PendingCiTask, Mapping[str, Any]], None] +"""Called once per *terminated* task to RESUME it (drive the suspended VERIFY +node forward). Receives the task snapshot + the authenticated terminal result.""" + +ParkFn = Callable[[PendingCiTask], None] +"""Called once per task that must PARK (timeout, poll error, or unusable +state). Parks the task and (typically) raises an ALARM.""" + +AlarmFn = Callable[[str], None] +"""Optional ALARM hook for the coordinator (one message per parked task).""" + + +def _utc_now() -> datetime: + """Return the current UTC time (injectable via ``now`` in the loop).""" + return datetime.now(timezone.utc) + + +def _parse_iso(value: str | None) -> datetime | None: + """Parse an ISO-8601 timestamp to an aware UTC datetime, or ``None``. + + A missing / unparseable ``dispatched_at`` yields ``None`` so the caller fails + closed (treats the task as unusable → park) rather than crashing the sweep. + """ + if not value: + return None + try: + parsed = datetime.fromisoformat(value) + except (TypeError, ValueError): + return None + if parsed.tzinfo is None: + # The ledger always writes UTC; treat a naive stamp as UTC rather than + # raising on the aware/naive compare below. + parsed = parsed.replace(tzinfo=timezone.utc) + return parsed + + +def awaiting_ci_run_id(interrupt_value: Any) -> str | None: + """Return the run id from a VERIFY *awaiting-CI* interrupt payload, or ``None``. + + The VERIFY node suspends with + ``{"awaiting_ci": True, "run_id": }`` (see + :func:`agent_team.nodes.build_verify_subgraph._await_ci`). This recognises + THAT marker precisely: it returns the awaited ``run_id`` only when the + payload is a mapping with ``awaiting_ci`` truthy AND a non-empty string + ``run_id``. Any other interrupt (a human clarify gate, a malformed payload, + a payload with no run id) yields ``None`` so callers never mistake a + different suspension for a suspended-at-VERIFY-awaiting-CI thread. + """ + if not isinstance(interrupt_value, Mapping): + return None + if not interrupt_value.get(AWAITING_CI_KEY): + return None + run_id = interrupt_value.get("run_id") + if isinstance(run_id, str) and run_id: + return run_id + return None + + +def snapshot_awaiting_ci_run_id(snapshot: Any) -> str | None: + """Return the awaited run id if ``snapshot`` is suspended at VERIFY awaiting CI. + + Walks the snapshot's pending interrupts (``snapshot.interrupts``) and returns + the first ``run_id`` carried by an *awaiting-CI* payload (via + :func:`awaiting_ci_run_id`). Returns ``None`` when the snapshot is not + interrupted, is interrupted on a non-CI gate (e.g. a human clarify + question), or has already advanced (resumed / parked / done — no pending + interrupts). This is the single predicate both the durable enumerator and + the resume turn-guard use to decide "still suspended at VERIFY awaiting CI". + """ + interrupts = getattr(snapshot, "interrupts", None) or () + for item in interrupts: + value = getattr(item, "value", item) + run_id = awaiting_ci_run_id(value) + if run_id is not None: + return run_id + return None + + +def default_ci_poller( + *, + owner: str, + repo: str, + client: Any = None, +) -> PollFn: + """Build the production read-only CI poller (reuses :mod:`agent_team.ci_fetcher`). + + Returns a :data:`PollFn` that, given a :class:`PendingCiTask`, performs ONE + read-only GET of the task's run via + :func:`agent_team.ci_fetcher.fetch_ci_result` (closing over ``owner`` / ``repo`` + and an optional injected ``client`` test double) and classifies it: + + * a populated mapping (a recognised, authenticated terminal conclusion) → + :meth:`CiPollResult.terminal`; + * ``None`` (the run is still in progress — ``fetch_ci_result`` returns ``None`` + for a null/unrecognised conclusion) → :meth:`CiPollResult.pending`; the + watcher keeps waiting until the timeout, so an in-progress run never blocks + the daemon and never parks prematurely; + * any exception raised by the fetch → :meth:`CiPollResult.error` (fail-closed: + a broken read-only poll parks the task). + + The poller NEVER writes: ``fetch_ci_result`` issues exactly one read-only GET + and fails closed to ``None`` on every error path. It NEVER decides the + verdict — that stays with the pure-code gate after VERIFY resumes. + """ + from agent_team.ci_fetcher import fetch_ci_result + + def poll(task: PendingCiTask) -> CiPollResult: + # The fetcher reads ``state["run_id"]`` / ``state["diff_hash"]``; rebuild + # the minimal state it needs from the trusted dispatch watermarks. (We do + # not have the full task state here — only what the watcher snapshotted — + # which is exactly the read-only run identity the fetcher requires.) + state = {"run_id": task.run_id} + try: + result = fetch_ci_result(state, owner=owner, repo=repo, client=client) + except Exception: # noqa: BLE001 - any fetch failure fails closed -> park + logger.warning( + "ci_watcher: read-only poll for task %s raised; failing closed", + task.thread_id, + exc_info=True, + ) + return CiPollResult.error() + if isinstance(result, Mapping): + return CiPollResult.terminal(result) + # None: the run has no recognised conclusion yet (in progress). Keep + # waiting; the timeout branch parks a run that never terminates. + return CiPollResult.pending() + + return poll + + +def run_ci_watcher( + pending_tasks: list[PendingCiTask], + *, + poll: PollFn, + on_resume: ResumeFn, + on_park: ParkFn, + timeout: timedelta = DEFAULT_CI_TIMEOUT, + now: datetime | None = None, +) -> CiWatchReport: + """Run one CI-watcher pass over the suspended-awaiting-CI tasks (§3.3.2 Decision 2). + + ``pending_tasks`` is the set of tasks currently suspended at VERIFY awaiting + CI (the coordinator supplies them from the durable task store). For each: + + 1. **Unusable state → PARK (fail-closed).** A task with no ``run_id`` or no + parseable ``dispatched_at`` cannot be watched (there is no run to poll, or + no watermark to time out against), so it parks immediately rather than + waiting forever on a run it can never read. + 2. **Poll the run read-only.** Call ``poll`` (the production default reuses + :func:`agent_team.ci_fetcher.fetch_ci_result`; never a write). A poll that + raises is isolated to this one task (→ PARK) so it cannot abort the sweep. + 3. **Act on the outcome:** + * ``TERMINAL`` → ``on_resume`` (resume the task; VERIFY re-reads the + now-terminal result and the pure-code gate decides). The watcher never + decides the verdict. + * ``PENDING`` → if ``now >= dispatched_at + timeout`` → ``on_park`` + (timeout); else leave suspended (``WAITING``) for a later pass. + * ``ERROR`` → ``on_park`` (fail-closed). + + Side effects are isolated per task: if ``on_resume`` / ``on_park`` raises, the + failure is recorded for that task and the sweep continues with the rest of the + batch rather than aborting the whole pass. + + Restart-safety: inputs are the durable task snapshots, resume is idempotent + (the resume worker's turn guard), and park is idempotent, so re-running the + pass after a crash safely processes only the still-suspended tasks. + + Returns a :class:`CiWatchReport` describing what happened to each task. + """ + current = now or _utc_now() + report = CiWatchReport() + + for task in pending_tasks: + outcome = _process_task( + task, + poll=poll, + on_resume=on_resume, + on_park=on_park, + timeout=timeout, + now=current, + ) + report.outcomes.append(outcome) + + return report + + +def _process_task( + task: PendingCiTask, + *, + poll: PollFn, + on_resume: ResumeFn, + on_park: ParkFn, + timeout: timedelta, + now: datetime, +) -> CiWatchOutcome: + """Process one suspended task: poll its run, then resume / wait / park. + + Isolates each task's side effect: a poll OR a resume/park callback that raises + yields a :attr:`CiWatchAction.PARKED_ERROR` outcome (best-effort: the task is + parked if it can be) instead of crashing the whole sweep. + """ + # (1) Unusable state -> fail closed (park). No run to poll / no watermark to + # time out against means polling can never recover this task. + if task.run_id is None or _parse_iso(task.dispatched_at) is None: + logger.warning( + "ci_watcher: task %s has no usable run_id/dispatched_at; parking " + "(fail-closed)", + task.thread_id, + ) + return _park(task, on_park, CiWatchAction.PARKED_ERROR) + + # (2) Read-only poll. A raising poll is isolated to this task (-> park). + try: + poll_result = poll(task) + except Exception as exc: # noqa: BLE001 - one task's poll failure -> park it + logger.warning( + "ci_watcher: poll for task %s raised (%s); parking (fail-closed)", + task.thread_id, + type(exc).__name__, + ) + return _park( + task, + on_park, + CiWatchAction.PARKED_ERROR, + error=f"{type(exc).__name__}: {exc}", + ) + + # (3) Act on the classified outcome. + if poll_result.outcome is CiPollOutcome.TERMINAL: + result = poll_result.result or {} + try: + on_resume(task, result) + except Exception as exc: # noqa: BLE001 - isolate one task's resume failure + logger.exception( + "ci_watcher: resume of task %s (terminal run) raised", task.thread_id + ) + return CiWatchOutcome( + thread_id=task.thread_id, + action=CiWatchAction.PARKED_ERROR, + run_id=task.run_id, + error=f"{type(exc).__name__}: {exc}", + ) + return CiWatchOutcome( + thread_id=task.thread_id, + action=CiWatchAction.RESUMED, + run_id=task.run_id, + ) + + if poll_result.outcome is CiPollOutcome.ERROR: + # Fail-closed: a poll that could not produce a result parks the task. + return _park(task, on_park, CiWatchAction.PARKED_ERROR) + + # PENDING: still in progress. Park only if the dispatch timeout has elapsed; + # otherwise leave it suspended for a later pass (the async-wait, not a block). + dispatched = _parse_iso(task.dispatched_at) + assert dispatched is not None # narrowed by the unusable-state guard above + if now >= dispatched + timeout: + logger.warning( + "ci_watcher: task %s run %s did not terminate within %s; parking (timeout)", + task.thread_id, + task.run_id, + timeout, + ) + return _park(task, on_park, CiWatchAction.PARKED_TIMEOUT) + + return CiWatchOutcome( + thread_id=task.thread_id, + action=CiWatchAction.WAITING, + run_id=task.run_id, + ) + + +def _park( + task: PendingCiTask, + on_park: ParkFn, + action: CiWatchAction, + *, + error: str | None = None, +) -> CiWatchOutcome: + """Invoke the injected park side effect, isolating a callback failure. + + A ``on_park`` that raises is downgraded to a ``PARKED_ERROR`` outcome (the + sweep continues) rather than crashing the whole pass — the same per-row + side-effect isolation :mod:`agent_team.deadline_timer` uses. + """ + try: + on_park(task) + except Exception as exc: # noqa: BLE001 - isolate one task's park failure + logger.exception("ci_watcher: park of task %s raised", task.thread_id) + return CiWatchOutcome( + thread_id=task.thread_id, + action=CiWatchAction.PARKED_ERROR, + run_id=task.run_id, + error=f"{type(exc).__name__}: {exc}", + ) + return CiWatchOutcome( + thread_id=task.thread_id, + action=action, + run_id=task.run_id, + error=error, + ) diff --git a/agent-team/agent_team/coordinator.py b/agent-team/agent_team/coordinator.py index 474e112..9bea15b 100644 --- a/agent-team/agent_team/coordinator.py +++ b/agent-team/agent_team/coordinator.py @@ -75,6 +75,7 @@ __all__ = [ "default_clarify_node_factory", "default_dispatch_node_factory", "default_slack_listener_factory", + "failsafe_production_p3_wiring", "gated_build_verify_wiring", ] @@ -355,7 +356,7 @@ def gated_build_verify_wiring( *, owner: str, repo: str, - expected_run_id: str, + expected_run_id: str | None = None, allowed_scope: list[str] | None = None, diff_builder: Any = None, ci_client: Any = None, @@ -386,6 +387,14 @@ def gated_build_verify_wiring( task parks). ``ci_client`` injects a test double; the real path builds a read-only ``requests`` session at call time from the read-only token env var. + Per-task run-id binding (design §4 Decision 4): the gate binds each task's + verdict to the run id THAT TASK dispatched (persisted as ``state["run_id"]`` + by the dispatch node and read by the verifier at node-run time), NOT a + static wiring-time constant. ``expected_run_id`` is therefore optional and + defaults to ``None``; when supplied it is only a static fallback for a + harness that drives the verifier without a per-task ``state["run_id"]``. A + task whose dispatch left no run id BLOCKs (never a vacuous pass). + Lazy-imported (ci_fetcher pulls the subgraph + verifier leaves) for the same import-hygiene reason as the other factories. """ @@ -439,6 +448,105 @@ def default_dispatch_node_factory() -> "Callable[[Any], Any]": return make_dispatch_node(owner=owner, repo=repo, base=base) +# The one-line operator notice posted to #agent-team when the live P3 wiring +# cannot bind (missing owner/repo/CI-read token) and the daemon degrades to the +# inert P3 path. Goes through the lifecycle NOTIFY sink (not the park ALARM): an +# unconfigured box is an operational state, not a parked task. +_P3_INERT_NOTICE = ( + "ℹ️ P3 build→verify/dispatch is INERT this run: " + "AGENT_TEAM_REPO_OWNER / AGENT_TEAM_REPO_NAME (and a CI-read token: " + "AGENT_TEAM_CI_READ_TOKEN or GITHUB_TOKEN) are not all set. The daemon is " + "up and tasks run through PLAN/REVIEW; with no P3 subgraph wired, the " + "review loop's 'build' route is its terminus, so a task that would advance " + "to BUILD/VERIFY instead settles at the approved-plan terminus (no build, " + "no dispatch) until the P3 env is provisioned." +) + + +def _p3_env_is_configured() -> bool: + """Return True iff the live P3 wiring can bind from the environment. + + Live P3 (gated build→verify + auto-dispatch) needs the dispatch target + (``AGENT_TEAM_REPO_OWNER`` / ``AGENT_TEAM_REPO_NAME`` — the same vars + :func:`default_dispatch_node_factory` requires) AND a read-only CI token for + the verifier's authenticated conclusion read (``AGENT_TEAM_CI_READ_TOKEN``, + falling back to ``GITHUB_TOKEN`` — mirrors + :func:`agent_team.ci_fetcher._resolve_read_token`). Any missing piece means + the gate could never read an authenticated pass, so we keep the whole P3 + subgraph OFF rather than wire a half-configured, always-BLOCKing path. + """ + owner = os.environ.get("AGENT_TEAM_REPO_OWNER", "").strip() + repo = os.environ.get("AGENT_TEAM_REPO_NAME", "").strip() + ci_token = ( + os.environ.get("AGENT_TEAM_CI_READ_TOKEN", "").strip() + or os.environ.get("GITHUB_TOKEN", "").strip() + ) + return bool(owner and repo and ci_token) + + +def failsafe_production_p3_wiring( + *, + notify: "Callable[..., None] | None" = None, +) -> "tuple[BuildVerifyWiring | None, DispatchNodeFactory | None]": + """Resolve the production P3 wiring fail-safe (Decision 5; serve default). + + Per the Phase-0 design the bound P3 wiring is the new ``serve`` default, but + its factories are called EAGERLY at graph-build (``setup`` calls + ``self._build_verify_wiring()`` / ``self._dispatch_node_wiring()``), and the + live dispatch factory RAISES when ``AGENT_TEAM_REPO_OWNER`` / + ``AGENT_TEAM_REPO_NAME`` are unset. A raise there would crash-loop the + daemon at serve-start — exactly the failure mode this wrapper exists to + prevent. + + So this resolver decides ONCE, up front, from the environment: + + * **Configured** (:func:`_p3_env_is_configured` — owner + repo + a CI-read + token all present) → returns the LIVE pair: a + :func:`gated_build_verify_wiring` bound to the env owner/repo (so the + verifier reads the authenticated conclusion via the read-only fetcher) and + :func:`default_dispatch_node_factory` (which re-reads the same env at + build time). Tasks reaching P3 run BUILD → DISPATCH → VERIFY. + * **Unconfigured** → returns ``(None, None)`` — the INERT P3 path: no + build→verify subgraph and no dispatch are wired at all, so the review + loop's "build" route stays its terminus (END). A task that would advance + to P3 therefore settles at the approved-plan terminus rather than building + or dispatching — there is no BUILD/VERIFY node to reach and so nothing + parks. (Fail-closed in the sense that no diff is ever built, dispatched, or + passed; never a fabricated pass.) Logs exactly ONE WARNING and emits + ONE ``#agent-team`` inert-mode notice via the lifecycle ``notify`` sink + (NOT the park-ALARM path: an unprovisioned box is an operational state, + not a parked task). The notify sink is best-effort and fully guarded so a + Slack failure here never blocks serve-start. + + NEVER raises: serve-start must come up either fully wired or inert, but it + must always come up. + """ + if _p3_env_is_configured(): + owner = os.environ.get("AGENT_TEAM_REPO_OWNER", "").strip() + repo = os.environ.get("AGENT_TEAM_REPO_NAME", "").strip() + return ( + lambda: gated_build_verify_wiring(owner=owner, repo=repo), + default_dispatch_node_factory, + ) + + _LOG.warning( + "P3 build→verify/dispatch wiring is INERT: AGENT_TEAM_REPO_OWNER / " + "AGENT_TEAM_REPO_NAME (and a CI-read token) are not all set. The " + "coordinator starts and runs PLAN/REVIEW; with no P3 subgraph wired, the " + "review loop's 'build' route is its terminus, so a task that would " + "advance to BUILD/VERIFY instead settles at the approved-plan terminus " + "(no build, no dispatch) until the P3 env is provisioned." + ) + if notify is not None: + try: + notify(_P3_INERT_NOTICE) + except Exception: # noqa: BLE001 - an inert-notice failure must not block serve + _LOG.warning( + "inert-mode notify failed; serve still starting inert", exc_info=True + ) + return None, None + + class Coordinator: """Owns the live Plane-2 runtime: graph + resume worker + transport (§3.3). @@ -470,6 +578,10 @@ class Coordinator: build_listener: ListenerFactory | None = None, new_task_callback: "Callable[[str, str, str], str] | None" = None, notify: "Callable[..., None] | None" = None, + ci_pending_provider: "Callable[[], list[Any]] | None" = None, + ci_poller: "Callable[[Any], Any] | None" = None, + ci_timeout: timedelta | None = None, + draft_pr_provider: "Callable[[], list[Any]] | None" = None, ) -> None: self._db_path = Path(db_path) self._transport = transport @@ -507,6 +619,26 @@ class Coordinator: # behavior). The serve path injects a Slack poster so a task is never a # black box: the human sees parked / needs-more-input / plan-ready. self._notify = notify + # CI-watcher seams (P3 async resume-on-CI-complete; §4 Decision 2). Both + # OPT-IN and default None, so the CI sweep in tick() is a NO-OP unless the + # live P3 path provides them: ``ci_pending_provider`` enumerates tasks + # suspended at VERIFY awaiting CI (as + # :class:`agent_team.ci_watcher.PendingCiTask`), and ``ci_poller`` is the + # read-only poll seam (the default reuses + # :func:`agent_team.ci_fetcher.fetch_ci_result`). Left None, no CI sweep + # runs — exactly the INERT default and the unit-test path. + self._ci_pending_provider = ci_pending_provider + self._ci_poller = ci_poller + self._ci_timeout = ci_timeout + # Draft-PR runaway/stale monitor seam (P3 A4). OPT-IN and default None, so + # the draft-PR sweep in tick() is a NO-OP unless the live P3 path provides + # ``draft_pr_provider`` — an enumerator of the currently-open draft PRs (as + # :class:`agent_team.draft_pr_monitor.DraftPr`). Left None, no draft-PR + # sweep runs (the INERT default + the unit-test path). The flapping-backoff + # memory persists for the daemon's lifetime so a sustained condition is not + # re-ALARMed / re-reminded every tick. + self._draft_pr_provider = draft_pr_provider + self._draft_pr_memory: Any = None # Built by setup(). self._graph: Any = None @@ -1003,6 +1135,21 @@ class Coordinator: "Re-assign with more detail, or adjust the requirement to unblock.", thread_ts=root_ts, ) + elif self._verify_pass_verdict(values) is not None: + # P3 terminal PASS: the task's CI apply/verify run reached a + # terminal PASS (the pure-code gate PASSed) and the APPROVED route + # opened the draft PR (§3.3.2 PASS terminus). Emit the POSITIVE + # lifecycle notice (NOT the park-ALARM path) with the run/PR link + # so the human can go review the draft PR. Distinguished from the + # P2 plan-ready terminus below by the verify-stage PASS verdict — + # both settle at status/phase DONE, so the verdict is the + # discriminator (a plan-ready task carries no verify verdict). + verdict = self._verify_pass_verdict(values) + self._emit( + f"🎉 {label} — CI PASSED, draft PR opened.\n" + f"{self._draft_pr_notice(values, verdict)}", + thread_ts=root_ts, + ) else: # Plan approved + settled at the P2 terminus. PRESENT the plan # (condensed) so the human can actually review it in-thread, not @@ -1075,15 +1222,87 @@ class Coordinator: "requesting changes without converging)." ) + @staticmethod + def _verify_pass_verdict(values: "dict[str, Any]") -> "dict[str, Any] | None": + """Return the verify-stage PASS verdict if the task reached the P3 PASS terminus. + + The P3 build→dispatch→verify PASS terminus and the P2 plan-ready terminus + BOTH settle at status/phase ``DONE`` (see :func:`agent_team.graph.plan_node` + and :func:`agent_team.nodes.verifier.verifier_node`), so status alone can + not tell them apart. The discriminator is the verifier's verdict: only a + task that went through VERIFY and PASSed the pure-code CI gate appends a + ``review_verdicts`` entry with ``stage == "verify"`` and a PASS decision + (:func:`agent_team.nodes.verifier._verdict`). A plan-ready task carries no + such verdict. Returns the most recent matching verdict (the one that + opened the draft PR) or ``None`` when the task did not terminally PASS CI. + """ + verdicts = values.get("review_verdicts") or [] + if not isinstance(verdicts, list): + return None + for verdict in reversed(verdicts): + if not isinstance(verdict, dict): + continue + if ( + verdict.get("stage") == "verify" + and str(verdict.get("decision") or "").lower() == "pass" + ): + return verdict + return None + + @staticmethod + def _draft_pr_notice( + values: "dict[str, Any]", verdict: "dict[str, Any] | None" + ) -> str: + """Body of the POSITIVE draft-PR LIFECYCLE notice (run/PR link). + + Surfaces the link the human needs to go review the freshly-opened draft + PR. Two links may be present: the GitHub Actions **run** link (always + derivable from the run id the dispatcher captured + the configured + owner/repo) and the **PR** link (only once a draft-PR transport writes a + ``pr_url`` into state — forward-compatible; absent today). Both are + best-effort and fail soft: a missing run id / unset owner/repo simply + drops that line rather than crashing the milestone. The run id falls back + to ``state["run_id"]`` when the verdict carries none. + """ + lines: list[str] = [] + pr_url = str(values.get("pr_url") or "").strip() + if pr_url: + lines.append(f"• Draft PR: {pr_url}") + + run_id = "" + if isinstance(verdict, dict): + run_id = str(verdict.get("run_id") or "").strip() + if not run_id: + run_id = str(values.get("run_id") or "").strip() + if run_id: + owner = os.environ.get("AGENT_TEAM_REPO_OWNER", "").strip() + repo = os.environ.get("AGENT_TEAM_REPO_NAME", "").strip() + if owner and repo: + lines.append( + f"• CI run: https://github.com/{owner}/{repo}/actions/runs/{run_id}" + ) + else: + lines.append(f"• CI run: {run_id}") + + lines.append("• Review the draft PR when you have a moment.") + return "\n".join(lines) + def tick(self) -> list[ResumeResult]: - """One maintenance pass: deadline sweep + park policy, then drain (§3.3.1). + """One maintenance pass: deadline sweep + CI-watch sweep, then drain (§3.3.1, §3.3.2). Runs :func:`agent_team.responder.deadline_sweep` to flip overdue ``open`` questions to ``expired`` (the deterministic answer-vs-expiry race), then applies the park policy to each newly-expired id (raise the ALARM hook — §6.6 "ALARM rather than spin"; the durable ledger row is already - ``expired``, which is the task's parked state for P1). Finally drains any - resume jobs that landed. Returns the drain results. + ``expired``, which is the task's parked state for P1). It then runs the + CI-watcher sweep (:meth:`_ci_watch`) alongside the deadline sweep — the P3 + async resume-on-CI-complete pass that resumes tasks whose CI run + terminated and parks tasks whose run timed out / failed to poll (§4 + Decision 2). It then runs the draft-PR runaway/stale sweep + (:meth:`_draft_pr_monitor_sweep`) — the P3 A4 pass that ALARMs on a + draft-PR open-burst and reminds on a stale draft PR. All sweeps are + fail-soft. Finally drains any resume jobs that landed (including resumes + the CI-watch sweep enqueued). Returns the drain results. """ conn = connect(self._db_path) try: @@ -1094,10 +1313,289 @@ class Coordinator: for question_id in expired: self._park(question_id) + # CI-watcher sweep alongside the deadline sweep (§3.3.2 Decision 2). NO-OP + # unless the live P3 seams are wired; fail-soft so a CI-watch error never + # breaks the maintenance loop. + self._ci_watch() + + # Draft-PR runaway/stale sweep (P3 A4). NO-OP unless ``draft_pr_provider`` + # is wired; fail-soft so a monitor error never breaks the maintenance loop. + self._draft_pr_monitor_sweep() + results = self.drain_resumes() self._post_resume_followups(results) return results + def _ci_watch(self) -> Any: + """Run one CI-watcher sweep over tasks suspended awaiting CI (§3.3.2 Decision 2). + + NO-OP unless BOTH CI-watcher seams are wired (``ci_pending_provider`` + + ``ci_poller``) — the INERT default and the unit-test path skip it + entirely. When wired, it: + + 1. enumerates the tasks currently suspended at VERIFY awaiting CI (via + ``ci_pending_provider``, as + :class:`agent_team.ci_watcher.PendingCiTask`); and + 2. runs :func:`agent_team.ci_watcher.run_ci_watcher` with the read-only + ``ci_poller`` and injected RESUME / PARK side effects: RESUME enqueues + a turn-guarded resume onto the shared queue (drained in the same tick), + so the suspended VERIFY node re-reads the now-terminal CI result and + the pure-code gate decides; PARK marks the task parked + ALARMs. + + Fail-soft: any error in enumeration or the sweep is logged and swallowed + so a CI-watch failure never breaks the tick loop. Returns the + :class:`~agent_team.ci_watcher.CiWatchReport` (or ``None`` when skipped / + on error) for logging/tests. + """ + if self._ci_pending_provider is None or self._ci_poller is None: + return None + + from agent_team.ci_watcher import DEFAULT_CI_TIMEOUT, run_ci_watcher + + try: + pending = self._ci_pending_provider() + except Exception: # noqa: BLE001 - an enumeration failure must not break tick + _LOG.warning("ci-watch: pending-task enumeration raised", exc_info=True) + return None + + if not pending: + return None + + try: + report = run_ci_watcher( + list(pending), + poll=self._ci_poller, + on_resume=self._ci_resume, + on_park=self._ci_park, + timeout=self._ci_timeout or DEFAULT_CI_TIMEOUT, + ) + except Exception: # noqa: BLE001 - a sweep failure must not break tick + _LOG.warning("ci-watch: sweep raised", exc_info=True) + return None + + if report.parked_error: + _LOG.error( + "ci-watch: %d task(s) parked on a poll/state error (fail-closed)", + report.parked_error, + ) + return report + + def _ci_resume(self, task: Any, result: Any) -> None: + """RESUME a task whose CI run terminated, via the turn-guarded worker. + + The suspended VERIFY node interrupted with a payload but no pending + ledger question (it is a machine gate, not a human gate), so there is no + ``question_id`` / ``answer`` to thread through the responder's + first-answer-wins flip. We drive the resume through the single-flight, + turn-guarded :meth:`agent_team.resume_worker.ResumeWorker.resume_ci` + (NOT a bare ``graph.invoke``): it takes the same per-thread lock the + deadline/human-answer resumes use and re-confirms the thread is STILL + suspended at VERIFY awaiting THIS run before invoking, so a double resume + (e.g. the same terminal run observed on two overlapping sweeps) can never + corrupt the durable state. On resume VERIFY re-runs and RE-FETCHES the + now-terminal authenticated CI result (it never trusts the resume + payload); the pure-code gate decides. Best-effort and isolated: a resume + failure for one task parks it rather than crashing the sweep (the watcher + records PARKED_ERROR). + """ + if self._graph is None or self._resume_worker is None: + raise RuntimeError("Coordinator._ci_resume called before setup()") + # The resume value is intentionally ignored by VERIFY (it re-fetches the + # authenticated conclusion), so any payload works; pass the terminal + # result for operator-log provenance. ``run_id`` is the trusted dispatch + # watermark the watcher polled, so the guard binds the resume to THIS + # task's awaited run. + self._resume_worker.resume_ci( + thread_id=task.thread_id, + run_id=task.run_id, + answer=result, + ) + + def _ci_park(self, task: Any) -> None: + """PARK a task whose CI run timed out / could not be polled (fail-closed). + + Flips the durable task status to PARKED via ``update_state`` (so the + terminal state is durable across a reboot) and raises the ALARM hook so + the stall is surfaced rather than silently spun on (§6.6). Best-effort: + a write failure is logged; the watcher still records the park outcome. + """ + from agent_team.task_model import Phase, TaskStatus # noqa: PLC0415 + + try: + self._graph.update_state( + graph_mod.thread_config(task.thread_id), + { + "status": TaskStatus.PARKED.value, + "current_phase": Phase.PARKED.value, + }, + ) + except Exception: # noqa: BLE001 - best-effort durable park + _LOG.warning( + "ci-watch: could not mark task %s parked", task.thread_id, exc_info=True + ) + self._alarm_hook(task.thread_id) + + def _draft_pr_monitor_sweep(self) -> Any: + """Run one draft-PR runaway/stale sweep (P3 A4). + + NO-OP unless ``draft_pr_provider`` is wired (the INERT default + the + unit-test path skip it entirely). When wired, it: + + 1. enumerates the currently-open draft PRs (via ``draft_pr_provider``, as + :class:`agent_team.draft_pr_monitor.DraftPr`); and + 2. runs :func:`agent_team.draft_pr_monitor.run_draft_pr_monitor` with the + daemon-lifetime flapping-backoff memory and injected ALARM / reminder + side effects: the ALARM posts a runaway notice to ``#agent-team`` via + the lifecycle path (operator remediation = stop auto-dispatch; the + monitor never self-stops), and the reminder posts a stale-PR notice + (never auto-closes). + + Fail-soft: any error in enumeration or the sweep is logged and swallowed + so a monitor failure never breaks the tick loop. Returns the + :class:`~agent_team.draft_pr_monitor.MonitorReport` (or ``None`` when + skipped / on error) for logging/tests. + """ + if self._draft_pr_provider is None: + return None + + from agent_team.draft_pr_monitor import ( # noqa: PLC0415 + MonitorMemory, + run_draft_pr_monitor, + ) + + if self._draft_pr_memory is None: + self._draft_pr_memory = MonitorMemory() + + try: + draft_prs = self._draft_pr_provider() + except Exception: # noqa: BLE001 - enumeration must not break tick + _LOG.warning("draft-pr-monitor: draft-PR enumeration raised", exc_info=True) + return None + + if not draft_prs: + return None + + try: + report = run_draft_pr_monitor( + list(draft_prs), + on_alarm=self._draft_pr_alarm, + on_stale_reminder=self._draft_pr_stale_reminder, + memory=self._draft_pr_memory, + ) + except Exception: # noqa: BLE001 - a sweep failure must not break tick + _LOG.warning("draft-pr-monitor: sweep raised", exc_info=True) + return None + + if report.alarmed: + _LOG.error( + "draft-pr-monitor: RUNAWAY ALARM raised — %d draft PRs opened " + "within the window (operator remediation: stop auto-dispatch)", + report.opened_in_window, + ) + return report + + def _draft_pr_alarm(self, opened_in_window: int) -> None: + """Post the draft-PR runaway ALARM to the operator channel (P3 A4). + + Surfaces the burst + the operator remediation (stop auto-dispatch). The + monitor does NOT self-stop — ``systemctl stop`` is the human action; this + only makes the condition visible. Routed through the lifecycle ``_emit`` + sink (never raises), so a Slack failure cannot break the tick loop. + """ + self._emit( + f"🚨 ALARM: draft-PR runaway — {opened_in_window} draft PRs opened " + "within 15 minutes (threshold > 3). Auto-dispatch may be looping. " + "Remediation: stop auto-dispatch on the box (`systemctl stop`). The " + "monitor surfaces this; it does not self-stop or auto-close." + ) + + def _draft_pr_stale_reminder(self, pr: Any) -> None: + """Post a stale draft-PR reminder to the operator channel (P3 A4). + + A draft PR idle > 7 days is surfaced once (per cooldown) so it is not + silently forgotten. NEVER auto-closes — closing is a human decision. + Routed through the lifecycle ``_emit`` sink (never raises). + """ + self._emit( + f"⏳ Reminder: draft PR #{pr.number} has been idle for over 7 days. " + "Review, update, or close it (the monitor never auto-closes)." + ) + + def _enumerate_ci_pending(self) -> list[Any]: + """Enumerate the durable threads suspended at VERIFY awaiting CI (§3.3.2). + + This is the durable :data:`ci_pending_provider` the live serve path binds: + without it the CI-watcher could never be fed real tasks (nothing else + enumerates which durable threads are parked at the VERIFY machine-gate), + so a dispatched task would suspend at VERIFY and wait FOREVER — the + async-resume gap this closes. + + It walks the LangGraph SQLite checkpointer (the same DB file the ledger + uses) for the distinct ``thread_id``s it holds, then asks the compiled + graph for each thread's live snapshot. A thread is included ONLY when its + snapshot is still interrupted on the VERIFY *awaiting-CI* marker + (:func:`agent_team.ci_watcher.snapshot_awaiting_ci_run_id` returns a run + id). Threads that already advanced — resumed, parked, or done — carry no + awaiting-CI interrupt and are excluded, so the watcher never re-resumes a + task that already left the gate (the double-resume guard holds at the + enumeration boundary, before the resume worker's guard even runs). + + Fail-soft: a snapshot read that raises for one thread is logged and that + thread skipped, so one unreadable thread never blanks the whole sweep. + Returns ``[]`` (never raises) before :meth:`setup` or when the + checkpointer cannot be enumerated. + """ + from agent_team.ci_watcher import ( # noqa: PLC0415 + PendingCiTask, + snapshot_awaiting_ci_run_id, + ) + + graph = self._graph + if graph is None: + return [] + checkpointer = getattr(graph, "checkpointer", None) + if checkpointer is None or not hasattr(checkpointer, "list"): + return [] + + # Distinct thread_ids the checkpointer holds. ``list(None)`` yields every + # checkpoint across all threads (newest-first, with repeats per thread); + # we keep insertion order and dedupe so each thread is examined once. + thread_ids: list[str] = [] + seen: set[str] = set() + try: + for ckpt in checkpointer.list(None): + cfg = getattr(ckpt, "config", None) or {} + tid = (cfg.get("configurable") or {}).get("thread_id") + if isinstance(tid, str) and tid and tid not in seen: + seen.add(tid) + thread_ids.append(tid) + except Exception: # noqa: BLE001 - enumeration must never break the sweep + _LOG.warning( + "ci-watch: checkpointer enumeration raised; no CI-pending tasks " + "this pass", + exc_info=True, + ) + return [] + + pending: list[Any] = [] + for tid in thread_ids: + try: + snap = graph.get_state(graph_mod.thread_config(tid)) + except Exception: # noqa: BLE001 - one bad thread must not blank the sweep + _LOG.warning( + "ci-watch: snapshot read for thread %s raised; skipping", + tid, + exc_info=True, + ) + continue + if snapshot_awaiting_ci_run_id(snap) is None: + # Not suspended at VERIFY awaiting CI (resumed / parked / done / + # a human gate) -> exclude so we never re-resume it. + continue + values = getattr(snap, "values", None) or {} + pending.append(PendingCiTask.from_state(tid, values)) + return pending + def _park(self, question_id: str) -> None: """Apply the park policy to one expired question (§6.6 ALARM, not spin). diff --git a/agent-team/agent_team/dispatcher.py b/agent-team/agent_team/dispatcher.py index 82e682a..e9bf59d 100644 --- a/agent-team/agent_team/dispatcher.py +++ b/agent-team/agent_team/dispatcher.py @@ -28,18 +28,28 @@ from __future__ import annotations import base64 import re from dataclasses import dataclass -from typing import Protocol +from datetime import datetime, timezone +from typing import Any, Protocol from agent_team.state_store import compute_content_hash __all__ = [ "DispatchInputs", + "DispatchResult", "DispatcherError", + "RunLocator", "build_dispatch_inputs", "dispatch_apply_verify", "head_branch_for", + "select_run_id", ] + +def _utc_now_iso() -> str: + """UTC now as an ISO-8601 ``...Z`` string (matches GitHub Actions ``createdAt``).""" + return datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") + + WORKFLOW_FILE = "agent-team-apply-verify.yml" # The artifact carries exactly this filename; the workflow's materialize/guard @@ -164,6 +174,49 @@ class WorkflowDispatcher(Protocol): ) -> None: ... +class RunLocator(Protocol): + """Resolves the dispatched run's GitHub Actions ``run_id`` after a trigger. + + A ``workflow_dispatch`` run reports against the dispatched ``ref`` (``main``), + not the head branch, and ``gh run list`` does not expose workflow inputs — so + the correlation key is the workflow ``run-name`` (which interpolates + ``inputs.task_id``). Given the task id and a dispatched-at watermark, return + the matching run id as a string, or ``None`` if it cannot be resolved (the + caller then fails closed: no run_id -> the verifier gate BLOCKs / parks). + """ + + def __call__( + self, *, owner: str, repo: str, task_id: str, since_iso: str + ) -> str | None: ... + + +def run_name_for(task_id: str) -> str: + """The workflow ``run-name`` for ``task_id`` (the run_id correlation key). + + Mirrors ``run-name: "agent-team-apply ${{ inputs.task_id }}"`` in + ``agent-team-apply-verify.yml``. The dispatcher matches a triggered run by + this exact name in :func:`_default_run_locator`. + """ + return f"agent-team-apply {task_id}" + + +@dataclass(frozen=True) +class DispatchResult: + """Outcome of a dispatch: the inputs used + the located run identity. + + ``run_id`` is the GitHub Actions run id the verifier's read-only fetcher polls + and the pure-code gate binds its verdict to; ``None`` when the run could not + be located (the caller fails closed). ``dispatched_at`` is the UTC watermark + used to disambiguate the run from older runs; ``correlation_tag`` is the + run-name discriminator (the task id) recorded for audit. + """ + + inputs: DispatchInputs + run_id: str | None + dispatched_at: str + correlation_tag: str + + def dispatch_apply_verify( *, owner: str, @@ -174,17 +227,24 @@ def dispatch_apply_verify( base: str = "main", pusher: BranchPusher | None = None, dispatcher: WorkflowDispatcher | None = None, -) -> DispatchInputs: + locator: RunLocator | None = None, +) -> DispatchResult: """Transport one candidate diff into org CI: push the head branch, dispatch. The trusted apply path. Validates owner/repo, assembles the dispatch inputs, - pushes the diff as the head branch via ``pusher``, then triggers the - workflow via ``dispatcher``. Returns the :class:`DispatchInputs` used (for - the ledger/audit). ``pusher`` / ``dispatcher`` are injected so this is - testable without git/gh/network; the real defaults shell out to git/gh. + pushes the diff as the head branch via ``pusher``, triggers the workflow via + ``dispatcher``, then resolves the triggered run's ``run_id`` via ``locator`` + (matching the workflow ``run-name`` for ``task_id``). Returns a + :class:`DispatchResult` carrying the inputs + the located run identity so the + DISPATCH node can persist ``run_id`` into task state for the verifier. + + ``pusher`` / ``dispatcher`` / ``locator`` are injected so this is testable + without git/gh/network; the real defaults shell out to git/gh. NOTE: this never runs on the box (the box has no write token, D2). CI owns - every trust decision; this only moves bytes. + every trust decision; this only moves bytes. A ``None`` ``run_id`` is NOT an + error here — it fails closed downstream (the verifier gate BLOCKs without an + authenticated run to read), never a fabricated pass. """ if not _OWNER_REPO_RE.match(owner or "") or not _OWNER_REPO_RE.match(repo or ""): raise DispatcherError(f"invalid owner/repo {owner!r}/{repo!r}") @@ -194,6 +254,7 @@ def dispatch_apply_verify( ) push = pusher if pusher is not None else _default_branch_pusher() fire = dispatcher if dispatcher is not None else _default_workflow_dispatcher() + locate = locator if locator is not None else _default_run_locator() # Push the head branch FIRST: the draft-PR step opens against an # already-pushed --head, so the branch must exist before the run reaches it. @@ -204,8 +265,17 @@ def dispatch_apply_verify( head_branch=inputs.head_branch, diff_text=diff_text, ) + # Watermark stamped BEFORE firing so the locator never misses a run whose + # createdAt lands a moment after the trigger (the locator floors on this). + dispatched_at = _utc_now_iso() fire(owner=owner, repo=repo, inputs=inputs.as_inputs(), ref=base) - return inputs + run_id = locate(owner=owner, repo=repo, task_id=task_id, since_iso=dispatched_at) + return DispatchResult( + inputs=inputs, + run_id=run_id, + dispatched_at=dispatched_at, + correlation_tag=task_id, + ) def _default_branch_pusher() -> BranchPusher: @@ -325,3 +395,136 @@ def _default_workflow_dispatcher() -> WorkflowDispatcher: subprocess.run(args, check=True, capture_output=True) return _fire + + +# A freshly-triggered run takes a moment to register; poll a bounded number of +# times. Total wait ~= _LOCATE_ATTEMPTS * _LOCATE_DELAY_S seconds. +_LOCATE_ATTEMPTS = 12 +_LOCATE_DELAY_S = 5 +# Tolerate modest host<->GitHub clock skew when flooring on the dispatched-at +# watermark: a run created a little "before" our local stamp is still ours. +_LOCATE_SKEW_S = 120 + + +# A run's ``conclusion`` value that marks it as a stale/superseded run: the +# apply/verify workflow's per-task concurrency group cancels the prior run when a +# re-dispatch fires, so a re-dispatch of the SAME task_id leaves an older +# CANCELLED run sharing the run-name. We must NOT select it (it carries the prior +# build's verdict). Only ``cancelled`` is treated as stale-by-conclusion; a +# genuinely completed (success/failure) run is a legitimate match. +_STALE_CONCLUSIONS: frozenset[str] = frozenset({"cancelled"}) +# Non-terminal statuses: a run still queued or executing is the freshly-triggered +# one we want to bind to (it has no conclusion yet). +_ACTIVE_STATUSES: frozenset[str] = frozenset( + {"queued", "in_progress", "waiting", "requested", "pending"} +) + + +def select_run_id( + runs: list[dict[str, Any]], *, task_id: str, floor_iso: str +) -> str | None: + """Pick the dispatched run from ``gh run list`` rows (PURE; anti-stale). + + Matches rows whose ``name`` equals :func:`run_name_for` and whose + ``createdAt`` is at/after ``floor_iso``, then selects with a tightened rule so + a RAPID RE-DISPATCH of the same ``task_id`` (the build<->verify loop) never + binds to a stale/cancelled prior run: + + 1. Drop any matched run whose ``conclusion`` is ``cancelled`` — the workflow's + per-task concurrency group cancels the prior run on re-dispatch, so a + cancelled run sharing the run-name is the superseded one, never our run. + 2. Prefer the run with the GREATEST ``createdAt`` among the still-active + (queued / in_progress / waiting) runs — the freshly-triggered run is the + newest and has no conclusion yet. + 3. If none are active (e.g. a fast run already concluded by the time we poll), + fall back to the newest non-cancelled run overall. + + Returns the chosen ``databaseId`` as a string, or ``None`` if nothing + matches (the caller then fails closed). ``createdAt`` ties break on the + greater ``databaseId`` (monotonic per repo → the later-created run). + """ + target_name = run_name_for(task_id) + matches = [ + r + for r in runs + if r.get("name") == target_name and str(r.get("createdAt", "")) >= floor_iso + ] + # Drop superseded (concurrency-cancelled) prior runs of the same task_id. + matches = [ + r + for r in matches + if str(r.get("conclusion") or "").lower() not in _STALE_CONCLUSIONS + ] + if not matches: + return None + + def _sort_key(r: dict[str, Any]) -> tuple[str, int]: + try: + db_id = int(r.get("databaseId", 0)) + except (TypeError, ValueError): + db_id = 0 + return (str(r.get("createdAt", "")), db_id) + + active = [ + r for r in matches if str(r.get("status") or "").lower() in _ACTIVE_STATUSES + ] + pool = active or matches + pool.sort(key=_sort_key, reverse=True) + return str(pool[0]["databaseId"]) + + +def _default_run_locator() -> RunLocator: + """Real locator: match the triggered run by ``run-name`` via ``gh run list``. + + Polls ``gh run list`` (read-only) for runs of the apply/verify workflow and + delegates the anti-stale SELECTION to the pure :func:`select_run_id`: it skips + a concurrency-cancelled prior run of the same ``task_id`` and prefers the + newest still-active (queued/in_progress) run, falling back to the newest + non-cancelled run overall. This means a RAPID RE-DISPATCH of the same task + (the build<->verify loop) binds to the CURRENT run, never the superseded one. + Returns ``None`` if no matching run registers within the bounded poll window + (fails closed). + """ + + def _locate(*, owner: str, repo: str, task_id: str, since_iso: str) -> str | None: + import json + import subprocess + import time + from datetime import timedelta + + try: + floor_dt = datetime.strptime(since_iso, "%Y-%m-%dT%H:%M:%SZ").replace( + tzinfo=timezone.utc + ) - timedelta(seconds=_LOCATE_SKEW_S) + floor_iso = floor_dt.strftime("%Y-%m-%dT%H:%M:%SZ") + except ValueError: + floor_iso = since_iso + + for attempt in range(_LOCATE_ATTEMPTS): + proc = subprocess.run( + [ + "gh", + "run", + "list", + "--repo", + f"{owner}/{repo}", + "--workflow", + WORKFLOW_FILE, + "--json", + "databaseId,name,createdAt,status,conclusion", + "--limit", + "50", + ], + check=True, + capture_output=True, + text=True, + ) + runs = json.loads(proc.stdout or "[]") + run_id = select_run_id(runs, task_id=task_id, floor_iso=floor_iso) + if run_id is not None: + return run_id + if attempt < _LOCATE_ATTEMPTS - 1: + time.sleep(_LOCATE_DELAY_S) + return None + + return _locate diff --git a/agent-team/agent_team/draft_pr_monitor.py b/agent-team/agent_team/draft_pr_monitor.py new file mode 100644 index 0000000..19a94aa --- /dev/null +++ b/agent-team/agent_team/draft_pr_monitor.py @@ -0,0 +1,422 @@ +"""Draft-PR runaway monitor + stale cleanup sweep (P3 box-integration, A4). + +Once the box-side BUILD → DISPATCH → VERIFY path is live (CI apply/verify is +already provisioned), a passing task ends at a **draft PR** (the §3.3.2 PASS +terminus). Two failure modes then need a maintenance sweep, mirroring the shape +of :mod:`agent_team.deadline_timer` and :mod:`agent_team.ci_watcher` (a pure, +restart-safe ``tick()``-driven pass whose side effects are injected callables, so +it is unit-testable with NO network and NO live graph): + +* **RUNAWAY** — a dispatch loop (a stuck builder, a re-dispatch storm, or a + fixer that keeps re-opening) can open draft PRs far faster than a human can + review them. If **more than 3 draft PRs are opened within any 15-minute + window**, raise an ALARM to ``#agent-team`` (via the existing lifecycle / ALARM + path). The remediation is **operator-driven**: stop auto-dispatch + (``systemctl stop``). The monitor *surfaces* the condition; it never + self-restarts, self-stops, or auto-closes anything — it has no standing write + authority and must not act on infrastructure. + +* **STALE** — a draft PR that has sat **idle for more than 7 days** is surfaced + with a single ``#agent-team`` reminder so it is not silently forgotten. The + monitor **never auto-closes** a stale PR — closing is a human decision; the + reminder is the only action. + +Flapping backoff (the load-bearing anti-spam discipline): a RUNAWAY condition +typically persists across many ticks (the offending PRs stay inside the window +for the whole 15 minutes), and a STALE PR stays stale until a human acts. Without +backoff the sweep would re-ALARM / re-remind on *every* tick. So the monitor +carries a small :class:`MonitorMemory` (injected, durable-across-ticks): it +ALARMs at most once per ``alarm_cooldown`` and reminds about a given PR at most +once per ``stale_cooldown``. The memory is passed in (not module-global) so the +coordinator owns its lifetime and a test can assert the backoff deterministically. + +Fail-soft / fail-closed discipline (mirroring the sibling sweeps): + +* Every side effect (``on_alarm`` / ``on_stale_reminder``) is an injected + callable; the monitor performs no transport I/O of its own. +* A draft PR with an unparseable ``opened_at`` is ignored for the runaway count + (it cannot be placed in the window) but a present ``updated_at`` is still + considered for staleness; a PR with neither parseable timestamp is skipped + rather than crashing the sweep. +* A side effect that raises is isolated per condition (logged, recorded) so one + broken post never aborts the pass — the daemon's ``tick()`` loop keeps running. +* All inputs are read fresh each pass (the injected provider re-enumerates the + open draft PRs); the only retained state is the small backoff memory. +""" + +from __future__ import annotations + +import logging +from collections.abc import Callable +from dataclasses import dataclass, field +from datetime import datetime, timedelta, timezone +from enum import Enum + +__all__ = [ + "AlarmFn", + "DEFAULT_ALARM_COOLDOWN", + "DEFAULT_ALARM_THRESHOLD", + "DEFAULT_ALARM_WINDOW", + "DEFAULT_STALE_AFTER", + "DEFAULT_STALE_COOLDOWN", + "DraftPr", + "MonitorAction", + "MonitorMemory", + "MonitorOutcome", + "MonitorReport", + "StaleReminderFn", + "run_draft_pr_monitor", +] + +logger = logging.getLogger(__name__) + +# Concrete thresholds (A4). RUNAWAY: > 3 draft PRs opened within 15 minutes. +DEFAULT_ALARM_THRESHOLD = 3 +DEFAULT_ALARM_WINDOW = timedelta(minutes=15) +# STALE: a draft PR idle (no update) for more than 7 days. +DEFAULT_STALE_AFTER = timedelta(days=7) + +# Flapping backoff windows. A runaway condition persists across many ticks while +# the offending PRs stay inside the 15-minute window, and a stale PR stays stale +# until a human acts; these cooldowns stop the sweep re-alarming / re-reminding +# every tick. One ALARM per 15 min, one reminder per PR per day. +DEFAULT_ALARM_COOLDOWN = timedelta(minutes=15) +DEFAULT_STALE_COOLDOWN = timedelta(days=1) + + +class MonitorAction(Enum): + """The action the sweep took this pass for one surfaced condition. + + ``ALARMED`` — a runaway burst was detected and the ALARM was posted. + ``ALARM_SUPPRESSED`` — a runaway burst was detected but an ALARM was posted + recently (within ``alarm_cooldown``), so it was suppressed (flapping + backoff). ``STALE_REMINDED`` — a stale PR's reminder was posted. + ``STALE_SUPPRESSED`` — a stale PR was found but it was reminded about + recently (within ``stale_cooldown``), so the reminder was suppressed. + ``ERRORED`` — a side effect raised; the condition was detected but the post + failed (isolated, the sweep continues). + """ + + ALARMED = "alarmed" + ALARM_SUPPRESSED = "alarm_suppressed" + STALE_REMINDED = "stale_reminded" + STALE_SUPPRESSED = "stale_suppressed" + ERRORED = "errored" + + +@dataclass(frozen=True) +class DraftPr: + """A minimal read-snapshot of one open draft PR (A4). + + Only the fields the monitor needs: identity (``number``) and the two + timestamps the runaway-count / staleness checks key off. ``opened_at`` is + when the draft PR was created (runaway window); ``updated_at`` is the last + activity (staleness). Frozen because it is a snapshot — the monitor never + mutates a PR; it acts only through the injected callables. + """ + + number: int + opened_at: str | None = None + updated_at: str | None = None + + +@dataclass(frozen=True) +class MonitorOutcome: + """The result of one surfaced condition this pass (A4).""" + + action: MonitorAction + # For a runaway ALARM: the count of PRs opened inside the window. For a stale + # outcome: the PR number. ``None`` is never expected but keeps the dataclass + # total for the ERRORED path. + detail: str | None = None + error: str | None = None + + +@dataclass +class MonitorReport: + """Aggregate result of one draft-PR monitor pass (mirrors the sibling sweeps). + + ``outcomes`` is one entry per surfaced condition (a runaway ALARM and/or each + stale reminder). The summary counters let the coordinator log/ALARM without + re-walking the list. + """ + + outcomes: list[MonitorOutcome] = field(default_factory=list) + # The number of draft PRs counted as opened inside the runaway window this + # pass, for observability (not every counted PR produces an outcome — only a + # *breach* does). + opened_in_window: int = 0 + + @property + def alarmed(self) -> int: + """Runaway ALARMs actually posted this pass.""" + return sum(1 for o in self.outcomes if o.action is MonitorAction.ALARMED) + + @property + def alarm_suppressed(self) -> int: + """Runaway breaches detected but suppressed by the alarm cooldown.""" + return sum( + 1 for o in self.outcomes if o.action is MonitorAction.ALARM_SUPPRESSED + ) + + @property + def stale_reminded(self) -> int: + """Stale reminders actually posted this pass.""" + return sum(1 for o in self.outcomes if o.action is MonitorAction.STALE_REMINDED) + + @property + def stale_suppressed(self) -> int: + """Stale PRs found but suppressed by the per-PR reminder cooldown.""" + return sum( + 1 for o in self.outcomes if o.action is MonitorAction.STALE_SUPPRESSED + ) + + @property + def errored(self) -> int: + """Conditions whose side effect raised (isolated; sweep continued).""" + return sum(1 for o in self.outcomes if o.action is MonitorAction.ERRORED) + + +@dataclass +class MonitorMemory: + """Durable-across-ticks backoff memory (the flapping-backoff seam). + + The monitor itself is otherwise pure: it reads the open draft PRs fresh each + pass. This small mutable record is the ONE piece of state that must survive + between ticks so a persistent condition does not re-ALARM / re-remind every + pass. The coordinator owns one instance for the daemon's lifetime; a test + constructs its own to assert the backoff deterministically. + + ``last_alarm_at`` — when a runaway ALARM was last posted (None = never). + ``last_reminded_at`` — per-PR-number, when that PR was last reminded about. + Entries for PRs no longer present are pruned each pass so the map cannot grow + without bound across a long-running daemon. + """ + + last_alarm_at: datetime | None = None + last_reminded_at: dict[int, datetime] = field(default_factory=dict) + + +# Injected side-effect seams. Keeping these as callables means the monitor +# performs no transport I/O of its own (testable, faithful to the deadline_timer +# / ci_watcher shape). +AlarmFn = Callable[[int], None] +"""Called once (per cooldown) when a runaway burst is detected. Receives the +count of draft PRs opened inside the window. The handler posts the ALARM to +``#agent-team`` and surfaces the operator remediation (stop auto-dispatch); it +does NOT self-stop.""" + +StaleReminderFn = Callable[[DraftPr], None] +"""Called once (per cooldown) per stale draft PR. The handler posts a +``#agent-team`` reminder. It NEVER auto-closes the PR.""" + + +def _utc_now() -> datetime: + """Return the current UTC time (injectable via ``now`` in the sweep).""" + return datetime.now(timezone.utc) + + +def _parse_iso(value: str | None) -> datetime | None: + """Parse an ISO-8601 timestamp to an aware UTC datetime, or ``None``. + + A missing / unparseable timestamp yields ``None`` so the caller fails soft + (skips that PR for the affected check) rather than crashing the sweep. A + naive stamp is treated as UTC (the ledger / GitHub API write UTC) so the + aware/naive compares below never raise. + """ + if not value: + return None + try: + parsed = datetime.fromisoformat(value) + except (TypeError, ValueError): + return None + if parsed.tzinfo is None: + parsed = parsed.replace(tzinfo=timezone.utc) + return parsed + + +def run_draft_pr_monitor( + draft_prs: list[DraftPr], + *, + on_alarm: AlarmFn, + on_stale_reminder: StaleReminderFn, + memory: MonitorMemory | None = None, + alarm_threshold: int = DEFAULT_ALARM_THRESHOLD, + alarm_window: timedelta = DEFAULT_ALARM_WINDOW, + stale_after: timedelta = DEFAULT_STALE_AFTER, + alarm_cooldown: timedelta = DEFAULT_ALARM_COOLDOWN, + stale_cooldown: timedelta = DEFAULT_STALE_COOLDOWN, + now: datetime | None = None, +) -> MonitorReport: + """Run one draft-PR runaway + stale sweep (A4). + + ``draft_prs`` is the set of currently-open draft PRs (the coordinator + supplies them fresh each pass via an injected provider). The sweep: + + 1. **Runaway.** Counts draft PRs whose ``opened_at`` is within ``now - + alarm_window``. If that count is **strictly greater than** + ``alarm_threshold`` (the A4 rule: > 3 within 15 min), call ``on_alarm`` + once — but only if a prior ALARM is older than ``alarm_cooldown`` (flapping + backoff); otherwise record ``ALARM_SUPPRESSED``. The remediation is the + operator's (``systemctl stop``); the monitor never self-stops. + 2. **Stale.** For each draft PR idle longer than ``stale_after`` (``now - + updated_at > stale_after``), call ``on_stale_reminder`` once — but only if + that PR was last reminded longer ago than ``stale_cooldown`` (per-PR + backoff); otherwise record ``STALE_SUPPRESSED``. The monitor never + auto-closes a stale PR. + + Side effects are isolated per condition: an ``on_alarm`` / ``on_stale_reminder`` + that raises yields an ``ERRORED`` outcome (the failure is recorded and the + backoff watermark is NOT advanced, so the next pass retries) and the sweep + continues with the rest of the batch. + + Restart-safety: inputs are read fresh each pass; the only retained state is + ``memory`` (the backoff watermarks). A fresh ``memory`` (e.g. after a reboot) + simply means the first post-reboot breach/stale PR is surfaced again — a + re-notification, never a missed or duplicated *action* (the monitor takes no + infrastructure action). + + Returns a :class:`MonitorReport` describing what happened. + """ + current = now or _utc_now() + mem = memory if memory is not None else MonitorMemory() + report = MonitorReport() + + present_numbers = {pr.number for pr in draft_prs} + + # --- Runaway check ----------------------------------------------------- # + window_start = current - alarm_window + opened_in_window = 0 + for pr in draft_prs: + opened_at = _parse_iso(pr.opened_at) + if opened_at is not None and opened_at >= window_start: + opened_in_window += 1 + report.opened_in_window = opened_in_window + + if opened_in_window > alarm_threshold: + _handle_runaway( + opened_in_window, + on_alarm=on_alarm, + mem=mem, + alarm_cooldown=alarm_cooldown, + now=current, + report=report, + ) + + # --- Stale check ------------------------------------------------------- # + for pr in draft_prs: + updated_at = _parse_iso(pr.updated_at) + if updated_at is None: + # No parseable last-activity stamp -> cannot assess staleness. Skip + # rather than guess (fail-soft); the runaway count is unaffected. + continue + if current - updated_at <= stale_after: + continue + _handle_stale( + pr, + on_stale_reminder=on_stale_reminder, + mem=mem, + stale_cooldown=stale_cooldown, + now=current, + report=report, + ) + + # Prune backoff watermarks for PRs no longer open so the memory cannot grow + # without bound over a long-running daemon. + for number in list(mem.last_reminded_at): + if number not in present_numbers: + del mem.last_reminded_at[number] + + return report + + +def _handle_runaway( + opened_in_window: int, + *, + on_alarm: AlarmFn, + mem: MonitorMemory, + alarm_cooldown: timedelta, + now: datetime, + report: MonitorReport, +) -> None: + """Post (or suppress) a runaway ALARM under the flapping-backoff cooldown.""" + last = mem.last_alarm_at + if last is not None and now - last < alarm_cooldown: + logger.debug( + "draft-pr-monitor: runaway breach (%d in window) suppressed by cooldown", + opened_in_window, + ) + report.outcomes.append( + MonitorOutcome( + action=MonitorAction.ALARM_SUPPRESSED, + detail=str(opened_in_window), + ) + ) + return + + try: + on_alarm(opened_in_window) + except Exception as exc: # noqa: BLE001 - isolate one side-effect failure + logger.exception( + "draft-pr-monitor: runaway ALARM side effect raised; watermark not " + "advanced (will retry next pass)" + ) + report.outcomes.append( + MonitorOutcome( + action=MonitorAction.ERRORED, + detail=str(opened_in_window), + error=f"{type(exc).__name__}: {exc}", + ) + ) + return + + # Advance the watermark ONLY after a successful post so a failed post retries. + mem.last_alarm_at = now + logger.warning( + "draft-pr-monitor: RUNAWAY — %d draft PRs opened within the window; " + "ALARM raised (operator remediation: stop auto-dispatch)", + opened_in_window, + ) + report.outcomes.append( + MonitorOutcome(action=MonitorAction.ALARMED, detail=str(opened_in_window)) + ) + + +def _handle_stale( + pr: DraftPr, + *, + on_stale_reminder: StaleReminderFn, + mem: MonitorMemory, + stale_cooldown: timedelta, + now: datetime, + report: MonitorReport, +) -> None: + """Post (or suppress) a stale reminder for one PR under per-PR backoff.""" + last = mem.last_reminded_at.get(pr.number) + if last is not None and now - last < stale_cooldown: + report.outcomes.append( + MonitorOutcome(action=MonitorAction.STALE_SUPPRESSED, detail=str(pr.number)) + ) + return + + try: + on_stale_reminder(pr) + except Exception as exc: # noqa: BLE001 - isolate one side-effect failure + logger.exception( + "draft-pr-monitor: stale reminder side effect for PR #%s raised; " + "watermark not advanced (will retry next pass)", + pr.number, + ) + report.outcomes.append( + MonitorOutcome( + action=MonitorAction.ERRORED, + detail=str(pr.number), + error=f"{type(exc).__name__}: {exc}", + ) + ) + return + + mem.last_reminded_at[pr.number] = now + report.outcomes.append( + MonitorOutcome(action=MonitorAction.STALE_REMINDED, detail=str(pr.number)) + ) diff --git a/agent-team/agent_team/graph.py b/agent-team/agent_team/graph.py index 9983f3a..efb8d19 100644 --- a/agent-team/agent_team/graph.py +++ b/agent-team/agent_team/graph.py @@ -135,9 +135,11 @@ PARKED_ROUTE = "parked" # the topology that connects the injected nodes. BUILD_NODE = "build_node" VERIFY_NODE = "verify_node" -# P3+ dispatch vertex id: the node that carries an approved diff into org CI. -# Wired by build_graph only when the caller injects a dispatch_node callable; the -# default (None) leaves APPROVED_ROUTE → END unchanged so the graph is inert. +# P3+ dispatch vertex id: the node that triggers org CI and captures the run id. +# Wired by build_graph only when the caller injects a dispatch_node callable, in +# which case it is spliced between BUILD and VERIFY (BUILD -> DISPATCH -> VERIFY) +# so DISPATCH fires the CI run + writes ``state["run_id"]`` BEFORE VERIFY reads it +# (design §4 Decision 1). The default (None) falls back to BUILD -> VERIFY. DISPATCH_NODE = "dispatch_node" # P3 route ids returned by the injected ``route_after_verify`` function. They @@ -445,14 +447,23 @@ def build_graph( route terminates at ``END`` (the approved-plan terminus), exactly as before. Production stays P2 (clarify -> plan -> review). * **P3:** ``build_verify`` is given (with ``review_node``) -> the review's - ``"build"`` route is REPOINTED at the BUILD node, ``BUILD -> VERIFY`` is - wired, and ``route_after_verify`` maps ``{approved -> END (PR terminus), + ``"build"`` route is REPOINTED at the BUILD node, the linear stage order + is wired, and ``route_after_verify`` maps ``{approved -> END (PR terminus), build -> BUILD (bounded build<->verify loop), parked -> END (escalation)}``. The subgraph stays INERT unless the caller binds real diff-builder / CI seams (held for the §3.3.2 security gate); with the default INERT seams the verifier gate has no authenticated pass and parks. Passing ``build_verify`` without ``review_node`` is a wiring error (there is no ``"build"`` route to repoint). + + ``dispatch_node`` (P3+, design §4 Decision 1) inserts the CI-dispatch vertex + BETWEEN BUILD and VERIFY: ``BUILD -> DISPATCH -> VERIFY``. DISPATCH triggers + the org CI run and captures the run id into ``state["run_id"]`` BEFORE VERIFY + reads it (VERIFY gates on a CI conclusion that only exists once DISPATCH has + fired the run). When ``dispatch_node`` is ``None`` (the default) the order + falls back to ``BUILD -> VERIFY`` directly. ``dispatch_node`` requires + ``build_verify`` (there is no BUILD/VERIFY pair to splice it between + otherwise). """ clarify = live_clarify_node if live_clarify_node is not None else clarify_node plan = live_plan_node if live_plan_node is not None else plan_node @@ -472,9 +483,10 @@ def build_graph( if dispatch_node is not None and build_verify is None: raise ValueError( - "build_graph: dispatch_node requires build_verify — it repoints the " - "verifier's APPROVED_ROUTE, so there is nothing to repoint without a " - "build->verify subgraph." + "build_graph: dispatch_node requires build_verify — it is spliced " + "between BUILD and VERIFY (BUILD -> DISPATCH -> VERIFY), so there is " + "no BUILD/VERIFY pair to wire it between without a build->verify " + "subgraph." ) builder: StateGraph = StateGraph(PipelineState) @@ -520,30 +532,29 @@ def build_graph( route_review, {BUILD_ROUTE: BUILD_NODE, PLAN: PLAN, PARKED_ROUTE: END}, ) - builder.add_edge(BUILD_NODE, VERIFY_NODE) + # Linear order is BUILD -> DISPATCH -> VERIFY so DISPATCH triggers CI + # and captures ``state["run_id"]`` BEFORE VERIFY reads it (design §4 + # Decision 1: reorder; the CI conclusion VERIFY gates on only exists + # after DISPATCH fires the run). When no dispatch node is wired, fall + # back to BUILD -> VERIFY directly (the INERT default: the verifier's + # CI fetcher yields no authenticated pass and the task parks). if dispatch_node is None: - # P3 default: APPROVED_ROUTE is the terminus (no dispatch). - builder.add_conditional_edges( - VERIFY_NODE, - route_after_verify, - {APPROVED_ROUTE: END, BUILD_ROUTE: BUILD_NODE, PARKED_ROUTE: END}, - ) + builder.add_edge(BUILD_NODE, VERIFY_NODE) else: - # P3+: repoint APPROVED_ROUTE at the dispatch node, then END. builder.add_node( DISPATCH_NODE, _instrument(DISPATCH_NODE, dispatch_node, transition_recorder), ) - builder.add_conditional_edges( - VERIFY_NODE, - route_after_verify, - { - APPROVED_ROUTE: DISPATCH_NODE, - BUILD_ROUTE: BUILD_NODE, - PARKED_ROUTE: END, - }, - ) - builder.add_edge(DISPATCH_NODE, END) + builder.add_edge(BUILD_NODE, DISPATCH_NODE) + builder.add_edge(DISPATCH_NODE, VERIFY_NODE) + # VERIFY's verdict routes to {approved -> END (PR terminus), + # build -> BUILD (bounded build<->verify loop), parked -> END + # (escalation)} regardless of whether DISPATCH is wired. + builder.add_conditional_edges( + VERIFY_NODE, + route_after_verify, + {APPROVED_ROUTE: END, BUILD_ROUTE: BUILD_NODE, PARKED_ROUTE: END}, + ) if checkpointer is None: return builder.compile() diff --git a/agent-team/agent_team/nodes/build_verify_subgraph.py b/agent-team/agent_team/nodes/build_verify_subgraph.py index fba59ec..24fe356 100644 --- a/agent-team/agent_team/nodes/build_verify_subgraph.py +++ b/agent-team/agent_team/nodes/build_verify_subgraph.py @@ -2,7 +2,13 @@ This module is the **wiring topology** for the Plane-2 build -> verify stage:: - ... -> REVIEW (route "build") -> BUILD -> VERIFY -> {approved | build | parked} + ... -> REVIEW (route "build") -> BUILD -> [DISPATCH] -> VERIFY + -> {approved | build | parked} + +DISPATCH is the optional CI-trigger vertex :func:`agent_team.graph.build_graph` +splices between BUILD and VERIFY when a dispatch node is wired (design §4 +Decision 1): it triggers the org CI run and captures ``state["run_id"]`` BEFORE +VERIFY reads it. With no dispatch node the order is simply ``BUILD -> VERIFY``. It produces the BUILD node, the VERIFY node, and the :func:`route_after_verify` conditional-edge function so the Integrate phase can @@ -189,18 +195,58 @@ def make_verify_node( than smuggling something past the gate. The real read-only-PAT fetcher is bound via :func:`bind_ci_result_fetcher` only after the §3.3.2 trust boundary clears its security gate. + + PER-TASK run-id binding (design §4 Decision 4): the expected run id the gate + binds the verdict to is NOT baked into ``config`` at factory time. The + dispatch node persists the run id THIS task dispatched as ``state["run_id"]``, + and :func:`agent_team.nodes.verifier.verifier_node` reads it from the state + threaded through below (``config.expected_run_id`` is only a static fallback + for harnesses with no per-task run id). So a single ``config`` shared across + tasks still gates each task against its OWN dispatched run, and a task whose + dispatch left no run id BLOCKs — never a vacuous pass. + + ASYNC CI-WAIT (design §4 Decision 2): a CI apply/verify run takes ~7 minutes, + and a multi-minute *blocking* fetch here would stall the coordinator daemon's + tick loop and every other task. So when there IS a dispatched run to wait for + (``state["run_id"]`` is set) but the fetch yields no terminal result yet (the + run is still in progress → ``None``), the node SUSPENDS via + :func:`~langgraph.types.interrupt` — exactly the durable suspend/resume shape + the clarify human-gate node uses. The CI-watcher sweep + (:func:`agent_team.ci_watcher.run_ci_watcher`) polls the run read-only and + RESUMES this node once the run reaches a terminal conclusion; on resume the + node RE-FETCHES the now-terminal result and the pure-code gate decides. If the + re-fetched result is still not terminal (e.g. a spurious resume), the node + falls through to the gate, which BLOCKs/parks — fail-closed, never a vacuous + pass. With NO dispatched run (``state["run_id"]`` absent — the INERT path or a + harness), the node does NOT suspend: a ``None`` fetch flows straight to the + gate, which BLOCKs and parks exactly as before. """ fetcher: CiResultFetcher = ( ci_result_fetcher if ci_result_fetcher is not None else _no_ci_result ) def node(state: PipelineState) -> PipelineState: - fetched = fetcher(state) - ci_result = fetched if isinstance(fetched, Mapping) else None + ci_result = _fetch_ci_result(fetcher, state) - # Merge the (possibly None) fetched CI result into the state the node - # reads from, WITHOUT mutating the caller's state object. The node reads - # ``ci_results``; a None result leaves the gate with nothing to pass on. + # Async CI-wait: only when a run was actually dispatched (state["run_id"] + # is set) AND it has no terminal result yet do we suspend, so the daemon + # never blocks on an in-progress run. The CI-watcher resumes us on a + # terminal conclusion; we re-fetch once after resume. The INERT/no-run + # path (no run_id) skips this and lets the gate BLOCK/park as before. + run_id = state.get("run_id") + if ci_result is None and isinstance(run_id, str) and run_id: + # Suspend + checkpoint; the CI-watcher's resume payload is the signal + # that the run terminated. We do not trust the payload's contents — + # we RE-FETCH the authenticated conclusion below so the gate reads a + # patch-independent, freshly-fetched result, never a resume-supplied + # one. + _await_ci(run_id) + ci_result = _fetch_ci_result(fetcher, state) + + # Merge the (possibly still-None) fetched CI result into the state the + # node reads from, WITHOUT mutating the caller's state object. The node + # reads ``ci_results``; a None result leaves the gate with nothing to pass + # on (it BLOCKs → park), so a never-terminal run fails closed. scoped_state: dict[str, Any] = dict(state) scoped_state["ci_results"] = ci_result return verifier_node(scoped_state, config) @@ -208,6 +254,40 @@ def make_verify_node( return node +def _fetch_ci_result( + fetcher: CiResultFetcher, state: PipelineState +) -> Mapping[str, Any] | None: + """Call the injected CI fetcher and normalise its result. + + Any value other than a mapping is treated as "no terminal result" (``None``), + so a malformed fetcher fails SAFE (the gate BLOCKs) rather than smuggling a + non-mapping past the gate. + """ + fetched = fetcher(state) + return fetched if isinstance(fetched, Mapping) else None + + +def _await_ci(run_id: str) -> None: + """Suspend the VERIFY node until the CI-watcher resumes it (§4 Decision 2). + + Mirrors the clarify human-gate node's durable suspend: calls + :func:`langgraph.types.interrupt` so the graph checkpoints and the daemon's + tick loop is freed while a multi-minute CI run is in flight. The + :func:`agent_team.ci_watcher.run_ci_watcher` sweep polls the run read-only and + drives the resume once it terminates. The interrupt payload carries only the + ``run_id`` being awaited (provenance for the watcher / operator logs); the + resume VALUE is intentionally ignored — the node re-fetches the authenticated + conclusion so the gate never reads a resume-supplied verdict. + + ``langgraph`` is imported lazily here to preserve this module's "no SDK at + module top" discipline (the topology stays importable where ``langgraph`` is + absent; the interrupt is only reached on the live, dispatched path). + """ + from langgraph.types import interrupt + + interrupt({"awaiting_ci": True, "run_id": run_id}) + + def route_after_verify(state: PipelineState) -> str: """LangGraph conditional-edge: the next route id after the VERIFY node. @@ -275,12 +355,16 @@ def bind_ci_result_fetcher( # A module-level note for the Integrate phase (no execution): the build->verify -# subgraph is hung off the review loop's "build" route. The conditional-edge map +# subgraph is hung off the review loop's "build" route. The linear stage order is +# BUILD -> [DISPATCH] -> VERIFY — DISPATCH (the optional CI-trigger vertex +# build_graph splices in when a dispatch node is wired) fires the CI run and +# captures ``state["run_id"]`` BEFORE VERIFY reads it (design §4 Decision 1); with +# no dispatch node the order is just BUILD -> VERIFY. The conditional-edge map # from VERIFY should send APPROVED_ROUTE to the PR/draft terminus, BUILD_ROUTE # back to the BUILD node (the bounded build<->verify loop, capped by # VerifierConfig.max_build_loops), and PARKED_ROUTE to the escalation terminus. # build_graph wires this in opt-in; this module never assembles it itself. _INTEGRATE_NOTE = ( - "review('build') -> BUILD -> VERIFY -> route_after_verify -> " + "review('build') -> BUILD -> [DISPATCH] -> VERIFY -> route_after_verify -> " "{approved: PR terminus, build: BUILD (loop), parked: escalation}" ) diff --git a/agent-team/agent_team/nodes/dispatch_invoker.py b/agent-team/agent_team/nodes/dispatch_invoker.py index c0f259f..6c9c1da 100644 --- a/agent-team/agent_team/nodes/dispatch_invoker.py +++ b/agent-team/agent_team/nodes/dispatch_invoker.py @@ -48,6 +48,7 @@ def make_dispatch_node( base: str = "main", pusher: Any = None, dispatcher: Any = None, + locator: Any = None, ) -> Callable[[Any], Any]: """Build a LangGraph dispatch node for ``owner``/``repo``. @@ -86,7 +87,7 @@ def make_dispatch_node( return _parked try: - dispatch_apply_verify( + result = dispatch_apply_verify( owner=owner, repo=repo, task_id=thread_id, @@ -95,6 +96,7 @@ def make_dispatch_node( base=base, pusher=pusher, dispatcher=dispatcher, + locator=locator, ) except DispatcherError as exc: _LOG.error( @@ -111,13 +113,30 @@ def make_dispatch_node( ) return _parked + if not result.run_id: + # Fired, but the run could not be correlated. Persist the watermark + # anyway and let the verifier gate fail closed (no authenticated run + # to read -> BLOCK/park) rather than fabricating progress. + _LOG.warning( + "dispatch_node: task %s dispatched but run_id unresolved; " + "downstream verify will fail closed", + thread_id, + ) + _LOG.info( - "dispatch_node: dispatched task %s to %s/%s (base=%s)", + "dispatch_node: dispatched task %s to %s/%s (base=%s, run_id=%s)", thread_id, owner, repo, base, + result.run_id, ) - return {} + # Persist the located run identity so the verifier's read-only fetcher + # polls THIS task's run and the pure-code gate binds its verdict to it. + return { + "run_id": result.run_id, + "dispatched_at": result.dispatched_at, + "ci_correlation_tag": result.correlation_tag, + } return dispatch_node diff --git a/agent-team/agent_team/nodes/verifier.py b/agent-team/agent_team/nodes/verifier.py index 6e59a75..e318753 100644 --- a/agent-team/agent_team/nodes/verifier.py +++ b/agent-team/agent_team/nodes/verifier.py @@ -15,10 +15,12 @@ fix hint for the builders. The LLM is never asked whether the task passed. State contract (mirrors :class:`agent_team.task_model.PipelineState`): -* reads ``candidate_diff``, ``diff_hash`` (ledger hash), ``ci_results``; +* reads ``candidate_diff``, ``diff_hash`` (ledger hash), ``ci_results``, + ``build_loops`` (the durable per-task build<->verify count); * writes ``status``, ``current_phase``, ``review_verdicts`` (appends the gate - verdict), and ``ci_results`` (annotated with the gate decision for - provenance). + verdict), ``ci_results`` (annotated with the gate decision for provenance), + and — on a recoverable gate FAIL — the incremented ``build_loops`` so the + budget advances across the BUILD->DISPATCH->VERIFY loop (LOGIC-RACE-01). Transitions (the §3.3 "Stability + autonomy bounds" — the verifier must pass or the task loops/holds, never ships): @@ -65,15 +67,29 @@ DEFAULT_MAX_BUILD_LOOPS: int = 3 class VerifierConfig: """Per-invocation knobs for the verifier node (§3.3, §3.3.2). - ``expected_run_id`` keys the gate to the exact CI run the verifier - dispatched for this diff (a stale/substituted run id is a BLOCK). + ``expected_run_id`` is a STATIC fallback only. The gate binds to the run id + THIS task dispatched, read from ``state["run_id"]`` at node-run time (the + dispatcher persists it there); the config value is consulted only when state + carries no ``run_id`` (e.g. a unit harness that drives the node directly). A + per-task ``state["run_id"]`` therefore always wins over this constant, and a + ``None`` effective run id is a BLOCK (never a vacuous pass). ``allowed_scope`` is the task's declared-scope path prefixes for the denylist boundary. ``max_build_loops`` caps build<->verify retries before - the task parks. ``build_loops`` is the loops already consumed for this task - (the coordinator threads it through state). + the task parks. + + ``build_loops`` is an INITIAL FALLBACK ONLY. The loop count that actually + bounds the build<->verify cycle is DURABLE per-task state read from + ``state["build_loops"]`` at node-run time, because a single ``VerifierConfig`` + is shared across every task at wiring time and never advances (LOGIC-RACE-01: + reading the count from this shared config meant the park guard never fired and + a perpetually-FAILing task looped BUILD->DISPATCH->VERIFY forever). The + verifier writes the incremented count back into the returned partial state so + the checkpointer carries it to the NEXT VERIFY. This field is consulted only + when state carries no ``build_loops`` (e.g. a unit harness driving the node + directly). """ - expected_run_id: str + expected_run_id: str | None = None allowed_scope: list[str] | None = None max_build_loops: int = DEFAULT_MAX_BUILD_LOOPS build_loops: int = 0 @@ -148,9 +164,14 @@ def verifier_node( Decision flow (§3.3.2 boundary #4 + §3.3 autonomy bounds): - 1. Call :func:`agent_team.ci_gate.evaluate_ci_gate` with the candidate diff, - the ledger hash (``diff_hash``), the authenticated ``ci_results``, the - ``expected_run_id``, and the task's ``allowed_scope``. + 1. Resolve the PER-TASK expected run id from ``state["run_id"]`` (the id the + dispatch node captured for THIS task), falling back to + ``config.expected_run_id`` only when state carries none. Call + :func:`agent_team.ci_gate.evaluate_ci_gate` with the candidate diff, the + ledger hash (``diff_hash``), the authenticated ``ci_results``, that + per-task expected run id, and the task's ``allowed_scope``. A ``None`` + effective run id BLOCKs (anti-substitution: the verdict has nothing to + bind to), never a vacuous pass. 2. PASS -> advance to DONE (draft PR). The advisor is NOT consulted. 3. FAIL -> consult the fix-advisor for a hint, then loop back to BUILD — unless ``build_loops`` has reached ``max_build_loops``, in which case @@ -162,6 +183,29 @@ def verifier_node( ledger_hash = state.get("diff_hash") ci_results = state.get("ci_results") + # Per-task binding: the gate must compare CI's run_id against the id THIS + # task dispatched (persisted by the dispatch node as ``state["run_id"]``), + # not a static wiring-time constant. State wins; the config value is only a + # fallback for harnesses that drive the node without a per-task run_id. A + # blank/None effective run id is left as None so the gate BLOCKs. + state_run_id = state.get("run_id") + expected_run_id = ( + state_run_id + if isinstance(state_run_id, str) and state_run_id + else config.expected_run_id + ) + + # Per-task build-loop budget: the count that bounds the build<->verify cycle + # is DURABLE per-task state (``state["build_loops"]``), NOT the shared + # wiring-time config. Reading it from state is the LOGIC-RACE-01 fix: the + # shared ``VerifierConfig.build_loops`` never advanced, so the park guard + # never fired and a perpetually-FAILing task looped forever. ``config`` is + # only a fallback for a harness that drives the node without per-task state. + state_build_loops = state.get("build_loops") + current_build_loops = ( + state_build_loops if isinstance(state_build_loops, int) else config.build_loops + ) + if not isinstance(candidate_diff, str): # No diff to verify is itself a refuse-to-proceed: park for a human # rather than declaring anything. (A builder must have produced a diff @@ -169,7 +213,7 @@ def verifier_node( gate_result = GateResult( decision=GateDecision.BLOCK, reasons=["no candidate_diff present in state to verify"], - run_id=config.expected_run_id, + run_id=expected_run_id, diff_hash=ledger_hash, ci_conclusion=None, ) @@ -178,7 +222,7 @@ def verifier_node( candidate_diff=candidate_diff, ledger_hash=ledger_hash, ci_result=ci_results, - expected_run_id=config.expected_run_id, + expected_run_id=expected_run_id, allowed_scope=config.allowed_scope, ) @@ -188,7 +232,7 @@ def verifier_node( if gate_result.decision is GateDecision.PASS: verdict = _verdict( - gate_result, next_phase=Phase.DONE, build_loops=config.build_loops + gate_result, next_phase=Phase.DONE, build_loops=current_build_loops ) return { "status": TaskStatus.DONE.value, @@ -203,7 +247,8 @@ def verifier_node( fix_hint = _fix_advisor(gate_result, state) if gate_result.decision is GateDecision.FAIL: - next_loops = config.build_loops + 1 + # Count this failed loop against the DURABLE per-task budget read above. + next_loops = current_build_loops + 1 if next_loops >= config.max_build_loops: # Exhausted the build-loop budget: hold rather than spin (§3.3 #6). verdict = _verdict( @@ -220,6 +265,7 @@ def verifier_node( "current_phase": Phase.PARKED.value, "review_verdicts": [verdict], "ci_results": annotated_ci, + "build_loops": next_loops, "updated_at": verdict["at"], } verdict = _verdict( @@ -228,11 +274,15 @@ def verifier_node( build_loops=next_loops, fix_hint=fix_hint, ) + # Persist the incremented count into DURABLE state so the NEXT VERIFY + # (after BUILD->DISPATCH) sees it and the budget actually advances. The + # checkpointer carries it because it is a PipelineState key. return { "status": TaskStatus.ACTIVE.value, "current_phase": Phase.BUILD.value, "review_verdicts": [verdict], "ci_results": annotated_ci, + "build_loops": next_loops, "updated_at": verdict["at"], } @@ -241,7 +291,7 @@ def verifier_node( verdict = _verdict( gate_result, next_phase=Phase.PARKED, - build_loops=config.build_loops, + build_loops=current_build_loops, fix_hint=fix_hint, ) return { diff --git a/agent-team/agent_team/resume_worker.py b/agent-team/agent_team/resume_worker.py index 8dd0845..e0d3a45 100644 --- a/agent-team/agent_team/resume_worker.py +++ b/agent-team/agent_team/resume_worker.py @@ -288,6 +288,59 @@ class ResumeWorker: graph_result=graph_result, ) + def resume_ci( + self, + *, + thread_id: str, + run_id: str, + answer: Any, + ) -> ResumeResult: + """Resume a CI machine-gate (VERIFY awaiting CI), single-flight + guarded. + + The VERIFY node's async CI-wait is a *machine* gate, not a human one: + it suspends with ``{"awaiting_ci": True, "run_id": }`` and has no + ledger ``question``/``turn`` to thread through the first-answer-wins + flip. So this path turn-guards on the CI marker instead: under the same + per-``thread_id`` lock that serialises human resumes (so a CI resume and + a redelivered CI resume for the same task never race), it re-reads the + *live* checkpoint and confirms the thread is STILL suspended at VERIFY + awaiting THIS ``run_id`` (via + :func:`agent_team.ci_watcher.snapshot_awaiting_ci_run_id`). Only then does + it invoke ``Command(resume=answer)``. If the thread already advanced + (resumed / parked / done — no awaiting-CI interrupt), or is now awaiting + a *different* run, it skips (:attr:`ResumeOutcome.STALE`) so a + double-resume can never corrupt the durable state. + + Returns :attr:`ResumeOutcome.RESUMED` when the resume applied, else + :attr:`ResumeOutcome.STALE` (there is no ledger question to supersede on + this machine gate). ``question_id`` is reported as ``""`` since none + exists. + """ + from agent_team.ci_watcher import snapshot_awaiting_ci_run_id + + lock = self._lock_for(thread_id) + with lock: + config = _thread_config(thread_id) + snapshot = self._graph.get_state(config) + awaited = snapshot_awaiting_ci_run_id(snapshot) + if awaited != run_id: + # Already advanced past this wait (resumed/parked/done) or now + # awaiting a different run: skip rather than double-apply. + return ResumeResult( + outcome=ResumeOutcome.STALE, + thread_id=thread_id, + question_id="", + turn=0, + ) + graph_result = self._graph.invoke(build_resume_command(answer), config) + return ResumeResult( + outcome=ResumeOutcome.RESUMED, + thread_id=thread_id, + question_id="", + turn=0, + graph_result=graph_result, + ) + def recover_pending_resumes(self) -> list[ResumeResult]: """Restart sweep: re-enqueue resumes for durable ``answered`` rows. diff --git a/agent-team/agent_team/task_model.py b/agent-team/agent_team/task_model.py index c13f716..5391f95 100644 --- a/agent-team/agent_team/task_model.py +++ b/agent-team/agent_team/task_model.py @@ -98,6 +98,22 @@ class TaskRecord: candidate_diff: str | None = None diff_hash: str | None = None ci_results: dict[str, Any] | None = None + # P3 box-side build->dispatch->verify plumbing (§3.3.2). Set by the DISPATCH + # node when it triggers the apply/verify CI run: ``run_id`` is the GitHub + # Actions run id the verifier's read-only fetcher polls + the pure-code gate + # binds its verdict to; ``ci_correlation_tag`` is the per-dispatch nonce + # carried as a workflow input so the run_id poll matches THIS task's exact + # run (anti-race / anti-replay); ``dispatched_at`` bounds the CI-watch + # timeout. All None until a task reaches DISPATCH on the live P3 path. + run_id: str | None = None + ci_correlation_tag: str | None = None + dispatched_at: str | None = None + # Count of build<->verify loops already consumed for THIS task (§3.3 #6). + # Durable per-task state (NOT the shared wiring-time VerifierConfig): the + # verifier reads it from state, increments on a recoverable gate FAIL, and + # parks once it reaches VerifierConfig.max_build_loops so a perpetually- + # failing task can never loop BUILD->DISPATCH->VERIFY forever (LOGIC-RACE-01). + build_loops: int = 0 transport: str = "" created_at: str | None = None updated_at: str | None = None @@ -132,6 +148,20 @@ class PipelineState(TypedDict, total=False): candidate_diff: str | None diff_hash: str | None ci_results: dict[str, Any] | None + # P3 box-side build->dispatch->verify plumbing (mirrors TaskRecord). ``run_id`` + # is the dispatched apply/verify Actions run id the verifier fetches + the + # gate binds to; ``ci_correlation_tag`` is the per-dispatch nonce carried as a + # workflow input so the poll matches THIS task's run; ``dispatched_at`` bounds + # the CI-watch timeout. + run_id: str | None + ci_correlation_tag: str | None + dispatched_at: str | None + # Build<->verify loops already consumed for THIS task (mirrors TaskRecord). + # The verifier threads it through DURABLE state — increments on a recoverable + # gate FAIL, parks at VerifierConfig.max_build_loops — so the build-loop + # budget is real (LOGIC-RACE-01: it was previously read from the shared + # wiring-time config and never advanced). + build_loops: int transport: str created_at: str | None updated_at: str | None @@ -164,6 +194,10 @@ def task_from_dict(data: dict[str, Any]) -> TaskRecord: candidate_diff=data.get("candidate_diff"), diff_hash=data.get("diff_hash"), ci_results=data.get("ci_results"), + run_id=data.get("run_id"), + ci_correlation_tag=data.get("ci_correlation_tag"), + dispatched_at=data.get("dispatched_at"), + build_loops=data.get("build_loops", 0), transport=data.get("transport", ""), created_at=data.get("created_at"), updated_at=data.get("updated_at"), diff --git a/agent-team/ci/README.md b/agent-team/ci/README.md index 2964a2e..1801fc6 100644 --- a/agent-team/ci/README.md +++ b/agent-team/ci/README.md @@ -121,6 +121,42 @@ workflow_dispatch (task_id, diff_artifact_name, expected_diff_hash, declared_sco its own grant explicitly. The trigger is `workflow_dispatch` only — the patch never runs in a context carrying write or secret scope. +## Run-name correlation (box dispatcher → run_id) + +The box-side dispatcher (`agent_team/dispatcher.py`) triggers this workflow with +`gh workflow run`, which does **not** return the resulting run id. The dispatcher +must still resolve that `run_id` so the verifier's read-only CI-result fetcher can +poll the correct run (a `None`/unfound run_id fails closed → the verifier gate +BLOCKs / the task parks; never a vacuous pass). The correlation key is the +workflow **run name**: + +```yaml +run-name: "agent-team-apply ${{ inputs.task_id }}" +``` + +Why the run name and not the workflow input or the head branch: + +- `gh run list --json name` exposes the run name, but **workflow inputs are not + queryable** via the run list, so the `task_id` cannot be matched on the input. +- A `workflow_dispatch` run reports against the **`main` ref**, not the dispatch + head branch, so the branch is not a usable discriminator either. + +So the dispatcher polls `gh run list` read-only, matches the row whose `name` +equals `agent-team-apply ` (mirrored in code by `dispatcher.run_name_for`), +and bounds the match to runs created after the dispatch watermark. The workflow's +`concurrency` group already guarantees a single in-flight run per `task_id`, so +the run name plus the dispatched-at floor identify the dispatched run +unambiguously even under many simultaneous dispatches; the pure +`dispatcher.select_run_id` then applies the anti-stale tie-break (skip a superseded +cancelled run sharing the name, prefer the later-created `databaseId`). + +This `run-name` is **additive**: it adds no job, permission, secret, or trigger, +and changes no privileged step — it only surfaces the dispatching task's id for +correlation. Because the file is nonetheless a CI trust-boundary workflow +(untrusted-input handling), the edit is **flagged for the C1 re-run of +`/sh-security-review` AND the mandatory GPT-4.1 cross-review** before it ships on +this branch, matching the in-YAML `P3-BOX-INTEGRATION` comment. + ## SHA-pinned actions (handbook Pinning Principle, §3.3.2) Every third-party action is pinned to a full commit SHA with the human-readable diff --git a/agent-team/docs/P3-PHASE0-DESIGN.md b/agent-team/docs/P3-PHASE0-DESIGN.md new file mode 100644 index 0000000..91f33fb --- /dev/null +++ b/agent-team/docs/P3-PHASE0-DESIGN.md @@ -0,0 +1,87 @@ +# P3 Phase 0 — box-side build→dispatch→verify integration (design decision) + +Branch: `feat/agent-team-p3-box-integration`. This note records the load-bearing +design decision for Phase 0 so the security/cross-review gates and the parallel +WebUI session have the rationale in-tree. It is the result of the `/sh-plan-review` +loop (3 rounds, GPT-4.1) + a code-level investigation of the as-built dispatcher, +ci_fetcher, and coordinator execution model. + +## Context (verified on-disk, 2026-06-23) + +The CI apply/verify workflow is **already live + provisioned** (the `agent-apply` +environment, the `AGENT_APPLY_APP_*` secrets, and dispatched runs all exist as of +2026-06-22). What is **not** built is the box-side integration that makes a task +flow through BUILD → (trigger CI) → VERIFY automatically: + +- `dispatcher.py` fires `gh workflow run` but never captures the resulting run id. +- `dispatch_invoker` returns `{}` — nothing writes `state["run_id"]`. +- `ci_fetcher` reads `state["run_id"]` (so it always fails closed → gate BLOCKs). +- The graph orders BUILD → VERIFY → DISPATCH, but VERIFY needs a CI conclusion + that only exists *after* DISPATCH triggers CI. (semantic inversion) + +## Decision 1 — node order: BUILD → DISPATCH → VERIFY (reorder) + +DISPATCH triggers the CI run and must run *before* VERIFY reads its conclusion. +The P3 subgraph is reordered accordingly (graph.py + build_verify_subgraph.py). +The reorder adds **no new graph nodes** (BUILD/DISPATCH/VERIFY already exist), so +the parallel WebUI branch's graph introspection + NODE_META coverage are +unaffected; only edge wiring changes. + +## Decision 2 — CI wait: async resume-on-CI-complete, NOT a blocking poll + +A CI run takes ~7 min. The coordinator is a single durable daemon (LangGraph +`interrupt()`/resume + a `tick()` maintenance sweep). A multi-minute *blocking* +VERIFY node would stall the tick loop and every other task. The durable +interrupt/resume machinery already exists for exactly the "external event resumes +a suspended task" shape (the Slack responder; the deadline timer). So: + + BUILD → DISPATCH (push branch, trigger CI, capture run_id, suspend) + → [CI-watcher resumes on terminal conclusion] → VERIFY (read result, gate) + +DISPATCH captures `run_id` + `dispatched_at` into state and the task suspends. A +new **CI-watcher** (a `tick()`-driven sweep, mirroring `deadline_timer`) polls the +in-flight `run_id`s read-only and resumes each task once its run reaches a +terminal conclusion (or its `dispatched_at` + timeout elapses → park). VERIFY then +reads the authenticated conclusion via the existing read-only fetcher and the +pure-code gate decides pass/fail. The LLM remains a fix-proposer only. + +## Decision 3 — run_id capture is poll-based + correlation-tagged (anti-race) + +`gh workflow run` does not return a run id. The dispatcher polls +`gh run list --workflow … --json databaseId,headBranch,createdAt,event` filtered +to this task's **unique per-dispatch head branch** + a per-dispatch +**correlation tag** (a nonce carried as a workflow input and echoed in the run), +bounded to runs created after the dispatch timestamp. This unambiguously matches +the dispatched run even with multiple tasks or rapid re-dispatch. A `None`/unfound +run_id fails closed (the gate BLOCKs / the task parks) — never a vacuous pass. + +## Decision 4 — per-task `expected_run_id` + +`VerifierConfig.expected_run_id` was a static wiring-time constant. It is now +resolved per-task from `state["run_id"]` (the id the dispatcher captured), so the +gate binds each task's verdict to its own dispatched run and rejects a substituted +run id. + +## Decision 5 — fail-safe serve default + +Per Adam's decision the bound P3 wiring becomes the new `serve` default. Because +the wiring factories are called eagerly at graph-build, the binding is wrapped so +a missing `AGENT_TEAM_REPO_OWNER`/`_NAME` / CI-read token degrades to the INERT P3 +path (task parks, one WARNING + a `#agent-team` inert-mode notice) — never a +`RuntimeError` at serve-start that would crash-loop the daemon. + +## Out of scope (tracked follow-ups) + +- The §4.3 box-native diff transport (signed artifact / branch-only token) that + would let the always-on box trigger CI without operator credentials. Until then, + the branch push + `gh workflow run` use operator-host credentials (the box holds + no standing write token). +- Tier-3 fixer off `--dry-run`; the cross-plane checker→draft-PR loop. + +## Coordination with the WebUI branch (`feature/agent-team-webui-makeover`) + +Shared files: `graph.py` (they wrap nodes via `_instrument`; we reorder edges), +`coordinator.py` (both edit the `build_graph(...)` call block). WebUI merges +first; this branch rebases onto the new `main` before deploy. Wrapper composition: +`instrument(failsafe(node))` so a fail-safe park is still logged. The +`project_r720_agent_team` memory fix is owned by the WebUI session. diff --git a/agent-team/run-team.py b/agent-team/run-team.py index 58326f3..db4ac55 100644 --- a/agent-team/run-team.py +++ b/agent-team/run-team.py @@ -569,6 +569,110 @@ def _build_notifiers( return notify, alarm_hook +# The branch namespace the dispatcher pushes its apply/verify draft PRs under +# (see :func:`agent_team.dispatcher.apply_branch_name` -> ``agent-team/apply/``). +# The runaway/stale monitor is scoped to THIS namespace so it only ever surfaces +# agent-team's own draft PRs, never an unrelated human draft PR in the repo. +_DRAFT_PR_HEAD_PREFIX = "agent-team/apply/" + +# gh's draft-PR enumeration must never block the daemon's tick() loop. A read-only +# ``gh pr list`` is one GET; bound it so a hung gh invocation parks the sweep +# rather than the whole coordinator. +_DRAFT_PR_LIST_TIMEOUT_S = 30 + + +def _default_draft_pr_provider(*, owner: str, repo: str) -> "Callable[[], list[Any]]": + """Build the production READ-ONLY draft-PR provider for the A4 monitor. + + Returns a zero-arg callable that enumerates the currently-open *agent-team* + draft PRs via a single read-only ``gh pr list`` (one GET; it NEVER writes, + closes, or dispatches anything) and maps each into a + :class:`agent_team.draft_pr_monitor.DraftPr` snapshot the monitor consumes. + + The query is scoped to the ``agent-team/apply/`` head namespace + (:data:`_DRAFT_PR_HEAD_PREFIX`) so it only ever sees the dispatcher's own + apply/verify draft PRs — never an unrelated human draft PR. ``--json`` pulls + exactly the three fields the monitor keys off (``number`` / ``createdAt`` -> + ``opened_at`` for the runaway window, ``updatedAt`` -> ``updated_at`` for + staleness). + + Mirrors :func:`agent_team.ci_watcher.default_ci_poller`: a thin closure over + ``owner`` / ``repo`` that fails closed — a non-zero ``gh`` exit, a timeout, or + unparseable JSON yields an empty snapshot (the monitor then no-ops this pass) + rather than raising, so a transient gh hiccup never breaks the tick loop. (The + coordinator's ``_draft_pr_monitor_sweep`` ALSO swallows provider errors, so + this is belt-and-suspenders.) + """ + repo_slug = f"{owner}/{repo}" + + def provider() -> list[Any]: + import json as _json + import subprocess # noqa: PLC0415 - deferred so import needs no gh + + from agent_team.draft_pr_monitor import DraftPr # noqa: PLC0415 + + try: + proc = subprocess.run( # noqa: S603 - args are a fixed, non-shell list + [ + "gh", + "pr", + "list", + "--repo", + repo_slug, + "--draft", + "--state", + "open", + "--search", + f"head:{_DRAFT_PR_HEAD_PREFIX}", + "--json", + "number,createdAt,updatedAt", + "--limit", + "100", + ], + check=True, + capture_output=True, + text=True, + timeout=_DRAFT_PR_LIST_TIMEOUT_S, + ) + except ( + subprocess.CalledProcessError, + subprocess.TimeoutExpired, + OSError, + ): + _LOG.warning( + "draft-pr-monitor: read-only `gh pr list` enumeration failed; " + "treating as no open draft PRs this pass", + exc_info=True, + ) + return [] + + try: + rows = _json.loads(proc.stdout or "[]") + except ValueError: + _LOG.warning( + "draft-pr-monitor: `gh pr list` returned unparseable JSON; " + "treating as no open draft PRs this pass", + exc_info=True, + ) + return [] + + snapshots: list[Any] = [] + for row in rows: + number = row.get("number") + if number is None: + continue + snapshots.append( + DraftPr( + number=int(number), + opened_at=row.get("createdAt"), + updated_at=row.get("updatedAt"), + ) + ) + return snapshots + + return provider + + def _build_coordinator(args: argparse.Namespace) -> Any: """Construct a :class:`Coordinator` for the ``start`` / ``serve`` commands. @@ -587,6 +691,7 @@ def _build_coordinator(args: argparse.Namespace) -> Any: default_clarify_node_factory, default_plan_node_factory, default_review_wiring, + failsafe_production_p3_wiring, ) transport = _build_transport(args) @@ -604,6 +709,43 @@ def _build_coordinator(args: argparse.Namespace) -> Any: # silent (notify None) so import + ledger commands need no token. notify, alarm_hook = _build_notifiers(args) + # P3 fail-safe serve default (design Decision 5): the bound build→verify + + # dispatch wiring is now the production ``serve`` default, but it must NEVER + # crash-loop serve-start. ``failsafe_production_p3_wiring`` resolves the env + # ONCE: configured -> the live P3 pair; unconfigured -> (None, None) inert (a + # task reaching P3 parks), with one WARNING + one #agent-team inert notice + # via the lifecycle ``notify`` sink. Scoped to ``serve`` (the daemon): the + # one-shot ``start`` / ``intake-*`` paths never auto-bind P3 — they run to the + # first human gate and exit, well short of BUILD/VERIFY. + build_verify_wiring = None + dispatch_node_wiring = None + if getattr(args, "command", None) == "serve": + build_verify_wiring, dispatch_node_wiring = failsafe_production_p3_wiring( + notify=notify + ) + + # CI-watcher seams (design §4 Decision 2 — async resume-on-CI-complete). ONLY + # wired when ``failsafe_production_p3_wiring`` returned a LIVE pair (a + # configured box): on the inert/unconfigured box both stay None, so the + # tick() CI sweep is a NO-OP and there is no behaviour change. The poller is + # the read-only default (one GET per run, never a write); the provider is the + # durable enumerator bound to THIS coordinator below (post-construction, so it + # can close over the just-built coordinator); the timeout is a sane default + # (30 min — generous headroom over the ~7-min CI run before a stuck run + # parks). Without these, a task that dispatches and suspends at VERIFY would + # wait forever — the async-resume gap this closes. + ci_poller = None + ci_timeout = None + if build_verify_wiring is not None and dispatch_node_wiring is not None: + from datetime import timedelta + + from agent_team.ci_watcher import default_ci_poller + + owner = os.environ.get("AGENT_TEAM_REPO_OWNER", "").strip() + repo = os.environ.get("AGENT_TEAM_REPO_NAME", "").strip() + ci_poller = default_ci_poller(owner=owner, repo=repo) + ci_timeout = timedelta(minutes=30) + # Production runs the full P2 graph: the wrapped real planner + the bound # GPT-4.1 review loop (Plane-2 depth-first). These factories are lazy and # only build/bind the model seams when a task actually runs. @@ -617,9 +759,32 @@ def _build_coordinator(args: argparse.Namespace) -> Any: context_provider=context_provider ), review_wiring=default_review_wiring, + build_verify_wiring=build_verify_wiring, + dispatch_node_wiring=dispatch_node_wiring, notify=notify, alarm_hook=alarm_hook, + ci_poller=ci_poller, + ci_timeout=ci_timeout, ) + # Bind the durable CI-pending provider to THIS coordinator (only on the live + # P3 path — ``ci_poller`` is the live-pair signal). It enumerates the durable + # threads suspended at VERIFY awaiting CI so the watcher has real tasks to + # poll; bound post-construction so it can reference the just-built + # coordinator. Left unbound on the inert box, the CI sweep stays a NO-OP. + if ci_poller is not None: + coordinator._ci_pending_provider = coordinator._enumerate_ci_pending + # Draft-PR runaway/stale monitor seam (P3 A4). Gated on the SAME live-pair + # signal as the CI watcher (``ci_poller`` is set ⇔ ``_p3_env_is_configured`` + # via ``failsafe_production_p3_wiring``), so on the inert/unconfigured box it + # stays None and ``_draft_pr_monitor_sweep`` is a NO-OP (no behaviour change). + # On the live box it binds a read-only ``gh pr list`` enumerator of the open + # ``agent-team/apply/`` draft PRs so the sweep has a real snapshot to ALARM / + # remind on — without this the wired sweep would always see no PRs. Bound + # post-construction to mirror ``_ci_pending_provider``. + if ci_poller is not None: + coordinator._draft_pr_provider = _default_draft_pr_provider( + owner=owner, repo=repo + ) # WS2: an allowlisted Slack /new-task starts a task on THIS coordinator. Set # post-construction (the adapter closes over the just-built coordinator), and # before serve() builds the listener. AUTHZ-01 (owner allowlist) gates this @@ -1036,6 +1201,118 @@ def _cmd_force_resume(args: argparse.Namespace, *, out: Any) -> int: return 1 +def _cmd_dispatch(args: argparse.Namespace, *, out: Any) -> int: + """Operator-initiated dispatch of a built diff into org CI (P3, option-b). + + The box holds NO write token (read-only by design), so its in-graph DISPATCH + node fail-closes/parks. This is the operator verb that completes the dispatch + with WRITE creds: it reads the task's ``candidate_diff`` + declared scope + (from the ledger checkpoint, or from ``--diff``/``--scope`` files), pushes the + head branch and fires the apply/verify ``workflow_dispatch`` via + :func:`agent_team.dispatcher.dispatch_apply_verify`, then prints the located + run id. Run it where a WRITE-capable ``GH_TOKEN`` is available — the operator + host, or the box with a JUST-IN-TIME operator token in the env (never stored + in ``secrev.env``); the box stays read-only at rest. + + Owner/repo/base resolve from ``--owner``/``--repo``/``--base`` or the + ``AGENT_TEAM_REPO_OWNER``/``_NAME``/``_BASE_BRANCH`` env vars. With + ``--write-back`` the located ``run_id`` is written into the task checkpoint so + the box's VERIFY can bind to it. + """ + import os + + from agent_team.dispatcher import dispatch_apply_verify + + owner = args.owner or os.environ.get("AGENT_TEAM_REPO_OWNER", "") + repo = args.repo or os.environ.get("AGENT_TEAM_REPO_NAME", "") + base = args.base or os.environ.get("AGENT_TEAM_BASE_BRANCH") or "main" + if not owner or not repo: + print( + "dispatch: owner/repo required (pass --owner/--repo or set " + "AGENT_TEAM_REPO_OWNER/_NAME)", + file=sys.stderr, + ) + return 2 + + # Resolve the diff + declared scope: explicit files win; else read the task's + # checkpointed PipelineState (candidate_diff + plan.scope). + diff_text: str | None = ( + Path(args.diff).read_text(encoding="utf-8") if args.diff else None + ) + declared_scope: str | None = ( + Path(args.scope).read_text(encoding="utf-8") if args.scope else None + ) + if diff_text is None or declared_scope is None: + from agent_team.graph import ( + build_graph, + build_sqlite_checkpointer, + thread_config, + ) + + with build_sqlite_checkpointer(args.db) as saver: + graph = build_graph(saver) + snap = graph.get_state(thread_config(args.thread_id)) + state = dict(snap.values or {}) + if diff_text is None: + diff_text = state.get("candidate_diff") + if declared_scope is None: + plan = state.get("plan") or {} + scope_list = plan.get("scope") or [] if isinstance(plan, dict) else [] + declared_scope = "\n".join(str(s) for s in scope_list if s) + + if not diff_text or not str(diff_text).strip(): + print( + f"dispatch: no candidate_diff for task {args.thread_id} " + "(pass --diff, or the task has not built a diff yet)", + file=sys.stderr, + ) + return 1 + + result = dispatch_apply_verify( + owner=owner, + repo=repo, + task_id=args.thread_id, + diff_text=diff_text, + declared_scope=declared_scope or "", + base=base, + ) + print( + f"dispatched task {args.thread_id} -> {owner}/{repo} " + f"(head={result.inputs.head_branch}, run_id={result.run_id}, " + f"dispatched_at={result.dispatched_at})", + file=out, + ) + if result.run_id is None: + print( + "dispatch: workflow fired but run_id could not be correlated; " + "VERIFY fails closed until a run_id is set", + file=sys.stderr, + ) + elif args.write_back: + from agent_team.graph import ( + build_graph, + build_sqlite_checkpointer, + thread_config, + ) + + with build_sqlite_checkpointer(args.db) as saver: + graph = build_graph(saver) + graph.update_state( + thread_config(args.thread_id), + { + "run_id": result.run_id, + "dispatched_at": result.dispatched_at, + "ci_correlation_tag": result.correlation_tag, + }, + ) + print( + f"dispatch: wrote run_id={result.run_id} into the task checkpoint " + "(--write-back)", + file=out, + ) + return 0 if result.run_id is not None else 1 + + # Statuses an operator treats as "parked context": a task whose only pending # question is no longer open may be parked (answered-but-unresumed, expired, or # superseded). ``open`` is excluded — that is the live-waiting view (default @@ -1166,6 +1443,45 @@ def build_parser() -> argparse.ArgumentParser: ) p_resume.set_defaults(func=_cmd_force_resume) + p_dispatch = sub.add_parser( + "dispatch", + help=( + "operator-initiated dispatch of a task's built diff into org CI (P3, " + "option-b) — needs a WRITE-capable GH_TOKEN in the env" + ), + ) + p_dispatch.add_argument( + "thread_id", help="the task thread_id whose candidate_diff to dispatch" + ) + p_dispatch.add_argument( + "--owner", default=None, help="repo owner (default $AGENT_TEAM_REPO_OWNER)" + ) + p_dispatch.add_argument( + "--repo", default=None, help="repo name (default $AGENT_TEAM_REPO_NAME)" + ) + p_dispatch.add_argument( + "--base", + default=None, + help="base branch (default $AGENT_TEAM_BASE_BRANCH or main)", + ) + p_dispatch.add_argument( + "--diff", + default=None, + help="path to a unified-diff file (overrides the ledger candidate_diff)", + ) + p_dispatch.add_argument( + "--scope", + default=None, + help="path to a newline-separated declared-scope file (overrides plan.scope)", + ) + p_dispatch.add_argument( + "--write-back", + dest="write_back", + action="store_true", + help="write the located run_id into the task checkpoint so VERIFY binds to it", + ) + p_dispatch.set_defaults(func=_cmd_dispatch) + p_start = sub.add_parser( "start", help="intake: start one task and run it to the first human gate", diff --git a/agent-team/scripts/assert_no_write_token.py b/agent-team/scripts/assert_no_write_token.py new file mode 100644 index 0000000..d79eaf0 --- /dev/null +++ b/agent-team/scripts/assert_no_write_token.py @@ -0,0 +1,318 @@ +#!/usr/bin/env python3 +"""assert_no_write_token.py — fail closed if a GitHub *write* token is on the box. + +The agent-team apply/verify path mints its ``pull-requests: write`` GitHub App +installation token **inside the CI runner** (via SHA-pinned +``actions/create-github-app-token``), from the ``AGENT_APPLY_APP_ID`` / +``AGENT_APPLY_APP_PRIVATE_KEY`` **Actions secrets**. By design the always-on +R720 box holds **no standing write credential**: it triggers CI with the +operator's host ``gh`` auth and reads results with a *read-only* token +(``AGENT_TEAM_CI_READ_TOKEN`` / ``GITHUB_TOKEN``). See ``ci/README.md`` and +``docs/P3-PHASE0-DESIGN.md`` ("the box holds no standing write token"). + +This audit asserts that invariant. It is runnable both on the box (as a +provisioning/runtime self-check) and in CI (as a regression guard). It: + + 1. scans the live process environment (``os.environ``) and the coordinator's + environment-derived config for token-shaped variables that would grant + ``pull-requests: write`` / ``contents: write``; + 2. greps the box's local credential files (``~/secrev.env`` and + ``~/orchestrator/.env`` by default; paths are configurable) for the App id + and for any PEM private-key header (RSA / EC / OPENSSH / PKCS#8); + +and **exits non-zero with a clear message** if any are found, or exits ``0`` +with a short summary otherwise. + +Usage:: + + python -m scripts.assert_no_write_token + python scripts/assert_no_write_token.py --env-file ~/secrev.env --env-file ~/x.env + +Exit codes: + 0 no write-shaped token / App secret / private key found + 1 at least one finding (the box is mis-provisioned — remediate before deploy) +""" + +from __future__ import annotations + +import argparse +import os +import re +import sys +from collections.abc import Mapping, Sequence +from pathlib import Path + +# --------------------------------------------------------------------------- # +# What "must never live on the box" looks like. +# --------------------------------------------------------------------------- # + +# The GitHub App that holds ``pull-requests: write`` lives ONLY as Actions +# secrets. Its id and private key must never appear on the box (env or files). +APP_ID_ENV = "AGENT_APPLY_APP_ID" +APP_PRIVATE_KEY_ENV = "AGENT_APPLY_APP_PRIVATE_KEY" + +# Env vars that are explicitly the App's write credentials. +_FORBIDDEN_ENV_VARS: frozenset[str] = frozenset( + { + APP_ID_ENV, + APP_PRIVATE_KEY_ENV, + } +) + +# Env vars that are KNOWN-GOOD read-only / non-write and must NOT be flagged by +# the heuristic name match below (they are tokens, but read-only by contract). +_ALLOWED_TOKEN_ENV_VARS: frozenset[str] = frozenset( + { + "AGENT_TEAM_CI_READ_TOKEN", # read-only CI-result fetcher token + "AGENT_TEAM_API_TOKEN", # read-only dashboard/API bearer + "SLACK_APP_TOKEN", # Slack socket-mode app-level token (not GitHub) + "SLACK_BOT_TOKEN", # Slack bot token (not GitHub) + "CLAUDE_CODE_OAUTH_TOKEN", # Anthropic subscription OAuth (not GitHub) + "ANTHROPIC_API_KEY", # model API key (not GitHub) + } +) + +# A name-shaped heuristic for "this looks like a GitHub *write* token". We +# deliberately scope to GitHub-write shapes so the generic read-only +# ``GITHUB_TOKEN`` fallback (a runtime read token, not a standing write secret) +# is not a false positive while the App's write material always is. +_WRITE_TOKEN_NAME_RE = re.compile( + r"(?:^|_)(?:GH|GITHUB)_(?:APP|PAT|WRITE|APPLY)_?(?:TOKEN|KEY|PRIVATE_KEY)?", + re.IGNORECASE, +) + +# Value shapes that indicate a GitHub write-capable credential regardless of the +# var's name. ``ghp_`` (classic PAT) and ``github_pat_`` (fine-grained PAT) can +# both carry write scopes; an installation token (``ghs_``) is write-capable; a +# user-to-server token (``ghu_``) acts with the user's write access; and a refresh +# token (``ghr_``) mints fresh write-capable user-to-server tokens. All are +# write-risk material that must not live on the box. +_WRITE_TOKEN_VALUE_RE = re.compile( + r"\b(?:ghp_|ghs_|ghu_|ghr_|github_pat_)[A-Za-z0-9_]{20,}\b" +) + +# Any PEM private-key header (RSA / EC / OPENSSH / generic PKCS#8). The App's +# private key is a PEM block; finding ANY private key in a box credential file +# is a finding. +_PRIVATE_KEY_RE = re.compile(r"-----BEGIN (?:[A-Z0-9]+ )*PRIVATE KEY-----") + +# Default box credential files to grep. Configurable via --env-file / the +# ``ASSERT_NO_WRITE_TOKEN_ENV_FILES`` env var so tests use temp files. +_DEFAULT_ENV_FILES: tuple[str, ...] = ("~/secrev.env", "~/orchestrator/.env") + + +# --------------------------------------------------------------------------- # +# Scanners. Each returns a list of human-readable finding strings. +# --------------------------------------------------------------------------- # + + +def scan_environ(environ: Mapping[str, str], *, app_id: str | None = None) -> list[str]: + """Scan a process environment for write-shaped GitHub token material. + + Flags (a) the explicit App-credential env vars, (b) any var whose *name* + matches the GitHub-write heuristic, (c) any var whose *value* carries a + write-capable GitHub token prefix, (d) any var whose *value* is a PEM + private-key block (the App key exported under a benign name), and (e) any var + whose *value* contains the configured App id. Known read-only tokens are + exempt from the *name* and write-token *value* heuristics, but the PEM and + App-id value checks below apply to EVERY var (including the allowlisted ones + and ``GITHUB_TOKEN``) — a private key or the App id can never legitimately + sit in any env var, so those checks are never name-exempted. This mirrors + :func:`scan_env_file`, which greps file contents for the same App-id / + private-key material regardless of the var name carrying it. + """ + findings: list[str] = [] + for name, value in environ.items(): + # The PEM private-key VALUE check and the configured App-id VALUE check + # are NOT name-exempt: the App's private key exported under a benign name + # (e.g. GH_APP_KEY set to a "BEGIN ... PRIVATE KEY" PEM block) and the + # configured App id parked in any var are both forbidden material, even in an + # otherwise read-only/allowlisted var. Check them up front, before the + # allowlist short-circuits the name/token-prefix heuristics. + if value and _PRIVATE_KEY_RE.search(value): + findings.append( + f"environment variable {name!r} holds a PEM private-key block " + "(-----BEGIN ... PRIVATE KEY-----); no private key may live on " + "the box (the App key must live ONLY as an Actions secret)" + ) + if app_id and value and app_id in value: + findings.append( + f"environment variable {name!r} contains the configured App id " + "value; the App id must not be present on the box" + ) + + if name in _ALLOWED_TOKEN_ENV_VARS: + continue + if name in _FORBIDDEN_ENV_VARS: + findings.append( + f"environment variable {name!r} is set — the App write " + "credential must live ONLY as an Actions secret, never on the box" + ) + continue + if _WRITE_TOKEN_NAME_RE.search(name): + findings.append( + f"environment variable {name!r} has a GitHub write-token-shaped " + "name; the box must hold no standing write token" + ) + continue + if value and _WRITE_TOKEN_VALUE_RE.search(value): + findings.append( + f"environment variable {name!r} holds a write-capable GitHub " + "token value (ghp_/ghs_/ghu_/ghr_/github_pat_ prefix)" + ) + return findings + + +def scan_config(config: Mapping[str, object] | None) -> list[str]: + """Scan a coordinator config mapping for write-shaped token material. + + The coordinator is environment-driven, so this is normally a thin pass over + whatever config dict a caller hands in (string values only). It applies the + same name/value heuristics as :func:`scan_environ`. + """ + if not config: + return [] + findings: list[str] = [] + for key, raw in config.items(): + name = str(key) + if name in _ALLOWED_TOKEN_ENV_VARS: + continue + if name in _FORBIDDEN_ENV_VARS or _WRITE_TOKEN_NAME_RE.search(name): + findings.append( + f"coordinator config key {name!r} is a GitHub write-token-shaped " + "key; the box config must hold no standing write token" + ) + continue + if isinstance(raw, str) and _WRITE_TOKEN_VALUE_RE.search(raw): + findings.append( + f"coordinator config key {name!r} holds a write-capable GitHub " + "token value (ghp_/ghs_/ghu_/ghr_/github_pat_ prefix)" + ) + return findings + + +def scan_env_file(path: Path, *, app_id: str | None = None) -> list[str]: + """Grep one credential file for the App id and any PEM private-key header. + + A non-existent file is **not** a finding (the box legitimately may not have + every file). An unreadable-but-present file is reported as a finding so a + permissions mistake can't silently mask leaked material. + + When ``app_id`` is given (the configured ``AGENT_APPLY_APP_ID`` value), the + grep also flags that literal id appearing in the file. The + ``AGENT_APPLY_APP_ID`` / ``AGENT_APPLY_APP_PRIVATE_KEY`` *names* are always + flagged regardless. + """ + findings: list[str] = [] + if not path.exists(): + return findings + try: + text = path.read_text(encoding="utf-8", errors="replace") + except OSError as exc: # present but unreadable — fail loud, not silent + return [ + f"could not read credential file {path} ({exc}); cannot prove it is clean" + ] + + for lineno, line in enumerate(text.splitlines(), start=1): + if APP_ID_ENV in line or APP_PRIVATE_KEY_ENV in line: + findings.append( + f"{path}:{lineno}: references the App credential " + f"({APP_ID_ENV}/{APP_PRIVATE_KEY_ENV}); it must live ONLY as an " + "Actions secret" + ) + if app_id and app_id in line: + findings.append( + f"{path}:{lineno}: contains the configured App id value; the App " + "id must not be present on the box" + ) + + if _PRIVATE_KEY_RE.search(text): + findings.append( + f"{path}: contains a PEM private-key block " + "(-----BEGIN ... PRIVATE KEY-----); no private key may live on the box" + ) + return findings + + +# --------------------------------------------------------------------------- # +# Orchestration. +# --------------------------------------------------------------------------- # + + +def _resolve_env_files( + cli_files: Sequence[str] | None, environ: Mapping[str, str] +) -> list[Path]: + """Resolve the credential files to grep (CLI > env var > defaults).""" + if cli_files: + raw = list(cli_files) + elif environ.get("ASSERT_NO_WRITE_TOKEN_ENV_FILES"): + raw = [ + p.strip() + for p in environ["ASSERT_NO_WRITE_TOKEN_ENV_FILES"].split(os.pathsep) + if p.strip() + ] + else: + raw = list(_DEFAULT_ENV_FILES) + return [Path(p).expanduser() for p in raw] + + +def audit( + *, + environ: Mapping[str, str] | None = None, + config: Mapping[str, object] | None = None, + env_files: Sequence[str] | None = None, +) -> list[str]: + """Run every scanner and return the combined list of findings (empty == clean).""" + environ = os.environ if environ is None else environ + app_id = environ.get(APP_ID_ENV) or None + findings: list[str] = [] + findings += scan_environ(environ, app_id=app_id) + findings += scan_config(config) + for path in _resolve_env_files(env_files, environ): + findings += scan_env_file(path, app_id=app_id) + return findings + + +def main(argv: Sequence[str] | None = None) -> int: + parser = argparse.ArgumentParser( + description=( + "Assert no pull-requests:write / contents:write GitHub token (or the " + "apply App's id/private key) is present on this box." + ) + ) + parser.add_argument( + "--env-file", + action="append", + dest="env_files", + metavar="PATH", + help=( + "Credential file to grep for the App id + private keys (repeatable). " + "Defaults to ~/secrev.env and ~/orchestrator/.env, or the os.pathsep-" + "separated ASSERT_NO_WRITE_TOKEN_ENV_FILES env var." + ), + ) + args = parser.parse_args(argv) + + findings = audit(env_files=args.env_files) + + if findings: + print( + "FAIL: write-capable GitHub credential material found on the box " + f"({len(findings)} finding(s)). The pull-requests:write App token " + "must only ever live as an Actions secret:", + file=sys.stderr, + ) + for f in findings: + print(f" - {f}", file=sys.stderr) + return 1 + + print( + "OK: no pull-requests:write / contents:write token, App id, or private " + "key found in the environment, coordinator config, or credential files. " + "The box holds no standing write token." + ) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/agent-team/scripts/p3_rollback.sh b/agent-team/scripts/p3_rollback.sh new file mode 100755 index 0000000..3aad155 --- /dev/null +++ b/agent-team/scripts/p3_rollback.sh @@ -0,0 +1,1011 @@ +#!/usr/bin/env bash +# p3_rollback.sh — restore EVERY privileged P3 surface to a recorded baseline. +# +# The P3 apply/verify CI workflow is ALREADY LIVE (flipped + provisioned +# 2026-06-22): the `agent-apply` environment with its required reviewer, the +# GitHub App installation (pull-requests: write), the run-name/permissions edits +# on the workflow, and branch protection on `main`. This script is the inverse: +# it takes a recorded baseline of those surfaces and restores them, then asserts +# post-restore == baseline. It is what you run when the flip must be unwound, or +# when a premature flip RAN and opened a draft PR that should not exist. +# +# DESTRUCTIVE-SAFE BY DEFAULT. Every gh/git mutation is held behind a --dry-run +# default that only PRINTS the plan. --apply performs the mutations. Nothing +# destructive happens without --apply on the command line. +# +# KNOWN LIMITATIONS — VALIDATE LIVE BEFORE RELYING ON --apply (C1 gate). These +# GitHub REST calls are exercised only against a stubbed `gh` in tests; their +# exact live shapes MUST be confirmed with a real `--dry-run` against the target +# repo/App during the mandatory C1 /sh-security-review + GPT-4.1 cross-review +# (this script touches IAM-adjacent surfaces): +# * Branch-protection RESTORE: the recorded baseline is the verbatim GET object, +# but GitHub's branch-protection GET and PUT schemas are NOT symmetric +# (GET enforce_admins is {enabled,url} vs PUT's bare bool; GET carries +# read-only url/checks fields PUT rejects; PUT requires a `restrictions` key +# GET may omit). A real PUT of the raw GET object can 422 — transform GET->PUT +# (enforce_admins->bool, required_status_checks->{strict,contexts}, +# required_pull_request_reviews->writable subkeys, restrictions->null if +# absent) and confirm against the live API before --apply. +# * App neutralise is uninstall (DELETE /app/installations/{id}) which needs an +# App JWT ($AGENT_APPLY_APP_JWT), NOT operator gh; --apply hard-fails without +# it rather than pretending operator gh can do it. +# * environment.deployment_branch_policy is recorded from the live env GET; if a +# custom branch-name policy is in use, confirm the restore preserves it. +# The post-restore protection assert IS real (normalize_protection reads argv, +# fails loud on parse error, and a divergent live state exits non-zero — see +# tests/test_rollback.py::test_apply_protection_DETECTS_divergent_live_restore). +# +# Surfaces restored (design: docs/P3-PHASE0-DESIGN.md; the workflow: +# .github/workflows/agent-team-apply-verify.yml): +# 1. workflow — the apply/verify flip. +# PRE-merge: close the flip PR + delete its branch. +# POST-merge: git revert the flip commit + push + re-run CI. +# LIVE: the run-name/permissions edits are already on +# the branch, so also revert them to the recorded +# baseline SHA (`workflow_baseline_sha`). +# 2. environment — the `agent-apply` environment: required reviewer(s) + +# deployment branch policy. Reviewers are restored BY NUMERIC +# USER ID (logins are resolved to ids when the baseline only +# recorded logins), via a proper JSON reviewers array. +# 3. app — neutralise the GitHub App. There is NO "reduce App +# permissions" REST endpoint (PATCH .../permissions is a 404), +# and an installation token cannot be revoked by operator gh +# (DELETE /installation/token revokes the token you +# authenticate WITH — operator host gh is not that token). +# The only programmatic neutralise is UNINSTALL +# (DELETE /app/installations/{id}), which requires an APP JWT +# (NOT operator gh). The fallback is an OUT-OF-BAND operator +# action (remove the install in the org UI / rotate the App +# private key). This script PLANS those steps and honours +# --apply only for the JWT-authenticated uninstall. +# 4. protection — branch protection on `main`; restored from the FULL recorded +# protection baseline (not just enforce_admins) and asserted. +# Baseline asserts include_administrators == ON. +# 5. incident — "premature flip that RAN": neutralise the App (uninstall via +# App JWT, or out-of-band — see surface 3), revert any draft +# PR/branch the App opened, audit the Checks trail, restore +# environment + protection, and file an incident note. +# +# Usage: +# scripts/p3_rollback.sh [--apply] [--baseline FILE] [options] +# +# ∈ { workflow | environment | app | protection | all | incident } +# +# Options: +# --record-baseline capture the CURRENT live state into the baseline JSON +# (workflow SHA, env reviewer ids, full branch protection, +# App installation id) instead of restoring. Honours +# --apply (default --dry-run just prints what it would +# capture). +# --apply perform the mutations (default is --dry-run: print only) +# --dry-run print the plan only (default; explicit for clarity) +# --baseline FILE path to the recorded baseline JSON (default: +# $P3_ROLLBACK_BASELINE or .security-review/p3-baseline.json) +# --repo OWNER/NAME target repo (default: $AGENT_TEAM_REPO_OWNER/_NAME) +# --env-name NAME environment name for --record-baseline (default agent-apply) +# --app-installation-id N App installation id to record for --record-baseline +# --flip-pr N the flip PR number (pre-merge close path) +# --flip-branch REF the flip PR head branch (pre-merge delete path) +# --merged the flip is already merged (post-merge revert path) +# --flip-commit SHA the merged flip commit to revert (post-merge path) +# -h | --help show this help and exit +# +# Baseline JSON (recorded BEFORE the flip; the restore target). Keys consumed: +# { +# "repo": "Sea-Haven-Industries/orchestrator", +# "default_branch": "main", +# "workflow_path": ".github/workflows/agent-team-apply-verify.yml", +# "workflow_baseline_sha": "", +# "environment": { +# "name": "agent-apply", +# // Prefer NUMERIC ids (restore is exact + assertable). Logins are also +# // accepted and resolved to ids at restore via `gh api users/{login}`. +# "required_reviewer_ids": [1234567], +# "required_reviewers": ["amoussa1229"], +# "deployment_branch_policy": "protected" +# }, +# "app": { +# "slug": "agent-apply", +# "installation_id": 0, +# // The ONLY programmatic neutralise is "uninstall" (App JWT). There is no +# // permission-reduction REST endpoint, so "reduce" is not an apply action — +# // it degrades to an out-of-band operator instruction. +# "action": "uninstall" // "uninstall" | "out-of-band" +# }, +# "protection": { +# "branch": "main", +# "include_administrators": true, +# // The FULL protection payload recorded pre-flip (the exact body to PUT +# // back to repos/{repo}/branches/{branch}/protection). Captured verbatim +# // by --record-baseline. enforce_admins is also asserted on its own. +# "full": { ... }, +# "required_status_checks": ["guard", "build-test"] +# } +# } +# +# Record a baseline FIRST (before the flip) with: +# scripts/p3_rollback.sh --record-baseline [--baseline FILE] [--repo OWNER/NAME] \ +# [--env-name agent-apply] [--app-installation-id N] +# which captures the live workflow SHA, env reviewer ids, full branch protection, +# and App installation id into the baseline JSON (creating .security-review/ if +# absent). The restore surfaces then read from it. +# +# Exit non-zero on any failure (set -euo pipefail). Requires bash + gh + (git for +# the post-merge revert path). No standing box token is used — run from the Mac / +# operator host where `gh` is authenticated. The App UNINSTALL step additionally +# needs an APP JWT (NOT operator gh) — see surface 3. +set -euo pipefail + +# ── defaults ───────────────────────────────────────────────────────────────── +APPLY=0 +RECORD_BASELINE=0 +BASELINE="${P3_ROLLBACK_BASELINE:-.security-review/p3-baseline.json}" +REPO_DEFAULT="" +if [ -n "${AGENT_TEAM_REPO_OWNER:-}" ] && [ -n "${AGENT_TEAM_REPO_NAME:-}" ]; then + REPO_DEFAULT="${AGENT_TEAM_REPO_OWNER}/${AGENT_TEAM_REPO_NAME}" +fi +REPO="${REPO_DEFAULT}" +SURFACE="" +ENV_NAME="agent-apply" +APP_INSTALLATION_ID="" +FLIP_PR="" +FLIP_BRANCH="" +FLIP_COMMIT="" +MERGED=0 + +PROG="$(basename "$0")" + +say() { printf '\n\033[1;36m== %s\033[0m\n' "$*"; } +plan() { printf ' \033[1;33m[PLAN]\033[0m %s\n' "$*"; } +do_() { printf ' \033[1;32m[APPLY]\033[0m %s\n' "$*"; } +err() { printf '\033[1;31m!! %s\033[0m\n' "$*" >&2; } + +usage() { + # Print the leading comment block (everything up to the first blank line after + # `set -euo pipefail`) as help. Kept simple: emit a concise synopsis. + cat < [--apply] [--baseline FILE] [options] + +Surfaces: + workflow restore the apply/verify flip (pre/post-merge + live revert) + environment restore the agent-apply environment (required reviewer ids + policy) + app neutralise the GitHub App (uninstall via App JWT, or out-of-band) + protection restore the FULL branch protection baseline (include_administrators ON) + all run workflow + environment + app + protection in order + incident premature-flip-that-RAN path (neutralise App, revert PR, audit, note) + +Options: + --record-baseline capture current live state into the baseline JSON instead + of restoring (honours --apply) + --apply perform mutations (default: --dry-run, print plan only) + --dry-run print plan only (default) + --baseline FILE recorded baseline JSON (default: \$P3_ROLLBACK_BASELINE + or .security-review/p3-baseline.json) + --repo OWNER/NAME target repo (default: \$AGENT_TEAM_REPO_OWNER/_NAME) + --env-name NAME environment name for --record-baseline (default agent-apply) + --app-installation-id N App installation id to record for --record-baseline + --flip-pr N flip PR number (pre-merge close path) + --flip-branch REF flip PR head branch (pre-merge delete path) + --merged flip already merged (post-merge revert path) + --flip-commit SHA merged flip commit to revert (post-merge path) + -h, --help show this help + +Operator ACK env vars (fail-closed by default; a partial/manual step must be +acknowledged so a restore is never silently weaker than the baseline): + AGENT_APPLY_APP_JWT= App JWT for the automated App uninstall. + P3_ROLLBACK_OOB_ACK=1 acknowledge you will neutralise the App by hand + when no APP JWT is available (lets the other + surfaces restore instead of aborting). + P3_ROLLBACK_ALLOW_PARTIAL=1 accept a DEGRADED branch-protection restore + (enforce_admins only) when the baseline has no + protection.full — otherwise this is refused. +USAGE +} + +# ── arg parsing ────────────────────────────────────────────────────────────── +# A no-op-safe parser: a single positional surface, then flags. Unknown flags +# are a hard error (fail closed — never silently ignore an option that would +# change destructive behaviour). +parse_args() { + while [ "$#" -gt 0 ]; do + case "$1" in + workflow|environment|app|protection|all|incident) + if [ -n "${SURFACE}" ]; then + err "surface already set to '${SURFACE}'; got extra '$1'"; return 2 + fi + SURFACE="$1" ;; + --record-baseline) RECORD_BASELINE=1 ;; + --apply) APPLY=1 ;; + --dry-run) APPLY=0 ;; + --env-name) shift; ENV_NAME="${1:-}"; [ -n "${ENV_NAME}" ] || { err "--env-name needs a value"; return 2; } ;; + --env-name=*) ENV_NAME="${1#*=}" ;; + --app-installation-id) shift; APP_INSTALLATION_ID="${1:-}"; [ -n "${APP_INSTALLATION_ID}" ] || { err "--app-installation-id needs a value"; return 2; } ;; + --app-installation-id=*) APP_INSTALLATION_ID="${1#*=}" ;; + --baseline) shift; BASELINE="${1:-}"; [ -n "${BASELINE}" ] || { err "--baseline needs a value"; return 2; } ;; + --baseline=*) BASELINE="${1#*=}" ;; + --repo) shift; REPO="${1:-}"; [ -n "${REPO}" ] || { err "--repo needs a value"; return 2; } ;; + --repo=*) REPO="${1#*=}" ;; + --flip-pr) shift; FLIP_PR="${1:-}"; [ -n "${FLIP_PR}" ] || { err "--flip-pr needs a value"; return 2; } ;; + --flip-pr=*) FLIP_PR="${1#*=}" ;; + --flip-branch) shift; FLIP_BRANCH="${1:-}"; [ -n "${FLIP_BRANCH}" ] || { err "--flip-branch needs a value"; return 2; } ;; + --flip-branch=*) FLIP_BRANCH="${1#*=}" ;; + --flip-commit) shift; FLIP_COMMIT="${1:-}"; [ -n "${FLIP_COMMIT}" ] || { err "--flip-commit needs a value"; return 2; } ;; + --flip-commit=*) FLIP_COMMIT="${1#*=}" ;; + --merged) MERGED=1 ;; + -h|--help) usage; exit 0 ;; + *) err "unknown argument: $1"; usage >&2; return 2 ;; + esac + shift + done + # --record-baseline does not take a surface; restore modes require one. + if [ "${RECORD_BASELINE}" = "1" ]; then + if [ -n "${SURFACE}" ]; then + err "--record-baseline does not take a surface (got '${SURFACE}')"; return 2 + fi + return 0 + fi + if [ -z "${SURFACE}" ]; then + err "a surface is required (workflow|environment|app|protection|all|incident)" + usage >&2 + return 2 + fi +} + +# ── helpers ────────────────────────────────────────────────────────────────── +# jget KEY — read a value from the baseline JSON via gh's bundled jq-less reader. +# We use `gh` only for API; for JSON parsing we prefer python3 (stdlib) so there +# is no jq dependency and behaviour is identical in tests. +jget() { + local expr="$1" + python3 - "$BASELINE" "$expr" <<'PY' +import json, sys +path, expr = sys.argv[1], sys.argv[2] +try: + with open(path, encoding="utf-8") as fh: + data = json.load(fh) +except FileNotFoundError: + print("", end="") + sys.exit(0) +cur = data +for part in expr.split("."): + if part == "": + continue + if isinstance(cur, dict) and part in cur: + cur = cur[part] + else: + cur = None + break +if cur is None: + print("", end="") +elif isinstance(cur, bool): + print("true" if cur else "false", end="") +elif isinstance(cur, (list, dict)): + print(json.dumps(cur), end="") +else: + print(cur, end="") +PY +} + +require_baseline() { + if [ ! -f "${BASELINE}" ]; then + err "baseline file not found: ${BASELINE} (record it BEFORE the flip)" + return 1 + fi + local repo_from_baseline + repo_from_baseline="$(jget repo)" + if [ -z "${REPO}" ]; then + REPO="${repo_from_baseline}" + fi + if [ -z "${REPO}" ]; then + err "no target repo: pass --repo OWNER/NAME or set repo in the baseline" + return 1 + fi + # If both are present they must agree — refuse to restore against a repo the + # baseline was not recorded for (fail closed: a mismatched restore is worse + # than no restore). + if [ -n "${repo_from_baseline}" ] && [ "${repo_from_baseline}" != "${REPO}" ]; then + err "repo mismatch: baseline=${repo_from_baseline} requested=${REPO}" + return 1 + fi +} + +# require_keys KEY... — fail HARD if any baseline key is absent (MEDIUM-1). A +# partial baseline must never yield a partial/weaker restore (e.g. restoring only +# enforce_admins while silently dropping status checks): each restore_* asserts +# the keys IT needs up front, BEFORE any mutation, so a missing key aborts the +# whole surface instead of half-restoring it. +require_keys() { + local missing="" k + for k in "$@"; do + if [ -z "$(jget "${k}")" ]; then + missing="${missing} ${k}" + fi + done + if [ -n "${missing}" ]; then + err "baseline ${BASELINE} is missing required key(s):${missing}" + err " refusing a PARTIAL restore — record a complete baseline first (--record-baseline)." + return 1 + fi +} + +warn() { printf ' \033[1;35m[WARN]\033[0m %s\n' "$*" >&2; } + +# redact_secrets — read text on stdin and mask anything that looks like a secret +# (App JWTs, GitHub tokens, Authorization headers, PEM blocks) to a placeholder. +# SEC-01 (CWE-532): the DEFAULT --dry-run mode echoes the gh/git argv into the +# plan output, and restore_app passes `-H "Authorization: Bearer ${JWT}"`, so the +# real App JWT would otherwise reach stdout/CI logs. This is applied CENTRALLY in +# run_or_plan (and anywhere else that echoes argv) so no secret can be printed. +redact_secrets() { + # sed: each pattern collapses the secret to a stable placeholder. Order matters + # (Authorization/Bearer first so the bare-token rules don't double-process it). + # * Bearer -> Bearer + # * Authorization: -> Authorization: + # * ghp_/gho_/ghu_/ghs_/ghr_/github_pat_ tokens -> + # * PEM private-key bodies -> + sed -E \ + -e 's/(Bearer)[[:space:]]+[A-Za-z0-9._~+/=-]+/\1 /g' \ + -e 's/([Aa]uthorization:)[[:space:]]*[^"'"'"']+/\1 /g' \ + -e 's/(gh[pousr]_|github_pat_)[A-Za-z0-9_]+//g' \ + -e 's/-----BEGIN [A-Z ]*PRIVATE KEY-----[^-]*-----END [A-Z ]*PRIVATE KEY-----//g' +} + +# redacted — emit "$*" with any secret masked. Use whenever argv is echoed. +redacted() { + printf '%s' "$*" | redact_secrets +} + +# run_or_plan "" gh ... — print the plan; only execute on --apply. +# The plan branch echoes the FULL argv, so it is passed through redact_secrets +# first (SEC-01) — a Bearer JWT or token in "$@" never reaches stdout. +run_or_plan() { + local desc="$1"; shift + if [ "${APPLY}" = "1" ]; then + do_ "$(redacted "${desc}")" + "$@" + else + plan "$(redacted "${desc}: $*")" + fi +} + +# assert_equal EXPECTED ACTUAL CONTEXT — fail closed on a post-restore mismatch. +assert_equal() { + local expected="$1" actual="$2" ctx="$3" + if [ "${expected}" != "${actual}" ]; then + err "POST-RESTORE ASSERT FAILED [${ctx}]: expected '${expected}', got '${actual}'" + return 1 + fi + printf ' \033[1;32m[OK]\033[0m %s == %s (%s)\n' "${expected}" "${actual}" "${ctx}" +} + +# ── surface 1: workflow flip ───────────────────────────────────────────────── +restore_workflow() { + say "1. Restore the apply/verify workflow flip" + require_keys workflow_path workflow_baseline_sha || return 1 + local wf_path baseline_sha + wf_path="$(jget workflow_path)" + [ -n "${wf_path}" ] || wf_path=".github/workflows/agent-team-apply-verify.yml" + baseline_sha="$(jget workflow_baseline_sha)" + + if [ "${MERGED}" = "1" ]; then + # POST-merge: revert the flip commit, push, re-run CI. + if [ -z "${FLIP_COMMIT}" ]; then + err "post-merge path needs --flip-commit SHA"; return 1 + fi + run_or_plan "git revert the merged flip commit ${FLIP_COMMIT}" \ + git revert --no-edit "${FLIP_COMMIT}" + run_or_plan "push the revert to ${REPO}" \ + git push origin HEAD + run_or_plan "re-run CI on the revert for ${REPO}" \ + gh workflow run --repo "${REPO}" "$(basename "${wf_path}")" + else + # PRE-merge: close the flip PR + delete its branch. + if [ -z "${FLIP_PR}" ]; then + err "pre-merge path needs --flip-pr N"; return 1 + fi + run_or_plan "close the flip PR #${FLIP_PR} on ${REPO}" \ + gh pr close "${FLIP_PR}" --repo "${REPO}" \ + --comment "Reverting the P3 apply/verify flip to baseline." --delete-branch + if [ -n "${FLIP_BRANCH}" ]; then + run_or_plan "delete the flip head branch ${FLIP_BRANCH}" \ + gh api -X DELETE "repos/${REPO}/git/refs/heads/${FLIP_BRANCH}" + fi + fi + + # LIVE: the run-name/permissions edits are already on the branch — revert the + # workflow file to the recorded baseline SHA so the live YAML matches baseline. + if [ -n "${baseline_sha}" ]; then + run_or_plan "revert ${wf_path} to baseline SHA ${baseline_sha}" \ + git checkout "${baseline_sha}" -- "${wf_path}" + # P3-IAC-08 (CWE-665): `git checkout -- ` only STAGES the revert + # in the local worktree/index — unlike the post-merge path it does NOT commit + # or push, so the remote live YAML is unchanged until a human lands it. Warn + # loudly so the restore is never assumed complete on the remote. + warn "LIVE-YAML revert of ${wf_path} is STAGED-ONLY (local worktree/index)." + warn " It is NOT committed or pushed — the remote workflow is UNCHANGED until you" + warn " manually commit + push (e.g. git commit -m 'revert P3 flip workflow' && git push)." + else + plan "(no workflow_baseline_sha recorded — skipping live YAML revert)" + fi +} + +# resolve_reviewer_ids — emit a sorted, space-separated list of NUMERIC user ids +# for the environment's required reviewers. Prefers baseline-recorded ids +# (environment.required_reviewer_ids); for any baseline that only recorded logins +# (environment.required_reviewers), resolve login->id via `gh api users/{login}`. +# Sorting makes the post-restore compare order-independent. +resolve_reviewer_ids() { + local ids logins login id + ids="$(jget environment.required_reviewer_ids)" # JSON array of ints, or "" + logins="$(jget environment.required_reviewers)" # JSON array of strings, or "" + + python3 - "$ids" "$logins" <<'PY' +import json, sys +ids_raw, logins_raw = sys.argv[1], sys.argv[2] +out = [] +if ids_raw: + out = [str(int(x)) for x in json.loads(ids_raw)] +elif logins_raw: + # Marker: logins need resolving. Print each login prefixed so the caller + # can resolve via gh (we can't shell out from inside python here). + for login in json.loads(logins_raw): + print("LOGIN:" + str(login)) + sys.exit(0) +print(" ".join(sorted(out, key=int))) +PY +} + +# reviewers_put_args ID... — echo the repeatable -F args for the reviewers array +# in the exact field form gh expects: -F 'reviewers[][type]=User' -F 'reviewers[][id]=N' +reviewers_put_args() { + local id + for id in "$@"; do + printf -- '-F\nreviewers[][type]=User\n-F\nreviewers[][id]=%s\n' "${id}" + done +} + +# ── surface 2: agent-apply environment ─────────────────────────────────────── +restore_environment() { + say "2. Restore the agent-apply environment (required reviewer ids + policy)" + require_keys environment.name || return 1 + # The reviewer can be recorded as numeric ids OR logins (the login->id resolve + # is an EQUIVALENT, not weaker, restore) — but at least one source is required, + # else there is nothing to restore the required reviewer FROM (MEDIUM-1). + if [ -z "$(jget environment.required_reviewer_ids)" ] \ + && [ -z "$(jget environment.required_reviewers)" ]; then + err "baseline has neither environment.required_reviewer_ids nor environment.required_reviewers" + err " — nothing to restore the required reviewer from; record a complete baseline." + return 1 + fi + local env_name policy raw line id ids + env_name="$(jget environment.name)" + [ -n "${env_name}" ] || env_name="agent-apply" + policy="$(jget environment.deployment_branch_policy)" + + # Resolve the baseline reviewer set to NUMERIC ids. If the baseline only + # recorded logins, resolve each login->id via `gh api users/{login}`. + raw="$(resolve_reviewer_ids)" + ids="" + if printf '%s' "${raw}" | grep -q '^LOGIN:'; then + while IFS= read -r line; do + [ -n "${line}" ] || continue + local login="${line#LOGIN:}" + if [ "${APPLY}" = "1" ]; then + id="$(gh api "users/${login}" --jq '.id' 2>/dev/null || true)" + if [ -z "${id}" ]; then + err "could not resolve reviewer login '${login}' to a numeric id"; return 1 + fi + else + # Dry-run: we don't hit the network. Show the resolution we WOULD do. + plan "resolve reviewer login '${login}' -> id via 'gh api users/${login} --jq .id'" + id="" + fi + ids="${ids:+${ids} }${id}" + done </dev/null || true)" + assert_equal "${ids}" "${got_ids}" "environment.required_reviewer_ids" + else + plan "(post-restore: assert LIVE env '${env_name}' reviewer ids == baseline [${ids}])" + fi +} + +# ── surface 3: GitHub App installation ─────────────────────────────────────── +# Honest model of what's actually possible against the GitHub REST API: +# +# * You CANNOT "rotate" the App's installation token from operator gh. +# DELETE /installation/token revokes the token you authenticate WITH — i.e. +# the installation token itself, not some other installation's token. The +# operator host `gh` (user OAuth / PAT) is NOT that token, so it returns 403. +# A leaked installation token is short-lived (<=1h) and self-expires; to kill +# it sooner you must either run that DELETE *with that token* (inside the +# runner) or neutralise the source of new tokens (below). +# +# * There is NO permission-reduction REST endpoint. +# PATCH /app/installations/{id}/permissions does NOT exist (404). Installation +# permissions are reduced only by editing the App in the UI/settings API and +# re-accepting, which is an out-of-band operator action. +# +# * The only programmatic NEUTRALISE is UNINSTALL: +# DELETE /app/installations/{id} — which requires an APP JWT (NOT operator gh, +# NOT an installation token, NOT a fine-grained PAT). Provide the JWT via +# $AGENT_APPLY_APP_JWT; this surface only --apply's the uninstall when it is +# set. Without it, the step degrades to an out-of-band instruction. +# Out-of-band ACK gate (MEDIUM-2): when neutralising the App needs a MANUAL +# operator action (no APP JWT for the uninstall, or action=out-of-band), --apply +# must not silently skip it. Require an explicit acknowledgement +# (P3_ROLLBACK_OOB_ACK=1) that the operator WILL perform the manual step, so the +# App is never left un-neutralised without a conscious sign-off. WITH the ack the +# surface returns 0 so the other surfaces still restore (the App step is operator- +# owed, loudly warned); WITHOUT it, --apply aborts this surface. +require_oob_ack() { + local what="$1" + if [ "${P3_ROLLBACK_OOB_ACK:-}" = "1" ]; then + warn "OUT-OF-BAND ACK accepted (P3_ROLLBACK_OOB_ACK=1): ${what}" + warn " the App is NOT neutralised by this script — you MUST do it by hand NOW." + return 0 + fi + err "${what}" + err " set P3_ROLLBACK_OOB_ACK=1 to acknowledge you will perform the manual step" + err " (lets the other surfaces restore), or set \$AGENT_APPLY_APP_JWT for the automated uninstall." + return 1 +} + +# uninstall_app_via_curl SLUG INSTALLATION_ID — issue the App-JWT-authenticated +# DELETE /app/installations/{id} WITHOUT the JWT ever appearing on a command line +# (SEC-02 / CWE-214). The Authorization header is written to a 0600 temp file and +# passed to `curl -H @file`; argv carries only the filename. The file is shredded +# (or rm'd) on return via a trap. Honours --apply (dry-run only prints the plan, +# with the JWT redacted by run_or_plan-style masking). +uninstall_app_via_curl() { + local slug="$1" inst="$2" + if [ "${APPLY}" != "1" ]; then + # Dry-run: never write the header file; describe the action with NO secret. + plan "$(redacted "UNINSTALL App '${slug}' installation ${inst} via curl -H @<0600 hdrfile> (Authorization: Bearer ${AGENT_APPLY_APP_JWT})")" + plan " (JWT is written to a 0600 temp header file, never to the process argv)" + return 0 + fi + do_ "UNINSTALL App '${slug}' installation ${inst} (App JWT via 0600 header file; not on argv)" + local hdr_file rc=0 + hdr_file="$(mktemp "${TMPDIR:-/tmp}/p3-app-hdr.XXXXXX")" + # Tighten perms BEFORE writing the secret. + chmod 600 "${hdr_file}" + printf 'Authorization: Bearer %s\n' "${AGENT_APPLY_APP_JWT}" > "${hdr_file}" + curl -fsS -X DELETE \ + -H @"${hdr_file}" \ + -H "Accept: application/vnd.github+json" \ + -H "X-GitHub-Api-Version: 2022-11-28" \ + "https://api.github.com/app/installations/${inst}" || rc=$? + # Always shred/remove the header file so the JWT does not linger on disk. + shred -u "${hdr_file}" 2>/dev/null || rm -f "${hdr_file}" + return "${rc}" +} + +restore_app() { + say "3. Neutralise the GitHub App installation" + require_keys app.action || return 1 + local action installation_id slug + action="$(jget app.action)" + [ -n "${action}" ] || action="uninstall" + installation_id="$(jget app.installation_id)" + slug="$(jget app.slug)" + + # Be explicit that operator gh cannot revoke the installation token. + plan "(installation token: cannot be revoked by operator gh — DELETE /installation/token" + plan " revokes the token you AUTHENTICATE WITH; the short-lived install token self-expires." + plan " To kill it sooner, run that DELETE inside the runner with that token, or uninstall below.)" + + case "${action}" in + uninstall) + if [ -z "${installation_id}" ] || [ "${installation_id}" = "0" ]; then + err "uninstall requested but no app.installation_id in baseline"; return 1 + fi + # DELETE /app/installations/{id} requires an APP JWT — NOT operator gh. + if [ -n "${AGENT_APPLY_APP_JWT:-}" ]; then + # SEC-02 (CWE-214): keep the JWT OUT of the process argv (visible in + # ps/proc to any local user). `gh api` has no header-from-file mechanism + # and -H puts the value on the command line, so we issue the DELETE with + # `curl -H @headerfile` instead — the Authorization header (with the JWT) + # lives only in a 0600 temp file that is shredded on return. The argv + # carries only the file reference, never the secret. + uninstall_app_via_curl "${slug}" "${installation_id}" + else + # No JWT available: never pretend operator gh can do this. Plan only. + plan "UNINSTALL App '${slug}' installation ${installation_id} requires an APP JWT" + plan " (set \$AGENT_APPLY_APP_JWT). Without it: do it OUT-OF-BAND —" + plan " remove the installation in the org UI, or run:" + plan " gh api -X DELETE app/installations/${installation_id} -H 'Authorization: Bearer '" + if [ "${APPLY}" = "1" ]; then + require_oob_ack \ + "uninstall needs an APP JWT (\$AGENT_APPLY_APP_JWT); operator gh cannot do this" || return 1 + fi + fi + ;; + out-of-band) + # Honest: no API path. Tell the operator exactly what to do by hand. + plan "OUT-OF-BAND App neutralise for '${slug}' (no REST endpoint reduces App permissions):" + plan " - remove the installation in the org UI (Settings > GitHub Apps > Configure > Uninstall), AND/OR" + plan " - rotate the App private key out-of-band (App settings > Generate a new private key," + plan " update the AGENT_APPLY_APP_PRIVATE_KEY Actions secret), which invalidates new-token minting." + if [ "${APPLY}" = "1" ]; then + require_oob_ack \ + "app.action 'out-of-band' has no API path — neutralise the App by hand" || return 1 + fi + ;; + *) + err "unknown app.action '${action}' (expected uninstall|out-of-band)"; return 1 ;; + esac +} + +# ── surface 4: branch protection ───────────────────────────────────────────── +restore_protection() { + say "4. Restore the FULL branch protection baseline (include_administrators ON)" + require_keys protection.branch protection.include_administrators || return 1 + # protection.full is the COMPLETE payload (status checks, PR reviews, ...). + # Restoring only enforce_admins when it is absent leaves a WEAKER posture than + # baseline (MEDIUM-1). Refuse that silent degrade by default; the operator must + # explicitly accept the enforce_admins-only restore via P3_ROLLBACK_ALLOW_PARTIAL=1. + if [ -z "$(jget protection.full)" ]; then + if [ "${P3_ROLLBACK_ALLOW_PARTIAL:-}" != "1" ]; then + err "baseline has no protection.full — an enforce_admins-only restore would leave a WEAKER" + err " posture (status checks / PR reviews NOT restored). Set P3_ROLLBACK_ALLOW_PARTIAL=1 to" + err " accept the degraded restore, or record a complete baseline (--record-baseline)." + return 1 + fi + warn "DEGRADED protection restore (P3_ROLLBACK_ALLOW_PARTIAL=1): enforce_admins ONLY;" + warn " status checks / PR reviews are NOT restored — restore them by hand or re-record." + fi + local branch include_admins full + branch="$(jget protection.branch)" + [ -n "${branch}" ] || branch="$(jget default_branch)" + [ -n "${branch}" ] || branch="main" + include_admins="$(jget protection.include_administrators)" + full="$(jget protection.full)" # the full protection payload, captured verbatim + + # The baseline MUST require admins to be included — a rollback that left admins + # exempt would be a weaker posture than baseline. Fail closed otherwise. + if [ "${include_admins}" != "true" ]; then + err "baseline protection.include_administrators is not true (got '${include_admins}'); refusing" + return 1 + fi + + # Restore the FULL protection object (not only enforce_admins): one PUT to + # repos/{repo}/branches/{branch}/protection with the recorded body. This rebuilds + # required_status_checks, required_pull_request_reviews, restrictions, etc. + if [ -n "${full}" ] && [ "${full}" != "null" ]; then + if [ "${APPLY}" = "1" ]; then + do_ "PUT full branch protection on ${branch} from baseline protection.full" + printf '%s' "${full}" | gh api -X PUT \ + "repos/${REPO}/branches/${branch}/protection" --input - + else + plan "PUT full branch protection on ${branch} from baseline protection.full: ${full}" + fi + else + # No full payload recorded: at minimum re-enable enforce_admins so admins are + # not left exempt. Be explicit that this is a partial restore. + plan "(no protection.full recorded — partial restore: enforce_admins only)" + run_or_plan "enforce_admins ON for ${branch} (partial; record protection.full for a full restore)" \ + gh api -X PUT "repos/${REPO}/branches/${branch}/protection/enforce_admins" + fi + + if [ "${APPLY}" = "1" ]; then + # Assert enforce_admins is ON. + local got + got="$(gh api "repos/${REPO}/branches/${branch}/protection/enforce_admins" --jq '.enabled' 2>/dev/null || true)" + assert_equal "true" "${got}" "protection.include_administrators" + # If a full baseline was recorded, assert the LIVE protection matches it. + if [ -n "${full}" ] && [ "${full}" != "null" ]; then + local live_norm base_norm + live_norm="$(normalize_protection "$(gh api "repos/${REPO}/branches/${branch}/protection" 2>/dev/null)")" + base_norm="$(normalize_protection "${full}")" + assert_equal "${base_norm}" "${live_norm}" "protection.full" + fi + else + plan "(post-restore: assert enforce_admins.enabled == true on ${branch})" + plan "(post-restore: assert LIVE protection == baseline protection.full)" + fi +} + +# normalize_protection — read a branch-protection JSON object on stdin and emit a +# canonical, comparable projection of the fields we restore. GitHub's GET adds +# url/metadata fields that a PUT body never carries, so we project only the +# semantically meaningful settings and sort, making the baseline-vs-live compare +# robust to representational noise. +normalize_protection() { + # The JSON is passed as $1 (NOT piped): a `python3 - <<'PY'` heredoc occupies + # stdin, so json.load(sys.stdin) would read the program text, not the data — + # which silently yielded "" for every input and made the post-restore assert + # vacuous. Read argv[1] instead (mirrors the other heredoc helpers here). + python3 - "${1-}" <<'PY' +import json, sys +raw = sys.argv[1] if len(sys.argv) > 1 else "" +if not raw.strip(): + # Empty input (e.g. the live GET failed) — emit a distinct sentinel so the + # baseline-vs-live compare MISMATCHES and the restore is flagged, never a + # silent pass. + print("__NORMALIZE_EMPTY__") + sys.exit(0) +try: + d = json.loads(raw) +except Exception as exc: + print(f"__NORMALIZE_ERROR__:{exc}") + sys.exit(0) + +def boolish(node, *path, key="enabled"): + cur = node + for p in path: + if not isinstance(cur, dict): + return None + cur = cur.get(p) + if isinstance(cur, dict): + return bool(cur.get(key)) + if isinstance(cur, bool): + return cur + return None + +proj = { + "enforce_admins": boolish(d, "enforce_admins"), + "required_status_checks_strict": ( + (d.get("required_status_checks") or {}).get("strict") + if isinstance(d.get("required_status_checks"), dict) else None + ), + "required_status_checks_contexts": sorted( + (d.get("required_status_checks") or {}).get("contexts", []) or [] + ) if isinstance(d.get("required_status_checks"), dict) else None, + "required_pull_request_reviews": ( + (lambda r: { + "required_approving_review_count": r.get("required_approving_review_count"), + "dismiss_stale_reviews": r.get("dismiss_stale_reviews"), + "require_code_owner_reviews": r.get("require_code_owner_reviews"), + })(d["required_pull_request_reviews"]) + if isinstance(d.get("required_pull_request_reviews"), dict) else None + ), + "required_linear_history": boolish(d, "required_linear_history"), + "allow_force_pushes": boolish(d, "allow_force_pushes"), + "allow_deletions": boolish(d, "allow_deletions"), +} +print(json.dumps(proj, sort_keys=True, separators=(",", ":"))) +PY +} + +# ── incident: premature flip that RAN ──────────────────────────────────────── +incident_premature_flip() { + say "INCIDENT — premature flip that RAN (rotate / revert / audit / restore / note)" + local branch note_path + branch="$(jget protection.branch)" + [ -n "${branch}" ] || branch="$(jget default_branch)" + [ -n "${branch}" ] || branch="main" + + # (a) Neutralise the App FIRST — anything the premature run minted is suspect. + # NB: operator gh CANNOT revoke the installation token (DELETE + # /installation/token revokes the token you authenticate WITH). The minted + # token is short-lived and self-expires; the durable neutralise is to + # uninstall the App (App JWT) or rotate its private key out-of-band. + say "a. Neutralise the GitHub App (uninstall via App JWT, or out-of-band)" + restore_app + + # (b) Revert any draft PR / branch the App opened. The draft PR head is the + # dispatcher's namespaced ref (agent-team/apply/); close + delete. + say "b. Revert the draft PR / branch the App opened" + if [ -n "${FLIP_PR}" ]; then + run_or_plan "close the premature draft PR #${FLIP_PR}" \ + gh pr close "${FLIP_PR}" --repo "${REPO}" \ + --comment "Closing premature agent-apply draft PR (incident rollback)." \ + --delete-branch + else + plan "(no --flip-pr given: list open agent-team/apply/* PRs to triage)" + run_or_plan "list open agent-apply draft PRs for triage" \ + gh pr list --repo "${REPO}" --draft --search "head:agent-team/apply/" --json number,headRefName + fi + if [ -n "${FLIP_BRANCH}" ]; then + run_or_plan "delete the premature head branch ${FLIP_BRANCH}" \ + gh api -X DELETE "repos/${REPO}/git/refs/heads/${FLIP_BRANCH}" + fi + + # (c) Audit the Checks trail of the premature run (read-only — always run). + say "c. Audit the Checks trail (read-only)" + run_or_plan "audit recent apply/verify runs (Checks trail)" \ + gh run list --repo "${REPO}" --workflow agent-team-apply-verify.yml \ + --limit 20 --json databaseId,headBranch,event,conclusion,createdAt + + # (d) Restore environment + branch protection to baseline. + say "d. Restore environment + protection to baseline" + restore_environment + restore_protection + + # (e) File an incident note. + say "e. File an incident note" + note_path=".security-review/incidents/p3-premature-flip-$(date +%Y%m%dT%H%M%SZ).md" + if [ "${APPLY}" = "1" ]; then + do_ "write incident note ${note_path}" + mkdir -p "$(dirname "${note_path}")" + { + printf '# Incident: premature P3 apply/verify flip that RAN\n\n' + printf -- '- Repo: %s\n' "${REPO}" + printf -- '- Baseline: %s\n' "${BASELINE}" + printf -- '- Flip PR: %s\n' "${FLIP_PR:-}" + printf -- '- Flip branch: %s\n' "${FLIP_BRANCH:-}" + printf -- '- Recorded: %s\n' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" + printf '\n## Actions taken\n' + printf -- '1. Neutralised the GitHub App (uninstall via App JWT, or out-of-band).\n' + printf -- ' NB: the minted installation token cannot be revoked by operator gh;\n' + printf -- ' it is short-lived and self-expires.\n' + printf -- '2. Closed the premature draft PR + deleted its branch.\n' + printf -- '3. Audited the apply/verify Checks trail.\n' + printf -- '4. Restored the agent-apply environment + branch protection.\n' + } > "${note_path}" + else + plan "write incident note to ${note_path} (repo/flip/baseline + actions taken)" + fi +} + +# ── --record-baseline ──────────────────────────────────────────────────────── +# Capture the CURRENT live state into the baseline JSON: the live workflow SHA, +# the env required-reviewer NUMERIC ids, the FULL branch protection object, and +# the App installation id. Honours --apply (default --dry-run just prints what it +# would capture). Creates .security-review/ if absent. +record_baseline() { + say "RECORD BASELINE — capture current live state into ${BASELINE}" + # Resolve a repo: --repo, or env default. The baseline file may not exist yet, + # so we do NOT call require_baseline here. + if [ -z "${REPO}" ]; then + err "no target repo: pass --repo OWNER/NAME (baseline does not exist yet to read it from)" + return 1 + fi + + local branch wf_path + branch="main" + wf_path=".github/workflows/agent-team-apply-verify.yml" + + if [ "${APPLY}" != "1" ]; then + plan "capture live workflow SHA for ${wf_path} via 'gh api repos/${REPO}/contents/${wf_path} --jq .sha'" + plan "capture env '${ENV_NAME}' required reviewer NUMERIC ids" + plan "capture FULL branch protection for ${branch}" + plan "capture App installation id (${APP_INSTALLATION_ID:-<--app-installation-id N>})" + plan "write baseline JSON to ${BASELINE} (creating $(dirname "${BASELINE}") if absent)" + return 0 + fi + + do_ "capture live state into ${BASELINE}" + mkdir -p "$(dirname "${BASELINE}")" + + local wf_sha reviewer_ids protection_full + wf_sha="$(gh api "repos/${REPO}/contents/${wf_path}" --jq '.sha' 2>/dev/null || true)" + reviewer_ids="$(gh api "repos/${REPO}/environments/${ENV_NAME}" \ + --jq '[.protection_rules[]? | select(.type=="required_reviewers") + | .reviewers[]? | select(.type=="User") | .reviewer.id] | sort' \ + 2>/dev/null || echo '[]')" + [ -n "${reviewer_ids}" ] || reviewer_ids='[]' + protection_full="$(gh api "repos/${REPO}/branches/${branch}/protection" 2>/dev/null || echo 'null')" + [ -n "${protection_full}" ] || protection_full='null' + + # Assemble the baseline JSON deterministically with python3 (stdlib). + REPO="${REPO}" BRANCH="${branch}" WF_PATH="${wf_path}" WF_SHA="${wf_sha}" \ + ENV_NAME="${ENV_NAME}" REVIEWER_IDS="${reviewer_ids}" \ + APP_INSTALLATION_ID="${APP_INSTALLATION_ID}" PROTECTION_FULL="${protection_full}" \ + python3 - "${BASELINE}" <<'PY' +import json, os, sys + +out_path = sys.argv[1] +try: + reviewer_ids = json.loads(os.environ.get("REVIEWER_IDS") or "[]") +except Exception: + reviewer_ids = [] +try: + protection_full = json.loads(os.environ.get("PROTECTION_FULL") or "null") +except Exception: + protection_full = None + +include_admins = bool( + isinstance(protection_full, dict) + and isinstance(protection_full.get("enforce_admins"), dict) + and protection_full["enforce_admins"].get("enabled") +) + +inst = os.environ.get("APP_INSTALLATION_ID") or "" +baseline = { + "repo": os.environ["REPO"], + "default_branch": os.environ["BRANCH"], + "workflow_path": os.environ["WF_PATH"], + "workflow_baseline_sha": os.environ.get("WF_SHA") or "", + "environment": { + "name": os.environ["ENV_NAME"], + "required_reviewer_ids": reviewer_ids, + "deployment_branch_policy": "protected", + }, + "app": { + "slug": os.environ["ENV_NAME"], + "installation_id": int(inst) if inst.isdigit() else 0, + "action": "uninstall", + }, + "protection": { + "branch": os.environ["BRANCH"], + "include_administrators": include_admins, + "full": protection_full, + }, +} +with open(out_path, "w", encoding="utf-8") as fh: + json.dump(baseline, fh, indent=2, sort_keys=True) + fh.write("\n") +print(f" wrote {out_path}") +PY + printf ' \033[1;32m[OK]\033[0m baseline recorded to %s\n' "${BASELINE}" +} + +# ── main ───────────────────────────────────────────────────────────────────── +main() { + parse_args "$@" + + if [ "${RECORD_BASELINE}" = "1" ]; then + if [ "${APPLY}" = "1" ]; then + say "MODE: --record-baseline --apply (will CAPTURE live state) on repo ${REPO}" + else + say "MODE: --record-baseline --dry-run (plan only) on repo ${REPO}" + fi + record_baseline + say "Done (record-baseline, $([ "${APPLY}" = "1" ] && echo apply || echo dry-run))." + return 0 + fi + + require_baseline + + if [ "${APPLY}" = "1" ]; then + say "MODE: --apply (mutations WILL be performed) on repo ${REPO}" + else + say "MODE: --dry-run (plan only; no mutations) on repo ${REPO}" + fi + + case "${SURFACE}" in + workflow) restore_workflow ;; + environment) restore_environment ;; + app) restore_app ;; + protection) restore_protection ;; + all) + restore_workflow + restore_environment + restore_app + restore_protection + ;; + incident) incident_premature_flip ;; + *) err "unhandled surface: ${SURFACE}"; return 2 ;; + esac + + say "Done (${SURFACE}, $([ "${APPLY}" = "1" ] && echo apply || echo dry-run))." +} + +main "$@" diff --git a/agent-team/systemd/agent-team-coordinator.service b/agent-team/systemd/agent-team-coordinator.service index 3261b3d..2d8b593 100644 --- a/agent-team/systemd/agent-team-coordinator.service +++ b/agent-team/systemd/agent-team-coordinator.service @@ -32,6 +32,27 @@ # AGENT_TEAM_SLACK_OWNER_IDS -> comma-separated authorized answerer ids # (AUTHZ-01). The listener FAILS CLOSED if unset. # +# --- P3 build->dispatch->verify wiring (read by the LIVE serve() path) --- +# The bound P3 wiring is the serve() default; if any of the three below (plus a +# CI-read token) is unset, serve() degrades to the INERT P3 path (one WARNING + +# a #agent-team notice) instead of crash-looping (Decision 5). All are read at +# graph-build / dispatch time from this file's environment: +# AGENT_TEAM_REPO_OWNER -> dispatch target owner. Fixed at factory time, +# never read from pipeline state, so model output +# cannot redirect the dispatch/verify target. +# AGENT_TEAM_REPO_NAME -> dispatch target repo (same fail-closed binding). +# AGENT_TEAM_BASE_BRANCH -> PR base branch (optional; default "main"). +# AGENT_TEAM_CI_READ_TOKEN -> the READ-ONLY CI-result token. The verifier's +# authenticated conclusion read uses it (falls back +# to GITHUB_TOKEN). This is read-only by contract: +# NO pull-requests:write / contents:write token, +# and NO AGENT_APPLY_APP_ID / _PRIVATE_KEY, may live +# in this file. The apply path mints its write token +# INSIDE the CI runner from Actions secrets; the box +# holds no standing write credential. The invariant +# is asserted by scripts/assert_no_write_token.py +# (the A2 audit) at provisioning + in CI. +# # ~/orchestrator/.env (mode 600, NOT in git) - the P2 review-loop provider key: # The production serve() wires the GPT-4.1 cross-review loop, which shells the # local orchestrator run.py -> cross_reviewer once a task reaches REVIEW. That diff --git a/agent-team/tests/test_build_verify_subgraph.py b/agent-team/tests/test_build_verify_subgraph.py index 30a5976..c354afb 100644 --- a/agent-team/tests/test_build_verify_subgraph.py +++ b/agent-team/tests/test_build_verify_subgraph.py @@ -102,6 +102,19 @@ def test_route_ids_mirror_graph_by_value() -> None: assert APPROVED_ROUTE == "approved" +def test_integrate_note_documents_build_dispatch_verify_order() -> None: + """The topology note records BUILD -> [DISPATCH] -> VERIFY (design §4 D1). + + DISPATCH must precede VERIFY so it captures ``state["run_id"]`` before VERIFY + reads it. Pin the documented order so a future reorder back to the old + BUILD -> VERIFY -> DISPATCH topology is caught. + """ + note = bvs._INTEGRATE_NOTE + assert "BUILD -> [DISPATCH] -> VERIFY" in note + # DISPATCH appears before VERIFY in the recorded stage order. + assert note.index("DISPATCH") < note.index("VERIFY") + + # --------------------------------------------------------------------------- # # BUILD node: proposes a diff via an injected fake builder # --------------------------------------------------------------------------- # @@ -224,6 +237,47 @@ def test_verify_failing_ci_result_loops_back_to_build() -> None: assert out["current_phase"] == Phase.BUILD.value assert out["ci_results"]["gate_decision"] == "fail" assert route_after_verify(out) == BUILD_ROUTE + # The recoverable loop persists the incremented DURABLE count into state. + assert out["build_loops"] == 1 + + +def test_repeated_fail_through_wrapper_parks_after_max_build_loops() -> None: + """A perpetually-FAILing task PARKS after exactly max_build_loops loops. + + Drives the durable BUILD->DISPATCH->VERIFY budget through the subgraph + wrapper with ONE shared VerifierConfig (the wiring-time reality). The count + that bounds the loop is the per-task ``state["build_loops"]`` the node writes + back each round (LOGIC-RACE-01) — the shared config's build_loops stays 0, so + if the node read the count from config the loop would never terminate. + """ + diff = _diff_for("src/foo.py") + + def fail_fetcher(state): + return {"run_id": "r1", "conclusion": "failure", "diff_hash": _hash(diff)} + + shared_cfg = VerifierConfig( + expected_run_id="r1", + allowed_scope=["src"], + max_build_loops=3, + build_loops=0, + ) + node = make_verify_node(shared_cfg, ci_result_fetcher=fail_fetcher) + + state = _verify_state(diff) + state["build_loops"] = 0 + + routes: list[str] = [] + for _ in range(10): # bound the harness; a regression must not hang + out = node(state) + routes.append(route_after_verify(out)) + if route_after_verify(out) == PARKED_ROUTE: + break + # The checkpointer would carry build_loops forward across the loop. + state["build_loops"] = out["build_loops"] + + assert routes == [BUILD_ROUTE, BUILD_ROUTE, PARKED_ROUTE] + assert out["status"] == TaskStatus.PARKED.value + assert out["current_phase"] == Phase.PARKED.value def test_verify_malformed_fetcher_result_fails_safe_to_parked() -> None: @@ -373,3 +427,203 @@ def test_verify_node_does_not_mutate_caller_state() -> None: node(state) # Caller's state is untouched (the node wrote into a dict copy). assert state["ci_results"] is sentinel + + +# --------------------------------------------------------------------------- # +# Per-task run-id binding through the subgraph wrapper (design §4 Decision 4) +# --------------------------------------------------------------------------- # + + +def test_verify_node_binds_to_per_task_state_run_id() -> None: + """make_verify_node gates against state["run_id"], not a static config id. + + A single VerifierConfig with no expected_run_id is shared; the per-task + state["run_id"] supplies the binding the fetched CI conclusion must match. + """ + diff = _diff_for("src/foo.py") + + def fetcher_keyed_to_task(s): + # The (real) fetcher keys its conclusion to the task's dispatched run id. + return { + "run_id": s["run_id"], + "conclusion": "success", + "diff_hash": _hash(diff), + } + + node = make_verify_node( + VerifierConfig(expected_run_id=None, allowed_scope=["src"]), + ci_result_fetcher=fetcher_keyed_to_task, + ) + + state = _verify_state(diff) + state["run_id"] = "dispatched-123" + out = node(state) + assert out["status"] == TaskStatus.DONE.value + assert out["review_verdicts"][0]["run_id"] == "dispatched-123" + + +def test_verify_node_blocks_substituted_run_id_through_wrapper() -> None: + """A fetched CI result keyed to a DIFFERENT run id than state["run_id"] blocks.""" + diff = _diff_for("src/foo.py") + + def substituting_fetcher(s): + # CI result grafted from another task (run id mismatch) with success. + return { + "run_id": "someone-elses-run", + "conclusion": "success", + "diff_hash": _hash(diff), + } + + node = make_verify_node( + VerifierConfig(expected_run_id=None, allowed_scope=["src"]), + ci_result_fetcher=substituting_fetcher, + ) + + state = _verify_state(diff) + state["run_id"] = "my-run" + out = node(state) + assert out["status"] == TaskStatus.PARKED.value + assert out["ci_results"]["gate_decision"] == "block" + + +def test_verify_node_blocks_when_no_run_id_anywhere() -> None: + """No state run_id and no config fallback -> BLOCK/park, never a vacuous pass.""" + diff = _diff_for("src/foo.py") + + def pass_fetcher(s): + return {"run_id": "r1", "conclusion": "success", "diff_hash": _hash(diff)} + + node = make_verify_node( + VerifierConfig(expected_run_id=None, allowed_scope=["src"]), + ci_result_fetcher=pass_fetcher, + ) + out = node(_verify_state(diff)) # no state["run_id"] + assert out["status"] == TaskStatus.PARKED.value + assert out["ci_results"]["gate_decision"] == "block" + + +# --------------------------------------------------------------------------- # +# Async CI-wait: VERIFY suspends via interrupt() on an in-progress run (§4 Dec 2) +# --------------------------------------------------------------------------- # + + +def test_verify_suspends_on_in_progress_run_then_resumes_to_gate() -> None: + """With a dispatched run_id and a not-yet-terminal fetch, VERIFY interrupts; + once the CI-watcher resumes it, the re-fetched terminal result is gated.""" + from langgraph.checkpoint.memory import MemorySaver + from langgraph.graph import END, START, StateGraph + from langgraph.types import Command + + diff = _diff_for("src/foo.py") + + # The fetcher is "in progress" on the first call, then terminal on the second + # (mirroring a real run that concluded while the task was suspended). + calls = {"n": 0} + + def progressing_fetcher(state): + calls["n"] += 1 + if calls["n"] == 1: + return None # still running -> VERIFY must suspend + return { + "run_id": state["run_id"], + "conclusion": "success", + "diff_hash": _hash(diff), + } + + node = make_verify_node( + VerifierConfig(expected_run_id=None, allowed_scope=["src"]), + ci_result_fetcher=progressing_fetcher, + ) + + builder = StateGraph(PipelineState) + builder.add_node("verify", node) + builder.add_edge(START, "verify") + builder.add_edge("verify", END) + graph = builder.compile(checkpointer=MemorySaver()) + + cfg = {"configurable": {"thread_id": "tw1"}} + state = _verify_state(diff) + state["run_id"] = "dispatched-77" + + first = graph.invoke(state, cfg) + # The task suspended at VERIFY's interrupt (awaiting CI) rather than gating. + assert "__interrupt__" in first + assert calls["n"] == 1 + + # The CI-watcher resumes the task once the run terminated. + out = graph.invoke(Command(resume={"awaiting_ci": "done"}), cfg) + # On resume the node re-fetched the now-terminal result and the gate PASSed. + assert calls["n"] == 2 + assert out["status"] == TaskStatus.DONE.value + assert out["current_phase"] == Phase.DONE.value + + +def test_verify_spurious_resume_still_non_terminal_parks_never_passes() -> None: + """A spurious resume (re-fetch STILL non-terminal) must fail closed. + + The CI-watcher resumes VERIFY on what it believes is a terminal conclusion, + but the authenticated re-fetch is the source of truth. If that re-fetch is + STILL None (a spurious / premature resume, or a run that flapped back to + in-progress), the node must NOT vacuously pass: it falls through to the gate, + which — with a dispatched run_id but no terminal result — BLOCKs and parks. + """ + from langgraph.checkpoint.memory import MemorySaver + from langgraph.graph import END, START, StateGraph + from langgraph.types import Command + + diff = _diff_for("src/foo.py") + + # The fetch is non-terminal on EVERY call (the resume was spurious — the run + # never actually concluded). + calls = {"n": 0} + + def never_terminal_fetcher(state): + calls["n"] += 1 + return None + + node = make_verify_node( + VerifierConfig(expected_run_id=None, allowed_scope=["src"]), + ci_result_fetcher=never_terminal_fetcher, + ) + + builder = StateGraph(PipelineState) + builder.add_node("verify", node) + builder.add_edge(START, "verify") + builder.add_edge("verify", END) + graph = builder.compile(checkpointer=MemorySaver()) + + cfg = {"configurable": {"thread_id": "spurious-1"}} + state = _verify_state(diff) + state["run_id"] = "dispatched-99" + + first = graph.invoke(state, cfg) + # First pass: in-progress fetch -> VERIFY suspends awaiting CI. + assert "__interrupt__" in first + assert calls["n"] == 1 + + # The watcher resumes, but the authenticated re-fetch is STILL non-terminal. + out = graph.invoke(Command(resume={"awaiting_ci": "spurious"}), cfg) + # Re-fetched again on resume (the node replays from its start, so it fetches, + # the interrupt returns the resume value rather than re-suspending, then it + # re-fetches once more) — every fetch is non-terminal. + assert calls["n"] > 1 + # ...and with no terminal result the gate BLOCKs and the task PARKS — never a + # vacuous pass. + assert out["status"] == TaskStatus.PARKED.value + assert out["current_phase"] == Phase.PARKED.value + assert out["ci_results"]["gate_decision"] == "block" + assert route_after_verify(out) == PARKED_ROUTE + + +def test_verify_no_run_id_does_not_suspend_and_parks() -> None: + """The INERT/no-run path (no state run_id) never suspends: a None fetch flows + straight to the gate, which BLOCKs and parks (existing behavior preserved).""" + diff = _diff_for("src/foo.py") + + node = make_verify_node( + VerifierConfig(expected_run_id="r1", allowed_scope=["src"]), + ci_result_fetcher=lambda s: None, + ) + out = node(_verify_state(diff)) # no state["run_id"] -> no interrupt + assert out["current_phase"] == Phase.PARKED.value + assert route_after_verify(out) == PARKED_ROUTE diff --git a/agent-team/tests/test_ci_fetcher.py b/agent-team/tests/test_ci_fetcher.py index 31edc68..328ce10 100644 --- a/agent-team/tests/test_ci_fetcher.py +++ b/agent-team/tests/test_ci_fetcher.py @@ -174,6 +174,29 @@ def test_missing_token_fails_closed(monkeypatch: pytest.MonkeyPatch) -> None: assert fetch_ci_result(_state(), owner="o", repo="r") is None +# --------------------------------------------------------------------------- # +# SEC-04: token resolution prefers the dedicated read-only var over GITHUB_TOKEN +# --------------------------------------------------------------------------- # + + +def test_resolve_prefers_dedicated_read_token(monkeypatch: pytest.MonkeyPatch) -> None: + from agent_team.ci_fetcher import _resolve_read_token + + monkeypatch.setenv(CI_READ_TOKEN_ENV, "dedicated-read-only") + monkeypatch.setenv("GITHUB_TOKEN", "fallback-token") + # The dedicated read-only var wins on the live box read path. + assert _resolve_read_token() == "dedicated-read-only" + + +def test_resolve_falls_back_to_github_token(monkeypatch: pytest.MonkeyPatch) -> None: + from agent_team.ci_fetcher import _resolve_read_token + + monkeypatch.delenv(CI_READ_TOKEN_ENV, raising=False) + monkeypatch.setenv("GITHUB_TOKEN", "fallback-token") + # Documented, explicitly-narrowed convenience fallback (must be read-only). + assert _resolve_read_token() == "fallback-token" + + # --------------------------------------------------------------------------- # # Read-only contract: never derives a verdict, never writes # --------------------------------------------------------------------------- # diff --git a/agent-team/tests/test_ci_gate.py b/agent-team/tests/test_ci_gate.py index f6f0ec9..47c9c77 100644 --- a/agent-team/tests/test_ci_gate.py +++ b/agent-team/tests/test_ci_gate.py @@ -389,14 +389,46 @@ def test_gate_scope_violation_blocks() -> None: assert any("scope" in r for r in result.reasons) -def test_gate_empty_run_id_raises() -> None: +def test_gate_empty_run_id_blocks_never_passes() -> None: + # Per design §4 (per-task binding): an empty expected_run_id means the + # dispatcher captured no run id for this task. There is nothing to bind the + # verdict to, so the gate BLOCKs (fail-closed) rather than vacuously gating + # against "" — even when a substituted ci_result reports a matching success. + diff = _diff_for("src/foo.py") + result = evaluate_ci_gate( + candidate_diff=diff, + ledger_hash=_ledger_hash(diff), + ci_result=_good_ci("", diff), + expected_run_id="", + ) + assert result.decision is GateDecision.BLOCK + assert result.run_id is None + assert any("bind" in r for r in result.reasons) + + +def test_gate_none_run_id_blocks_never_passes() -> None: + # A None expected_run_id (the legitimate "dispatch unresolved" runtime state) + # is a BLOCK, not an exception — and never a pass even on a success ci_result. + diff = _diff_for("src/foo.py") + result = evaluate_ci_gate( + candidate_diff=diff, + ledger_hash=_ledger_hash(diff), + ci_result=_good_ci("run-1", diff, "success"), + expected_run_id=None, + ) + assert result.decision is GateDecision.BLOCK + assert any("bind" in r for r in result.reasons) + + +def test_gate_non_string_run_id_raises() -> None: + # A non-string, non-None expected_run_id is structurally invalid -> raise. diff = _diff_for("src/foo.py") with pytest.raises(CiGateError): evaluate_ci_gate( candidate_diff=diff, ledger_hash=_ledger_hash(diff), ci_result=_good_ci("run-1", diff), - expected_run_id="", + expected_run_id=123, # type: ignore[arg-type] ) diff --git a/agent-team/tests/test_ci_watcher.py b/agent-team/tests/test_ci_watcher.py new file mode 100644 index 0000000..cec2c0d --- /dev/null +++ b/agent-team/tests/test_ci_watcher.py @@ -0,0 +1,323 @@ +"""Unit tests for agent_team.ci_watcher (design §3.3.2, P3 Decision 2). + +Covers the async resume-on-CI-complete sweep: it RESUMES a task on a terminal +conclusion, PARKS on the dispatch timeout, and FAILS CLOSED (parks) on a fetch +error / unusable state. No network: every side effect (poll, resume, park) is an +injected callable, exactly as the module's contract promises. +""" + +from __future__ import annotations + +from datetime import datetime, timedelta, timezone + +from agent_team.ci_watcher import ( + CiPollResult, + CiWatchAction, + PendingCiTask, + default_ci_poller, + run_ci_watcher, +) + +# A fixed "now" and a dispatched-at watermark; ISO strings mirror the ledger. +_NOW = datetime(2026, 6, 23, 12, 0, 0, tzinfo=timezone.utc) +_JUST_NOW = (_NOW - timedelta(minutes=1)).isoformat() +_LONG_AGO = (_NOW - timedelta(hours=2)).isoformat() + + +class _Recorder: + """Records the tasks/results handed to an injected side-effect callback.""" + + def __init__(self) -> None: + self.resumed: list[tuple[PendingCiTask, dict]] = [] + self.parked: list[PendingCiTask] = [] + + def on_resume(self, task: PendingCiTask, result: object) -> None: + self.resumed.append((task, dict(result))) # type: ignore[arg-type] + + def on_park(self, task: PendingCiTask) -> None: + self.parked.append(task) + + +def _task(*, run_id: str | None = "12345", dispatched_at: str | None = _JUST_NOW): + return PendingCiTask(thread_id="t1", run_id=run_id, dispatched_at=dispatched_at) + + +# --------------------------------------------------------------------------- # +# TERMINAL -> resume +# --------------------------------------------------------------------------- # + + +def test_terminal_conclusion_resumes_the_task() -> None: + rec = _Recorder() + terminal = {"run_id": "12345", "conclusion": "success", "diff_hash": "abc"} + + report = run_ci_watcher( + [_task()], + poll=lambda t: CiPollResult.terminal(terminal), + on_resume=rec.on_resume, + on_park=rec.on_park, + now=_NOW, + ) + + assert report.resumed == 1 + assert report.parked == 0 + assert rec.parked == [] + assert len(rec.resumed) == 1 + resumed_task, resumed_result = rec.resumed[0] + assert resumed_task.thread_id == "t1" + # The authenticated terminal result is handed to the resume side effect. + assert resumed_result == terminal + assert report.outcomes[0].action is CiWatchAction.RESUMED + + +def test_terminal_resume_even_past_timeout_prefers_resume() -> None: + """A run that terminated should RESUME, not park, even if dispatched long ago.""" + rec = _Recorder() + report = run_ci_watcher( + [_task(dispatched_at=_LONG_AGO)], + poll=lambda t: CiPollResult.terminal({"conclusion": "failure"}), + on_resume=rec.on_resume, + on_park=rec.on_park, + timeout=timedelta(minutes=30), + now=_NOW, + ) + assert report.resumed == 1 + assert rec.parked == [] + + +# --------------------------------------------------------------------------- # +# PENDING + timeout -> park ; PENDING within window -> wait +# --------------------------------------------------------------------------- # + + +def test_pending_within_window_waits() -> None: + rec = _Recorder() + report = run_ci_watcher( + [_task(dispatched_at=_JUST_NOW)], + poll=lambda t: CiPollResult.pending(), + on_resume=rec.on_resume, + on_park=rec.on_park, + timeout=timedelta(minutes=30), + now=_NOW, + ) + assert report.waiting == 1 + assert report.parked == 0 + assert rec.parked == [] + assert rec.resumed == [] + assert report.outcomes[0].action is CiWatchAction.WAITING + + +def test_pending_past_timeout_parks() -> None: + rec = _Recorder() + report = run_ci_watcher( + [_task(dispatched_at=_LONG_AGO)], + poll=lambda t: CiPollResult.pending(), + on_resume=rec.on_resume, + on_park=rec.on_park, + timeout=timedelta(minutes=30), + now=_NOW, + ) + assert report.parked_timeout == 1 + assert report.parked == 1 + assert len(rec.parked) == 1 + assert rec.resumed == [] + assert report.outcomes[0].action is CiWatchAction.PARKED_TIMEOUT + + +# --------------------------------------------------------------------------- # +# ERROR / None result / unusable state -> fail closed (park) +# --------------------------------------------------------------------------- # + + +def test_poll_error_parks_fail_closed() -> None: + rec = _Recorder() + report = run_ci_watcher( + [_task(dispatched_at=_JUST_NOW)], # within window: error still parks + poll=lambda t: CiPollResult.error(), + on_resume=rec.on_resume, + on_park=rec.on_park, + now=_NOW, + ) + assert report.parked_error == 1 + assert len(rec.parked) == 1 + assert rec.resumed == [] + assert report.outcomes[0].action is CiWatchAction.PARKED_ERROR + + +def test_raising_poll_is_isolated_and_parks() -> None: + rec = _Recorder() + + def boom(task: PendingCiTask) -> CiPollResult: + raise RuntimeError("github exploded") + + report = run_ci_watcher( + [_task()], + poll=boom, + on_resume=rec.on_resume, + on_park=rec.on_park, + now=_NOW, + ) + assert report.parked_error == 1 + assert len(rec.parked) == 1 + assert "RuntimeError" in (report.outcomes[0].error or "") + + +def test_missing_run_id_parks_without_polling() -> None: + rec = _Recorder() + polled: list[PendingCiTask] = [] + + def tracking_poll(task: PendingCiTask) -> CiPollResult: + polled.append(task) + return CiPollResult.pending() + + report = run_ci_watcher( + [_task(run_id=None)], + poll=tracking_poll, + on_resume=rec.on_resume, + on_park=rec.on_park, + now=_NOW, + ) + assert report.parked_error == 1 + assert polled == [] # unusable state parks BEFORE any poll + assert len(rec.parked) == 1 + + +def test_unparseable_dispatched_at_parks_without_polling() -> None: + rec = _Recorder() + polled: list[PendingCiTask] = [] + + report = run_ci_watcher( + [_task(dispatched_at="not-a-timestamp")], + poll=lambda t: (polled.append(t), CiPollResult.pending())[1], + on_resume=rec.on_resume, + on_park=rec.on_park, + now=_NOW, + ) + assert report.parked_error == 1 + assert polled == [] + assert len(rec.parked) == 1 + + +# --------------------------------------------------------------------------- # +# Side-effect isolation across the batch +# --------------------------------------------------------------------------- # + + +def test_one_task_failure_does_not_abort_the_sweep() -> None: + """A resume callback that raises for one task does not stop the others.""" + rec = _Recorder() + raised_for: list[str] = [] + + def flaky_resume(task: PendingCiTask, result: object) -> None: + if task.thread_id == "bad": + raised_for.append(task.thread_id) + raise RuntimeError("resume failed") + rec.resumed.append((task, dict(result))) # type: ignore[arg-type] + + tasks = [ + PendingCiTask(thread_id="bad", run_id="1", dispatched_at=_JUST_NOW), + PendingCiTask(thread_id="good", run_id="2", dispatched_at=_JUST_NOW), + ] + report = run_ci_watcher( + tasks, + poll=lambda t: CiPollResult.terminal({"conclusion": "success"}), + on_resume=flaky_resume, + on_park=rec.on_park, + now=_NOW, + ) + assert report.examined == 2 + # The bad task is recorded as a PARKED_ERROR; the good one resumed. + actions = {o.thread_id: o.action for o in report.outcomes} + assert actions["bad"] is CiWatchAction.PARKED_ERROR + assert actions["good"] is CiWatchAction.RESUMED + assert raised_for == ["bad"] + + +def test_park_callback_failure_downgrades_to_error_outcome() -> None: + def park_boom(task: PendingCiTask) -> None: + raise RuntimeError("park write failed") + + report = run_ci_watcher( + [_task(dispatched_at=_LONG_AGO)], + poll=lambda t: CiPollResult.pending(), + on_resume=lambda t, r: None, + on_park=park_boom, + timeout=timedelta(minutes=30), + now=_NOW, + ) + # The intended action was a timeout-park, but the park callback raised, so the + # outcome is recorded as PARKED_ERROR (the sweep continues either way). + assert report.outcomes[0].action is CiWatchAction.PARKED_ERROR + assert "RuntimeError" in (report.outcomes[0].error or "") + + +# --------------------------------------------------------------------------- # +# default_ci_poller — reuses ci_fetcher read-only, classifies the result +# --------------------------------------------------------------------------- # + + +class _FakeResponse: + def __init__(self, status: int, body: dict) -> None: + self.status_code = status + self._body = body + + def json(self) -> dict: + return self._body + + +class _FakeClient: + """A read-only ``requests``-like client: records GETs, never writes.""" + + def __init__(self, response: _FakeResponse) -> None: + self._response = response + self.gets: list[str] = [] + + def get(self, url: str, *, timeout: float) -> _FakeResponse: + self.gets.append(url) + return self._response + + +def test_default_poller_terminal_on_success_conclusion() -> None: + client = _FakeClient(_FakeResponse(200, {"id": 12345, "conclusion": "success"})) + poll = default_ci_poller(owner="o", repo="r", client=client) + result = poll(_task(run_id="12345")) + assert result.outcome.value == "terminal" + assert result.result is not None + assert result.result["conclusion"] == "success" + # Exactly one read-only GET issued. + assert len(client.gets) == 1 + + +def test_default_poller_pending_on_in_progress_run() -> None: + # An in-progress run has conclusion=None -> fetch_ci_result returns None. + client = _FakeClient(_FakeResponse(200, {"id": 12345, "conclusion": None})) + poll = default_ci_poller(owner="o", repo="r", client=client) + result = poll(_task(run_id="12345")) + assert result.outcome.value == "pending" + assert result.result is None + + +def test_default_poller_pending_on_http_error_status() -> None: + # A 404/5xx makes fetch_ci_result return None; the poller classifies it as + # pending (the timeout branch in the sweep is the fail-closed backstop, and a + # never-resolving run parks on timeout). + client = _FakeClient(_FakeResponse(404, {})) + poll = default_ci_poller(owner="o", repo="r", client=client) + result = poll(_task(run_id="12345")) + assert result.outcome.value == "pending" + + +def test_default_poller_error_when_fetch_raises() -> None: + class _BoomClient: + def get(self, url: str, *, timeout: float): + raise RuntimeError("network down") + + # fetch_ci_result itself swallows GET errors to None, so the poller sees + # pending; but a defensive wrapper still classifies a raised fetch as error. + # Here we drive the error path by passing a client whose .get raises AND + # bypass fetch_ci_result's own swallow via a poller that re-raises is not + # possible; instead assert the read-only GET error degrades to pending (the + # timeout backstop parks it). This documents the boundary. + poll = default_ci_poller(owner="o", repo="r", client=_BoomClient()) + result = poll(_task(run_id="12345")) + assert result.outcome.value == "pending" diff --git a/agent-team/tests/test_coordinator.py b/agent-team/tests/test_coordinator.py index 15afb11..375036e 100644 --- a/agent-team/tests/test_coordinator.py +++ b/agent-team/tests/test_coordinator.py @@ -18,8 +18,9 @@ Every test injects: from __future__ import annotations import queue -from datetime import timedelta +from datetime import datetime, timedelta, timezone from pathlib import Path +from types import SimpleNamespace from typing import Any import pytest @@ -87,6 +88,7 @@ def _make_coordinator( alarm_hook: Any = None, build_plan_node: Any = None, notify: Any = None, + draft_pr_provider: Any = None, ) -> Coordinator: """Build a Coordinator wired with an in-memory saver + the stub clarify node.""" saver = _Saver() @@ -100,6 +102,7 @@ def _make_coordinator( deadline_window=deadline_window, alarm_hook=alarm_hook, notify=notify, + draft_pr_provider=draft_pr_provider, ) @@ -476,6 +479,83 @@ def test_plan_ready_milestone_threads_under_root_and_presents_plan( assert "Summary:" in message # the plan is actually presented +def test_verify_pass_emits_draft_pr_lifecycle_notice( + db_path: Path, monkeypatch: pytest.MonkeyPatch +) -> None: + """A terminal CI PASS posts the POSITIVE draft-PR notice (run link), NOT the park ALARM. + + Unit B3: when a task's CI run reaches terminal PASS and the draft PR is + opened (the APPROVED route — recorded as a verify-stage PASS verdict on a + task settled at DONE), ``_post_resume_followups`` emits a ``#agent-team`` + LIFECYCLE notice via the positive notify sink — distinct from the P2 + plan-ready terminus (which also settles at DONE) and from the park-ALARM + path. The notice carries the GitHub Actions run link built from the captured + run id + the configured owner/repo. + """ + monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "Sea-Haven-Industries") + monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "demo-repo") + + posted: list[tuple[str, str | None]] = [] + alarmed: list[str] = [] + coord = _make_coordinator( + db_path, + notify=lambda message, thread_ts=None: posted.append((message, thread_ts)), + alarm_hook=alarmed.append, + ) + coord.setup() + thread_id = coord.start_task( + task_text="ship the widget", transport_name="slack", slack_thread_ts="ROOT.TS" + ) + + # Inject the settled P3 verify-PASS terminal state (CI terminal PASS + draft + # PR opened on the APPROVED route): DONE, with a verify-stage PASS verdict + # carrying the dispatched run id. update_state clears the pending interrupt, + # so _post_resume_followups treats the thread as settled (no open question). + coord.graph.update_state( + graph_mod.thread_config(thread_id), + { + "status": "done", + "current_phase": "done", + "run_id": "987654321", + "review_verdicts": [ + {"stage": "verify", "decision": "pass", "run_id": "987654321"} + ], + }, + ) + + coord._post_resume_followups([SimpleNamespace(thread_id=thread_id)]) + + notices = [p for p in posted if "draft PR opened" in p[0]] + assert len(notices) == 1 + message, thread_ts = notices[0] + assert thread_ts == "ROOT.TS" # threaded under the task root + assert "CI PASSED" in message + # The positive path: a run link, NOT the park-ALARM path. + assert ( + "https://github.com/Sea-Haven-Industries/demo-repo/actions/runs/987654321" + in message + ) + assert alarmed == [] + # It must NOT be misreported as the P2 plan-ready terminus. + assert not any("plan ready for review" in p[0] for p in posted) + + +def test_verify_pass_verdict_discriminates_from_plan_ready() -> None: + """The verify-PASS discriminator ignores a plan-ready DONE (no verify verdict).""" + # A plan-ready terminus settles at DONE but carries no verify verdict. + assert Coordinator._verify_pass_verdict({"review_verdicts": []}) is None + # A non-PASS verify verdict (e.g. a loop-back FAIL) is not the PASS terminus. + assert ( + Coordinator._verify_pass_verdict( + {"review_verdicts": [{"stage": "verify", "decision": "fail"}]} + ) + is None + ) + # A verify-stage PASS verdict is the P3 PASS terminus. + verdict = {"stage": "verify", "decision": "pass", "run_id": "42"} + assert Coordinator._verify_pass_verdict({"review_verdicts": [verdict]}) == verdict + + # --------------------------------------------------------------------------- # # tick — deadline sweep + park ALARM + drain # --------------------------------------------------------------------------- # @@ -518,6 +598,105 @@ def test_tick_drains_pending_resume(db_path: Path) -> None: assert state["status"] == "done" +# --------------------------------------------------------------------------- # +# tick — draft-PR runaway/stale sweep wiring (P3 A4) +# --------------------------------------------------------------------------- # + + +def test_draft_pr_sweep_is_noop_without_provider(db_path: Path) -> None: + """The INERT default: no ``draft_pr_provider`` means the sweep does nothing.""" + posted: list[Any] = [] + coord = _make_coordinator(db_path, notify=lambda m, **k: posted.append(m)) + coord.setup() + assert coord._draft_pr_monitor_sweep() is None + assert posted == [] + + +def test_draft_pr_sweep_alarms_on_runaway_via_emit(db_path: Path) -> None: + """A runaway burst routes a ``#agent-team`` ALARM through the notify sink.""" + from agent_team.draft_pr_monitor import DraftPr + + now = datetime.now(timezone.utc).isoformat() + burst = [ + DraftPr(number=i, opened_at=now, updated_at=now) for i in range(4) + ] # > 3 within window + posted: list[str] = [] + coord = _make_coordinator( + db_path, + notify=lambda m, **k: posted.append(m), + draft_pr_provider=lambda: burst, + ) + coord.setup() + + report = coord._draft_pr_monitor_sweep() + assert report is not None + assert report.alarmed == 1 + assert any("draft-PR runaway" in m for m in posted) + # The ALARM surfaces the operator remediation, never a self-stop. + assert any("systemctl stop" in m for m in posted) + + +def test_draft_pr_sweep_reminds_on_stale_via_emit(db_path: Path) -> None: + """A stale (>7d idle) draft PR routes a single reminder through the sink.""" + from agent_team.draft_pr_monitor import DraftPr + + stale_iso = (datetime.now(timezone.utc) - timedelta(days=10)).isoformat() + posted: list[str] = [] + coord = _make_coordinator( + db_path, + notify=lambda m, **k: posted.append(m), + draft_pr_provider=lambda: [DraftPr(number=7, updated_at=stale_iso)], + ) + coord.setup() + + report = coord._draft_pr_monitor_sweep() + assert report is not None + assert report.stale_reminded == 1 + assert any("draft PR #7" in m and "7 days" in m for m in posted) + + +def test_draft_pr_sweep_failsoft_on_enumeration_error(db_path: Path) -> None: + """An enumeration that raises is swallowed (returns None) — tick never breaks.""" + + def _boom() -> list[Any]: + raise RuntimeError("github unreachable") + + coord = _make_coordinator(db_path, draft_pr_provider=_boom) + coord.setup() + # Must not raise; sweep returns None on a failed enumeration. + assert coord._draft_pr_monitor_sweep() is None + # And a full tick (which calls the sweep) likewise survives. + assert coord.tick() == [] + + +def test_draft_pr_sweep_memory_persists_across_ticks(db_path: Path) -> None: + """The flapping-backoff memory persists for the daemon's lifetime. + + A sustained runaway condition ALARMs once, then is suppressed on the next + sweep (the per-daemon ``MonitorMemory`` carries the cooldown watermark). + """ + from agent_team.draft_pr_monitor import DraftPr + + now = datetime.now(timezone.utc).isoformat() + burst = [DraftPr(number=i, opened_at=now, updated_at=now) for i in range(5)] + posted: list[str] = [] + coord = _make_coordinator( + db_path, + notify=lambda m, **k: posted.append(m), + draft_pr_provider=lambda: burst, + ) + coord.setup() + + first = coord._draft_pr_monitor_sweep() + second = coord._draft_pr_monitor_sweep() + assert first is not None and second is not None + assert first.alarmed == 1 + assert second.alarmed == 0 + assert second.alarm_suppressed == 1 + # Exactly one ALARM posted despite two sweeps of the same condition. + assert sum("draft-PR runaway" in m for m in posted) == 1 + + # --------------------------------------------------------------------------- # # recover — startup convergence (§3.3.1) # --------------------------------------------------------------------------- # @@ -1318,3 +1497,165 @@ def test_parked_message_infers_phase_when_current_phase_is_parked( coord._post_resume_followups([_resume_result("abcd1234ef00")]) assert "Reached phase: review" in msgs[0] assert "Reached phase: parked" not in msgs[0] + + +# --------------------------------------------------------------------------- # +# CI-watcher sweep in tick (§3.3.2 Decision 2) — async resume-on-CI-complete +# --------------------------------------------------------------------------- # + + +def test_tick_ci_watch_noop_without_seams(db_path: Path) -> None: + """With no CI-watcher seams wired, tick() runs no CI sweep (the default).""" + coord = _make_coordinator(db_path) + coord.setup() + # _ci_watch returns None (skipped) when the seams are absent. + assert coord._ci_watch() is None + # tick still works as before. + assert coord.tick() == [] + + +def _ci_coordinator(db_path: Path, *, pending, poller): + """A Coordinator wired with CI-watcher seams + an in-memory saver.""" + saver = _Saver() + return Coordinator( + db_path=db_path, + transport=FakeTransport(), + build_clarify_node=lambda: graph_mod.clarify_node, + build_checkpointer=lambda _path: saver, + ci_pending_provider=lambda: pending, + ci_poller=poller, + ) + + +class _AwaitingCiSnap: + """A StateSnapshot stand-in interrupted at VERIFY awaiting a given run.""" + + def __init__(self, run_id: str) -> None: + self.interrupts = ( + SimpleNamespace(value={"awaiting_ci": True, "run_id": run_id}), + ) + self.next = ("verify",) + self.values: dict[str, Any] = {} + + +def test_tick_ci_watch_resumes_on_terminal_conclusion(db_path: Path) -> None: + """A terminated run drives a graph resume of the suspended VERIFY task. + + The resume MUST go through the turn-guarded resume worker (single-flight + + awaiting-CI guard), NOT a bare ``graph.invoke``: the guard re-reads the live + snapshot and only invokes while the thread is still suspended at VERIFY + awaiting THIS run, so a double-resume cannot corrupt state. + """ + from agent_team.ci_watcher import CiPollResult, PendingCiTask + + task = PendingCiTask( + thread_id="t-ci", run_id="999", dispatched_at="2026-06-23T11:00:00+00:00" + ) + coord = _ci_coordinator( + db_path, + pending=[task], + poller=lambda t: CiPollResult.terminal( + {"run_id": "999", "conclusion": "success", "diff_hash": "h"} + ), + ) + coord.setup() + + # Thread is genuinely suspended at VERIFY awaiting run 999, so the worker's + # awaiting-CI guard passes and the resume applies exactly once. + coord._graph.get_state = lambda _cfg: _AwaitingCiSnap("999") # type: ignore[assignment] + invoked: list[tuple[Any, Any]] = [] + coord._graph.invoke = lambda inp, cfg: invoked.append((inp, cfg)) # type: ignore[assignment] + + report = coord._ci_watch() + assert report is not None + assert report.resumed == 1 + # The graph was driven forward for the suspended task's thread. + assert len(invoked) == 1 + _inp, cfg = invoked[0] + assert cfg["configurable"]["thread_id"] == "t-ci" + + +def test_ci_resume_skips_when_thread_already_advanced(db_path: Path) -> None: + """A double-resume is a no-op: if the thread already left the CI gate (no + awaiting-CI interrupt — resumed/parked/done), the turn-guarded path skips the + bare invoke so state cannot be corrupted.""" + from agent_team.ci_watcher import CiPollResult, PendingCiTask + + task = PendingCiTask( + thread_id="t-gone", run_id="5", dispatched_at="2026-06-23T11:00:00+00:00" + ) + coord = _ci_coordinator( + db_path, + pending=[task], + poller=lambda t: CiPollResult.terminal({"run_id": "5", "conclusion": "ok"}), + ) + coord.setup() + + # Already advanced: no pending interrupts -> guard must skip the invoke. + coord._graph.get_state = lambda _cfg: SimpleNamespace(interrupts=(), next=()) # type: ignore[assignment] + invoked: list[Any] = [] + coord._graph.invoke = lambda inp, cfg: invoked.append((inp, cfg)) # type: ignore[assignment] + + report = coord._ci_watch() + assert report is not None + # The watcher still records it as RESUMED (the resume callable returned + # without raising), but the bare invoke never fired — the guard held. + assert invoked == [] + + +def test_tick_ci_watch_parks_on_timeout(db_path: Path) -> None: + """A run that never terminates within the timeout parks the task + ALARMs.""" + from agent_team.ci_watcher import CiPollResult, PendingCiTask + + alarms: list[str] = [] + task = PendingCiTask( + thread_id="t-slow", run_id="42", dispatched_at="2000-01-01T00:00:00+00:00" + ) + saver = _Saver() + coord = Coordinator( + db_path=db_path, + transport=FakeTransport(), + build_clarify_node=lambda: graph_mod.clarify_node, + build_checkpointer=lambda _path: saver, + ci_pending_provider=lambda: [task], + ci_poller=lambda t: CiPollResult.pending(), + alarm_hook=alarms.append, + ) + coord.setup() + + updated: list[tuple[Any, Any]] = [] + coord._graph.update_state = lambda cfg, delta: updated.append((cfg, delta)) # type: ignore[assignment] + + report = coord._ci_watch() + assert report is not None + assert report.parked_timeout == 1 + # The task was durably parked and an ALARM was raised. + assert updated and updated[0][1]["status"] == "parked" + assert alarms == ["t-slow"] + + +def test_tick_ci_watch_parks_fail_closed_on_error(db_path: Path) -> None: + """A poll error parks the task (fail-closed), even within the timeout window.""" + from agent_team.ci_watcher import CiPollResult, PendingCiTask + + alarms: list[str] = [] + task = PendingCiTask( + thread_id="t-err", run_id="7", dispatched_at="2026-06-23T11:59:00+00:00" + ) + saver = _Saver() + coord = Coordinator( + db_path=db_path, + transport=FakeTransport(), + build_clarify_node=lambda: graph_mod.clarify_node, + build_checkpointer=lambda _path: saver, + ci_pending_provider=lambda: [task], + ci_poller=lambda t: CiPollResult.error(), + alarm_hook=alarms.append, + ) + coord.setup() + coord._graph.update_state = lambda cfg, delta: None # type: ignore[assignment] + + report = coord._ci_watch() + assert report is not None + assert report.parked_error == 1 + assert alarms == ["t-err"] diff --git a/agent-team/tests/test_dispatcher.py b/agent-team/tests/test_dispatcher.py index a6802ef..6008abe 100644 --- a/agent-team/tests/test_dispatcher.py +++ b/agent-team/tests/test_dispatcher.py @@ -15,9 +15,12 @@ from agent_team.dispatcher import ( MAX_DIFF_BYTES, DispatcherError, DispatchInputs, + DispatchResult, build_dispatch_inputs, dispatch_apply_verify, head_branch_for, + run_name_for, + select_run_id, ) from agent_team.state_store import compute_content_hash @@ -116,7 +119,16 @@ def test_dispatch_pushes_then_fires_with_correct_inputs() -> None: order.append("dispatch") fired.calls.append({"owner": owner, "repo": repo, "inputs": inputs, "ref": ref}) - di = dispatch_apply_verify( + located = _Recorder() + + def locator(*, owner, repo, task_id, since_iso): + order.append("locate") + located.calls.append( + {"owner": owner, "repo": repo, "task_id": task_id, "since_iso": since_iso} + ) + return "27990718108" + + result = dispatch_apply_verify( owner="Sea-Haven-Industries", repo="orchestrator", task_id=TASK, @@ -124,12 +136,18 @@ def test_dispatch_pushes_then_fires_with_correct_inputs() -> None: declared_scope=SCOPE, pusher=pusher, dispatcher=dispatcher, + locator=locator, ) - assert isinstance(di, DispatchInputs) - # Branch is pushed BEFORE the workflow is dispatched (the draft-PR step opens - # against an already-pushed --head). - assert order == ["push", "dispatch"] + assert isinstance(result, DispatchResult) + assert isinstance(result.inputs, DispatchInputs) + # run_id is captured from the locator and surfaced for the verifier. + assert result.run_id == "27990718108" + assert result.correlation_tag == TASK + assert result.dispatched_at # stamped, non-empty + # Push BEFORE dispatch BEFORE locate (the run can only be located after it is + # triggered, and the branch must exist before the run reaches the PR step). + assert order == ["push", "dispatch", "locate"] assert pushed.calls[0]["head"] == f"agent-team/apply/{TASK}" assert pushed.calls[0]["diff"] == DIFF # The dispatch carries all six inputs, including the b64 diff + head branch. @@ -138,6 +156,31 @@ def test_dispatch_pushes_then_fires_with_correct_inputs() -> None: assert base64.b64decode(inputs["diff_b64"]).decode("utf-8") == DIFF assert inputs["expected_diff_hash"] == compute_content_hash(DIFF.encode("utf-8")) assert fired.calls[0]["ref"] == "main" + # The locator is keyed by THIS task and the dispatched-at watermark. + assert located.calls[0]["task_id"] == TASK + assert located.calls[0]["since_iso"] == result.dispatched_at + + +def test_dispatch_returns_none_run_id_when_locator_cannot_resolve() -> None: + # A fired-but-unlocatable run fails closed (None run_id); never raises here. + result = dispatch_apply_verify( + owner="o", + repo="r", + task_id=TASK, + diff_text=DIFF, + declared_scope=SCOPE, + pusher=lambda **_k: None, + dispatcher=lambda **_k: None, + locator=lambda **_k: None, + ) + assert isinstance(result, DispatchResult) + assert result.run_id is None + assert result.dispatched_at # still stamped for the CI-watch timeout + + +def test_run_name_for_matches_workflow_run_name_convention() -> None: + # Mirrors run-name: "agent-team-apply ${{ inputs.task_id }}" in the workflow. + assert run_name_for(TASK) == f"agent-team-apply {TASK}" def test_dispatch_does_not_fire_if_push_fails() -> None: @@ -177,3 +220,111 @@ def test_unsafe_owner_repo_rejected(owner: str, repo: str) -> None: pusher=lambda **_k: None, dispatcher=lambda **_k: None, ) + + +# --------------------------------------------------------------------------- # +# select_run_id (pure; anti-stale on rapid re-dispatch of the SAME task_id) +# --------------------------------------------------------------------------- # + + +def _row(db_id, *, created, status="completed", conclusion=None): + """A minimal ``gh run list`` row for the apply/verify run-name of TASK.""" + return { + "databaseId": db_id, + "name": run_name_for(TASK), + "createdAt": created, + "status": status, + "conclusion": conclusion, + } + + +def test_select_run_id_skips_cancelled_prior_run_on_re_dispatch() -> None: + # Rapid re-dispatch of the SAME task_id: the concurrency group cancelled the + # OLDER run, and a NEWER run is now in progress. We must bind to the newer, + # active run — never the older cancelled one (it carries the prior verdict). + older_cancelled = _row( + 100, created="2026-06-23T10:00:00Z", status="completed", conclusion="cancelled" + ) + newer_active = _row( + 200, created="2026-06-23T10:05:00Z", status="in_progress", conclusion=None + ) + runs = [newer_active, older_cancelled] + chosen = select_run_id(runs, task_id=TASK, floor_iso="2026-06-23T09:58:00Z") + assert chosen == "200" + + +def test_select_run_id_skips_cancelled_even_when_it_is_newest() -> None: + # Defensive: a cancelled run is NEVER selected even if its createdAt is the + # greatest — it is the superseded run, not ours. + active = _row(300, created="2026-06-23T10:00:00Z", status="queued", conclusion=None) + newest_cancelled = _row( + 400, created="2026-06-23T10:10:00Z", status="completed", conclusion="cancelled" + ) + runs = [active, newest_cancelled] + chosen = select_run_id(runs, task_id=TASK, floor_iso="2026-06-23T09:58:00Z") + assert chosen == "300" + + +def test_select_run_id_prefers_newest_active_over_older_completed() -> None: + # An older legitimately-completed run plus a newer active run -> the active, + # newest run wins (the freshly-triggered one with no conclusion yet). + older_done = _row( + 500, created="2026-06-23T10:00:00Z", status="completed", conclusion="success" + ) + newer_active = _row( + 600, created="2026-06-23T10:05:00Z", status="in_progress", conclusion=None + ) + chosen = select_run_id( + [older_done, newer_active], task_id=TASK, floor_iso="2026-06-23T09:58:00Z" + ) + assert chosen == "600" + + +def test_select_run_id_falls_back_to_newest_non_cancelled_when_none_active() -> None: + # No active runs (e.g. a fast run already concluded by the time we poll): + # fall back to the newest NON-cancelled run overall. + older = _row( + 700, created="2026-06-23T10:00:00Z", status="completed", conclusion="success" + ) + newer = _row( + 800, created="2026-06-23T10:05:00Z", status="completed", conclusion="failure" + ) + cancelled = _row( + 900, created="2026-06-23T10:09:00Z", status="completed", conclusion="cancelled" + ) + chosen = select_run_id( + [older, newer, cancelled], task_id=TASK, floor_iso="2026-06-23T09:58:00Z" + ) + assert chosen == "800" + + +def test_select_run_id_respects_created_floor_and_run_name() -> None: + # Below-floor runs and other-task runs are not matched. + below_floor = _row( + 1000, created="2026-06-23T09:00:00Z", status="in_progress", conclusion=None + ) + other_task = { + "databaseId": 1100, + "name": "agent-team-apply other-task", + "createdAt": "2026-06-23T10:00:00Z", + "status": "in_progress", + "conclusion": None, + } + assert ( + select_run_id( + [below_floor, other_task], task_id=TASK, floor_iso="2026-06-23T09:58:00Z" + ) + is None + ) + + +def test_select_run_id_returns_none_when_only_cancelled_matches() -> None: + # If the only matching run is cancelled, there is nothing to bind to -> None + # (the caller fails closed: no run_id -> verify BLOCKs/parks). + only_cancelled = _row( + 1200, created="2026-06-23T10:00:00Z", status="completed", conclusion="cancelled" + ) + assert ( + select_run_id([only_cancelled], task_id=TASK, floor_iso="2026-06-23T09:58:00Z") + is None + ) diff --git a/agent-team/tests/test_graph.py b/agent-team/tests/test_graph.py index b4b58be..afb8a15 100644 --- a/agent-team/tests/test_graph.py +++ b/agent-team/tests/test_graph.py @@ -616,6 +616,128 @@ def test_p3_graph_inert_default_parks_at_verify(restore_review_invoker) -> None: assert final["status"] == TaskStatus.PARKED.value +def _p3_graph_with_dispatch(review_text: str, *, ci_result_fetcher, dispatch_node): + """Compile a P3+ graph: review -> build -> DISPATCH -> verify. + + Same as ``_p3_graph`` but splices a ``dispatch_node`` between BUILD and + VERIFY so the reordered topology (design §4 Decision 1) can be driven end to + end — DISPATCH writes ``state["run_id"]`` before VERIFY reads it. + """ + from agent_team.nodes import review_loop + from agent_team.nodes.build_verify_subgraph import ( + make_build_node, + make_verify_node, + route_after_verify, + ) + from agent_team.nodes.verifier import VerifierConfig + + review_loop.set_review_invoker(lambda prompt, **kw: review_text) + + def fake_builder(*, plan, config): + return _p3_diff() + + build_node = make_build_node(diff_builder=fake_builder) + # No static expected_run_id: the gate must bind to the run id DISPATCH wrote. + verify_node = make_verify_node( + VerifierConfig(expected_run_id=None, allowed_scope=["src"]), + ci_result_fetcher=ci_result_fetcher, + ) + return build_graph( + checkpointer=_Saver(), + live_plan_node=_p3_plan_stub, + review_node=review_loop.bind_review_node(), + route_review=review_loop.route_after_review, + build_verify=(build_node, verify_node, route_after_verify), + dispatch_node=dispatch_node, + ) + + +def test_build_graph_dispatch_node_requires_build_verify() -> None: + """dispatch_node without build_verify is a wiring error (nothing to splice).""" + with pytest.raises(ValueError, match="dispatch_node requires build_verify"): + build_graph(dispatch_node=lambda state: {}) + + +def test_p3_dispatch_runs_before_verify_and_supplies_run_id( + restore_review_invoker, +) -> None: + """BUILD -> DISPATCH -> VERIFY: DISPATCH captures run_id BEFORE VERIFY reads it. + + The verify node is wired with NO static expected_run_id, so the only way the + authenticated-pass gate can bind a verdict is if DISPATCH wrote + ``state["run_id"]`` first. The fetcher keys its conclusion to the dispatched + run id and asserts it observes that id — proving DISPATCH ran before VERIFY. + """ + from agent_team.state_store import compute_content_hash + + observed: dict[str, object] = {} + dispatched_run_id = "r-dispatched-007" + + def dispatch_node(state): + # Mirror dispatch_invoker's contract: persist the located run identity so + # the downstream verifier binds the gate to THIS task's dispatched run. + return {"run_id": dispatched_run_id, "dispatched_at": "2026-06-23T00:00:00Z"} + + def pass_fetcher(state): + # The fetcher only sees state["run_id"] if DISPATCH already ran. + observed["run_id"] = state.get("run_id") + diff_hash = compute_content_hash(_p3_diff().encode("utf-8")) + return { + "run_id": state.get("run_id"), + "conclusion": "success", + "diff_hash": diff_hash, + } + + graph = _p3_graph_with_dispatch( + "VERDICT: APPROVE\nlooks solid", + ci_result_fetcher=pass_fetcher, + dispatch_node=dispatch_node, + ) + thread_id, _ = start_task(graph, transport="slack") + final = resume_task(graph, thread_id=thread_id, answer="scope is X") + + # VERIFY observed the run id DISPATCH wrote -> DISPATCH ran first. + assert observed["run_id"] == dispatched_run_id + # And the per-task-bound authenticated pass cleared the gate -> DONE terminus. + assert final["current_phase"] == Phase.DONE.value + assert final["status"] == TaskStatus.DONE.value + assert final["run_id"] == dispatched_run_id + + +def test_p3_dispatch_node_present_in_graph_topology(restore_review_invoker) -> None: + """The DISPATCH vertex is wired between BUILD and VERIFY when injected.""" + from agent_team.graph import BUILD_NODE, DISPATCH_NODE, VERIFY_NODE + + graph = _p3_graph_with_dispatch( + "VERDICT: APPROVE\nlooks solid", + ci_result_fetcher=lambda state: None, + dispatch_node=lambda state: {"run_id": "r"}, + ) + g = graph.get_graph() + nodes = set(g.nodes) + assert {BUILD_NODE, DISPATCH_NODE, VERIFY_NODE} <= nodes + + # The linear order is BUILD -> DISPATCH -> VERIFY (no direct BUILD -> VERIFY). + edges = {(e.source, e.target) for e in g.edges} + assert (BUILD_NODE, DISPATCH_NODE) in edges + assert (DISPATCH_NODE, VERIFY_NODE) in edges + assert (BUILD_NODE, VERIFY_NODE) not in edges + + +def test_p3_no_dispatch_falls_back_to_build_then_verify(restore_review_invoker) -> None: + """Without a dispatch node the order falls back to BUILD -> VERIFY directly.""" + from agent_team.graph import BUILD_NODE, DISPATCH_NODE, VERIFY_NODE + + graph = _p3_graph("VERDICT: APPROVE\nlooks solid", ci_result_fetcher=lambda s: None) + g = graph.get_graph() + nodes = set(g.nodes) + assert {BUILD_NODE, VERIFY_NODE} <= nodes + assert DISPATCH_NODE not in nodes + + edges = {(e.source, e.target) for e in g.edges} + assert (BUILD_NODE, VERIFY_NODE) in edges + + def test_p3_graph_route_constants_mirror_subgraph_by_value() -> None: """graph.py's P3 route ids match the subgraph module by value (no cycle).""" from agent_team.nodes import build_verify_subgraph as bvs diff --git a/agent-team/tests/test_no_write_token.py b/agent-team/tests/test_no_write_token.py new file mode 100644 index 0000000..e9d0af4 --- /dev/null +++ b/agent-team/tests/test_no_write_token.py @@ -0,0 +1,395 @@ +"""Unit tests for ``scripts/assert_no_write_token.py`` (P3 no-write-token audit). + +The audit asserts the box invariant from ``docs/P3-PHASE0-DESIGN.md`` and +``ci/README.md``: the ``pull-requests: write`` GitHub App credential +(``AGENT_APPLY_APP_ID`` / ``AGENT_APPLY_APP_PRIVATE_KEY``) lives ONLY as an +Actions secret, and the always-on box holds **no standing write token**. + +These tests never touch real box paths: they grep temp files and a patched +``os.environ``. ``scripts/`` is not a package, so (matching the ``run-team.py`` +idiom) the module is loaded from its file path via :mod:`importlib`. +""" + +from __future__ import annotations + +import importlib.util +import os +from pathlib import Path +from types import ModuleType + +import pytest + +_SCRIPT_PATH = ( + Path(__file__).resolve().parents[1] / "scripts" / "assert_no_write_token.py" +) + +# A sample PEM private key header for each algorithm the design calls out. We +# assert the regex covers RSA / EC / OPENSSH (and a bare PKCS#8 block). The PEM +# markers are ASSEMBLED at runtime (not written as contiguous literals) so this +# test data does not itself trip the repo's gitleaks pre-push backstop — the +# values are still full PEM blocks at runtime, which is what the detector sees. +_BEGIN = "-----BEGIN " +_END = "-----END " +_PEM_BODY = "\nMIIB...redacted...\n" + + +def _pem(alg: str) -> str: + label = f"{alg} PRIVATE KEY-----" if alg else "PRIVATE KEY-----" + return f"{_BEGIN}{label}{_PEM_BODY}{_END}{label}" + + +_RSA_KEY = _pem("RSA") +_EC_KEY = _pem("EC") +_OPENSSH_KEY = _pem("OPENSSH") +_PKCS8_KEY = _pem("") + + +def _load_script() -> ModuleType: + """Load the (non-package) audit script from its file path.""" + spec = importlib.util.spec_from_file_location("assert_no_write_token", _SCRIPT_PATH) + assert spec is not None and spec.loader is not None + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +@pytest.fixture(scope="module") +def mod() -> ModuleType: + """The loaded audit module (loaded once per test module).""" + return _load_script() + + +# --------------------------------------------------------------------------- # +# scan_environ +# --------------------------------------------------------------------------- # + + +def test_clean_environ_has_no_findings(mod: ModuleType) -> None: + env = {"PATH": "/usr/bin", "AGENT_TEAM_REPO_OWNER": "Sea-Haven-Industries"} + assert mod.scan_environ(env) == [] + + +def test_forbidden_app_id_env_var_is_flagged(mod: ModuleType) -> None: + findings = mod.scan_environ({"AGENT_APPLY_APP_ID": "123456"}) + assert len(findings) == 1 + assert "AGENT_APPLY_APP_ID" in findings[0] + + +def test_forbidden_app_private_key_env_var_is_flagged(mod: ModuleType) -> None: + # The var trips both the forbidden-name check AND the PEM-value check (SEC-03) + # — both are legitimate findings; assert it is flagged (not an exact count). + findings = mod.scan_environ({"AGENT_APPLY_APP_PRIVATE_KEY": _RSA_KEY}) + assert findings + assert all("AGENT_APPLY_APP_PRIVATE_KEY" in fn for fn in findings) + assert any("PRIVATE KEY" in fn for fn in findings) + + +def test_write_token_shaped_name_is_flagged(mod: ModuleType) -> None: + for name in ( + "GITHUB_APP_TOKEN", + "GH_PAT_TOKEN", + "GITHUB_WRITE_TOKEN", + "GH_APPLY_KEY", + ): + findings = mod.scan_environ({name: "whatever"}) + assert findings, f"{name} should be flagged by the write-name heuristic" + + +def test_write_capable_token_value_is_flagged_regardless_of_name( + mod: ModuleType, +) -> None: + # A benign-looking var name but a write-capable PAT value. + findings = mod.scan_environ({"SOME_VAR": "ghp_" + "a" * 36}) + assert len(findings) == 1 + assert "SOME_VAR" in findings[0] + + findings = mod.scan_environ({"OTHER": "github_pat_" + "b" * 30}) + assert len(findings) == 1 + + findings = mod.scan_environ({"INSTALL": "ghs_" + "c" * 36}) + assert len(findings) == 1 + + # user-to-server (ghu_) and refresh (ghr_) tokens are also write-risk. + findings = mod.scan_environ({"U2S": "ghu_" + "d" * 36}) + assert len(findings) == 1 + + findings = mod.scan_environ({"REFRESH": "ghr_" + "e" * 36}) + assert len(findings) == 1 + + +def test_read_only_tokens_are_not_flagged(mod: ModuleType) -> None: + # Known read-only / non-GitHub tokens must not trip the audit. + env = { + "AGENT_TEAM_CI_READ_TOKEN": "ghp_" + + "r" * 36, # read-only by contract, allowlisted + "AGENT_TEAM_API_TOKEN": "secret-bearer", + "SLACK_APP_TOKEN": "xapp-1-abc", + "SLACK_BOT_TOKEN": "xoxb-abc", + "CLAUDE_CODE_OAUTH_TOKEN": "oauth-abc", + "ANTHROPIC_API_KEY": "sk-ant-abc", + } + assert mod.scan_environ(env) == [] + + +def test_plain_github_token_fallback_is_not_flagged_by_name(mod: ModuleType) -> None: + # The read-only GITHUB_TOKEN runtime fallback should not match the *name* + # heuristic (it carries no GH_*_WRITE/APP/PAT shape). + assert mod.scan_environ({"GITHUB_TOKEN": "a-non-write-shaped-value"}) == [] + + +def test_github_token_value_is_still_write_value_scanned(mod: ModuleType) -> None: + # SEC-04: GITHUB_TOKEN is name-exempt from the write *name* heuristic, but a + # write-capable token VALUE parked in it must still be flagged. + findings = mod.scan_environ({"GITHUB_TOKEN": "ghp_" + "a" * 36}) + assert len(findings) == 1 + assert "GITHUB_TOKEN" in findings[0] + + +# SEC-03: a PEM private key exported under a *benign* env name must be flagged +# (the App key value check is NOT name-exempt). +def test_pem_private_key_under_benign_env_name_is_flagged(mod: ModuleType) -> None: + findings = mod.scan_environ({"GH_APP_KEY": _RSA_KEY}) + assert any("PRIVATE KEY" in fn for fn in findings) + assert any("GH_APP_KEY" in fn for fn in findings) + + +@pytest.mark.parametrize("key", [_RSA_KEY, _EC_KEY, _OPENSSH_KEY, _PKCS8_KEY]) +def test_pem_private_key_value_is_flagged_for_each_algo( + mod: ModuleType, key: str +) -> None: + # Even a totally unremarkable var name carrying any PEM algorithm is flagged. + findings = mod.scan_environ({"HARMLESS": key}) + assert any("PRIVATE KEY" in fn for fn in findings) + + +def test_pem_private_key_in_allowlisted_env_var_is_still_flagged( + mod: ModuleType, +) -> None: + # The value-level PEM check is not exempted by the read-only allowlist. + findings = mod.scan_environ({"AGENT_TEAM_CI_READ_TOKEN": _PKCS8_KEY}) + assert any("PRIVATE KEY" in fn for fn in findings) + + +# SEC-03: the configured App-id appearing as an env VALUE (under any name) must +# be flagged — mirroring scan_env_file's App-id grep. +def test_configured_app_id_value_in_env_is_flagged(mod: ModuleType) -> None: + findings = mod.scan_environ({"SOME_VAR": "installed-987654-here"}, app_id="987654") + assert any("App id value" in fn for fn in findings) + assert any("SOME_VAR" in fn for fn in findings) + + +def test_app_id_value_check_is_skipped_without_app_id(mod: ModuleType) -> None: + # No configured app_id -> the value-substring check does not fire. + assert mod.scan_environ({"SOME_VAR": "987654"}) == [] + + +def test_audit_flags_app_id_value_in_environ(mod: ModuleType, tmp_path: Path) -> None: + # End-to-end: AGENT_APPLY_APP_ID in environ both flags itself and is matched + # against other env VALUES (a benign var carrying the same id is flagged). + clean = tmp_path / "secrev.env" + clean.write_text("ok\n") + findings = mod.audit( + environ={"AGENT_APPLY_APP_ID": "424242", "BENIGN": "id-is-424242"}, + env_files=[str(clean)], + ) + assert any("AGENT_APPLY_APP_ID" in fn for fn in findings) + assert any("BENIGN" in fn and "App id value" in fn for fn in findings) + + +# --------------------------------------------------------------------------- # +# scan_config +# --------------------------------------------------------------------------- # + + +def test_scan_config_none_and_empty(mod: ModuleType) -> None: + assert mod.scan_config(None) == [] + assert mod.scan_config({}) == [] + + +def test_scan_config_flags_forbidden_key(mod: ModuleType) -> None: + findings = mod.scan_config({"AGENT_APPLY_APP_PRIVATE_KEY": _EC_KEY}) + assert len(findings) == 1 + assert "AGENT_APPLY_APP_PRIVATE_KEY" in findings[0] + + +def test_scan_config_flags_write_value(mod: ModuleType) -> None: + findings = mod.scan_config({"token": "ghp_" + "z" * 36}) + assert len(findings) == 1 + + +def test_scan_config_ignores_non_string_values(mod: ModuleType) -> None: + assert mod.scan_config({"timeout": 15, "enabled": True}) == [] + + +# --------------------------------------------------------------------------- # +# scan_env_file +# --------------------------------------------------------------------------- # + + +def test_missing_file_is_not_a_finding(mod: ModuleType, tmp_path: Path) -> None: + assert mod.scan_env_file(tmp_path / "nope.env") == [] + + +def test_clean_file_is_not_a_finding(mod: ModuleType, tmp_path: Path) -> None: + f = tmp_path / "secrev.env" + f.write_text("CLAUDE_CODE_OAUTH_TOKEN=abc\nAGENT_TEAM_CI_READ_TOKEN=def\n") + assert mod.scan_env_file(f) == [] + + +@pytest.mark.parametrize("key", [_RSA_KEY, _EC_KEY, _OPENSSH_KEY, _PKCS8_KEY]) +def test_private_key_header_is_flagged_for_each_algo( + mod: ModuleType, tmp_path: Path, key: str +) -> None: + f = tmp_path / "orchestrator.env" + f.write_text("HARMLESS=1\n" + key + "\n") + findings = mod.scan_env_file(f) + assert any("PRIVATE KEY" in fn for fn in findings) + + +def test_app_credential_name_in_file_is_flagged( + mod: ModuleType, tmp_path: Path +) -> None: + f = tmp_path / "secrev.env" + f.write_text("AGENT_APPLY_APP_ID=987654\n") + findings = mod.scan_env_file(f) + assert any("AGENT_APPLY_APP_ID" in fn for fn in findings) + + +def test_configured_app_id_value_in_file_is_flagged( + mod: ModuleType, tmp_path: Path +) -> None: + f = tmp_path / "secrev.env" + # The literal id value leaking under any var name. + f.write_text("SOMETHING=installed-as-987654-here\n") + findings = mod.scan_env_file(f, app_id="987654") + assert any("App id value" in fn for fn in findings) + + +def test_unreadable_present_file_is_a_finding( + mod: ModuleType, tmp_path: Path, monkeypatch: pytest.MonkeyPatch +) -> None: + f = tmp_path / "secrev.env" + f.write_text("ok\n") + + def _boom(*_a: object, **_k: object) -> str: + raise OSError("permission denied") + + monkeypatch.setattr(Path, "read_text", _boom) + findings = mod.scan_env_file(f) + assert len(findings) == 1 + assert "could not read" in findings[0] + + +# --------------------------------------------------------------------------- # +# audit() orchestration + file resolution +# --------------------------------------------------------------------------- # + + +def test_audit_clean_box_returns_empty(mod: ModuleType, tmp_path: Path) -> None: + clean = tmp_path / "secrev.env" + clean.write_text("CLAUDE_CODE_OAUTH_TOKEN=abc\n") + findings = mod.audit( + environ={"PATH": "/usr/bin"}, + config={"timeout": 15}, + env_files=[str(clean)], + ) + assert findings == [] + + +def test_audit_aggregates_findings_across_scanners( + mod: ModuleType, tmp_path: Path +) -> None: + leaky = tmp_path / "orchestrator.env" + leaky.write_text(_OPENSSH_KEY + "\n") + findings = mod.audit( + environ={"AGENT_APPLY_APP_ID": "42", "GH_PAT_TOKEN": "x"}, + config={"token": "ghp_" + "q" * 36}, + env_files=[str(leaky)], + ) + # env (2) + config (1) + file (1) + assert len(findings) >= 4 + + +def test_audit_uses_configured_app_id_from_environ( + mod: ModuleType, tmp_path: Path +) -> None: + f = tmp_path / "secrev.env" + f.write_text("LEAK=value-555-leaked\n") + findings = mod.audit( + environ={"AGENT_APPLY_APP_ID": "555"}, + env_files=[str(f)], + ) + # AGENT_APPLY_APP_ID in environ is itself a finding, and "555" in the file + # is a second. + assert any("AGENT_APPLY_APP_ID" in fn for fn in findings) + assert any("App id value" in fn for fn in findings) + + +def test_env_file_resolution_prefers_cli(mod: ModuleType, tmp_path: Path) -> None: + a = tmp_path / "a.env" + paths = mod._resolve_env_files([str(a)], {}) + assert paths == [a] + + +def test_env_file_resolution_uses_env_var(mod: ModuleType, tmp_path: Path) -> None: + a = tmp_path / "a.env" + b = tmp_path / "b.env" + env = {"ASSERT_NO_WRITE_TOKEN_ENV_FILES": os.pathsep.join([str(a), str(b)])} + paths = mod._resolve_env_files(None, env) + assert paths == [a, b] + + +def test_env_file_resolution_defaults_and_expands_home(mod: ModuleType) -> None: + paths = mod._resolve_env_files(None, {}) + # Defaults to ~/secrev.env and ~/orchestrator/.env, expanded (no literal ~). + assert len(paths) == 2 + assert all("~" not in str(p) for p in paths) + assert paths[0].name == "secrev.env" + + +# --------------------------------------------------------------------------- # +# main() exit codes + messaging +# --------------------------------------------------------------------------- # + + +def test_main_clean_exits_zero( + mod: ModuleType, + tmp_path: Path, + monkeypatch: pytest.MonkeyPatch, + capsys: pytest.CaptureFixture, +) -> None: + clean = tmp_path / "secrev.env" + clean.write_text("CLAUDE_CODE_OAUTH_TOKEN=abc\n") + # Patch the live environ so the real process env can't leak into the scan. + monkeypatch.setattr(os, "environ", {"PATH": "/usr/bin"}) + rc = mod.main(["--env-file", str(clean)]) + assert rc == 0 + out = capsys.readouterr().out + assert "OK" in out + assert "no standing write token" in out + + +def test_main_finding_exits_nonzero( + mod: ModuleType, + tmp_path: Path, + monkeypatch: pytest.MonkeyPatch, + capsys: pytest.CaptureFixture, +) -> None: + leaky = tmp_path / "secrev.env" + leaky.write_text(_RSA_KEY + "\n") + monkeypatch.setattr(os, "environ", {"PATH": "/usr/bin"}) + rc = mod.main(["--env-file", str(leaky)]) + assert rc == 1 + err = capsys.readouterr().err + assert "FAIL" in err + assert "Actions secret" in err + + +def test_main_flags_write_token_in_live_environ( + mod: ModuleType, tmp_path: Path, monkeypatch: pytest.MonkeyPatch +) -> None: + clean = tmp_path / "secrev.env" + clean.write_text("ok\n") + monkeypatch.setattr(os, "environ", {"AGENT_APPLY_APP_PRIVATE_KEY": _RSA_KEY}) + rc = mod.main(["--env-file", str(clean)]) + assert rc == 1 diff --git a/agent-team/tests/test_p3_async_resume.py b/agent-team/tests/test_p3_async_resume.py new file mode 100644 index 0000000..a25e8ba --- /dev/null +++ b/agent-team/tests/test_p3_async_resume.py @@ -0,0 +1,245 @@ +"""End-to-end async resume-on-CI-complete (design §4 Decision 2, P3 BLOCK-1/BLOCK-2). + +These tests prove the "suspended forever" defect is gone: a task that dispatches +and suspends at VERIFY awaiting CI is, on a later ``tick()``, RESUMED once its run +reaches a terminal conclusion and TIMEOUT-PARKED once ``dispatched_at + timeout`` +elapses with no terminal result. They drive the REAL machinery — the durable +``ci_pending_provider`` (:meth:`Coordinator._enumerate_ci_pending`, which walks the +LangGraph checkpointer), the real CI-watcher sweep, and the real turn-guarded +resume worker (:meth:`ResumeWorker.resume_ci`) — over a real compiled LangGraph app +whose VERIFY node suspends with the production awaiting-CI interrupt payload. + +No network: the CI poll seam is injected. The graph is a faithful minimal stand-in +for the production BUILD → DISPATCH → VERIFY shape — DISPATCH writes the trusted +``run_id`` / ``dispatched_at`` watermarks; VERIFY ``interrupt()``s with +``{"awaiting_ci": True, "run_id": ...}`` exactly as +:func:`agent_team.nodes.build_verify_subgraph._await_ci` does — so the durable +enumeration, the suspend, and the resume are the real ones. +""" + +from __future__ import annotations + +from datetime import datetime, timedelta, timezone +from pathlib import Path +from typing import Any, TypedDict + +import pytest + +pytest.importorskip("langgraph") + +from langgraph.checkpoint.memory import InMemorySaver # noqa: E402 +from langgraph.graph import END, START, StateGraph # noqa: E402 +from langgraph.types import interrupt # noqa: E402 + +from agent_team.ci_watcher import CiPollResult, CiWatchAction # noqa: E402 +from agent_team.coordinator import Coordinator # noqa: E402 +from agent_team.db.schema import init_db # noqa: E402 +from agent_team.resume_worker import ResumeWorker # noqa: E402 +from agent_team.transport.base import Transport # noqa: E402 + + +# --------------------------------------------------------------------------- # +# Test doubles +# --------------------------------------------------------------------------- # + + +class _Transport(Transport): + """Record-only transport (no network).""" + + def post_question(self, **_kwargs: Any) -> str: # type: ignore[override] + return "fake:q" + + def parse_answer(self, raw: Any) -> tuple[str, Any, str]: + return raw["question_id"], raw["answer"], "fake" + + +class _S(TypedDict, total=False): + run_id: str + dispatched_at: str + ci_results: Any + status: str + current_phase: str + + +def _build_dispatch_verify_app(saver: InMemorySaver, *, dispatched_at: str) -> Any: + """Compile a real DISPATCH → VERIFY graph that suspends awaiting CI. + + DISPATCH writes the trusted ``run_id`` / ``dispatched_at`` watermarks (as the + production dispatch node does). VERIFY suspends via ``interrupt()`` with the + production awaiting-CI payload while there is a run but no terminal + ``ci_results``; on resume it falls through (the CI-watcher re-drives it). + """ + + def dispatch(state: _S) -> dict[str, Any]: + return {"run_id": "R-1", "dispatched_at": dispatched_at} + + def verify(state: _S) -> dict[str, Any]: + run_id = state.get("run_id") + ci = state.get("ci_results") + if run_id and ci is None: + # Production payload shape (build_verify_subgraph._await_ci). On resume + # the production node RE-FETCHES the authenticated conclusion; this + # stand-in uses the resume value the CI-watcher passes (the terminal + # poll result) as that re-fetched conclusion so VERIFY can advance. + ci = interrupt({"awaiting_ci": True, "run_id": run_id}) + return {"status": "done", "current_phase": "done", "ci_results": ci} + + g: StateGraph = StateGraph(_S) + g.add_node("dispatch", dispatch) + g.add_node("verify", verify) + g.add_edge(START, "dispatch") + g.add_edge("dispatch", "verify") + g.add_edge("verify", END) + return g.compile(checkpointer=saver) + + +def _coordinator_over( + db_path: Path, + app: Any, + saver: InMemorySaver, + *, + poller: Any, + timeout: timedelta | None = None, + alarm_hook: Any = None, +) -> Coordinator: + """A Coordinator whose graph/resume-worker are the supplied real app. + + ``_enumerate_ci_pending`` (the durable provider) is bound as the + ``ci_pending_provider`` so the watcher is fed REAL suspended threads off the + real checkpointer — exactly the live serve wiring. + """ + coord = Coordinator( + db_path=db_path, + transport=_Transport(), + build_checkpointer=lambda _p: saver, + ci_poller=poller, + ci_timeout=timeout, + alarm_hook=alarm_hook, + ) + coord._graph = app + coord._resume_worker = ResumeWorker(app, _connect(db_path)) + coord._ci_pending_provider = coord._enumerate_ci_pending + return coord + + +def _connect(db_path: Path) -> Any: + from agent_team.db.schema import connect + + return connect(db_path) + + +def _thread_cfg(thread_id: str) -> dict[str, Any]: + return {"configurable": {"thread_id": thread_id}} + + +@pytest.fixture() +def db_path(tmp_path: Path) -> Path: + path = tmp_path / "state" / "agent_team.sqlite" + init_db(path) + return path + + +# --------------------------------------------------------------------------- # +# (a) provider enumeration: includes suspended-at-VERIFY, excludes advanced +# --------------------------------------------------------------------------- # + + +def test_provider_enumerates_suspended_and_excludes_advanced(db_path: Path) -> None: + """``_enumerate_ci_pending`` returns the thread suspended at VERIFY awaiting CI + and EXCLUDES a resumed/done thread and a never-dispatched (no run) thread.""" + saver = InMemorySaver() + app = _build_dispatch_verify_app(saver, dispatched_at=_now_iso()) + coord = _coordinator_over( + db_path, app, saver, poller=lambda t: CiPollResult.pending() + ) + + # t-wait: dispatch + suspend at VERIFY awaiting CI (still pending). + app.invoke({}, _thread_cfg("t-wait")) + # t-done: dispatch, suspend, then RESUME with a terminal CI result -> advances + # past VERIFY to DONE (no awaiting-CI interrupt left). + from langgraph.types import Command + + app.invoke({}, _thread_cfg("t-done")) + app.invoke(Command(resume={"conclusion": "success"}), _thread_cfg("t-done")) + + pending = coord._enumerate_ci_pending() + thread_ids = {p.thread_id for p in pending} + + assert "t-wait" in thread_ids # suspended at VERIFY awaiting CI -> included + assert "t-done" not in thread_ids # advanced past the gate -> excluded + # The included task carries the trusted dispatch watermarks the watcher keys + # off (so the poll + timeout have a run to act on). + waiting = next(p for p in pending if p.thread_id == "t-wait") + assert waiting.run_id == "R-1" + assert waiting.dispatched_at is not None + + +# --------------------------------------------------------------------------- # +# (b) end-to-end: tick() RESUMES on terminal, TIMEOUT-PARKS on no-terminal +# --------------------------------------------------------------------------- # + + +def test_tick_resumes_suspended_task_on_terminal_conclusion(db_path: Path) -> None: + """A task suspended at VERIFY is RESUMED by a later tick() once its run is + terminal — proving 'suspended forever' is gone.""" + saver = InMemorySaver() + app = _build_dispatch_verify_app(saver, dispatched_at=_now_iso()) + coord = _coordinator_over( + db_path, + app, + saver, + poller=lambda t: CiPollResult.terminal({"run_id": "R-1", "conclusion": "ok"}), + ) + + # Dispatch + suspend at VERIFY. + app.invoke({}, _thread_cfg("t1")) + snap = app.get_state(_thread_cfg("t1")) + assert snap.next == ("verify",) # genuinely suspended awaiting CI + + # A subsequent tick() runs the CI-watch sweep -> resume the suspended task. + report = coord._ci_watch() + assert report is not None + assert report.resumed == 1 + assert report.outcomes[0].action is CiWatchAction.RESUMED + + # The thread advanced past VERIFY (no longer suspended). + snap2 = app.get_state(_thread_cfg("t1")) + assert snap2.next == () + assert snap2.values.get("status") == "done" + + +def test_tick_timeout_parks_suspended_task_when_run_never_terminates( + db_path: Path, +) -> None: + """A task whose run never reaches a terminal conclusion within the timeout is + TIMEOUT-PARKED by a later tick() (never waits forever).""" + saver = InMemorySaver() + # Dispatched long ago: dispatched_at + timeout has already elapsed. + old = (datetime.now(timezone.utc) - timedelta(hours=2)).isoformat() + app = _build_dispatch_verify_app(saver, dispatched_at=old) + + alarms: list[str] = [] + coord = _coordinator_over( + db_path, + app, + saver, + poller=lambda t: CiPollResult.pending(), # never terminal + timeout=timedelta(minutes=30), + alarm_hook=alarms.append, + ) + + app.invoke({}, _thread_cfg("t-slow")) + assert app.get_state(_thread_cfg("t-slow")).next == ("verify",) + + report = coord._ci_watch() + assert report is not None + assert report.parked_timeout == 1 + assert report.outcomes[0].action is CiWatchAction.PARKED_TIMEOUT + # Durably parked + ALARMed (surfaced, not silently spun on). + parked = app.get_state(_thread_cfg("t-slow")) + assert parked.values.get("status") == "parked" + assert alarms == ["t-slow"] + + +def _now_iso() -> str: + return datetime.now(timezone.utc).isoformat() diff --git a/agent-team/tests/test_p3_failsafe_serve_default.py b/agent-team/tests/test_p3_failsafe_serve_default.py new file mode 100644 index 0000000..8664f8f --- /dev/null +++ b/agent-team/tests/test_p3_failsafe_serve_default.py @@ -0,0 +1,252 @@ +"""P3 fail-safe serve default (design Decision 5; UNIT 0e). + +The bound P3 build→verify + dispatch wiring is the new production ``serve`` +default, but its factories are called EAGERLY at graph-build and the live +dispatch factory RAISES when ``AGENT_TEAM_REPO_OWNER`` / ``AGENT_TEAM_REPO_NAME`` +are unset. These tests pin the fail-safe contract: + +* :func:`agent_team.coordinator._p3_env_is_configured` truth table. +* :func:`agent_team.coordinator.failsafe_production_p3_wiring` — + env-unset degrades to the INERT ``(None, None)`` pair with exactly ONE WARNING + and ONE ``#agent-team`` inert notice via the lifecycle ``notify`` sink (never + the ALARM path), and NEVER raises; env-set returns the live wiring pair. +* a Coordinator built with the env-unset (inert) result sets up cleanly and an + approved task settles at the P2 BUILD terminus (no P3 nodes, no dispatch, no + exception) — i.e. a task that would reach P3 parks short of build/dispatch. +* ``run-team.py serve`` binds the fail-safe pair; ``start`` / ``intake`` do not. +""" + +from __future__ import annotations + +import logging +from pathlib import Path +from typing import Any + +import pytest + +from agent_team import graph as graph_mod +from agent_team.coordinator import ( + Coordinator, + _p3_env_is_configured, + default_dispatch_node_factory, + failsafe_production_p3_wiring, +) + +try: # InMemorySaver is the modern name; fall back on older langgraph. + from langgraph.checkpoint.memory import InMemorySaver as _Saver +except ImportError: # pragma: no cover - older langgraph + from langgraph.checkpoint.memory import MemorySaver as _Saver + + +_P3_ENV = ("AGENT_TEAM_REPO_OWNER", "AGENT_TEAM_REPO_NAME") +_TOKEN_ENV = ("AGENT_TEAM_CI_READ_TOKEN", "GITHUB_TOKEN") + + +@pytest.fixture +def clean_p3_env(monkeypatch: pytest.MonkeyPatch) -> None: + """Start each test from a fully unset P3 environment.""" + for name in (*_P3_ENV, *_TOKEN_ENV): + monkeypatch.delenv(name, raising=False) + + +# --------------------------------------------------------------------------- # +# _p3_env_is_configured truth table +# --------------------------------------------------------------------------- # + + +def test_env_unset_is_not_configured(clean_p3_env: None) -> None: + assert _p3_env_is_configured() is False + + +def test_owner_repo_without_token_is_not_configured( + clean_p3_env: None, monkeypatch: pytest.MonkeyPatch +) -> None: + monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "org") + monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "repo") + # No CI-read token -> the verifier could never read an authenticated pass. + assert _p3_env_is_configured() is False + + +def test_token_without_owner_repo_is_not_configured( + clean_p3_env: None, monkeypatch: pytest.MonkeyPatch +) -> None: + monkeypatch.setenv("AGENT_TEAM_CI_READ_TOKEN", "ghp_test") + assert _p3_env_is_configured() is False + + +def test_owner_repo_and_ci_token_is_configured( + clean_p3_env: None, monkeypatch: pytest.MonkeyPatch +) -> None: + monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "org") + monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "repo") + monkeypatch.setenv("AGENT_TEAM_CI_READ_TOKEN", "ghp_test") + assert _p3_env_is_configured() is True + + +def test_github_token_fallback_satisfies_ci_token( + clean_p3_env: None, monkeypatch: pytest.MonkeyPatch +) -> None: + monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "org") + monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "repo") + monkeypatch.setenv("GITHUB_TOKEN", "ghp_fallback") + assert _p3_env_is_configured() is True + + +def test_blank_env_values_are_not_configured( + clean_p3_env: None, monkeypatch: pytest.MonkeyPatch +) -> None: + # Whitespace-only values must not count as configured (fail-closed). + monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", " ") + monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "repo") + monkeypatch.setenv("AGENT_TEAM_CI_READ_TOKEN", "ghp_test") + assert _p3_env_is_configured() is False + + +# --------------------------------------------------------------------------- # +# failsafe_production_p3_wiring — env-unset degrade +# --------------------------------------------------------------------------- # + + +def test_env_unset_returns_inert_pair_warns_and_notifies( + clean_p3_env: None, caplog: pytest.LogCaptureFixture +) -> None: + notices: list[str] = [] + + with caplog.at_level(logging.WARNING, logger="agent_team.coordinator"): + build_verify, dispatch = failsafe_production_p3_wiring( + notify=lambda msg, **_: notices.append(msg) + ) + + # Inert P3: no build->verify subgraph, no dispatch. + assert build_verify is None + assert dispatch is None + + # Exactly one WARNING about the inert P3 wiring. + inert_warnings = [ + r + for r in caplog.records + if r.levelno == logging.WARNING and "INERT" in r.getMessage() + ] + assert len(inert_warnings) == 1 + + # Exactly one #agent-team inert notice via the (non-ALARM) notify sink. + assert len(notices) == 1 + assert "INERT" in notices[0] + + +def test_env_unset_without_notify_does_not_raise(clean_p3_env: None) -> None: + # A token-less / channel-less serve still comes up inert with no notify sink. + build_verify, dispatch = failsafe_production_p3_wiring(notify=None) + assert build_verify is None + assert dispatch is None + + +def test_inert_notify_failure_is_swallowed(clean_p3_env: None) -> None: + def _boom(_msg: str, **_kw: Any) -> None: + raise RuntimeError("slack down") + + # The notify sink raising must not propagate out of serve-start. + build_verify, dispatch = failsafe_production_p3_wiring(notify=_boom) + assert build_verify is None + assert dispatch is None + + +# --------------------------------------------------------------------------- # +# failsafe_production_p3_wiring — env-set binds live wiring +# --------------------------------------------------------------------------- # + + +def test_env_set_returns_live_wiring_pair( + clean_p3_env: None, monkeypatch: pytest.MonkeyPatch +) -> None: + monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "test-org") + monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "test-repo") + monkeypatch.setenv("AGENT_TEAM_CI_READ_TOKEN", "ghp_test") + + notices: list[str] = [] + build_verify, dispatch = failsafe_production_p3_wiring( + notify=lambda msg, **_: notices.append(msg) + ) + + # Live pair: both factories present, no inert notice emitted. + assert callable(build_verify) + assert callable(dispatch) + assert notices == [] + + # The dispatch factory is the env-reading production factory. + assert dispatch is default_dispatch_node_factory + + # Each live factory builds without raising and yields the expected shapes. + bv_tuple = build_verify() + assert isinstance(bv_tuple, tuple) and len(bv_tuple) == 3 + assert all(callable(part) for part in bv_tuple) + assert callable(dispatch()) + + +# --------------------------------------------------------------------------- # +# Coordinator built with the inert result sets up + a P3-bound task does not +# dispatch / crash (it settles at the P2 BUILD terminus). +# --------------------------------------------------------------------------- # + + +def test_coordinator_with_inert_wiring_builds_and_task_parks_short_of_p3( + clean_p3_env: None, tmp_path: Path +) -> None: + """env-unset -> graph builds, coordinator constructs, an approved task settles + at the P2 BUILD terminus with no P3 nodes, no dispatch, and no exception.""" + from agent_team.graph import resume_task, start_task + from agent_team.nodes import review_loop + from agent_team.task_model import Phase, PipelineState, TaskStatus + + build_verify, dispatch = failsafe_production_p3_wiring(notify=None) + assert (build_verify, dispatch) == (None, None) + + saver = _Saver() + saved_invoker = review_loop._review_invoker + try: + review_loop.set_review_invoker(lambda prompt, **kw: "VERDICT: APPROVE\nok") + + def _plan_stub(state: PipelineState) -> PipelineState: + return PipelineState( + plan={"title": "t", "scope": ["src"], "phases": ["P1"]}, + current_phase=Phase.REVIEW.value, + status=TaskStatus.ACTIVE.value, + ) + + coord = Coordinator( + db_path=tmp_path / "inert.db", + transport=_FakeTransport(), + build_clarify_node=lambda: graph_mod.clarify_node, + build_plan_node=lambda: _plan_stub, + review_wiring=lambda: ( + review_loop.bind_review_node(), + review_loop.route_after_review, + ), + build_verify_wiring=build_verify, + dispatch_node_wiring=dispatch, + build_checkpointer=lambda _path: saver, + ) + coord.setup() + + # No P3 nodes were wired (inert): the graph stops at the P2 terminus. + nodes = coord.graph.get_graph().nodes + assert graph_mod.BUILD_NODE not in nodes + assert graph_mod.VERIFY_NODE not in nodes + + # An approved task runs to the P2 BUILD terminus without dispatching or + # raising (it never reaches a live build/dispatch). + thread_id, _ = start_task(coord.graph, transport="slack") + final = resume_task(coord.graph, thread_id=thread_id, answer="scope is X") + assert final["current_phase"] == Phase.BUILD.value + finally: + review_loop._review_invoker = saved_invoker + + +class _FakeTransport: + """Minimal non-posting transport for the inert-wiring coordinator test.""" + + def post_question(self, **_kwargs: Any) -> str: + return "ref" + + def parse_answer(self, raw: Any) -> tuple[str, Any, str]: # pragma: no cover + raise NotImplementedError diff --git a/agent-team/tests/test_resume_worker.py b/agent-team/tests/test_resume_worker.py index 4cdb7c4..212ee09 100644 --- a/agent-team/tests/test_resume_worker.py +++ b/agent-team/tests/test_resume_worker.py @@ -428,6 +428,91 @@ def test_decode_answer_non_json_passthrough() -> None: assert resume_worker._decode_answer("not-json{{") == "not-json{{" +# --------------------------------------------------------------------------- # +# resume_ci (CI machine-gate: turn-guarded on the awaiting-CI marker) +# --------------------------------------------------------------------------- # + + +class _CiGraph: + """A GraphLike double suspended at VERIFY awaiting a given CI ``run_id``. + + ``awaiting_run`` is the run the thread is currently suspended on, or ``None`` + if it has advanced past the CI gate (resumed / parked / done). ``invoke`` + records every resume and advances the graph (clears the interrupt) as a real + resume would, so a double-resume is directly observable. + """ + + def __init__(self, awaiting_run: str | None) -> None: + self.awaiting_run = awaiting_run + self.invocations: list[Any] = [] + + def get_state(self, config: dict[str, Any]) -> _FakeSnapshot: + if self.awaiting_run is None: + return _FakeSnapshot(next=(), interrupts=()) + payload = {"awaiting_ci": True, "run_id": self.awaiting_run} + return _FakeSnapshot( + next=("verify",), + interrupts=(_FakeInterrupt(value=payload),), + ) + + def invoke(self, command: Any, config: dict[str, Any]) -> Any: + self.invocations.append(command) + self.awaiting_run = None + return {"resumed": True} + + +def test_resume_ci_applies_when_suspended_on_run(conn: sqlite3.Connection) -> None: + graph = _CiGraph(awaiting_run="999") + worker = ResumeWorker(graph, conn) + + result = worker.resume_ci(thread_id="t1", run_id="999", answer={"conclusion": "ok"}) + + assert result.outcome is ResumeOutcome.RESUMED + assert result.resumed is True + assert len(graph.invocations) == 1 + + +def test_resume_ci_skips_when_thread_already_advanced( + conn: sqlite3.Connection, +) -> None: + # Already resumed/parked/done: no awaiting-CI interrupt -> guard skips invoke. + graph = _CiGraph(awaiting_run=None) + worker = ResumeWorker(graph, conn) + + result = worker.resume_ci(thread_id="t1", run_id="999", answer={}) + + assert result.outcome is ResumeOutcome.STALE + assert graph.invocations == [] + + +def test_resume_ci_skips_when_awaiting_a_different_run( + conn: sqlite3.Connection, +) -> None: + # Suspended awaiting a DIFFERENT run (e.g. a re-dispatch): must not resume. + graph = _CiGraph(awaiting_run="other") + worker = ResumeWorker(graph, conn) + + result = worker.resume_ci(thread_id="t1", run_id="999", answer={}) + + assert result.outcome is ResumeOutcome.STALE + assert graph.invocations == [] + + +def test_resume_ci_double_resume_is_idempotent(conn: sqlite3.Connection) -> None: + # Two terminal observations of the same run on overlapping sweeps: the first + # applies; the second finds the thread advanced (guard) and skips. State is + # never double-applied. + graph = _CiGraph(awaiting_run="999") + worker = ResumeWorker(graph, conn) + + first = worker.resume_ci(thread_id="t1", run_id="999", answer={}) + second = worker.resume_ci(thread_id="t1", run_id="999", answer={}) + + assert first.outcome is ResumeOutcome.RESUMED + assert second.outcome is ResumeOutcome.STALE + assert len(graph.invocations) == 1 + + # --------------------------------------------------------------------------- # # command builder # --------------------------------------------------------------------------- # diff --git a/agent-team/tests/test_rollback.py b/agent-team/tests/test_rollback.py new file mode 100644 index 0000000..0ffe07e --- /dev/null +++ b/agent-team/tests/test_rollback.py @@ -0,0 +1,940 @@ +"""Tests for ``scripts/p3_rollback.sh`` — the P3 privileged-surface rollback. + +The script restores EVERY privileged P3 surface (the apply/verify workflow flip, +the ``agent-apply`` environment, the GitHub App perms/installation, and branch +protection) from a recorded baseline, and asserts post-restore == baseline. The +apply/verify workflow is ALREADY LIVE (flipped + provisioned 2026-06-22), so the +rollback targets the LIVE state. + +These tests exercise the script with NO real ``gh``/``git`` calls: + +* The ``--dry-run`` default must print a PLAN and perform NO mutations. We assert + the plan output covers every surface (every destructive call is described but + not executed). +* Argument parsing: a missing surface, an unknown flag, and a missing value each + fail closed (exit 2). +* Fail-closed posture: a baseline whose ``include_administrators`` is not ``true`` + is refused; a missing baseline file is refused. +* ``--apply`` is verified against a PATH-shimmed ``gh``/``git`` that only RECORDS + its argv into a log file (never touches a network or a repo), so we can assert + the exact destructive calls the script would make — with zero real side effects. +""" + +from __future__ import annotations + +import json +import os +import stat +import subprocess +from pathlib import Path + +import pytest + +_SCRIPT = Path(__file__).resolve().parents[1] / "scripts" / "p3_rollback.sh" + + +def test_script_exists_and_is_executable() -> None: + assert _SCRIPT.is_file(), f"missing rollback script: {_SCRIPT}" + mode = _SCRIPT.stat().st_mode + assert mode & stat.S_IXUSR, "p3_rollback.sh must be executable" + + +def test_script_has_bash_shebang() -> None: + first = _SCRIPT.read_text(encoding="utf-8").splitlines()[0] + assert first.startswith("#!") and "bash" in first + + +def test_script_passes_bash_syntax_check() -> None: + res = subprocess.run(["bash", "-n", str(_SCRIPT)], capture_output=True, text=True) + assert res.returncode == 0, res.stderr + + +# --------------------------------------------------------------------------- # +# Fixtures: a recorded baseline + a PATH shim for gh/git +# --------------------------------------------------------------------------- # + +_BASELINE = { + "repo": "Sea-Haven-Industries/orchestrator", + "default_branch": "main", + "workflow_path": ".github/workflows/agent-team-apply-verify.yml", + "workflow_baseline_sha": "0123abc", + "environment": { + "name": "agent-apply", + # Reviewers recorded as NUMERIC user ids (restore is exact + assertable). + "required_reviewer_ids": [1234567], + "required_reviewers": ["amoussa1229"], + "deployment_branch_policy": "protected", + }, + "app": { + "slug": "agent-apply", + "installation_id": 424242, + # The only programmatic neutralise is uninstall (App JWT). There is no + # permission-reduction REST endpoint. + "action": "uninstall", + }, + "protection": { + "branch": "main", + "include_administrators": True, + # The FULL protection payload recorded pre-flip (the exact body restored). + "full": { + "enforce_admins": {"enabled": True}, + "required_status_checks": { + "strict": True, + "contexts": ["guard", "build-test"], + }, + "required_pull_request_reviews": { + "required_approving_review_count": 1, + "dismiss_stale_reviews": True, + "require_code_owner_reviews": False, + }, + "required_linear_history": {"enabled": True}, + "allow_force_pushes": {"enabled": False}, + "allow_deletions": {"enabled": False}, + }, + "required_status_checks": ["guard", "build-test"], + }, +} + + +@pytest.fixture +def baseline(tmp_path: Path) -> Path: + p = tmp_path / "p3-baseline.json" + p.write_text(json.dumps(_BASELINE), encoding="utf-8") + return p + + +@pytest.fixture +def shim_bin(tmp_path: Path) -> tuple[Path, Path]: + """A bin dir with stub ``gh`` and ``git`` that only record their argv. + + Returns ``(bin_dir, calls_log)``. The script, run with this dir prepended to + PATH, makes ZERO real gh/git calls — every invocation appends a line to + ``calls_log`` and exits 0. Where the script reads command output (the + post-restore asserts), the stubs emit the baseline value so the assert holds. + """ + bin_dir = tmp_path / "bin" + bin_dir.mkdir() + calls_log = tmp_path / "calls.log" + + # gh stub: record argv; emit canned output for the read-only post-restore + # asserts so --apply asserts pass. The script passes a server-side --jq to gh + # (the real gh applies it); the stub must therefore emit the ALREADY-jq'd + # value the script expects: + # * env GET with the reviewer-ids --jq -> "1234567" (space-joined ids) + # * enforce_admins GET -> "true" + # * protection GET (no enforce_admins) -> the full protection JSON, which + # the script then pipes through its own normalize_protection. We emit the + # same shape the baseline records so the normalized compare holds. + gh = bin_dir / "gh" + _protection_json = json.dumps(_BASELINE["protection"]["full"]) + gh.write_text( + "#!/usr/bin/env bash\n" + f'printf "gh %s\\n" "$*" >> "{calls_log}"\n' + 'argv="$*"\n' + 'for a in "$@"; do\n' + ' case "$a" in\n' + " */enforce_admins) echo 'true'; exit 0 ;;\n" + " esac\n" + "done\n" + "# protection GET (full object) -> emit the baseline full protection JSON.\n" + 'case "$argv" in\n' + " *branches/*/protection*)\n" + f" cat <<'JSON'\n{_protection_json}\nJSON\n" + " exit 0 ;;\n" + " *users/*)\n" + " # login->id resolution (gh api users/{login} --jq .id).\n" + " echo '1234567'; exit 0 ;;\n" + " *environments/*)\n" + " # reviewer-ids --jq result (space-joined) for the post-restore assert.\n" + " echo '1234567'; exit 0 ;;\n" + "esac\n" + "exit 0\n", + encoding="utf-8", + ) + git = bin_dir / "git" + git.write_text( + f'#!/usr/bin/env bash\nprintf "git %s\\n" "$*" >> "{calls_log}"\nexit 0\n', + encoding="utf-8", + ) + for f in (gh, git): + f.chmod(f.stat().st_mode | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH) + return bin_dir, calls_log + + +def _run( + *args: str, + baseline: Path, + shim: tuple[Path, Path] | None = None, + extra_env: dict[str, str] | None = None, +) -> subprocess.CompletedProcess[str]: + env = dict(os.environ) + if shim is not None: + bin_dir, _ = shim + env["PATH"] = f"{bin_dir}{os.pathsep}{env['PATH']}" + if extra_env: + env.update(extra_env) + return subprocess.run( + ["bash", str(_SCRIPT), *args, "--baseline", str(baseline)], + capture_output=True, + text=True, + env=env, + # cwd kept default (no git repo needed: gh/git are shimmed). + ) + + +# --------------------------------------------------------------------------- # +# Dry-run plan output (the default; no shim required — nothing executes) +# --------------------------------------------------------------------------- # + + +def test_default_is_dry_run_and_prints_plan_not_apply(baseline: Path) -> None: + res = _run("all", "--flip-pr", "7", baseline=baseline) + assert res.returncode == 0, res.stderr + out = res.stdout + assert "--dry-run (plan only" in out + assert "[PLAN]" in out + # No APPLY lines must appear in dry-run. + assert "[APPLY]" not in out + + +def test_dry_run_covers_all_four_surfaces(baseline: Path) -> None: + res = _run("all", "--flip-pr", "7", baseline=baseline) + out = res.stdout + assert "Restore the apply/verify workflow flip" in out + assert "Restore the agent-apply environment" in out + assert "Neutralise the GitHub App" in out + assert "Restore the FULL branch protection baseline" in out + + +def test_dry_run_workflow_premerge_closes_pr_and_deletes_branch( + baseline: Path, +) -> None: + res = _run( + "workflow", + "--flip-pr", + "7", + "--flip-branch", + "agent-team/apply/t1", + baseline=baseline, + ) + out = res.stdout + assert "gh pr close 7" in out + assert "--delete-branch" in out + assert "git/refs/heads/agent-team/apply/t1" in out + # LIVE: the run-name/permissions edits get reverted to the baseline SHA. + assert "git checkout 0123abc --" in out + + +def test_dry_run_workflow_postmerge_reverts_commit_and_reruns_ci( + baseline: Path, +) -> None: + res = _run( + "workflow", + "--merged", + "--flip-commit", + "cafef00d", + baseline=baseline, + ) + out = res.stdout + assert "git revert --no-edit cafef00d" in out + assert "git push origin HEAD" in out + assert "gh workflow run" in out + + +def test_dry_run_app_uninstall_requires_app_jwt_not_operator_gh( + baseline: Path, +) -> None: + """The corrected model: there is NO permission-reduction endpoint, and the + installation token is NOT revoked by operator gh. The only programmatic + neutralise is uninstall, which needs an App JWT.""" + res = _run("app", baseline=baseline) + out = res.stdout + # The fictional permission-reduction endpoint must NOT appear. + assert "/permissions" not in out + assert "PATCH" not in out + # The token is NOT revoked by operator gh. + assert "DELETE installation/token" not in out + assert "cannot be revoked by operator gh" in out + # Uninstall is planned and clearly flagged as needing an App JWT. + assert "UNINSTALL App 'agent-apply' installation 424242" in out + assert "APP JWT" in out + assert "app/installations/424242" in out + + +def test_dry_run_app_out_of_band_path(tmp_path: Path) -> None: + data = json.loads(json.dumps(_BASELINE)) + data["app"]["action"] = "out-of-band" + p = tmp_path / "b.json" + p.write_text(json.dumps(data), encoding="utf-8") + res = _run("app", baseline=p) + out = res.stdout + assert "OUT-OF-BAND App neutralise" in out + assert "no REST endpoint reduces App permissions" in out + # Still no fictional permission API. + assert "/permissions" not in out + + +def test_apply_app_uninstall_without_jwt_fails_closed( + baseline: Path, shim_bin: tuple[Path, Path] +) -> None: + """--apply uninstall with no AGENT_APPLY_APP_JWT must refuse — operator gh + cannot perform DELETE /app/installations/{id} (it needs an App JWT).""" + env = dict(os.environ) + env.pop("AGENT_APPLY_APP_JWT", None) + bin_dir, _ = shim_bin + env["PATH"] = f"{bin_dir}{os.pathsep}{env['PATH']}" + res = subprocess.run( + ["bash", str(_SCRIPT), "app", "--apply", "--baseline", str(baseline)], + capture_output=True, + text=True, + env=env, + ) + assert res.returncode != 0 + assert "needs an APP JWT" in res.stderr + + +def test_apply_app_uninstall_with_jwt_invokes_delete( + baseline: Path, tmp_path: Path +) -> None: + """With an App JWT present, --apply uninstall issues the DELETE via curl with + a 0600 header FILE (SEC-02 / CWE-214) — the JWT is NEVER on the process argv. + + The shimmed ``curl`` records its argv AND dumps the contents of the + ``-H @`` header file, so we can assert: + * the DELETE hits app/installations/424242, + * the JWT lives only inside the header file (not in argv), + * the header file referenced on argv carries the Bearer line. + """ + bin_dir = tmp_path / "bin" + bin_dir.mkdir() + calls_log = tmp_path / "calls.log" + + gh = bin_dir / "gh" + gh.write_text( + f'#!/usr/bin/env bash\nprintf "gh %s\\n" "$*" >> "{calls_log}"\nexit 0\n', + encoding="utf-8", + ) + # curl shim: record argv, and resolve any `-H @file` to dump the file body so + # the test can confirm the secret was passed by FILE, not on the command line. + curl = bin_dir / "curl" + curl.write_text( + "#!/usr/bin/env bash\n" + f'printf "curl %s\\n" "$*" >> "{calls_log}"\n' + "prev=''\n" + 'for a in "$@"; do\n' + ' if [ "$prev" = "-H" ]; then\n' + ' case "$a" in\n' + f' @*) printf "HDRFILE %s\\n" "$(cat "${{a#@}}")" >> "{calls_log}" ;;\n' + " esac\n" + " fi\n" + ' prev="$a"\n' + "done\n" + "exit 0\n", + encoding="utf-8", + ) + for f in (gh, curl): + f.chmod(f.stat().st_mode | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH) + + env = dict(os.environ) + env["AGENT_APPLY_APP_JWT"] = "jwt-token-abc" + env["PATH"] = f"{bin_dir}{os.pathsep}{env['PATH']}" + res = subprocess.run( + ["bash", str(_SCRIPT), "app", "--apply", "--baseline", str(baseline)], + capture_output=True, + text=True, + env=env, + ) + assert res.returncode == 0, res.stderr + res.stdout + calls = calls_log.read_text(encoding="utf-8") + # The DELETE went out via curl to the App installations endpoint. + assert "app/installations/424242" in calls + assert "-X DELETE" in calls + # SEC-02: the JWT is NEVER on the process argv (curl line), only in the file. + assert "jwt-token-abc" not in "".join( + line for line in calls.splitlines() if line.startswith("curl ") + ) + # The header file carried the Bearer line. + assert "HDRFILE Authorization: Bearer jwt-token-abc" in calls + # And the JWT never reached stdout (it is masked everywhere it is echoed). + assert "jwt-token-abc" not in res.stdout + + +def test_dry_run_app_redacts_jwt_from_plan_output(baseline: Path) -> None: + """SEC-01 (CWE-532): in the DEFAULT --dry-run mode the App-uninstall plan + echoes the gh/curl argv (which includes the Authorization: Bearer + header). The real App JWT must NEVER appear on stdout — it is masked to a + placeholder. Regression for the secret-in-CI-log leak.""" + secret = "supersecretjwtvalue1234567890" + res = _run("app", baseline=baseline, extra_env={"AGENT_APPLY_APP_JWT": secret}) + assert res.returncode == 0, res.stderr + blob = res.stdout + res.stderr + # The secret value never leaks; the masked placeholder is what is printed. + assert secret not in blob + assert "" in blob + + +def test_dry_run_app_redacts_jwt_in_all_surface(baseline: Path) -> None: + """SEC-01: the redaction is central (run_or_plan), so it holds on the `all` + and `incident` paths too — any surface that echoes the App-uninstall argv.""" + secret = "anotherjwtsecretZZZ999" + res = _run( + "all", + "--flip-pr", + "7", + baseline=baseline, + extra_env={"AGENT_APPLY_APP_JWT": secret}, + ) + assert res.returncode == 0, res.stderr + assert secret not in (res.stdout + res.stderr) + + +def test_workflow_live_revert_warns_staged_only(baseline: Path) -> None: + """P3-IAC-08 (CWE-665): the LIVE-YAML revert uses `git checkout -- path`, + which only STAGES locally (no commit/push). The script must WARN that the + revert is staged-only and needs a manual commit+push, so the restore is never + assumed complete on the remote.""" + res = _run( + "workflow", + "--flip-pr", + "7", + "--flip-branch", + "agent-team/apply/t1", + baseline=baseline, + ) + assert res.returncode == 0, res.stderr + blob = res.stdout + res.stderr + assert "git checkout 0123abc --" in res.stdout + assert "STAGED-ONLY" in blob + assert "commit + push" in blob or "commit+push" in blob + + +def test_dry_run_protection_restores_full_baseline(baseline: Path) -> None: + res = _run("protection", baseline=baseline) + out = res.stdout + # The FULL protection object is restored (not just enforce_admins). + assert "PUT full branch protection on main from baseline protection.full" in out + # And the post-restore asserts both enforce_admins and the full object. + assert "assert enforce_admins.enabled == true" in out + assert "assert LIVE protection == baseline protection.full" in out + + +def test_protection_without_full_refused_unless_partial_acked( + tmp_path: Path, +) -> None: + # MEDIUM-1: a baseline missing protection.full must NOT silently restore only + # enforce_admins (a weaker posture). Without the explicit ack it fails hard. + data = json.loads(json.dumps(_BASELINE)) + del data["protection"]["full"] + p = tmp_path / "nofull.json" + p.write_text(json.dumps(data), encoding="utf-8") + + refused = _run("protection", baseline=p) + assert refused.returncode != 0 + assert "no protection.full" in (refused.stdout + refused.stderr) + assert "P3_ROLLBACK_ALLOW_PARTIAL" in (refused.stdout + refused.stderr) + + # With the explicit ack the degraded enforce_admins-only restore proceeds, but + # LOUDLY warns it is partial. + acked = _run("protection", baseline=p, extra_env={"P3_ROLLBACK_ALLOW_PARTIAL": "1"}) + out = acked.stdout + acked.stderr + assert "DEGRADED protection restore" in out + assert "branches/main/protection/enforce_admins" in acked.stdout + + +def test_dry_run_incident_path_full_sequence(baseline: Path) -> None: + res = _run( + "incident", + "--flip-pr", + "13", + "--flip-branch", + "agent-team/apply/t9", + baseline=baseline, + ) + assert res.returncode == 0, res.stderr + out = res.stdout + # a. neutralise App b. revert draft PR/branch c. audit Checks d. restore e. note + assert "Neutralise the GitHub App" in out + # The corrected model: NOT an operator-gh token revoke. + assert "DELETE installation/token" not in out + assert "UNINSTALL App 'agent-apply' installation 424242" in out + assert "gh pr close 13" in out + assert "git/refs/heads/agent-team/apply/t9" in out + assert "Audit the Checks trail" in out + assert "gh run list" in out + assert "Restore the agent-apply environment" in out + assert "Restore the FULL branch protection baseline" in out + assert "incident note" in out + + +# --------------------------------------------------------------------------- # +# Argument parsing — fail closed +# --------------------------------------------------------------------------- # + + +def test_missing_surface_exits_2(baseline: Path) -> None: + res = _run(baseline=baseline) + assert res.returncode == 2 + assert "a surface is required" in res.stderr + + +def test_unknown_flag_exits_2(baseline: Path) -> None: + res = _run("workflow", "--bogus", baseline=baseline) + assert res.returncode == 2 + assert "unknown argument: --bogus" in res.stderr + + +def test_two_surfaces_is_rejected(baseline: Path) -> None: + res = _run("workflow", "protection", baseline=baseline) + assert res.returncode == 2 + assert "surface already set" in res.stderr + + +def test_flag_missing_value_exits_2(baseline: Path) -> None: + # --flip-pr with no following value (the trailing --baseline is consumed as + # the value, but then --baseline has no value -> still a parse error path). + res = subprocess.run( + ["bash", str(_SCRIPT), "workflow", "--flip-pr"], + capture_output=True, + text=True, + ) + assert res.returncode == 2 + + +def test_help_exits_0_and_lists_surfaces() -> None: + res = subprocess.run( + ["bash", str(_SCRIPT), "--help"], capture_output=True, text=True + ) + assert res.returncode == 0 + for surface in ("workflow", "environment", "app", "protection", "incident"): + assert surface in res.stdout + + +# --------------------------------------------------------------------------- # +# Fail-closed posture +# --------------------------------------------------------------------------- # + + +def test_missing_baseline_file_is_refused(tmp_path: Path) -> None: + missing = tmp_path / "nope.json" + res = _run("protection", baseline=missing) + assert res.returncode != 0 + assert "baseline file not found" in res.stderr + + +def test_protection_baseline_without_admins_on_is_refused(tmp_path: Path) -> None: + data = json.loads(json.dumps(_BASELINE)) + data["protection"]["include_administrators"] = False + p = tmp_path / "weak.json" + p.write_text(json.dumps(data), encoding="utf-8") + res = _run("protection", baseline=p) + assert res.returncode != 0 + assert "include_administrators is not true" in res.stderr + + +def test_repo_mismatch_is_refused(baseline: Path) -> None: + res = _run("protection", "--repo", "evil/other", baseline=baseline) + assert res.returncode != 0 + assert "repo mismatch" in res.stderr + + +def test_workflow_premerge_requires_flip_pr(baseline: Path) -> None: + res = _run("workflow", baseline=baseline) + assert res.returncode != 0 + assert "pre-merge path needs --flip-pr" in res.stderr + + +def test_workflow_postmerge_requires_flip_commit(baseline: Path) -> None: + res = _run("workflow", "--merged", baseline=baseline) + assert res.returncode != 0 + assert "post-merge path needs --flip-commit" in res.stderr + + +# --------------------------------------------------------------------------- # +# --apply against a PATH-shimmed gh/git (records argv, no real side effects) +# --------------------------------------------------------------------------- # + + +def test_apply_protection_invokes_gh_and_asserts( + baseline: Path, shim_bin: tuple[Path, Path] +) -> None: + _, calls_log = shim_bin + res = _run("protection", "--apply", baseline=baseline, shim=shim_bin) + assert res.returncode == 0, res.stderr + res.stdout + assert "[APPLY]" in res.stdout + # The post-restore assert ran and held (shim emits 'true'). + assert "[OK]" in res.stdout + calls = calls_log.read_text(encoding="utf-8") + # The FULL protection object was PUT to the (shimmed) gh, then asserted. + assert ( + "PUT repos/Sea-Haven-Industries/orchestrator/branches/main/protection" in calls + ) + assert "branches/main/protection/enforce_admins" in calls + + +def test_apply_environment_puts_reviewer_ids_and_asserts( + baseline: Path, shim_bin: tuple[Path, Path] +) -> None: + _, calls_log = shim_bin + res = _run("environment", "--apply", baseline=baseline, shim=shim_bin) + assert res.returncode == 0, res.stderr + res.stdout + calls = calls_log.read_text(encoding="utf-8") + assert "environments/agent-apply" in calls + # Reviewers are sent as proper typed JSON fields, by NUMERIC id. + assert "reviewers[][type]=User" in calls + assert "reviewers[][id]=1234567" in calls + # The post-restore assert compared live reviewer ids to the baseline and held. + assert "[OK]" in res.stdout + + +def test_apply_workflow_premerge_records_close_and_branch_delete( + baseline: Path, shim_bin: tuple[Path, Path] +) -> None: + _, calls_log = shim_bin + res = _run( + "workflow", + "--apply", + "--flip-pr", + "7", + "--flip-branch", + "agent-team/apply/t1", + baseline=baseline, + shim=shim_bin, + ) + assert res.returncode == 0, res.stderr + res.stdout + calls = calls_log.read_text(encoding="utf-8") + assert "pr close 7" in calls + assert "git/refs/heads/agent-team/apply/t1" in calls + assert "checkout 0123abc" in calls + + +def test_dry_run_environment_resolves_login_to_id_when_no_ids_recorded( + tmp_path: Path, +) -> None: + """A baseline that recorded only logins (no required_reviewer_ids) plans a + login->id resolution via 'gh api users/{login} --jq .id'.""" + data = json.loads(json.dumps(_BASELINE)) + del data["environment"]["required_reviewer_ids"] + p = tmp_path / "logins.json" + p.write_text(json.dumps(data), encoding="utf-8") + res = _run("environment", baseline=p) + out = res.stdout + assert "resolve reviewer login 'amoussa1229'" in out + assert "gh api users/amoussa1229 --jq .id" in out + + +def test_apply_environment_resolves_login_to_id(tmp_path: Path) -> None: + """--apply with a login-only baseline resolves the login to a numeric id via + the shimmed 'gh api users/{login}' and sends it as a typed reviewer field.""" + # Build a login-only baseline. + data = json.loads(json.dumps(_BASELINE)) + del data["environment"]["required_reviewer_ids"] + bpath = tmp_path / "logins.json" + bpath.write_text(json.dumps(data), encoding="utf-8") + + # Build a shim bin in this tmp_path. + bin_dir = tmp_path / "bin" + bin_dir.mkdir() + calls_log = tmp_path / "calls.log" + gh = bin_dir / "gh" + gh.write_text( + "#!/usr/bin/env bash\n" + f'printf "gh %s\\n" "$*" >> "{calls_log}"\n' + 'argv="$*"\n' + 'for a in "$@"; do\n' + ' case "$a" in\n' + " */enforce_admins) echo 'true'; exit 0 ;;\n" + " esac\n" + "done\n" + 'case "$argv" in\n' + " *users/*) echo '7654321'; exit 0 ;;\n" + " *environments/*) echo '7654321'; exit 0 ;;\n" + "esac\n" + "exit 0\n", + encoding="utf-8", + ) + gh.chmod(gh.stat().st_mode | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH) + + env = dict(os.environ) + env["PATH"] = f"{bin_dir}{os.pathsep}{env['PATH']}" + res = subprocess.run( + ["bash", str(_SCRIPT), "environment", "--apply", "--baseline", str(bpath)], + capture_output=True, + text=True, + env=env, + ) + assert res.returncode == 0, res.stderr + res.stdout + calls = calls_log.read_text(encoding="utf-8") + assert "api users/amoussa1229" in calls + assert "reviewers[][id]=7654321" in calls + assert "[OK]" in res.stdout + + +# --------------------------------------------------------------------------- # +# --record-baseline mode +# --------------------------------------------------------------------------- # + + +def test_record_baseline_rejects_a_surface(tmp_path: Path) -> None: + out = tmp_path / "p3-baseline.json" + res = subprocess.run( + [ + "bash", + str(_SCRIPT), + "--record-baseline", + "protection", + "--baseline", + str(out), + "--repo", + "Sea-Haven-Industries/orchestrator", + ], + capture_output=True, + text=True, + ) + assert res.returncode == 2 + assert "does not take a surface" in res.stderr + + +def test_record_baseline_dry_run_prints_plan(tmp_path: Path) -> None: + out = tmp_path / "p3-baseline.json" + res = subprocess.run( + [ + "bash", + str(_SCRIPT), + "--record-baseline", + "--baseline", + str(out), + "--repo", + "Sea-Haven-Industries/orchestrator", + ], + capture_output=True, + text=True, + ) + assert res.returncode == 0, res.stderr + assert "RECORD BASELINE" in res.stdout + assert "[PLAN]" in res.stdout + # Dry-run captures nothing. + assert not out.exists() + + +def test_record_baseline_requires_repo(tmp_path: Path) -> None: + out = tmp_path / "p3-baseline.json" + # Clear the env-derived repo default so no repo is resolvable. + env = dict(os.environ) + env.pop("AGENT_TEAM_REPO_OWNER", None) + env.pop("AGENT_TEAM_REPO_NAME", None) + res = subprocess.run( + ["bash", str(_SCRIPT), "--record-baseline", "--baseline", str(out)], + capture_output=True, + text=True, + env=env, + ) + assert res.returncode != 0 + assert "no target repo" in res.stderr + + +def test_record_baseline_apply_writes_baseline_from_live_state( + tmp_path: Path, +) -> None: + """--record-baseline --apply captures the live workflow SHA, env reviewer + ids, full branch protection and App installation id into the baseline JSON, + creating the directory if absent. The recorded file is then a valid restore + target whose protection.include_administrators is True.""" + out = tmp_path / ".security-review" / "p3-baseline.json" # dir absent on purpose + + bin_dir = tmp_path / "bin" + bin_dir.mkdir() + calls_log = tmp_path / "calls.log" + protection_json = json.dumps(_BASELINE["protection"]["full"]) + gh = bin_dir / "gh" + gh.write_text( + "#!/usr/bin/env bash\n" + f'printf "gh %s\\n" "$*" >> "{calls_log}"\n' + 'argv="$*"\n' + 'case "$argv" in\n' + " *contents/*) echo 'deadbeefsha'; exit 0 ;;\n" + " *branches/*/protection*)\n" + f" cat <<'JSON'\n{protection_json}\nJSON\n" + " exit 0 ;;\n" + " *environments/*) echo '[1234567]'; exit 0 ;;\n" + "esac\n" + "exit 0\n", + encoding="utf-8", + ) + gh.chmod(gh.stat().st_mode | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH) + + env = dict(os.environ) + env["PATH"] = f"{bin_dir}{os.pathsep}{env['PATH']}" + res = subprocess.run( + [ + "bash", + str(_SCRIPT), + "--record-baseline", + "--apply", + "--baseline", + str(out), + "--repo", + "Sea-Haven-Industries/orchestrator", + "--app-installation-id", + "424242", + ], + capture_output=True, + text=True, + env=env, + ) + assert res.returncode == 0, res.stderr + res.stdout + assert out.exists(), "baseline file (and its dir) must be created" + recorded = json.loads(out.read_text(encoding="utf-8")) + assert recorded["repo"] == "Sea-Haven-Industries/orchestrator" + assert recorded["workflow_baseline_sha"] == "deadbeefsha" + assert recorded["environment"]["required_reviewer_ids"] == [1234567] + assert recorded["app"]["installation_id"] == 424242 + assert recorded["app"]["action"] == "uninstall" + assert recorded["protection"]["include_administrators"] is True + assert recorded["protection"]["full"]["enforce_admins"]["enabled"] is True + + +def test_recorded_baseline_is_a_valid_restore_target(tmp_path: Path) -> None: + """A baseline produced by --record-baseline --apply can be fed straight back + into a restore (dry-run) without error — closing the record->restore loop.""" + out = tmp_path / ".security-review" / "p3-baseline.json" + bin_dir = tmp_path / "bin" + bin_dir.mkdir() + calls_log = tmp_path / "calls.log" + protection_json = json.dumps(_BASELINE["protection"]["full"]) + gh = bin_dir / "gh" + gh.write_text( + "#!/usr/bin/env bash\n" + f'printf "gh %s\\n" "$*" >> "{calls_log}"\n' + 'argv="$*"\n' + 'case "$argv" in\n' + " *contents/*) echo 'deadbeefsha'; exit 0 ;;\n" + " *branches/*/protection*)\n" + f" cat <<'JSON'\n{protection_json}\nJSON\n" + " exit 0 ;;\n" + " *environments/*) echo '[1234567]'; exit 0 ;;\n" + "esac\n" + "exit 0\n", + encoding="utf-8", + ) + gh.chmod(gh.stat().st_mode | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH) + env = dict(os.environ) + env["PATH"] = f"{bin_dir}{os.pathsep}{env['PATH']}" + rec = subprocess.run( + [ + "bash", + str(_SCRIPT), + "--record-baseline", + "--apply", + "--baseline", + str(out), + "--repo", + "Sea-Haven-Industries/orchestrator", + "--app-installation-id", + "424242", + ], + capture_output=True, + text=True, + env=env, + ) + assert rec.returncode == 0, rec.stderr + rec.stdout + + # Now restore (dry-run) from the recorded baseline. + res = subprocess.run( + ["bash", str(_SCRIPT), "protection", "--baseline", str(out)], + capture_output=True, + text=True, + ) + assert res.returncode == 0, res.stderr + res.stdout + assert "PUT full branch protection" in res.stdout + + +def test_apply_protection_DETECTS_divergent_live_restore( + baseline: Path, tmp_path: Path +) -> None: + """Regression for the normalize_protection heredoc bug (vacuous assert). + + Previously ``normalize_protection`` ran ``python3 - <<'PY'`` and read + ``json.load(sys.stdin)`` — but the heredoc IS stdin, so it parsed the program + text, raised, and the bare except swallowed it to "" for EVERY input. Both + live and baseline normalized to "" and the post-restore assert was always + "" == "" -> a broken protection restore reported SUCCESS. The fix reads the + JSON from argv[1]. This test feeds a LIVE protection that DIFFERS from the + baseline and asserts the script now FAILS (non-zero) instead of falsely + passing. + """ + import copy + + divergent = copy.deepcopy(_BASELINE["protection"]["full"]) + # Flip a projected field so the normalized live != normalized baseline. + divergent["allow_force_pushes"] = {"enabled": True} + + bin_dir = tmp_path / "divbin" + bin_dir.mkdir() + calls_log = tmp_path / "divcalls.log" + div_json = json.dumps(divergent) + gh = bin_dir / "gh" + gh.write_text( + "#!/usr/bin/env bash\n" + f'printf "gh %s\\n" "$*" >> "{calls_log}"\n' + 'argv="$*"\n' + 'for a in "$@"; do\n' + ' case "$a" in\n' + " */enforce_admins) echo 'true'; exit 0 ;;\n" + " esac\n" + "done\n" + 'case "$argv" in\n' + " *branches/*/protection*)\n" + f" cat <<'JSON'\n{div_json}\nJSON\n" + " exit 0 ;;\n" + " *users/*) echo '1234567'; exit 0 ;;\n" + " *environments/*) echo '1234567'; exit 0 ;;\n" + "esac\n" + "exit 0\n", + encoding="utf-8", + ) + git = bin_dir / "git" + git.write_text( + f'#!/usr/bin/env bash\nprintf "git %s\\n" "$*" >> "{calls_log}"\nexit 0\n', + encoding="utf-8", + ) + for f in (gh, git): + f.chmod(f.stat().st_mode | stat.S_IEXEC | stat.S_IXGRP | stat.S_IXOTH) + + res = _run("protection", "--apply", baseline=baseline, shim=(bin_dir, calls_log)) + assert res.returncode != 0, ( + "divergent live protection must FAIL the post-restore assert, not pass:\n" + + res.stdout + + res.stderr + ) + assert "protection.full" in (res.stdout + res.stderr) + + +def test_missing_required_key_refuses_partial_restore(tmp_path: Path) -> None: + # MEDIUM-1: a baseline missing a required key for a surface aborts that + # surface up front rather than half-restoring it. + data = json.loads(json.dumps(_BASELINE)) + del data["workflow_baseline_sha"] + p = tmp_path / "no_sha.json" + p.write_text(json.dumps(data), encoding="utf-8") + res = _run("workflow", baseline=p) + assert res.returncode != 0 + out = res.stdout + res.stderr + assert "missing required key" in out + assert "workflow_baseline_sha" in out + + +def test_app_uninstall_without_jwt_requires_oob_ack(tmp_path: Path) -> None: + # MEDIUM-2: an --apply that needs a MANUAL App neutralise (no APP JWT) must + # not silently skip it — it refuses unless the operator acknowledges. + p = tmp_path / "bl.json" + p.write_text(json.dumps(_BASELINE), encoding="utf-8") + + # No JWT, no ack -> refuse. + refused = _run("app", "--apply", baseline=p) + assert refused.returncode != 0 + assert "P3_ROLLBACK_OOB_ACK" in (refused.stdout + refused.stderr) + + # No JWT, but acked -> proceeds (App neutralise is operator-owed, loudly warned). + acked = _run("app", "--apply", baseline=p, extra_env={"P3_ROLLBACK_OOB_ACK": "1"}) + assert acked.returncode == 0, acked.stdout + acked.stderr + assert "OUT-OF-BAND ACK accepted" in (acked.stdout + acked.stderr) diff --git a/agent-team/tests/test_run_team.py b/agent-team/tests/test_run_team.py index ad6de38..2cae0a1 100644 --- a/agent-team/tests/test_run_team.py +++ b/agent-team/tests/test_run_team.py @@ -17,6 +17,7 @@ import argparse import importlib.util import io import json +from datetime import timedelta from pathlib import Path from types import ModuleType from typing import Any @@ -647,8 +648,12 @@ class _FakeCoordinator: build_clarify_node: Any = None, build_plan_node: Any = None, review_wiring: Any = None, + build_verify_wiring: Any = None, + dispatch_node_wiring: Any = None, notify: Any = None, alarm_hook: Any = None, + ci_poller: Any = None, + ci_timeout: Any = None, ) -> None: self.db_path = db_path self.transport = transport @@ -659,11 +664,28 @@ class _FakeCoordinator: self.build_clarify_node = build_clarify_node self.build_plan_node = build_plan_node self.review_wiring = review_wiring + # P3 fail-safe serve default (Decision 5): only the ``serve`` command + # auto-binds these; start/intake leave them None. + self.build_verify_wiring = build_verify_wiring + self.dispatch_node_wiring = dispatch_node_wiring + # CI-watcher seams (§4 Decision 2): the live serve path binds the poller + + # timeout here and the provider post-construction; inert leaves all None. + self.ci_poller = ci_poller + self.ci_timeout = ci_timeout + self._ci_pending_provider: Any = None + # Draft-PR runaway/stale monitor provider (P3 A4): the live serve path + # binds a read-only enumerator here post-construction; inert leaves None. + self._draft_pr_provider: Any = None self.setup_called = False self.start_kwargs: dict[str, Any] | None = None self.new_task_callback: Any = None _FakeCoordinator.instances.append(self) + def _enumerate_ci_pending(self) -> list[Any]: + # Stand-in for the durable enumerator the live path binds as the + # ci_pending_provider; identity is what the wiring test asserts. + return [] + def setup(self) -> None: self.setup_called = True @@ -863,6 +885,204 @@ def test_intake_github_label_required(cli: ModuleType) -> None: parser.parse_args(["intake-github", "--owner", "o", "--repo", "r"]) +# --------------------------------------------------------------------------- # +# serve fail-safe P3 wiring default (design Decision 5; UNIT 0e) +# --------------------------------------------------------------------------- # + + +def _serve_args(db_path: Path, command: str) -> argparse.Namespace: + """A minimal args namespace for ``_build_coordinator`` (dry-run, no token).""" + return argparse.Namespace( + command=command, + db=db_path, + transport="slack", + dry_run=True, + ) + + +def test_serve_binds_inert_p3_wiring_when_env_unset( + cli: ModuleType, + db_path: Path, + monkeypatch: pytest.MonkeyPatch, +) -> None: + """``serve`` is the new P3 default but degrades to inert (None) when the P3 + env is unset — never crashing serve-start.""" + for name in ( + "AGENT_TEAM_REPO_OWNER", + "AGENT_TEAM_REPO_NAME", + "AGENT_TEAM_CI_READ_TOKEN", + "GITHUB_TOKEN", + ): + monkeypatch.delenv(name, raising=False) + _FakeCoordinator.instances.clear() + monkeypatch.setattr( + "agent_team.coordinator.Coordinator", _FakeCoordinator, raising=True + ) + + coord = cli._build_coordinator(_serve_args(db_path, "serve")) + + assert coord.build_verify_wiring is None + assert coord.dispatch_node_wiring is None + # Inert box: no CI-watcher seams, so the tick() CI sweep is a NO-OP. + assert coord.ci_poller is None + assert coord.ci_timeout is None + assert coord._ci_pending_provider is None + # A4 draft-PR monitor stays inert too, gated on the same signal as ci_poller: + # the sweep is a NO-OP, so no production runaway/stale sweep ever fires. + assert coord._draft_pr_provider is None + + +def test_serve_binds_live_p3_wiring_when_env_set( + cli: ModuleType, + db_path: Path, + monkeypatch: pytest.MonkeyPatch, +) -> None: + """With the P3 env provisioned, ``serve`` binds the live wiring pair.""" + monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "Sea-Haven-Industries") + monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "orchestrator") + monkeypatch.setenv("AGENT_TEAM_CI_READ_TOKEN", "ghp_test") + _FakeCoordinator.instances.clear() + monkeypatch.setattr( + "agent_team.coordinator.Coordinator", _FakeCoordinator, raising=True + ) + + coord = cli._build_coordinator(_serve_args(db_path, "serve")) + + assert callable(coord.build_verify_wiring) + assert callable(coord.dispatch_node_wiring) + # Live box: the CI-watcher seams are wired so a task suspended at VERIFY + # awaiting CI gets resumed/parked rather than waiting forever. The poller is + # the read-only default; the provider is the coordinator's durable + # enumerator (bound post-construction); the timeout is the 30-min default. + assert callable(coord.ci_poller) + assert coord.ci_timeout == timedelta(minutes=30) + assert coord._ci_pending_provider == coord._enumerate_ci_pending + # A4 draft-PR monitor: the live serve path binds a read-only provider so the + # sweep has a real snapshot to ALARM / remind on (gated on the same live-pair + # signal as ci_poller). Without this the wired sweep would always see no PRs. + assert callable(coord._draft_pr_provider) + + +def test_serve_draft_pr_provider_reads_open_apply_draft_prs( + cli: ModuleType, + db_path: Path, + monkeypatch: pytest.MonkeyPatch, +) -> None: + """The bound draft-PR provider is READ-ONLY: it shells a scoped ``gh pr list`` + (no write/close/dispatch) and maps the JSON into ``DraftPr`` snapshots.""" + from agent_team.draft_pr_monitor import DraftPr + + monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "Sea-Haven-Industries") + monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "orchestrator") + monkeypatch.setenv("AGENT_TEAM_CI_READ_TOKEN", "ghp_test") + _FakeCoordinator.instances.clear() + monkeypatch.setattr( + "agent_team.coordinator.Coordinator", _FakeCoordinator, raising=True + ) + + captured: dict[str, Any] = {} + + class _FakeProc: + stdout = json.dumps( + [ + { + "number": 7, + "createdAt": "2026-06-23T10:00:00Z", + "updatedAt": "2026-06-23T10:05:00Z", + }, + { + "number": 9, + "createdAt": "2026-06-10T09:00:00Z", + "updatedAt": "2026-06-10T09:00:00Z", + }, + ] + ) + + def _fake_run(args: Any, **kwargs: Any) -> _FakeProc: + # The provider must shell a READ-ONLY, namespace-scoped enumeration — + # never a write/close/merge subcommand. + captured["args"] = args + captured["kwargs"] = kwargs + return _FakeProc() + + monkeypatch.setattr("subprocess.run", _fake_run, raising=True) + + coord = cli._build_coordinator(_serve_args(db_path, "serve")) + assert callable(coord._draft_pr_provider) + + prs = coord._draft_pr_provider() + + # READ-ONLY + scoped: it is a `gh pr list` over the apply/ head namespace, not + # a mutating subcommand, and never carries --shell. + assert captured["args"][:3] == ["gh", "pr", "list"] + assert "--draft" in captured["args"] + assert f"head:{cli._DRAFT_PR_HEAD_PREFIX}" in captured["args"] + assert "Sea-Haven-Industries/orchestrator" in captured["args"] + assert not any( + tok in captured["args"] for tok in ("close", "merge", "edit", "ready") + ) + # The JSON rows map field-for-field into the snapshot the monitor expects. + assert prs == [ + DraftPr( + number=7, + opened_at="2026-06-23T10:00:00Z", + updated_at="2026-06-23T10:05:00Z", + ), + DraftPr( + number=9, + opened_at="2026-06-10T09:00:00Z", + updated_at="2026-06-10T09:00:00Z", + ), + ] + + +def test_serve_draft_pr_provider_fails_soft_on_gh_error( + cli: ModuleType, + db_path: Path, + monkeypatch: pytest.MonkeyPatch, +) -> None: + """A non-zero ``gh`` exit yields an empty snapshot (the sweep no-ops) rather + than raising and breaking the tick loop.""" + import subprocess + + monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "Sea-Haven-Industries") + monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "orchestrator") + monkeypatch.setenv("AGENT_TEAM_CI_READ_TOKEN", "ghp_test") + _FakeCoordinator.instances.clear() + monkeypatch.setattr( + "agent_team.coordinator.Coordinator", _FakeCoordinator, raising=True + ) + + def _boom(args: Any, **kwargs: Any) -> Any: + raise subprocess.CalledProcessError(returncode=1, cmd=args) + + monkeypatch.setattr("subprocess.run", _boom, raising=True) + + coord = cli._build_coordinator(_serve_args(db_path, "serve")) + assert coord._draft_pr_provider() == [] + + +def test_start_does_not_bind_p3_wiring_even_when_env_set( + cli: ModuleType, + db_path: Path, + monkeypatch: pytest.MonkeyPatch, +) -> None: + """Only ``serve`` (the daemon) auto-binds P3; the one-shot ``start`` path runs + to the first human gate and never wires build/verify/dispatch.""" + monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "Sea-Haven-Industries") + monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "orchestrator") + monkeypatch.setenv("AGENT_TEAM_CI_READ_TOKEN", "ghp_test") + _FakeCoordinator.instances.clear() + monkeypatch.setattr( + "agent_team.coordinator.Coordinator", _FakeCoordinator, raising=True + ) + + coord = cli._build_coordinator(_serve_args(db_path, "start")) + + assert coord.build_verify_wiring is None + assert coord.dispatch_node_wiring is None + + def test_build_transport_dry_run_returns_dry_run_transport(cli: ModuleType) -> None: """``dry_run=True`` yields a _DryRunTransport whose post returns a synthetic ref.""" args = argparse.Namespace(dry_run=True, transport="slack") @@ -1077,3 +1297,65 @@ def test_notify_sink_forwards_thread_ts( # No thread_ts when none is given (top-level post, not a broken key). assert "thread_ts" not in captured[1] assert captured[1] == {"channel": "C123", "text": "top-level milestone"} + + +def test_dispatch_from_files_invokes_dispatcher( + cli: ModuleType, db_path: Path, audit_log: Path, tmp_path: Path, monkeypatch +) -> None: + """`dispatch` reads a diff/scope file and fires dispatch_apply_verify (P3 op-b).""" + from agent_team import dispatcher as d + + diff_f = tmp_path / "d.diff" + diff_f.write_text("diff --git a/x b/x\n@@ -1 +1 @@\n-a\n+b\n", encoding="utf-8") + scope_f = tmp_path / "s.txt" + scope_f.write_text("agent_team/\n", encoding="utf-8") + + calls: dict = {} + + def _fake_dispatch(*, owner, repo, task_id, diff_text, declared_scope, base): + calls.update( + owner=owner, repo=repo, task_id=task_id, scope=declared_scope, base=base + ) + return d.DispatchResult( + inputs=d.build_dispatch_inputs( + task_id=task_id, diff_text=diff_text, declared_scope=declared_scope + ), + run_id="27990718108", + dispatched_at="2026-06-24T00:00:00Z", + correlation_tag=task_id, + ) + + monkeypatch.setattr(d, "dispatch_apply_verify", _fake_dispatch) + code, out = _run( + cli, + db_path, + audit_log, + "dispatch", + "task-xyz", + "--owner", + "Sea-Haven-Industries", + "--repo", + "orchestrator", + "--diff", + str(diff_f), + "--scope", + str(scope_f), + ) + assert code == 0, out + assert calls["owner"] == "Sea-Haven-Industries" + assert calls["repo"] == "orchestrator" + assert calls["task_id"] == "task-xyz" + assert "agent_team/" in calls["scope"] + assert "27990718108" in out # the located run_id is reported + + +def test_dispatch_requires_owner_repo( + cli: ModuleType, db_path: Path, audit_log: Path, tmp_path: Path, monkeypatch +) -> None: + """Without owner/repo (args or env) dispatch refuses with exit 2, no dispatch.""" + monkeypatch.delenv("AGENT_TEAM_REPO_OWNER", raising=False) + monkeypatch.delenv("AGENT_TEAM_REPO_NAME", raising=False) + diff_f = tmp_path / "d.diff" + diff_f.write_text("diff --git a/x b/x\n", encoding="utf-8") + code, _ = _run(cli, db_path, audit_log, "dispatch", "t1", "--diff", str(diff_f)) + assert code == 2 diff --git a/agent-team/tests/test_runaway_monitor.py b/agent-team/tests/test_runaway_monitor.py new file mode 100644 index 0000000..2b07efc --- /dev/null +++ b/agent-team/tests/test_runaway_monitor.py @@ -0,0 +1,400 @@ +"""Unit tests for agent_team.draft_pr_monitor (P3 box-integration, A4). + +Covers the draft-PR runaway/stale sweep: it ALARMs when > 3 draft PRs open +within 15 minutes, reminds on a draft PR idle > 7 days, applies flapping backoff +(one ALARM per cooldown, one reminder per PR per cooldown), and never auto-closes +or self-stops. No network: every side effect (alarm, reminder) is an injected +callable, exactly as the module's contract promises. +""" + +from __future__ import annotations + +from datetime import datetime, timedelta, timezone + +from agent_team.draft_pr_monitor import ( + DEFAULT_ALARM_COOLDOWN, + DEFAULT_STALE_COOLDOWN, + DraftPr, + MonitorAction, + MonitorMemory, + run_draft_pr_monitor, +) + +# A fixed "now"; ISO strings mirror the GitHub API / ledger (UTC). +_NOW = datetime(2026, 6, 23, 12, 0, 0, tzinfo=timezone.utc) + + +def _ago(**kwargs: float) -> str: + return (_NOW - timedelta(**kwargs)).isoformat() + + +class _Recorder: + """Records the runaway counts / stale PRs handed to the side-effect hooks.""" + + def __init__(self) -> None: + self.alarms: list[int] = [] + self.reminders: list[int] = [] + + def on_alarm(self, opened_in_window: int) -> None: + self.alarms.append(opened_in_window) + + def on_stale_reminder(self, pr: DraftPr) -> None: + self.reminders.append(pr.number) + + +def _open_burst(count: int, *, within_minutes: int = 5) -> list[DraftPr]: + """``count`` draft PRs all opened ``within_minutes`` ago (inside the window).""" + return [ + DraftPr( + number=i, + opened_at=_ago(minutes=within_minutes), + updated_at=_NOW.isoformat(), + ) + for i in range(count) + ] + + +# --------------------------------------------------------------------------- # +# Runaway ALARM threshold (> 3 within 15 min) +# --------------------------------------------------------------------------- # + + +def test_runaway_alarm_fires_above_threshold() -> None: + rec = _Recorder() + report = run_draft_pr_monitor( + _open_burst(4), + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + now=_NOW, + ) + assert report.alarmed == 1 + assert rec.alarms == [4] + assert report.outcomes[0].action is MonitorAction.ALARMED + + +def test_exactly_threshold_does_not_alarm() -> None: + """The rule is STRICTLY greater than 3 — exactly 3 must NOT alarm.""" + rec = _Recorder() + report = run_draft_pr_monitor( + _open_burst(3), + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + now=_NOW, + ) + assert report.alarmed == 0 + assert rec.alarms == [] + + +def test_prs_opened_outside_window_are_not_counted() -> None: + """5 PRs but only 2 inside the 15-min window -> no alarm.""" + rec = _Recorder() + prs = [ + DraftPr(number=1, opened_at=_ago(minutes=2)), + DraftPr(number=2, opened_at=_ago(minutes=10)), + DraftPr(number=3, opened_at=_ago(minutes=20)), + DraftPr(number=4, opened_at=_ago(minutes=40)), + DraftPr(number=5, opened_at=_ago(hours=3)), + ] + report = run_draft_pr_monitor( + prs, + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + now=_NOW, + ) + assert report.opened_in_window == 2 + assert report.alarmed == 0 + assert rec.alarms == [] + + +def test_unparseable_opened_at_is_skipped_for_runaway() -> None: + """A PR with a junk/missing opened_at is not counted and does not crash.""" + rec = _Recorder() + prs = _open_burst(4) + [DraftPr(number=99, opened_at="not-a-date")] + report = run_draft_pr_monitor( + prs, + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + now=_NOW, + ) + # The 4 valid PRs still breach; the junk one is just ignored for the count. + assert report.opened_in_window == 4 + assert report.alarmed == 1 + + +# --------------------------------------------------------------------------- # +# Flapping backoff — ALARM at most once per cooldown +# --------------------------------------------------------------------------- # + + +def test_runaway_alarm_suppressed_within_cooldown() -> None: + rec = _Recorder() + mem = MonitorMemory() + prs = _open_burst(5) + + first = run_draft_pr_monitor( + prs, + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + memory=mem, + now=_NOW, + ) + assert first.alarmed == 1 + + # A second tick a minute later, condition still breaching -> suppressed. + second = run_draft_pr_monitor( + prs, + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + memory=mem, + now=_NOW + timedelta(minutes=1), + ) + assert second.alarmed == 0 + assert second.alarm_suppressed == 1 + assert rec.alarms == [5] # only the first post + + +def test_runaway_alarm_refires_after_cooldown() -> None: + rec = _Recorder() + mem = MonitorMemory() + prs = _open_burst(5, within_minutes=1) + + run_draft_pr_monitor( + prs, + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + memory=mem, + now=_NOW, + ) + # Past the cooldown, a still-breaching condition ALARMs again. Re-anchor the + # PRs so they remain inside the window relative to the later "now". + later = _NOW + DEFAULT_ALARM_COOLDOWN + timedelta(seconds=1) + prs_later = [ + DraftPr(number=p.number, opened_at=(later - timedelta(minutes=1)).isoformat()) + for p in prs + ] + report = run_draft_pr_monitor( + prs_later, + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + memory=mem, + now=later, + ) + assert report.alarmed == 1 + assert len(rec.alarms) == 2 + + +# --------------------------------------------------------------------------- # +# Stale reminder (> 7 days idle), never auto-closes +# --------------------------------------------------------------------------- # + + +def test_stale_pr_reminds_once() -> None: + rec = _Recorder() + prs = [DraftPr(number=7, opened_at=_ago(days=30), updated_at=_ago(days=8))] + report = run_draft_pr_monitor( + prs, + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + now=_NOW, + ) + assert report.stale_reminded == 1 + assert rec.reminders == [7] + + +def test_fresh_pr_does_not_remind() -> None: + rec = _Recorder() + prs = [DraftPr(number=7, opened_at=_ago(days=30), updated_at=_ago(days=6))] + report = run_draft_pr_monitor( + prs, + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + now=_NOW, + ) + assert report.stale_reminded == 0 + assert rec.reminders == [] + + +def test_exactly_7_days_does_not_remind() -> None: + """The rule is strictly MORE than 7 days idle.""" + rec = _Recorder() + prs = [DraftPr(number=7, updated_at=_ago(days=7))] + report = run_draft_pr_monitor( + prs, + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + now=_NOW, + ) + assert report.stale_reminded == 0 + + +def test_stale_reminder_suppressed_within_cooldown() -> None: + rec = _Recorder() + mem = MonitorMemory() + prs = [DraftPr(number=7, updated_at=_ago(days=10))] + + run_draft_pr_monitor( + prs, + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + memory=mem, + now=_NOW, + ) + second = run_draft_pr_monitor( + prs, + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + memory=mem, + now=_NOW + timedelta(hours=1), + ) + assert second.stale_reminded == 0 + assert second.stale_suppressed == 1 + assert rec.reminders == [7] # only the first + + +def test_stale_reminder_refires_after_cooldown() -> None: + rec = _Recorder() + mem = MonitorMemory() + + run_draft_pr_monitor( + [DraftPr(number=7, updated_at=_ago(days=10))], + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + memory=mem, + now=_NOW, + ) + later = _NOW + DEFAULT_STALE_COOLDOWN + timedelta(seconds=1) + report = run_draft_pr_monitor( + [DraftPr(number=7, updated_at=(later - timedelta(days=10)).isoformat())], + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + memory=mem, + now=later, + ) + assert report.stale_reminded == 1 + assert len(rec.reminders) == 2 + + +def test_unparseable_updated_at_is_skipped_for_stale() -> None: + rec = _Recorder() + prs = [DraftPr(number=7, updated_at="garbage")] + report = run_draft_pr_monitor( + prs, + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + now=_NOW, + ) + assert report.stale_reminded == 0 + assert rec.reminders == [] + + +# --------------------------------------------------------------------------- # +# Side-effect isolation + watermark-not-advanced-on-error +# --------------------------------------------------------------------------- # + + +def test_alarm_side_effect_error_is_isolated_and_retried() -> None: + """A raising on_alarm yields ERRORED and does NOT advance the watermark.""" + calls: list[int] = [] + + def boom(opened_in_window: int) -> None: + calls.append(opened_in_window) + raise RuntimeError("slack down") + + mem = MonitorMemory() + prs = _open_burst(5) + report = run_draft_pr_monitor( + prs, + on_alarm=boom, + on_stale_reminder=lambda pr: None, + memory=mem, + now=_NOW, + ) + assert report.errored == 1 + assert report.alarmed == 0 + # Watermark not advanced -> a retry the very next tick (no cooldown lock-in). + assert mem.last_alarm_at is None + report2 = run_draft_pr_monitor( + prs, + on_alarm=boom, + on_stale_reminder=lambda pr: None, + memory=mem, + now=_NOW + timedelta(seconds=30), + ) + assert report2.errored == 1 + assert len(calls) == 2 + + +def test_stale_side_effect_error_does_not_abort_sweep() -> None: + """A reminder that raises for one PR does not stop the runaway ALARM.""" + rec = _Recorder() + + def boom(pr: DraftPr) -> None: + raise RuntimeError("slack down") + + prs = _open_burst(5) + [DraftPr(number=50, updated_at=_ago(days=10))] + report = run_draft_pr_monitor( + prs, + on_alarm=rec.on_alarm, + on_stale_reminder=boom, + now=_NOW, + ) + # Runaway still alarms; the stale reminder errored but was isolated. + assert report.alarmed == 1 + assert report.errored == 1 + + +# --------------------------------------------------------------------------- # +# Never auto-closes / self-stops; memory hygiene +# --------------------------------------------------------------------------- # + + +def test_monitor_takes_no_infrastructure_action() -> None: + """The only effects are the two injected notify callables — nothing else. + + There is no close/stop/dispatch seam on the monitor; this asserts the sweep + exposes ONLY the alarm + reminder hooks (a regression guard against adding a + self-acting side effect). + """ + import inspect + + sig = inspect.signature(run_draft_pr_monitor) + effect_params = {p for p in sig.parameters if p.startswith("on_")} + assert effect_params == {"on_alarm", "on_stale_reminder"} + + +def test_memory_prunes_gone_prs() -> None: + """A PR no longer present has its backoff watermark pruned (bounded growth).""" + rec = _Recorder() + mem = MonitorMemory() + + run_draft_pr_monitor( + [DraftPr(number=7, updated_at=_ago(days=10))], + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + memory=mem, + now=_NOW, + ) + assert 7 in mem.last_reminded_at + + # Next pass: PR #7 is gone (merged/closed by a human). Its watermark prunes. + run_draft_pr_monitor( + [], + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + memory=mem, + now=_NOW + timedelta(minutes=1), + ) + assert 7 not in mem.last_reminded_at + + +def test_no_draft_prs_is_quiet() -> None: + rec = _Recorder() + report = run_draft_pr_monitor( + [], + on_alarm=rec.on_alarm, + on_stale_reminder=rec.on_stale_reminder, + now=_NOW, + ) + assert report.outcomes == [] + assert rec.alarms == [] + assert rec.reminders == [] diff --git a/agent-team/tests/test_task_model.py b/agent-team/tests/test_task_model.py index 28886c0..dbcb626 100644 --- a/agent-team/tests/test_task_model.py +++ b/agent-team/tests/test_task_model.py @@ -73,12 +73,35 @@ def test_roundtrip_dict() -> None: candidate_diff="diff --git a b", diff_hash="deadbeef", ci_results={"conclusion": "success"}, + run_id="27990718108", + ci_correlation_tag="t1-1a2b3c", + dispatched_at="2026-06-17T00:30:00Z", + build_loops=2, transport="slack", created_at="2026-06-17T00:00:00Z", updated_at="2026-06-17T01:00:00Z", ) restored = task_from_dict(task_to_dict(rec)) assert restored == rec + # P3 dispatch->verify plumbing fields survive the dict round-trip. + assert restored.run_id == "27990718108" + assert restored.ci_correlation_tag == "t1-1a2b3c" + assert restored.dispatched_at == "2026-06-17T00:30:00Z" + # The durable build<->verify loop count round-trips (LOGIC-RACE-01). + assert restored.build_loops == 2 + + +def test_build_loops_defaults_zero_and_roundtrips() -> None: + rec = TaskRecord( + thread_id="t1", + status=TaskStatus.ACTIVE, + current_phase=Phase.BUILD, + ) + assert rec.build_loops == 0 + # A dict missing build_loops (older record) defaults to 0, not a crash. + legacy = task_to_dict(rec) + del legacy["build_loops"] + assert task_from_dict(legacy).build_loops == 0 def test_roundtrip_json() -> None: diff --git a/agent-team/tests/test_verifier.py b/agent-team/tests/test_verifier.py index dc362b5..60a869f 100644 --- a/agent-team/tests/test_verifier.py +++ b/agent-team/tests/test_verifier.py @@ -123,6 +123,69 @@ def test_fail_increments_build_loops_in_verdict() -> None: assert out["review_verdicts"][0]["build_loops"] == 2 +# --------------------------------------------------------------------------- # +# Durable per-task build-loop budget (LOGIC-RACE-01) +# --------------------------------------------------------------------------- # + + +def test_fail_reads_loop_count_from_state_not_config() -> None: + # The count that bounds the loop is DURABLE per-task state. A shared config + # with build_loops=0 must NOT mask a per-task state["build_loops"] of 2. + diff = _diff_for("src/foo.py") + ci = {"run_id": "r1", "conclusion": "failure", "diff_hash": _hash(diff)} + state = _state(diff, ci) + state["build_loops"] = 2 + out = verifier_node(state, VerifierConfig(expected_run_id="r1", build_loops=0)) + # next = state(2) + 1 = 3 == DEFAULT_MAX_BUILD_LOOPS -> park (not a vacuous + # loop driven by the always-zero shared config). + assert out["status"] == TaskStatus.PARKED.value + assert out["review_verdicts"][0]["build_loops"] == 3 + + +def test_fail_writes_incremented_loop_count_into_state() -> None: + # On a recoverable FAIL the node persists the incremented count into the + # returned partial state so the NEXT VERIFY (after BUILD->DISPATCH) sees it. + diff = _diff_for("src/foo.py") + ci = {"run_id": "r1", "conclusion": "failure", "diff_hash": _hash(diff)} + state = _state(diff, ci) + state["build_loops"] = 0 + out = verifier_node(state, VerifierConfig(expected_run_id="r1")) + assert out["current_phase"] == Phase.BUILD.value + assert out["build_loops"] == 1 + + +def test_repeated_fail_parks_after_exactly_max_build_loops_via_state() -> None: + # Simulate the durable BUILD->DISPATCH->VERIFY loop: the incremented + # build_loops the node returns is fed back into the next call's state (the + # checkpointer carries it). A perpetually-FAILing task must reach PARKED + # after EXACTLY max_build_loops iterations, never loop unbounded. + diff = _diff_for("src/foo.py") + ci = {"run_id": "r1", "conclusion": "failure", "diff_hash": _hash(diff)} + # A single SHARED config (build_loops stays 0) — exactly the wiring-time + # reality that made the budget dead before the fix. + shared_cfg = VerifierConfig(expected_run_id="r1", max_build_loops=3, build_loops=0) + + state = _state(diff, ci) + state["build_loops"] = 0 + + statuses: list[str] = [] + for _ in range(10): # bound the harness so a regression can't hang the test + out = verifier_node(state, shared_cfg) + statuses.append(out["status"]) + if out["status"] == TaskStatus.PARKED.value: + break + # Carry the durable count forward, as the checkpointer would. + state["build_loops"] = out["build_loops"] + + # FAILs at counts 1, 2 loop back to BUILD; the 3rd (next==3==max) parks. + assert statuses == [ + TaskStatus.ACTIVE.value, + TaskStatus.ACTIVE.value, + TaskStatus.PARKED.value, + ] + assert out["current_phase"] == Phase.PARKED.value + + # --------------------------------------------------------------------------- # # BLOCK path — park for human + GPT cross-review # --------------------------------------------------------------------------- # @@ -202,3 +265,80 @@ def test_allowed_scope_threaded_to_gate() -> None: cfg = VerifierConfig(expected_run_id="r1", allowed_scope=["src/"]) out = verifier_node(_state(diff, ci), cfg) assert out["status"] == TaskStatus.PARKED.value # out-of-scope -> BLOCK -> park + + +# --------------------------------------------------------------------------- # +# Per-task run-id binding (design §4 Decision 4) +# --------------------------------------------------------------------------- # + + +def test_state_run_id_overrides_static_config_run_id() -> None: + # The gate must bind to the run id THIS task dispatched (state["run_id"]), + # not the static config constant. A CI conclusion keyed to the per-task + # run id passes even though config carries a different (stale) run id. + diff = _diff_for("src/foo.py") + state = _state( + diff, {"run_id": "task-run", "conclusion": "success", "diff_hash": _hash(diff)} + ) + state["run_id"] = "task-run" + # config.expected_run_id is a DIFFERENT, stale value — state must win. + out = verifier_node(state, VerifierConfig(expected_run_id="stale-wiring-run")) + assert out["status"] == TaskStatus.DONE.value + assert out["review_verdicts"][0]["run_id"] == "task-run" + + +def test_substituted_run_id_is_rejected() -> None: + # Anti-substitution: a CI result whose run_id != state["run_id"] is a BLOCK, + # even with a success conclusion (someone tried to graft a passing run from + # another task onto this one). + diff = _diff_for("src/foo.py") + state = _state( + diff, + {"run_id": "other-task-run", "conclusion": "success", "diff_hash": _hash(diff)}, + ) + state["run_id"] = "my-task-run" + out = verifier_node(state, VerifierConfig(expected_run_id=None)) + assert out["status"] == TaskStatus.PARKED.value + assert out["ci_results"]["gate_decision"] == "block" + assert any("run-id mismatch" in r for r in out["review_verdicts"][0]["reasons"]) + + +def test_two_concurrent_tasks_each_gate_against_own_run_id() -> None: + # Two tasks share ONE VerifierConfig but each gates against its OWN + # state["run_id"]. Task A's CI matches A's run id (PASS); task B's CI is + # keyed to A's run id (substitution) so B BLOCKs. + shared_cfg = VerifierConfig(expected_run_id=None) + + diff_a = _diff_for("src/a.py") + state_a = _state( + diff_a, {"run_id": "run-A", "conclusion": "success", "diff_hash": _hash(diff_a)} + ) + state_a["run_id"] = "run-A" + + diff_b = _diff_for("src/b.py") + # B's fetched CI is wrongly keyed to run-A (a leaked/substituted run). + state_b = _state( + diff_b, {"run_id": "run-A", "conclusion": "success", "diff_hash": _hash(diff_b)} + ) + state_b["run_id"] = "run-B" + + out_a = verifier_node(state_a, shared_cfg) + out_b = verifier_node(state_b, shared_cfg) + + assert out_a["status"] == TaskStatus.DONE.value + assert out_a["review_verdicts"][0]["run_id"] == "run-A" + assert out_b["status"] == TaskStatus.PARKED.value + assert out_b["review_verdicts"][0]["run_id"] == "run-B" + + +def test_none_run_id_at_gate_time_blocks_never_passes() -> None: + # No per-task run_id in state AND no config fallback -> BLOCK (park), never a + # vacuous pass, even when CI reports success. + diff = _diff_for("src/foo.py") + state = _state( + diff, {"run_id": "", "conclusion": "success", "diff_hash": _hash(diff)} + ) + # state has no "run_id" key; config fallback is None. + out = verifier_node(state, VerifierConfig(expected_run_id=None)) + assert out["status"] == TaskStatus.PARKED.value + assert out["ci_results"]["gate_decision"] == "block" diff --git a/agent-team/tests/test_ws3_dispatch_invoker.py b/agent-team/tests/test_ws3_dispatch_invoker.py index 164051e..22dda10 100644 --- a/agent-team/tests/test_ws3_dispatch_invoker.py +++ b/agent-team/tests/test_ws3_dispatch_invoker.py @@ -27,6 +27,26 @@ import pytest from agent_team.nodes.dispatch_invoker import DispatchNodeFactory, make_dispatch_node from agent_team.task_model import Phase, TaskStatus +_FAKE_RUN_ID = "27990718108" + + +@pytest.fixture(autouse=True) +def _stub_run_locator(monkeypatch: pytest.MonkeyPatch) -> None: + """Stub the real ``gh run list`` locator so no test shells out / sleeps. + + ``dispatch_apply_verify`` resolves the dispatched run via ``_default_run_locator`` + when no ``locator`` is injected; the real one polls ``gh`` with retry sleeps. + Replace it module-wide with a fast fake returning a fixed run id so every + ``make_dispatch_node`` call (incl. factory paths) stays hermetic. + """ + from agent_team import dispatcher as _dispatcher + + monkeypatch.setattr( + _dispatcher, + "_default_run_locator", + lambda: lambda **_kw: _FAKE_RUN_ID, + ) + # --------------------------------------------------------------------------- # # Fake BranchPusher / WorkflowDispatcher (the two injectable seams in dispatcher) @@ -76,16 +96,23 @@ _VALID_STATE: dict[str, Any] = { # --------------------------------------------------------------------------- # -def test_dispatch_node_happy_path_returns_partial_state() -> None: - """On success the node returns {} (partial state update — DONE comes from graph).""" +def test_dispatch_node_happy_path_persists_run_id() -> None: + """On success the node persists the located run identity into state. + + The dispatched run's id (plus the dispatched-at watermark and correlation + tag) flows into PipelineState so the verifier's read-only fetcher polls THIS + task's run and the pure-code gate binds its verdict to it. + """ pusher, dispatcher = _fake_seams() node = make_dispatch_node( owner="org", repo="repo", pusher=pusher, dispatcher=dispatcher ) result = node(_VALID_STATE) - # Remote implementation returns {} on success (graph topology marks DONE). assert isinstance(result, dict) + assert result["run_id"] == _FAKE_RUN_ID + assert result["dispatched_at"] # stamped + assert result["ci_correlation_tag"] == "task-abc" def test_dispatch_node_injects_owner_repo_at_factory_time() -> None: @@ -145,6 +172,60 @@ def test_dispatch_node_flattens_scope_list_to_string() -> None: assert "agent_team/" in scope_str +def test_dispatch_node_unresolved_run_id_persists_watermark_and_verify_fails_closed( + monkeypatch: pytest.MonkeyPatch, +) -> None: + """The ``if not result.run_id`` branch: dispatch fired but the run could not + be correlated. The node must NOT park (the dispatch succeeded) — it persists + the dispatched-at watermark + correlation tag with ``run_id=None`` so the + downstream verifier gate fails closed (BLOCK/park), never fabricating a pass. + """ + from agent_team import dispatcher as _dispatcher + + # Override the autouse locator stub: this run cannot be located -> None. + monkeypatch.setattr(_dispatcher, "_default_run_locator", lambda: lambda **_kw: None) + + pusher, dispatcher = _fake_seams() + node = make_dispatch_node( + owner="org", repo="repo", pusher=pusher, dispatcher=dispatcher + ) + result = node(_VALID_STATE) + + # Not parked: the dispatch itself succeeded; only correlation failed. + assert result.get("status") != TaskStatus.PARKED.value + assert result["run_id"] is None + # The watermark + correlation tag are still persisted for the CI-watch timeout + # and for audit, even though the run id is unresolved. + assert result["dispatched_at"] + assert result["ci_correlation_tag"] == "task-abc" + + # Feed the dispatch output into the verifier: a None run_id has nothing for + # the gate to bind to, so it BLOCKs and parks — fail closed, never a pass. + from agent_team.nodes.build_verify_subgraph import make_verify_node + from agent_team.nodes.verifier import VerifierConfig + + verify_state: dict[str, Any] = { + "thread_id": "task-abc", + "candidate_diff": _VALID_STATE["candidate_diff"], + "diff_hash": "deadbeef", + "ci_results": None, + "run_id": result["run_id"], # None, carried from dispatch + } + + def _pass_fetcher(_s: Any) -> Any: + # Even a 'success' CI result cannot rescue a missing run_id: the gate has + # no run identity to bind the verdict to. + return {"run_id": "whatever", "conclusion": "success", "diff_hash": "deadbeef"} + + verify_node = make_verify_node( + VerifierConfig(expected_run_id=None, allowed_scope=["agent_team/"]), + ci_result_fetcher=_pass_fetcher, + ) + out = verify_node(verify_state) + assert out["status"] == TaskStatus.PARKED.value + assert out["ci_results"]["gate_decision"] == "block" + + # --------------------------------------------------------------------------- # # make_dispatch_node — fail closed (never dispatches incomplete input) # --------------------------------------------------------------------------- # diff --git a/docs/provisioning/OPERATOR-RUNBOOK.md b/docs/provisioning/OPERATOR-RUNBOOK.md index 98f1f38..cec3cbf 100644 --- a/docs/provisioning/OPERATOR-RUNBOOK.md +++ b/docs/provisioning/OPERATOR-RUNBOOK.md @@ -308,6 +308,114 @@ follow the escalation ladder below. --- +## Incident 7 — P3 rollback (unwind the apply/verify flip) + +The P3 apply/verify CI workflow is **already LIVE** (flipped + provisioned +2026-06-22): the `agent-apply` environment with its required reviewer, the GitHub +App installation (`pull-requests: write`), the run-name/permissions edits on +`.github/workflows/agent-team-apply-verify.yml`, and `enforce_admins` branch +protection on `main`. `agent-team/scripts/p3_rollback.sh` is the inverse — it +restores those four privileged surfaces to a recorded baseline and asserts +post-restore == baseline (fail-closed: a mismatched restore aborts non-zero). + +> **Run from the operator host (the Mac), NOT the box.** No standing box token is +> used; the script relies on `gh` being authenticated on the operator host (plus +> `git` for the post-merge revert path). The R720 daemon holds no write token by +> design (P3-PHASE0-DESIGN.md, "Out of scope"). + +**When to trigger:** +- The flip must be unwound (P3 apply/verify is being backed out), OR +- A **premature flip that RAN** — an apply/verify run executed before it should + have and opened a draft PR / branch that must not exist (use the `incident` + surface, below). + +**Prerequisites:** +- A baseline JSON recorded **BEFORE** the flip (the restore target). Default path + `.security-review/p3-baseline.json` or `$P3_ROLLBACK_BASELINE`, override with + `--baseline FILE`. The script refuses to run if the baseline file is missing + (`record it BEFORE the flip`) — there is no inferred baseline. +- Target repo from the baseline's `repo` key, `--repo OWNER/NAME`, or + `$AGENT_TEAM_REPO_OWNER/$AGENT_TEAM_REPO_NAME`. If both the baseline and `--repo` + are set they must agree (fail-closed on mismatch). + +**The four privileged surfaces it restores:** +1. **`workflow`** — the apply/verify flip itself. Pre-merge: close the flip PR + + delete its branch (`--flip-pr N [--flip-branch REF]`). Post-merge: `git revert` + the flip commit + push + re-run CI (`--merged --flip-commit SHA`). LIVE: also + reverts the workflow YAML to the recorded `workflow_baseline_sha` so the live + file matches baseline. +2. **`environment`** — the `agent-apply` environment: re-PUTs the recorded required + reviewer(s) + deployment branch policy, then asserts the environment exists. +3. **`app`** — the GitHub App: **always rotates (revokes) the installation token + first** (any minted token is now suspect), then `reduce`s permissions to the + baseline or `uninstall`s the installation per `app.action`. +4. **`protection`** — branch protection on `main`: re-enables `enforce_admins` + (`include_administrators`) and asserts it is ON. **Refuses to run if the + baseline's `include_administrators` is not `true`** — a rollback must never + leave a weaker posture than baseline. + +`all` runs surfaces 1→4 in order. + +**How to run — always `--dry-run` first, then `--apply`:** + +The script is **destructive-safe by default**: with no `--apply` it only PRINTS +the plan (`[PLAN] ...` lines). `--apply` is the only thing that performs +mutations. Inspect the dry-run plan, confirm it targets the right repo + baseline, +then re-run with `--apply`. + +```bash +cd ~/Documents/repositories/orchestrator/agent-team # operator host, gh authed + +# 1. Dry-run: print the plan only, no mutations. +scripts/p3_rollback.sh all --baseline .security-review/p3-baseline.json + +# 2. Apply, once the plan looks right (pre-merge flip example): +scripts/p3_rollback.sh all --apply \ + --baseline .security-review/p3-baseline.json \ + --flip-pr 123 --flip-branch agent-team/apply/flip + +# Post-merge workflow revert instead of pre-merge close: +scripts/p3_rollback.sh workflow --apply --merged --flip-commit + +# A single surface at a time is fine too: +scripts/p3_rollback.sh protection --apply +``` + +**Premature-flip-that-RAN incident (the `incident` surface):** + +Use this when a flip executed prematurely and opened a draft PR/branch. It runs +the full incident sequence — rotate, revert, audit (read-only), restore, note: + +```bash +# Dry-run first (the audit/list steps run read-only either way): +scripts/p3_rollback.sh incident --baseline .security-review/p3-baseline.json + +# Apply, with the premature draft PR if known: +scripts/p3_rollback.sh incident --apply \ + --flip-pr --flip-branch agent-team/apply/ +``` + +It performs, in order: +1. **Rotate the App installation token first** — anything the premature run minted + is suspect. +2. **Revert the draft PR / branch the App opened** — close `--flip-pr` + delete its + branch; if no `--flip-pr` is given it lists open `agent-team/apply/*` draft PRs + to triage. +3. **Audit the Checks trail** (read-only — always runs) — recent + `agent-team-apply-verify.yml` runs with conclusions. +4. **Restore the `agent-apply` environment + branch protection** to baseline + (surfaces 2 + 4). +5. **File an incident note** under + `.security-review/incidents/p3-premature-flip-.md` recording the + repo, baseline, flip PR/branch, and actions taken. + +After any rollback, confirm the post-restore `[OK]` asserts printed (the script +exits non-zero if any assert failed), and record the action — for the `incident` +path the note is written automatically; otherwise note it on the Jira ticket per +the escalation ladder. + +--- + ## Escalation ladder (design §5 / §6.6, resolves Q3) Anything that does not clear on the first ALARM escalates — but **ALARM-only in @@ -337,3 +445,88 @@ non-CLI recoveries note it on the Jira ticket. PROVISIONING-RUNBOOK steps. - Update `project_r720_agent_team` memory if the incident revealed a durable fact (a new failure mode, a config that must change). + +## P3 box env wiring (build → dispatch → verify) + +The coordinator's environment is loaded from `EnvironmentFile=-/home/adam/secrev.env` +(declared in the unit; the leading `-` makes it optional so a missing file does +not fail the unit). The **live** P3 path — gated build → dispatch (trigger CI, +capture `run_id`) → verify (read the CI conclusion) — reads four variables from +that file at graph-build / dispatch time: + +- `AGENT_TEAM_REPO_OWNER` — dispatch/verify target owner (fixed at factory time, + never read from pipeline state, so model output cannot redirect the target). +- `AGENT_TEAM_REPO_NAME` — dispatch/verify target repo (same fail-closed binding). +- `AGENT_TEAM_BASE_BRANCH` — PR base branch; optional, defaults to `main`. +- `AGENT_TEAM_CI_READ_TOKEN` — the **read-only** CI-result token used for the + verifier's authenticated conclusion read (falls back to `GITHUB_TOKEN`). + +If owner, repo, or the CI-read token is unset, `serve()` degrades to the INERT P3 +path (one WARNING + a `#agent-team` notice) rather than crash-looping the daemon +(Phase-0 Decision 5). No `pull-requests:write` / `contents:write` token and no +`AGENT_APPLY_APP_ID` / `AGENT_APPLY_APP_PRIVATE_KEY` may live in `~/secrev.env`: +the apply path mints its write token **inside** the CI runner from Actions +secrets, so the box holds no standing write credential. That invariant is +enforced by `scripts/assert_no_write_token.py` (the A2 audit) at provisioning and +in CI — run it before any deploy. + +### Verifying the vars load + +After installing/editing `~/secrev.env` and `systemctl daemon-reload` + +`systemctl restart agent-team-coordinator.service`, confirm the P3 vars reached +the **running coordinator process**. + +> ⚠️ Do NOT use `systemctl show -p Environment` — it lists only inline +> `Environment=` directives and does **NOT** show vars loaded from +> `EnvironmentFile=` (which is how `~/secrev.env` is loaded). It comes back empty +> even when the vars are correctly loaded, so it is misleading here. + +Read the actual process environment instead (requires sudo to read another +process's `environ`): + +``` +MP=$(systemctl show agent-team-coordinator.service -p MainPID --value) +sudo tr '\0' '\n' < /proc/$MP/environ | grep -E '^AGENT_TEAM_REPO|^AGENT_TEAM_BASE' +``` + +The output should list `AGENT_TEAM_REPO_OWNER`, `AGENT_TEAM_REPO_NAME`, and +`AGENT_TEAM_BASE_BRANCH` (if set). To confirm the live P3 wiring actually bound +(not the INERT path), check the code resolver directly: + +``` +cd ~/orchestrator/agent-team && set -a && source ~/secrev.env && set +a \ + && .venv/bin/python -c "from agent_team.coordinator import _p3_env_is_configured; print(_p3_env_is_configured())" +``` + +`True` means the live build+verify path is bound; `False` means the daemon is on +the INERT P3 path (fix `~/secrev.env`, reload, restart). The CI-read token is +satisfied by `AGENT_TEAM_CI_READ_TOKEN` or the read-only `GITHUB_TOKEN` fallback; +treat any token value in process output as sensitive. The box must hold NO +`AGENT_APPLY_APP_ID` / `AGENT_APPLY_APP_PRIVATE_KEY` / write token — verify with +`python scripts/assert_no_write_token.py` (and note its scope-detection caveat in +the script header: a fine-grained token's write capability is only definitively +confirmed by a live `POST /git/refs` probe returning `403`). + +### Operator-initiated dispatch (P3 option-b) + +The box is read-only, so its in-graph DISPATCH node fail-closes/parks — it never +pushes or triggers CI. Completing a dispatch is an explicit operator step with a +**just-in-time** write token (never stored in `~/secrev.env`): + +``` +# On the box (where the task's candidate_diff lives in the ledger), with a +# WRITE-capable token provided for THIS invocation only: +cd ~/orchestrator/agent-team +GH_TOKEN= \ + .venv/bin/python run-team.py dispatch --write-back +``` + +This reads the task's `candidate_diff` + declared scope from the checkpoint, +pushes the head branch, fires the `agent-team-apply-verify` `workflow_dispatch`, +prints the located `run_id`, and (`--write-back`) writes it into the task +checkpoint so the box's VERIFY binds to that run. CI then runs guard → build-test +→ pure-code gate; the privileged `gate-and-pr` job pauses at the `agent-apply` +environment for your **required-reviewer approval** before the draft PR opens. +Alternatively pass `--diff FILE --scope FILE` to dispatch a diff without reading +the ledger. The token is consumed by `gh`/`git` for the one command and never +persisted; the box returns to read-only at rest. diff --git a/docs/provisioning/P3-LIVE-FLIP-PLAN.md b/docs/provisioning/P3-LIVE-FLIP-PLAN.md index 9f5c68e..bced8cb 100644 --- a/docs/provisioning/P3-LIVE-FLIP-PLAN.md +++ b/docs/provisioning/P3-LIVE-FLIP-PLAN.md @@ -2,8 +2,21 @@ Formal phased plan to take the agent-team Plane-2 pipeline from **clarify+plan only** to **producing reviewable draft PRs**, while keeping the always-on R720 box -read-only and the apply path zero-AWS. Status as of 2026-06-22: **NOT STARTED** -(P3 is built but inert). This plan is the input to `/sh-plan-review` before any build. +read-only and the apply path zero-AWS. + +**Status as of 2026-06-23: PARTIALLY LIVE.** The split-CI apply/verify workflow +and its provisioning (the scoped GitHub App, the `agent-apply` environment + Actions +secrets, and the dispatched runs) are **LIVE since 2026-06-22** — i.e. the +workflow-authoring + provisioning phases below (Phases 1/1b/2) are **DONE**. The +remaining work is the **box-side build → dispatch → verify integration** on +`feat/agent-team-p3-box-integration` (BUILD → DISPATCH → VERIFY wiring; +operator-initiated dispatch; run-name correlation for run_id capture; async +CI-watch resume-on-complete; fail-safe serve default) — see `docs/P3-PHASE0-DESIGN.md` +for the recorded design decisions. The two remaining **human gates** are: (C1) re-run +`/sh-security-review` + the GPT-4.1 cross-family review against the *enabled* workflow ++ the bound box-side wiring (Phase 3 / B5), and (D) deploy → smoke-test → merge +(Phase 4). This plan was the input to `/sh-plan-review` before the build began; that +loop is complete. > **Prerequisite reading:** `docs/r720-agent-team-design.md` §3.3.2 (CI-as-verifier > trust boundary, "B4"), `PROVISIONING-RUNBOOK.md` (the P3-live-flip section), @@ -13,16 +26,29 @@ read-only and the apply path zero-AWS. Status as of 2026-06-22: **NOT STARTED** ## 1. Objective & current state -**Today (inert):** the pipeline runs `INTAKE → CLARIFIER → PLANNER → REVIEW`, but -`serve` passes `build_verify_wiring=None`, the Tier-3 fixer is `--dry-run` only, and -`agent-team/ci/agent-team-apply-verify.yml` has its privileged steps disabled with -`if: ${{ false }}` and `pull-requests:write` / `environment:` commented out. So it -clarifies + plans but writes no code and opens no PR. +**Live infrastructure (since 2026-06-22):** the split-CI apply/verify workflow +(`agent-team/ci/agent-team-apply-verify.yml`) is authored, enabled, and **provisioned** +— the scoped **GitHub App** (`pull-requests:write` + minimal contents), the +`agent-apply` **GitHub Actions Environment** (Adam as required reviewer + branch +protection), and the App credentials as **Actions secrets** all exist, and dispatched +runs have executed against it. Org CI can build/test/security-review an untrusted diff +in a sandbox, a pure-code gate confirms green from authenticated Checks-API results, +and the privileged job opens a **draft PR** — all without the box ever holding a write +token. -**After P3:** the pipeline can emit a diff, have **org CI** build/test/security-review -it in an untrusted sandbox, a pure-code gate confirm green from authenticated -Checks-API results, and a scoped **GitHub App** open a **draft PR** for human review. -The box never gains a write token. +**Not yet wired (the remaining build, on `feat/agent-team-p3-box-integration`):** the +**box-side integration** that makes a task flow BUILD → DISPATCH (trigger that live CI, +capture the run_id) → VERIFY (read its conclusion) automatically. As-built, `serve` +binds the fail-safe gated P3 wiring (degrading to the INERT path if the dispatch target +/ CI-read token is unset — Phase-0 Decision 5), dispatch is **operator-initiated** +(branch push + `gh workflow run` on operator-host credentials; the box holds no standing +write token), run_id capture is **run-name / correlation-tag** based, and the CI wait is +**async resume-on-complete** via a `tick()`-driven CI-watcher (not a blocking poll). The +Tier-3 fixer remains `--dry-run` only until the Phase-4 smoke test passes. + +**After the box-side integration + the remaining human gates:** the pipeline emits a +diff end-to-end into a reviewable **draft PR** for human review. The box never gains a +write token. ## 2. Locked decisions (carried in — do not relitigate here) @@ -35,16 +61,19 @@ The box never gains a write token. ## 3. Hard gates (must clear before the flip — these block everything) -| Gate | Why | Owner | -|---|---|---| -| `/sh-plan-review` on THIS plan | adversarial plan audit before build | me → GPT-4.1 | -| `/sh-security-review` on the apply/verify CI surface | auth + untrusted-input + CI trust boundary | me | -| GPT-4.1 cross-family review on the apply/verify CI + any permission change | mandatory for the trust-boundary / permissions surface | orchestrator | -| ~~`GH_TOKEN`→`GITHUB_TOKEN` resolved~~ ✅ DONE 2026-06-22 | transport reads `GITHUB_TOKEN`; box now has a `GITHUB_TOKEN` alias of `GH_TOKEN` in `~/secrev.env` | me | +| Gate | Why | Owner | Status | +|---|---|---|---| +| `/sh-plan-review` on THIS plan | adversarial plan audit before build | me → GPT-4.1 | ✅ DONE (Round 1 + 2, 2026-06-22) | +| `/sh-security-review` on the apply/verify CI surface | auth + untrusted-input + CI trust boundary | me | ⏳ RE-RUN on the *enabled* workflow + bound box-side wiring (C1 / Phase 3 / B5) | +| GPT-4.1 cross-family review on the apply/verify CI + any permission change | mandatory for the trust-boundary / permissions surface | orchestrator | ⏳ RE-RUN on the *enabled* workflow + bound box-side wiring (C1 / Phase 3 / B5) | +| ~~`GH_TOKEN`→`GITHUB_TOKEN` resolved~~ ✅ DONE 2026-06-22 | transport reads `GITHUB_TOKEN`; box now has a `GITHUB_TOKEN` alias of `GH_TOKEN` in `~/secrev.env` | me | ✅ DONE | -The flip does NOT proceed until `/sh-security-review` AND the GPT-4.1 cross-review on -the CI surface both pass — **and these gates are re-run against the ACTUAL enabled -workflow (Phase 3), not only the inert version** (see Phase 3). +**The apply/verify CI workflow + its provisioning are already LIVE (2026-06-22).** The +two `/sh-security-review` + GPT-4.1 cross-review gates were satisfied against the +authored workflow during build; per B5 they are **RE-RUN against the ACTUAL enabled +workflow AND the bound box-side build→dispatch→verify wiring** before the box-side +integration is deployed/merged (Phase 3). The box-side integration does NOT deploy/merge +until that re-run passes — **hard stop** (see Phase 3). > **Plan-review disposition (GPT-4.1 cross-family, 2026-06-22 — REQUEST CHANGES).** Findings > folded into §4 and Phases 1/1b: expanded denylist vectors, runner-trust, concrete @@ -117,40 +146,44 @@ workflow (Phase 3), not only the inert version** (see Phase 3). ## 5. Phases -### Phase 0 — Plan review & pre-reqs 🤖/🧑 +### Phase 0 — Plan review & pre-reqs 🤖/🧑 — ✅ DONE +> **DONE.** Plan review complete; pre-reqs resolved. (Task label: **C0**.) - [x] Run `/sh-plan-review` on this doc; fold BLOCK/FIX items in. (Round 1 + Round 2 done; this doc is the result.) -- [ ] Confirm a clean revert point (git tag main; Hyper-V snapshot of sh-secrev). +- [x] Confirm a clean revert point (git tag main; Hyper-V snapshot of sh-secrev). **Snapshot retention:** keep the pre-P3 snapshot until P3 has run clean for one full cycle (Phase 4 DoD), then prune — recorded here so it is not an open-ended snapshot. - [x] Resolve `GH_TOKEN`→`GITHUB_TOKEN` (box `~/secrev.env` now has a `GITHUB_TOKEN` alias). - **Rollback:** none (no state changed). -### Phase 1 — Author/verify the split-CI apply/verify workflow 🤖 (review-gated) -- [ ] Reconcile `agent-team/ci/agent-team-apply-verify.yml` with §4. **First confirm** the +### Phase 1 — Author/verify the split-CI apply/verify workflow 🤖 (review-gated) — ✅ DONE (2026-06-22) +> **DONE.** The split-CI apply/verify workflow is authored, enabled, and provisioned. +> The deliverables below were built and shipped; they are retained for the record and +> for the Phase-3 re-run gates to verify against. (Task label: this is part of **C2**.) +- [x] Reconcile `agent-team/ci/agent-team-apply-verify.yml` with §4. **First confirm** the PR-#17 controls are present (canonicalized denylist, egress restriction, SHA-pins, empty-hash fail-closed, Checks-API consumption); only then add the new §4 items. -- [ ] **Add denylist vectors** (§4.2): submodules/`.gitmodules`, git hooks/`core.hooksPath`/`.husky`, +- [x] **Add denylist vectors** (§4.2): submodules/`.gitmodules`, git hooks/`core.hooksPath`/`.husky`, `.gitattributes` filters, lockfile postinstall/preinstall, generated/build artifacts. Add a test suite proving canonicalization resists symlink/rename/traversal. -- [ ] **Runner-trust assertion** (§4.1): test that privileged jobs cannot run on a +- [x] **Runner-trust assertion** (§4.1): test that privileged jobs cannot run on a self-hosted/user-provided runner. -- [ ] **Concretize + threat-model the diff transport** (§4.3): pick content-addressed signed +- [x] **Concretize + threat-model the diff transport** (§4.3): pick content-addressed signed artifact (shared HMAC secret) or short-lived branch-only token; add per-task nonce anti-replay; document and test it. -- [ ] **Gate-weakening detector** (§4.5): CI step that fails on a diff adding +- [x] **Gate-weakening detector** (§4.5): CI step that fails on a diff adding `noqa`/`type: ignore`/skip/xfail/excludes/`--no-verify` or editing the gate config. The detector's pattern list is **reviewed/expanded each time a new bypass vector is found** (FIX) — record the list in code with a comment pointing here, and update it + memory when a vector is added (same discipline as the denylist below). -- [ ] **PR-metadata sanitization** (§4.6) and **ledger anti-tamper** (§4.7) implemented + tested. -- [ ] `agent_team/ci_fetcher.py` (read-only Checks-API fetcher; fails closed) + +- [x] **PR-metadata sanitization** (§4.6) and **ledger anti-tamper** (§4.7) implemented + tested. +- [x] `agent_team/ci_fetcher.py` (read-only Checks-API fetcher; fails closed) + `ci_gate.py` (pure-code green). Add a mechanism for the gate to **discover the correct required check names per repo/branch** (avoid hardcoded check-name drift across repos), **with a test that exercises discovery against every intended target repo/branch** (FIX). -- [ ] **Memory/doc update when denylist vectors change** (FIX): adding a denylist vector (here +- [x] **Memory/doc update when denylist vectors change** (FIX): adding a denylist vector (here or later) updates `project_r720_agent_team` memory + the Confluence host page in the SAME change — the denied set is operational/security-critical, not tribal knowledge. -- [ ] **Deploy-before-merge enforcement — CONCRETE, CI-enforced (B1). REQUIRED Phase-1 +- [x] **Deploy-before-merge enforcement — CONCRETE, CI-enforced (B1). REQUIRED Phase-1 deliverable; the flip does not proceed without it (no manual fallback).** "Deploy" of this change = the privileged path is *proven on the live box before the workflow PR merges*. Build, in this phase: @@ -171,63 +204,92 @@ workflow (Phase 3), not only the inert version** (see Phase 3). the Phase-1 `/sh-security-review` + cross-review hard stop. The check's implementation is itself a Phase-1 build task — that it is not yet physically built is expected for a pre-build plan; what matters is it is non-optional and flip-blocking, enforced at the Phase-1 hard stop.) -- [ ] **Concrete "no write token on the box" audit (B-QUESTION → a real check):** a +- [x] **Concrete "no write token on the box" audit (B-QUESTION → a real check):** a test/script asserting the box env + coordinator config hold no `pull-requests:write` / contents-write token (grep the live env names + assert the App token is only an Actions secret), runnable on the box and in CI. Not a prose claim. -- [ ] `/sh-security-review` + GPT-4.1 cross-review on this surface. **Hard stop until both pass.** - (NOTE: these are RE-RUN on the *enabled* workflow in Phase 3 — see B5 there.) -- **Rollback:** workflow file stays inert (`if: ${{ false }}` not yet flipped); delete the file. +- [x] `/sh-security-review` + GPT-4.1 cross-review on this surface (against the authored + workflow during build). **Hard stop until both pass.** (These are RE-RUN on the *enabled* + workflow + the bound box-side wiring in Phase 3 — the C1 human gate, see B5 there. That + re-run is the genuine outstanding gate; the in-build pass is satisfied.) +- **Rollback:** the apply/verify workflow is now LIVE; rollback is no longer "delete an inert + file" — use the Phase-1b tested rollback script (revert the workflow SHA → inert, restore + the environment + branch protection from baseline). See Phase 1b / Phase 3 rollback. -### Phase 1b — Recovery for an accidentally-merged/applied privileged change 🤖/🧑 -- [ ] **Rollback as a TESTED SCRIPT covering ALL privileged surfaces (B3)** — not a one-time +### Phase 1b — Recovery for an accidentally-merged/applied privileged change 🤖/🧑 — ✅ DONE (2026-06-22) +> **DONE.** The privileged surfaces are LIVE, so their recovery tooling shipped with them. +> The tested rollback script + the runaway/stale-PR monitoring below are built and in place; +> the Phase-3 rollback exercises the script against the live state. (Task label: part of **C2**.) +- [x] **Rollback as a TESTED SCRIPT covering ALL privileged surfaces (B3)** — not a one-time manual exercise. One scripted, re-runnable rollback that, per surface, restores from a recorded baseline: (1) the apply/verify **workflow** (revert the SHA → inert), (2) the **`agent-apply` environment** (required-reviewer + protection rules), (3) the **GitHub App permissions/installation** (rotate token, reduce/uninstall), (4) **branch protection**. The script asserts the post-restore state matches the baseline. Exercised in Phase 3 rollback AND re-runnable on demand. -- [ ] **Premature-flip-that-RAN rollback (B2).** Distinct from an accidental *merge*: cover the +- [x] **Premature-flip-that-RAN rollback (B2).** Distinct from an accidental *merge*: cover the case where `if:${{ false }}` is flipped early (or the environment gate is misconfigured) and the privileged job **actually runs** — incident steps: rotate the GitHub App token immediately, close/revert any draft PR (or branch) it opened, confirm via the Checks/PR audit trail exactly what ran in the window, restore the environment + branch protection from baseline, and file the incident. This is the "it executed" path, not just "it merged." -- [ ] Add light **monitoring on draft-PR creation rate** (runaway-volume alarm) and an +- [x] Add light **monitoring on draft-PR creation rate** (runaway-volume alarm) and an **orphaned/stale draft-PR cleanup** step. -### Phase 2 — Provision the GitHub App + environment 🧑 OPERATOR (browser/admin) -- [ ] Create a dedicated **GitHub App** with **`pull-requests:write`** (+ minimal contents to +### Phase 2 — Provision the GitHub App + environment 🧑 OPERATOR (browser/admin) — ✅ DONE (2026-06-22) +> **DONE.** Provisioning is LIVE: the scoped GitHub App, the `agent-apply` environment, and +> the App-credential Actions secrets all exist, and dispatched runs have executed against +> them. (Task label: **C0** complete.) No write token landed on the box — the App token lives +> only as an Actions secret; this is enforced by `scripts/assert_no_write_token.py`. +- [x] Create a dedicated **GitHub App** with **`pull-requests:write`** (+ minimal contents to open a branch/PR); install on the org. Token lives in **CI**, never on the box. -- [ ] Create the **`agent-apply` GitHub Actions Environment** with a **required reviewer** +- [x] Create the **`agent-apply` GitHub Actions Environment** with a **required reviewer** (Adam) + branch-protection so the privileged job cannot run unreviewed. -- [ ] Store the App credentials as repo/org **Actions secrets** (not on the box). -- **Rollback:** uninstall the App; delete the environment + secrets. +- [x] Store the App credentials as repo/org **Actions secrets** (not on the box). +- **Rollback:** uninstall the App; delete the environment + secrets (via the Phase-1b script). -### Phase 3 — Bind the live wiring (still gated by the environment) 🤖 -- [ ] **PRECONDITION (B4): Phase 2 fully complete + verified before ANY flip.** Do not proceed - until the GitHub App exists with `pull-requests:write` (+ minimal contents) and is installed, - the `agent-apply` environment exists with Adam as required reviewer + branch protection, and - the App credentials are stored as Actions secrets (NOT on the box). Verify each before the - next step; the flip is blocked otherwise. -- [ ] In the workflow: uncomment `permissions: pull-requests: write` and - `environment: agent-apply`; flip the two `if: ${{ false }}` → enabled. -- [ ] **RE-RUN BOTH GATES ON THE ENABLED WORKFLOW (B5).** `/sh-security-review` + the GPT-4.1 - cross-family review are run again against the *actual enabled* `agent-team-apply-verify.yml` - (permissions live, `if:` true) and the bound `gated_build_verify_wiring` — NOT only the inert - Phase-1 version. **Hard stop until both pass on the enabled file.** (Permissions changed → - the mandatory cross-family review is independently required here against the real diff.) -- [ ] Bind `agent_team.coordinator.gated_build_verify_wiring(...)` (real diff builder + - read-only CI-result fetcher) so a leaf calls it only **after** the gate clears. -- [ ] Set the box-side apply env vars the live path reads (read-only CI-result token + - dispatch target). Confirm **no** write token lands on the box (run the Phase-1 no-write-token audit). -- [ ] **Incremental docs (B6):** update `OPERATOR-RUNBOOK.md` + memory NOW that the flip is live +### Phase 3 — Bind the box-side build→dispatch→verify wiring 🤖 — IN PROGRESS (the remaining build) +> **Workflow flip already DONE (2026-06-22):** `permissions: pull-requests: write` and +> `environment: agent-apply` are uncommented and the two `if: ${{ false }}` are enabled — the +> workflow is LIVE. The PRECONDITION (B4) below is **satisfied** (Phase 2 provisioning is +> complete + verified). The remaining Phase-3 work is the **box-side integration** built on +> `feat/agent-team-p3-box-integration` (BUILD → DISPATCH → VERIFY; operator-initiated dispatch; +> run-name correlation; async CI-watch; fail-safe serve default — see `docs/P3-PHASE0-DESIGN.md`) +> plus the C1 re-run gates and the box env wiring. **C1 (the re-run gates) and D (deploy/smoke/ +> merge, Phase 4) remain the outstanding HUMAN gates.** +- [x] **PRECONDITION (B4): Phase 2 fully complete + verified before ANY flip.** ✅ SATISFIED — + the GitHub App exists with `pull-requests:write` (+ minimal contents) and is installed, the + `agent-apply` environment exists with Adam as required reviewer + branch protection, and the + App credentials are stored as Actions secrets (NOT on the box). +- [x] In the workflow: uncomment `permissions: pull-requests: write` and + `environment: agent-apply`; flip the two `if: ${{ false }}` → enabled. ✅ DONE 2026-06-22. +- [ ] **C1 — RE-RUN BOTH GATES ON THE ENABLED WORKFLOW + BOUND WIRING (B5). 🧑 OUTSTANDING HUMAN + GATE.** `/sh-security-review` + the GPT-4.1 cross-family review are run again against the + *actual enabled* `agent-team-apply-verify.yml` (permissions live, `if:` true) **and** the + bound box-side `gated_build_verify_wiring` (BUILD → DISPATCH → VERIFY) — NOT only the inert + Phase-1 version. **Hard stop until both pass.** (Permissions are live + the dispatch/verify + wiring is new → the mandatory cross-family review is independently required here against the + real diff.) +- [ ] Bind `agent_team.coordinator.gated_build_verify_wiring(...)` (real diff builder + dispatch + with run-name correlation run_id capture + read-only CI-result fetcher) as the **fail-safe + `serve` default** (Decision 5: degrade to the INERT P3 path if the dispatch target / CI-read + token is unset, never crash-loop). The leaf path is BUILD → DISPATCH (trigger the live CI, + suspend) → [CI-watcher resumes on terminal conclusion] → VERIFY (pure-code gate). *(Built on + `feat/agent-team-p3-box-integration`.)* +- [ ] Set the box-side apply env vars the live path reads (read-only CI-result token + dispatch + target: `AGENT_TEAM_REPO_OWNER`/`_NAME`/`_BASE_BRANCH`/`_CI_READ_TOKEN`). Confirm **no** write + token lands on the box (run the no-write-token audit, `scripts/assert_no_write_token.py`). +- [ ] **Incremental docs (B6):** update `OPERATOR-RUNBOOK.md` + memory as the box-side flip lands (do not wait for Phase 6) — what the apply path can/can't do, the denylist, the rollback. + *(The P3 box env wiring is already documented in `docs/provisioning/OPERATOR-RUNBOOK.md`.)* - **Rollback:** run the Phase-1b tested rollback script (re-set `if: ${{ false }}`, re-comment `environment:`, set `build_verify_wiring=None`, restart the coordinator). **Exercise it once here** to prove it works before relying on it. -### Phase 4 — Smoke test to a first draft PR 🧑/🤖 +### Phase 4 — Deploy, smoke test to a first draft PR, merge 🧑/🤖 — D (OUTSTANDING HUMAN GATE) +> **D — the remaining HUMAN gate.** After C1 passes, deploy the box-side integration to the live +> coordinator (`sh-secrev`, via `/sh-deploy-r720`), drive the smoke test below, then merge. Gated +> on Phase 3 (box-side wiring bound + C1 re-run gates green). - [ ] Drive one trivial, in-scope task end-to-end → confirm: untrusted job builds/tests with no secrets, denylist rejects an out-of-scope diff, pure-code gate gates on real Checks results, privileged job opens a **draft PR** with required checks attached, **nothing merged**. @@ -282,8 +344,13 @@ workflow (Phase 3), not only the inert version** (see Phase 3). | Runaway PR volume | start with Tier-3 only + one finding at a time; required-reviewer environment gates each | ## 8. Definition of done -- [ ] `/sh-plan-review`, `/sh-security-review`, and GPT-4.1 cross-review on the CI surface all passed. -- [ ] Phase-4 smoke test produced a draft PR; nothing auto-merged; rollback exercised once. -- [ ] No write token on the box (verified); apply path is zero-AWS. -- [ ] Docs + Confluence + memory updated. -- [ ] Snapshot retained until P3 runs clean for one cycle, then pruned. +- [x] `/sh-plan-review` passed; `/sh-security-review` + GPT-4.1 cross-review passed on the CI + surface during build. ⏳ **C1 outstanding:** both are RE-RUN against the *enabled* workflow + + the bound box-side wiring before the box-side integration deploys/merges (Phase 3 / B5). +- [ ] **D:** Phase-4 deploy + smoke test produced a draft PR; nothing auto-merged; rollback + exercised once. +- [x] No write token on the box (verified via `scripts/assert_no_write_token.py`); apply path is + zero-AWS. *(Re-confirm after the box-side env vars are set in Phase 3.)* +- [ ] Docs + Confluence + memory updated *(this plan + `OPERATOR-RUNBOOK.md` reflect the live + infra; Confluence + memory final reconciliation is Phase 6)*. +- [x] Snapshot retained until P3 runs clean for one cycle, then pruned.