Merge pull request #56 from Sea-Haven-Industries/feat/agent-team-p3-box-integration

feat(agent-team): P3 box-side build→dispatch→verify integration
This commit is contained in:
Adam Moussa 2026-06-23 21:01:00 -04:00 • committed by GitHub
commit addf23e883
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
39 changed files with 8302 additions and 179 deletions

View file

@ -55,6 +55,18 @@
name: agent-team-apply-verify
# Run name surfaces the dispatching task's thread_id so the box-side dispatcher
# can correlate the triggered run back to its task via `gh run list --json name`
# (workflow inputs are NOT queryable; a workflow_dispatch run reports against the
# `main` ref, not the head branch — so the task_id in the run name is the
# correlation key). The `concurrency` group below already guarantees ONE in-flight
# run per task_id, so this name + the dispatched-at watermark match the run
# unambiguously even under many simultaneous task dispatches.
# P3-BOX-INTEGRATION (feat/agent-team-p3-box-integration): additive run-name only;
# no privilege/permission/trigger change. Flagged for the C1 re-run of
# /sh-security-review + GPT-4.1 cross-review on this trust-boundary workflow.
run-name: "agent-team-apply ${{ inputs.task_id }}"
# Manual / API trigger only. The trusted, separate apply path (which owns the
# GitHub App write token) invokes this with the candidate-diff
# artifact + the ledger-recorded hash + the declared scope. There is NO

View file

@ -2,7 +2,7 @@
The durable, human-gated agentic SDLC pipeline for the R720 (`sh-secrev` VM),
design: `../docs/r720-agent-team-design.md`. A task flows INTAKE → CLARIFY (the
human gate) → PLAN → REVIEW, and (opt-in, deploy-gated) → BUILD → VERIFY → draft
human gate) → PLAN → REVIEW, and (P3) → BUILD → DISPATCH → VERIFY → draft
PR. Every stage is durable and resumable (LangGraph + a SQLite checkpointer);
the human gate suspends on `interrupt()` and resumes on a real answer.
@ -11,20 +11,25 @@ INTAKE → CLARIFY (Claude, human gate) → PLAN (Claude) → REVIEW (GPT-4.1)
▲ │
└── loop-back ───┤
approve/escalate → END
(P3, opt-in + INERT until the CI gate clears):
approve → BUILD (DeepSeek) → VERIFY (ci_gate) → draft PR
(P3 box path — gated behind the C1 re-review):
approve → BUILD (DeepSeek) → DISPATCH (push branch, trigger CI,
capture run_id, suspend) → [CI-watcher resumes on terminal
conclusion] → VERIFY (ci_gate) → draft PR
```
> **Status (2026-06-18).** P1 (human gate) + P2 (planner + adversarial review
> **Status (2026-06-23).** P1 (human gate) + P2 (planner + adversarial review
> loop) + the live runtime (coordinator, Slack/GitHub/Claude-Code transports,
> intake) + P3-inert (build/verify subgraph, opt-in) + P4 (more transports +
> GitHub-issue intake) are **built, reviewed, and merged to main** (~795 tests).
> **Nothing is provisioned**: not rsync'd to the box, no live tokens, no
> systemd, no live CI. Production default runs **P2** (no builders).
> **Deploy-gated / not yet built:** the P3 *live* CI apply/verify + OIDC role
> (held behind `/sh-security-review` + the mandatory GPT-4.1 cross-review), and
> all provisioning. See the project memory `project_r720_agent_team` and §7 of
> the design for the phased rollout.
> intake) + P3 (build/dispatch/verify subgraph) + P4 (more transports +
> GitHub-issue intake) are **built, reviewed, and merged to main**.
> **The CI apply/verify workflow is LIVE + provisioned** (the `agent-apply`
> environment, the `AGENT_APPLY_APP_*` secrets, and dispatched runs all exist as
> of 2026-06-22). **The box-side integration that drives it** (run_id capture,
> the BUILD → DISPATCH → VERIFY reorder, async CI-watch, and the fail-safe bound
> `serve` default) is built on this branch but **gated behind the C1 re-review**
> — the `/sh-security-review` + mandatory GPT-4.1 cross-review of the new
> CI-trigger boundary — before it ships to the box. See the project memory
> `project_r720_agent_team` and §7 of the design (and `docs/P3-PHASE0-DESIGN.md`)
> for the phased rollout.
## Layout
@ -62,11 +67,17 @@ agent-team/
planner.py # plan (Claude)
review_loop.py + review_loop_llm.py # adversarial review (GPT-4.1 via orchestrator)
builders.py + builders_llm.py # candidate diff (DeepSeek) — INERT, proposes only
verifier.py + verifier_llm.py # ci_gate sole PASS authority; LLM = fix-proposer
build_verify_subgraph.py # P3 BUILD→VERIFY topology (opt-in)
verifier.py + verifier_llm.py # ci_gate sole PASS authority; LLM = fix-proposer;
# binds expected_run_id per-task from state["run_id"]
build_verify_subgraph.py # P3 BUILD→DISPATCH→VERIFY topology + the tick()-driven
# CI-watcher that resumes a suspended task on the
# dispatched run's terminal conclusion (or parks on timeout)
handbook.py # WS5 load_handbook_conventions (handbook seam,
# fail-safe → "" if dir missing); planner context
dispatch_invoker.py # WS3 auto-dispatch node — INERT (NOT wired live)
dispatch_invoker.py # P3 DISPATCH node — pushes the per-dispatch head branch,
# triggers CI (gh workflow run), and captures run_id +
# dispatched_at into state via a correlation-tagged poll
# of gh run list (fails closed on an unfound run)
transport/ # one adapter contract + a live impl per channel
base.py # Transport ABC + QuestionSet / NormalizedAnswer
slack_adapter.py + slack_live.py + slack_listener.py # Block Kit + Socket Mode + /new-task
@ -75,7 +86,7 @@ agent-team/
web/ # React/Vite/TypeScript status-dashboard SPA (React Flow map,
# task list, click-through task history); built to web/dist
scripts/ # deploy-r720-ws-rollout.sh — attended WS0–WS5 UPDATE of the box
ci/ # §3.3.2 split-job CI apply/verify workflow (DEPLOY-GATED)
ci/ # §3.3.2 split-job CI apply/verify workflow (LIVE since 2026-06-22)
systemd/ # agent-team-coordinator.service + agent-team-status.service
DEPLOY-R720.md # provisioning runbook (snapshot-first, rsync, tokens, demo)
tests/ # pytest, one module per source module + sim harness
@ -149,7 +160,7 @@ P1–P4 pipeline. What is **live** vs **inert** after the rollout:
| WS1 | `invoker_multi.py` (in-process GPT-4.1 / DeepSeek / Gemini via the orchestrator's `models.py`) + `api.py` (FastAPI HTTP API, bearer auth via `AGENT_TEAM_API_TOKEN`, binds `127.0.0.1:8765`, `/docs`+`/openapi` disabled, concurrency-capped) | `bind_multi_invoker()` wired in `run-team.py` `_cmd_serve` (LIVE); the **HTTP API is a separate opt-in process** (`api.serve()`), NOT started by the coordinator |
| WS5 | `nodes/handbook.py` `load_handbook_conventions` (reads `SEA_HAVEN_HANDBOOK_DIR` or `~/.sea-haven/engineering-handbook`, fail-safe → `""`); `retriever.py` `save_memory` writes to a `_box-drafts/` review queue | LIVE — the planner prompt receives the handbook via the `context_provider` seam in `run-team.py` `_build_coordinator` |
| WS2/WS0/WS4 | Slack `/new-task` slash command (AUTHZ-01 owner-allowlist gated) → `Coordinator.set_new_task_callback`; the `sea-haven-claude-plugin/` (CLAUDE.md, settings, `/delegate` `UserPromptSubmit` hook) | LIVE (`/new-task` wired in `serve`); the plugin/HTTP-API path is opt-in |
| WS3 | `nodes/dispatch_invoker.py` (auto-dispatch LangGraph node) + graph/coordinator wiring | **INERT — NOT wired live.** The `agent-apply` GitHub Environment human-approval gate is KEPT; the P3 dispatch/build-verify path stays inert pending per-task `run_id` plumbing + a CI-boundary security re-review |
| WS3 / P3 | `nodes/dispatch_invoker.py` (DISPATCH LangGraph node) + the BUILD → DISPATCH → VERIFY reorder, the `tick()`-driven CI-watcher, per-task `run_id` plumbing, and the bound `serve` default | **Built on `feat/agent-team-p3-box-integration`, gated behind the C1 re-review** before it ships to the box. The `agent-apply` GitHub Environment human-approval gate is KEPT; dispatch is **operator-initiated** (the box holds no standing write token — the branch push + `gh workflow run` use operator-host credentials). The bound P3 wiring is the new fail-safe `serve` default: a missing `AGENT_TEAM_REPO_OWNER`/`_NAME` or CI-read token degrades to the INERT P3 path (task parks + a `#agent-team` notice), never a serve-start crash |
The HTTP API endpoints: `POST /tasks` (start a task), `GET /tasks/{thread_id}`
(status), `POST /orchestrator/invoke` (one-shot model invoke). See

View file

@ -101,10 +101,23 @@ _ALLOWED_CONCLUSIONS: frozenset[str] = frozenset(
}
)
# The dedicated read-only token env var (preferred), falling back to the generic
# GITHUB_TOKEN (the idiom github_live uses). The token MUST be read-only — the
# fetcher only ever GETs; a write-scoped token here would be unnecessary blast
# radius (provisioning issues a read-only fine-grained PAT / read-only var).
# The dedicated read-only token env var (PREFERRED on the live box), falling back
# to the generic GITHUB_TOKEN (the idiom github_live uses). The token MUST be
# read-only — the fetcher only ever GETs; a write-scoped token here would be
# unnecessary blast radius (provisioning issues a read-only fine-grained PAT /
# read-only var).
#
# SEC-04 (CWE-269): the GITHUB_TOKEN fallback is a *documented, explicitly-
# narrowed convenience* for the github_live idiom — on the live box GITHUB_TOKEN
# MUST be read-only (contents:read). Provisioning issues the dedicated
# AGENT_TEAM_CI_READ_TOKEN for this read path and that name is the preferred
# source; GITHUB_TOKEN is only the fallback. Because the no-standing-write-token
# audit (scripts/assert_no_write_token.py) name-exempts GITHUB_TOKEN from the
# write-token *name* heuristic, it deliberately does NOT exempt GITHUB_TOKEN from
# the write-token *value* scan: a write-scoped PAT/installation token
# (ghp_/ghs_/...) parked in GITHUB_TOKEN on the box is still flagged there. This
# fetcher itself is unconditionally read-only (single GET, no write verbs), so
# even a mistakenly write-scoped token is never exercised for a write here.
CI_READ_TOKEN_ENV = "AGENT_TEAM_CI_READ_TOKEN"
_FALLBACK_TOKEN_ENV = "GITHUB_TOKEN"
@ -118,9 +131,13 @@ _DEFAULT_TIMEOUT_S = 15.0
def _resolve_read_token() -> str | None:
"""Resolve the read-only GitHub token at call time, or ``None``.
Prefers :data:`CI_READ_TOKEN_ENV`, falls back to ``GITHUB_TOKEN``. Returns
``None`` when neither is set so the fetcher fails closed (the caller maps a
missing token to a ``None`` result -> gate BLOCK), rather than raising.
PREFERS the dedicated read-only :data:`CI_READ_TOKEN_ENV`
(``AGENT_TEAM_CI_READ_TOKEN``); falls back to ``GITHUB_TOKEN`` only as the
documented, explicitly-narrowed convenience described above — on the live box
``GITHUB_TOKEN`` MUST be read-only (the no-standing-write-token audit
value-scans it, and this fetcher only ever issues a single read-only GET).
Returns ``None`` when neither is set so the fetcher fails closed (the caller
maps a missing token to a ``None`` result -> gate BLOCK), rather than raising.
"""
return os.environ.get(CI_READ_TOKEN_ENV) or os.environ.get(_FALLBACK_TOKEN_ENV)

View file

@ -473,7 +473,7 @@ def evaluate_ci_gate(
candidate_diff: str,
ledger_hash: str | None,
ci_result: Mapping[str, Any] | None,
expected_run_id: str,
expected_run_id: str | None,
allowed_scope: Sequence[str] | None = None,
) -> GateResult:
"""Make the deterministic pass/fail/block decision (§3.3.2 boundary #4).
@ -486,26 +486,47 @@ def evaluate_ci_gate(
read-only PAT (GitHub Checks/Actions API). The gate reads only
``run_id``, ``conclusion``, and (optionally) ``diff_hash`` from it; it
**never** reads a patch-written success file/artifact.
* ``expected_run_id`` — the run id the verifier dispatched for this exact
diff; the conclusion must be keyed to it (a stale/substituted run id is a
BLOCK).
* ``expected_run_id`` — the run id THIS task dispatched the apply/verify
workflow under (per-task, sourced from ``state["run_id"]`` at node-run
time, NOT a static wiring-time constant). The conclusion must be keyed to
it (a stale/substituted run id is a BLOCK). A ``None``/empty value means
the dispatcher captured no run id to bind to (e.g. the run-id poll fell
through) — there is nothing to anchor the verdict against, so the gate
BLOCKs (refuse-to-proceed, never a vacuous pass).
* ``allowed_scope`` — optional declared-scope prefixes for the task.
Decision order (a trust violation always wins over a CI verdict):
1. **Denylist / scope** — any violation -> :data:`GateDecision.BLOCK`.
2. **Diff-hash integrity** — hash mismatch (ledger or CI-verified) ->
1. **Missing binding** — a ``None``/empty ``expected_run_id`` ->
:data:`GateDecision.BLOCK` (no per-task run to bind the verdict to).
2. **Denylist / scope** — any violation -> :data:`GateDecision.BLOCK`.
3. **Diff-hash integrity** — hash mismatch (ledger or CI-verified) ->
``BLOCK``.
3. **Authenticated conclusion** — missing result, a ``run_id`` that does not
4. **Authenticated conclusion** — missing result, a ``run_id`` that does not
match ``expected_run_id``, or an unrecognised/ambiguous conclusion ->
``BLOCK``; a recognised failure -> :data:`GateDecision.FAIL`; ``success``
-> :data:`GateDecision.PASS`.
Returns a :class:`GateResult` with the decision and the reasons behind it.
Raises :class:`CiGateError` on structurally invalid inputs.
Raises :class:`CiGateError` on structurally invalid inputs (a non-string
``expected_run_id`` is structurally invalid; ``None``/empty is a BLOCK, not
an exception, because it is the legitimate "no run captured" runtime state).
"""
if not isinstance(expected_run_id, str) or not expected_run_id:
raise CiGateError("expected_run_id must be a non-empty string")
if expected_run_id is not None and not isinstance(expected_run_id, str):
raise CiGateError("expected_run_id must be a string or None")
# (0) Missing per-task binding: the dispatcher captured no run id for this
# task (None/empty). There is nothing to anchor the verdict to, so refuse to
# proceed rather than gating against an empty string (which a substituted
# ``ci_result`` with no/empty run_id could otherwise vacuously satisfy).
if not expected_run_id:
return GateResult(
decision=GateDecision.BLOCK,
reasons=["no per-task run_id to bind the verdict to (dispatch unresolved)"],
run_id=None,
diff_hash=ledger_hash,
ci_conclusion=None,
)
reasons: list[str] = []

View file

@ -0,0 +1,550 @@
"""CI-watcher sweep — async resume-on-CI-complete for suspended VERIFY (design §3.3.2, P3 Decision 2).
A CI apply/verify run takes ~7 minutes. The coordinator is a single durable
daemon: a multi-minute *blocking* VERIFY node would stall the ``tick()`` loop and
every other task. So the §4 Decision-2 shape is **async resume-on-CI-complete**,
NOT a blocking poll:
BUILD → DISPATCH (push branch, trigger CI, capture run_id, suspend at VERIFY)
→ [CI-watcher resumes on terminal conclusion] → VERIFY (read result, gate)
This module is that CI-watcher: a ``tick()``-driven sweep that **mirrors
:mod:`agent_team.deadline_timer`'s shape** (a pure, restart-safe maintenance pass
whose side effects are injected callables, so it is unit-testable with no network
and no live graph). For each task suspended awaiting CI it:
1. reads the task's ``run_id`` / ``dispatched_at`` (written ONLY by the trusted
dispatch node — never by an LLM/builder/verifier node, mirroring the
``ci_fetcher`` TRUST SOURCE note);
2. polls that run **read-only** via an injected poller (the production default
reuses :func:`agent_team.ci_fetcher.fetch_ci_result` — a single read-only GET;
it NEVER writes, dispatches, or applies anything); and
3. acts on the poll outcome:
* **TERMINAL** (the run concluded with a recognised conclusion) → RESUME the
suspended task. VERIFY then re-reads the now-terminal authenticated result
and the pure-code gate (:mod:`agent_team.ci_gate`) decides PASS/FAIL. The
watcher itself NEVER decides the verdict — it only un-suspends the task.
* **PENDING** (the run has no conclusion yet) → if ``dispatched_at + timeout``
has elapsed, PARK the task (a CI run that never terminates must not wait
forever); otherwise leave it suspended for a later pass.
* **ERROR** (the read-only poll raised / could not produce a result) → PARK
the task. **Fail-closed:** a broken poll parks for a human rather than
spinning or fabricating progress.
Fail-closed discipline (mirroring ``ci_fetcher`` / ``deadline_timer``):
* A task whose state carries no usable ``run_id`` or ``dispatched_at`` is
**parked immediately** (a suspended-awaiting-CI task with no run to watch is
unrecoverable by polling — it can never resume on a conclusion that has no run
id to read). This is the "None result" fail-closed case.
* Any poll exception is isolated per task (caught → PARK that one task) so one
task's poll failure never aborts the whole sweep or crashes the daemon.
* All inputs are read fresh from the durable task state each pass; the watcher
holds no in-memory-only state, so a reboot mid-sweep simply re-runs the
remaining suspended tasks on the next pass. RESUME is idempotent via the
resume worker's turn guard; PARK is idempotent (a parked task is no longer
surfaced as awaiting-CI).
Side effects (RESUME, PARK, ALARM) are injected as callables — exactly what makes
the terminal/timeout/error branches testable without a live graph, a network, or
a real coordinator.
"""
from __future__ import annotations
import logging
from collections.abc import Callable, Mapping
from dataclasses import dataclass, field
from datetime import datetime, timedelta, timezone
from enum import Enum
from typing import Any
__all__ = [
"AWAITING_CI_KEY",
"AlarmFn",
"CiPollOutcome",
"CiPollResult",
"CiWatchAction",
"CiWatchOutcome",
"CiWatchReport",
"ParkFn",
"PendingCiTask",
"PollFn",
"ResumeFn",
"awaiting_ci_run_id",
"default_ci_poller",
"run_ci_watcher",
"snapshot_awaiting_ci_run_id",
]
logger = logging.getLogger(__name__)
# The marker key the VERIFY node stamps on its ``interrupt()`` payload when it
# suspends awaiting a CI conclusion (see
# :func:`agent_team.nodes.build_verify_subgraph._await_ci`, whose payload is
# ``{"awaiting_ci": True, "run_id": <run>}``). Both the durable pending-task
# enumerator (which threads are suspended-at-VERIFY?) and the resume turn-guard
# (is this thread STILL suspended awaiting THIS run?) key off this single marker
# so they recognise a suspended-at-VERIFY thread the exact same way.
AWAITING_CI_KEY = "awaiting_ci"
# Conservative default: a CI apply/verify run that has not reached a terminal
# conclusion within this window is treated as stuck and the task parks (§4
# Decision 2 "dispatched_at + timeout elapses -> park"). A normal run is ~7 min;
# 30 min leaves generous headroom for queueing / retries before we ALARM.
DEFAULT_CI_TIMEOUT = timedelta(minutes=30)
class CiPollOutcome(Enum):
"""The classification of one read-only CI poll (the poller's verdict).
``TERMINAL`` — the run reached a recognised, authenticated conclusion (the
poll ``result`` mapping is populated); the task should RESUME so VERIFY can
gate it. ``PENDING`` — the run is still in progress (no conclusion yet); keep
waiting unless the timeout has elapsed. ``ERROR`` — the poll could not produce
a result (a fetch error / fail-closed ``None``); the task PARKS.
"""
TERMINAL = "terminal"
PENDING = "pending"
ERROR = "error"
@dataclass(frozen=True)
class CiPollResult:
"""The outcome of one read-only poll of a task's CI run.
``outcome`` is the :class:`CiPollOutcome`; ``result`` carries the
authenticated CI conclusion mapping (``run_id`` / ``conclusion`` /
``diff_hash``) when (and only when) ``outcome`` is ``TERMINAL`` — the same
DATA shape :func:`agent_team.ci_gate.evaluate_ci_gate` consumes. The watcher
never derives a verdict from it; it only decides resume-vs-wait-vs-park.
"""
outcome: CiPollOutcome
result: Mapping[str, Any] | None = None
@classmethod
def terminal(cls, result: Mapping[str, Any]) -> CiPollResult:
"""A run that reached a recognised terminal conclusion."""
return cls(outcome=CiPollOutcome.TERMINAL, result=result)
@classmethod
def pending(cls) -> CiPollResult:
"""A run still in progress (no authenticated conclusion yet)."""
return cls(outcome=CiPollOutcome.PENDING)
@classmethod
def error(cls) -> CiPollResult:
"""A poll that could not produce a result (fail-closed)."""
return cls(outcome=CiPollOutcome.ERROR)
class CiWatchAction(Enum):
"""The action the sweep actually took for one suspended-awaiting-CI task.
``RESUMED`` — the run terminated, the task was resumed so VERIFY can gate it.
``PARKED_TIMEOUT`` — the run never terminated within the timeout; the task
parked. ``PARKED_ERROR`` — the poll failed / the state was unusable; the task
parked (fail-closed). ``WAITING`` — the run is still in progress and within
the timeout; the task stays suspended for a later pass.
"""
RESUMED = "resumed"
PARKED_TIMEOUT = "parked_timeout"
PARKED_ERROR = "parked_error"
WAITING = "waiting"
@dataclass(frozen=True)
class PendingCiTask:
"""A minimal read-snapshot of one task suspended awaiting CI (§3.3.2).
Only the fields the watcher needs: identity (``thread_id``) and the trusted
dispatch watermarks (``run_id`` / ``dispatched_at``) the poll + timeout key
off. Frozen because it is a snapshot — the watcher never mutates a task
in-place; it acts through the injected resume/park callables.
"""
thread_id: str
run_id: str | None = None
dispatched_at: str | None = None
@classmethod
def from_state(cls, thread_id: str, state: Mapping[str, Any]) -> PendingCiTask:
"""Build a :class:`PendingCiTask` from a durable task ``state`` mapping."""
run_id = state.get("run_id")
dispatched_at = state.get("dispatched_at")
return cls(
thread_id=thread_id,
run_id=str(run_id) if run_id is not None and str(run_id) != "" else None,
dispatched_at=(
str(dispatched_at)
if dispatched_at is not None and str(dispatched_at) != ""
else None
),
)
@dataclass(frozen=True)
class CiWatchOutcome:
"""The result of processing one suspended-awaiting-CI task this pass."""
thread_id: str
action: CiWatchAction
run_id: str | None = None
error: str | None = None
@dataclass
class CiWatchReport:
"""Aggregate result of one CI-watcher pass (mirrors :class:`~agent_team.deadline_timer.TimerLoopReport`).
``outcomes`` is one entry per suspended task examined. The summary counters
let the coordinator decide whether to ALARM (any ``parked_error``) without
re-walking the list.
"""
outcomes: list[CiWatchOutcome] = field(default_factory=list)
@property
def examined(self) -> int:
"""Number of suspended-awaiting-CI tasks examined this pass."""
return len(self.outcomes)
@property
def resumed(self) -> int:
"""Tasks whose run terminated and were resumed."""
return sum(1 for o in self.outcomes if o.action is CiWatchAction.RESUMED)
@property
def parked_timeout(self) -> int:
"""Tasks parked because their run never terminated within the timeout."""
return sum(1 for o in self.outcomes if o.action is CiWatchAction.PARKED_TIMEOUT)
@property
def parked_error(self) -> int:
"""Tasks parked because the poll failed / the state was unusable."""
return sum(1 for o in self.outcomes if o.action is CiWatchAction.PARKED_ERROR)
@property
def waiting(self) -> int:
"""Tasks still in progress and left suspended for a later pass."""
return sum(1 for o in self.outcomes if o.action is CiWatchAction.WAITING)
@property
def parked(self) -> int:
"""Total tasks parked this pass (timeout + error)."""
return self.parked_timeout + self.parked_error
# Injected seams. Keeping these as callables means the watcher performs no
# transport / graph / network I/O of its own (testable, and faithful to the
# deadline_timer shape).
PollFn = Callable[[PendingCiTask], CiPollResult]
"""Poll one task's CI run read-only and classify it. The production default
(:func:`default_ci_poller`) reuses :func:`agent_team.ci_fetcher.fetch_ci_result`
(a single read-only GET; never a write)."""
ResumeFn = Callable[[PendingCiTask, Mapping[str, Any]], None]
"""Called once per *terminated* task to RESUME it (drive the suspended VERIFY
node forward). Receives the task snapshot + the authenticated terminal result."""
ParkFn = Callable[[PendingCiTask], None]
"""Called once per task that must PARK (timeout, poll error, or unusable
state). Parks the task and (typically) raises an ALARM."""
AlarmFn = Callable[[str], None]
"""Optional ALARM hook for the coordinator (one message per parked task)."""
def _utc_now() -> datetime:
"""Return the current UTC time (injectable via ``now`` in the loop)."""
return datetime.now(timezone.utc)
def _parse_iso(value: str | None) -> datetime | None:
"""Parse an ISO-8601 timestamp to an aware UTC datetime, or ``None``.
A missing / unparseable ``dispatched_at`` yields ``None`` so the caller fails
closed (treats the task as unusable → park) rather than crashing the sweep.
"""
if not value:
return None
try:
parsed = datetime.fromisoformat(value)
except (TypeError, ValueError):
return None
if parsed.tzinfo is None:
# The ledger always writes UTC; treat a naive stamp as UTC rather than
# raising on the aware/naive compare below.
parsed = parsed.replace(tzinfo=timezone.utc)
return parsed
def awaiting_ci_run_id(interrupt_value: Any) -> str | None:
"""Return the run id from a VERIFY *awaiting-CI* interrupt payload, or ``None``.
The VERIFY node suspends with
``{"awaiting_ci": True, "run_id": <run>}`` (see
:func:`agent_team.nodes.build_verify_subgraph._await_ci`). This recognises
THAT marker precisely: it returns the awaited ``run_id`` only when the
payload is a mapping with ``awaiting_ci`` truthy AND a non-empty string
``run_id``. Any other interrupt (a human clarify gate, a malformed payload,
a payload with no run id) yields ``None`` so callers never mistake a
different suspension for a suspended-at-VERIFY-awaiting-CI thread.
"""
if not isinstance(interrupt_value, Mapping):
return None
if not interrupt_value.get(AWAITING_CI_KEY):
return None
run_id = interrupt_value.get("run_id")
if isinstance(run_id, str) and run_id:
return run_id
return None
def snapshot_awaiting_ci_run_id(snapshot: Any) -> str | None:
"""Return the awaited run id if ``snapshot`` is suspended at VERIFY awaiting CI.
Walks the snapshot's pending interrupts (``snapshot.interrupts``) and returns
the first ``run_id`` carried by an *awaiting-CI* payload (via
:func:`awaiting_ci_run_id`). Returns ``None`` when the snapshot is not
interrupted, is interrupted on a non-CI gate (e.g. a human clarify
question), or has already advanced (resumed / parked / done — no pending
interrupts). This is the single predicate both the durable enumerator and
the resume turn-guard use to decide "still suspended at VERIFY awaiting CI".
"""
interrupts = getattr(snapshot, "interrupts", None) or ()
for item in interrupts:
value = getattr(item, "value", item)
run_id = awaiting_ci_run_id(value)
if run_id is not None:
return run_id
return None
def default_ci_poller(
*,
owner: str,
repo: str,
client: Any = None,
) -> PollFn:
"""Build the production read-only CI poller (reuses :mod:`agent_team.ci_fetcher`).
Returns a :data:`PollFn` that, given a :class:`PendingCiTask`, performs ONE
read-only GET of the task's run via
:func:`agent_team.ci_fetcher.fetch_ci_result` (closing over ``owner`` / ``repo``
and an optional injected ``client`` test double) and classifies it:
* a populated mapping (a recognised, authenticated terminal conclusion) →
:meth:`CiPollResult.terminal`;
* ``None`` (the run is still in progress — ``fetch_ci_result`` returns ``None``
for a null/unrecognised conclusion) → :meth:`CiPollResult.pending`; the
watcher keeps waiting until the timeout, so an in-progress run never blocks
the daemon and never parks prematurely;
* any exception raised by the fetch → :meth:`CiPollResult.error` (fail-closed:
a broken read-only poll parks the task).
The poller NEVER writes: ``fetch_ci_result`` issues exactly one read-only GET
and fails closed to ``None`` on every error path. It NEVER decides the
verdict — that stays with the pure-code gate after VERIFY resumes.
"""
from agent_team.ci_fetcher import fetch_ci_result
def poll(task: PendingCiTask) -> CiPollResult:
# The fetcher reads ``state["run_id"]`` / ``state["diff_hash"]``; rebuild
# the minimal state it needs from the trusted dispatch watermarks. (We do
# not have the full task state here — only what the watcher snapshotted —
# which is exactly the read-only run identity the fetcher requires.)
state = {"run_id": task.run_id}
try:
result = fetch_ci_result(state, owner=owner, repo=repo, client=client)
except Exception: # noqa: BLE001 - any fetch failure fails closed -> park
logger.warning(
"ci_watcher: read-only poll for task %s raised; failing closed",
task.thread_id,
exc_info=True,
)
return CiPollResult.error()
if isinstance(result, Mapping):
return CiPollResult.terminal(result)
# None: the run has no recognised conclusion yet (in progress). Keep
# waiting; the timeout branch parks a run that never terminates.
return CiPollResult.pending()
return poll
def run_ci_watcher(
pending_tasks: list[PendingCiTask],
*,
poll: PollFn,
on_resume: ResumeFn,
on_park: ParkFn,
timeout: timedelta = DEFAULT_CI_TIMEOUT,
now: datetime | None = None,
) -> CiWatchReport:
"""Run one CI-watcher pass over the suspended-awaiting-CI tasks (§3.3.2 Decision 2).
``pending_tasks`` is the set of tasks currently suspended at VERIFY awaiting
CI (the coordinator supplies them from the durable task store). For each:
1. **Unusable state → PARK (fail-closed).** A task with no ``run_id`` or no
parseable ``dispatched_at`` cannot be watched (there is no run to poll, or
no watermark to time out against), so it parks immediately rather than
waiting forever on a run it can never read.
2. **Poll the run read-only.** Call ``poll`` (the production default reuses
:func:`agent_team.ci_fetcher.fetch_ci_result`; never a write). A poll that
raises is isolated to this one task (→ PARK) so it cannot abort the sweep.
3. **Act on the outcome:**
* ``TERMINAL`` → ``on_resume`` (resume the task; VERIFY re-reads the
now-terminal result and the pure-code gate decides). The watcher never
decides the verdict.
* ``PENDING`` → if ``now >= dispatched_at + timeout`` → ``on_park``
(timeout); else leave suspended (``WAITING``) for a later pass.
* ``ERROR`` → ``on_park`` (fail-closed).
Side effects are isolated per task: if ``on_resume`` / ``on_park`` raises, the
failure is recorded for that task and the sweep continues with the rest of the
batch rather than aborting the whole pass.
Restart-safety: inputs are the durable task snapshots, resume is idempotent
(the resume worker's turn guard), and park is idempotent, so re-running the
pass after a crash safely processes only the still-suspended tasks.
Returns a :class:`CiWatchReport` describing what happened to each task.
"""
current = now or _utc_now()
report = CiWatchReport()
for task in pending_tasks:
outcome = _process_task(
task,
poll=poll,
on_resume=on_resume,
on_park=on_park,
timeout=timeout,
now=current,
)
report.outcomes.append(outcome)
return report
def _process_task(
task: PendingCiTask,
*,
poll: PollFn,
on_resume: ResumeFn,
on_park: ParkFn,
timeout: timedelta,
now: datetime,
) -> CiWatchOutcome:
"""Process one suspended task: poll its run, then resume / wait / park.
Isolates each task's side effect: a poll OR a resume/park callback that raises
yields a :attr:`CiWatchAction.PARKED_ERROR` outcome (best-effort: the task is
parked if it can be) instead of crashing the whole sweep.
"""
# (1) Unusable state -> fail closed (park). No run to poll / no watermark to
# time out against means polling can never recover this task.
if task.run_id is None or _parse_iso(task.dispatched_at) is None:
logger.warning(
"ci_watcher: task %s has no usable run_id/dispatched_at; parking "
"(fail-closed)",
task.thread_id,
)
return _park(task, on_park, CiWatchAction.PARKED_ERROR)
# (2) Read-only poll. A raising poll is isolated to this task (-> park).
try:
poll_result = poll(task)
except Exception as exc: # noqa: BLE001 - one task's poll failure -> park it
logger.warning(
"ci_watcher: poll for task %s raised (%s); parking (fail-closed)",
task.thread_id,
type(exc).__name__,
)
return _park(
task,
on_park,
CiWatchAction.PARKED_ERROR,
error=f"{type(exc).__name__}: {exc}",
)
# (3) Act on the classified outcome.
if poll_result.outcome is CiPollOutcome.TERMINAL:
result = poll_result.result or {}
try:
on_resume(task, result)
except Exception as exc: # noqa: BLE001 - isolate one task's resume failure
logger.exception(
"ci_watcher: resume of task %s (terminal run) raised", task.thread_id
)
return CiWatchOutcome(
thread_id=task.thread_id,
action=CiWatchAction.PARKED_ERROR,
run_id=task.run_id,
error=f"{type(exc).__name__}: {exc}",
)
return CiWatchOutcome(
thread_id=task.thread_id,
action=CiWatchAction.RESUMED,
run_id=task.run_id,
)
if poll_result.outcome is CiPollOutcome.ERROR:
# Fail-closed: a poll that could not produce a result parks the task.
return _park(task, on_park, CiWatchAction.PARKED_ERROR)
# PENDING: still in progress. Park only if the dispatch timeout has elapsed;
# otherwise leave it suspended for a later pass (the async-wait, not a block).
dispatched = _parse_iso(task.dispatched_at)
assert dispatched is not None # narrowed by the unusable-state guard above
if now >= dispatched + timeout:
logger.warning(
"ci_watcher: task %s run %s did not terminate within %s; parking (timeout)",
task.thread_id,
task.run_id,
timeout,
)
return _park(task, on_park, CiWatchAction.PARKED_TIMEOUT)
return CiWatchOutcome(
thread_id=task.thread_id,
action=CiWatchAction.WAITING,
run_id=task.run_id,
)
def _park(
task: PendingCiTask,
on_park: ParkFn,
action: CiWatchAction,
*,
error: str | None = None,
) -> CiWatchOutcome:
"""Invoke the injected park side effect, isolating a callback failure.
A ``on_park`` that raises is downgraded to a ``PARKED_ERROR`` outcome (the
sweep continues) rather than crashing the whole pass — the same per-row
side-effect isolation :mod:`agent_team.deadline_timer` uses.
"""
try:
on_park(task)
except Exception as exc: # noqa: BLE001 - isolate one task's park failure
logger.exception("ci_watcher: park of task %s raised", task.thread_id)
return CiWatchOutcome(
thread_id=task.thread_id,
action=CiWatchAction.PARKED_ERROR,
run_id=task.run_id,
error=f"{type(exc).__name__}: {exc}",
)
return CiWatchOutcome(
thread_id=task.thread_id,
action=action,
run_id=task.run_id,
error=error,
)

View file

@ -75,6 +75,7 @@ __all__ = [
"default_clarify_node_factory",
"default_dispatch_node_factory",
"default_slack_listener_factory",
"failsafe_production_p3_wiring",
"gated_build_verify_wiring",
]
@ -355,7 +356,7 @@ def gated_build_verify_wiring(
*,
owner: str,
repo: str,
expected_run_id: str,
expected_run_id: str | None = None,
allowed_scope: list[str] | None = None,
diff_builder: Any = None,
ci_client: Any = None,
@ -386,6 +387,14 @@ def gated_build_verify_wiring(
task parks). ``ci_client`` injects a test double; the real path builds a
read-only ``requests`` session at call time from the read-only token env var.
Per-task run-id binding (design §4 Decision 4): the gate binds each task's
verdict to the run id THAT TASK dispatched (persisted as ``state["run_id"]``
by the dispatch node and read by the verifier at node-run time), NOT a
static wiring-time constant. ``expected_run_id`` is therefore optional and
defaults to ``None``; when supplied it is only a static fallback for a
harness that drives the verifier without a per-task ``state["run_id"]``. A
task whose dispatch left no run id BLOCKs (never a vacuous pass).
Lazy-imported (ci_fetcher pulls the subgraph + verifier leaves) for the same
import-hygiene reason as the other factories.
"""
@ -439,6 +448,105 @@ def default_dispatch_node_factory() -> "Callable[[Any], Any]":
return make_dispatch_node(owner=owner, repo=repo, base=base)
# The one-line operator notice posted to #agent-team when the live P3 wiring
# cannot bind (missing owner/repo/CI-read token) and the daemon degrades to the
# inert P3 path. Goes through the lifecycle NOTIFY sink (not the park ALARM): an
# unconfigured box is an operational state, not a parked task.
_P3_INERT_NOTICE = (
"ℹ️ P3 build→verify/dispatch is INERT this run: "
"AGENT_TEAM_REPO_OWNER / AGENT_TEAM_REPO_NAME (and a CI-read token: "
"AGENT_TEAM_CI_READ_TOKEN or GITHUB_TOKEN) are not all set. The daemon is "
"up and tasks run through PLAN/REVIEW; with no P3 subgraph wired, the "
"review loop's 'build' route is its terminus, so a task that would advance "
"to BUILD/VERIFY instead settles at the approved-plan terminus (no build, "
"no dispatch) until the P3 env is provisioned."
)
def _p3_env_is_configured() -> bool:
"""Return True iff the live P3 wiring can bind from the environment.
Live P3 (gated build→verify + auto-dispatch) needs the dispatch target
(``AGENT_TEAM_REPO_OWNER`` / ``AGENT_TEAM_REPO_NAME`` — the same vars
:func:`default_dispatch_node_factory` requires) AND a read-only CI token for
the verifier's authenticated conclusion read (``AGENT_TEAM_CI_READ_TOKEN``,
falling back to ``GITHUB_TOKEN`` — mirrors
:func:`agent_team.ci_fetcher._resolve_read_token`). Any missing piece means
the gate could never read an authenticated pass, so we keep the whole P3
subgraph OFF rather than wire a half-configured, always-BLOCKing path.
"""
owner = os.environ.get("AGENT_TEAM_REPO_OWNER", "").strip()
repo = os.environ.get("AGENT_TEAM_REPO_NAME", "").strip()
ci_token = (
os.environ.get("AGENT_TEAM_CI_READ_TOKEN", "").strip()
or os.environ.get("GITHUB_TOKEN", "").strip()
)
return bool(owner and repo and ci_token)
def failsafe_production_p3_wiring(
*,
notify: "Callable[..., None] | None" = None,
) -> "tuple[BuildVerifyWiring | None, DispatchNodeFactory | None]":
"""Resolve the production P3 wiring fail-safe (Decision 5; serve default).
Per the Phase-0 design the bound P3 wiring is the new ``serve`` default, but
its factories are called EAGERLY at graph-build (``setup`` calls
``self._build_verify_wiring()`` / ``self._dispatch_node_wiring()``), and the
live dispatch factory RAISES when ``AGENT_TEAM_REPO_OWNER`` /
``AGENT_TEAM_REPO_NAME`` are unset. A raise there would crash-loop the
daemon at serve-start — exactly the failure mode this wrapper exists to
prevent.
So this resolver decides ONCE, up front, from the environment:
* **Configured** (:func:`_p3_env_is_configured` — owner + repo + a CI-read
token all present) → returns the LIVE pair: a
:func:`gated_build_verify_wiring` bound to the env owner/repo (so the
verifier reads the authenticated conclusion via the read-only fetcher) and
:func:`default_dispatch_node_factory` (which re-reads the same env at
build time). Tasks reaching P3 run BUILD → DISPATCH → VERIFY.
* **Unconfigured** → returns ``(None, None)`` — the INERT P3 path: no
build→verify subgraph and no dispatch are wired at all, so the review
loop's "build" route stays its terminus (END). A task that would advance
to P3 therefore settles at the approved-plan terminus rather than building
or dispatching — there is no BUILD/VERIFY node to reach and so nothing
parks. (Fail-closed in the sense that no diff is ever built, dispatched, or
passed; never a fabricated pass.) Logs exactly ONE WARNING and emits
ONE ``#agent-team`` inert-mode notice via the lifecycle ``notify`` sink
(NOT the park-ALARM path: an unprovisioned box is an operational state,
not a parked task). The notify sink is best-effort and fully guarded so a
Slack failure here never blocks serve-start.
NEVER raises: serve-start must come up either fully wired or inert, but it
must always come up.
"""
if _p3_env_is_configured():
owner = os.environ.get("AGENT_TEAM_REPO_OWNER", "").strip()
repo = os.environ.get("AGENT_TEAM_REPO_NAME", "").strip()
return (
lambda: gated_build_verify_wiring(owner=owner, repo=repo),
default_dispatch_node_factory,
)
_LOG.warning(
"P3 build→verify/dispatch wiring is INERT: AGENT_TEAM_REPO_OWNER / "
"AGENT_TEAM_REPO_NAME (and a CI-read token) are not all set. The "
"coordinator starts and runs PLAN/REVIEW; with no P3 subgraph wired, the "
"review loop's 'build' route is its terminus, so a task that would "
"advance to BUILD/VERIFY instead settles at the approved-plan terminus "
"(no build, no dispatch) until the P3 env is provisioned."
)
if notify is not None:
try:
notify(_P3_INERT_NOTICE)
except Exception: # noqa: BLE001 - an inert-notice failure must not block serve
_LOG.warning(
"inert-mode notify failed; serve still starting inert", exc_info=True
)
return None, None
class Coordinator:
"""Owns the live Plane-2 runtime: graph + resume worker + transport (§3.3).
@ -470,6 +578,10 @@ class Coordinator:
build_listener: ListenerFactory | None = None,
new_task_callback: "Callable[[str, str, str], str] | None" = None,
notify: "Callable[..., None] | None" = None,
ci_pending_provider: "Callable[[], list[Any]] | None" = None,
ci_poller: "Callable[[Any], Any] | None" = None,
ci_timeout: timedelta | None = None,
draft_pr_provider: "Callable[[], list[Any]] | None" = None,
) -> None:
self._db_path = Path(db_path)
self._transport = transport
@ -507,6 +619,26 @@ class Coordinator:
# behavior). The serve path injects a Slack poster so a task is never a
# black box: the human sees parked / needs-more-input / plan-ready.
self._notify = notify
# CI-watcher seams (P3 async resume-on-CI-complete; §4 Decision 2). Both
# OPT-IN and default None, so the CI sweep in tick() is a NO-OP unless the
# live P3 path provides them: ``ci_pending_provider`` enumerates tasks
# suspended at VERIFY awaiting CI (as
# :class:`agent_team.ci_watcher.PendingCiTask`), and ``ci_poller`` is the
# read-only poll seam (the default reuses
# :func:`agent_team.ci_fetcher.fetch_ci_result`). Left None, no CI sweep
# runs — exactly the INERT default and the unit-test path.
self._ci_pending_provider = ci_pending_provider
self._ci_poller = ci_poller
self._ci_timeout = ci_timeout
# Draft-PR runaway/stale monitor seam (P3 A4). OPT-IN and default None, so
# the draft-PR sweep in tick() is a NO-OP unless the live P3 path provides
# ``draft_pr_provider`` — an enumerator of the currently-open draft PRs (as
# :class:`agent_team.draft_pr_monitor.DraftPr`). Left None, no draft-PR
# sweep runs (the INERT default + the unit-test path). The flapping-backoff
# memory persists for the daemon's lifetime so a sustained condition is not
# re-ALARMed / re-reminded every tick.
self._draft_pr_provider = draft_pr_provider
self._draft_pr_memory: Any = None
# Built by setup().
self._graph: Any = None
@ -1003,6 +1135,21 @@ class Coordinator:
"Re-assign with more detail, or adjust the requirement to unblock.",
thread_ts=root_ts,
)
elif self._verify_pass_verdict(values) is not None:
# P3 terminal PASS: the task's CI apply/verify run reached a
# terminal PASS (the pure-code gate PASSed) and the APPROVED route
# opened the draft PR (§3.3.2 PASS terminus). Emit the POSITIVE
# lifecycle notice (NOT the park-ALARM path) with the run/PR link
# so the human can go review the draft PR. Distinguished from the
# P2 plan-ready terminus below by the verify-stage PASS verdict —
# both settle at status/phase DONE, so the verdict is the
# discriminator (a plan-ready task carries no verify verdict).
verdict = self._verify_pass_verdict(values)
self._emit(
f"🎉 {label} — CI PASSED, draft PR opened.\n"
f"{self._draft_pr_notice(values, verdict)}",
thread_ts=root_ts,
)
else:
# Plan approved + settled at the P2 terminus. PRESENT the plan
# (condensed) so the human can actually review it in-thread, not
@ -1075,15 +1222,87 @@ class Coordinator:
"requesting changes without converging)."
)
@staticmethod
def _verify_pass_verdict(values: "dict[str, Any]") -> "dict[str, Any] | None":
"""Return the verify-stage PASS verdict if the task reached the P3 PASS terminus.
The P3 build→dispatch→verify PASS terminus and the P2 plan-ready terminus
BOTH settle at status/phase ``DONE`` (see :func:`agent_team.graph.plan_node`
and :func:`agent_team.nodes.verifier.verifier_node`), so status alone can
not tell them apart. The discriminator is the verifier's verdict: only a
task that went through VERIFY and PASSed the pure-code CI gate appends a
``review_verdicts`` entry with ``stage == "verify"`` and a PASS decision
(:func:`agent_team.nodes.verifier._verdict`). A plan-ready task carries no
such verdict. Returns the most recent matching verdict (the one that
opened the draft PR) or ``None`` when the task did not terminally PASS CI.
"""
verdicts = values.get("review_verdicts") or []
if not isinstance(verdicts, list):
return None
for verdict in reversed(verdicts):
if not isinstance(verdict, dict):
continue
if (
verdict.get("stage") == "verify"
and str(verdict.get("decision") or "").lower() == "pass"
):
return verdict
return None
@staticmethod
def _draft_pr_notice(
values: "dict[str, Any]", verdict: "dict[str, Any] | None"
) -> str:
"""Body of the POSITIVE draft-PR LIFECYCLE notice (run/PR link).
Surfaces the link the human needs to go review the freshly-opened draft
PR. Two links may be present: the GitHub Actions **run** link (always
derivable from the run id the dispatcher captured + the configured
owner/repo) and the **PR** link (only once a draft-PR transport writes a
``pr_url`` into state — forward-compatible; absent today). Both are
best-effort and fail soft: a missing run id / unset owner/repo simply
drops that line rather than crashing the milestone. The run id falls back
to ``state["run_id"]`` when the verdict carries none.
"""
lines: list[str] = []
pr_url = str(values.get("pr_url") or "").strip()
if pr_url:
lines.append(f"• Draft PR: {pr_url}")
run_id = ""
if isinstance(verdict, dict):
run_id = str(verdict.get("run_id") or "").strip()
if not run_id:
run_id = str(values.get("run_id") or "").strip()
if run_id:
owner = os.environ.get("AGENT_TEAM_REPO_OWNER", "").strip()
repo = os.environ.get("AGENT_TEAM_REPO_NAME", "").strip()
if owner and repo:
lines.append(
f"• CI run: https://github.com/{owner}/{repo}/actions/runs/{run_id}"
)
else:
lines.append(f"• CI run: {run_id}")
lines.append("• Review the draft PR when you have a moment.")
return "\n".join(lines)
def tick(self) -> list[ResumeResult]:
"""One maintenance pass: deadline sweep + park policy, then drain (§3.3.1).
"""One maintenance pass: deadline sweep + CI-watch sweep, then drain (§3.3.1, §3.3.2).
Runs :func:`agent_team.responder.deadline_sweep` to flip overdue ``open``
questions to ``expired`` (the deterministic answer-vs-expiry race), then
applies the park policy to each newly-expired id (raise the ALARM hook —
§6.6 "ALARM rather than spin"; the durable ledger row is already
``expired``, which is the task's parked state for P1). Finally drains any
resume jobs that landed. Returns the drain results.
``expired``, which is the task's parked state for P1). It then runs the
CI-watcher sweep (:meth:`_ci_watch`) alongside the deadline sweep — the P3
async resume-on-CI-complete pass that resumes tasks whose CI run
terminated and parks tasks whose run timed out / failed to poll (§4
Decision 2). It then runs the draft-PR runaway/stale sweep
(:meth:`_draft_pr_monitor_sweep`) — the P3 A4 pass that ALARMs on a
draft-PR open-burst and reminds on a stale draft PR. All sweeps are
fail-soft. Finally drains any resume jobs that landed (including resumes
the CI-watch sweep enqueued). Returns the drain results.
"""
conn = connect(self._db_path)
try:
@ -1094,10 +1313,289 @@ class Coordinator:
for question_id in expired:
self._park(question_id)
# CI-watcher sweep alongside the deadline sweep (§3.3.2 Decision 2). NO-OP
# unless the live P3 seams are wired; fail-soft so a CI-watch error never
# breaks the maintenance loop.
self._ci_watch()
# Draft-PR runaway/stale sweep (P3 A4). NO-OP unless ``draft_pr_provider``
# is wired; fail-soft so a monitor error never breaks the maintenance loop.
self._draft_pr_monitor_sweep()
results = self.drain_resumes()
self._post_resume_followups(results)
return results
def _ci_watch(self) -> Any:
"""Run one CI-watcher sweep over tasks suspended awaiting CI (§3.3.2 Decision 2).
NO-OP unless BOTH CI-watcher seams are wired (``ci_pending_provider`` +
``ci_poller``) — the INERT default and the unit-test path skip it
entirely. When wired, it:
1. enumerates the tasks currently suspended at VERIFY awaiting CI (via
``ci_pending_provider``, as
:class:`agent_team.ci_watcher.PendingCiTask`); and
2. runs :func:`agent_team.ci_watcher.run_ci_watcher` with the read-only
``ci_poller`` and injected RESUME / PARK side effects: RESUME enqueues
a turn-guarded resume onto the shared queue (drained in the same tick),
so the suspended VERIFY node re-reads the now-terminal CI result and
the pure-code gate decides; PARK marks the task parked + ALARMs.
Fail-soft: any error in enumeration or the sweep is logged and swallowed
so a CI-watch failure never breaks the tick loop. Returns the
:class:`~agent_team.ci_watcher.CiWatchReport` (or ``None`` when skipped /
on error) for logging/tests.
"""
if self._ci_pending_provider is None or self._ci_poller is None:
return None
from agent_team.ci_watcher import DEFAULT_CI_TIMEOUT, run_ci_watcher
try:
pending = self._ci_pending_provider()
except Exception: # noqa: BLE001 - an enumeration failure must not break tick
_LOG.warning("ci-watch: pending-task enumeration raised", exc_info=True)
return None
if not pending:
return None
try:
report = run_ci_watcher(
list(pending),
poll=self._ci_poller,
on_resume=self._ci_resume,
on_park=self._ci_park,
timeout=self._ci_timeout or DEFAULT_CI_TIMEOUT,
)
except Exception: # noqa: BLE001 - a sweep failure must not break tick
_LOG.warning("ci-watch: sweep raised", exc_info=True)
return None
if report.parked_error:
_LOG.error(
"ci-watch: %d task(s) parked on a poll/state error (fail-closed)",
report.parked_error,
)
return report
def _ci_resume(self, task: Any, result: Any) -> None:
"""RESUME a task whose CI run terminated, via the turn-guarded worker.
The suspended VERIFY node interrupted with a payload but no pending
ledger question (it is a machine gate, not a human gate), so there is no
``question_id`` / ``answer`` to thread through the responder's
first-answer-wins flip. We drive the resume through the single-flight,
turn-guarded :meth:`agent_team.resume_worker.ResumeWorker.resume_ci`
(NOT a bare ``graph.invoke``): it takes the same per-thread lock the
deadline/human-answer resumes use and re-confirms the thread is STILL
suspended at VERIFY awaiting THIS run before invoking, so a double resume
(e.g. the same terminal run observed on two overlapping sweeps) can never
corrupt the durable state. On resume VERIFY re-runs and RE-FETCHES the
now-terminal authenticated CI result (it never trusts the resume
payload); the pure-code gate decides. Best-effort and isolated: a resume
failure for one task parks it rather than crashing the sweep (the watcher
records PARKED_ERROR).
"""
if self._graph is None or self._resume_worker is None:
raise RuntimeError("Coordinator._ci_resume called before setup()")
# The resume value is intentionally ignored by VERIFY (it re-fetches the
# authenticated conclusion), so any payload works; pass the terminal
# result for operator-log provenance. ``run_id`` is the trusted dispatch
# watermark the watcher polled, so the guard binds the resume to THIS
# task's awaited run.
self._resume_worker.resume_ci(
thread_id=task.thread_id,
run_id=task.run_id,
answer=result,
)
def _ci_park(self, task: Any) -> None:
"""PARK a task whose CI run timed out / could not be polled (fail-closed).
Flips the durable task status to PARKED via ``update_state`` (so the
terminal state is durable across a reboot) and raises the ALARM hook so
the stall is surfaced rather than silently spun on (§6.6). Best-effort:
a write failure is logged; the watcher still records the park outcome.
"""
from agent_team.task_model import Phase, TaskStatus # noqa: PLC0415
try:
self._graph.update_state(
graph_mod.thread_config(task.thread_id),
{
"status": TaskStatus.PARKED.value,
"current_phase": Phase.PARKED.value,
},
)
except Exception: # noqa: BLE001 - best-effort durable park
_LOG.warning(
"ci-watch: could not mark task %s parked", task.thread_id, exc_info=True
)
self._alarm_hook(task.thread_id)
def _draft_pr_monitor_sweep(self) -> Any:
"""Run one draft-PR runaway/stale sweep (P3 A4).
NO-OP unless ``draft_pr_provider`` is wired (the INERT default + the
unit-test path skip it entirely). When wired, it:
1. enumerates the currently-open draft PRs (via ``draft_pr_provider``, as
:class:`agent_team.draft_pr_monitor.DraftPr`); and
2. runs :func:`agent_team.draft_pr_monitor.run_draft_pr_monitor` with the
daemon-lifetime flapping-backoff memory and injected ALARM / reminder
side effects: the ALARM posts a runaway notice to ``#agent-team`` via
the lifecycle path (operator remediation = stop auto-dispatch; the
monitor never self-stops), and the reminder posts a stale-PR notice
(never auto-closes).
Fail-soft: any error in enumeration or the sweep is logged and swallowed
so a monitor failure never breaks the tick loop. Returns the
:class:`~agent_team.draft_pr_monitor.MonitorReport` (or ``None`` when
skipped / on error) for logging/tests.
"""
if self._draft_pr_provider is None:
return None
from agent_team.draft_pr_monitor import ( # noqa: PLC0415
MonitorMemory,
run_draft_pr_monitor,
)
if self._draft_pr_memory is None:
self._draft_pr_memory = MonitorMemory()
try:
draft_prs = self._draft_pr_provider()
except Exception: # noqa: BLE001 - enumeration must not break tick
_LOG.warning("draft-pr-monitor: draft-PR enumeration raised", exc_info=True)
return None
if not draft_prs:
return None
try:
report = run_draft_pr_monitor(
list(draft_prs),
on_alarm=self._draft_pr_alarm,
on_stale_reminder=self._draft_pr_stale_reminder,
memory=self._draft_pr_memory,
)
except Exception: # noqa: BLE001 - a sweep failure must not break tick
_LOG.warning("draft-pr-monitor: sweep raised", exc_info=True)
return None
if report.alarmed:
_LOG.error(
"draft-pr-monitor: RUNAWAY ALARM raised — %d draft PRs opened "
"within the window (operator remediation: stop auto-dispatch)",
report.opened_in_window,
)
return report
def _draft_pr_alarm(self, opened_in_window: int) -> None:
"""Post the draft-PR runaway ALARM to the operator channel (P3 A4).
Surfaces the burst + the operator remediation (stop auto-dispatch). The
monitor does NOT self-stop — ``systemctl stop`` is the human action; this
only makes the condition visible. Routed through the lifecycle ``_emit``
sink (never raises), so a Slack failure cannot break the tick loop.
"""
self._emit(
f"🚨 ALARM: draft-PR runaway — {opened_in_window} draft PRs opened "
"within 15 minutes (threshold > 3). Auto-dispatch may be looping. "
"Remediation: stop auto-dispatch on the box (`systemctl stop`). The "
"monitor surfaces this; it does not self-stop or auto-close."
)
def _draft_pr_stale_reminder(self, pr: Any) -> None:
"""Post a stale draft-PR reminder to the operator channel (P3 A4).
A draft PR idle > 7 days is surfaced once (per cooldown) so it is not
silently forgotten. NEVER auto-closes — closing is a human decision.
Routed through the lifecycle ``_emit`` sink (never raises).
"""
self._emit(
f"⏳ Reminder: draft PR #{pr.number} has been idle for over 7 days. "
"Review, update, or close it (the monitor never auto-closes)."
)
def _enumerate_ci_pending(self) -> list[Any]:
"""Enumerate the durable threads suspended at VERIFY awaiting CI (§3.3.2).
This is the durable :data:`ci_pending_provider` the live serve path binds:
without it the CI-watcher could never be fed real tasks (nothing else
enumerates which durable threads are parked at the VERIFY machine-gate),
so a dispatched task would suspend at VERIFY and wait FOREVER — the
async-resume gap this closes.
It walks the LangGraph SQLite checkpointer (the same DB file the ledger
uses) for the distinct ``thread_id``s it holds, then asks the compiled
graph for each thread's live snapshot. A thread is included ONLY when its
snapshot is still interrupted on the VERIFY *awaiting-CI* marker
(:func:`agent_team.ci_watcher.snapshot_awaiting_ci_run_id` returns a run
id). Threads that already advanced — resumed, parked, or done — carry no
awaiting-CI interrupt and are excluded, so the watcher never re-resumes a
task that already left the gate (the double-resume guard holds at the
enumeration boundary, before the resume worker's guard even runs).
Fail-soft: a snapshot read that raises for one thread is logged and that
thread skipped, so one unreadable thread never blanks the whole sweep.
Returns ``[]`` (never raises) before :meth:`setup` or when the
checkpointer cannot be enumerated.
"""
from agent_team.ci_watcher import ( # noqa: PLC0415
PendingCiTask,
snapshot_awaiting_ci_run_id,
)
graph = self._graph
if graph is None:
return []
checkpointer = getattr(graph, "checkpointer", None)
if checkpointer is None or not hasattr(checkpointer, "list"):
return []
# Distinct thread_ids the checkpointer holds. ``list(None)`` yields every
# checkpoint across all threads (newest-first, with repeats per thread);
# we keep insertion order and dedupe so each thread is examined once.
thread_ids: list[str] = []
seen: set[str] = set()
try:
for ckpt in checkpointer.list(None):
cfg = getattr(ckpt, "config", None) or {}
tid = (cfg.get("configurable") or {}).get("thread_id")
if isinstance(tid, str) and tid and tid not in seen:
seen.add(tid)
thread_ids.append(tid)
except Exception: # noqa: BLE001 - enumeration must never break the sweep
_LOG.warning(
"ci-watch: checkpointer enumeration raised; no CI-pending tasks "
"this pass",
exc_info=True,
)
return []
pending: list[Any] = []
for tid in thread_ids:
try:
snap = graph.get_state(graph_mod.thread_config(tid))
except Exception: # noqa: BLE001 - one bad thread must not blank the sweep
_LOG.warning(
"ci-watch: snapshot read for thread %s raised; skipping",
tid,
exc_info=True,
)
continue
if snapshot_awaiting_ci_run_id(snap) is None:
# Not suspended at VERIFY awaiting CI (resumed / parked / done /
# a human gate) -> exclude so we never re-resume it.
continue
values = getattr(snap, "values", None) or {}
pending.append(PendingCiTask.from_state(tid, values))
return pending
def _park(self, question_id: str) -> None:
"""Apply the park policy to one expired question (§6.6 ALARM, not spin).

View file

@ -28,18 +28,28 @@ from __future__ import annotations
import base64
import re
from dataclasses import dataclass
from typing import Protocol
from datetime import datetime, timezone
from typing import Any, Protocol
from agent_team.state_store import compute_content_hash
__all__ = [
"DispatchInputs",
"DispatchResult",
"DispatcherError",
"RunLocator",
"build_dispatch_inputs",
"dispatch_apply_verify",
"head_branch_for",
"select_run_id",
]
def _utc_now_iso() -> str:
"""UTC now as an ISO-8601 ``...Z`` string (matches GitHub Actions ``createdAt``)."""
return datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
WORKFLOW_FILE = "agent-team-apply-verify.yml"
# The artifact carries exactly this filename; the workflow's materialize/guard
@ -164,6 +174,49 @@ class WorkflowDispatcher(Protocol):
) -> None: ...
class RunLocator(Protocol):
"""Resolves the dispatched run's GitHub Actions ``run_id`` after a trigger.
A ``workflow_dispatch`` run reports against the dispatched ``ref`` (``main``),
not the head branch, and ``gh run list`` does not expose workflow inputs — so
the correlation key is the workflow ``run-name`` (which interpolates
``inputs.task_id``). Given the task id and a dispatched-at watermark, return
the matching run id as a string, or ``None`` if it cannot be resolved (the
caller then fails closed: no run_id -> the verifier gate BLOCKs / parks).
"""
def __call__(
self, *, owner: str, repo: str, task_id: str, since_iso: str
) -> str | None: ...
def run_name_for(task_id: str) -> str:
"""The workflow ``run-name`` for ``task_id`` (the run_id correlation key).
Mirrors ``run-name: "agent-team-apply ${{ inputs.task_id }}"`` in
``agent-team-apply-verify.yml``. The dispatcher matches a triggered run by
this exact name in :func:`_default_run_locator`.
"""
return f"agent-team-apply {task_id}"
@dataclass(frozen=True)
class DispatchResult:
"""Outcome of a dispatch: the inputs used + the located run identity.
``run_id`` is the GitHub Actions run id the verifier's read-only fetcher polls
and the pure-code gate binds its verdict to; ``None`` when the run could not
be located (the caller fails closed). ``dispatched_at`` is the UTC watermark
used to disambiguate the run from older runs; ``correlation_tag`` is the
run-name discriminator (the task id) recorded for audit.
"""
inputs: DispatchInputs
run_id: str | None
dispatched_at: str
correlation_tag: str
def dispatch_apply_verify(
*,
owner: str,
@ -174,17 +227,24 @@ def dispatch_apply_verify(
base: str = "main",
pusher: BranchPusher | None = None,
dispatcher: WorkflowDispatcher | None = None,
) -> DispatchInputs:
locator: RunLocator | None = None,
) -> DispatchResult:
"""Transport one candidate diff into org CI: push the head branch, dispatch.
The trusted apply path. Validates owner/repo, assembles the dispatch inputs,
pushes the diff as the head branch via ``pusher``, then triggers the
workflow via ``dispatcher``. Returns the :class:`DispatchInputs` used (for
the ledger/audit). ``pusher`` / ``dispatcher`` are injected so this is
testable without git/gh/network; the real defaults shell out to git/gh.
pushes the diff as the head branch via ``pusher``, triggers the workflow via
``dispatcher``, then resolves the triggered run's ``run_id`` via ``locator``
(matching the workflow ``run-name`` for ``task_id``). Returns a
:class:`DispatchResult` carrying the inputs + the located run identity so the
DISPATCH node can persist ``run_id`` into task state for the verifier.
``pusher`` / ``dispatcher`` / ``locator`` are injected so this is testable
without git/gh/network; the real defaults shell out to git/gh.
NOTE: this never runs on the box (the box has no write token, D2). CI owns
every trust decision; this only moves bytes.
every trust decision; this only moves bytes. A ``None`` ``run_id`` is NOT an
error here — it fails closed downstream (the verifier gate BLOCKs without an
authenticated run to read), never a fabricated pass.
"""
if not _OWNER_REPO_RE.match(owner or "") or not _OWNER_REPO_RE.match(repo or ""):
raise DispatcherError(f"invalid owner/repo {owner!r}/{repo!r}")
@ -194,6 +254,7 @@ def dispatch_apply_verify(
)
push = pusher if pusher is not None else _default_branch_pusher()
fire = dispatcher if dispatcher is not None else _default_workflow_dispatcher()
locate = locator if locator is not None else _default_run_locator()
# Push the head branch FIRST: the draft-PR step opens against an
# already-pushed --head, so the branch must exist before the run reaches it.
@ -204,8 +265,17 @@ def dispatch_apply_verify(
head_branch=inputs.head_branch,
diff_text=diff_text,
)
# Watermark stamped BEFORE firing so the locator never misses a run whose
# createdAt lands a moment after the trigger (the locator floors on this).
dispatched_at = _utc_now_iso()
fire(owner=owner, repo=repo, inputs=inputs.as_inputs(), ref=base)
return inputs
run_id = locate(owner=owner, repo=repo, task_id=task_id, since_iso=dispatched_at)
return DispatchResult(
inputs=inputs,
run_id=run_id,
dispatched_at=dispatched_at,
correlation_tag=task_id,
)
def _default_branch_pusher() -> BranchPusher:
@ -325,3 +395,136 @@ def _default_workflow_dispatcher() -> WorkflowDispatcher:
subprocess.run(args, check=True, capture_output=True)
return _fire
# A freshly-triggered run takes a moment to register; poll a bounded number of
# times. Total wait ~= _LOCATE_ATTEMPTS * _LOCATE_DELAY_S seconds.
_LOCATE_ATTEMPTS = 12
_LOCATE_DELAY_S = 5
# Tolerate modest host<->GitHub clock skew when flooring on the dispatched-at
# watermark: a run created a little "before" our local stamp is still ours.
_LOCATE_SKEW_S = 120
# A run's ``conclusion`` value that marks it as a stale/superseded run: the
# apply/verify workflow's per-task concurrency group cancels the prior run when a
# re-dispatch fires, so a re-dispatch of the SAME task_id leaves an older
# CANCELLED run sharing the run-name. We must NOT select it (it carries the prior
# build's verdict). Only ``cancelled`` is treated as stale-by-conclusion; a
# genuinely completed (success/failure) run is a legitimate match.
_STALE_CONCLUSIONS: frozenset[str] = frozenset({"cancelled"})
# Non-terminal statuses: a run still queued or executing is the freshly-triggered
# one we want to bind to (it has no conclusion yet).
_ACTIVE_STATUSES: frozenset[str] = frozenset(
{"queued", "in_progress", "waiting", "requested", "pending"}
)
def select_run_id(
runs: list[dict[str, Any]], *, task_id: str, floor_iso: str
) -> str | None:
"""Pick the dispatched run from ``gh run list`` rows (PURE; anti-stale).
Matches rows whose ``name`` equals :func:`run_name_for` and whose
``createdAt`` is at/after ``floor_iso``, then selects with a tightened rule so
a RAPID RE-DISPATCH of the same ``task_id`` (the build<->verify loop) never
binds to a stale/cancelled prior run:
1. Drop any matched run whose ``conclusion`` is ``cancelled`` — the workflow's
per-task concurrency group cancels the prior run on re-dispatch, so a
cancelled run sharing the run-name is the superseded one, never our run.
2. Prefer the run with the GREATEST ``createdAt`` among the still-active
(queued / in_progress / waiting) runs — the freshly-triggered run is the
newest and has no conclusion yet.
3. If none are active (e.g. a fast run already concluded by the time we poll),
fall back to the newest non-cancelled run overall.
Returns the chosen ``databaseId`` as a string, or ``None`` if nothing
matches (the caller then fails closed). ``createdAt`` ties break on the
greater ``databaseId`` (monotonic per repo → the later-created run).
"""
target_name = run_name_for(task_id)
matches = [
r
for r in runs
if r.get("name") == target_name and str(r.get("createdAt", "")) >= floor_iso
]
# Drop superseded (concurrency-cancelled) prior runs of the same task_id.
matches = [
r
for r in matches
if str(r.get("conclusion") or "").lower() not in _STALE_CONCLUSIONS
]
if not matches:
return None
def _sort_key(r: dict[str, Any]) -> tuple[str, int]:
try:
db_id = int(r.get("databaseId", 0))
except (TypeError, ValueError):
db_id = 0
return (str(r.get("createdAt", "")), db_id)
active = [
r for r in matches if str(r.get("status") or "").lower() in _ACTIVE_STATUSES
]
pool = active or matches
pool.sort(key=_sort_key, reverse=True)
return str(pool[0]["databaseId"])
def _default_run_locator() -> RunLocator:
"""Real locator: match the triggered run by ``run-name`` via ``gh run list``.
Polls ``gh run list`` (read-only) for runs of the apply/verify workflow and
delegates the anti-stale SELECTION to the pure :func:`select_run_id`: it skips
a concurrency-cancelled prior run of the same ``task_id`` and prefers the
newest still-active (queued/in_progress) run, falling back to the newest
non-cancelled run overall. This means a RAPID RE-DISPATCH of the same task
(the build<->verify loop) binds to the CURRENT run, never the superseded one.
Returns ``None`` if no matching run registers within the bounded poll window
(fails closed).
"""
def _locate(*, owner: str, repo: str, task_id: str, since_iso: str) -> str | None:
import json
import subprocess
import time
from datetime import timedelta
try:
floor_dt = datetime.strptime(since_iso, "%Y-%m-%dT%H:%M:%SZ").replace(
tzinfo=timezone.utc
) - timedelta(seconds=_LOCATE_SKEW_S)
floor_iso = floor_dt.strftime("%Y-%m-%dT%H:%M:%SZ")
except ValueError:
floor_iso = since_iso
for attempt in range(_LOCATE_ATTEMPTS):
proc = subprocess.run(
[
"gh",
"run",
"list",
"--repo",
f"{owner}/{repo}",
"--workflow",
WORKFLOW_FILE,
"--json",
"databaseId,name,createdAt,status,conclusion",
"--limit",
"50",
],
check=True,
capture_output=True,
text=True,
)
runs = json.loads(proc.stdout or "[]")
run_id = select_run_id(runs, task_id=task_id, floor_iso=floor_iso)
if run_id is not None:
return run_id
if attempt < _LOCATE_ATTEMPTS - 1:
time.sleep(_LOCATE_DELAY_S)
return None
return _locate

View file

@ -0,0 +1,422 @@
"""Draft-PR runaway monitor + stale cleanup sweep (P3 box-integration, A4).
Once the box-side BUILD → DISPATCH → VERIFY path is live (CI apply/verify is
already provisioned), a passing task ends at a **draft PR** (the §3.3.2 PASS
terminus). Two failure modes then need a maintenance sweep, mirroring the shape
of :mod:`agent_team.deadline_timer` and :mod:`agent_team.ci_watcher` (a pure,
restart-safe ``tick()``-driven pass whose side effects are injected callables, so
it is unit-testable with NO network and NO live graph):
* **RUNAWAY** — a dispatch loop (a stuck builder, a re-dispatch storm, or a
fixer that keeps re-opening) can open draft PRs far faster than a human can
review them. If **more than 3 draft PRs are opened within any 15-minute
window**, raise an ALARM to ``#agent-team`` (via the existing lifecycle / ALARM
path). The remediation is **operator-driven**: stop auto-dispatch
(``systemctl stop``). The monitor *surfaces* the condition; it never
self-restarts, self-stops, or auto-closes anything — it has no standing write
authority and must not act on infrastructure.
* **STALE** — a draft PR that has sat **idle for more than 7 days** is surfaced
with a single ``#agent-team`` reminder so it is not silently forgotten. The
monitor **never auto-closes** a stale PR — closing is a human decision; the
reminder is the only action.
Flapping backoff (the load-bearing anti-spam discipline): a RUNAWAY condition
typically persists across many ticks (the offending PRs stay inside the window
for the whole 15 minutes), and a STALE PR stays stale until a human acts. Without
backoff the sweep would re-ALARM / re-remind on *every* tick. So the monitor
carries a small :class:`MonitorMemory` (injected, durable-across-ticks): it
ALARMs at most once per ``alarm_cooldown`` and reminds about a given PR at most
once per ``stale_cooldown``. The memory is passed in (not module-global) so the
coordinator owns its lifetime and a test can assert the backoff deterministically.
Fail-soft / fail-closed discipline (mirroring the sibling sweeps):
* Every side effect (``on_alarm`` / ``on_stale_reminder``) is an injected
callable; the monitor performs no transport I/O of its own.
* A draft PR with an unparseable ``opened_at`` is ignored for the runaway count
(it cannot be placed in the window) but a present ``updated_at`` is still
considered for staleness; a PR with neither parseable timestamp is skipped
rather than crashing the sweep.
* A side effect that raises is isolated per condition (logged, recorded) so one
broken post never aborts the pass — the daemon's ``tick()`` loop keeps running.
* All inputs are read fresh each pass (the injected provider re-enumerates the
open draft PRs); the only retained state is the small backoff memory.
"""
from __future__ import annotations
import logging
from collections.abc import Callable
from dataclasses import dataclass, field
from datetime import datetime, timedelta, timezone
from enum import Enum
__all__ = [
"AlarmFn",
"DEFAULT_ALARM_COOLDOWN",
"DEFAULT_ALARM_THRESHOLD",
"DEFAULT_ALARM_WINDOW",
"DEFAULT_STALE_AFTER",
"DEFAULT_STALE_COOLDOWN",
"DraftPr",
"MonitorAction",
"MonitorMemory",
"MonitorOutcome",
"MonitorReport",
"StaleReminderFn",
"run_draft_pr_monitor",
]
logger = logging.getLogger(__name__)
# Concrete thresholds (A4). RUNAWAY: > 3 draft PRs opened within 15 minutes.
DEFAULT_ALARM_THRESHOLD = 3
DEFAULT_ALARM_WINDOW = timedelta(minutes=15)
# STALE: a draft PR idle (no update) for more than 7 days.
DEFAULT_STALE_AFTER = timedelta(days=7)
# Flapping backoff windows. A runaway condition persists across many ticks while
# the offending PRs stay inside the 15-minute window, and a stale PR stays stale
# until a human acts; these cooldowns stop the sweep re-alarming / re-reminding
# every tick. One ALARM per 15 min, one reminder per PR per day.
DEFAULT_ALARM_COOLDOWN = timedelta(minutes=15)
DEFAULT_STALE_COOLDOWN = timedelta(days=1)
class MonitorAction(Enum):
"""The action the sweep took this pass for one surfaced condition.
``ALARMED`` — a runaway burst was detected and the ALARM was posted.
``ALARM_SUPPRESSED`` — a runaway burst was detected but an ALARM was posted
recently (within ``alarm_cooldown``), so it was suppressed (flapping
backoff). ``STALE_REMINDED`` — a stale PR's reminder was posted.
``STALE_SUPPRESSED`` — a stale PR was found but it was reminded about
recently (within ``stale_cooldown``), so the reminder was suppressed.
``ERRORED`` — a side effect raised; the condition was detected but the post
failed (isolated, the sweep continues).
"""
ALARMED = "alarmed"
ALARM_SUPPRESSED = "alarm_suppressed"
STALE_REMINDED = "stale_reminded"
STALE_SUPPRESSED = "stale_suppressed"
ERRORED = "errored"
@dataclass(frozen=True)
class DraftPr:
"""A minimal read-snapshot of one open draft PR (A4).
Only the fields the monitor needs: identity (``number``) and the two
timestamps the runaway-count / staleness checks key off. ``opened_at`` is
when the draft PR was created (runaway window); ``updated_at`` is the last
activity (staleness). Frozen because it is a snapshot — the monitor never
mutates a PR; it acts only through the injected callables.
"""
number: int
opened_at: str | None = None
updated_at: str | None = None
@dataclass(frozen=True)
class MonitorOutcome:
"""The result of one surfaced condition this pass (A4)."""
action: MonitorAction
# For a runaway ALARM: the count of PRs opened inside the window. For a stale
# outcome: the PR number. ``None`` is never expected but keeps the dataclass
# total for the ERRORED path.
detail: str | None = None
error: str | None = None
@dataclass
class MonitorReport:
"""Aggregate result of one draft-PR monitor pass (mirrors the sibling sweeps).
``outcomes`` is one entry per surfaced condition (a runaway ALARM and/or each
stale reminder). The summary counters let the coordinator log/ALARM without
re-walking the list.
"""
outcomes: list[MonitorOutcome] = field(default_factory=list)
# The number of draft PRs counted as opened inside the runaway window this
# pass, for observability (not every counted PR produces an outcome — only a
# *breach* does).
opened_in_window: int = 0
@property
def alarmed(self) -> int:
"""Runaway ALARMs actually posted this pass."""
return sum(1 for o in self.outcomes if o.action is MonitorAction.ALARMED)
@property
def alarm_suppressed(self) -> int:
"""Runaway breaches detected but suppressed by the alarm cooldown."""
return sum(
1 for o in self.outcomes if o.action is MonitorAction.ALARM_SUPPRESSED
)
@property
def stale_reminded(self) -> int:
"""Stale reminders actually posted this pass."""
return sum(1 for o in self.outcomes if o.action is MonitorAction.STALE_REMINDED)
@property
def stale_suppressed(self) -> int:
"""Stale PRs found but suppressed by the per-PR reminder cooldown."""
return sum(
1 for o in self.outcomes if o.action is MonitorAction.STALE_SUPPRESSED
)
@property
def errored(self) -> int:
"""Conditions whose side effect raised (isolated; sweep continued)."""
return sum(1 for o in self.outcomes if o.action is MonitorAction.ERRORED)
@dataclass
class MonitorMemory:
"""Durable-across-ticks backoff memory (the flapping-backoff seam).
The monitor itself is otherwise pure: it reads the open draft PRs fresh each
pass. This small mutable record is the ONE piece of state that must survive
between ticks so a persistent condition does not re-ALARM / re-remind every
pass. The coordinator owns one instance for the daemon's lifetime; a test
constructs its own to assert the backoff deterministically.
``last_alarm_at`` — when a runaway ALARM was last posted (None = never).
``last_reminded_at`` — per-PR-number, when that PR was last reminded about.
Entries for PRs no longer present are pruned each pass so the map cannot grow
without bound across a long-running daemon.
"""
last_alarm_at: datetime | None = None
last_reminded_at: dict[int, datetime] = field(default_factory=dict)
# Injected side-effect seams. Keeping these as callables means the monitor
# performs no transport I/O of its own (testable, faithful to the deadline_timer
# / ci_watcher shape).
AlarmFn = Callable[[int], None]
"""Called once (per cooldown) when a runaway burst is detected. Receives the
count of draft PRs opened inside the window. The handler posts the ALARM to
``#agent-team`` and surfaces the operator remediation (stop auto-dispatch); it
does NOT self-stop."""
StaleReminderFn = Callable[[DraftPr], None]
"""Called once (per cooldown) per stale draft PR. The handler posts a
``#agent-team`` reminder. It NEVER auto-closes the PR."""
def _utc_now() -> datetime:
"""Return the current UTC time (injectable via ``now`` in the sweep)."""
return datetime.now(timezone.utc)
def _parse_iso(value: str | None) -> datetime | None:
"""Parse an ISO-8601 timestamp to an aware UTC datetime, or ``None``.
A missing / unparseable timestamp yields ``None`` so the caller fails soft
(skips that PR for the affected check) rather than crashing the sweep. A
naive stamp is treated as UTC (the ledger / GitHub API write UTC) so the
aware/naive compares below never raise.
"""
if not value:
return None
try:
parsed = datetime.fromisoformat(value)
except (TypeError, ValueError):
return None
if parsed.tzinfo is None:
parsed = parsed.replace(tzinfo=timezone.utc)
return parsed
def run_draft_pr_monitor(
draft_prs: list[DraftPr],
*,
on_alarm: AlarmFn,
on_stale_reminder: StaleReminderFn,
memory: MonitorMemory | None = None,
alarm_threshold: int = DEFAULT_ALARM_THRESHOLD,
alarm_window: timedelta = DEFAULT_ALARM_WINDOW,
stale_after: timedelta = DEFAULT_STALE_AFTER,
alarm_cooldown: timedelta = DEFAULT_ALARM_COOLDOWN,
stale_cooldown: timedelta = DEFAULT_STALE_COOLDOWN,
now: datetime | None = None,
) -> MonitorReport:
"""Run one draft-PR runaway + stale sweep (A4).
``draft_prs`` is the set of currently-open draft PRs (the coordinator
supplies them fresh each pass via an injected provider). The sweep:
1. **Runaway.** Counts draft PRs whose ``opened_at`` is within ``now -
alarm_window``. If that count is **strictly greater than**
``alarm_threshold`` (the A4 rule: > 3 within 15 min), call ``on_alarm``
once — but only if a prior ALARM is older than ``alarm_cooldown`` (flapping
backoff); otherwise record ``ALARM_SUPPRESSED``. The remediation is the
operator's (``systemctl stop``); the monitor never self-stops.
2. **Stale.** For each draft PR idle longer than ``stale_after`` (``now -
updated_at > stale_after``), call ``on_stale_reminder`` once — but only if
that PR was last reminded longer ago than ``stale_cooldown`` (per-PR
backoff); otherwise record ``STALE_SUPPRESSED``. The monitor never
auto-closes a stale PR.
Side effects are isolated per condition: an ``on_alarm`` / ``on_stale_reminder``
that raises yields an ``ERRORED`` outcome (the failure is recorded and the
backoff watermark is NOT advanced, so the next pass retries) and the sweep
continues with the rest of the batch.
Restart-safety: inputs are read fresh each pass; the only retained state is
``memory`` (the backoff watermarks). A fresh ``memory`` (e.g. after a reboot)
simply means the first post-reboot breach/stale PR is surfaced again — a
re-notification, never a missed or duplicated *action* (the monitor takes no
infrastructure action).
Returns a :class:`MonitorReport` describing what happened.
"""
current = now or _utc_now()
mem = memory if memory is not None else MonitorMemory()
report = MonitorReport()
present_numbers = {pr.number for pr in draft_prs}
# --- Runaway check ----------------------------------------------------- #
window_start = current - alarm_window
opened_in_window = 0
for pr in draft_prs:
opened_at = _parse_iso(pr.opened_at)
if opened_at is not None and opened_at >= window_start:
opened_in_window += 1
report.opened_in_window = opened_in_window
if opened_in_window > alarm_threshold:
_handle_runaway(
opened_in_window,
on_alarm=on_alarm,
mem=mem,
alarm_cooldown=alarm_cooldown,
now=current,
report=report,
)
# --- Stale check ------------------------------------------------------- #
for pr in draft_prs:
updated_at = _parse_iso(pr.updated_at)
if updated_at is None:
# No parseable last-activity stamp -> cannot assess staleness. Skip
# rather than guess (fail-soft); the runaway count is unaffected.
continue
if current - updated_at <= stale_after:
continue
_handle_stale(
pr,
on_stale_reminder=on_stale_reminder,
mem=mem,
stale_cooldown=stale_cooldown,
now=current,
report=report,
)
# Prune backoff watermarks for PRs no longer open so the memory cannot grow
# without bound over a long-running daemon.
for number in list(mem.last_reminded_at):
if number not in present_numbers:
del mem.last_reminded_at[number]
return report
def _handle_runaway(
opened_in_window: int,
*,
on_alarm: AlarmFn,
mem: MonitorMemory,
alarm_cooldown: timedelta,
now: datetime,
report: MonitorReport,
) -> None:
"""Post (or suppress) a runaway ALARM under the flapping-backoff cooldown."""
last = mem.last_alarm_at
if last is not None and now - last < alarm_cooldown:
logger.debug(
"draft-pr-monitor: runaway breach (%d in window) suppressed by cooldown",
opened_in_window,
)
report.outcomes.append(
MonitorOutcome(
action=MonitorAction.ALARM_SUPPRESSED,
detail=str(opened_in_window),
)
)
return
try:
on_alarm(opened_in_window)
except Exception as exc: # noqa: BLE001 - isolate one side-effect failure
logger.exception(
"draft-pr-monitor: runaway ALARM side effect raised; watermark not "
"advanced (will retry next pass)"
)
report.outcomes.append(
MonitorOutcome(
action=MonitorAction.ERRORED,
detail=str(opened_in_window),
error=f"{type(exc).__name__}: {exc}",
)
)
return
# Advance the watermark ONLY after a successful post so a failed post retries.
mem.last_alarm_at = now
logger.warning(
"draft-pr-monitor: RUNAWAY — %d draft PRs opened within the window; "
"ALARM raised (operator remediation: stop auto-dispatch)",
opened_in_window,
)
report.outcomes.append(
MonitorOutcome(action=MonitorAction.ALARMED, detail=str(opened_in_window))
)
def _handle_stale(
pr: DraftPr,
*,
on_stale_reminder: StaleReminderFn,
mem: MonitorMemory,
stale_cooldown: timedelta,
now: datetime,
report: MonitorReport,
) -> None:
"""Post (or suppress) a stale reminder for one PR under per-PR backoff."""
last = mem.last_reminded_at.get(pr.number)
if last is not None and now - last < stale_cooldown:
report.outcomes.append(
MonitorOutcome(action=MonitorAction.STALE_SUPPRESSED, detail=str(pr.number))
)
return
try:
on_stale_reminder(pr)
except Exception as exc: # noqa: BLE001 - isolate one side-effect failure
logger.exception(
"draft-pr-monitor: stale reminder side effect for PR #%s raised; "
"watermark not advanced (will retry next pass)",
pr.number,
)
report.outcomes.append(
MonitorOutcome(
action=MonitorAction.ERRORED,
detail=str(pr.number),
error=f"{type(exc).__name__}: {exc}",
)
)
return
mem.last_reminded_at[pr.number] = now
report.outcomes.append(
MonitorOutcome(action=MonitorAction.STALE_REMINDED, detail=str(pr.number))
)

View file

@ -135,9 +135,11 @@ PARKED_ROUTE = "parked"
# the topology that connects the injected nodes.
BUILD_NODE = "build_node"
VERIFY_NODE = "verify_node"
# P3+ dispatch vertex id: the node that carries an approved diff into org CI.
# Wired by build_graph only when the caller injects a dispatch_node callable; the
# default (None) leaves APPROVED_ROUTE → END unchanged so the graph is inert.
# P3+ dispatch vertex id: the node that triggers org CI and captures the run id.
# Wired by build_graph only when the caller injects a dispatch_node callable, in
# which case it is spliced between BUILD and VERIFY (BUILD -> DISPATCH -> VERIFY)
# so DISPATCH fires the CI run + writes ``state["run_id"]`` BEFORE VERIFY reads it
# (design §4 Decision 1). The default (None) falls back to BUILD -> VERIFY.
DISPATCH_NODE = "dispatch_node"
# P3 route ids returned by the injected ``route_after_verify`` function. They
@ -445,14 +447,23 @@ def build_graph(
route terminates at ``END`` (the approved-plan terminus), exactly as
before. Production stays P2 (clarify -> plan -> review).
* **P3:** ``build_verify`` is given (with ``review_node``) -> the review's
``"build"`` route is REPOINTED at the BUILD node, ``BUILD -> VERIFY`` is
wired, and ``route_after_verify`` maps ``{approved -> END (PR terminus),
``"build"`` route is REPOINTED at the BUILD node, the linear stage order
is wired, and ``route_after_verify`` maps ``{approved -> END (PR terminus),
build -> BUILD (bounded build<->verify loop), parked -> END (escalation)}``.
The subgraph stays INERT unless the caller binds real diff-builder / CI
seams (held for the §3.3.2 security gate); with the default INERT seams the
verifier gate has no authenticated pass and parks. Passing
``build_verify`` without ``review_node`` is a wiring error (there is no
``"build"`` route to repoint).
``dispatch_node`` (P3+, design §4 Decision 1) inserts the CI-dispatch vertex
BETWEEN BUILD and VERIFY: ``BUILD -> DISPATCH -> VERIFY``. DISPATCH triggers
the org CI run and captures the run id into ``state["run_id"]`` BEFORE VERIFY
reads it (VERIFY gates on a CI conclusion that only exists once DISPATCH has
fired the run). When ``dispatch_node`` is ``None`` (the default) the order
falls back to ``BUILD -> VERIFY`` directly. ``dispatch_node`` requires
``build_verify`` (there is no BUILD/VERIFY pair to splice it between
otherwise).
"""
clarify = live_clarify_node if live_clarify_node is not None else clarify_node
plan = live_plan_node if live_plan_node is not None else plan_node
@ -472,9 +483,10 @@ def build_graph(
if dispatch_node is not None and build_verify is None:
raise ValueError(
"build_graph: dispatch_node requires build_verify — it repoints the "
"verifier's APPROVED_ROUTE, so there is nothing to repoint without a "
"build->verify subgraph."
"build_graph: dispatch_node requires build_verify — it is spliced "
"between BUILD and VERIFY (BUILD -> DISPATCH -> VERIFY), so there is "
"no BUILD/VERIFY pair to wire it between without a build->verify "
"subgraph."
)
builder: StateGraph = StateGraph(PipelineState)
@ -520,30 +532,29 @@ def build_graph(
route_review,
{BUILD_ROUTE: BUILD_NODE, PLAN: PLAN, PARKED_ROUTE: END},
)
builder.add_edge(BUILD_NODE, VERIFY_NODE)
# Linear order is BUILD -> DISPATCH -> VERIFY so DISPATCH triggers CI
# and captures ``state["run_id"]`` BEFORE VERIFY reads it (design §4
# Decision 1: reorder; the CI conclusion VERIFY gates on only exists
# after DISPATCH fires the run). When no dispatch node is wired, fall
# back to BUILD -> VERIFY directly (the INERT default: the verifier's
# CI fetcher yields no authenticated pass and the task parks).
if dispatch_node is None:
# P3 default: APPROVED_ROUTE is the terminus (no dispatch).
builder.add_conditional_edges(
VERIFY_NODE,
route_after_verify,
{APPROVED_ROUTE: END, BUILD_ROUTE: BUILD_NODE, PARKED_ROUTE: END},
)
builder.add_edge(BUILD_NODE, VERIFY_NODE)
else:
# P3+: repoint APPROVED_ROUTE at the dispatch node, then END.
builder.add_node(
DISPATCH_NODE,
_instrument(DISPATCH_NODE, dispatch_node, transition_recorder),
)
builder.add_conditional_edges(
VERIFY_NODE,
route_after_verify,
{
APPROVED_ROUTE: DISPATCH_NODE,
BUILD_ROUTE: BUILD_NODE,
PARKED_ROUTE: END,
},
)
builder.add_edge(DISPATCH_NODE, END)
builder.add_edge(BUILD_NODE, DISPATCH_NODE)
builder.add_edge(DISPATCH_NODE, VERIFY_NODE)
# VERIFY's verdict routes to {approved -> END (PR terminus),
# build -> BUILD (bounded build<->verify loop), parked -> END
# (escalation)} regardless of whether DISPATCH is wired.
builder.add_conditional_edges(
VERIFY_NODE,
route_after_verify,
{APPROVED_ROUTE: END, BUILD_ROUTE: BUILD_NODE, PARKED_ROUTE: END},
)
if checkpointer is None:
return builder.compile()

View file

@ -2,7 +2,13 @@
This module is the **wiring topology** for the Plane-2 build -> verify stage::
... -> REVIEW (route "build") -> BUILD -> VERIFY -> {approved | build | parked}
... -> REVIEW (route "build") -> BUILD -> [DISPATCH] -> VERIFY
-> {approved | build | parked}
DISPATCH is the optional CI-trigger vertex :func:`agent_team.graph.build_graph`
splices between BUILD and VERIFY when a dispatch node is wired (design §4
Decision 1): it triggers the org CI run and captures ``state["run_id"]`` BEFORE
VERIFY reads it. With no dispatch node the order is simply ``BUILD -> VERIFY``.
It produces the BUILD node, the VERIFY node, and the
:func:`route_after_verify` conditional-edge function so the Integrate phase can
@ -189,18 +195,58 @@ def make_verify_node(
than smuggling something past the gate. The real read-only-PAT fetcher is
bound via :func:`bind_ci_result_fetcher` only after the §3.3.2 trust boundary
clears its security gate.
PER-TASK run-id binding (design §4 Decision 4): the expected run id the gate
binds the verdict to is NOT baked into ``config`` at factory time. The
dispatch node persists the run id THIS task dispatched as ``state["run_id"]``,
and :func:`agent_team.nodes.verifier.verifier_node` reads it from the state
threaded through below (``config.expected_run_id`` is only a static fallback
for harnesses with no per-task run id). So a single ``config`` shared across
tasks still gates each task against its OWN dispatched run, and a task whose
dispatch left no run id BLOCKs — never a vacuous pass.
ASYNC CI-WAIT (design §4 Decision 2): a CI apply/verify run takes ~7 minutes,
and a multi-minute *blocking* fetch here would stall the coordinator daemon's
tick loop and every other task. So when there IS a dispatched run to wait for
(``state["run_id"]`` is set) but the fetch yields no terminal result yet (the
run is still in progress → ``None``), the node SUSPENDS via
:func:`~langgraph.types.interrupt` — exactly the durable suspend/resume shape
the clarify human-gate node uses. The CI-watcher sweep
(:func:`agent_team.ci_watcher.run_ci_watcher`) polls the run read-only and
RESUMES this node once the run reaches a terminal conclusion; on resume the
node RE-FETCHES the now-terminal result and the pure-code gate decides. If the
re-fetched result is still not terminal (e.g. a spurious resume), the node
falls through to the gate, which BLOCKs/parks — fail-closed, never a vacuous
pass. With NO dispatched run (``state["run_id"]`` absent — the INERT path or a
harness), the node does NOT suspend: a ``None`` fetch flows straight to the
gate, which BLOCKs and parks exactly as before.
"""
fetcher: CiResultFetcher = (
ci_result_fetcher if ci_result_fetcher is not None else _no_ci_result
)
def node(state: PipelineState) -> PipelineState:
fetched = fetcher(state)
ci_result = fetched if isinstance(fetched, Mapping) else None
ci_result = _fetch_ci_result(fetcher, state)
# Merge the (possibly None) fetched CI result into the state the node
# reads from, WITHOUT mutating the caller's state object. The node reads
# ``ci_results``; a None result leaves the gate with nothing to pass on.
# Async CI-wait: only when a run was actually dispatched (state["run_id"]
# is set) AND it has no terminal result yet do we suspend, so the daemon
# never blocks on an in-progress run. The CI-watcher resumes us on a
# terminal conclusion; we re-fetch once after resume. The INERT/no-run
# path (no run_id) skips this and lets the gate BLOCK/park as before.
run_id = state.get("run_id")
if ci_result is None and isinstance(run_id, str) and run_id:
# Suspend + checkpoint; the CI-watcher's resume payload is the signal
# that the run terminated. We do not trust the payload's contents —
# we RE-FETCH the authenticated conclusion below so the gate reads a
# patch-independent, freshly-fetched result, never a resume-supplied
# one.
_await_ci(run_id)
ci_result = _fetch_ci_result(fetcher, state)
# Merge the (possibly still-None) fetched CI result into the state the
# node reads from, WITHOUT mutating the caller's state object. The node
# reads ``ci_results``; a None result leaves the gate with nothing to pass
# on (it BLOCKs → park), so a never-terminal run fails closed.
scoped_state: dict[str, Any] = dict(state)
scoped_state["ci_results"] = ci_result
return verifier_node(scoped_state, config)
@ -208,6 +254,40 @@ def make_verify_node(
return node
def _fetch_ci_result(
fetcher: CiResultFetcher, state: PipelineState
) -> Mapping[str, Any] | None:
"""Call the injected CI fetcher and normalise its result.
Any value other than a mapping is treated as "no terminal result" (``None``),
so a malformed fetcher fails SAFE (the gate BLOCKs) rather than smuggling a
non-mapping past the gate.
"""
fetched = fetcher(state)
return fetched if isinstance(fetched, Mapping) else None
def _await_ci(run_id: str) -> None:
"""Suspend the VERIFY node until the CI-watcher resumes it (§4 Decision 2).
Mirrors the clarify human-gate node's durable suspend: calls
:func:`langgraph.types.interrupt` so the graph checkpoints and the daemon's
tick loop is freed while a multi-minute CI run is in flight. The
:func:`agent_team.ci_watcher.run_ci_watcher` sweep polls the run read-only and
drives the resume once it terminates. The interrupt payload carries only the
``run_id`` being awaited (provenance for the watcher / operator logs); the
resume VALUE is intentionally ignored — the node re-fetches the authenticated
conclusion so the gate never reads a resume-supplied verdict.
``langgraph`` is imported lazily here to preserve this module's "no SDK at
module top" discipline (the topology stays importable where ``langgraph`` is
absent; the interrupt is only reached on the live, dispatched path).
"""
from langgraph.types import interrupt
interrupt({"awaiting_ci": True, "run_id": run_id})
def route_after_verify(state: PipelineState) -> str:
"""LangGraph conditional-edge: the next route id after the VERIFY node.
@ -275,12 +355,16 @@ def bind_ci_result_fetcher(
# A module-level note for the Integrate phase (no execution): the build->verify
# subgraph is hung off the review loop's "build" route. The conditional-edge map
# subgraph is hung off the review loop's "build" route. The linear stage order is
# BUILD -> [DISPATCH] -> VERIFY — DISPATCH (the optional CI-trigger vertex
# build_graph splices in when a dispatch node is wired) fires the CI run and
# captures ``state["run_id"]`` BEFORE VERIFY reads it (design §4 Decision 1); with
# no dispatch node the order is just BUILD -> VERIFY. The conditional-edge map
# from VERIFY should send APPROVED_ROUTE to the PR/draft terminus, BUILD_ROUTE
# back to the BUILD node (the bounded build<->verify loop, capped by
# VerifierConfig.max_build_loops), and PARKED_ROUTE to the escalation terminus.
# build_graph wires this in opt-in; this module never assembles it itself.
_INTEGRATE_NOTE = (
"review('build') -> BUILD -> VERIFY -> route_after_verify -> "
"review('build') -> BUILD -> [DISPATCH] -> VERIFY -> route_after_verify -> "
"{approved: PR terminus, build: BUILD (loop), parked: escalation}"
)

View file

@ -48,6 +48,7 @@ def make_dispatch_node(
base: str = "main",
pusher: Any = None,
dispatcher: Any = None,
locator: Any = None,
) -> Callable[[Any], Any]:
"""Build a LangGraph dispatch node for ``owner``/``repo``.
@ -86,7 +87,7 @@ def make_dispatch_node(
return _parked
try:
dispatch_apply_verify(
result = dispatch_apply_verify(
owner=owner,
repo=repo,
task_id=thread_id,
@ -95,6 +96,7 @@ def make_dispatch_node(
base=base,
pusher=pusher,
dispatcher=dispatcher,
locator=locator,
)
except DispatcherError as exc:
_LOG.error(
@ -111,13 +113,30 @@ def make_dispatch_node(
)
return _parked
if not result.run_id:
# Fired, but the run could not be correlated. Persist the watermark
# anyway and let the verifier gate fail closed (no authenticated run
# to read -> BLOCK/park) rather than fabricating progress.
_LOG.warning(
"dispatch_node: task %s dispatched but run_id unresolved; "
"downstream verify will fail closed",
thread_id,
)
_LOG.info(
"dispatch_node: dispatched task %s to %s/%s (base=%s)",
"dispatch_node: dispatched task %s to %s/%s (base=%s, run_id=%s)",
thread_id,
owner,
repo,
base,
result.run_id,
)
return {}
# Persist the located run identity so the verifier's read-only fetcher
# polls THIS task's run and the pure-code gate binds its verdict to it.
return {
"run_id": result.run_id,
"dispatched_at": result.dispatched_at,
"ci_correlation_tag": result.correlation_tag,
}
return dispatch_node

View file

@ -15,10 +15,12 @@ fix hint for the builders. The LLM is never asked whether the task passed.
State contract (mirrors :class:`agent_team.task_model.PipelineState`):
* reads ``candidate_diff``, ``diff_hash`` (ledger hash), ``ci_results``;
* reads ``candidate_diff``, ``diff_hash`` (ledger hash), ``ci_results``,
``build_loops`` (the durable per-task build<->verify count);
* writes ``status``, ``current_phase``, ``review_verdicts`` (appends the gate
verdict), and ``ci_results`` (annotated with the gate decision for
provenance).
verdict), ``ci_results`` (annotated with the gate decision for provenance),
and — on a recoverable gate FAIL — the incremented ``build_loops`` so the
budget advances across the BUILD->DISPATCH->VERIFY loop (LOGIC-RACE-01).
Transitions (the §3.3 "Stability + autonomy bounds" — the verifier must pass or
the task loops/holds, never ships):
@ -65,15 +67,29 @@ DEFAULT_MAX_BUILD_LOOPS: int = 3
class VerifierConfig:
"""Per-invocation knobs for the verifier node (§3.3, §3.3.2).
``expected_run_id`` keys the gate to the exact CI run the verifier
dispatched for this diff (a stale/substituted run id is a BLOCK).
``expected_run_id`` is a STATIC fallback only. The gate binds to the run id
THIS task dispatched, read from ``state["run_id"]`` at node-run time (the
dispatcher persists it there); the config value is consulted only when state
carries no ``run_id`` (e.g. a unit harness that drives the node directly). A
per-task ``state["run_id"]`` therefore always wins over this constant, and a
``None`` effective run id is a BLOCK (never a vacuous pass).
``allowed_scope`` is the task's declared-scope path prefixes for the
denylist boundary. ``max_build_loops`` caps build<->verify retries before
the task parks. ``build_loops`` is the loops already consumed for this task
(the coordinator threads it through state).
the task parks.
``build_loops`` is an INITIAL FALLBACK ONLY. The loop count that actually
bounds the build<->verify cycle is DURABLE per-task state read from
``state["build_loops"]`` at node-run time, because a single ``VerifierConfig``
is shared across every task at wiring time and never advances (LOGIC-RACE-01:
reading the count from this shared config meant the park guard never fired and
a perpetually-FAILing task looped BUILD->DISPATCH->VERIFY forever). The
verifier writes the incremented count back into the returned partial state so
the checkpointer carries it to the NEXT VERIFY. This field is consulted only
when state carries no ``build_loops`` (e.g. a unit harness driving the node
directly).
"""
expected_run_id: str
expected_run_id: str | None = None
allowed_scope: list[str] | None = None
max_build_loops: int = DEFAULT_MAX_BUILD_LOOPS
build_loops: int = 0
@ -148,9 +164,14 @@ def verifier_node(
Decision flow (§3.3.2 boundary #4 + §3.3 autonomy bounds):
1. Call :func:`agent_team.ci_gate.evaluate_ci_gate` with the candidate diff,
the ledger hash (``diff_hash``), the authenticated ``ci_results``, the
``expected_run_id``, and the task's ``allowed_scope``.
1. Resolve the PER-TASK expected run id from ``state["run_id"]`` (the id the
dispatch node captured for THIS task), falling back to
``config.expected_run_id`` only when state carries none. Call
:func:`agent_team.ci_gate.evaluate_ci_gate` with the candidate diff, the
ledger hash (``diff_hash``), the authenticated ``ci_results``, that
per-task expected run id, and the task's ``allowed_scope``. A ``None``
effective run id BLOCKs (anti-substitution: the verdict has nothing to
bind to), never a vacuous pass.
2. PASS -> advance to DONE (draft PR). The advisor is NOT consulted.
3. FAIL -> consult the fix-advisor for a hint, then loop back to BUILD —
unless ``build_loops`` has reached ``max_build_loops``, in which case
@ -162,6 +183,29 @@ def verifier_node(
ledger_hash = state.get("diff_hash")
ci_results = state.get("ci_results")
# Per-task binding: the gate must compare CI's run_id against the id THIS
# task dispatched (persisted by the dispatch node as ``state["run_id"]``),
# not a static wiring-time constant. State wins; the config value is only a
# fallback for harnesses that drive the node without a per-task run_id. A
# blank/None effective run id is left as None so the gate BLOCKs.
state_run_id = state.get("run_id")
expected_run_id = (
state_run_id
if isinstance(state_run_id, str) and state_run_id
else config.expected_run_id
)
# Per-task build-loop budget: the count that bounds the build<->verify cycle
# is DURABLE per-task state (``state["build_loops"]``), NOT the shared
# wiring-time config. Reading it from state is the LOGIC-RACE-01 fix: the
# shared ``VerifierConfig.build_loops`` never advanced, so the park guard
# never fired and a perpetually-FAILing task looped forever. ``config`` is
# only a fallback for a harness that drives the node without per-task state.
state_build_loops = state.get("build_loops")
current_build_loops = (
state_build_loops if isinstance(state_build_loops, int) else config.build_loops
)
if not isinstance(candidate_diff, str):
# No diff to verify is itself a refuse-to-proceed: park for a human
# rather than declaring anything. (A builder must have produced a diff
@ -169,7 +213,7 @@ def verifier_node(
gate_result = GateResult(
decision=GateDecision.BLOCK,
reasons=["no candidate_diff present in state to verify"],
run_id=config.expected_run_id,
run_id=expected_run_id,
diff_hash=ledger_hash,
ci_conclusion=None,
)
@ -178,7 +222,7 @@ def verifier_node(
candidate_diff=candidate_diff,
ledger_hash=ledger_hash,
ci_result=ci_results,
expected_run_id=config.expected_run_id,
expected_run_id=expected_run_id,
allowed_scope=config.allowed_scope,
)
@ -188,7 +232,7 @@ def verifier_node(
if gate_result.decision is GateDecision.PASS:
verdict = _verdict(
gate_result, next_phase=Phase.DONE, build_loops=config.build_loops
gate_result, next_phase=Phase.DONE, build_loops=current_build_loops
)
return {
"status": TaskStatus.DONE.value,
@ -203,7 +247,8 @@ def verifier_node(
fix_hint = _fix_advisor(gate_result, state)
if gate_result.decision is GateDecision.FAIL:
next_loops = config.build_loops + 1
# Count this failed loop against the DURABLE per-task budget read above.
next_loops = current_build_loops + 1
if next_loops >= config.max_build_loops:
# Exhausted the build-loop budget: hold rather than spin (§3.3 #6).
verdict = _verdict(
@ -220,6 +265,7 @@ def verifier_node(
"current_phase": Phase.PARKED.value,
"review_verdicts": [verdict],
"ci_results": annotated_ci,
"build_loops": next_loops,
"updated_at": verdict["at"],
}
verdict = _verdict(
@ -228,11 +274,15 @@ def verifier_node(
build_loops=next_loops,
fix_hint=fix_hint,
)
# Persist the incremented count into DURABLE state so the NEXT VERIFY
# (after BUILD->DISPATCH) sees it and the budget actually advances. The
# checkpointer carries it because it is a PipelineState key.
return {
"status": TaskStatus.ACTIVE.value,
"current_phase": Phase.BUILD.value,
"review_verdicts": [verdict],
"ci_results": annotated_ci,
"build_loops": next_loops,
"updated_at": verdict["at"],
}
@ -241,7 +291,7 @@ def verifier_node(
verdict = _verdict(
gate_result,
next_phase=Phase.PARKED,
build_loops=config.build_loops,
build_loops=current_build_loops,
fix_hint=fix_hint,
)
return {

View file

@ -288,6 +288,59 @@ class ResumeWorker:
graph_result=graph_result,
)
def resume_ci(
self,
*,
thread_id: str,
run_id: str,
answer: Any,
) -> ResumeResult:
"""Resume a CI machine-gate (VERIFY awaiting CI), single-flight + guarded.
The VERIFY node's async CI-wait is a *machine* gate, not a human one:
it suspends with ``{"awaiting_ci": True, "run_id": <run>}`` and has no
ledger ``question``/``turn`` to thread through the first-answer-wins
flip. So this path turn-guards on the CI marker instead: under the same
per-``thread_id`` lock that serialises human resumes (so a CI resume and
a redelivered CI resume for the same task never race), it re-reads the
*live* checkpoint and confirms the thread is STILL suspended at VERIFY
awaiting THIS ``run_id`` (via
:func:`agent_team.ci_watcher.snapshot_awaiting_ci_run_id`). Only then does
it invoke ``Command(resume=answer)``. If the thread already advanced
(resumed / parked / done — no awaiting-CI interrupt), or is now awaiting
a *different* run, it skips (:attr:`ResumeOutcome.STALE`) so a
double-resume can never corrupt the durable state.
Returns :attr:`ResumeOutcome.RESUMED` when the resume applied, else
:attr:`ResumeOutcome.STALE` (there is no ledger question to supersede on
this machine gate). ``question_id`` is reported as ``""`` since none
exists.
"""
from agent_team.ci_watcher import snapshot_awaiting_ci_run_id
lock = self._lock_for(thread_id)
with lock:
config = _thread_config(thread_id)
snapshot = self._graph.get_state(config)
awaited = snapshot_awaiting_ci_run_id(snapshot)
if awaited != run_id:
# Already advanced past this wait (resumed/parked/done) or now
# awaiting a different run: skip rather than double-apply.
return ResumeResult(
outcome=ResumeOutcome.STALE,
thread_id=thread_id,
question_id="",
turn=0,
)
graph_result = self._graph.invoke(build_resume_command(answer), config)
return ResumeResult(
outcome=ResumeOutcome.RESUMED,
thread_id=thread_id,
question_id="",
turn=0,
graph_result=graph_result,
)
def recover_pending_resumes(self) -> list[ResumeResult]:
"""Restart sweep: re-enqueue resumes for durable ``answered`` rows.

View file

@ -98,6 +98,22 @@ class TaskRecord:
candidate_diff: str | None = None
diff_hash: str | None = None
ci_results: dict[str, Any] | None = None
# P3 box-side build->dispatch->verify plumbing (§3.3.2). Set by the DISPATCH
# node when it triggers the apply/verify CI run: ``run_id`` is the GitHub
# Actions run id the verifier's read-only fetcher polls + the pure-code gate
# binds its verdict to; ``ci_correlation_tag`` is the per-dispatch nonce
# carried as a workflow input so the run_id poll matches THIS task's exact
# run (anti-race / anti-replay); ``dispatched_at`` bounds the CI-watch
# timeout. All None until a task reaches DISPATCH on the live P3 path.
run_id: str | None = None
ci_correlation_tag: str | None = None
dispatched_at: str | None = None
# Count of build<->verify loops already consumed for THIS task (§3.3 #6).
# Durable per-task state (NOT the shared wiring-time VerifierConfig): the
# verifier reads it from state, increments on a recoverable gate FAIL, and
# parks once it reaches VerifierConfig.max_build_loops so a perpetually-
# failing task can never loop BUILD->DISPATCH->VERIFY forever (LOGIC-RACE-01).
build_loops: int = 0
transport: str = ""
created_at: str | None = None
updated_at: str | None = None
@ -132,6 +148,20 @@ class PipelineState(TypedDict, total=False):
candidate_diff: str | None
diff_hash: str | None
ci_results: dict[str, Any] | None
# P3 box-side build->dispatch->verify plumbing (mirrors TaskRecord). ``run_id``
# is the dispatched apply/verify Actions run id the verifier fetches + the
# gate binds to; ``ci_correlation_tag`` is the per-dispatch nonce carried as a
# workflow input so the poll matches THIS task's run; ``dispatched_at`` bounds
# the CI-watch timeout.
run_id: str | None
ci_correlation_tag: str | None
dispatched_at: str | None
# Build<->verify loops already consumed for THIS task (mirrors TaskRecord).
# The verifier threads it through DURABLE state — increments on a recoverable
# gate FAIL, parks at VerifierConfig.max_build_loops — so the build-loop
# budget is real (LOGIC-RACE-01: it was previously read from the shared
# wiring-time config and never advanced).
build_loops: int
transport: str
created_at: str | None
updated_at: str | None
@ -164,6 +194,10 @@ def task_from_dict(data: dict[str, Any]) -> TaskRecord:
candidate_diff=data.get("candidate_diff"),
diff_hash=data.get("diff_hash"),
ci_results=data.get("ci_results"),
run_id=data.get("run_id"),
ci_correlation_tag=data.get("ci_correlation_tag"),
dispatched_at=data.get("dispatched_at"),
build_loops=data.get("build_loops", 0),
transport=data.get("transport", ""),
created_at=data.get("created_at"),
updated_at=data.get("updated_at"),

View file

@ -121,6 +121,42 @@ workflow_dispatch (task_id, diff_artifact_name, expected_diff_hash, declared_sco
its own grant explicitly. The trigger is `workflow_dispatch` only — the patch
never runs in a context carrying write or secret scope.
## Run-name correlation (box dispatcher → run_id)
The box-side dispatcher (`agent_team/dispatcher.py`) triggers this workflow with
`gh workflow run`, which does **not** return the resulting run id. The dispatcher
must still resolve that `run_id` so the verifier's read-only CI-result fetcher can
poll the correct run (a `None`/unfound run_id fails closed → the verifier gate
BLOCKs / the task parks; never a vacuous pass). The correlation key is the
workflow **run name**:
```yaml
run-name: "agent-team-apply ${{ inputs.task_id }}"
```
Why the run name and not the workflow input or the head branch:
- `gh run list --json name` exposes the run name, but **workflow inputs are not
queryable** via the run list, so the `task_id` cannot be matched on the input.
- A `workflow_dispatch` run reports against the **`main` ref**, not the dispatch
head branch, so the branch is not a usable discriminator either.
So the dispatcher polls `gh run list` read-only, matches the row whose `name`
equals `agent-team-apply <task_id>` (mirrored in code by `dispatcher.run_name_for`),
and bounds the match to runs created after the dispatch watermark. The workflow's
`concurrency` group already guarantees a single in-flight run per `task_id`, so
the run name plus the dispatched-at floor identify the dispatched run
unambiguously even under many simultaneous dispatches; the pure
`dispatcher.select_run_id` then applies the anti-stale tie-break (skip a superseded
cancelled run sharing the name, prefer the later-created `databaseId`).
This `run-name` is **additive**: it adds no job, permission, secret, or trigger,
and changes no privileged step — it only surfaces the dispatching task's id for
correlation. Because the file is nonetheless a CI trust-boundary workflow
(untrusted-input handling), the edit is **flagged for the C1 re-run of
`/sh-security-review` AND the mandatory GPT-4.1 cross-review** before it ships on
this branch, matching the in-YAML `P3-BOX-INTEGRATION` comment.
## SHA-pinned actions (handbook Pinning Principle, §3.3.2)
Every third-party action is pinned to a full commit SHA with the human-readable

View file

@ -0,0 +1,87 @@
# P3 Phase 0 — box-side build→dispatch→verify integration (design decision)
Branch: `feat/agent-team-p3-box-integration`. This note records the load-bearing
design decision for Phase 0 so the security/cross-review gates and the parallel
WebUI session have the rationale in-tree. It is the result of the `/sh-plan-review`
loop (3 rounds, GPT-4.1) + a code-level investigation of the as-built dispatcher,
ci_fetcher, and coordinator execution model.
## Context (verified on-disk, 2026-06-23)
The CI apply/verify workflow is **already live + provisioned** (the `agent-apply`
environment, the `AGENT_APPLY_APP_*` secrets, and dispatched runs all exist as of
2026-06-22). What is **not** built is the box-side integration that makes a task
flow through BUILD → (trigger CI) → VERIFY automatically:
- `dispatcher.py` fires `gh workflow run` but never captures the resulting run id.
- `dispatch_invoker` returns `{}` — nothing writes `state["run_id"]`.
- `ci_fetcher` reads `state["run_id"]` (so it always fails closed → gate BLOCKs).
- The graph orders BUILD → VERIFY → DISPATCH, but VERIFY needs a CI conclusion
that only exists *after* DISPATCH triggers CI. (semantic inversion)
## Decision 1 — node order: BUILD → DISPATCH → VERIFY (reorder)
DISPATCH triggers the CI run and must run *before* VERIFY reads its conclusion.
The P3 subgraph is reordered accordingly (graph.py + build_verify_subgraph.py).
The reorder adds **no new graph nodes** (BUILD/DISPATCH/VERIFY already exist), so
the parallel WebUI branch's graph introspection + NODE_META coverage are
unaffected; only edge wiring changes.
## Decision 2 — CI wait: async resume-on-CI-complete, NOT a blocking poll
A CI run takes ~7 min. The coordinator is a single durable daemon (LangGraph
`interrupt()`/resume + a `tick()` maintenance sweep). A multi-minute *blocking*
VERIFY node would stall the tick loop and every other task. The durable
interrupt/resume machinery already exists for exactly the "external event resumes
a suspended task" shape (the Slack responder; the deadline timer). So:
BUILD → DISPATCH (push branch, trigger CI, capture run_id, suspend)
→ [CI-watcher resumes on terminal conclusion] → VERIFY (read result, gate)
DISPATCH captures `run_id` + `dispatched_at` into state and the task suspends. A
new **CI-watcher** (a `tick()`-driven sweep, mirroring `deadline_timer`) polls the
in-flight `run_id`s read-only and resumes each task once its run reaches a
terminal conclusion (or its `dispatched_at` + timeout elapses → park). VERIFY then
reads the authenticated conclusion via the existing read-only fetcher and the
pure-code gate decides pass/fail. The LLM remains a fix-proposer only.
## Decision 3 — run_id capture is poll-based + correlation-tagged (anti-race)
`gh workflow run` does not return a run id. The dispatcher polls
`gh run list --workflow … --json databaseId,headBranch,createdAt,event` filtered
to this task's **unique per-dispatch head branch** + a per-dispatch
**correlation tag** (a nonce carried as a workflow input and echoed in the run),
bounded to runs created after the dispatch timestamp. This unambiguously matches
the dispatched run even with multiple tasks or rapid re-dispatch. A `None`/unfound
run_id fails closed (the gate BLOCKs / the task parks) — never a vacuous pass.
## Decision 4 — per-task `expected_run_id`
`VerifierConfig.expected_run_id` was a static wiring-time constant. It is now
resolved per-task from `state["run_id"]` (the id the dispatcher captured), so the
gate binds each task's verdict to its own dispatched run and rejects a substituted
run id.
## Decision 5 — fail-safe serve default
Per Adam's decision the bound P3 wiring becomes the new `serve` default. Because
the wiring factories are called eagerly at graph-build, the binding is wrapped so
a missing `AGENT_TEAM_REPO_OWNER`/`_NAME` / CI-read token degrades to the INERT P3
path (task parks, one WARNING + a `#agent-team` inert-mode notice) — never a
`RuntimeError` at serve-start that would crash-loop the daemon.
## Out of scope (tracked follow-ups)
- The §4.3 box-native diff transport (signed artifact / branch-only token) that
would let the always-on box trigger CI without operator credentials. Until then,
the branch push + `gh workflow run` use operator-host credentials (the box holds
no standing write token).
- Tier-3 fixer off `--dry-run`; the cross-plane checker→draft-PR loop.
## Coordination with the WebUI branch (`feature/agent-team-webui-makeover`)
Shared files: `graph.py` (they wrap nodes via `_instrument`; we reorder edges),
`coordinator.py` (both edit the `build_graph(...)` call block). WebUI merges
first; this branch rebases onto the new `main` before deploy. Wrapper composition:
`instrument(failsafe(node))` so a fail-safe park is still logged. The
`project_r720_agent_team` memory fix is owned by the WebUI session.

View file

@ -569,6 +569,110 @@ def _build_notifiers(
return notify, alarm_hook
# The branch namespace the dispatcher pushes its apply/verify draft PRs under
# (see :func:`agent_team.dispatcher.apply_branch_name` -> ``agent-team/apply/``).
# The runaway/stale monitor is scoped to THIS namespace so it only ever surfaces
# agent-team's own draft PRs, never an unrelated human draft PR in the repo.
_DRAFT_PR_HEAD_PREFIX = "agent-team/apply/"
# gh's draft-PR enumeration must never block the daemon's tick() loop. A read-only
# ``gh pr list`` is one GET; bound it so a hung gh invocation parks the sweep
# rather than the whole coordinator.
_DRAFT_PR_LIST_TIMEOUT_S = 30
def _default_draft_pr_provider(*, owner: str, repo: str) -> "Callable[[], list[Any]]":
"""Build the production READ-ONLY draft-PR provider for the A4 monitor.
Returns a zero-arg callable that enumerates the currently-open *agent-team*
draft PRs via a single read-only ``gh pr list`` (one GET; it NEVER writes,
closes, or dispatches anything) and maps each into a
:class:`agent_team.draft_pr_monitor.DraftPr` snapshot the monitor consumes.
The query is scoped to the ``agent-team/apply/`` head namespace
(:data:`_DRAFT_PR_HEAD_PREFIX`) so it only ever sees the dispatcher's own
apply/verify draft PRs — never an unrelated human draft PR. ``--json`` pulls
exactly the three fields the monitor keys off (``number`` / ``createdAt`` ->
``opened_at`` for the runaway window, ``updatedAt`` -> ``updated_at`` for
staleness).
Mirrors :func:`agent_team.ci_watcher.default_ci_poller`: a thin closure over
``owner`` / ``repo`` that fails closed — a non-zero ``gh`` exit, a timeout, or
unparseable JSON yields an empty snapshot (the monitor then no-ops this pass)
rather than raising, so a transient gh hiccup never breaks the tick loop. (The
coordinator's ``_draft_pr_monitor_sweep`` ALSO swallows provider errors, so
this is belt-and-suspenders.)
"""
repo_slug = f"{owner}/{repo}"
def provider() -> list[Any]:
import json as _json
import subprocess # noqa: PLC0415 - deferred so import needs no gh
from agent_team.draft_pr_monitor import DraftPr # noqa: PLC0415
try:
proc = subprocess.run( # noqa: S603 - args are a fixed, non-shell list
[
"gh",
"pr",
"list",
"--repo",
repo_slug,
"--draft",
"--state",
"open",
"--search",
f"head:{_DRAFT_PR_HEAD_PREFIX}",
"--json",
"number,createdAt,updatedAt",
"--limit",
"100",
],
check=True,
capture_output=True,
text=True,
timeout=_DRAFT_PR_LIST_TIMEOUT_S,
)
except (
subprocess.CalledProcessError,
subprocess.TimeoutExpired,
OSError,
):
_LOG.warning(
"draft-pr-monitor: read-only `gh pr list` enumeration failed; "
"treating as no open draft PRs this pass",
exc_info=True,
)
return []
try:
rows = _json.loads(proc.stdout or "[]")
except ValueError:
_LOG.warning(
"draft-pr-monitor: `gh pr list` returned unparseable JSON; "
"treating as no open draft PRs this pass",
exc_info=True,
)
return []
snapshots: list[Any] = []
for row in rows:
number = row.get("number")
if number is None:
continue
snapshots.append(
DraftPr(
number=int(number),
opened_at=row.get("createdAt"),
updated_at=row.get("updatedAt"),
)
)
return snapshots
return provider
def _build_coordinator(args: argparse.Namespace) -> Any:
"""Construct a :class:`Coordinator` for the ``start`` / ``serve`` commands.
@ -587,6 +691,7 @@ def _build_coordinator(args: argparse.Namespace) -> Any:
default_clarify_node_factory,
default_plan_node_factory,
default_review_wiring,
failsafe_production_p3_wiring,
)
transport = _build_transport(args)
@ -604,6 +709,43 @@ def _build_coordinator(args: argparse.Namespace) -> Any:
# silent (notify None) so import + ledger commands need no token.
notify, alarm_hook = _build_notifiers(args)
# P3 fail-safe serve default (design Decision 5): the bound build→verify +
# dispatch wiring is now the production ``serve`` default, but it must NEVER
# crash-loop serve-start. ``failsafe_production_p3_wiring`` resolves the env
# ONCE: configured -> the live P3 pair; unconfigured -> (None, None) inert (a
# task reaching P3 parks), with one WARNING + one #agent-team inert notice
# via the lifecycle ``notify`` sink. Scoped to ``serve`` (the daemon): the
# one-shot ``start`` / ``intake-*`` paths never auto-bind P3 — they run to the
# first human gate and exit, well short of BUILD/VERIFY.
build_verify_wiring = None
dispatch_node_wiring = None
if getattr(args, "command", None) == "serve":
build_verify_wiring, dispatch_node_wiring = failsafe_production_p3_wiring(
notify=notify
)
# CI-watcher seams (design §4 Decision 2 — async resume-on-CI-complete). ONLY
# wired when ``failsafe_production_p3_wiring`` returned a LIVE pair (a
# configured box): on the inert/unconfigured box both stay None, so the
# tick() CI sweep is a NO-OP and there is no behaviour change. The poller is
# the read-only default (one GET per run, never a write); the provider is the
# durable enumerator bound to THIS coordinator below (post-construction, so it
# can close over the just-built coordinator); the timeout is a sane default
# (30 min — generous headroom over the ~7-min CI run before a stuck run
# parks). Without these, a task that dispatches and suspends at VERIFY would
# wait forever — the async-resume gap this closes.
ci_poller = None
ci_timeout = None
if build_verify_wiring is not None and dispatch_node_wiring is not None:
from datetime import timedelta
from agent_team.ci_watcher import default_ci_poller
owner = os.environ.get("AGENT_TEAM_REPO_OWNER", "").strip()
repo = os.environ.get("AGENT_TEAM_REPO_NAME", "").strip()
ci_poller = default_ci_poller(owner=owner, repo=repo)
ci_timeout = timedelta(minutes=30)
# Production runs the full P2 graph: the wrapped real planner + the bound
# GPT-4.1 review loop (Plane-2 depth-first). These factories are lazy and
# only build/bind the model seams when a task actually runs.
@ -617,9 +759,32 @@ def _build_coordinator(args: argparse.Namespace) -> Any:
context_provider=context_provider
),
review_wiring=default_review_wiring,
build_verify_wiring=build_verify_wiring,
dispatch_node_wiring=dispatch_node_wiring,
notify=notify,
alarm_hook=alarm_hook,
ci_poller=ci_poller,
ci_timeout=ci_timeout,
)
# Bind the durable CI-pending provider to THIS coordinator (only on the live
# P3 path — ``ci_poller`` is the live-pair signal). It enumerates the durable
# threads suspended at VERIFY awaiting CI so the watcher has real tasks to
# poll; bound post-construction so it can reference the just-built
# coordinator. Left unbound on the inert box, the CI sweep stays a NO-OP.
if ci_poller is not None:
coordinator._ci_pending_provider = coordinator._enumerate_ci_pending
# Draft-PR runaway/stale monitor seam (P3 A4). Gated on the SAME live-pair
# signal as the CI watcher (``ci_poller`` is set ⇔ ``_p3_env_is_configured``
# via ``failsafe_production_p3_wiring``), so on the inert/unconfigured box it
# stays None and ``_draft_pr_monitor_sweep`` is a NO-OP (no behaviour change).
# On the live box it binds a read-only ``gh pr list`` enumerator of the open
# ``agent-team/apply/`` draft PRs so the sweep has a real snapshot to ALARM /
# remind on — without this the wired sweep would always see no PRs. Bound
# post-construction to mirror ``_ci_pending_provider``.
if ci_poller is not None:
coordinator._draft_pr_provider = _default_draft_pr_provider(
owner=owner, repo=repo
)
# WS2: an allowlisted Slack /new-task starts a task on THIS coordinator. Set
# post-construction (the adapter closes over the just-built coordinator), and
# before serve() builds the listener. AUTHZ-01 (owner allowlist) gates this
@ -1036,6 +1201,118 @@ def _cmd_force_resume(args: argparse.Namespace, *, out: Any) -> int:
return 1
def _cmd_dispatch(args: argparse.Namespace, *, out: Any) -> int:
"""Operator-initiated dispatch of a built diff into org CI (P3, option-b).
The box holds NO write token (read-only by design), so its in-graph DISPATCH
node fail-closes/parks. This is the operator verb that completes the dispatch
with WRITE creds: it reads the task's ``candidate_diff`` + declared scope
(from the ledger checkpoint, or from ``--diff``/``--scope`` files), pushes the
head branch and fires the apply/verify ``workflow_dispatch`` via
:func:`agent_team.dispatcher.dispatch_apply_verify`, then prints the located
run id. Run it where a WRITE-capable ``GH_TOKEN`` is available — the operator
host, or the box with a JUST-IN-TIME operator token in the env (never stored
in ``secrev.env``); the box stays read-only at rest.
Owner/repo/base resolve from ``--owner``/``--repo``/``--base`` or the
``AGENT_TEAM_REPO_OWNER``/``_NAME``/``_BASE_BRANCH`` env vars. With
``--write-back`` the located ``run_id`` is written into the task checkpoint so
the box's VERIFY can bind to it.
"""
import os
from agent_team.dispatcher import dispatch_apply_verify
owner = args.owner or os.environ.get("AGENT_TEAM_REPO_OWNER", "")
repo = args.repo or os.environ.get("AGENT_TEAM_REPO_NAME", "")
base = args.base or os.environ.get("AGENT_TEAM_BASE_BRANCH") or "main"
if not owner or not repo:
print(
"dispatch: owner/repo required (pass --owner/--repo or set "
"AGENT_TEAM_REPO_OWNER/_NAME)",
file=sys.stderr,
)
return 2
# Resolve the diff + declared scope: explicit files win; else read the task's
# checkpointed PipelineState (candidate_diff + plan.scope).
diff_text: str | None = (
Path(args.diff).read_text(encoding="utf-8") if args.diff else None
)
declared_scope: str | None = (
Path(args.scope).read_text(encoding="utf-8") if args.scope else None
)
if diff_text is None or declared_scope is None:
from agent_team.graph import (
build_graph,
build_sqlite_checkpointer,
thread_config,
)
with build_sqlite_checkpointer(args.db) as saver:
graph = build_graph(saver)
snap = graph.get_state(thread_config(args.thread_id))
state = dict(snap.values or {})
if diff_text is None:
diff_text = state.get("candidate_diff")
if declared_scope is None:
plan = state.get("plan") or {}
scope_list = plan.get("scope") or [] if isinstance(plan, dict) else []
declared_scope = "\n".join(str(s) for s in scope_list if s)
if not diff_text or not str(diff_text).strip():
print(
f"dispatch: no candidate_diff for task {args.thread_id} "
"(pass --diff, or the task has not built a diff yet)",
file=sys.stderr,
)
return 1
result = dispatch_apply_verify(
owner=owner,
repo=repo,
task_id=args.thread_id,
diff_text=diff_text,
declared_scope=declared_scope or "",
base=base,
)
print(
f"dispatched task {args.thread_id} -> {owner}/{repo} "
f"(head={result.inputs.head_branch}, run_id={result.run_id}, "
f"dispatched_at={result.dispatched_at})",
file=out,
)
if result.run_id is None:
print(
"dispatch: workflow fired but run_id could not be correlated; "
"VERIFY fails closed until a run_id is set",
file=sys.stderr,
)
elif args.write_back:
from agent_team.graph import (
build_graph,
build_sqlite_checkpointer,
thread_config,
)
with build_sqlite_checkpointer(args.db) as saver:
graph = build_graph(saver)
graph.update_state(
thread_config(args.thread_id),
{
"run_id": result.run_id,
"dispatched_at": result.dispatched_at,
"ci_correlation_tag": result.correlation_tag,
},
)
print(
f"dispatch: wrote run_id={result.run_id} into the task checkpoint "
"(--write-back)",
file=out,
)
return 0 if result.run_id is not None else 1
# Statuses an operator treats as "parked context": a task whose only pending
# question is no longer open may be parked (answered-but-unresumed, expired, or
# superseded). ``open`` is excluded — that is the live-waiting view (default
@ -1166,6 +1443,45 @@ def build_parser() -> argparse.ArgumentParser:
)
p_resume.set_defaults(func=_cmd_force_resume)
p_dispatch = sub.add_parser(
"dispatch",
help=(
"operator-initiated dispatch of a task's built diff into org CI (P3, "
"option-b) — needs a WRITE-capable GH_TOKEN in the env"
),
)
p_dispatch.add_argument(
"thread_id", help="the task thread_id whose candidate_diff to dispatch"
)
p_dispatch.add_argument(
"--owner", default=None, help="repo owner (default $AGENT_TEAM_REPO_OWNER)"
)
p_dispatch.add_argument(
"--repo", default=None, help="repo name (default $AGENT_TEAM_REPO_NAME)"
)
p_dispatch.add_argument(
"--base",
default=None,
help="base branch (default $AGENT_TEAM_BASE_BRANCH or main)",
)
p_dispatch.add_argument(
"--diff",
default=None,
help="path to a unified-diff file (overrides the ledger candidate_diff)",
)
p_dispatch.add_argument(
"--scope",
default=None,
help="path to a newline-separated declared-scope file (overrides plan.scope)",
)
p_dispatch.add_argument(
"--write-back",
dest="write_back",
action="store_true",
help="write the located run_id into the task checkpoint so VERIFY binds to it",
)
p_dispatch.set_defaults(func=_cmd_dispatch)
p_start = sub.add_parser(
"start",
help="intake: start one task and run it to the first human gate",

View file

@ -0,0 +1,318 @@
#!/usr/bin/env python3
"""assert_no_write_token.py — fail closed if a GitHub *write* token is on the box.
The agent-team apply/verify path mints its ``pull-requests: write`` GitHub App
installation token **inside the CI runner** (via SHA-pinned
``actions/create-github-app-token``), from the ``AGENT_APPLY_APP_ID`` /
``AGENT_APPLY_APP_PRIVATE_KEY`` **Actions secrets**. By design the always-on
R720 box holds **no standing write credential**: it triggers CI with the
operator's host ``gh`` auth and reads results with a *read-only* token
(``AGENT_TEAM_CI_READ_TOKEN`` / ``GITHUB_TOKEN``). See ``ci/README.md`` and
``docs/P3-PHASE0-DESIGN.md`` ("the box holds no standing write token").
This audit asserts that invariant. It is runnable both on the box (as a
provisioning/runtime self-check) and in CI (as a regression guard). It:
1. scans the live process environment (``os.environ``) and the coordinator's
environment-derived config for token-shaped variables that would grant
``pull-requests: write`` / ``contents: write``;
2. greps the box's local credential files (``~/secrev.env`` and
``~/orchestrator/.env`` by default; paths are configurable) for the App id
and for any PEM private-key header (RSA / EC / OPENSSH / PKCS#8);
and **exits non-zero with a clear message** if any are found, or exits ``0``
with a short summary otherwise.
Usage::
python -m scripts.assert_no_write_token
python scripts/assert_no_write_token.py --env-file ~/secrev.env --env-file ~/x.env
Exit codes:
0 no write-shaped token / App secret / private key found
1 at least one finding (the box is mis-provisioned — remediate before deploy)
"""
from __future__ import annotations
import argparse
import os
import re
import sys
from collections.abc import Mapping, Sequence
from pathlib import Path
# --------------------------------------------------------------------------- #
# What "must never live on the box" looks like.
# --------------------------------------------------------------------------- #
# The GitHub App that holds ``pull-requests: write`` lives ONLY as Actions
# secrets. Its id and private key must never appear on the box (env or files).
APP_ID_ENV = "AGENT_APPLY_APP_ID"
APP_PRIVATE_KEY_ENV = "AGENT_APPLY_APP_PRIVATE_KEY"
# Env vars that are explicitly the App's write credentials.
_FORBIDDEN_ENV_VARS: frozenset[str] = frozenset(
{
APP_ID_ENV,
APP_PRIVATE_KEY_ENV,
}
)
# Env vars that are KNOWN-GOOD read-only / non-write and must NOT be flagged by
# the heuristic name match below (they are tokens, but read-only by contract).
_ALLOWED_TOKEN_ENV_VARS: frozenset[str] = frozenset(
{
"AGENT_TEAM_CI_READ_TOKEN", # read-only CI-result fetcher token
"AGENT_TEAM_API_TOKEN", # read-only dashboard/API bearer
"SLACK_APP_TOKEN", # Slack socket-mode app-level token (not GitHub)
"SLACK_BOT_TOKEN", # Slack bot token (not GitHub)
"CLAUDE_CODE_OAUTH_TOKEN", # Anthropic subscription OAuth (not GitHub)
"ANTHROPIC_API_KEY", # model API key (not GitHub)
}
)
# A name-shaped heuristic for "this looks like a GitHub *write* token". We
# deliberately scope to GitHub-write shapes so the generic read-only
# ``GITHUB_TOKEN`` fallback (a runtime read token, not a standing write secret)
# is not a false positive while the App's write material always is.
_WRITE_TOKEN_NAME_RE = re.compile(
r"(?:^|_)(?:GH|GITHUB)_(?:APP|PAT|WRITE|APPLY)_?(?:TOKEN|KEY|PRIVATE_KEY)?",
re.IGNORECASE,
)
# Value shapes that indicate a GitHub write-capable credential regardless of the
# var's name. ``ghp_`` (classic PAT) and ``github_pat_`` (fine-grained PAT) can
# both carry write scopes; an installation token (``ghs_``) is write-capable; a
# user-to-server token (``ghu_``) acts with the user's write access; and a refresh
# token (``ghr_``) mints fresh write-capable user-to-server tokens. All are
# write-risk material that must not live on the box.
_WRITE_TOKEN_VALUE_RE = re.compile(
r"\b(?:ghp_|ghs_|ghu_|ghr_|github_pat_)[A-Za-z0-9_]{20,}\b"
)
# Any PEM private-key header (RSA / EC / OPENSSH / generic PKCS#8). The App's
# private key is a PEM block; finding ANY private key in a box credential file
# is a finding.
_PRIVATE_KEY_RE = re.compile(r"-----BEGIN (?:[A-Z0-9]+ )*PRIVATE KEY-----")
# Default box credential files to grep. Configurable via --env-file / the
# ``ASSERT_NO_WRITE_TOKEN_ENV_FILES`` env var so tests use temp files.
_DEFAULT_ENV_FILES: tuple[str, ...] = ("~/secrev.env", "~/orchestrator/.env")
# --------------------------------------------------------------------------- #
# Scanners. Each returns a list of human-readable finding strings.
# --------------------------------------------------------------------------- #
def scan_environ(environ: Mapping[str, str], *, app_id: str | None = None) -> list[str]:
"""Scan a process environment for write-shaped GitHub token material.
Flags (a) the explicit App-credential env vars, (b) any var whose *name*
matches the GitHub-write heuristic, (c) any var whose *value* carries a
write-capable GitHub token prefix, (d) any var whose *value* is a PEM
private-key block (the App key exported under a benign name), and (e) any var
whose *value* contains the configured App id. Known read-only tokens are
exempt from the *name* and write-token *value* heuristics, but the PEM and
App-id value checks below apply to EVERY var (including the allowlisted ones
and ``GITHUB_TOKEN``) — a private key or the App id can never legitimately
sit in any env var, so those checks are never name-exempted. This mirrors
:func:`scan_env_file`, which greps file contents for the same App-id /
private-key material regardless of the var name carrying it.
"""
findings: list[str] = []
for name, value in environ.items():
# The PEM private-key VALUE check and the configured App-id VALUE check
# are NOT name-exempt: the App's private key exported under a benign name
# (e.g. GH_APP_KEY set to a "BEGIN ... PRIVATE KEY" PEM block) and the
# configured App id parked in any var are both forbidden material, even in an
# otherwise read-only/allowlisted var. Check them up front, before the
# allowlist short-circuits the name/token-prefix heuristics.
if value and _PRIVATE_KEY_RE.search(value):
findings.append(
f"environment variable {name!r} holds a PEM private-key block "
"(-----BEGIN ... PRIVATE KEY-----); no private key may live on "
"the box (the App key must live ONLY as an Actions secret)"
)
if app_id and value and app_id in value:
findings.append(
f"environment variable {name!r} contains the configured App id "
"value; the App id must not be present on the box"
)
if name in _ALLOWED_TOKEN_ENV_VARS:
continue
if name in _FORBIDDEN_ENV_VARS:
findings.append(
f"environment variable {name!r} is set — the App write "
"credential must live ONLY as an Actions secret, never on the box"
)
continue
if _WRITE_TOKEN_NAME_RE.search(name):
findings.append(
f"environment variable {name!r} has a GitHub write-token-shaped "
"name; the box must hold no standing write token"
)
continue
if value and _WRITE_TOKEN_VALUE_RE.search(value):
findings.append(
f"environment variable {name!r} holds a write-capable GitHub "
"token value (ghp_/ghs_/ghu_/ghr_/github_pat_ prefix)"
)
return findings
def scan_config(config: Mapping[str, object] | None) -> list[str]:
"""Scan a coordinator config mapping for write-shaped token material.
The coordinator is environment-driven, so this is normally a thin pass over
whatever config dict a caller hands in (string values only). It applies the
same name/value heuristics as :func:`scan_environ`.
"""
if not config:
return []
findings: list[str] = []
for key, raw in config.items():
name = str(key)
if name in _ALLOWED_TOKEN_ENV_VARS:
continue
if name in _FORBIDDEN_ENV_VARS or _WRITE_TOKEN_NAME_RE.search(name):
findings.append(
f"coordinator config key {name!r} is a GitHub write-token-shaped "
"key; the box config must hold no standing write token"
)
continue
if isinstance(raw, str) and _WRITE_TOKEN_VALUE_RE.search(raw):
findings.append(
f"coordinator config key {name!r} holds a write-capable GitHub "
"token value (ghp_/ghs_/ghu_/ghr_/github_pat_ prefix)"
)
return findings
def scan_env_file(path: Path, *, app_id: str | None = None) -> list[str]:
"""Grep one credential file for the App id and any PEM private-key header.
A non-existent file is **not** a finding (the box legitimately may not have
every file). An unreadable-but-present file is reported as a finding so a
permissions mistake can't silently mask leaked material.
When ``app_id`` is given (the configured ``AGENT_APPLY_APP_ID`` value), the
grep also flags that literal id appearing in the file. The
``AGENT_APPLY_APP_ID`` / ``AGENT_APPLY_APP_PRIVATE_KEY`` *names* are always
flagged regardless.
"""
findings: list[str] = []
if not path.exists():
return findings
try:
text = path.read_text(encoding="utf-8", errors="replace")
except OSError as exc: # present but unreadable — fail loud, not silent
return [
f"could not read credential file {path} ({exc}); cannot prove it is clean"
]
for lineno, line in enumerate(text.splitlines(), start=1):
if APP_ID_ENV in line or APP_PRIVATE_KEY_ENV in line:
findings.append(
f"{path}:{lineno}: references the App credential "
f"({APP_ID_ENV}/{APP_PRIVATE_KEY_ENV}); it must live ONLY as an "
"Actions secret"
)
if app_id and app_id in line:
findings.append(
f"{path}:{lineno}: contains the configured App id value; the App "
"id must not be present on the box"
)
if _PRIVATE_KEY_RE.search(text):
findings.append(
f"{path}: contains a PEM private-key block "
"(-----BEGIN ... PRIVATE KEY-----); no private key may live on the box"
)
return findings
# --------------------------------------------------------------------------- #
# Orchestration.
# --------------------------------------------------------------------------- #
def _resolve_env_files(
cli_files: Sequence[str] | None, environ: Mapping[str, str]
) -> list[Path]:
"""Resolve the credential files to grep (CLI > env var > defaults)."""
if cli_files:
raw = list(cli_files)
elif environ.get("ASSERT_NO_WRITE_TOKEN_ENV_FILES"):
raw = [
p.strip()
for p in environ["ASSERT_NO_WRITE_TOKEN_ENV_FILES"].split(os.pathsep)
if p.strip()
]
else:
raw = list(_DEFAULT_ENV_FILES)
return [Path(p).expanduser() for p in raw]
def audit(
*,
environ: Mapping[str, str] | None = None,
config: Mapping[str, object] | None = None,
env_files: Sequence[str] | None = None,
) -> list[str]:
"""Run every scanner and return the combined list of findings (empty == clean)."""
environ = os.environ if environ is None else environ
app_id = environ.get(APP_ID_ENV) or None
findings: list[str] = []
findings += scan_environ(environ, app_id=app_id)
findings += scan_config(config)
for path in _resolve_env_files(env_files, environ):
findings += scan_env_file(path, app_id=app_id)
return findings
def main(argv: Sequence[str] | None = None) -> int:
parser = argparse.ArgumentParser(
description=(
"Assert no pull-requests:write / contents:write GitHub token (or the "
"apply App's id/private key) is present on this box."
)
)
parser.add_argument(
"--env-file",
action="append",
dest="env_files",
metavar="PATH",
help=(
"Credential file to grep for the App id + private keys (repeatable). "
"Defaults to ~/secrev.env and ~/orchestrator/.env, or the os.pathsep-"
"separated ASSERT_NO_WRITE_TOKEN_ENV_FILES env var."
),
)
args = parser.parse_args(argv)
findings = audit(env_files=args.env_files)
if findings:
print(
"FAIL: write-capable GitHub credential material found on the box "
f"({len(findings)} finding(s)). The pull-requests:write App token "
"must only ever live as an Actions secret:",
file=sys.stderr,
)
for f in findings:
print(f" - {f}", file=sys.stderr)
return 1
print(
"OK: no pull-requests:write / contents:write token, App id, or private "
"key found in the environment, coordinator config, or credential files. "
"The box holds no standing write token."
)
return 0
if __name__ == "__main__":
raise SystemExit(main())

1011
agent-team/scripts/p3_rollback.sh Executable file

File diff suppressed because it is too large Load diff

View file

@ -32,6 +32,27 @@
# AGENT_TEAM_SLACK_OWNER_IDS -> comma-separated authorized answerer ids
# (AUTHZ-01). The listener FAILS CLOSED if unset.
#
# --- P3 build->dispatch->verify wiring (read by the LIVE serve() path) ---
# The bound P3 wiring is the serve() default; if any of the three below (plus a
# CI-read token) is unset, serve() degrades to the INERT P3 path (one WARNING +
# a #agent-team notice) instead of crash-looping (Decision 5). All are read at
# graph-build / dispatch time from this file's environment:
# AGENT_TEAM_REPO_OWNER -> dispatch target owner. Fixed at factory time,
# never read from pipeline state, so model output
# cannot redirect the dispatch/verify target.
# AGENT_TEAM_REPO_NAME -> dispatch target repo (same fail-closed binding).
# AGENT_TEAM_BASE_BRANCH -> PR base branch (optional; default "main").
# AGENT_TEAM_CI_READ_TOKEN -> the READ-ONLY CI-result token. The verifier's
# authenticated conclusion read uses it (falls back
# to GITHUB_TOKEN). This is read-only by contract:
# NO pull-requests:write / contents:write token,
# and NO AGENT_APPLY_APP_ID / _PRIVATE_KEY, may live
# in this file. The apply path mints its write token
# INSIDE the CI runner from Actions secrets; the box
# holds no standing write credential. The invariant
# is asserted by scripts/assert_no_write_token.py
# (the A2 audit) at provisioning + in CI.
#
# ~/orchestrator/.env (mode 600, NOT in git) - the P2 review-loop provider key:
# The production serve() wires the GPT-4.1 cross-review loop, which shells the
# local orchestrator run.py -> cross_reviewer once a task reaches REVIEW. That

View file

@ -102,6 +102,19 @@ def test_route_ids_mirror_graph_by_value() -> None:
assert APPROVED_ROUTE == "approved"
def test_integrate_note_documents_build_dispatch_verify_order() -> None:
"""The topology note records BUILD -> [DISPATCH] -> VERIFY (design §4 D1).
DISPATCH must precede VERIFY so it captures ``state["run_id"]`` before VERIFY
reads it. Pin the documented order so a future reorder back to the old
BUILD -> VERIFY -> DISPATCH topology is caught.
"""
note = bvs._INTEGRATE_NOTE
assert "BUILD -> [DISPATCH] -> VERIFY" in note
# DISPATCH appears before VERIFY in the recorded stage order.
assert note.index("DISPATCH") < note.index("VERIFY")
# --------------------------------------------------------------------------- #
# BUILD node: proposes a diff via an injected fake builder
# --------------------------------------------------------------------------- #
@ -224,6 +237,47 @@ def test_verify_failing_ci_result_loops_back_to_build() -> None:
assert out["current_phase"] == Phase.BUILD.value
assert out["ci_results"]["gate_decision"] == "fail"
assert route_after_verify(out) == BUILD_ROUTE
# The recoverable loop persists the incremented DURABLE count into state.
assert out["build_loops"] == 1
def test_repeated_fail_through_wrapper_parks_after_max_build_loops() -> None:
"""A perpetually-FAILing task PARKS after exactly max_build_loops loops.
Drives the durable BUILD->DISPATCH->VERIFY budget through the subgraph
wrapper with ONE shared VerifierConfig (the wiring-time reality). The count
that bounds the loop is the per-task ``state["build_loops"]`` the node writes
back each round (LOGIC-RACE-01) — the shared config's build_loops stays 0, so
if the node read the count from config the loop would never terminate.
"""
diff = _diff_for("src/foo.py")
def fail_fetcher(state):
return {"run_id": "r1", "conclusion": "failure", "diff_hash": _hash(diff)}
shared_cfg = VerifierConfig(
expected_run_id="r1",
allowed_scope=["src"],
max_build_loops=3,
build_loops=0,
)
node = make_verify_node(shared_cfg, ci_result_fetcher=fail_fetcher)
state = _verify_state(diff)
state["build_loops"] = 0
routes: list[str] = []
for _ in range(10): # bound the harness; a regression must not hang
out = node(state)
routes.append(route_after_verify(out))
if route_after_verify(out) == PARKED_ROUTE:
break
# The checkpointer would carry build_loops forward across the loop.
state["build_loops"] = out["build_loops"]
assert routes == [BUILD_ROUTE, BUILD_ROUTE, PARKED_ROUTE]
assert out["status"] == TaskStatus.PARKED.value
assert out["current_phase"] == Phase.PARKED.value
def test_verify_malformed_fetcher_result_fails_safe_to_parked() -> None:
@ -373,3 +427,203 @@ def test_verify_node_does_not_mutate_caller_state() -> None:
node(state)
# Caller's state is untouched (the node wrote into a dict copy).
assert state["ci_results"] is sentinel
# --------------------------------------------------------------------------- #
# Per-task run-id binding through the subgraph wrapper (design §4 Decision 4)
# --------------------------------------------------------------------------- #
def test_verify_node_binds_to_per_task_state_run_id() -> None:
"""make_verify_node gates against state["run_id"], not a static config id.
A single VerifierConfig with no expected_run_id is shared; the per-task
state["run_id"] supplies the binding the fetched CI conclusion must match.
"""
diff = _diff_for("src/foo.py")
def fetcher_keyed_to_task(s):
# The (real) fetcher keys its conclusion to the task's dispatched run id.
return {
"run_id": s["run_id"],
"conclusion": "success",
"diff_hash": _hash(diff),
}
node = make_verify_node(
VerifierConfig(expected_run_id=None, allowed_scope=["src"]),
ci_result_fetcher=fetcher_keyed_to_task,
)
state = _verify_state(diff)
state["run_id"] = "dispatched-123"
out = node(state)
assert out["status"] == TaskStatus.DONE.value
assert out["review_verdicts"][0]["run_id"] == "dispatched-123"
def test_verify_node_blocks_substituted_run_id_through_wrapper() -> None:
"""A fetched CI result keyed to a DIFFERENT run id than state["run_id"] blocks."""
diff = _diff_for("src/foo.py")
def substituting_fetcher(s):
# CI result grafted from another task (run id mismatch) with success.
return {
"run_id": "someone-elses-run",
"conclusion": "success",
"diff_hash": _hash(diff),
}
node = make_verify_node(
VerifierConfig(expected_run_id=None, allowed_scope=["src"]),
ci_result_fetcher=substituting_fetcher,
)
state = _verify_state(diff)
state["run_id"] = "my-run"
out = node(state)
assert out["status"] == TaskStatus.PARKED.value
assert out["ci_results"]["gate_decision"] == "block"
def test_verify_node_blocks_when_no_run_id_anywhere() -> None:
"""No state run_id and no config fallback -> BLOCK/park, never a vacuous pass."""
diff = _diff_for("src/foo.py")
def pass_fetcher(s):
return {"run_id": "r1", "conclusion": "success", "diff_hash": _hash(diff)}
node = make_verify_node(
VerifierConfig(expected_run_id=None, allowed_scope=["src"]),
ci_result_fetcher=pass_fetcher,
)
out = node(_verify_state(diff)) # no state["run_id"]
assert out["status"] == TaskStatus.PARKED.value
assert out["ci_results"]["gate_decision"] == "block"
# --------------------------------------------------------------------------- #
# Async CI-wait: VERIFY suspends via interrupt() on an in-progress run (§4 Dec 2)
# --------------------------------------------------------------------------- #
def test_verify_suspends_on_in_progress_run_then_resumes_to_gate() -> None:
"""With a dispatched run_id and a not-yet-terminal fetch, VERIFY interrupts;
once the CI-watcher resumes it, the re-fetched terminal result is gated."""
from langgraph.checkpoint.memory import MemorySaver
from langgraph.graph import END, START, StateGraph
from langgraph.types import Command
diff = _diff_for("src/foo.py")
# The fetcher is "in progress" on the first call, then terminal on the second
# (mirroring a real run that concluded while the task was suspended).
calls = {"n": 0}
def progressing_fetcher(state):
calls["n"] += 1
if calls["n"] == 1:
return None # still running -> VERIFY must suspend
return {
"run_id": state["run_id"],
"conclusion": "success",
"diff_hash": _hash(diff),
}
node = make_verify_node(
VerifierConfig(expected_run_id=None, allowed_scope=["src"]),
ci_result_fetcher=progressing_fetcher,
)
builder = StateGraph(PipelineState)
builder.add_node("verify", node)
builder.add_edge(START, "verify")
builder.add_edge("verify", END)
graph = builder.compile(checkpointer=MemorySaver())
cfg = {"configurable": {"thread_id": "tw1"}}
state = _verify_state(diff)
state["run_id"] = "dispatched-77"
first = graph.invoke(state, cfg)
# The task suspended at VERIFY's interrupt (awaiting CI) rather than gating.
assert "__interrupt__" in first
assert calls["n"] == 1
# The CI-watcher resumes the task once the run terminated.
out = graph.invoke(Command(resume={"awaiting_ci": "done"}), cfg)
# On resume the node re-fetched the now-terminal result and the gate PASSed.
assert calls["n"] == 2
assert out["status"] == TaskStatus.DONE.value
assert out["current_phase"] == Phase.DONE.value
def test_verify_spurious_resume_still_non_terminal_parks_never_passes() -> None:
"""A spurious resume (re-fetch STILL non-terminal) must fail closed.
The CI-watcher resumes VERIFY on what it believes is a terminal conclusion,
but the authenticated re-fetch is the source of truth. If that re-fetch is
STILL None (a spurious / premature resume, or a run that flapped back to
in-progress), the node must NOT vacuously pass: it falls through to the gate,
which — with a dispatched run_id but no terminal result — BLOCKs and parks.
"""
from langgraph.checkpoint.memory import MemorySaver
from langgraph.graph import END, START, StateGraph
from langgraph.types import Command
diff = _diff_for("src/foo.py")
# The fetch is non-terminal on EVERY call (the resume was spurious — the run
# never actually concluded).
calls = {"n": 0}
def never_terminal_fetcher(state):
calls["n"] += 1
return None
node = make_verify_node(
VerifierConfig(expected_run_id=None, allowed_scope=["src"]),
ci_result_fetcher=never_terminal_fetcher,
)
builder = StateGraph(PipelineState)
builder.add_node("verify", node)
builder.add_edge(START, "verify")
builder.add_edge("verify", END)
graph = builder.compile(checkpointer=MemorySaver())
cfg = {"configurable": {"thread_id": "spurious-1"}}
state = _verify_state(diff)
state["run_id"] = "dispatched-99"
first = graph.invoke(state, cfg)
# First pass: in-progress fetch -> VERIFY suspends awaiting CI.
assert "__interrupt__" in first
assert calls["n"] == 1
# The watcher resumes, but the authenticated re-fetch is STILL non-terminal.
out = graph.invoke(Command(resume={"awaiting_ci": "spurious"}), cfg)
# Re-fetched again on resume (the node replays from its start, so it fetches,
# the interrupt returns the resume value rather than re-suspending, then it
# re-fetches once more) — every fetch is non-terminal.
assert calls["n"] > 1
# ...and with no terminal result the gate BLOCKs and the task PARKS — never a
# vacuous pass.
assert out["status"] == TaskStatus.PARKED.value
assert out["current_phase"] == Phase.PARKED.value
assert out["ci_results"]["gate_decision"] == "block"
assert route_after_verify(out) == PARKED_ROUTE
def test_verify_no_run_id_does_not_suspend_and_parks() -> None:
"""The INERT/no-run path (no state run_id) never suspends: a None fetch flows
straight to the gate, which BLOCKs and parks (existing behavior preserved)."""
diff = _diff_for("src/foo.py")
node = make_verify_node(
VerifierConfig(expected_run_id="r1", allowed_scope=["src"]),
ci_result_fetcher=lambda s: None,
)
out = node(_verify_state(diff)) # no state["run_id"] -> no interrupt
assert out["current_phase"] == Phase.PARKED.value
assert route_after_verify(out) == PARKED_ROUTE

View file

@ -174,6 +174,29 @@ def test_missing_token_fails_closed(monkeypatch: pytest.MonkeyPatch) -> None:
assert fetch_ci_result(_state(), owner="o", repo="r") is None
# --------------------------------------------------------------------------- #
# SEC-04: token resolution prefers the dedicated read-only var over GITHUB_TOKEN
# --------------------------------------------------------------------------- #
def test_resolve_prefers_dedicated_read_token(monkeypatch: pytest.MonkeyPatch) -> None:
from agent_team.ci_fetcher import _resolve_read_token
monkeypatch.setenv(CI_READ_TOKEN_ENV, "dedicated-read-only")
monkeypatch.setenv("GITHUB_TOKEN", "fallback-token")
# The dedicated read-only var wins on the live box read path.
assert _resolve_read_token() == "dedicated-read-only"
def test_resolve_falls_back_to_github_token(monkeypatch: pytest.MonkeyPatch) -> None:
from agent_team.ci_fetcher import _resolve_read_token
monkeypatch.delenv(CI_READ_TOKEN_ENV, raising=False)
monkeypatch.setenv("GITHUB_TOKEN", "fallback-token")
# Documented, explicitly-narrowed convenience fallback (must be read-only).
assert _resolve_read_token() == "fallback-token"
# --------------------------------------------------------------------------- #
# Read-only contract: never derives a verdict, never writes
# --------------------------------------------------------------------------- #

View file

@ -389,14 +389,46 @@ def test_gate_scope_violation_blocks() -> None:
assert any("scope" in r for r in result.reasons)
def test_gate_empty_run_id_raises() -> None:
def test_gate_empty_run_id_blocks_never_passes() -> None:
# Per design §4 (per-task binding): an empty expected_run_id means the
# dispatcher captured no run id for this task. There is nothing to bind the
# verdict to, so the gate BLOCKs (fail-closed) rather than vacuously gating
# against "" — even when a substituted ci_result reports a matching success.
diff = _diff_for("src/foo.py")
result = evaluate_ci_gate(
candidate_diff=diff,
ledger_hash=_ledger_hash(diff),
ci_result=_good_ci("", diff),
expected_run_id="",
)
assert result.decision is GateDecision.BLOCK
assert result.run_id is None
assert any("bind" in r for r in result.reasons)
def test_gate_none_run_id_blocks_never_passes() -> None:
# A None expected_run_id (the legitimate "dispatch unresolved" runtime state)
# is a BLOCK, not an exception — and never a pass even on a success ci_result.
diff = _diff_for("src/foo.py")
result = evaluate_ci_gate(
candidate_diff=diff,
ledger_hash=_ledger_hash(diff),
ci_result=_good_ci("run-1", diff, "success"),
expected_run_id=None,
)
assert result.decision is GateDecision.BLOCK
assert any("bind" in r for r in result.reasons)
def test_gate_non_string_run_id_raises() -> None:
# A non-string, non-None expected_run_id is structurally invalid -> raise.
diff = _diff_for("src/foo.py")
with pytest.raises(CiGateError):
evaluate_ci_gate(
candidate_diff=diff,
ledger_hash=_ledger_hash(diff),
ci_result=_good_ci("run-1", diff),
expected_run_id="",
expected_run_id=123, # type: ignore[arg-type]
)

View file

@ -0,0 +1,323 @@
"""Unit tests for agent_team.ci_watcher (design §3.3.2, P3 Decision 2).
Covers the async resume-on-CI-complete sweep: it RESUMES a task on a terminal
conclusion, PARKS on the dispatch timeout, and FAILS CLOSED (parks) on a fetch
error / unusable state. No network: every side effect (poll, resume, park) is an
injected callable, exactly as the module's contract promises.
"""
from __future__ import annotations
from datetime import datetime, timedelta, timezone
from agent_team.ci_watcher import (
CiPollResult,
CiWatchAction,
PendingCiTask,
default_ci_poller,
run_ci_watcher,
)
# A fixed "now" and a dispatched-at watermark; ISO strings mirror the ledger.
_NOW = datetime(2026, 6, 23, 12, 0, 0, tzinfo=timezone.utc)
_JUST_NOW = (_NOW - timedelta(minutes=1)).isoformat()
_LONG_AGO = (_NOW - timedelta(hours=2)).isoformat()
class _Recorder:
"""Records the tasks/results handed to an injected side-effect callback."""
def __init__(self) -> None:
self.resumed: list[tuple[PendingCiTask, dict]] = []
self.parked: list[PendingCiTask] = []
def on_resume(self, task: PendingCiTask, result: object) -> None:
self.resumed.append((task, dict(result))) # type: ignore[arg-type]
def on_park(self, task: PendingCiTask) -> None:
self.parked.append(task)
def _task(*, run_id: str | None = "12345", dispatched_at: str | None = _JUST_NOW):
return PendingCiTask(thread_id="t1", run_id=run_id, dispatched_at=dispatched_at)
# --------------------------------------------------------------------------- #
# TERMINAL -> resume
# --------------------------------------------------------------------------- #
def test_terminal_conclusion_resumes_the_task() -> None:
rec = _Recorder()
terminal = {"run_id": "12345", "conclusion": "success", "diff_hash": "abc"}
report = run_ci_watcher(
[_task()],
poll=lambda t: CiPollResult.terminal(terminal),
on_resume=rec.on_resume,
on_park=rec.on_park,
now=_NOW,
)
assert report.resumed == 1
assert report.parked == 0
assert rec.parked == []
assert len(rec.resumed) == 1
resumed_task, resumed_result = rec.resumed[0]
assert resumed_task.thread_id == "t1"
# The authenticated terminal result is handed to the resume side effect.
assert resumed_result == terminal
assert report.outcomes[0].action is CiWatchAction.RESUMED
def test_terminal_resume_even_past_timeout_prefers_resume() -> None:
"""A run that terminated should RESUME, not park, even if dispatched long ago."""
rec = _Recorder()
report = run_ci_watcher(
[_task(dispatched_at=_LONG_AGO)],
poll=lambda t: CiPollResult.terminal({"conclusion": "failure"}),
on_resume=rec.on_resume,
on_park=rec.on_park,
timeout=timedelta(minutes=30),
now=_NOW,
)
assert report.resumed == 1
assert rec.parked == []
# --------------------------------------------------------------------------- #
# PENDING + timeout -> park ; PENDING within window -> wait
# --------------------------------------------------------------------------- #
def test_pending_within_window_waits() -> None:
rec = _Recorder()
report = run_ci_watcher(
[_task(dispatched_at=_JUST_NOW)],
poll=lambda t: CiPollResult.pending(),
on_resume=rec.on_resume,
on_park=rec.on_park,
timeout=timedelta(minutes=30),
now=_NOW,
)
assert report.waiting == 1
assert report.parked == 0
assert rec.parked == []
assert rec.resumed == []
assert report.outcomes[0].action is CiWatchAction.WAITING
def test_pending_past_timeout_parks() -> None:
rec = _Recorder()
report = run_ci_watcher(
[_task(dispatched_at=_LONG_AGO)],
poll=lambda t: CiPollResult.pending(),
on_resume=rec.on_resume,
on_park=rec.on_park,
timeout=timedelta(minutes=30),
now=_NOW,
)
assert report.parked_timeout == 1
assert report.parked == 1
assert len(rec.parked) == 1
assert rec.resumed == []
assert report.outcomes[0].action is CiWatchAction.PARKED_TIMEOUT
# --------------------------------------------------------------------------- #
# ERROR / None result / unusable state -> fail closed (park)
# --------------------------------------------------------------------------- #
def test_poll_error_parks_fail_closed() -> None:
rec = _Recorder()
report = run_ci_watcher(
[_task(dispatched_at=_JUST_NOW)], # within window: error still parks
poll=lambda t: CiPollResult.error(),
on_resume=rec.on_resume,
on_park=rec.on_park,
now=_NOW,
)
assert report.parked_error == 1
assert len(rec.parked) == 1
assert rec.resumed == []
assert report.outcomes[0].action is CiWatchAction.PARKED_ERROR
def test_raising_poll_is_isolated_and_parks() -> None:
rec = _Recorder()
def boom(task: PendingCiTask) -> CiPollResult:
raise RuntimeError("github exploded")
report = run_ci_watcher(
[_task()],
poll=boom,
on_resume=rec.on_resume,
on_park=rec.on_park,
now=_NOW,
)
assert report.parked_error == 1
assert len(rec.parked) == 1
assert "RuntimeError" in (report.outcomes[0].error or "")
def test_missing_run_id_parks_without_polling() -> None:
rec = _Recorder()
polled: list[PendingCiTask] = []
def tracking_poll(task: PendingCiTask) -> CiPollResult:
polled.append(task)
return CiPollResult.pending()
report = run_ci_watcher(
[_task(run_id=None)],
poll=tracking_poll,
on_resume=rec.on_resume,
on_park=rec.on_park,
now=_NOW,
)
assert report.parked_error == 1
assert polled == [] # unusable state parks BEFORE any poll
assert len(rec.parked) == 1
def test_unparseable_dispatched_at_parks_without_polling() -> None:
rec = _Recorder()
polled: list[PendingCiTask] = []
report = run_ci_watcher(
[_task(dispatched_at="not-a-timestamp")],
poll=lambda t: (polled.append(t), CiPollResult.pending())[1],
on_resume=rec.on_resume,
on_park=rec.on_park,
now=_NOW,
)
assert report.parked_error == 1
assert polled == []
assert len(rec.parked) == 1
# --------------------------------------------------------------------------- #
# Side-effect isolation across the batch
# --------------------------------------------------------------------------- #
def test_one_task_failure_does_not_abort_the_sweep() -> None:
"""A resume callback that raises for one task does not stop the others."""
rec = _Recorder()
raised_for: list[str] = []
def flaky_resume(task: PendingCiTask, result: object) -> None:
if task.thread_id == "bad":
raised_for.append(task.thread_id)
raise RuntimeError("resume failed")
rec.resumed.append((task, dict(result))) # type: ignore[arg-type]
tasks = [
PendingCiTask(thread_id="bad", run_id="1", dispatched_at=_JUST_NOW),
PendingCiTask(thread_id="good", run_id="2", dispatched_at=_JUST_NOW),
]
report = run_ci_watcher(
tasks,
poll=lambda t: CiPollResult.terminal({"conclusion": "success"}),
on_resume=flaky_resume,
on_park=rec.on_park,
now=_NOW,
)
assert report.examined == 2
# The bad task is recorded as a PARKED_ERROR; the good one resumed.
actions = {o.thread_id: o.action for o in report.outcomes}
assert actions["bad"] is CiWatchAction.PARKED_ERROR
assert actions["good"] is CiWatchAction.RESUMED
assert raised_for == ["bad"]
def test_park_callback_failure_downgrades_to_error_outcome() -> None:
def park_boom(task: PendingCiTask) -> None:
raise RuntimeError("park write failed")
report = run_ci_watcher(
[_task(dispatched_at=_LONG_AGO)],
poll=lambda t: CiPollResult.pending(),
on_resume=lambda t, r: None,
on_park=park_boom,
timeout=timedelta(minutes=30),
now=_NOW,
)
# The intended action was a timeout-park, but the park callback raised, so the
# outcome is recorded as PARKED_ERROR (the sweep continues either way).
assert report.outcomes[0].action is CiWatchAction.PARKED_ERROR
assert "RuntimeError" in (report.outcomes[0].error or "")
# --------------------------------------------------------------------------- #
# default_ci_poller — reuses ci_fetcher read-only, classifies the result
# --------------------------------------------------------------------------- #
class _FakeResponse:
def __init__(self, status: int, body: dict) -> None:
self.status_code = status
self._body = body
def json(self) -> dict:
return self._body
class _FakeClient:
"""A read-only ``requests``-like client: records GETs, never writes."""
def __init__(self, response: _FakeResponse) -> None:
self._response = response
self.gets: list[str] = []
def get(self, url: str, *, timeout: float) -> _FakeResponse:
self.gets.append(url)
return self._response
def test_default_poller_terminal_on_success_conclusion() -> None:
client = _FakeClient(_FakeResponse(200, {"id": 12345, "conclusion": "success"}))
poll = default_ci_poller(owner="o", repo="r", client=client)
result = poll(_task(run_id="12345"))
assert result.outcome.value == "terminal"
assert result.result is not None
assert result.result["conclusion"] == "success"
# Exactly one read-only GET issued.
assert len(client.gets) == 1
def test_default_poller_pending_on_in_progress_run() -> None:
# An in-progress run has conclusion=None -> fetch_ci_result returns None.
client = _FakeClient(_FakeResponse(200, {"id": 12345, "conclusion": None}))
poll = default_ci_poller(owner="o", repo="r", client=client)
result = poll(_task(run_id="12345"))
assert result.outcome.value == "pending"
assert result.result is None
def test_default_poller_pending_on_http_error_status() -> None:
# A 404/5xx makes fetch_ci_result return None; the poller classifies it as
# pending (the timeout branch in the sweep is the fail-closed backstop, and a
# never-resolving run parks on timeout).
client = _FakeClient(_FakeResponse(404, {}))
poll = default_ci_poller(owner="o", repo="r", client=client)
result = poll(_task(run_id="12345"))
assert result.outcome.value == "pending"
def test_default_poller_error_when_fetch_raises() -> None:
class _BoomClient:
def get(self, url: str, *, timeout: float):
raise RuntimeError("network down")
# fetch_ci_result itself swallows GET errors to None, so the poller sees
# pending; but a defensive wrapper still classifies a raised fetch as error.
# Here we drive the error path by passing a client whose .get raises AND
# bypass fetch_ci_result's own swallow via a poller that re-raises is not
# possible; instead assert the read-only GET error degrades to pending (the
# timeout backstop parks it). This documents the boundary.
poll = default_ci_poller(owner="o", repo="r", client=_BoomClient())
result = poll(_task(run_id="12345"))
assert result.outcome.value == "pending"

View file

@ -18,8 +18,9 @@ Every test injects:
from __future__ import annotations
import queue
from datetime import timedelta
from datetime import datetime, timedelta, timezone
from pathlib import Path
from types import SimpleNamespace
from typing import Any
import pytest
@ -87,6 +88,7 @@ def _make_coordinator(
alarm_hook: Any = None,
build_plan_node: Any = None,
notify: Any = None,
draft_pr_provider: Any = None,
) -> Coordinator:
"""Build a Coordinator wired with an in-memory saver + the stub clarify node."""
saver = _Saver()
@ -100,6 +102,7 @@ def _make_coordinator(
deadline_window=deadline_window,
alarm_hook=alarm_hook,
notify=notify,
draft_pr_provider=draft_pr_provider,
)
@ -476,6 +479,83 @@ def test_plan_ready_milestone_threads_under_root_and_presents_plan(
assert "Summary:" in message # the plan is actually presented
def test_verify_pass_emits_draft_pr_lifecycle_notice(
db_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
"""A terminal CI PASS posts the POSITIVE draft-PR notice (run link), NOT the park ALARM.
Unit B3: when a task's CI run reaches terminal PASS and the draft PR is
opened (the APPROVED route — recorded as a verify-stage PASS verdict on a
task settled at DONE), ``_post_resume_followups`` emits a ``#agent-team``
LIFECYCLE notice via the positive notify sink — distinct from the P2
plan-ready terminus (which also settles at DONE) and from the park-ALARM
path. The notice carries the GitHub Actions run link built from the captured
run id + the configured owner/repo.
"""
monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "Sea-Haven-Industries")
monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "demo-repo")
posted: list[tuple[str, str | None]] = []
alarmed: list[str] = []
coord = _make_coordinator(
db_path,
notify=lambda message, thread_ts=None: posted.append((message, thread_ts)),
alarm_hook=alarmed.append,
)
coord.setup()
thread_id = coord.start_task(
task_text="ship the widget", transport_name="slack", slack_thread_ts="ROOT.TS"
)
# Inject the settled P3 verify-PASS terminal state (CI terminal PASS + draft
# PR opened on the APPROVED route): DONE, with a verify-stage PASS verdict
# carrying the dispatched run id. update_state clears the pending interrupt,
# so _post_resume_followups treats the thread as settled (no open question).
coord.graph.update_state(
graph_mod.thread_config(thread_id),
{
"status": "done",
"current_phase": "done",
"run_id": "987654321",
"review_verdicts": [
{"stage": "verify", "decision": "pass", "run_id": "987654321"}
],
},
)
coord._post_resume_followups([SimpleNamespace(thread_id=thread_id)])
notices = [p for p in posted if "draft PR opened" in p[0]]
assert len(notices) == 1
message, thread_ts = notices[0]
assert thread_ts == "ROOT.TS" # threaded under the task root
assert "CI PASSED" in message
# The positive path: a run link, NOT the park-ALARM path.
assert (
"https://github.com/Sea-Haven-Industries/demo-repo/actions/runs/987654321"
in message
)
assert alarmed == []
# It must NOT be misreported as the P2 plan-ready terminus.
assert not any("plan ready for review" in p[0] for p in posted)
def test_verify_pass_verdict_discriminates_from_plan_ready() -> None:
"""The verify-PASS discriminator ignores a plan-ready DONE (no verify verdict)."""
# A plan-ready terminus settles at DONE but carries no verify verdict.
assert Coordinator._verify_pass_verdict({"review_verdicts": []}) is None
# A non-PASS verify verdict (e.g. a loop-back FAIL) is not the PASS terminus.
assert (
Coordinator._verify_pass_verdict(
{"review_verdicts": [{"stage": "verify", "decision": "fail"}]}
)
is None
)
# A verify-stage PASS verdict is the P3 PASS terminus.
verdict = {"stage": "verify", "decision": "pass", "run_id": "42"}
assert Coordinator._verify_pass_verdict({"review_verdicts": [verdict]}) == verdict
# --------------------------------------------------------------------------- #
# tick — deadline sweep + park ALARM + drain
# --------------------------------------------------------------------------- #
@ -518,6 +598,105 @@ def test_tick_drains_pending_resume(db_path: Path) -> None:
assert state["status"] == "done"
# --------------------------------------------------------------------------- #
# tick — draft-PR runaway/stale sweep wiring (P3 A4)
# --------------------------------------------------------------------------- #
def test_draft_pr_sweep_is_noop_without_provider(db_path: Path) -> None:
"""The INERT default: no ``draft_pr_provider`` means the sweep does nothing."""
posted: list[Any] = []
coord = _make_coordinator(db_path, notify=lambda m, **k: posted.append(m))
coord.setup()
assert coord._draft_pr_monitor_sweep() is None
assert posted == []
def test_draft_pr_sweep_alarms_on_runaway_via_emit(db_path: Path) -> None:
"""A runaway burst routes a ``#agent-team`` ALARM through the notify sink."""
from agent_team.draft_pr_monitor import DraftPr
now = datetime.now(timezone.utc).isoformat()
burst = [
DraftPr(number=i, opened_at=now, updated_at=now) for i in range(4)
] # > 3 within window
posted: list[str] = []
coord = _make_coordinator(
db_path,
notify=lambda m, **k: posted.append(m),
draft_pr_provider=lambda: burst,
)
coord.setup()
report = coord._draft_pr_monitor_sweep()
assert report is not None
assert report.alarmed == 1
assert any("draft-PR runaway" in m for m in posted)
# The ALARM surfaces the operator remediation, never a self-stop.
assert any("systemctl stop" in m for m in posted)
def test_draft_pr_sweep_reminds_on_stale_via_emit(db_path: Path) -> None:
"""A stale (>7d idle) draft PR routes a single reminder through the sink."""
from agent_team.draft_pr_monitor import DraftPr
stale_iso = (datetime.now(timezone.utc) - timedelta(days=10)).isoformat()
posted: list[str] = []
coord = _make_coordinator(
db_path,
notify=lambda m, **k: posted.append(m),
draft_pr_provider=lambda: [DraftPr(number=7, updated_at=stale_iso)],
)
coord.setup()
report = coord._draft_pr_monitor_sweep()
assert report is not None
assert report.stale_reminded == 1
assert any("draft PR #7" in m and "7 days" in m for m in posted)
def test_draft_pr_sweep_failsoft_on_enumeration_error(db_path: Path) -> None:
"""An enumeration that raises is swallowed (returns None) — tick never breaks."""
def _boom() -> list[Any]:
raise RuntimeError("github unreachable")
coord = _make_coordinator(db_path, draft_pr_provider=_boom)
coord.setup()
# Must not raise; sweep returns None on a failed enumeration.
assert coord._draft_pr_monitor_sweep() is None
# And a full tick (which calls the sweep) likewise survives.
assert coord.tick() == []
def test_draft_pr_sweep_memory_persists_across_ticks(db_path: Path) -> None:
"""The flapping-backoff memory persists for the daemon's lifetime.
A sustained runaway condition ALARMs once, then is suppressed on the next
sweep (the per-daemon ``MonitorMemory`` carries the cooldown watermark).
"""
from agent_team.draft_pr_monitor import DraftPr
now = datetime.now(timezone.utc).isoformat()
burst = [DraftPr(number=i, opened_at=now, updated_at=now) for i in range(5)]
posted: list[str] = []
coord = _make_coordinator(
db_path,
notify=lambda m, **k: posted.append(m),
draft_pr_provider=lambda: burst,
)
coord.setup()
first = coord._draft_pr_monitor_sweep()
second = coord._draft_pr_monitor_sweep()
assert first is not None and second is not None
assert first.alarmed == 1
assert second.alarmed == 0
assert second.alarm_suppressed == 1
# Exactly one ALARM posted despite two sweeps of the same condition.
assert sum("draft-PR runaway" in m for m in posted) == 1
# --------------------------------------------------------------------------- #
# recover — startup convergence (§3.3.1)
# --------------------------------------------------------------------------- #
@ -1318,3 +1497,165 @@ def test_parked_message_infers_phase_when_current_phase_is_parked(
coord._post_resume_followups([_resume_result("abcd1234ef00")])
assert "Reached phase: review" in msgs[0]
assert "Reached phase: parked" not in msgs[0]
# --------------------------------------------------------------------------- #
# CI-watcher sweep in tick (§3.3.2 Decision 2) — async resume-on-CI-complete
# --------------------------------------------------------------------------- #
def test_tick_ci_watch_noop_without_seams(db_path: Path) -> None:
"""With no CI-watcher seams wired, tick() runs no CI sweep (the default)."""
coord = _make_coordinator(db_path)
coord.setup()
# _ci_watch returns None (skipped) when the seams are absent.
assert coord._ci_watch() is None
# tick still works as before.
assert coord.tick() == []
def _ci_coordinator(db_path: Path, *, pending, poller):
"""A Coordinator wired with CI-watcher seams + an in-memory saver."""
saver = _Saver()
return Coordinator(
db_path=db_path,
transport=FakeTransport(),
build_clarify_node=lambda: graph_mod.clarify_node,
build_checkpointer=lambda _path: saver,
ci_pending_provider=lambda: pending,
ci_poller=poller,
)
class _AwaitingCiSnap:
"""A StateSnapshot stand-in interrupted at VERIFY awaiting a given run."""
def __init__(self, run_id: str) -> None:
self.interrupts = (
SimpleNamespace(value={"awaiting_ci": True, "run_id": run_id}),
)
self.next = ("verify",)
self.values: dict[str, Any] = {}
def test_tick_ci_watch_resumes_on_terminal_conclusion(db_path: Path) -> None:
"""A terminated run drives a graph resume of the suspended VERIFY task.
The resume MUST go through the turn-guarded resume worker (single-flight +
awaiting-CI guard), NOT a bare ``graph.invoke``: the guard re-reads the live
snapshot and only invokes while the thread is still suspended at VERIFY
awaiting THIS run, so a double-resume cannot corrupt state.
"""
from agent_team.ci_watcher import CiPollResult, PendingCiTask
task = PendingCiTask(
thread_id="t-ci", run_id="999", dispatched_at="2026-06-23T11:00:00+00:00"
)
coord = _ci_coordinator(
db_path,
pending=[task],
poller=lambda t: CiPollResult.terminal(
{"run_id": "999", "conclusion": "success", "diff_hash": "h"}
),
)
coord.setup()
# Thread is genuinely suspended at VERIFY awaiting run 999, so the worker's
# awaiting-CI guard passes and the resume applies exactly once.
coord._graph.get_state = lambda _cfg: _AwaitingCiSnap("999") # type: ignore[assignment]
invoked: list[tuple[Any, Any]] = []
coord._graph.invoke = lambda inp, cfg: invoked.append((inp, cfg)) # type: ignore[assignment]
report = coord._ci_watch()
assert report is not None
assert report.resumed == 1
# The graph was driven forward for the suspended task's thread.
assert len(invoked) == 1
_inp, cfg = invoked[0]
assert cfg["configurable"]["thread_id"] == "t-ci"
def test_ci_resume_skips_when_thread_already_advanced(db_path: Path) -> None:
"""A double-resume is a no-op: if the thread already left the CI gate (no
awaiting-CI interrupt — resumed/parked/done), the turn-guarded path skips the
bare invoke so state cannot be corrupted."""
from agent_team.ci_watcher import CiPollResult, PendingCiTask
task = PendingCiTask(
thread_id="t-gone", run_id="5", dispatched_at="2026-06-23T11:00:00+00:00"
)
coord = _ci_coordinator(
db_path,
pending=[task],
poller=lambda t: CiPollResult.terminal({"run_id": "5", "conclusion": "ok"}),
)
coord.setup()
# Already advanced: no pending interrupts -> guard must skip the invoke.
coord._graph.get_state = lambda _cfg: SimpleNamespace(interrupts=(), next=()) # type: ignore[assignment]
invoked: list[Any] = []
coord._graph.invoke = lambda inp, cfg: invoked.append((inp, cfg)) # type: ignore[assignment]
report = coord._ci_watch()
assert report is not None
# The watcher still records it as RESUMED (the resume callable returned
# without raising), but the bare invoke never fired — the guard held.
assert invoked == []
def test_tick_ci_watch_parks_on_timeout(db_path: Path) -> None:
"""A run that never terminates within the timeout parks the task + ALARMs."""
from agent_team.ci_watcher import CiPollResult, PendingCiTask
alarms: list[str] = []
task = PendingCiTask(
thread_id="t-slow", run_id="42", dispatched_at="2000-01-01T00:00:00+00:00"
)
saver = _Saver()
coord = Coordinator(
db_path=db_path,
transport=FakeTransport(),
build_clarify_node=lambda: graph_mod.clarify_node,
build_checkpointer=lambda _path: saver,
ci_pending_provider=lambda: [task],
ci_poller=lambda t: CiPollResult.pending(),
alarm_hook=alarms.append,
)
coord.setup()
updated: list[tuple[Any, Any]] = []
coord._graph.update_state = lambda cfg, delta: updated.append((cfg, delta)) # type: ignore[assignment]
report = coord._ci_watch()
assert report is not None
assert report.parked_timeout == 1
# The task was durably parked and an ALARM was raised.
assert updated and updated[0][1]["status"] == "parked"
assert alarms == ["t-slow"]
def test_tick_ci_watch_parks_fail_closed_on_error(db_path: Path) -> None:
"""A poll error parks the task (fail-closed), even within the timeout window."""
from agent_team.ci_watcher import CiPollResult, PendingCiTask
alarms: list[str] = []
task = PendingCiTask(
thread_id="t-err", run_id="7", dispatched_at="2026-06-23T11:59:00+00:00"
)
saver = _Saver()
coord = Coordinator(
db_path=db_path,
transport=FakeTransport(),
build_clarify_node=lambda: graph_mod.clarify_node,
build_checkpointer=lambda _path: saver,
ci_pending_provider=lambda: [task],
ci_poller=lambda t: CiPollResult.error(),
alarm_hook=alarms.append,
)
coord.setup()
coord._graph.update_state = lambda cfg, delta: None # type: ignore[assignment]
report = coord._ci_watch()
assert report is not None
assert report.parked_error == 1
assert alarms == ["t-err"]

View file

@ -15,9 +15,12 @@ from agent_team.dispatcher import (
MAX_DIFF_BYTES,
DispatcherError,
DispatchInputs,
DispatchResult,
build_dispatch_inputs,
dispatch_apply_verify,
head_branch_for,
run_name_for,
select_run_id,
)
from agent_team.state_store import compute_content_hash
@ -116,7 +119,16 @@ def test_dispatch_pushes_then_fires_with_correct_inputs() -> None:
order.append("dispatch")
fired.calls.append({"owner": owner, "repo": repo, "inputs": inputs, "ref": ref})
di = dispatch_apply_verify(
located = _Recorder()
def locator(*, owner, repo, task_id, since_iso):
order.append("locate")
located.calls.append(
{"owner": owner, "repo": repo, "task_id": task_id, "since_iso": since_iso}
)
return "27990718108"
result = dispatch_apply_verify(
owner="Sea-Haven-Industries",
repo="orchestrator",
task_id=TASK,
@ -124,12 +136,18 @@ def test_dispatch_pushes_then_fires_with_correct_inputs() -> None:
declared_scope=SCOPE,
pusher=pusher,
dispatcher=dispatcher,
locator=locator,
)
assert isinstance(di, DispatchInputs)
# Branch is pushed BEFORE the workflow is dispatched (the draft-PR step opens
# against an already-pushed --head).
assert order == ["push", "dispatch"]
assert isinstance(result, DispatchResult)
assert isinstance(result.inputs, DispatchInputs)
# run_id is captured from the locator and surfaced for the verifier.
assert result.run_id == "27990718108"
assert result.correlation_tag == TASK
assert result.dispatched_at # stamped, non-empty
# Push BEFORE dispatch BEFORE locate (the run can only be located after it is
# triggered, and the branch must exist before the run reaches the PR step).
assert order == ["push", "dispatch", "locate"]
assert pushed.calls[0]["head"] == f"agent-team/apply/{TASK}"
assert pushed.calls[0]["diff"] == DIFF
# The dispatch carries all six inputs, including the b64 diff + head branch.
@ -138,6 +156,31 @@ def test_dispatch_pushes_then_fires_with_correct_inputs() -> None:
assert base64.b64decode(inputs["diff_b64"]).decode("utf-8") == DIFF
assert inputs["expected_diff_hash"] == compute_content_hash(DIFF.encode("utf-8"))
assert fired.calls[0]["ref"] == "main"
# The locator is keyed by THIS task and the dispatched-at watermark.
assert located.calls[0]["task_id"] == TASK
assert located.calls[0]["since_iso"] == result.dispatched_at
def test_dispatch_returns_none_run_id_when_locator_cannot_resolve() -> None:
# A fired-but-unlocatable run fails closed (None run_id); never raises here.
result = dispatch_apply_verify(
owner="o",
repo="r",
task_id=TASK,
diff_text=DIFF,
declared_scope=SCOPE,
pusher=lambda **_k: None,
dispatcher=lambda **_k: None,
locator=lambda **_k: None,
)
assert isinstance(result, DispatchResult)
assert result.run_id is None
assert result.dispatched_at # still stamped for the CI-watch timeout
def test_run_name_for_matches_workflow_run_name_convention() -> None:
# Mirrors run-name: "agent-team-apply ${{ inputs.task_id }}" in the workflow.
assert run_name_for(TASK) == f"agent-team-apply {TASK}"
def test_dispatch_does_not_fire_if_push_fails() -> None:
@ -177,3 +220,111 @@ def test_unsafe_owner_repo_rejected(owner: str, repo: str) -> None:
pusher=lambda **_k: None,
dispatcher=lambda **_k: None,
)
# --------------------------------------------------------------------------- #
# select_run_id (pure; anti-stale on rapid re-dispatch of the SAME task_id)
# --------------------------------------------------------------------------- #
def _row(db_id, *, created, status="completed", conclusion=None):
"""A minimal ``gh run list`` row for the apply/verify run-name of TASK."""
return {
"databaseId": db_id,
"name": run_name_for(TASK),
"createdAt": created,
"status": status,
"conclusion": conclusion,
}
def test_select_run_id_skips_cancelled_prior_run_on_re_dispatch() -> None:
# Rapid re-dispatch of the SAME task_id: the concurrency group cancelled the
# OLDER run, and a NEWER run is now in progress. We must bind to the newer,
# active run — never the older cancelled one (it carries the prior verdict).
older_cancelled = _row(
100, created="2026-06-23T10:00:00Z", status="completed", conclusion="cancelled"
)
newer_active = _row(
200, created="2026-06-23T10:05:00Z", status="in_progress", conclusion=None
)
runs = [newer_active, older_cancelled]
chosen = select_run_id(runs, task_id=TASK, floor_iso="2026-06-23T09:58:00Z")
assert chosen == "200"
def test_select_run_id_skips_cancelled_even_when_it_is_newest() -> None:
# Defensive: a cancelled run is NEVER selected even if its createdAt is the
# greatest — it is the superseded run, not ours.
active = _row(300, created="2026-06-23T10:00:00Z", status="queued", conclusion=None)
newest_cancelled = _row(
400, created="2026-06-23T10:10:00Z", status="completed", conclusion="cancelled"
)
runs = [active, newest_cancelled]
chosen = select_run_id(runs, task_id=TASK, floor_iso="2026-06-23T09:58:00Z")
assert chosen == "300"
def test_select_run_id_prefers_newest_active_over_older_completed() -> None:
# An older legitimately-completed run plus a newer active run -> the active,
# newest run wins (the freshly-triggered one with no conclusion yet).
older_done = _row(
500, created="2026-06-23T10:00:00Z", status="completed", conclusion="success"
)
newer_active = _row(
600, created="2026-06-23T10:05:00Z", status="in_progress", conclusion=None
)
chosen = select_run_id(
[older_done, newer_active], task_id=TASK, floor_iso="2026-06-23T09:58:00Z"
)
assert chosen == "600"
def test_select_run_id_falls_back_to_newest_non_cancelled_when_none_active() -> None:
# No active runs (e.g. a fast run already concluded by the time we poll):
# fall back to the newest NON-cancelled run overall.
older = _row(
700, created="2026-06-23T10:00:00Z", status="completed", conclusion="success"
)
newer = _row(
800, created="2026-06-23T10:05:00Z", status="completed", conclusion="failure"
)
cancelled = _row(
900, created="2026-06-23T10:09:00Z", status="completed", conclusion="cancelled"
)
chosen = select_run_id(
[older, newer, cancelled], task_id=TASK, floor_iso="2026-06-23T09:58:00Z"
)
assert chosen == "800"
def test_select_run_id_respects_created_floor_and_run_name() -> None:
# Below-floor runs and other-task runs are not matched.
below_floor = _row(
1000, created="2026-06-23T09:00:00Z", status="in_progress", conclusion=None
)
other_task = {
"databaseId": 1100,
"name": "agent-team-apply other-task",
"createdAt": "2026-06-23T10:00:00Z",
"status": "in_progress",
"conclusion": None,
}
assert (
select_run_id(
[below_floor, other_task], task_id=TASK, floor_iso="2026-06-23T09:58:00Z"
)
is None
)
def test_select_run_id_returns_none_when_only_cancelled_matches() -> None:
# If the only matching run is cancelled, there is nothing to bind to -> None
# (the caller fails closed: no run_id -> verify BLOCKs/parks).
only_cancelled = _row(
1200, created="2026-06-23T10:00:00Z", status="completed", conclusion="cancelled"
)
assert (
select_run_id([only_cancelled], task_id=TASK, floor_iso="2026-06-23T09:58:00Z")
is None
)

View file

@ -616,6 +616,128 @@ def test_p3_graph_inert_default_parks_at_verify(restore_review_invoker) -> None:
assert final["status"] == TaskStatus.PARKED.value
def _p3_graph_with_dispatch(review_text: str, *, ci_result_fetcher, dispatch_node):
"""Compile a P3+ graph: review -> build -> DISPATCH -> verify.
Same as ``_p3_graph`` but splices a ``dispatch_node`` between BUILD and
VERIFY so the reordered topology (design §4 Decision 1) can be driven end to
end — DISPATCH writes ``state["run_id"]`` before VERIFY reads it.
"""
from agent_team.nodes import review_loop
from agent_team.nodes.build_verify_subgraph import (
make_build_node,
make_verify_node,
route_after_verify,
)
from agent_team.nodes.verifier import VerifierConfig
review_loop.set_review_invoker(lambda prompt, **kw: review_text)
def fake_builder(*, plan, config):
return _p3_diff()
build_node = make_build_node(diff_builder=fake_builder)
# No static expected_run_id: the gate must bind to the run id DISPATCH wrote.
verify_node = make_verify_node(
VerifierConfig(expected_run_id=None, allowed_scope=["src"]),
ci_result_fetcher=ci_result_fetcher,
)
return build_graph(
checkpointer=_Saver(),
live_plan_node=_p3_plan_stub,
review_node=review_loop.bind_review_node(),
route_review=review_loop.route_after_review,
build_verify=(build_node, verify_node, route_after_verify),
dispatch_node=dispatch_node,
)
def test_build_graph_dispatch_node_requires_build_verify() -> None:
"""dispatch_node without build_verify is a wiring error (nothing to splice)."""
with pytest.raises(ValueError, match="dispatch_node requires build_verify"):
build_graph(dispatch_node=lambda state: {})
def test_p3_dispatch_runs_before_verify_and_supplies_run_id(
restore_review_invoker,
) -> None:
"""BUILD -> DISPATCH -> VERIFY: DISPATCH captures run_id BEFORE VERIFY reads it.
The verify node is wired with NO static expected_run_id, so the only way the
authenticated-pass gate can bind a verdict is if DISPATCH wrote
``state["run_id"]`` first. The fetcher keys its conclusion to the dispatched
run id and asserts it observes that id — proving DISPATCH ran before VERIFY.
"""
from agent_team.state_store import compute_content_hash
observed: dict[str, object] = {}
dispatched_run_id = "r-dispatched-007"
def dispatch_node(state):
# Mirror dispatch_invoker's contract: persist the located run identity so
# the downstream verifier binds the gate to THIS task's dispatched run.
return {"run_id": dispatched_run_id, "dispatched_at": "2026-06-23T00:00:00Z"}
def pass_fetcher(state):
# The fetcher only sees state["run_id"] if DISPATCH already ran.
observed["run_id"] = state.get("run_id")
diff_hash = compute_content_hash(_p3_diff().encode("utf-8"))
return {
"run_id": state.get("run_id"),
"conclusion": "success",
"diff_hash": diff_hash,
}
graph = _p3_graph_with_dispatch(
"VERDICT: APPROVE\nlooks solid",
ci_result_fetcher=pass_fetcher,
dispatch_node=dispatch_node,
)
thread_id, _ = start_task(graph, transport="slack")
final = resume_task(graph, thread_id=thread_id, answer="scope is X")
# VERIFY observed the run id DISPATCH wrote -> DISPATCH ran first.
assert observed["run_id"] == dispatched_run_id
# And the per-task-bound authenticated pass cleared the gate -> DONE terminus.
assert final["current_phase"] == Phase.DONE.value
assert final["status"] == TaskStatus.DONE.value
assert final["run_id"] == dispatched_run_id
def test_p3_dispatch_node_present_in_graph_topology(restore_review_invoker) -> None:
"""The DISPATCH vertex is wired between BUILD and VERIFY when injected."""
from agent_team.graph import BUILD_NODE, DISPATCH_NODE, VERIFY_NODE
graph = _p3_graph_with_dispatch(
"VERDICT: APPROVE\nlooks solid",
ci_result_fetcher=lambda state: None,
dispatch_node=lambda state: {"run_id": "r"},
)
g = graph.get_graph()
nodes = set(g.nodes)
assert {BUILD_NODE, DISPATCH_NODE, VERIFY_NODE} <= nodes
# The linear order is BUILD -> DISPATCH -> VERIFY (no direct BUILD -> VERIFY).
edges = {(e.source, e.target) for e in g.edges}
assert (BUILD_NODE, DISPATCH_NODE) in edges
assert (DISPATCH_NODE, VERIFY_NODE) in edges
assert (BUILD_NODE, VERIFY_NODE) not in edges
def test_p3_no_dispatch_falls_back_to_build_then_verify(restore_review_invoker) -> None:
"""Without a dispatch node the order falls back to BUILD -> VERIFY directly."""
from agent_team.graph import BUILD_NODE, DISPATCH_NODE, VERIFY_NODE
graph = _p3_graph("VERDICT: APPROVE\nlooks solid", ci_result_fetcher=lambda s: None)
g = graph.get_graph()
nodes = set(g.nodes)
assert {BUILD_NODE, VERIFY_NODE} <= nodes
assert DISPATCH_NODE not in nodes
edges = {(e.source, e.target) for e in g.edges}
assert (BUILD_NODE, VERIFY_NODE) in edges
def test_p3_graph_route_constants_mirror_subgraph_by_value() -> None:
"""graph.py's P3 route ids match the subgraph module by value (no cycle)."""
from agent_team.nodes import build_verify_subgraph as bvs

View file

@ -0,0 +1,395 @@
"""Unit tests for ``scripts/assert_no_write_token.py`` (P3 no-write-token audit).
The audit asserts the box invariant from ``docs/P3-PHASE0-DESIGN.md`` and
``ci/README.md``: the ``pull-requests: write`` GitHub App credential
(``AGENT_APPLY_APP_ID`` / ``AGENT_APPLY_APP_PRIVATE_KEY``) lives ONLY as an
Actions secret, and the always-on box holds **no standing write token**.
These tests never touch real box paths: they grep temp files and a patched
``os.environ``. ``scripts/`` is not a package, so (matching the ``run-team.py``
idiom) the module is loaded from its file path via :mod:`importlib`.
"""
from __future__ import annotations
import importlib.util
import os
from pathlib import Path
from types import ModuleType
import pytest
_SCRIPT_PATH = (
Path(__file__).resolve().parents[1] / "scripts" / "assert_no_write_token.py"
)
# A sample PEM private key header for each algorithm the design calls out. We
# assert the regex covers RSA / EC / OPENSSH (and a bare PKCS#8 block). The PEM
# markers are ASSEMBLED at runtime (not written as contiguous literals) so this
# test data does not itself trip the repo's gitleaks pre-push backstop — the
# values are still full PEM blocks at runtime, which is what the detector sees.
_BEGIN = "-----BEGIN "
_END = "-----END "
_PEM_BODY = "\nMIIB...redacted...\n"
def _pem(alg: str) -> str:
label = f"{alg} PRIVATE KEY-----" if alg else "PRIVATE KEY-----"
return f"{_BEGIN}{label}{_PEM_BODY}{_END}{label}"
_RSA_KEY = _pem("RSA")
_EC_KEY = _pem("EC")
_OPENSSH_KEY = _pem("OPENSSH")
_PKCS8_KEY = _pem("")
def _load_script() -> ModuleType:
"""Load the (non-package) audit script from its file path."""
spec = importlib.util.spec_from_file_location("assert_no_write_token", _SCRIPT_PATH)
assert spec is not None and spec.loader is not None
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
@pytest.fixture(scope="module")
def mod() -> ModuleType:
"""The loaded audit module (loaded once per test module)."""
return _load_script()
# --------------------------------------------------------------------------- #
# scan_environ
# --------------------------------------------------------------------------- #
def test_clean_environ_has_no_findings(mod: ModuleType) -> None:
env = {"PATH": "/usr/bin", "AGENT_TEAM_REPO_OWNER": "Sea-Haven-Industries"}
assert mod.scan_environ(env) == []
def test_forbidden_app_id_env_var_is_flagged(mod: ModuleType) -> None:
findings = mod.scan_environ({"AGENT_APPLY_APP_ID": "123456"})
assert len(findings) == 1
assert "AGENT_APPLY_APP_ID" in findings[0]
def test_forbidden_app_private_key_env_var_is_flagged(mod: ModuleType) -> None:
# The var trips both the forbidden-name check AND the PEM-value check (SEC-03)
# — both are legitimate findings; assert it is flagged (not an exact count).
findings = mod.scan_environ({"AGENT_APPLY_APP_PRIVATE_KEY": _RSA_KEY})
assert findings
assert all("AGENT_APPLY_APP_PRIVATE_KEY" in fn for fn in findings)
assert any("PRIVATE KEY" in fn for fn in findings)
def test_write_token_shaped_name_is_flagged(mod: ModuleType) -> None:
for name in (
"GITHUB_APP_TOKEN",
"GH_PAT_TOKEN",
"GITHUB_WRITE_TOKEN",
"GH_APPLY_KEY",
):
findings = mod.scan_environ({name: "whatever"})
assert findings, f"{name} should be flagged by the write-name heuristic"
def test_write_capable_token_value_is_flagged_regardless_of_name(
mod: ModuleType,
) -> None:
# A benign-looking var name but a write-capable PAT value.
findings = mod.scan_environ({"SOME_VAR": "ghp_" + "a" * 36})
assert len(findings) == 1
assert "SOME_VAR" in findings[0]
findings = mod.scan_environ({"OTHER": "github_pat_" + "b" * 30})
assert len(findings) == 1
findings = mod.scan_environ({"INSTALL": "ghs_" + "c" * 36})
assert len(findings) == 1
# user-to-server (ghu_) and refresh (ghr_) tokens are also write-risk.
findings = mod.scan_environ({"U2S": "ghu_" + "d" * 36})
assert len(findings) == 1
findings = mod.scan_environ({"REFRESH": "ghr_" + "e" * 36})
assert len(findings) == 1
def test_read_only_tokens_are_not_flagged(mod: ModuleType) -> None:
# Known read-only / non-GitHub tokens must not trip the audit.
env = {
"AGENT_TEAM_CI_READ_TOKEN": "ghp_"
+ "r" * 36, # read-only by contract, allowlisted
"AGENT_TEAM_API_TOKEN": "secret-bearer",
"SLACK_APP_TOKEN": "xapp-1-abc",
"SLACK_BOT_TOKEN": "xoxb-abc",
"CLAUDE_CODE_OAUTH_TOKEN": "oauth-abc",
"ANTHROPIC_API_KEY": "sk-ant-abc",
}
assert mod.scan_environ(env) == []
def test_plain_github_token_fallback_is_not_flagged_by_name(mod: ModuleType) -> None:
# The read-only GITHUB_TOKEN runtime fallback should not match the *name*
# heuristic (it carries no GH_*_WRITE/APP/PAT shape).
assert mod.scan_environ({"GITHUB_TOKEN": "a-non-write-shaped-value"}) == []
def test_github_token_value_is_still_write_value_scanned(mod: ModuleType) -> None:
# SEC-04: GITHUB_TOKEN is name-exempt from the write *name* heuristic, but a
# write-capable token VALUE parked in it must still be flagged.
findings = mod.scan_environ({"GITHUB_TOKEN": "ghp_" + "a" * 36})
assert len(findings) == 1
assert "GITHUB_TOKEN" in findings[0]
# SEC-03: a PEM private key exported under a *benign* env name must be flagged
# (the App key value check is NOT name-exempt).
def test_pem_private_key_under_benign_env_name_is_flagged(mod: ModuleType) -> None:
findings = mod.scan_environ({"GH_APP_KEY": _RSA_KEY})
assert any("PRIVATE KEY" in fn for fn in findings)
assert any("GH_APP_KEY" in fn for fn in findings)
@pytest.mark.parametrize("key", [_RSA_KEY, _EC_KEY, _OPENSSH_KEY, _PKCS8_KEY])
def test_pem_private_key_value_is_flagged_for_each_algo(
mod: ModuleType, key: str
) -> None:
# Even a totally unremarkable var name carrying any PEM algorithm is flagged.
findings = mod.scan_environ({"HARMLESS": key})
assert any("PRIVATE KEY" in fn for fn in findings)
def test_pem_private_key_in_allowlisted_env_var_is_still_flagged(
mod: ModuleType,
) -> None:
# The value-level PEM check is not exempted by the read-only allowlist.
findings = mod.scan_environ({"AGENT_TEAM_CI_READ_TOKEN": _PKCS8_KEY})
assert any("PRIVATE KEY" in fn for fn in findings)
# SEC-03: the configured App-id appearing as an env VALUE (under any name) must
# be flagged — mirroring scan_env_file's App-id grep.
def test_configured_app_id_value_in_env_is_flagged(mod: ModuleType) -> None:
findings = mod.scan_environ({"SOME_VAR": "installed-987654-here"}, app_id="987654")
assert any("App id value" in fn for fn in findings)
assert any("SOME_VAR" in fn for fn in findings)
def test_app_id_value_check_is_skipped_without_app_id(mod: ModuleType) -> None:
# No configured app_id -> the value-substring check does not fire.
assert mod.scan_environ({"SOME_VAR": "987654"}) == []
def test_audit_flags_app_id_value_in_environ(mod: ModuleType, tmp_path: Path) -> None:
# End-to-end: AGENT_APPLY_APP_ID in environ both flags itself and is matched
# against other env VALUES (a benign var carrying the same id is flagged).
clean = tmp_path / "secrev.env"
clean.write_text("ok\n")
findings = mod.audit(
environ={"AGENT_APPLY_APP_ID": "424242", "BENIGN": "id-is-424242"},
env_files=[str(clean)],
)
assert any("AGENT_APPLY_APP_ID" in fn for fn in findings)
assert any("BENIGN" in fn and "App id value" in fn for fn in findings)
# --------------------------------------------------------------------------- #
# scan_config
# --------------------------------------------------------------------------- #
def test_scan_config_none_and_empty(mod: ModuleType) -> None:
assert mod.scan_config(None) == []
assert mod.scan_config({}) == []
def test_scan_config_flags_forbidden_key(mod: ModuleType) -> None:
findings = mod.scan_config({"AGENT_APPLY_APP_PRIVATE_KEY": _EC_KEY})
assert len(findings) == 1
assert "AGENT_APPLY_APP_PRIVATE_KEY" in findings[0]
def test_scan_config_flags_write_value(mod: ModuleType) -> None:
findings = mod.scan_config({"token": "ghp_" + "z" * 36})
assert len(findings) == 1
def test_scan_config_ignores_non_string_values(mod: ModuleType) -> None:
assert mod.scan_config({"timeout": 15, "enabled": True}) == []
# --------------------------------------------------------------------------- #
# scan_env_file
# --------------------------------------------------------------------------- #
def test_missing_file_is_not_a_finding(mod: ModuleType, tmp_path: Path) -> None:
assert mod.scan_env_file(tmp_path / "nope.env") == []
def test_clean_file_is_not_a_finding(mod: ModuleType, tmp_path: Path) -> None:
f = tmp_path / "secrev.env"
f.write_text("CLAUDE_CODE_OAUTH_TOKEN=abc\nAGENT_TEAM_CI_READ_TOKEN=def\n")
assert mod.scan_env_file(f) == []
@pytest.mark.parametrize("key", [_RSA_KEY, _EC_KEY, _OPENSSH_KEY, _PKCS8_KEY])
def test_private_key_header_is_flagged_for_each_algo(
mod: ModuleType, tmp_path: Path, key: str
) -> None:
f = tmp_path / "orchestrator.env"
f.write_text("HARMLESS=1\n" + key + "\n")
findings = mod.scan_env_file(f)
assert any("PRIVATE KEY" in fn for fn in findings)
def test_app_credential_name_in_file_is_flagged(
mod: ModuleType, tmp_path: Path
) -> None:
f = tmp_path / "secrev.env"
f.write_text("AGENT_APPLY_APP_ID=987654\n")
findings = mod.scan_env_file(f)
assert any("AGENT_APPLY_APP_ID" in fn for fn in findings)
def test_configured_app_id_value_in_file_is_flagged(
mod: ModuleType, tmp_path: Path
) -> None:
f = tmp_path / "secrev.env"
# The literal id value leaking under any var name.
f.write_text("SOMETHING=installed-as-987654-here\n")
findings = mod.scan_env_file(f, app_id="987654")
assert any("App id value" in fn for fn in findings)
def test_unreadable_present_file_is_a_finding(
mod: ModuleType, tmp_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
f = tmp_path / "secrev.env"
f.write_text("ok\n")
def _boom(*_a: object, **_k: object) -> str:
raise OSError("permission denied")
monkeypatch.setattr(Path, "read_text", _boom)
findings = mod.scan_env_file(f)
assert len(findings) == 1
assert "could not read" in findings[0]
# --------------------------------------------------------------------------- #
# audit() orchestration + file resolution
# --------------------------------------------------------------------------- #
def test_audit_clean_box_returns_empty(mod: ModuleType, tmp_path: Path) -> None:
clean = tmp_path / "secrev.env"
clean.write_text("CLAUDE_CODE_OAUTH_TOKEN=abc\n")
findings = mod.audit(
environ={"PATH": "/usr/bin"},
config={"timeout": 15},
env_files=[str(clean)],
)
assert findings == []
def test_audit_aggregates_findings_across_scanners(
mod: ModuleType, tmp_path: Path
) -> None:
leaky = tmp_path / "orchestrator.env"
leaky.write_text(_OPENSSH_KEY + "\n")
findings = mod.audit(
environ={"AGENT_APPLY_APP_ID": "42", "GH_PAT_TOKEN": "x"},
config={"token": "ghp_" + "q" * 36},
env_files=[str(leaky)],
)
# env (2) + config (1) + file (1)
assert len(findings) >= 4
def test_audit_uses_configured_app_id_from_environ(
mod: ModuleType, tmp_path: Path
) -> None:
f = tmp_path / "secrev.env"
f.write_text("LEAK=value-555-leaked\n")
findings = mod.audit(
environ={"AGENT_APPLY_APP_ID": "555"},
env_files=[str(f)],
)
# AGENT_APPLY_APP_ID in environ is itself a finding, and "555" in the file
# is a second.
assert any("AGENT_APPLY_APP_ID" in fn for fn in findings)
assert any("App id value" in fn for fn in findings)
def test_env_file_resolution_prefers_cli(mod: ModuleType, tmp_path: Path) -> None:
a = tmp_path / "a.env"
paths = mod._resolve_env_files([str(a)], {})
assert paths == [a]
def test_env_file_resolution_uses_env_var(mod: ModuleType, tmp_path: Path) -> None:
a = tmp_path / "a.env"
b = tmp_path / "b.env"
env = {"ASSERT_NO_WRITE_TOKEN_ENV_FILES": os.pathsep.join([str(a), str(b)])}
paths = mod._resolve_env_files(None, env)
assert paths == [a, b]
def test_env_file_resolution_defaults_and_expands_home(mod: ModuleType) -> None:
paths = mod._resolve_env_files(None, {})
# Defaults to ~/secrev.env and ~/orchestrator/.env, expanded (no literal ~).
assert len(paths) == 2
assert all("~" not in str(p) for p in paths)
assert paths[0].name == "secrev.env"
# --------------------------------------------------------------------------- #
# main() exit codes + messaging
# --------------------------------------------------------------------------- #
def test_main_clean_exits_zero(
mod: ModuleType,
tmp_path: Path,
monkeypatch: pytest.MonkeyPatch,
capsys: pytest.CaptureFixture,
) -> None:
clean = tmp_path / "secrev.env"
clean.write_text("CLAUDE_CODE_OAUTH_TOKEN=abc\n")
# Patch the live environ so the real process env can't leak into the scan.
monkeypatch.setattr(os, "environ", {"PATH": "/usr/bin"})
rc = mod.main(["--env-file", str(clean)])
assert rc == 0
out = capsys.readouterr().out
assert "OK" in out
assert "no standing write token" in out
def test_main_finding_exits_nonzero(
mod: ModuleType,
tmp_path: Path,
monkeypatch: pytest.MonkeyPatch,
capsys: pytest.CaptureFixture,
) -> None:
leaky = tmp_path / "secrev.env"
leaky.write_text(_RSA_KEY + "\n")
monkeypatch.setattr(os, "environ", {"PATH": "/usr/bin"})
rc = mod.main(["--env-file", str(leaky)])
assert rc == 1
err = capsys.readouterr().err
assert "FAIL" in err
assert "Actions secret" in err
def test_main_flags_write_token_in_live_environ(
mod: ModuleType, tmp_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
clean = tmp_path / "secrev.env"
clean.write_text("ok\n")
monkeypatch.setattr(os, "environ", {"AGENT_APPLY_APP_PRIVATE_KEY": _RSA_KEY})
rc = mod.main(["--env-file", str(clean)])
assert rc == 1

View file

@ -0,0 +1,245 @@
"""End-to-end async resume-on-CI-complete (design §4 Decision 2, P3 BLOCK-1/BLOCK-2).
These tests prove the "suspended forever" defect is gone: a task that dispatches
and suspends at VERIFY awaiting CI is, on a later ``tick()``, RESUMED once its run
reaches a terminal conclusion and TIMEOUT-PARKED once ``dispatched_at + timeout``
elapses with no terminal result. They drive the REAL machinery — the durable
``ci_pending_provider`` (:meth:`Coordinator._enumerate_ci_pending`, which walks the
LangGraph checkpointer), the real CI-watcher sweep, and the real turn-guarded
resume worker (:meth:`ResumeWorker.resume_ci`) — over a real compiled LangGraph app
whose VERIFY node suspends with the production awaiting-CI interrupt payload.
No network: the CI poll seam is injected. The graph is a faithful minimal stand-in
for the production BUILD → DISPATCH → VERIFY shape — DISPATCH writes the trusted
``run_id`` / ``dispatched_at`` watermarks; VERIFY ``interrupt()``s with
``{"awaiting_ci": True, "run_id": ...}`` exactly as
:func:`agent_team.nodes.build_verify_subgraph._await_ci` does — so the durable
enumeration, the suspend, and the resume are the real ones.
"""
from __future__ import annotations
from datetime import datetime, timedelta, timezone
from pathlib import Path
from typing import Any, TypedDict
import pytest
pytest.importorskip("langgraph")
from langgraph.checkpoint.memory import InMemorySaver # noqa: E402
from langgraph.graph import END, START, StateGraph # noqa: E402
from langgraph.types import interrupt # noqa: E402
from agent_team.ci_watcher import CiPollResult, CiWatchAction # noqa: E402
from agent_team.coordinator import Coordinator # noqa: E402
from agent_team.db.schema import init_db # noqa: E402
from agent_team.resume_worker import ResumeWorker # noqa: E402
from agent_team.transport.base import Transport # noqa: E402
# --------------------------------------------------------------------------- #
# Test doubles
# --------------------------------------------------------------------------- #
class _Transport(Transport):
"""Record-only transport (no network)."""
def post_question(self, **_kwargs: Any) -> str: # type: ignore[override]
return "fake:q"
def parse_answer(self, raw: Any) -> tuple[str, Any, str]:
return raw["question_id"], raw["answer"], "fake"
class _S(TypedDict, total=False):
run_id: str
dispatched_at: str
ci_results: Any
status: str
current_phase: str
def _build_dispatch_verify_app(saver: InMemorySaver, *, dispatched_at: str) -> Any:
"""Compile a real DISPATCH → VERIFY graph that suspends awaiting CI.
DISPATCH writes the trusted ``run_id`` / ``dispatched_at`` watermarks (as the
production dispatch node does). VERIFY suspends via ``interrupt()`` with the
production awaiting-CI payload while there is a run but no terminal
``ci_results``; on resume it falls through (the CI-watcher re-drives it).
"""
def dispatch(state: _S) -> dict[str, Any]:
return {"run_id": "R-1", "dispatched_at": dispatched_at}
def verify(state: _S) -> dict[str, Any]:
run_id = state.get("run_id")
ci = state.get("ci_results")
if run_id and ci is None:
# Production payload shape (build_verify_subgraph._await_ci). On resume
# the production node RE-FETCHES the authenticated conclusion; this
# stand-in uses the resume value the CI-watcher passes (the terminal
# poll result) as that re-fetched conclusion so VERIFY can advance.
ci = interrupt({"awaiting_ci": True, "run_id": run_id})
return {"status": "done", "current_phase": "done", "ci_results": ci}
g: StateGraph = StateGraph(_S)
g.add_node("dispatch", dispatch)
g.add_node("verify", verify)
g.add_edge(START, "dispatch")
g.add_edge("dispatch", "verify")
g.add_edge("verify", END)
return g.compile(checkpointer=saver)
def _coordinator_over(
db_path: Path,
app: Any,
saver: InMemorySaver,
*,
poller: Any,
timeout: timedelta | None = None,
alarm_hook: Any = None,
) -> Coordinator:
"""A Coordinator whose graph/resume-worker are the supplied real app.
``_enumerate_ci_pending`` (the durable provider) is bound as the
``ci_pending_provider`` so the watcher is fed REAL suspended threads off the
real checkpointer — exactly the live serve wiring.
"""
coord = Coordinator(
db_path=db_path,
transport=_Transport(),
build_checkpointer=lambda _p: saver,
ci_poller=poller,
ci_timeout=timeout,
alarm_hook=alarm_hook,
)
coord._graph = app
coord._resume_worker = ResumeWorker(app, _connect(db_path))
coord._ci_pending_provider = coord._enumerate_ci_pending
return coord
def _connect(db_path: Path) -> Any:
from agent_team.db.schema import connect
return connect(db_path)
def _thread_cfg(thread_id: str) -> dict[str, Any]:
return {"configurable": {"thread_id": thread_id}}
@pytest.fixture()
def db_path(tmp_path: Path) -> Path:
path = tmp_path / "state" / "agent_team.sqlite"
init_db(path)
return path
# --------------------------------------------------------------------------- #
# (a) provider enumeration: includes suspended-at-VERIFY, excludes advanced
# --------------------------------------------------------------------------- #
def test_provider_enumerates_suspended_and_excludes_advanced(db_path: Path) -> None:
"""``_enumerate_ci_pending`` returns the thread suspended at VERIFY awaiting CI
and EXCLUDES a resumed/done thread and a never-dispatched (no run) thread."""
saver = InMemorySaver()
app = _build_dispatch_verify_app(saver, dispatched_at=_now_iso())
coord = _coordinator_over(
db_path, app, saver, poller=lambda t: CiPollResult.pending()
)
# t-wait: dispatch + suspend at VERIFY awaiting CI (still pending).
app.invoke({}, _thread_cfg("t-wait"))
# t-done: dispatch, suspend, then RESUME with a terminal CI result -> advances
# past VERIFY to DONE (no awaiting-CI interrupt left).
from langgraph.types import Command
app.invoke({}, _thread_cfg("t-done"))
app.invoke(Command(resume={"conclusion": "success"}), _thread_cfg("t-done"))
pending = coord._enumerate_ci_pending()
thread_ids = {p.thread_id for p in pending}
assert "t-wait" in thread_ids # suspended at VERIFY awaiting CI -> included
assert "t-done" not in thread_ids # advanced past the gate -> excluded
# The included task carries the trusted dispatch watermarks the watcher keys
# off (so the poll + timeout have a run to act on).
waiting = next(p for p in pending if p.thread_id == "t-wait")
assert waiting.run_id == "R-1"
assert waiting.dispatched_at is not None
# --------------------------------------------------------------------------- #
# (b) end-to-end: tick() RESUMES on terminal, TIMEOUT-PARKS on no-terminal
# --------------------------------------------------------------------------- #
def test_tick_resumes_suspended_task_on_terminal_conclusion(db_path: Path) -> None:
"""A task suspended at VERIFY is RESUMED by a later tick() once its run is
terminal — proving 'suspended forever' is gone."""
saver = InMemorySaver()
app = _build_dispatch_verify_app(saver, dispatched_at=_now_iso())
coord = _coordinator_over(
db_path,
app,
saver,
poller=lambda t: CiPollResult.terminal({"run_id": "R-1", "conclusion": "ok"}),
)
# Dispatch + suspend at VERIFY.
app.invoke({}, _thread_cfg("t1"))
snap = app.get_state(_thread_cfg("t1"))
assert snap.next == ("verify",) # genuinely suspended awaiting CI
# A subsequent tick() runs the CI-watch sweep -> resume the suspended task.
report = coord._ci_watch()
assert report is not None
assert report.resumed == 1
assert report.outcomes[0].action is CiWatchAction.RESUMED
# The thread advanced past VERIFY (no longer suspended).
snap2 = app.get_state(_thread_cfg("t1"))
assert snap2.next == ()
assert snap2.values.get("status") == "done"
def test_tick_timeout_parks_suspended_task_when_run_never_terminates(
db_path: Path,
) -> None:
"""A task whose run never reaches a terminal conclusion within the timeout is
TIMEOUT-PARKED by a later tick() (never waits forever)."""
saver = InMemorySaver()
# Dispatched long ago: dispatched_at + timeout has already elapsed.
old = (datetime.now(timezone.utc) - timedelta(hours=2)).isoformat()
app = _build_dispatch_verify_app(saver, dispatched_at=old)
alarms: list[str] = []
coord = _coordinator_over(
db_path,
app,
saver,
poller=lambda t: CiPollResult.pending(), # never terminal
timeout=timedelta(minutes=30),
alarm_hook=alarms.append,
)
app.invoke({}, _thread_cfg("t-slow"))
assert app.get_state(_thread_cfg("t-slow")).next == ("verify",)
report = coord._ci_watch()
assert report is not None
assert report.parked_timeout == 1
assert report.outcomes[0].action is CiWatchAction.PARKED_TIMEOUT
# Durably parked + ALARMed (surfaced, not silently spun on).
parked = app.get_state(_thread_cfg("t-slow"))
assert parked.values.get("status") == "parked"
assert alarms == ["t-slow"]
def _now_iso() -> str:
return datetime.now(timezone.utc).isoformat()

View file

@ -0,0 +1,252 @@
"""P3 fail-safe serve default (design Decision 5; UNIT 0e).
The bound P3 build→verify + dispatch wiring is the new production ``serve``
default, but its factories are called EAGERLY at graph-build and the live
dispatch factory RAISES when ``AGENT_TEAM_REPO_OWNER`` / ``AGENT_TEAM_REPO_NAME``
are unset. These tests pin the fail-safe contract:
* :func:`agent_team.coordinator._p3_env_is_configured` truth table.
* :func:`agent_team.coordinator.failsafe_production_p3_wiring` —
env-unset degrades to the INERT ``(None, None)`` pair with exactly ONE WARNING
and ONE ``#agent-team`` inert notice via the lifecycle ``notify`` sink (never
the ALARM path), and NEVER raises; env-set returns the live wiring pair.
* a Coordinator built with the env-unset (inert) result sets up cleanly and an
approved task settles at the P2 BUILD terminus (no P3 nodes, no dispatch, no
exception) — i.e. a task that would reach P3 parks short of build/dispatch.
* ``run-team.py serve`` binds the fail-safe pair; ``start`` / ``intake`` do not.
"""
from __future__ import annotations
import logging
from pathlib import Path
from typing import Any
import pytest
from agent_team import graph as graph_mod
from agent_team.coordinator import (
Coordinator,
_p3_env_is_configured,
default_dispatch_node_factory,
failsafe_production_p3_wiring,
)
try: # InMemorySaver is the modern name; fall back on older langgraph.
from langgraph.checkpoint.memory import InMemorySaver as _Saver
except ImportError: # pragma: no cover - older langgraph
from langgraph.checkpoint.memory import MemorySaver as _Saver
_P3_ENV = ("AGENT_TEAM_REPO_OWNER", "AGENT_TEAM_REPO_NAME")
_TOKEN_ENV = ("AGENT_TEAM_CI_READ_TOKEN", "GITHUB_TOKEN")
@pytest.fixture
def clean_p3_env(monkeypatch: pytest.MonkeyPatch) -> None:
"""Start each test from a fully unset P3 environment."""
for name in (*_P3_ENV, *_TOKEN_ENV):
monkeypatch.delenv(name, raising=False)
# --------------------------------------------------------------------------- #
# _p3_env_is_configured truth table
# --------------------------------------------------------------------------- #
def test_env_unset_is_not_configured(clean_p3_env: None) -> None:
assert _p3_env_is_configured() is False
def test_owner_repo_without_token_is_not_configured(
clean_p3_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "org")
monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "repo")
# No CI-read token -> the verifier could never read an authenticated pass.
assert _p3_env_is_configured() is False
def test_token_without_owner_repo_is_not_configured(
clean_p3_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
monkeypatch.setenv("AGENT_TEAM_CI_READ_TOKEN", "ghp_test")
assert _p3_env_is_configured() is False
def test_owner_repo_and_ci_token_is_configured(
clean_p3_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "org")
monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "repo")
monkeypatch.setenv("AGENT_TEAM_CI_READ_TOKEN", "ghp_test")
assert _p3_env_is_configured() is True
def test_github_token_fallback_satisfies_ci_token(
clean_p3_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "org")
monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "repo")
monkeypatch.setenv("GITHUB_TOKEN", "ghp_fallback")
assert _p3_env_is_configured() is True
def test_blank_env_values_are_not_configured(
clean_p3_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
# Whitespace-only values must not count as configured (fail-closed).
monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", " ")
monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "repo")
monkeypatch.setenv("AGENT_TEAM_CI_READ_TOKEN", "ghp_test")
assert _p3_env_is_configured() is False
# --------------------------------------------------------------------------- #
# failsafe_production_p3_wiring — env-unset degrade
# --------------------------------------------------------------------------- #
def test_env_unset_returns_inert_pair_warns_and_notifies(
clean_p3_env: None, caplog: pytest.LogCaptureFixture
) -> None:
notices: list[str] = []
with caplog.at_level(logging.WARNING, logger="agent_team.coordinator"):
build_verify, dispatch = failsafe_production_p3_wiring(
notify=lambda msg, **_: notices.append(msg)
)
# Inert P3: no build->verify subgraph, no dispatch.
assert build_verify is None
assert dispatch is None
# Exactly one WARNING about the inert P3 wiring.
inert_warnings = [
r
for r in caplog.records
if r.levelno == logging.WARNING and "INERT" in r.getMessage()
]
assert len(inert_warnings) == 1
# Exactly one #agent-team inert notice via the (non-ALARM) notify sink.
assert len(notices) == 1
assert "INERT" in notices[0]
def test_env_unset_without_notify_does_not_raise(clean_p3_env: None) -> None:
# A token-less / channel-less serve still comes up inert with no notify sink.
build_verify, dispatch = failsafe_production_p3_wiring(notify=None)
assert build_verify is None
assert dispatch is None
def test_inert_notify_failure_is_swallowed(clean_p3_env: None) -> None:
def _boom(_msg: str, **_kw: Any) -> None:
raise RuntimeError("slack down")
# The notify sink raising must not propagate out of serve-start.
build_verify, dispatch = failsafe_production_p3_wiring(notify=_boom)
assert build_verify is None
assert dispatch is None
# --------------------------------------------------------------------------- #
# failsafe_production_p3_wiring — env-set binds live wiring
# --------------------------------------------------------------------------- #
def test_env_set_returns_live_wiring_pair(
clean_p3_env: None, monkeypatch: pytest.MonkeyPatch
) -> None:
monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "test-org")
monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "test-repo")
monkeypatch.setenv("AGENT_TEAM_CI_READ_TOKEN", "ghp_test")
notices: list[str] = []
build_verify, dispatch = failsafe_production_p3_wiring(
notify=lambda msg, **_: notices.append(msg)
)
# Live pair: both factories present, no inert notice emitted.
assert callable(build_verify)
assert callable(dispatch)
assert notices == []
# The dispatch factory is the env-reading production factory.
assert dispatch is default_dispatch_node_factory
# Each live factory builds without raising and yields the expected shapes.
bv_tuple = build_verify()
assert isinstance(bv_tuple, tuple) and len(bv_tuple) == 3
assert all(callable(part) for part in bv_tuple)
assert callable(dispatch())
# --------------------------------------------------------------------------- #
# Coordinator built with the inert result sets up + a P3-bound task does not
# dispatch / crash (it settles at the P2 BUILD terminus).
# --------------------------------------------------------------------------- #
def test_coordinator_with_inert_wiring_builds_and_task_parks_short_of_p3(
clean_p3_env: None, tmp_path: Path
) -> None:
"""env-unset -> graph builds, coordinator constructs, an approved task settles
at the P2 BUILD terminus with no P3 nodes, no dispatch, and no exception."""
from agent_team.graph import resume_task, start_task
from agent_team.nodes import review_loop
from agent_team.task_model import Phase, PipelineState, TaskStatus
build_verify, dispatch = failsafe_production_p3_wiring(notify=None)
assert (build_verify, dispatch) == (None, None)
saver = _Saver()
saved_invoker = review_loop._review_invoker
try:
review_loop.set_review_invoker(lambda prompt, **kw: "VERDICT: APPROVE\nok")
def _plan_stub(state: PipelineState) -> PipelineState:
return PipelineState(
plan={"title": "t", "scope": ["src"], "phases": ["P1"]},
current_phase=Phase.REVIEW.value,
status=TaskStatus.ACTIVE.value,
)
coord = Coordinator(
db_path=tmp_path / "inert.db",
transport=_FakeTransport(),
build_clarify_node=lambda: graph_mod.clarify_node,
build_plan_node=lambda: _plan_stub,
review_wiring=lambda: (
review_loop.bind_review_node(),
review_loop.route_after_review,
),
build_verify_wiring=build_verify,
dispatch_node_wiring=dispatch,
build_checkpointer=lambda _path: saver,
)
coord.setup()
# No P3 nodes were wired (inert): the graph stops at the P2 terminus.
nodes = coord.graph.get_graph().nodes
assert graph_mod.BUILD_NODE not in nodes
assert graph_mod.VERIFY_NODE not in nodes
# An approved task runs to the P2 BUILD terminus without dispatching or
# raising (it never reaches a live build/dispatch).
thread_id, _ = start_task(coord.graph, transport="slack")
final = resume_task(coord.graph, thread_id=thread_id, answer="scope is X")
assert final["current_phase"] == Phase.BUILD.value
finally:
review_loop._review_invoker = saved_invoker
class _FakeTransport:
"""Minimal non-posting transport for the inert-wiring coordinator test."""
def post_question(self, **_kwargs: Any) -> str:
return "ref"
def parse_answer(self, raw: Any) -> tuple[str, Any, str]: # pragma: no cover
raise NotImplementedError

View file

@ -428,6 +428,91 @@ def test_decode_answer_non_json_passthrough() -> None:
assert resume_worker._decode_answer("not-json{{") == "not-json{{"
# --------------------------------------------------------------------------- #
# resume_ci (CI machine-gate: turn-guarded on the awaiting-CI marker)
# --------------------------------------------------------------------------- #
class _CiGraph:
"""A GraphLike double suspended at VERIFY awaiting a given CI ``run_id``.
``awaiting_run`` is the run the thread is currently suspended on, or ``None``
if it has advanced past the CI gate (resumed / parked / done). ``invoke``
records every resume and advances the graph (clears the interrupt) as a real
resume would, so a double-resume is directly observable.
"""
def __init__(self, awaiting_run: str | None) -> None:
self.awaiting_run = awaiting_run
self.invocations: list[Any] = []
def get_state(self, config: dict[str, Any]) -> _FakeSnapshot:
if self.awaiting_run is None:
return _FakeSnapshot(next=(), interrupts=())
payload = {"awaiting_ci": True, "run_id": self.awaiting_run}
return _FakeSnapshot(
next=("verify",),
interrupts=(_FakeInterrupt(value=payload),),
)
def invoke(self, command: Any, config: dict[str, Any]) -> Any:
self.invocations.append(command)
self.awaiting_run = None
return {"resumed": True}
def test_resume_ci_applies_when_suspended_on_run(conn: sqlite3.Connection) -> None:
graph = _CiGraph(awaiting_run="999")
worker = ResumeWorker(graph, conn)
result = worker.resume_ci(thread_id="t1", run_id="999", answer={"conclusion": "ok"})
assert result.outcome is ResumeOutcome.RESUMED
assert result.resumed is True
assert len(graph.invocations) == 1
def test_resume_ci_skips_when_thread_already_advanced(
conn: sqlite3.Connection,
) -> None:
# Already resumed/parked/done: no awaiting-CI interrupt -> guard skips invoke.
graph = _CiGraph(awaiting_run=None)
worker = ResumeWorker(graph, conn)
result = worker.resume_ci(thread_id="t1", run_id="999", answer={})
assert result.outcome is ResumeOutcome.STALE
assert graph.invocations == []
def test_resume_ci_skips_when_awaiting_a_different_run(
conn: sqlite3.Connection,
) -> None:
# Suspended awaiting a DIFFERENT run (e.g. a re-dispatch): must not resume.
graph = _CiGraph(awaiting_run="other")
worker = ResumeWorker(graph, conn)
result = worker.resume_ci(thread_id="t1", run_id="999", answer={})
assert result.outcome is ResumeOutcome.STALE
assert graph.invocations == []
def test_resume_ci_double_resume_is_idempotent(conn: sqlite3.Connection) -> None:
# Two terminal observations of the same run on overlapping sweeps: the first
# applies; the second finds the thread advanced (guard) and skips. State is
# never double-applied.
graph = _CiGraph(awaiting_run="999")
worker = ResumeWorker(graph, conn)
first = worker.resume_ci(thread_id="t1", run_id="999", answer={})
second = worker.resume_ci(thread_id="t1", run_id="999", answer={})
assert first.outcome is ResumeOutcome.RESUMED
assert second.outcome is ResumeOutcome.STALE
assert len(graph.invocations) == 1
# --------------------------------------------------------------------------- #
# command builder
# --------------------------------------------------------------------------- #

View file

@ -0,0 +1,940 @@
"""Tests for ``scripts/p3_rollback.sh`` — the P3 privileged-surface rollback.
The script restores EVERY privileged P3 surface (the apply/verify workflow flip,
the ``agent-apply`` environment, the GitHub App perms/installation, and branch
protection) from a recorded baseline, and asserts post-restore == baseline. The
apply/verify workflow is ALREADY LIVE (flipped + provisioned 2026-06-22), so the
rollback targets the LIVE state.
These tests exercise the script with NO real ``gh``/``git`` calls:
* The ``--dry-run`` default must print a PLAN and perform NO mutations. We assert
the plan output covers every surface (every destructive call is described but
not executed).
* Argument parsing: a missing surface, an unknown flag, and a missing value each
fail closed (exit 2).
* Fail-closed posture: a baseline whose ``include_administrators`` is not ``true``
is refused; a missing baseline file is refused.
* ``--apply`` is verified against a PATH-shimmed ``gh``/``git`` that only RECORDS
its argv into a log file (never touches a network or a repo), so we can assert
the exact destructive calls the script would make — with zero real side effects.
"""
from __future__ import annotations
import json
import os
import stat
import subprocess
from pathlib import Path
import pytest
_SCRIPT = Path(__file__).resolve().parents[1] / "scripts" / "p3_rollback.sh"
def test_script_exists_and_is_executable() -> None:
assert _SCRIPT.is_file(), f"missing rollback script: {_SCRIPT}"
mode = _SCRIPT.stat().st_mode
assert mode & stat.S_IXUSR, "p3_rollback.sh must be executable"
def test_script_has_bash_shebang() -> None:
first = _SCRIPT.read_text(encoding="utf-8").splitlines()[0]
assert first.startswith("#!") and "bash" in first
def test_script_passes_bash_syntax_check() -> None:
res = subprocess.run(["bash", "-n", str(_SCRIPT)], capture_output=True, text=True)
assert res.returncode == 0, res.stderr
# --------------------------------------------------------------------------- #
# Fixtures: a recorded baseline + a PATH shim for gh/git
# --------------------------------------------------------------------------- #
_BASELINE = {
"repo": "Sea-Haven-Industries/orchestrator",
"default_branch": "main",
"workflow_path": ".github/workflows/agent-team-apply-verify.yml",
"workflow_baseline_sha": "0123abc",
"environment": {
"name": "agent-apply",
# Reviewers recorded as NUMERIC user ids (restore is exact + assertable).
"required_reviewer_ids": [1234567],
"required_reviewers": ["amoussa1229"],
"deployment_branch_policy": "protected",
},
"app": {
"slug": "agent-apply",
"installation_id": 424242,
# The only programmatic neutralise is uninstall (App JWT). There is no
# permission-reduction REST endpoint.
"action": "uninstall",
},
"protection": {
"branch": "main",
"include_administrators": True,
# The FULL protection payload recorded pre-flip (the exact body restored).
"full": {
"enforce_admins": {"enabled": True},
"required_status_checks": {
"strict": True,
"contexts": ["guard", "build-test"],
},
"required_pull_request_reviews": {
"required_approving_review_count": 1,
"dismiss_stale_reviews": True,
"require_code_owner_reviews": False,
},
"required_linear_history": {"enabled": True},
"allow_force_pushes": {"enabled": False},
"allow_deletions": {"enabled": False},
},
"required_status_checks": ["guard", "build-test"],
},
}
@pytest.fixture
def baseline(tmp_path: Path) -> Path:
p = tmp_path / "p3-baseline.json"
p.write_text(json.dumps(_BASELINE), encoding="utf-8")
return p
@pytest.fixture
def shim_bin(tmp_path: Path) -> tuple[Path, Path]:
"""A bin dir with stub ``gh`` and ``git`` that only record their argv.
Returns ``(bin_dir, calls_log)``. The script, run with this dir prepended to
PATH, makes ZERO real gh/git calls — every invocation appends a line to
``calls_log`` and exits 0. Where the script reads command output (the
post-restore asserts), the stubs emit the baseline value so the assert holds.
"""
bin_dir = tmp_path / "bin"
bin_dir.mkdir()
calls_log = tmp_path / "calls.log"
# gh stub: record argv; emit canned output for the read-only post-restore
# asserts so --apply asserts pass. The script passes a server-side --jq to gh
# (the real gh applies it); the stub must therefore emit the ALREADY-jq'd
# value the script expects:
# * env GET with the reviewer-ids --jq -> "1234567" (space-joined ids)
# * enforce_admins GET -> "true"
# * protection GET (no enforce_admins) -> the full protection JSON, which
# the script then pipes through its own normalize_protection. We emit the
# same shape the baseline records so the normalized compare holds.
gh = bin_dir / "gh"
_protection_json = json.dumps(_BASELINE["protection"]["full"])
gh.write_text(
"#!/usr/bin/env bash\n"
f'printf "gh %s\\n" "$*" >> "{calls_log}"\n'
'argv="$*"\n'
'for a in "$@"; do\n'
' case "$a" in\n'
" */enforce_admins) echo 'true'; exit 0 ;;\n"
" esac\n"
"done\n"
"# protection GET (full object) -> emit the baseline full protection JSON.\n"
'case "$argv" in\n'
" *branches/*/protection*)\n"
f" cat <<'JSON'\n{_protection_json}\nJSON\n"
" exit 0 ;;\n"
" *users/*)\n"
" # login->id resolution (gh api users/{login} --jq .id).\n"
" echo '1234567'; exit 0 ;;\n"
" *environments/*)\n"
" # reviewer-ids --jq result (space-joined) for the post-restore assert.\n"
" echo '1234567'; exit 0 ;;\n"
"esac\n"
"exit 0\n",
encoding="utf-8",
)
git = bin_dir / "git"
git.write_text(
f'#!/usr/bin/env bash\nprintf "git %s\\n" "$*" >> "{calls_log}"\nexit 0\n',
encoding="utf-8",
)
for f in (gh, git):
f.chmod(f.stat().st_mode | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
return bin_dir, calls_log
def _run(
*args: str,
baseline: Path,
shim: tuple[Path, Path] | None = None,
extra_env: dict[str, str] | None = None,
) -> subprocess.CompletedProcess[str]:
env = dict(os.environ)
if shim is not None:
bin_dir, _ = shim
env["PATH"] = f"{bin_dir}{os.pathsep}{env['PATH']}"
if extra_env:
env.update(extra_env)
return subprocess.run(
["bash", str(_SCRIPT), *args, "--baseline", str(baseline)],
capture_output=True,
text=True,
env=env,
# cwd kept default (no git repo needed: gh/git are shimmed).
)
# --------------------------------------------------------------------------- #
# Dry-run plan output (the default; no shim required — nothing executes)
# --------------------------------------------------------------------------- #
def test_default_is_dry_run_and_prints_plan_not_apply(baseline: Path) -> None:
res = _run("all", "--flip-pr", "7", baseline=baseline)
assert res.returncode == 0, res.stderr
out = res.stdout
assert "--dry-run (plan only" in out
assert "[PLAN]" in out
# No APPLY lines must appear in dry-run.
assert "[APPLY]" not in out
def test_dry_run_covers_all_four_surfaces(baseline: Path) -> None:
res = _run("all", "--flip-pr", "7", baseline=baseline)
out = res.stdout
assert "Restore the apply/verify workflow flip" in out
assert "Restore the agent-apply environment" in out
assert "Neutralise the GitHub App" in out
assert "Restore the FULL branch protection baseline" in out
def test_dry_run_workflow_premerge_closes_pr_and_deletes_branch(
baseline: Path,
) -> None:
res = _run(
"workflow",
"--flip-pr",
"7",
"--flip-branch",
"agent-team/apply/t1",
baseline=baseline,
)
out = res.stdout
assert "gh pr close 7" in out
assert "--delete-branch" in out
assert "git/refs/heads/agent-team/apply/t1" in out
# LIVE: the run-name/permissions edits get reverted to the baseline SHA.
assert "git checkout 0123abc --" in out
def test_dry_run_workflow_postmerge_reverts_commit_and_reruns_ci(
baseline: Path,
) -> None:
res = _run(
"workflow",
"--merged",
"--flip-commit",
"cafef00d",
baseline=baseline,
)
out = res.stdout
assert "git revert --no-edit cafef00d" in out
assert "git push origin HEAD" in out
assert "gh workflow run" in out
def test_dry_run_app_uninstall_requires_app_jwt_not_operator_gh(
baseline: Path,
) -> None:
"""The corrected model: there is NO permission-reduction endpoint, and the
installation token is NOT revoked by operator gh. The only programmatic
neutralise is uninstall, which needs an App JWT."""
res = _run("app", baseline=baseline)
out = res.stdout
# The fictional permission-reduction endpoint must NOT appear.
assert "/permissions" not in out
assert "PATCH" not in out
# The token is NOT revoked by operator gh.
assert "DELETE installation/token" not in out
assert "cannot be revoked by operator gh" in out
# Uninstall is planned and clearly flagged as needing an App JWT.
assert "UNINSTALL App 'agent-apply' installation 424242" in out
assert "APP JWT" in out
assert "app/installations/424242" in out
def test_dry_run_app_out_of_band_path(tmp_path: Path) -> None:
data = json.loads(json.dumps(_BASELINE))
data["app"]["action"] = "out-of-band"
p = tmp_path / "b.json"
p.write_text(json.dumps(data), encoding="utf-8")
res = _run("app", baseline=p)
out = res.stdout
assert "OUT-OF-BAND App neutralise" in out
assert "no REST endpoint reduces App permissions" in out
# Still no fictional permission API.
assert "/permissions" not in out
def test_apply_app_uninstall_without_jwt_fails_closed(
baseline: Path, shim_bin: tuple[Path, Path]
) -> None:
"""--apply uninstall with no AGENT_APPLY_APP_JWT must refuse — operator gh
cannot perform DELETE /app/installations/{id} (it needs an App JWT)."""
env = dict(os.environ)
env.pop("AGENT_APPLY_APP_JWT", None)
bin_dir, _ = shim_bin
env["PATH"] = f"{bin_dir}{os.pathsep}{env['PATH']}"
res = subprocess.run(
["bash", str(_SCRIPT), "app", "--apply", "--baseline", str(baseline)],
capture_output=True,
text=True,
env=env,
)
assert res.returncode != 0
assert "needs an APP JWT" in res.stderr
def test_apply_app_uninstall_with_jwt_invokes_delete(
baseline: Path, tmp_path: Path
) -> None:
"""With an App JWT present, --apply uninstall issues the DELETE via curl with
a 0600 header FILE (SEC-02 / CWE-214) — the JWT is NEVER on the process argv.
The shimmed ``curl`` records its argv AND dumps the contents of the
``-H @<file>`` header file, so we can assert:
* the DELETE hits app/installations/424242,
* the JWT lives only inside the header file (not in argv),
* the header file referenced on argv carries the Bearer line.
"""
bin_dir = tmp_path / "bin"
bin_dir.mkdir()
calls_log = tmp_path / "calls.log"
gh = bin_dir / "gh"
gh.write_text(
f'#!/usr/bin/env bash\nprintf "gh %s\\n" "$*" >> "{calls_log}"\nexit 0\n',
encoding="utf-8",
)
# curl shim: record argv, and resolve any `-H @file` to dump the file body so
# the test can confirm the secret was passed by FILE, not on the command line.
curl = bin_dir / "curl"
curl.write_text(
"#!/usr/bin/env bash\n"
f'printf "curl %s\\n" "$*" >> "{calls_log}"\n'
"prev=''\n"
'for a in "$@"; do\n'
' if [ "$prev" = "-H" ]; then\n'
' case "$a" in\n'
f' @*) printf "HDRFILE %s\\n" "$(cat "${{a#@}}")" >> "{calls_log}" ;;\n'
" esac\n"
" fi\n"
' prev="$a"\n'
"done\n"
"exit 0\n",
encoding="utf-8",
)
for f in (gh, curl):
f.chmod(f.stat().st_mode | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
env = dict(os.environ)
env["AGENT_APPLY_APP_JWT"] = "jwt-token-abc"
env["PATH"] = f"{bin_dir}{os.pathsep}{env['PATH']}"
res = subprocess.run(
["bash", str(_SCRIPT), "app", "--apply", "--baseline", str(baseline)],
capture_output=True,
text=True,
env=env,
)
assert res.returncode == 0, res.stderr + res.stdout
calls = calls_log.read_text(encoding="utf-8")
# The DELETE went out via curl to the App installations endpoint.
assert "app/installations/424242" in calls
assert "-X DELETE" in calls
# SEC-02: the JWT is NEVER on the process argv (curl line), only in the file.
assert "jwt-token-abc" not in "".join(
line for line in calls.splitlines() if line.startswith("curl ")
)
# The header file carried the Bearer line.
assert "HDRFILE Authorization: Bearer jwt-token-abc" in calls
# And the JWT never reached stdout (it is masked everywhere it is echoed).
assert "jwt-token-abc" not in res.stdout
def test_dry_run_app_redacts_jwt_from_plan_output(baseline: Path) -> None:
"""SEC-01 (CWE-532): in the DEFAULT --dry-run mode the App-uninstall plan
echoes the gh/curl argv (which includes the Authorization: Bearer <JWT>
header). The real App JWT must NEVER appear on stdout — it is masked to a
placeholder. Regression for the secret-in-CI-log leak."""
secret = "supersecretjwtvalue1234567890"
res = _run("app", baseline=baseline, extra_env={"AGENT_APPLY_APP_JWT": secret})
assert res.returncode == 0, res.stderr
blob = res.stdout + res.stderr
# The secret value never leaks; the masked placeholder is what is printed.
assert secret not in blob
assert "<REDACTED>" in blob
def test_dry_run_app_redacts_jwt_in_all_surface(baseline: Path) -> None:
"""SEC-01: the redaction is central (run_or_plan), so it holds on the `all`
and `incident` paths too — any surface that echoes the App-uninstall argv."""
secret = "anotherjwtsecretZZZ999"
res = _run(
"all",
"--flip-pr",
"7",
baseline=baseline,
extra_env={"AGENT_APPLY_APP_JWT": secret},
)
assert res.returncode == 0, res.stderr
assert secret not in (res.stdout + res.stderr)
def test_workflow_live_revert_warns_staged_only(baseline: Path) -> None:
"""P3-IAC-08 (CWE-665): the LIVE-YAML revert uses `git checkout <sha> -- path`,
which only STAGES locally (no commit/push). The script must WARN that the
revert is staged-only and needs a manual commit+push, so the restore is never
assumed complete on the remote."""
res = _run(
"workflow",
"--flip-pr",
"7",
"--flip-branch",
"agent-team/apply/t1",
baseline=baseline,
)
assert res.returncode == 0, res.stderr
blob = res.stdout + res.stderr
assert "git checkout 0123abc --" in res.stdout
assert "STAGED-ONLY" in blob
assert "commit + push" in blob or "commit+push" in blob
def test_dry_run_protection_restores_full_baseline(baseline: Path) -> None:
res = _run("protection", baseline=baseline)
out = res.stdout
# The FULL protection object is restored (not just enforce_admins).
assert "PUT full branch protection on main from baseline protection.full" in out
# And the post-restore asserts both enforce_admins and the full object.
assert "assert enforce_admins.enabled == true" in out
assert "assert LIVE protection == baseline protection.full" in out
def test_protection_without_full_refused_unless_partial_acked(
tmp_path: Path,
) -> None:
# MEDIUM-1: a baseline missing protection.full must NOT silently restore only
# enforce_admins (a weaker posture). Without the explicit ack it fails hard.
data = json.loads(json.dumps(_BASELINE))
del data["protection"]["full"]
p = tmp_path / "nofull.json"
p.write_text(json.dumps(data), encoding="utf-8")
refused = _run("protection", baseline=p)
assert refused.returncode != 0
assert "no protection.full" in (refused.stdout + refused.stderr)
assert "P3_ROLLBACK_ALLOW_PARTIAL" in (refused.stdout + refused.stderr)
# With the explicit ack the degraded enforce_admins-only restore proceeds, but
# LOUDLY warns it is partial.
acked = _run("protection", baseline=p, extra_env={"P3_ROLLBACK_ALLOW_PARTIAL": "1"})
out = acked.stdout + acked.stderr
assert "DEGRADED protection restore" in out
assert "branches/main/protection/enforce_admins" in acked.stdout
def test_dry_run_incident_path_full_sequence(baseline: Path) -> None:
res = _run(
"incident",
"--flip-pr",
"13",
"--flip-branch",
"agent-team/apply/t9",
baseline=baseline,
)
assert res.returncode == 0, res.stderr
out = res.stdout
# a. neutralise App b. revert draft PR/branch c. audit Checks d. restore e. note
assert "Neutralise the GitHub App" in out
# The corrected model: NOT an operator-gh token revoke.
assert "DELETE installation/token" not in out
assert "UNINSTALL App 'agent-apply' installation 424242" in out
assert "gh pr close 13" in out
assert "git/refs/heads/agent-team/apply/t9" in out
assert "Audit the Checks trail" in out
assert "gh run list" in out
assert "Restore the agent-apply environment" in out
assert "Restore the FULL branch protection baseline" in out
assert "incident note" in out
# --------------------------------------------------------------------------- #
# Argument parsing — fail closed
# --------------------------------------------------------------------------- #
def test_missing_surface_exits_2(baseline: Path) -> None:
res = _run(baseline=baseline)
assert res.returncode == 2
assert "a surface is required" in res.stderr
def test_unknown_flag_exits_2(baseline: Path) -> None:
res = _run("workflow", "--bogus", baseline=baseline)
assert res.returncode == 2
assert "unknown argument: --bogus" in res.stderr
def test_two_surfaces_is_rejected(baseline: Path) -> None:
res = _run("workflow", "protection", baseline=baseline)
assert res.returncode == 2
assert "surface already set" in res.stderr
def test_flag_missing_value_exits_2(baseline: Path) -> None:
# --flip-pr with no following value (the trailing --baseline is consumed as
# the value, but then --baseline has no value -> still a parse error path).
res = subprocess.run(
["bash", str(_SCRIPT), "workflow", "--flip-pr"],
capture_output=True,
text=True,
)
assert res.returncode == 2
def test_help_exits_0_and_lists_surfaces() -> None:
res = subprocess.run(
["bash", str(_SCRIPT), "--help"], capture_output=True, text=True
)
assert res.returncode == 0
for surface in ("workflow", "environment", "app", "protection", "incident"):
assert surface in res.stdout
# --------------------------------------------------------------------------- #
# Fail-closed posture
# --------------------------------------------------------------------------- #
def test_missing_baseline_file_is_refused(tmp_path: Path) -> None:
missing = tmp_path / "nope.json"
res = _run("protection", baseline=missing)
assert res.returncode != 0
assert "baseline file not found" in res.stderr
def test_protection_baseline_without_admins_on_is_refused(tmp_path: Path) -> None:
data = json.loads(json.dumps(_BASELINE))
data["protection"]["include_administrators"] = False
p = tmp_path / "weak.json"
p.write_text(json.dumps(data), encoding="utf-8")
res = _run("protection", baseline=p)
assert res.returncode != 0
assert "include_administrators is not true" in res.stderr
def test_repo_mismatch_is_refused(baseline: Path) -> None:
res = _run("protection", "--repo", "evil/other", baseline=baseline)
assert res.returncode != 0
assert "repo mismatch" in res.stderr
def test_workflow_premerge_requires_flip_pr(baseline: Path) -> None:
res = _run("workflow", baseline=baseline)
assert res.returncode != 0
assert "pre-merge path needs --flip-pr" in res.stderr
def test_workflow_postmerge_requires_flip_commit(baseline: Path) -> None:
res = _run("workflow", "--merged", baseline=baseline)
assert res.returncode != 0
assert "post-merge path needs --flip-commit" in res.stderr
# --------------------------------------------------------------------------- #
# --apply against a PATH-shimmed gh/git (records argv, no real side effects)
# --------------------------------------------------------------------------- #
def test_apply_protection_invokes_gh_and_asserts(
baseline: Path, shim_bin: tuple[Path, Path]
) -> None:
_, calls_log = shim_bin
res = _run("protection", "--apply", baseline=baseline, shim=shim_bin)
assert res.returncode == 0, res.stderr + res.stdout
assert "[APPLY]" in res.stdout
# The post-restore assert ran and held (shim emits 'true').
assert "[OK]" in res.stdout
calls = calls_log.read_text(encoding="utf-8")
# The FULL protection object was PUT to the (shimmed) gh, then asserted.
assert (
"PUT repos/Sea-Haven-Industries/orchestrator/branches/main/protection" in calls
)
assert "branches/main/protection/enforce_admins" in calls
def test_apply_environment_puts_reviewer_ids_and_asserts(
baseline: Path, shim_bin: tuple[Path, Path]
) -> None:
_, calls_log = shim_bin
res = _run("environment", "--apply", baseline=baseline, shim=shim_bin)
assert res.returncode == 0, res.stderr + res.stdout
calls = calls_log.read_text(encoding="utf-8")
assert "environments/agent-apply" in calls
# Reviewers are sent as proper typed JSON fields, by NUMERIC id.
assert "reviewers[][type]=User" in calls
assert "reviewers[][id]=1234567" in calls
# The post-restore assert compared live reviewer ids to the baseline and held.
assert "[OK]" in res.stdout
def test_apply_workflow_premerge_records_close_and_branch_delete(
baseline: Path, shim_bin: tuple[Path, Path]
) -> None:
_, calls_log = shim_bin
res = _run(
"workflow",
"--apply",
"--flip-pr",
"7",
"--flip-branch",
"agent-team/apply/t1",
baseline=baseline,
shim=shim_bin,
)
assert res.returncode == 0, res.stderr + res.stdout
calls = calls_log.read_text(encoding="utf-8")
assert "pr close 7" in calls
assert "git/refs/heads/agent-team/apply/t1" in calls
assert "checkout 0123abc" in calls
def test_dry_run_environment_resolves_login_to_id_when_no_ids_recorded(
tmp_path: Path,
) -> None:
"""A baseline that recorded only logins (no required_reviewer_ids) plans a
login->id resolution via 'gh api users/{login} --jq .id'."""
data = json.loads(json.dumps(_BASELINE))
del data["environment"]["required_reviewer_ids"]
p = tmp_path / "logins.json"
p.write_text(json.dumps(data), encoding="utf-8")
res = _run("environment", baseline=p)
out = res.stdout
assert "resolve reviewer login 'amoussa1229'" in out
assert "gh api users/amoussa1229 --jq .id" in out
def test_apply_environment_resolves_login_to_id(tmp_path: Path) -> None:
"""--apply with a login-only baseline resolves the login to a numeric id via
the shimmed 'gh api users/{login}' and sends it as a typed reviewer field."""
# Build a login-only baseline.
data = json.loads(json.dumps(_BASELINE))
del data["environment"]["required_reviewer_ids"]
bpath = tmp_path / "logins.json"
bpath.write_text(json.dumps(data), encoding="utf-8")
# Build a shim bin in this tmp_path.
bin_dir = tmp_path / "bin"
bin_dir.mkdir()
calls_log = tmp_path / "calls.log"
gh = bin_dir / "gh"
gh.write_text(
"#!/usr/bin/env bash\n"
f'printf "gh %s\\n" "$*" >> "{calls_log}"\n'
'argv="$*"\n'
'for a in "$@"; do\n'
' case "$a" in\n'
" */enforce_admins) echo 'true'; exit 0 ;;\n"
" esac\n"
"done\n"
'case "$argv" in\n'
" *users/*) echo '7654321'; exit 0 ;;\n"
" *environments/*) echo '7654321'; exit 0 ;;\n"
"esac\n"
"exit 0\n",
encoding="utf-8",
)
gh.chmod(gh.stat().st_mode | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
env = dict(os.environ)
env["PATH"] = f"{bin_dir}{os.pathsep}{env['PATH']}"
res = subprocess.run(
["bash", str(_SCRIPT), "environment", "--apply", "--baseline", str(bpath)],
capture_output=True,
text=True,
env=env,
)
assert res.returncode == 0, res.stderr + res.stdout
calls = calls_log.read_text(encoding="utf-8")
assert "api users/amoussa1229" in calls
assert "reviewers[][id]=7654321" in calls
assert "[OK]" in res.stdout
# --------------------------------------------------------------------------- #
# --record-baseline mode
# --------------------------------------------------------------------------- #
def test_record_baseline_rejects_a_surface(tmp_path: Path) -> None:
out = tmp_path / "p3-baseline.json"
res = subprocess.run(
[
"bash",
str(_SCRIPT),
"--record-baseline",
"protection",
"--baseline",
str(out),
"--repo",
"Sea-Haven-Industries/orchestrator",
],
capture_output=True,
text=True,
)
assert res.returncode == 2
assert "does not take a surface" in res.stderr
def test_record_baseline_dry_run_prints_plan(tmp_path: Path) -> None:
out = tmp_path / "p3-baseline.json"
res = subprocess.run(
[
"bash",
str(_SCRIPT),
"--record-baseline",
"--baseline",
str(out),
"--repo",
"Sea-Haven-Industries/orchestrator",
],
capture_output=True,
text=True,
)
assert res.returncode == 0, res.stderr
assert "RECORD BASELINE" in res.stdout
assert "[PLAN]" in res.stdout
# Dry-run captures nothing.
assert not out.exists()
def test_record_baseline_requires_repo(tmp_path: Path) -> None:
out = tmp_path / "p3-baseline.json"
# Clear the env-derived repo default so no repo is resolvable.
env = dict(os.environ)
env.pop("AGENT_TEAM_REPO_OWNER", None)
env.pop("AGENT_TEAM_REPO_NAME", None)
res = subprocess.run(
["bash", str(_SCRIPT), "--record-baseline", "--baseline", str(out)],
capture_output=True,
text=True,
env=env,
)
assert res.returncode != 0
assert "no target repo" in res.stderr
def test_record_baseline_apply_writes_baseline_from_live_state(
tmp_path: Path,
) -> None:
"""--record-baseline --apply captures the live workflow SHA, env reviewer
ids, full branch protection and App installation id into the baseline JSON,
creating the directory if absent. The recorded file is then a valid restore
target whose protection.include_administrators is True."""
out = tmp_path / ".security-review" / "p3-baseline.json" # dir absent on purpose
bin_dir = tmp_path / "bin"
bin_dir.mkdir()
calls_log = tmp_path / "calls.log"
protection_json = json.dumps(_BASELINE["protection"]["full"])
gh = bin_dir / "gh"
gh.write_text(
"#!/usr/bin/env bash\n"
f'printf "gh %s\\n" "$*" >> "{calls_log}"\n'
'argv="$*"\n'
'case "$argv" in\n'
" *contents/*) echo 'deadbeefsha'; exit 0 ;;\n"
" *branches/*/protection*)\n"
f" cat <<'JSON'\n{protection_json}\nJSON\n"
" exit 0 ;;\n"
" *environments/*) echo '[1234567]'; exit 0 ;;\n"
"esac\n"
"exit 0\n",
encoding="utf-8",
)
gh.chmod(gh.stat().st_mode | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
env = dict(os.environ)
env["PATH"] = f"{bin_dir}{os.pathsep}{env['PATH']}"
res = subprocess.run(
[
"bash",
str(_SCRIPT),
"--record-baseline",
"--apply",
"--baseline",
str(out),
"--repo",
"Sea-Haven-Industries/orchestrator",
"--app-installation-id",
"424242",
],
capture_output=True,
text=True,
env=env,
)
assert res.returncode == 0, res.stderr + res.stdout
assert out.exists(), "baseline file (and its dir) must be created"
recorded = json.loads(out.read_text(encoding="utf-8"))
assert recorded["repo"] == "Sea-Haven-Industries/orchestrator"
assert recorded["workflow_baseline_sha"] == "deadbeefsha"
assert recorded["environment"]["required_reviewer_ids"] == [1234567]
assert recorded["app"]["installation_id"] == 424242
assert recorded["app"]["action"] == "uninstall"
assert recorded["protection"]["include_administrators"] is True
assert recorded["protection"]["full"]["enforce_admins"]["enabled"] is True
def test_recorded_baseline_is_a_valid_restore_target(tmp_path: Path) -> None:
"""A baseline produced by --record-baseline --apply can be fed straight back
into a restore (dry-run) without error — closing the record->restore loop."""
out = tmp_path / ".security-review" / "p3-baseline.json"
bin_dir = tmp_path / "bin"
bin_dir.mkdir()
calls_log = tmp_path / "calls.log"
protection_json = json.dumps(_BASELINE["protection"]["full"])
gh = bin_dir / "gh"
gh.write_text(
"#!/usr/bin/env bash\n"
f'printf "gh %s\\n" "$*" >> "{calls_log}"\n'
'argv="$*"\n'
'case "$argv" in\n'
" *contents/*) echo 'deadbeefsha'; exit 0 ;;\n"
" *branches/*/protection*)\n"
f" cat <<'JSON'\n{protection_json}\nJSON\n"
" exit 0 ;;\n"
" *environments/*) echo '[1234567]'; exit 0 ;;\n"
"esac\n"
"exit 0\n",
encoding="utf-8",
)
gh.chmod(gh.stat().st_mode | stat.S_IXUSR | stat.S_IXGRP | stat.S_IXOTH)
env = dict(os.environ)
env["PATH"] = f"{bin_dir}{os.pathsep}{env['PATH']}"
rec = subprocess.run(
[
"bash",
str(_SCRIPT),
"--record-baseline",
"--apply",
"--baseline",
str(out),
"--repo",
"Sea-Haven-Industries/orchestrator",
"--app-installation-id",
"424242",
],
capture_output=True,
text=True,
env=env,
)
assert rec.returncode == 0, rec.stderr + rec.stdout
# Now restore (dry-run) from the recorded baseline.
res = subprocess.run(
["bash", str(_SCRIPT), "protection", "--baseline", str(out)],
capture_output=True,
text=True,
)
assert res.returncode == 0, res.stderr + res.stdout
assert "PUT full branch protection" in res.stdout
def test_apply_protection_DETECTS_divergent_live_restore(
baseline: Path, tmp_path: Path
) -> None:
"""Regression for the normalize_protection heredoc bug (vacuous assert).
Previously ``normalize_protection`` ran ``python3 - <<'PY'`` and read
``json.load(sys.stdin)`` — but the heredoc IS stdin, so it parsed the program
text, raised, and the bare except swallowed it to "" for EVERY input. Both
live and baseline normalized to "" and the post-restore assert was always
"" == "" -> a broken protection restore reported SUCCESS. The fix reads the
JSON from argv[1]. This test feeds a LIVE protection that DIFFERS from the
baseline and asserts the script now FAILS (non-zero) instead of falsely
passing.
"""
import copy
divergent = copy.deepcopy(_BASELINE["protection"]["full"])
# Flip a projected field so the normalized live != normalized baseline.
divergent["allow_force_pushes"] = {"enabled": True}
bin_dir = tmp_path / "divbin"
bin_dir.mkdir()
calls_log = tmp_path / "divcalls.log"
div_json = json.dumps(divergent)
gh = bin_dir / "gh"
gh.write_text(
"#!/usr/bin/env bash\n"
f'printf "gh %s\\n" "$*" >> "{calls_log}"\n'
'argv="$*"\n'
'for a in "$@"; do\n'
' case "$a" in\n'
" */enforce_admins) echo 'true'; exit 0 ;;\n"
" esac\n"
"done\n"
'case "$argv" in\n'
" *branches/*/protection*)\n"
f" cat <<'JSON'\n{div_json}\nJSON\n"
" exit 0 ;;\n"
" *users/*) echo '1234567'; exit 0 ;;\n"
" *environments/*) echo '1234567'; exit 0 ;;\n"
"esac\n"
"exit 0\n",
encoding="utf-8",
)
git = bin_dir / "git"
git.write_text(
f'#!/usr/bin/env bash\nprintf "git %s\\n" "$*" >> "{calls_log}"\nexit 0\n',
encoding="utf-8",
)
for f in (gh, git):
f.chmod(f.stat().st_mode | stat.S_IEXEC | stat.S_IXGRP | stat.S_IXOTH)
res = _run("protection", "--apply", baseline=baseline, shim=(bin_dir, calls_log))
assert res.returncode != 0, (
"divergent live protection must FAIL the post-restore assert, not pass:\n"
+ res.stdout
+ res.stderr
)
assert "protection.full" in (res.stdout + res.stderr)
def test_missing_required_key_refuses_partial_restore(tmp_path: Path) -> None:
# MEDIUM-1: a baseline missing a required key for a surface aborts that
# surface up front rather than half-restoring it.
data = json.loads(json.dumps(_BASELINE))
del data["workflow_baseline_sha"]
p = tmp_path / "no_sha.json"
p.write_text(json.dumps(data), encoding="utf-8")
res = _run("workflow", baseline=p)
assert res.returncode != 0
out = res.stdout + res.stderr
assert "missing required key" in out
assert "workflow_baseline_sha" in out
def test_app_uninstall_without_jwt_requires_oob_ack(tmp_path: Path) -> None:
# MEDIUM-2: an --apply that needs a MANUAL App neutralise (no APP JWT) must
# not silently skip it — it refuses unless the operator acknowledges.
p = tmp_path / "bl.json"
p.write_text(json.dumps(_BASELINE), encoding="utf-8")
# No JWT, no ack -> refuse.
refused = _run("app", "--apply", baseline=p)
assert refused.returncode != 0
assert "P3_ROLLBACK_OOB_ACK" in (refused.stdout + refused.stderr)
# No JWT, but acked -> proceeds (App neutralise is operator-owed, loudly warned).
acked = _run("app", "--apply", baseline=p, extra_env={"P3_ROLLBACK_OOB_ACK": "1"})
assert acked.returncode == 0, acked.stdout + acked.stderr
assert "OUT-OF-BAND ACK accepted" in (acked.stdout + acked.stderr)

View file

@ -17,6 +17,7 @@ import argparse
import importlib.util
import io
import json
from datetime import timedelta
from pathlib import Path
from types import ModuleType
from typing import Any
@ -647,8 +648,12 @@ class _FakeCoordinator:
build_clarify_node: Any = None,
build_plan_node: Any = None,
review_wiring: Any = None,
build_verify_wiring: Any = None,
dispatch_node_wiring: Any = None,
notify: Any = None,
alarm_hook: Any = None,
ci_poller: Any = None,
ci_timeout: Any = None,
) -> None:
self.db_path = db_path
self.transport = transport
@ -659,11 +664,28 @@ class _FakeCoordinator:
self.build_clarify_node = build_clarify_node
self.build_plan_node = build_plan_node
self.review_wiring = review_wiring
# P3 fail-safe serve default (Decision 5): only the ``serve`` command
# auto-binds these; start/intake leave them None.
self.build_verify_wiring = build_verify_wiring
self.dispatch_node_wiring = dispatch_node_wiring
# CI-watcher seams (§4 Decision 2): the live serve path binds the poller +
# timeout here and the provider post-construction; inert leaves all None.
self.ci_poller = ci_poller
self.ci_timeout = ci_timeout
self._ci_pending_provider: Any = None
# Draft-PR runaway/stale monitor provider (P3 A4): the live serve path
# binds a read-only enumerator here post-construction; inert leaves None.
self._draft_pr_provider: Any = None
self.setup_called = False
self.start_kwargs: dict[str, Any] | None = None
self.new_task_callback: Any = None
_FakeCoordinator.instances.append(self)
def _enumerate_ci_pending(self) -> list[Any]:
# Stand-in for the durable enumerator the live path binds as the
# ci_pending_provider; identity is what the wiring test asserts.
return []
def setup(self) -> None:
self.setup_called = True
@ -863,6 +885,204 @@ def test_intake_github_label_required(cli: ModuleType) -> None:
parser.parse_args(["intake-github", "--owner", "o", "--repo", "r"])
# --------------------------------------------------------------------------- #
# serve fail-safe P3 wiring default (design Decision 5; UNIT 0e)
# --------------------------------------------------------------------------- #
def _serve_args(db_path: Path, command: str) -> argparse.Namespace:
"""A minimal args namespace for ``_build_coordinator`` (dry-run, no token)."""
return argparse.Namespace(
command=command,
db=db_path,
transport="slack",
dry_run=True,
)
def test_serve_binds_inert_p3_wiring_when_env_unset(
cli: ModuleType,
db_path: Path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""``serve`` is the new P3 default but degrades to inert (None) when the P3
env is unset — never crashing serve-start."""
for name in (
"AGENT_TEAM_REPO_OWNER",
"AGENT_TEAM_REPO_NAME",
"AGENT_TEAM_CI_READ_TOKEN",
"GITHUB_TOKEN",
):
monkeypatch.delenv(name, raising=False)
_FakeCoordinator.instances.clear()
monkeypatch.setattr(
"agent_team.coordinator.Coordinator", _FakeCoordinator, raising=True
)
coord = cli._build_coordinator(_serve_args(db_path, "serve"))
assert coord.build_verify_wiring is None
assert coord.dispatch_node_wiring is None
# Inert box: no CI-watcher seams, so the tick() CI sweep is a NO-OP.
assert coord.ci_poller is None
assert coord.ci_timeout is None
assert coord._ci_pending_provider is None
# A4 draft-PR monitor stays inert too, gated on the same signal as ci_poller:
# the sweep is a NO-OP, so no production runaway/stale sweep ever fires.
assert coord._draft_pr_provider is None
def test_serve_binds_live_p3_wiring_when_env_set(
cli: ModuleType,
db_path: Path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""With the P3 env provisioned, ``serve`` binds the live wiring pair."""
monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "Sea-Haven-Industries")
monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "orchestrator")
monkeypatch.setenv("AGENT_TEAM_CI_READ_TOKEN", "ghp_test")
_FakeCoordinator.instances.clear()
monkeypatch.setattr(
"agent_team.coordinator.Coordinator", _FakeCoordinator, raising=True
)
coord = cli._build_coordinator(_serve_args(db_path, "serve"))
assert callable(coord.build_verify_wiring)
assert callable(coord.dispatch_node_wiring)
# Live box: the CI-watcher seams are wired so a task suspended at VERIFY
# awaiting CI gets resumed/parked rather than waiting forever. The poller is
# the read-only default; the provider is the coordinator's durable
# enumerator (bound post-construction); the timeout is the 30-min default.
assert callable(coord.ci_poller)
assert coord.ci_timeout == timedelta(minutes=30)
assert coord._ci_pending_provider == coord._enumerate_ci_pending
# A4 draft-PR monitor: the live serve path binds a read-only provider so the
# sweep has a real snapshot to ALARM / remind on (gated on the same live-pair
# signal as ci_poller). Without this the wired sweep would always see no PRs.
assert callable(coord._draft_pr_provider)
def test_serve_draft_pr_provider_reads_open_apply_draft_prs(
cli: ModuleType,
db_path: Path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""The bound draft-PR provider is READ-ONLY: it shells a scoped ``gh pr list``
(no write/close/dispatch) and maps the JSON into ``DraftPr`` snapshots."""
from agent_team.draft_pr_monitor import DraftPr
monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "Sea-Haven-Industries")
monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "orchestrator")
monkeypatch.setenv("AGENT_TEAM_CI_READ_TOKEN", "ghp_test")
_FakeCoordinator.instances.clear()
monkeypatch.setattr(
"agent_team.coordinator.Coordinator", _FakeCoordinator, raising=True
)
captured: dict[str, Any] = {}
class _FakeProc:
stdout = json.dumps(
[
{
"number": 7,
"createdAt": "2026-06-23T10:00:00Z",
"updatedAt": "2026-06-23T10:05:00Z",
},
{
"number": 9,
"createdAt": "2026-06-10T09:00:00Z",
"updatedAt": "2026-06-10T09:00:00Z",
},
]
)
def _fake_run(args: Any, **kwargs: Any) -> _FakeProc:
# The provider must shell a READ-ONLY, namespace-scoped enumeration —
# never a write/close/merge subcommand.
captured["args"] = args
captured["kwargs"] = kwargs
return _FakeProc()
monkeypatch.setattr("subprocess.run", _fake_run, raising=True)
coord = cli._build_coordinator(_serve_args(db_path, "serve"))
assert callable(coord._draft_pr_provider)
prs = coord._draft_pr_provider()
# READ-ONLY + scoped: it is a `gh pr list` over the apply/ head namespace, not
# a mutating subcommand, and never carries --shell.
assert captured["args"][:3] == ["gh", "pr", "list"]
assert "--draft" in captured["args"]
assert f"head:{cli._DRAFT_PR_HEAD_PREFIX}" in captured["args"]
assert "Sea-Haven-Industries/orchestrator" in captured["args"]
assert not any(
tok in captured["args"] for tok in ("close", "merge", "edit", "ready")
)
# The JSON rows map field-for-field into the snapshot the monitor expects.
assert prs == [
DraftPr(
number=7,
opened_at="2026-06-23T10:00:00Z",
updated_at="2026-06-23T10:05:00Z",
),
DraftPr(
number=9,
opened_at="2026-06-10T09:00:00Z",
updated_at="2026-06-10T09:00:00Z",
),
]
def test_serve_draft_pr_provider_fails_soft_on_gh_error(
cli: ModuleType,
db_path: Path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""A non-zero ``gh`` exit yields an empty snapshot (the sweep no-ops) rather
than raising and breaking the tick loop."""
import subprocess
monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "Sea-Haven-Industries")
monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "orchestrator")
monkeypatch.setenv("AGENT_TEAM_CI_READ_TOKEN", "ghp_test")
_FakeCoordinator.instances.clear()
monkeypatch.setattr(
"agent_team.coordinator.Coordinator", _FakeCoordinator, raising=True
)
def _boom(args: Any, **kwargs: Any) -> Any:
raise subprocess.CalledProcessError(returncode=1, cmd=args)
monkeypatch.setattr("subprocess.run", _boom, raising=True)
coord = cli._build_coordinator(_serve_args(db_path, "serve"))
assert coord._draft_pr_provider() == []
def test_start_does_not_bind_p3_wiring_even_when_env_set(
cli: ModuleType,
db_path: Path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""Only ``serve`` (the daemon) auto-binds P3; the one-shot ``start`` path runs
to the first human gate and never wires build/verify/dispatch."""
monkeypatch.setenv("AGENT_TEAM_REPO_OWNER", "Sea-Haven-Industries")
monkeypatch.setenv("AGENT_TEAM_REPO_NAME", "orchestrator")
monkeypatch.setenv("AGENT_TEAM_CI_READ_TOKEN", "ghp_test")
_FakeCoordinator.instances.clear()
monkeypatch.setattr(
"agent_team.coordinator.Coordinator", _FakeCoordinator, raising=True
)
coord = cli._build_coordinator(_serve_args(db_path, "start"))
assert coord.build_verify_wiring is None
assert coord.dispatch_node_wiring is None
def test_build_transport_dry_run_returns_dry_run_transport(cli: ModuleType) -> None:
"""``dry_run=True`` yields a _DryRunTransport whose post returns a synthetic ref."""
args = argparse.Namespace(dry_run=True, transport="slack")
@ -1077,3 +1297,65 @@ def test_notify_sink_forwards_thread_ts(
# No thread_ts when none is given (top-level post, not a broken key).
assert "thread_ts" not in captured[1]
assert captured[1] == {"channel": "C123", "text": "top-level milestone"}
def test_dispatch_from_files_invokes_dispatcher(
cli: ModuleType, db_path: Path, audit_log: Path, tmp_path: Path, monkeypatch
) -> None:
"""`dispatch` reads a diff/scope file and fires dispatch_apply_verify (P3 op-b)."""
from agent_team import dispatcher as d
diff_f = tmp_path / "d.diff"
diff_f.write_text("diff --git a/x b/x\n@@ -1 +1 @@\n-a\n+b\n", encoding="utf-8")
scope_f = tmp_path / "s.txt"
scope_f.write_text("agent_team/\n", encoding="utf-8")
calls: dict = {}
def _fake_dispatch(*, owner, repo, task_id, diff_text, declared_scope, base):
calls.update(
owner=owner, repo=repo, task_id=task_id, scope=declared_scope, base=base
)
return d.DispatchResult(
inputs=d.build_dispatch_inputs(
task_id=task_id, diff_text=diff_text, declared_scope=declared_scope
),
run_id="27990718108",
dispatched_at="2026-06-24T00:00:00Z",
correlation_tag=task_id,
)
monkeypatch.setattr(d, "dispatch_apply_verify", _fake_dispatch)
code, out = _run(
cli,
db_path,
audit_log,
"dispatch",
"task-xyz",
"--owner",
"Sea-Haven-Industries",
"--repo",
"orchestrator",
"--diff",
str(diff_f),
"--scope",
str(scope_f),
)
assert code == 0, out
assert calls["owner"] == "Sea-Haven-Industries"
assert calls["repo"] == "orchestrator"
assert calls["task_id"] == "task-xyz"
assert "agent_team/" in calls["scope"]
assert "27990718108" in out # the located run_id is reported
def test_dispatch_requires_owner_repo(
cli: ModuleType, db_path: Path, audit_log: Path, tmp_path: Path, monkeypatch
) -> None:
"""Without owner/repo (args or env) dispatch refuses with exit 2, no dispatch."""
monkeypatch.delenv("AGENT_TEAM_REPO_OWNER", raising=False)
monkeypatch.delenv("AGENT_TEAM_REPO_NAME", raising=False)
diff_f = tmp_path / "d.diff"
diff_f.write_text("diff --git a/x b/x\n", encoding="utf-8")
code, _ = _run(cli, db_path, audit_log, "dispatch", "t1", "--diff", str(diff_f))
assert code == 2

View file

@ -0,0 +1,400 @@
"""Unit tests for agent_team.draft_pr_monitor (P3 box-integration, A4).
Covers the draft-PR runaway/stale sweep: it ALARMs when > 3 draft PRs open
within 15 minutes, reminds on a draft PR idle > 7 days, applies flapping backoff
(one ALARM per cooldown, one reminder per PR per cooldown), and never auto-closes
or self-stops. No network: every side effect (alarm, reminder) is an injected
callable, exactly as the module's contract promises.
"""
from __future__ import annotations
from datetime import datetime, timedelta, timezone
from agent_team.draft_pr_monitor import (
DEFAULT_ALARM_COOLDOWN,
DEFAULT_STALE_COOLDOWN,
DraftPr,
MonitorAction,
MonitorMemory,
run_draft_pr_monitor,
)
# A fixed "now"; ISO strings mirror the GitHub API / ledger (UTC).
_NOW = datetime(2026, 6, 23, 12, 0, 0, tzinfo=timezone.utc)
def _ago(**kwargs: float) -> str:
return (_NOW - timedelta(**kwargs)).isoformat()
class _Recorder:
"""Records the runaway counts / stale PRs handed to the side-effect hooks."""
def __init__(self) -> None:
self.alarms: list[int] = []
self.reminders: list[int] = []
def on_alarm(self, opened_in_window: int) -> None:
self.alarms.append(opened_in_window)
def on_stale_reminder(self, pr: DraftPr) -> None:
self.reminders.append(pr.number)
def _open_burst(count: int, *, within_minutes: int = 5) -> list[DraftPr]:
"""``count`` draft PRs all opened ``within_minutes`` ago (inside the window)."""
return [
DraftPr(
number=i,
opened_at=_ago(minutes=within_minutes),
updated_at=_NOW.isoformat(),
)
for i in range(count)
]
# --------------------------------------------------------------------------- #
# Runaway ALARM threshold (> 3 within 15 min)
# --------------------------------------------------------------------------- #
def test_runaway_alarm_fires_above_threshold() -> None:
rec = _Recorder()
report = run_draft_pr_monitor(
_open_burst(4),
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
now=_NOW,
)
assert report.alarmed == 1
assert rec.alarms == [4]
assert report.outcomes[0].action is MonitorAction.ALARMED
def test_exactly_threshold_does_not_alarm() -> None:
"""The rule is STRICTLY greater than 3 — exactly 3 must NOT alarm."""
rec = _Recorder()
report = run_draft_pr_monitor(
_open_burst(3),
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
now=_NOW,
)
assert report.alarmed == 0
assert rec.alarms == []
def test_prs_opened_outside_window_are_not_counted() -> None:
"""5 PRs but only 2 inside the 15-min window -> no alarm."""
rec = _Recorder()
prs = [
DraftPr(number=1, opened_at=_ago(minutes=2)),
DraftPr(number=2, opened_at=_ago(minutes=10)),
DraftPr(number=3, opened_at=_ago(minutes=20)),
DraftPr(number=4, opened_at=_ago(minutes=40)),
DraftPr(number=5, opened_at=_ago(hours=3)),
]
report = run_draft_pr_monitor(
prs,
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
now=_NOW,
)
assert report.opened_in_window == 2
assert report.alarmed == 0
assert rec.alarms == []
def test_unparseable_opened_at_is_skipped_for_runaway() -> None:
"""A PR with a junk/missing opened_at is not counted and does not crash."""
rec = _Recorder()
prs = _open_burst(4) + [DraftPr(number=99, opened_at="not-a-date")]
report = run_draft_pr_monitor(
prs,
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
now=_NOW,
)
# The 4 valid PRs still breach; the junk one is just ignored for the count.
assert report.opened_in_window == 4
assert report.alarmed == 1
# --------------------------------------------------------------------------- #
# Flapping backoff — ALARM at most once per cooldown
# --------------------------------------------------------------------------- #
def test_runaway_alarm_suppressed_within_cooldown() -> None:
rec = _Recorder()
mem = MonitorMemory()
prs = _open_burst(5)
first = run_draft_pr_monitor(
prs,
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
memory=mem,
now=_NOW,
)
assert first.alarmed == 1
# A second tick a minute later, condition still breaching -> suppressed.
second = run_draft_pr_monitor(
prs,
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
memory=mem,
now=_NOW + timedelta(minutes=1),
)
assert second.alarmed == 0
assert second.alarm_suppressed == 1
assert rec.alarms == [5] # only the first post
def test_runaway_alarm_refires_after_cooldown() -> None:
rec = _Recorder()
mem = MonitorMemory()
prs = _open_burst(5, within_minutes=1)
run_draft_pr_monitor(
prs,
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
memory=mem,
now=_NOW,
)
# Past the cooldown, a still-breaching condition ALARMs again. Re-anchor the
# PRs so they remain inside the window relative to the later "now".
later = _NOW + DEFAULT_ALARM_COOLDOWN + timedelta(seconds=1)
prs_later = [
DraftPr(number=p.number, opened_at=(later - timedelta(minutes=1)).isoformat())
for p in prs
]
report = run_draft_pr_monitor(
prs_later,
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
memory=mem,
now=later,
)
assert report.alarmed == 1
assert len(rec.alarms) == 2
# --------------------------------------------------------------------------- #
# Stale reminder (> 7 days idle), never auto-closes
# --------------------------------------------------------------------------- #
def test_stale_pr_reminds_once() -> None:
rec = _Recorder()
prs = [DraftPr(number=7, opened_at=_ago(days=30), updated_at=_ago(days=8))]
report = run_draft_pr_monitor(
prs,
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
now=_NOW,
)
assert report.stale_reminded == 1
assert rec.reminders == [7]
def test_fresh_pr_does_not_remind() -> None:
rec = _Recorder()
prs = [DraftPr(number=7, opened_at=_ago(days=30), updated_at=_ago(days=6))]
report = run_draft_pr_monitor(
prs,
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
now=_NOW,
)
assert report.stale_reminded == 0
assert rec.reminders == []
def test_exactly_7_days_does_not_remind() -> None:
"""The rule is strictly MORE than 7 days idle."""
rec = _Recorder()
prs = [DraftPr(number=7, updated_at=_ago(days=7))]
report = run_draft_pr_monitor(
prs,
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
now=_NOW,
)
assert report.stale_reminded == 0
def test_stale_reminder_suppressed_within_cooldown() -> None:
rec = _Recorder()
mem = MonitorMemory()
prs = [DraftPr(number=7, updated_at=_ago(days=10))]
run_draft_pr_monitor(
prs,
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
memory=mem,
now=_NOW,
)
second = run_draft_pr_monitor(
prs,
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
memory=mem,
now=_NOW + timedelta(hours=1),
)
assert second.stale_reminded == 0
assert second.stale_suppressed == 1
assert rec.reminders == [7] # only the first
def test_stale_reminder_refires_after_cooldown() -> None:
rec = _Recorder()
mem = MonitorMemory()
run_draft_pr_monitor(
[DraftPr(number=7, updated_at=_ago(days=10))],
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
memory=mem,
now=_NOW,
)
later = _NOW + DEFAULT_STALE_COOLDOWN + timedelta(seconds=1)
report = run_draft_pr_monitor(
[DraftPr(number=7, updated_at=(later - timedelta(days=10)).isoformat())],
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
memory=mem,
now=later,
)
assert report.stale_reminded == 1
assert len(rec.reminders) == 2
def test_unparseable_updated_at_is_skipped_for_stale() -> None:
rec = _Recorder()
prs = [DraftPr(number=7, updated_at="garbage")]
report = run_draft_pr_monitor(
prs,
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
now=_NOW,
)
assert report.stale_reminded == 0
assert rec.reminders == []
# --------------------------------------------------------------------------- #
# Side-effect isolation + watermark-not-advanced-on-error
# --------------------------------------------------------------------------- #
def test_alarm_side_effect_error_is_isolated_and_retried() -> None:
"""A raising on_alarm yields ERRORED and does NOT advance the watermark."""
calls: list[int] = []
def boom(opened_in_window: int) -> None:
calls.append(opened_in_window)
raise RuntimeError("slack down")
mem = MonitorMemory()
prs = _open_burst(5)
report = run_draft_pr_monitor(
prs,
on_alarm=boom,
on_stale_reminder=lambda pr: None,
memory=mem,
now=_NOW,
)
assert report.errored == 1
assert report.alarmed == 0
# Watermark not advanced -> a retry the very next tick (no cooldown lock-in).
assert mem.last_alarm_at is None
report2 = run_draft_pr_monitor(
prs,
on_alarm=boom,
on_stale_reminder=lambda pr: None,
memory=mem,
now=_NOW + timedelta(seconds=30),
)
assert report2.errored == 1
assert len(calls) == 2
def test_stale_side_effect_error_does_not_abort_sweep() -> None:
"""A reminder that raises for one PR does not stop the runaway ALARM."""
rec = _Recorder()
def boom(pr: DraftPr) -> None:
raise RuntimeError("slack down")
prs = _open_burst(5) + [DraftPr(number=50, updated_at=_ago(days=10))]
report = run_draft_pr_monitor(
prs,
on_alarm=rec.on_alarm,
on_stale_reminder=boom,
now=_NOW,
)
# Runaway still alarms; the stale reminder errored but was isolated.
assert report.alarmed == 1
assert report.errored == 1
# --------------------------------------------------------------------------- #
# Never auto-closes / self-stops; memory hygiene
# --------------------------------------------------------------------------- #
def test_monitor_takes_no_infrastructure_action() -> None:
"""The only effects are the two injected notify callables — nothing else.
There is no close/stop/dispatch seam on the monitor; this asserts the sweep
exposes ONLY the alarm + reminder hooks (a regression guard against adding a
self-acting side effect).
"""
import inspect
sig = inspect.signature(run_draft_pr_monitor)
effect_params = {p for p in sig.parameters if p.startswith("on_")}
assert effect_params == {"on_alarm", "on_stale_reminder"}
def test_memory_prunes_gone_prs() -> None:
"""A PR no longer present has its backoff watermark pruned (bounded growth)."""
rec = _Recorder()
mem = MonitorMemory()
run_draft_pr_monitor(
[DraftPr(number=7, updated_at=_ago(days=10))],
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
memory=mem,
now=_NOW,
)
assert 7 in mem.last_reminded_at
# Next pass: PR #7 is gone (merged/closed by a human). Its watermark prunes.
run_draft_pr_monitor(
[],
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
memory=mem,
now=_NOW + timedelta(minutes=1),
)
assert 7 not in mem.last_reminded_at
def test_no_draft_prs_is_quiet() -> None:
rec = _Recorder()
report = run_draft_pr_monitor(
[],
on_alarm=rec.on_alarm,
on_stale_reminder=rec.on_stale_reminder,
now=_NOW,
)
assert report.outcomes == []
assert rec.alarms == []
assert rec.reminders == []

View file

@ -73,12 +73,35 @@ def test_roundtrip_dict() -> None:
candidate_diff="diff --git a b",
diff_hash="deadbeef",
ci_results={"conclusion": "success"},
run_id="27990718108",
ci_correlation_tag="t1-1a2b3c",
dispatched_at="2026-06-17T00:30:00Z",
build_loops=2,
transport="slack",
created_at="2026-06-17T00:00:00Z",
updated_at="2026-06-17T01:00:00Z",
)
restored = task_from_dict(task_to_dict(rec))
assert restored == rec
# P3 dispatch->verify plumbing fields survive the dict round-trip.
assert restored.run_id == "27990718108"
assert restored.ci_correlation_tag == "t1-1a2b3c"
assert restored.dispatched_at == "2026-06-17T00:30:00Z"
# The durable build<->verify loop count round-trips (LOGIC-RACE-01).
assert restored.build_loops == 2
def test_build_loops_defaults_zero_and_roundtrips() -> None:
rec = TaskRecord(
thread_id="t1",
status=TaskStatus.ACTIVE,
current_phase=Phase.BUILD,
)
assert rec.build_loops == 0
# A dict missing build_loops (older record) defaults to 0, not a crash.
legacy = task_to_dict(rec)
del legacy["build_loops"]
assert task_from_dict(legacy).build_loops == 0
def test_roundtrip_json() -> None:

View file

@ -123,6 +123,69 @@ def test_fail_increments_build_loops_in_verdict() -> None:
assert out["review_verdicts"][0]["build_loops"] == 2
# --------------------------------------------------------------------------- #
# Durable per-task build-loop budget (LOGIC-RACE-01)
# --------------------------------------------------------------------------- #
def test_fail_reads_loop_count_from_state_not_config() -> None:
# The count that bounds the loop is DURABLE per-task state. A shared config
# with build_loops=0 must NOT mask a per-task state["build_loops"] of 2.
diff = _diff_for("src/foo.py")
ci = {"run_id": "r1", "conclusion": "failure", "diff_hash": _hash(diff)}
state = _state(diff, ci)
state["build_loops"] = 2
out = verifier_node(state, VerifierConfig(expected_run_id="r1", build_loops=0))
# next = state(2) + 1 = 3 == DEFAULT_MAX_BUILD_LOOPS -> park (not a vacuous
# loop driven by the always-zero shared config).
assert out["status"] == TaskStatus.PARKED.value
assert out["review_verdicts"][0]["build_loops"] == 3
def test_fail_writes_incremented_loop_count_into_state() -> None:
# On a recoverable FAIL the node persists the incremented count into the
# returned partial state so the NEXT VERIFY (after BUILD->DISPATCH) sees it.
diff = _diff_for("src/foo.py")
ci = {"run_id": "r1", "conclusion": "failure", "diff_hash": _hash(diff)}
state = _state(diff, ci)
state["build_loops"] = 0
out = verifier_node(state, VerifierConfig(expected_run_id="r1"))
assert out["current_phase"] == Phase.BUILD.value
assert out["build_loops"] == 1
def test_repeated_fail_parks_after_exactly_max_build_loops_via_state() -> None:
# Simulate the durable BUILD->DISPATCH->VERIFY loop: the incremented
# build_loops the node returns is fed back into the next call's state (the
# checkpointer carries it). A perpetually-FAILing task must reach PARKED
# after EXACTLY max_build_loops iterations, never loop unbounded.
diff = _diff_for("src/foo.py")
ci = {"run_id": "r1", "conclusion": "failure", "diff_hash": _hash(diff)}
# A single SHARED config (build_loops stays 0) — exactly the wiring-time
# reality that made the budget dead before the fix.
shared_cfg = VerifierConfig(expected_run_id="r1", max_build_loops=3, build_loops=0)
state = _state(diff, ci)
state["build_loops"] = 0
statuses: list[str] = []
for _ in range(10): # bound the harness so a regression can't hang the test
out = verifier_node(state, shared_cfg)
statuses.append(out["status"])
if out["status"] == TaskStatus.PARKED.value:
break
# Carry the durable count forward, as the checkpointer would.
state["build_loops"] = out["build_loops"]
# FAILs at counts 1, 2 loop back to BUILD; the 3rd (next==3==max) parks.
assert statuses == [
TaskStatus.ACTIVE.value,
TaskStatus.ACTIVE.value,
TaskStatus.PARKED.value,
]
assert out["current_phase"] == Phase.PARKED.value
# --------------------------------------------------------------------------- #
# BLOCK path — park for human + GPT cross-review
# --------------------------------------------------------------------------- #
@ -202,3 +265,80 @@ def test_allowed_scope_threaded_to_gate() -> None:
cfg = VerifierConfig(expected_run_id="r1", allowed_scope=["src/"])
out = verifier_node(_state(diff, ci), cfg)
assert out["status"] == TaskStatus.PARKED.value # out-of-scope -> BLOCK -> park
# --------------------------------------------------------------------------- #
# Per-task run-id binding (design §4 Decision 4)
# --------------------------------------------------------------------------- #
def test_state_run_id_overrides_static_config_run_id() -> None:
# The gate must bind to the run id THIS task dispatched (state["run_id"]),
# not the static config constant. A CI conclusion keyed to the per-task
# run id passes even though config carries a different (stale) run id.
diff = _diff_for("src/foo.py")
state = _state(
diff, {"run_id": "task-run", "conclusion": "success", "diff_hash": _hash(diff)}
)
state["run_id"] = "task-run"
# config.expected_run_id is a DIFFERENT, stale value — state must win.
out = verifier_node(state, VerifierConfig(expected_run_id="stale-wiring-run"))
assert out["status"] == TaskStatus.DONE.value
assert out["review_verdicts"][0]["run_id"] == "task-run"
def test_substituted_run_id_is_rejected() -> None:
# Anti-substitution: a CI result whose run_id != state["run_id"] is a BLOCK,
# even with a success conclusion (someone tried to graft a passing run from
# another task onto this one).
diff = _diff_for("src/foo.py")
state = _state(
diff,
{"run_id": "other-task-run", "conclusion": "success", "diff_hash": _hash(diff)},
)
state["run_id"] = "my-task-run"
out = verifier_node(state, VerifierConfig(expected_run_id=None))
assert out["status"] == TaskStatus.PARKED.value
assert out["ci_results"]["gate_decision"] == "block"
assert any("run-id mismatch" in r for r in out["review_verdicts"][0]["reasons"])
def test_two_concurrent_tasks_each_gate_against_own_run_id() -> None:
# Two tasks share ONE VerifierConfig but each gates against its OWN
# state["run_id"]. Task A's CI matches A's run id (PASS); task B's CI is
# keyed to A's run id (substitution) so B BLOCKs.
shared_cfg = VerifierConfig(expected_run_id=None)
diff_a = _diff_for("src/a.py")
state_a = _state(
diff_a, {"run_id": "run-A", "conclusion": "success", "diff_hash": _hash(diff_a)}
)
state_a["run_id"] = "run-A"
diff_b = _diff_for("src/b.py")
# B's fetched CI is wrongly keyed to run-A (a leaked/substituted run).
state_b = _state(
diff_b, {"run_id": "run-A", "conclusion": "success", "diff_hash": _hash(diff_b)}
)
state_b["run_id"] = "run-B"
out_a = verifier_node(state_a, shared_cfg)
out_b = verifier_node(state_b, shared_cfg)
assert out_a["status"] == TaskStatus.DONE.value
assert out_a["review_verdicts"][0]["run_id"] == "run-A"
assert out_b["status"] == TaskStatus.PARKED.value
assert out_b["review_verdicts"][0]["run_id"] == "run-B"
def test_none_run_id_at_gate_time_blocks_never_passes() -> None:
# No per-task run_id in state AND no config fallback -> BLOCK (park), never a
# vacuous pass, even when CI reports success.
diff = _diff_for("src/foo.py")
state = _state(
diff, {"run_id": "", "conclusion": "success", "diff_hash": _hash(diff)}
)
# state has no "run_id" key; config fallback is None.
out = verifier_node(state, VerifierConfig(expected_run_id=None))
assert out["status"] == TaskStatus.PARKED.value
assert out["ci_results"]["gate_decision"] == "block"

View file

@ -27,6 +27,26 @@ import pytest
from agent_team.nodes.dispatch_invoker import DispatchNodeFactory, make_dispatch_node
from agent_team.task_model import Phase, TaskStatus
_FAKE_RUN_ID = "27990718108"
@pytest.fixture(autouse=True)
def _stub_run_locator(monkeypatch: pytest.MonkeyPatch) -> None:
"""Stub the real ``gh run list`` locator so no test shells out / sleeps.
``dispatch_apply_verify`` resolves the dispatched run via ``_default_run_locator``
when no ``locator`` is injected; the real one polls ``gh`` with retry sleeps.
Replace it module-wide with a fast fake returning a fixed run id so every
``make_dispatch_node`` call (incl. factory paths) stays hermetic.
"""
from agent_team import dispatcher as _dispatcher
monkeypatch.setattr(
_dispatcher,
"_default_run_locator",
lambda: lambda **_kw: _FAKE_RUN_ID,
)
# --------------------------------------------------------------------------- #
# Fake BranchPusher / WorkflowDispatcher (the two injectable seams in dispatcher)
@ -76,16 +96,23 @@ _VALID_STATE: dict[str, Any] = {
# --------------------------------------------------------------------------- #
def test_dispatch_node_happy_path_returns_partial_state() -> None:
"""On success the node returns {} (partial state update — DONE comes from graph)."""
def test_dispatch_node_happy_path_persists_run_id() -> None:
"""On success the node persists the located run identity into state.
The dispatched run's id (plus the dispatched-at watermark and correlation
tag) flows into PipelineState so the verifier's read-only fetcher polls THIS
task's run and the pure-code gate binds its verdict to it.
"""
pusher, dispatcher = _fake_seams()
node = make_dispatch_node(
owner="org", repo="repo", pusher=pusher, dispatcher=dispatcher
)
result = node(_VALID_STATE)
# Remote implementation returns {} on success (graph topology marks DONE).
assert isinstance(result, dict)
assert result["run_id"] == _FAKE_RUN_ID
assert result["dispatched_at"] # stamped
assert result["ci_correlation_tag"] == "task-abc"
def test_dispatch_node_injects_owner_repo_at_factory_time() -> None:
@ -145,6 +172,60 @@ def test_dispatch_node_flattens_scope_list_to_string() -> None:
assert "agent_team/" in scope_str
def test_dispatch_node_unresolved_run_id_persists_watermark_and_verify_fails_closed(
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""The ``if not result.run_id`` branch: dispatch fired but the run could not
be correlated. The node must NOT park (the dispatch succeeded) — it persists
the dispatched-at watermark + correlation tag with ``run_id=None`` so the
downstream verifier gate fails closed (BLOCK/park), never fabricating a pass.
"""
from agent_team import dispatcher as _dispatcher
# Override the autouse locator stub: this run cannot be located -> None.
monkeypatch.setattr(_dispatcher, "_default_run_locator", lambda: lambda **_kw: None)
pusher, dispatcher = _fake_seams()
node = make_dispatch_node(
owner="org", repo="repo", pusher=pusher, dispatcher=dispatcher
)
result = node(_VALID_STATE)
# Not parked: the dispatch itself succeeded; only correlation failed.
assert result.get("status") != TaskStatus.PARKED.value
assert result["run_id"] is None
# The watermark + correlation tag are still persisted for the CI-watch timeout
# and for audit, even though the run id is unresolved.
assert result["dispatched_at"]
assert result["ci_correlation_tag"] == "task-abc"
# Feed the dispatch output into the verifier: a None run_id has nothing for
# the gate to bind to, so it BLOCKs and parks — fail closed, never a pass.
from agent_team.nodes.build_verify_subgraph import make_verify_node
from agent_team.nodes.verifier import VerifierConfig
verify_state: dict[str, Any] = {
"thread_id": "task-abc",
"candidate_diff": _VALID_STATE["candidate_diff"],
"diff_hash": "deadbeef",
"ci_results": None,
"run_id": result["run_id"], # None, carried from dispatch
}
def _pass_fetcher(_s: Any) -> Any:
# Even a 'success' CI result cannot rescue a missing run_id: the gate has
# no run identity to bind the verdict to.
return {"run_id": "whatever", "conclusion": "success", "diff_hash": "deadbeef"}
verify_node = make_verify_node(
VerifierConfig(expected_run_id=None, allowed_scope=["agent_team/"]),
ci_result_fetcher=_pass_fetcher,
)
out = verify_node(verify_state)
assert out["status"] == TaskStatus.PARKED.value
assert out["ci_results"]["gate_decision"] == "block"
# --------------------------------------------------------------------------- #
# make_dispatch_node — fail closed (never dispatches incomplete input)
# --------------------------------------------------------------------------- #

View file

@ -308,6 +308,114 @@ follow the escalation ladder below.
---
## Incident 7 — P3 rollback (unwind the apply/verify flip)
The P3 apply/verify CI workflow is **already LIVE** (flipped + provisioned
2026-06-22): the `agent-apply` environment with its required reviewer, the GitHub
App installation (`pull-requests: write`), the run-name/permissions edits on
`.github/workflows/agent-team-apply-verify.yml`, and `enforce_admins` branch
protection on `main`. `agent-team/scripts/p3_rollback.sh` is the inverse — it
restores those four privileged surfaces to a recorded baseline and asserts
post-restore == baseline (fail-closed: a mismatched restore aborts non-zero).
> **Run from the operator host (the Mac), NOT the box.** No standing box token is
> used; the script relies on `gh` being authenticated on the operator host (plus
> `git` for the post-merge revert path). The R720 daemon holds no write token by
> design (P3-PHASE0-DESIGN.md, "Out of scope").
**When to trigger:**
- The flip must be unwound (P3 apply/verify is being backed out), OR
- A **premature flip that RAN** — an apply/verify run executed before it should
have and opened a draft PR / branch that must not exist (use the `incident`
surface, below).
**Prerequisites:**
- A baseline JSON recorded **BEFORE** the flip (the restore target). Default path
`.security-review/p3-baseline.json` or `$P3_ROLLBACK_BASELINE`, override with
`--baseline FILE`. The script refuses to run if the baseline file is missing
(`record it BEFORE the flip`) — there is no inferred baseline.
- Target repo from the baseline's `repo` key, `--repo OWNER/NAME`, or
`$AGENT_TEAM_REPO_OWNER/$AGENT_TEAM_REPO_NAME`. If both the baseline and `--repo`
are set they must agree (fail-closed on mismatch).
**The four privileged surfaces it restores:**
1. **`workflow`** — the apply/verify flip itself. Pre-merge: close the flip PR +
delete its branch (`--flip-pr N [--flip-branch REF]`). Post-merge: `git revert`
the flip commit + push + re-run CI (`--merged --flip-commit SHA`). LIVE: also
reverts the workflow YAML to the recorded `workflow_baseline_sha` so the live
file matches baseline.
2. **`environment`** — the `agent-apply` environment: re-PUTs the recorded required
reviewer(s) + deployment branch policy, then asserts the environment exists.
3. **`app`** — the GitHub App: **always rotates (revokes) the installation token
first** (any minted token is now suspect), then `reduce`s permissions to the
baseline or `uninstall`s the installation per `app.action`.
4. **`protection`** — branch protection on `main`: re-enables `enforce_admins`
(`include_administrators`) and asserts it is ON. **Refuses to run if the
baseline's `include_administrators` is not `true`** — a rollback must never
leave a weaker posture than baseline.
`all` runs surfaces 1→4 in order.
**How to run — always `--dry-run` first, then `--apply`:**
The script is **destructive-safe by default**: with no `--apply` it only PRINTS
the plan (`[PLAN] ...` lines). `--apply` is the only thing that performs
mutations. Inspect the dry-run plan, confirm it targets the right repo + baseline,
then re-run with `--apply`.
```bash
cd ~/Documents/repositories/orchestrator/agent-team # operator host, gh authed
# 1. Dry-run: print the plan only, no mutations.
scripts/p3_rollback.sh all --baseline .security-review/p3-baseline.json
# 2. Apply, once the plan looks right (pre-merge flip example):
scripts/p3_rollback.sh all --apply \
--baseline .security-review/p3-baseline.json \
--flip-pr 123 --flip-branch agent-team/apply/flip
# Post-merge workflow revert instead of pre-merge close:
scripts/p3_rollback.sh workflow --apply --merged --flip-commit <SHA>
# A single surface at a time is fine too:
scripts/p3_rollback.sh protection --apply
```
**Premature-flip-that-RAN incident (the `incident` surface):**
Use this when a flip executed prematurely and opened a draft PR/branch. It runs
the full incident sequence — rotate, revert, audit (read-only), restore, note:
```bash
# Dry-run first (the audit/list steps run read-only either way):
scripts/p3_rollback.sh incident --baseline .security-review/p3-baseline.json
# Apply, with the premature draft PR if known:
scripts/p3_rollback.sh incident --apply \
--flip-pr <N> --flip-branch agent-team/apply/<task_id>
```
It performs, in order:
1. **Rotate the App installation token first** — anything the premature run minted
is suspect.
2. **Revert the draft PR / branch the App opened** — close `--flip-pr` + delete its
branch; if no `--flip-pr` is given it lists open `agent-team/apply/*` draft PRs
to triage.
3. **Audit the Checks trail** (read-only — always runs) — recent
`agent-team-apply-verify.yml` runs with conclusions.
4. **Restore the `agent-apply` environment + branch protection** to baseline
(surfaces 2 + 4).
5. **File an incident note** under
`.security-review/incidents/p3-premature-flip-<timestamp>.md` recording the
repo, baseline, flip PR/branch, and actions taken.
After any rollback, confirm the post-restore `[OK]` asserts printed (the script
exits non-zero if any assert failed), and record the action — for the `incident`
path the note is written automatically; otherwise note it on the Jira ticket per
the escalation ladder.
---
## Escalation ladder (design §5 / §6.6, resolves Q3)
Anything that does not clear on the first ALARM escalates — but **ALARM-only in
@ -337,3 +445,88 @@ non-CLI recoveries note it on the Jira ticket.
PROVISIONING-RUNBOOK steps.
- Update `project_r720_agent_team` memory if the incident revealed a durable
fact (a new failure mode, a config that must change).
## P3 box env wiring (build → dispatch → verify)
The coordinator's environment is loaded from `EnvironmentFile=-/home/adam/secrev.env`
(declared in the unit; the leading `-` makes it optional so a missing file does
not fail the unit). The **live** P3 path — gated build → dispatch (trigger CI,
capture `run_id`) → verify (read the CI conclusion) — reads four variables from
that file at graph-build / dispatch time:
- `AGENT_TEAM_REPO_OWNER` — dispatch/verify target owner (fixed at factory time,
never read from pipeline state, so model output cannot redirect the target).
- `AGENT_TEAM_REPO_NAME` — dispatch/verify target repo (same fail-closed binding).
- `AGENT_TEAM_BASE_BRANCH` — PR base branch; optional, defaults to `main`.
- `AGENT_TEAM_CI_READ_TOKEN` — the **read-only** CI-result token used for the
verifier's authenticated conclusion read (falls back to `GITHUB_TOKEN`).
If owner, repo, or the CI-read token is unset, `serve()` degrades to the INERT P3
path (one WARNING + a `#agent-team` notice) rather than crash-looping the daemon
(Phase-0 Decision 5). No `pull-requests:write` / `contents:write` token and no
`AGENT_APPLY_APP_ID` / `AGENT_APPLY_APP_PRIVATE_KEY` may live in `~/secrev.env`:
the apply path mints its write token **inside** the CI runner from Actions
secrets, so the box holds no standing write credential. That invariant is
enforced by `scripts/assert_no_write_token.py` (the A2 audit) at provisioning and
in CI — run it before any deploy.
### Verifying the vars load
After installing/editing `~/secrev.env` and `systemctl daemon-reload` +
`systemctl restart agent-team-coordinator.service`, confirm the P3 vars reached
the **running coordinator process**.
> ⚠️ Do NOT use `systemctl show -p Environment` — it lists only inline
> `Environment=` directives and does **NOT** show vars loaded from
> `EnvironmentFile=` (which is how `~/secrev.env` is loaded). It comes back empty
> even when the vars are correctly loaded, so it is misleading here.
Read the actual process environment instead (requires sudo to read another
process's `environ`):
```
MP=$(systemctl show agent-team-coordinator.service -p MainPID --value)
sudo tr '\0' '\n' < /proc/$MP/environ | grep -E '^AGENT_TEAM_REPO|^AGENT_TEAM_BASE'
```
The output should list `AGENT_TEAM_REPO_OWNER`, `AGENT_TEAM_REPO_NAME`, and
`AGENT_TEAM_BASE_BRANCH` (if set). To confirm the live P3 wiring actually bound
(not the INERT path), check the code resolver directly:
```
cd ~/orchestrator/agent-team && set -a && source ~/secrev.env && set +a \
&& .venv/bin/python -c "from agent_team.coordinator import _p3_env_is_configured; print(_p3_env_is_configured())"
```
`True` means the live build+verify path is bound; `False` means the daemon is on
the INERT P3 path (fix `~/secrev.env`, reload, restart). The CI-read token is
satisfied by `AGENT_TEAM_CI_READ_TOKEN` or the read-only `GITHUB_TOKEN` fallback;
treat any token value in process output as sensitive. The box must hold NO
`AGENT_APPLY_APP_ID` / `AGENT_APPLY_APP_PRIVATE_KEY` / write token — verify with
`python scripts/assert_no_write_token.py` (and note its scope-detection caveat in
the script header: a fine-grained token's write capability is only definitively
confirmed by a live `POST /git/refs` probe returning `403`).
### Operator-initiated dispatch (P3 option-b)
The box is read-only, so its in-graph DISPATCH node fail-closes/parks — it never
pushes or triggers CI. Completing a dispatch is an explicit operator step with a
**just-in-time** write token (never stored in `~/secrev.env`):
```
# On the box (where the task's candidate_diff lives in the ledger), with a
# WRITE-capable token provided for THIS invocation only:
cd ~/orchestrator/agent-team
GH_TOKEN=<operator pull-requests+contents:write token> \
.venv/bin/python run-team.py dispatch <thread_id> --write-back
```
This reads the task's `candidate_diff` + declared scope from the checkpoint,
pushes the head branch, fires the `agent-team-apply-verify` `workflow_dispatch`,
prints the located `run_id`, and (`--write-back`) writes it into the task
checkpoint so the box's VERIFY binds to that run. CI then runs guard → build-test
→ pure-code gate; the privileged `gate-and-pr` job pauses at the `agent-apply`
environment for your **required-reviewer approval** before the draft PR opens.
Alternatively pass `--diff FILE --scope FILE` to dispatch a diff without reading
the ledger. The token is consumed by `gh`/`git` for the one command and never
persisted; the box returns to read-only at rest.

View file

@ -2,8 +2,21 @@
Formal phased plan to take the agent-team Plane-2 pipeline from **clarify+plan only**
to **producing reviewable draft PRs**, while keeping the always-on R720 box
read-only and the apply path zero-AWS. Status as of 2026-06-22: **NOT STARTED**
(P3 is built but inert). This plan is the input to `/sh-plan-review` before any build.
read-only and the apply path zero-AWS.
**Status as of 2026-06-23: PARTIALLY LIVE.** The split-CI apply/verify workflow
and its provisioning (the scoped GitHub App, the `agent-apply` environment + Actions
secrets, and the dispatched runs) are **LIVE since 2026-06-22** — i.e. the
workflow-authoring + provisioning phases below (Phases 1/1b/2) are **DONE**. The
remaining work is the **box-side build → dispatch → verify integration** on
`feat/agent-team-p3-box-integration` (BUILD → DISPATCH → VERIFY wiring;
operator-initiated dispatch; run-name correlation for run_id capture; async
CI-watch resume-on-complete; fail-safe serve default) — see `docs/P3-PHASE0-DESIGN.md`
for the recorded design decisions. The two remaining **human gates** are: (C1) re-run
`/sh-security-review` + the GPT-4.1 cross-family review against the *enabled* workflow
+ the bound box-side wiring (Phase 3 / B5), and (D) deploy → smoke-test → merge
(Phase 4). This plan was the input to `/sh-plan-review` before the build began; that
loop is complete.
> **Prerequisite reading:** `docs/r720-agent-team-design.md` §3.3.2 (CI-as-verifier
> trust boundary, "B4"), `PROVISIONING-RUNBOOK.md` (the P3-live-flip section),
@ -13,16 +26,29 @@ read-only and the apply path zero-AWS. Status as of 2026-06-22: **NOT STARTED**
## 1. Objective & current state
**Today (inert):** the pipeline runs `INTAKE → CLARIFIER → PLANNER → REVIEW`, but
`serve` passes `build_verify_wiring=None`, the Tier-3 fixer is `--dry-run` only, and
`agent-team/ci/agent-team-apply-verify.yml` has its privileged steps disabled with
`if: ${{ false }}` and `pull-requests:write` / `environment:` commented out. So it
clarifies + plans but writes no code and opens no PR.
**Live infrastructure (since 2026-06-22):** the split-CI apply/verify workflow
(`agent-team/ci/agent-team-apply-verify.yml`) is authored, enabled, and **provisioned**
— the scoped **GitHub App** (`pull-requests:write` + minimal contents), the
`agent-apply` **GitHub Actions Environment** (Adam as required reviewer + branch
protection), and the App credentials as **Actions secrets** all exist, and dispatched
runs have executed against it. Org CI can build/test/security-review an untrusted diff
in a sandbox, a pure-code gate confirms green from authenticated Checks-API results,
and the privileged job opens a **draft PR** — all without the box ever holding a write
token.
**After P3:** the pipeline can emit a diff, have **org CI** build/test/security-review
it in an untrusted sandbox, a pure-code gate confirm green from authenticated
Checks-API results, and a scoped **GitHub App** open a **draft PR** for human review.
The box never gains a write token.
**Not yet wired (the remaining build, on `feat/agent-team-p3-box-integration`):** the
**box-side integration** that makes a task flow BUILD → DISPATCH (trigger that live CI,
capture the run_id) → VERIFY (read its conclusion) automatically. As-built, `serve`
binds the fail-safe gated P3 wiring (degrading to the INERT path if the dispatch target
/ CI-read token is unset — Phase-0 Decision 5), dispatch is **operator-initiated**
(branch push + `gh workflow run` on operator-host credentials; the box holds no standing
write token), run_id capture is **run-name / correlation-tag** based, and the CI wait is
**async resume-on-complete** via a `tick()`-driven CI-watcher (not a blocking poll). The
Tier-3 fixer remains `--dry-run` only until the Phase-4 smoke test passes.
**After the box-side integration + the remaining human gates:** the pipeline emits a
diff end-to-end into a reviewable **draft PR** for human review. The box never gains a
write token.
## 2. Locked decisions (carried in — do not relitigate here)
@ -35,16 +61,19 @@ The box never gains a write token.
## 3. Hard gates (must clear before the flip — these block everything)
| Gate | Why | Owner |
|---|---|---|
| `/sh-plan-review` on THIS plan | adversarial plan audit before build | me → GPT-4.1 |
| `/sh-security-review` on the apply/verify CI surface | auth + untrusted-input + CI trust boundary | me |
| GPT-4.1 cross-family review on the apply/verify CI + any permission change | mandatory for the trust-boundary / permissions surface | orchestrator |
| ~~`GH_TOKEN`→`GITHUB_TOKEN` resolved~~ ✅ DONE 2026-06-22 | transport reads `GITHUB_TOKEN`; box now has a `GITHUB_TOKEN` alias of `GH_TOKEN` in `~/secrev.env` | me |
| Gate | Why | Owner | Status |
|---|---|---|---|
| `/sh-plan-review` on THIS plan | adversarial plan audit before build | me → GPT-4.1 | ✅ DONE (Round 1 + 2, 2026-06-22) |
| `/sh-security-review` on the apply/verify CI surface | auth + untrusted-input + CI trust boundary | me | ⏳ RE-RUN on the *enabled* workflow + bound box-side wiring (C1 / Phase 3 / B5) |
| GPT-4.1 cross-family review on the apply/verify CI + any permission change | mandatory for the trust-boundary / permissions surface | orchestrator | ⏳ RE-RUN on the *enabled* workflow + bound box-side wiring (C1 / Phase 3 / B5) |
| ~~`GH_TOKEN`→`GITHUB_TOKEN` resolved~~ ✅ DONE 2026-06-22 | transport reads `GITHUB_TOKEN`; box now has a `GITHUB_TOKEN` alias of `GH_TOKEN` in `~/secrev.env` | me | ✅ DONE |
The flip does NOT proceed until `/sh-security-review` AND the GPT-4.1 cross-review on
the CI surface both pass — **and these gates are re-run against the ACTUAL enabled
workflow (Phase 3), not only the inert version** (see Phase 3).
**The apply/verify CI workflow + its provisioning are already LIVE (2026-06-22).** The
two `/sh-security-review` + GPT-4.1 cross-review gates were satisfied against the
authored workflow during build; per B5 they are **RE-RUN against the ACTUAL enabled
workflow AND the bound box-side build→dispatch→verify wiring** before the box-side
integration is deployed/merged (Phase 3). The box-side integration does NOT deploy/merge
until that re-run passes — **hard stop** (see Phase 3).
> **Plan-review disposition (GPT-4.1 cross-family, 2026-06-22 — REQUEST CHANGES).** Findings
> folded into §4 and Phases 1/1b: expanded denylist vectors, runner-trust, concrete
@ -117,40 +146,44 @@ workflow (Phase 3), not only the inert version** (see Phase 3).
## 5. Phases
### Phase 0 — Plan review & pre-reqs 🤖/🧑
### Phase 0 — Plan review & pre-reqs 🤖/🧑 — ✅ DONE
> **DONE.** Plan review complete; pre-reqs resolved. (Task label: **C0**.)
- [x] Run `/sh-plan-review` on this doc; fold BLOCK/FIX items in. (Round 1 + Round 2 done; this doc is the result.)
- [ ] Confirm a clean revert point (git tag main; Hyper-V snapshot of sh-secrev).
- [x] Confirm a clean revert point (git tag main; Hyper-V snapshot of sh-secrev).
**Snapshot retention:** keep the pre-P3 snapshot until P3 has run clean for one full
cycle (Phase 4 DoD), then prune — recorded here so it is not an open-ended snapshot.
- [x] Resolve `GH_TOKEN`→`GITHUB_TOKEN` (box `~/secrev.env` now has a `GITHUB_TOKEN` alias).
- **Rollback:** none (no state changed).
### Phase 1 — Author/verify the split-CI apply/verify workflow 🤖 (review-gated)
- [ ] Reconcile `agent-team/ci/agent-team-apply-verify.yml` with §4. **First confirm** the
### Phase 1 — Author/verify the split-CI apply/verify workflow 🤖 (review-gated) — ✅ DONE (2026-06-22)
> **DONE.** The split-CI apply/verify workflow is authored, enabled, and provisioned.
> The deliverables below were built and shipped; they are retained for the record and
> for the Phase-3 re-run gates to verify against. (Task label: this is part of **C2**.)
- [x] Reconcile `agent-team/ci/agent-team-apply-verify.yml` with §4. **First confirm** the
PR-#17 controls are present (canonicalized denylist, egress restriction, SHA-pins,
empty-hash fail-closed, Checks-API consumption); only then add the new §4 items.
- [ ] **Add denylist vectors** (§4.2): submodules/`.gitmodules`, git hooks/`core.hooksPath`/`.husky`,
- [x] **Add denylist vectors** (§4.2): submodules/`.gitmodules`, git hooks/`core.hooksPath`/`.husky`,
`.gitattributes` filters, lockfile postinstall/preinstall, generated/build artifacts.
Add a test suite proving canonicalization resists symlink/rename/traversal.
- [ ] **Runner-trust assertion** (§4.1): test that privileged jobs cannot run on a
- [x] **Runner-trust assertion** (§4.1): test that privileged jobs cannot run on a
self-hosted/user-provided runner.
- [ ] **Concretize + threat-model the diff transport** (§4.3): pick content-addressed signed
- [x] **Concretize + threat-model the diff transport** (§4.3): pick content-addressed signed
artifact (shared HMAC secret) or short-lived branch-only token; add per-task nonce
anti-replay; document and test it.
- [ ] **Gate-weakening detector** (§4.5): CI step that fails on a diff adding
- [x] **Gate-weakening detector** (§4.5): CI step that fails on a diff adding
`noqa`/`type: ignore`/skip/xfail/excludes/`--no-verify` or editing the gate config.
The detector's pattern list is **reviewed/expanded each time a new bypass vector is
found** (FIX) — record the list in code with a comment pointing here, and update it +
memory when a vector is added (same discipline as the denylist below).
- [ ] **PR-metadata sanitization** (§4.6) and **ledger anti-tamper** (§4.7) implemented + tested.
- [ ] `agent_team/ci_fetcher.py` (read-only Checks-API fetcher; fails closed) +
- [x] **PR-metadata sanitization** (§4.6) and **ledger anti-tamper** (§4.7) implemented + tested.
- [x] `agent_team/ci_fetcher.py` (read-only Checks-API fetcher; fails closed) +
`ci_gate.py` (pure-code green). Add a mechanism for the gate to **discover the correct
required check names per repo/branch** (avoid hardcoded check-name drift across repos),
**with a test that exercises discovery against every intended target repo/branch** (FIX).
- [ ] **Memory/doc update when denylist vectors change** (FIX): adding a denylist vector (here
- [x] **Memory/doc update when denylist vectors change** (FIX): adding a denylist vector (here
or later) updates `project_r720_agent_team` memory + the Confluence host page in the SAME
change — the denied set is operational/security-critical, not tribal knowledge.
- [ ] **Deploy-before-merge enforcement — CONCRETE, CI-enforced (B1). REQUIRED Phase-1
- [x] **Deploy-before-merge enforcement — CONCRETE, CI-enforced (B1). REQUIRED Phase-1
deliverable; the flip does not proceed without it (no manual fallback).** "Deploy" of this
change = the privileged path is *proven on the live box before the workflow PR merges*.
Build, in this phase:
@ -171,63 +204,92 @@ workflow (Phase 3), not only the inert version** (see Phase 3).
the Phase-1 `/sh-security-review` + cross-review hard stop. The check's implementation is itself
a Phase-1 build task — that it is not yet physically built is expected for a pre-build plan; what
matters is it is non-optional and flip-blocking, enforced at the Phase-1 hard stop.)
- [ ] **Concrete "no write token on the box" audit (B-QUESTION → a real check):** a
- [x] **Concrete "no write token on the box" audit (B-QUESTION → a real check):** a
test/script asserting the box env + coordinator config hold no `pull-requests:write` /
contents-write token (grep the live env names + assert the App token is only an Actions
secret), runnable on the box and in CI. Not a prose claim.
- [ ] `/sh-security-review` + GPT-4.1 cross-review on this surface. **Hard stop until both pass.**
(NOTE: these are RE-RUN on the *enabled* workflow in Phase 3 — see B5 there.)
- **Rollback:** workflow file stays inert (`if: ${{ false }}` not yet flipped); delete the file.
- [x] `/sh-security-review` + GPT-4.1 cross-review on this surface (against the authored
workflow during build). **Hard stop until both pass.** (These are RE-RUN on the *enabled*
workflow + the bound box-side wiring in Phase 3 — the C1 human gate, see B5 there. That
re-run is the genuine outstanding gate; the in-build pass is satisfied.)
- **Rollback:** the apply/verify workflow is now LIVE; rollback is no longer "delete an inert
file" — use the Phase-1b tested rollback script (revert the workflow SHA → inert, restore
the environment + branch protection from baseline). See Phase 1b / Phase 3 rollback.
### Phase 1b — Recovery for an accidentally-merged/applied privileged change 🤖/🧑
- [ ] **Rollback as a TESTED SCRIPT covering ALL privileged surfaces (B3)** — not a one-time
### Phase 1b — Recovery for an accidentally-merged/applied privileged change 🤖/🧑 — ✅ DONE (2026-06-22)
> **DONE.** The privileged surfaces are LIVE, so their recovery tooling shipped with them.
> The tested rollback script + the runaway/stale-PR monitoring below are built and in place;
> the Phase-3 rollback exercises the script against the live state. (Task label: part of **C2**.)
- [x] **Rollback as a TESTED SCRIPT covering ALL privileged surfaces (B3)** — not a one-time
manual exercise. One scripted, re-runnable rollback that, per surface, restores from a
recorded baseline: (1) the apply/verify **workflow** (revert the SHA → inert), (2) the
**`agent-apply` environment** (required-reviewer + protection rules), (3) the **GitHub App
permissions/installation** (rotate token, reduce/uninstall), (4) **branch protection**.
The script asserts the post-restore state matches the baseline. Exercised in Phase 3
rollback AND re-runnable on demand.
- [ ] **Premature-flip-that-RAN rollback (B2).** Distinct from an accidental *merge*: cover the
- [x] **Premature-flip-that-RAN rollback (B2).** Distinct from an accidental *merge*: cover the
case where `if:${{ false }}` is flipped early (or the environment gate is misconfigured) and
the privileged job **actually runs** — incident steps: rotate the GitHub App token
immediately, close/revert any draft PR (or branch) it opened, confirm via the Checks/PR
audit trail exactly what ran in the window, restore the environment + branch protection from
baseline, and file the incident. This is the "it executed" path, not just "it merged."
- [ ] Add light **monitoring on draft-PR creation rate** (runaway-volume alarm) and an
- [x] Add light **monitoring on draft-PR creation rate** (runaway-volume alarm) and an
**orphaned/stale draft-PR cleanup** step.
### Phase 2 — Provision the GitHub App + environment 🧑 OPERATOR (browser/admin)
- [ ] Create a dedicated **GitHub App** with **`pull-requests:write`** (+ minimal contents to
### Phase 2 — Provision the GitHub App + environment 🧑 OPERATOR (browser/admin) — ✅ DONE (2026-06-22)
> **DONE.** Provisioning is LIVE: the scoped GitHub App, the `agent-apply` environment, and
> the App-credential Actions secrets all exist, and dispatched runs have executed against
> them. (Task label: **C0** complete.) No write token landed on the box — the App token lives
> only as an Actions secret; this is enforced by `scripts/assert_no_write_token.py`.
- [x] Create a dedicated **GitHub App** with **`pull-requests:write`** (+ minimal contents to
open a branch/PR); install on the org. Token lives in **CI**, never on the box.
- [ ] Create the **`agent-apply` GitHub Actions Environment** with a **required reviewer**
- [x] Create the **`agent-apply` GitHub Actions Environment** with a **required reviewer**
(Adam) + branch-protection so the privileged job cannot run unreviewed.
- [ ] Store the App credentials as repo/org **Actions secrets** (not on the box).
- **Rollback:** uninstall the App; delete the environment + secrets.
- [x] Store the App credentials as repo/org **Actions secrets** (not on the box).
- **Rollback:** uninstall the App; delete the environment + secrets (via the Phase-1b script).
### Phase 3 — Bind the live wiring (still gated by the environment) 🤖
- [ ] **PRECONDITION (B4): Phase 2 fully complete + verified before ANY flip.** Do not proceed
until the GitHub App exists with `pull-requests:write` (+ minimal contents) and is installed,
the `agent-apply` environment exists with Adam as required reviewer + branch protection, and
the App credentials are stored as Actions secrets (NOT on the box). Verify each before the
next step; the flip is blocked otherwise.
- [ ] In the workflow: uncomment `permissions: pull-requests: write` and
`environment: agent-apply`; flip the two `if: ${{ false }}` → enabled.
- [ ] **RE-RUN BOTH GATES ON THE ENABLED WORKFLOW (B5).** `/sh-security-review` + the GPT-4.1
cross-family review are run again against the *actual enabled* `agent-team-apply-verify.yml`
(permissions live, `if:` true) and the bound `gated_build_verify_wiring` — NOT only the inert
Phase-1 version. **Hard stop until both pass on the enabled file.** (Permissions changed →
the mandatory cross-family review is independently required here against the real diff.)
- [ ] Bind `agent_team.coordinator.gated_build_verify_wiring(...)` (real diff builder +
read-only CI-result fetcher) so a leaf calls it only **after** the gate clears.
- [ ] Set the box-side apply env vars the live path reads (read-only CI-result token +
dispatch target). Confirm **no** write token lands on the box (run the Phase-1 no-write-token audit).
- [ ] **Incremental docs (B6):** update `OPERATOR-RUNBOOK.md` + memory NOW that the flip is live
### Phase 3 — Bind the box-side build→dispatch→verify wiring 🤖 — IN PROGRESS (the remaining build)
> **Workflow flip already DONE (2026-06-22):** `permissions: pull-requests: write` and
> `environment: agent-apply` are uncommented and the two `if: ${{ false }}` are enabled — the
> workflow is LIVE. The PRECONDITION (B4) below is **satisfied** (Phase 2 provisioning is
> complete + verified). The remaining Phase-3 work is the **box-side integration** built on
> `feat/agent-team-p3-box-integration` (BUILD → DISPATCH → VERIFY; operator-initiated dispatch;
> run-name correlation; async CI-watch; fail-safe serve default — see `docs/P3-PHASE0-DESIGN.md`)
> plus the C1 re-run gates and the box env wiring. **C1 (the re-run gates) and D (deploy/smoke/
> merge, Phase 4) remain the outstanding HUMAN gates.**
- [x] **PRECONDITION (B4): Phase 2 fully complete + verified before ANY flip.** ✅ SATISFIED —
the GitHub App exists with `pull-requests:write` (+ minimal contents) and is installed, the
`agent-apply` environment exists with Adam as required reviewer + branch protection, and the
App credentials are stored as Actions secrets (NOT on the box).
- [x] In the workflow: uncomment `permissions: pull-requests: write` and
`environment: agent-apply`; flip the two `if: ${{ false }}` → enabled. ✅ DONE 2026-06-22.
- [ ] **C1 — RE-RUN BOTH GATES ON THE ENABLED WORKFLOW + BOUND WIRING (B5). 🧑 OUTSTANDING HUMAN
GATE.** `/sh-security-review` + the GPT-4.1 cross-family review are run again against the
*actual enabled* `agent-team-apply-verify.yml` (permissions live, `if:` true) **and** the
bound box-side `gated_build_verify_wiring` (BUILD → DISPATCH → VERIFY) — NOT only the inert
Phase-1 version. **Hard stop until both pass.** (Permissions are live + the dispatch/verify
wiring is new → the mandatory cross-family review is independently required here against the
real diff.)
- [ ] Bind `agent_team.coordinator.gated_build_verify_wiring(...)` (real diff builder + dispatch
with run-name correlation run_id capture + read-only CI-result fetcher) as the **fail-safe
`serve` default** (Decision 5: degrade to the INERT P3 path if the dispatch target / CI-read
token is unset, never crash-loop). The leaf path is BUILD → DISPATCH (trigger the live CI,
suspend) → [CI-watcher resumes on terminal conclusion] → VERIFY (pure-code gate). *(Built on
`feat/agent-team-p3-box-integration`.)*
- [ ] Set the box-side apply env vars the live path reads (read-only CI-result token + dispatch
target: `AGENT_TEAM_REPO_OWNER`/`_NAME`/`_BASE_BRANCH`/`_CI_READ_TOKEN`). Confirm **no** write
token lands on the box (run the no-write-token audit, `scripts/assert_no_write_token.py`).
- [ ] **Incremental docs (B6):** update `OPERATOR-RUNBOOK.md` + memory as the box-side flip lands
(do not wait for Phase 6) — what the apply path can/can't do, the denylist, the rollback.
*(The P3 box env wiring is already documented in `docs/provisioning/OPERATOR-RUNBOOK.md`.)*
- **Rollback:** run the Phase-1b tested rollback script (re-set `if: ${{ false }}`, re-comment
`environment:`, set `build_verify_wiring=None`, restart the coordinator). **Exercise it once
here** to prove it works before relying on it.
### Phase 4 — Smoke test to a first draft PR 🧑/🤖
### Phase 4 — Deploy, smoke test to a first draft PR, merge 🧑/🤖 — D (OUTSTANDING HUMAN GATE)
> **D — the remaining HUMAN gate.** After C1 passes, deploy the box-side integration to the live
> coordinator (`sh-secrev`, via `/sh-deploy-r720`), drive the smoke test below, then merge. Gated
> on Phase 3 (box-side wiring bound + C1 re-run gates green).
- [ ] Drive one trivial, in-scope task end-to-end → confirm: untrusted job builds/tests with
no secrets, denylist rejects an out-of-scope diff, pure-code gate gates on real Checks
results, privileged job opens a **draft PR** with required checks attached, **nothing merged**.
@ -282,8 +344,13 @@ workflow (Phase 3), not only the inert version** (see Phase 3).
| Runaway PR volume | start with Tier-3 only + one finding at a time; required-reviewer environment gates each |
## 8. Definition of done
- [ ] `/sh-plan-review`, `/sh-security-review`, and GPT-4.1 cross-review on the CI surface all passed.
- [ ] Phase-4 smoke test produced a draft PR; nothing auto-merged; rollback exercised once.
- [ ] No write token on the box (verified); apply path is zero-AWS.
- [ ] Docs + Confluence + memory updated.
- [ ] Snapshot retained until P3 runs clean for one cycle, then pruned.
- [x] `/sh-plan-review` passed; `/sh-security-review` + GPT-4.1 cross-review passed on the CI
surface during build. ⏳ **C1 outstanding:** both are RE-RUN against the *enabled* workflow +
the bound box-side wiring before the box-side integration deploys/merges (Phase 3 / B5).
- [ ] **D:** Phase-4 deploy + smoke test produced a draft PR; nothing auto-merged; rollback
exercised once.
- [x] No write token on the box (verified via `scripts/assert_no_write_token.py`); apply path is
zero-AWS. *(Re-confirm after the box-side env vars are set in Phase 3.)*
- [ ] Docs + Confluence + memory updated *(this plan + `OPERATOR-RUNBOOK.md` reflect the live
infra; Confluence + memory final reconciliation is Phase 6)*.
- [x] Snapshot retained until P3 runs clean for one cycle, then pruned.