This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/agent-team
Adam Moussa 3d97139300
feat(agent-team): P3-live CI apply/verify hardening + ci_fetcher (gate-passed, provisioning-gated) (#17)
* feat(agent-team): read-only CI-result fetcher for P3 verify gate (opt-in, inert)

ci_fetcher.py: fail-closed CiResultFetcher reading the GitHub Actions run
conclusion via a read-only PAT (AGENT_TEAM_CI_READ_TOKEN→GITHUB_TOKEN), returns
{run_id,conclusion,diff_hash} or None on any error. Data-fetcher only — ci_gate
owns the verdict; never writes, no OIDC/AWS, never reads patch artifacts.
coordinator gains opt-in gated_build_verify_wiring() composing it via
bind_ci_result_fetcher; NOT wired into the default run-team.py path. 20 tests.

* harden(agent-team): P3 apply/verify workflow — GitHub App token, CWE-94, fail-closed

Decision-1 auth model: gate-and-pr uses a GitHub App installation token
(pull-requests:write) behind the agent-apply environment; ALL OIDC/id-token/AWS
removed. Hardening: task_id env-indirection (CWE-94 — GitHub expands ${{ }} into
the run shell before exec, so %s/quoting is insufficient); run-id pinning on both
download-artifact; post-build denied-path check (build-hook writes into denied
paths fail the job); empty-hash fail-closed in BOTH the embedded gate (fixed a
real ''=='' pass bug) and ci_gate.py. App-token + draft-PR steps stay if:${{ false }}
until provisioning (App + environment + branch protection). +17 tests.

* harden(agent-team): apply P3-live security-gate fixes (GPT-4.1 xreview + sh-security-review)

BLOCK-1/FIX-4: gate-and-pr re-comments pull-requests:write + environment:agent-apply
(provisioning-time uncomment) and gains needs.guard/build-test=='success' job guard —
zero privilege until provisioning. BLOCK-2/3+FIX-5: ci_fetcher validates run_id (^[0-9]{1,20}$),
owner/repo (^[A-Za-z0-9_.-]{1,100}$), and fetched_id (int) — fail closed, no SSRF/path
injection. FIX-1: conclusion allowlist. FIX-3: api_root removed from public builder (no
injectable endpoint). INJ-02: post-build denied-path check uses NUL-delimited git output +
explicit rename parsing, no backslash mangling, non-UTF8=violation. INJ-03: all three trust-
control denylists unified to one 22-entry union + drift-guard test. Q1: documented run_id/
diff_hash trust source (dispatcher/ledger only). 884 tests, ruff clean. Privileged steps stay
if:${{ false }} until provisioning.

* build(security-review): prune .claude worktrees from deterministic scanners

Agent worktrees under .claude/worktrees/ are full repo copies; the cfn-lint
find|xargs template scan overflowed ('command line cannot be assembled') and the
pre-push hook fail-closed to BLOCK whenever a worktree was present. Prune .claude
in the cfn-lint find + semgrep/checkov excludes, and gitignore .claude/ so it is
never scanned or committed. Unblocks main-tree pushes during parallel agent work.
2026-06-18 15:53:26 -04:00
..
.security-review Resolve security-review BLOCK: CI-guard bypasses, denylist parity, force-resume 2026-06-17 15:16:12 -04:00
agent_team feat(agent-team): P3-live CI apply/verify hardening + ci_fetcher (gate-passed, provisioning-gated) (#17) 2026-06-18 15:53:26 -04:00
ci feat(agent-team): P3-live CI apply/verify hardening + ci_fetcher (gate-passed, provisioning-gated) (#17) 2026-06-18 15:53:26 -04:00
systemd docs(agent-team): R720 deploy runbook + coordinator systemd unit 2026-06-18 12:56:42 -04:00
tests feat(agent-team): P3-live CI apply/verify hardening + ci_fetcher (gate-passed, provisioning-gated) (#17) 2026-06-18 15:53:26 -04:00
.gitignore Plane 2 foundation: interfaces, SQLite schemas, state-store, billing seam 2026-06-17 15:16:12 -04:00
DEPLOY-R720.md docs(agent-team): R720 deploy runbook + coordinator systemd unit 2026-06-18 12:56:42 -04:00
README.md docs(agent-team): refresh README to current built state (P1-P4) 2026-06-18 14:19:00 -04:00
run-team.py feat(secrev): Plane-1 Phase 2 — coordinator + dependency-cve checker (#16) 2026-06-18 15:08:58 -04:00

agent-team — R720 Plane-2 SDLC pipeline

The durable, human-gated agentic SDLC pipeline for the R720 (sh-secrev VM), design: ../docs/r720-agent-team-design.md. A task flows INTAKE → CLARIFY (the human gate) → PLAN → REVIEW, and (opt-in, deploy-gated) → BUILD → VERIFY → draft PR. Every stage is durable and resumable (LangGraph + a SQLite checkpointer); the human gate suspends on interrupt() and resumes on a real answer.

INTAKE → CLARIFY (Claude, human gate) → PLAN (Claude) → REVIEW (GPT-4.1)
                                            ▲                │
                                            └── loop-back ───┤
                                                approve/escalate → END
        (P3, opt-in + INERT until the CI gate clears):
                              approve → BUILD (DeepSeek) → VERIFY (ci_gate) → draft PR

Status (2026-06-18). P1 (human gate) + P2 (planner + adversarial review loop) + the live runtime (coordinator, Slack/GitHub/Claude-Code transports, intake) + P3-inert (build/verify subgraph, opt-in) + P4 (more transports + GitHub-issue intake) are built, reviewed, and merged to main (~795 tests). Nothing is provisioned: not rsync'd to the box, no live tokens, no systemd, no live CI. Production default runs P2 (no builders). Deploy-gated / not yet built: the P3 live CI apply/verify + OIDC role (held behind /sh-security-review + the mandatory GPT-4.1 cross-review), and all provisioning. See the project memory project_r720_agent_team and §7 of the design for the phased rollout.

Layout

agent-team/
  run-team.py                  # operator CLI: init-db, list/show/answer/expire/
                               #   force-resume (ledger), start (intake), serve (daemon),
                               #   intake-github
  agent_team/                  # importable package (snake_case)
    coordinator.py             # the live runtime keystone: invoker→graph→ResumeWorker;
                               #   start_task / submit_answer / drain / tick / recover / serve
    graph.py                   # LangGraph wiring: P1 (intake→clarify→plan) + opt-in P2
                               #   review loop + opt-in P3 build/verify subgraph
    invoker.py                 # §3.1 real Claude path (subscription-OAuth / API / Bedrock)
    billing.py                 # §3.1 claude_invoke billing-mode seam
    ci_gate.py                 # §3.3.2 pure-code authenticated-Checks PASS/FAIL gate
    task_model.py / state_store.py
    db/{schema.py,schema.sql}  # SQLite ledger DDL + BEGIN IMMEDIATE compare-and-set
    ledger.py / responder.py / resume_worker.py / deadline_timer.py / recovery.py
    operator_cli.py
    nodes/                     # pipeline stages + their model bindings
      clarifier.py + clarifier_llm.py     # human gate (Claude)
      planner.py                          # plan (Claude)
      review_loop.py + review_loop_llm.py # adversarial review (GPT-4.1 via orchestrator)
      builders.py + builders_llm.py       # candidate diff (DeepSeek) — INERT, proposes only
      verifier.py + verifier_llm.py       # ci_gate sole PASS authority; LLM = fix-proposer
      build_verify_subgraph.py            # P3 BUILD→VERIFY topology (opt-in)
    transport/                 # one adapter contract + a live impl per channel
      base.py                  # Transport ABC + QuestionSet / NormalizedAnswer
      slack_adapter.py + slack_live.py + slack_listener.py   # Block Kit + Socket Mode
      github_adapter.py + github_live.py + github_intake.py  # issue-comment + issue intake
      claude_code_adapter.py + claude_code_live.py           # file-drop responder
  ci/                          # §3.3.2 split-job CI apply/verify workflow (DEPLOY-GATED)
  systemd/                     # agent-team-coordinator.service (not installed)
  DEPLOY-R720.md               # provisioning runbook (snapshot-first, rsync, tokens, demo)
  tests/                       # pytest, one module per source module + sim harness

The top directory is kebab-case (agent-team/); the importable package is snake_case (agent_team/), per the engineering handbook.

Key design points

  • Durable human gate (§3.3.1). The pending_questions ledger is the single source of truth for the question lifecycle. Every race (duplicate answers, transport redelivery, answer-vs-timeout) resolves via one atomic compare-and-set against status, inside a BEGIN IMMEDIATE transaction — first-answer-wins (rowcount == 1), late/duplicate ignored. The LangGraph SqliteSaver checkpointer shares the same DB file.
  • Fail-safe model seams. Every node treats model output as untrusted and fails SAFE: garbage never clears the 98% clarifier gate, never auto-approves a plan, never fabricates a build success, and the verifier's ci_gate is the sole PASS authority (the LLM is structurally a fix-proposer only).
  • Inbound auth (§3.3.1). The Slack Socket Mode listener authorizes the sender against an owner allowlist (AGENT_TEAM_SLACK_OWNER_IDS, fail-closed) on top of the open-status CAS anti-replay.
  • Billing seam (§3.1). Claude runs under subscription OAuth on the box; GPT-4.1 (review) and DeepSeek (builders) route through the local orchestrator run.py. Switching Claude billing is a config flip.

Running the tests

cd agent-team
python3 -m pytest -q          # conftest puts the package on sys.path; no install needed

Deploy

Deploy-gated. See DEPLOY-R720.md for the provisioning runbook (VM snapshot first, rsync, venv deps, ~/secrev.env tokens, init-db, systemd, and the live P1 exit-criteria demo). Secrets are never committed.