This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/agent-team
Adam Moussa 02ff708593
feat(agent-team): P5 cross-plane loop — checker finding -> pipeline task (opt-in) (#18)
* feat(agent-team): P5 checker-finding intake module + tests

Add agent_team.transport.checker_intake: turn a confirmed, at/above-threshold
Plane-1 checker FINDING into one Plane-2 pipeline remediation task via the
committed coordinator intake entry (start_task), mirroring github_intake.

- select_findings: status==confirmed AND severity>=threshold (default high);
  unverified/suppressed/below-threshold dropped; unknown threshold rejected.
- finding_identity: stable de-dup key (finding id, else content-hash). In-memory
  set, best-effort, NOT durable across restart (ledger table is the follow-up).
- finding_task_text/_sanitize: every repo-controlled field (title, proof, repo)
  is newline/control-char neutralised and length-bounded before it reaches the
  task text or operator log (log-injection hygiene).
- load_report_findings/ingest_reports: read the exact checker report JSON shape
  (top-level object with findings[]; bare array and dir-of-*.json also accepted).

28 hermetic unit tests (stub coordinator, in-memory findings / temp reports).

* feat(agent-team): wire opt-in intake-checker run-team subcommand

Expose the P5 cross-plane loop only as a manual run-team subcommand
(intake-checker --report PATH [--threshold] [--transport] [--dry-run]),
mirroring how intake-github is exposed. NOT wired into the always-on serve
path: the loop stays opt-in/inert by default.
2026-06-18 15:56:52 -04:00
..
.security-review Resolve security-review BLOCK: CI-guard bypasses, denylist parity, force-resume 2026-06-17 15:16:12 -04:00
agent_team feat(agent-team): P5 cross-plane loop — checker finding -> pipeline task (opt-in) (#18) 2026-06-18 15:56:52 -04:00
ci feat(agent-team): P3-live CI apply/verify hardening + ci_fetcher (gate-passed, provisioning-gated) (#17) 2026-06-18 15:53:26 -04:00
systemd docs(agent-team): R720 deploy runbook + coordinator systemd unit 2026-06-18 12:56:42 -04:00
tests feat(agent-team): P5 cross-plane loop — checker finding -> pipeline task (opt-in) (#18) 2026-06-18 15:56:52 -04:00
.gitignore Plane 2 foundation: interfaces, SQLite schemas, state-store, billing seam 2026-06-17 15:16:12 -04:00
DEPLOY-R720.md docs(agent-team): R720 deploy runbook + coordinator systemd unit 2026-06-18 12:56:42 -04:00
README.md docs(agent-team): refresh README to current built state (P1-P4) 2026-06-18 14:19:00 -04:00
run-team.py feat(agent-team): P5 cross-plane loop — checker finding -> pipeline task (opt-in) (#18) 2026-06-18 15:56:52 -04:00

agent-team — R720 Plane-2 SDLC pipeline

The durable, human-gated agentic SDLC pipeline for the R720 (sh-secrev VM), design: ../docs/r720-agent-team-design.md. A task flows INTAKE → CLARIFY (the human gate) → PLAN → REVIEW, and (opt-in, deploy-gated) → BUILD → VERIFY → draft PR. Every stage is durable and resumable (LangGraph + a SQLite checkpointer); the human gate suspends on interrupt() and resumes on a real answer.

INTAKE → CLARIFY (Claude, human gate) → PLAN (Claude) → REVIEW (GPT-4.1)
                                            ▲                │
                                            └── loop-back ───┤
                                                approve/escalate → END
        (P3, opt-in + INERT until the CI gate clears):
                              approve → BUILD (DeepSeek) → VERIFY (ci_gate) → draft PR

Status (2026-06-18). P1 (human gate) + P2 (planner + adversarial review loop) + the live runtime (coordinator, Slack/GitHub/Claude-Code transports, intake) + P3-inert (build/verify subgraph, opt-in) + P4 (more transports + GitHub-issue intake) are built, reviewed, and merged to main (~795 tests). Nothing is provisioned: not rsync'd to the box, no live tokens, no systemd, no live CI. Production default runs P2 (no builders). Deploy-gated / not yet built: the P3 live CI apply/verify + OIDC role (held behind /sh-security-review + the mandatory GPT-4.1 cross-review), and all provisioning. See the project memory project_r720_agent_team and §7 of the design for the phased rollout.

Layout

agent-team/
  run-team.py                  # operator CLI: init-db, list/show/answer/expire/
                               #   force-resume (ledger), start (intake), serve (daemon),
                               #   intake-github
  agent_team/                  # importable package (snake_case)
    coordinator.py             # the live runtime keystone: invoker→graph→ResumeWorker;
                               #   start_task / submit_answer / drain / tick / recover / serve
    graph.py                   # LangGraph wiring: P1 (intake→clarify→plan) + opt-in P2
                               #   review loop + opt-in P3 build/verify subgraph
    invoker.py                 # §3.1 real Claude path (subscription-OAuth / API / Bedrock)
    billing.py                 # §3.1 claude_invoke billing-mode seam
    ci_gate.py                 # §3.3.2 pure-code authenticated-Checks PASS/FAIL gate
    task_model.py / state_store.py
    db/{schema.py,schema.sql}  # SQLite ledger DDL + BEGIN IMMEDIATE compare-and-set
    ledger.py / responder.py / resume_worker.py / deadline_timer.py / recovery.py
    operator_cli.py
    nodes/                     # pipeline stages + their model bindings
      clarifier.py + clarifier_llm.py     # human gate (Claude)
      planner.py                          # plan (Claude)
      review_loop.py + review_loop_llm.py # adversarial review (GPT-4.1 via orchestrator)
      builders.py + builders_llm.py       # candidate diff (DeepSeek) — INERT, proposes only
      verifier.py + verifier_llm.py       # ci_gate sole PASS authority; LLM = fix-proposer
      build_verify_subgraph.py            # P3 BUILD→VERIFY topology (opt-in)
    transport/                 # one adapter contract + a live impl per channel
      base.py                  # Transport ABC + QuestionSet / NormalizedAnswer
      slack_adapter.py + slack_live.py + slack_listener.py   # Block Kit + Socket Mode
      github_adapter.py + github_live.py + github_intake.py  # issue-comment + issue intake
      claude_code_adapter.py + claude_code_live.py           # file-drop responder
  ci/                          # §3.3.2 split-job CI apply/verify workflow (DEPLOY-GATED)
  systemd/                     # agent-team-coordinator.service (not installed)
  DEPLOY-R720.md               # provisioning runbook (snapshot-first, rsync, tokens, demo)
  tests/                       # pytest, one module per source module + sim harness

The top directory is kebab-case (agent-team/); the importable package is snake_case (agent_team/), per the engineering handbook.

Key design points

  • Durable human gate (§3.3.1). The pending_questions ledger is the single source of truth for the question lifecycle. Every race (duplicate answers, transport redelivery, answer-vs-timeout) resolves via one atomic compare-and-set against status, inside a BEGIN IMMEDIATE transaction — first-answer-wins (rowcount == 1), late/duplicate ignored. The LangGraph SqliteSaver checkpointer shares the same DB file.
  • Fail-safe model seams. Every node treats model output as untrusted and fails SAFE: garbage never clears the 98% clarifier gate, never auto-approves a plan, never fabricates a build success, and the verifier's ci_gate is the sole PASS authority (the LLM is structurally a fix-proposer only).
  • Inbound auth (§3.3.1). The Slack Socket Mode listener authorizes the sender against an owner allowlist (AGENT_TEAM_SLACK_OWNER_IDS, fail-closed) on top of the open-status CAS anti-replay.
  • Billing seam (§3.1). Claude runs under subscription OAuth on the box; GPT-4.1 (review) and DeepSeek (builders) route through the local orchestrator run.py. Switching Claude billing is a config flip.

Running the tests

cd agent-team
python3 -m pytest -q          # conftest puts the package on sys.path; no install needed

Deploy

Deploy-gated. See DEPLOY-R720.md for the provisioning runbook (VM snapshot first, rsync, venv deps, ~/secrev.env tokens, init-db, systemd, and the live P1 exit-criteria demo). Secrets are never committed.