This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/agent-team/README.md

217 lines
15 KiB
Markdown
Raw Normal View History

# agent-team — R720 Plane-2 SDLC pipeline
The durable, human-gated agentic SDLC pipeline for the R720 (`sh-secrev` VM),
design: `../docs/r720-agent-team-design.md`. A task flows INTAKE → CLARIFY (the
first human gate) → PLAN → REVIEW, and (P3) → BUILD → DISPATCH → VERIFY → draft
PR. Every stage is durable and resumable (LangGraph + a SQLite checkpointer);
both human gates suspend on `interrupt()` and resume on a real answer.
There are now **two human gates**: the **clarifier** (CLARIFY asks question-sets
until confident) and the **plan-decision gate** (the dead-end when PLAN ⇄ REVIEW
cannot auto-converge). When the review loop hits its revision cap (or the planner
salvages only a partial plan), the pipeline no longer terminally PARKs — it
suspends on a resumable `interrupt()` and the coordinator posts the plan +
reviewer findings to Slack `#agent-team`, threaded under the task root, for the
owner to decide.
```
INTAKE → CLARIFY (Claude, human gate) → PLAN (Claude) → REVIEW (GPT-4.1)
▲ │
└── loop-back ───┤
approve → BUILD/END
review-cap / partial plan
→ PLAN-DECISION GATE (human)
approve → BUILD/END
request changes → PLAN
abandon → FAILED
(P3 box path — gated behind the C1 re-review):
approve → BUILD (DeepSeek) → DISPATCH (push branch, trigger CI,
capture run_id, suspend) → [CI-watcher resumes on terminal
conclusion] → VERIFY (ci_gate) → draft PR
```
> **Status (2026-06-23).** P1 (human gate) + P2 (planner + adversarial review
> loop) + the live runtime (coordinator, Slack/GitHub/Claude-Code transports,
> intake) + P3 (build/dispatch/verify subgraph) + P4 (more transports +
> GitHub-issue intake) are **built, reviewed, and merged to main**.
> **The CI apply/verify workflow is LIVE + provisioned** (the `agent-apply`
> environment, the `AGENT_APPLY_APP_*` secrets, and dispatched runs all exist as
> of 2026-06-22). **The box-side integration that drives it** (run_id capture,
> the BUILD → DISPATCH → VERIFY reorder, async CI-watch, and the fail-safe bound
> `serve` default) is built on this branch but **gated behind the C1 re-review**
> — the `/sh-security-review` + mandatory GPT-4.1 cross-review of the new
> CI-trigger boundary — before it ships to the box. See the project memory
> `project_r720_agent_team` and §7 of the design (and `docs/P3-PHASE0-DESIGN.md`)
> for the phased rollout.
## Layout
```
agent-team/
run-team.py # operator CLI: init-db, list/show/answer/expire/
# force-resume (ledger), start (intake), serve (daemon),
# intake-github
agent_team/ # importable package (snake_case)
coordinator.py # the live runtime keystone: invoker→graph→ResumeWorker;
# start_task / submit_answer / drain / tick / recover / serve
graph.py # LangGraph wiring: P1 (intake→clarify→plan) + opt-in P2
# review loop + opt-in P3 build/verify subgraph
invoker.py # §3.1 real Claude path (subscription-OAuth / API / Bedrock)
invoker_multi.py # WS1 in-process non-Claude invokers (GPT-4.1 / DeepSeek /
# Gemini via the orchestrator's models.py); bind_multi_invoker()
api.py # WS1 FastAPI HTTP API (bearer auth, 127.0.0.1:8765) — SEPARATE
# opt-in process (api.serve()), NOT started by the coordinator
dashboard.py # read-only LAN status dashboard (FastAPI, 0.0.0.0:8770):
# /api/state /api/topology /api/task/{id} + serves web/dist SPA
topology.py # pipeline map derived from the compiled LangGraph
# (get_graph() + NODE_META sidecar) — new agents appear auto
status_page.py # read-only DATA LAYER for /api/state (build_snapshot /
# snapshot_to_dict); HTML rendering retired in the makeover
billing.py # §3.1 claude_invoke billing-mode seam
ci_gate.py # §3.3.2 pure-code authenticated-Checks PASS/FAIL gate
task_model.py / state_store.py
db/{schema.py,schema.sql} # SQLite ledger DDL + BEGIN IMMEDIATE compare-and-set
db/transitions.py # task_transitions recorder (per-task pipeline history,
# idempotent + fail-soft; written by instrumented graph nodes)
ledger.py / responder.py / resume_worker.py / deadline_timer.py / recovery.py
operator_cli.py
nodes/ # pipeline stages + their model bindings
clarifier.py + clarifier_llm.py # human gate (Claude)
planner.py # plan (Claude); per-call max_turns=4 +
# classified retry-once (reliability fix)
review_loop.py + review_loop_llm.py # adversarial review (GPT-4.1 via orchestrator)
builders.py + builders_llm.py # candidate diff (DeepSeek) — INERT, proposes only
verifier.py + verifier_llm.py # ci_gate sole PASS authority; LLM = fix-proposer;
# binds expected_run_id per-task from state["run_id"]
build_verify_subgraph.py # P3 BUILD→DISPATCH→VERIFY topology + the tick()-driven
# CI-watcher that resumes a suspended task on the
# dispatched run's terminal conclusion (or parks on timeout)
handbook.py # WS5 load_handbook_conventions (handbook seam,
# fail-safe → "" if dir missing); planner context
dispatch_invoker.py # P3 DISPATCH node — pushes the per-dispatch head branch,
# triggers CI (gh workflow run), and captures run_id +
# dispatched_at into state via a correlation-tagged poll
# of gh run list (fails closed on an unfound run)
transport/ # one adapter contract + a live impl per channel
base.py # Transport ABC + QuestionSet / NormalizedAnswer
slack_adapter.py + slack_live.py + slack_listener.py # Block Kit + Socket Mode + /new-task
github_adapter.py + github_live.py + github_intake.py # issue-comment + issue intake
claude_code_adapter.py + claude_code_live.py # file-drop responder
web/ # React/Vite/TypeScript status-dashboard SPA (React Flow map,
# task list, click-through task history); built to web/dist
scripts/ # deploy-r720-ws-rollout.sh — attended WS0–WS5 UPDATE of the box
ci/ # §3.3.2 split-job CI apply/verify workflow (LIVE since 2026-06-22)
systemd/ # agent-team-coordinator.service + agent-team-status.service
DEPLOY-R720.md # provisioning runbook (snapshot-first, rsync, tokens, demo)
tests/ # pytest, one module per source module + sim harness
```
The Claude Code plugin lives in a sibling top-level dir, `../sea-haven-claude-plugin/`
(CLAUDE.md, settings.template.json, `hooks/user_prompt_submit.py`): a
`UserPromptSubmit` hook that forwards `/delegate <task>` prompts from Claude Code
to the HTTP API's `POST /tasks` (env `AGENT_TEAM_API_URL` / `AGENT_TEAM_API_TOKEN`).
The top directory is kebab-case (`agent-team/`); the importable package is
snake_case (`agent_team/`), per the engineering handbook.
## Key design points
- **Durable human gates (§3.3.1).** The `pending_questions` ledger is the single
source of truth for the question lifecycle, with a `kind` discriminator
(`clarify` | `plan_decision`) marking which gate a row belongs to (schema v4,
idempotent additive migration). Every race (duplicate answers, transport
redelivery, answer-vs-timeout) resolves via one atomic compare-and-set against
`status`, inside a `BEGIN IMMEDIATE` transaction — first-answer-wins
(`rowcount == 1`), late/duplicate ignored. **Single-open-gate invariant:** a
thread holds at most one open question at a time (the clarifier row is answered
before the plan stage runs), so clarifier and plan-decision gates can never be
open simultaneously for one thread. The LangGraph `SqliteSaver` checkpointer
shares the same DB file.
- **Plan-review decision gate.** When PLAN ⇄ REVIEW cannot auto-converge
(review-revision cap) or only a partial plan is salvaged, the coordinator posts
the plan (`_summarize_plan`) + reviewer findings (`_summarize_blocker`) to Slack
`#agent-team` and the task owner decides via three verbs — **Approve** (settle
the plan → BUILD), **Request changes** (loop back to the planner with the notes
folded into review feedback), **Abandon** (FAILED). The decision arrives via
Block Kit buttons, a notes modal, or a free-text thread reply; **free-text
prose that isn't a recognized approve/abandon verb defaults to request-changes**
(carrying the full reply as the notes) so a change request can never be
misread as an accidental approve or abandon. Bounded by `MAX_PLAN_GATE_VISITS`
(= 3) so the human loop always terminates.
- **Planner reliability.** The planner's single-shot Claude call runs with
`max_turns=4` (tools stay disabled) so it has room to finish emitting its JSON
rather than exhausting the default 1-turn budget mid-reply, plus a classified
retry-once: a *transient* failure (turn-cap exhaustion or an empty reply) is
retried exactly once; a *deterministic* failure (malformed JSON, missing
phases) fails fast.
- **Fail-safe model seams.** Every node treats model output as untrusted and
fails SAFE: garbage never clears the 98% clarifier gate, never auto-approves a
plan, never fabricates a build success, and the verifier's `ci_gate` is the
**sole** PASS authority (the LLM is structurally a fix-proposer only).
- **Inbound auth (§3.3.1).** The Slack Socket Mode listener authorizes the
**sender** against an owner allowlist (`AGENT_TEAM_SLACK_OWNER_IDS`,
fail-closed) on top of the open-status CAS anti-replay.
- **Billing seam (§3.1).** Claude runs under subscription OAuth on the box;
GPT-4.1 (review) and DeepSeek (builders) route through the local orchestrator
`run.py`. Switching Claude billing is a config flip.
## Status dashboard (WebUI)
A read-only LAN dashboard (FastAPI, `0.0.0.0:8770`, no auth, `mode=ro` ledger
opens) for watching the pipeline. Served by `agent_team.dashboard` (systemd unit
`agent-team-status.service`):
- **Live pipeline map** — a React Flow graph **auto-laid-out from the real
LangGraph** (`topology.py` introspects `compiled.get_graph()` + a `NODE_META`
display sidecar). Adding an agent node in `graph.py` makes it appear on the map
with no manual coordinates; nodes group into **trees** (processes) branching off
`intake`. Node color = live state; loop-back edges (review→plan, verify→build)
render dashed.
- **Click-through task history** — selecting a task opens a timeline of its journey
through each node (entry/exit timestamps, per-node duration, per-stage cost, Q&A,
verdicts, plan), backed by the `task_transitions` ledger (schema v3) written by
the coordinator's instrumented graph nodes (`db/transitions.py`, fail-soft).
Endpoints: `GET /api/state` (live overview + per-node state), `GET /api/topology`
(map nodes/edges/trees), `GET /api/task/{thread_id}` (one task's history;
`thread_id` is validated `^[A-Za-z0-9_-]{1,64}$`). The legacy stdlib HTML page was
retired; `status_page.py` remains as the `/api/state` data layer.
**Build (on the Mac — the box Node is too old for Vite 5+):**
```
cd agent-team/web
npm ci && npm run build # -> web/dist (gitignored), rsynced to the VM
npm test # Vitest + React Testing Library
```
In dev, `npm run dev` proxies `/api` to a locally running dashboard
(`AGENT_TEAM_DASH`, default `http://127.0.0.1:8770`).
## WS0–WS5 rollout glossary
The "WS-rollout" (workstreams 0–5) layered HTTP/integration surfaces onto the
P1–P4 pipeline. What is **live** vs **inert** after the rollout:
| WS | What it adds | Live? |
|---|---|---|
| WS1 | `invoker_multi.py` (in-process GPT-4.1 / DeepSeek / Gemini via the orchestrator's `models.py`) + `api.py` (FastAPI HTTP API, bearer auth via `AGENT_TEAM_API_TOKEN`, binds `127.0.0.1:8765`, `/docs`+`/openapi` disabled, concurrency-capped) | `bind_multi_invoker()` wired in `run-team.py` `_cmd_serve` (LIVE); the **HTTP API is a separate opt-in process** (`api.serve()`), NOT started by the coordinator |
| WS5 | `nodes/handbook.py` `load_handbook_conventions` (reads `SEA_HAVEN_HANDBOOK_DIR` or `~/.sea-haven/engineering-handbook`, fail-safe → `""`); `retriever.py` `save_memory` writes to a `_box-drafts/` review queue | LIVE — the planner prompt receives the handbook via the `context_provider` seam in `run-team.py` `_build_coordinator` |
| WS2/WS0/WS4 | Slack `/new-task` slash command (AUTHZ-01 owner-allowlist gated) → `Coordinator.set_new_task_callback`; the `sea-haven-claude-plugin/` (CLAUDE.md, settings, `/delegate` `UserPromptSubmit` hook) | LIVE (`/new-task` wired in `serve`); the plugin/HTTP-API path is opt-in |
| WS3 / P3 | `nodes/dispatch_invoker.py` (DISPATCH LangGraph node) + the BUILD → DISPATCH → VERIFY reorder, the `tick()`-driven CI-watcher, per-task `run_id` plumbing, and the bound `serve` default | **Built on `feat/agent-team-p3-box-integration`, gated behind the C1 re-review** before it ships to the box. The `agent-apply` GitHub Environment human-approval gate is KEPT; dispatch is **operator-initiated** (the box holds no standing write token — the branch push + `gh workflow run` use operator-host credentials). The bound P3 wiring is the new fail-safe `serve` default: a missing `AGENT_TEAM_REPO_OWNER`/`_NAME` or CI-read token degrades to the INERT P3 path (task parks + a `#agent-team` notice), never a serve-start crash |
The HTTP API endpoints: `POST /tasks` (start a task), `GET /tasks/{thread_id}`
(status), `POST /orchestrator/invoke` (one-shot model invoke). See
`DEPLOY-R720.md` for the WS-rollout deploy (`scripts/deploy-r720-ws-rollout.sh`).
## Running the tests
```
cd agent-team
python3 -m pytest -q # conftest puts the package on sys.path; no install needed
```
## Deploy
Deploy-gated. See `DEPLOY-R720.md` for the provisioning runbook (VM snapshot
first, rsync, venv deps, `~/secrev.env` tokens, `init-db`, systemd, and the live
P1 exit-criteria demo). Secrets are never committed.