This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/agent-team/README.md

128 lines
8.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# agent-team — R720 Plane-2 SDLC pipeline
The durable, human-gated agentic SDLC pipeline for the R720 (`sh-secrev` VM),
design: `../docs/r720-agent-team-design.md`. A task flows INTAKE → CLARIFY (the
human gate) → PLAN → REVIEW, and (opt-in, deploy-gated) → BUILD → VERIFY → draft
PR. Every stage is durable and resumable (LangGraph + a SQLite checkpointer);
the human gate suspends on `interrupt()` and resumes on a real answer.
```
INTAKE → CLARIFY (Claude, human gate) → PLAN (Claude) → REVIEW (GPT-4.1)
▲ │
└── loop-back ───┤
approve/escalate → END
(P3, opt-in + INERT until the CI gate clears):
approve → BUILD (DeepSeek) → VERIFY (ci_gate) → draft PR
```
> **Status (2026-06-18).** P1 (human gate) + P2 (planner + adversarial review
> loop) + the live runtime (coordinator, Slack/GitHub/Claude-Code transports,
> intake) + P3-inert (build/verify subgraph, opt-in) + P4 (more transports +
> GitHub-issue intake) are **built, reviewed, and merged to main** (~795 tests).
> **Nothing is provisioned**: not rsync'd to the box, no live tokens, no
> systemd, no live CI. Production default runs **P2** (no builders).
> **Deploy-gated / not yet built:** the P3 *live* CI apply/verify + OIDC role
> (held behind `/sh-security-review` + the mandatory GPT-4.1 cross-review), and
> all provisioning. See the project memory `project_r720_agent_team` and §7 of
> the design for the phased rollout.
## Layout
```
agent-team/
run-team.py # operator CLI: init-db, list/show/answer/expire/
# force-resume (ledger), start (intake), serve (daemon),
# intake-github
agent_team/ # importable package (snake_case)
coordinator.py # the live runtime keystone: invoker→graph→ResumeWorker;
# start_task / submit_answer / drain / tick / recover / serve
graph.py # LangGraph wiring: P1 (intake→clarify→plan) + opt-in P2
# review loop + opt-in P3 build/verify subgraph
invoker.py # §3.1 real Claude path (subscription-OAuth / API / Bedrock)
invoker_multi.py # WS1 in-process non-Claude invokers (GPT-4.1 / DeepSeek /
# Gemini via the orchestrator's models.py); bind_multi_invoker()
api.py # WS1 FastAPI HTTP API (bearer auth, 127.0.0.1:8765) — SEPARATE
# opt-in process (api.serve()), NOT started by the coordinator
billing.py # §3.1 claude_invoke billing-mode seam
ci_gate.py # §3.3.2 pure-code authenticated-Checks PASS/FAIL gate
task_model.py / state_store.py
db/{schema.py,schema.sql} # SQLite ledger DDL + BEGIN IMMEDIATE compare-and-set
ledger.py / responder.py / resume_worker.py / deadline_timer.py / recovery.py
operator_cli.py
nodes/ # pipeline stages + their model bindings
clarifier.py + clarifier_llm.py # human gate (Claude)
planner.py # plan (Claude)
review_loop.py + review_loop_llm.py # adversarial review (GPT-4.1 via orchestrator)
builders.py + builders_llm.py # candidate diff (DeepSeek) — INERT, proposes only
verifier.py + verifier_llm.py # ci_gate sole PASS authority; LLM = fix-proposer
build_verify_subgraph.py # P3 BUILD→VERIFY topology (opt-in)
handbook.py # WS5 load_handbook_conventions (handbook seam,
# fail-safe → "" if dir missing); planner context
dispatch_invoker.py # WS3 auto-dispatch node — INERT (NOT wired live)
transport/ # one adapter contract + a live impl per channel
base.py # Transport ABC + QuestionSet / NormalizedAnswer
slack_adapter.py + slack_live.py + slack_listener.py # Block Kit + Socket Mode + /new-task
github_adapter.py + github_live.py + github_intake.py # issue-comment + issue intake
claude_code_adapter.py + claude_code_live.py # file-drop responder
scripts/ # deploy-r720-ws-rollout.sh — attended WS0–WS5 UPDATE of the box
ci/ # §3.3.2 split-job CI apply/verify workflow (DEPLOY-GATED)
systemd/ # agent-team-coordinator.service (not installed)
DEPLOY-R720.md # provisioning runbook (snapshot-first, rsync, tokens, demo)
tests/ # pytest, one module per source module + sim harness
```
The Claude Code plugin lives in a sibling top-level dir, `../sea-haven-claude-plugin/`
(CLAUDE.md, settings.template.json, `hooks/user_prompt_submit.py`): a
`UserPromptSubmit` hook that forwards `/delegate <task>` prompts from Claude Code
to the HTTP API's `POST /tasks` (env `AGENT_TEAM_API_URL` / `AGENT_TEAM_API_TOKEN`).
The top directory is kebab-case (`agent-team/`); the importable package is
snake_case (`agent_team/`), per the engineering handbook.
## Key design points
- **Durable human gate (§3.3.1).** The `pending_questions` ledger is the single
source of truth for the question lifecycle. Every race (duplicate answers,
transport redelivery, answer-vs-timeout) resolves via one atomic
compare-and-set against `status`, inside a `BEGIN IMMEDIATE` transaction —
first-answer-wins (`rowcount == 1`), late/duplicate ignored. The LangGraph
`SqliteSaver` checkpointer shares the same DB file.
- **Fail-safe model seams.** Every node treats model output as untrusted and
fails SAFE: garbage never clears the 98% clarifier gate, never auto-approves a
plan, never fabricates a build success, and the verifier's `ci_gate` is the
**sole** PASS authority (the LLM is structurally a fix-proposer only).
- **Inbound auth (§3.3.1).** The Slack Socket Mode listener authorizes the
**sender** against an owner allowlist (`AGENT_TEAM_SLACK_OWNER_IDS`,
fail-closed) on top of the open-status CAS anti-replay.
- **Billing seam (§3.1).** Claude runs under subscription OAuth on the box;
GPT-4.1 (review) and DeepSeek (builders) route through the local orchestrator
`run.py`. Switching Claude billing is a config flip.
## WS0–WS5 rollout glossary
The "WS-rollout" (workstreams 0–5) layered HTTP/integration surfaces onto the
P1–P4 pipeline. What is **live** vs **inert** after the rollout:
| WS | What it adds | Live? |
|---|---|---|
| WS1 | `invoker_multi.py` (in-process GPT-4.1 / DeepSeek / Gemini via the orchestrator's `models.py`) + `api.py` (FastAPI HTTP API, bearer auth via `AGENT_TEAM_API_TOKEN`, binds `127.0.0.1:8765`, `/docs`+`/openapi` disabled, concurrency-capped) | `bind_multi_invoker()` wired in `run-team.py` `_cmd_serve` (LIVE); the **HTTP API is a separate opt-in process** (`api.serve()`), NOT started by the coordinator |
| WS5 | `nodes/handbook.py` `load_handbook_conventions` (reads `SEA_HAVEN_HANDBOOK_DIR` or `~/.sea-haven/engineering-handbook`, fail-safe → `""`); `retriever.py` `save_memory` writes to a `_box-drafts/` review queue | LIVE — the planner prompt receives the handbook via the `context_provider` seam in `run-team.py` `_build_coordinator` |
| WS2/WS0/WS4 | Slack `/new-task` slash command (AUTHZ-01 owner-allowlist gated) → `Coordinator.set_new_task_callback`; the `sea-haven-claude-plugin/` (CLAUDE.md, settings, `/delegate` `UserPromptSubmit` hook) | LIVE (`/new-task` wired in `serve`); the plugin/HTTP-API path is opt-in |
| WS3 | `nodes/dispatch_invoker.py` (auto-dispatch LangGraph node) + graph/coordinator wiring | **INERT — NOT wired live.** The `agent-apply` GitHub Environment human-approval gate is KEPT; the P3 dispatch/build-verify path stays inert pending per-task `run_id` plumbing + a CI-boundary security re-review |
The HTTP API endpoints: `POST /tasks` (start a task), `GET /tasks/{thread_id}`
(status), `POST /orchestrator/invoke` (one-shot model invoke). See
`DEPLOY-R720.md` for the WS-rollout deploy (`scripts/deploy-r720-ws-rollout.sh`).
## Running the tests
```
cd agent-team
python3 -m pytest -q # conftest puts the package on sys.path; no install needed
```
## Deploy
Deploy-gated. See `DEPLOY-R720.md` for the provisioning runbook (VM snapshot
first, rsync, venv deps, `~/secrev.env` tokens, `init-db`, systemd, and the live
P1 exit-criteria demo). Secrets are never committed.