The #60 max_turns/tools fix stopped the build-node crash but exposed the next gap: with tools enabled the agentic builder reads files and its final text is NARRATION (observed live: candidate_diff = 'Let me read the key source files to get exact signatures bef…'), not a unified diff. That non-diff dispatched to CI, could not be applied, and the task parked at verify. - builders.py: default_diff_builder now extracts the unified diff from the response via _extract_unified_diff — prefers a fenced ```diff block, else slices from the first 'diff --git' header, dropping surrounding prose. A reply with NO diff header raises BuildError so narration fails closed (the node fails the task with a clear reason) instead of dispatching a bogus diff. - builders.py: _render_build_prompt now instructs the model that its FINAL message must be ONLY the unified diff in a single ```diff fenced block, no narration before/after. - tests: extract from fenced/bare diff with narration; reject narration-only (BuildError); default_diff_builder returns the clean diff from a narrated response and raises on a prose-only reply. Full suite 1503 passed; ruff clean. |
||
|---|---|---|
| .. | ||
| .security-review | ||
| agent_team | ||
| ci | ||
| docs | ||
| scripts | ||
| slack | ||
| systemd | ||
| tests | ||
| web | ||
| .gitignore | ||
| DEPLOY-R720.md | ||
| README.md | ||
| run-team.py | ||
agent-team — R720 Plane-2 SDLC pipeline
The durable, human-gated agentic SDLC pipeline for the R720 (sh-secrev VM),
design: ../docs/r720-agent-team-design.md. A task flows INTAKE → CLARIFY (the
first human gate) → PLAN → REVIEW, and (P3) → BUILD → DISPATCH → VERIFY → draft
PR. Every stage is durable and resumable (LangGraph + a SQLite checkpointer);
both human gates suspend on interrupt() and resume on a real answer.
There are now two human gates: the clarifier (CLARIFY asks question-sets
until confident) and the plan-decision gate (the dead-end when PLAN ⇄ REVIEW
cannot auto-converge). When the review loop hits its revision cap (or the planner
salvages only a partial plan), the pipeline no longer terminally PARKs — it
suspends on a resumable interrupt() and the coordinator posts the plan +
reviewer findings to Slack #agent-team, threaded under the task root, for the
owner to decide.
INTAKE → CLARIFY (Claude, human gate) → PLAN (Claude) → REVIEW (GPT-4.1)
▲ │
└── loop-back ───┤
approve → BUILD/END
review-cap / partial plan
→ PLAN-DECISION GATE (human)
approve → BUILD/END
request changes → PLAN
abandon → FAILED
(P3 box path — gated behind the C1 re-review):
approve → BUILD (DeepSeek) → DISPATCH (push branch, trigger CI,
capture run_id, suspend) → [CI-watcher resumes on terminal
conclusion] → VERIFY (ci_gate) → draft PR
Status (2026-06-23). P1 (human gate) + P2 (planner + adversarial review loop) + the live runtime (coordinator, Slack/GitHub/Claude-Code transports, intake) + P3 (build/dispatch/verify subgraph) + P4 (more transports + GitHub-issue intake) are built, reviewed, and merged to main. The CI apply/verify workflow is LIVE + provisioned (the
agent-applyenvironment, theAGENT_APPLY_APP_*secrets, and dispatched runs all exist as of 2026-06-22). The box-side integration that drives it (run_id capture, the BUILD → DISPATCH → VERIFY reorder, async CI-watch, and the fail-safe boundservedefault) is built on this branch but gated behind the C1 re-review — the/sh-security-review+ mandatory GPT-4.1 cross-review of the new CI-trigger boundary — before it ships to the box. See the project memoryproject_r720_agent_teamand §7 of the design (anddocs/P3-PHASE0-DESIGN.md) for the phased rollout.
Layout
agent-team/
run-team.py # operator CLI: init-db, list/show/answer/expire/
# force-resume (ledger), start (intake), serve (daemon),
# intake-github
agent_team/ # importable package (snake_case)
coordinator.py # the live runtime keystone: invoker→graph→ResumeWorker;
# start_task / submit_answer / drain / tick / recover / serve
graph.py # LangGraph wiring: P1 (intake→clarify→plan) + opt-in P2
# review loop + opt-in P3 build/verify subgraph
invoker.py # §3.1 real Claude path (subscription-OAuth / API / Bedrock)
invoker_multi.py # WS1 in-process non-Claude invokers (GPT-4.1 / DeepSeek /
# Gemini via the orchestrator's models.py); bind_multi_invoker()
api.py # WS1 FastAPI HTTP API (bearer auth, 127.0.0.1:8765) — SEPARATE
# opt-in process (api.serve()), NOT started by the coordinator
dashboard.py # read-only LAN status dashboard (FastAPI, 0.0.0.0:8770):
# /api/state /api/topology /api/task/{id} + serves web/dist SPA
topology.py # pipeline map derived from the compiled LangGraph
# (get_graph() + NODE_META sidecar) — new agents appear auto
status_page.py # read-only DATA LAYER for /api/state (build_snapshot /
# snapshot_to_dict); HTML rendering retired in the makeover
billing.py # §3.1 claude_invoke billing-mode seam
ci_gate.py # §3.3.2 pure-code authenticated-Checks PASS/FAIL gate
task_model.py / state_store.py
db/{schema.py,schema.sql} # SQLite ledger DDL + BEGIN IMMEDIATE compare-and-set
db/transitions.py # task_transitions recorder (per-task pipeline history,
# idempotent + fail-soft; written by instrumented graph nodes)
ledger.py / responder.py / resume_worker.py / deadline_timer.py / recovery.py
operator_cli.py
nodes/ # pipeline stages + their model bindings
clarifier.py + clarifier_llm.py # human gate (Claude)
planner.py # plan (Claude); per-call max_turns=4 +
# classified retry-once (reliability fix)
review_loop.py + review_loop_llm.py # adversarial review (GPT-4.1 via orchestrator)
builders.py + builders_llm.py # candidate diff (DeepSeek) — INERT, proposes only
verifier.py + verifier_llm.py # ci_gate sole PASS authority; LLM = fix-proposer;
# binds expected_run_id per-task from state["run_id"]
build_verify_subgraph.py # P3 BUILD→DISPATCH→VERIFY topology + the tick()-driven
# CI-watcher that resumes a suspended task on the
# dispatched run's terminal conclusion (or parks on timeout)
handbook.py # WS5 load_handbook_conventions (handbook seam,
# fail-safe → "" if dir missing); planner context
dispatch_invoker.py # P3 DISPATCH node — pushes the per-dispatch head branch,
# triggers CI (gh workflow run), and captures run_id +
# dispatched_at into state via a correlation-tagged poll
# of gh run list (fails closed on an unfound run)
transport/ # one adapter contract + a live impl per channel
base.py # Transport ABC + QuestionSet / NormalizedAnswer
slack_adapter.py + slack_live.py + slack_listener.py # Block Kit + Socket Mode + /new-task
github_adapter.py + github_live.py + github_intake.py # issue-comment + issue intake
claude_code_adapter.py + claude_code_live.py # file-drop responder
web/ # React/Vite/TypeScript status-dashboard SPA (React Flow map,
# task list, click-through task history); built to web/dist
scripts/ # deploy-r720-ws-rollout.sh — attended WS0–WS5 UPDATE of the box
ci/ # §3.3.2 split-job CI apply/verify workflow (LIVE since 2026-06-22)
systemd/ # agent-team-coordinator.service + agent-team-status.service
DEPLOY-R720.md # provisioning runbook (snapshot-first, rsync, tokens, demo)
tests/ # pytest, one module per source module + sim harness
The Claude Code plugin lives in a sibling top-level dir, ../sea-haven-claude-plugin/
(CLAUDE.md, settings.template.json, hooks/user_prompt_submit.py): a
UserPromptSubmit hook that forwards /delegate <task> prompts from Claude Code
to the HTTP API's POST /tasks (env AGENT_TEAM_API_URL / AGENT_TEAM_API_TOKEN).
The top directory is kebab-case (agent-team/); the importable package is
snake_case (agent_team/), per the engineering handbook.
Key design points
- Durable human gates (§3.3.1). The
pending_questionsledger is the single source of truth for the question lifecycle, with akinddiscriminator (clarify|plan_decision) marking which gate a row belongs to (schema v4, idempotent additive migration). Every race (duplicate answers, transport redelivery, answer-vs-timeout) resolves via one atomic compare-and-set againststatus, inside aBEGIN IMMEDIATEtransaction — first-answer-wins (rowcount == 1), late/duplicate ignored. Single-open-gate invariant: a thread holds at most one open question at a time (the clarifier row is answered before the plan stage runs), so clarifier and plan-decision gates can never be open simultaneously for one thread. The LangGraphSqliteSavercheckpointer shares the same DB file. - Plan-review decision gate. When PLAN ⇄ REVIEW cannot auto-converge
(review-revision cap) or only a partial plan is salvaged, the coordinator posts
the plan (
_summarize_plan) + reviewer findings (_summarize_blocker) to Slack#agent-teamand the task owner decides via three verbs — Approve (settle the plan → BUILD), Request changes (loop back to the planner with the notes folded into review feedback), Abandon (FAILED). The decision arrives via Block Kit buttons, a notes modal, or a free-text thread reply; free-text prose that isn't a recognized approve/abandon verb defaults to request-changes (carrying the full reply as the notes) so a change request can never be misread as an accidental approve or abandon. Bounded byMAX_PLAN_GATE_VISITS(= 3) so the human loop always terminates. - Planner reliability. The planner's single-shot Claude call runs with
max_turns=4(tools stay disabled) so it has room to finish emitting its JSON rather than exhausting the default 1-turn budget mid-reply, plus a classified retry-once: a transient failure (turn-cap exhaustion or an empty reply) is retried exactly once; a deterministic failure (malformed JSON, missing phases) fails fast. - Fail-safe model seams. Every node treats model output as untrusted and
fails SAFE: garbage never clears the 98% clarifier gate, never auto-approves a
plan, never fabricates a build success, and the verifier's
ci_gateis the sole PASS authority (the LLM is structurally a fix-proposer only). - Inbound auth (§3.3.1). The Slack Socket Mode listener authorizes the
sender against an owner allowlist (
AGENT_TEAM_SLACK_OWNER_IDS, fail-closed) on top of the open-status CAS anti-replay. - Billing seam (§3.1). Claude runs under subscription OAuth on the box;
GPT-4.1 (review) and DeepSeek (builders) route through the local orchestrator
run.py. Switching Claude billing is a config flip.
Status dashboard (WebUI)
A read-only LAN dashboard (FastAPI, 0.0.0.0:8770, no auth, mode=ro ledger
opens) for watching the pipeline. Served by agent_team.dashboard (systemd unit
agent-team-status.service):
- Live pipeline map — a React Flow graph auto-laid-out from the real
LangGraph (
topology.pyintrospectscompiled.get_graph()+ aNODE_METAdisplay sidecar). Adding an agent node ingraph.pymakes it appear on the map with no manual coordinates; nodes group into trees (processes) branching offintake. Node color = live state; loop-back edges (review→plan, verify→build) render dashed. - Click-through task history — selecting a task opens a timeline of its journey
through each node (entry/exit timestamps, per-node duration, per-stage cost, Q&A,
verdicts, plan), backed by the
task_transitionsledger (schema v3) written by the coordinator's instrumented graph nodes (db/transitions.py, fail-soft).
Endpoints: GET /api/state (live overview + per-node state), GET /api/topology
(map nodes/edges/trees), GET /api/task/{thread_id} (one task's history;
thread_id is validated ^[A-Za-z0-9_-]{1,64}$). The legacy stdlib HTML page was
retired; status_page.py remains as the /api/state data layer.
Build (on the Mac — the box Node is too old for Vite 5+):
cd agent-team/web
npm ci && npm run build # -> web/dist (gitignored), rsynced to the VM
npm test # Vitest + React Testing Library
In dev, npm run dev proxies /api to a locally running dashboard
(AGENT_TEAM_DASH, default http://127.0.0.1:8770).
WS0–WS5 rollout glossary
The "WS-rollout" (workstreams 0–5) layered HTTP/integration surfaces onto the P1–P4 pipeline. What is live vs inert after the rollout:
| WS | What it adds | Live? |
|---|---|---|
| WS1 | invoker_multi.py (in-process GPT-4.1 / DeepSeek / Gemini via the orchestrator's models.py) + api.py (FastAPI HTTP API, bearer auth via AGENT_TEAM_API_TOKEN, binds 127.0.0.1:8765, /docs+/openapi disabled, concurrency-capped) |
bind_multi_invoker() wired in run-team.py _cmd_serve (LIVE); the HTTP API is a separate opt-in process (api.serve()), NOT started by the coordinator |
| WS5 | nodes/handbook.py load_handbook_conventions (reads SEA_HAVEN_HANDBOOK_DIR or ~/.sea-haven/engineering-handbook, fail-safe → ""); retriever.py save_memory writes to a _box-drafts/ review queue |
LIVE — the planner prompt receives the handbook via the context_provider seam in run-team.py _build_coordinator |
| WS2/WS0/WS4 | Slack /new-task slash command (AUTHZ-01 owner-allowlist gated) → Coordinator.set_new_task_callback; the sea-haven-claude-plugin/ (CLAUDE.md, settings, /delegate UserPromptSubmit hook) |
LIVE (/new-task wired in serve); the plugin/HTTP-API path is opt-in |
| WS3 / P3 | nodes/dispatch_invoker.py (DISPATCH LangGraph node) + the BUILD → DISPATCH → VERIFY reorder, the tick()-driven CI-watcher, per-task run_id plumbing, and the bound serve default |
Built on feat/agent-team-p3-box-integration, gated behind the C1 re-review before it ships to the box. The agent-apply GitHub Environment human-approval gate is KEPT; dispatch is operator-initiated (the box holds no standing write token — the branch push + gh workflow run use operator-host credentials). The bound P3 wiring is the new fail-safe serve default: a missing AGENT_TEAM_REPO_OWNER/_NAME or CI-read token degrades to the INERT P3 path (task parks + a #agent-team notice), never a serve-start crash |
The HTTP API endpoints: POST /tasks (start a task), GET /tasks/{thread_id}
(status), POST /orchestrator/invoke (one-shot model invoke). See
DEPLOY-R720.md for the WS-rollout deploy (scripts/deploy-r720-ws-rollout.sh).
Running the tests
cd agent-team
python3 -m pytest -q # conftest puts the package on sys.path; no install needed
Deploy
Deploy-gated. See DEPLOY-R720.md for the provisioning runbook (VM snapshot
first, rsync, venv deps, ~/secrev.env tokens, init-db, systemd, and the live
P1 exit-criteria demo). Secrets are never committed.