This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/agent-team
Adam Moussa ad6f31c115 fix(agent-team): wire the DeepSeek mechanical-edit builder as the live diff_builder (#60 root fix)
The real #60 root cause: failsafe_production_p3_wiring never bound a diff_builder,
so the build node fell back to builders.default_diff_builder (Claude). The Claude
agentic builder (max_turns + read-only tools) then exhausted its turn cap reading
files before it could emit a diff — 'Reached maximum number of turns (8)' live.

The intended builder already exists: builders_llm.as_diff_builder() routes the
build through DeepSeek fast_coder as a SINGLE model completion (with the
orchestrator's retrieval context + defensive _extract_diff + fail-safe no-op).
A single completion has no agent turn loop, so it CANNOT exhaust turns. Confirmed
get_fast_coder() works in the daemon environment.

- coordinator.py: failsafe_production_p3_wiring (P3-configured branch) now binds
  builders_llm.as_diff_builder() as the diff_builder. Without it the build node
  silently used the turn-exhausting Claude fallback.
- builders.py: default_diff_builder reverted to a safe SINGLE-SHOT, TOOL-LESS
  fallback (max_turns=4, no allowed_tools/budget) — tools are what consumed the
  turns; this fallback is no longer the live builder. Keeps _extract_unified_diff.
- tests: failsafe binds a real (callable) diff_builder when P3 is configured;
  default_diff_builder is single-shot + tool-less.

Full suite 1504 passed; ruff clean.
2026-06-24 14:58:48 -04:00
..
.security-review Resolve security-review BLOCK: CI-guard bypasses, denylist parity, force-resume 2026-06-17 15:16:12 -04:00
agent_team fix(agent-team): wire the DeepSeek mechanical-edit builder as the live diff_builder (#60 root fix) 2026-06-24 14:58:48 -04:00
ci feat(agent-team): P3 Phases A/B/E — safety tooling, wiring, docs 2026-06-23 19:52:04 -04:00
docs feat(agent-team): capture dispatched run_id for P3 box-side verify 2026-06-23 19:52:04 -04:00
scripts test(agent-team): assemble PEM test data at runtime (avoid gitleaks FP) 2026-06-23 19:52:04 -04:00
slack feat(agent-team): 👍-acknowledge received Slack answers (reactions:write) 2026-06-23 15:49:40 -04:00
systemd feat(agent-team): P3 Phases A/B/E — safety tooling, wiring, docs 2026-06-23 19:52:04 -04:00
tests fix(agent-team): wire the DeepSeek mechanical-edit builder as the live diff_builder (#60 root fix) 2026-06-24 14:58:48 -04:00
web feat(agent-team): cleaner retry-loop display + human-readable task history 2026-06-24 12:42:55 -04:00
.gitignore Plane 2 foundation: interfaces, SQLite schemas, state-store, billing seam 2026-06-17 15:16:12 -04:00
DEPLOY-R720.md docs(agent-team): document the plan-review decision gate (Phase C) 2026-06-24 11:30:54 -04:00
README.md docs(agent-team): document the plan-review decision gate (Phase C) 2026-06-24 11:30:54 -04:00
run-team.py feat(agent-team): operator dispatch command + runbook fixes 2026-06-23 20:49:09 -04:00

agent-team — R720 Plane-2 SDLC pipeline

The durable, human-gated agentic SDLC pipeline for the R720 (sh-secrev VM), design: ../docs/r720-agent-team-design.md. A task flows INTAKE → CLARIFY (the first human gate) → PLAN → REVIEW, and (P3) → BUILD → DISPATCH → VERIFY → draft PR. Every stage is durable and resumable (LangGraph + a SQLite checkpointer); both human gates suspend on interrupt() and resume on a real answer.

There are now two human gates: the clarifier (CLARIFY asks question-sets until confident) and the plan-decision gate (the dead-end when PLAN ⇄ REVIEW cannot auto-converge). When the review loop hits its revision cap (or the planner salvages only a partial plan), the pipeline no longer terminally PARKs — it suspends on a resumable interrupt() and the coordinator posts the plan + reviewer findings to Slack #agent-team, threaded under the task root, for the owner to decide.

INTAKE → CLARIFY (Claude, human gate) → PLAN (Claude) → REVIEW (GPT-4.1)
                                            ▲                │
                                            └── loop-back ───┤
                                                approve → BUILD/END
                                                review-cap / partial plan
                                                  → PLAN-DECISION GATE (human)
                                                      approve → BUILD/END
                                                      request changes → PLAN
                                                      abandon → FAILED
        (P3 box path — gated behind the C1 re-review):
            approve → BUILD (DeepSeek) → DISPATCH (push branch, trigger CI,
                      capture run_id, suspend) → [CI-watcher resumes on terminal
                      conclusion] → VERIFY (ci_gate) → draft PR

Status (2026-06-23). P1 (human gate) + P2 (planner + adversarial review loop) + the live runtime (coordinator, Slack/GitHub/Claude-Code transports, intake) + P3 (build/dispatch/verify subgraph) + P4 (more transports + GitHub-issue intake) are built, reviewed, and merged to main. The CI apply/verify workflow is LIVE + provisioned (the agent-apply environment, the AGENT_APPLY_APP_* secrets, and dispatched runs all exist as of 2026-06-22). The box-side integration that drives it (run_id capture, the BUILD → DISPATCH → VERIFY reorder, async CI-watch, and the fail-safe bound serve default) is built on this branch but gated behind the C1 re-review — the /sh-security-review + mandatory GPT-4.1 cross-review of the new CI-trigger boundary — before it ships to the box. See the project memory project_r720_agent_team and §7 of the design (and docs/P3-PHASE0-DESIGN.md) for the phased rollout.

Layout

agent-team/
  run-team.py                  # operator CLI: init-db, list/show/answer/expire/
                               #   force-resume (ledger), start (intake), serve (daemon),
                               #   intake-github
  agent_team/                  # importable package (snake_case)
    coordinator.py             # the live runtime keystone: invoker→graph→ResumeWorker;
                               #   start_task / submit_answer / drain / tick / recover / serve
    graph.py                   # LangGraph wiring: P1 (intake→clarify→plan) + opt-in P2
                               #   review loop + opt-in P3 build/verify subgraph
    invoker.py                 # §3.1 real Claude path (subscription-OAuth / API / Bedrock)
    invoker_multi.py           # WS1 in-process non-Claude invokers (GPT-4.1 / DeepSeek /
                               #   Gemini via the orchestrator's models.py); bind_multi_invoker()
    api.py                     # WS1 FastAPI HTTP API (bearer auth, 127.0.0.1:8765) — SEPARATE
                               #   opt-in process (api.serve()), NOT started by the coordinator
    dashboard.py               # read-only LAN status dashboard (FastAPI, 0.0.0.0:8770):
                               #   /api/state /api/topology /api/task/{id} + serves web/dist SPA
    topology.py                # pipeline map derived from the compiled LangGraph
                               #   (get_graph() + NODE_META sidecar) — new agents appear auto
    status_page.py             # read-only DATA LAYER for /api/state (build_snapshot /
                               #   snapshot_to_dict); HTML rendering retired in the makeover
    billing.py                 # §3.1 claude_invoke billing-mode seam
    ci_gate.py                 # §3.3.2 pure-code authenticated-Checks PASS/FAIL gate
    task_model.py / state_store.py
    db/{schema.py,schema.sql}  # SQLite ledger DDL + BEGIN IMMEDIATE compare-and-set
    db/transitions.py          # task_transitions recorder (per-task pipeline history,
                               #   idempotent + fail-soft; written by instrumented graph nodes)
    ledger.py / responder.py / resume_worker.py / deadline_timer.py / recovery.py
    operator_cli.py
    nodes/                     # pipeline stages + their model bindings
      clarifier.py + clarifier_llm.py     # human gate (Claude)
      planner.py                          # plan (Claude); per-call max_turns=4 +
                                          #   classified retry-once (reliability fix)
      review_loop.py + review_loop_llm.py # adversarial review (GPT-4.1 via orchestrator)
      builders.py + builders_llm.py       # candidate diff (DeepSeek) — INERT, proposes only
      verifier.py + verifier_llm.py       # ci_gate sole PASS authority; LLM = fix-proposer;
                                          #   binds expected_run_id per-task from state["run_id"]
      build_verify_subgraph.py            # P3 BUILD→DISPATCH→VERIFY topology + the tick()-driven
                                          #   CI-watcher that resumes a suspended task on the
                                          #   dispatched run's terminal conclusion (or parks on timeout)
      handbook.py                         # WS5 load_handbook_conventions (handbook seam,
                                          #   fail-safe → "" if dir missing); planner context
      dispatch_invoker.py                 # P3 DISPATCH node — pushes the per-dispatch head branch,
                                          #   triggers CI (gh workflow run), and captures run_id +
                                          #   dispatched_at into state via a correlation-tagged poll
                                          #   of gh run list (fails closed on an unfound run)
    transport/                 # one adapter contract + a live impl per channel
      base.py                  # Transport ABC + QuestionSet / NormalizedAnswer
      slack_adapter.py + slack_live.py + slack_listener.py   # Block Kit + Socket Mode + /new-task
      github_adapter.py + github_live.py + github_intake.py  # issue-comment + issue intake
      claude_code_adapter.py + claude_code_live.py           # file-drop responder
  web/                         # React/Vite/TypeScript status-dashboard SPA (React Flow map,
                               #   task list, click-through task history); built to web/dist
  scripts/                     # deploy-r720-ws-rollout.sh — attended WS0–WS5 UPDATE of the box
  ci/                          # §3.3.2 split-job CI apply/verify workflow (LIVE since 2026-06-22)
  systemd/                     # agent-team-coordinator.service + agent-team-status.service
  DEPLOY-R720.md               # provisioning runbook (snapshot-first, rsync, tokens, demo)
  tests/                       # pytest, one module per source module + sim harness

The Claude Code plugin lives in a sibling top-level dir, ../sea-haven-claude-plugin/ (CLAUDE.md, settings.template.json, hooks/user_prompt_submit.py): a UserPromptSubmit hook that forwards /delegate <task> prompts from Claude Code to the HTTP API's POST /tasks (env AGENT_TEAM_API_URL / AGENT_TEAM_API_TOKEN).

The top directory is kebab-case (agent-team/); the importable package is snake_case (agent_team/), per the engineering handbook.

Key design points

  • Durable human gates (§3.3.1). The pending_questions ledger is the single source of truth for the question lifecycle, with a kind discriminator (clarify | plan_decision) marking which gate a row belongs to (schema v4, idempotent additive migration). Every race (duplicate answers, transport redelivery, answer-vs-timeout) resolves via one atomic compare-and-set against status, inside a BEGIN IMMEDIATE transaction — first-answer-wins (rowcount == 1), late/duplicate ignored. Single-open-gate invariant: a thread holds at most one open question at a time (the clarifier row is answered before the plan stage runs), so clarifier and plan-decision gates can never be open simultaneously for one thread. The LangGraph SqliteSaver checkpointer shares the same DB file.
  • Plan-review decision gate. When PLAN ⇄ REVIEW cannot auto-converge (review-revision cap) or only a partial plan is salvaged, the coordinator posts the plan (_summarize_plan) + reviewer findings (_summarize_blocker) to Slack #agent-team and the task owner decides via three verbs — Approve (settle the plan → BUILD), Request changes (loop back to the planner with the notes folded into review feedback), Abandon (FAILED). The decision arrives via Block Kit buttons, a notes modal, or a free-text thread reply; free-text prose that isn't a recognized approve/abandon verb defaults to request-changes (carrying the full reply as the notes) so a change request can never be misread as an accidental approve or abandon. Bounded by MAX_PLAN_GATE_VISITS (= 3) so the human loop always terminates.
  • Planner reliability. The planner's single-shot Claude call runs with max_turns=4 (tools stay disabled) so it has room to finish emitting its JSON rather than exhausting the default 1-turn budget mid-reply, plus a classified retry-once: a transient failure (turn-cap exhaustion or an empty reply) is retried exactly once; a deterministic failure (malformed JSON, missing phases) fails fast.
  • Fail-safe model seams. Every node treats model output as untrusted and fails SAFE: garbage never clears the 98% clarifier gate, never auto-approves a plan, never fabricates a build success, and the verifier's ci_gate is the sole PASS authority (the LLM is structurally a fix-proposer only).
  • Inbound auth (§3.3.1). The Slack Socket Mode listener authorizes the sender against an owner allowlist (AGENT_TEAM_SLACK_OWNER_IDS, fail-closed) on top of the open-status CAS anti-replay.
  • Billing seam (§3.1). Claude runs under subscription OAuth on the box; GPT-4.1 (review) and DeepSeek (builders) route through the local orchestrator run.py. Switching Claude billing is a config flip.

Status dashboard (WebUI)

A read-only LAN dashboard (FastAPI, 0.0.0.0:8770, no auth, mode=ro ledger opens) for watching the pipeline. Served by agent_team.dashboard (systemd unit agent-team-status.service):

  • Live pipeline map — a React Flow graph auto-laid-out from the real LangGraph (topology.py introspects compiled.get_graph() + a NODE_META display sidecar). Adding an agent node in graph.py makes it appear on the map with no manual coordinates; nodes group into trees (processes) branching off intake. Node color = live state; loop-back edges (review→plan, verify→build) render dashed.
  • Click-through task history — selecting a task opens a timeline of its journey through each node (entry/exit timestamps, per-node duration, per-stage cost, Q&A, verdicts, plan), backed by the task_transitions ledger (schema v3) written by the coordinator's instrumented graph nodes (db/transitions.py, fail-soft).

Endpoints: GET /api/state (live overview + per-node state), GET /api/topology (map nodes/edges/trees), GET /api/task/{thread_id} (one task's history; thread_id is validated ^[A-Za-z0-9_-]{1,64}$). The legacy stdlib HTML page was retired; status_page.py remains as the /api/state data layer.

Build (on the Mac — the box Node is too old for Vite 5+):

cd agent-team/web
npm ci && npm run build          # -> web/dist (gitignored), rsynced to the VM
npm test                         # Vitest + React Testing Library

In dev, npm run dev proxies /api to a locally running dashboard (AGENT_TEAM_DASH, default http://127.0.0.1:8770).

WS0–WS5 rollout glossary

The "WS-rollout" (workstreams 0–5) layered HTTP/integration surfaces onto the P1–P4 pipeline. What is live vs inert after the rollout:

WS What it adds Live?
WS1 invoker_multi.py (in-process GPT-4.1 / DeepSeek / Gemini via the orchestrator's models.py) + api.py (FastAPI HTTP API, bearer auth via AGENT_TEAM_API_TOKEN, binds 127.0.0.1:8765, /docs+/openapi disabled, concurrency-capped) bind_multi_invoker() wired in run-team.py _cmd_serve (LIVE); the HTTP API is a separate opt-in process (api.serve()), NOT started by the coordinator
WS5 nodes/handbook.py load_handbook_conventions (reads SEA_HAVEN_HANDBOOK_DIR or ~/.sea-haven/engineering-handbook, fail-safe → ""); retriever.py save_memory writes to a _box-drafts/ review queue LIVE — the planner prompt receives the handbook via the context_provider seam in run-team.py _build_coordinator
WS2/WS0/WS4 Slack /new-task slash command (AUTHZ-01 owner-allowlist gated) → Coordinator.set_new_task_callback; the sea-haven-claude-plugin/ (CLAUDE.md, settings, /delegate UserPromptSubmit hook) LIVE (/new-task wired in serve); the plugin/HTTP-API path is opt-in
WS3 / P3 nodes/dispatch_invoker.py (DISPATCH LangGraph node) + the BUILD → DISPATCH → VERIFY reorder, the tick()-driven CI-watcher, per-task run_id plumbing, and the bound serve default Built on feat/agent-team-p3-box-integration, gated behind the C1 re-review before it ships to the box. The agent-apply GitHub Environment human-approval gate is KEPT; dispatch is operator-initiated (the box holds no standing write token — the branch push + gh workflow run use operator-host credentials). The bound P3 wiring is the new fail-safe serve default: a missing AGENT_TEAM_REPO_OWNER/_NAME or CI-read token degrades to the INERT P3 path (task parks + a #agent-team notice), never a serve-start crash

The HTTP API endpoints: POST /tasks (start a task), GET /tasks/{thread_id} (status), POST /orchestrator/invoke (one-shot model invoke). See DEPLOY-R720.md for the WS-rollout deploy (scripts/deploy-r720-ws-rollout.sh).

Running the tests

cd agent-team
python3 -m pytest -q          # conftest puts the package on sys.path; no install needed

Deploy

Deploy-gated. See DEPLOY-R720.md for the provisioning runbook (VM snapshot first, rsync, venv deps, ~/secrev.env tokens, init-db, systemd, and the live P1 exit-criteria demo). Secrets are never committed.