Codifies the hardened deploy-before-merge procedure for ongoing agent-team code changes, generalizing the one-off deploy-r720-ws-rollout.sh. Prevents the two self-inflicted live crash-loops: - Whole agent_team/ package rsync (never per-file, which misplaces e.g. nodes/planner.py at the package root -> ImportError/TypeError crash-loop). - Snapshot HARD GATE (Adam's Hyper-V step; Claude ssh reaches only the guest) + ledger backup before any change. - Pre-restart import sanity, then mandatory verify-after (is-active==active, NRestarts didn't climb, ~6 threads, clean journal) with rollback guidance on failure. Idempotent, fails loudly. Optional SYNC_DEPS / SYNC_HANDBOOK / RESTART_STATUS. Drives the new /sh-deploy-r720 skill. |
||
|---|---|---|
| .. | ||
| .security-review | ||
| agent_team | ||
| ci | ||
| scripts | ||
| slack | ||
| systemd | ||
| tests | ||
| .gitignore | ||
| DEPLOY-R720.md | ||
| README.md | ||
| run-team.py | ||
agent-team — R720 Plane-2 SDLC pipeline
The durable, human-gated agentic SDLC pipeline for the R720 (sh-secrev VM),
design: ../docs/r720-agent-team-design.md. A task flows INTAKE → CLARIFY (the
human gate) → PLAN → REVIEW, and (opt-in, deploy-gated) → BUILD → VERIFY → draft
PR. Every stage is durable and resumable (LangGraph + a SQLite checkpointer);
the human gate suspends on interrupt() and resumes on a real answer.
INTAKE → CLARIFY (Claude, human gate) → PLAN (Claude) → REVIEW (GPT-4.1)
▲ │
└── loop-back ───┤
approve/escalate → END
(P3, opt-in + INERT until the CI gate clears):
approve → BUILD (DeepSeek) → VERIFY (ci_gate) → draft PR
Status (2026-06-18). P1 (human gate) + P2 (planner + adversarial review loop) + the live runtime (coordinator, Slack/GitHub/Claude-Code transports, intake) + P3-inert (build/verify subgraph, opt-in) + P4 (more transports + GitHub-issue intake) are built, reviewed, and merged to main (~795 tests). Nothing is provisioned: not rsync'd to the box, no live tokens, no systemd, no live CI. Production default runs P2 (no builders). Deploy-gated / not yet built: the P3 live CI apply/verify + OIDC role (held behind
/sh-security-review+ the mandatory GPT-4.1 cross-review), and all provisioning. See the project memoryproject_r720_agent_teamand §7 of the design for the phased rollout.
Layout
agent-team/
run-team.py # operator CLI: init-db, list/show/answer/expire/
# force-resume (ledger), start (intake), serve (daemon),
# intake-github
agent_team/ # importable package (snake_case)
coordinator.py # the live runtime keystone: invoker→graph→ResumeWorker;
# start_task / submit_answer / drain / tick / recover / serve
graph.py # LangGraph wiring: P1 (intake→clarify→plan) + opt-in P2
# review loop + opt-in P3 build/verify subgraph
invoker.py # §3.1 real Claude path (subscription-OAuth / API / Bedrock)
invoker_multi.py # WS1 in-process non-Claude invokers (GPT-4.1 / DeepSeek /
# Gemini via the orchestrator's models.py); bind_multi_invoker()
api.py # WS1 FastAPI HTTP API (bearer auth, 127.0.0.1:8765) — SEPARATE
# opt-in process (api.serve()), NOT started by the coordinator
billing.py # §3.1 claude_invoke billing-mode seam
ci_gate.py # §3.3.2 pure-code authenticated-Checks PASS/FAIL gate
task_model.py / state_store.py
db/{schema.py,schema.sql} # SQLite ledger DDL + BEGIN IMMEDIATE compare-and-set
ledger.py / responder.py / resume_worker.py / deadline_timer.py / recovery.py
operator_cli.py
nodes/ # pipeline stages + their model bindings
clarifier.py + clarifier_llm.py # human gate (Claude)
planner.py # plan (Claude)
review_loop.py + review_loop_llm.py # adversarial review (GPT-4.1 via orchestrator)
builders.py + builders_llm.py # candidate diff (DeepSeek) — INERT, proposes only
verifier.py + verifier_llm.py # ci_gate sole PASS authority; LLM = fix-proposer
build_verify_subgraph.py # P3 BUILD→VERIFY topology (opt-in)
handbook.py # WS5 load_handbook_conventions (handbook seam,
# fail-safe → "" if dir missing); planner context
dispatch_invoker.py # WS3 auto-dispatch node — INERT (NOT wired live)
transport/ # one adapter contract + a live impl per channel
base.py # Transport ABC + QuestionSet / NormalizedAnswer
slack_adapter.py + slack_live.py + slack_listener.py # Block Kit + Socket Mode + /new-task
github_adapter.py + github_live.py + github_intake.py # issue-comment + issue intake
claude_code_adapter.py + claude_code_live.py # file-drop responder
scripts/ # deploy-r720-ws-rollout.sh — attended WS0–WS5 UPDATE of the box
ci/ # §3.3.2 split-job CI apply/verify workflow (DEPLOY-GATED)
systemd/ # agent-team-coordinator.service (not installed)
DEPLOY-R720.md # provisioning runbook (snapshot-first, rsync, tokens, demo)
tests/ # pytest, one module per source module + sim harness
The Claude Code plugin lives in a sibling top-level dir, ../sea-haven-claude-plugin/
(CLAUDE.md, settings.template.json, hooks/user_prompt_submit.py): a
UserPromptSubmit hook that forwards /delegate <task> prompts from Claude Code
to the HTTP API's POST /tasks (env AGENT_TEAM_API_URL / AGENT_TEAM_API_TOKEN).
The top directory is kebab-case (agent-team/); the importable package is
snake_case (agent_team/), per the engineering handbook.
Key design points
- Durable human gate (§3.3.1). The
pending_questionsledger is the single source of truth for the question lifecycle. Every race (duplicate answers, transport redelivery, answer-vs-timeout) resolves via one atomic compare-and-set againststatus, inside aBEGIN IMMEDIATEtransaction — first-answer-wins (rowcount == 1), late/duplicate ignored. The LangGraphSqliteSavercheckpointer shares the same DB file. - Fail-safe model seams. Every node treats model output as untrusted and
fails SAFE: garbage never clears the 98% clarifier gate, never auto-approves a
plan, never fabricates a build success, and the verifier's
ci_gateis the sole PASS authority (the LLM is structurally a fix-proposer only). - Inbound auth (§3.3.1). The Slack Socket Mode listener authorizes the
sender against an owner allowlist (
AGENT_TEAM_SLACK_OWNER_IDS, fail-closed) on top of the open-status CAS anti-replay. - Billing seam (§3.1). Claude runs under subscription OAuth on the box;
GPT-4.1 (review) and DeepSeek (builders) route through the local orchestrator
run.py. Switching Claude billing is a config flip.
WS0–WS5 rollout glossary
The "WS-rollout" (workstreams 0–5) layered HTTP/integration surfaces onto the P1–P4 pipeline. What is live vs inert after the rollout:
| WS | What it adds | Live? |
|---|---|---|
| WS1 | invoker_multi.py (in-process GPT-4.1 / DeepSeek / Gemini via the orchestrator's models.py) + api.py (FastAPI HTTP API, bearer auth via AGENT_TEAM_API_TOKEN, binds 127.0.0.1:8765, /docs+/openapi disabled, concurrency-capped) |
bind_multi_invoker() wired in run-team.py _cmd_serve (LIVE); the HTTP API is a separate opt-in process (api.serve()), NOT started by the coordinator |
| WS5 | nodes/handbook.py load_handbook_conventions (reads SEA_HAVEN_HANDBOOK_DIR or ~/.sea-haven/engineering-handbook, fail-safe → ""); retriever.py save_memory writes to a _box-drafts/ review queue |
LIVE — the planner prompt receives the handbook via the context_provider seam in run-team.py _build_coordinator |
| WS2/WS0/WS4 | Slack /new-task slash command (AUTHZ-01 owner-allowlist gated) → Coordinator.set_new_task_callback; the sea-haven-claude-plugin/ (CLAUDE.md, settings, /delegate UserPromptSubmit hook) |
LIVE (/new-task wired in serve); the plugin/HTTP-API path is opt-in |
| WS3 | nodes/dispatch_invoker.py (auto-dispatch LangGraph node) + graph/coordinator wiring |
INERT — NOT wired live. The agent-apply GitHub Environment human-approval gate is KEPT; the P3 dispatch/build-verify path stays inert pending per-task run_id plumbing + a CI-boundary security re-review |
The HTTP API endpoints: POST /tasks (start a task), GET /tasks/{thread_id}
(status), POST /orchestrator/invoke (one-shot model invoke). See
DEPLOY-R720.md for the WS-rollout deploy (scripts/deploy-r720-ws-rollout.sh).
Running the tests
cd agent-team
python3 -m pytest -q # conftest puts the package on sys.path; no install needed
Deploy
Deploy-gated. See DEPLOY-R720.md for the provisioning runbook (VM snapshot
first, rsync, venv deps, ~/secrev.env tokens, init-db, systemd, and the live
P1 exit-criteria demo). Secrets are never committed.