Security follow-up from the per-PR review (non-blocking, defense-in-depth): - Replace the save_memory name blocklist with an allowlist regex (^[A-Za-z0-9][A-Za-z0-9._-]*$, max 128) so dot-only/hidden/backslash/NUL/ over-long names are rejected outright, not written as malformed-but-contained files. - Write via os.open(..., O_NOFOLLOW): the open fails (ELOOP) if the final path component is a pre-planted symlink, closing the TOCTOU where a symlink in _box-drafts/ could redirect the write outside the dir. O_CREAT|O_TRUNC keeps overwrite-on-resave for regular files. Tests: adds allowlist-rejection + symlink-refusal cases (21 pass). |
||
|---|---|---|
| .. | ||
| .security-review | ||
| agent_team | ||
| ci | ||
| slack | ||
| systemd | ||
| tests | ||
| .gitignore | ||
| DEPLOY-R720.md | ||
| README.md | ||
| run-team.py | ||
agent-team — R720 Plane-2 SDLC pipeline
The durable, human-gated agentic SDLC pipeline for the R720 (sh-secrev VM),
design: ../docs/r720-agent-team-design.md. A task flows INTAKE → CLARIFY (the
human gate) → PLAN → REVIEW, and (opt-in, deploy-gated) → BUILD → VERIFY → draft
PR. Every stage is durable and resumable (LangGraph + a SQLite checkpointer);
the human gate suspends on interrupt() and resumes on a real answer.
INTAKE → CLARIFY (Claude, human gate) → PLAN (Claude) → REVIEW (GPT-4.1)
▲ │
└── loop-back ───┤
approve/escalate → END
(P3, opt-in + INERT until the CI gate clears):
approve → BUILD (DeepSeek) → VERIFY (ci_gate) → draft PR
Status (2026-06-18). P1 (human gate) + P2 (planner + adversarial review loop) + the live runtime (coordinator, Slack/GitHub/Claude-Code transports, intake) + P3-inert (build/verify subgraph, opt-in) + P4 (more transports + GitHub-issue intake) are built, reviewed, and merged to main (~795 tests). Nothing is provisioned: not rsync'd to the box, no live tokens, no systemd, no live CI. Production default runs P2 (no builders). Deploy-gated / not yet built: the P3 live CI apply/verify + OIDC role (held behind
/sh-security-review+ the mandatory GPT-4.1 cross-review), and all provisioning. See the project memoryproject_r720_agent_teamand §7 of the design for the phased rollout.
Layout
agent-team/
run-team.py # operator CLI: init-db, list/show/answer/expire/
# force-resume (ledger), start (intake), serve (daemon),
# intake-github
agent_team/ # importable package (snake_case)
coordinator.py # the live runtime keystone: invoker→graph→ResumeWorker;
# start_task / submit_answer / drain / tick / recover / serve
graph.py # LangGraph wiring: P1 (intake→clarify→plan) + opt-in P2
# review loop + opt-in P3 build/verify subgraph
invoker.py # §3.1 real Claude path (subscription-OAuth / API / Bedrock)
billing.py # §3.1 claude_invoke billing-mode seam
ci_gate.py # §3.3.2 pure-code authenticated-Checks PASS/FAIL gate
task_model.py / state_store.py
db/{schema.py,schema.sql} # SQLite ledger DDL + BEGIN IMMEDIATE compare-and-set
ledger.py / responder.py / resume_worker.py / deadline_timer.py / recovery.py
operator_cli.py
nodes/ # pipeline stages + their model bindings
clarifier.py + clarifier_llm.py # human gate (Claude)
planner.py # plan (Claude)
review_loop.py + review_loop_llm.py # adversarial review (GPT-4.1 via orchestrator)
builders.py + builders_llm.py # candidate diff (DeepSeek) — INERT, proposes only
verifier.py + verifier_llm.py # ci_gate sole PASS authority; LLM = fix-proposer
build_verify_subgraph.py # P3 BUILD→VERIFY topology (opt-in)
transport/ # one adapter contract + a live impl per channel
base.py # Transport ABC + QuestionSet / NormalizedAnswer
slack_adapter.py + slack_live.py + slack_listener.py # Block Kit + Socket Mode
github_adapter.py + github_live.py + github_intake.py # issue-comment + issue intake
claude_code_adapter.py + claude_code_live.py # file-drop responder
ci/ # §3.3.2 split-job CI apply/verify workflow (DEPLOY-GATED)
systemd/ # agent-team-coordinator.service (not installed)
DEPLOY-R720.md # provisioning runbook (snapshot-first, rsync, tokens, demo)
tests/ # pytest, one module per source module + sim harness
The top directory is kebab-case (agent-team/); the importable package is
snake_case (agent_team/), per the engineering handbook.
Key design points
- Durable human gate (§3.3.1). The
pending_questionsledger is the single source of truth for the question lifecycle. Every race (duplicate answers, transport redelivery, answer-vs-timeout) resolves via one atomic compare-and-set againststatus, inside aBEGIN IMMEDIATEtransaction — first-answer-wins (rowcount == 1), late/duplicate ignored. The LangGraphSqliteSavercheckpointer shares the same DB file. - Fail-safe model seams. Every node treats model output as untrusted and
fails SAFE: garbage never clears the 98% clarifier gate, never auto-approves a
plan, never fabricates a build success, and the verifier's
ci_gateis the sole PASS authority (the LLM is structurally a fix-proposer only). - Inbound auth (§3.3.1). The Slack Socket Mode listener authorizes the
sender against an owner allowlist (
AGENT_TEAM_SLACK_OWNER_IDS, fail-closed) on top of the open-status CAS anti-replay. - Billing seam (§3.1). Claude runs under subscription OAuth on the box;
GPT-4.1 (review) and DeepSeek (builders) route through the local orchestrator
run.py. Switching Claude billing is a config flip.
Running the tests
cd agent-team
python3 -m pytest -q # conftest puts the package on sys.path; no install needed
Deploy
Deploy-gated. See DEPLOY-R720.md for the provisioning runbook (VM snapshot
first, rsync, venv deps, ~/secrev.env tokens, init-db, systemd, and the live
P1 exit-criteria demo). Secrets are never committed.