* feat(agent-team): P3-flip Phase 1 — expand denylist vectors (§4.2) + runner-trust assertion (§4.1) First controls of the P3-live-flip Phase-1 CI hardening (workflow stays INERT; this only tightens the trust boundary). Whole Phase-1 surface is gated by /sh-security-review + GPT-4.1 cross-review before any flip. §4.2 — expand the trust-control denylist with direct code-execution / supply-chain vectors, kept byte-identical across all three copies (ci_gate.DENYLIST_GLOBS + the guard + post-build inline DENY_GLOBS), drift-guarded: .gitmodules, .husky/**, .githooks/**, .gitattributes, .npmrc, and generated/build artifacts (__generated__, *.generated.*, dist/**, build/**, *.min.js). Deliberate: lockfiles are NOT wholesale denied — lockfile-postinstall RCE is already contained by the credential-less egress-blocked build sandbox, and the Tier-3 dep-CVE fixer rewrites lockfiles to produce its draft PRs; a blanket deny would make it un-shippable. Flagged in-code for the security gate. Direct code-execution config (hooks/filters/npmrc/submodules) is the actual §4.2 RCE surface. §4.1 — runner-trust: assert no job (esp. the privileged gate-and-pr) can run on a self-hosted/user-provided runner; all must be GitHub-hosted. 998 tests pass, ruff clean. * feat(agent-team): P3-flip Phase 1 — gate-weakening detector (§4.5) A diff that ADDS a lint/type/coverage/security suppression (noqa, type: ignore, pragma: no cover, nosec, nosemgrep), a test skip/xfail, or a hook bypass (--no-verify) could make CI pass falsely. The pure-code gate now flags these via gate_weakening_violations() and BLOCKs in evaluate_ci_gate as a top-priority trust violation (step 1b, alongside the denylist) — regardless of the authenticated CI conclusion. A build cannot pass itself by disabling its own checks; flagged diffs escalate to a human. Only ADDED lines are inspected (removing a suppression is fine). 1015 tests pass, ruff clean. * feat(agent-team): P3-flip — diff transport (§4.3) + flip privileged apply path live Completes the box->CI diff handoff and flips the apply/verify privileged job live (gated behind the agent-apply environment's required reviewer). Transport (§4.3): the read-only box (D2) emits a diff but holds no write token. - New credential-less `materialize` job decodes the untrusted `diff_b64` dispatch input via env (CWE-94), fail-closed re-hashes it against `expected_diff_hash`, and uploads it as the named artifact so guard/build-test download it same-run. guard now `needs: materialize`. - New `dispatcher.py` (the trusted apply path, operator/Mac-side — never the box): pushes the diff as a head branch then `gh workflow run`s the workflow. Pure input-assembly (sha256 == sha256sum, b64 round-trip, head ref) is unit-tested; git/gh are injected seams. Push-before-dispatch; fail-closed on empty diff/scope, unsafe task_id/owner/repo. Flip: gate-and-pr binds `environment: agent-apply` (required reviewer amoussa1229) + grants exactly `pull-requests: write`; the App-token + draft-PR steps run only on `steps.gate.outputs.gate == 'pass'` (no more if:false); the draft PR opens with an explicit `--head`; task_id/head_branch charset-validated (§4.6). Updated the hardening tests from inert-state to live-state assertions + added transport tests. 1039 tests, ruff clean, workflow YAML valid. NOTE: workflow only runs on manual workflow_dispatch and the privileged job is held at the required-reviewer gate, so nothing privileged runs unapproved. * fix(agent-team): P3-flip — address GPT-4.1 cross-review (size bound, ref-traversal guard) - BLOCK: cap candidate diff at 40 KB in the dispatcher (the diff rides a base64 workflow_dispatch input; GitHub caps inputs at ~64 KB so an oversized diff cannot dispatch at all) + a defense-in-depth decoded-size bound in materialize. - FIX: harden the draft-PR HEAD_BRANCH guard to reject leading/trailing slash, '..' segments, and '//' (CWE-88 git ref-traversal), not just bad charset. - NIT: document the mandatory invariants on gate-and-pr (required-reviewer environment must stay; runs-on must stay GitHub-hosted). - QUESTION (lockfiles): answered in-code — the build-test sandbox is credential-less + egress-blocked, so lockfile-postinstall RCE is contained. Tests added for all guards. 1042 tests, ruff clean, YAML valid. * fix(agent-team): P3-flip — resolve /sh-security-review findings (LOGIC-1/2/3) High-recall fan-out (injection/logic/iac+secrets) + proof-or-kill on the LIVE apply path found 3 real issues the cross-review missed; all fixed: - LOGIC-2 (HIGH, was a live hole): build-test ran `ruff check . || echo` / `pytest -q || echo`, swallowing failures so the job was always 'success' and the gate would open draft PRs on RED builds. ruff/pytest now run authoritatively under set -e (pytest exit 5 'no tests' is the only non-fatal case); the exit code IS the build-test conclusion the gate keys on. - LOGIC-1 (verified!=shipped): the dispatcher used `git apply` + `git add -A`, staging stray untracked content into the pushed PR head. Now `git apply --index` stages exactly the diff, so the head tree is precisely base+diff — bound to the bytes CI hash-verified. - LOGIC-3 (§4.5 on the live path): gate-weakening was enforced only box-side; added a gate-weakening check to the guard job so the live PR-opening path rejects a diff that adds suppressions/skips, even on a green build. Injection / secrets / least-privilege / flip-correctness / no-untrusted-checkout all came back clean. 1044 tests, ruff clean, YAML valid. |
||
|---|---|---|
| .. | ||
| .security-review | ||
| agent_team | ||
| ci | ||
| slack | ||
| systemd | ||
| tests | ||
| .gitignore | ||
| DEPLOY-R720.md | ||
| README.md | ||
| run-team.py | ||
agent-team — R720 Plane-2 SDLC pipeline
The durable, human-gated agentic SDLC pipeline for the R720 (sh-secrev VM),
design: ../docs/r720-agent-team-design.md. A task flows INTAKE → CLARIFY (the
human gate) → PLAN → REVIEW, and (opt-in, deploy-gated) → BUILD → VERIFY → draft
PR. Every stage is durable and resumable (LangGraph + a SQLite checkpointer);
the human gate suspends on interrupt() and resumes on a real answer.
INTAKE → CLARIFY (Claude, human gate) → PLAN (Claude) → REVIEW (GPT-4.1)
▲ │
└── loop-back ───┤
approve/escalate → END
(P3, opt-in + INERT until the CI gate clears):
approve → BUILD (DeepSeek) → VERIFY (ci_gate) → draft PR
Status (2026-06-18). P1 (human gate) + P2 (planner + adversarial review loop) + the live runtime (coordinator, Slack/GitHub/Claude-Code transports, intake) + P3-inert (build/verify subgraph, opt-in) + P4 (more transports + GitHub-issue intake) are built, reviewed, and merged to main (~795 tests). Nothing is provisioned: not rsync'd to the box, no live tokens, no systemd, no live CI. Production default runs P2 (no builders). Deploy-gated / not yet built: the P3 live CI apply/verify + OIDC role (held behind
/sh-security-review+ the mandatory GPT-4.1 cross-review), and all provisioning. See the project memoryproject_r720_agent_teamand §7 of the design for the phased rollout.
Layout
agent-team/
run-team.py # operator CLI: init-db, list/show/answer/expire/
# force-resume (ledger), start (intake), serve (daemon),
# intake-github
agent_team/ # importable package (snake_case)
coordinator.py # the live runtime keystone: invoker→graph→ResumeWorker;
# start_task / submit_answer / drain / tick / recover / serve
graph.py # LangGraph wiring: P1 (intake→clarify→plan) + opt-in P2
# review loop + opt-in P3 build/verify subgraph
invoker.py # §3.1 real Claude path (subscription-OAuth / API / Bedrock)
billing.py # §3.1 claude_invoke billing-mode seam
ci_gate.py # §3.3.2 pure-code authenticated-Checks PASS/FAIL gate
task_model.py / state_store.py
db/{schema.py,schema.sql} # SQLite ledger DDL + BEGIN IMMEDIATE compare-and-set
ledger.py / responder.py / resume_worker.py / deadline_timer.py / recovery.py
operator_cli.py
nodes/ # pipeline stages + their model bindings
clarifier.py + clarifier_llm.py # human gate (Claude)
planner.py # plan (Claude)
review_loop.py + review_loop_llm.py # adversarial review (GPT-4.1 via orchestrator)
builders.py + builders_llm.py # candidate diff (DeepSeek) — INERT, proposes only
verifier.py + verifier_llm.py # ci_gate sole PASS authority; LLM = fix-proposer
build_verify_subgraph.py # P3 BUILD→VERIFY topology (opt-in)
transport/ # one adapter contract + a live impl per channel
base.py # Transport ABC + QuestionSet / NormalizedAnswer
slack_adapter.py + slack_live.py + slack_listener.py # Block Kit + Socket Mode
github_adapter.py + github_live.py + github_intake.py # issue-comment + issue intake
claude_code_adapter.py + claude_code_live.py # file-drop responder
ci/ # §3.3.2 split-job CI apply/verify workflow (DEPLOY-GATED)
systemd/ # agent-team-coordinator.service (not installed)
DEPLOY-R720.md # provisioning runbook (snapshot-first, rsync, tokens, demo)
tests/ # pytest, one module per source module + sim harness
The top directory is kebab-case (agent-team/); the importable package is
snake_case (agent_team/), per the engineering handbook.
Key design points
- Durable human gate (§3.3.1). The
pending_questionsledger is the single source of truth for the question lifecycle. Every race (duplicate answers, transport redelivery, answer-vs-timeout) resolves via one atomic compare-and-set againststatus, inside aBEGIN IMMEDIATEtransaction — first-answer-wins (rowcount == 1), late/duplicate ignored. The LangGraphSqliteSavercheckpointer shares the same DB file. - Fail-safe model seams. Every node treats model output as untrusted and
fails SAFE: garbage never clears the 98% clarifier gate, never auto-approves a
plan, never fabricates a build success, and the verifier's
ci_gateis the sole PASS authority (the LLM is structurally a fix-proposer only). - Inbound auth (§3.3.1). The Slack Socket Mode listener authorizes the
sender against an owner allowlist (
AGENT_TEAM_SLACK_OWNER_IDS, fail-closed) on top of the open-status CAS anti-replay. - Billing seam (§3.1). Claude runs under subscription OAuth on the box;
GPT-4.1 (review) and DeepSeek (builders) route through the local orchestrator
run.py. Switching Claude billing is a config flip.
Running the tests
cd agent-team
python3 -m pytest -q # conftest puts the package on sys.path; no install needed
Deploy
Deploy-gated. See DEPLOY-R720.md for the provisioning runbook (VM snapshot
first, rsync, venv deps, ~/secrev.env tokens, init-db, systemd, and the live
P1 exit-criteria demo). Secrets are never committed.