* docs(agent-team): fold round-2 GPT-4.1 plan-review findings into P3-live-flip plan Round-2 cross-review (REQUEST CHANGES) folded: - B1 deploy-before-merge made a concrete CI-enforced gate (required status check fed by a box-exercised dispatch dry-run), not prose. - B2 added rollback for a prematurely-flipped privileged job that actually RAN (token rotate, revert opened PR/branch, audit the window) — distinct from an accidental merge. - B3 rollback is a tested, re-runnable script over ALL privileged surfaces (workflow, environment, App perms, branch protection), not a one-time manual run. - B4 Phase 3 explicitly gated on Phase 2 being fully provisioned + verified. - B5 /sh-security-review + GPT-4.1 cross-review RE-RUN on the actual enabled workflow before the flip, not only the inert version. - B6 docs/memory updated incrementally at each privileged step; Phase 6 is the final reconciliation pass. Plus FIX/NIT/QUESTION: gate-weakening pattern review, check-name discovery test, memory update on denylist change, snapshot retention in Phase 0, draft-PR notification (Phase 4) + conservative-rollout controls (Phase 5) made concrete. GH_TOKEN->GITHUB_TOKEN gate marked done (box alias added). * docs(agent-team): close B1 — deploy-before-merge is a committed hard gate, not a manual fallback Round-3 re-review resolved B2-B6 but flagged B1 still-open: the prior wording left a 'enforced by hand until the check exists' escape hatch. Reframe B1 as a REQUIRED Phase-1 build deliverable that blocks the flip (no manual fallback), with the required-status-check + admin-bypass-disabled branch protection, and a dry-run faithfulness note (same workflow file/jobs as live, only the privileged if: differs). * docs(agent-team): close B1 ordering — admin-bypass disabled before any flip via the B4 precondition Round-4 confirmatory re-review noted (c) admin-bypass-disable sits in Phase 2 while the gate is framed Phase-1. Clarify there is no ordering window: the flip (Phase 3) is gated on Phase 2 completion (B4), so branch protection incl. admin-bypass-disable is necessarily in place before any flip. The check's implementation being a Phase-1 build task (not yet physically built) is expected for a pre-build plan; it is non-optional and flip-blocking, enforced at the Phase-1 hard stop. Stopping the plan-review cycle here per the project's '3 cycles, residual is build-time' rule.
21 KiB
P3-LIVE-FLIP PLAN — agent-team build → verify → draft-PR
Formal phased plan to take the agent-team Plane-2 pipeline from clarify+plan only
to producing reviewable draft PRs, while keeping the always-on R720 box
read-only and the apply path zero-AWS. Status as of 2026-06-22: NOT STARTED
(P3 is built but inert). This plan is the input to /sh-plan-review before any build.
Prerequisite reading:
docs/r720-agent-team-design.md§3.3.2 (CI-as-verifier trust boundary, "B4"),PROVISIONING-RUNBOOK.md(the P3-live-flip section), and memoryproject_r720_agent_team(locked decisions).
1. Objective & current state
Today (inert): the pipeline runs INTAKE → CLARIFIER → PLANNER → REVIEW, but
serve passes build_verify_wiring=None, the Tier-3 fixer is --dry-run only, and
agent-team/ci/agent-team-apply-verify.yml has its privileged steps disabled with
if: ${{ false }} and pull-requests:write / environment: commented out. So it
clarifies + plans but writes no code and opens no PR.
After P3: the pipeline can emit a diff, have org CI build/test/security-review it in an untrusted sandbox, a pure-code gate confirm green from authenticated Checks-API results, and a scoped GitHub App open a draft PR for human review. The box never gains a write token.
2. Locked decisions (carried in — do not relitigate here)
- D-OIDC: apply path uses a GitHub App
pull-requests:writetoken, ZERO AWS, no OIDC. - D2: box stays read-only / no standing write token; org CI does the applying.
- B4: the LLM-proposed diff is untrusted code; the CI trust boundary (below) is mandatory.
- Output is draft PRs only — nothing auto-merges; human approval is the merge gate.
- (Separate, not part of this plan:
aws-postureresident access via step-ca + IAM Roles Anywhere — that IAM is already cross-review-approved and is its own sub-task.)
3. Hard gates (must clear before the flip — these block everything)
| Gate | Why | Owner |
|---|---|---|
/sh-plan-review on THIS plan |
adversarial plan audit before build | me → GPT-4.1 |
/sh-security-review on the apply/verify CI surface |
auth + untrusted-input + CI trust boundary | me |
| GPT-4.1 cross-family review on the apply/verify CI + any permission change | mandatory for the trust-boundary / permissions surface | orchestrator |
GH_TOKEN→GITHUB_TOKEN resolved |
transport reads GITHUB_TOKEN; box now has a GITHUB_TOKEN alias of GH_TOKEN in ~/secrev.env |
me |
The flip does NOT proceed until /sh-security-review AND the GPT-4.1 cross-review on
the CI surface both pass — and these gates are re-run against the ACTUAL enabled
workflow (Phase 3), not only the inert version (see Phase 3).
Plan-review disposition (GPT-4.1 cross-family, 2026-06-22 — REQUEST CHANGES). Findings folded into §4 and Phases 1/1b: expanded denylist vectors, runner-trust, concrete diff-transport + threat model, gate-weakening detection, PR-metadata sanitization, ledger anti-tamper, recovery for a merged-privileged change, required-check-name discovery, deploy-before-merge enforcement, no-write-token audit, draft-PR rate monitoring + stale-PR cleanup. Several items the reviewer marked BLOCK are already implemented in PR #17's CI (canonicalization, egress, SHA-pin, empty-hash fail-closed, Checks-API) — Phase 1 verifies them rather than rebuilding. Open QUESTIONs to answer when building: how human reviewers are notified of new draft PRs, and how "every deployable repo has CI" is enforced for targets.
Plan-review disposition — ROUND 2 (GPT-4.1 cross-family, 2026-06-22 eve — REQUEST CHANGES). Re-review after the durable-intake work merged (#32) and the coordinator went live. Six BLOCK items folded into the phases below: (B1) deploy-before-merge made a CONCRETE CI-enforced gate, not prose (Phase 1 + Phase 3); (B2) rollback added for a prematurely-flipped privileged job that RUNS — not only an accidental merge (Phase 1b); (B3) rollback for ALL privileged surfaces is a tested script, exercised, not one-time manual (Phase 1b); (B4) Phase 3 is explicitly gated on Phase 2 being fully provisioned + verified (Phase 3 precondition); (B5)
/sh-security-review+ GPT-4.1 cross-review are RE-RUN on the actual enabled workflow before the flip (Phase 3); (B6) docs/memory updated INCREMENTALLY at each privileged step, not deferred to Phase 6. FIX/NIT also folded: periodic review of gate-weakening patterns, a check-name-discovery test across targets, a memory update when denylist vectors change, snapshot retention in Phase 0, and the draft-PR notification QUESTION assigned in Phase 4. (Findings-to-verify, not gospel: the reviewer under-read §3 — cross-review IS already a hard gate; B5 is the genuine delta = re-run on the enabled file.)
4. The CI trust boundary (design B4 — what the workflow must enforce)
Already implemented in the merged P3-live CI hardening (PR #17) — Phase 1 VERIFIES, does not rebuild: denylist path canonicalization + symlink/rename/traversal resistance, egress restriction on the untrusted job, SHA-pinned actions, empty/missing-hash fail-closed, and authenticated Checks-API result consumption. The GPT-4.1 plan-review (2026-06-22) flagged these as "missing" because the plan under-restated them; confirm each against the actual
agent-team/ci/agent-team-apply-verify.yml+ci_fetcher.py/ci_gate.pyrather than re-authoring.
- Split CI. An untrusted build/test job:
contents: readonly, no secrets / no OIDC / no write token, egress-restricted (verify/audit the restriction, don't just assert it). A separate privileged job that never checks out the patch code (nopull_request_target+ head checkout) opens the draft PR. Privileged jobs MUST run only on GitHub-hosted runners — assert no self-hosted/user-provided runner can pick them up. - Denylist (reject/escalate, never auto-build). Beyond
.github/workflows/**, IAM/permission IaC, branch-protection /CODEOWNERS/ Dependabot config, and out-of-scope files, the denylist MUST also cover these RCE/priv-esc vectors:.gitmodules/ submodule changes, git hooks (.git/hooks,core.hooksPath,.husky/**),.gitattributes(filter/clean-smudge process), lockfiles + package-manager postinstall/preinstall hooks, and generated/build-artifact files (codegen output is not reviewable as source). Path matching is canonicalized (PR #17) so symlinks/renames/traversal can't slip a denied path past. - Diff-transport integrity (concretely specified, threat-modeled). The box records the
diff content hash in its ledger; CI verifies the hash before apply. The diff
reaches CI as a content-addressed signed artifact (HMAC/keyed digest the box and the
privileged job share via an Actions secret) OR a short-lived, branch-only token
scoped to a single ref — the box never holds a write token. Threat model the path:
tamper-in-transit (defeated by hash+signature verify), replay of an old diff (defeated by
per-task nonce + the
status='open'/one-shot ledger state), and a hostile artifact name. - Pure-code green gate. Pass/fail is owned by a pure-code gate reading authenticated Checks-API results (run id + diff hash). The LLM verifier may propose fixes but can never declare a build green.
- Gate-weakening detection. A diff that lowers a gate — adds
# noqa,# type: ignore, testskip/xfail, coverage/lint excludes,--no-verify, or edits the gate config itself — is flagged and escalated (a build can't make itself pass by disabling the checks). - PR-metadata sanitization. The draft-PR title / body / comments are sanitized so a hostile diff or LLM output can't exfiltrate secrets/env or inject content into the PR text.
- Ledger anti-tamper. The diff-hash ledger entry is integrity-protected (the existing atomic-write + integrity-check substrate; verify the hash row can't be silently rewritten between record and apply).
- Merge gate. Draft PR + required checks +
/sh-security-review+ Claude Code App review + human approval.
5. Phases
Phase 0 — Plan review & pre-reqs 🤖/🧑
- Run
/sh-plan-reviewon this doc; fold BLOCK/FIX items in. (Round 1 + Round 2 done; this doc is the result.) - Confirm a clean revert point (git tag main; Hyper-V snapshot of sh-secrev). Snapshot retention: keep the pre-P3 snapshot until P3 has run clean for one full cycle (Phase 4 DoD), then prune — recorded here so it is not an open-ended snapshot.
- Resolve
GH_TOKEN→GITHUB_TOKEN(box~/secrev.envnow has aGITHUB_TOKENalias). - Rollback: none (no state changed).
Phase 1 — Author/verify the split-CI apply/verify workflow 🤖 (review-gated)
- Reconcile
agent-team/ci/agent-team-apply-verify.ymlwith §4. First confirm the PR-#17 controls are present (canonicalized denylist, egress restriction, SHA-pins, empty-hash fail-closed, Checks-API consumption); only then add the new §4 items. - Add denylist vectors (§4.2): submodules/
.gitmodules, git hooks/core.hooksPath/.husky,.gitattributesfilters, lockfile postinstall/preinstall, generated/build artifacts. Add a test suite proving canonicalization resists symlink/rename/traversal. - Runner-trust assertion (§4.1): test that privileged jobs cannot run on a self-hosted/user-provided runner.
- Concretize + threat-model the diff transport (§4.3): pick content-addressed signed artifact (shared HMAC secret) or short-lived branch-only token; add per-task nonce anti-replay; document and test it.
- Gate-weakening detector (§4.5): CI step that fails on a diff adding
noqa/type: ignore/skip/xfail/excludes/--no-verifyor editing the gate config. The detector's pattern list is reviewed/expanded each time a new bypass vector is found (FIX) — record the list in code with a comment pointing here, and update it + memory when a vector is added (same discipline as the denylist below). - PR-metadata sanitization (§4.6) and ledger anti-tamper (§4.7) implemented + tested.
agent_team/ci_fetcher.py(read-only Checks-API fetcher; fails closed) +ci_gate.py(pure-code green). Add a mechanism for the gate to discover the correct required check names per repo/branch (avoid hardcoded check-name drift across repos), with a test that exercises discovery against every intended target repo/branch (FIX).- Memory/doc update when denylist vectors change (FIX): adding a denylist vector (here
or later) updates
project_r720_agent_teammemory + the Confluence host page in the SAME change — the denied set is operational/security-critical, not tribal knowledge. - Deploy-before-merge enforcement — CONCRETE, CI-enforced (B1). REQUIRED Phase-1
deliverable; the flip does not proceed without it (no manual fallback). "Deploy" of this
change = the privileged path is proven on the live box before the workflow PR merges.
Build, in this phase:
- (a) the box exercises a dispatch dry-run of the apply/verify workflow (privileged steps
still
if:${{ false }}) that emits a signed "exercised-on-box" artifact/status. Dry-run faithfulness (NIT): it runs the SAME workflow file and the SAME untrusted build/verify + pure-code-gate jobs as the live path — ONLY the privilegedif:differs — so the dry-run is a faithful proof of the path, not a separate mock. - (b) a required status check on the workflow-file PR consumes that proof, so the PR is un-mergeable until the box has run it. - (c) branch-protection set so the required checks cannot be bypassed by admins ("do not allow bypassing the above settings" / include-administrators) — no--adminmerge of the privileged flip. Applied in Phase 2; no ordering window because the flip (Phase 3) is gated on Phase 2 completion (the B4 precondition), so admin-bypass is already disabled before any flip is possible. The Phase-1 required-status-check (a/b) and the Phase-2 branch-protection (c) together are the gate; the flip cannot happen until BOTH are in place. This is the gate; until it is built+green, the flip is blocked. (Owner: 🤖 build; verified in the Phase-1/sh-security-review+ cross-review hard stop. The check's implementation is itself a Phase-1 build task — that it is not yet physically built is expected for a pre-build plan; what matters is it is non-optional and flip-blocking, enforced at the Phase-1 hard stop.) - Concrete "no write token on the box" audit (B-QUESTION → a real check): a
test/script asserting the box env + coordinator config hold no
pull-requests:write/ contents-write token (grep the live env names + assert the App token is only an Actions secret), runnable on the box and in CI. Not a prose claim. /sh-security-review+ GPT-4.1 cross-review on this surface. Hard stop until both pass. (NOTE: these are RE-RUN on the enabled workflow in Phase 3 — see B5 there.)- Rollback: workflow file stays inert (
if: ${{ false }}not yet flipped); delete the file.
Phase 1b — Recovery for an accidentally-merged/applied privileged change 🤖/🧑
- Rollback as a TESTED SCRIPT covering ALL privileged surfaces (B3) — not a one-time
manual exercise. One scripted, re-runnable rollback that, per surface, restores from a
recorded baseline: (1) the apply/verify workflow (revert the SHA → inert), (2) the
agent-applyenvironment (required-reviewer + protection rules), (3) the GitHub App permissions/installation (rotate token, reduce/uninstall), (4) branch protection. The script asserts the post-restore state matches the baseline. Exercised in Phase 3 rollback AND re-runnable on demand. - Premature-flip-that-RAN rollback (B2). Distinct from an accidental merge: cover the
case where
if:${{ false }}is flipped early (or the environment gate is misconfigured) and the privileged job actually runs — incident steps: rotate the GitHub App token immediately, close/revert any draft PR (or branch) it opened, confirm via the Checks/PR audit trail exactly what ran in the window, restore the environment + branch protection from baseline, and file the incident. This is the "it executed" path, not just "it merged." - Add light monitoring on draft-PR creation rate (runaway-volume alarm) and an orphaned/stale draft-PR cleanup step.
Phase 2 — Provision the GitHub App + environment 🧑 OPERATOR (browser/admin)
- Create a dedicated GitHub App with
pull-requests:write(+ minimal contents to open a branch/PR); install on the org. Token lives in CI, never on the box. - Create the
agent-applyGitHub Actions Environment with a required reviewer (Adam) + branch-protection so the privileged job cannot run unreviewed. - Store the App credentials as repo/org Actions secrets (not on the box).
- Rollback: uninstall the App; delete the environment + secrets.
Phase 3 — Bind the live wiring (still gated by the environment) 🤖
- PRECONDITION (B4): Phase 2 fully complete + verified before ANY flip. Do not proceed
until the GitHub App exists with
pull-requests:write(+ minimal contents) and is installed, theagent-applyenvironment exists with Adam as required reviewer + branch protection, and the App credentials are stored as Actions secrets (NOT on the box). Verify each before the next step; the flip is blocked otherwise. - In the workflow: uncomment
permissions: pull-requests: writeandenvironment: agent-apply; flip the twoif: ${{ false }}→ enabled. - RE-RUN BOTH GATES ON THE ENABLED WORKFLOW (B5).
/sh-security-review+ the GPT-4.1 cross-family review are run again against the actual enabledagent-team-apply-verify.yml(permissions live,if:true) and the boundgated_build_verify_wiring— NOT only the inert Phase-1 version. Hard stop until both pass on the enabled file. (Permissions changed → the mandatory cross-family review is independently required here against the real diff.) - Bind
agent_team.coordinator.gated_build_verify_wiring(...)(real diff builder + read-only CI-result fetcher) so a leaf calls it only after the gate clears. - Set the box-side apply env vars the live path reads (read-only CI-result token + dispatch target). Confirm no write token lands on the box (run the Phase-1 no-write-token audit).
- Incremental docs (B6): update
OPERATOR-RUNBOOK.md+ memory NOW that the flip is live (do not wait for Phase 6) — what the apply path can/can't do, the denylist, the rollback. - Rollback: run the Phase-1b tested rollback script (re-set
if: ${{ false }}, re-commentenvironment:, setbuild_verify_wiring=None, restart the coordinator). Exercise it once here to prove it works before relying on it.
Phase 4 — Smoke test to a first draft PR 🧑/🤖
- Drive one trivial, in-scope task end-to-end → confirm: untrusted job builds/tests with no secrets, denylist rejects an out-of-scope diff, pure-code gate gates on real Checks results, privileged job opens a draft PR with required checks attached, nothing merged.
- Flip the Tier-3 fixer off
--dry-runonly after the smoke test passes; verify a dependency-CVE bump produces a draft PR. - Draft-PR reviewer notification (QUESTION resolved/assigned): decide + wire how a new
draft PR pings the human reviewer — default = the existing Slack ALARM path posts a
#agent-teamnotice with the PR link (Adam owns this decision; recorded so it is not left open). - Rollback: close the draft PR; Phase-3 rollback.
Phase 5 — Enable the cross-plane loop 🤖
- Allow confirmed Plane-1 checker findings (
intake-checker) to flow into pipeline tasks that end in draft-PR fixes (start conservative: highest-severity, one at a time). - Conservative-rollout controls (QUESTION resolved): enforce "one at a time, highest
severity" concretely — a cap on concurrently-open cross-plane draft PRs (start = 1) read by
the intake-checker leaf, the severity floor in config, and a
#agent-teamlog line per promotion so the rate is observable. Not just prose. - Rollback: revert intake-checker wiring to report-only.
Phase 6 — Final docs & memory consolidation 🤖
B6: docs/memory are updated INCREMENTALLY at each privileged step above (Phase 3 runbook+memory on flip; denylist/gate-weakening changes update memory in the same change). Phase 6 is the FINAL reconciliation pass, not the first time docs are touched.
- Reconcile
OPERATOR-RUNBOOK.md, the Confluence host page (20054017) + runbook (22609921), andproject_r720_agent_teammemory: P3 is LIVE, what the apply path can/can't do, the denylist, the merge gate, the rollback — confirming the incremental updates are complete + consistent.
6. What changes (and what does NOT)
Changes:
- The pipeline produces draft PRs for well-scoped work (dep bumps, small fixes, docs, checker-surfaced issues). Your role shifts from author → reviewer/approver for that class.
- Tier-3 fixer goes live (dep-CVE bumps → draft PRs).
- Plane-1 findings can become draft-PR fixes (cross-plane loop).
Does NOT change:
- Box stays read-only; CI applies, not the box. No AWS / no OIDC in the apply path.
- Denylisted, high-blast-radius changes (workflows, IAM, branch protection, CODEOWNERS, Dependabot) are never auto-built — always escalated.
- Human gates stay: the clarifier question-set AND the PR approval are both yours.
- Drafts, not merges — you remain the merge authority.
7. Risks & mitigations
| Risk | Mitigation |
|---|---|
| Untrusted diff exfiltrates secrets / escalates in CI | split CI: build job has no secrets/OIDC/write, egress-restricted; privileged job never checks out patch code |
| Malicious diff edits CI/IAM/branch-protection to self-escalate | denylist → reject/escalate, never auto-build |
| LLM "declares" a broken build green | pure-code gate reads authenticated Checks-API only; LLM can't set status |
| Diff tampered between box and CI | diff-hash recorded in ledger, verified before apply |
| Standing write capability on the always-on box | there is none — App token lives in CI; box holds only a read-only CI-result token |
| Runaway PR volume | start with Tier-3 only + one finding at a time; required-reviewer environment gates each |
8. Definition of done
/sh-plan-review,/sh-security-review, and GPT-4.1 cross-review on the CI surface all passed.- Phase-4 smoke test produced a draft PR; nothing auto-merged; rollback exercised once.
- No write token on the box (verified); apply path is zero-AWS.
- Docs + Confluence + memory updated.
- Snapshot retained until P3 runs clean for one cycle, then pruned.