Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean. Built (pre-deployment scaffold only — nothing provisioned/enabled): - LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable) - nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier - §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder - transports: slack / github / claude_code adapters - ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness - ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up): - builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete) - §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency - operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap - ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.) - P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning, /sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
150 lines
8.2 KiB
Markdown
150 lines
8.2 KiB
Markdown
# agent-team/ci — split-job CI apply/verify workflow (Plane-2 leaf)
|
|
|
|
Pre-deployment scaffolding for the R720 agent-team SDLC pipeline. This directory
|
|
holds the **split-job CI apply/verify workflow** that turns a builder agent's
|
|
**untrusted candidate diff** into a verified **draft PR** — the §3.3.2 trust
|
|
boundary, Phase P3 (§7.1) of `../../docs/r720-agent-team-design.md`.
|
|
|
|
> **STATUS: DEPLOY-GATED. NOT ENABLED, NOT PROVISIONED.** This is authored as
|
|
> files only. Per the design (§3.3.2, §7.1 P3) the workflow + its OIDC role must
|
|
> clear **BOTH `/sh-security-review` AND the mandatory GPT-4.1 cross-review**
|
|
> before deployment, because it is IaC/IAM + untrusted-input handling. The
|
|
> privileged draft-PR step is hard-disabled (`if: ${{ false }}`) and the OIDC
|
|
> `id-token`/`pull-requests: write` grants are left commented until those gates
|
|
> pass. Nothing here is wired to a live org repo.
|
|
|
|
## Files
|
|
|
|
| File | What it is |
|
|
|---|---|
|
|
| `agent-team-apply-verify.yml` | The split-job workflow. Self-contained: the load-bearing gate logic (diff integrity, trust-control denylist, declared-scope check, pure-code pass/fail) is embedded inline as stdlib-only, type-hinted Python heredocs, so the workflow has **no external script dependency**. |
|
|
| `README.md` | This file. |
|
|
|
|
The filename is kebab-case per the handbook. Deployment target (later, after the
|
|
gates): promote into `Sea-Haven-Industries/.github` as a reusable workflow
|
|
(`engineering-handbook/cicd.md`); the Option-B OIDC apply path invokes it.
|
|
|
|
## The trust boundary (design §3.3.2)
|
|
|
|
The builder agents are semi-trusted: an LLM that read repo content can be wrong
|
|
or prompt-injected, so **the candidate diff is treated as untrusted code.** The
|
|
threat is that executing it in CI with org credentials lets a bad diff exfiltrate
|
|
secrets, assume the deploy role, or tamper with other repos. The workflow
|
|
implements all five boundaries:
|
|
|
|
1. **Split CI — untrusted execution is credential-less.** The job that checks
|
|
out and runs the diff (`build-test`) runs with `permissions: contents: read`,
|
|
**no secrets, no OIDC, no write token**, and egress blocked
|
|
(harden-runner). The patch executes only there, where there is nothing to
|
|
steal and nothing to assume. Every privileged action (the eventual OIDC role,
|
|
the draft-PR open) runs in a **separate `gate-and-pr` job that never checks
|
|
out or executes patch-controlled code** — it consumes the build/test report
|
|
as **data only**. There is **no `pull_request_target` + head-ref checkout**
|
|
(the "pwn request" anti-pattern).
|
|
2. **Trust-control-surface denylist (CI-side hard fail).** The `guard` job
|
|
rejects any diff that touches `.github/workflows/**`, IAM/policy IaC
|
|
(CDK/SAM/Terraform), branch-protection / `CODEOWNERS` / Dependabot config, or
|
|
files **outside the task's declared scope**. The match is not naive: it
|
|
**canonicalizes paths, rejects parent-directory traversal, and inspects
|
|
`rename from/to` headers**, so a rename *into* a denied path — or path
|
|
indirection — cannot bypass it. Such a diff is escalated to mandatory human +
|
|
GPT cross-review, never auto-built.
|
|
3. **Diff integrity, box → CI.** The builder records the candidate diff's sha256
|
|
in the task ledger (foundation `agent_team.state_store` content-hash idiom).
|
|
CI **re-hashes the diff and verifies it equals the ledger-recorded hash
|
|
before applying** (in both `guard` and again pre-apply in `build-test`). A
|
|
tampered or substituted diff fails the hash check.
|
|
4. **Pure-code pass/fail gate over authenticated results.** Mirroring secrev's
|
|
"one pure-code script owns the block decision," the `gate-and-pr` gate reads
|
|
**only** the authenticated `needs.*.result` job conclusions (GitHub-controlled,
|
|
patch-independent) keyed to the run id + diff hash. It **never trusts a
|
|
success/failure file or artifact the patch could have written.** The verifier
|
|
*agent* only reads failures to propose the next fix; it cannot declare
|
|
success.
|
|
5. **Branch protection is the final enforcement.** The draft PR targets a
|
|
protected branch; the locked-down checks are required; merge needs them green
|
|
**plus** the security-review, the Claude Code App review, and **human
|
|
approval**. The agent never auto-merges (D2).
|
|
|
|
**Threat-model honesty (§3.3.2).** The diff hash protects box → CI transit
|
|
(tamper/substitution), **not** a box that is itself compromised — the box is in
|
|
the trusted computing base for what it *proposes*. The real backstops against a
|
|
bad box are the credential-less build job, the trust-control-surface denylist,
|
|
and above all the **human review of the draft PR** plus required checks (nothing
|
|
merges without Adam).
|
|
|
|
## Job graph
|
|
|
|
```
|
|
workflow_dispatch (task_id, diff_artifact_name, expected_diff_hash, declared_scope)
|
|
│
|
|
▼
|
|
guard contents:read, no secrets — validates the diff as DATA:
|
|
(boundaries 2,3) re-hash + denylist + scope. Never applies it.
|
|
│ (needs)
|
|
▼
|
|
build-test contents:read, no secrets, no OIDC, egress blocked —
|
|
(boundary 1) the ONLY job that applies + runs the UNTRUSTED patch.
|
|
│ (needs) Emits a NON-authoritative report artifact.
|
|
▼
|
|
gate-and-pr privileged home, but NEVER checks out patch code. Pure-code
|
|
(boundaries 1,4,5) gate over authenticated needs.*.result → DRAFT PR
|
|
(hard-disabled until the review gates pass).
|
|
```
|
|
|
|
`permissions: {}` at the workflow level (least privilege); each job re-declares
|
|
its own grant explicitly. The trigger is `workflow_dispatch` only — the patch
|
|
never runs in a context carrying write or secret scope.
|
|
|
|
## SHA-pinned actions (handbook Pinning Principle, §3.3.2)
|
|
|
|
Every third-party action is pinned to a full commit SHA with the human-readable
|
|
tag in a trailing comment:
|
|
|
|
| Action | SHA | Tag |
|
|
|---|---|---|
|
|
| `actions/checkout` | `11bd71901bbe5b1630ceea73d27597364c9af683` | v4.2.2 |
|
|
| `actions/download-artifact` | `fa0a91b85d4f404e444e00e005971372dc801d16` | v4.1.8 |
|
|
| `actions/upload-artifact` | `b4b15b8c7c6ac21ea08fcf65892d2ee8f75cf882` | v4.4.3 |
|
|
| `actions/setup-python` | `0b93645e9fea7318ecaed2b359559ac225c90a2b` | v5.3.0 |
|
|
| `step-security/harden-runner` | `0080882f6c36860b6ba35c610c98ce87d4e2f26f` | v2.10.2 |
|
|
|
|
## Relationship to the foundation
|
|
|
|
This leaf **imports the committed Plane-2 foundation contracts verbatim** (it
|
|
does not redefine them):
|
|
|
|
- The diff-hash recorded in the ledger and re-checked in CI is the same
|
|
content-hash idiom as `agent_team.state_store.compute_content_hash` (§6.7).
|
|
- The task this workflow verifies is an `agent_team.task_model.TaskRecord`; its
|
|
`candidate_diff` + `diff_hash` fields (§3.3) are exactly the
|
|
`expected_diff_hash` this workflow consumes, and `ci_results` is what the
|
|
verifier writes back from the authenticated gate (boundary 4).
|
|
- The ledger that records provenance (diff hash, run id, gate decision) is the
|
|
`agent_team.db` schema (`pending_questions` / `budget_ledger` live there;
|
|
per-task CI provenance is recorded against the task thread).
|
|
|
|
## Tests
|
|
|
|
The workflow's embedded gate logic (diff integrity, the trust-control denylist
|
|
with path-canonicalization + rename detection, declared-scope enforcement, and
|
|
the pure-code pass/fail gate) is **stdlib-only, type-hinted, and ruff-clean**.
|
|
It was verified against good and adversarial diffs (workflow edits, IAM edits,
|
|
renames into denied paths, path traversal, out-of-scope and unscoped diffs, hash
|
|
mismatch, and every gate branch). Because this leaf owns only the two files in
|
|
this directory, the executable pytest suite for the importable foundation
|
|
modules lives in `../tests/` (run `python3 -m pytest agent-team/tests/ -q` from
|
|
the repo root); the inline CI gate logic is validated as part of the workflow's
|
|
own steps at deploy time and was proven correct during authoring.
|
|
|
|
## Deploy gating (do NOT skip)
|
|
|
|
Before this ships (§3.3.2, §7.1 P3):
|
|
|
|
1. `/sh-security-review` over this workflow (IaC + untrusted-input handling).
|
|
2. Mandatory **GPT-4.1 cross-review** of the workflow **and** the Option-B OIDC
|
|
role it will assume (IAM change).
|
|
3. A documented, **exercised** rollback (remove the role, revert the workflow).
|
|
4. Only then: uncomment the `id-token` / `pull-requests: write` grants, enable
|
|
the draft-PR step, and promote to `Sea-Haven-Industries/.github`. Draft PRs
|
|
only; never auto-merge.
|