This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/agent-team/ci/README.md
Adam Moussa 15a416d31a Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.

Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled

KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven

Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 15:16:12 -04:00

150 lines
8.2 KiB
Markdown

# agent-team/ci — split-job CI apply/verify workflow (Plane-2 leaf)
Pre-deployment scaffolding for the R720 agent-team SDLC pipeline. This directory
holds the **split-job CI apply/verify workflow** that turns a builder agent's
**untrusted candidate diff** into a verified **draft PR** — the §3.3.2 trust
boundary, Phase P3 (§7.1) of `../../docs/r720-agent-team-design.md`.
> **STATUS: DEPLOY-GATED. NOT ENABLED, NOT PROVISIONED.** This is authored as
> files only. Per the design (§3.3.2, §7.1 P3) the workflow + its OIDC role must
> clear **BOTH `/sh-security-review` AND the mandatory GPT-4.1 cross-review**
> before deployment, because it is IaC/IAM + untrusted-input handling. The
> privileged draft-PR step is hard-disabled (`if: ${{ false }}`) and the OIDC
> `id-token`/`pull-requests: write` grants are left commented until those gates
> pass. Nothing here is wired to a live org repo.
## Files
| File | What it is |
|---|---|
| `agent-team-apply-verify.yml` | The split-job workflow. Self-contained: the load-bearing gate logic (diff integrity, trust-control denylist, declared-scope check, pure-code pass/fail) is embedded inline as stdlib-only, type-hinted Python heredocs, so the workflow has **no external script dependency**. |
| `README.md` | This file. |
The filename is kebab-case per the handbook. Deployment target (later, after the
gates): promote into `Sea-Haven-Industries/.github` as a reusable workflow
(`engineering-handbook/cicd.md`); the Option-B OIDC apply path invokes it.
## The trust boundary (design §3.3.2)
The builder agents are semi-trusted: an LLM that read repo content can be wrong
or prompt-injected, so **the candidate diff is treated as untrusted code.** The
threat is that executing it in CI with org credentials lets a bad diff exfiltrate
secrets, assume the deploy role, or tamper with other repos. The workflow
implements all five boundaries:
1. **Split CI — untrusted execution is credential-less.** The job that checks
out and runs the diff (`build-test`) runs with `permissions: contents: read`,
**no secrets, no OIDC, no write token**, and egress blocked
(harden-runner). The patch executes only there, where there is nothing to
steal and nothing to assume. Every privileged action (the eventual OIDC role,
the draft-PR open) runs in a **separate `gate-and-pr` job that never checks
out or executes patch-controlled code** — it consumes the build/test report
as **data only**. There is **no `pull_request_target` + head-ref checkout**
(the "pwn request" anti-pattern).
2. **Trust-control-surface denylist (CI-side hard fail).** The `guard` job
rejects any diff that touches `.github/workflows/**`, IAM/policy IaC
(CDK/SAM/Terraform), branch-protection / `CODEOWNERS` / Dependabot config, or
files **outside the task's declared scope**. The match is not naive: it
**canonicalizes paths, rejects parent-directory traversal, and inspects
`rename from/to` headers**, so a rename *into* a denied path — or path
indirection — cannot bypass it. Such a diff is escalated to mandatory human +
GPT cross-review, never auto-built.
3. **Diff integrity, box → CI.** The builder records the candidate diff's sha256
in the task ledger (foundation `agent_team.state_store` content-hash idiom).
CI **re-hashes the diff and verifies it equals the ledger-recorded hash
before applying** (in both `guard` and again pre-apply in `build-test`). A
tampered or substituted diff fails the hash check.
4. **Pure-code pass/fail gate over authenticated results.** Mirroring secrev's
"one pure-code script owns the block decision," the `gate-and-pr` gate reads
**only** the authenticated `needs.*.result` job conclusions (GitHub-controlled,
patch-independent) keyed to the run id + diff hash. It **never trusts a
success/failure file or artifact the patch could have written.** The verifier
*agent* only reads failures to propose the next fix; it cannot declare
success.
5. **Branch protection is the final enforcement.** The draft PR targets a
protected branch; the locked-down checks are required; merge needs them green
**plus** the security-review, the Claude Code App review, and **human
approval**. The agent never auto-merges (D2).
**Threat-model honesty (§3.3.2).** The diff hash protects box → CI transit
(tamper/substitution), **not** a box that is itself compromised — the box is in
the trusted computing base for what it *proposes*. The real backstops against a
bad box are the credential-less build job, the trust-control-surface denylist,
and above all the **human review of the draft PR** plus required checks (nothing
merges without Adam).
## Job graph
```
workflow_dispatch (task_id, diff_artifact_name, expected_diff_hash, declared_scope)
│
▼
guard contents:read, no secrets — validates the diff as DATA:
(boundaries 2,3) re-hash + denylist + scope. Never applies it.
│ (needs)
▼
build-test contents:read, no secrets, no OIDC, egress blocked —
(boundary 1) the ONLY job that applies + runs the UNTRUSTED patch.
│ (needs) Emits a NON-authoritative report artifact.
▼
gate-and-pr privileged home, but NEVER checks out patch code. Pure-code
(boundaries 1,4,5) gate over authenticated needs.*.result → DRAFT PR
(hard-disabled until the review gates pass).
```
`permissions: {}` at the workflow level (least privilege); each job re-declares
its own grant explicitly. The trigger is `workflow_dispatch` only — the patch
never runs in a context carrying write or secret scope.
## SHA-pinned actions (handbook Pinning Principle, §3.3.2)
Every third-party action is pinned to a full commit SHA with the human-readable
tag in a trailing comment:
| Action | SHA | Tag |
|---|---|---|
| `actions/checkout` | `11bd71901bbe5b1630ceea73d27597364c9af683` | v4.2.2 |
| `actions/download-artifact` | `fa0a91b85d4f404e444e00e005971372dc801d16` | v4.1.8 |
| `actions/upload-artifact` | `b4b15b8c7c6ac21ea08fcf65892d2ee8f75cf882` | v4.4.3 |
| `actions/setup-python` | `0b93645e9fea7318ecaed2b359559ac225c90a2b` | v5.3.0 |
| `step-security/harden-runner` | `0080882f6c36860b6ba35c610c98ce87d4e2f26f` | v2.10.2 |
## Relationship to the foundation
This leaf **imports the committed Plane-2 foundation contracts verbatim** (it
does not redefine them):
- The diff-hash recorded in the ledger and re-checked in CI is the same
content-hash idiom as `agent_team.state_store.compute_content_hash` (§6.7).
- The task this workflow verifies is an `agent_team.task_model.TaskRecord`; its
`candidate_diff` + `diff_hash` fields (§3.3) are exactly the
`expected_diff_hash` this workflow consumes, and `ci_results` is what the
verifier writes back from the authenticated gate (boundary 4).
- The ledger that records provenance (diff hash, run id, gate decision) is the
`agent_team.db` schema (`pending_questions` / `budget_ledger` live there;
per-task CI provenance is recorded against the task thread).
## Tests
The workflow's embedded gate logic (diff integrity, the trust-control denylist
with path-canonicalization + rename detection, declared-scope enforcement, and
the pure-code pass/fail gate) is **stdlib-only, type-hinted, and ruff-clean**.
It was verified against good and adversarial diffs (workflow edits, IAM edits,
renames into denied paths, path traversal, out-of-scope and unscoped diffs, hash
mismatch, and every gate branch). Because this leaf owns only the two files in
this directory, the executable pytest suite for the importable foundation
modules lives in `../tests/` (run `python3 -m pytest agent-team/tests/ -q` from
the repo root); the inline CI gate logic is validated as part of the workflow's
own steps at deploy time and was proven correct during authoring.
## Deploy gating (do NOT skip)
Before this ships (§3.3.2, §7.1 P3):
1. `/sh-security-review` over this workflow (IaC + untrusted-input handling).
2. Mandatory **GPT-4.1 cross-review** of the workflow **and** the Option-B OIDC
role it will assume (IAM change).
3. A documented, **exercised** rollback (remove the role, revert the workflow).
4. Only then: uncomment the `id-token` / `pull-requests: write` grants, enable
the draft-PR step, and promote to `Sea-Haven-Industries/.github`. Draft PRs
only; never auto-merge.