Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.
Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled
KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven
Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 14:17:44 -04:00
|
|
|
"""Unit tests for agent_team.ci_gate — the pure-code pass/fail gate (§3.3.2)."""
|
|
|
|
|
|
|
|
|
|
from __future__ import annotations
|
|
|
|
|
|
|
|
|
|
import pytest
|
|
|
|
|
|
|
|
|
|
from agent_team.ci_gate import (
|
|
|
|
|
DENYLIST_GLOBS,
|
|
|
|
|
CiGateError,
|
|
|
|
|
GateDecision,
|
|
|
|
|
denylist_violations,
|
|
|
|
|
diff_touched_paths,
|
|
|
|
|
evaluate_ci_gate,
|
feat(agent-team): P3-flip Phase 1 — CI trust-boundary hardening (WIP, gated) (#34)
* feat(agent-team): P3-flip Phase 1 — expand denylist vectors (§4.2) + runner-trust assertion (§4.1)
First controls of the P3-live-flip Phase-1 CI hardening (workflow stays INERT;
this only tightens the trust boundary). Whole Phase-1 surface is gated by
/sh-security-review + GPT-4.1 cross-review before any flip.
§4.2 — expand the trust-control denylist with direct code-execution / supply-chain
vectors, kept byte-identical across all three copies (ci_gate.DENYLIST_GLOBS + the
guard + post-build inline DENY_GLOBS), drift-guarded:
.gitmodules, .husky/**, .githooks/**, .gitattributes, .npmrc, and generated/build
artifacts (__generated__, *.generated.*, dist/**, build/**, *.min.js).
Deliberate: lockfiles are NOT wholesale denied — lockfile-postinstall RCE is already
contained by the credential-less egress-blocked build sandbox, and the Tier-3 dep-CVE
fixer rewrites lockfiles to produce its draft PRs; a blanket deny would make it
un-shippable. Flagged in-code for the security gate. Direct code-execution config
(hooks/filters/npmrc/submodules) is the actual §4.2 RCE surface.
§4.1 — runner-trust: assert no job (esp. the privileged gate-and-pr) can run on a
self-hosted/user-provided runner; all must be GitHub-hosted.
998 tests pass, ruff clean.
* feat(agent-team): P3-flip Phase 1 — gate-weakening detector (§4.5)
A diff that ADDS a lint/type/coverage/security suppression (noqa, type: ignore,
pragma: no cover, nosec, nosemgrep), a test skip/xfail, or a hook bypass
(--no-verify) could make CI pass falsely. The pure-code gate now flags these via
gate_weakening_violations() and BLOCKs in evaluate_ci_gate as a top-priority trust
violation (step 1b, alongside the denylist) — regardless of the authenticated CI
conclusion. A build cannot pass itself by disabling its own checks; flagged diffs
escalate to a human. Only ADDED lines are inspected (removing a suppression is fine).
1015 tests pass, ruff clean.
* feat(agent-team): P3-flip — diff transport (§4.3) + flip privileged apply path live
Completes the box->CI diff handoff and flips the apply/verify privileged job
live (gated behind the agent-apply environment's required reviewer).
Transport (§4.3): the read-only box (D2) emits a diff but holds no write token.
- New credential-less `materialize` job decodes the untrusted `diff_b64`
dispatch input via env (CWE-94), fail-closed re-hashes it against
`expected_diff_hash`, and uploads it as the named artifact so guard/build-test
download it same-run. guard now `needs: materialize`.
- New `dispatcher.py` (the trusted apply path, operator/Mac-side — never the
box): pushes the diff as a head branch then `gh workflow run`s the workflow.
Pure input-assembly (sha256 == sha256sum, b64 round-trip, head ref) is
unit-tested; git/gh are injected seams. Push-before-dispatch; fail-closed on
empty diff/scope, unsafe task_id/owner/repo.
Flip: gate-and-pr binds `environment: agent-apply` (required reviewer
amoussa1229) + grants exactly `pull-requests: write`; the App-token + draft-PR
steps run only on `steps.gate.outputs.gate == 'pass'` (no more if:false); the
draft PR opens with an explicit `--head`; task_id/head_branch charset-validated
(§4.6). Updated the hardening tests from inert-state to live-state assertions +
added transport tests. 1039 tests, ruff clean, workflow YAML valid.
NOTE: workflow only runs on manual workflow_dispatch and the privileged job is
held at the required-reviewer gate, so nothing privileged runs unapproved.
* fix(agent-team): P3-flip — address GPT-4.1 cross-review (size bound, ref-traversal guard)
- BLOCK: cap candidate diff at 40 KB in the dispatcher (the diff rides a base64
workflow_dispatch input; GitHub caps inputs at ~64 KB so an oversized diff
cannot dispatch at all) + a defense-in-depth decoded-size bound in materialize.
- FIX: harden the draft-PR HEAD_BRANCH guard to reject leading/trailing slash,
'..' segments, and '//' (CWE-88 git ref-traversal), not just bad charset.
- NIT: document the mandatory invariants on gate-and-pr (required-reviewer
environment must stay; runs-on must stay GitHub-hosted).
- QUESTION (lockfiles): answered in-code — the build-test sandbox is
credential-less + egress-blocked, so lockfile-postinstall RCE is contained.
Tests added for all guards. 1042 tests, ruff clean, YAML valid.
* fix(agent-team): P3-flip — resolve /sh-security-review findings (LOGIC-1/2/3)
High-recall fan-out (injection/logic/iac+secrets) + proof-or-kill on the LIVE
apply path found 3 real issues the cross-review missed; all fixed:
- LOGIC-2 (HIGH, was a live hole): build-test ran `ruff check . || echo` /
`pytest -q || echo`, swallowing failures so the job was always 'success' and
the gate would open draft PRs on RED builds. ruff/pytest now run
authoritatively under set -e (pytest exit 5 'no tests' is the only non-fatal
case); the exit code IS the build-test conclusion the gate keys on.
- LOGIC-1 (verified!=shipped): the dispatcher used `git apply` + `git add -A`,
staging stray untracked content into the pushed PR head. Now `git apply
--index` stages exactly the diff, so the head tree is precisely base+diff —
bound to the bytes CI hash-verified.
- LOGIC-3 (§4.5 on the live path): gate-weakening was enforced only box-side;
added a gate-weakening check to the guard job so the live PR-opening path
rejects a diff that adds suppressions/skips, even on a green build.
Injection / secrets / least-privilege / flip-correctness / no-untrusted-checkout
all came back clean. 1044 tests, ruff clean, YAML valid.
2026-06-22 18:51:52 -04:00
|
|
|
gate_weakening_violations,
|
Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.
Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled
KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven
Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 14:17:44 -04:00
|
|
|
verify_diff_hash,
|
|
|
|
|
)
|
|
|
|
|
from agent_team.state_store import compute_content_hash
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _diff_for(*paths: str) -> str:
|
|
|
|
|
"""Build a minimal unified diff touching ``paths`` (no rename)."""
|
|
|
|
|
chunks = []
|
|
|
|
|
for p in paths:
|
|
|
|
|
chunks.append(
|
|
|
|
|
f"diff --git a/{p} b/{p}\n--- a/{p}\n+++ b/{p}\n@@ -1 +1 @@\n-old\n+new\n"
|
|
|
|
|
)
|
|
|
|
|
return "".join(chunks)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _ledger_hash(diff: str) -> str:
|
|
|
|
|
return compute_content_hash(diff.encode("utf-8"))
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _good_ci(run_id: str, diff: str, conclusion: str = "success") -> dict:
|
|
|
|
|
return {
|
|
|
|
|
"run_id": run_id,
|
|
|
|
|
"conclusion": conclusion,
|
|
|
|
|
"diff_hash": _ledger_hash(diff),
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
# diff_touched_paths
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_touched_paths_basic() -> None:
|
|
|
|
|
diff = _diff_for("src/foo.py", "tests/test_foo.py")
|
|
|
|
|
assert diff_touched_paths(diff) == ["src/foo.py", "tests/test_foo.py"]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_touched_paths_strips_git_prefix_and_dedups() -> None:
|
|
|
|
|
diff = "diff --git a/pkg/mod.py b/pkg/mod.py\n@@ @@\n+x\n"
|
|
|
|
|
assert diff_touched_paths(diff) == ["pkg/mod.py"]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_touched_paths_includes_rename_lines() -> None:
|
|
|
|
|
diff = (
|
|
|
|
|
"diff --git a/safe.txt b/.github/workflows/evil.yml\n"
|
|
|
|
|
"similarity index 100%\n"
|
|
|
|
|
"rename from safe.txt\n"
|
|
|
|
|
"rename to .github/workflows/evil.yml\n"
|
|
|
|
|
)
|
|
|
|
|
paths = diff_touched_paths(diff)
|
|
|
|
|
assert ".github/workflows/evil.yml" in paths
|
|
|
|
|
assert "safe.txt" in paths
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_touched_paths_non_string_raises() -> None:
|
|
|
|
|
with pytest.raises(CiGateError):
|
|
|
|
|
diff_touched_paths(None) # type: ignore[arg-type]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
# denylist_violations
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_clean_diff_has_no_violations() -> None:
|
|
|
|
|
diff = _diff_for("src/foo.py", "README.md")
|
|
|
|
|
assert denylist_violations(diff) == []
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
|
|
|
"path",
|
|
|
|
|
[
|
|
|
|
|
".github/workflows/ci.yml",
|
|
|
|
|
".github/actions/deploy/action.yml",
|
|
|
|
|
".github/CODEOWNERS",
|
|
|
|
|
"CODEOWNERS",
|
|
|
|
|
".github/dependabot.yml",
|
|
|
|
|
"infra/cdk.json",
|
|
|
|
|
"service/template.yaml",
|
|
|
|
|
"deploy/iam/role.json",
|
|
|
|
|
"stacks/policies/admin.json",
|
|
|
|
|
"modules/main.tf",
|
feat(agent-team): P3-flip Phase 1 — CI trust-boundary hardening (WIP, gated) (#34)
* feat(agent-team): P3-flip Phase 1 — expand denylist vectors (§4.2) + runner-trust assertion (§4.1)
First controls of the P3-live-flip Phase-1 CI hardening (workflow stays INERT;
this only tightens the trust boundary). Whole Phase-1 surface is gated by
/sh-security-review + GPT-4.1 cross-review before any flip.
§4.2 — expand the trust-control denylist with direct code-execution / supply-chain
vectors, kept byte-identical across all three copies (ci_gate.DENYLIST_GLOBS + the
guard + post-build inline DENY_GLOBS), drift-guarded:
.gitmodules, .husky/**, .githooks/**, .gitattributes, .npmrc, and generated/build
artifacts (__generated__, *.generated.*, dist/**, build/**, *.min.js).
Deliberate: lockfiles are NOT wholesale denied — lockfile-postinstall RCE is already
contained by the credential-less egress-blocked build sandbox, and the Tier-3 dep-CVE
fixer rewrites lockfiles to produce its draft PRs; a blanket deny would make it
un-shippable. Flagged in-code for the security gate. Direct code-execution config
(hooks/filters/npmrc/submodules) is the actual §4.2 RCE surface.
§4.1 — runner-trust: assert no job (esp. the privileged gate-and-pr) can run on a
self-hosted/user-provided runner; all must be GitHub-hosted.
998 tests pass, ruff clean.
* feat(agent-team): P3-flip Phase 1 — gate-weakening detector (§4.5)
A diff that ADDS a lint/type/coverage/security suppression (noqa, type: ignore,
pragma: no cover, nosec, nosemgrep), a test skip/xfail, or a hook bypass
(--no-verify) could make CI pass falsely. The pure-code gate now flags these via
gate_weakening_violations() and BLOCKs in evaluate_ci_gate as a top-priority trust
violation (step 1b, alongside the denylist) — regardless of the authenticated CI
conclusion. A build cannot pass itself by disabling its own checks; flagged diffs
escalate to a human. Only ADDED lines are inspected (removing a suppression is fine).
1015 tests pass, ruff clean.
* feat(agent-team): P3-flip — diff transport (§4.3) + flip privileged apply path live
Completes the box->CI diff handoff and flips the apply/verify privileged job
live (gated behind the agent-apply environment's required reviewer).
Transport (§4.3): the read-only box (D2) emits a diff but holds no write token.
- New credential-less `materialize` job decodes the untrusted `diff_b64`
dispatch input via env (CWE-94), fail-closed re-hashes it against
`expected_diff_hash`, and uploads it as the named artifact so guard/build-test
download it same-run. guard now `needs: materialize`.
- New `dispatcher.py` (the trusted apply path, operator/Mac-side — never the
box): pushes the diff as a head branch then `gh workflow run`s the workflow.
Pure input-assembly (sha256 == sha256sum, b64 round-trip, head ref) is
unit-tested; git/gh are injected seams. Push-before-dispatch; fail-closed on
empty diff/scope, unsafe task_id/owner/repo.
Flip: gate-and-pr binds `environment: agent-apply` (required reviewer
amoussa1229) + grants exactly `pull-requests: write`; the App-token + draft-PR
steps run only on `steps.gate.outputs.gate == 'pass'` (no more if:false); the
draft PR opens with an explicit `--head`; task_id/head_branch charset-validated
(§4.6). Updated the hardening tests from inert-state to live-state assertions +
added transport tests. 1039 tests, ruff clean, workflow YAML valid.
NOTE: workflow only runs on manual workflow_dispatch and the privileged job is
held at the required-reviewer gate, so nothing privileged runs unapproved.
* fix(agent-team): P3-flip — address GPT-4.1 cross-review (size bound, ref-traversal guard)
- BLOCK: cap candidate diff at 40 KB in the dispatcher (the diff rides a base64
workflow_dispatch input; GitHub caps inputs at ~64 KB so an oversized diff
cannot dispatch at all) + a defense-in-depth decoded-size bound in materialize.
- FIX: harden the draft-PR HEAD_BRANCH guard to reject leading/trailing slash,
'..' segments, and '//' (CWE-88 git ref-traversal), not just bad charset.
- NIT: document the mandatory invariants on gate-and-pr (required-reviewer
environment must stay; runs-on must stay GitHub-hosted).
- QUESTION (lockfiles): answered in-code — the build-test sandbox is
credential-less + egress-blocked, so lockfile-postinstall RCE is contained.
Tests added for all guards. 1042 tests, ruff clean, YAML valid.
* fix(agent-team): P3-flip — resolve /sh-security-review findings (LOGIC-1/2/3)
High-recall fan-out (injection/logic/iac+secrets) + proof-or-kill on the LIVE
apply path found 3 real issues the cross-review missed; all fixed:
- LOGIC-2 (HIGH, was a live hole): build-test ran `ruff check . || echo` /
`pytest -q || echo`, swallowing failures so the job was always 'success' and
the gate would open draft PRs on RED builds. ruff/pytest now run
authoritatively under set -e (pytest exit 5 'no tests' is the only non-fatal
case); the exit code IS the build-test conclusion the gate keys on.
- LOGIC-1 (verified!=shipped): the dispatcher used `git apply` + `git add -A`,
staging stray untracked content into the pushed PR head. Now `git apply
--index` stages exactly the diff, so the head tree is precisely base+diff —
bound to the bytes CI hash-verified.
- LOGIC-3 (§4.5 on the live path): gate-weakening was enforced only box-side;
added a gate-weakening check to the guard job so the live PR-opening path
rejects a diff that adds suppressions/skips, even on a green build.
Injection / secrets / least-privilege / flip-correctness / no-untrusted-checkout
all came back clean. 1044 tests, ruff clean, YAML valid.
2026-06-22 18:51:52 -04:00
|
|
|
# P3-flip §4.2 — direct code-execution / supply-chain vectors.
|
|
|
|
|
".gitmodules",
|
|
|
|
|
"vendor/.gitmodules",
|
|
|
|
|
".husky/pre-commit",
|
|
|
|
|
"frontend/.husky/commit-msg",
|
|
|
|
|
".githooks/pre-push",
|
|
|
|
|
".gitattributes",
|
|
|
|
|
"pkg/.gitattributes",
|
|
|
|
|
".npmrc",
|
|
|
|
|
"web/.npmrc",
|
|
|
|
|
"src/__generated__/schema.ts",
|
|
|
|
|
"api/types.generated.ts",
|
|
|
|
|
"web/dist/bundle.js",
|
|
|
|
|
"service/build/output.o",
|
|
|
|
|
"assets/app.min.js",
|
Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.
Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled
KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven
Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 14:17:44 -04:00
|
|
|
],
|
|
|
|
|
)
|
|
|
|
|
def test_denylisted_paths_flagged(path: str) -> None:
|
|
|
|
|
violations = denylist_violations(_diff_for(path))
|
|
|
|
|
assert violations, f"expected {path!r} to be denylisted"
|
|
|
|
|
|
|
|
|
|
|
feat(agent-team): P3-flip Phase 1 — CI trust-boundary hardening (WIP, gated) (#34)
* feat(agent-team): P3-flip Phase 1 — expand denylist vectors (§4.2) + runner-trust assertion (§4.1)
First controls of the P3-live-flip Phase-1 CI hardening (workflow stays INERT;
this only tightens the trust boundary). Whole Phase-1 surface is gated by
/sh-security-review + GPT-4.1 cross-review before any flip.
§4.2 — expand the trust-control denylist with direct code-execution / supply-chain
vectors, kept byte-identical across all three copies (ci_gate.DENYLIST_GLOBS + the
guard + post-build inline DENY_GLOBS), drift-guarded:
.gitmodules, .husky/**, .githooks/**, .gitattributes, .npmrc, and generated/build
artifacts (__generated__, *.generated.*, dist/**, build/**, *.min.js).
Deliberate: lockfiles are NOT wholesale denied — lockfile-postinstall RCE is already
contained by the credential-less egress-blocked build sandbox, and the Tier-3 dep-CVE
fixer rewrites lockfiles to produce its draft PRs; a blanket deny would make it
un-shippable. Flagged in-code for the security gate. Direct code-execution config
(hooks/filters/npmrc/submodules) is the actual §4.2 RCE surface.
§4.1 — runner-trust: assert no job (esp. the privileged gate-and-pr) can run on a
self-hosted/user-provided runner; all must be GitHub-hosted.
998 tests pass, ruff clean.
* feat(agent-team): P3-flip Phase 1 — gate-weakening detector (§4.5)
A diff that ADDS a lint/type/coverage/security suppression (noqa, type: ignore,
pragma: no cover, nosec, nosemgrep), a test skip/xfail, or a hook bypass
(--no-verify) could make CI pass falsely. The pure-code gate now flags these via
gate_weakening_violations() and BLOCKs in evaluate_ci_gate as a top-priority trust
violation (step 1b, alongside the denylist) — regardless of the authenticated CI
conclusion. A build cannot pass itself by disabling its own checks; flagged diffs
escalate to a human. Only ADDED lines are inspected (removing a suppression is fine).
1015 tests pass, ruff clean.
* feat(agent-team): P3-flip — diff transport (§4.3) + flip privileged apply path live
Completes the box->CI diff handoff and flips the apply/verify privileged job
live (gated behind the agent-apply environment's required reviewer).
Transport (§4.3): the read-only box (D2) emits a diff but holds no write token.
- New credential-less `materialize` job decodes the untrusted `diff_b64`
dispatch input via env (CWE-94), fail-closed re-hashes it against
`expected_diff_hash`, and uploads it as the named artifact so guard/build-test
download it same-run. guard now `needs: materialize`.
- New `dispatcher.py` (the trusted apply path, operator/Mac-side — never the
box): pushes the diff as a head branch then `gh workflow run`s the workflow.
Pure input-assembly (sha256 == sha256sum, b64 round-trip, head ref) is
unit-tested; git/gh are injected seams. Push-before-dispatch; fail-closed on
empty diff/scope, unsafe task_id/owner/repo.
Flip: gate-and-pr binds `environment: agent-apply` (required reviewer
amoussa1229) + grants exactly `pull-requests: write`; the App-token + draft-PR
steps run only on `steps.gate.outputs.gate == 'pass'` (no more if:false); the
draft PR opens with an explicit `--head`; task_id/head_branch charset-validated
(§4.6). Updated the hardening tests from inert-state to live-state assertions +
added transport tests. 1039 tests, ruff clean, workflow YAML valid.
NOTE: workflow only runs on manual workflow_dispatch and the privileged job is
held at the required-reviewer gate, so nothing privileged runs unapproved.
* fix(agent-team): P3-flip — address GPT-4.1 cross-review (size bound, ref-traversal guard)
- BLOCK: cap candidate diff at 40 KB in the dispatcher (the diff rides a base64
workflow_dispatch input; GitHub caps inputs at ~64 KB so an oversized diff
cannot dispatch at all) + a defense-in-depth decoded-size bound in materialize.
- FIX: harden the draft-PR HEAD_BRANCH guard to reject leading/trailing slash,
'..' segments, and '//' (CWE-88 git ref-traversal), not just bad charset.
- NIT: document the mandatory invariants on gate-and-pr (required-reviewer
environment must stay; runs-on must stay GitHub-hosted).
- QUESTION (lockfiles): answered in-code — the build-test sandbox is
credential-less + egress-blocked, so lockfile-postinstall RCE is contained.
Tests added for all guards. 1042 tests, ruff clean, YAML valid.
* fix(agent-team): P3-flip — resolve /sh-security-review findings (LOGIC-1/2/3)
High-recall fan-out (injection/logic/iac+secrets) + proof-or-kill on the LIVE
apply path found 3 real issues the cross-review missed; all fixed:
- LOGIC-2 (HIGH, was a live hole): build-test ran `ruff check . || echo` /
`pytest -q || echo`, swallowing failures so the job was always 'success' and
the gate would open draft PRs on RED builds. ruff/pytest now run
authoritatively under set -e (pytest exit 5 'no tests' is the only non-fatal
case); the exit code IS the build-test conclusion the gate keys on.
- LOGIC-1 (verified!=shipped): the dispatcher used `git apply` + `git add -A`,
staging stray untracked content into the pushed PR head. Now `git apply
--index` stages exactly the diff, so the head tree is precisely base+diff —
bound to the bytes CI hash-verified.
- LOGIC-3 (§4.5 on the live path): gate-weakening was enforced only box-side;
added a gate-weakening check to the guard job so the live PR-opening path
rejects a diff that adds suppressions/skips, even on a green build.
Injection / secrets / least-privilege / flip-correctness / no-untrusted-checkout
all came back clean. 1044 tests, ruff clean, YAML valid.
2026-06-22 18:51:52 -04:00
|
|
|
def test_lockfiles_are_NOT_denylisted_so_the_tier3_fixer_can_ship() -> None:
|
|
|
|
|
"""Deliberate (§4.2 note): lockfile RCE is contained by the credential-less
|
|
|
|
|
egress-blocked build sandbox, and the dep-CVE fixer rewrites lockfiles to
|
|
|
|
|
produce draft PRs — so lockfiles are intentionally NOT on the denylist."""
|
|
|
|
|
for lock in ("package-lock.json", "web/yarn.lock", "poetry.lock", "Cargo.lock"):
|
|
|
|
|
assert denylist_violations(_diff_for(lock)) == [], (
|
|
|
|
|
f"{lock!r} must not be denylisted (would break the Tier-3 fixer)"
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.
Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled
KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven
Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 14:17:44 -04:00
|
|
|
def test_rename_into_workflow_is_flagged() -> None:
|
|
|
|
|
diff = (
|
|
|
|
|
"diff --git a/safe.txt b/.github/workflows/evil.yml\n"
|
|
|
|
|
"rename from safe.txt\n"
|
|
|
|
|
"rename to .github/workflows/evil.yml\n"
|
|
|
|
|
)
|
|
|
|
|
violations = denylist_violations(diff)
|
|
|
|
|
assert any(".github/workflows/evil.yml" in v for v in violations)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_path_escape_via_dotdot_flagged() -> None:
|
|
|
|
|
diff = "diff --git a/x b/../../etc/passwd\n@@ @@\n+x\n"
|
|
|
|
|
violations = denylist_violations(diff)
|
|
|
|
|
assert any("escapes repo root" in v for v in violations)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_allowed_scope_blocks_out_of_scope_path() -> None:
|
|
|
|
|
diff = _diff_for("src/in_scope.py", "other/out_of_scope.py")
|
|
|
|
|
violations = denylist_violations(diff, allowed_scope=["src/"])
|
|
|
|
|
assert any("declared scope" in v for v in violations)
|
|
|
|
|
# In-scope path alone is clean.
|
|
|
|
|
assert (
|
|
|
|
|
denylist_violations(_diff_for("src/in_scope.py"), allowed_scope=["src/"]) == []
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_allowed_scope_exact_file_prefix() -> None:
|
|
|
|
|
diff = _diff_for("pkg/exact.py")
|
|
|
|
|
assert denylist_violations(diff, allowed_scope=["pkg/exact.py"]) == []
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_denylist_globs_exported_nonempty() -> None:
|
|
|
|
|
assert isinstance(DENYLIST_GLOBS, tuple)
|
|
|
|
|
assert ".github/workflows/**" in DENYLIST_GLOBS
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
# verify_diff_hash
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_hash_matches_ledger() -> None:
|
|
|
|
|
diff = _diff_for("a.py")
|
|
|
|
|
assert verify_diff_hash(diff, ledger_hash=_ledger_hash(diff)) is True
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_hash_mismatch_ledger() -> None:
|
|
|
|
|
diff = _diff_for("a.py")
|
|
|
|
|
assert verify_diff_hash(diff, ledger_hash="deadbeef") is False
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_hash_none_ledger_fails() -> None:
|
|
|
|
|
diff = _diff_for("a.py")
|
|
|
|
|
assert verify_diff_hash(diff, ledger_hash=None) is False
|
|
|
|
|
|
|
|
|
|
|
feat(agent-team): P3-live CI apply/verify hardening + ci_fetcher (gate-passed, provisioning-gated) (#17)
* feat(agent-team): read-only CI-result fetcher for P3 verify gate (opt-in, inert)
ci_fetcher.py: fail-closed CiResultFetcher reading the GitHub Actions run
conclusion via a read-only PAT (AGENT_TEAM_CI_READ_TOKEN→GITHUB_TOKEN), returns
{run_id,conclusion,diff_hash} or None on any error. Data-fetcher only — ci_gate
owns the verdict; never writes, no OIDC/AWS, never reads patch artifacts.
coordinator gains opt-in gated_build_verify_wiring() composing it via
bind_ci_result_fetcher; NOT wired into the default run-team.py path. 20 tests.
* harden(agent-team): P3 apply/verify workflow — GitHub App token, CWE-94, fail-closed
Decision-1 auth model: gate-and-pr uses a GitHub App installation token
(pull-requests:write) behind the agent-apply environment; ALL OIDC/id-token/AWS
removed. Hardening: task_id env-indirection (CWE-94 — GitHub expands ${{ }} into
the run shell before exec, so %s/quoting is insufficient); run-id pinning on both
download-artifact; post-build denied-path check (build-hook writes into denied
paths fail the job); empty-hash fail-closed in BOTH the embedded gate (fixed a
real ''=='' pass bug) and ci_gate.py. App-token + draft-PR steps stay if:${{ false }}
until provisioning (App + environment + branch protection). +17 tests.
* harden(agent-team): apply P3-live security-gate fixes (GPT-4.1 xreview + sh-security-review)
BLOCK-1/FIX-4: gate-and-pr re-comments pull-requests:write + environment:agent-apply
(provisioning-time uncomment) and gains needs.guard/build-test=='success' job guard —
zero privilege until provisioning. BLOCK-2/3+FIX-5: ci_fetcher validates run_id (^[0-9]{1,20}$),
owner/repo (^[A-Za-z0-9_.-]{1,100}$), and fetched_id (int) — fail closed, no SSRF/path
injection. FIX-1: conclusion allowlist. FIX-3: api_root removed from public builder (no
injectable endpoint). INJ-02: post-build denied-path check uses NUL-delimited git output +
explicit rename parsing, no backslash mangling, non-UTF8=violation. INJ-03: all three trust-
control denylists unified to one 22-entry union + drift-guard test. Q1: documented run_id/
diff_hash trust source (dispatcher/ledger only). 884 tests, ruff clean. Privileged steps stay
if:${{ false }} until provisioning.
* build(security-review): prune .claude worktrees from deterministic scanners
Agent worktrees under .claude/worktrees/ are full repo copies; the cfn-lint
find|xargs template scan overflowed ('command line cannot be assembled') and the
pre-push hook fail-closed to BLOCK whenever a worktree was present. Prune .claude
in the cfn-lint find + semgrep/checkov excludes, and gitignore .claude/ so it is
never scanned or committed. Unblocks main-tree pushes during parallel agent work.
2026-06-18 15:53:26 -04:00
|
|
|
def test_hash_empty_ledger_fails_closed() -> None:
|
|
|
|
|
# An EMPTY expected hash must never be treated as a match (empty != real
|
|
|
|
|
# sha). Fail-closed: there is nothing to bind to.
|
|
|
|
|
diff = _diff_for("a.py")
|
|
|
|
|
assert verify_diff_hash(diff, ledger_hash="") is False
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_hash_empty_ci_verified_does_not_silently_pass() -> None:
|
|
|
|
|
# A real ledger hash but an EMPTY ci_verified hash must not match (the empty
|
|
|
|
|
# string is not the recomputed sha).
|
|
|
|
|
diff = _diff_for("a.py")
|
|
|
|
|
h = _ledger_hash(diff)
|
|
|
|
|
assert verify_diff_hash(diff, ledger_hash=h, ci_verified_hash="") is False
|
|
|
|
|
|
|
|
|
|
|
Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.
Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled
KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven
Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 14:17:44 -04:00
|
|
|
def test_hash_ci_verified_must_also_match() -> None:
|
|
|
|
|
diff = _diff_for("a.py")
|
|
|
|
|
h = _ledger_hash(diff)
|
|
|
|
|
assert verify_diff_hash(diff, ledger_hash=h, ci_verified_hash=h) is True
|
|
|
|
|
assert verify_diff_hash(diff, ledger_hash=h, ci_verified_hash="nope") is False
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_hash_non_string_raises() -> None:
|
|
|
|
|
with pytest.raises(CiGateError):
|
|
|
|
|
verify_diff_hash(123, ledger_hash="x") # type: ignore[arg-type]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
# evaluate_ci_gate — the block decision
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_gate_pass_on_authenticated_success() -> None:
|
|
|
|
|
diff = _diff_for("src/foo.py")
|
|
|
|
|
result = evaluate_ci_gate(
|
|
|
|
|
candidate_diff=diff,
|
|
|
|
|
ledger_hash=_ledger_hash(diff),
|
|
|
|
|
ci_result=_good_ci("run-1", diff, "success"),
|
|
|
|
|
expected_run_id="run-1",
|
|
|
|
|
)
|
|
|
|
|
assert result.decision is GateDecision.PASS
|
|
|
|
|
assert result.passed is True
|
|
|
|
|
assert result.ci_conclusion == "success"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_gate_fail_on_authenticated_failure() -> None:
|
|
|
|
|
diff = _diff_for("src/foo.py")
|
|
|
|
|
result = evaluate_ci_gate(
|
|
|
|
|
candidate_diff=diff,
|
|
|
|
|
ledger_hash=_ledger_hash(diff),
|
|
|
|
|
ci_result=_good_ci("run-1", diff, "failure"),
|
|
|
|
|
expected_run_id="run-1",
|
|
|
|
|
)
|
|
|
|
|
assert result.decision is GateDecision.FAIL
|
|
|
|
|
assert result.ci_conclusion == "failure"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
|
|
|
"conclusion", ["timed_out", "cancelled", "startup_failure", "action_required"]
|
|
|
|
|
)
|
|
|
|
|
def test_gate_recognised_failures(conclusion: str) -> None:
|
|
|
|
|
diff = _diff_for("src/foo.py")
|
|
|
|
|
result = evaluate_ci_gate(
|
|
|
|
|
candidate_diff=diff,
|
|
|
|
|
ledger_hash=_ledger_hash(diff),
|
|
|
|
|
ci_result=_good_ci("run-1", diff, conclusion),
|
|
|
|
|
expected_run_id="run-1",
|
|
|
|
|
)
|
|
|
|
|
assert result.decision is GateDecision.FAIL
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_gate_blocks_denylisted_diff_even_if_ci_success() -> None:
|
|
|
|
|
# A diff that touches the trust-control surface BLOCKs regardless of CI.
|
|
|
|
|
diff = _diff_for(".github/workflows/ci.yml")
|
|
|
|
|
result = evaluate_ci_gate(
|
|
|
|
|
candidate_diff=diff,
|
|
|
|
|
ledger_hash=_ledger_hash(diff),
|
|
|
|
|
ci_result=_good_ci("run-1", diff, "success"),
|
|
|
|
|
expected_run_id="run-1",
|
|
|
|
|
)
|
|
|
|
|
assert result.decision is GateDecision.BLOCK
|
|
|
|
|
assert result.blocked is True
|
|
|
|
|
assert any("denylisted" in r for r in result.reasons)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_gate_blocks_on_hash_mismatch() -> None:
|
|
|
|
|
diff = _diff_for("src/foo.py")
|
|
|
|
|
result = evaluate_ci_gate(
|
|
|
|
|
candidate_diff=diff,
|
|
|
|
|
ledger_hash="deadbeef", # does not match recomputed hash
|
|
|
|
|
ci_result=_good_ci("run-1", diff, "success"),
|
|
|
|
|
expected_run_id="run-1",
|
|
|
|
|
)
|
|
|
|
|
assert result.decision is GateDecision.BLOCK
|
|
|
|
|
assert any("hash" in r for r in result.reasons)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_gate_blocks_on_run_id_mismatch() -> None:
|
|
|
|
|
diff = _diff_for("src/foo.py")
|
|
|
|
|
result = evaluate_ci_gate(
|
|
|
|
|
candidate_diff=diff,
|
|
|
|
|
ledger_hash=_ledger_hash(diff),
|
|
|
|
|
ci_result=_good_ci("run-OTHER", diff, "success"),
|
|
|
|
|
expected_run_id="run-1",
|
|
|
|
|
)
|
|
|
|
|
assert result.decision is GateDecision.BLOCK
|
|
|
|
|
assert any("run-id" in r for r in result.reasons)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_gate_blocks_on_missing_ci_result() -> None:
|
|
|
|
|
diff = _diff_for("src/foo.py")
|
|
|
|
|
result = evaluate_ci_gate(
|
|
|
|
|
candidate_diff=diff,
|
|
|
|
|
ledger_hash=_ledger_hash(diff),
|
|
|
|
|
ci_result=None,
|
|
|
|
|
expected_run_id="run-1",
|
|
|
|
|
)
|
|
|
|
|
assert result.decision is GateDecision.BLOCK
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@pytest.mark.parametrize("conclusion", [None, "neutral", "skipped", "in_progress"])
|
|
|
|
|
def test_gate_blocks_on_ambiguous_conclusion(conclusion) -> None:
|
|
|
|
|
# Anything not an explicit success or recognised failure must NOT silently
|
|
|
|
|
# pass — it BLOCKs (refuse-to-proceed).
|
|
|
|
|
diff = _diff_for("src/foo.py")
|
|
|
|
|
ci = _good_ci("run-1", diff, "success")
|
|
|
|
|
ci["conclusion"] = conclusion
|
|
|
|
|
result = evaluate_ci_gate(
|
|
|
|
|
candidate_diff=diff,
|
|
|
|
|
ledger_hash=_ledger_hash(diff),
|
|
|
|
|
ci_result=ci,
|
|
|
|
|
expected_run_id="run-1",
|
|
|
|
|
)
|
|
|
|
|
assert result.decision is GateDecision.BLOCK
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_gate_ignores_patch_written_success_field() -> None:
|
|
|
|
|
# The gate reads only the authenticated conclusion. A patch-controlled
|
|
|
|
|
# "passed" flag must not flip a failing run to pass.
|
|
|
|
|
diff = _diff_for("src/foo.py")
|
|
|
|
|
ci = _good_ci("run-1", diff, "failure")
|
|
|
|
|
ci["passed"] = True # attacker-controlled artifact field
|
|
|
|
|
ci["success"] = "true"
|
|
|
|
|
result = evaluate_ci_gate(
|
|
|
|
|
candidate_diff=diff,
|
|
|
|
|
ledger_hash=_ledger_hash(diff),
|
|
|
|
|
ci_result=ci,
|
|
|
|
|
expected_run_id="run-1",
|
|
|
|
|
)
|
|
|
|
|
assert result.decision is GateDecision.FAIL
|
|
|
|
|
|
|
|
|
|
|
feat(agent-team): P3-live CI apply/verify hardening + ci_fetcher (gate-passed, provisioning-gated) (#17)
* feat(agent-team): read-only CI-result fetcher for P3 verify gate (opt-in, inert)
ci_fetcher.py: fail-closed CiResultFetcher reading the GitHub Actions run
conclusion via a read-only PAT (AGENT_TEAM_CI_READ_TOKEN→GITHUB_TOKEN), returns
{run_id,conclusion,diff_hash} or None on any error. Data-fetcher only — ci_gate
owns the verdict; never writes, no OIDC/AWS, never reads patch artifacts.
coordinator gains opt-in gated_build_verify_wiring() composing it via
bind_ci_result_fetcher; NOT wired into the default run-team.py path. 20 tests.
* harden(agent-team): P3 apply/verify workflow — GitHub App token, CWE-94, fail-closed
Decision-1 auth model: gate-and-pr uses a GitHub App installation token
(pull-requests:write) behind the agent-apply environment; ALL OIDC/id-token/AWS
removed. Hardening: task_id env-indirection (CWE-94 — GitHub expands ${{ }} into
the run shell before exec, so %s/quoting is insufficient); run-id pinning on both
download-artifact; post-build denied-path check (build-hook writes into denied
paths fail the job); empty-hash fail-closed in BOTH the embedded gate (fixed a
real ''=='' pass bug) and ci_gate.py. App-token + draft-PR steps stay if:${{ false }}
until provisioning (App + environment + branch protection). +17 tests.
* harden(agent-team): apply P3-live security-gate fixes (GPT-4.1 xreview + sh-security-review)
BLOCK-1/FIX-4: gate-and-pr re-comments pull-requests:write + environment:agent-apply
(provisioning-time uncomment) and gains needs.guard/build-test=='success' job guard —
zero privilege until provisioning. BLOCK-2/3+FIX-5: ci_fetcher validates run_id (^[0-9]{1,20}$),
owner/repo (^[A-Za-z0-9_.-]{1,100}$), and fetched_id (int) — fail closed, no SSRF/path
injection. FIX-1: conclusion allowlist. FIX-3: api_root removed from public builder (no
injectable endpoint). INJ-02: post-build denied-path check uses NUL-delimited git output +
explicit rename parsing, no backslash mangling, non-UTF8=violation. INJ-03: all three trust-
control denylists unified to one 22-entry union + drift-guard test. Q1: documented run_id/
diff_hash trust source (dispatcher/ledger only). 884 tests, ruff clean. Privileged steps stay
if:${{ false }} until provisioning.
* build(security-review): prune .claude worktrees from deterministic scanners
Agent worktrees under .claude/worktrees/ are full repo copies; the cfn-lint
find|xargs template scan overflowed ('command line cannot be assembled') and the
pre-push hook fail-closed to BLOCK whenever a worktree was present. Prune .claude
in the cfn-lint find + semgrep/checkov excludes, and gitignore .claude/ so it is
never scanned or committed. Unblocks main-tree pushes during parallel agent work.
2026-06-18 15:53:26 -04:00
|
|
|
def test_gate_blocks_on_empty_ledger_hash() -> None:
|
|
|
|
|
# Fail-closed: an empty ledger hash binds to nothing -> BLOCK, never a pass,
|
|
|
|
|
# even with an authenticated CI success.
|
|
|
|
|
diff = _diff_for("src/foo.py")
|
|
|
|
|
result = evaluate_ci_gate(
|
|
|
|
|
candidate_diff=diff,
|
|
|
|
|
ledger_hash="",
|
|
|
|
|
ci_result=_good_ci("run-1", diff, "success"),
|
|
|
|
|
expected_run_id="run-1",
|
|
|
|
|
)
|
|
|
|
|
assert result.decision is GateDecision.BLOCK
|
|
|
|
|
assert any("hash" in r for r in result.reasons)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_gate_blocks_on_empty_ci_diff_hash() -> None:
|
|
|
|
|
# An empty CI-verified diff_hash must not silently pass the integrity check.
|
|
|
|
|
diff = _diff_for("src/foo.py")
|
|
|
|
|
ci = _good_ci("run-1", diff, "success")
|
|
|
|
|
ci["diff_hash"] = ""
|
|
|
|
|
result = evaluate_ci_gate(
|
|
|
|
|
candidate_diff=diff,
|
|
|
|
|
ledger_hash=_ledger_hash(diff),
|
|
|
|
|
ci_result=ci,
|
|
|
|
|
expected_run_id="run-1",
|
|
|
|
|
)
|
|
|
|
|
assert result.decision is GateDecision.BLOCK
|
|
|
|
|
|
|
|
|
|
|
Add Plane-2 leaf scaffold (pipeline graph, nodes, HITL, transports, CI)
Consolidates the 18 leaf modules from the r720-plane2-scaffold workflow onto
the foundation commit. Full suite: 535 passed, 1 skipped; ruff + format clean.
Built (pre-deployment scaffold only — nothing provisioned/enabled):
- LangGraph pipeline graph.py (INTAKE->CLARIFY->PLAN, interrupt()/resume, checkpointer-injectable)
- nodes: clarifier (98% gate), planner, review_loop (GPT-4.1), builders->candidate diff, verifier
- §3.3.1 HITL: ledger ops, resume_worker, deadline_timer, recovery sweep, responder
- transports: slack / github / claude_code adapters
- ci_gate (pure-code pass/fail), operator_cli, run-team.py entry, P1 sim harness
- ci/agent-team-apply-verify.yml (split untrusted/privileged jobs) — authored, disabled
KNOWN OPEN FINDINGS (verifier/cross-review, not yet fixed — see follow-up):
- builders denylist: 4 execution-proven bypasses (delete, mode-change, copy-to, out-of-scope delete)
- §3.3.1 CAS: BEGIN IMMEDIATE outside try/except; shared-connection txn nesting unsafe under concurrency
- operator_cli: missing re-deliver/force-resume; audit-after-mutate ordering gap
- ci yaml: GPT-4.1 cross-review PASS w/ 4 FIX items (symlink path escape, etc.)
- P1 sim harness models the ledger layer, not real LangGraph interrupt/resume; P1 exit criteria not yet truly proven
Deploy-gated (NOT done): IAM/step-ca/Roles Anywhere/confluence-bot provisioning,
/sh-security-review sign-off, live Slack/CI, rsync, live dry-runs, Adam approval.
2026-06-17 14:17:44 -04:00
|
|
|
def test_gate_blocks_when_ci_diff_hash_mismatch() -> None:
|
|
|
|
|
# CI verified a different diff than the ledger recorded -> BLOCK.
|
|
|
|
|
diff = _diff_for("src/foo.py")
|
|
|
|
|
ci = _good_ci("run-1", diff, "success")
|
|
|
|
|
ci["diff_hash"] = "tampered"
|
|
|
|
|
result = evaluate_ci_gate(
|
|
|
|
|
candidate_diff=diff,
|
|
|
|
|
ledger_hash=_ledger_hash(diff),
|
|
|
|
|
ci_result=ci,
|
|
|
|
|
expected_run_id="run-1",
|
|
|
|
|
)
|
|
|
|
|
assert result.decision is GateDecision.BLOCK
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_gate_scope_violation_blocks() -> None:
|
|
|
|
|
diff = _diff_for("src/foo.py", "unrelated/bar.py")
|
|
|
|
|
result = evaluate_ci_gate(
|
|
|
|
|
candidate_diff=diff,
|
|
|
|
|
ledger_hash=_ledger_hash(diff),
|
|
|
|
|
ci_result=_good_ci("run-1", diff, "success"),
|
|
|
|
|
expected_run_id="run-1",
|
|
|
|
|
allowed_scope=["src/"],
|
|
|
|
|
)
|
|
|
|
|
assert result.decision is GateDecision.BLOCK
|
|
|
|
|
assert any("scope" in r for r in result.reasons)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_gate_empty_run_id_raises() -> None:
|
|
|
|
|
diff = _diff_for("src/foo.py")
|
|
|
|
|
with pytest.raises(CiGateError):
|
|
|
|
|
evaluate_ci_gate(
|
|
|
|
|
candidate_diff=diff,
|
|
|
|
|
ledger_hash=_ledger_hash(diff),
|
|
|
|
|
ci_result=_good_ci("run-1", diff),
|
|
|
|
|
expected_run_id="",
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_gate_result_carries_provenance() -> None:
|
|
|
|
|
diff = _diff_for("src/foo.py")
|
|
|
|
|
h = _ledger_hash(diff)
|
|
|
|
|
result = evaluate_ci_gate(
|
|
|
|
|
candidate_diff=diff,
|
|
|
|
|
ledger_hash=h,
|
|
|
|
|
ci_result=_good_ci("run-7", diff, "success"),
|
|
|
|
|
expected_run_id="run-7",
|
|
|
|
|
)
|
|
|
|
|
assert result.run_id == "run-7"
|
|
|
|
|
assert result.diff_hash == h
|
|
|
|
|
assert result.reasons # always records why
|
feat(agent-team): P3-flip Phase 1 — CI trust-boundary hardening (WIP, gated) (#34)
* feat(agent-team): P3-flip Phase 1 — expand denylist vectors (§4.2) + runner-trust assertion (§4.1)
First controls of the P3-live-flip Phase-1 CI hardening (workflow stays INERT;
this only tightens the trust boundary). Whole Phase-1 surface is gated by
/sh-security-review + GPT-4.1 cross-review before any flip.
§4.2 — expand the trust-control denylist with direct code-execution / supply-chain
vectors, kept byte-identical across all three copies (ci_gate.DENYLIST_GLOBS + the
guard + post-build inline DENY_GLOBS), drift-guarded:
.gitmodules, .husky/**, .githooks/**, .gitattributes, .npmrc, and generated/build
artifacts (__generated__, *.generated.*, dist/**, build/**, *.min.js).
Deliberate: lockfiles are NOT wholesale denied — lockfile-postinstall RCE is already
contained by the credential-less egress-blocked build sandbox, and the Tier-3 dep-CVE
fixer rewrites lockfiles to produce its draft PRs; a blanket deny would make it
un-shippable. Flagged in-code for the security gate. Direct code-execution config
(hooks/filters/npmrc/submodules) is the actual §4.2 RCE surface.
§4.1 — runner-trust: assert no job (esp. the privileged gate-and-pr) can run on a
self-hosted/user-provided runner; all must be GitHub-hosted.
998 tests pass, ruff clean.
* feat(agent-team): P3-flip Phase 1 — gate-weakening detector (§4.5)
A diff that ADDS a lint/type/coverage/security suppression (noqa, type: ignore,
pragma: no cover, nosec, nosemgrep), a test skip/xfail, or a hook bypass
(--no-verify) could make CI pass falsely. The pure-code gate now flags these via
gate_weakening_violations() and BLOCKs in evaluate_ci_gate as a top-priority trust
violation (step 1b, alongside the denylist) — regardless of the authenticated CI
conclusion. A build cannot pass itself by disabling its own checks; flagged diffs
escalate to a human. Only ADDED lines are inspected (removing a suppression is fine).
1015 tests pass, ruff clean.
* feat(agent-team): P3-flip — diff transport (§4.3) + flip privileged apply path live
Completes the box->CI diff handoff and flips the apply/verify privileged job
live (gated behind the agent-apply environment's required reviewer).
Transport (§4.3): the read-only box (D2) emits a diff but holds no write token.
- New credential-less `materialize` job decodes the untrusted `diff_b64`
dispatch input via env (CWE-94), fail-closed re-hashes it against
`expected_diff_hash`, and uploads it as the named artifact so guard/build-test
download it same-run. guard now `needs: materialize`.
- New `dispatcher.py` (the trusted apply path, operator/Mac-side — never the
box): pushes the diff as a head branch then `gh workflow run`s the workflow.
Pure input-assembly (sha256 == sha256sum, b64 round-trip, head ref) is
unit-tested; git/gh are injected seams. Push-before-dispatch; fail-closed on
empty diff/scope, unsafe task_id/owner/repo.
Flip: gate-and-pr binds `environment: agent-apply` (required reviewer
amoussa1229) + grants exactly `pull-requests: write`; the App-token + draft-PR
steps run only on `steps.gate.outputs.gate == 'pass'` (no more if:false); the
draft PR opens with an explicit `--head`; task_id/head_branch charset-validated
(§4.6). Updated the hardening tests from inert-state to live-state assertions +
added transport tests. 1039 tests, ruff clean, workflow YAML valid.
NOTE: workflow only runs on manual workflow_dispatch and the privileged job is
held at the required-reviewer gate, so nothing privileged runs unapproved.
* fix(agent-team): P3-flip — address GPT-4.1 cross-review (size bound, ref-traversal guard)
- BLOCK: cap candidate diff at 40 KB in the dispatcher (the diff rides a base64
workflow_dispatch input; GitHub caps inputs at ~64 KB so an oversized diff
cannot dispatch at all) + a defense-in-depth decoded-size bound in materialize.
- FIX: harden the draft-PR HEAD_BRANCH guard to reject leading/trailing slash,
'..' segments, and '//' (CWE-88 git ref-traversal), not just bad charset.
- NIT: document the mandatory invariants on gate-and-pr (required-reviewer
environment must stay; runs-on must stay GitHub-hosted).
- QUESTION (lockfiles): answered in-code — the build-test sandbox is
credential-less + egress-blocked, so lockfile-postinstall RCE is contained.
Tests added for all guards. 1042 tests, ruff clean, YAML valid.
* fix(agent-team): P3-flip — resolve /sh-security-review findings (LOGIC-1/2/3)
High-recall fan-out (injection/logic/iac+secrets) + proof-or-kill on the LIVE
apply path found 3 real issues the cross-review missed; all fixed:
- LOGIC-2 (HIGH, was a live hole): build-test ran `ruff check . || echo` /
`pytest -q || echo`, swallowing failures so the job was always 'success' and
the gate would open draft PRs on RED builds. ruff/pytest now run
authoritatively under set -e (pytest exit 5 'no tests' is the only non-fatal
case); the exit code IS the build-test conclusion the gate keys on.
- LOGIC-1 (verified!=shipped): the dispatcher used `git apply` + `git add -A`,
staging stray untracked content into the pushed PR head. Now `git apply
--index` stages exactly the diff, so the head tree is precisely base+diff —
bound to the bytes CI hash-verified.
- LOGIC-3 (§4.5 on the live path): gate-weakening was enforced only box-side;
added a gate-weakening check to the guard job so the live PR-opening path
rejects a diff that adds suppressions/skips, even on a green build.
Injection / secrets / least-privilege / flip-correctness / no-untrusted-checkout
all came back clean. 1044 tests, ruff clean, YAML valid.
2026-06-22 18:51:52 -04:00
|
|
|
|
|
|
|
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
# §4.5 gate-weakening detector
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _diff_adding(added_line: str, path: str = "src/mod.py") -> str:
|
|
|
|
|
"""A unified diff that ADDS one line (plus an unchanged context line)."""
|
|
|
|
|
return (
|
|
|
|
|
f"diff --git a/{path} b/{path}\n--- a/{path}\n+++ b/{path}\n"
|
|
|
|
|
f"@@ -1 +1,2 @@\n unchanged\n+{added_line}\n"
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
|
|
|
"added_line",
|
|
|
|
|
[
|
|
|
|
|
"x = 1 # noqa",
|
|
|
|
|
"y = 2 # noqa: E501",
|
|
|
|
|
"z = bad() # type: ignore",
|
|
|
|
|
"z = bad() # type:ignore[arg-type]",
|
|
|
|
|
"def f(): # pragma: no cover",
|
|
|
|
|
"def f(): # pragma: no-cover",
|
|
|
|
|
"subprocess.run(cmd) # nosec",
|
|
|
|
|
"eval(x) # nosemgrep",
|
|
|
|
|
" git commit --no-verify",
|
|
|
|
|
"@pytest.mark.skip(reason='flaky')",
|
|
|
|
|
"@pytest.mark.xfail",
|
|
|
|
|
" pytest.skip('todo')",
|
|
|
|
|
" self.skipTest('later')",
|
|
|
|
|
],
|
|
|
|
|
)
|
|
|
|
|
def test_gate_weakening_added_line_flagged(added_line: str) -> None:
|
|
|
|
|
violations = gate_weakening_violations(_diff_adding(added_line))
|
|
|
|
|
assert violations, f"expected {added_line!r} to be flagged as gate-weakening"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_gate_weakening_clean_diff_has_no_violations() -> None:
|
|
|
|
|
assert gate_weakening_violations(_diff_for("src/foo.py")) == []
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_gate_weakening_only_inspects_added_lines() -> None:
|
|
|
|
|
# A diff that REMOVES a suppression (line starts with '-') must not be flagged.
|
|
|
|
|
removing = (
|
|
|
|
|
"diff --git a/m.py b/m.py\n--- a/m.py\n+++ b/m.py\n"
|
|
|
|
|
"@@ -1,2 +1 @@\n-x = 1 # noqa\n unchanged\n"
|
|
|
|
|
)
|
|
|
|
|
assert gate_weakening_violations(removing) == []
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_gate_weakening_rejects_non_str() -> None:
|
|
|
|
|
with pytest.raises(CiGateError):
|
|
|
|
|
gate_weakening_violations(None) # type: ignore[arg-type]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_gate_blocks_weakening_diff_even_if_ci_success() -> None:
|
|
|
|
|
# A diff that adds a coverage suppression BLOCKs even with a green CI run.
|
|
|
|
|
diff = _diff_adding("def f(): # pragma: no cover")
|
|
|
|
|
result = evaluate_ci_gate(
|
|
|
|
|
candidate_diff=diff,
|
|
|
|
|
ledger_hash=_ledger_hash(diff),
|
|
|
|
|
ci_result=_good_ci("run-w", diff, "success"),
|
|
|
|
|
expected_run_id="run-w",
|
|
|
|
|
)
|
|
|
|
|
assert result.decision is GateDecision.BLOCK
|
|
|
|
|
assert result.blocked is True
|
|
|
|
|
assert any("gate-weakening" in r for r in result.reasons)
|