* feat(agent-team): P3-flip Phase 1 — expand denylist vectors (§4.2) + runner-trust assertion (§4.1) First controls of the P3-live-flip Phase-1 CI hardening (workflow stays INERT; this only tightens the trust boundary). Whole Phase-1 surface is gated by /sh-security-review + GPT-4.1 cross-review before any flip. §4.2 — expand the trust-control denylist with direct code-execution / supply-chain vectors, kept byte-identical across all three copies (ci_gate.DENYLIST_GLOBS + the guard + post-build inline DENY_GLOBS), drift-guarded: .gitmodules, .husky/**, .githooks/**, .gitattributes, .npmrc, and generated/build artifacts (__generated__, *.generated.*, dist/**, build/**, *.min.js). Deliberate: lockfiles are NOT wholesale denied — lockfile-postinstall RCE is already contained by the credential-less egress-blocked build sandbox, and the Tier-3 dep-CVE fixer rewrites lockfiles to produce its draft PRs; a blanket deny would make it un-shippable. Flagged in-code for the security gate. Direct code-execution config (hooks/filters/npmrc/submodules) is the actual §4.2 RCE surface. §4.1 — runner-trust: assert no job (esp. the privileged gate-and-pr) can run on a self-hosted/user-provided runner; all must be GitHub-hosted. 998 tests pass, ruff clean. * feat(agent-team): P3-flip Phase 1 — gate-weakening detector (§4.5) A diff that ADDS a lint/type/coverage/security suppression (noqa, type: ignore, pragma: no cover, nosec, nosemgrep), a test skip/xfail, or a hook bypass (--no-verify) could make CI pass falsely. The pure-code gate now flags these via gate_weakening_violations() and BLOCKs in evaluate_ci_gate as a top-priority trust violation (step 1b, alongside the denylist) — regardless of the authenticated CI conclusion. A build cannot pass itself by disabling its own checks; flagged diffs escalate to a human. Only ADDED lines are inspected (removing a suppression is fine). 1015 tests pass, ruff clean. * feat(agent-team): P3-flip — diff transport (§4.3) + flip privileged apply path live Completes the box->CI diff handoff and flips the apply/verify privileged job live (gated behind the agent-apply environment's required reviewer). Transport (§4.3): the read-only box (D2) emits a diff but holds no write token. - New credential-less `materialize` job decodes the untrusted `diff_b64` dispatch input via env (CWE-94), fail-closed re-hashes it against `expected_diff_hash`, and uploads it as the named artifact so guard/build-test download it same-run. guard now `needs: materialize`. - New `dispatcher.py` (the trusted apply path, operator/Mac-side — never the box): pushes the diff as a head branch then `gh workflow run`s the workflow. Pure input-assembly (sha256 == sha256sum, b64 round-trip, head ref) is unit-tested; git/gh are injected seams. Push-before-dispatch; fail-closed on empty diff/scope, unsafe task_id/owner/repo. Flip: gate-and-pr binds `environment: agent-apply` (required reviewer amoussa1229) + grants exactly `pull-requests: write`; the App-token + draft-PR steps run only on `steps.gate.outputs.gate == 'pass'` (no more if:false); the draft PR opens with an explicit `--head`; task_id/head_branch charset-validated (§4.6). Updated the hardening tests from inert-state to live-state assertions + added transport tests. 1039 tests, ruff clean, workflow YAML valid. NOTE: workflow only runs on manual workflow_dispatch and the privileged job is held at the required-reviewer gate, so nothing privileged runs unapproved. * fix(agent-team): P3-flip — address GPT-4.1 cross-review (size bound, ref-traversal guard) - BLOCK: cap candidate diff at 40 KB in the dispatcher (the diff rides a base64 workflow_dispatch input; GitHub caps inputs at ~64 KB so an oversized diff cannot dispatch at all) + a defense-in-depth decoded-size bound in materialize. - FIX: harden the draft-PR HEAD_BRANCH guard to reject leading/trailing slash, '..' segments, and '//' (CWE-88 git ref-traversal), not just bad charset. - NIT: document the mandatory invariants on gate-and-pr (required-reviewer environment must stay; runs-on must stay GitHub-hosted). - QUESTION (lockfiles): answered in-code — the build-test sandbox is credential-less + egress-blocked, so lockfile-postinstall RCE is contained. Tests added for all guards. 1042 tests, ruff clean, YAML valid. * fix(agent-team): P3-flip — resolve /sh-security-review findings (LOGIC-1/2/3) High-recall fan-out (injection/logic/iac+secrets) + proof-or-kill on the LIVE apply path found 3 real issues the cross-review missed; all fixed: - LOGIC-2 (HIGH, was a live hole): build-test ran `ruff check . || echo` / `pytest -q || echo`, swallowing failures so the job was always 'success' and the gate would open draft PRs on RED builds. ruff/pytest now run authoritatively under set -e (pytest exit 5 'no tests' is the only non-fatal case); the exit code IS the build-test conclusion the gate keys on. - LOGIC-1 (verified!=shipped): the dispatcher used `git apply` + `git add -A`, staging stray untracked content into the pushed PR head. Now `git apply --index` stages exactly the diff, so the head tree is precisely base+diff — bound to the bytes CI hash-verified. - LOGIC-3 (§4.5 on the live path): gate-weakening was enforced only box-side; added a gate-weakening check to the guard job so the live PR-opening path rejects a diff that adds suppressions/skips, even on a green build. Injection / secrets / least-privilege / flip-correctness / no-untrusted-checkout all came back clean. 1044 tests, ruff clean, YAML valid.
179 lines
6 KiB
Python
179 lines
6 KiB
Python
"""Unit tests for agent_team.dispatcher — the trusted apply-path transport (§4.3).
|
|
|
|
Fully hermetic: the branch-push and workflow-dispatch side effects are injected
|
|
fakes, so no git, no ``gh``, and no network are exercised. The tests pin the
|
|
pure input-assembly + validation contract and the push-before-dispatch order.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import base64
|
|
|
|
import pytest
|
|
|
|
from agent_team.dispatcher import (
|
|
MAX_DIFF_BYTES,
|
|
DispatcherError,
|
|
DispatchInputs,
|
|
build_dispatch_inputs,
|
|
dispatch_apply_verify,
|
|
head_branch_for,
|
|
)
|
|
from agent_team.state_store import compute_content_hash
|
|
|
|
TASK = "0a1b2c3d4e5f6071"
|
|
DIFF = "diff --git a/README.md b/README.md\n--- a/README.md\n+++ b/README.md\n@@ -1 +1,2 @@\n title\n+added line\n"
|
|
SCOPE = "README.md\ndocs/**"
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
# build_dispatch_inputs (pure)
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
def test_build_dispatch_inputs_binds_hash_and_b64() -> None:
|
|
di = build_dispatch_inputs(task_id=TASK, diff_text=DIFF, declared_scope=SCOPE)
|
|
# expected_diff_hash is the plain sha256 the workflow's sha256sum reproduces.
|
|
assert di.expected_diff_hash == compute_content_hash(DIFF.encode("utf-8"))
|
|
# diff_b64 round-trips back to the exact diff bytes.
|
|
assert base64.b64decode(di.diff_b64).decode("utf-8") == DIFF
|
|
assert di.head_branch == f"agent-team/apply/{TASK}"
|
|
assert di.diff_artifact_name == f"agent-team-diff-{TASK}"
|
|
assert di.declared_scope == SCOPE
|
|
|
|
|
|
def test_as_inputs_keys_match_the_six_workflow_inputs() -> None:
|
|
di = build_dispatch_inputs(task_id=TASK, diff_text=DIFF, declared_scope=SCOPE)
|
|
assert set(di.as_inputs()) == {
|
|
"task_id",
|
|
"diff_artifact_name",
|
|
"expected_diff_hash",
|
|
"declared_scope",
|
|
"diff_b64",
|
|
"head_branch",
|
|
}
|
|
|
|
|
|
@pytest.mark.parametrize("bad_diff", ["", " ", "\n\n"])
|
|
def test_empty_diff_rejected(bad_diff: str) -> None:
|
|
with pytest.raises(DispatcherError):
|
|
build_dispatch_inputs(task_id=TASK, diff_text=bad_diff, declared_scope=SCOPE)
|
|
|
|
|
|
def test_oversized_diff_rejected() -> None:
|
|
# The diff rides a base64 workflow_dispatch input (GitHub ~64 KB cap); a diff
|
|
# over MAX_DIFF_BYTES must fail closed in the dispatcher, not be dispatched.
|
|
big = "diff --git a/x b/x\n" + "+" + ("x" * (MAX_DIFF_BYTES + 1)) + "\n"
|
|
with pytest.raises(DispatcherError):
|
|
build_dispatch_inputs(task_id=TASK, diff_text=big, declared_scope=SCOPE)
|
|
|
|
|
|
@pytest.mark.parametrize("bad_scope", ["", " "])
|
|
def test_empty_scope_rejected(bad_scope: str) -> None:
|
|
# An empty declared scope would let a diff touch ANY path — fail closed.
|
|
with pytest.raises(DispatcherError):
|
|
build_dispatch_inputs(task_id=TASK, diff_text=DIFF, declared_scope=bad_scope)
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"bad_task", ["", "has space", "semi;colon", "../escape", "a/b", "x" * 201]
|
|
)
|
|
def test_unsafe_task_id_rejected(bad_task: str) -> None:
|
|
with pytest.raises(DispatcherError):
|
|
head_branch_for(bad_task)
|
|
with pytest.raises(DispatcherError):
|
|
build_dispatch_inputs(task_id=bad_task, diff_text=DIFF, declared_scope=SCOPE)
|
|
|
|
|
|
# --------------------------------------------------------------------------- #
|
|
# dispatch_apply_verify (injected seams)
|
|
# --------------------------------------------------------------------------- #
|
|
|
|
|
|
class _Recorder:
|
|
def __init__(self) -> None:
|
|
self.calls: list[dict] = []
|
|
|
|
|
|
def test_dispatch_pushes_then_fires_with_correct_inputs() -> None:
|
|
order: list[str] = []
|
|
pushed = _Recorder()
|
|
fired = _Recorder()
|
|
|
|
def pusher(*, owner, repo, base, head_branch, diff_text):
|
|
order.append("push")
|
|
pushed.calls.append(
|
|
{
|
|
"owner": owner,
|
|
"repo": repo,
|
|
"base": base,
|
|
"head": head_branch,
|
|
"diff": diff_text,
|
|
}
|
|
)
|
|
|
|
def dispatcher(*, owner, repo, inputs, ref):
|
|
order.append("dispatch")
|
|
fired.calls.append({"owner": owner, "repo": repo, "inputs": inputs, "ref": ref})
|
|
|
|
di = dispatch_apply_verify(
|
|
owner="Sea-Haven-Industries",
|
|
repo="orchestrator",
|
|
task_id=TASK,
|
|
diff_text=DIFF,
|
|
declared_scope=SCOPE,
|
|
pusher=pusher,
|
|
dispatcher=dispatcher,
|
|
)
|
|
|
|
assert isinstance(di, DispatchInputs)
|
|
# Branch is pushed BEFORE the workflow is dispatched (the draft-PR step opens
|
|
# against an already-pushed --head).
|
|
assert order == ["push", "dispatch"]
|
|
assert pushed.calls[0]["head"] == f"agent-team/apply/{TASK}"
|
|
assert pushed.calls[0]["diff"] == DIFF
|
|
# The dispatch carries all six inputs, including the b64 diff + head branch.
|
|
inputs = fired.calls[0]["inputs"]
|
|
assert inputs["head_branch"] == f"agent-team/apply/{TASK}"
|
|
assert base64.b64decode(inputs["diff_b64"]).decode("utf-8") == DIFF
|
|
assert inputs["expected_diff_hash"] == compute_content_hash(DIFF.encode("utf-8"))
|
|
assert fired.calls[0]["ref"] == "main"
|
|
|
|
|
|
def test_dispatch_does_not_fire_if_push_fails() -> None:
|
|
fired = _Recorder()
|
|
|
|
def failing_pusher(**_kw):
|
|
raise RuntimeError("push failed")
|
|
|
|
def dispatcher(**kw):
|
|
fired.calls.append(kw)
|
|
|
|
with pytest.raises(RuntimeError):
|
|
dispatch_apply_verify(
|
|
owner="o",
|
|
repo="r",
|
|
task_id=TASK,
|
|
diff_text=DIFF,
|
|
declared_scope=SCOPE,
|
|
pusher=failing_pusher,
|
|
dispatcher=dispatcher,
|
|
)
|
|
# A failed push must NOT dispatch a run (no orphan run against a missing head).
|
|
assert fired.calls == []
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"owner,repo", [("", "r"), ("o", ""), ("bad owner", "r"), ("o", "r/x")]
|
|
)
|
|
def test_unsafe_owner_repo_rejected(owner: str, repo: str) -> None:
|
|
with pytest.raises(DispatcherError):
|
|
dispatch_apply_verify(
|
|
owner=owner,
|
|
repo=repo,
|
|
task_id=TASK,
|
|
diff_text=DIFF,
|
|
declared_scope=SCOPE,
|
|
pusher=lambda **_k: None,
|
|
dispatcher=lambda **_k: None,
|
|
)
|