This repository has been archived on 2026-08-04. You can view files and clone it, but cannot push or open issues or pull requests.
orchestrator/security-review
Adam Moussa 3d97139300
feat(agent-team): P3-live CI apply/verify hardening + ci_fetcher (gate-passed, provisioning-gated) (#17)
* feat(agent-team): read-only CI-result fetcher for P3 verify gate (opt-in, inert)

ci_fetcher.py: fail-closed CiResultFetcher reading the GitHub Actions run
conclusion via a read-only PAT (AGENT_TEAM_CI_READ_TOKEN→GITHUB_TOKEN), returns
{run_id,conclusion,diff_hash} or None on any error. Data-fetcher only — ci_gate
owns the verdict; never writes, no OIDC/AWS, never reads patch artifacts.
coordinator gains opt-in gated_build_verify_wiring() composing it via
bind_ci_result_fetcher; NOT wired into the default run-team.py path. 20 tests.

* harden(agent-team): P3 apply/verify workflow — GitHub App token, CWE-94, fail-closed

Decision-1 auth model: gate-and-pr uses a GitHub App installation token
(pull-requests:write) behind the agent-apply environment; ALL OIDC/id-token/AWS
removed. Hardening: task_id env-indirection (CWE-94 — GitHub expands ${{ }} into
the run shell before exec, so %s/quoting is insufficient); run-id pinning on both
download-artifact; post-build denied-path check (build-hook writes into denied
paths fail the job); empty-hash fail-closed in BOTH the embedded gate (fixed a
real ''=='' pass bug) and ci_gate.py. App-token + draft-PR steps stay if:${{ false }}
until provisioning (App + environment + branch protection). +17 tests.

* harden(agent-team): apply P3-live security-gate fixes (GPT-4.1 xreview + sh-security-review)

BLOCK-1/FIX-4: gate-and-pr re-comments pull-requests:write + environment:agent-apply
(provisioning-time uncomment) and gains needs.guard/build-test=='success' job guard —
zero privilege until provisioning. BLOCK-2/3+FIX-5: ci_fetcher validates run_id (^[0-9]{1,20}$),
owner/repo (^[A-Za-z0-9_.-]{1,100}$), and fetched_id (int) — fail closed, no SSRF/path
injection. FIX-1: conclusion allowlist. FIX-3: api_root removed from public builder (no
injectable endpoint). INJ-02: post-build denied-path check uses NUL-delimited git output +
explicit rename parsing, no backslash mangling, non-UTF8=violation. INJ-03: all three trust-
control denylists unified to one 22-entry union + drift-guard test. Q1: documented run_id/
diff_hash trust source (dispatcher/ledger only). 884 tests, ruff clean. Privileged steps stay
if:${{ false }} until provisioning.

* build(security-review): prune .claude worktrees from deterministic scanners

Agent worktrees under .claude/worktrees/ are full repo copies; the cfn-lint
find|xargs template scan overflowed ('command line cannot be assembled') and the
pre-push hook fail-closed to BLOCK whenever a worktree was present. Prune .claude
in the cfn-lint find + semgrep/checkov excludes, and gitignore .claude/ so it is
never scanned or committed. Unblocks main-tree pushes during parallel agent work.
2026-06-18 15:53:26 -04:00
..
checkers feat(secrev): Plane-1 Phase 2 — coordinator + dependency-cve checker (#16) 2026-06-18 15:08:58 -04:00
hooks Add machine-level suppressions to the repo-sourced hooks 2026-06-16 15:33:43 -04:00
lib refactor(secrev): factor shared sweep substrate out of nightly_sweep.sh (Plane-1 Phase 0) 2026-06-18 14:06:31 -04:00
skill Make security-review hooks and skill installable from the repo 2026-06-16 15:00:00 -04:00
systemd Raise nightly agentic budget $20 -> $120 for full per-night coverage 2026-06-17 13:18:46 -04:00
checker_coordinator.sh feat(secrev): Plane-1 Phase 2 — coordinator + dependency-cve checker (#16) 2026-06-18 15:08:58 -04:00
DEPLOY-R720.md Raise nightly agentic budget $20 -> $120 for full per-night coverage 2026-06-17 13:18:46 -04:00
finding.schema.json Make security-review hooks and skill installable from the repo 2026-06-16 15:00:00 -04:00
install-hooks.sh Make security-review hooks and skill installable from the repo 2026-06-16 15:00:00 -04:00
nightly_sweep.sh refactor(secrev): factor shared sweep substrate out of nightly_sweep.sh (Plane-1 Phase 0) 2026-06-18 14:06:31 -04:00
README.md Raise nightly agentic budget $20 -> $120 for full per-night coverage 2026-06-17 13:18:46 -04:00
review.sh feat(agent-team): P3-live CI apply/verify hardening + ci_fetcher (gate-passed, provisioning-gated) (#17) 2026-06-18 15:53:26 -04:00
run_headless.py Add headless detector fan-out + proof-or-kill verifier runner 2026-06-16 15:00:10 -04:00
sweep-targets.txt Rewrite nightly sweep as two-tier clean-clone auto-discovery 2026-06-16 15:00:00 -04:00

security-review

The Sea Haven security-review gate. One pure-code script (review.sh) is the decision-maker; everything else (hooks, the interactive skill, the headless runner, the nightly sweep) is a trigger that feeds it. See memory project-security-review-agent for the full design.

Pieces

  • review.sh — merges deterministic-scanner findings + agent findings, dedups, applies suppressions (justification required), and makes the block decision (no agent decides). Exit 1 = BLOCK.
  • hooks/pre-commit, hooks/pre-push + install-hooks.sh — fast --scanners-only gates. Install once globally for every repo, or per-repo (see below).
  • skill/sh-security-review.md — the interactive agentic detector/verifier prompt (Path A, Max-covered). finding.schema.json — the structured finding contract both paths emit. install-hooks.sh --global links these into ~/.claude/ (this repo is the source of truth).
  • run_headless.py — the Path B headless detector fan-out + proof-or-kill verifier (Claude Agent SDK, subscription OAuth). Self-contained: prompts are inline, so the VM needs no ~/.claude assets to run it.
  • nightly_sweep.sh + systemd/ — the unattended two-tier sweep on the sh-secrev VM (R720).

Triggers (one script, many entry points)

  • On-demand (primary): run /sh-security-review in a Claude Code session (Max-covered), have it write its schema JSON, then review.sh --agent-findings out.json <repo> to gate. Required before pushing payments/auth/IaC/input-handling changes (see the global CLAUDE.md security-review rule).
  • Pre-commit / pre-push: the global git hooks run deterministic scanners automatically.
  • Nightly: the VM sweep (Path B) is the unattended backstop.

Installing the hooks

# Global — gate EVERY repo on this machine, and link the skill + schema into ~/.claude:
security-review/install-hooks.sh --global

# Per-repo — for a repo that sets its own core.hooksPath (e.g. husky) and would shadow the global hook:
security-review/install-hooks.sh /path/to/repo

The global mode sets git config --global core.hooksPath ~/.config/git/hooks. Skip a repo with a .security-review-skip file at its root; bypass once with git push --no-verify. Caveat: a repo with its own local core.hooksPath overrides the global hook — install per-repo there. See memory reference_global_security_review_hook.

Suppressing a false positive. A written justification is required and is surfaced in the report. The hooks resolve a suppressions file in this order:

  1. Machine-level (preferred), kept out of repo history: ${SH_SECURITY_SUPPRESSIONS_DIR:-~/.config/sea-haven/security-review}/<repo-basename>/suppressions.json (override the base dir with SH_SECURITY_SUPPRESSIONS_DIR). Keeps a suppression from becoming a permanent in-history "ignore."
  2. Repo-local fallback: <repo>/.security-review/suppressions.json (used only if no machine-level file exists).

Same JSON either place: {"suppressions":[{"id":"<review.sh finding id>","justification":"…"}]}. Caveat: machine-level files are keyed by repo basename, so two repos sharing a name collide — fine for the current single-namespace layout under ~/Documents/repositories.

No CI — by design

There is no CI wiring for this gate. For a solo dev the git hooks + nightly VM sweep are the backstop, so the parked CI drafts (ci/*.yml, CI-BACKSTOP-NOTES.md) were removed; recover them from git history (the commit that deleted security-review/ci/) if the team ever goes multi-dev. The orchestrator repo's own ci.yaml (ruff + tests) is unrelated and stays.

Scanners

review.sh runs whatever is installed and logs the rest with install commands (no silent skips): semgrep (p/security-audit + p/secrets + p/javascript), gitleaks (git-mode — scans committed history, respects .gitignore), checkov, cfn-lint, pip-audit, npm audit. Each is normalized into the finding schema. Install the full set:

pipx install semgrep pip-audit checkov   # SAST / vulnerable Python deps / IaC misconfig
brew install gitleaks                      # hardcoded secrets
# cfn-lint via pip; Node.js provides npm audit

Nightly sweep (Path B) — two-tier, clean-clone auto-discovery

nightly_sweep.sh runs on the sh-secrev Ubuntu VM (R720) and needs no per-repo wiring. It:

  1. Discovers every non-archived Sea-Haven-Industries repo via the GitHub REST API (curl + a read-only GH_TOKEN; no gh CLI dependency) and mirrors each as a shallow clean clone (git clone --depth=1, default branch from the API) into ~/repo-mirrors. Scanning server-side clones — not developer working trees — structurally keeps local gitignored .env secrets out of scope.
  2. Tier 1 (every repo, every night, $0 Claude): review.sh --scanners-only over every mirror — complete deterministic baseline coverage.
  3. Tier 2 (bounded agentic): run_headless.py over a deterministic round-robin rotation that fits TOTAL_BUDGET_USD, with a persistent cycle pointer so every repo gets a deep pass within MAX_CYCLE_NIGHTS. This bounds the draw on the shared Max limits (a clean night never deep-scans all repos). A COVERAGE ALARM fires if the rotation falls behind.

It is ALARM-only: a clean night posts nothing. See memory feedback_cloudwatch_alarms.

Skip / override

  • A repo is skipped if it commits a .security-review-skip marker or is listed in the central skip file (~/.secrev-skip.txt, one repo name per line). Repos skipped via their own committed marker are logged in the report so a sensitive repo can't silently self-exclude.
  • TARGETS="/path/a /path/b" overrides discovery entirely (scan explicit paths, no cloning).
  • The canary corpus (~/security-review-testbed) is always scanned agentically first as the anti-complacency check — independent of the skip filter.

Anti-complacency + guards

  • Canary check: the testbed MUST block AND surface ≥ CANARY_FLOOR (default 10) confirmed crit/high. Otherwise → COMPLACENCY ALARM. The corpus now includes Node + .NET fixtures (see the testbed key); re-tune the floor after the first VM canary run reports the expanded recall number.
  • Budget ceiling: TOTAL_BUDGET_USD (default 120 — full deep-pass coverage of every repo per night) caps aggregate agentic spend; PER_TARGET_BUDGET_USD (default 12) caps each repo; MAX_AGENTIC_PER_NIGHT (default 0 = unlimited) optionally caps wall-clock. With the SDK-billing split deferred (memory reference-claude-subscription-billing), spend draws from the Max subscription limits, so the two-tier design keeps full coverage cheap and bounds the agentic draw.
  • Two-model hook (optional, off): ENABLE_XMODEL_HOOK=1 re-checks each confirmed CRITICAL with the orchestrator cross-family reviewer (GPT-4.1) and flags disagreement; skips gracefully, never fails the sweep.
  • Redaction: secret-shaped values are masked in the Slack ALARM string; on-disk reports are mode 600.

Secrets / env (~/secrev.env, mode 600)

  • CLAUDE_CODE_OAUTH_TOKEN — required (run_headless.py pops ANTHROPIC_API_KEY).
  • GH_TOKEN — read-only fine-grained PAT scoped to the org (Contents + Metadata: read-only, nothing else) for discovery + cloning. Never give this unattended box a write-capable token.
  • SLACK_WEBHOOK_URL — alarms (plain incoming-webhook). ~/orchestrator/.env → OPENAI_API_KEY (xmodel hook only).
  • Reports + per-target JSON land under ~/sweep-reports/<UTC-date>/.

Install the timer

Units are in systemd/; full runbook is DEPLOY-R720.md. On the VM:

sudo cp systemd/sea-haven-secrev.{service,timer} /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now sea-haven-secrev.timer   # the timer drives it; do not enable the .service
sudo systemctl start sea-haven-secrev.service        # optional one-off smoke test

The timer fires nightly at ~02:00 local (Persistent=true catches missed runs after downtime).