feat: re-home cross-family reviewer CLI, retire Path B host artifacts (#5)

* feat: add router-less cross-family reviewer CLI (cross_review.py)

Re-homes the archived orchestrator repo's GPT-4.1 cross_reviewer as a
direct OpenAI SDK CLI: verbatim system prompt, same model id, ported
3-attempt exponential-backoff retry. Reads OPENAI_API_KEY from the
environment or the gitignored repo-root .env. Lazy openai import so
--help works without the package.

* feat: create machine-level suppressions dir on --global hook install

install-hooks.sh --global now mkdir -p's
${SH_SECURITY_SUPPRESSIONS_DIR:-~/.config/sea-haven/security-review} and
states the convention: machine-level <dir>/<repo-basename>/suppressions.json
is preferred over repo-local .security-review/suppressions.json, with
review.sh merging both when run without --suppressions.

* docs: describe sh-build-review as cross-family GPT-4.1 pass via cross_review.py

* chore: retire Path B VM host artifacts, repoint xmodel hook to cross_review.py

- delete systemd/ units and DEPLOY-R720.md (VM destroyed; recoverable
  from git history)
- nightly_sweep.sh: retained for reference — header notes the VM-based
  Path B sweep is retired; ENABLE_XMODEL_HOOK now calls this repo's
  cross_review.py instead of the archived orchestrator's run.py
- sweep-targets.txt: drop the stale ~/orchestrator warning block
- README: document the Claude Code web cloud routines
  (repo-scanner-nightly-sweep 08:00 ET, repo-checkers-plane1 07:30 ET,
  ALARM-only to #repo-scanner) and add the cross_review.py section

* docs(iam): repoint cross-family review invocations to cross_review.py

* fix(ci): exclude canary corpus from ruff, add import smoke test

CI has been red repo-wide: the latest ruff wants to reformat the
intentionally-vulnerable canary fixtures (whose line numbers are keyed
in canary-meta/KEY.md), and pytest --collect-only exits 5 with zero
tests. Exclude canary/ via ruff.toml and add a root-level import smoke
test so collection is non-empty and imports cross_review.py.
This commit is contained in:
Adam Moussa 2026-07-14 19:24:01 -04:00 • committed by GitHub
parent 5408649943
commit 58aa07d8be
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
16 changed files with 226 additions and 371 deletions

View file

@ -1,144 +0,0 @@
# Phase 3 — Path B deployment (R720 VM)
Status: **host built; headless runner built + validated; two-tier auto-discovery nightly sweep built.**
Pending: provision the read-only `GH_TOKEN` and run one live VM dry-run to validate the clone-mirror path
end-to-end. **CI was removed by design** — the git hooks + this nightly sweep are the backstop. See memory
`project-security-review-agent`.
Path B is the unattended backstop that shares one pure-code gate (`review.sh`) with the interactive Path A
(`/sh-security-review`). This file is the operator runbook for the box that runs it.
## Host
- **Hypervisor:** R720 at `10.10.60.40` (Windows Server 2022, Hyper-V role).
- **VM:** `sh-secrev`, always-on Ubuntu 24.04 (kernel 6.8), Gen2, 4GB / 2 vCPU / 40GB dynamic vhdx.
- **Reach it:** `ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120` (key-only, NOPASSWD sudo).
Operate on the VM, not from the Mac against the host by hand.
## What is installed on the VM
- **Deterministic scanners:** semgrep, gitleaks, checkov, pip-audit, cfn-lint. **Node 18** (`npm audit`).
- **`claude` CLI** (Node) — the subscription-auth path for Path B.
- **Python 3.12 venv** at `~/orchestrator/.venv` with `claude-agent-sdk`.
- **Repo:** `~/orchestrator/` (rsync from the Mac, `.env` excluded — NOT a git clone). After editing the
sweep locally, re-sync: `rsync -av --exclude .env --exclude .venv ~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/`.
- **Testbed corpus:** `~/security-review-testbed` (also rsync'd; includes the Node + .NET fixtures).
- **No `gh` CLI required** — discovery uses the GitHub REST API via `curl`. `run_headless.py` is
self-contained (detector/verifier prompts are inline), so the VM needs no `~/.claude` assets to run.
## Auth, billing, and the read-only GitHub token
### Claude (subscription OAuth)
- Token from `claude setup-token`, stored in `~/secrev.env` as `CLAUDE_CODE_OAUTH_TOKEN` (mode 600, NOT in git).
- The 2026-06-15 SDK-billing split was **deferred**, so automated SDK usage draws from the Max 20x
subscription's normal usage limits — the same pool as interactive Claude Code. The two-tier sweep below
is what keeps that draw bounded. See memory `reference-claude-subscription-billing`.
- **CRITICAL:** a raw `ANTHROPIC_API_KEY` would silently win and meter to API rates — it must NOT be set on
this host. `run_headless.py` pops it defensively and refuses to run without `CLAUDE_CODE_OAUTH_TOKEN`.
### GitHub (`GH_TOKEN`, read-only — REQUIRED for auto-discovery)
The nightly sweep enumerates and clones org repos with a **fine-grained, read-only PAT**. Never give this
always-on box a write-capable token.
1. github.com → Settings → Developer settings → **Fine-grained personal access tokens** → Generate new.
2. **Resource owner:** Sea-Haven-Industries. **Repository access:** All repositories.
3. **Permissions:** Repository → **Contents: Read-only**, **Metadata: Read-only** (auto). Nothing else.
4. Set an expiry (e.g. 90 days; calendar a rotation). Generate and copy the `github_pat_...` value.
5. On the VM, append it to `~/secrev.env` and lock the file down:
```
echo 'GH_TOKEN=github_pat_xxxxxxxx' >> ~/secrev.env && chmod 600 ~/secrev.env
```
6. Verify (should print repo names, not a 401):
```
set -a; . ~/secrev.env; set +a
curl -fsS -H "Authorization: Bearer $GH_TOKEN" \
"https://api.github.com/orgs/Sea-Haven-Industries/repos?per_page=3" | jq '.[].full_name'
```
### Non-Claude provider keys
The GPT-4.1 critical tiebreak (optional) uses keys in `~/orchestrator/.env` (mode 600, gitignored,
auto-loaded by `run.py`). They bill to their own provider accounts — keep them out of `~/secrev.env`.
## The headless runner: `run_headless.py`
Runs the 6 fresh-context detectors + proof-or-kill verifier unattended via the Agent SDK; emits the
finding-schema JSON that `review.sh --agent-findings` consumes. Read-only tools, hermetic
(`setting_sources=[]`), fails toward over-reporting. CLI:
```
CLAUDE_CODE_OAUTH_TOKEN=... python3 run_headless.py TARGET_DIR \
[--scope "src infra web"] [--out findings.json] [--model claude-...] \
[--detectors injection,authz,...] [--concurrency 3] [--max-turns 40] \
[--detector-budget-usd 2.0] [--total-budget-usd 12.0]
```
When the total budget is exhausted the verifier is skipped and remaining candidates stay `unverified` —
never silently dropped. Manual single-repo run:
```
cd ~/orchestrator
set -a; . ~/secrev.env; set +a
.venv/bin/python security-review/run_headless.py ~/security-review-testbed --out /tmp/agent.json
security-review/review.sh --agent-findings /tmp/agent.json ~/security-review-testbed
```
## Nightly two-tier, clean-clone auto-discovery sweep
`nightly_sweep.sh` needs **no per-repo wiring**. Each night it:
1. **Discovers** every non-archived Sea-Haven-Industries repo via the REST API (`curl` + `GH_TOKEN`) and
**mirrors** each as a shallow clean clone (`git clone --depth=1`, default branch from the API
`default_branch`) into `~/repo-mirrors`. The token is injected only for the fetch and scrubbed from the
on-disk remote afterward. Clean clones contain no developer-local gitignored `.env`, so live secrets
stay out of scope by construction.
2. **Canary first:** scans `~/security-review-testbed` agentically (anti-complacency) — must block and meet
the recall floor, else COMPLACENCY ALARM.
3. **Tier 1 (every repo, $0 Claude):** `review.sh --scanners-only` over every mirror.
4. **Tier 2 (bounded agentic):** `run_headless.py` over a deterministic round-robin rotation that fits
`TOTAL_BUDGET_USD`, with a persistent cycle pointer (`~/sweep-reports/.rotation-state.json`) so every
repo gets a deep pass within `MAX_CYCLE_NIGHTS`; a COVERAGE ALARM fires if it falls behind.
ALARM-only (a clean night posts nothing). Secret-shaped values are redacted from the Slack string; reports
under `~/sweep-reports/<UTC-date>/` are mode 600.
### Config (env / systemd `Environment=`)
`GH_ORG` (Sea-Haven-Industries) · `MIRROR_DIR` (~/repo-mirrors) · `TOTAL_BUDGET_USD` (120) ·
`PER_TARGET_BUDGET_USD` (12) · `CANARY_FLOOR` (10) · `MAX_CYCLE_NIGHTS` (4) · `MAX_AGENTIC_PER_NIGHT`
(0 = unlimited) · `CENTRAL_SKIP_FILE` (~/.secrev-skip.txt) · `ENABLE_XMODEL_HOOK` (0) ·
`TARGETS` (manual override — scan explicit paths, no discovery).
### Skip a repo
Commit a `.security-review-skip` at its root, **or** add its name to `~/.secrev-skip.txt`. Marker-skips are
logged in the report (a sensitive repo cannot silently self-exclude).
### Manual dry-run (do this once after provisioning `GH_TOKEN`)
```
cd ~/orchestrator
set -a; . ~/secrev.env; set +a
./security-review/nightly_sweep.sh
# Watch: discovery count, mirrors, canary block+recall, tier1 over all repos, tier2 rotation, clean exit.
# Then re-tune CANARY_FLOOR to the reported recall, and confirm the ALARM path with a forced failure.
```
### Install the timer
```
sudo cp security-review/systemd/sea-haven-secrev.{service,timer} /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now sea-haven-secrev.timer # the timer drives it; do not enable the .service
systemctl list-timers sea-haven-secrev.timer
```
Fires nightly ~02:00 local (`Persistent=true` catches missed runs). `TimeoutStartSec=21600` (6h) bounds a
hang without killing a healthy long night; spend is capped by `TOTAL_BUDGET_USD`.
## Anti-complacency reinforcements
- **Canary:** the testbed (now Python/IaC/React + Node + .NET planted vulns) is scanned every night; a
recall drop or non-block is a COMPLACENCY ALARM.
- **Coverage:** the rotation pointer + `MAX_CYCLE_NIGHTS` guarantee every repo gets a deep pass on a cadence,
with a COVERAGE ALARM if it slips — no silent incomplete coverage.
- **Two-model disagreement (optional):** `ENABLE_XMODEL_HOOK=1` re-checks confirmed criticals with GPT-4.1.
## Remaining (deferred by design)
- **Persistent budget/telemetry ledger:** cross-run spend tracking beyond the per-run + nightly caps (optional).
- **Phase 4 roster growth** (compliance/drift sweep, CVE agent, optional auto-fixer) — only per a real job.
- **Phase 5 remediation:** harden findings as real repos surface them (payments-dashboard first).
- **Confluence:** document `sh-secrev` as standing infrastructure (always-on VM holding a read-only org PAT,
pulling all org repos nightly) in the IT host/LAN inventory.

View file

@ -16,16 +16,18 @@ See memory `project-security-review-agent` for the full design.
- `skill/sh-security-review.md` — the interactive agentic detector/verifier prompt (Path A, Max-covered).
`finding.schema.json` — the structured finding contract both paths emit. `install-hooks.sh --global`
links these into `~/.claude/` (this repo is the source of truth).
- `run_headless.py` — the Path B headless detector fan-out + proof-or-kill verifier (Claude Agent SDK,
subscription OAuth). Self-contained: prompts are inline, so the VM needs no `~/.claude` assets to run it.
- `nightly_sweep.sh` + `systemd/` — the unattended two-tier sweep on the `sh-secrev` VM (R720).
- `run_headless.py` — the headless detector fan-out + proof-or-kill verifier (Claude Agent SDK,
subscription OAuth). Self-contained: prompts are inline, so the host needs no `~/.claude` assets to run it.
- `cross_review.py` — the mandatory cross-family GPT-4.1 reviewer CLI (see its section below).
- `nightly_sweep.sh` — the retired VM-based two-tier sweep, retained for reference; the automated
sweep now runs as Claude Code web cloud routines (see below).
## Triggers (one script, many entry points)
- **On-demand (primary):** run `/sh-security-review` in a Claude Code session (Max-covered), have it
write its schema JSON, then `review.sh --agent-findings out.json <repo>` to gate. Required before
pushing payments/auth/IaC/input-handling changes (see the global CLAUDE.md security-review rule).
- **Pre-commit / pre-push:** the global git hooks run deterministic scanners automatically.
- **Nightly:** the VM sweep (Path B) is the unattended backstop.
- **Scheduled:** the Claude Code web cloud routines are the unattended backstop (see below).
### Installing the hooks
```
@ -40,6 +42,11 @@ The global mode sets `git config --global core.hooksPath ~/.config/git/hooks`. S
its own local `core.hooksPath` overrides the global hook — install per-repo there. See memory
`reference_global_security_review_hook`.
- `--global` also creates the machine-level suppressions dir
(`${SH_SECURITY_SUPPRESSIONS_DIR:-~/.config/sea-haven/security-review}`): the per-repo file
`<dir>/<repo-basename>/suppressions.json` is preferred by the hooks over repo-local
`.security-review/suppressions.json`, and `review.sh` merges both when run without `--suppressions`.
**Suppressing a false positive.** A written justification is required and is surfaced in the report.
When no `--suppressions FILE` is passed, `review.sh` auto-resolves and **merges** suppressions from
two locations (an explicit `--suppressions` still overrides both):
@ -50,10 +57,10 @@ two locations (an explicit `--suppressions` still overrides both):
2. **Repo-local:** `<repo>/.security-review/suppressions.json` (tracked; travels with the repo).
Both files' `.suppressions[]` are concatenated (machine-level first, so it wins any id collision).
This means every entry point — the pre-push hook, the nightly sweep, on-demand/CI runs, and the
This means every entry point — the pre-push hook, the scheduled sweeps, on-demand/CI runs, and the
Open SWE daily-report automation — resolves suppressions identically. **Portability note:** hosts that
cannot see the Mac's `~/.config` (the Open SWE automation, the R720 VM) get only the tracked repo-local
file, so a suppression that must be honored off-Mac has to live repo-local. Fail-safe: if a file is
cannot see the Mac's `~/.config` (the Open SWE automation, the cloud routines) get only the tracked
repo-local file, so a suppression that must be honored off-Mac has to live repo-local. Fail-safe: if a file is
present but unparseable, nothing is suppressed (the gate blocks).
> **⚠️ Trust model — automated scanners must target TRUSTED repos only.** The repo-local
@ -87,63 +94,33 @@ brew install gitleaks # hardcoded secrets
# cfn-lint via pip; Node.js provides npm audit
```
## Scheduled execution — migrating from the VM to Claude Code web routines
The unattended runs are moving off the `sh-secrev` VM into **Claude Code web scheduled routines** (which
post ALARM-only to Slack `#repo-scanner`): one routine for the agentic two-tier sweep, and a second that
runs the deterministic `checker_coordinator.sh` (the script owns the findings + ALARM decision; the
routine relays its output verbatim, never re-judging). `nightly_sweep.sh` / `checker_coordinator.sh` and
the `systemd/` units below remain the source of truth and the VM-deployment path; the VM timers are being
retired once the routines are validated.
## Scheduled execution — Claude Code web cloud routines
The automated sweep runs as **Claude Code web cloud routines**, both **ALARM-only** to Slack
`#repo-scanner` (a clean run posts nothing — see memory `feedback_cloudwatch_alarms`):
## Nightly sweep (Path B) — two-tier, clean-clone auto-discovery
`nightly_sweep.sh` runs on the `sh-secrev` Ubuntu VM (R720) and needs **no per-repo wiring**. It:
- **repo-scanner-nightly-sweep** — daily 08:00 ET: the agentic two-tier sweep (deterministic
scanners over every repo, plus the bounded agentic detector/verifier rotation).
- **repo-checkers-plane1** — daily 07:30 ET: the deterministic `checker_coordinator.sh` (the script
owns the findings + ALARM decision; the routine relays its output verbatim, never re-judging).
1. **Discovers** every non-archived Sea-Haven-Industries repo via the GitHub REST API (`curl` + a
read-only `GH_TOKEN`; no `gh` CLI dependency) and **mirrors** each as a shallow clean clone
(`git clone --depth=1`, default branch from the API) into `~/repo-mirrors`. Scanning server-side
clones — not developer working trees — structurally keeps local gitignored `.env` secrets out of scope.
2. **Tier 1 (every repo, every night, $0 Claude):** `review.sh --scanners-only` over every mirror —
complete deterministic baseline coverage.
3. **Tier 2 (bounded agentic):** `run_headless.py` over a deterministic **round-robin rotation** that
fits `TOTAL_BUDGET_USD`, with a persistent cycle pointer so every repo gets a deep pass within
`MAX_CYCLE_NIGHTS`. This bounds the draw on the shared Max limits (a clean night never deep-scans all
repos). A `COVERAGE ALARM` fires if the rotation falls behind.
The VM-based Path B sweep is **retired** (the sweep VM was destroyed); its host artifacts (the
`systemd/` units and the `DEPLOY-R720.md` runbook) are deleted — recover them from git history if
ever needed. `nightly_sweep.sh` is retained in-repo as the reference implementation of the two-tier
design: GitHub REST API auto-discovery into shallow clean clones, Tier 1 `review.sh --scanners-only`
over every mirror, Tier 2 bounded agentic `run_headless.py` round-robin rotation, canary
anti-complacency floor, budget ceilings, and secret redaction. Its optional `ENABLE_XMODEL_HOOK=1`
critical tiebreak now calls this repo's `cross_review.py` (default off). Target scoping stays
trusted-repos-only — see `sweep-targets.txt`.
It is **ALARM-only**: a clean night posts nothing. See memory `feedback_cloudwatch_alarms`.
## cross_review.py — mandatory cross-family reviewer
The GPT-4.1 cross-family reviewer CLI, re-homed here from the archived orchestrator repo as a
router-less direct OpenAI SDK call. It is **mandatory** for IAM/policy and Lambda-handler-signature
changes (per the global instructions), and is the reasoning backend for the `sh-plan-review`,
`sh-security-audit`, and `sh-build-review` gates.
### Skip / override
- A repo is skipped if it commits a `.security-review-skip` marker **or** is listed in the central skip
file (`~/.secrev-skip.txt`, one repo name per line). Repos skipped via their own committed marker are
**logged in the report** so a sensitive repo can't silently self-exclude.
- `TARGETS="/path/a /path/b"` overrides discovery entirely (scan explicit paths, no cloning).
- The canary corpus (`~/security-review-testbed`) is **always** scanned agentically first as the
anti-complacency check — independent of the skip filter.
### Anti-complacency + guards
- **Canary check:** the testbed MUST block AND surface ≥ `CANARY_FLOOR` (default 10) confirmed crit/high.
Otherwise → COMPLACENCY ALARM. The corpus now includes Node + .NET fixtures (see the testbed key);
re-tune the floor after the first VM canary run reports the expanded recall number.
- **Budget ceiling:** `TOTAL_BUDGET_USD` (default 120 — full deep-pass coverage of every repo per night) caps aggregate agentic spend; `PER_TARGET_BUDGET_USD`
(default 12) caps each repo; `MAX_AGENTIC_PER_NIGHT` (default 0 = unlimited) optionally caps wall-clock.
With the SDK-billing split deferred (memory `reference-claude-subscription-billing`), spend draws from
the Max subscription limits, so the two-tier design keeps full coverage cheap and bounds the agentic draw.
- **Two-model hook (optional, off):** `ENABLE_XMODEL_HOOK=1` re-checks each confirmed CRITICAL with the
orchestrator cross-family reviewer (GPT-4.1) and flags disagreement; skips gracefully, never fails the sweep.
- **Redaction:** secret-shaped values are masked in the Slack ALARM string; on-disk reports are mode 600.
### Secrets / env (`~/secrev.env`, mode 600)
- `CLAUDE_CODE_OAUTH_TOKEN` — required (`run_headless.py` pops `ANTHROPIC_API_KEY`).
- `GH_TOKEN` — **read-only fine-grained PAT** scoped to the org (Contents + Metadata: read-only, nothing
else) for discovery + cloning. Never give this unattended box a write-capable token.
- `SLACK_WEBHOOK_URL` — alarms (plain incoming-webhook). `~/orchestrator/.env` → `OPENAI_API_KEY` (xmodel hook only).
- Reports + per-target JSON land under `~/sweep-reports/<UTC-date>/`.
### Install the timer
Units are in `systemd/`; full runbook is `DEPLOY-R720.md`. On the VM:
```
sudo cp systemd/sea-haven-secrev.{service,timer} /etc/systemd/system/
sudo systemctl daemon-reload
sudo systemctl enable --now sea-haven-secrev.timer # the timer drives it; do not enable the .service
sudo systemctl start sea-haven-secrev.service # optional one-off smoke test
python3 ~/Documents/repositories/seahaven/security-review/cross_review.py "<task>"
```
The timer fires nightly at ~02:00 local (`Persistent=true` catches missed runs after downtime).
Setup: `OPENAI_API_KEY` in the environment or in the repo-root `.env` (gitignored — never commit
it); `pip install -r requirements.txt` provides the `openai` SDK.

139
cross_review.py Executable file
View file

@ -0,0 +1,139 @@
#!/usr/bin/env python3
"""cross_review.py — Sea Haven cross-family (GPT-4.1) review CLI.
Re-homed from the archived Sea-Haven-Industries/orchestrator repo's
`cross_reviewer` agent. Router-less: one direct OpenAI SDK call, no langchain,
no routing logic. This is the mandatory cross-family reviewer for IAM/policy
and Lambda-handler-signature changes, and the reasoning backend for the
sh-plan-review / sh-security-audit / sh-build-review gates.
Usage:
python3 cross_review.py "Review this diff for breaking changes: <diff>"
Auth: reads OPENAI_API_KEY from the environment, falling back to the
gitignored .env at the repo root. Never prints the key.
"""
import argparse
import os
import random
import sys
import time
# Model ID — single source of truth (ported from orchestrator/models.py).
DEFAULT_MODEL = "gpt-4.1"
TEMPERATURE = 0.2
MAX_ATTEMPTS = 3
# System prompt ported verbatim from the orchestrator's cross_reviewer agent
# (orchestrator/agents.py, CROSS_REVIEWER_PROMPT).
CROSS_REVIEWER_PROMPT = """You are a cross-family code review agent. You provide an independent review perspective.
Review the provided code for bugs, security issues, and improvements.
Categorize findings as BLOCK, FIX, or NIT. Be concise.
Focus on issues that might be missed by the primary development team."""
def load_api_key() -> str:
"""OPENAI_API_KEY from the environment, else hand-parsed from repo-root .env."""
key = os.environ.get("OPENAI_API_KEY")
if key:
return key
env_path = os.path.join(os.path.dirname(os.path.abspath(__file__)), ".env")
try:
with open(env_path, encoding="utf-8") as f:
for line in f:
line = line.strip()
if not line or line.startswith("#") or "=" not in line:
continue
name, _, value = line.partition("=")
if name.strip() == "OPENAI_API_KEY":
return value.strip().strip("'\"")
except OSError:
pass
return ""
def run_review(task: str, model: str, api_key: str) -> str:
"""One review call with retries on transient API errors (3 attempts,
exponential backoff with jitter — ported from orchestrator/models.py
with_retries)."""
# Lazy import so --help works without the openai package installed.
import openai
retriable = (
openai.APIConnectionError,
openai.RateLimitError,
openai.InternalServerError,
)
client = openai.OpenAI(api_key=api_key)
for attempt in range(1, MAX_ATTEMPTS + 1):
try:
response = client.chat.completions.create(
model=model,
temperature=TEMPERATURE,
messages=[
{"role": "system", "content": CROSS_REVIEWER_PROMPT},
{"role": "user", "content": task},
],
)
content = response.choices[0].message.content
if not content:
raise RuntimeError("model returned an empty review")
return content
except retriable as exc:
if attempt == MAX_ATTEMPTS:
raise
delay = (2 ** (attempt - 1)) + random.uniform(0, 1)
print(
f"cross_review: transient API error ({type(exc).__name__}), "
f"retrying in {delay:.1f}s (attempt {attempt}/{MAX_ATTEMPTS})",
file=sys.stderr,
)
time.sleep(delay)
raise RuntimeError("unreachable")
def main() -> int:
parser = argparse.ArgumentParser(
description=(
"Cross-family GPT-4.1 review (direct OpenAI SDK, no router). "
"Mandatory for IAM/policy and Lambda-handler-signature changes."
)
)
parser.add_argument("task", help="review task: a diff, change description, or plan")
parser.add_argument(
"--model",
default=DEFAULT_MODEL,
help=f"OpenAI model id (default: {DEFAULT_MODEL})",
)
args = parser.parse_args()
api_key = load_api_key()
if not api_key:
print(
"cross_review: OPENAI_API_KEY not set and not found in repo-root .env",
file=sys.stderr,
)
return 2
try:
review = run_review(args.task, args.model, api_key)
except ImportError:
print(
"cross_review: the 'openai' package is not installed "
"(pip install -r requirements.txt)",
file=sys.stderr,
)
return 2
except Exception as exc: # noqa: BLE001 — CLI boundary: report and exit nonzero
print(
f"cross_review: review failed: {type(exc).__name__}: {exc}", file=sys.stderr
)
return 1
print(review)
return 0
if __name__ == "__main__":
sys.exit(main())

View file

@ -177,6 +177,6 @@ snapshot-restore away from gone.
## Process note
Per global instructions this IAM change ALSO requires the GPT-4.1 cross-family review run via
`python3 ~/Documents/repositories/orchestrator/run.py "<diff/describe>"` (it is an IAM
`python3 ~/Documents/repositories/seahaven/security-review/cross_review.py "<task>"` (it is an IAM
role/policy + trust-anchor change). This packet is the input to that review; aws-posture is not
built until the review is recorded.

View file

@ -27,4 +27,5 @@ box**.
| `step-ca-config-sketch.md` | Internal CA config + systemd-timer auto-renewal of the short-lived leaf. |
Per global instructions this IAM change also requires the GPT-4.1 cross-family review via
`orchestrator/run.py`; this directory is that review's input.
`python3 ~/Documents/repositories/seahaven/security-review/cross_review.py "<task>"`; this
directory is that review's input.

View file

@ -18,6 +18,7 @@ set -euo pipefail
SRC="$(cd "$(dirname "$0")" && pwd)"
GLOBAL_HOOKS="$HOME/.config/git/hooks"
CLAUDE_DIR="$HOME/.claude"
SUP_DIR="${SH_SECURITY_SUPPRESSIONS_DIR:-$HOME/.config/sea-haven/security-review}"
usage() { grep '^#' "$0" | sed 's/^# \{0,1\}//'; }
@ -40,6 +41,9 @@ case "${1:-}" in
--global)
echo "== Installing Sea Haven security-review hooks globally =="
mkdir -p "$GLOBAL_HOOKS"
mkdir -p "$SUP_DIR"
echo " suppressions: machine-level per-repo file $SUP_DIR/<repo-basename>/suppressions.json is preferred by the hooks over repo-local .security-review/suppressions.json"
echo " suppressions: review.sh merges both locations when run without --suppressions"
# Warn (don't clobber silently) if a different global hooksPath is already set.
CURRENT="$(git config --global --get core.hooksPath || true)"

View file

@ -1,5 +1,9 @@
#!/usr/bin/env bash
# nightly_sweep.sh — Sea Haven Path B nightly security sweep (R720 / sh-secrev VM).
# nightly_sweep.sh — Sea Haven nightly security sweep (RETIRED — retained for reference).
#
# The VM-based Path B sweep is retired (the sweep VM was destroyed); the automated
# sweep now runs as Claude Code web cloud routines (see README.md). This script is
# kept as the reference implementation of the two-tier sweep design.
#
# TWO-TIER, CLEAN-CLONE AUTO-DISCOVERY (no per-repo wiring):
# Discovery: enumerate ALL Sea-Haven-Industries org repos via the GitHub REST API
@ -27,7 +31,7 @@
#
# Contract notes:
# - run_headless.py REQUIRES CLAUDE_CODE_OAUTH_TOKEN and pops ANTHROPIC_API_KEY.
# Source ~/secrev.env before invoking (the systemd unit does this via EnvironmentFile).
# Source the sweep env file before invoking.
# - review.sh re-derives the block decision (exit 1 = BLOCK). This script makes NO
# block decision itself; it only reports.
#
@ -48,9 +52,9 @@
# (default: 0 = unlimited, bounded only by TOTAL_BUDGET_USD)
# REPORT_ROOT base dir for logs+JSON (default: ~/sweep-reports)
# SLACK_WEBHOOK_URL incoming-webhook URL; if unset, alarms are logged only
# ENABLE_XMODEL_HOOK 1 to run the cross-family critical tiebreak (default: 0)
# ORCHESTRATOR_DIR orchestrator repo root (default: ~/orchestrator)
# VENV_PY python in the SDK venv (default: ~/orchestrator/.venv/bin/python)
# ENABLE_XMODEL_HOOK 1 to run the cross-family critical tiebreak via this repo's
# cross_review.py (default: 0)
# VENV_PY python interpreter (default: python3 on PATH)
#
# Exit: 0 = sweep completed (whether or not it alarmed); 2 = setup/usage error.
set -euo pipefail
@ -60,16 +64,16 @@ log() { echo "[nightly_sweep] $*" >&2; }
die() { echo "[nightly_sweep] FATAL: $*" >&2; exit 2; }
# --- Shared substrate (discovery / mirror / budget / rotation / Slack / canary) -
# Factored out so secrev and the R720 agent-team reuse one implementation, WITHOUT
# changing any secrev behavior. The functions close over this script's globals by name
# Factored out so the sweep and the checker coordinator reuse one implementation, WITHOUT
# changing any sweep behavior. The functions close over this script's globals by name
# (bash dynamic scoping); see lib/sweep_substrate.sh for the read/mutate contract.
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# shellcheck source=lib/sweep_substrate.sh
. "$HERE/lib/sweep_substrate.sh"
# --- Config + defaults --------------------------------------------------------
ORCHESTRATOR_DIR="${ORCHESTRATOR_DIR:-$HOME/orchestrator}"
VENV_PY="${VENV_PY:-$ORCHESTRATOR_DIR/.venv/bin/python}"
VENV_PY="${VENV_PY:-$(command -v python3)}"
CROSS_REVIEW="$HERE/cross_review.py"
RUN_HEADLESS="$HERE/run_headless.py"
REVIEW_SH="$HERE/review.sh"
GH_ORG="${GH_ORG:-Sea-Haven-Industries}"
@ -90,7 +94,7 @@ command -v git >/dev/null || die "git is required"
[ -x "$VENV_PY" ] || die "venv python not found/executable: $VENV_PY"
[ -f "$RUN_HEADLESS" ] || die "run_headless.py not found: $RUN_HEADLESS"
[ -x "$REVIEW_SH" ] || die "review.sh not found/executable: $REVIEW_SH"
[ -n "${CLAUDE_CODE_OAUTH_TOKEN:-}" ] || die "CLAUDE_CODE_OAUTH_TOKEN not set (source ~/secrev.env)"
[ -n "${CLAUDE_CODE_OAUTH_TOKEN:-}" ] || die "CLAUDE_CODE_OAUTH_TOKEN not set (source the sweep env file)"
UTC_DATE="$(date -u +%Y-%m-%d)"
UTC_STAMP="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
@ -129,8 +133,8 @@ skip_reason() { # name dir
xmodel_check_criticals() {
local label="$1" result_json="$2"
[ "$ENABLE_XMODEL_HOOK" = "1" ] || return 0
[ -n "${OPENAI_API_KEY:-}" ] || { log " xmodel hook: OPENAI_API_KEY unset — skipping"; return 0; }
"$VENV_PY" -c 'import langchain_openai' >/dev/null 2>&1 || { log " xmodel hook: deps missing — skipping"; return 0; }
[ -f "$CROSS_REVIEW" ] || { log " xmodel hook: cross_review.py missing — skipping"; return 0; }
"$VENV_PY" -c 'import openai' >/dev/null 2>&1 || { log " xmodel hook: openai package missing — skipping"; return 0; }
local crits n; crits="$(jq -c '[.findings[]? | select(.status=="confirmed" and .severity=="critical")]' "$result_json" 2>/dev/null || echo '[]')"
n="$(echo "$crits" | jq 'length')"; [ "${n:-0}" -gt 0 ] || return 0
log " xmodel hook: re-checking $n confirmed critical(s) for $label"
@ -138,11 +142,11 @@ xmodel_check_criticals() {
while [ "$i" -lt "$n" ]; do
local summary; summary="$(echo "$crits" | jq -r --argjson i "$i" '.[$i] | "\(.cwe // "n/a") \(.file):\(.line // 0) — \(.title // .id) :: \(.data_flow // "")"')"
local verdict
if verdict="$(cd "$ORCHESTRATOR_DIR" && "$VENV_PY" run.py "Independently assess whether this is a real exploitable vulnerability (yes/no) and why: $summary" 2>>"$REPORT_DIR/xmodel.log")"; then
if verdict="$("$VENV_PY" "$CROSS_REVIEW" "Independently assess whether this is a real exploitable vulnerability (yes/no) and why: $summary" 2>>"$REPORT_DIR/xmodel.log")"; then
if echo "$verdict" | grep -qiE '(^|[^a-z])no([^a-z]|$)|not (a |an )?(real |exploitable )?vuln'; then
XMODEL_LINES+=( "DISAGREEMENT on $label critical: $summary (cross_reviewer says NOT a vuln)" )
fi
else log " xmodel hook: run.py failed for a critical (logged) — continuing"; fi
else log " xmodel hook: cross_review.py failed for a critical (logged) — continuing"; fi
i=$((i+1))
done
}
@ -350,7 +354,7 @@ $ALARM_BODY
Coverage: tier1 scanners ${#SCANNABLE[@]} repos · tier2 agentic $AGENTIC_DONE this night ($REMAINING pending) · canary_ok=$CANARY_OK
Spend: \$$TOTAL_SPEND (ceiling \$$TOTAL_BUDGET_USD)
Reports + JSON: \`$REPORT_DIR\` (on sh-secrev VM)"
Reports + JSON: \`$REPORT_DIR\`"
SLACK_TEXT="$(echo "$SLACK_TEXT" | redact)"
log "ALARM conditions present — composing Slack post"

View file

@ -1,5 +1,6 @@
# Path B headless agentic runner (run_headless.py). Scanners (semgrep, gitleaks,
# Headless agentic runner (run_headless.py). Scanners (semgrep, gitleaks,
# checkov, cfn-lint, pip-audit) and npm are external binaries, installed separately
# (see README). The optional cross-model hook (ENABLE_XMODEL_HOOK=1) additionally
# needs langchain-openai, intentionally not pinned here since it is off by default.
# (see README).
claude-agent-sdk
# Cross-family reviewer CLI (cross_review.py) — direct OpenAI SDK, no langchain.
openai

4
ruff.toml Normal file
View file

@ -0,0 +1,4 @@
# canary/ is the intentionally-vulnerable scanner-recall corpus: it must stay
# byte-stable (canary-meta/KEY.md keys findings to its line numbers) and linting
# deliberately-bad code is meaningless, so ruff skips it entirely.
extend-exclude = ["canary"]

View file

@ -83,7 +83,7 @@ with file:line, the numbered data-flow trace, and the proof. List unverified and
nothing is silently dropped. The reader reviews proofs, not raw code.
## Relationship to existing tooling
- `sh-build-review` — single deep Fable pass over a change surface (depth, one reasoner).
- `sh-build-review` — single deep cross-family pass (GPT-4.1 via `cross_review.py` in this repo) over a change surface (depth, one out-of-family reasoner).
- `sh-security-review` — high-recall fan-out + proof-or-kill verifier (breadth + anti-complacency).
- IAM/policy and Lambda-signature changes still require the mandatory cross-family review per global instructions; this does not replace it.

View file

@ -19,8 +19,5 @@
# TODO (Phase 5): add the first real hardened repo here once it is cloned on the
# VM (candidates: payments-dashboard / proposal-system / procurement-ingest).
#
# Deliberately EMPTY by default = canary-only nights. Do NOT scan ~/orchestrator:
# it holds ~/orchestrator/.env with live provider API keys, which the agentic
# detector could surface into sweep reports / Slack. Only add repos with no
# Deliberately EMPTY by default = canary-only nights. Only add repos with no
# plaintext secrets (or scrub/exclude secret files first).
# ~/orchestrator

View file

@ -1,52 +0,0 @@
# sea-haven-checkers.service — Plane-1 nightly checker coordinator (sh-secrev VM, user adam).
#
# Runs security-review/checker_coordinator.sh: the read-only Plane-1 checkers
# (compliance-drift, dependency-cve, doc-drift, plan-groomer) under ONE shared
# budget ledger + versioned rotation/coverage, ALARM-only to Slack. Reuses the
# secrev sweep's substrate ($MIRROR_DIR clones, budget discipline) — no re-clone.
#
# Install (on the VM, as root):
# sudo cp sea-haven-checkers.service /etc/systemd/system/
# sudo cp sea-haven-checkers.timer /etc/systemd/system/
# sudo systemctl daemon-reload
# sudo systemctl enable --now sea-haven-checkers.timer # the timer drives it
# systemctl list-timers sea-haven-checkers.timer
#
# Secrets/config come from the EnvironmentFiles (leading '-' = optional):
# ~/secrev.env -> CLAUDE_CODE_OAUTH_TOKEN, GH_TOKEN, SLACK_WEBHOOK_URL
# ~/orchestrator/.env -> OPENAI/etc. (only if a checker shells the cross-model run.py)
#
# COORDINATOR_SKIP_ROLES excludes roles whose creds are NOT provisioned:
# - aws-posture needs IAM Roles Anywhere / step-ca (not provisioned)
# confluence-doc is ONLINE (2026-06-22): authenticates via the OAuth 2.0
# client-credentials service account in ~/secrev.env (CONFLUENCE_BASE_URL +
# CONFLUENCE_OAUTH_CLIENT_ID/_SECRET); PAGE_MAP_FILE points at the IT page-ID map
# generated from the live space. Remove a name from the skip list once its
# credential is provisioned to bring that checker online.
[Unit]
Description=Sea Haven agent-team Plane-1 nightly checker coordinator
After=network-online.target
Wants=network-online.target
[Service]
Type=oneshot
User=adam
WorkingDirectory=/home/adam/orchestrator/security-review
EnvironmentFile=-/home/adam/secrev.env
EnvironmentFile=-/home/adam/orchestrator/.env
Environment=GH_ORG=Sea-Haven-Industries
Environment=COORDINATOR_SKIP_ROLES=aws-posture
# IT page-ID map for confluence-doc (generated from the live space; regenerate
# periodically as pages change).
Environment=PAGE_MAP_FILE=/home/adam/confluence-page-map.json
# Tune the shared ceiling without editing the script (uncomment to override):
# Environment=TOTAL_BUDGET_USD=120
# Environment=MAX_CYCLE_NIGHTS=4
ExecStart=/home/adam/orchestrator/security-review/checker_coordinator.sh
# Bounded so a hung checker cannot run forever; spend is capped by TOTAL_BUDGET_USD.
TimeoutStartSec=10800
Nice=10
[Install]
WantedBy=multi-user.target

View file

@ -1,20 +0,0 @@
# sea-haven-checkers.timer — fires the Plane-1 checker coordinator nightly.
#
# 03:30 UTC — ~90 min after the sea-haven-secrev sweep (02:00) so the two do not
# contend on $MIRROR_DIR or the shared Claude subscription pool at the same instant.
# Persistent=true → if the VM was off, it runs at next boot. RandomizedDelaySec
# spreads load off an exact-minute spike.
#
# Install: see the header of sea-haven-checkers.service.
[Unit]
Description=Run the Sea Haven Plane-1 checker coordinator nightly (~03:30 UTC)
[Timer]
OnCalendar=*-*-* 03:30:00
Persistent=true
RandomizedDelaySec=600
Unit=sea-haven-checkers.service
[Install]
WantedBy=timers.target

View file

@ -1,49 +0,0 @@
# sea-haven-secrev.service — Path B nightly security sweep (sh-secrev VM, user adam).
#
# Install (on the VM, as root):
# sudo cp sea-haven-secrev.service /etc/systemd/system/
# sudo cp sea-haven-secrev.timer /etc/systemd/system/
# sudo systemctl daemon-reload
# sudo systemctl enable --now sea-haven-secrev.timer # timer drives the run; do NOT enable the .service
# systemctl list-timers sea-haven-secrev.timer # confirm next run
# sudo systemctl start sea-haven-secrev.service # optional: run once now to smoke-test
# journalctl -u sea-haven-secrev.service -e # logs (also under ~/sweep-reports/<date>/)
#
# Secrets come from the EnvironmentFiles (the leading '-' = optional, no failure if absent):
# ~/secrev.env -> CLAUDE_CODE_OAUTH_TOKEN (required by run_headless.py),
# GH_TOKEN (read-only fine-grained PAT — REQUIRED for org auto-discovery),
# SLACK_WEBHOOK_URL
# ~/orchestrator/.env -> OPENAI_API_KEY etc. (only needed if ENABLE_XMODEL_HOOK=1)
#
# GH_TOKEN must be a fine-grained PAT scoped to the Sea-Haven-Industries org with READ-ONLY
# Contents (and Metadata) permission — nothing else. It enumerates repos and clones them into
# ~/repo-mirrors. Never give this unattended box a write-capable token.
[Unit]
Description=Sea Haven Path B nightly security sweep
After=network-online.target
Wants=network-online.target
[Service]
Type=oneshot
User=adam
WorkingDirectory=/home/adam/orchestrator
EnvironmentFile=-/home/adam/secrev.env
EnvironmentFile=-/home/adam/orchestrator/.env
# Tune ceilings/targets here without editing the script (uncomment to override defaults):
# Environment=TOTAL_BUDGET_USD=120
# Environment=PER_TARGET_BUDGET_USD=12
# Environment=CANARY_FLOOR=10
# Environment=MAX_CYCLE_NIGHTS=4
# Environment=MAX_AGENTIC_PER_NIGHT=0
# Environment=GH_ORG=Sea-Haven-Industries
# Environment=MIRROR_DIR=/home/adam/repo-mirrors
# Environment=ENABLE_XMODEL_HOOK=0
ExecStart=/home/adam/orchestrator/security-review/nightly_sweep.sh
# Two-tier sweep (scanners over every repo + a budget-bounded agentic rotation) runs for hours;
# 6h ceiling bounds a hang without killing a healthy long night. Spend is capped by TOTAL_BUDGET_USD.
TimeoutStartSec=21600
Nice=10
[Install]
WantedBy=multi-user.target

View file

@ -1,20 +0,0 @@
# sea-haven-secrev.timer — fires the nightly sweep at ~02:00 local, user adam.
#
# Install: see the header of sea-haven-secrev.service. In short:
# sudo systemctl enable --now sea-haven-secrev.timer
# systemctl list-timers sea-haven-secrev.timer
#
# Persistent=true → if the VM was off at 02:00, the sweep runs at next boot.
# RandomizedDelaySec spreads load off an exact-minute spike.
[Unit]
Description=Run the Sea Haven Path B security sweep nightly (~02:00)
[Timer]
OnCalendar=*-*-* 02:00:00
Persistent=true
RandomizedDelaySec=600
Unit=sea-haven-secrev.service
[Install]
WantedBy=timers.target

13
test_imports.py Normal file
View file

@ -0,0 +1,13 @@
"""Import smoke test.
CI runs `pytest --collect-only` as an import check; collection imports this
module, which imports the repo's Python entry points. No API calls are made
(cross_review imports the openai SDK lazily, inside the call path).
"""
import cross_review
def test_cross_review_importable():
assert cross_review.DEFAULT_MODEL == "gpt-4.1"
assert "BLOCK" in cross_review.CROSS_REVIEWER_PROMPT