diff --git a/security-review/DEPLOY-R720.md b/security-review/DEPLOY-R720.md index 24fde9e..4f1259b 100644 --- a/security-review/DEPLOY-R720.md +++ b/security-review/DEPLOY-R720.md @@ -1,54 +1,144 @@ -# Phase 3 — Path B deployment (R720 host + CI backstop) +# Phase 3 — Path B deployment (R720 VM) -Status: **design + scripts; NOT deployed.** Needs the R720 (hostname SH-NVR, deprecated NVR, no GPU) -and Adam's hands-on involvement. Nothing here has been run against the box. See memory +Status: **host built; headless runner built + validated; two-tier auto-discovery nightly sweep built.** +Pending: provision the read-only `GH_TOKEN` and run one live VM dry-run to validate the clone-mirror path +end-to-end. **CI was removed by design** — the git hooks + this nightly sweep are the backstop. See memory `project-security-review-agent`. -## What Phase 3 delivers -1. The orchestrator hosted on the R720 as an always-on service (the Path B execution engine). -2. A **headless agentic runner** (`run_headless.py`) so the detector fan-out + proof-or-kill verifier - can run unattended (interactive `/sh-security-review` can't run in CI/cron). -3. CI backstop (`ci/security-review.yml`) running the same `review.sh` on every PR. -4. Nightly full-repo sweep + planted canary vulns + two-model disagreement on criticals. -5. Budget/telemetry ceiling against the Max-20x **$200/mo Agent SDK credit pool** - (see memory `reference-claude-subscription-billing`). +Path B is the unattended backstop that shares one pure-code gate (`review.sh`) with the interactive Path A +(`/sh-security-review`). This file is the operator runbook for the box that runs it. -## The load-bearing piece to BUILD first: `run_headless.py` -The interactive skill orchestrates subagents via the Claude Code Agent tool (Max-covered). Headless, -that must become explicit API/SDK calls. `run_headless.py` should: -- take `--scope` + target dir, read the in-scope files; -- run the 6 detectors (injection/authz/secrets-crypto/iac-iam/web-client/logic) as parallel calls, - each fresh context, using the SAME prompts as `~/.claude/commands/sh-security-review.md`; -- run the proof-or-kill verifier pass; -- emit the finding schema JSON that `review.sh --agent-findings` consumes. -- **Billing:** authenticate via the **Claude Agent SDK with the subscription** to spend the $200/mo - pool before API overflow (NOT a raw key) — see the billing memory. Non-Claude cross-family tiebreak - stays on Bedrock/API. ⚠️ This code cannot be validated until it runs with real creds; do not mark - done until tested. +## Host -## R720 host setup (Windows Server 2022, no GPU) -Recommended: run the Python orchestrator in **WSL2 or Docker** rather than native Windows Python. -1. `git clone https://github.com/amoussa1229/orchestrator` onto the box. -2. Python 3.12 + `pip install -r requirements.txt` (+ the Agent SDK). -3. Install the deterministic scanners (semgrep/gitleaks/checkov/cfn-lint/pip-audit) on the box. -4. Secrets on the box (NOT in git): Agent SDK subscription auth, Bedrock creds, Slack token, repo PATs. -5. Run as a service (NSSM on Windows, or systemd inside WSL2). Expose to the harness as an - MCP/HTTP service (mirror the existing unifi MCP pattern) OR keep it as a local scheduled job. -6. Scheduler: nightly full-repo sweep → `review.sh` → Slack report (ALARM-style, see - memory `feedback_cloudwatch_alarms`: only notify on findings, not clean runs). +- **Hypervisor:** R720 at `10.10.60.40` (Windows Server 2022, Hyper-V role). +- **VM:** `sh-secrev`, always-on Ubuntu 24.04 (kernel 6.8), Gen2, 4GB / 2 vCPU / 40GB dynamic vhdx. +- **Reach it:** `ssh -i ~/.ssh/r720_seahaven adam@10.10.60.120` (key-only, NOPASSWD sudo). -## Anti-complacency reinforcements (Phase 3) -- **Canary vulns:** keep the testbed's planted vulns in a fixture the nightly sweep also scans; if the - agent ever misses a known canary, that's a complacency signal → alert. -- **Two-model disagreement:** on confirmed criticals, run the cross-family reviewer (orchestrator - GPT-4.1) and flag disagreement for human review. -- **Budget guard:** track output tokens vs the $200/mo pool; stop/alert before overflow. +Operate on the VM, not from the Mac against the host by hand. -## Open dependencies (Adam) -- R720 reachable + WSL2/Docker chosen. -- GitHub repo/org secrets: `ANTHROPIC_API_KEY` (or SDK auth), set `vars.SECURITY_REVIEW_AGENTIC=1` - once `run_headless.py` is tested. -- Decide: promote `ci/security-review.yml` into Sea-Haven-Industries/.github as a reusable workflow - (per engineering-handbook/cicd.md) vs per-repo copy. -- checkov severity filtering (it emitted 51-55 best-practice mediums on payments-dashboard — tune - before nightly or the Slack reports are noise). +## What is installed on the VM + +- **Deterministic scanners:** semgrep, gitleaks, checkov, pip-audit, cfn-lint. **Node 18** (`npm audit`). +- **`claude` CLI** (Node) — the subscription-auth path for Path B. +- **Python 3.12 venv** at `~/orchestrator/.venv` with `claude-agent-sdk`. +- **Repo:** `~/orchestrator/` (rsync from the Mac, `.env` excluded — NOT a git clone). After editing the + sweep locally, re-sync: `rsync -av --exclude .env --exclude .venv ~/Documents/repositories/orchestrator/ adam@10.10.60.120:orchestrator/`. +- **Testbed corpus:** `~/security-review-testbed` (also rsync'd; includes the Node + .NET fixtures). +- **No `gh` CLI required** — discovery uses the GitHub REST API via `curl`. `run_headless.py` is + self-contained (detector/verifier prompts are inline), so the VM needs no `~/.claude` assets to run. + +## Auth, billing, and the read-only GitHub token + +### Claude (subscription OAuth) +- Token from `claude setup-token`, stored in `~/secrev.env` as `CLAUDE_CODE_OAUTH_TOKEN` (mode 600, NOT in git). +- The 2026-06-15 SDK-billing split was **deferred**, so automated SDK usage draws from the Max 20x + subscription's normal usage limits — the same pool as interactive Claude Code. The two-tier sweep below + is what keeps that draw bounded. See memory `reference-claude-subscription-billing`. +- **CRITICAL:** a raw `ANTHROPIC_API_KEY` would silently win and meter to API rates — it must NOT be set on + this host. `run_headless.py` pops it defensively and refuses to run without `CLAUDE_CODE_OAUTH_TOKEN`. + +### GitHub (`GH_TOKEN`, read-only — REQUIRED for auto-discovery) +The nightly sweep enumerates and clones org repos with a **fine-grained, read-only PAT**. Never give this +always-on box a write-capable token. + +1. github.com → Settings → Developer settings → **Fine-grained personal access tokens** → Generate new. +2. **Resource owner:** Sea-Haven-Industries. **Repository access:** All repositories. +3. **Permissions:** Repository → **Contents: Read-only**, **Metadata: Read-only** (auto). Nothing else. +4. Set an expiry (e.g. 90 days; calendar a rotation). Generate and copy the `github_pat_...` value. +5. On the VM, append it to `~/secrev.env` and lock the file down: + ``` + echo 'GH_TOKEN=github_pat_xxxxxxxx' >> ~/secrev.env && chmod 600 ~/secrev.env + ``` +6. Verify (should print repo names, not a 401): + ``` + set -a; . ~/secrev.env; set +a + curl -fsS -H "Authorization: Bearer $GH_TOKEN" \ + "https://api.github.com/orgs/Sea-Haven-Industries/repos?per_page=3" | jq '.[].full_name' + ``` + +### Non-Claude provider keys +The GPT-4.1 critical tiebreak (optional) uses keys in `~/orchestrator/.env` (mode 600, gitignored, +auto-loaded by `run.py`). They bill to their own provider accounts — keep them out of `~/secrev.env`. + +## The headless runner: `run_headless.py` + +Runs the 6 fresh-context detectors + proof-or-kill verifier unattended via the Agent SDK; emits the +finding-schema JSON that `review.sh --agent-findings` consumes. Read-only tools, hermetic +(`setting_sources=[]`), fails toward over-reporting. CLI: + +``` +CLAUDE_CODE_OAUTH_TOKEN=... python3 run_headless.py TARGET_DIR \ + [--scope "src infra web"] [--out findings.json] [--model claude-...] \ + [--detectors injection,authz,...] [--concurrency 3] [--max-turns 40] \ + [--detector-budget-usd 2.0] [--total-budget-usd 12.0] +``` +When the total budget is exhausted the verifier is skipped and remaining candidates stay `unverified` — +never silently dropped. Manual single-repo run: +``` +cd ~/orchestrator +set -a; . ~/secrev.env; set +a +.venv/bin/python security-review/run_headless.py ~/security-review-testbed --out /tmp/agent.json +security-review/review.sh --agent-findings /tmp/agent.json ~/security-review-testbed +``` + +## Nightly two-tier, clean-clone auto-discovery sweep + +`nightly_sweep.sh` needs **no per-repo wiring**. Each night it: + +1. **Discovers** every non-archived Sea-Haven-Industries repo via the REST API (`curl` + `GH_TOKEN`) and + **mirrors** each as a shallow clean clone (`git clone --depth=1`, default branch from the API + `default_branch`) into `~/repo-mirrors`. The token is injected only for the fetch and scrubbed from the + on-disk remote afterward. Clean clones contain no developer-local gitignored `.env`, so live secrets + stay out of scope by construction. +2. **Canary first:** scans `~/security-review-testbed` agentically (anti-complacency) — must block and meet + the recall floor, else COMPLACENCY ALARM. +3. **Tier 1 (every repo, $0 Claude):** `review.sh --scanners-only` over every mirror. +4. **Tier 2 (bounded agentic):** `run_headless.py` over a deterministic round-robin rotation that fits + `TOTAL_BUDGET_USD`, with a persistent cycle pointer (`~/sweep-reports/.rotation-state.json`) so every + repo gets a deep pass within `MAX_CYCLE_NIGHTS`; a COVERAGE ALARM fires if it falls behind. + +ALARM-only (a clean night posts nothing). Secret-shaped values are redacted from the Slack string; reports +under `~/sweep-reports//` are mode 600. + +### Config (env / systemd `Environment=`) +`GH_ORG` (Sea-Haven-Industries) · `MIRROR_DIR` (~/repo-mirrors) · `TOTAL_BUDGET_USD` (20) · +`PER_TARGET_BUDGET_USD` (12) · `CANARY_FLOOR` (10) · `MAX_CYCLE_NIGHTS` (4) · `MAX_AGENTIC_PER_NIGHT` +(0 = unlimited) · `CENTRAL_SKIP_FILE` (~/.secrev-skip.txt) · `ENABLE_XMODEL_HOOK` (0) · +`TARGETS` (manual override — scan explicit paths, no discovery). + +### Skip a repo +Commit a `.security-review-skip` at its root, **or** add its name to `~/.secrev-skip.txt`. Marker-skips are +logged in the report (a sensitive repo cannot silently self-exclude). + +### Manual dry-run (do this once after provisioning `GH_TOKEN`) +``` +cd ~/orchestrator +set -a; . ~/secrev.env; set +a +./security-review/nightly_sweep.sh +# Watch: discovery count, mirrors, canary block+recall, tier1 over all repos, tier2 rotation, clean exit. +# Then re-tune CANARY_FLOOR to the reported recall, and confirm the ALARM path with a forced failure. +``` + +### Install the timer +``` +sudo cp security-review/systemd/sea-haven-secrev.{service,timer} /etc/systemd/system/ +sudo systemctl daemon-reload +sudo systemctl enable --now sea-haven-secrev.timer # the timer drives it; do not enable the .service +systemctl list-timers sea-haven-secrev.timer +``` +Fires nightly ~02:00 local (`Persistent=true` catches missed runs). `TimeoutStartSec=21600` (6h) bounds a +hang without killing a healthy long night; spend is capped by `TOTAL_BUDGET_USD`. + +## Anti-complacency reinforcements +- **Canary:** the testbed (now Python/IaC/React + Node + .NET planted vulns) is scanned every night; a + recall drop or non-block is a COMPLACENCY ALARM. +- **Coverage:** the rotation pointer + `MAX_CYCLE_NIGHTS` guarantee every repo gets a deep pass on a cadence, + with a COVERAGE ALARM if it slips — no silent incomplete coverage. +- **Two-model disagreement (optional):** `ENABLE_XMODEL_HOOK=1` re-checks confirmed criticals with GPT-4.1. + +## Remaining (deferred by design) +- **Persistent budget/telemetry ledger:** cross-run spend tracking beyond the per-run + nightly caps (optional). +- **Phase 4 roster growth** (compliance/drift sweep, CVE agent, optional auto-fixer) — only per a real job. +- **Phase 5 remediation:** harden findings as real repos surface them (payments-dashboard first). +- **Confluence:** document `sh-secrev` as standing infrastructure (always-on VM holding a read-only org PAT, + pulling all org repos nightly) in the IT host/LAN inventory. diff --git a/security-review/README.md b/security-review/README.md index b542167..dc23536 100644 --- a/security-review/README.md +++ b/security-review/README.md @@ -1,32 +1,108 @@ # security-review -The Sea Haven security-review gate. One pure-code script (`review.sh`), many triggers. +The Sea Haven security-review gate. One pure-code script (`review.sh`) is the decision-maker; everything +else (hooks, the interactive skill, the headless runner, the nightly sweep) is a trigger that feeds it. See memory `project-security-review-agent` for the full design. ## Pieces -- `review.sh` — merges deterministic-scanner findings + agent findings, dedups, applies - suppressions (justification required), and makes the **block decision** (no agent decides). Exit 1 = BLOCK. -- `hooks/pre-commit` + `install-hooks.sh` — fast scanners-only hook for a target repo. -- The agentic detector/verifier pass is the interactive `/sh-security-review` slash command - (`~/.claude/commands/sh-security-review.md`), schema at `~/.claude/security-review/finding.schema.json`. +- `review.sh` — merges deterministic-scanner findings + agent findings, dedups, applies suppressions + (justification required), and makes the **block decision** (no agent decides). Exit 1 = BLOCK. +- `hooks/pre-commit`, `hooks/pre-push` + `install-hooks.sh` — fast `--scanners-only` gates. Install once + globally for every repo, or per-repo (see below). +- `skill/sh-security-review.md` — the interactive agentic detector/verifier prompt (Path A, Max-covered). + `finding.schema.json` — the structured finding contract both paths emit. `install-hooks.sh --global` + links these into `~/.claude/` (this repo is the source of truth). +- `run_headless.py` — the Path B headless detector fan-out + proof-or-kill verifier (Claude Agent SDK, + subscription OAuth). Self-contained: prompts are inline, so the VM needs no `~/.claude` assets to run it. +- `nightly_sweep.sh` + `systemd/` — the unattended two-tier sweep on the `sh-secrev` VM (R720). ## Triggers (one script, many entry points) - **On-demand (primary):** run `/sh-security-review` in a Claude Code session (Max-covered), have it - write its schema JSON, then `review.sh --agent-findings out.json ` to gate. -- **Pre-commit:** `install-hooks.sh ` — fast deterministic scanners abort the commit early. -- **CI (Phase 3):** the same `review.sh` runs headless as the unbypassable backstop. + write its schema JSON, then `review.sh --agent-findings out.json ` to gate. Required before + pushing payments/auth/IaC/input-handling changes (see the global CLAUDE.md security-review rule). +- **Pre-commit / pre-push:** the global git hooks run deterministic scanners automatically. +- **Nightly:** the VM sweep (Path B) is the unattended backstop. + +### Installing the hooks +``` +# Global — gate EVERY repo on this machine, and link the skill + schema into ~/.claude: +security-review/install-hooks.sh --global + +# Per-repo — for a repo that sets its own core.hooksPath (e.g. husky) and would shadow the global hook: +security-review/install-hooks.sh /path/to/repo +``` +The global mode sets `git config --global core.hooksPath ~/.config/git/hooks`. Skip a repo with a +`.security-review-skip` file at its root; suppress a specific false positive in the repo's +`.security-review/suppressions.json` (a written justification is required and is surfaced); bypass once +with `git push --no-verify`. Caveat: a repo with its own local `core.hooksPath` overrides the global hook +— install per-repo there. See memory `reference_global_security_review_hook`. + +## No CI — by design +There is **no CI** wiring for this gate. For a solo dev the git hooks + nightly VM sweep are the backstop, +so the parked CI drafts (`ci/*.yml`, `CI-BACKSTOP-NOTES.md`) were removed; recover them from git history +(the commit that deleted `security-review/ci/`) if the team ever goes multi-dev. The orchestrator repo's +own `ci.yaml` (ruff + tests) is unrelated and stays. ## Scanners -`review.sh` runs whatever is installed and logs the rest with install commands (no silent skips). -Currently wired: `cfn-lint`. To give the deterministic layer teeth, install: +`review.sh` runs whatever is installed and logs the rest with install commands (no silent skips): +`semgrep` (`p/security-audit` + `p/secrets` + `p/javascript`), `gitleaks` (git-mode — scans committed +history, respects `.gitignore`), `checkov`, `cfn-lint`, `pip-audit`, `npm audit`. Each is normalized into +the finding schema. Install the full set: ``` -pipx install semgrep # SAST: injection / authz / xss -brew install gitleaks # hardcoded secrets -pipx install pip-audit # vulnerable Python deps -pipx install checkov # IaC / IAM misconfig +pipx install semgrep pip-audit checkov # SAST / vulnerable Python deps / IaC misconfig +brew install gitleaks # hardcoded secrets +# cfn-lint via pip; Node.js provides npm audit ``` -Each needs a small normalizer added to `review.sh` (map its JSON to the finding schema) when installed. -## Status -review.sh gate validated against the local testbed: 16 findings, correct dedup of distinct -same-CWE findings, suppression-without-justification rejected and surfaced, BLOCK on confirmed crit/high. +## Nightly sweep (Path B) — two-tier, clean-clone auto-discovery +`nightly_sweep.sh` runs on the `sh-secrev` Ubuntu VM (R720) and needs **no per-repo wiring**. It: + +1. **Discovers** every non-archived Sea-Haven-Industries repo via the GitHub REST API (`curl` + a + read-only `GH_TOKEN`; no `gh` CLI dependency) and **mirrors** each as a shallow clean clone + (`git clone --depth=1`, default branch from the API) into `~/repo-mirrors`. Scanning server-side + clones — not developer working trees — structurally keeps local gitignored `.env` secrets out of scope. +2. **Tier 1 (every repo, every night, $0 Claude):** `review.sh --scanners-only` over every mirror — + complete deterministic baseline coverage. +3. **Tier 2 (bounded agentic):** `run_headless.py` over a deterministic **round-robin rotation** that + fits `TOTAL_BUDGET_USD`, with a persistent cycle pointer so every repo gets a deep pass within + `MAX_CYCLE_NIGHTS`. This bounds the draw on the shared Max limits (a clean night never deep-scans all + repos). A `COVERAGE ALARM` fires if the rotation falls behind. + +It is **ALARM-only**: a clean night posts nothing. See memory `feedback_cloudwatch_alarms`. + +### Skip / override +- A repo is skipped if it commits a `.security-review-skip` marker **or** is listed in the central skip + file (`~/.secrev-skip.txt`, one repo name per line). Repos skipped via their own committed marker are + **logged in the report** so a sensitive repo can't silently self-exclude. +- `TARGETS="/path/a /path/b"` overrides discovery entirely (scan explicit paths, no cloning). +- The canary corpus (`~/security-review-testbed`) is **always** scanned agentically first as the + anti-complacency check — independent of the skip filter. + +### Anti-complacency + guards +- **Canary check:** the testbed MUST block AND surface ≥ `CANARY_FLOOR` (default 10) confirmed crit/high. + Otherwise → COMPLACENCY ALARM. The corpus now includes Node + .NET fixtures (see the testbed key); + re-tune the floor after the first VM canary run reports the expanded recall number. +- **Budget ceiling:** `TOTAL_BUDGET_USD` (default 20) caps aggregate agentic spend; `PER_TARGET_BUDGET_USD` + (default 12) caps each repo; `MAX_AGENTIC_PER_NIGHT` (default 0 = unlimited) optionally caps wall-clock. + With the SDK-billing split deferred (memory `reference-claude-subscription-billing`), spend draws from + the Max subscription limits, so the two-tier design keeps full coverage cheap and bounds the agentic draw. +- **Two-model hook (optional, off):** `ENABLE_XMODEL_HOOK=1` re-checks each confirmed CRITICAL with the + orchestrator cross-family reviewer (GPT-4.1) and flags disagreement; skips gracefully, never fails the sweep. +- **Redaction:** secret-shaped values are masked in the Slack ALARM string; on-disk reports are mode 600. + +### Secrets / env (`~/secrev.env`, mode 600) +- `CLAUDE_CODE_OAUTH_TOKEN` — required (`run_headless.py` pops `ANTHROPIC_API_KEY`). +- `GH_TOKEN` — **read-only fine-grained PAT** scoped to the org (Contents + Metadata: read-only, nothing + else) for discovery + cloning. Never give this unattended box a write-capable token. +- `SLACK_WEBHOOK_URL` — alarms (plain incoming-webhook). `~/orchestrator/.env` → `OPENAI_API_KEY` (xmodel hook only). +- Reports + per-target JSON land under `~/sweep-reports//`. + +### Install the timer +Units are in `systemd/`; full runbook is `DEPLOY-R720.md`. On the VM: +``` +sudo cp systemd/sea-haven-secrev.{service,timer} /etc/systemd/system/ +sudo systemctl daemon-reload +sudo systemctl enable --now sea-haven-secrev.timer # the timer drives it; do not enable the .service +sudo systemctl start sea-haven-secrev.service # optional one-off smoke test +``` +The timer fires nightly at ~02:00 local (`Persistent=true` catches missed runs after downtime).