mirror of
https://github.com/Sea-Haven-Industries/security-review.git
synced 2026-09-30 08:03:17 +00:00
* feat: add router-less cross-family reviewer CLI (cross_review.py)
Re-homes the archived orchestrator repo's GPT-4.1 cross_reviewer as a
direct OpenAI SDK CLI: verbatim system prompt, same model id, ported
3-attempt exponential-backoff retry. Reads OPENAI_API_KEY from the
environment or the gitignored repo-root .env. Lazy openai import so
--help works without the package.
* feat: create machine-level suppressions dir on --global hook install
install-hooks.sh --global now mkdir -p's
${SH_SECURITY_SUPPRESSIONS_DIR:-~/.config/sea-haven/security-review} and
states the convention: machine-level <dir>/<repo-basename>/suppressions.json
is preferred over repo-local .security-review/suppressions.json, with
review.sh merging both when run without --suppressions.
* docs: describe sh-build-review as cross-family GPT-4.1 pass via cross_review.py
* chore: retire Path B VM host artifacts, repoint xmodel hook to cross_review.py
- delete systemd/ units and DEPLOY-R720.md (VM destroyed; recoverable
from git history)
- nightly_sweep.sh: retained for reference — header notes the VM-based
Path B sweep is retired; ENABLE_XMODEL_HOOK now calls this repo's
cross_review.py instead of the archived orchestrator's run.py
- sweep-targets.txt: drop the stale ~/orchestrator warning block
- README: document the Claude Code web cloud routines
(repo-scanner-nightly-sweep 08:00 ET, repo-checkers-plane1 07:30 ET,
ALARM-only to #repo-scanner) and add the cross_review.py section
* docs(iam): repoint cross-family review invocations to cross_review.py
* fix(ci): exclude canary corpus from ruff, add import smoke test
CI has been red repo-wide: the latest ruff wants to reformat the
intentionally-vulnerable canary fixtures (whose line numbers are keyed
in canary-meta/KEY.md), and pytest --collect-only exits 5 with zero
tests. Exclude canary/ via ruff.toml and add a root-level import smoke
test so collection is non-empty and imports cross_review.py.
93 lines
6.2 KiB
Markdown
93 lines
6.2 KiB
Markdown
---
|
|
name: sh-security-review
|
|
description: High-recall agentic security review with anti-complacency structure. Opus fans out N narrow fresh-context detectors over the target, a separate fresh-context verifier demands proof-of-exploit or downgrades to unverified, then findings are emitted in the structured schema with a block decision. Returns severity-ranked findings each carrying a concrete proof.
|
|
---
|
|
|
|
# Sea Haven Security Review (detector fan-out + proof-or-kill verifier)
|
|
|
|
A high-recall security review built to resist reviewer complacency. Instead of one model judging the
|
|
whole surface (that's `sh-build-review`), this fans out **N narrow detectors, each fresh context and
|
|
adversarial**, then a **separate verifier** with the opposing incentive demands a concrete
|
|
proof-of-exploit for every candidate or downgrades it to `unverified`. Output is the structured
|
|
finding schema (`~/.claude/security-review/finding.schema.json`) plus a block decision.
|
|
|
|
This is the interactive (Path A) entry point and runs under the Max subscription. The same prompts and
|
|
schema are reused by the automated `review.sh` (Path B) later. See memory `project-security-review-agent`.
|
|
|
|
## When to Use
|
|
- Security review of a branch, working tree, file, or whole repo
|
|
- Before pushing payments/auth/IaC/input-handling changes
|
|
- As the high-recall pass; complements `/security-review`, `/code-review ultra`, and `sh-build-review`
|
|
|
|
## Arguments
|
|
- Optional target: a path, file, branch, or diff. Default: the current working tree / branch diff vs main.
|
|
- Optional `--scope <dirs>`: restrict the scan (e.g. `src/ infra/ web/`).
|
|
|
|
## Mechanism (Opus runs these)
|
|
|
|
Opus is the main loop. It scopes the target, runs the detector fan-out and verifier as subagents
|
|
(each with **fresh context** so no "we already passed 10 files" approval prior accumulates), then
|
|
applies the gate rule and reports. Opus does NOT soften the gate; the block condition is mechanical.
|
|
|
|
### 1. Scope
|
|
Identify the files in scope (respect `--scope`; never scan an answer key or `.git`). Note the languages
|
|
present (Python/Lambda, .NET, React/JS, SAM/CDK IaC) so detectors apply the right checklist.
|
|
|
|
### 2. Detector fan-out (parallel, fresh context, closed checklist, adversarial)
|
|
Spawn these detectors as **separate parallel subagents** (`subagent_type: general-purpose`). Each gets
|
|
ONLY its category, the schema, and the adversarial framing. Each must, per checklist item, either cite a
|
|
specific safe line OR file a finding — no open "looks fine" judgment.
|
|
|
|
Detectors:
|
|
- **injection** — SQL/command/template injection, unsafe deserialization, SSRF, path/file traversal, XXE
|
|
- **authz** — broken object-level auth/IDOR, missing access checks, missing webhook/Slack signature verification, auth bypass
|
|
- **secrets-crypto** — hardcoded secrets/keys, weak/again crypto (MD5/SHA1/unsalted), sensitive data in logs/errors/responses, wrong SSM-vs-Secrets-Manager
|
|
- **iac-iam** — wildcard IAM actions/resources, public buckets/endpoints, open security-group ingress (0.0.0.0/0), missing encryption, over-broad trust policies
|
|
- **web-client** — XSS (incl. dangerouslySetInnerHTML), CSRF, open redirect, client-side secret exposure
|
|
- **logic** — broken multi-step invariants, race conditions, missing tenant isolation, auth-state confusion
|
|
|
|
Detector prompt template (fill `{CATEGORY}`, `{CHECKLIST}`, `{SCOPE}`):
|
|
```
|
|
You are a hostile {CATEGORY} security auditor for Sea Haven. Assume this code is hostile and the author
|
|
missed something. Audit ONLY {CATEGORY} issues in: {SCOPE}. Read the actual files.
|
|
For each checklist item, either name a specific line that is safe, OR file a finding. Do not hand-wave.
|
|
Checklist: {CHECKLIST}
|
|
Return a JSON array of findings, each: {id, title, claimed_severity (critical|high|medium|low|info),
|
|
cwe (CWE-####), file, line, category:"{CATEGORY}", data_flow (numbered source->sink trace),
|
|
proof:{input, outcome, test|null}, recommendation}. If none, return []. No prose outside the JSON.
|
|
```
|
|
|
|
### 3. Verifier (separate subagent, fresh context, proof-or-kill)
|
|
Spawn one or more **verifier** subagents with the OPPOSING incentive. The verifier did not find these and
|
|
is rewarded for killing weak claims. For each candidate it demands a concrete, plausible proof.
|
|
```
|
|
You are a skeptical exploitation verifier. You did NOT find these; your job is to REFUTE weak claims.
|
|
For each candidate finding, decide: is there a concrete, plausible proof-of-exploit (a specific malicious
|
|
input and the specific bad outcome, consistent with the code)?
|
|
- If yes: status="confirmed", keep severity = claimed_severity, tighten the proof.
|
|
- If no / speculative / not reachable: status="unverified" and severity="unverified". Default to unverified
|
|
when uncertain. A confident assertion without a demonstrable input is NOT proof.
|
|
Return the findings array with status, severity, and the verified proof. No prose outside the JSON.
|
|
```
|
|
|
|
### 4. Merge + gate (mechanical, Opus does not soften)
|
|
- Dedup by (file, cwe, nearby line); keep the highest accepted severity.
|
|
- `summary.confirmed_critical` / `confirmed_high` = counts of status=confirmed at that severity.
|
|
- `summary.block = true` if any unsuppressed confirmed critical/high exists.
|
|
- A finding may be suppressed ONLY with a written `suppression_justification`, which is surfaced in the report.
|
|
No silent dismissal: a critical/high with no proof becomes `unverified` (still listed), never deleted.
|
|
|
|
### 5. Report
|
|
Emit the schema JSON, then a human summary: BLOCK/PASS, confirmed findings highest-severity-first, each
|
|
with file:line, the numbered data-flow trace, and the proof. List unverified and suppressed separately so
|
|
nothing is silently dropped. The reader reviews proofs, not raw code.
|
|
|
|
## Relationship to existing tooling
|
|
- `sh-build-review` — single deep cross-family pass (GPT-4.1 via `cross_review.py` in this repo) over a change surface (depth, one out-of-family reasoner).
|
|
- `sh-security-review` — high-recall fan-out + proof-or-kill verifier (breadth + anti-complacency).
|
|
- IAM/policy and Lambda-signature changes still require the mandatory cross-family review per global instructions; this does not replace it.
|
|
|
|
## Output
|
|
- Structured findings (schema) + block decision
|
|
- Confirmed findings with proofs, unverified and suppressed listed separately
|
|
- HARD STOP / BLOCK surfaced when `summary.block` is true
|