checkov: high-signal exposure/access checks -> high, best-practice noise -> low (was 51 undifferentiated mediums). npm audit wired for Node dep CVEs. Report collapses the low/info tail to a count. Adds ci/security-review.yml (PR backstop) and DEPLOY-R720.md (Phase 3 host runbook).
3.5 KiB
3.5 KiB
Phase 3 — Path B deployment (R720 host + CI backstop)
Status: design + scripts; NOT deployed. Needs the R720 (hostname SH-NVR, deprecated NVR, no GPU)
and Adam's hands-on involvement. Nothing here has been run against the box. See memory
project-security-review-agent.
What Phase 3 delivers
- The orchestrator hosted on the R720 as an always-on service (the Path B execution engine).
- A headless agentic runner (
run_headless.py) so the detector fan-out + proof-or-kill verifier can run unattended (interactive/sh-security-reviewcan't run in CI/cron). - CI backstop (
ci/security-review.yml) running the samereview.shon every PR. - Nightly full-repo sweep + planted canary vulns + two-model disagreement on criticals.
- Budget/telemetry ceiling against the Max-20x $200/mo Agent SDK credit pool
(see memory
reference-claude-subscription-billing).
The load-bearing piece to BUILD first: run_headless.py
The interactive skill orchestrates subagents via the Claude Code Agent tool (Max-covered). Headless,
that must become explicit API/SDK calls. run_headless.py should:
- take
--scope+ target dir, read the in-scope files; - run the 6 detectors (injection/authz/secrets-crypto/iac-iam/web-client/logic) as parallel calls,
each fresh context, using the SAME prompts as
~/.claude/commands/sh-security-review.md; - run the proof-or-kill verifier pass;
- emit the finding schema JSON that
review.sh --agent-findingsconsumes. - Billing: authenticate via the Claude Agent SDK with the subscription to spend the $200/mo pool before API overflow (NOT a raw key) — see the billing memory. Non-Claude cross-family tiebreak stays on Bedrock/API. ⚠️ This code cannot be validated until it runs with real creds; do not mark done until tested.
R720 host setup (Windows Server 2022, no GPU)
Recommended: run the Python orchestrator in WSL2 or Docker rather than native Windows Python.
git clone https://github.com/amoussa1229/orchestratoronto the box.- Python 3.12 +
pip install -r requirements.txt(+ the Agent SDK). - Install the deterministic scanners (semgrep/gitleaks/checkov/cfn-lint/pip-audit) on the box.
- Secrets on the box (NOT in git): Agent SDK subscription auth, Bedrock creds, Slack token, repo PATs.
- Run as a service (NSSM on Windows, or systemd inside WSL2). Expose to the harness as an MCP/HTTP service (mirror the existing unifi MCP pattern) OR keep it as a local scheduled job.
- Scheduler: nightly full-repo sweep →
review.sh→ Slack report (ALARM-style, see memoryfeedback_cloudwatch_alarms: only notify on findings, not clean runs).
Anti-complacency reinforcements (Phase 3)
- Canary vulns: keep the testbed's planted vulns in a fixture the nightly sweep also scans; if the agent ever misses a known canary, that's a complacency signal → alert.
- Two-model disagreement: on confirmed criticals, run the cross-family reviewer (orchestrator GPT-4.1) and flag disagreement for human review.
- Budget guard: track output tokens vs the $200/mo pool; stop/alert before overflow.
Open dependencies (Adam)
- R720 reachable + WSL2/Docker chosen.
- GitHub repo/org secrets:
ANTHROPIC_API_KEY(or SDK auth), setvars.SECURITY_REVIEW_AGENTIC=1oncerun_headless.pyis tested. - Decide: promote
ci/security-review.ymlinto Sea-Haven-Industries/.github as a reusable workflow (per engineering-handbook/cicd.md) vs per-repo copy. - checkov severity filtering (it emitted 51-55 best-practice mediums on payments-dashboard — tune before nightly or the Slack reports are noise).