# Phase 3 — Path B deployment (R720 host + CI backstop) Status: **design + scripts; NOT deployed.** Needs the R720 (hostname SH-NVR, deprecated NVR, no GPU) and Adam's hands-on involvement. Nothing here has been run against the box. See memory `project-security-review-agent`. ## What Phase 3 delivers 1. The orchestrator hosted on the R720 as an always-on service (the Path B execution engine). 2. A **headless agentic runner** (`run_headless.py`) so the detector fan-out + proof-or-kill verifier can run unattended (interactive `/sh-security-review` can't run in CI/cron). 3. CI backstop (`ci/security-review.yml`) running the same `review.sh` on every PR. 4. Nightly full-repo sweep + planted canary vulns + two-model disagreement on criticals. 5. Budget/telemetry ceiling against the Max-20x **$200/mo Agent SDK credit pool** (see memory `reference-claude-subscription-billing`). ## The load-bearing piece to BUILD first: `run_headless.py` The interactive skill orchestrates subagents via the Claude Code Agent tool (Max-covered). Headless, that must become explicit API/SDK calls. `run_headless.py` should: - take `--scope` + target dir, read the in-scope files; - run the 6 detectors (injection/authz/secrets-crypto/iac-iam/web-client/logic) as parallel calls, each fresh context, using the SAME prompts as `~/.claude/commands/sh-security-review.md`; - run the proof-or-kill verifier pass; - emit the finding schema JSON that `review.sh --agent-findings` consumes. - **Billing:** authenticate via the **Claude Agent SDK with the subscription** to spend the $200/mo pool before API overflow (NOT a raw key) — see the billing memory. Non-Claude cross-family tiebreak stays on Bedrock/API. ⚠️ This code cannot be validated until it runs with real creds; do not mark done until tested. ## R720 host setup (Windows Server 2022, no GPU) Recommended: run the Python orchestrator in **WSL2 or Docker** rather than native Windows Python. 1. `git clone https://github.com/amoussa1229/orchestrator` onto the box. 2. Python 3.12 + `pip install -r requirements.txt` (+ the Agent SDK). 3. Install the deterministic scanners (semgrep/gitleaks/checkov/cfn-lint/pip-audit) on the box. 4. Secrets on the box (NOT in git): Agent SDK subscription auth, Bedrock creds, Slack token, repo PATs. 5. Run as a service (NSSM on Windows, or systemd inside WSL2). Expose to the harness as an MCP/HTTP service (mirror the existing unifi MCP pattern) OR keep it as a local scheduled job. 6. Scheduler: nightly full-repo sweep → `review.sh` → Slack report (ALARM-style, see memory `feedback_cloudwatch_alarms`: only notify on findings, not clean runs). ## Anti-complacency reinforcements (Phase 3) - **Canary vulns:** keep the testbed's planted vulns in a fixture the nightly sweep also scans; if the agent ever misses a known canary, that's a complacency signal → alert. - **Two-model disagreement:** on confirmed criticals, run the cross-family reviewer (orchestrator GPT-4.1) and flag disagreement for human review. - **Budget guard:** track output tokens vs the $200/mo pool; stop/alert before overflow. ## Open dependencies (Adam) - R720 reachable + WSL2/Docker chosen. - GitHub repo/org secrets: `ANTHROPIC_API_KEY` (or SDK auth), set `vars.SECURITY_REVIEW_AGENTIC=1` once `run_headless.py` is tested. - Decide: promote `ci/security-review.yml` into Sea-Haven-Industries/.github as a reusable workflow (per engineering-handbook/cicd.md) vs per-repo copy. - checkov severity filtering (it emitted 51-55 best-practice mediums on payments-dashboard — tune before nightly or the Slack reports are noise).