Commit graph

6 commits

Author SHA1 Message Date
1a1facd945 refactor(secrev): factor shared sweep substrate out of nightly_sweep.sh (Plane-1 Phase 0)
Extract discovery/mirror/budget-ledger/rotation/Slack-ALARM(redaction)/canary into security-review/lib/sweep_substrate.sh (sourceable, bash, stdlib only); nightly_sweep.sh (415->362) now sources it. Zero behavior change proven: shellcheck -x clean, offline two-tier dry-run byte-identical before/after, token discipline (read-only PAT, REST-only, no gh CLI, origin scrubbed) preserved, review.sh untouched. Revert point: 73e35f3. Foundation both planes' scheduled side reuses.
2026-06-18 14:06:31 -04:00
dbe2f8c186 Raise nightly agentic budget $20 -> $120 for full per-night coverage
First-run data (2026-06-17) showed the $20 ceiling covered only the
canary + 5 of ~22 scannable repos before pausing the rotation, leaving
16 repos un-deep-scanned that night. Raise TOTAL_BUDGET_USD default to
$120 so every repo gets a deep agentic pass each night (~22 x ~$5 +
canary, with headroom). Spend draws on the Max subscription pool; the
per-target cap ($12) and round-robin rotation are unchanged, so this is
a ceiling raise, not a per-repo cost change. Updates README, DEPLOY, and
the systemd Environment example to match.
2026-06-17 13:18:46 -04:00
15a16f8b88 Tune nightly defaults from first-run data
First full VM run: canary recall 18 with the expanded Node/.NET corpus, and a real
agentic repo cost ~$3.50 (vs the $1.4 testbed). Raise CANARY_FLOOR 10->14 (catches a
language-blindness recall collapse to ~9 while keeping margin under 18) and MAX_CYCLE_NIGHTS
4->6 (at ~4-6 agentic repos/night a full 22-repo rotation takes ~4-5 nights; 4 would
false-fire the coverage alarm). Both stay env-overridable and re-tunable as data accrues.
2026-06-16 18:02:57 -04:00
f4bb2bce8a Make scan_scanners return 0 so a passing repo doesn't trip set -e
scan_scanners communicates results via globals (T1_BLOCK/T1_CRIT/T1_HIGH); its last
statement was a bare [ $rc -eq 1 ] && T1_BLOCK=1. On a PASS (rc=0) that test is false,
so the function returned non-zero and set -e killed the whole sweep at the first passing
repo in tier 1 (right after the unbound-variable fix let it get that far). Add an explicit
return 0. Audited the rest of the tier1/tier2/summary path; scan_agentic and the others
already end on a zero-status command.
2026-06-16 16:43:35 -04:00
6d65c54b58 Fix unbound-variable crash in the nightly sweep tier-1 loop
scan_scanners referenced ${slug} inside the SAME local statement that defines it
(local target=... slug=$2 result_json=...${slug}...). bash expands the local's
arguments before the builtin assigns them, so under set -u ${slug} is unbound and
the sweep died right after the canary, before tier 1 ever ran. Split the local so
slug exists first. Fix the same latent self-reference in mirror_repo (dir=...$name),
which only worked by accident because the discovery loop left a global $name.
2026-06-16 16:00:21 -04:00
537e83975b Rewrite nightly sweep as two-tier clean-clone auto-discovery
Replace the opt-in sweep-targets allowlist with zero-wiring discovery: enumerate org
repos via the GitHub REST API (curl + read-only GH_TOKEN, no gh dependency) and mirror
each as a shallow clean clone (git clone --depth=1, default branch from the API) into
~/repo-mirrors. Scanning server-side clones keeps local .env secrets out of scope.

Tier 1 runs deterministic scanners over every repo nightly ($0 Claude); tier 2 runs the
agentic pass over a budget-bounded round-robin rotation with a persistent cycle pointer,
so the draw on the shared Max limits stays bounded and coverage never goes silently
incomplete (COVERAGE ALARM if the rotation falls behind). Skip = committed marker or
central list (marker-skips logged). Redact secrets from Slack; reports mode 600. Raise
the systemd timeout to 6h for the longer two-tier run.
2026-06-16 15:00:00 -04:00