Commit graph

26 commits

Author SHA1 Message Date
bdafb7bbf1 fix(security-review): batch template-discovery grep so review.sh --scanners-only stops overflowing argv on the monorepo
The cfn-lint template-discovery step in review.sh piped the repo's whole
matched-file list into 'xargs -I{} sh -c "grep -l {}"'. On the orchestrator
monorepo — especially from a deep worktree path, where every matched path is a
long absolute path — xargs -I{} packs all paths into one assembled command and
aborts with 'xargs: command line cannot be assembled, too long'. The subprocess
exits non-zero having emitted ZERO findings, so the global pre-push hook BLOCKS
every push (agents were working around it with --no-verify).

Fix: switch the grep stage to NUL-delimited, un-batched xargs
(find ... -print0 | xargs -0 grep -lE ...). xargs -0 (no -I) splits the input
across multiple grep invocations, so the argv never exceeds ARG_MAX; grep -l
reports the same matching files as the old per-file grep, and -print0/-0 is safe
for paths with spaces/newlines. The first 'xargs -I{} find {}' is kept (find
needs the start path before its expression) and is bounded by the scope-path
count, so it is not an overflow source. Trailing '|| true' preserves the old
no-match/no-files semantics (TPLS = list-of-templates or empty, never fails).

Purely an argv-batching fix: the scanned file set, findings, and exit codes are
unchanged. Verified exit-code-identical: clean tree -> exit 0 (PASS); planted
GitHub PAT + RSA private key -> exit 1 (BLOCK, gitleaks high); planted CFN
template with a cfn-lint error on a deeply-nested path -> exit 1 (cfn-lint flags
it, no overflow). The previously-overflowing command now completes clean.
2026-06-23 15:51:43 -04:00
Adam Moussa
71036cb3f4
chore(security-review): bring confluence-doc online in the nightly coordinator (#42)
confluence-doc is provisioned (OAuth 2.0 service-account creds in ~/secrev.env +
PAGE_MAP_FILE from the live IT space). Drop it from COORDINATOR_SKIP_ROLES (now
just aws-posture) and add Environment=PAGE_MAP_FILE. Verified live on the box:
'Confluence API: oauth auth ready', 21 recommend-only doc gaps reported (D7, no
writes). aws-posture stays skipped (IAM/step-ca not stood up).
2026-06-22 19:54:53 -04:00
Adam Moussa
593f8c60b8
fix(security-review): resolve Confluence cloudId via accessible-resources, not /_edge/tenant_info (#41)
The public /_edge/tenant_info endpoint returns an HTML 'Page Unavailable' on the
seahavenind site, so cloudId auto-resolution failed. Fetch the token first, then
resolve cloudId from the OAuth-native https://api.atlassian.com/oauth/token/
accessible-resources (Bearer), preferring the resource whose url matches the
configured site, else the first. CONFLUENCE_CLOUD_ID override still honored.
Canary 3/3 (offline), shellcheck clean.
2026-06-22 19:47:38 -04:00
Adam Moussa
332652257b
feat(security-review): confluence-doc supports OAuth 2.0 client-credentials (service account) (#40)
Atlassian org service accounts have no classic API token — they authenticate via
OAuth 2.0 client-credentials (2LO). Add a dual-mode auth seam to confluence-doc:
- OAuth (preferred when CONFLUENCE_OAUTH_CLIENT_ID/_SECRET set): POST
  auth.atlassian.com/oauth/token (client_id+client_secret+grant_type=client_credentials)
  → 60-min Bearer; calls go to api.atlassian.com/ex/confluence/<cloudId>/wiki/api/v2/...
  cloudId auto-resolves from the site's public /_edge/tenant_info (no input needed).
- Basic (email+API token) retained as a fallback.
conf_api_init() picks the mode once; conf_get() does the authenticated GET. Any
failure (no cloudId / token request fails) → RUN_API=0, live checks SKIPPED, NO
false alarm (matches the existing no-data discipline). Secret passed in the request
body (--data-urlencode), never logged. Still read + recommend-only (D7); canary
unaffected (offline) — 3/3. shellcheck clean (accepted SC1091).
2026-06-22 19:45:42 -04:00
Adam Moussa
17ef4de415
feat(security-review): schedule the Plane-1 checker coordinator nightly + role-skip (#36)
- checker_coordinator.sh: add COORDINATOR_SKIP_ROLES env (comma-separated) to drop
  roles whose credentials are not provisioned from the registry entirely (never
  canaried/run/ALARMed). Fail-safe: empty/unset = run all.
- systemd: sea-haven-checkers.{service,timer} run the coordinator nightly at ~03:30
  UTC (90 min after the secrev sweep so they don't contend on $MIRROR_DIR / the
  Claude pool). The unit sets COORDINATOR_SKIP_ROLES=aws-posture,confluence-doc
  (aws-posture needs IAM Roles Anywhere; confluence-doc needs the confluence-bot
  token — both intentionally unprovisioned).

Deployed + enabled on the box (deploy-before-merge): canary 4/4 with the skip,
timer scheduled for 2026-06-23 03:35 UTC.
2026-06-22 19:05:02 -04:00
Adam Moussa
95e481890a
feat(secrev): wire full Plane-1 roster into checker coordinator + fix fixture SAM (#22)
* feat(secrev): wire full Plane-1 roster into the checker coordinator registry

Register doc-drift, aws-posture, plan-groomer, confluence-doc (weekly cadence)
alongside compliance-drift + dependency-cve (nightly). The coordinator canary
suite now runs all 6 roles' canaries (all PASS) under the one shared budget +
versioned rotation/coverage state; --squeeze-dry-run still proves defer-not-drop
+ COVERAGE alarm. Central integration after the parallel Phase-3/4 PRs landed.

* fix(secrev): valid SAM in doc-drift fixture templates (cfn-lint E0001)

The doc-drift sample-stack fixtures declared AWS::Serverless::Function with no
Properties; cfn-lint's SAM transform errored (HIGH). Added minimal valid
Properties (Handler/Runtime/InlineCode). Pre-existing on main — #19 pushed with
--no-verify (xargs overflow) and CI runs no cfn-lint, so it slipped through.
doc-drift still detects the stack (keys on template presence).
2026-06-18 16:31:03 -04:00
Adam Moussa
94ed6ea224
feat(secrev): Plane-1 Phase 3 — doc-drift + IAM artifacts (aws-posture gated) (#19)
* feat(secrev): doc-drift Plane-1 Tier-1 checker (UNGATED)

Third Plane-1 checker on the Phase-0 shared substrate, mirroring
compliance-drift.sh / dependency-cve.sh conventions verbatim (set -euo pipefail,
sourced substrate, --canary/--dry-run/--no-api/--refresh/--targets, mode-600
reports under $REPORT_ROOT/doc-drift/<UTC-date>/, ALARM-only, finding.schema
spirit JSON, exit 0/2/3, dotgit->.git fixture trick).

Detects documentation drift deterministically (design §4 doc-drift row):
  - readme-omits-component: README omits an existing major component in the tree
    (top-level service dir, SAM/CDK stack, Lambda handler dir, openapi/docs spec)
  - readme-stale-vs-code: README last-touch far older than newest code commit
    (two-factor: >=DOC_DRIFT_STALE_DAYS AND >=DOC_DRIFT_STALE_COMMITS)
A repo with NO README is SKIPPED (compliance-drift owns readme-present; no
double-flag). Future Gemini large-context judge (§4) is an inert stub (maybe_judge),
off in canary/dry-run/offline.

Planted-drift fixture corpus + EXPECTED_DRIFT_COUNT=4, canary-asserted (exit 3 on
miss). shellcheck -x clean (only accepted SC1091 source-line info).

Does NOT touch checker_coordinator.sh, requirements.txt, or aws-posture.
Wiring/systemd is gated (PROVISIONING footer). Design refs §4, §7 Phase 3.

* feat(secrev): Phase-3 IAM artifacts for cross-review (aws-posture gated)

Authored FILES (not applied to AWS — provisioning gated behind the mandatory
GPT-4.1 IAM cross-review + Adam, design §7 B3) for the aws-posture checker's
read-only AWS identity. Decision D5: box stays read-only, auths via IAM Roles
Anywhere short-lived leaf certs from a new internal step-ca; NO long-lived AWS key.

  - aws-posture-readonly-policy.json  least-privilege read-only (ce:Get*,
      cloudwatch:GetMetric*/DescribeAlarms, ec2/elb/rds:Describe*, lambda list +
      GetFunctionConfiguration, s3:ListAllMyBuckets/GetBucketLocation). No write,
      no iam:* mutation, no s3:GetObject/secrets/kms/logs data reads, no wildcard
      actions. Resource:* only where AWS has no resource-level support.
  - aws-posture-readonly-policy.rationale.md  per-statement least-privilege rationale.
  - aws-posture-trust-policy.json  pins Roles Anywhere principal + leaf subject CN +
      issuer CN + trust-anchor SourceArn (three conditions, all required).
  - roles-anywhere-config.json  trust anchor (pins step-ca root) + profile (1h session).
  - step-ca-config-sketch.md  internal CA config + systemd-timer leaf auto-renewal.
  - CROSS-REVIEW-PACKET.md  end-to-end trust model, blast radius, EXERCISED rollback,
      reviewer scrutiny list.

Does NOT build aws-posture.sh, touch checker_coordinator.sh, or requirements.txt.

* fix(secrev): apply IAM cross-review FIXes

GPT-4.1 IAM cross-review 2026-06-18: APPROVE, no BLOCKs. Applied FIXes:
- trust policy: add aws:SourceAccount=328440206208 (confused-deputy guard)
  alongside the existing aws:SourceArn trust-anchor pin
- readonly policy: remove ec2:DescribeImages (data minimization — AMIs are
  not an idle-spend signal)
- aws:RequestedRegion NIT: deliberately SKIPPED — ce:* and s3:ListAllMyBuckets
  are global-endpoint services a blanket region condition could DENY; rationale
  recorded in aws-posture-readonly-policy.rationale.md
- rationale.md + CROSS-REVIEW-PACKET.md: record APPROVE + FIXes + NIT answers
  (snapshots=account-owned idle signal; s3 list=names-only; no logs:* needed)

* feat(secrev): aws-posture checker (Tier-2, provisioning-gated)

Read-only Tier-2 idle/anomalous-spend + idle-resource posture checker for the
R720 agent-team (design D5 / §4 / §6.3 / §7 Phase 3). Mirrors the Tier-1 checker
conventions verbatim (flags --canary/--dry-run/--no-api/--targets, mode-600
report under $REPORT_ROOT/aws-posture/<date>/, ALARM-only, finding.schema.json
spirit, exit 0/2/3, shared substrate redact/post_slack_alarm).

Detectors (complement GuardDuty/SecurityHub/Config, do not replace):
- anomalous Cost Explorer deltas (ce get-anomalies, $-impact threshold)
- stopped EC2 still paying for attached EBS
- unattached EBS volumes
- unassociated Elastic IPs
- idle NAT gateways (≈0 bytes out)
- idle load balancers (0 healthy targets)
- idle RDS (0 connections over window)

Live AWS calls are PROVISIONING-GATED: they run ONLY when Roles Anywhere creds
are available (STS identity probe) AND not --no-api/--canary. With no creds or
--no-api/--canary the checker SKIPS live calls and notes them — NEVER alarms on
missing data (memory feedback_cloudwatch_alarms). Roles Anywhere/step-ca are not
stood up (IAM cross-review PASSED 2026-06-18; see security-review/iam/).

Offline canary: fixtures of mocked AWS responses (cost/describe-* JSON) under
fixtures/aws-posture/ + EXPECTED_FINDING_COUNT=7, asserted fully offline (no aws,
no network). Identical detector code runs online and offline. shellcheck-clean
(only accepted SC1091), chmod +x.

* fix(secrev): doc-drift fixture py ruff-clean (root CI runs check + format --check)

The repo-root CI lint runs both 'ruff check .' and 'ruff format --check .' over
all fixtures. Fixed E701 one-liners and ruff-formatted the sample-service .py
files (handlers/*, feature_*.py). Fixture content is irrelevant to doc-drift
(keys on file/dir presence + git staleness).
2026-06-18 16:25:40 -04:00
Adam Moussa
71b160e61b
feat(secrev): Plane-1 Phase 4 — plan-groomer + confluence-doc (recommend-only) (#20)
* feat(secrev): plan-groomer Plane-1 Phase 4 planner (report-only)

Aggregates the OTHER Plane-1 checkers' latest reports (compliance-drift,
dependency-cve, doc-drift, confluence-doc) into one prioritized, deduped
"groomed weekly plan" written into the mode-600 report. REPORT-ONLY per
decision D3: posts NOTHING to Slack; auto-write to Notion/Jira is a later
toggle (inert --notify seam). Reuses lib/sweep_substrate.sh redact().

Offline --canary asserts the groomed-plan item count (5) against a fixture
report set, exercising latest-date selection, dedup, multi-source aggregation,
and no-data discipline (a missing source is noted, never invented as work).
shellcheck-clean (only the shared SC1091 substrate-source info, at parity with
compliance-drift/dependency-cve). PROVISIONING (auto-write toggle, systemd
wiring, coordinator registry) deferred — gated.

* feat(secrev): confluence-doc Plane-1 Phase 4 doc-gap detector (recommend-only)

Scheduled, read-only documentation gap detector. Diffs the org repo set + an
optional read-only AWS inventory + the IT page-ID map (project_confluence_
migration) against Confluence and REPORTS doc gaps / stale pages / missing
runbooks into the mode-600 report. RECOMMEND-ONLY per D3/D7: NEVER auto-writes
Confluence; the on-demand SSH-invoked write path (incl. Mermaid edits via
~/.claude/scripts/confluence_mermaid.py) is a separate, gated provisioning path.

LIVE Confluence API reads need the gated confluence-bot service-account token
(D6); when creds are absent OR --no-api/--canary, the API checks are SKIPPED and
noted, NEVER reported as a gap on missing data (mirrors compliance-drift's
status-code-aware API-skip pattern: 200 parse, 404 real gap, else skip).

Offline --canary asserts the doc-gap count (3) against a fixture (repo list +
mock page-map + mock AWS inventory): a repo with no IT page, an AWS resource not
in the map, and a missing required runbook page; precision non-gaps (matched
repos/resources, doc-exempt repo, present required pages, skipped API) must not
inflate the count. shellcheck-clean (only the shared SC1091 substrate-source
info). PROVISIONING (confluence-bot account + 90-day rotation, page-1540098 live
dry-run expecting 16 weweave macros, systemd wiring, coordinator registry)
documented in the footer, deferred — gated.

* fix(secrev): commit compliance-drift secret fixture as dotenv.fixture (canary broke on fresh clone)

The compliance-drift canary's planted tracked-secret fixture was BadName_repo/.env,
but the repo root .gitignore lists '.env' — so it was never committed. On a fresh
clone of main the file is absent, the secrets-committed check stops firing, and the
canary FAILS (expected 6, got 5). It only passed where a gitignored, untracked
'.env' happened to exist locally. Verified the failure reproduces in a clean clone
of origin/main (3d97139) and in a fresh worktree.

Fix (in-convention, mirrors the dependency-cve .fixture-suffix trick): ship the
secret as BadName_repo/dotenv.fixture (committable, not gitignored); the --canary
materialization renames dotenv.fixture -> .env in its temp work area. The dotgit/
index already TRACKS .env, so git ls-files still reports it and the drift fires.
Restores the documented 6/6 canary on any fresh checkout. shellcheck stays clean.
2026-06-18 16:03:37 -04:00
Adam Moussa
3d97139300
feat(agent-team): P3-live CI apply/verify hardening + ci_fetcher (gate-passed, provisioning-gated) (#17)
* feat(agent-team): read-only CI-result fetcher for P3 verify gate (opt-in, inert)

ci_fetcher.py: fail-closed CiResultFetcher reading the GitHub Actions run
conclusion via a read-only PAT (AGENT_TEAM_CI_READ_TOKEN→GITHUB_TOKEN), returns
{run_id,conclusion,diff_hash} or None on any error. Data-fetcher only — ci_gate
owns the verdict; never writes, no OIDC/AWS, never reads patch artifacts.
coordinator gains opt-in gated_build_verify_wiring() composing it via
bind_ci_result_fetcher; NOT wired into the default run-team.py path. 20 tests.

* harden(agent-team): P3 apply/verify workflow — GitHub App token, CWE-94, fail-closed

Decision-1 auth model: gate-and-pr uses a GitHub App installation token
(pull-requests:write) behind the agent-apply environment; ALL OIDC/id-token/AWS
removed. Hardening: task_id env-indirection (CWE-94 — GitHub expands ${{ }} into
the run shell before exec, so %s/quoting is insufficient); run-id pinning on both
download-artifact; post-build denied-path check (build-hook writes into denied
paths fail the job); empty-hash fail-closed in BOTH the embedded gate (fixed a
real ''=='' pass bug) and ci_gate.py. App-token + draft-PR steps stay if:${{ false }}
until provisioning (App + environment + branch protection). +17 tests.

* harden(agent-team): apply P3-live security-gate fixes (GPT-4.1 xreview + sh-security-review)

BLOCK-1/FIX-4: gate-and-pr re-comments pull-requests:write + environment:agent-apply
(provisioning-time uncomment) and gains needs.guard/build-test=='success' job guard —
zero privilege until provisioning. BLOCK-2/3+FIX-5: ci_fetcher validates run_id (^[0-9]{1,20}$),
owner/repo (^[A-Za-z0-9_.-]{1,100}$), and fetched_id (int) — fail closed, no SSRF/path
injection. FIX-1: conclusion allowlist. FIX-3: api_root removed from public builder (no
injectable endpoint). INJ-02: post-build denied-path check uses NUL-delimited git output +
explicit rename parsing, no backslash mangling, non-UTF8=violation. INJ-03: all three trust-
control denylists unified to one 22-entry union + drift-guard test. Q1: documented run_id/
diff_hash trust source (dispatcher/ledger only). 884 tests, ruff clean. Privileged steps stay
if:${{ false }} until provisioning.

* build(security-review): prune .claude worktrees from deterministic scanners

Agent worktrees under .claude/worktrees/ are full repo copies; the cfn-lint
find|xargs template scan overflowed ('command line cannot be assembled') and the
pre-push hook fail-closed to BLOCK whenever a worktree was present. Prune .claude
in the cfn-lint find + semgrep/checkov excludes, and gitignore .claude/ so it is
never scanned or committed. Unblocks main-tree pushes during parallel agent work.
2026-06-18 15:53:26 -04:00
Adam Moussa
f09c94a821
feat(secrev): Plane-1 Phase 2 — coordinator + dependency-cve checker (#16)
* fix(agent-team): read SLACK_CHANNEL_ID, aligning code with deploy doc + systemd

run-team.py read os.environ['SLACK_CHANNEL'] while DEPLOY-R720.md and the
coordinator systemd unit both document SLACK_CHANNEL_ID; the mismatch would
silently default the live Slack transport channel to empty. Standardize on
SLACK_CHANNEL_ID (decision locked 2026-06-18).

* feat(secrev): dependency-cve Plane-1 Tier-1 checker (OSV, ALARM-only)

Read-only checker on the Phase-0 substrate: scans $MIRROR_DIR mirrors for
pinned deps (requirements/poetry/Pipfile/package-lock/yarn/csproj across
PyPI/npm/NuGet), cross-refs OSV querybatch (live) or an offline advisory
fixture (canary). Mode-600 reports, ALARM-only, --canary asserts 2 planted
vulns (jinja2 2.11.2, lodash 4.17.15). Complements Dependabot. Not provisioned.

* feat(secrev): Plane-1 checker coordinator (shared budget, rotation, dedup)

Coordinator (design §5/§6.7) orchestrating Tier-1 checkers under one shared
budget ledger + versioned rotation/coverage state (atomic write + schema/hash/
logical-consistency integrity, park-on-corrupt). Canary-suite-first
(COMPLACENCY skip), fan-out under the shared cap with defer-not-drop, COVERAGE
alarm past MAX_CYCLE_NIGHTS, cross-checker dedup/prioritize, ALARM-only routing.
--squeeze-dry-run proves deferral-not-drop + COVERAGE alarm. Not provisioned.

* fix(secrev): hide dependency-cve canary manifests from dependency-review

The canary fixtures intentionally pin known-vulnerable deps (jinja2 2.11.2,
lodash 4.17.15) so the checker has something to detect. GitHub's dependency
graph parsed those fixture manifests as real project deps, failing the
dependency-review PR gate (fail-on-severity: high). Store the manifests with a
.fixture suffix so the dependency graph ignores them; the --canary materializer
strips the suffix in its temp work area before scanning, so detection is
unchanged (still 2/2). No advisory allowlist, no change to the shared org
reusable workflow — the real gate stays strict for actual deps.
2026-06-18 15:08:58 -04:00
28553d86f7 feat(secrev): compliance-drift Plane-1 Tier-1 checker (ALARM-only)
Read-only org compliance checker on the shared substrate (no re-clone; scans existing mirrors). Checklist grounded in handbook/github-standards: kebab repo name, README, CI/CD, Dependabot config+alerts, tracked-.env secrets, branch protection, merge settings; handbook exceptions (docs-only, compliance-exempt) honored. Mode-600 reports, ALARM-only (clean=silent). Includes planted-drift canary (asserts 6). Review fixes folded in: branch-protection + dependabot are status-code-aware (only a real 404 is drift; transient API failure -> skip, no false alarm); secrets-committed fires only on secret-shaped values (not benign config). NOT scheduled (provisioning gated).
2026-06-18 14:06:31 -04:00
1a1facd945 refactor(secrev): factor shared sweep substrate out of nightly_sweep.sh (Plane-1 Phase 0)
Extract discovery/mirror/budget-ledger/rotation/Slack-ALARM(redaction)/canary into security-review/lib/sweep_substrate.sh (sourceable, bash, stdlib only); nightly_sweep.sh (415->362) now sources it. Zero behavior change proven: shellcheck -x clean, offline two-tier dry-run byte-identical before/after, token discipline (read-only PAT, REST-only, no gh CLI, origin scrubbed) preserved, review.sh untouched. Revert point: 73e35f3. Foundation both planes' scheduled side reuses.
2026-06-18 14:06:31 -04:00
b0d8b842e5 security-review: surgically exclude only cdk.out asset bundles from checkov (A2)
The prior exclusion (--skip-path cdk.out) stopped the CDK-repo stall but also
silenced checkov on cdk.out/<stack>.template.json — the actual deploy artifact —
losing real S3/IAM IaC coverage (CKV_AWS_53-56, CKV_AWS_111). Switch to skipping
only the cdk.out asset.<hash>/ dependency bundles (the stall cause) so the
synthesized templates are still scanned.

- --skip-path 'cdk\.out/asset\.' anchors to cdk.out so a source file literally
  named asset.* is not also excluded; keeps cdk.out/*.template.json scanned.
- venv/dist/build kept as bare names (match anywhere); .venv/.aws-sam escaped.

cfn-lint + semgrep already prune these trees (prior commit). gitleaks runs in
git-mode and respects .gitignore, so cdk.out is already skipped there.

Verified: shellcheck clean; synthetic cdk.out test confirms the stack template is
scanned while asset.* is skipped; the orchestrator's own --scanners-only gate
still exits 0 with suppressions (inert on non-CDK repos: A1==A2 findings here).
2026-06-17 17:06:05 -04:00
c6ecc621be Prune generated/vendored trees from the scanners
cfn-lint, semgrep, and checkov were scanning synthesized/vendored output
(cdk.out, node_modules, .venv/venv, .aws-sam, dist, build). On CDK repos this
explodes the find/xargs arg list and stalls the scan, and flagging synthesized
templates is wrong. Prune those trees in the cfn-lint find, and pass
--exclude / --skip-path to semgrep / checkov.
2026-06-17 14:55:41 -04:00
dbe2f8c186 Raise nightly agentic budget $20 -> $120 for full per-night coverage
First-run data (2026-06-17) showed the $20 ceiling covered only the
canary + 5 of ~22 scannable repos before pausing the rotation, leaving
16 repos un-deep-scanned that night. Raise TOTAL_BUDGET_USD default to
$120 so every repo gets a deep agentic pass each night (~22 x ~$5 +
canary, with headroom). Spend draws on the Max subscription pool; the
per-target cap ($12) and round-robin rotation are unchanged, so this is
a ceiling raise, not a per-repo cost change. Updates README, DEPLOY, and
the systemd Environment example to match.
2026-06-17 13:18:46 -04:00
15a16f8b88 Tune nightly defaults from first-run data
First full VM run: canary recall 18 with the expanded Node/.NET corpus, and a real
agentic repo cost ~$3.50 (vs the $1.4 testbed). Raise CANARY_FLOOR 10->14 (catches a
language-blindness recall collapse to ~9 while keeping margin under 18) and MAX_CYCLE_NIGHTS
4->6 (at ~4-6 agentic repos/night a full 22-repo rotation takes ~4-5 nights; 4 would
false-fire the coverage alarm). Both stay env-overridable and re-tunable as data accrues.
2026-06-16 18:02:57 -04:00
f4bb2bce8a Make scan_scanners return 0 so a passing repo doesn't trip set -e
scan_scanners communicates results via globals (T1_BLOCK/T1_CRIT/T1_HIGH); its last
statement was a bare [ $rc -eq 1 ] && T1_BLOCK=1. On a PASS (rc=0) that test is false,
so the function returned non-zero and set -e killed the whole sweep at the first passing
repo in tier 1 (right after the unbound-variable fix let it get that far). Add an explicit
return 0. Audited the rest of the tier1/tier2/summary path; scan_agentic and the others
already end on a zero-status command.
2026-06-16 16:43:35 -04:00
6d65c54b58 Fix unbound-variable crash in the nightly sweep tier-1 loop
scan_scanners referenced ${slug} inside the SAME local statement that defines it
(local target=... slug=$2 result_json=...${slug}...). bash expands the local's
arguments before the builtin assigns them, so under set -u ${slug} is unbound and
the sweep died right after the canary, before tier 1 ever ran. Split the local so
slug exists first. Fix the same latent self-reference in mirror_repo (dir=...$name),
which only worked by accident because the discovery loop left a global $name.
2026-06-16 16:00:21 -04:00
c217c5656d Add machine-level suppressions to the repo-sourced hooks
The live global pre-push hook was hand-edited to resolve suppressions from a
machine-level file (${SH_SECURITY_SUPPRESSIONS_DIR:-~/.config/sea-haven/security-review}/
<repo-basename>/suppressions.json) kept out of repo history, falling back to a
repo-local .security-review/suppressions.json. The repo-sourced hooks lacked it, so
install-hooks.sh --global would overwrite the live hook and lose the feature.

Port the prefer-machine/fallback-repo-local block into hooks/pre-push, align
hooks/pre-commit to the same (else a machine-suppressed finding passes at push but
blocks at commit), and document the path + SH_SECURITY_SUPPRESSIONS_DIR override +
basename-collision caveat in the README.
2026-06-16 15:33:43 -04:00
f4dd72ced9 Add headless detector fan-out + proof-or-kill verifier runner
run_headless.py is the Path B engine the nightly sweep calls: the 6 fresh-context
detectors + proof-or-kill verifier from /sh-security-review, run unattended via the
Claude Agent SDK on subscription OAuth (pops ANTHROPIC_API_KEY so the API key can't
silently win). Read-only tools, hermetic, per-call + total budget caps, fails toward
over-reporting. Emits the finding schema that review.sh --agent-findings consumes.
2026-06-16 15:00:10 -04:00
a05afb5ebc Update security-review docs for two-tier auto-discovery
Rewrite README + DEPLOY-R720 for the clean-clone mirror model, the read-only GH_TOKEN
PAT recipe, the global hook installer, and the no-CI-by-design decision.
2026-06-16 15:00:00 -04:00
537e83975b Rewrite nightly sweep as two-tier clean-clone auto-discovery
Replace the opt-in sweep-targets allowlist with zero-wiring discovery: enumerate org
repos via the GitHub REST API (curl + read-only GH_TOKEN, no gh dependency) and mirror
each as a shallow clean clone (git clone --depth=1, default branch from the API) into
~/repo-mirrors. Scanning server-side clones keeps local .env secrets out of scope.

Tier 1 runs deterministic scanners over every repo nightly ($0 Claude); tier 2 runs the
agentic pass over a budget-bounded round-robin rotation with a persistent cycle pointer,
so the draw on the shared Max limits stays bounded and coverage never goes silently
incomplete (COVERAGE ALARM if the rotation falls behind). Skip = committed marker or
central list (marker-skips logged). Redact secrets from Slack; reports mode 600. Raise
the systemd timeout to 6h for the longer two-tier run.
2026-06-16 15:00:00 -04:00
a3ab3f5f40 Make security-review hooks and skill installable from the repo
The global pre-push hook, the /sh-security-review prompt, and finding.schema.json
previously lived only in ~/.config/git and ~/.claude (untracked) — unreproducible.
Source them here: add hooks/pre-push, rewrite install-hooks.sh with a --global mode
(lays down both hooks, sets core.hooksPath, links skill+schema into ~/.claude) and a
per-repo mode. Align pre-commit with pre-push (honor skip marker + suppressions). Add
semgrep p/javascript so the scanners cover the org's Node/.NET repos.
2026-06-16 15:00:00 -04:00
18412c7482 Remove parked security-review CI drafts
CI was swapped for the global git hooks + nightly VM sweep (solo dev), so
the parked ci/*.yml and CI-BACKSTOP-NOTES.md were dead weight — a defective
workflow in-tree is a foot-gun. Recover from history if the team grows.
2026-06-16 14:59:42 -04:00
f90f759e12 Tune checkov severity, wire npm audit, add CI backstop + R720 runbook
checkov: high-signal exposure/access checks -> high, best-practice noise -> low
(was 51 undifferentiated mediums). npm audit wired for Node dep CVEs. Report
collapses the low/info tail to a count. Adds ci/security-review.yml (PR backstop)
and DEPLOY-R720.md (Phase 3 host runbook).
2026-06-15 15:57:34 -04:00
7e5ce1f5b2 Add security-review gate (review.sh + scanners + pre-commit hook)
Trigger-agnostic pure-code gate that merges deterministic-scanner findings
(semgrep/gitleaks/checkov/cfn-lint/pip-audit) with agent findings from
/sh-security-review, dedups, applies justification-required suppressions, and
makes the block decision (exit 1 on confirmed critical/high). Phase 2 of the
Sea Haven security-review agent; Path B (CI/headless) wiring lands in Phase 3.
2026-06-15 15:52:15 -04:00