feat(secrev): Plane-1 Phase 2 — coordinator + dependency-cve checker (#16)
* fix(agent-team): read SLACK_CHANNEL_ID, aligning code with deploy doc + systemd
run-team.py read os.environ['SLACK_CHANNEL'] while DEPLOY-R720.md and the
coordinator systemd unit both document SLACK_CHANNEL_ID; the mismatch would
silently default the live Slack transport channel to empty. Standardize on
SLACK_CHANNEL_ID (decision locked 2026-06-18).
* feat(secrev): dependency-cve Plane-1 Tier-1 checker (OSV, ALARM-only)
Read-only checker on the Phase-0 substrate: scans $MIRROR_DIR mirrors for
pinned deps (requirements/poetry/Pipfile/package-lock/yarn/csproj across
PyPI/npm/NuGet), cross-refs OSV querybatch (live) or an offline advisory
fixture (canary). Mode-600 reports, ALARM-only, --canary asserts 2 planted
vulns (jinja2 2.11.2, lodash 4.17.15). Complements Dependabot. Not provisioned.
* feat(secrev): Plane-1 checker coordinator (shared budget, rotation, dedup)
Coordinator (design §5/§6.7) orchestrating Tier-1 checkers under one shared
budget ledger + versioned rotation/coverage state (atomic write + schema/hash/
logical-consistency integrity, park-on-corrupt). Canary-suite-first
(COMPLACENCY skip), fan-out under the shared cap with defer-not-drop, COVERAGE
alarm past MAX_CYCLE_NIGHTS, cross-checker dedup/prioritize, ALARM-only routing.
--squeeze-dry-run proves deferral-not-drop + COVERAGE alarm. Not provisioned.
* fix(secrev): hide dependency-cve canary manifests from dependency-review
The canary fixtures intentionally pin known-vulnerable deps (jinja2 2.11.2,
lodash 4.17.15) so the checker has something to detect. GitHub's dependency
graph parsed those fixture manifests as real project deps, failing the
dependency-review PR gate (fail-on-severity: high). Store the manifests with a
.fixture suffix so the dependency graph ignores them; the --canary materializer
strips the suffix in its temp work area before scanning, so detection is
unchanged (still 2/2). No advisory allowlist, no change to the shared org
reusable workflow — the real gate stays strict for actual deps.
2026-06-18 15:08:58 -04:00
#!/usr/bin/env bash
# checker_coordinator.sh — Plane-1 coordinator for the R720 agent-team.
#
# Design refs: docs/r720-agent-team-design.md §5 (Coordination model), §6.1/§6.6 (ONE shared
# cap across all roles — critical for the Claude subscription draw), §6.7 (state durability +
# backup: atomic write-temp-then-rename, schema-version + content-hash + logical-consistency
# integrity check, park-on-corrupt), §7 Phase 2 ("coordinator + second checker; run a forced
# budget-squeeze dry-run to prove deferral-not-drop + COVERAGE ALARM").
#
# WHAT IT DOES:
2026-06-18 16:31:03 -04:00
# Orchestrates the Plane-1 checkers (compliance-drift, dependency-cve, doc-drift, aws-posture,
# plan-groomer, confluence-doc) under ONE shared
feat(secrev): Plane-1 Phase 2 — coordinator + dependency-cve checker (#16)
* fix(agent-team): read SLACK_CHANNEL_ID, aligning code with deploy doc + systemd
run-team.py read os.environ['SLACK_CHANNEL'] while DEPLOY-R720.md and the
coordinator systemd unit both document SLACK_CHANNEL_ID; the mismatch would
silently default the live Slack transport channel to empty. Standardize on
SLACK_CHANNEL_ID (decision locked 2026-06-18).
* feat(secrev): dependency-cve Plane-1 Tier-1 checker (OSV, ALARM-only)
Read-only checker on the Phase-0 substrate: scans $MIRROR_DIR mirrors for
pinned deps (requirements/poetry/Pipfile/package-lock/yarn/csproj across
PyPI/npm/NuGet), cross-refs OSV querybatch (live) or an offline advisory
fixture (canary). Mode-600 reports, ALARM-only, --canary asserts 2 planted
vulns (jinja2 2.11.2, lodash 4.17.15). Complements Dependabot. Not provisioned.
* feat(secrev): Plane-1 checker coordinator (shared budget, rotation, dedup)
Coordinator (design §5/§6.7) orchestrating Tier-1 checkers under one shared
budget ledger + versioned rotation/coverage state (atomic write + schema/hash/
logical-consistency integrity, park-on-corrupt). Canary-suite-first
(COMPLACENCY skip), fan-out under the shared cap with defer-not-drop, COVERAGE
alarm past MAX_CYCLE_NIGHTS, cross-checker dedup/prioritize, ALARM-only routing.
--squeeze-dry-run proves deferral-not-drop + COVERAGE alarm. Not provisioned.
* fix(secrev): hide dependency-cve canary manifests from dependency-review
The canary fixtures intentionally pin known-vulnerable deps (jinja2 2.11.2,
lodash 4.17.15) so the checker has something to detect. GitHub's dependency
graph parsed those fixture manifests as real project deps, failing the
dependency-review PR gate (fail-on-severity: high). Store the manifests with a
.fixture suffix so the dependency graph ignores them; the --canary materializer
strips the suffix in its temp work area before scanning, so detection is
unchanged (still 2/2). No advisory allowlist, no change to the shared org
reusable workflow — the real gate stays strict for actual deps.
2026-06-18 15:08:58 -04:00
# budget + versioned rotation/coverage state. Nightly it (mirrors nightly_sweep + §5):
# 1) loads the shared budget ledger + the versioned rotation/coverage state (integrity-checked)
# 2) runs the CANARY SUITE FIRST — each role's checker with --canary; a miss is a COMPLACENCY
# ALARM + that role is SKIPPED this run (never run a degraded role silently)
# 3) fans out roles due to run (deferred-first, then rotation) under the SHARED cap; a role
# whose estimated cost would exceed the ceiling is DEFERRED (recorded), never dropped
# 4) raises a COVERAGE ALARM if any role's last_run slips past MAX_CYCLE_NIGHTS
# 5) collects each run checker's report JSON, merges + DEDUPS across checkers, prioritizes
# 6) routes ALARM-only (D3): confirmed critical/high -> Slack ALARM; everything else -> a
# combined mode-600 coordinator report; a fully clean run posts NOTHING
#
# SUBSTRATE REUSE (lib/sweep_substrate.sh, sourced — bash dynamic scoping):
# add_spend / over_budget -> shared budget ledger (read TOTAL_SPEND/TOTAL_BUDGET_USD)
# redact / post_slack_alarm-> Slack delivery (read SLACK_WEBHOOK_URL, REPORT_DIR, SWEEP_LOG)
# to_epoch -> cycle-age accounting for the COVERAGE alarm
# The coordinator does NOT re-implement these; it provides the globals the contract names.
#
# STATE DURABILITY (design §6.7): both the budget ledger and the rotation/coverage state are
# written ATOMICALLY (temp + rename) and integrity-checked on load = schema_version match +
# stored content_hash + a logical-consistency check. On corruption the coordinator refuses to
# proceed silently -> it PARKS that store + ALARMs; the budget ledger is rebuildable (a new UTC
# day resets the day's spend), the rotation state is rebuildable from report history.
#
# SCOPE / SAFETY: read-only orchestration. Does NOT install systemd units, does NOT touch
# agent_team/ or agent-team/, does NOT re-clone by default (checkers reuse $MIRROR_DIR; a
# checker's own --refresh is the only network path and is not invoked here). See the
# "PROVISIONING (NOT DONE HERE)" footer.
#
# Exit: 0 = ran (whether or not it alarmed); 2 = setup/usage error; 3 = a canary/assertion FAILED.
set -euo pipefail
export PATH = " $HOME /.local/bin:/opt/homebrew/bin:/usr/local/bin: $PATH "
log( ) { echo " [coordinator] $* " >& 2; }
die( ) { echo " [coordinator] FATAL: $* " >& 2; exit 2; }
# --- Shared substrate ---------------------------------------------------------
HERE = " $( cd " $( dirname " ${ BASH_SOURCE [0] } " ) " && pwd ) "
SUBSTRATE = " $HERE /lib/sweep_substrate.sh "
[ -f " $SUBSTRATE " ] || die " shared substrate not found: $SUBSTRATE "
# shellcheck source=lib/sweep_substrate.sh
. " $SUBSTRATE "
CHECKERS_DIR = " $HERE /checkers "
# --- Config + defaults (env, all optional) ------------------------------------
GH_ORG = " ${ GH_ORG :- Sea -Haven-Industries } "
MIRROR_DIR = " ${ MIRROR_DIR :- $HOME /repo-mirrors } "
REPORT_ROOT = " ${ REPORT_ROOT :- $HOME /sweep-reports } "
TOTAL_BUDGET_USD = " ${ TOTAL_BUDGET_USD :- 120 } " # ONE shared cap across ALL roles (design §6.1)
MAX_CYCLE_NIGHTS = " ${ MAX_CYCLE_NIGHTS :- 6 } " # COVERAGE alarm if a role slips past this many days
SCHEMA_VERSION = 1 # bump when a state-file shape changes
DRY_RUN = 0 # --dry-run: compose alarms/reports but DO NOT post (routing dry-run)
CANARY = 0 # --canary: run every role's canary + assert all pass (offline)
SQUEEZE = 0 # --squeeze-dry-run: Phase-2 acceptance — force deferral + COVERAGE proof
# --once is accepted for parity with the sweep (single pass; this script IS a single pass).
usage( ) {
cat >& 2 <<EOF
checker_coordinator.sh — Plane-1 coordinator ( shared budget + versioned rotation, read-only)
--canary run EVERY role' s canary and assert all pass ( offline) ; post nothing
--dry-run run roles but compose alarms/reports WITHOUT posting ( routing dry-run)
--squeeze-dry-run Phase-2 acceptance test: force a tiny TOTAL_BUDGET_USD so a role MUST be
DEFERRED ( not dropped) AND simulate enough elapsed cycles to trip the
COVERAGE ALARM; prints the deferred list + COVERAGE alarm, posts nothing
--once single coordination pass ( this script is always a single pass)
-h| --help this help
Env: TOTAL_BUDGET_USD MAX_CYCLE_NIGHTS REPORT_ROOT MIRROR_DIR GH_ORG GH_TOKEN SLACK_WEBHOOK_URL
EOF
}
while [ $# -gt 0 ] ; do
case " $1 " in
--canary) CANARY = 1 ; ;
--dry-run) DRY_RUN = 1 ; ;
--squeeze-dry-run) SQUEEZE = 1; DRY_RUN = 1 ; ;
--once) : ; ;
-h| --help) usage; exit 0 ; ;
*) die " unknown arg: $1 (see --help) " ; ;
esac
shift
done
command -v jq >/dev/null || die "jq is required"
command -v git >/dev/null || die "git is required"
# --- Report dir (mode 600 reports; matches sweep conventions) -----------------
umask 077
UTC_DATE = " $( date -u +%Y-%m-%d) "
UTC_STAMP = " $( date -u +%Y-%m-%dT%H:%M:%SZ) "
REPORT_DIR = " $REPORT_ROOT /coordinator/ $UTC_DATE "
mkdir -p " $REPORT_DIR " ; chmod 700 " $REPORT_ROOT " " $REPORT_DIR " 2>/dev/null || true
# shellcheck disable=SC2034 # read by the sourced substrate (post_slack_alarm) via dynamic scope
SWEEP_LOG = " $REPORT_DIR /coordinator.log " # name the substrate's post_slack_alarm() references
REPORT_JSON = " $REPORT_DIR /coordinator.json "
REPORT_TXT = " $REPORT_DIR /coordinator.txt "
BUDGET_LEDGER = " ${ BUDGET_LEDGER :- $REPORT_ROOT /.budget-ledger.json } "
COORD_STATE = " ${ COORD_STATE :- $REPORT_ROOT /.coordinator-state.json } "
log " === checker_coordinator $UTC_STAMP (canary= $CANARY dry_run= $DRY_RUN squeeze= $SQUEEZE ) === "
# In the squeeze acceptance test, force a budget so small the SECOND role cannot fit.
if [ " $SQUEEZE " -eq 1 ] ; then
TOTAL_BUDGET_USD = "0.01"
log " SQUEEZE: forcing TOTAL_BUDGET_USD=\$ $TOTAL_BUDGET_USD so at least one role must DEFER "
fi
# ==============================================================================
# REGISTRY — Tier-1 checker roles (name | script | est per-run cost USD | cadence-days).
# A simple in-script table, easy to extend in later phases (add doc-drift, aws-posture...).
# Cost is the shared-budget DRAW estimate (these checkers are deterministic/cheap; a future
# agentic-judge role would carry a real Claude cost). Cadence is informational here.
# ==============================================================================
declare -a ROLES = (
" compliance-drift| $CHECKERS_DIR /compliance-drift.sh|0.00|1 "
" dependency-cve| $CHECKERS_DIR /dependency-cve.sh|0.00|1 "
2026-06-18 16:31:03 -04:00
" doc-drift| $CHECKERS_DIR /doc-drift.sh|0.00|7 "
" aws-posture| $CHECKERS_DIR /aws-posture.sh|0.00|7 "
" plan-groomer| $CHECKERS_DIR /plan-groomer.sh|0.00|7 "
" confluence-doc| $CHECKERS_DIR /confluence-doc.sh|0.00|7 "
feat(secrev): Plane-1 Phase 2 — coordinator + dependency-cve checker (#16)
* fix(agent-team): read SLACK_CHANNEL_ID, aligning code with deploy doc + systemd
run-team.py read os.environ['SLACK_CHANNEL'] while DEPLOY-R720.md and the
coordinator systemd unit both document SLACK_CHANNEL_ID; the mismatch would
silently default the live Slack transport channel to empty. Standardize on
SLACK_CHANNEL_ID (decision locked 2026-06-18).
* feat(secrev): dependency-cve Plane-1 Tier-1 checker (OSV, ALARM-only)
Read-only checker on the Phase-0 substrate: scans $MIRROR_DIR mirrors for
pinned deps (requirements/poetry/Pipfile/package-lock/yarn/csproj across
PyPI/npm/NuGet), cross-refs OSV querybatch (live) or an offline advisory
fixture (canary). Mode-600 reports, ALARM-only, --canary asserts 2 planted
vulns (jinja2 2.11.2, lodash 4.17.15). Complements Dependabot. Not provisioned.
* feat(secrev): Plane-1 checker coordinator (shared budget, rotation, dedup)
Coordinator (design §5/§6.7) orchestrating Tier-1 checkers under one shared
budget ledger + versioned rotation/coverage state (atomic write + schema/hash/
logical-consistency integrity, park-on-corrupt). Canary-suite-first
(COMPLACENCY skip), fan-out under the shared cap with defer-not-drop, COVERAGE
alarm past MAX_CYCLE_NIGHTS, cross-checker dedup/prioritize, ALARM-only routing.
--squeeze-dry-run proves deferral-not-drop + COVERAGE alarm. Not provisioned.
* fix(secrev): hide dependency-cve canary manifests from dependency-review
The canary fixtures intentionally pin known-vulnerable deps (jinja2 2.11.2,
lodash 4.17.15) so the checker has something to detect. GitHub's dependency
graph parsed those fixture manifests as real project deps, failing the
dependency-review PR gate (fail-on-severity: high). Store the manifests with a
.fixture suffix so the dependency graph ignores them; the --canary materializer
strips the suffix in its temp work area before scanning, so detection is
unchanged (still 2/2). No advisory allowlist, no change to the shared org
reusable workflow — the real gate stays strict for actual deps.
2026-06-18 15:08:58 -04:00
)
role_field( ) { echo " $1 " | cut -d'|' -f" $2 " ; }
# In SQUEEZE mode, assign non-zero costs so the shared cap is meaningful: the first role fits,
# the second cannot — proving deferral-not-drop deterministically regardless of real cost.
if [ " $SQUEEZE " -eq 1 ] ; then
ROLES = (
" compliance-drift| $CHECKERS_DIR /compliance-drift.sh|0.008|1 "
" dependency-cve| $CHECKERS_DIR /dependency-cve.sh|0.008|1 "
)
fi
2026-06-22 19:05:02 -04:00
# Operator role-skip (COORDINATOR_SKIP_ROLES="aws-posture,confluence-doc"): remove
# roles whose backing credentials are not provisioned (aws-posture needs IAM Roles
# Anywhere; confluence-doc needs the confluence-bot token). A skipped role is
# dropped from the registry entirely — never canaried, run, or ALARMed — so the
# nightly schedule only exercises credential-ready checkers. Empty/unset = run all.
if [ -n " ${ COORDINATOR_SKIP_ROLES :- } " ] ; then
declare -a _kept = ( )
for entry in " ${ ROLES [@] } " ; do
_name = " $( role_field " $entry " 1) "
case " , ${ COORDINATOR_SKIP_ROLES } , " in
*" , ${ _name } , " *) log " SKIP role ' $_name ' (COORDINATOR_SKIP_ROLES) " ; ;
*) _kept += ( " $entry " ) ; ;
esac
done
ROLES = ( ${ _kept [@]+ " ${ _kept [@] } " } )
fi
feat(secrev): Plane-1 Phase 2 — coordinator + dependency-cve checker (#16)
* fix(agent-team): read SLACK_CHANNEL_ID, aligning code with deploy doc + systemd
run-team.py read os.environ['SLACK_CHANNEL'] while DEPLOY-R720.md and the
coordinator systemd unit both document SLACK_CHANNEL_ID; the mismatch would
silently default the live Slack transport channel to empty. Standardize on
SLACK_CHANNEL_ID (decision locked 2026-06-18).
* feat(secrev): dependency-cve Plane-1 Tier-1 checker (OSV, ALARM-only)
Read-only checker on the Phase-0 substrate: scans $MIRROR_DIR mirrors for
pinned deps (requirements/poetry/Pipfile/package-lock/yarn/csproj across
PyPI/npm/NuGet), cross-refs OSV querybatch (live) or an offline advisory
fixture (canary). Mode-600 reports, ALARM-only, --canary asserts 2 planted
vulns (jinja2 2.11.2, lodash 4.17.15). Complements Dependabot. Not provisioned.
* feat(secrev): Plane-1 checker coordinator (shared budget, rotation, dedup)
Coordinator (design §5/§6.7) orchestrating Tier-1 checkers under one shared
budget ledger + versioned rotation/coverage state (atomic write + schema/hash/
logical-consistency integrity, park-on-corrupt). Canary-suite-first
(COMPLACENCY skip), fan-out under the shared cap with defer-not-drop, COVERAGE
alarm past MAX_CYCLE_NIGHTS, cross-checker dedup/prioritize, ALARM-only routing.
--squeeze-dry-run proves deferral-not-drop + COVERAGE alarm. Not provisioned.
* fix(secrev): hide dependency-cve canary manifests from dependency-review
The canary fixtures intentionally pin known-vulnerable deps (jinja2 2.11.2,
lodash 4.17.15) so the checker has something to detect. GitHub's dependency
graph parsed those fixture manifests as real project deps, failing the
dependency-review PR gate (fail-on-severity: high). Store the manifests with a
.fixture suffix so the dependency graph ignores them; the --canary materializer
strips the suffix in its temp work area before scanning, so detection is
unchanged (still 2/2). No advisory allowlist, no change to the shared org
reusable workflow — the real gate stays strict for actual deps.
2026-06-18 15:08:58 -04:00
# ==============================================================================
# DURABLE STATE (design §6.7): atomic write-temp-then-rename + integrity check.
# Integrity = schema_version match + stored content_hash + logical-consistency.
# content_hash is computed over the state WITHOUT its own hash field (canonical jq -S -c).
# ==============================================================================
state_hash( ) { # state_json_without_hash -> hex
if command -v sha256sum >/dev/null 2>& 1; then echo " $1 " | jq -S -cj 'del(.content_hash)' | sha256sum | cut -d' ' -f1
elif command -v shasum >/dev/null 2>& 1; then echo " $1 " | jq -S -cj 'del(.content_hash)' | shasum -a 256 | cut -d' ' -f1
else echo " $1 " | jq -S -cj 'del(.content_hash)' | cksum | cut -d' ' -f1; fi
}
atomic_write_state( ) { # path json
local path = " $1 " json = " $2 " h tmp
h = " $( state_hash " $json " ) "
json = " $( echo " $json " | jq -c --arg h " $h " '.content_hash=$h' ) "
tmp = " $( mktemp " ${ path } .XXXXXX " ) "
printf '%s\n' " $json " > " $tmp "
chmod 600 " $tmp " 2>/dev/null || true
mv -f " $tmp " " $path " # rename is atomic on the same filesystem
}
# Verify integrity; echo "ok" or a reason. schema + hash + logical-consistency.
verify_state( ) { # path expected_schema -> "ok" | reason
local path = " $1 " want = " $2 " json sv stored calc
json = " $( cat " $path " 2>/dev/null) " || { echo "unreadable" ; return ; }
echo " $json " | jq -e 'type=="object"' >/dev/null 2>& 1 || { echo "not-json-object" ; return ; }
sv = " $( echo " $json " | jq -r '.schema_version // empty' ) "
[ " $sv " = " $want " ] || { echo " schema-mismatch(got= ${ sv :- none } want= $want ) " ; return ; }
stored = " $( echo " $json " | jq -r '.content_hash // empty' ) "
[ -n " $stored " ] || { echo "missing-content-hash" ; return ; }
calc = " $( state_hash " $json " ) "
[ " $stored " = " $calc " ] || { echo "content-hash-mismatch" ; return ; }
echo "ok"
}
declare -a STATE_ALARMS = ( )
# --- Budget ledger: {schema_version, day, spend, content_hash}. New UTC day resets spend. ----
TOTAL_SPEND = "0"
load_budget_ledger( ) {
if [ -f " $BUDGET_LEDGER " ] ; then
local v; v = " $( verify_state " $BUDGET_LEDGER " " $SCHEMA_VERSION " ) "
if [ " $v " != "ok" ] ; then
STATE_ALARMS += ( " *STATE ALARM*: budget ledger corrupt ( $v ) — rebuilt for $UTC_DATE (rebuildable; a new UTC day resets spend). " )
log " budget ledger integrity FAIL: $v — rebuilding (park-on-corrupt, design §6.7) "
TOTAL_SPEND = "0"
else
local day; day = " $( jq -r '.day // empty' " $BUDGET_LEDGER " ) "
if [ " $day " = " $UTC_DATE " ] ; then TOTAL_SPEND = " $( jq -r '.spend // 0' " $BUDGET_LEDGER " ) "
else log " budget ledger from $day — new UTC day, resetting day spend " ; TOTAL_SPEND = "0" ; fi
fi
fi
log " budget: shared cap \$ $TOTAL_BUDGET_USD , day spend so far \$ $TOTAL_SPEND ( $UTC_DATE ) "
}
save_budget_ledger( ) {
atomic_write_state " $BUDGET_LEDGER " \
" $( jq -n --argjson sv " $SCHEMA_VERSION " --arg day " $UTC_DATE " --argjson sp " $TOTAL_SPEND " \
'{schema_version:$sv, day:$day, spend:$sp}' ) "
}
# --- Coordinator state: {schema_version, cycle_start, last_run:{role:date}, deferred:[], content_hash} ---
declare -A LAST_RUN = ( ) ; declare -a DEFERRED = ( ) ; CYCLE_START = " $UTC_DATE "
load_coord_state( ) {
if [ -f " $COORD_STATE " ] ; then
local v; v = " $( verify_state " $COORD_STATE " " $SCHEMA_VERSION " ) "
if [ " $v " != "ok" ] ; then
STATE_ALARMS += ( " *STATE ALARM*: coordinator state corrupt ( $v ) — rebuilt (rebuildable from report history; rotation restarts). " )
log " coordinator state integrity FAIL: $v — rebuilding (park-on-corrupt, design §6.7) "
return
fi
CYCLE_START = " $( jq -r '.cycle_start // empty' " $COORD_STATE " ) " ; [ -n " $CYCLE_START " ] || CYCLE_START = " $UTC_DATE "
while IFS = $'\t' read -r role date; do [ -n " $role " ] && LAST_RUN[ " $role " ] = " $date " ; done \
< <( jq -r '(.last_run // {}) | to_entries[] | "\(.key)\t\(.value)"' " $COORD_STATE " )
while IFS = read -r role; do [ -n " $role " ] && DEFERRED += ( " $role " ) ; done \
< <( jq -r '(.deferred // [])[]' " $COORD_STATE " )
fi
}
save_coord_state( ) {
local lr = "{}"
for role in " ${ !LAST_RUN[@] } " ; do
lr = " $( echo " $lr " | jq -c --arg k " $role " --arg v " ${ LAST_RUN [ $role ] } " '.[$k]=$v' ) "
done
local df = "[]"
if [ " ${# DEFERRED [@] } " -gt 0 ] ; then df = " $( printf '%s\n' " ${ DEFERRED [@] } " | jq -R . | jq -cs 'unique' ) " ; fi
atomic_write_state " $COORD_STATE " \
" $( jq -n --argjson sv " $SCHEMA_VERSION " --arg cs " $CYCLE_START " --argjson lr " $lr " --argjson df " $df " \
'{schema_version:$sv, cycle_start:$cs, last_run:$lr, deferred:$df}' ) "
}
load_budget_ledger
load_coord_state
# In the squeeze test, backdate cycle_start + a role's last_run so the COVERAGE ALARM trips
# deterministically (simulate enough elapsed cycles). This proves the COVERAGE path without
# waiting MAX_CYCLE_NIGHTS real days.
if [ " $SQUEEZE " -eq 1 ] ; then
OLD_DATE = " $( to_epoch " $UTC_DATE " ) " ; OLD_DATE = $(( OLD_DATE - ( MAX_CYCLE_NIGHTS + 2 ) * 86400 ))
# portable epoch -> YYYY-MM-DD
OLD_DATE_STR = " $( date -u -d " @ $OLD_DATE " +%Y-%m-%d 2>/dev/null || date -u -r " $OLD_DATE " +%Y-%m-%d 2>/dev/null || echo " $UTC_DATE " ) "
CYCLE_START = " $OLD_DATE_STR "
LAST_RUN[ "dependency-cve" ] = " $OLD_DATE_STR " # this role has not run in > MAX_CYCLE_NIGHTS
log " SQUEEZE: backdated cycle_start + dependency-cve last_run to $OLD_DATE_STR (> ${ MAX_CYCLE_NIGHTS } d) to trip COVERAGE "
fi
# ==============================================================================
# 1) CANARY SUITE FIRST — each role's checker --canary; a miss = COMPLACENCY ALARM + skip.
# ==============================================================================
declare -a ALARM_LINES = ( ) ; declare -A CANARY_OK = ( )
for entry in " ${ ROLES [@] } " ; do
role = " $( role_field " $entry " 1) " ; script = " $( role_field " $entry " 2) "
if [ ! -x " $script " ] && [ ! -f " $script " ] ; then
CANARY_OK[ " $role " ] = 0
ALARM_LINES += ( " *COMPLACENCY ALARM*: role ' $role ' checker missing ( $script ) — skipped. " )
continue
fi
set +e
bash " $script " --canary >" $REPORT_DIR / $role .canary.log " 2>& 1
rc = $?
set -e
if [ " $rc " -eq 0 ] ; then
CANARY_OK[ " $role " ] = 1; log " canary PASS: $role "
else
CANARY_OK[ " $role " ] = 0
ALARM_LINES += ( " *COMPLACENCY ALARM*: role ' $role ' canary FAILED (rc= $rc ) — skipped this run. See \` $REPORT_DIR / $role .canary.log\`. " )
log " canary FAIL: $role (rc= $rc ) — will SKIP this role "
fi
done
# --canary mode: assert every role's canary passed, then stop (offline; post nothing).
if [ " $CANARY " -eq 1 ] ; then
fail = 0
for entry in " ${ ROLES [@] } " ; do
role = " $( role_field " $entry " 1) "
[ " ${ CANARY_OK [ $role ] :- 0 } " -eq 1 ] || { echo " [coordinator] CANARY FAIL: role ' $role ' did not pass " >& 2; fail = 1; }
done
if [ " $fail " -ne 0 ] ; then
echo "[coordinator] CANARY SUITE FAILED — at least one role's canary did not pass." >& 2
exit 3
fi
log " canary suite PASS: all ${# ROLES [@] } role(s) green. "
exit 0
fi
# ==============================================================================
# 2) FAN-OUT under the SHARED cap. Order: DEFERRED roles first, then by rotation
# (oldest last_run first). A role whose est cost would exceed the ceiling is DEFERRED
# (recorded), never dropped. A degraded (canary-failed) role is skipped.
# ==============================================================================
# Build the run order: deferred-first, then never-run, then oldest-last_run.
order_roles( ) {
local entry role lr key
for entry in " ${ ROLES [@] } " ; do
role = " $( role_field " $entry " 1) "
# is it currently deferred?
if printf '%s\n' ${ DEFERRED [@]+ " ${ DEFERRED [@] } " } | grep -qxF " $role " ; then
echo " 0000000000| $role " ; continue
fi
lr = " ${ LAST_RUN [ $role ] :- } "
if [ -z " $lr " ] ; then key = "0000000001" ; else key = " $( to_epoch " $lr " ) " ; fi
echo " $key | $role "
done | sort -n | cut -d'|' -f2
}
declare -a NEW_DEFERRED = ( ) ; declare -a RAN_ROLES = ( )
declare -a RUN_REPORT_JSONS = ( )
while IFS = read -r role; do
[ -n " $role " ] || continue
# find the registry entry
entry = "" ; for e in " ${ ROLES [@] } " ; do [ " $( role_field " $e " 1) " = " $role " ] && entry = " $e " ; done
[ -n " $entry " ] || continue
script = " $( role_field " $entry " 2) " ; cost = " $( role_field " $entry " 3) "
# Skip degraded roles (canary failed) — never run silently degraded.
if [ " ${ CANARY_OK [ $role ] :- 0 } " -ne 1 ] ; then
log " skip $role : canary not green (already alarmed) "
continue
fi
# Budget headroom check: would this role's est cost push us over the SHARED ceiling?
projected = " $( jq -n --argjson s " $TOTAL_SPEND " --argjson c " $cost " '$s + $c' ) "
if jq -n --argjson p " $projected " --argjson cap " $TOTAL_BUDGET_USD " -e '$cap > 0 and $p > $cap' >/dev/null 2>& 1; then
NEW_DEFERRED += ( " $role " )
log " DEFER $role : est \$ $cost would exceed shared cap \$ $TOTAL_BUDGET_USD (spend \$ $TOTAL_SPEND ) — DEFERRED, not dropped "
ALARM_LINES += ( " * $role * DEFERRED: est \$ $cost over shared cap \$ $TOTAL_BUDGET_USD (day spend \$ $TOTAL_SPEND ). Will run next eligible night. " )
continue
fi
# Run the checker in --dry-run (the coordinator owns routing; checkers must not post).
# In SQUEEZE mode (synthetic acceptance test, may run on a box without $MIRROR_DIR) point the
# checker at its own fixture via --targets so the "ran" role succeeds deterministically; this
# keeps the deferral/COVERAGE proof self-contained. Normal runs use the real mirror set.
log " --- run role: $role (est \$ $cost ) --- "
set +e
if [ " $SQUEEZE " -eq 1 ] ; then
bash " $script " --dry-run --no-api --targets " $CHECKERS_DIR /fixtures/ $role /clean-repo " \
>" $REPORT_DIR / $role .run.log " 2>& 1
else
bash " $script " --dry-run >" $REPORT_DIR / $role .run.log " 2>& 1
fi
rc = $?
set -e
if [ " $rc " -ne 0 ] ; then
ALARM_LINES += ( " * $role *: checker run error (rc= $rc ). See \` $REPORT_DIR / $role .run.log\`. " )
log " $role run error rc= $rc (logged) — NOT collecting its report (avoid stale/partial findings) "
else
# Collect the checker's own report JSON (REPORT_ROOT/<role>/<date>/<role>.json) only on a
# clean run — a failed run could leave a stale report from an earlier (e.g. canary) pass,
# and folding that in would misattribute findings.
src = " $REPORT_ROOT / $role / $UTC_DATE / $role .json "
if [ -f " $src " ] ; then rj = " $REPORT_DIR / $role .json " ; cp -f " $src " " $rj " ; RUN_REPORT_JSONS += ( " $rj " ) ; fi
fi
# Account spend, record last_run, drop from deferred.
add_spend " $cost "
LAST_RUN[ " $role " ] = " $UTC_DATE "
RAN_ROLES += ( " $role " )
done < <( order_roles)
# New deferral set = roles deferred this run, plus any previously-deferred role we did NOT run.
for role in ${ DEFERRED [@]+ " ${ DEFERRED [@] } " } ; do
printf '%s\n' ${ RAN_ROLES [@]+ " ${ RAN_ROLES [@] } " } | grep -qxF " $role " && continue
printf '%s\n' ${ NEW_DEFERRED [@]+ " ${ NEW_DEFERRED [@] } " } | grep -qxF " $role " && continue
NEW_DEFERRED += ( " $role " )
done
DEFERRED = ( ${ NEW_DEFERRED [@]+ " ${ NEW_DEFERRED [@] } " } )
log " ran: ${ RAN_ROLES [*] :- none } | deferred: ${ DEFERRED [*] :- none } | day spend \$ $TOTAL_SPEND /\$ $TOTAL_BUDGET_USD "
# ==============================================================================
# 3) COVERAGE ALARM — any role whose last_run is older than MAX_CYCLE_NIGHTS days
# (or never run and deferred that long) is behind (design §5).
# ==============================================================================
NOW_EPOCH = " $( to_epoch " $UTC_DATE " ) "
for entry in " ${ ROLES [@] } " ; do
role = " $( role_field " $entry " 1) "
lr = " ${ LAST_RUN [ $role ] :- } "
if [ -z " $lr " ] ; then ref = " $CYCLE_START " ; else ref = " $lr " ; fi
age = $(( ( NOW_EPOCH - $( to_epoch " $ref " ) ) / 86400 ))
if [ " $age " -ge " $MAX_CYCLE_NIGHTS " ] ; then
ALARM_LINES += ( " *COVERAGE ALARM*: role ' $role ' not run in ${ age } d (last= ${ lr :- never , cycle since $CYCLE_START } , max $MAX_CYCLE_NIGHTS ). Deferred= $( printf '%s\n' ${ DEFERRED [@]+ " ${ DEFERRED [@] } " } | grep -qxF " $role " && echo yes || echo no) . Raise budget or check failures. " )
log " COVERAGE ALARM: $role age ${ age } d >= $MAX_CYCLE_NIGHTS "
fi
done
# Persist state (atomic + hashed). Even in dry-run we persist so rotation advances; the
# squeeze test runs dry, so guard: in SQUEEZE we do NOT persist (it is a synthetic scenario).
if [ " $SQUEEZE " -eq 0 ] ; then
save_budget_ledger
save_coord_state
else
log "SQUEEZE: synthetic scenario — NOT persisting state."
fi
# Fold any state-integrity alarms in.
for x in ${ STATE_ALARMS [@]+ " ${ STATE_ALARMS [@] } " } ; do ALARM_LINES += ( " $x " ) ; done
# ==============================================================================
# 4) COLLECT + DEDUP + PRIORITIZE across the run checkers' reports.
# DEDUP rule: same (repo + check + title) OR identical finding id -> one. Sort by severity.
# ==============================================================================
ALL_FINDINGS = "[]"
if [ " ${# RUN_REPORT_JSONS [@] } " -gt 0 ] ; then
ALL_FINDINGS = " $( jq -s '
[ .[ ] .findings[ ] ? ]
| unique_by( .id) # identical id -> one
| unique_by( [ .repo, .check, .title] ) # same repo+check+title -> one
| sort_by( { critical:0, high:1, medium:2, low:3, info:4, unverified:5} [ .severity] // 6 )
' "${RUN_REPORT_JSONS[@]}" 2>/dev/null || echo ' [ ] ' ) "
fi
N_FIND = " $( echo " $ALL_FINDINGS " | jq 'length' ) "
N_CRITHIGH = " $( echo " $ALL_FINDINGS " | jq '[.[]|select(.severity=="critical" or .severity=="high")] | length' ) "
declare -a CRITHIGH_LINES = ( )
while IFS = read -r line; do [ -n " $line " ] && CRITHIGH_LINES += ( " $line " ) ; done < <(
echo " $ALL_FINDINGS " | jq -r ' .[ ] | select ( .severity= = "critical" or .severity= = "high" )
| "*\(.repo)* [\(.severity)] \(.title)" ' )
# ==============================================================================
# 5) ASSEMBLE the combined coordinator report (JSON + text), mode 600.
# ==============================================================================
DEFERRED_JSON = "[]" ; [ " ${# DEFERRED [@] } " -gt 0 ] && DEFERRED_JSON = " $( printf '%s\n' " ${ DEFERRED [@] } " | jq -R . | jq -cs .) "
RAN_JSON = "[]" ; [ " ${# RAN_ROLES [@] } " -gt 0 ] && RAN_JSON = " $( printf '%s\n' " ${ RAN_ROLES [@] } " | jq -R . | jq -cs .) "
ALARMS_JSON = "[]" ; [ " ${# ALARM_LINES [@] } " -gt 0 ] && ALARMS_JSON = " $( printf '%s\n' " ${ ALARM_LINES [@] } " | jq -R . | jq -cs .) "
jq -n \
--arg ts " $UTC_STAMP " --arg org " $GH_ORG " \
--argjson cap " $TOTAL_BUDGET_USD " --argjson spend " $TOTAL_SPEND " \
--argjson ran " $RAN_JSON " --argjson deferred " $DEFERRED_JSON " \
--argjson findings " $ALL_FINDINGS " --argjson alarms " $ALARMS_JSON " \
' { coordinator:"plane1" , generated:$ts , org:$org ,
shared_budget_usd:$cap , day_spend_usd:$spend ,
ran_roles:$ran , deferred_roles:$deferred ,
finding_count:( $findings | length) ,
crit_high:( [ $findings [ ] | select ( .severity= = "critical" or .severity= = "high" ) ] | length) ,
findings:$findings , alarms:$alarms } ' > " $REPORT_JSON "
{
echo " plane-1 coordinator report — $UTC_STAMP "
echo " org= $GH_ORG shared_cap=\$ $TOTAL_BUDGET_USD day_spend=\$ $TOTAL_SPEND "
echo " ran: ${ RAN_ROLES [*] :- none } "
echo " deferred (NOT dropped): ${ DEFERRED [*] :- none } "
echo " findings: $N_FIND ( $N_CRITHIGH crit/high) "
echo
echo " $ALL_FINDINGS " | jq -r '.[] | "• [\(.severity)] \(.repo): \(.title)"'
if [ " ${# ALARM_LINES [@] } " -gt 0 ] ; then
echo; echo "alarms:" ; printf ' - %s\n' " ${ ALARM_LINES [@] } "
fi
} > " $REPORT_TXT "
chmod 600 " $REPORT_JSON " " $REPORT_TXT " 2>/dev/null || true
log " report: $REPORT_JSON ( $N_FIND finding(s), ${# ALARM_LINES [@] } alarm line(s)) "
# ==============================================================================
# 6) ROUTE (ALARM-only, D3): confirmed crit/high OR any alarm line -> Slack ALARM;
# everything else -> the mode-600 report only; a fully clean run posts NOTHING.
# ==============================================================================
ALARM = 0
[ " $N_CRITHIGH " -gt 0 ] && ALARM = 1
[ " ${# ALARM_LINES [@] } " -gt 0 ] && ALARM = 1
# Squeeze acceptance: print the proof lines explicitly to stdout.
if [ " $SQUEEZE " -eq 1 ] ; then
echo "=== SQUEEZE ACCEPTANCE (Phase-2) ==="
echo " DEFERRED (not dropped): ${ DEFERRED [*] :- none } "
printf '%s\n' ${ ALARM_LINES [@]+ " ${ ALARM_LINES [@] } " } | grep -E 'COVERAGE ALARM|DEFERRED' || true
echo "===================================="
fi
if [ " $ALARM " -ne 1 ] ; then
log "clean run — no crit/high findings, no alarm conditions. Posting NOTHING (ALARM-only policy)."
exit 0
fi
ALARM_BODY = ""
[ " ${# CRITHIGH_LINES [@] } " -gt 0 ] && ALARM_BODY = " $( printf '%s\n' " ${ CRITHIGH_LINES [@] } " | sed 's/^/• /' ) "
META_BODY = " $( printf '%s\n' ${ ALARM_LINES [@]+ " ${ ALARM_LINES [@] } " } | sed 's/^/• /' ) "
SLACK_TEXT = " :satellite_antenna: *Sea Haven Plane-1 coordinator — ALARM* ( $UTC_STAMP )
ran: ${ RAN_ROLES [*] :- none } · deferred: ${ DEFERRED [*] :- none } · spend \$ $TOTAL_SPEND /\$ $TOTAL_BUDGET_USD
$N_CRITHIGH confirmed crit/high finding( s) :
$ALARM_BODY
coordination alarms:
$META_BODY
Combined report ( mode 600) : \` $REPORT_JSON \` ( on R720) "
SLACK_TEXT = " $( echo " $SLACK_TEXT " | redact) "
echo " $SLACK_TEXT " >& 2
if [ " $DRY_RUN " -eq 1 ] ; then
log "DRY-RUN: alarm composed but NOT posted (routing dry-run, design §7 Phase 2)."
exit 0
fi
post_slack_alarm " $SLACK_TEXT "
exit 0
# ==============================================================================
# PROVISIONING (NOT DONE HERE — gated, Phase 6):
# - No systemd unit / timer is installed by this script. Wiring it into the live
# sea-haven-secrev schedule (or a sibling timer) is provisioning and is gated.
# - This coordinator runs ONLY the Plane-1 Tier-1 checkers (compliance-drift,
# dependency-cve). doc-drift / aws-posture / planner / fixer are later phases.
# - It does NOT re-clone (checkers reuse $MIRROR_DIR); a checker's own --refresh is the
# only network path and is not invoked here.
# - It does NOT touch agent_team/ or agent-team/, and installs no systemd units.
# - Confluence + project_r720_agent_team memory updates are docs-as-you-go obligations.
# ==============================================================================