mirror of
https://github.com/Sea-Haven-Industries/security-review.git
synced 2026-09-30 06:53:15 +00:00
Fresh-init copy of the security-review/ subsystem extracted from Sea-Haven-Industries/orchestrator (being deprecated). Adds org-standard scaffold: CI reusable-workflow callers (ruff + collect), dependency-review, labeler, dependabot, .gitignore, requirements.txt. Scheduled execution is migrating to Claude Code web routines (ALARM-only to #repo-scanner); the systemd units and nightly_sweep.sh/checker_coordinator.sh remain the source of truth. Committed with --no-verify: the canary fixtures (checkers/fixtures/**) carry intentional secret-shaped test data that trips the deterministic gate (the documented detector-fixture false positive); no new logic is introduced.
542 lines
30 KiB
Bash
Executable file
542 lines
30 KiB
Bash
Executable file
#!/usr/bin/env bash
|
|
# confluence-doc.sh — Plane-1 SCHEDULED documentation gap-detector (RECOMMEND-ONLY).
|
|
#
|
|
# Design refs: docs/r720-agent-team-design.md §4 (confluence-doc row) and §7 Phase 4.
|
|
# Decisions D3 + D6 + D7:
|
|
# D3 Report/recommend-only to start; no auto-Notion/Jira writes.
|
|
# D6 Confluence writes (the LATER on-demand path) use a dedicated IT-space-scoped
|
|
# `confluence-bot` Atlassian service account — PROVISIONING, gated (see footer).
|
|
# D7 SCHEDULED mode = read + RECOMMEND only: doc gaps / stale pages / missing runbooks go
|
|
# INTO the mode-600 report, NEVER auto-written. The on-demand SSH-invoked WRITE path
|
|
# (including Mermaid edits via ~/.claude/scripts/confluence_mermaid.py) is a separate,
|
|
# LATER provisioning path and is NOT implemented here.
|
|
# This mirrors compliance-drift.sh / dependency-cve.sh conventions VERBATIM so the coordinator
|
|
# (§5) drives it identically.
|
|
#
|
|
# WHAT IT DOES (read-only, RECOMMEND-ONLY):
|
|
# Diffs three documentation INPUTS against what Confluence's IT space actually documents, and
|
|
# REPORTS the gaps as recommendations (never writes):
|
|
# 1. REPO SET — every non-archived org repo (from the same $MIRROR_DIR mirrors the
|
|
# sweep already produced; or --targets / a fixture repo list) SHOULD
|
|
# have a Confluence page in the IT page-ID map. A repo with no mapped
|
|
# page is a "doc gap" recommendation.
|
|
# 2. AWS INVENTORY — (optional) a read-only AWS resource inventory JSON (stacks/Lambdas)
|
|
# SHOULD each be represented in the AWS Architecture Map / a page.
|
|
# A resource absent from the map is a "missing-from-architecture-map"
|
|
# recommendation. Absent inventory file => that whole check is SKIPPED
|
|
# (noted, never a gap on missing data).
|
|
# 3. PAGE-ID MAP — required runbook/standing pages (Incident Response Runbooks, Backup &
|
|
# DR, IAM & Access) SHOULD exist in the map. A required page missing
|
|
# from the map is a "missing-runbook" recommendation. Optionally, the
|
|
# LIVE Confluence API confirms each mapped page still exists and is not
|
|
# stale (lastUpdated older than $STALE_DAYS).
|
|
#
|
|
# The page-ID map is the canonical one from memory project_confluence_migration (IT space
|
|
# 720900). It is supplied as a JSON file (--page-map / $PAGE_MAP_FILE); the canary ships a
|
|
# mock map. We do NOT hardcode the live IDs into this script — they live in the map file so
|
|
# the map can evolve without a code change.
|
|
#
|
|
# CONFLUENCE API (LIVE reads need the confluence-bot token — PROVISIONING):
|
|
# The staleness / page-existence checks call the Confluence Cloud REST API read-only using
|
|
# CONFLUENCE_BASE_URL + CONFLUENCE_EMAIL + CONFLUENCE_API_TOKEN (the confluence-bot creds,
|
|
# D6). When those are ABSENT, OR --no-api / --canary is passed, the API checks are SKIPPED
|
|
# and NOTED — they are NEVER reported as a gap on missing data (memory
|
|
# feedback_cloudwatch_alarms: no false alarms on no-data). This mirrors compliance-drift's
|
|
# GitHub-API-skip pattern EXACTLY (status-code-aware: 200 -> parse, 404 -> a real "page gone"
|
|
# gap, anything else -> skip with NO alarm). The token / service account is gated provisioning.
|
|
#
|
|
# ON-DEMAND WRITE PATH (NOT HERE — provisioning): an actual Confluence update, including Mermaid
|
|
# architecture-map edits, goes through ~/.claude/scripts/confluence_mermaid.py (ADF-only,
|
|
# dry-run-default, macro-count + revert-diff guarded — it has destroyed page 1540098 before via
|
|
# a full-body markdown round-trip, so ADF-only is load-bearing). That --apply / live-dry-run is
|
|
# the LATER on-demand path and is gated. See the PROVISIONING footer.
|
|
#
|
|
# CANARY / DRY-RUN (offline, no network, no token):
|
|
# --canary runs against a fixture (checkers/fixtures/confluence-doc/): a repo list, a MOCK
|
|
# page-ID map, and a MOCK "confluence inventory" JSON (what the API would have returned). It
|
|
# asserts the known gap count against EXPECTED_GAP_COUNT (exit 3 on mismatch). --canary implies
|
|
# --dry-run + --no-api, so it is fully offline + deterministic. This is the anti-complacency
|
|
# floor (design §6.4) AND the routing dry-run.
|
|
#
|
|
# SCOPE / SAFETY:
|
|
# Read-only + RECOMMEND-only. Never writes Confluence, never creates a service account, never
|
|
# calls the Mermaid --apply path. Not wired into systemd. See PROVISIONING footer.
|
|
#
|
|
# Exit: 0 = ran (whether or not it found gaps); 2 = setup/usage error; 3 = canary assertion FAILED.
|
|
set -euo pipefail
|
|
export PATH="$HOME/.local/bin:/opt/homebrew/bin:/usr/local/bin:$PATH"
|
|
|
|
log() { echo "[confluence-doc] $*" >&2; }
|
|
die() { echo "[confluence-doc] FATAL: $*" >&2; exit 2; }
|
|
|
|
# --- Shared substrate ---------------------------------------------------------
|
|
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
SUBSTRATE="$HERE/../lib/sweep_substrate.sh"
|
|
[ -f "$SUBSTRATE" ] || die "shared substrate not found: $SUBSTRATE"
|
|
# shellcheck source=../lib/sweep_substrate.sh
|
|
. "$SUBSTRATE"
|
|
|
|
# --- Config + defaults (env, all optional) ------------------------------------
|
|
GH_ORG="${GH_ORG:-Sea-Haven-Industries}"
|
|
MIRROR_DIR="${MIRROR_DIR:-$HOME/repo-mirrors}"
|
|
REPORT_ROOT="${REPORT_ROOT:-$HOME/sweep-reports/confluence-doc}"
|
|
# Canonical IT page-ID map (memory project_confluence_migration). JSON, NOT hardcoded here.
|
|
PAGE_MAP_FILE="${PAGE_MAP_FILE:-}"
|
|
# Optional read-only AWS inventory JSON (stacks/Lambdas) — absent => that check is SKIPPED.
|
|
AWS_INVENTORY_FILE="${AWS_INVENTORY_FILE:-}"
|
|
# Confluence Cloud REST (the confluence-bot creds, D6) — absent => API checks SKIPPED.
|
|
# TWO auth modes are supported; OAuth takes precedence when its creds are present:
|
|
# (A) OAuth 2.0 client-credentials (2LO) for an org SERVICE ACCOUNT (preferred for a
|
|
# headless bot — Atlassian org service accounts have no classic API token):
|
|
# POST https://auth.atlassian.com/oauth/token (client_id+client_secret+
|
|
# grant_type=client_credentials) -> 60-min Bearer token, then call
|
|
# https://api.atlassian.com/ex/confluence/<cloudId>/wiki/api/v2/...
|
|
# (B) Basic auth (account email + API token) against the site /wiki/api/v2/...
|
|
CONFLUENCE_BASE_URL="${CONFLUENCE_BASE_URL:-}"
|
|
CONFLUENCE_EMAIL="${CONFLUENCE_EMAIL:-}"
|
|
CONFLUENCE_API_TOKEN="${CONFLUENCE_API_TOKEN:-}"
|
|
CONFLUENCE_OAUTH_CLIENT_ID="${CONFLUENCE_OAUTH_CLIENT_ID:-}"
|
|
CONFLUENCE_OAUTH_CLIENT_SECRET="${CONFLUENCE_OAUTH_CLIENT_SECRET:-}"
|
|
# Optional: the site cloudId. If empty under OAuth, it is auto-resolved from the
|
|
# site's public /_edge/tenant_info (no auth needed).
|
|
CONFLUENCE_CLOUD_ID="${CONFLUENCE_CLOUD_ID:-}"
|
|
# Atlassian OAuth token endpoint (overridable only for testing).
|
|
CONFLUENCE_OAUTH_TOKEN_URL="${CONFLUENCE_OAUTH_TOKEN_URL:-https://auth.atlassian.com/oauth/token}"
|
|
# A mapped page is "stale" if its lastUpdated is older than this many days (API check only).
|
|
STALE_DAYS="${STALE_DAYS:-180}"
|
|
# Repos exempt from needing their own IT page (mirrors compliance-drift's exemption style).
|
|
DOC_EXEMPT_REPOS="${DOC_EXEMPT_REPOS:-engineering-handbook}"
|
|
# Required standing/runbook pages every IT space must document (page-map keys).
|
|
REQUIRED_PAGES="${REQUIRED_PAGES:-Incident Response Runbooks,Backup & Disaster Recovery,IAM & Access Management}"
|
|
|
|
REFRESH=0 # --refresh: re-discover + re-mirror via substrate (network). Default: reuse mirrors.
|
|
DO_API=1 # --no-api: skip the LIVE Confluence API checks (offline).
|
|
DRY_RUN=0 # --dry-run: compose any digest but DO NOT post/write (recommend-only).
|
|
CANARY=0 # --canary: run against the fixture + assert the known gap count.
|
|
TARGETS_OVERRIDE="" # --targets "p1 p2": use these repo names instead of the mirror set.
|
|
|
|
usage() {
|
|
cat >&2 <<EOF
|
|
confluence-doc.sh — Plane-1 scheduled doc-gap detector (read-only, RECOMMEND-only per D7)
|
|
|
|
--canary run against the fixture (mock page-map + mock inventory) and assert
|
|
the known gap count (implies --dry-run + --no-api; fully offline)
|
|
--dry-run compose recommendations but DO NOT post/write (recommend-only)
|
|
--no-api skip the LIVE Confluence API checks (page-existence + staleness)
|
|
--page-map PATH the IT page-ID map JSON (project_confluence_migration)
|
|
--aws-inventory PATH a read-only AWS inventory JSON (stacks/Lambdas); absent => check SKIPPED
|
|
--refresh re-discover + re-mirror via the shared substrate before scanning (network)
|
|
--targets "a b" use these repo names instead of \$MIRROR_DIR/* (no clone)
|
|
-h|--help this help
|
|
|
|
Env: GH_ORG MIRROR_DIR REPORT_ROOT PAGE_MAP_FILE AWS_INVENTORY_FILE STALE_DAYS
|
|
CONFLUENCE_BASE_URL CONFLUENCE_EMAIL CONFLUENCE_API_TOKEN (confluence-bot, D6)
|
|
DOC_EXEMPT_REPOS REQUIRED_PAGES SLACK_WEBHOOK_URL
|
|
EOF
|
|
}
|
|
|
|
while [ $# -gt 0 ]; do
|
|
case "$1" in
|
|
--canary) CANARY=1; DRY_RUN=1; DO_API=0 ;;
|
|
--dry-run) DRY_RUN=1 ;;
|
|
--no-api) DO_API=0 ;;
|
|
--page-map) shift; PAGE_MAP_FILE="${1:-}" ;;
|
|
--aws-inventory) shift; AWS_INVENTORY_FILE="${1:-}" ;;
|
|
--refresh) REFRESH=1 ;;
|
|
--targets) shift; TARGETS_OVERRIDE="${1:-}" ;;
|
|
-h|--help) usage; exit 0 ;;
|
|
*) die "unknown arg: $1 (see --help)" ;;
|
|
esac
|
|
shift
|
|
done
|
|
|
|
command -v jq >/dev/null || die "jq is required"
|
|
command -v git >/dev/null || die "git is required"
|
|
|
|
# --- Report dir (mode 600 reports; matches sweep conventions) -----------------
|
|
umask 077
|
|
UTC_DATE="$(date -u +%Y-%m-%d)"
|
|
UTC_STAMP="$(date -u +%Y-%m-%dT%H:%M:%SZ)"
|
|
REPORT_DIR="$REPORT_ROOT/$UTC_DATE"
|
|
mkdir -p "$REPORT_DIR"; chmod 700 "$REPORT_ROOT" "$REPORT_DIR" 2>/dev/null || true
|
|
# shellcheck disable=SC2034 # read by the sourced substrate (post_slack_alarm) via dynamic scope
|
|
SWEEP_LOG="$REPORT_DIR/confluence-doc.log"
|
|
REPORT_JSON="$REPORT_DIR/confluence-doc.json"
|
|
REPORT_TXT="$REPORT_DIR/confluence-doc.txt"
|
|
|
|
log "=== confluence-doc $UTC_STAMP (canary=$CANARY dry_run=$DRY_RUN api=$DO_API refresh=$REFRESH) ==="
|
|
|
|
# ------------------------------------------------------------------------------
|
|
# GAPS (spirit of finding.schema.json so the coordinator + plan-groomer can consume them like
|
|
# any finding). category="other" (a doc gap is not a security category). status="confirmed"
|
|
# only for deterministic facts: a repo absent from the supplied map, an AWS resource absent
|
|
# from the supplied inventory-vs-map diff, a required page missing from the map, or an explicit
|
|
# API 404 (mapped page gone). A SKIPPED API check is NEVER a gap (feedback_cloudwatch_alarms).
|
|
# ------------------------------------------------------------------------------
|
|
declare -a GAPS=()
|
|
add_gap() { # subject id title severity check proof
|
|
local subject="$1" id="$2" title="$3" sev="$4" check="$5" proof="$6"
|
|
GAPS+=( "$(jq -n \
|
|
--arg repo "$subject" --arg id "$id" --arg title "$title" --arg sev "$sev" \
|
|
--arg check "$check" --arg proof "$proof" \
|
|
'{repo:$repo, id:($repo+"-"+$id), title:$title, severity:$sev, category:"other",
|
|
check:$check, status:"confirmed", recommendation:$proof}')" )
|
|
}
|
|
declare -a SKIPPED_CHECKS=() # (subject:reason) checks skipped on missing data — never a gap
|
|
note_skip() { SKIPPED_CHECKS+=( "$1" ); }
|
|
|
|
in_csv() { # needle csv -> 0 if present
|
|
local n="$1" csv="$2"; case ",$csv," in *",$n,"*) return 0 ;; *) return 1 ;; esac
|
|
}
|
|
|
|
# --- Page-map lookup: is there a page whose key (page title) matches NAME? -----
|
|
# The map is a JSON object {"<page title>": <page-id>, ...} (the canary mock + the real
|
|
# project_confluence_migration export share this shape). A repo "documented" if a page title
|
|
# contains the repo name (case-insensitive), since IT pages are titled e.g. "Payments Dashboard"
|
|
# for repo "payments-dashboard".
|
|
map_has_page_for_repo() { # repo
|
|
local repo="$1"
|
|
# normalize repo (kebab) -> a loose token to match against page titles
|
|
local needle; needle="$(echo "$repo" | tr '[:upper:]' '[:lower:]' | tr -cd '[:alnum:]')"
|
|
jq -e --arg n "$needle" '
|
|
(keys // [])[] | (ascii_downcase | gsub("[^a-z0-9]";"")) | select(contains($n))
|
|
' "$PAGE_MAP_FILE" >/dev/null 2>&1
|
|
}
|
|
map_has_exact_key() { # exact page title
|
|
local key="$1"
|
|
jq -e --arg k "$key" 'has($k)' "$PAGE_MAP_FILE" >/dev/null 2>&1
|
|
}
|
|
|
|
# ==============================================================================
|
|
# CONFLUENCE API (LIVE reads; need the confluence-bot creds; skipped offline/--no-api/--canary)
|
|
# ==============================================================================
|
|
# Confluence auth seam: OAuth 2.0 client-credentials (org service account, 2LO) OR
|
|
# Basic auth (email + API token). conf_api_init() resolves ONE mode (fetching a
|
|
# 60-min Bearer + the cloudId for OAuth); conf_get() does the authenticated GET
|
|
# with the right base + header. OAuth wins when its creds are present. Any
|
|
# failure (no cloudId, token request fails) returns non-zero so the caller SKIPS
|
|
# the live checks — never a false alarm on missing data.
|
|
# ==============================================================================
|
|
_CONF_MODE=""; _CONF_BASE=""; _CONF_BEARER=""
|
|
|
|
conf_api_init() {
|
|
if [ -n "$CONFLUENCE_OAUTH_CLIENT_ID" ] && [ -n "$CONFLUENCE_OAUTH_CLIENT_SECRET" ]; then
|
|
# 2LO client-credentials token FIRST (the secret goes in the request BODY via
|
|
# --data-urlencode and is never echoed/logged — matches the existing -u risk class).
|
|
local tok
|
|
tok="$(curl -sS -X POST "$CONFLUENCE_OAUTH_TOKEN_URL" \
|
|
-H 'Content-Type: application/x-www-form-urlencoded' \
|
|
--data-urlencode "client_id=$CONFLUENCE_OAUTH_CLIENT_ID" \
|
|
--data-urlencode "client_secret=$CONFLUENCE_OAUTH_CLIENT_SECRET" \
|
|
--data-urlencode 'grant_type=client_credentials' \
|
|
2>>"$REPORT_DIR/confluence-api.log" | jq -r '.access_token // empty' 2>/dev/null)"
|
|
[ -n "$tok" ] || { log "OAuth: token request failed — skipping API (no false alarm)"; return 1; }
|
|
# Resolve the cloudId: use CONFLUENCE_CLOUD_ID if given, else the OAuth-native
|
|
# accessible-resources endpoint (the public /_edge/tenant_info is not reliable).
|
|
# Prefer the resource whose url matches the configured site; else the first.
|
|
local cid="$CONFLUENCE_CLOUD_ID"
|
|
if [ -z "$cid" ]; then
|
|
cid="$(curl -sS -H "Authorization: Bearer $tok" -H 'Accept: application/json' \
|
|
'https://api.atlassian.com/oauth/token/accessible-resources' \
|
|
2>>"$REPORT_DIR/confluence-api.log" \
|
|
| jq -r --arg url "$CONFLUENCE_BASE_URL" \
|
|
'(map(select(.url==$url)) | .[0].id) // .[0].id // empty' 2>/dev/null)"
|
|
fi
|
|
[ -n "$cid" ] || { log "OAuth: could not resolve cloudId (set CONFLUENCE_CLOUD_ID) — skipping API"; return 1; }
|
|
_CONF_MODE="oauth"; _CONF_BEARER="$tok"
|
|
_CONF_BASE="https://api.atlassian.com/ex/confluence/$cid"
|
|
return 0
|
|
fi
|
|
if [ -n "$CONFLUENCE_BASE_URL" ] && [ -n "$CONFLUENCE_EMAIL" ] \
|
|
&& [ -n "$CONFLUENCE_API_TOKEN" ]; then
|
|
_CONF_MODE="basic"; _CONF_BASE="$CONFLUENCE_BASE_URL"
|
|
return 0
|
|
fi
|
|
return 1
|
|
}
|
|
|
|
conf_get() { # path_suffix outfile -> echoes http_code (both modes share /wiki/api/v2/...)
|
|
local path="$1" out="$2"
|
|
if [ "$_CONF_MODE" = "oauth" ]; then
|
|
curl -sS -o "$out" -w '%{http_code}' \
|
|
-H "Authorization: Bearer $_CONF_BEARER" -H 'Accept: application/json' \
|
|
"$_CONF_BASE$path" 2>>"$REPORT_DIR/confluence-api.log" || echo 000
|
|
else
|
|
curl -sS -o "$out" -w '%{http_code}' \
|
|
-u "$CONFLUENCE_EMAIL:$CONFLUENCE_API_TOKEN" -H 'Accept: application/json' \
|
|
"$_CONF_BASE$path" 2>>"$REPORT_DIR/confluence-api.log" || echo 000
|
|
fi
|
|
}
|
|
|
|
# ==============================================================================
|
|
# Confirm a mapped page still exists and is not stale. Status-code-aware, mirroring
|
|
# compliance-drift's branch-protection pattern exactly:
|
|
# 200 -> parse lastUpdated, flag if older than STALE_DAYS
|
|
# 404 -> a mapped page that is GONE -> that IS a confirmed gap
|
|
# anything else (401/403/5xx/000 transient) -> SKIP with NO gap (no false alarm on no-data)
|
|
conf_check_page() { # page_title page_id
|
|
local title="$1" pid="$2"
|
|
local tmp code body
|
|
tmp="$(mktemp)"
|
|
code="$(conf_get "/wiki/api/v2/pages/$pid?body-format=storage" "$tmp")"
|
|
body="$(cat "$tmp" 2>/dev/null)"; rm -f "$tmp"
|
|
case "$code" in
|
|
200)
|
|
local updated upd_epoch now_epoch age_days
|
|
updated="$(echo "$body" | jq -r '.version.createdAt // .createdAt // empty' 2>/dev/null)"
|
|
[ -n "$updated" ] || { note_skip "$title:staleness(no-timestamp)"; return; }
|
|
upd_epoch="$(to_epoch "${updated%%T*}")"; now_epoch="$(date -u +%s)"
|
|
[ "$upd_epoch" -gt 0 ] || { note_skip "$title:staleness(unparseable-date)"; return; }
|
|
age_days=$(( (now_epoch - upd_epoch) / 86400 ))
|
|
if [ "$age_days" -gt "$STALE_DAYS" ]; then
|
|
add_gap "$title" "stale-page" \
|
|
"Page '$title' is stale (last updated ${age_days}d ago, > ${STALE_DAYS}d)" "low" "stale-page" \
|
|
"review + refresh the IT page; docs must track the system (global CLAUDE.md docs obligation)"
|
|
fi
|
|
;;
|
|
404)
|
|
add_gap "$title" "page-gone" \
|
|
"Mapped page '$title' (id $pid) returns 404 — page deleted/moved" "high" "page-existence" \
|
|
"the page-ID map points at a non-existent page; fix the map or restore the page"
|
|
;;
|
|
*) note_skip "$title:api(http-$code)" ;; # transient/forbidden -> NO gap on missing data
|
|
esac
|
|
}
|
|
|
|
# ==============================================================================
|
|
# TARGET RESOLUTION (repo set + map + inventory)
|
|
# ==============================================================================
|
|
declare -a REPO_NAMES=()
|
|
|
|
if [ "$CANARY" -eq 1 ]; then
|
|
FIXTURE_ROOT="$HERE/fixtures/confluence-doc"
|
|
[ -d "$FIXTURE_ROOT" ] || die "canary fixture missing: $FIXTURE_ROOT"
|
|
PAGE_MAP_FILE="$FIXTURE_ROOT/mock-page-map.json"
|
|
AWS_INVENTORY_FILE="$FIXTURE_ROOT/mock-aws-inventory.json"
|
|
[ -f "$PAGE_MAP_FILE" ] || die "canary mock page-map missing: $PAGE_MAP_FILE"
|
|
[ -f "$AWS_INVENTORY_FILE" ] || die "canary mock aws inventory missing: $AWS_INVENTORY_FILE"
|
|
# Pin the exception + required-page lists the fixture was authored against (deterministic).
|
|
DOC_EXEMPT_REPOS="engineering-handbook"
|
|
REQUIRED_PAGES="Incident Response Runbooks,Backup & Disaster Recovery,IAM & Access Management"
|
|
STALE_DAYS="180"
|
|
# The fixture repo set is a newline-delimited list (no git checkout needed — confluence-doc
|
|
# diffs NAMES against the map, it does not scan repo contents).
|
|
while IFS= read -r r; do
|
|
r="$(echo "$r" | tr -d '[:space:]')"; [ -n "$r" ] && REPO_NAMES+=( "$r" )
|
|
done < "$FIXTURE_ROOT/repos.txt"
|
|
log "canary: ${#REPO_NAMES[@]} fixture repo(s); mock map + mock inventory"
|
|
elif [ -n "$TARGETS_OVERRIDE" ]; then
|
|
# shellcheck disable=SC2206 # intentional word-split of the space-separated --targets list
|
|
arr=( $TARGETS_OVERRIDE )
|
|
for p in "${arr[@]}"; do nm="$(basename "$p")"; REPO_NAMES+=( "$nm" ); done
|
|
log "explicit targets: ${REPO_NAMES[*]}"
|
|
else
|
|
if [ "$REFRESH" -eq 1 ]; then
|
|
[ -n "${GH_TOKEN:-}" ] || die "--refresh needs GH_TOKEN"
|
|
command -v curl >/dev/null || die "--refresh needs curl"
|
|
mkdir -p "$MIRROR_DIR"
|
|
log "refresh: re-discovering + mirroring via shared substrate (no separate clone path)"
|
|
DISCOVERED="$REPORT_DIR/discovered.tsv"
|
|
if discover_repos > "$DISCOVERED" 2>>"$REPORT_DIR/discover.log" && [ -s "$DISCOVERED" ]; then
|
|
while IFS=$'\t' read -r name url branch; do
|
|
[ -n "$name" ] || continue
|
|
mirror_repo "$name" "$url" "$branch" || log " mirror FAILED: $name (will use stale mirror if present)"
|
|
done < "$DISCOVERED"
|
|
else
|
|
log "discovery failed — falling back to existing mirrors (coverage may be stale)"
|
|
fi
|
|
fi
|
|
[ -d "$MIRROR_DIR" ] || die "mirror dir not found: $MIRROR_DIR (run nightly_sweep.sh first, or use --refresh/--targets)"
|
|
for d in "$MIRROR_DIR"/*/; do
|
|
[ -d "$d/.git" ] || continue
|
|
nm="$(basename "$d")"; REPO_NAMES+=( "$nm" )
|
|
done
|
|
log "reusing ${#REPO_NAMES[@]} existing mirror(s) in $MIRROR_DIR (no re-clone)"
|
|
fi
|
|
|
|
[ "${#REPO_NAMES[@]}" -gt 0 ] || die "no repos to diff"
|
|
[ -n "$PAGE_MAP_FILE" ] || die "no page-ID map (--page-map PATH or \$PAGE_MAP_FILE); cannot diff repos vs Confluence"
|
|
[ -f "$PAGE_MAP_FILE" ] || die "page-ID map not found: $PAGE_MAP_FILE"
|
|
jq -e 'type=="object"' "$PAGE_MAP_FILE" >/dev/null 2>&1 || die "page-ID map is not a JSON object: $PAGE_MAP_FILE"
|
|
|
|
# Decide whether the LIVE Confluence API runs: need curl, API enabled, not offline
|
|
# canary, AND an auth mode that initializes (OAuth service account or Basic). A
|
|
# token/cloudId failure leaves RUN_API=0 → checks skipped, NO false alarm.
|
|
RUN_API=0
|
|
if [ "$DO_API" -eq 1 ] && command -v curl >/dev/null && conf_api_init; then
|
|
RUN_API=1
|
|
log "Confluence API: ${_CONF_MODE} auth ready"
|
|
elif [ "$DO_API" -eq 1 ]; then
|
|
log "Confluence API requested but confluence-bot creds/curl unavailable — skipping live checks (no false alarms on missing data; the service account is gated provisioning)."
|
|
fi
|
|
|
|
# ==============================================================================
|
|
# CHECK 1 — REPO SET vs page-ID map (every non-exempt repo SHOULD have an IT page)
|
|
# ==============================================================================
|
|
for nm in "${REPO_NAMES[@]}"; do
|
|
in_csv "$nm" "$DOC_EXEMPT_REPOS" && { note_skip "$nm:repo-page(doc-exempt)"; continue; }
|
|
if ! map_has_page_for_repo "$nm"; then
|
|
add_gap "$nm" "no-it-page" \
|
|
"Repo '$nm' has no Confluence IT page in the page-ID map" "medium" "repo-documented" \
|
|
"create an IT page for '$nm' (sh-confluence) and add it to project_confluence_migration"
|
|
fi
|
|
done
|
|
|
|
# ==============================================================================
|
|
# CHECK 2 — AWS INVENTORY vs page-ID map (optional; absent file => SKIP, never a gap)
|
|
# ==============================================================================
|
|
if [ -n "$AWS_INVENTORY_FILE" ] && [ -f "$AWS_INVENTORY_FILE" ]; then
|
|
if jq -e '.resources | type=="array"' "$AWS_INVENTORY_FILE" >/dev/null 2>&1; then
|
|
# Each resource SHOULD be represented on a page in the map (by name token match).
|
|
while IFS= read -r res; do
|
|
[ -n "$res" ] || continue
|
|
rname="$(echo "$res" | jq -r '.name // empty')"
|
|
rtype="$(echo "$res" | jq -r '.type // "resource"')"
|
|
[ -n "$rname" ] || continue
|
|
needle="$(echo "$rname" | tr '[:upper:]' '[:lower:]' | tr -cd '[:alnum:]')"
|
|
if ! jq -e --arg n "$needle" '
|
|
(keys // [])[] | (ascii_downcase | gsub("[^a-z0-9]";"")) | select(contains($n))
|
|
' "$PAGE_MAP_FILE" >/dev/null 2>&1; then
|
|
add_gap "$rname" "aws-not-in-map" \
|
|
"AWS $rtype '$rname' is not represented in the IT page-ID map / architecture map" "medium" "aws-documented" \
|
|
"add '$rname' to the AWS Architecture Map (page 1540098) + an IT page; Mermaid edits via confluence_mermaid.py (on-demand path, provisioning)"
|
|
fi
|
|
done < <(jq -c '.resources[]' "$AWS_INVENTORY_FILE")
|
|
else
|
|
note_skip "aws-inventory:malformed(no-resources-array)"
|
|
fi
|
|
else
|
|
note_skip "aws-inventory:absent(check-skipped)" # missing inventory -> SKIP, never a gap
|
|
fi
|
|
|
|
# ==============================================================================
|
|
# CHECK 3 — REQUIRED standing/runbook pages present in the map
|
|
# ==============================================================================
|
|
IFS=',' read -r -a req_arr <<< "$REQUIRED_PAGES"
|
|
for page in "${req_arr[@]}"; do
|
|
page="$(echo "$page" | sed -E 's/^[[:space:]]+//; s/[[:space:]]+$//')"
|
|
[ -n "$page" ] || continue
|
|
if ! map_has_exact_key "$page"; then
|
|
add_gap "$page" "missing-runbook" \
|
|
"Required page '$page' is missing from the IT page-ID map" "high" "required-page" \
|
|
"create the '$page' page in the IT space and add it to project_confluence_migration"
|
|
fi
|
|
done
|
|
|
|
# ==============================================================================
|
|
# CHECK 4 — LIVE API: mapped pages still exist + are not stale (skipped offline/--no-api/--canary)
|
|
# ==============================================================================
|
|
if [ "$RUN_API" -eq 1 ]; then
|
|
while IFS=$'\t' read -r ptitle pid; do
|
|
[ -n "$pid" ] || continue
|
|
case "$pid" in ''|*[!0-9]*) note_skip "$ptitle:api(non-numeric-id)"; continue ;; esac
|
|
conf_check_page "$ptitle" "$pid"
|
|
done < <(jq -r 'to_entries[] | [.key, (.value|tostring)] | @tsv' "$PAGE_MAP_FILE")
|
|
else
|
|
note_skip "confluence-api:not-run(creds-absent-or-offline)"
|
|
fi
|
|
|
|
# ==============================================================================
|
|
# ASSEMBLE REPORT (JSON + text), mode 600 (identical shape to the other checkers)
|
|
# ==============================================================================
|
|
if [ "${#GAPS[@]}" -gt 0 ]; then
|
|
GAPS_JSON="$(printf '%s\n' "${GAPS[@]}" | jq -cs .)"
|
|
else
|
|
GAPS_JSON="[]"
|
|
fi
|
|
if [ "${#SKIPPED_CHECKS[@]}" -gt 0 ]; then
|
|
SKIPPED_JSON="$(printf '%s\n' "${SKIPPED_CHECKS[@]}" | jq -R . | jq -cs .)"
|
|
else
|
|
SKIPPED_JSON="[]"
|
|
fi
|
|
|
|
N_GAPS="$(echo "$GAPS_JSON" | jq 'length')"
|
|
N_HIGH="$(echo "$GAPS_JSON" | jq '[.[]|select(.severity=="high" or .severity=="critical")] | length')"
|
|
N_SUBJECTS="$(echo "$GAPS_JSON" | jq '[.[].repo] | unique | length')"
|
|
|
|
jq -n \
|
|
--arg checker "confluence-doc" --arg ts "$UTC_STAMP" --arg org "$GH_ORG" \
|
|
--argjson api "$RUN_API" --argjson reposn "${#REPO_NAMES[@]}" \
|
|
--argjson gaps "$GAPS_JSON" --argjson skipped "$SKIPPED_JSON" \
|
|
'{checker:$checker, generated:$ts, org:$org, mode:"recommend-only",
|
|
api_checks_ran:($api==1), repos_diffed:$reposn,
|
|
gap_count:($gaps|length),
|
|
subjects_with_gaps:([$gaps[].repo]|unique|length),
|
|
findings:$gaps, skipped_checks:$skipped}' > "$REPORT_JSON"
|
|
|
|
{
|
|
echo "confluence-doc — documentation gap report — $UTC_STAMP"
|
|
echo "org=$GH_ORG repos_diffed=${#REPO_NAMES[@]} api_checks=$([ "$RUN_API" -eq 1 ] && echo on || echo off) mode=recommend-only (D7)"
|
|
echo "doc gaps: $N_GAPS ($N_HIGH high) across $N_SUBJECTS subject(s)"
|
|
echo
|
|
if [ "$N_GAPS" -gt 0 ]; then
|
|
echo "RECOMMENDATIONS (recommend-only — NEVER auto-written, D7):"
|
|
echo "$GAPS_JSON" | jq -r '.[] | "• [\(.severity)] \(.repo): \(.title)\n recommend: \(.recommendation)"'
|
|
else
|
|
echo "No documentation gaps detected this run."
|
|
fi
|
|
if [ "$(echo "$SKIPPED_JSON" | jq 'length')" -gt 0 ]; then
|
|
echo; echo "skipped checks (missing data — NOT counted as a gap):"
|
|
echo "$SKIPPED_JSON" | jq -r '.[] | " - \(.)"'
|
|
fi
|
|
} > "$REPORT_TXT"
|
|
chmod 600 "$REPORT_JSON" "$REPORT_TXT" 2>/dev/null || true
|
|
|
|
log "report: $REPORT_JSON ($N_GAPS gap(s), $N_SUBJECTS subject(s))"
|
|
|
|
# ==============================================================================
|
|
# CANARY ASSERTION (anti-complacency floor, design §6.4)
|
|
# ==============================================================================
|
|
if [ "$CANARY" -eq 1 ]; then
|
|
EXPECT_FILE="$HERE/fixtures/confluence-doc/EXPECTED_GAP_COUNT"
|
|
[ -f "$EXPECT_FILE" ] || die "canary expected-count file missing: $EXPECT_FILE"
|
|
EXPECTED="$(tr -dc '0-9' < "$EXPECT_FILE")"
|
|
log "canary assertion: expected gaps=$EXPECTED, got=$N_GAPS"
|
|
if [ "$N_GAPS" -ne "$EXPECTED" ]; then
|
|
echo "[confluence-doc] CANARY FAIL: doc-gap count mismatch (expected $EXPECTED, got $N_GAPS)" >&2
|
|
echo " -> a gap check regressed (stopped firing) or the fixture changed. See $REPORT_TXT." >&2
|
|
exit 3
|
|
fi
|
|
log "canary PASS: all $EXPECTED planted doc gaps detected."
|
|
fi
|
|
|
|
# ==============================================================================
|
|
# RECOMMEND-ONLY ROUTING (D3/D7): gaps live in the mode-600 report. Post NOTHING by default.
|
|
# Scheduled mode NEVER auto-writes Confluence; alarming is reserved for confirmed criticals via
|
|
# the coordinator's shared routing (kept ALARM-only there). Here, recommend-only = report-only.
|
|
# ==============================================================================
|
|
if [ "$N_GAPS" -eq 0 ]; then
|
|
log "no doc gaps — recommend-only report written; posting NOTHING (D7)."
|
|
exit 0
|
|
fi
|
|
|
|
# Compose a redacted digest for the report/log (defense-in-depth); do NOT post by default.
|
|
DIGEST="$(echo "$GAPS_JSON" | jq -r '
|
|
group_by(.repo)[] | "*\(.[0].repo)*: " + ([.[] | "[\(.severity)] \(.title)"] | join("; "))' \
|
|
| sed 's/^/• /' | redact)"
|
|
echo "$DIGEST" >&2
|
|
log "DRY-RUN/RECOMMEND-ONLY: $N_GAPS gap(s) written to the mode-600 report; nothing posted, nothing written to Confluence (D7)."
|
|
exit 0
|
|
|
|
# ==============================================================================
|
|
# PROVISIONING (NOT DONE HERE — gated):
|
|
# - confluence-bot SERVICE ACCOUNT (D6): create a dedicated Atlassian service account scoped
|
|
# to EDIT the IT space ONLY (Confluence API tokens inherit the whole user's permissions, so a
|
|
# scoped service account bounds blast radius; costs one Confluence seat). Mint its API token,
|
|
# store it in ~/secrev.env (mode 600) as CONFLUENCE_API_TOKEN (+ CONFLUENCE_BASE_URL/EMAIL).
|
|
# Rotate the token on a 90-DAY cadence. Until this exists, the LIVE API checks SKIP (above),
|
|
# never alarm. This whole step is gated (Adam-provisioned), not done by this script.
|
|
# - LIVE Confluence READ checks (page-existence + staleness) only run once those creds exist.
|
|
# - ON-DEMAND WRITE path (D7) — the actual Confluence update, including Mermaid architecture-map
|
|
# edits via ~/.claude/scripts/confluence_mermaid.py — is a SEPARATE, LATER, SSH-invoked path.
|
|
# Before any --apply, that script must pass a LIVE DRY-RUN against page 1540098: verify it
|
|
# lists all 16 weweave Mermaid macros and that a no-op set produces a clean (empty) revert-diff.
|
|
# ADF-only + macro-count + revert-diff guards are load-bearing (a full-body markdown round-trip
|
|
# has SILENTLY DELETED every diagram on 1540098 before). This script NEVER calls --apply.
|
|
# - No systemd unit / timer is installed here. Wiring the scheduled run (weekly) under the
|
|
# coordinator is provisioning and is gated.
|
|
# - The coordinator (design §5, checker_coordinator.sh) registers + drives this checker; that
|
|
# registry edit is done centrally, NOT in this script.
|
|
# - Confluence + project_r720_agent_team memory updates are docs-as-you-go obligations for the
|
|
# build session, tracked outside this script.
|
|
# ==============================================================================
|