Commit graph

4 commits

Author SHA1 Message Date
Adam Moussa
ea54cb1e60 Fix Grafana bootstrap: grafana-cli --homepath + valid plugin version
Found on first boot (cloud-init errored, grafana-server never started):
- grafana-cli needs --homepath=/usr/share/grafana or it can't find config
  defaults; under set -e that aborted the whole bootstrap.
- pinned plugin version 2.18.2 doesn't exist (conflated with awswrangler's
  version) — grafana-athena-datasource latest is 3.2.0.

Verified by running the corrected bootstrap on the instance via SSM: plugin
installs, grafana-server active, /api/health 200, ALB target healthy.

NOTE: the running instance was repaired in-place (the user-data change updated
the launch template but did not replace the instance). The committed user-data
is now correct, so a fresh launch boots clean — a one-time clean instance
replacement should validate that before prod sign-off.
2026-05-28 18:43:08 -04:00
Adam Moussa
4870784fbd Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5)
The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the
imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the
office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the
syslog-server pattern). Instance in private subnets, reachable only from the ALB
SG, administered via SSM Session Manager (no SSH/key pair).

grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office
rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue
read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com
alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a
BucketDeployment that uploads grafana/ to the S3 config prefix.

grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write
grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning +
dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo
edits land without an instance rebuild.

Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid
apm-wo so the Slack 📊 button resolves): 7 panels — category distribution,
escalation summary, action/routine, escalations-by-site, trend time-series over
dt (the new capability), filterable WO table (5 template vars, escalation row
coloring, CSV export, no APM links), and the mismatch panel. Datasource uid
"athena" pinned in the provisioning yaml.

Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443
(caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB,
no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM
backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
Adam Moussa
c3a2c936d1 Add Slack post + interactions Lambdas with drill-down modals (Phase 4)
Two push surfaces (no App Home) + interactive drill-down, per CLAUDE.md.

Block Kit (blockkit.py, pure/offline): build_daily_summary (header, vs-yesterday
deltas, escalation breakdown with 3rd highlighted, action/routine, top sites,
mismatch callout, category drill buttons + 📊 Open dashboard link, footer),
build_escalation_alert (one @here, returns None on zero-3rd — suppression), and
build_wo_modal (views.open payload, capped under Slack's 100-block limit).

Lambdas: slack_post/handler.py (classifier-invoked: read today/yesterday
summary.json, post daily summary, conditionally post the batched alert from
details.json) and slack_post/interactions.py (API Gateway: verify Slack
signature, filter details.json, views.open the WO modal within the 3s trigger_id
window). slackio.py centralizes Secrets Manager creds, the SSM dashboard URL,
signature verification, and analytics/ reads — keeping blockkit pure.

Classifier: emit analytics/dt=*/details.json (per-WO index for the modals) and
async-invoke slack-post after the snapshot write (best-effort; a Slack failure
never fails classification).

CDK: slack-post + interactions Lambdas (Docker-bundled slack_sdk), HTTP API on
apm-wo.seahaven.com (wildcard ACM cert + Route53 alias; signature-verified, so
the route is unauthenticated by design), SSM /apm-wo-analysis/grafana-base-url,
and scoped IAM (read analytics/, read the Slack secret + dashboard param;
classifier granted lambda:InvokeFunction on slack-post). Slack creds live in one
Secrets Manager secret apm-wo-analysis/slack-credentials {botToken, signingSecret,
channelId}; cdk.json gains cert/zone/domain context.

WO drill-downs link to Grafana only — no APM deep-links (per decision).

Deliverables for test time: slack/manifest.yaml (app manifest, interactivity
request_url = apm-wo.seahaven.com).

Tests: tests/test_blockkit.py (30 offline cases — deltas, zero-3rd None, <100
blocks under large inputs, modal truncation/overflow, dashboard URL) and Phase 4
assertions in test_pipeline_synth.py (both Lambdas, the API route/domain/alias,
and no broad/write IAM on the Slack roles). 49/49 tests pass; cdk synth green.
2026-05-28 17:48:51 -04:00
Adam Moussa
58b91bda70 Scaffold apm-wo-analysis repository
Stand up the Phase 0 CDK scaffold for the daily APM work-order
analysis pipeline: two-stack CDK app (pipeline + grafana), classifier
and slack-post Lambda packages, dashboards-as-code, the local
drop-folder uploader, and a classifier smoke-test placeholder.

Wire CI/CD to the org reusable workflows: ci.yaml -> ci-python-sam
(ruff + cdk synth) and deploy.yaml -> cd-cdk (OIDC, cdk deploy --all).
Pin aws-cdk-lib==2.253.1; Lambdas target Python 3.12 / arm64.

Rewrite .gitignore to the org Python-CDK standard so the source-of-
truth files (CLAUDE.md, docs/, .claude/agents) are tracked while build
artifacts (.venv, cdk.out, caches) stay ignored.

Domain logic, stack resources, and dashboards are stubbed and filled
in across Phases 1-5 (docs/BUILD.md). cdk synth is green for both
stacks; ruff check/format pass.
2026-05-28 16:13:13 -04:00