apm-wo-analysis/cdk/assets/grafana_userdata.sh

95 lines
3 KiB
Bash
Raw Normal View History

Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
#!/bin/bash
# Grafana OSS bootstrap for the apm-wo-analysis dashboard host (Amazon Linux 2023,
# ARM64). Idempotent enough to re-run. Config + dashboards are pulled from S3
# (the repo is the source of truth); a systemd timer re-syncs dashboards so panel
# updates ship by re-uploading to S3 — no instance rebuild.
#
# Templated by CDK: __CONFIG_BUCKET__ / __CONFIG_PREFIX__ / __PLUGIN_VERSION__.
set -euxo pipefail
CONFIG_BUCKET="__CONFIG_BUCKET__"
CONFIG_PREFIX="__CONFIG_PREFIX__"
PLUGIN_VERSION="__PLUGIN_VERSION__"
# --- Grafana OSS repo + install ---
cat >/etc/yum.repos.d/grafana.repo <<'REPO'
[grafana]
name=grafana
baseurl=https://rpm.grafana.com
repo_gpgcheck=1
enabled=1
gpgcheck=1
gpgkey=https://rpm.grafana.com/gpg.key
sslverify=1
REPO
dnf install -y grafana
# --- Athena datasource plugin (pinned for reproducibility) ---
# --homepath is required or grafana-cli can't find its config defaults.
grafana-cli --homepath=/usr/share/grafana --pluginsDir=/var/lib/grafana/plugins \
plugins install grafana-athena-datasource "${PLUGIN_VERSION}"
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
# --- grafana.ini: behind the ALB at grafana.seahaven.com, kiosk-friendly ---
cat >/etc/grafana/grafana.ini <<'INI'
[server]
protocol = http
http_port = 3000
root_url = https://grafana.seahaven.com/
enforce_domain = false
[security]
# Behind an office-IP-restricted ALB; allow embedding for the kiosk wall display.
allow_embedding = true
cookie_secure = true
[users]
default_theme = dark
[analytics]
reporting_enabled = false
check_for_updates = false
INI
# --- sync provisioning + dashboards from S3 (repo is source of truth) ---
sync_config() {
aws s3 sync "s3://${CONFIG_BUCKET}/${CONFIG_PREFIX}/provisioning/" /etc/grafana/provisioning/ --delete --exact-timestamps
aws s3 sync "s3://${CONFIG_BUCKET}/${CONFIG_PREFIX}/dashboards/" /var/lib/grafana/dashboards/ --delete --exact-timestamps
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
chown -R grafana:grafana /etc/grafana/provisioning /var/lib/grafana/dashboards
}
mkdir -p /var/lib/grafana/dashboards
sync_config
systemctl daemon-reload
systemctl enable --now grafana-server
# --- systemd timer: re-sync dashboards every 15 min so repo edits land without a rebuild ---
cat >/usr/local/bin/grafana-config-sync.sh <<SYNC
#!/bin/bash
set -euo pipefail
aws s3 sync "s3://${CONFIG_BUCKET}/${CONFIG_PREFIX}/provisioning/" /etc/grafana/provisioning/ --delete --exact-timestamps
aws s3 sync "s3://${CONFIG_BUCKET}/${CONFIG_PREFIX}/dashboards/" /var/lib/grafana/dashboards/ --delete --exact-timestamps
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
chown -R grafana:grafana /etc/grafana/provisioning /var/lib/grafana/dashboards
SYNC
chmod +x /usr/local/bin/grafana-config-sync.sh
cat >/etc/systemd/system/grafana-config-sync.service <<'SVC'
[Unit]
Description=Sync apm-wo Grafana config/dashboards from S3
[Service]
Type=oneshot
ExecStart=/usr/local/bin/grafana-config-sync.sh
SVC
cat >/etc/systemd/system/grafana-config-sync.timer <<'TIMER'
[Unit]
Description=Periodic apm-wo Grafana config sync
[Timer]
OnBootSec=5min
OnUnitActiveSec=15min
[Install]
WantedBy=timers.target
TIMER
systemctl daemon-reload
systemctl enable --now grafana-config-sync.timer