apm-wo-analysis/grafana/dashboards/apm-work-orders.json

806 lines
27 KiB
JSON
Raw Normal View History

{
"uid": "apm-wo",
"title": "APM Work Orders",
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"description": "Daily APM work-order analysis — breakdown, escalations, trend, filterable WO table, mismatches. Built in Phase 5 by the grafana-author agent. Athena datasource uid='athena'; data grain is one row per WO per daily snapshot partitioned by dt.",
"tags": ["apm", "work-orders"],
"timezone": "browser",
"schemaVersion": 39,
"version": 1,
"refresh": "1h",
"time": {
"from": "now-30d",
"to": "now"
},
"templating": {
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"list": [
{
"name": "dt",
"label": "Snapshot Date",
"description": "Single daily snapshot partition. Defaults to the latest available dt. Most panels filter WHERE dt = '$dt'.",
"type": "query",
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"query": {
"rawSQL": "SELECT DISTINCT dt FROM apm_wo_analysis.apm_wo_snapshots ORDER BY dt DESC",
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"format": "table",
"connectionArgs": {
"catalog": "AwsDataCatalog",
"database": "apm_wo_analysis",
"region": "us-east-1"
}
},
"refresh": 1,
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"sort": 0,
"multi": false,
"includeAll": false,
"current": {},
"options": [],
"hide": 0
},
{
"name": "site",
"label": "Site",
"description": "Multi-select. WO table applies: AND ('${site:raw}' = 'All' OR site IN (${site:singlequote}))",
"type": "query",
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"query": {
"rawSQL": "SELECT DISTINCT site FROM apm_wo_analysis.apm_wo_snapshots WHERE dt='$dt' AND site<>'' ORDER BY 1",
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"format": "table",
"connectionArgs": {
"catalog": "AwsDataCatalog",
"database": "apm_wo_analysis",
"region": "us-east-1"
}
},
"refresh": 1,
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"sort": 1,
"multi": true,
"includeAll": true,
"current": {},
"options": [],
"hide": 0
},
{
"name": "department",
"label": "Department",
"description": "Multi-select. WO table applies: AND ('${department:raw}' = 'All' OR department IN (${department:singlequote}))",
"type": "query",
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"query": {
"rawSQL": "SELECT DISTINCT department FROM apm_wo_analysis.apm_wo_snapshots WHERE dt='$dt' AND department<>'' ORDER BY 1",
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"format": "table",
"connectionArgs": {
"catalog": "AwsDataCatalog",
"database": "apm_wo_analysis",
"region": "us-east-1"
}
},
"refresh": 1,
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"sort": 1,
"multi": true,
"includeAll": true,
"current": {},
"options": [],
"hide": 0
},
{
"name": "category",
"label": "Category",
"description": "Multi-select. WO table applies: AND ('${category:raw}' = 'All' OR category IN (${category:singlequote}))",
"type": "query",
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"query": {
"rawSQL": "SELECT DISTINCT category FROM apm_wo_analysis.apm_wo_snapshots WHERE dt='$dt' AND category<>'' ORDER BY 1",
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"format": "table",
"connectionArgs": {
"catalog": "AwsDataCatalog",
"database": "apm_wo_analysis",
"region": "us-east-1"
}
},
"refresh": 1,
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"sort": 1,
"multi": true,
"includeAll": true,
"current": {},
"options": [],
"hide": 0
},
{
"name": "wo_status",
"label": "WO Status",
"description": "Multi-select. WO table applies: AND ('${wo_status:raw}' = 'All' OR wo_status IN (${wo_status:singlequote}))",
"type": "query",
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"query": {
"rawSQL": "SELECT DISTINCT wo_status FROM apm_wo_analysis.apm_wo_snapshots WHERE dt='$dt' AND wo_status<>'' ORDER BY 1",
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"format": "table",
"connectionArgs": {
"catalog": "AwsDataCatalog",
"database": "apm_wo_analysis",
"region": "us-east-1"
}
},
"refresh": 1,
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"sort": 1,
"multi": true,
"includeAll": true,
"current": {},
"options": [],
"hide": 0
},
{
"name": "hold_reason",
"label": "Hold Reason",
"description": "Multi-select. WO table applies: AND ('${hold_reason:raw}' = 'All' OR hold_reason IN (${hold_reason:singlequote}))",
"type": "query",
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"query": {
"rawSQL": "SELECT DISTINCT hold_reason FROM apm_wo_analysis.apm_wo_snapshots WHERE dt='$dt' AND hold_reason<>'' ORDER BY 1",
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"format": "table",
"connectionArgs": {
"catalog": "AwsDataCatalog",
"database": "apm_wo_analysis",
"region": "us-east-1"
}
},
"refresh": 1,
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"sort": 1,
"multi": true,
"includeAll": true,
"current": {},
"options": [],
"hide": 0
}
]
},
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"panels": [
{
"id": 1,
"type": "barchart",
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"title": "Category Distribution",
"description": "Count of work orders by classification category for the selected snapshot date. Sorted descending by volume. Filter using the template variables above.",
"gridPos": { "x": 0, "y": 0, "w": 14, "h": 9 },
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"targets": [
{
"refId": "A",
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"rawSQL": "SELECT category, COUNT(*) AS wos FROM apm_wo_analysis.apm_wo_snapshots WHERE dt='$dt' GROUP BY 1 ORDER BY 2 DESC",
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"format": "table",
"connectionArgs": {
"catalog": "AwsDataCatalog",
"database": "apm_wo_analysis",
"region": "us-east-1"
}
}
],
"options": {
"orientation": "horizontal",
"barRadius": 0,
"groupWidth": 0.7,
"showValue": "always",
"stacking": "none",
"tooltip": { "mode": "single", "sort": "none" },
"legend": { "showLegend": false, "displayMode": "list", "placement": "bottom" }
},
"fieldConfig": {
"defaults": {
"color": { "mode": "palette-classic" },
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"fillOpacity": 80,
"gradientMode": "none",
"hideFrom": { "legend": false, "tooltip": false, "viz": false },
"lineWidth": 1,
"scaleDistribution": { "type": "linear" },
"thresholdsStyle": { "mode": "off" }
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{ "color": "green", "value": null },
{ "color": "red", "value": 80 }
]
}
},
"overrides": []
}
},
{
"id": 2,
"type": "piechart",
"title": "Escalation Summary",
"description": "Distribution across escalation categories (1st, 2nd, 3rd Escalation, SIM Ticket, Other Escalation) for the selected snapshot date. 3rd Escalation is highlighted red.",
"gridPos": { "x": 14, "y": 0, "w": 10, "h": 9 },
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"targets": [
{
"refId": "A",
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"rawSQL": "SELECT category, COUNT(*) AS escalations FROM apm_wo_analysis.apm_wo_snapshots WHERE dt='$dt' AND is_escalation=true GROUP BY 1 ORDER BY 2 DESC",
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"format": "table",
"connectionArgs": {
"catalog": "AwsDataCatalog",
"database": "apm_wo_analysis",
"region": "us-east-1"
}
}
],
"options": {
"pieType": "pie",
"displayLabels": ["name", "value"],
"tooltip": { "mode": "single", "sort": "none" },
"legend": { "showLegend": true, "displayMode": "table", "placement": "right", "values": ["value", "percent"] }
},
"fieldConfig": {
"defaults": {
"color": { "mode": "palette-classic" },
"custom": {
"hideFrom": { "legend": false, "tooltip": false, "viz": false }
},
"mappings": []
},
"overrides": [
{
"matcher": { "id": "byName", "options": "3rd Escalation" },
"properties": [
{ "id": "color", "value": { "fixedColor": "red", "mode": "fixed" } }
]
},
{
"matcher": { "id": "byName", "options": "2nd Escalation" },
"properties": [
{ "id": "color", "value": { "fixedColor": "orange", "mode": "fixed" } }
]
},
{
"matcher": { "id": "byName", "options": "1st Escalation" },
"properties": [
{ "id": "color", "value": { "fixedColor": "yellow", "mode": "fixed" } }
]
},
{
"matcher": { "id": "byName", "options": "SIM Ticket" },
"properties": [
{ "id": "color", "value": { "fixedColor": "purple", "mode": "fixed" } }
]
}
]
}
},
{
"id": 3,
"type": "piechart",
"title": "Action-Needed vs Routine",
"description": "Action-needed WOs require attention (escalations, scheduling, vendor/report waits, status inquiries, vendor no-shows). Routine WOs are on track. Counts are for the selected snapshot date.",
"gridPos": { "x": 0, "y": 9, "w": 8, "h": 8 },
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"targets": [
{
"refId": "A",
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"rawSQL": "SELECT CASE WHEN is_action=true THEN 'Action Needed' ELSE 'Routine' END AS action_type, COUNT(*) AS wos FROM apm_wo_analysis.apm_wo_snapshots WHERE dt='$dt' GROUP BY 1",
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"format": "table",
"connectionArgs": {
"catalog": "AwsDataCatalog",
"database": "apm_wo_analysis",
"region": "us-east-1"
}
}
],
"options": {
"pieType": "donut",
"displayLabels": ["name", "percent"],
"tooltip": { "mode": "single", "sort": "none" },
"legend": { "showLegend": true, "displayMode": "list", "placement": "bottom", "values": ["value", "percent"] }
},
"fieldConfig": {
"defaults": {
"color": { "mode": "palette-classic" },
"custom": {
"hideFrom": { "legend": false, "tooltip": false, "viz": false }
},
"mappings": []
},
"overrides": [
{
"matcher": { "id": "byName", "options": "Action Needed" },
"properties": [
{ "id": "color", "value": { "fixedColor": "semi-dark-orange", "mode": "fixed" } }
]
},
{
"matcher": { "id": "byName", "options": "Routine" },
"properties": [
{ "id": "color", "value": { "fixedColor": "green", "mode": "fixed" } }
]
}
]
}
},
{
"id": 4,
"type": "barchart",
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"title": "Escalations by Site",
"description": "Total escalations per site for the selected snapshot date. Includes all escalation categories.",
"gridPos": { "x": 8, "y": 9, "w": 16, "h": 8 },
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"targets": [
{
"refId": "A",
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"rawSQL": "SELECT site, COUNT(*) AS escalations FROM apm_wo_analysis.apm_wo_snapshots WHERE dt='$dt' AND is_escalation=true AND site<>'' GROUP BY 1 ORDER BY 2 DESC",
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"format": "table",
"connectionArgs": {
"catalog": "AwsDataCatalog",
"database": "apm_wo_analysis",
"region": "us-east-1"
}
}
],
"options": {
"orientation": "vertical",
"barRadius": 0,
"groupWidth": 0.7,
"showValue": "always",
"stacking": "none",
"tooltip": { "mode": "single", "sort": "none" },
"legend": { "showLegend": false, "displayMode": "list", "placement": "bottom" },
"xTickLabelRotation": -45,
"xTickLabelMaxLength": 12
},
"fieldConfig": {
"defaults": {
"color": { "mode": "fixed", "fixedColor": "semi-dark-red" },
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "Escalations",
"axisPlacement": "auto",
"fillOpacity": 80,
"gradientMode": "none",
"hideFrom": { "legend": false, "tooltip": false, "viz": false },
"lineWidth": 1,
"scaleDistribution": { "type": "linear" },
"thresholdsStyle": { "mode": "off" }
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{ "color": "green", "value": null }
]
}
},
"overrides": []
}
},
{
"id": 5,
"type": "timeseries",
"title": "Trend Over Time — Escalations & Action-Needed per Day",
"description": "Runs across all partitions (no $dt filter) to show daily escalation and action-needed volume over time. Use the dashboard time range picker to zoom in. This is the capability the legacy Sheet never had.",
"gridPos": { "x": 0, "y": 17, "w": 24, "h": 9 },
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"targets": [
{
"refId": "A",
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"rawSQL": "SELECT date_parse(dt, '%Y-%m-%d') AS time, SUM(CAST(is_escalation AS INTEGER)) AS escalations, SUM(CAST(is_action AS INTEGER)) AS action_needed FROM apm_wo_analysis.apm_wo_snapshots GROUP BY 1 ORDER BY 1",
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"format": "timeSeries",
"connectionArgs": {
"catalog": "AwsDataCatalog",
"database": "apm_wo_analysis",
"region": "us-east-1"
}
}
],
"options": {
"tooltip": { "mode": "multi", "sort": "desc" },
"legend": { "showLegend": true, "displayMode": "list", "placement": "bottom", "calcs": ["mean", "max", "last"] }
},
"fieldConfig": {
"defaults": {
"color": { "mode": "palette-classic" },
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "Work Orders",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": { "legend": false, "tooltip": false, "viz": false },
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 2,
"pointSize": 5,
"scaleDistribution": { "type": "linear" },
"showPoints": "auto",
"spanNulls": false,
"stacking": { "group": "A", "mode": "none" },
"thresholdsStyle": { "mode": "off" }
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{ "color": "green", "value": null }
]
},
"unit": "short"
},
"overrides": [
{
"matcher": { "id": "byName", "options": "escalations" },
"properties": [
{ "id": "color", "value": { "fixedColor": "semi-dark-red", "mode": "fixed" } },
{ "id": "displayName", "value": "Escalations" }
]
},
{
"matcher": { "id": "byName", "options": "action_needed" },
"properties": [
{ "id": "color", "value": { "fixedColor": "semi-dark-orange", "mode": "fixed" } },
{ "id": "displayName", "value": "Action Needed" }
]
}
]
}
},
{
"id": 6,
"type": "table",
"title": "Work Order Detail Table",
"description": "Filterable, exportable table of all work orders for the selected snapshot date. Multi-value template variables are applied via: AND ('${variable:raw}' = 'All' OR column IN (${variable:singlequote})). Escalation rows are color-coded: 3rd Escalation = red, 2nd Escalation = orange, 1st Escalation = yellow. WO numbers are plain text (no APM deep-link per project decision). Use the Download CSV button (table header menu) to export. last_comment is wrapped for readability.",
"gridPos": { "x": 0, "y": 26, "w": 24, "h": 14 },
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"targets": [
{
"refId": "A",
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"rawSQL": "SELECT wo_number, site, department, category, wo_status, hold_reason, wo_description, last_comment FROM apm_wo_analysis.apm_wo_snapshots WHERE dt='$dt' AND ('${site:raw}' = 'All' OR site IN (${site:singlequote})) AND ('${department:raw}' = 'All' OR department IN (${department:singlequote})) AND ('${category:raw}' = 'All' OR category IN (${category:singlequote})) AND ('${wo_status:raw}' = 'All' OR wo_status IN (${wo_status:singlequote})) AND ('${hold_reason:raw}' = 'All' OR hold_reason IN (${hold_reason:singlequote})) ORDER BY category, site",
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"format": "table",
"connectionArgs": {
"catalog": "AwsDataCatalog",
"database": "apm_wo_analysis",
"region": "us-east-1"
}
}
],
"options": {
"frameIndex": 0,
"showHeader": true,
"sortBy": [],
"footer": {
"show": false,
"reducer": ["sum"],
"fields": "",
"enablePagination": false
}
},
"fieldConfig": {
"defaults": {
"color": { "mode": "thresholds" },
"custom": {
"align": "left",
"cellOptions": { "type": "auto" },
"inspect": false,
"filterable": true,
"minWidth": 80,
"width": 0
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{ "color": "text", "value": null }
]
}
},
"overrides": [
{
"matcher": { "id": "byName", "options": "wo_number" },
"properties": [
{ "id": "displayName", "value": "WO #" },
{ "id": "custom.width", "value": 100 }
]
},
{
"matcher": { "id": "byName", "options": "site" },
"properties": [
{ "id": "displayName", "value": "Site" },
{ "id": "custom.width", "value": 80 }
]
},
{
"matcher": { "id": "byName", "options": "department" },
"properties": [
{ "id": "displayName", "value": "Dept" },
{ "id": "custom.width", "value": 80 }
]
},
{
"matcher": { "id": "byName", "options": "category" },
"properties": [
{ "id": "displayName", "value": "Category" },
{ "id": "custom.width", "value": 160 },
{
"id": "mappings",
"value": [
{
"type": "value",
"options": {
"3rd Escalation": {
"color": "dark-red",
"index": 0
},
"2nd Escalation": {
"color": "dark-orange",
"index": 1
},
"1st Escalation": {
"color": "dark-yellow",
"index": 2
},
"SIM Ticket": {
"color": "dark-purple",
"index": 3
},
"Other Escalation": {
"color": "orange",
"index": 4
}
}
}
]
},
{
"id": "custom.cellOptions",
"value": { "type": "color-text" }
}
]
},
{
"matcher": { "id": "byName", "options": "wo_status" },
"properties": [
{ "id": "displayName", "value": "Status" },
{ "id": "custom.width", "value": 80 }
]
},
{
"matcher": { "id": "byName", "options": "hold_reason" },
"properties": [
{ "id": "displayName", "value": "Hold Reason" },
{ "id": "custom.width", "value": 120 }
]
},
{
"matcher": { "id": "byName", "options": "wo_description" },
"properties": [
{ "id": "displayName", "value": "Description" },
{ "id": "custom.width", "value": 220 }
]
},
{
"matcher": { "id": "byName", "options": "last_comment" },
"properties": [
{ "id": "displayName", "value": "Last Comment" },
{
"id": "custom.cellOptions",
"value": { "type": "auto", "wrapText": true }
},
{ "id": "custom.width", "value": 420 },
{ "id": "custom.minWidth", "value": 200 }
]
}
]
}
},
{
"id": 7,
"type": "table",
"title": "Mismatch Panel — Comment vs Structured-State Contradictions",
"description": "Work orders where the classified comment intent contradicts the structured WO state (e.g. comment says 'completed' but WO is on a REPORT/VENDOR/SCHEDULING hold, or comment says 'scheduled' while status is IP). Non-empty mismatch column only. Use these rows for manual review before the next export.",
"gridPos": { "x": 0, "y": 40, "w": 24, "h": 8 },
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"targets": [
{
"refId": "A",
"datasource": {
"type": "grafana-athena-datasource",
"uid": "athena"
},
"rawSQL": "SELECT wo_number, site, category, wo_status, hold_reason, mismatch FROM apm_wo_analysis.apm_wo_snapshots WHERE dt='$dt' AND mismatch<>'' ORDER BY category, site",
Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5) The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the syslog-server pattern). Instance in private subnets, reachable only from the ALB SG, administered via SSM Session Manager (no SSH/key pair). grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a BucketDeployment that uploads grafana/ to the S3 config prefix. grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning + dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo edits land without an instance rebuild. Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid apm-wo so the Slack 📊 button resolves): 7 panels — category distribution, escalation summary, action/routine, escalations-by-site, trend time-series over dt (the new capability), filterable WO table (5 template vars, escalation row coloring, CSV export, no APM links), and the mismatch panel. Datasource uid "athena" pinned in the provisioning yaml. Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443 (caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB, no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
"format": "table",
"connectionArgs": {
"catalog": "AwsDataCatalog",
"database": "apm_wo_analysis",
"region": "us-east-1"
}
}
],
"options": {
"frameIndex": 0,
"showHeader": true,
"sortBy": [],
"footer": {
"show": false,
"reducer": ["sum"],
"fields": "",
"enablePagination": false
}
},
"fieldConfig": {
"defaults": {
"color": { "mode": "thresholds" },
"custom": {
"align": "left",
"cellOptions": { "type": "auto" },
"inspect": false,
"filterable": true,
"minWidth": 80,
"width": 0
},
"mappings": [],
"thresholds": {
"mode": "absolute",
"steps": [
{ "color": "text", "value": null }
]
}
},
"overrides": [
{
"matcher": { "id": "byName", "options": "wo_number" },
"properties": [
{ "id": "displayName", "value": "WO #" },
{ "id": "custom.width", "value": 100 }
]
},
{
"matcher": { "id": "byName", "options": "site" },
"properties": [
{ "id": "displayName", "value": "Site" },
{ "id": "custom.width", "value": 80 }
]
},
{
"matcher": { "id": "byName", "options": "category" },
"properties": [
{ "id": "displayName", "value": "Category" },
{ "id": "custom.width", "value": 160 },
{
"id": "mappings",
"value": [
{
"type": "value",
"options": {
"3rd Escalation": {
"color": "dark-red",
"index": 0
},
"2nd Escalation": {
"color": "dark-orange",
"index": 1
},
"1st Escalation": {
"color": "dark-yellow",
"index": 2
},
"SIM Ticket": {
"color": "dark-purple",
"index": 3
},
"Other Escalation": {
"color": "orange",
"index": 4
}
}
}
]
},
{
"id": "custom.cellOptions",
"value": { "type": "color-text" }
}
]
},
{
"matcher": { "id": "byName", "options": "wo_status" },
"properties": [
{ "id": "displayName", "value": "WO Status" },
{ "id": "custom.width", "value": 90 }
]
},
{
"matcher": { "id": "byName", "options": "hold_reason" },
"properties": [
{ "id": "displayName", "value": "Hold Reason" },
{ "id": "custom.width", "value": 120 }
]
},
{
"matcher": { "id": "byName", "options": "mismatch" },
"properties": [
{ "id": "displayName", "value": "Mismatch Description" },
{
"id": "custom.cellOptions",
"value": { "type": "auto", "wrapText": true }
},
{ "id": "custom.width", "value": 500 },
{ "id": "custom.minWidth", "value": 200 },
{ "id": "color", "value": { "mode": "fixed", "fixedColor": "semi-dark-orange" } }
]
}
]
}
}
]
}