* feat(infra): migrate pipeline and Grafana to HCP Terraform (PLAT-75) Move apm-wo-analysis into seahaven-prod under workspace apm-wo-analysis-prod with in-repo hcptf/githubdeploy IAM, stub Lambdas, and GitHub Actions zip CD. * chore(iam): add Checkov skip comments for HCP IAM documents Pre-push HIGH findings are the DLM snapshot describe, tagged EC2 creates, exec boundary DescribeLogGroups star, and the drop-uploader user policy.
10 KiB
apm-wo-analysis — Operational Runbook
Operational procedures and incident response for the daily APM work-order
analysis pipeline. Account 011934824531 (seahaven-prod) / us-east-1.
HCP workspace apm-wo-analysis-prod. See README.md
for architecture and resource detail.
Admin access: the Grafana box is SSM Session Manager only (no SSH/key pair):
aws ssm start-session --target <instance-id>. AWS API via the office_mac /
deploy roles. Grafana UI is office-IP-restricted at the ALB.
1. Incident Runbook: Missing daily APM WO analysis
Severity: High — a day's WO analysis is missing: escalations (incl. 3rd-escalation alerts) aren't surfaced and the dashboard has no new snapshot. Recoverable by re-processing the export; not permanent data loss.
Detection
- Alarm
apm-wo-analysis-classifier-invocations(Sum Invocations over 86400 s< 1,treat_missing_data=breaching) — a silent day, including a missing export. - No daily summary in the WO Slack channel by the usual time (or no 3rd-escalation alert on a day one's expected).
- Grafana shows no
dt = todayin the Snapshot Date dropdown / panels empty for today. - Messages in
apm-wo-analysis-classifier-dlq(SQS) — strongest signal the classifier failed after it did run. - CloudWatch errors in
/aws/lambda/apm-wo-analysis-classifieror-slack-post.
Context
| Item | Value |
|---|---|
| Stacks | HCP workspace apm-wo-analysis-prod |
| Lambdas | apm-wo-analysis-classifier, -slack-post, -slack-interactions |
| S3 | apm-wo-analysis-exports-011934824531 — raw/, analytics/dt=…/, meta/dt=…/ |
| SQS DLQ | apm-wo-analysis-classifier-dlq |
| Glue / Athena | db apm_wo_analysis, table apm_wo_snapshots, workgroup apm-wo-analysis |
| External | Slack, Anthropic API (Haiku fallback) |
| Secrets (names) | apm-wo-analysis/slack-credentials, apm-wo-analysis/anthropic-api-key |
| SSM | /apm-wo-analysis/grafana-base-url |
Triage
- Export uploaded?
aws s3 ls s3://apm-wo-analysis-exports-011934824531/raw/— today's file present? Absent → upstream (§2.1), not the pipeline. - Classifier ran/failed?
/aws/lambda/apm-wo-analysis-classifierlogs; peek the DLQ:aws sqs receive-message --queue-url <dlq-url> --max-number-of-messages 1. - Outputs written?
aws s3 ls .../analytics/dt=<today>/(Parquet) and.../meta/dt=<today>/(summary.json,details.json). - slack-post ran/failed?
/aws/lambda/apm-wo-analysis-slack-postlogs —SlackApiError(invalid_auth,not_in_channel,invalid_blocks)? - Slack creds valid + bot still in channel? (
apm-wo-analysis/slack-credentials). - IAM/throttle: grep logs for
AccessDenied/ throttling. - Recent change? Any merge/deploy to
mainjust before the failure.
Resolution (by root cause)
- Export not uploaded →
aws s3 cp <export>.xlsx s3://apm-wo-analysis-exports-011934824531/raw/; then check the drop-folder agent (§2.1). - Classifier failed (DLQ) → read the DLQ message; fix; reprocess by re-uploading the export to
raw/(overwrite_partitionsmakes same-dtidempotent). - Outputs present, no Slack post → re-invoke:
If Slack auth was the cause → rotateaws lambda invoke --function-name apm-wo-analysis-slack-post \ --payload "$(printf '{"dt":"<YYYY-MM-DD>"}' | base64)" /tmp/out.jsonapm-wo-analysis/slack-credentials(and/or/invitethe bot), then re-invoke (secret read per-call; no redeploy). - Code regression → identify the PR, revert/hotfix, redeploy via CI (push to
main). - Data present, Grafana empty → datasource/dashboard issue (§2.3); check the instance via SSM.
- Anthropic/Haiku down → non-fatal (deterministic path still classifies ~95%); set classifier env
APM_HAIKU_FALLBACK=offto bypass.
Post-Incident
- Verify re-process → summary posts + Grafana shows today's
dt. - Check for other missed days (gaps in
analytics/dt=…/meta/) and reprocess each. - Update README/Confluence if knowledge changed; add a memory entry; add a test if code caused it.
2. Operational Procedures
2.1 How the export gets uploaded
The curated daily APM filter-view export (~350 WOs, .xlsx/.csv) reaches S3 by
direct upload or a local drop-folder — never SES/email. Any object under
raw/ with a .xlsx/.csv suffix triggers the classifier.
- Direct:
aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-011934824531/raw/ - Drop-folder (zero-touch): a launchd agent (
com.seahaven.apm-wo-uploader) watches~/apm-wo-drop/, uploads new files toraw/using the scopedapm-wo-dropprofile (IAM userapm-wo-drop-uploader—s3:PutObjectonraw/*only), and archives them locally. The runnable script lives at~/.local/bin/apm-wo-uploader.sh(it and the watched folder must be outside~/Documents— macOS TCC sandbox).
Verify the agent:
launchctl list | grep apm-wo-uploader # present + last exit 0
tail -f ~/.local/log/apm-wo-uploader.log # per-run logging (outside WatchPaths)
Common upstream issues: agent unloaded (launchctl load -w …plist); script
moved back into ~/Documents (TCC blocks it — LastExitStatus=32256); apm-wo-drop
access key expired/rotated (aws configure --profile apm-wo-drop).
2.2 Grafana OS / app patching cadence
The Grafana EC2 box (t4g.small, Amazon Linux 2023, ARM64) is the only
patch-bearing piece — everything else is serverless. It's **reproducible from
terraform/templates/grafana_userdata.sh.tftpl, so the preferred patch path is a clean
instance replacement rather than long-lived in-place drift.
- OS (recommended monthly + on critical CVEs): via SSM —
sudo dnf upgrade --security -y && sudo reboot(Session Manager, or an SSM Run Command / Patch Manager maintenance window — TBD: not yet automated). - Grafana OSS:
sudo dnf upgrade grafana -y && sudo systemctl restart grafana-server(installed from the pinnedrpm.grafana.comrepo). - Athena datasource plugin: pinned to
3.2.0interraform/variables.tf(athena_plugin_version). Bump there, then replace the instance (AMI/user-data areignore_changes; taint/replace to pick up user-data edits). - Preferred = clean replacement: terminate the instance; HCP apply relaunches
it from the latest AL2023 AMI and re-runs user-data only if
ami/user_dataignore_changes is lifted or the instance is replaced. The root volume isDeleteOnTermination=false, so detach/reuse or restoregrafana.db(§2.4) if local settings must persist.
2.3 Dashboard-JSON redeploy
Source of truth is grafana/dashboards/apm-work-orders.json in this repo
(uid apm-wo); the running instance is never the source of truth
(allowUiUpdates: false — UI edits are reverted on the next sync).
Flow: edit JSON in repo → GitHub Actions deploy.yaml syncs grafana/ to
s3://…/grafana-config/ → the instance
syncs S3 → /var/lib/grafana/dashboards/ (on boot + a 15-min systemd timer)
→ Grafana's file provider polls every 60 s and reloads.
Apply immediately (skip the timer) via SSM:
sudo /usr/local/bin/grafana-config-sync.sh # pulls grafana-config/ from S3
# Grafana file provider picks up the dashboard within ~60s
Gotchas:
- The sync uses
aws s3 sync --exact-timestamps— required so same-size edits (e.g. a one-char query change) actually propagate. - Datasource/provisioning changes (
grafana/provisioning/*.yaml) are loaded at startup — after syncing,sudo systemctl restart grafana-server(a dashboard-only change does not need a restart). - Athena query key is
rawSQL(capital); datasourceauthType: default; template varsrefresh: 1. (See README "Notes / Gotchas".)
2.4 Backup & restore (config + EBS / grafana.db)
Two distinct layers:
Config (dashboards, datasources, provisioning) — fully reproducible from
git (grafana/ → S3 grafana-config/). Restore: re-run deploy.yaml Grafana
sync (or grafana-config-sync.sh on the box). No snapshot needed.
Local state (/var/lib/grafana/grafana.db) — Grafana's SQLite (admin user,
any API keys, org prefs). Lives on the gp3 root volume (encrypted,
DeleteOnTermination=false). Backed up by a daily DLM snapshot (07:00 UTC,
7 retained) of the instance (tag apm-grafana-backup=true), policy in the
grafana stack.
Restore from snapshot:
# find the latest DLM snapshot
aws ec2 describe-snapshots --owner-ids self \
--filters "Name=tag:aws:dlm:lifecycle-policy-id,Values=*" \
--query 'reverse(sort_by(Snapshots,&StartTime))[0].SnapshotId' --output text
# create a volume from it and attach to a replacement instance, OR mount it and
# copy /var/lib/grafana/grafana.db onto the new instance, then:
sudo systemctl restart grafana-server
Because dashboards + datasource are provisioned from code, the only thing the
snapshot uniquely protects is grafana.db (admin/login state) — low stakes; a
fresh instance + provisioning recovers everything else.
Note: the Grafana admin auth model is unsettled (TBD) — the password was reset ad-hoc during build/testing. Decide the intended model (fixed admin password in Secrets Manager / SSO / anonymous view-only for the kiosk) and document it here.
2.5 Common failures (quick index)
| Symptom | Likely cause | Go to |
|---|---|---|
| No daily Slack post / no new dashboard day | export not uploaded, classifier failed (DLQ), or slack-post failed | §1 |
| Dashboard loads but all panels "No data" | datasource/auth, rawSQL, template refresh, or JSON in analytics/ prefix |
§2.3, README gotchas |
| Grafana unreachable | instance down / ALB unhealthy / office IP changed (officeCidrs) |
SSM triage; aws elbv2 describe-target-health |
| Slack modal click does nothing / error | apm-wo-analysis-slack-interactions, API Gateway, or signing-secret mismatch |
/aws/lambda/apm-wo-analysis-slack-interactions logs |
Exports never arrive in raw/ |
drop-folder agent unloaded / TCC / expired key | §2.1 |
| Deploy not applying | OIDC role, empty DEPLOY_ROLE_ARN, or HCP apply role |
GitHub Environment prod; HCP run |
Maintained in-repo (docs/RUNBOOK.md) and mirrored to Confluence. Update both
when operational knowledge changes.