apm-wo-analysis/docs/RUNBOOK.md
Adam Moussa e4fc32599f
Some checks failed
Deploy / deploy (push) Has been cancelled
fix(pipeline): page when classifier has no daily invocation
2026-08-13 17:39:11 -04:00

10 KiB

apm-wo-analysis — Operational Runbook

Operational procedures and incident response for the daily APM work-order analysis pipeline. Account 328440206208 / us-east-1. Stacks apm-wo-analysis-pipeline and apm-wo-analysis-grafana. See README.md for architecture and resource detail.

Admin access: the Grafana box is SSM Session Manager only (no SSH/key pair): aws ssm start-session --target <instance-id>. AWS API via the office_mac / deploy roles. Grafana UI is office-IP-restricted at the ALB.


1. Incident Runbook: Missing daily APM WO analysis

Severity: High — a day's WO analysis is missing: escalations (incl. 3rd-escalation alerts) aren't surfaced and the dashboard has no new snapshot. Recoverable by re-processing the export; not permanent data loss.

Detection

  • Alarm apm-wo-analysis-classifier-invocations (Sum Invocations over 86400 s < 1, treat_missing_data=breaching) — a silent day, including a missing export.
  • No daily summary in the WO Slack channel by the usual time (or no 3rd-escalation alert on a day one's expected).
  • Grafana shows no dt = today in the Snapshot Date dropdown / panels empty for today.
  • Messages in apm-wo-analysis-classifier-dlq (SQS) — strongest signal the classifier failed after it did run.
  • CloudWatch errors in /aws/lambda/apm-wo-analysis-classifier or -slack-post.

Context

Item Value
Stacks apm-wo-analysis-pipeline, apm-wo-analysis-grafana
Lambdas apm-wo-analysis-classifier, -slack-post, -slack-interactions
S3 apm-wo-analysis-exports-328440206208 — raw/, analytics/dt=…/, meta/dt=…/
SQS DLQ apm-wo-analysis-classifier-dlq
Glue / Athena db apm_wo_analysis, table apm_wo_snapshots, workgroup apm-wo-analysis
External Slack, Anthropic API (Haiku fallback)
Secrets (names) apm-wo-analysis/slack-credentials, apm-wo-analysis/anthropic-api-key
SSM /apm-wo-analysis/grafana-base-url

Triage

  1. Export uploaded? aws s3 ls s3://apm-wo-analysis-exports-328440206208/raw/ — today's file present? Absent → upstream (§2.1), not the pipeline.
  2. Classifier ran/failed? /aws/lambda/apm-wo-analysis-classifier logs; peek the DLQ: aws sqs receive-message --queue-url <dlq-url> --max-number-of-messages 1.
  3. Outputs written? aws s3 ls .../analytics/dt=<today>/ (Parquet) and .../meta/dt=<today>/ (summary.json, details.json).
  4. slack-post ran/failed? /aws/lambda/apm-wo-analysis-slack-post logs — SlackApiError (invalid_auth, not_in_channel, invalid_blocks)?
  5. Slack creds valid + bot still in channel? (apm-wo-analysis/slack-credentials).
  6. IAM/throttle: grep logs for AccessDenied / throttling.
  7. Recent change? Any merge/deploy to main just before the failure.

Resolution (by root cause)

  1. Export not uploaded → aws s3 cp <export>.xlsx s3://apm-wo-analysis-exports-328440206208/raw/; then check the drop-folder agent (§2.1).
  2. Classifier failed (DLQ) → read the DLQ message; fix; reprocess by re-uploading the export to raw/ (overwrite_partitions makes same-dt idempotent).
  3. Outputs present, no Slack post → re-invoke:
    aws lambda invoke --function-name apm-wo-analysis-slack-post \
      --payload "$(printf '{"dt":"<YYYY-MM-DD>"}' | base64)" /tmp/out.json
    
    If Slack auth was the cause → rotate apm-wo-analysis/slack-credentials (and/or /invite the bot), then re-invoke (secret read per-call; no redeploy).
  4. Code regression → identify the PR, revert/hotfix, redeploy via CI (push to main).
  5. Data present, Grafana empty → datasource/dashboard issue (§2.3); check the instance via SSM.
  6. Anthropic/Haiku down → non-fatal (deterministic path still classifies ~95%); set classifier env APM_HAIKU_FALLBACK=off to bypass.

Post-Incident

  • Verify re-process → summary posts + Grafana shows today's dt.
  • Check for other missed days (gaps in analytics/dt=…/meta/) and reprocess each.
  • Update README/Confluence if knowledge changed; add a memory entry; add a test if code caused it.

2. Operational Procedures

2.1 How the export gets uploaded

The curated daily APM filter-view export (~350 WOs, .xlsx/.csv) reaches S3 by direct upload or a local drop-folder — never SES/email. Any object under raw/ with a .xlsx/.csv suffix triggers the classifier.

  • Direct: aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-328440206208/raw/
  • Drop-folder (zero-touch): a launchd agent (com.seahaven.apm-wo-uploader) watches ~/apm-wo-drop/, uploads new files to raw/ using the scoped apm-wo-drop profile (IAM user apm-wo-drop-uploader — s3:PutObject on raw/* only), and archives them locally. The runnable script lives at ~/.local/bin/apm-wo-uploader.sh (it and the watched folder must be outside ~/Documents — macOS TCC sandbox).

Verify the agent:

launchctl list | grep apm-wo-uploader          # present + last exit 0
tail -f ~/.local/log/apm-wo-uploader.log       # per-run logging (outside WatchPaths)

Common upstream issues: agent unloaded (launchctl load -w …plist); script moved back into ~/Documents (TCC blocks it — LastExitStatus=32256); apm-wo-drop access key expired/rotated (aws configure --profile apm-wo-drop).

2.2 Grafana OS / app patching cadence

The Grafana EC2 box (t4g.small, Amazon Linux 2023, ARM64) is the only patch-bearing piece — everything else is serverless. It's reproducible from cdk/assets/grafana_userdata.sh, so the preferred patch path is a clean instance replacement rather than long-lived in-place drift.

  • OS (recommended monthly + on critical CVEs): via SSM — sudo dnf upgrade --security -y && sudo reboot (Session Manager, or an SSM Run Command / Patch Manager maintenance window — TBD: not yet automated).
  • Grafana OSS: sudo dnf upgrade grafana -y && sudo systemctl restart grafana-server (installed from the pinned rpm.grafana.com repo).
  • Athena datasource plugin: pinned to 3.2.0 in cdk/cdk.json (athenaPluginVersion). Bump there, then redeploy/replace the instance.
  • Preferred = clean replacement: terminate the instance; cdk deploy apm-wo-analysis-grafana relaunches it from the latest AL2023 AMI and re-runs user-data (fresh Grafana + plugin + config sync). The root volume is DeleteOnTermination=false, so detach/reuse or restore grafana.db (§2.4) if local settings must persist. Validates the committed user-data from a cold boot.

Outstanding: one clean instance replacement is owed to validate cold-boot user-data and apply root-volume encryption (encryption can't be added in place).

2.3 Dashboard-JSON redeploy

Source of truth is grafana/dashboards/apm-work-orders.json in this repo (uid apm-wo); the running instance is never the source of truth (allowUiUpdates: false — UI edits are reverted on the next sync).

Flow: edit JSON in repo → cdk deploy apm-wo-analysis-grafana (the BucketDeployment uploads grafana/ to s3://…/grafana-config/) → the instance syncs S3 → /var/lib/grafana/dashboards/ (on boot + a 15-min systemd timer) → Grafana's file provider polls every 60 s and reloads.

Apply immediately (skip the timer) via SSM:

sudo /usr/local/bin/grafana-config-sync.sh    # pulls grafana-config/ from S3
# Grafana file provider picks up the dashboard within ~60s

Gotchas:

  • The sync uses aws s3 sync --exact-timestamps — required so same-size edits (e.g. a one-char query change) actually propagate.
  • Datasource/provisioning changes (grafana/provisioning/*.yaml) are loaded at startup — after syncing, sudo systemctl restart grafana-server (a dashboard-only change does not need a restart).
  • Athena query key is rawSQL (capital); datasource authType: default; template vars refresh: 1. (See README "Notes / Gotchas".)

2.4 Backup & restore (config + EBS / grafana.db)

Two distinct layers:

Config (dashboards, datasources, provisioning) — fully reproducible from git (grafana/ → S3 grafana-config/). Restore: cdk deploy apm-wo-analysis-grafana (or grafana-config-sync.sh on the box). No snapshot needed.

Local state (/var/lib/grafana/grafana.db) — Grafana's SQLite (admin user, any API keys, org prefs). Lives on the gp3 root volume (encrypted, DeleteOnTermination=false). Backed up by a daily DLM snapshot (07:00 UTC, 7 retained) of the instance (tag apm-grafana-backup=true), policy in the grafana stack.

Restore from snapshot:

# find the latest DLM snapshot
aws ec2 describe-snapshots --owner-ids self \
  --filters "Name=tag:aws:dlm:lifecycle-policy-id,Values=*" \
  --query 'reverse(sort_by(Snapshots,&StartTime))[0].SnapshotId' --output text
# create a volume from it and attach to a replacement instance, OR mount it and
# copy /var/lib/grafana/grafana.db onto the new instance, then:
sudo systemctl restart grafana-server

Because dashboards + datasource are provisioned from code, the only thing the snapshot uniquely protects is grafana.db (admin/login state) — low stakes; a fresh instance + provisioning recovers everything else.

Note: the Grafana admin auth model is unsettled (TBD) — the password was reset ad-hoc during build/testing. Decide the intended model (fixed admin password in Secrets Manager / SSO / anonymous view-only for the kiosk) and document it here.

2.5 Common failures (quick index)

Symptom Likely cause Go to
No daily Slack post / no new dashboard day export not uploaded, classifier failed (DLQ), or slack-post failed §1
Dashboard loads but all panels "No data" datasource/auth, rawSQL, template refresh, or JSON in analytics/ prefix §2.3, README gotchas
Grafana unreachable instance down / ALB unhealthy / office IP changed (officeCidrs) SSM triage; aws elbv2 describe-target-health
Slack modal click does nothing / error apm-wo-analysis-slack-interactions, API Gateway, or signing-secret mismatch /aws/lambda/apm-wo-analysis-slack-interactions logs
Exports never arrive in raw/ drop-folder agent unloaded / TCC / expired key §2.1
Deploy not applying OIDC role, CloudFormation rollback, Docker bundling CloudFormation events; CI logs

Maintained in-repo (docs/RUNBOOK.md) and mirrored to Confluence. Update both when operational knowledge changes.