# apm-wo-analysis — Operational Runbook Operational procedures and incident response for the daily APM work-order analysis pipeline. Account **328440206208** / **us-east-1**. Stacks `apm-wo-analysis-pipeline` and `apm-wo-analysis-grafana`. See [`README.md`](../README.md) for architecture and resource detail. **Admin access:** the Grafana box is **SSM Session Manager only** (no SSH/key pair): `aws ssm start-session --target `. AWS API via the `office_mac` / deploy roles. Grafana UI is office-IP-restricted at the ALB. --- ## 1. Incident Runbook: Missing daily APM WO analysis **Severity: High** — a day's WO analysis is missing: escalations (incl. 3rd-escalation alerts) aren't surfaced and the dashboard has no new snapshot. Recoverable by re-processing the export; not permanent data loss. ### Detection - No **daily summary** in the WO Slack channel by the usual time (or no 3rd-escalation alert on a day one's expected). - **Grafana** shows no `dt = today` in the Snapshot Date dropdown / panels empty for today. - **Messages in `apm-wo-analysis-classifier-dlq`** (SQS) — strongest signal the classifier failed. - CloudWatch errors in `/aws/lambda/apm-wo-analysis-classifier` or `-slack-post`. ### Context | Item | Value | |---|---| | Stacks | `apm-wo-analysis-pipeline`, `apm-wo-analysis-grafana` | | Lambdas | `apm-wo-analysis-classifier`, `-slack-post`, `-slack-interactions` | | S3 | `apm-wo-analysis-exports-328440206208` — `raw/`, `analytics/dt=…/`, `meta/dt=…/` | | SQS DLQ | `apm-wo-analysis-classifier-dlq` | | Glue / Athena | db `apm_wo_analysis`, table `apm_wo_snapshots`, workgroup `apm-wo-analysis` | | External | Slack, Anthropic API (Haiku fallback) | | Secrets (names) | `apm-wo-analysis/slack-credentials`, `apm-wo-analysis/anthropic-api-key` | | SSM | `/apm-wo-analysis/grafana-base-url` | ### Triage 1. **Export uploaded?** `aws s3 ls s3://apm-wo-analysis-exports-328440206208/raw/` — today's file present? Absent → upstream (§2.1), not the pipeline. 2. **Classifier ran/failed?** `/aws/lambda/apm-wo-analysis-classifier` logs; peek the DLQ: `aws sqs receive-message --queue-url --max-number-of-messages 1`. 3. **Outputs written?** `aws s3 ls .../analytics/dt=/` (Parquet) and `.../meta/dt=/` (`summary.json`, `details.json`). 4. **slack-post ran/failed?** `/aws/lambda/apm-wo-analysis-slack-post` logs — `SlackApiError` (`invalid_auth`, `not_in_channel`, `invalid_blocks`)? 5. **Slack creds** valid + bot still in channel? (`apm-wo-analysis/slack-credentials`). 6. **IAM/throttle:** grep logs for `AccessDenied` / throttling. 7. **Recent change?** Any merge/deploy to `main` just before the failure. ### Resolution (by root cause) 1. **Export not uploaded** → `aws s3 cp .xlsx s3://apm-wo-analysis-exports-328440206208/raw/`; then check the drop-folder agent (§2.1). 2. **Classifier failed (DLQ)** → read the DLQ message; fix; **reprocess by re-uploading the export to `raw/`** (`overwrite_partitions` makes same-`dt` idempotent). 3. **Outputs present, no Slack post** → re-invoke: ```bash aws lambda invoke --function-name apm-wo-analysis-slack-post \ --payload "$(printf '{"dt":""}' | base64)" /tmp/out.json ``` If Slack auth was the cause → rotate `apm-wo-analysis/slack-credentials` (and/or `/invite` the bot), then re-invoke (secret read per-call; no redeploy). 4. **Code regression** → identify the PR, revert/hotfix, redeploy via CI (push to `main`). 5. **Data present, Grafana empty** → datasource/dashboard issue (§2.3); check the instance via SSM. 6. **Anthropic/Haiku down** → non-fatal (deterministic path still classifies ~95%); set classifier env `APM_HAIKU_FALLBACK=off` to bypass. ### Post-Incident - Verify re-process → summary posts + Grafana shows today's `dt`. - Check for **other missed days** (gaps in `analytics/dt=…`/`meta/`) and reprocess each. - Update README/Confluence if knowledge changed; add a memory entry; add a test if code caused it. --- ## 2. Operational Procedures ### 2.1 How the export gets uploaded The curated daily APM filter-view export (~350 WOs, `.xlsx`/`.csv`) reaches S3 by **direct upload or a local drop-folder** — never SES/email. Any object under `raw/` with a `.xlsx`/`.csv` suffix triggers the classifier. - **Direct:** `aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-328440206208/raw/` - **Drop-folder (zero-touch):** a launchd agent (`com.seahaven.apm-wo-uploader`) watches `~/apm-wo-drop/`, uploads new files to `raw/` using the scoped `apm-wo-drop` profile (IAM user **`apm-wo-drop-uploader`** — `s3:PutObject` on `raw/*` only), and archives them locally. The runnable script lives at `~/.local/bin/apm-wo-uploader.sh` (it and the watched folder **must** be outside `~/Documents` — macOS TCC sandbox). **Verify the agent:** ```bash launchctl list | grep apm-wo-uploader # present + last exit 0 tail -f ~/apm-wo-drop/uploader.log # per-run logging (TBD: confirm log path) ``` **Common upstream issues:** agent unloaded (`launchctl load -w …plist`); script moved back into `~/Documents` (TCC blocks it — `LastExitStatus=32256`); `apm-wo-drop` access key expired/rotated (`aws configure --profile apm-wo-drop`). ### 2.2 Grafana OS / app patching cadence The Grafana EC2 box (`t4g.small`, **Amazon Linux 2023**, ARM64) is the only patch-bearing piece — everything else is serverless. It's **reproducible from `cdk/assets/grafana_userdata.sh`**, so the preferred patch path is a **clean instance replacement** rather than long-lived in-place drift. - **OS (recommended monthly + on critical CVEs):** via SSM — `sudo dnf upgrade --security -y && sudo reboot` (Session Manager, or an SSM Run Command / Patch Manager maintenance window — **TBD: not yet automated**). - **Grafana OSS:** `sudo dnf upgrade grafana -y && sudo systemctl restart grafana-server` (installed from the pinned `rpm.grafana.com` repo). - **Athena datasource plugin:** pinned to **`3.2.0`** in `cdk/cdk.json` (`athenaPluginVersion`). Bump there, then redeploy/replace the instance. - **Preferred = clean replacement:** terminate the instance; `cdk deploy apm-wo-analysis-grafana` relaunches it from the latest AL2023 AMI and re-runs user-data (fresh Grafana + plugin + config sync). The root volume is `DeleteOnTermination=false`, so detach/reuse or restore `grafana.db` (§2.4) if local settings must persist. Validates the committed user-data from a cold boot. > **Outstanding:** one clean instance replacement is owed to validate cold-boot > user-data and apply root-volume encryption (encryption can't be added in place). ### 2.3 Dashboard-JSON redeploy Source of truth is **`grafana/dashboards/apm-work-orders.json`** in this repo (uid **`apm-wo`**); the running instance is never the source of truth (`allowUiUpdates: false` — UI edits are reverted on the next sync). **Flow:** edit JSON in repo → `cdk deploy apm-wo-analysis-grafana` (the `BucketDeployment` uploads `grafana/` to `s3://…/grafana-config/`) → the instance syncs S3 → `/var/lib/grafana/dashboards/` (on boot + a **15-min systemd timer**) → Grafana's file provider polls every **60 s** and reloads. **Apply immediately** (skip the timer) via SSM: ```bash sudo /usr/local/bin/grafana-config-sync.sh # pulls grafana-config/ from S3 # Grafana file provider picks up the dashboard within ~60s ``` **Gotchas:** - The sync uses `aws s3 sync --exact-timestamps` — required so **same-size edits** (e.g. a one-char query change) actually propagate. - **Datasource/provisioning** changes (`grafana/provisioning/*.yaml`) are loaded at **startup** — after syncing, `sudo systemctl restart grafana-server` (a dashboard-only change does **not** need a restart). - Athena query key is **`rawSQL`** (capital); datasource `authType: default`; template vars `refresh: 1`. (See README "Notes / Gotchas".) ### 2.4 Backup & restore (config + EBS / grafana.db) Two distinct layers: **Config (dashboards, datasources, provisioning)** — fully **reproducible from git** (`grafana/` → S3 `grafana-config/`). *Restore:* `cdk deploy apm-wo-analysis-grafana` (or `grafana-config-sync.sh` on the box). No snapshot needed. **Local state (`/var/lib/grafana/grafana.db`)** — Grafana's SQLite (admin user, any API keys, org prefs). Lives on the **gp3 root volume** (encrypted, `DeleteOnTermination=false`). Backed up by a **daily DLM snapshot** (07:00 UTC, 7 retained) of the instance (tag `apm-grafana-backup=true`), policy in the grafana stack. *Restore from snapshot:* ```bash # find the latest DLM snapshot aws ec2 describe-snapshots --owner-ids self \ --filters "Name=tag:aws:dlm:lifecycle-policy-id,Values=*" \ --query 'reverse(sort_by(Snapshots,&StartTime))[0].SnapshotId' --output text # create a volume from it and attach to a replacement instance, OR mount it and # copy /var/lib/grafana/grafana.db onto the new instance, then: sudo systemctl restart grafana-server ``` Because dashboards + datasource are provisioned from code, the only thing the snapshot uniquely protects is `grafana.db` (admin/login state) — low stakes; a fresh instance + provisioning recovers everything else. > **Note:** the Grafana **admin auth model is unsettled (TBD)** — the password was > reset ad-hoc during build/testing. Decide the intended model (fixed admin > password in Secrets Manager / SSO / anonymous view-only for the kiosk) and > document it here. ### 2.5 Common failures (quick index) | Symptom | Likely cause | Go to | |---|---|---| | No daily Slack post / no new dashboard day | export not uploaded, classifier failed (DLQ), or slack-post failed | §1 | | Dashboard loads but all panels "No data" | datasource/auth, `rawSQL`, template `refresh`, or JSON in `analytics/` prefix | §2.3, README gotchas | | Grafana unreachable | instance down / ALB unhealthy / office IP changed (`officeCidrs`) | SSM triage; `aws elbv2 describe-target-health` | | Slack modal click does nothing / error | `apm-wo-analysis-slack-interactions`, API Gateway, or signing-secret mismatch | `/aws/lambda/apm-wo-analysis-slack-interactions` logs | | Exports never arrive in `raw/` | drop-folder agent unloaded / TCC / expired key | §2.1 | | Deploy not applying | OIDC role, CloudFormation rollback, Docker bundling | CloudFormation events; CI logs | --- *Maintained in-repo (`docs/RUNBOOK.md`) and mirrored to Confluence. Update both when operational knowledge changes.*