diff --git a/docs/RUNBOOK.md b/docs/RUNBOOK.md new file mode 100644 index 0000000..60084ab --- /dev/null +++ b/docs/RUNBOOK.md @@ -0,0 +1,187 @@ +# apm-wo-analysis — Operational Runbook + +Operational procedures and incident response for the daily APM work-order +analysis pipeline. Account **328440206208** / **us-east-1**. Stacks +`apm-wo-analysis-pipeline` and `apm-wo-analysis-grafana`. See [`README.md`](../README.md) +for architecture and resource detail. + +**Admin access:** the Grafana box is **SSM Session Manager only** (no SSH/key pair): +`aws ssm start-session --target `. AWS API via the `office_mac` / +deploy roles. Grafana UI is office-IP-restricted at the ALB. + +--- + +## 1. Incident Runbook: Missing daily APM WO analysis + +**Severity: High** — a day's WO analysis is missing: escalations (incl. 3rd-escalation +alerts) aren't surfaced and the dashboard has no new snapshot. Recoverable by +re-processing the export; not permanent data loss. + +### Detection +- No **daily summary** in the WO Slack channel by the usual time (or no 3rd-escalation alert on a day one's expected). +- **Grafana** shows no `dt = today` in the Snapshot Date dropdown / panels empty for today. +- **Messages in `apm-wo-analysis-classifier-dlq`** (SQS) — strongest signal the classifier failed. +- CloudWatch errors in `/aws/lambda/apm-wo-analysis-classifier` or `-slack-post`. + +### Context +| Item | Value | +|---|---| +| Stacks | `apm-wo-analysis-pipeline`, `apm-wo-analysis-grafana` | +| Lambdas | `apm-wo-analysis-classifier`, `-slack-post`, `-slack-interactions` | +| S3 | `apm-wo-analysis-exports-328440206208` — `raw/`, `analytics/dt=…/`, `meta/dt=…/` | +| SQS DLQ | `apm-wo-analysis-classifier-dlq` | +| Glue / Athena | db `apm_wo_analysis`, table `apm_wo_snapshots`, workgroup `apm-wo-analysis` | +| External | Slack, Anthropic API (Haiku fallback) | +| Secrets (names) | `apm-wo-analysis/slack-credentials`, `apm-wo-analysis/anthropic-api-key` | +| SSM | `/apm-wo-analysis/grafana-base-url` | + +### Triage +1. **Export uploaded?** `aws s3 ls s3://apm-wo-analysis-exports-328440206208/raw/` — today's file present? Absent → upstream (§2.1), not the pipeline. +2. **Classifier ran/failed?** `/aws/lambda/apm-wo-analysis-classifier` logs; peek the DLQ: `aws sqs receive-message --queue-url --max-number-of-messages 1`. +3. **Outputs written?** `aws s3 ls .../analytics/dt=/` (Parquet) and `.../meta/dt=/` (`summary.json`, `details.json`). +4. **slack-post ran/failed?** `/aws/lambda/apm-wo-analysis-slack-post` logs — `SlackApiError` (`invalid_auth`, `not_in_channel`, `invalid_blocks`)? +5. **Slack creds** valid + bot still in channel? (`apm-wo-analysis/slack-credentials`). +6. **IAM/throttle:** grep logs for `AccessDenied` / throttling. +7. **Recent change?** Any merge/deploy to `main` just before the failure. + +### Resolution (by root cause) +1. **Export not uploaded** → `aws s3 cp .xlsx s3://apm-wo-analysis-exports-328440206208/raw/`; then check the drop-folder agent (§2.1). +2. **Classifier failed (DLQ)** → read the DLQ message; fix; **reprocess by re-uploading the export to `raw/`** (`overwrite_partitions` makes same-`dt` idempotent). +3. **Outputs present, no Slack post** → re-invoke: + ```bash + aws lambda invoke --function-name apm-wo-analysis-slack-post \ + --payload "$(printf '{"dt":""}' | base64)" /tmp/out.json + ``` + If Slack auth was the cause → rotate `apm-wo-analysis/slack-credentials` (and/or `/invite` the bot), then re-invoke (secret read per-call; no redeploy). +4. **Code regression** → identify the PR, revert/hotfix, redeploy via CI (push to `main`). +5. **Data present, Grafana empty** → datasource/dashboard issue (§2.3); check the instance via SSM. +6. **Anthropic/Haiku down** → non-fatal (deterministic path still classifies ~95%); set classifier env `APM_HAIKU_FALLBACK=off` to bypass. + +### Post-Incident +- Verify re-process → summary posts + Grafana shows today's `dt`. +- Check for **other missed days** (gaps in `analytics/dt=…`/`meta/`) and reprocess each. +- Update README/Confluence if knowledge changed; add a memory entry; add a test if code caused it. + +--- + +## 2. Operational Procedures + +### 2.1 How the export gets uploaded + +The curated daily APM filter-view export (~350 WOs, `.xlsx`/`.csv`) reaches S3 by +**direct upload or a local drop-folder** — never SES/email. Any object under +`raw/` with a `.xlsx`/`.csv` suffix triggers the classifier. + +- **Direct:** `aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-328440206208/raw/` +- **Drop-folder (zero-touch):** a launchd agent (`com.seahaven.apm-wo-uploader`) + watches `~/apm-wo-drop/`, uploads new files to `raw/` using the scoped + `apm-wo-drop` profile (IAM user **`apm-wo-drop-uploader`** — `s3:PutObject` on + `raw/*` only), and archives them locally. The runnable script lives at + `~/.local/bin/apm-wo-uploader.sh` (it and the watched folder **must** be outside + `~/Documents` — macOS TCC sandbox). + +**Verify the agent:** +```bash +launchctl list | grep apm-wo-uploader # present + last exit 0 +tail -f ~/apm-wo-drop/uploader.log # per-run logging (TBD: confirm log path) +``` +**Common upstream issues:** agent unloaded (`launchctl load -w …plist`); script +moved back into `~/Documents` (TCC blocks it — `LastExitStatus=32256`); `apm-wo-drop` +access key expired/rotated (`aws configure --profile apm-wo-drop`). + +### 2.2 Grafana OS / app patching cadence + +The Grafana EC2 box (`t4g.small`, **Amazon Linux 2023**, ARM64) is the only +patch-bearing piece — everything else is serverless. It's **reproducible from +`cdk/assets/grafana_userdata.sh`**, so the preferred patch path is a **clean +instance replacement** rather than long-lived in-place drift. + +- **OS (recommended monthly + on critical CVEs):** via SSM — + `sudo dnf upgrade --security -y && sudo reboot` (Session Manager, or an SSM + Run Command / Patch Manager maintenance window — **TBD: not yet automated**). +- **Grafana OSS:** `sudo dnf upgrade grafana -y && sudo systemctl restart grafana-server` + (installed from the pinned `rpm.grafana.com` repo). +- **Athena datasource plugin:** pinned to **`3.2.0`** in `cdk/cdk.json` + (`athenaPluginVersion`). Bump there, then redeploy/replace the instance. +- **Preferred = clean replacement:** terminate the instance; `cdk deploy + apm-wo-analysis-grafana` relaunches it from the latest AL2023 AMI and re-runs + user-data (fresh Grafana + plugin + config sync). The root volume is + `DeleteOnTermination=false`, so detach/reuse or restore `grafana.db` (§2.4) if + local settings must persist. Validates the committed user-data from a cold boot. + +> **Outstanding:** one clean instance replacement is owed to validate cold-boot +> user-data and apply root-volume encryption (encryption can't be added in place). + +### 2.3 Dashboard-JSON redeploy + +Source of truth is **`grafana/dashboards/apm-work-orders.json`** in this repo +(uid **`apm-wo`**); the running instance is never the source of truth +(`allowUiUpdates: false` — UI edits are reverted on the next sync). + +**Flow:** edit JSON in repo → `cdk deploy apm-wo-analysis-grafana` (the +`BucketDeployment` uploads `grafana/` to `s3://…/grafana-config/`) → the instance +syncs S3 → `/var/lib/grafana/dashboards/` (on boot + a **15-min systemd timer**) +→ Grafana's file provider polls every **60 s** and reloads. + +**Apply immediately** (skip the timer) via SSM: +```bash +sudo /usr/local/bin/grafana-config-sync.sh # pulls grafana-config/ from S3 +# Grafana file provider picks up the dashboard within ~60s +``` +**Gotchas:** +- The sync uses `aws s3 sync --exact-timestamps` — required so **same-size edits** + (e.g. a one-char query change) actually propagate. +- **Datasource/provisioning** changes (`grafana/provisioning/*.yaml`) are loaded + at **startup** — after syncing, `sudo systemctl restart grafana-server` (a + dashboard-only change does **not** need a restart). +- Athena query key is **`rawSQL`** (capital); datasource `authType: default`; + template vars `refresh: 1`. (See README "Notes / Gotchas".) + +### 2.4 Backup & restore (config + EBS / grafana.db) + +Two distinct layers: + +**Config (dashboards, datasources, provisioning)** — fully **reproducible from +git** (`grafana/` → S3 `grafana-config/`). *Restore:* `cdk deploy +apm-wo-analysis-grafana` (or `grafana-config-sync.sh` on the box). No snapshot needed. + +**Local state (`/var/lib/grafana/grafana.db`)** — Grafana's SQLite (admin user, +any API keys, org prefs). Lives on the **gp3 root volume** (encrypted, +`DeleteOnTermination=false`). Backed up by a **daily DLM snapshot** (07:00 UTC, +7 retained) of the instance (tag `apm-grafana-backup=true`), policy in the +grafana stack. + +*Restore from snapshot:* +```bash +# find the latest DLM snapshot +aws ec2 describe-snapshots --owner-ids self \ + --filters "Name=tag:aws:dlm:lifecycle-policy-id,Values=*" \ + --query 'reverse(sort_by(Snapshots,&StartTime))[0].SnapshotId' --output text +# create a volume from it and attach to a replacement instance, OR mount it and +# copy /var/lib/grafana/grafana.db onto the new instance, then: +sudo systemctl restart grafana-server +``` +Because dashboards + datasource are provisioned from code, the only thing the +snapshot uniquely protects is `grafana.db` (admin/login state) — low stakes; a +fresh instance + provisioning recovers everything else. + +> **Note:** the Grafana **admin auth model is unsettled (TBD)** — the password was +> reset ad-hoc during build/testing. Decide the intended model (fixed admin +> password in Secrets Manager / SSO / anonymous view-only for the kiosk) and +> document it here. + +### 2.5 Common failures (quick index) + +| Symptom | Likely cause | Go to | +|---|---|---| +| No daily Slack post / no new dashboard day | export not uploaded, classifier failed (DLQ), or slack-post failed | §1 | +| Dashboard loads but all panels "No data" | datasource/auth, `rawSQL`, template `refresh`, or JSON in `analytics/` prefix | §2.3, README gotchas | +| Grafana unreachable | instance down / ALB unhealthy / office IP changed (`officeCidrs`) | SSM triage; `aws elbv2 describe-target-health` | +| Slack modal click does nothing / error | `apm-wo-analysis-slack-interactions`, API Gateway, or signing-secret mismatch | `/aws/lambda/apm-wo-analysis-slack-interactions` logs | +| Exports never arrive in `raw/` | drop-folder agent unloaded / TCC / expired key | §2.1 | +| Deploy not applying | OIDC role, CloudFormation rollback, Docker bundling | CloudFormation events; CI logs | + +--- + +*Maintained in-repo (`docs/RUNBOOK.md`) and mirrored to Confluence. Update both +when operational knowledge changes.*