apm-wo-analysis/docs/RUNBOOK.md

188 lines
10 KiB
Markdown
Raw Normal View History

# apm-wo-analysis — Operational Runbook
Operational procedures and incident response for the daily APM work-order
analysis pipeline. Account **328440206208** / **us-east-1**. Stacks
`apm-wo-analysis-pipeline` and `apm-wo-analysis-grafana`. See [`README.md`](../README.md)
for architecture and resource detail.
**Admin access:** the Grafana box is **SSM Session Manager only** (no SSH/key pair):
`aws ssm start-session --target <instance-id>`. AWS API via the `office_mac` /
deploy roles. Grafana UI is office-IP-restricted at the ALB.
---
## 1. Incident Runbook: Missing daily APM WO analysis
**Severity: High** — a day's WO analysis is missing: escalations (incl. 3rd-escalation
alerts) aren't surfaced and the dashboard has no new snapshot. Recoverable by
re-processing the export; not permanent data loss.
### Detection
- No **daily summary** in the WO Slack channel by the usual time (or no 3rd-escalation alert on a day one's expected).
- **Grafana** shows no `dt = today` in the Snapshot Date dropdown / panels empty for today.
- **Messages in `apm-wo-analysis-classifier-dlq`** (SQS) — strongest signal the classifier failed.
- CloudWatch errors in `/aws/lambda/apm-wo-analysis-classifier` or `-slack-post`.
### Context
| Item | Value |
|---|---|
| Stacks | `apm-wo-analysis-pipeline`, `apm-wo-analysis-grafana` |
| Lambdas | `apm-wo-analysis-classifier`, `-slack-post`, `-slack-interactions` |
| S3 | `apm-wo-analysis-exports-328440206208` — `raw/`, `analytics/dt=…/`, `meta/dt=…/` |
| SQS DLQ | `apm-wo-analysis-classifier-dlq` |
| Glue / Athena | db `apm_wo_analysis`, table `apm_wo_snapshots`, workgroup `apm-wo-analysis` |
| External | Slack, Anthropic API (Haiku fallback) |
| Secrets (names) | `apm-wo-analysis/slack-credentials`, `apm-wo-analysis/anthropic-api-key` |
| SSM | `/apm-wo-analysis/grafana-base-url` |
### Triage
1. **Export uploaded?** `aws s3 ls s3://apm-wo-analysis-exports-328440206208/raw/` — today's file present? Absent → upstream (§2.1), not the pipeline.
2. **Classifier ran/failed?** `/aws/lambda/apm-wo-analysis-classifier` logs; peek the DLQ: `aws sqs receive-message --queue-url <dlq-url> --max-number-of-messages 1`.
3. **Outputs written?** `aws s3 ls .../analytics/dt=<today>/` (Parquet) and `.../meta/dt=<today>/` (`summary.json`, `details.json`).
4. **slack-post ran/failed?** `/aws/lambda/apm-wo-analysis-slack-post` logs — `SlackApiError` (`invalid_auth`, `not_in_channel`, `invalid_blocks`)?
5. **Slack creds** valid + bot still in channel? (`apm-wo-analysis/slack-credentials`).
6. **IAM/throttle:** grep logs for `AccessDenied` / throttling.
7. **Recent change?** Any merge/deploy to `main` just before the failure.
### Resolution (by root cause)
1. **Export not uploaded** → `aws s3 cp <export>.xlsx s3://apm-wo-analysis-exports-328440206208/raw/`; then check the drop-folder agent (§2.1).
2. **Classifier failed (DLQ)** → read the DLQ message; fix; **reprocess by re-uploading the export to `raw/`** (`overwrite_partitions` makes same-`dt` idempotent).
3. **Outputs present, no Slack post** → re-invoke:
```bash
aws lambda invoke --function-name apm-wo-analysis-slack-post \
--payload "$(printf '{"dt":"<YYYY-MM-DD>"}' | base64)" /tmp/out.json
```
If Slack auth was the cause → rotate `apm-wo-analysis/slack-credentials` (and/or `/invite` the bot), then re-invoke (secret read per-call; no redeploy).
4. **Code regression** → identify the PR, revert/hotfix, redeploy via CI (push to `main`).
5. **Data present, Grafana empty** → datasource/dashboard issue (§2.3); check the instance via SSM.
6. **Anthropic/Haiku down** → non-fatal (deterministic path still classifies ~95%); set classifier env `APM_HAIKU_FALLBACK=off` to bypass.
### Post-Incident
- Verify re-process → summary posts + Grafana shows today's `dt`.
- Check for **other missed days** (gaps in `analytics/dt=…`/`meta/`) and reprocess each.
- Update README/Confluence if knowledge changed; add a memory entry; add a test if code caused it.
---
## 2. Operational Procedures
### 2.1 How the export gets uploaded
The curated daily APM filter-view export (~350 WOs, `.xlsx`/`.csv`) reaches S3 by
**direct upload or a local drop-folder** — never SES/email. Any object under
`raw/` with a `.xlsx`/`.csv` suffix triggers the classifier.
- **Direct:** `aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-328440206208/raw/`
- **Drop-folder (zero-touch):** a launchd agent (`com.seahaven.apm-wo-uploader`)
watches `~/apm-wo-drop/`, uploads new files to `raw/` using the scoped
`apm-wo-drop` profile (IAM user **`apm-wo-drop-uploader`** — `s3:PutObject` on
`raw/*` only), and archives them locally. The runnable script lives at
`~/.local/bin/apm-wo-uploader.sh` (it and the watched folder **must** be outside
`~/Documents` — macOS TCC sandbox).
**Verify the agent:**
```bash
launchctl list | grep apm-wo-uploader # present + last exit 0
tail -f ~/apm-wo-drop/uploader.log # per-run logging (TBD: confirm log path)
```
**Common upstream issues:** agent unloaded (`launchctl load -w …plist`); script
moved back into `~/Documents` (TCC blocks it — `LastExitStatus=32256`); `apm-wo-drop`
access key expired/rotated (`aws configure --profile apm-wo-drop`).
### 2.2 Grafana OS / app patching cadence
The Grafana EC2 box (`t4g.small`, **Amazon Linux 2023**, ARM64) is the only
patch-bearing piece — everything else is serverless. It's **reproducible from
`cdk/assets/grafana_userdata.sh`**, so the preferred patch path is a **clean
instance replacement** rather than long-lived in-place drift.
- **OS (recommended monthly + on critical CVEs):** via SSM —
`sudo dnf upgrade --security -y && sudo reboot` (Session Manager, or an SSM
Run Command / Patch Manager maintenance window — **TBD: not yet automated**).
- **Grafana OSS:** `sudo dnf upgrade grafana -y && sudo systemctl restart grafana-server`
(installed from the pinned `rpm.grafana.com` repo).
- **Athena datasource plugin:** pinned to **`3.2.0`** in `cdk/cdk.json`
(`athenaPluginVersion`). Bump there, then redeploy/replace the instance.
- **Preferred = clean replacement:** terminate the instance; `cdk deploy
apm-wo-analysis-grafana` relaunches it from the latest AL2023 AMI and re-runs
user-data (fresh Grafana + plugin + config sync). The root volume is
`DeleteOnTermination=false`, so detach/reuse or restore `grafana.db` (§2.4) if
local settings must persist. Validates the committed user-data from a cold boot.
> **Outstanding:** one clean instance replacement is owed to validate cold-boot
> user-data and apply root-volume encryption (encryption can't be added in place).
### 2.3 Dashboard-JSON redeploy
Source of truth is **`grafana/dashboards/apm-work-orders.json`** in this repo
(uid **`apm-wo`**); the running instance is never the source of truth
(`allowUiUpdates: false` — UI edits are reverted on the next sync).
**Flow:** edit JSON in repo → `cdk deploy apm-wo-analysis-grafana` (the
`BucketDeployment` uploads `grafana/` to `s3://…/grafana-config/`) → the instance
syncs S3 → `/var/lib/grafana/dashboards/` (on boot + a **15-min systemd timer**)
→ Grafana's file provider polls every **60 s** and reloads.
**Apply immediately** (skip the timer) via SSM:
```bash
sudo /usr/local/bin/grafana-config-sync.sh # pulls grafana-config/ from S3
# Grafana file provider picks up the dashboard within ~60s
```
**Gotchas:**
- The sync uses `aws s3 sync --exact-timestamps` — required so **same-size edits**
(e.g. a one-char query change) actually propagate.
- **Datasource/provisioning** changes (`grafana/provisioning/*.yaml`) are loaded
at **startup** — after syncing, `sudo systemctl restart grafana-server` (a
dashboard-only change does **not** need a restart).
- Athena query key is **`rawSQL`** (capital); datasource `authType: default`;
template vars `refresh: 1`. (See README "Notes / Gotchas".)
### 2.4 Backup & restore (config + EBS / grafana.db)
Two distinct layers:
**Config (dashboards, datasources, provisioning)** — fully **reproducible from
git** (`grafana/` → S3 `grafana-config/`). *Restore:* `cdk deploy
apm-wo-analysis-grafana` (or `grafana-config-sync.sh` on the box). No snapshot needed.
**Local state (`/var/lib/grafana/grafana.db`)** — Grafana's SQLite (admin user,
any API keys, org prefs). Lives on the **gp3 root volume** (encrypted,
`DeleteOnTermination=false`). Backed up by a **daily DLM snapshot** (07:00 UTC,
7 retained) of the instance (tag `apm-grafana-backup=true`), policy in the
grafana stack.
*Restore from snapshot:*
```bash
# find the latest DLM snapshot
aws ec2 describe-snapshots --owner-ids self \
--filters "Name=tag:aws:dlm:lifecycle-policy-id,Values=*" \
--query 'reverse(sort_by(Snapshots,&StartTime))[0].SnapshotId' --output text
# create a volume from it and attach to a replacement instance, OR mount it and
# copy /var/lib/grafana/grafana.db onto the new instance, then:
sudo systemctl restart grafana-server
```
Because dashboards + datasource are provisioned from code, the only thing the
snapshot uniquely protects is `grafana.db` (admin/login state) — low stakes; a
fresh instance + provisioning recovers everything else.
> **Note:** the Grafana **admin auth model is unsettled (TBD)** — the password was
> reset ad-hoc during build/testing. Decide the intended model (fixed admin
> password in Secrets Manager / SSO / anonymous view-only for the kiosk) and
> document it here.
### 2.5 Common failures (quick index)
| Symptom | Likely cause | Go to |
|---|---|---|
| No daily Slack post / no new dashboard day | export not uploaded, classifier failed (DLQ), or slack-post failed | §1 |
| Dashboard loads but all panels "No data" | datasource/auth, `rawSQL`, template `refresh`, or JSON in `analytics/` prefix | §2.3, README gotchas |
| Grafana unreachable | instance down / ALB unhealthy / office IP changed (`officeCidrs`) | SSM triage; `aws elbv2 describe-target-health` |
| Slack modal click does nothing / error | `apm-wo-analysis-slack-interactions`, API Gateway, or signing-secret mismatch | `/aws/lambda/apm-wo-analysis-slack-interactions` logs |
| Exports never arrive in `raw/` | drop-folder agent unloaded / TCC / expired key | §2.1 |
| Deploy not applying | OIDC role, CloudFormation rollback, Docker bundling | CloudFormation events; CI logs |
---
*Maintained in-repo (`docs/RUNBOOK.md`) and mirrored to Confluence. Update both
when operational knowledge changes.*