mirror of
https://github.com/Sea-Haven-Industries/apm-wo-analysis.git
synced 2026-09-30 06:33:14 +00:00
188 lines
10 KiB
Markdown
188 lines
10 KiB
Markdown
# apm-wo-analysis — Operational Runbook
|
|
|
|
Operational procedures and incident response for the daily APM work-order
|
|
analysis pipeline. Account **328440206208** / **us-east-1**. Stacks
|
|
`apm-wo-analysis-pipeline` and `apm-wo-analysis-grafana`. See [`README.md`](../README.md)
|
|
for architecture and resource detail.
|
|
|
|
**Admin access:** the Grafana box is **SSM Session Manager only** (no SSH/key pair):
|
|
`aws ssm start-session --target <instance-id>`. AWS API via the `office_mac` /
|
|
deploy roles. Grafana UI is office-IP-restricted at the ALB.
|
|
|
|
---
|
|
|
|
## 1. Incident Runbook: Missing daily APM WO analysis
|
|
|
|
**Severity: High** — a day's WO analysis is missing: escalations (incl. 3rd-escalation
|
|
alerts) aren't surfaced and the dashboard has no new snapshot. Recoverable by
|
|
re-processing the export; not permanent data loss.
|
|
|
|
### Detection
|
|
- Alarm **`apm-wo-analysis-classifier-invocations`** (Sum Invocations over 86400 s `< 1`, `treat_missing_data=breaching`) — a silent day, including a missing export.
|
|
- No **daily summary** in the WO Slack channel by the usual time (or no 3rd-escalation alert on a day one's expected).
|
|
- **Grafana** shows no `dt = today` in the Snapshot Date dropdown / panels empty for today.
|
|
- **Messages in `apm-wo-analysis-classifier-dlq`** (SQS) — strongest signal the classifier failed after it *did* run.
|
|
- CloudWatch errors in `/aws/lambda/apm-wo-analysis-classifier` or `-slack-post`.
|
|
|
|
### Context
|
|
| Item | Value |
|
|
|---|---|
|
|
| Stacks | `apm-wo-analysis-pipeline`, `apm-wo-analysis-grafana` |
|
|
| Lambdas | `apm-wo-analysis-classifier`, `-slack-post`, `-slack-interactions` |
|
|
| S3 | `apm-wo-analysis-exports-328440206208` — `raw/`, `analytics/dt=…/`, `meta/dt=…/` |
|
|
| SQS DLQ | `apm-wo-analysis-classifier-dlq` |
|
|
| Glue / Athena | db `apm_wo_analysis`, table `apm_wo_snapshots`, workgroup `apm-wo-analysis` |
|
|
| External | Slack, Anthropic API (Haiku fallback) |
|
|
| Secrets (names) | `apm-wo-analysis/slack-credentials`, `apm-wo-analysis/anthropic-api-key` |
|
|
| SSM | `/apm-wo-analysis/grafana-base-url` |
|
|
|
|
### Triage
|
|
1. **Export uploaded?** `aws s3 ls s3://apm-wo-analysis-exports-328440206208/raw/` — today's file present? Absent → upstream (§2.1), not the pipeline.
|
|
2. **Classifier ran/failed?** `/aws/lambda/apm-wo-analysis-classifier` logs; peek the DLQ: `aws sqs receive-message --queue-url <dlq-url> --max-number-of-messages 1`.
|
|
3. **Outputs written?** `aws s3 ls .../analytics/dt=<today>/` (Parquet) and `.../meta/dt=<today>/` (`summary.json`, `details.json`).
|
|
4. **slack-post ran/failed?** `/aws/lambda/apm-wo-analysis-slack-post` logs — `SlackApiError` (`invalid_auth`, `not_in_channel`, `invalid_blocks`)?
|
|
5. **Slack creds** valid + bot still in channel? (`apm-wo-analysis/slack-credentials`).
|
|
6. **IAM/throttle:** grep logs for `AccessDenied` / throttling.
|
|
7. **Recent change?** Any merge/deploy to `main` just before the failure.
|
|
|
|
### Resolution (by root cause)
|
|
1. **Export not uploaded** → `aws s3 cp <export>.xlsx s3://apm-wo-analysis-exports-328440206208/raw/`; then check the drop-folder agent (§2.1).
|
|
2. **Classifier failed (DLQ)** → read the DLQ message; fix; **reprocess by re-uploading the export to `raw/`** (`overwrite_partitions` makes same-`dt` idempotent).
|
|
3. **Outputs present, no Slack post** → re-invoke:
|
|
```bash
|
|
aws lambda invoke --function-name apm-wo-analysis-slack-post \
|
|
--payload "$(printf '{"dt":"<YYYY-MM-DD>"}' | base64)" /tmp/out.json
|
|
```
|
|
If Slack auth was the cause → rotate `apm-wo-analysis/slack-credentials` (and/or `/invite` the bot), then re-invoke (secret read per-call; no redeploy).
|
|
4. **Code regression** → identify the PR, revert/hotfix, redeploy via CI (push to `main`).
|
|
5. **Data present, Grafana empty** → datasource/dashboard issue (§2.3); check the instance via SSM.
|
|
6. **Anthropic/Haiku down** → non-fatal (deterministic path still classifies ~95%); set classifier env `APM_HAIKU_FALLBACK=off` to bypass.
|
|
|
|
### Post-Incident
|
|
- Verify re-process → summary posts + Grafana shows today's `dt`.
|
|
- Check for **other missed days** (gaps in `analytics/dt=…`/`meta/`) and reprocess each.
|
|
- Update README/Confluence if knowledge changed; add a memory entry; add a test if code caused it.
|
|
|
|
---
|
|
|
|
## 2. Operational Procedures
|
|
|
|
### 2.1 How the export gets uploaded
|
|
|
|
The curated daily APM filter-view export (~350 WOs, `.xlsx`/`.csv`) reaches S3 by
|
|
**direct upload or a local drop-folder** — never SES/email. Any object under
|
|
`raw/` with a `.xlsx`/`.csv` suffix triggers the classifier.
|
|
|
|
- **Direct:** `aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-328440206208/raw/`
|
|
- **Drop-folder (zero-touch):** a launchd agent (`com.seahaven.apm-wo-uploader`)
|
|
watches `~/apm-wo-drop/`, uploads new files to `raw/` using the scoped
|
|
`apm-wo-drop` profile (IAM user **`apm-wo-drop-uploader`** — `s3:PutObject` on
|
|
`raw/*` only), and archives them locally. The runnable script lives at
|
|
`~/.local/bin/apm-wo-uploader.sh` (it and the watched folder **must** be outside
|
|
`~/Documents` — macOS TCC sandbox).
|
|
|
|
**Verify the agent:**
|
|
```bash
|
|
launchctl list | grep apm-wo-uploader # present + last exit 0
|
|
tail -f ~/.local/log/apm-wo-uploader.log # per-run logging (outside WatchPaths)
|
|
```
|
|
**Common upstream issues:** agent unloaded (`launchctl load -w …plist`); script
|
|
moved back into `~/Documents` (TCC blocks it — `LastExitStatus=32256`); `apm-wo-drop`
|
|
access key expired/rotated (`aws configure --profile apm-wo-drop`).
|
|
|
|
### 2.2 Grafana OS / app patching cadence
|
|
|
|
The Grafana EC2 box (`t4g.small`, **Amazon Linux 2023**, ARM64) is the only
|
|
patch-bearing piece — everything else is serverless. It's **reproducible from
|
|
`cdk/assets/grafana_userdata.sh`**, so the preferred patch path is a **clean
|
|
instance replacement** rather than long-lived in-place drift.
|
|
|
|
- **OS (recommended monthly + on critical CVEs):** via SSM —
|
|
`sudo dnf upgrade --security -y && sudo reboot` (Session Manager, or an SSM
|
|
Run Command / Patch Manager maintenance window — **TBD: not yet automated**).
|
|
- **Grafana OSS:** `sudo dnf upgrade grafana -y && sudo systemctl restart grafana-server`
|
|
(installed from the pinned `rpm.grafana.com` repo).
|
|
- **Athena datasource plugin:** pinned to **`3.2.0`** in `cdk/cdk.json`
|
|
(`athenaPluginVersion`). Bump there, then redeploy/replace the instance.
|
|
- **Preferred = clean replacement:** terminate the instance; `cdk deploy
|
|
apm-wo-analysis-grafana` relaunches it from the latest AL2023 AMI and re-runs
|
|
user-data (fresh Grafana + plugin + config sync). The root volume is
|
|
`DeleteOnTermination=false`, so detach/reuse or restore `grafana.db` (§2.4) if
|
|
local settings must persist. Validates the committed user-data from a cold boot.
|
|
|
|
> **Outstanding:** one clean instance replacement is owed to validate cold-boot
|
|
> user-data and apply root-volume encryption (encryption can't be added in place).
|
|
|
|
### 2.3 Dashboard-JSON redeploy
|
|
|
|
Source of truth is **`grafana/dashboards/apm-work-orders.json`** in this repo
|
|
(uid **`apm-wo`**); the running instance is never the source of truth
|
|
(`allowUiUpdates: false` — UI edits are reverted on the next sync).
|
|
|
|
**Flow:** edit JSON in repo → `cdk deploy apm-wo-analysis-grafana` (the
|
|
`BucketDeployment` uploads `grafana/` to `s3://…/grafana-config/`) → the instance
|
|
syncs S3 → `/var/lib/grafana/dashboards/` (on boot + a **15-min systemd timer**)
|
|
→ Grafana's file provider polls every **60 s** and reloads.
|
|
|
|
**Apply immediately** (skip the timer) via SSM:
|
|
```bash
|
|
sudo /usr/local/bin/grafana-config-sync.sh # pulls grafana-config/ from S3
|
|
# Grafana file provider picks up the dashboard within ~60s
|
|
```
|
|
**Gotchas:**
|
|
- The sync uses `aws s3 sync --exact-timestamps` — required so **same-size edits**
|
|
(e.g. a one-char query change) actually propagate.
|
|
- **Datasource/provisioning** changes (`grafana/provisioning/*.yaml`) are loaded
|
|
at **startup** — after syncing, `sudo systemctl restart grafana-server` (a
|
|
dashboard-only change does **not** need a restart).
|
|
- Athena query key is **`rawSQL`** (capital); datasource `authType: default`;
|
|
template vars `refresh: 1`. (See README "Notes / Gotchas".)
|
|
|
|
### 2.4 Backup & restore (config + EBS / grafana.db)
|
|
|
|
Two distinct layers:
|
|
|
|
**Config (dashboards, datasources, provisioning)** — fully **reproducible from
|
|
git** (`grafana/` → S3 `grafana-config/`). *Restore:* `cdk deploy
|
|
apm-wo-analysis-grafana` (or `grafana-config-sync.sh` on the box). No snapshot needed.
|
|
|
|
**Local state (`/var/lib/grafana/grafana.db`)** — Grafana's SQLite (admin user,
|
|
any API keys, org prefs). Lives on the **gp3 root volume** (encrypted,
|
|
`DeleteOnTermination=false`). Backed up by a **daily DLM snapshot** (07:00 UTC,
|
|
7 retained) of the instance (tag `apm-grafana-backup=true`), policy in the
|
|
grafana stack.
|
|
|
|
*Restore from snapshot:*
|
|
```bash
|
|
# find the latest DLM snapshot
|
|
aws ec2 describe-snapshots --owner-ids self \
|
|
--filters "Name=tag:aws:dlm:lifecycle-policy-id,Values=*" \
|
|
--query 'reverse(sort_by(Snapshots,&StartTime))[0].SnapshotId' --output text
|
|
# create a volume from it and attach to a replacement instance, OR mount it and
|
|
# copy /var/lib/grafana/grafana.db onto the new instance, then:
|
|
sudo systemctl restart grafana-server
|
|
```
|
|
Because dashboards + datasource are provisioned from code, the only thing the
|
|
snapshot uniquely protects is `grafana.db` (admin/login state) — low stakes; a
|
|
fresh instance + provisioning recovers everything else.
|
|
|
|
> **Note:** the Grafana **admin auth model is unsettled (TBD)** — the password was
|
|
> reset ad-hoc during build/testing. Decide the intended model (fixed admin
|
|
> password in Secrets Manager / SSO / anonymous view-only for the kiosk) and
|
|
> document it here.
|
|
|
|
### 2.5 Common failures (quick index)
|
|
|
|
| Symptom | Likely cause | Go to |
|
|
|---|---|---|
|
|
| No daily Slack post / no new dashboard day | export not uploaded, classifier failed (DLQ), or slack-post failed | §1 |
|
|
| Dashboard loads but all panels "No data" | datasource/auth, `rawSQL`, template `refresh`, or JSON in `analytics/` prefix | §2.3, README gotchas |
|
|
| Grafana unreachable | instance down / ALB unhealthy / office IP changed (`officeCidrs`) | SSM triage; `aws elbv2 describe-target-health` |
|
|
| Slack modal click does nothing / error | `apm-wo-analysis-slack-interactions`, API Gateway, or signing-secret mismatch | `/aws/lambda/apm-wo-analysis-slack-interactions` logs |
|
|
| Exports never arrive in `raw/` | drop-folder agent unloaded / TCC / expired key | §2.1 |
|
|
| Deploy not applying | OIDC role, CloudFormation rollback, Docker bundling | CloudFormation events; CI logs |
|
|
|
|
---
|
|
|
|
*Maintained in-repo (`docs/RUNBOOK.md`) and mirrored to Confluence. Update both
|
|
when operational knowledge changes.*
|