apm-wo-analysis/README.md
Adam Moussa 60e0b878e3 Refresh README to full operational doc (Phase 6)
Rewrite to the Sea Haven operational template with real resource names from both
stacks: AWS Resources + Lambda Functions tables, Configuration (Secrets/SSM/env/
context), Operations (verify, logs, classifier DLQ, reprocess, Grafana admin),
Documentation, and Notes/Gotchas incl. the deploy-time lessons + a known-debt
list. Corrects stale bits: ALB is internet-facing + office-IP-restricted (not
internal); classifier uses partition projection (no runtime Glue registration);
summary/details JSON live under meta/ not analytics/; grafana uses an instance
role. Adds the resources missing from the old README (classifier DLQ, slack-
interactions Lambda, HTTP API, meta/ + grafana-config/ prefixes, DLM backup,
encrypted volume).
2026-05-29 12:27:05 -04:00

258 lines
15 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# apm-wo-analysis
Daily analysis of Amazon **APM work-order** "Last Comment" data for Sea Haven
facility ops. A curated daily filter-view export (~350 work orders) is classified
on two axes (comment intent + structured `WO Status`/`Hold Reason`), pushed to
Slack as a daily summary plus a batched 3rd-escalation alert, and surfaced in a
self-hosted Grafana dashboard. Replaces a legacy Google Apps Script +
versioned-Google-Sheet workflow.
This is a **sibling concern** to the `apm@` email pipeline in `procurement-ingest`
— it consumes a different feed (the manual export) and does **not** read those
tables. There is deliberately **no DynamoDB**: this is an analytics workload
backed by S3 + Athena (Grafana cannot query DynamoDB).
- **Account / region:** 328440206208 / us-east-1
- **IaC:** CDK (Python), `aws-cdk-lib==2.253.1`. Lambdas Python 3.12, ARM64.
## Architecture
```
APM export (xlsx/csv)
→ S3 raw/ (direct `aws s3 cp` OR local launchd drop-folder)
→ classifier Lambda (HTML strip + two-axis classify, Haiku fallback)
├→ S3 analytics/dt=YYYY-MM-DD/ (per-WO snapshot, Parquet)
│ → Glue table (partition projection) → Athena → Grafana (EC2, office-IP, kiosk)
├→ S3 meta/dt=YYYY-MM-DD/ (summary.json + details.json — NOT in the table prefix)
└→ async-invoke slack-post Lambda
→ daily summary post [📊 Open dashboard button]
→ standalone batched 3rd-escalation alert (suppressed when zero)
→ drill-down modals via apm-wo.seahaven.com (signature-verified)
```
Two CDK stacks (`cdk/app.py` instantiates both):
| Stack | Resources |
|---|---|
| `apm-wo-analysis-pipeline` | S3 exports bucket, drop-uploader IAM user, classifier + slack-post + slack-interactions Lambdas, classifier DLQ, Glue DB + projection table, Athena workgroup, HTTP API (`apm-wo.seahaven.com`), SSM param, scoped IAM |
| `apm-wo-analysis-grafana` | EC2 (Grafana OSS), internet-facing **office-IP-restricted** ALB (`grafana.seahaven.com`), security groups, instance IAM role, Route53 alias, daily DLM snapshot, dashboards-as-code S3 deployment |
## AWS Resources
| Resource | Name | Purpose |
|---|---|---|
| S3 bucket | `apm-wo-analysis-exports-328440206208` | Single bucket. Prefixes: `raw/` (incoming, 90-day expiry), `analytics/` (Parquet snapshots, kept), `meta/` (summary/details JSON), `athena-results/` (query output, 30-day expiry), `grafana-config/` (dashboards-as-code). SSE-S3, BPA-all, enforce-SSL, `RETAIN`. |
| IAM user | `apm-wo-drop-uploader` | Drop-folder identity; `s3:PutObject` on `raw/*` only. Access key created out-of-band, stored in local `apm-wo-drop` profile. |
| Glue database | `apm_wo_analysis` | Analytics catalog. |
| Glue table | `apm_wo_snapshots` | External Parquet table over `analytics/`, **partition projection** on `dt` (date, `2026-01-01..NOW`) — no crawler, no `MSCK`. 17-column snapshot schema. |
| Athena workgroup | `apm-wo-analysis` | Enforced result location `athena-results/`, SSE-S3. |
| SQS queue | `apm-wo-analysis-classifier-dlq` | Dead-letter for failed classifier async invocations (14-day retention). |
| SSM parameter | `/apm-wo-analysis/grafana-base-url` | Grafana dashboard URL for the 📊 button / modal links (ops-editable, no redeploy). |
| HTTP API + domain | `apm-wo.seahaven.com` → `POST /slack/interactions` | Slack interactivity endpoint. Stage throttled 10 rps / 20 burst. `*.seahaven.com` ACM cert; Route53 alias. |
| EC2 instance | Grafana (`t4g.small`, AL2023, ARM64) | Self-hosted Grafana OSS in `seahaven-vpc` private subnets, IMDSv2-only, SSM-managed. gp3 20 GB **encrypted**, `DeleteOnTermination=false`, tagged `apm-grafana-backup`. |
| ALB | Grafana ALB (`grafana.seahaven.com`) | Internet-facing, HTTPS 443, SG admits **only office CIDRs** (`47.21.61.4/32`, `96.250.164.146/32`); forwards to instance:3000, health `/api/health`. |
| DLM policy | Grafana volume backup | Daily snapshot (07:00 UTC) of the tagged instance, 7 retained. |
| Route53 | `apm-wo.seahaven.com`, `grafana.seahaven.com` | Aliases in zone `Z06652411XKH89KTZD3XA` (`seahaven.com`). |
## Lambda Functions
All Python 3.12, ARM64, explicit LogGroup (`/aws/lambda/<name>`, 60-day retention).
| Function | Trigger | Purpose |
|---|---|---|
| `apm-wo-analysis-classifier` | S3 `ObjectCreated` on `raw/*.xlsx|.csv` | Parse export, two-axis classify each non-blank-comment row, write per-WO Parquet to `analytics/dt=…/` and `summary.json`/`details.json` to `meta/dt=…/`, then async-invoke slack-post. 512 MB / 120 s. AWS-managed SDK-for-pandas layer (`AWSSDKPandas-Python312-Arm64:27`); DLQ attached. |
| `apm-wo-analysis-slack-post` | Async-invoked by the classifier (`{"dt": …}`) | Read today's + yesterday's `meta/.../summary.json`, post the daily summary, and (only when `third_escalation_count > 0`) the batched 3rd-escalation alert from `details.json`. 256 MB / 30 s. |
| `apm-wo-analysis-slack-interactions` | HTTP API `POST /slack/interactions` | Verify the Slack request signature, read `meta/.../details.json`, and `views.open` a filtered WO-list modal within Slack's 3 s `trigger_id` window. 256 MB / 30 s. |
## Configuration
### Secrets Manager (names only — created out-of-band, never in CloudFormation)
| Secret | Purpose |
|---|---|
| `apm-wo-analysis/anthropic-api-key` | Claude Haiku fallback for ambiguous free-text comments. |
| `apm-wo-analysis/slack-credentials` | JSON `{ botToken, signingSecret, channelId }` for the reused Slack app. |
### SSM Parameters
| Parameter | Purpose |
|---|---|
| `/apm-wo-analysis/grafana-base-url` | Grafana dashboard deep-link base (`https://grafana.seahaven.com/d/apm-wo/...`). |
### Environment Variables (non-secret)
- **classifier:** `APM_HAIKU_FALLBACK` (`on`/`off`), `SLACK_POST_FUNCTION_NAME`.
- **slack-post / slack-interactions:** `SLACK_SECRET_NAME`, `DASHBOARD_URL_PARAM`, `ANALYTICS_BUCKET`.
### GitHub repo secret
- `AWS_DEPLOY_ROLE_ARN` — the OIDC deploy role `githubdeploy-apm-wo-analysis`.
### CDK context (`cdk/cdk.json`)
`wildcardCertArn`, `hostedZoneId`/`hostedZoneName`, `slackInteractionsDomain`, `grafanaDomain`, `grafanaVpcId`/`grafanaAzs`/`grafana{Public,Private}SubnetIds`, `officeCidrs`, `athenaPluginVersion` (`3.2.0`).
## The classification model
**Always two-axis, never comment-only.** The legacy script's central flaw was
reading only the comment while ignoring `WO Status` + `Hold Reason`, which left
~17% in "Other". The two-axis model cuts that to ~9% before any AI — and on the
real 347-row export the implementation lands "Other" at **5.2%** (18 rows) with
**18 mismatches** flagged.
- **Axis 1 — comment intent:** regex over the HTML-stripped `Last Comment`,
most-specific first (escalations → SIM ticket → vendor no-show → scheduling →
reports → completion → … → other).
- **Axis 2 — structured state:** `Hold Reason` → category, and `WO Status`
signals (`RCAN`→Cancelled, `H` corroborates On Hold, `IP`/`R`/`RR` in-flight).
- **Resolution:** comment intent wins when confident → else structured state →
else `Other`. A **Claude Haiku** fallback (Secrets Manager key) is reserved for
ambiguous free-text with no structured signal.
- **Mismatch detector (a feature):** flags when comment intent contradicts
structured state (e.g. "completed" while `WO Status` is `IP`). Surfaced, never
suppressed.
Authoritative spec: [`CLAUDE.md`](./CLAUDE.md). Implementation:
`lambdas/classifier/classify.py` (logic) and `lambdas/classifier/handler.py`
(S3 → Parquet + JSON + invoke).
## Repository layout
```
cdk/
app.py CDK entry point — instantiates both stacks
cdk.json context: cert, zone, subnets, office CIDRs, plugin version
requirements.txt aws-cdk-lib==2.253.1, constructs>=10.6.0
assets/
grafana_userdata.sh EC2 bootstrap: install Grafana + Athena plugin, S3 config sync
stacks/
pipeline_stack.py S3, Lambdas, DLQ, Glue, Athena, HTTP API, IAM
grafana_stack.py VPC import, EC2, ALB, SG, Route53, instance role, DLM
lambdas/
classifier/ S3-triggered: parse → two-axis classify → Parquet + meta JSON
slack_post/ blockkit.py (builders), handler.py (post), interactions.py
(modals), slackio.py (Secrets/SSM/S3/signature)
grafana/
provisioning/ Athena datasource + dashboard provider (as code)
dashboards/ apm-work-orders.json (uid apm-wo) — source of truth
slack/manifest.yaml Slack app manifest (interactivity request URL)
scripts/ local drop-folder uploader + launchd plist
tests/ classifier smoke test + offline synth/blockkit assertions
docs/BUILD.md phased, end-to-end build guide
```
## Ingestion (no email)
The export reaches S3 by **direct upload or a local drop-folder**, never SES/email.
- **Direct:** `aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-328440206208/raw/`
- **Drop-folder (optional zero-touch):** a launchd agent (`scripts/apm-wo-uploader.sh`
+ `scripts/com.seahaven.apm-wo-uploader.plist`) that watches `~/apm-wo-drop/`,
uploads new `.xlsx`/`.csv` files to `raw/`, and archives them locally. Uploads
with the scoped `apm-wo-drop` profile (IAM user `apm-wo-drop-uploader`).
Install — the runnable copy and watched folder **must** live outside `~/Documents`
(macOS TCC sandbox; a repo-path script fails silently with `LastExitStatus=32256`):
```bash
install -d "$HOME/.local/bin" "$HOME/apm-wo-drop"
cp scripts/apm-wo-uploader.sh "$HOME/.local/bin/apm-wo-uploader.sh"
chmod +x "$HOME/.local/bin/apm-wo-uploader.sh"
cp scripts/com.seahaven.apm-wo-uploader.plist "$HOME/Library/LaunchAgents/"
launchctl load -w "$HOME/Library/LaunchAgents/com.seahaven.apm-wo-uploader.plist"
aws configure --profile apm-wo-drop # one-time, with the uploader's access key
```
Any `.xlsx`/`.csv` landing under `raw/` invokes the classifier.
## Local development
- pyenv Python 3.12. `ruff check` + `ruff format --check` before pushing (hook-enforced).
- Tests: `python -m pytest tests/ -q` (60 tests — classifier smoke test against a real
export + offline `cdk.assertions` synth checks + Block Kit builders). No AWS needed.
- Smoke-test the classifier against a **real export** before declaring any
classification change done: `~/Downloads/_documents/Sheet1-1.xlsx`.
- `cdk synth` must pass in CI before merge (**Docker required** — Lambda deps are
bundled for ARM64).
## Deployment
CI/CD via the org reusable workflows (no manual prod deploys in steady state):
- **CI** (`.github/workflows/ci.yaml`) → `ci-python-sam.yaml@main`: ruff + `cdk synth`. Runs on PRs into `main`.
- **Deploy** (`.github/workflows/deploy.yaml`) → `cd-cdk.yaml@main`: OIDC assume-role, `cdk deploy --all`, single-flight concurrency. Runs on push to `main`.
Stack name/region/account: `apm-wo-analysis-{pipeline,grafana}` / us-east-1 / 328440206208.
OIDC deploy role `githubdeploy-apm-wo-analysis` must exist before the first deploy.
> **Note:** CD is **temporarily disabled** (deploy job gated `if: ${{ false }}` on
> the phase-0 branch) while the Phase 0–5 stack is merged into `main`, to avoid a
> deploy on every merge. **Re-enable as the first Phase 6 step** by reverting that
> commit. Both stacks are already deployed manually and validated in prod.
Manual deploy (emergency/reference; pipeline first so the bucket/table exist):
```bash
cd cdk && pip install -r requirements.txt
cdk deploy apm-wo-analysis-pipeline # S3, Glue, Athena, Lambdas, DLQ, API, IAM
cdk deploy apm-wo-analysis-grafana # EC2, ALB, SG, Route53, DLM, dashboards
```
## Operations
**Verify it's working** — drop a real export into `raw/`, then within ~1 min:
- `analytics/dt=<today>/*.parquet` and `meta/dt=<today>/{summary,details}.json` appear,
- the daily summary posts to the WO Slack channel (+ alert if any 3rd escalations),
- the Grafana dashboard renders over the office network (`grafana.seahaven.com`).
**Logs:** CloudWatch `/aws/lambda/apm-wo-analysis-{classifier,slack-post,slack-interactions}` (60-day retention).
**Failure modes:**
- Classifier failure (malformed export, transient error) → after Lambda retries, the event lands in **`apm-wo-analysis-classifier-dlq`**. Check the DLQ if a day's data is missing.
- Slack post failure is best-effort and does **not** fail classification (data still lands in S3).
- Grafana down → check the instance via **SSM Session Manager** (no SSH); `systemctl status grafana-server`; ALB target health.
**Reprocess a day:** re-upload the same export to `raw/` — the classifier uses
`overwrite_partitions`, so a same-day re-run replaces that `dt` partition idempotently.
**Grafana admin:** access is office-IP-restricted at the ALB; the instance is
SSM-only. Dashboards are provisioned from `grafana-config/` in S3 (synced on boot
and by a 15-min systemd timer); **edit dashboards in-repo, not in the UI**
(`allowUiUpdates: false`). `grafana.db` lives on the RETAIN'd gp3 volume and is
snapshotted daily by DLM.
## Documentation
- **Confluence "AWS Architecture Map"** (IT space, page **1540098**): a Mermaid
subgraph for this stack — **TBD** (Phase 6; previously blocked on the Atlassian
MCP, now unblocked).
- **Slack Apps Inventory** (Confluence page 524569): add the "APM Work Orders" app — **TBD**.
- This README + `docs/BUILD.md` (phased build guide) + `CLAUDE.md` (domain spec).
## Notes / Gotchas
- **Partition date** comes from the **S3 event time**, not the Lambda wall-clock —
stable across retries and the midnight boundary.
- **`meta/` vs `analytics/`:** summary/details JSON must stay **out** of `analytics/` —
Athena reads every object in the table prefix as Parquet and chokes on JSON.
- **Classifier deps** (awswrangler/pandas/pyarrow/numpy) come from the AWS-managed
SDK-for-pandas **layer** — bundling them blows Lambda's 250 MB unzipped limit.
Verify the pinned layer ARN/version on region or runtime changes.
- **Grafana dashboard contract:** dashboard `uid` must stay `apm-wo` (the Slack 📊
button deep-links to `d/apm-wo`). Athena plugin query key is `rawSQL` (capital).
Datasource `authType: default` (the EC2 instance role; `ec2_iam_role` is rejected
by the plugin). Config sync uses `aws s3 sync --exact-timestamps` (plain sync
skips same-size edits). Template vars use `refresh: 1` (on load).
- **No Client VPN exists** — "VPN-only" Grafana is realized as **office-IP SG
restriction**. `seahaven-vpc` has a single NAT (one AZ) for instance egress.
- **Slack interactions endpoint** is unauthenticated at the gateway **by design**;
the Lambda verifies the Slack signature (replay window + HMAC). Stage-throttled.
### Known operational debt
- Re-enable CD (revert the phase-0 disable) once the stack is merged.
- One **clean instance replacement** is owed to validate the committed user-data
from a cold boot and to apply root-volume encryption (can't encrypt in place).
- Deferred review NITs: `print()`→`logging`, source-IP logging on signature
failure, S3 versioning, the `'${site:raw}'` WO-table SQL tidy, and a
CIDR-maintenance note in the runbook.
## Status
Phases 2–5 implemented, **deployed to prod and validated end-to-end** (classifier,
Slack post + alert + modal, Grafana dashboard). Stacked PRs **#6→#11** are open and
unmerged; cross-review (#2/#4/#5) and `/security-review` of the two public endpoints
are **cleared**. Phase 6 (this docs pass + Confluence + runbook) is in progress on
`feature/phase-6-docs`.