apm-wo-analysis/README.md
2026-05-28 18:06:10 -04:00

168 lines
8 KiB
Markdown

# apm-wo-analysis
Daily analysis of Amazon **APM work-order** "Last Comment" data for Sea Haven
facility ops. A curated daily filter-view export (~350 work orders) is classified
on two axes, pushed to Slack, and surfaced in a self-hosted Grafana dashboard.
Replaces a legacy Google Apps Script + versioned-Google-Sheet workflow.
This is a **sibling concern** to the `apm@` email pipeline in `procurement-ingest`
— it consumes a different feed (the manual export) and does **not** read those
tables. There is deliberately **no DynamoDB**: this is an analytics workload
backed by S3 + Athena (Grafana cannot query DynamoDB).
- **Account / region:** 328440206208 / us-east-1
- **IaC:** CDK (Python), `aws-cdk-lib==2.253.1`. Lambdas Python 3.12, ARM64.
## Architecture
```
APM export (xlsx/csv)
→ S3 raw/ (direct upload OR local launchd drop-folder)
→ classifier Lambda (HTML strip + two-axis classify, Haiku fallback)
→ S3 analytics/dt=YYYY-MM-DD/ (per-WO daily snapshot, Parquet)
→ Glue table → Athena → Grafana (self-hosted EC2, VPN-only, kiosk)
→ slack-post Lambda (reads today + yesterday partitions)
→ daily summary post [📊 Open dashboard button]
→ standalone batched 3rd-escalation alert (suppressed if zero)
```
Two CDK stacks:
| Stack | Resources |
|---|---|
| `apm-wo-analysis-pipeline` | S3 exports bucket, classifier + slack-post Lambdas, Glue database, Athena workgroup, IAM |
| `apm-wo-analysis-grafana` | EC2 (Grafana OSS), internal ALB, security group, Route53, Athena datasource IAM role |
## The classification model
**Always two-axis, never comment-only.** The legacy script's central flaw was
reading only the comment while ignoring `WO Status` + `Hold Reason`, which left
~17% in "Other". The two-axis model cuts that to ~9% before any AI — and on the
real 347-row export the current implementation lands "Other" at **5.2%** (18
rows) with **18 mismatches** flagged.
- **Axis 1 — comment intent:** regex over the HTML-stripped `Last Comment`,
most-specific first (escalations → SIM ticket → vendor no-show → scheduling →
reports → completion → … → other).
- **Axis 2 — structured state:** `Hold Reason` → category and `WO Status`
signals (`RCAN`→Cancelled, `H` corroborates On Hold, `IP`/`R`/`RR` in-flight).
- **Resolution:** comment intent wins when confident → else structured state →
else `Other`. A **Claude Haiku** fallback (Secrets Manager) is reserved for
ambiguous free-text with no structured signal.
- **Mismatch detector (a feature):** flags when comment intent contradicts
structured state. Surfaced, never suppressed.
The authoritative spec lives in [`CLAUDE.md`](./CLAUDE.md); the implementation is
in `lambdas/classifier/classify.py` (owned by the `classifier-engineer` agent).
The S3-triggered `lambdas/classifier/handler.py` parses each export, classifies
every non-blank-comment row, writes a per-WO Parquet snapshot to
`analytics/dt=YYYY-MM-DD/` (registering the Glue partition via awswrangler), and
emits a `summary.json` for the Phase 4 slack-post Lambda.
## Repository layout
```
cdk/
app.py CDK entry point — instantiates both stacks
cdk.json
requirements.txt aws-cdk-lib==2.253.1, constructs>=10.6.0
stacks/
pipeline_stack.py S3, Lambdas, Glue, Athena, IAM
grafana_stack.py VPC import, EC2, ALB, SG, Route53, datasource role
lambdas/
classifier/ S3-triggered: parse → two-axis classify → Parquet
slack_post/ builds + posts the daily summary and alert
grafana/
provisioning/ Athena datasource + dashboard provider (as code)
dashboards/ committed dashboard JSON (source of truth)
scripts/ local drop-folder uploader + launchd plist
tests/ classifier smoke test
docs/BUILD.md phased, end-to-end build guide
```
## Configuration
| Where | What |
|---|---|
| **Secrets Manager** | `apm-wo-analysis/anthropic-api-key` (Haiku fallback); `apm-wo-analysis/slack-credentials` = `{ botToken, signingSecret, channelId }` (reused Slack app). |
| **SSM Parameter Store** | `/apm-wo-analysis/grafana-base-url` (the 📊 dashboard button / modal overflow link; ops-editable). |
| **GitHub repo secret** | `AWS_DEPLOY_ROLE_ARN` — the OIDC deploy role `githubdeploy-apm-wo-analysis`. |
No secrets in Lambda environment variables.
## Ingestion (no email)
The export reaches S3 by **direct upload or a local drop-folder**, never SES/email.
- **Direct:** `aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-328440206208/raw/`
- **Drop-folder (optional zero-touch):** a launchd agent (`scripts/apm-wo-uploader.sh`
+ `scripts/com.seahaven.apm-wo-uploader.plist`) that watches `~/apm-wo-drop/`,
uploads new `.xlsx`/`.csv` files to `raw/`, and archives them to `uploaded/`.
It uploads with the scoped `apm-wo-drop` AWS profile (IAM user
`apm-wo-drop-uploader` — `s3:PutObject` on `raw/*` only).
Install (the runnable copy **must** live outside `~/Documents` — macOS TCC
sandbox; a repo-path script fails silently with `LastExitStatus=32256`):
```bash
install -d "$HOME/.local/bin" "$HOME/apm-wo-drop"
cp scripts/apm-wo-uploader.sh "$HOME/.local/bin/apm-wo-uploader.sh"
chmod +x "$HOME/.local/bin/apm-wo-uploader.sh"
cp scripts/com.seahaven.apm-wo-uploader.plist "$HOME/Library/LaunchAgents/"
launchctl load -w "$HOME/Library/LaunchAgents/com.seahaven.apm-wo-uploader.plist"
```
Re-copy the script to `~/.local/bin` after editing the repo source. Configure
the profile once with the uploader's access key:
`aws configure --profile apm-wo-drop`.
The classifier Lambda is S3-triggered on the `raw/` prefix regardless of path
(any `.xlsx`/`.csv` landing under `raw/` invokes it).
## Deployment
CI/CD via the org reusable workflows (no manual prod deploys):
- **CI** (`.github/workflows/ci.yaml`) → `ci-python-sam.yaml@main`: ruff + `cdk synth`.
- **Deploy** (`.github/workflows/deploy.yaml`) → `cd-cdk.yaml@main`: OIDC assume-role,
`cdk deploy --all`, single-flight concurrency.
The OIDC deploy role must exist **before** the first deploy. Deploy order:
```bash
cd cdk && pip install -r requirements.txt
cdk deploy apm-wo-analysis-pipeline # S3, Glue, Athena, Lambdas, IAM
cdk deploy apm-wo-analysis-grafana # EC2, ALB, SG, Route53, datasource role
```
## Local development
- pyenv Python 3.12; `ruff check` + `ruff format --check` before pushing (hook-enforced).
- Smoke-test the classifier against a **real export** before declaring any
classification change done: `~/Downloads/_documents/Sheet1-1.xlsx`.
- `cdk synth` must pass in CI before merge.
## Status
**Phase 5 — Grafana (in review); build complete, docs remain.** Build-out per
[`docs/BUILD.md`](./docs/BUILD.md): ingestion → classifier → Glue/Athena → Slack
→ Grafana → docs.
- **Phase 0** scaffold — merged-pending (PR #6).
- **Phase 1** ingestion (S3 bucket, drop-folder uploader, OIDC deploy role) —
deployed; PR #7 open.
- **Phase 2** classifier Lambda + Glue database + S3 `raw/` trigger — implemented
and `cdk synth`-green; PR #8 open (stacked on Phase 1, not yet deployed).
- **Phase 3** `apm_wo_snapshots` projection table + Athena workgroup —
implemented and `cdk synth`-green; PR #9 open (stacked on Phase 2). The
classifier writes pure Parquet and holds no Glue access (projection handles
partitions).
- **Phase 4** Slack post + interactions Lambdas (daily summary, batched
3rd-escalation alert, drill-down modals on `apm-wo.seahaven.com`) — implemented
and `cdk synth`-green; PR #10 open (stacked on Phase 3). App manifest in
`slack/manifest.yaml`.
- **Phase 5** self-hosted Grafana (EC2 + ALB on `grafana.seahaven.com`,
office-IP-restricted, 7-panel dashboard as code) — implemented and
`cdk synth`-green; PR #11 open (stacked on Phase 4).
- **Phase 6** (final docs / Confluence / runbook) — not started.
All phase PRs are stacked (#6→#7→#8→#9→#10→#11) and **unmerged**; nothing is
deployed yet. Cross-review and `/security-review` are outstanding across the stack.