apm-wo-analysis/README.md
2026-05-28 17:49:57 -04:00

162 lines
7.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# apm-wo-analysis
Daily analysis of Amazon **APM work-order** "Last Comment" data for Sea Haven
facility ops. A curated daily filter-view export (~350 work orders) is classified
on two axes, pushed to Slack, and surfaced in a self-hosted Grafana dashboard.
Replaces a legacy Google Apps Script + versioned-Google-Sheet workflow.
This is a **sibling concern** to the `apm@` email pipeline in `procurement-ingest`
— it consumes a different feed (the manual export) and does **not** read those
tables. There is deliberately **no DynamoDB**: this is an analytics workload
backed by S3 + Athena (Grafana cannot query DynamoDB).
- **Account / region:** 328440206208 / us-east-1
- **IaC:** CDK (Python), `aws-cdk-lib==2.253.1`. Lambdas Python 3.12, ARM64.
## Architecture
```
APM export (xlsx/csv)
→ S3 raw/ (direct upload OR local launchd drop-folder)
→ classifier Lambda (HTML strip + two-axis classify, Haiku fallback)
→ S3 analytics/dt=YYYY-MM-DD/ (per-WO daily snapshot, Parquet)
→ Glue table → Athena → Grafana (self-hosted EC2, VPN-only, kiosk)
→ slack-post Lambda (reads today + yesterday partitions)
→ daily summary post [📊 Open dashboard button]
→ standalone batched 3rd-escalation alert (suppressed if zero)
```
Two CDK stacks:
| Stack | Resources |
|---|---|
| `apm-wo-analysis-pipeline` | S3 exports bucket, classifier + slack-post Lambdas, Glue database, Athena workgroup, IAM |
| `apm-wo-analysis-grafana` | EC2 (Grafana OSS), internal ALB, security group, Route53, Athena datasource IAM role |
## The classification model
**Always two-axis, never comment-only.** The legacy script's central flaw was
reading only the comment while ignoring `WO Status` + `Hold Reason`, which left
~17% in "Other". The two-axis model cuts that to ~9% before any AI — and on the
real 347-row export the current implementation lands "Other" at **5.2%** (18
rows) with **18 mismatches** flagged.
- **Axis 1 — comment intent:** regex over the HTML-stripped `Last Comment`,
most-specific first (escalations → SIM ticket → vendor no-show → scheduling →
reports → completion → … → other).
- **Axis 2 — structured state:** `Hold Reason` → category and `WO Status`
signals (`RCAN`→Cancelled, `H` corroborates On Hold, `IP`/`R`/`RR` in-flight).
- **Resolution:** comment intent wins when confident → else structured state →
else `Other`. A **Claude Haiku** fallback (Secrets Manager) is reserved for
ambiguous free-text with no structured signal.
- **Mismatch detector (a feature):** flags when comment intent contradicts
structured state. Surfaced, never suppressed.
The authoritative spec lives in [`CLAUDE.md`](./CLAUDE.md); the implementation is
in `lambdas/classifier/classify.py` (owned by the `classifier-engineer` agent).
The S3-triggered `lambdas/classifier/handler.py` parses each export, classifies
every non-blank-comment row, writes a per-WO Parquet snapshot to
`analytics/dt=YYYY-MM-DD/` (registering the Glue partition via awswrangler), and
emits a `summary.json` for the Phase 4 slack-post Lambda.
## Repository layout
```
cdk/
app.py CDK entry point — instantiates both stacks
cdk.json
requirements.txt aws-cdk-lib==2.253.1, constructs>=10.6.0
stacks/
pipeline_stack.py S3, Lambdas, Glue, Athena, IAM
grafana_stack.py VPC import, EC2, ALB, SG, Route53, datasource role
lambdas/
classifier/ S3-triggered: parse → two-axis classify → Parquet
slack_post/ builds + posts the daily summary and alert
grafana/
provisioning/ Athena datasource + dashboard provider (as code)
dashboards/ committed dashboard JSON (source of truth)
scripts/ local drop-folder uploader + launchd plist
tests/ classifier smoke test
docs/BUILD.md phased, end-to-end build guide
```
## Configuration
| Where | What |
|---|---|
| **Secrets Manager** | `apm-wo-analysis/anthropic-api-key` (Haiku fallback); `apm-wo-analysis/slack-credentials` = `{ botToken, signingSecret, channelId }` (reused Slack app). |
| **SSM Parameter Store** | `/apm-wo-analysis/grafana-base-url` (the 📊 dashboard button / modal overflow link; ops-editable). |
| **GitHub repo secret** | `AWS_DEPLOY_ROLE_ARN` — the OIDC deploy role `githubdeploy-apm-wo-analysis`. |
No secrets in Lambda environment variables.
## Ingestion (no email)
The export reaches S3 by **direct upload or a local drop-folder**, never SES/email.
- **Direct:** `aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-328440206208/raw/`
- **Drop-folder (optional zero-touch):** a launchd agent (`scripts/apm-wo-uploader.sh`
+ `scripts/com.seahaven.apm-wo-uploader.plist`) that watches `~/apm-wo-drop/`,
uploads new `.xlsx`/`.csv` files to `raw/`, and archives them to `uploaded/`.
It uploads with the scoped `apm-wo-drop` AWS profile (IAM user
`apm-wo-drop-uploader` — `s3:PutObject` on `raw/*` only).
Install (the runnable copy **must** live outside `~/Documents` — macOS TCC
sandbox; a repo-path script fails silently with `LastExitStatus=32256`):
```bash
install -d "$HOME/.local/bin" "$HOME/apm-wo-drop"
cp scripts/apm-wo-uploader.sh "$HOME/.local/bin/apm-wo-uploader.sh"
chmod +x "$HOME/.local/bin/apm-wo-uploader.sh"
cp scripts/com.seahaven.apm-wo-uploader.plist "$HOME/Library/LaunchAgents/"
launchctl load -w "$HOME/Library/LaunchAgents/com.seahaven.apm-wo-uploader.plist"
```
Re-copy the script to `~/.local/bin` after editing the repo source. Configure
the profile once with the uploader's access key:
`aws configure --profile apm-wo-drop`.
The classifier Lambda is S3-triggered on the `raw/` prefix regardless of path
(any `.xlsx`/`.csv` landing under `raw/` invokes it).
## Deployment
CI/CD via the org reusable workflows (no manual prod deploys):
- **CI** (`.github/workflows/ci.yaml`) → `ci-python-sam.yaml@main`: ruff + `cdk synth`.
- **Deploy** (`.github/workflows/deploy.yaml`) → `cd-cdk.yaml@main`: OIDC assume-role,
`cdk deploy --all`, single-flight concurrency.
The OIDC deploy role must exist **before** the first deploy. Deploy order:
```bash
cd cdk && pip install -r requirements.txt
cdk deploy apm-wo-analysis-pipeline # S3, Glue, Athena, Lambdas, IAM
cdk deploy apm-wo-analysis-grafana # EC2, ALB, SG, Route53, datasource role
```
## Local development
- pyenv Python 3.12; `ruff check` + `ruff format --check` before pushing (hook-enforced).
- Smoke-test the classifier against a **real export** before declaring any
classification change done: `~/Downloads/_documents/Sheet1-1.xlsx`.
- `cdk synth` must pass in CI before merge.
## Status
**Phase 4 — Slack surfaces (in review).** Build-out proceeds per
[`docs/BUILD.md`](./docs/BUILD.md): ingestion → classifier → Glue/Athena → Slack
→ Grafana → docs.
- **Phase 0** scaffold — merged-pending (PR #6).
- **Phase 1** ingestion (S3 bucket, drop-folder uploader, OIDC deploy role) —
deployed; PR #7 open.
- **Phase 2** classifier Lambda + Glue database + S3 `raw/` trigger — implemented
and `cdk synth`-green; PR #8 open (stacked on Phase 1, not yet deployed).
- **Phase 3** `apm_wo_snapshots` projection table + Athena workgroup —
implemented and `cdk synth`-green; PR #9 open (stacked on Phase 2). The
classifier writes pure Parquet and holds no Glue access (projection handles
partitions).
- **Phase 4** Slack post + interactions Lambdas (daily summary, batched
3rd-escalation alert, drill-down modals on `apm-wo.seahaven.com`) — implemented
and `cdk synth`-green; PR #10 open (stacked on Phase 3). App manifest in
`slack/manifest.yaml`.
- **Phases 5–6** (Grafana, final docs) — not started.