2026-05-28 16:13:01 -04:00
|
|
|
|
# apm-wo-analysis
|
|
|
|
|
|
|
|
|
|
|
|
Daily analysis of Amazon **APM work-order** "Last Comment" data for Sea Haven
|
|
|
|
|
|
facility ops. A curated daily filter-view export (~350 work orders) is classified
|
|
|
|
|
|
on two axes, pushed to Slack, and surfaced in a self-hosted Grafana dashboard.
|
|
|
|
|
|
Replaces a legacy Google Apps Script + versioned-Google-Sheet workflow.
|
|
|
|
|
|
|
|
|
|
|
|
This is a **sibling concern** to the `apm@` email pipeline in `procurement-ingest`
|
|
|
|
|
|
— it consumes a different feed (the manual export) and does **not** read those
|
|
|
|
|
|
tables. There is deliberately **no DynamoDB**: this is an analytics workload
|
|
|
|
|
|
backed by S3 + Athena (Grafana cannot query DynamoDB).
|
|
|
|
|
|
|
|
|
|
|
|
- **Account / region:** 328440206208 / us-east-1
|
|
|
|
|
|
- **IaC:** CDK (Python), `aws-cdk-lib==2.253.1`. Lambdas Python 3.12, ARM64.
|
|
|
|
|
|
|
|
|
|
|
|
## Architecture
|
|
|
|
|
|
|
|
|
|
|
|
```
|
|
|
|
|
|
APM export (xlsx/csv)
|
|
|
|
|
|
→ S3 raw/ (direct upload OR local launchd drop-folder)
|
|
|
|
|
|
→ classifier Lambda (HTML strip + two-axis classify, Haiku fallback)
|
|
|
|
|
|
→ S3 analytics/dt=YYYY-MM-DD/ (per-WO daily snapshot, Parquet)
|
|
|
|
|
|
→ Glue table → Athena → Grafana (self-hosted EC2, VPN-only, kiosk)
|
|
|
|
|
|
→ slack-post Lambda (reads today + yesterday partitions)
|
|
|
|
|
|
→ daily summary post [📊 Open dashboard button]
|
|
|
|
|
|
→ standalone batched 3rd-escalation alert (suppressed if zero)
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
Two CDK stacks:
|
|
|
|
|
|
|
|
|
|
|
|
| Stack | Resources |
|
|
|
|
|
|
|---|---|
|
|
|
|
|
|
| `apm-wo-analysis-pipeline` | S3 exports bucket, classifier + slack-post Lambdas, Glue database, Athena workgroup, IAM |
|
|
|
|
|
|
| `apm-wo-analysis-grafana` | EC2 (Grafana OSS), internal ALB, security group, Route53, Athena datasource IAM role |
|
|
|
|
|
|
|
|
|
|
|
|
## The classification model
|
|
|
|
|
|
|
|
|
|
|
|
**Always two-axis, never comment-only.** The legacy script's central flaw was
|
|
|
|
|
|
reading only the comment while ignoring `WO Status` + `Hold Reason`, which left
|
2026-05-28 17:13:11 -04:00
|
|
|
|
~17% in "Other". The two-axis model cuts that to ~9% before any AI — and on the
|
|
|
|
|
|
real 347-row export the current implementation lands "Other" at **5.2%** (18
|
|
|
|
|
|
rows) with **18 mismatches** flagged.
|
2026-05-28 16:13:01 -04:00
|
|
|
|
|
|
|
|
|
|
- **Axis 1 — comment intent:** regex over the HTML-stripped `Last Comment`,
|
|
|
|
|
|
most-specific first (escalations → SIM ticket → vendor no-show → scheduling →
|
|
|
|
|
|
reports → completion → … → other).
|
|
|
|
|
|
- **Axis 2 — structured state:** `Hold Reason` → category and `WO Status`
|
|
|
|
|
|
signals (`RCAN`→Cancelled, `H` corroborates On Hold, `IP`/`R`/`RR` in-flight).
|
|
|
|
|
|
- **Resolution:** comment intent wins when confident → else structured state →
|
|
|
|
|
|
else `Other`. A **Claude Haiku** fallback (Secrets Manager) is reserved for
|
|
|
|
|
|
ambiguous free-text with no structured signal.
|
|
|
|
|
|
- **Mismatch detector (a feature):** flags when comment intent contradicts
|
|
|
|
|
|
structured state. Surfaced, never suppressed.
|
|
|
|
|
|
|
|
|
|
|
|
The authoritative spec lives in [`CLAUDE.md`](./CLAUDE.md); the implementation is
|
|
|
|
|
|
in `lambdas/classifier/classify.py` (owned by the `classifier-engineer` agent).
|
2026-05-28 17:13:11 -04:00
|
|
|
|
The S3-triggered `lambdas/classifier/handler.py` parses each export, classifies
|
|
|
|
|
|
every non-blank-comment row, writes a per-WO Parquet snapshot to
|
|
|
|
|
|
`analytics/dt=YYYY-MM-DD/` (registering the Glue partition via awswrangler), and
|
|
|
|
|
|
emits a `summary.json` for the Phase 4 slack-post Lambda.
|
2026-05-28 16:13:01 -04:00
|
|
|
|
|
|
|
|
|
|
## Repository layout
|
|
|
|
|
|
|
|
|
|
|
|
```
|
|
|
|
|
|
cdk/
|
|
|
|
|
|
app.py CDK entry point — instantiates both stacks
|
|
|
|
|
|
cdk.json
|
|
|
|
|
|
requirements.txt aws-cdk-lib==2.253.1, constructs>=10.6.0
|
|
|
|
|
|
stacks/
|
|
|
|
|
|
pipeline_stack.py S3, Lambdas, Glue, Athena, IAM
|
|
|
|
|
|
grafana_stack.py VPC import, EC2, ALB, SG, Route53, datasource role
|
|
|
|
|
|
lambdas/
|
|
|
|
|
|
classifier/ S3-triggered: parse → two-axis classify → Parquet
|
|
|
|
|
|
slack_post/ builds + posts the daily summary and alert
|
|
|
|
|
|
grafana/
|
|
|
|
|
|
provisioning/ Athena datasource + dashboard provider (as code)
|
|
|
|
|
|
dashboards/ committed dashboard JSON (source of truth)
|
|
|
|
|
|
scripts/ local drop-folder uploader + launchd plist
|
|
|
|
|
|
tests/ classifier smoke test
|
|
|
|
|
|
docs/BUILD.md phased, end-to-end build guide
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
## Configuration
|
|
|
|
|
|
|
|
|
|
|
|
| Where | What |
|
|
|
|
|
|
|---|---|
|
|
|
|
|
|
| **Secrets Manager** | `apm-wo-analysis/anthropic-api-key` (Haiku fallback). Slack bot token reused from the `payments-dashboard` app. |
|
|
|
|
|
|
| **SSM Parameter Store** | operational config (Slack channel ID, schedule expressions, deep-link base URL). |
|
|
|
|
|
|
| **GitHub repo secret** | `AWS_DEPLOY_ROLE_ARN` — the OIDC deploy role `githubdeploy-apm-wo-analysis`. |
|
|
|
|
|
|
|
|
|
|
|
|
No secrets in Lambda environment variables.
|
|
|
|
|
|
|
|
|
|
|
|
## Ingestion (no email)
|
|
|
|
|
|
|
|
|
|
|
|
The export reaches S3 by **direct upload or a local drop-folder**, never SES/email.
|
|
|
|
|
|
|
|
|
|
|
|
- **Direct:** `aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-328440206208/raw/`
|
Add drop-folder ingestion and scoped uploader IAM user
Complete Phase 1 ingestion. Add a least-privilege IAM user
(apm-wo-drop-uploader) to the pipeline stack, scoped to s3:PutObject
on the raw/ prefix only — the local launchd uploader authenticates as
this user via a dedicated profile, so a laptop credential leak cannot
read, list, or touch the analytics data.
Replace the scaffold uploader stub with the hardened stampli-pattern
script (lockfile, logging, timestamped archive, notifications, settle
delay) and align names to the convention (~/apm-wo-drop, ~/.local/bin,
com.seahaven.apm-wo-uploader). The plist sets PATH/HOME because launchd
runs with a stripped environment and otherwise cannot find aws.
The exports bucket already shipped in the Phase 0 scaffold, so the code
delta here is the uploader identity and tooling.
2026-05-28 16:32:56 -04:00
|
|
|
|
- **Drop-folder (optional zero-touch):** a launchd agent (`scripts/apm-wo-uploader.sh`
|
|
|
|
|
|
+ `scripts/com.seahaven.apm-wo-uploader.plist`) that watches `~/apm-wo-drop/`,
|
|
|
|
|
|
uploads new `.xlsx`/`.csv` files to `raw/`, and archives them to `uploaded/`.
|
|
|
|
|
|
It uploads with the scoped `apm-wo-drop` AWS profile (IAM user
|
|
|
|
|
|
`apm-wo-drop-uploader` — `s3:PutObject` on `raw/*` only).
|
|
|
|
|
|
|
|
|
|
|
|
Install (the runnable copy **must** live outside `~/Documents` — macOS TCC
|
|
|
|
|
|
sandbox; a repo-path script fails silently with `LastExitStatus=32256`):
|
|
|
|
|
|
```bash
|
|
|
|
|
|
install -d "$HOME/.local/bin" "$HOME/apm-wo-drop"
|
|
|
|
|
|
cp scripts/apm-wo-uploader.sh "$HOME/.local/bin/apm-wo-uploader.sh"
|
|
|
|
|
|
chmod +x "$HOME/.local/bin/apm-wo-uploader.sh"
|
|
|
|
|
|
cp scripts/com.seahaven.apm-wo-uploader.plist "$HOME/Library/LaunchAgents/"
|
|
|
|
|
|
launchctl load -w "$HOME/Library/LaunchAgents/com.seahaven.apm-wo-uploader.plist"
|
|
|
|
|
|
```
|
|
|
|
|
|
Re-copy the script to `~/.local/bin` after editing the repo source. Configure
|
|
|
|
|
|
the profile once with the uploader's access key:
|
|
|
|
|
|
`aws configure --profile apm-wo-drop`.
|
|
|
|
|
|
|
|
|
|
|
|
The classifier Lambda is S3-triggered on the `raw/` prefix regardless of path
|
2026-05-28 17:13:11 -04:00
|
|
|
|
(any `.xlsx`/`.csv` landing under `raw/` invokes it).
|
2026-05-28 16:13:01 -04:00
|
|
|
|
|
|
|
|
|
|
## Deployment
|
|
|
|
|
|
|
|
|
|
|
|
CI/CD via the org reusable workflows (no manual prod deploys):
|
|
|
|
|
|
|
|
|
|
|
|
- **CI** (`.github/workflows/ci.yaml`) → `ci-python-sam.yaml@main`: ruff + `cdk synth`.
|
|
|
|
|
|
- **Deploy** (`.github/workflows/deploy.yaml`) → `cd-cdk.yaml@main`: OIDC assume-role,
|
|
|
|
|
|
`cdk deploy --all`, single-flight concurrency.
|
|
|
|
|
|
|
|
|
|
|
|
The OIDC deploy role must exist **before** the first deploy. Deploy order:
|
|
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
|
cd cdk && pip install -r requirements.txt
|
|
|
|
|
|
cdk deploy apm-wo-analysis-pipeline # S3, Glue, Athena, Lambdas, IAM
|
|
|
|
|
|
cdk deploy apm-wo-analysis-grafana # EC2, ALB, SG, Route53, datasource role
|
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
|
|
## Local development
|
|
|
|
|
|
|
|
|
|
|
|
- pyenv Python 3.12; `ruff check` + `ruff format --check` before pushing (hook-enforced).
|
|
|
|
|
|
- Smoke-test the classifier against a **real export** before declaring any
|
|
|
|
|
|
classification change done: `~/Downloads/_documents/Sheet1-1.xlsx`.
|
|
|
|
|
|
- `cdk synth` must pass in CI before merge.
|
|
|
|
|
|
|
|
|
|
|
|
## Status
|
|
|
|
|
|
|
2026-05-28 17:13:11 -04:00
|
|
|
|
**Phase 2 — classifier (in review).** Build-out proceeds per
|
|
|
|
|
|
[`docs/BUILD.md`](./docs/BUILD.md): ingestion → classifier → Glue/Athena → Slack
|
|
|
|
|
|
→ Grafana → docs.
|
|
|
|
|
|
|
|
|
|
|
|
- **Phase 0** scaffold — merged-pending (PR #6).
|
|
|
|
|
|
- **Phase 1** ingestion (S3 bucket, drop-folder uploader, OIDC deploy role) —
|
|
|
|
|
|
deployed; PR #7 open.
|
|
|
|
|
|
- **Phase 2** classifier Lambda + Glue database + S3 `raw/` trigger — implemented
|
|
|
|
|
|
and `cdk synth`-green; PR #8 open (stacked on Phase 1, not yet deployed).
|
|
|
|
|
|
- **Phases 3–6** (Glue/Athena query layer, Slack, Grafana, final docs) — not
|
|
|
|
|
|
started.
|