8 KiB
apm-wo-analysis
Daily analysis of Amazon APM work-order "Last Comment" data for Sea Haven facility ops. A curated daily filter-view export (~350 work orders) is classified on two axes, pushed to Slack, and surfaced in a self-hosted Grafana dashboard. Replaces a legacy Google Apps Script + versioned-Google-Sheet workflow.
This is a sibling concern to the apm@ email pipeline in procurement-ingest
— it consumes a different feed (the manual export) and does not read those
tables. There is deliberately no DynamoDB: this is an analytics workload
backed by S3 + Athena (Grafana cannot query DynamoDB).
- Account / region: 328440206208 / us-east-1
- IaC: CDK (Python),
aws-cdk-lib==2.253.1. Lambdas Python 3.12, ARM64.
Architecture
APM export (xlsx/csv)
→ S3 raw/ (direct upload OR local launchd drop-folder)
→ classifier Lambda (HTML strip + two-axis classify, Haiku fallback)
→ S3 analytics/dt=YYYY-MM-DD/ (per-WO daily snapshot, Parquet)
→ Glue table → Athena → Grafana (self-hosted EC2, VPN-only, kiosk)
→ slack-post Lambda (reads today + yesterday partitions)
→ daily summary post [📊 Open dashboard button]
→ standalone batched 3rd-escalation alert (suppressed if zero)
Two CDK stacks:
| Stack | Resources |
|---|---|
apm-wo-analysis-pipeline |
S3 exports bucket, classifier + slack-post Lambdas, Glue database, Athena workgroup, IAM |
apm-wo-analysis-grafana |
EC2 (Grafana OSS), internal ALB, security group, Route53, Athena datasource IAM role |
The classification model
Always two-axis, never comment-only. The legacy script's central flaw was
reading only the comment while ignoring WO Status + Hold Reason, which left
~17% in "Other". The two-axis model cuts that to ~9% before any AI — and on the
real 347-row export the current implementation lands "Other" at 5.2% (18
rows) with 18 mismatches flagged.
- Axis 1 — comment intent: regex over the HTML-stripped
Last Comment, most-specific first (escalations → SIM ticket → vendor no-show → scheduling → reports → completion → … → other). - Axis 2 — structured state:
Hold Reason→ category andWO Statussignals (RCAN→Cancelled,Hcorroborates On Hold,IP/R/RRin-flight). - Resolution: comment intent wins when confident → else structured state →
else
Other. A Claude Haiku fallback (Secrets Manager) is reserved for ambiguous free-text with no structured signal. - Mismatch detector (a feature): flags when comment intent contradicts structured state. Surfaced, never suppressed.
The authoritative spec lives in CLAUDE.md; the implementation is
in lambdas/classifier/classify.py (owned by the classifier-engineer agent).
The S3-triggered lambdas/classifier/handler.py parses each export, classifies
every non-blank-comment row, writes a per-WO Parquet snapshot to
analytics/dt=YYYY-MM-DD/ (registering the Glue partition via awswrangler), and
emits a summary.json for the Phase 4 slack-post Lambda.
Repository layout
cdk/
app.py CDK entry point — instantiates both stacks
cdk.json
requirements.txt aws-cdk-lib==2.253.1, constructs>=10.6.0
stacks/
pipeline_stack.py S3, Lambdas, Glue, Athena, IAM
grafana_stack.py VPC import, EC2, ALB, SG, Route53, datasource role
lambdas/
classifier/ S3-triggered: parse → two-axis classify → Parquet
slack_post/ builds + posts the daily summary and alert
grafana/
provisioning/ Athena datasource + dashboard provider (as code)
dashboards/ committed dashboard JSON (source of truth)
scripts/ local drop-folder uploader + launchd plist
tests/ classifier smoke test
docs/BUILD.md phased, end-to-end build guide
Configuration
| Where | What |
|---|---|
| Secrets Manager | apm-wo-analysis/anthropic-api-key (Haiku fallback); apm-wo-analysis/slack-credentials = { botToken, signingSecret, channelId } (reused Slack app). |
| SSM Parameter Store | /apm-wo-analysis/grafana-base-url (the 📊 dashboard button / modal overflow link; ops-editable). |
| GitHub repo secret | AWS_DEPLOY_ROLE_ARN — the OIDC deploy role githubdeploy-apm-wo-analysis. |
No secrets in Lambda environment variables.
Ingestion (no email)
The export reaches S3 by direct upload or a local drop-folder, never SES/email.
-
Direct:
aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-328440206208/raw/ -
Drop-folder (optional zero-touch): a launchd agent (
scripts/apm-wo-uploader.shscripts/com.seahaven.apm-wo-uploader.plist) that watches~/apm-wo-drop/, uploads new.xlsx/.csvfiles toraw/, and archives them touploaded/. It uploads with the scopedapm-wo-dropAWS profile (IAM userapm-wo-drop-uploader—s3:PutObjectonraw/*only).
Install (the runnable copy must live outside
~/Documents— macOS TCC sandbox; a repo-path script fails silently withLastExitStatus=32256):install -d "$HOME/.local/bin" "$HOME/apm-wo-drop" cp scripts/apm-wo-uploader.sh "$HOME/.local/bin/apm-wo-uploader.sh" chmod +x "$HOME/.local/bin/apm-wo-uploader.sh" cp scripts/com.seahaven.apm-wo-uploader.plist "$HOME/Library/LaunchAgents/" launchctl load -w "$HOME/Library/LaunchAgents/com.seahaven.apm-wo-uploader.plist"Re-copy the script to
~/.local/binafter editing the repo source. Configure the profile once with the uploader's access key:aws configure --profile apm-wo-drop.
The classifier Lambda is S3-triggered on the raw/ prefix regardless of path
(any .xlsx/.csv landing under raw/ invokes it).
Deployment
CI/CD via the org reusable workflows (no manual prod deploys):
- CI (
.github/workflows/ci.yaml) →ci-python-sam.yaml@main: ruff +cdk synth. - Deploy (
.github/workflows/deploy.yaml) →cd-cdk.yaml@main: OIDC assume-role,cdk deploy --all, single-flight concurrency.
The OIDC deploy role must exist before the first deploy. Deploy order:
cd cdk && pip install -r requirements.txt
cdk deploy apm-wo-analysis-pipeline # S3, Glue, Athena, Lambdas, IAM
cdk deploy apm-wo-analysis-grafana # EC2, ALB, SG, Route53, datasource role
Local development
- pyenv Python 3.12;
ruff check+ruff format --checkbefore pushing (hook-enforced). - Smoke-test the classifier against a real export before declaring any
classification change done:
~/Downloads/_documents/Sheet1-1.xlsx. cdk synthmust pass in CI before merge.
Status
Phase 5 — Grafana (in review); build complete, docs remain. Build-out per
docs/BUILD.md: ingestion → classifier → Glue/Athena → Slack
→ Grafana → docs.
- Phase 0 scaffold — merged-pending (PR #6).
- Phase 1 ingestion (S3 bucket, drop-folder uploader, OIDC deploy role) — deployed; PR #7 open.
- Phase 2 classifier Lambda + Glue database + S3
raw/trigger — implemented andcdk synth-green; PR #8 open (stacked on Phase 1, not yet deployed). - Phase 3
apm_wo_snapshotsprojection table + Athena workgroup — implemented andcdk synth-green; PR #9 open (stacked on Phase 2). The classifier writes pure Parquet and holds no Glue access (projection handles partitions). - Phase 4 Slack post + interactions Lambdas (daily summary, batched
3rd-escalation alert, drill-down modals on
apm-wo.seahaven.com) — implemented andcdk synth-green; PR #10 open (stacked on Phase 3). App manifest inslack/manifest.yaml. - Phase 5 self-hosted Grafana (EC2 + ALB on
grafana.seahaven.com, office-IP-restricted, 7-panel dashboard as code) — implemented andcdk synth-green; PR #11 open (stacked on Phase 4). - Phase 6 (final docs / Confluence / runbook) — not started.
All phase PRs are stacked (#6→#7→#8→#9→#10→#11) and unmerged; nothing is
deployed yet. Cross-review and /security-review are outstanding across the stack.