Two push surfaces (no App Home) + interactive drill-down, per CLAUDE.md.
Block Kit (blockkit.py, pure/offline): build_daily_summary (header, vs-yesterday
deltas, escalation breakdown with 3rd highlighted, action/routine, top sites,
mismatch callout, category drill buttons + 📊 Open dashboard link, footer),
build_escalation_alert (one @here, returns None on zero-3rd — suppression), and
build_wo_modal (views.open payload, capped under Slack's 100-block limit).
Lambdas: slack_post/handler.py (classifier-invoked: read today/yesterday
summary.json, post daily summary, conditionally post the batched alert from
details.json) and slack_post/interactions.py (API Gateway: verify Slack
signature, filter details.json, views.open the WO modal within the 3s trigger_id
window). slackio.py centralizes Secrets Manager creds, the SSM dashboard URL,
signature verification, and analytics/ reads — keeping blockkit pure.
Classifier: emit analytics/dt=*/details.json (per-WO index for the modals) and
async-invoke slack-post after the snapshot write (best-effort; a Slack failure
never fails classification).
CDK: slack-post + interactions Lambdas (Docker-bundled slack_sdk), HTTP API on
apm-wo.seahaven.com (wildcard ACM cert + Route53 alias; signature-verified, so
the route is unauthenticated by design), SSM /apm-wo-analysis/grafana-base-url,
and scoped IAM (read analytics/, read the Slack secret + dashboard param;
classifier granted lambda:InvokeFunction on slack-post). Slack creds live in one
Secrets Manager secret apm-wo-analysis/slack-credentials {botToken, signingSecret,
channelId}; cdk.json gains cert/zone/domain context.
WO drill-downs link to Grafana only — no APM deep-links (per decision).
Deliverables for test time: slack/manifest.yaml (app manifest, interactivity
request_url = apm-wo.seahaven.com).
Tests: tests/test_blockkit.py (30 offline cases — deltas, zero-3rd None, <100
blocks under large inputs, modal truncation/overflow, dashboard URL) and Phase 4
assertions in test_pipeline_synth.py (both Lambdas, the API route/domain/alias,
and no broad/write IAM on the Slack roles). 49/49 tests pass; cdk synth green.
|
||
|---|---|---|
| .claude/agents | ||
| .github | ||
| cdk | ||
| docs | ||
| grafana | ||
| lambdas | ||
| scripts | ||
| slack | ||
| tests | ||
| .gitignore | ||
| CLAUDE.md | ||
| README.md | ||
apm-wo-analysis
Daily analysis of Amazon APM work-order "Last Comment" data for Sea Haven facility ops. A curated daily filter-view export (~350 work orders) is classified on two axes, pushed to Slack, and surfaced in a self-hosted Grafana dashboard. Replaces a legacy Google Apps Script + versioned-Google-Sheet workflow.
This is a sibling concern to the apm@ email pipeline in procurement-ingest
— it consumes a different feed (the manual export) and does not read those
tables. There is deliberately no DynamoDB: this is an analytics workload
backed by S3 + Athena (Grafana cannot query DynamoDB).
- Account / region: 328440206208 / us-east-1
- IaC: CDK (Python),
aws-cdk-lib==2.253.1. Lambdas Python 3.12, ARM64.
Architecture
APM export (xlsx/csv)
→ S3 raw/ (direct upload OR local launchd drop-folder)
→ classifier Lambda (HTML strip + two-axis classify, Haiku fallback)
→ S3 analytics/dt=YYYY-MM-DD/ (per-WO daily snapshot, Parquet)
→ Glue table → Athena → Grafana (self-hosted EC2, VPN-only, kiosk)
→ slack-post Lambda (reads today + yesterday partitions)
→ daily summary post [📊 Open dashboard button]
→ standalone batched 3rd-escalation alert (suppressed if zero)
Two CDK stacks:
| Stack | Resources |
|---|---|
apm-wo-analysis-pipeline |
S3 exports bucket, classifier + slack-post Lambdas, Glue database, Athena workgroup, IAM |
apm-wo-analysis-grafana |
EC2 (Grafana OSS), internal ALB, security group, Route53, Athena datasource IAM role |
The classification model
Always two-axis, never comment-only. The legacy script's central flaw was
reading only the comment while ignoring WO Status + Hold Reason, which left
~17% in "Other". The two-axis model cuts that to ~9% before any AI — and on the
real 347-row export the current implementation lands "Other" at 5.2% (18
rows) with 18 mismatches flagged.
- Axis 1 — comment intent: regex over the HTML-stripped
Last Comment, most-specific first (escalations → SIM ticket → vendor no-show → scheduling → reports → completion → … → other). - Axis 2 — structured state:
Hold Reason→ category andWO Statussignals (RCAN→Cancelled,Hcorroborates On Hold,IP/R/RRin-flight). - Resolution: comment intent wins when confident → else structured state →
else
Other. A Claude Haiku fallback (Secrets Manager) is reserved for ambiguous free-text with no structured signal. - Mismatch detector (a feature): flags when comment intent contradicts structured state. Surfaced, never suppressed.
The authoritative spec lives in CLAUDE.md; the implementation is
in lambdas/classifier/classify.py (owned by the classifier-engineer agent).
The S3-triggered lambdas/classifier/handler.py parses each export, classifies
every non-blank-comment row, writes a per-WO Parquet snapshot to
analytics/dt=YYYY-MM-DD/ (registering the Glue partition via awswrangler), and
emits a summary.json for the Phase 4 slack-post Lambda.
Repository layout
cdk/
app.py CDK entry point — instantiates both stacks
cdk.json
requirements.txt aws-cdk-lib==2.253.1, constructs>=10.6.0
stacks/
pipeline_stack.py S3, Lambdas, Glue, Athena, IAM
grafana_stack.py VPC import, EC2, ALB, SG, Route53, datasource role
lambdas/
classifier/ S3-triggered: parse → two-axis classify → Parquet
slack_post/ builds + posts the daily summary and alert
grafana/
provisioning/ Athena datasource + dashboard provider (as code)
dashboards/ committed dashboard JSON (source of truth)
scripts/ local drop-folder uploader + launchd plist
tests/ classifier smoke test
docs/BUILD.md phased, end-to-end build guide
Configuration
| Where | What |
|---|---|
| Secrets Manager | apm-wo-analysis/anthropic-api-key (Haiku fallback). Slack bot token reused from the payments-dashboard app. |
| SSM Parameter Store | operational config (Slack channel ID, schedule expressions, deep-link base URL). |
| GitHub repo secret | AWS_DEPLOY_ROLE_ARN — the OIDC deploy role githubdeploy-apm-wo-analysis. |
No secrets in Lambda environment variables.
Ingestion (no email)
The export reaches S3 by direct upload or a local drop-folder, never SES/email.
-
Direct:
aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-328440206208/raw/ -
Drop-folder (optional zero-touch): a launchd agent (
scripts/apm-wo-uploader.shscripts/com.seahaven.apm-wo-uploader.plist) that watches~/apm-wo-drop/, uploads new.xlsx/.csvfiles toraw/, and archives them touploaded/. It uploads with the scopedapm-wo-dropAWS profile (IAM userapm-wo-drop-uploader—s3:PutObjectonraw/*only).
Install (the runnable copy must live outside
~/Documents— macOS TCC sandbox; a repo-path script fails silently withLastExitStatus=32256):install -d "$HOME/.local/bin" "$HOME/apm-wo-drop" cp scripts/apm-wo-uploader.sh "$HOME/.local/bin/apm-wo-uploader.sh" chmod +x "$HOME/.local/bin/apm-wo-uploader.sh" cp scripts/com.seahaven.apm-wo-uploader.plist "$HOME/Library/LaunchAgents/" launchctl load -w "$HOME/Library/LaunchAgents/com.seahaven.apm-wo-uploader.plist"Re-copy the script to
~/.local/binafter editing the repo source. Configure the profile once with the uploader's access key:aws configure --profile apm-wo-drop.
The classifier Lambda is S3-triggered on the raw/ prefix regardless of path
(any .xlsx/.csv landing under raw/ invokes it).
Deployment
CI/CD via the org reusable workflows (no manual prod deploys):
- CI (
.github/workflows/ci.yaml) →ci-python-sam.yaml@main: ruff +cdk synth. - Deploy (
.github/workflows/deploy.yaml) →cd-cdk.yaml@main: OIDC assume-role,cdk deploy --all, single-flight concurrency.
The OIDC deploy role must exist before the first deploy. Deploy order:
cd cdk && pip install -r requirements.txt
cdk deploy apm-wo-analysis-pipeline # S3, Glue, Athena, Lambdas, IAM
cdk deploy apm-wo-analysis-grafana # EC2, ALB, SG, Route53, datasource role
Local development
- pyenv Python 3.12;
ruff check+ruff format --checkbefore pushing (hook-enforced). - Smoke-test the classifier against a real export before declaring any
classification change done:
~/Downloads/_documents/Sheet1-1.xlsx. cdk synthmust pass in CI before merge.
Status
Phase 3 — analytics dataset (in review). Build-out proceeds per
docs/BUILD.md: ingestion → classifier → Glue/Athena → Slack
→ Grafana → docs.
- Phase 0 scaffold — merged-pending (PR #6).
- Phase 1 ingestion (S3 bucket, drop-folder uploader, OIDC deploy role) — deployed; PR #7 open.
- Phase 2 classifier Lambda + Glue database + S3
raw/trigger — implemented andcdk synth-green; PR #8 open (stacked on Phase 1, not yet deployed). - Phase 3
apm_wo_snapshotsprojection table + Athena workgroup — implemented andcdk synth-green; PR #9 open (stacked on Phase 2). The classifier writes pure Parquet and holds no Glue access (projection handles partitions). - Phases 4–6 (Slack, Grafana, final docs) — not started.