apm-wo-analysis/README.md
Adam Moussa 7befc8f0d7 Add drop-folder ingestion and scoped uploader IAM user
Complete Phase 1 ingestion. Add a least-privilege IAM user
(apm-wo-drop-uploader) to the pipeline stack, scoped to s3:PutObject
on the raw/ prefix only — the local launchd uploader authenticates as
this user via a dedicated profile, so a laptop credential leak cannot
read, list, or touch the analytics data.

Replace the scaffold uploader stub with the hardened stampli-pattern
script (lockfile, logging, timestamped archive, notifications, settle
delay) and align names to the convention (~/apm-wo-drop, ~/.local/bin,
com.seahaven.apm-wo-uploader). The plist sets PATH/HOME because launchd
runs with a stripped environment and otherwise cannot find aws.

The exports bucket already shipped in the Phase 0 scaffold, so the code
delta here is the uploader identity and tooling.
2026-05-28 16:32:56 -04:00

6.3 KiB

apm-wo-analysis

Daily analysis of Amazon APM work-order "Last Comment" data for Sea Haven facility ops. A curated daily filter-view export (~350 work orders) is classified on two axes, pushed to Slack, and surfaced in a self-hosted Grafana dashboard. Replaces a legacy Google Apps Script + versioned-Google-Sheet workflow.

This is a sibling concern to the apm@ email pipeline in procurement-ingest — it consumes a different feed (the manual export) and does not read those tables. There is deliberately no DynamoDB: this is an analytics workload backed by S3 + Athena (Grafana cannot query DynamoDB).

  • Account / region: 328440206208 / us-east-1
  • IaC: CDK (Python), aws-cdk-lib==2.253.1. Lambdas Python 3.12, ARM64.

Architecture

APM export (xlsx/csv)
  → S3 raw/  (direct upload OR local launchd drop-folder)
      → classifier Lambda (HTML strip + two-axis classify, Haiku fallback)
          → S3 analytics/dt=YYYY-MM-DD/  (per-WO daily snapshot, Parquet)
                → Glue table → Athena → Grafana (self-hosted EC2, VPN-only, kiosk)
          → slack-post Lambda (reads today + yesterday partitions)
                → daily summary post  [📊 Open dashboard button]
                → standalone batched 3rd-escalation alert (suppressed if zero)

Two CDK stacks:

Stack Resources
apm-wo-analysis-pipeline S3 exports bucket, classifier + slack-post Lambdas, Glue database, Athena workgroup, IAM
apm-wo-analysis-grafana EC2 (Grafana OSS), internal ALB, security group, Route53, Athena datasource IAM role

The classification model

Always two-axis, never comment-only. The legacy script's central flaw was reading only the comment while ignoring WO Status + Hold Reason, which left ~17% in "Other". The two-axis model cuts that to ~9% before any AI.

  • Axis 1 — comment intent: regex over the HTML-stripped Last Comment, most-specific first (escalations → SIM ticket → vendor no-show → scheduling → reports → completion → … → other).
  • Axis 2 — structured state: Hold Reason → category and WO Status signals (RCAN→Cancelled, H corroborates On Hold, IP/R/RR in-flight).
  • Resolution: comment intent wins when confident → else structured state → else Other. A Claude Haiku fallback (Secrets Manager) is reserved for ambiguous free-text with no structured signal.
  • Mismatch detector (a feature): flags when comment intent contradicts structured state. Surfaced, never suppressed.

The authoritative spec lives in CLAUDE.md; the implementation is in lambdas/classifier/classify.py (owned by the classifier-engineer agent).

Repository layout

cdk/
  app.py                 CDK entry point — instantiates both stacks
  cdk.json
  requirements.txt       aws-cdk-lib==2.253.1, constructs>=10.6.0
  stacks/
    pipeline_stack.py    S3, Lambdas, Glue, Athena, IAM
    grafana_stack.py     VPC import, EC2, ALB, SG, Route53, datasource role
lambdas/
  classifier/            S3-triggered: parse → two-axis classify → Parquet
  slack_post/            builds + posts the daily summary and alert
grafana/
  provisioning/          Athena datasource + dashboard provider (as code)
  dashboards/            committed dashboard JSON (source of truth)
scripts/                 local drop-folder uploader + launchd plist
tests/                   classifier smoke test
docs/BUILD.md            phased, end-to-end build guide

Configuration

Where What
Secrets Manager apm-wo-analysis/anthropic-api-key (Haiku fallback). Slack bot token reused from the payments-dashboard app.
SSM Parameter Store operational config (Slack channel ID, schedule expressions, deep-link base URL).
GitHub repo secret AWS_DEPLOY_ROLE_ARN — the OIDC deploy role githubdeploy-apm-wo-analysis.

No secrets in Lambda environment variables.

Ingestion (no email)

The export reaches S3 by direct upload or a local drop-folder, never SES/email.

  • Direct: aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-328440206208/raw/

  • Drop-folder (optional zero-touch): a launchd agent (scripts/apm-wo-uploader.sh

    • scripts/com.seahaven.apm-wo-uploader.plist) that watches ~/apm-wo-drop/, uploads new .xlsx/.csv files to raw/, and archives them to uploaded/. It uploads with the scoped apm-wo-drop AWS profile (IAM user apm-wo-drop-uploader — s3:PutObject on raw/* only).

    Install (the runnable copy must live outside ~/Documents — macOS TCC sandbox; a repo-path script fails silently with LastExitStatus=32256):

    install -d "$HOME/.local/bin" "$HOME/apm-wo-drop"
    cp scripts/apm-wo-uploader.sh "$HOME/.local/bin/apm-wo-uploader.sh"
    chmod +x "$HOME/.local/bin/apm-wo-uploader.sh"
    cp scripts/com.seahaven.apm-wo-uploader.plist "$HOME/Library/LaunchAgents/"
    launchctl load -w "$HOME/Library/LaunchAgents/com.seahaven.apm-wo-uploader.plist"
    

    Re-copy the script to ~/.local/bin after editing the repo source. Configure the profile once with the uploader's access key: aws configure --profile apm-wo-drop.

The classifier Lambda is S3-triggered on the raw/ prefix regardless of path (added in Phase 2 — uploads currently land in raw/ and wait).

Deployment

CI/CD via the org reusable workflows (no manual prod deploys):

  • CI (.github/workflows/ci.yaml) → ci-python-sam.yaml@main: ruff + cdk synth.
  • Deploy (.github/workflows/deploy.yaml) → cd-cdk.yaml@main: OIDC assume-role, cdk deploy --all, single-flight concurrency.

The OIDC deploy role must exist before the first deploy. Deploy order:

cd cdk && pip install -r requirements.txt
cdk deploy apm-wo-analysis-pipeline   # S3, Glue, Athena, Lambdas, IAM
cdk deploy apm-wo-analysis-grafana    # EC2, ALB, SG, Route53, datasource role

Local development

  • pyenv Python 3.12; ruff check + ruff format --check before pushing (hook-enforced).
  • Smoke-test the classifier against a real export before declaring any classification change done: ~/Downloads/_documents/Sheet1-1.xlsx.
  • cdk synth must pass in CI before merge.

Status

Scaffold (Phase 0). Build-out proceeds per docs/BUILD.md: ingestion → classifier → Glue/Athena → Slack → Grafana → docs.