No description
Find a file
Adam Moussa 6ee218eea3 Add analytics dataset: projection table + Athena workgroup (Phase 3)
CDK-define the apm_wo_snapshots Glue table with partition projection over
analytics/ (dt as projected date partition, 2026-01-01..NOW). Projection means
no crawler, no MSCK REPAIR, and — critically — the classifier needs no Glue
catalog access at all.

pipeline_stack.py: glue.CfnTable (Parquet SerDe, 17-column schema mirroring the
classifier's snapshot incl. contractor_description and the two boolean flags) +
athena.CfnWorkGroup `apm-wo-analysis` (enforced result location, SSE-S3) + a
30-day lifecycle rule on athena-results/ (disposable query output in a RETAIN
bucket). Trim the classifier role: drop the entire Glue policy statement.

handler.py: stop registering the table at runtime — drop database=/table= from
to_parquet so the classifier writes pure Parquet; partition projection handles
the rest. Keeps overwrite_partitions for idempotent same-day re-uploads.

test_pipeline_synth.py: offline synth assertions (bundling skipped) — projection
properties, column schema/types, workgroup result enforcement, and that no IAM
policy grants glue:* to the classifier.

cdk synth green; 11/11 tests pass.
2026-05-28 17:24:36 -04:00
.claude/agents Scaffold apm-wo-analysis repository 2026-05-28 16:13:13 -04:00
.github Scaffold apm-wo-analysis repository 2026-05-28 16:13:13 -04:00
cdk Add analytics dataset: projection table + Athena workgroup (Phase 3) 2026-05-28 17:24:36 -04:00
docs Scaffold apm-wo-analysis repository 2026-05-28 16:13:13 -04:00
grafana Scaffold apm-wo-analysis repository 2026-05-28 16:13:13 -04:00
lambdas Add analytics dataset: projection table + Athena workgroup (Phase 3) 2026-05-28 17:24:36 -04:00
scripts Add drop-folder ingestion and scoped uploader IAM user 2026-05-28 16:32:56 -04:00
tests Add analytics dataset: projection table + Athena workgroup (Phase 3) 2026-05-28 17:24:36 -04:00
.gitignore Scaffold apm-wo-analysis repository 2026-05-28 16:13:13 -04:00
CLAUDE.md Scaffold apm-wo-analysis repository 2026-05-28 16:13:13 -04:00
README.md Update README for Phase 2 classifier 2026-05-28 17:13:11 -04:00

apm-wo-analysis

Daily analysis of Amazon APM work-order "Last Comment" data for Sea Haven facility ops. A curated daily filter-view export (~350 work orders) is classified on two axes, pushed to Slack, and surfaced in a self-hosted Grafana dashboard. Replaces a legacy Google Apps Script + versioned-Google-Sheet workflow.

This is a sibling concern to the apm@ email pipeline in procurement-ingest — it consumes a different feed (the manual export) and does not read those tables. There is deliberately no DynamoDB: this is an analytics workload backed by S3 + Athena (Grafana cannot query DynamoDB).

  • Account / region: 328440206208 / us-east-1
  • IaC: CDK (Python), aws-cdk-lib==2.253.1. Lambdas Python 3.12, ARM64.

Architecture

APM export (xlsx/csv)
  → S3 raw/  (direct upload OR local launchd drop-folder)
      → classifier Lambda (HTML strip + two-axis classify, Haiku fallback)
          → S3 analytics/dt=YYYY-MM-DD/  (per-WO daily snapshot, Parquet)
                → Glue table → Athena → Grafana (self-hosted EC2, VPN-only, kiosk)
          → slack-post Lambda (reads today + yesterday partitions)
                → daily summary post  [📊 Open dashboard button]
                → standalone batched 3rd-escalation alert (suppressed if zero)

Two CDK stacks:

Stack Resources
apm-wo-analysis-pipeline S3 exports bucket, classifier + slack-post Lambdas, Glue database, Athena workgroup, IAM
apm-wo-analysis-grafana EC2 (Grafana OSS), internal ALB, security group, Route53, Athena datasource IAM role

The classification model

Always two-axis, never comment-only. The legacy script's central flaw was reading only the comment while ignoring WO Status + Hold Reason, which left ~17% in "Other". The two-axis model cuts that to ~9% before any AI — and on the real 347-row export the current implementation lands "Other" at 5.2% (18 rows) with 18 mismatches flagged.

  • Axis 1 — comment intent: regex over the HTML-stripped Last Comment, most-specific first (escalations → SIM ticket → vendor no-show → scheduling → reports → completion → … → other).
  • Axis 2 — structured state: Hold Reason → category and WO Status signals (RCAN→Cancelled, H corroborates On Hold, IP/R/RR in-flight).
  • Resolution: comment intent wins when confident → else structured state → else Other. A Claude Haiku fallback (Secrets Manager) is reserved for ambiguous free-text with no structured signal.
  • Mismatch detector (a feature): flags when comment intent contradicts structured state. Surfaced, never suppressed.

The authoritative spec lives in CLAUDE.md; the implementation is in lambdas/classifier/classify.py (owned by the classifier-engineer agent). The S3-triggered lambdas/classifier/handler.py parses each export, classifies every non-blank-comment row, writes a per-WO Parquet snapshot to analytics/dt=YYYY-MM-DD/ (registering the Glue partition via awswrangler), and emits a summary.json for the Phase 4 slack-post Lambda.

Repository layout

cdk/
  app.py                 CDK entry point — instantiates both stacks
  cdk.json
  requirements.txt       aws-cdk-lib==2.253.1, constructs>=10.6.0
  stacks/
    pipeline_stack.py    S3, Lambdas, Glue, Athena, IAM
    grafana_stack.py     VPC import, EC2, ALB, SG, Route53, datasource role
lambdas/
  classifier/            S3-triggered: parse → two-axis classify → Parquet
  slack_post/            builds + posts the daily summary and alert
grafana/
  provisioning/          Athena datasource + dashboard provider (as code)
  dashboards/            committed dashboard JSON (source of truth)
scripts/                 local drop-folder uploader + launchd plist
tests/                   classifier smoke test
docs/BUILD.md            phased, end-to-end build guide

Configuration

Where What
Secrets Manager apm-wo-analysis/anthropic-api-key (Haiku fallback). Slack bot token reused from the payments-dashboard app.
SSM Parameter Store operational config (Slack channel ID, schedule expressions, deep-link base URL).
GitHub repo secret AWS_DEPLOY_ROLE_ARN — the OIDC deploy role githubdeploy-apm-wo-analysis.

No secrets in Lambda environment variables.

Ingestion (no email)

The export reaches S3 by direct upload or a local drop-folder, never SES/email.

  • Direct: aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-328440206208/raw/

  • Drop-folder (optional zero-touch): a launchd agent (scripts/apm-wo-uploader.sh

    • scripts/com.seahaven.apm-wo-uploader.plist) that watches ~/apm-wo-drop/, uploads new .xlsx/.csv files to raw/, and archives them to uploaded/. It uploads with the scoped apm-wo-drop AWS profile (IAM user apm-wo-drop-uploader — s3:PutObject on raw/* only).

    Install (the runnable copy must live outside ~/Documents — macOS TCC sandbox; a repo-path script fails silently with LastExitStatus=32256):

    install -d "$HOME/.local/bin" "$HOME/apm-wo-drop"
    cp scripts/apm-wo-uploader.sh "$HOME/.local/bin/apm-wo-uploader.sh"
    chmod +x "$HOME/.local/bin/apm-wo-uploader.sh"
    cp scripts/com.seahaven.apm-wo-uploader.plist "$HOME/Library/LaunchAgents/"
    launchctl load -w "$HOME/Library/LaunchAgents/com.seahaven.apm-wo-uploader.plist"
    

    Re-copy the script to ~/.local/bin after editing the repo source. Configure the profile once with the uploader's access key: aws configure --profile apm-wo-drop.

The classifier Lambda is S3-triggered on the raw/ prefix regardless of path (any .xlsx/.csv landing under raw/ invokes it).

Deployment

CI/CD via the org reusable workflows (no manual prod deploys):

  • CI (.github/workflows/ci.yaml) → ci-python-sam.yaml@main: ruff + cdk synth.
  • Deploy (.github/workflows/deploy.yaml) → cd-cdk.yaml@main: OIDC assume-role, cdk deploy --all, single-flight concurrency.

The OIDC deploy role must exist before the first deploy. Deploy order:

cd cdk && pip install -r requirements.txt
cdk deploy apm-wo-analysis-pipeline   # S3, Glue, Athena, Lambdas, IAM
cdk deploy apm-wo-analysis-grafana    # EC2, ALB, SG, Route53, datasource role

Local development

  • pyenv Python 3.12; ruff check + ruff format --check before pushing (hook-enforced).
  • Smoke-test the classifier against a real export before declaring any classification change done: ~/Downloads/_documents/Sheet1-1.xlsx.
  • cdk synth must pass in CI before merge.

Status

Phase 2 — classifier (in review). Build-out proceeds per docs/BUILD.md: ingestion → classifier → Glue/Athena → Slack → Grafana → docs.

  • Phase 0 scaffold — merged-pending (PR #6).
  • Phase 1 ingestion (S3 bucket, drop-folder uploader, OIDC deploy role) — deployed; PR #7 open.
  • Phase 2 classifier Lambda + Glue database + S3 raw/ trigger — implemented and cdk synth-green; PR #8 open (stacked on Phase 1, not yet deployed).
  • Phases 3–6 (Glue/Athena query layer, Slack, Grafana, final docs) — not started.