No description
Find a file
Adam Moussa f86b4ba1c5
Merge pull request #6 from Sea-Haven-Industries/feature/phase-0-scaffold
Phase 0: scaffold apm-wo-analysis repository
2026-05-29 13:14:37 -04:00
.claude/agents Scaffold apm-wo-analysis repository 2026-05-28 16:13:13 -04:00
.github Temporarily disable CD deploy during stack merges 2026-05-29 11:43:48 -04:00
cdk Scaffold apm-wo-analysis repository 2026-05-28 16:13:13 -04:00
docs Scaffold apm-wo-analysis repository 2026-05-28 16:13:13 -04:00
grafana Scaffold apm-wo-analysis repository 2026-05-28 16:13:13 -04:00
lambdas Scaffold apm-wo-analysis repository 2026-05-28 16:13:13 -04:00
scripts Scaffold apm-wo-analysis repository 2026-05-28 16:13:13 -04:00
tests Scaffold apm-wo-analysis repository 2026-05-28 16:13:13 -04:00
.gitignore Scaffold apm-wo-analysis repository 2026-05-28 16:13:13 -04:00
CLAUDE.md Scaffold apm-wo-analysis repository 2026-05-28 16:13:13 -04:00
README.md Add project README 2026-05-28 16:13:01 -04:00

apm-wo-analysis

Daily analysis of Amazon APM work-order "Last Comment" data for Sea Haven facility ops. A curated daily filter-view export (~350 work orders) is classified on two axes, pushed to Slack, and surfaced in a self-hosted Grafana dashboard. Replaces a legacy Google Apps Script + versioned-Google-Sheet workflow.

This is a sibling concern to the apm@ email pipeline in procurement-ingest — it consumes a different feed (the manual export) and does not read those tables. There is deliberately no DynamoDB: this is an analytics workload backed by S3 + Athena (Grafana cannot query DynamoDB).

  • Account / region: 328440206208 / us-east-1
  • IaC: CDK (Python), aws-cdk-lib==2.253.1. Lambdas Python 3.12, ARM64.

Architecture

APM export (xlsx/csv)
  → S3 raw/  (direct upload OR local launchd drop-folder)
      → classifier Lambda (HTML strip + two-axis classify, Haiku fallback)
          → S3 analytics/dt=YYYY-MM-DD/  (per-WO daily snapshot, Parquet)
                → Glue table → Athena → Grafana (self-hosted EC2, VPN-only, kiosk)
          → slack-post Lambda (reads today + yesterday partitions)
                → daily summary post  [📊 Open dashboard button]
                → standalone batched 3rd-escalation alert (suppressed if zero)

Two CDK stacks:

Stack Resources
apm-wo-analysis-pipeline S3 exports bucket, classifier + slack-post Lambdas, Glue database, Athena workgroup, IAM
apm-wo-analysis-grafana EC2 (Grafana OSS), internal ALB, security group, Route53, Athena datasource IAM role

The classification model

Always two-axis, never comment-only. The legacy script's central flaw was reading only the comment while ignoring WO Status + Hold Reason, which left ~17% in "Other". The two-axis model cuts that to ~9% before any AI.

  • Axis 1 — comment intent: regex over the HTML-stripped Last Comment, most-specific first (escalations → SIM ticket → vendor no-show → scheduling → reports → completion → … → other).
  • Axis 2 — structured state: Hold Reason → category and WO Status signals (RCAN→Cancelled, H corroborates On Hold, IP/R/RR in-flight).
  • Resolution: comment intent wins when confident → else structured state → else Other. A Claude Haiku fallback (Secrets Manager) is reserved for ambiguous free-text with no structured signal.
  • Mismatch detector (a feature): flags when comment intent contradicts structured state. Surfaced, never suppressed.

The authoritative spec lives in CLAUDE.md; the implementation is in lambdas/classifier/classify.py (owned by the classifier-engineer agent).

Repository layout

cdk/
  app.py                 CDK entry point — instantiates both stacks
  cdk.json
  requirements.txt       aws-cdk-lib==2.253.1, constructs>=10.6.0
  stacks/
    pipeline_stack.py    S3, Lambdas, Glue, Athena, IAM
    grafana_stack.py     VPC import, EC2, ALB, SG, Route53, datasource role
lambdas/
  classifier/            S3-triggered: parse → two-axis classify → Parquet
  slack_post/            builds + posts the daily summary and alert
grafana/
  provisioning/          Athena datasource + dashboard provider (as code)
  dashboards/            committed dashboard JSON (source of truth)
scripts/                 local drop-folder uploader + launchd plist
tests/                   classifier smoke test
docs/BUILD.md            phased, end-to-end build guide

Configuration

Where What
Secrets Manager apm-wo-analysis/anthropic-api-key (Haiku fallback). Slack bot token reused from the payments-dashboard app.
SSM Parameter Store operational config (Slack channel ID, schedule expressions, deep-link base URL).
GitHub repo secret AWS_DEPLOY_ROLE_ARN — the OIDC deploy role githubdeploy-apm-wo-analysis.

No secrets in Lambda environment variables.

Ingestion (no email)

The export reaches S3 by direct upload or a local drop-folder, never SES/email.

  • Direct: aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-328440206208/raw/
  • Drop-folder (optional): the launchd agent in scripts/. Both the script and the watched folder must live outside ~/Documents (macOS TCC sandbox).

The classifier Lambda is S3-triggered on the raw/ prefix regardless of path.

Deployment

CI/CD via the org reusable workflows (no manual prod deploys):

  • CI (.github/workflows/ci.yaml) → ci-python-sam.yaml@main: ruff + cdk synth.
  • Deploy (.github/workflows/deploy.yaml) → cd-cdk.yaml@main: OIDC assume-role, cdk deploy --all, single-flight concurrency.

The OIDC deploy role must exist before the first deploy. Deploy order:

cd cdk && pip install -r requirements.txt
cdk deploy apm-wo-analysis-pipeline   # S3, Glue, Athena, Lambdas, IAM
cdk deploy apm-wo-analysis-grafana    # EC2, ALB, SG, Route53, datasource role

Local development

  • pyenv Python 3.12; ruff check + ruff format --check before pushing (hook-enforced).
  • Smoke-test the classifier against a real export before declaring any classification change done: ~/Downloads/_documents/Sheet1-1.xlsx.
  • cdk synth must pass in CI before merge.

Status

Scaffold (Phase 0). Build-out proceeds per docs/BUILD.md: ingestion → classifier → Glue/Athena → Slack → Grafana → docs.