apm-wo-analysis/CLAUDE.md
Adam Moussa b92c25faf2
Some checks failed
Deploy / deploy (push) Has been cancelled
docs: CLAUDE.md cdk pin follows handbook Pinning Principle, not a hardcoded version (#22)
2026-06-05 18:44:12 -04:00

7.6 KiB

Project Instructions — apm-wo-analysis

Project-specific context and rules. Supplements the global ~/.claude/CLAUDE.md (Sea Haven standards) and the engineering handbook. Where this file is silent, the global rules and handbook apply.


What this is

Daily analysis of Amazon APM work-order "Last Comment" data for Sea Haven facility ops. Replaces a legacy Google Apps Script + versioned-Google-Sheet workflow. A curated daily filter-view export (~350 WOs) is classified on two axes, then pushed to Slack and surfaced in a Grafana dashboard.

This is a separate, sibling concern to the apm@ email pipeline in procurement-ingest. That pipeline event-sources notification emails into a WorkOrders table; this repo consumes a different feed (the manual export) and does not read those tables. See the apm-wo-comment-analysis and procurement-ingest memories.


Architecture

APM export (xlsx/csv)
  → S3 raw/  (direct upload OR local launchd drop-folder)
      → classifier Lambda (Python, ARM64: HTML strip + two-axis classify, Haiku fallback)
          → S3 analytics/dt=YYYY-MM-DD/  (per-WO daily snapshot, Parquet)
                → Glue table → Athena → Grafana (self-hosted EC2, VPN-only, kiosk)
          → slack-post Lambda (reads today + yesterday partitions)
                → daily summary post  [📊 Open dashboard button]
                → standalone batched 3rd-escalation alert (suppressed if zero)
  • Account / region: 328440206208 / us-east-1
  • IaC: CDK (Python). aws-cdk-lib pinned exact (==) in cdk/requirements.txt, kept current by Dependabot per the handbook Pinning Principle (no blanket ignores) — don't hardcode a version number here, read the requirements file. Lambdas Python 3.12, ARM64.
  • No DynamoDB — deliberate. This is an analytics workload and Grafana cannot query DynamoDB; S3 + Athena is the store. Do not "helpfully" add a table.

Ingestion (no email)

The export reaches S3 by direct upload or a local drop-folder, never SES/email.

  • Direct: console/CLI aws s3 cp into raw/.
  • Drop-folder: a launchd agent that uploads files dropped in a local folder, mirroring the stampli-drop-folder pattern. The launchd script must live outside ~/Documents (macOS TCC sandbox — see the macos-tcc-launchd memory). The watched drop folder must also sit outside ~/Documents.
  • The classifier Lambda is S3-triggered on the raw/ prefix regardless of how the file arrives.

The classification model (the core domain knowledge)

Always two-axis. Never comment-only. The legacy script's central flaw was reading only the comment and ignoring WO Status + Hold Reason; that left ~17% in "Other" and mis-stated dozens of WOs. The two-axis model cut "Other" to ~9% before any AI.

Axis 1 — comment intent (regex over the HTML-stripped Last Comment), most-specific first: 3rd / 2nd / 1st Escalation · SIM Ticket · Vendor No-Show · Weekly WO Scheduled · Schedule Confirmed · Awaiting Scheduling · Report / Docs Needed · Awaiting Report / Invoice · Completed / Pending Close · Acknowledgement / No-op · Cancelled · On Hold · Rescheduled · Avetta Project Created · Other Escalation · Status Inquiry.

Axis 2 — structured state:

  • Hold Reason → category: SCHEDULING→Awaiting Scheduling, REPORT→Report / Docs Needed, VENDOR/PARTS/ORDER→Awaiting Vendor / Parts, VERIFY→Verification Needed, RESOURCE→Resource Hold, NOEQUIP→No Equipment.
  • WO Status: RCAN→Cancelled; H corroborates On Hold; IP/R/RR are in-flight states.

Resolution: comment intent wins when confident → else structured state → else Other. Reserve a Claude Haiku fallback (Secrets Manager key) strictly for ambiguous free-text with no structured signal ("Copy", "Cant close", "Uplift request submitted").

Mismatch detector (a feature): flag when comment intent contradicts structured state — e.g. comment "completed/scheduled" while on a REPORT/SCHEDULING/VENDOR hold, or "completed" while WO Status is IP. Surface these; do not suppress.

Escalation categories: 1st / 2nd / 3rd Escalation, SIM Ticket, Other Escalation. Action-needed categories: escalations + Awaiting Scheduling + Report / Docs Needed + Awaiting Report / Invoice + Awaiting Vendor / Parts + Status Inquiry + Vendor No-Show. Everything else is routine.


Export schema (13 columns)

WO Number, WO Description, Equipment Code, Organization (site code, e.g. ABQ5/ACY9), Due Date, Department (SSP/BBM/RME/AMOC), WO Status (IP/R/H/RCAN/RR), Hold Reason (REPORT/SCHEDULING/VENDOR/RESOURCE/ORDER/PARTS/VERIFY/NOEQUIP), Last Comment (HTML-wrapped), Last Comment By, Last Comment Date, Contractor, Contractor Description. ~350 rows/day; blank-comment rows are excluded from the classified total.

Analytics grain: one row per WO per daily snapshot, partitioned by dt. Aggregates are GROUP BY views in Athena; the WO table reads rows directly.


Slack surfaces

  • Two push surfaces, no App Home. Daily summary post + standalone batched 3rd-escalation alert.
  • Suppress the alert entirely on zero-3rd days (consistent with the "only notify on ALARM, never OK/recovery" preference). One @here per batch.
  • Daily post includes the 📊 Open dashboard link button to Grafana.
  • Long WO lists go in modals, never the channel (keeps under the 100-block limit).
  • Reuse an existing Slack app/bot token where possible (the payments-slackAppHome setup in payments-dashboard is the pattern; bot token in SSM).

Grafana

  • Self-hosted on EC2, VPN-only, kiosk-able for wall display. Athena datasource via IAM role (no static keys).
  • Dashboards as code in grafana/dashboards/; the running instance is never the source of truth — round-trip edits back to repo JSON.
  • Must provide: breakdown parity with the legacy Sheet, trend time-series (the new capability), a filterable/exportable WO table with APM deep-links, and the mismatch panel.
  • This EC2 box is the only non-serverless piece here: it carries OS + Grafana patching and a config/dashboard backup obligation. Keep it reproducible.

Repo-specific rules

  • kebab-case everything (repo, stack, bucket, Lambda, role). Buckets apm-wo-analysis-*-328440206208.
  • Secrets → Secrets Manager (apm-wo-analysis/anthropic-api-key); operational config → SSM.
  • Mandatory cross-review (cross_reviewer, GPT-4.1) on any IAM/policy change or Lambda handler-signature change before merge. Flag as outstanding if the orchestrator is unavailable.
  • Smoke-test the classifier against a real export before declaring any classification change done (see the pre-action-smoke-test preference).
  • README + the Confluence "AWS Architecture Map" (page 1540098) updated in the same work as any architecture change. Update the apm-wo-comment-analysis project memory on status/resource changes.

Repo agents

Repo-specific subagents live in .claude/agents/:

  • classifier-engineer — owns classification accuracy and the "Other"-reduction goal.
  • slack-blockkit-designer — owns the daily post / alert / modal Block Kit.
  • grafana-author — owns the dashboards-as-code JSON and Athena SQL.

Most build work (CDK, Lambda code, git, deploys) stays native. Delegate the IAM/handler review to cross_reviewer; use the sh-* skills at provisioning, review, and documentation boundaries.


Local dev

  • pyenv Python 3.12; ruff check + ruff format --check before pushing (hook-enforced).
  • Sample export for smoke-tests: ~/Downloads/_documents/Sheet1-1.xlsx (raw single-sheet, HTML-wrapped comments).
  • cdk synth must pass in CI before merge; no manual prod deploys.