apm-wo-analysis/CLAUDE.md

115 lines
7.3 KiB
Markdown
Raw Normal View History

# Project Instructions — apm-wo-analysis
Project-specific context and rules. Supplements the global `~/.claude/CLAUDE.md` (Sea Haven standards) and the engineering handbook. Where this file is silent, the global rules and handbook apply.
---
## What this is
Daily analysis of Amazon **APM work-order** "Last Comment" data for Sea Haven facility ops. Replaces a legacy Google Apps Script + versioned-Google-Sheet workflow. A curated daily **filter-view export** (~350 WOs) is classified on two axes, then pushed to Slack and surfaced in a Grafana dashboard.
This is a **separate, sibling concern** to the `apm@` email pipeline in `procurement-ingest`. That pipeline event-sources notification emails into a `WorkOrders` table; this repo consumes a different feed (the manual export) and does **not** read those tables. See the `apm-wo-comment-analysis` and `procurement-ingest` memories.
---
## Architecture
```
APM export (xlsx/csv)
→ S3 raw/ (direct upload OR local launchd drop-folder)
→ classifier Lambda (Python, ARM64: HTML strip + two-axis classify, Haiku fallback)
→ S3 analytics/dt=YYYY-MM-DD/ (per-WO daily snapshot, Parquet)
→ Glue table → Athena → Grafana (self-hosted EC2, VPN-only, kiosk)
→ slack-post Lambda (reads today + yesterday partitions)
→ daily summary post [📊 Open dashboard button]
→ standalone batched 3rd-escalation alert (suppressed if zero)
```
- **Account / region:** seahaven-prod 011934824531 / us-east-1
- **IaC:** HCP Terraform (Python Lambdas 3.12, ARM64). `cdk/` is frozen leftover from the mgmt stack; do not deploy it.
- **No DynamoDB** — deliberate. This is an analytics workload and Grafana cannot query DynamoDB; S3 + Athena is the store. Do not "helpfully" add a table.
---
## Ingestion (no email)
The export reaches S3 by **direct upload or a local drop-folder**, never SES/email.
- Direct: console/CLI `aws s3 cp` into `raw/`.
- Drop-folder: a launchd agent that uploads files dropped in a local folder, mirroring the `stampli-drop-folder` pattern. **The launchd script must live outside `~/Documents`** (macOS TCC sandbox — see the `macos-tcc-launchd` memory). The watched drop folder must also sit outside `~/Documents`.
- The classifier Lambda is S3-triggered on the `raw/` prefix regardless of how the file arrives.
---
## The classification model (the core domain knowledge)
**Always two-axis. Never comment-only.** The legacy script's central flaw was reading only the comment and ignoring `WO Status` + `Hold Reason`; that left ~17% in "Other" and mis-stated dozens of WOs. The two-axis model cut "Other" to ~9% before any AI.
**Axis 1 — comment intent** (regex over the HTML-stripped `Last Comment`), most-specific first:
3rd / 2nd / 1st Escalation · SIM Ticket · Vendor No-Show · Weekly WO Scheduled · Schedule Confirmed · Awaiting Scheduling · Report / Docs Needed · Awaiting Report / Invoice · Completed / Pending Close · Acknowledgement / No-op · Cancelled · On Hold · Rescheduled · Avetta Project Created · Other Escalation · Status Inquiry.
**Axis 2 — structured state:**
- `Hold Reason` → category: `SCHEDULING`→Awaiting Scheduling, `REPORT`→Report / Docs Needed, `VENDOR`/`PARTS`/`ORDER`→Awaiting Vendor / Parts, `VERIFY`→Verification Needed, `RESOURCE`→Resource Hold, `NOEQUIP`→No Equipment.
- `WO Status`: `RCAN`→Cancelled; `H` corroborates On Hold; `IP`/`R`/`RR` are in-flight states.
**Resolution:** comment intent wins when confident → else structured state → else `Other`. Reserve a **Claude Haiku** fallback (Secrets Manager key) strictly for ambiguous free-text with no structured signal ("Copy", "Cant close", "Uplift request submitted").
**Mismatch detector (a feature):** flag when comment intent contradicts structured state — e.g. comment "completed/scheduled" while on a `REPORT`/`SCHEDULING`/`VENDOR` hold, or "completed" while `WO Status` is `IP`. Surface these; do not suppress.
**Escalation categories:** 1st / 2nd / 3rd Escalation, SIM Ticket, Other Escalation.
**Action-needed categories:** escalations + Awaiting Scheduling + Report / Docs Needed + Awaiting Report / Invoice + Awaiting Vendor / Parts + Status Inquiry + Vendor No-Show. Everything else is routine.
---
## Export schema (13 columns)
`WO Number, WO Description, Equipment Code, Organization` (site code, e.g. ABQ5/ACY9), `Due Date, Department` (SSP/BBM/RME/AMOC), `WO Status` (IP/R/H/RCAN/RR), `Hold Reason` (REPORT/SCHEDULING/VENDOR/RESOURCE/ORDER/PARTS/VERIFY/NOEQUIP), `Last Comment` (HTML-wrapped), `Last Comment By, Last Comment Date, Contractor, Contractor Description`. ~350 rows/day; blank-comment rows are excluded from the classified total.
**Analytics grain:** one row per WO per daily snapshot, partitioned by `dt`. Aggregates are `GROUP BY` views in Athena; the WO table reads rows directly.
---
## Slack surfaces
- **Two push surfaces, no App Home.** Daily summary post + standalone batched 3rd-escalation alert.
- **Suppress the alert entirely on zero-3rd days** (consistent with the "only notify on ALARM, never OK/recovery" preference). One `@here` per batch.
- Daily post includes the **📊 Open dashboard** link button to Grafana.
- Long WO lists go in **modals**, never the channel (keeps under the 100-block limit).
- Reuse an existing Slack app/bot token where possible (the `payments-slackAppHome` setup in `payments-dashboard` is the pattern; bot token in SSM).
---
## Grafana
- **Self-hosted on EC2, VPN-only**, kiosk-able for wall display. Athena datasource via IAM role (no static keys).
- **Dashboards as code** in `grafana/dashboards/`; the running instance is never the source of truth — round-trip edits back to repo JSON.
- Must provide: breakdown parity with the legacy Sheet, trend time-series (the new capability), a filterable/exportable WO table with APM deep-links, and the mismatch panel.
- This EC2 box is the only non-serverless piece here: it carries OS + Grafana patching and a config/dashboard backup obligation. Keep it reproducible.
---
## Repo-specific rules
- Global Sea Haven rules and the engineering handbook apply (naming, secrets placement, cross-review gates, README/Confluence updates, ruff before push).
- Buckets: `apm-wo-analysis-*-011934824531`. Anthropic API key secret: `apm-wo-analysis/anthropic-api-key`.
- Smoke-test the classifier against a **real export** before declaring any classification change done (see the pre-action-smoke-test preference).
- Confluence page for architecture changes: "AWS Architecture Map" (page 1540098). Project memory: `apm-wo-comment-analysis`.
---
## Repo agents
Repo-specific subagents live in `.claude/agents/`:
- **classifier-engineer** — owns classification accuracy and the "Other"-reduction goal.
- **slack-blockkit-designer** — owns the daily post / alert / modal Block Kit.
- **grafana-author** — owns the dashboards-as-code JSON and Athena SQL.
Most build work (CDK, Lambda code, git, deploys) stays native. For the IAM/handler review, run `cross_review.py` (`~/Documents/repositories/seahaven/security-review/`); use the `sh-*` skills at provisioning, review, and documentation boundaries.
---
## Local dev
- pyenv Python 3.12.
- Sample export for smoke-tests: `~/Downloads/_documents/Sheet1-1.xlsx` (raw single-sheet, HTML-wrapped comments).
- `cdk synth` must pass in CI before merge; no manual prod deploys.