* feat(infra): migrate pipeline and Grafana to HCP Terraform (PLAT-75) Move apm-wo-analysis into seahaven-prod under workspace apm-wo-analysis-prod with in-repo hcptf/githubdeploy IAM, stub Lambdas, and GitHub Actions zip CD. * chore(iam): add Checkov skip comments for HCP IAM documents Pre-push HIGH findings are the DLM snapshot describe, tagged EC2 creates, exec boundary DescribeLogGroups star, and the drop-uploader user policy.
7.3 KiB
Project Instructions — apm-wo-analysis
Project-specific context and rules. Supplements the global ~/.claude/CLAUDE.md (Sea Haven standards) and the engineering handbook. Where this file is silent, the global rules and handbook apply.
What this is
Daily analysis of Amazon APM work-order "Last Comment" data for Sea Haven facility ops. Replaces a legacy Google Apps Script + versioned-Google-Sheet workflow. A curated daily filter-view export (~350 WOs) is classified on two axes, then pushed to Slack and surfaced in a Grafana dashboard.
This is a separate, sibling concern to the apm@ email pipeline in procurement-ingest. That pipeline event-sources notification emails into a WorkOrders table; this repo consumes a different feed (the manual export) and does not read those tables. See the apm-wo-comment-analysis and procurement-ingest memories.
Architecture
APM export (xlsx/csv)
→ S3 raw/ (direct upload OR local launchd drop-folder)
→ classifier Lambda (Python, ARM64: HTML strip + two-axis classify, Haiku fallback)
→ S3 analytics/dt=YYYY-MM-DD/ (per-WO daily snapshot, Parquet)
→ Glue table → Athena → Grafana (self-hosted EC2, VPN-only, kiosk)
→ slack-post Lambda (reads today + yesterday partitions)
→ daily summary post [📊 Open dashboard button]
→ standalone batched 3rd-escalation alert (suppressed if zero)
- Account / region: seahaven-prod 011934824531 / us-east-1
- IaC: HCP Terraform (Python Lambdas 3.12, ARM64).
cdk/is frozen leftover from the mgmt stack; do not deploy it. - No DynamoDB — deliberate. This is an analytics workload and Grafana cannot query DynamoDB; S3 + Athena is the store. Do not "helpfully" add a table.
Ingestion (no email)
The export reaches S3 by direct upload or a local drop-folder, never SES/email.
- Direct: console/CLI
aws s3 cpintoraw/. - Drop-folder: a launchd agent that uploads files dropped in a local folder, mirroring the
stampli-drop-folderpattern. The launchd script must live outside~/Documents(macOS TCC sandbox — see themacos-tcc-launchdmemory). The watched drop folder must also sit outside~/Documents. - The classifier Lambda is S3-triggered on the
raw/prefix regardless of how the file arrives.
The classification model (the core domain knowledge)
Always two-axis. Never comment-only. The legacy script's central flaw was reading only the comment and ignoring WO Status + Hold Reason; that left ~17% in "Other" and mis-stated dozens of WOs. The two-axis model cut "Other" to ~9% before any AI.
Axis 1 — comment intent (regex over the HTML-stripped Last Comment), most-specific first:
3rd / 2nd / 1st Escalation · SIM Ticket · Vendor No-Show · Weekly WO Scheduled · Schedule Confirmed · Awaiting Scheduling · Report / Docs Needed · Awaiting Report / Invoice · Completed / Pending Close · Acknowledgement / No-op · Cancelled · On Hold · Rescheduled · Avetta Project Created · Other Escalation · Status Inquiry.
Axis 2 — structured state:
Hold Reason→ category:SCHEDULING→Awaiting Scheduling,REPORT→Report / Docs Needed,VENDOR/PARTS/ORDER→Awaiting Vendor / Parts,VERIFY→Verification Needed,RESOURCE→Resource Hold,NOEQUIP→No Equipment.WO Status:RCAN→Cancelled;Hcorroborates On Hold;IP/R/RRare in-flight states.
Resolution: comment intent wins when confident → else structured state → else Other. Reserve a Claude Haiku fallback (Secrets Manager key) strictly for ambiguous free-text with no structured signal ("Copy", "Cant close", "Uplift request submitted").
Mismatch detector (a feature): flag when comment intent contradicts structured state — e.g. comment "completed/scheduled" while on a REPORT/SCHEDULING/VENDOR hold, or "completed" while WO Status is IP. Surface these; do not suppress.
Escalation categories: 1st / 2nd / 3rd Escalation, SIM Ticket, Other Escalation. Action-needed categories: escalations + Awaiting Scheduling + Report / Docs Needed + Awaiting Report / Invoice + Awaiting Vendor / Parts + Status Inquiry + Vendor No-Show. Everything else is routine.
Export schema (13 columns)
WO Number, WO Description, Equipment Code, Organization (site code, e.g. ABQ5/ACY9), Due Date, Department (SSP/BBM/RME/AMOC), WO Status (IP/R/H/RCAN/RR), Hold Reason (REPORT/SCHEDULING/VENDOR/RESOURCE/ORDER/PARTS/VERIFY/NOEQUIP), Last Comment (HTML-wrapped), Last Comment By, Last Comment Date, Contractor, Contractor Description. ~350 rows/day; blank-comment rows are excluded from the classified total.
Analytics grain: one row per WO per daily snapshot, partitioned by dt. Aggregates are GROUP BY views in Athena; the WO table reads rows directly.
Slack surfaces
- Two push surfaces, no App Home. Daily summary post + standalone batched 3rd-escalation alert.
- Suppress the alert entirely on zero-3rd days (consistent with the "only notify on ALARM, never OK/recovery" preference). One
@hereper batch. - Daily post includes the 📊 Open dashboard link button to Grafana.
- Long WO lists go in modals, never the channel (keeps under the 100-block limit).
- Reuse an existing Slack app/bot token where possible (the
payments-slackAppHomesetup inpayments-dashboardis the pattern; bot token in SSM).
Grafana
- Self-hosted on EC2, VPN-only, kiosk-able for wall display. Athena datasource via IAM role (no static keys).
- Dashboards as code in
grafana/dashboards/; the running instance is never the source of truth — round-trip edits back to repo JSON. - Must provide: breakdown parity with the legacy Sheet, trend time-series (the new capability), a filterable/exportable WO table with APM deep-links, and the mismatch panel.
- This EC2 box is the only non-serverless piece here: it carries OS + Grafana patching and a config/dashboard backup obligation. Keep it reproducible.
Repo-specific rules
- Global Sea Haven rules and the engineering handbook apply (naming, secrets placement, cross-review gates, README/Confluence updates, ruff before push).
- Buckets:
apm-wo-analysis-*-011934824531. Anthropic API key secret:apm-wo-analysis/anthropic-api-key. - Smoke-test the classifier against a real export before declaring any classification change done (see the pre-action-smoke-test preference).
- Confluence page for architecture changes: "AWS Architecture Map" (page 1540098). Project memory:
apm-wo-comment-analysis.
Repo agents
Repo-specific subagents live in .claude/agents/:
- classifier-engineer — owns classification accuracy and the "Other"-reduction goal.
- slack-blockkit-designer — owns the daily post / alert / modal Block Kit.
- grafana-author — owns the dashboards-as-code JSON and Athena SQL.
Most build work (CDK, Lambda code, git, deploys) stays native. For the IAM/handler review, run cross_review.py (~/Documents/repositories/seahaven/security-review/); use the sh-* skills at provisioning, review, and documentation boundaries.
Local dev
- pyenv Python 3.12.
- Sample export for smoke-tests:
~/Downloads/_documents/Sheet1-1.xlsx(raw single-sheet, HTML-wrapped comments). cdk synthmust pass in CI before merge; no manual prod deploys.