Merge pull request #6 from Sea-Haven-Industries/feature/phase-0-scaffold

Phase 0: scaffold apm-wo-analysis repository
This commit is contained in:
Adam Moussa 2026-05-29 13:14:37 -04:00 • committed by GitHub
commit f86b4ba1c5
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
27 changed files with 900 additions and 0 deletions

View file

@ -0,0 +1,27 @@
---
name: classifier-engineer
description: Use this agent to build, tune, or validate the APM work-order comment classification logic. Triggers: a new APM export reveals comments landing in "Other", a bucket needs adding/refining, the Hold Reason / WO Status cross-reference needs adjusting, or the classifier output needs smoke-testing against a real export before merge. It owns classification accuracy and the "Other"-reduction goal.
tools: Read, Edit, Write, Bash, Grep, Glob
model: sonnet
---
You are the classification engineer for the APM Work Order comment analysis pipeline. You own one thing: the accuracy of `classify()` and everything it depends on. Read the project `CLAUDE.md` first for the full taxonomy and mappings.
## The model is two-axis, never comment-only
The single biggest mistake the legacy Google Apps Script made was reading only the `Last Comment` text and ignoring `WO Status` and `Hold Reason`. Never regress to that. Every classification decision considers:
1. **Comment intent** — regex/keyword signals from the stripped comment text.
2. **Structured state** — `Hold Reason` (REPORT/SCHEDULING/VENDOR/PARTS/ORDER/VERIFY/RESOURCE/NOEQUIP) and `WO Status` (IP/R/H/RCAN/RR).
Resolution policy: comment intent wins when confident; fall back to structured state when the comment is silent; only then `Other`. Always emit a `mismatch` flag when comment intent contradicts structured state (e.g. comment says "completed" but `WO Status` is still `IP`, or "schedule confirmed" while on a `SCHEDULING`/`REPORT`/`VENDOR` hold). The mismatch detector is a feature, not noise.
## How you work
- **Always smoke-test against a real export before declaring anything done.** The canonical fixture is the latest `Sheet1-1.xlsx`-style export. Run the classifier in Python, print the full distribution and the per-bucket "Other" residual, and compare before/after. Never claim an accuracy improvement you haven't measured.
- When a comment lands in `Other`, first check whether `Hold Reason`/`WO Status` already resolves it deterministically. Reach for a Haiku fallback only for genuinely ambiguous free-text (e.g. "Copy", "Cant close", "Uplift request submitted") where no structured signal exists.
- Quantify every change: report Other count and %, which buckets moved, and any new mismatches surfaced. Target is single-digit "Other" %.
- Keep classification rules explainable and ordered (most-specific first). Prefer a readable rule ladder over a clever single regex.
- Strip HTML from comments before matching (comments arrive `<html>...</html>`-wrapped).
## Guardrails
- Do not invent buckets without checking the canonical taxonomy in `CLAUDE.md`; if a new bucket is warranted, propose it with the supporting comment samples and counts.
- Changes to the Lambda handler signature or its IAM require the mandatory `cross_reviewer` pass (see project `CLAUDE.md`). Flag when your change crosses that line.
- Never silently drop rows. Blank-comment rows are excluded from the classified total by design; say so when reporting counts.

View file

@ -0,0 +1,27 @@
---
name: grafana-author
description: Use this agent to author or maintain the Grafana dashboard-as-code and the Athena SQL behind it. Triggers: adding/editing a panel, writing or tuning an Athena query over the analytics dataset, adding template variables/filters, or wiring chart-to-table drill-down. Owns the dashboards-as-code JSON and the Athena/Glue query layer.
tools: Read, Edit, Write, Bash, Grep, Glob
model: sonnet
---
You are the Grafana + Athena author for the APM Work Order analysis dashboard. You own the provisioned dashboard JSON and the SQL that feeds it. Read the project `CLAUDE.md` for the dataset schema and the hosting decisions.
## Locked decisions (do not relitigate without being asked)
- **Self-hosted Grafana on EC2, VPN-only.** Datasource is **Athena over the S3 analytics dataset** (not DynamoDB, not CloudWatch). No Timestream.
- **Dashboards are code.** Every panel lives as provisioned JSON checked into the repo (`grafana/dashboards/`). Never treat a hand-edited panel in the running instance as the source of truth — round-trip changes back into the repo JSON so the box is reproducible.
- The dataset grain is **one row per WO per daily snapshot**, partitioned by `dt`. Aggregate panels roll up with `GROUP BY` over partitions; the WO table reads rows directly.
## What the dashboard must cover
- Breakdown parity with the legacy Sheet: category distribution, escalation summary, action vs routine, escalation-by-site.
- What the Sheet never had: **trend time-series** (escalations/day, Other %/day, per-site over time) using the `dt` partition.
- A **filterable WO table**: per-column filters, template variables (`$site`, `$department`, `$category`, `$status`, `$hold_reason`), WO-number cell links into APM, conditional escalation-row coloring, CSV export.
- A **mismatch panel** (comment vs structured-state divergences).
- Chart-to-table drill-down via data links setting the table's variable.
## How you work
- Validate Athena SQL with partition projection in mind; keep scans cheap (filter on `dt`). At ~350 rows/day cost is pennies, but write tidy partitioned queries anyway.
- Set panel refresh to match the once-daily data cadence (hourly is plenty; do not hammer Athena every few seconds on the kiosk display).
- Use a Grafana IAM role for the Athena datasource, never static keys.
- Keep the long free-text `last_comment` readable: enable cell text-wrap or an inspect/detail panel rather than letting it truncate silently.
- When you change a panel, update the committed JSON and note how to redeploy it to the instance.

View file

@ -0,0 +1,25 @@
---
name: slack-blockkit-designer
description: Use this agent to build or iterate the Slack Block Kit surfaces for this repo — the daily summary post, the batched 3rd-escalation alert, and the drill-down modal. Triggers: changing what the daily post shows, adjusting the alert, adding buttons/filters, or validating block counts and Slack API limits before shipping.
tools: Read, Edit, Write, Bash, WebFetch
model: sonnet
---
You are the Slack surface designer for the APM Work Order analysis pipeline. You own the Block Kit JSON and how the daily push reads in the channel. Read the project `CLAUDE.md` for the surface decisions already locked.
## Locked surface decisions (do not relitigate without being asked)
- **Two push surfaces, no App Home.** A daily summary post + a standalone batched 3rd-escalation alert.
- **The 3rd-escalation alert is suppressed entirely on zero-3rd days** (matches the "only notify on ALARM" preference). One `@here` per batch, never one ping per WO.
- The daily post carries a **"📊 Open dashboard"** link button to Grafana.
- Long WO lists live in **modals**, never dumped into the channel — this is how the 100-block-per-surface limit stays a non-issue.
## How you work
- **Validate against the real export, not mock data.** Generate the JSON from the latest classified export so counts and site rollups are real, then report the block count per surface.
- Respect Slack limits: ≤100 blocks per message/view, section `text` ≤3000 chars, ≤10 fields per section (each ≤2000), button text short. If a list could exceed limits, page it ("+N more") and say what was truncated — never silently cut.
- Keep the daily post scannable: header, context line with vs-yesterday deltas, escalation + action/routine fields, top sites, mismatch callout, action buttons, footer context. Don't add panels that belong in Grafana.
- Make WO numbers clickable deep-links into APM (the real URL format is in `CLAUDE.md`; if it's still a placeholder, flag it).
- Output JSON the user can paste straight into Slack's Block Kit Builder, and render a plain-text mock so the layout is reviewable without Slack.
## Style
- Follow the visual-iteration preference: small incremental changes per pass (~15-20%), not wholesale redesigns.
- Honest about fakes: if deltas or links are stubbed pending real data/config, say so explicitly.

41
.github/dependabot.yml vendored Normal file
View file

@ -0,0 +1,41 @@
version: 2
updates:
- package-ecosystem: "pip"
directory: "/cdk"
schedule:
interval: "weekly"
groups:
minor-and-patch:
update-types:
- "minor"
- "patch"
- package-ecosystem: "pip"
directory: "/lambdas/classifier"
schedule:
interval: "weekly"
groups:
minor-and-patch:
update-types:
- "minor"
- "patch"
- package-ecosystem: "pip"
directory: "/lambdas/slack_post"
schedule:
interval: "weekly"
groups:
minor-and-patch:
update-types:
- "minor"
- "patch"
- package-ecosystem: "github-actions"
directory: "/"
schedule:
interval: "weekly"
groups:
minor-and-patch:
update-types:
- "minor"
- "patch"

15
.github/workflows/ci.yaml vendored Normal file
View file

@ -0,0 +1,15 @@
name: CI
on:
pull_request:
branches: [main]
jobs:
ci:
uses: Sea-Haven-Industries/.github/.github/workflows/ci-python-sam.yaml@main
with:
python-version: "3.12"
source-dirs: "cdk lambdas tests"
run-sam-validate: false
run-cdk-synth: true
cdk-dir: cdk
run-tests: false

27
.github/workflows/deploy.yaml vendored Normal file
View file

@ -0,0 +1,27 @@
name: Deploy
on:
push:
branches: [main]
permissions:
id-token: write
contents: read
concurrency:
group: deploy
cancel-in-progress: false
jobs:
deploy:
# TEMPORARILY DISABLED while the Phase 0–5 stack is merged into main, so each
# merge to main does NOT trigger `cdk deploy --all`. Re-enable when starting
# Phase 6 by removing this `if:` line (revert this commit). The pipeline +
# grafana stacks are already deployed (manually) and validated in prod.
if: ${{ false }}
uses: Sea-Haven-Industries/.github/.github/workflows/cd-cdk.yaml@main
with:
python-version: "3.12"
region: us-east-1
cdk-dir: cdk
secrets:
deploy-role-arn: ${{ secrets.AWS_DEPLOY_ROLE_ARN }}

23
.gitignore vendored Normal file
View file

@ -0,0 +1,23 @@
# Python
__pycache__/
*.py[cod]
*.egg-info/
.venv/
venv/
# CDK
cdk.out/
cdk.context.json
# Local env / secrets
.env
.env.*
# Test / lint caches
.pytest_cache/
.ruff_cache/
.coverage
htmlcov/
# OS
.DS_Store

115
CLAUDE.md Normal file
View file

@ -0,0 +1,115 @@
# Project Instructions — apm-wo-analysis
Project-specific context and rules. Supplements the global `~/.claude/CLAUDE.md` (Sea Haven standards) and the engineering handbook. Where this file is silent, the global rules and handbook apply.
---
## What this is
Daily analysis of Amazon **APM work-order** "Last Comment" data for Sea Haven facility ops. Replaces a legacy Google Apps Script + versioned-Google-Sheet workflow. A curated daily **filter-view export** (~350 WOs) is classified on two axes, then pushed to Slack and surfaced in a Grafana dashboard.
This is a **separate, sibling concern** to the `apm@` email pipeline in `procurement-ingest`. That pipeline event-sources notification emails into a `WorkOrders` table; this repo consumes a different feed (the manual export) and does **not** read those tables. See the `apm-wo-comment-analysis` and `procurement-ingest` memories.
---
## Architecture
```
APM export (xlsx/csv)
→ S3 raw/ (direct upload OR local launchd drop-folder)
→ classifier Lambda (Python, ARM64: HTML strip + two-axis classify, Haiku fallback)
→ S3 analytics/dt=YYYY-MM-DD/ (per-WO daily snapshot, Parquet)
→ Glue table → Athena → Grafana (self-hosted EC2, VPN-only, kiosk)
→ slack-post Lambda (reads today + yesterday partitions)
→ daily summary post [📊 Open dashboard button]
→ standalone batched 3rd-escalation alert (suppressed if zero)
```
- **Account / region:** 328440206208 / us-east-1
- **IaC:** CDK (Python). `aws-cdk-lib` pinned **==2.253.1**. Lambdas Python 3.12, ARM64.
- **No DynamoDB** — deliberate. This is an analytics workload and Grafana cannot query DynamoDB; S3 + Athena is the store. Do not "helpfully" add a table.
---
## Ingestion (no email)
The export reaches S3 by **direct upload or a local drop-folder**, never SES/email.
- Direct: console/CLI `aws s3 cp` into `raw/`.
- Drop-folder: a launchd agent that uploads files dropped in a local folder, mirroring the `stampli-drop-folder` pattern. **The launchd script must live outside `~/Documents`** (macOS TCC sandbox — see the `macos-tcc-launchd` memory). The watched drop folder must also sit outside `~/Documents`.
- The classifier Lambda is S3-triggered on the `raw/` prefix regardless of how the file arrives.
---
## The classification model (the core domain knowledge)
**Always two-axis. Never comment-only.** The legacy script's central flaw was reading only the comment and ignoring `WO Status` + `Hold Reason`; that left ~17% in "Other" and mis-stated dozens of WOs. The two-axis model cut "Other" to ~9% before any AI.
**Axis 1 — comment intent** (regex over the HTML-stripped `Last Comment`), most-specific first:
3rd / 2nd / 1st Escalation · SIM Ticket · Vendor No-Show · Weekly WO Scheduled · Schedule Confirmed · Awaiting Scheduling · Report / Docs Needed · Awaiting Report / Invoice · Completed / Pending Close · Acknowledgement / No-op · Cancelled · On Hold · Rescheduled · Avetta Project Created · Other Escalation · Status Inquiry.
**Axis 2 — structured state:**
- `Hold Reason` → category: `SCHEDULING`→Awaiting Scheduling, `REPORT`→Report / Docs Needed, `VENDOR`/`PARTS`/`ORDER`→Awaiting Vendor / Parts, `VERIFY`→Verification Needed, `RESOURCE`→Resource Hold, `NOEQUIP`→No Equipment.
- `WO Status`: `RCAN`→Cancelled; `H` corroborates On Hold; `IP`/`R`/`RR` are in-flight states.
**Resolution:** comment intent wins when confident → else structured state → else `Other`. Reserve a **Claude Haiku** fallback (Secrets Manager key) strictly for ambiguous free-text with no structured signal ("Copy", "Cant close", "Uplift request submitted").
**Mismatch detector (a feature):** flag when comment intent contradicts structured state — e.g. comment "completed/scheduled" while on a `REPORT`/`SCHEDULING`/`VENDOR` hold, or "completed" while `WO Status` is `IP`. Surface these; do not suppress.
**Escalation categories:** 1st / 2nd / 3rd Escalation, SIM Ticket, Other Escalation.
**Action-needed categories:** escalations + Awaiting Scheduling + Report / Docs Needed + Awaiting Report / Invoice + Awaiting Vendor / Parts + Status Inquiry + Vendor No-Show. Everything else is routine.
---
## Export schema (13 columns)
`WO Number, WO Description, Equipment Code, Organization` (site code, e.g. ABQ5/ACY9), `Due Date, Department` (SSP/BBM/RME/AMOC), `WO Status` (IP/R/H/RCAN/RR), `Hold Reason` (REPORT/SCHEDULING/VENDOR/RESOURCE/ORDER/PARTS/VERIFY/NOEQUIP), `Last Comment` (HTML-wrapped), `Last Comment By, Last Comment Date, Contractor, Contractor Description`. ~350 rows/day; blank-comment rows are excluded from the classified total.
**Analytics grain:** one row per WO per daily snapshot, partitioned by `dt`. Aggregates are `GROUP BY` views in Athena; the WO table reads rows directly.
---
## Slack surfaces
- **Two push surfaces, no App Home.** Daily summary post + standalone batched 3rd-escalation alert.
- **Suppress the alert entirely on zero-3rd days** (consistent with the "only notify on ALARM, never OK/recovery" preference). One `@here` per batch.
- Daily post includes the **📊 Open dashboard** link button to Grafana.
- Long WO lists go in **modals**, never the channel (keeps under the 100-block limit).
- Reuse an existing Slack app/bot token where possible (the `payments-slackAppHome` setup in `payments-dashboard` is the pattern; bot token in SSM).
---
## Grafana
- **Self-hosted on EC2, VPN-only**, kiosk-able for wall display. Athena datasource via IAM role (no static keys).
- **Dashboards as code** in `grafana/dashboards/`; the running instance is never the source of truth — round-trip edits back to repo JSON.
- Must provide: breakdown parity with the legacy Sheet, trend time-series (the new capability), a filterable/exportable WO table with APM deep-links, and the mismatch panel.
- This EC2 box is the only non-serverless piece here: it carries OS + Grafana patching and a config/dashboard backup obligation. Keep it reproducible.
---
## Repo-specific rules
- **kebab-case** everything (repo, stack, bucket, Lambda, role). Buckets `apm-wo-analysis-*-328440206208`.
- **Secrets → Secrets Manager** (`apm-wo-analysis/anthropic-api-key`); operational config → SSM.
- **Mandatory cross-review** (`cross_reviewer`, GPT-4.1) on any IAM/policy change or Lambda handler-signature change before merge. Flag as outstanding if the orchestrator is unavailable.
- Smoke-test the classifier against a **real export** before declaring any classification change done (see the pre-action-smoke-test preference).
- README + the Confluence "AWS Architecture Map" (page 1540098) updated **in the same work** as any architecture change. Update the `apm-wo-comment-analysis` project memory on status/resource changes.
---
## Repo agents
Repo-specific subagents live in `.claude/agents/`:
- **classifier-engineer** — owns classification accuracy and the "Other"-reduction goal.
- **slack-blockkit-designer** — owns the daily post / alert / modal Block Kit.
- **grafana-author** — owns the dashboards-as-code JSON and Athena SQL.
Most build work (CDK, Lambda code, git, deploys) stays native. Delegate the IAM/handler review to `cross_reviewer`; use the `sh-*` skills at provisioning, review, and documentation boundaries.
---
## Local dev
- pyenv Python 3.12; `ruff check` + `ruff format --check` before pushing (hook-enforced).
- Sample export for smoke-tests: `~/Downloads/_documents/Sheet1-1.xlsx` (raw single-sheet, HTML-wrapped comments).
- `cdk synth` must pass in CI before merge; no manual prod deploys.

34
cdk/app.py Normal file
View file

@ -0,0 +1,34 @@
#!/usr/bin/env python3
"""CDK app entry point for apm-wo-analysis.
Two stacks, both pinned to account 328440206208 / us-east-1:
- apm-wo-analysis-pipeline : S3, classifier + slack-post Lambdas, Glue, Athena, IAM
- apm-wo-analysis-grafana : self-hosted Grafana on EC2 (VPN-only)
See docs/BUILD.md for the phased build-out.
"""
import aws_cdk as cdk
from stacks.grafana_stack import GrafanaStack
from stacks.pipeline_stack import PipelineStack
ENV = cdk.Environment(account="328440206208", region="us-east-1")
app = cdk.App()
PipelineStack(
app,
"apm-wo-analysis-pipeline",
stack_name="apm-wo-analysis-pipeline",
env=ENV,
)
GrafanaStack(
app,
"apm-wo-analysis-grafana",
stack_name="apm-wo-analysis-grafana",
env=ENV,
)
app.synth()

6
cdk/cdk.json Normal file
View file

@ -0,0 +1,6 @@
{
"app": "python3 app.py",
"context": {
"@aws-cdk/core:bootstrapQualifier": "hnb659fds"
}
}

2
cdk/requirements.txt Normal file
View file

@ -0,0 +1,2 @@
aws-cdk-lib==2.253.1
constructs>=10.6.0

0
cdk/stacks/__init__.py Normal file
View file

View file

@ -0,0 +1,22 @@
"""Grafana stack: VPC import, EC2, ALB, SG, Route53, Athena datasource role.
Scaffold — resources are added in Phase 5 of docs/BUILD.md. The running box is
the only non-serverless piece here (self-hosted Grafana OSS on a t4g.small,
ARM64, VPN-only) and carries an OS/Grafana patching + config-backup obligation.
Dashboards are provisioned as code from grafana/ — the running instance is never
the source of truth.
"""
from aws_cdk import Stack
from constructs import Construct
class GrafanaStack(Stack):
def __init__(self, scope: Construct, construct_id: str, **kwargs) -> None:
super().__init__(scope, construct_id, **kwargs)
# Phase 5 — EC2 (Amazon Linux 2023, Grafana OSS via user-data),
# internal ALB (HTTPS, *.seahaven.com ACM cert),
# SG ingress from VPN/office CIDRs only,
# Route53 alias grafana.seahaven.com,
# instance role: Athena + Glue + S3 read (no static keys). TODO

View file

@ -0,0 +1,49 @@
"""Pipeline stack: S3, classifier + slack-post Lambdas, Glue, Athena, IAM.
Scaffold — the exports bucket (Phase 1 of docs/BUILD.md) is included so the
stack synthesizes to something real. The classifier Lambda (Phase 2), Glue
database + Athena workgroup with partition projection (Phase 3), and the
slack-post Lambda + IAM (Phase 4) are added in their respective phases.
Lambda defaults when added: Python 3.12, ARM64, explicit LogGroup with 60-day
retention. No DynamoDB — this is an S3 + Athena analytics workload (see CLAUDE.md).
"""
from aws_cdk import (
Duration,
RemovalPolicy,
Stack,
)
from aws_cdk import (
aws_s3 as s3,
)
from constructs import Construct
class PipelineStack(Stack):
def __init__(self, scope: Construct, construct_id: str, **kwargs) -> None:
super().__init__(scope, construct_id, **kwargs)
# Phase 1 — single exports bucket.
# Prefixes: raw/ (incoming), analytics/ (per-WO snapshots), athena-results/.
self.exports_bucket = s3.Bucket(
self,
"Exports",
bucket_name=f"apm-wo-analysis-exports-{self.account}",
encryption=s3.BucketEncryption.S3_MANAGED,
block_public_access=s3.BlockPublicAccess.BLOCK_ALL,
enforce_ssl=True,
removal_policy=RemovalPolicy.RETAIN,
lifecycle_rules=[
s3.LifecycleRule(
id="expire-raw-exports",
prefix="raw/",
expiration=Duration.days(90),
)
],
)
# Phase 2 — classifier Lambda, S3-triggered on the raw/ prefix. TODO
# Phase 3 — Glue database `apm_wo_analysis` + Athena workgroup
# (partition projection on dt; no crawler). TODO
# Phase 4 — slack-post Lambda + scoped IAM. TODO

219
docs/BUILD.md Normal file
View file

@ -0,0 +1,219 @@
# apm-wo-analysis — Step-by-Step Build Guide
End-to-end build instructions. Read `../CLAUDE.md` first for the domain model and locked decisions. Account `328440206208`, region `us-east-1`, all names kebab-case.
> Convention gates that apply throughout: OIDC deploy role created **before** any CD; secrets in Secrets Manager; `cross_reviewer` on IAM/handler diffs; `ruff` clean + `cdk synth` green before push; README + Confluence + memory updated as part of the work, not after.
## Target repo layout
```
apm-wo-analysis/
├── CLAUDE.md
├── README.md
├── cdk/
│ ├── app.py
│ ├── cdk.json
│ ├── requirements.txt # aws-cdk-lib==2.253.1, constructs>=10.6.0
│ └── stacks/
│ ├── pipeline_stack.py # S3, Lambdas, Glue, Athena, IAM, schedule
│ └── grafana_stack.py # VPC import, EC2, ALB, SG, Route53, datasource role
├── lambdas/
│ ├── classifier/ # S3-triggered: parse → classify → write parquet
│ │ ├── handler.py
│ │ ├── classify.py # the two-axis model (the core logic)
│ │ └── requirements.txt # awswrangler, openpyxl, anthropic
│ └── slack_post/ # builds + posts daily summary and alert
│ ├── handler.py
│ └── blockkit.py
├── grafana/
│ ├── provisioning/
│ │ ├── datasources/athena.yaml
│ │ └── dashboards/apm.yaml
│ └── dashboards/apm-work-orders.json
├── scripts/
│ ├── drop_folder_upload.sh # local launchd uploader (lives in repo, deployed outside ~/Documents)
│ └── com.seahaven.apm-wo-drop.plist
├── tests/
│ └── test_classify.py # smoke test against a sample export
└── .github/
├── workflows/{ci.yaml,deploy.yaml}
└── dependabot.yml
```
---
## Phase 0 — Repo provisioning
**0.1 Lock reuse patterns first.** Run the `Explore` agent over `payments-dashboard` (the `payments-slackAppHome` Lambda + how its Slack bot token is read from SSM) and the org `.github` reusable workflows (`ci-python-sam.yaml`, `cd-cdk.yaml`) so the new repo matches them exactly. Optionally scaffold with the `sh-bootstrap` skill.
**0.2 Create the repo.**
```bash
gh repo create Sea-Haven-Industries/apm-wo-analysis --private
git init && git branch -M main
git remote add origin git@github.com:Sea-Haven-Industries/apm-wo-analysis.git
```
**0.3 OIDC deploy role FIRST** (before any CD). Create `githubdeploy-apm-wo-analysis`, trust scoped to `repo:Sea-Haven-Industries/apm-wo-analysis:*`, with permissions to deploy the two stacks (CloudFormation, plus the resource services they create). Mirror the trust policy from `githubdeploy-procurement-ingest`. Store the ARN as the repo secret `AWS_DEPLOY_ROLE_ARN`.
**0.4 CDK scaffold.**
```bash
mkdir -p cdk/stacks lambdas tests grafana
printf 'aws-cdk-lib==2.253.1\nconstructs>=10.6.0\n' > cdk/requirements.txt
```
`cdk/app.py` instantiates `PipelineStack` and `GrafanaStack`, both pinned to env `328440206208`/`us-east-1`.
**0.5 CI/CD + Dependabot.** `.github/workflows/ci.yaml` calls the reusable `ci-python-sam.yaml@main` (lint `lambdas cdk` + `cdk synth`); `deploy.yaml` calls `cd-cdk.yaml@main` (Python 3.12, cdk dir `cdk`, OIDC, `cdk deploy --all`, concurrency single). Add `dependabot.yml` (pip for `cdk/` and each `lambdas/*`, github-actions).
**0.6 Branch protection** on `main`: require the CI check + the Claude Code App review.
**Validate:** open a throwaway PR with the empty scaffold; CI `cdk synth` is green.
---
## Phase 1 — Ingestion (direct S3 / local drop-folder, NO email)
**1.1 Bucket** in `pipeline_stack.py`:
```python
bucket = s3.Bucket(self, "Exports",
bucket_name=f"apm-wo-analysis-exports-{self.account}",
removal_policy=RETAIN, encryption=s3.BucketEncryption.S3_MANAGED,
block_public_access=s3.BlockPublicAccess.BLOCK_ALL,
lifecycle_rules=[s3.LifecycleRule(prefix="raw/", expiration=Duration.days(90))])
```
Prefixes: `raw/` (incoming exports), `analytics/` (per-WO snapshots), `athena-results/`.
**1.2 Direct upload path** (always available):
```bash
aws s3 cp ./Sheet1-1.xlsx s3://apm-wo-analysis-exports-328440206208/raw/
```
**1.3 Local drop-folder path** (optional zero-touch). Mirror `stampli-drop-folder`. **The script and the watched folder must live OUTSIDE `~/Documents`** (TCC sandbox — see `macos-tcc-launchd` memory). Suggested: folder `~/APM-WO-Drop`, script `~/Library/Application Support/seahaven/apm-wo-drop/upload.sh`, plist `~/Library/LaunchAgents/com.seahaven.apm-wo-drop.plist`.
- `upload.sh`: on a new file in the drop folder, `aws s3 cp` it to `raw/`, then move it to a local `processed/` subfolder.
- plist: `WatchPaths` = the drop folder; `ProgramArguments` = the script. `launchctl load` it.
- Uses a least-privilege local IAM user/profile with `s3:PutObject` to `raw/` only.
**Validate:** drop or upload a real export; confirm the object lands in `raw/`.
---
## Phase 2 — Classifier Lambda
**2.1 `lambdas/classifier/classify.py`** — port the two-axis model exactly as specified in `CLAUDE.md`:
- `strip_html(text)` — unwrap `<html>…</html>` and entities.
- `comment_intent(text)` — ordered rule ladder (most-specific first), returns a bucket or `None`.
- `HOLD_TO_CAT` map + `WO Status` rules (`RCAN`→Cancelled, etc.).
- `classify(status, hold, comment)` → `(final_category, mismatch_reason | None)` with resolution policy comment-intent → structured → `Other`.
- Haiku fallback (Secrets Manager `apm-wo-analysis/anthropic-api-key`) invoked **only** for `Other` rows with blank Hold Reason and a non-trivial comment.
**2.2 `lambdas/classifier/handler.py`** — S3-triggered on `raw/`:
1. Read the object, parse with `openpyxl` (xlsx) / csv.
2. For each row: strip HTML, `classify()`, derive `is_escalation`, `is_action`, `mismatch`.
3. Write a per-WO snapshot to `analytics/dt=YYYY-MM-DD/` as Parquet via `awswrangler.s3.to_parquet(..., dataset=True, partition_cols=["dt"], database="apm_wo_analysis", table="apm_wo_snapshots")` (also registers the Glue partition).
4. Emit a small `analytics/dt=YYYY-MM-DD/summary.json` (counts per category, escalation total, action/routine, top sites, mismatch list) for the Slack Lambda to read cheaply.
5. On completion, async-invoke the slack-post Lambda (or fire an EventBridge event).
**2.3 Runtime:** Python 3.12, ARM64, 512 MB, 120 s. Layer: `awswrangler` (AWS SDK for pandas) ARM64 managed layer. IAM: read `raw/`, write `analytics/`, read the Anthropic secret, Glue `CreatePartition`/`BatchCreatePartition`.
**2.4 Cross-review (MANDATORY):** the handler signature + IAM policy diff go through `cross_reviewer` before merge:
```bash
python3 ~/Documents/repositories/orchestrator/run.py "Review this diff for breaking changes: <classifier IAM policy + handler contract>"
```
**Validate (smoke test):** `pytest tests/test_classify.py` runs `classify()` over the sample export and asserts "Other" ≤ ~10% and that known fixtures land in the right buckets. Use the `classifier-engineer` agent for tuning.
---
## Phase 3 — Analytics dataset (Glue + Athena)
**3.1 Glue database** `apm_wo_analysis` (CDK `glue.CfnDatabase`).
**3.2 Table** `apm_wo_snapshots` over `s3://apm-wo-analysis-exports-328440206208/analytics/`, columns matching the snapshot (wo_number, description, equipment_code, site, due_date, department, wo_status, hold_reason, last_comment, last_comment_by, last_comment_date, contractor, category, is_escalation, is_action, mismatch), partitioned by `dt` with **partition projection** enabled:
```
projection.enabled = true
projection.dt.type = date
projection.dt.format = yyyy-MM-dd
projection.dt.range = 2026-01-01,NOW
storage.location.template = s3://.../analytics/dt=${dt}/
```
Projection means no crawler and no `MSCK REPAIR`.
**3.3 Athena workgroup** `apm-wo-analysis` with result location `s3://.../athena-results/` and result encryption.
**Validate:**
```sql
SELECT category, count(*) FROM apm_wo_analysis.apm_wo_snapshots
WHERE dt = current_date GROUP BY 1 ORDER BY 2 DESC;
```
returns today's breakdown matching the smoke-test distribution.
---
## Phase 4 — Slack post + alert Lambda
**4.1 `lambdas/slack_post/blockkit.py`** — build the daily summary and the batched 3rd-escalation alert (port the prototype). Daily post: header, context line with vs-yesterday deltas, escalation + action/routine fields, top sites, mismatch callout, action buttons incl. **📊 Open dashboard** (`url` → `https://grafana.seahaven.com/d/apm-wo/...?from=now-30d&to=now`), footer. Alert: batched, one `@here`, **return `None` / skip post when zero 3rd escalations**.
**4.2 `lambdas/slack_post/handler.py`** — triggered after the classifier:
1. Read `analytics/dt=today/summary.json` and `dt=yesterday/summary.json` (deltas).
2. Post the daily summary to the WO channel via the reused bot token (SSM, same pattern as `payments-dashboard`).
3. If 3rd-escalation count > 0, post the standalone alert; else post nothing.
4. Drill-down `block_actions` (category/site buttons → `views.open` modal) handled by an API Gateway endpoint with Slack signature verification (or defer modals to v1.1 and rely on the Grafana table).
**4.3 IAM:** read `analytics/`, read the Slack token secret/param. **Cross-review** the IAM diff.
**Validate:** point at a test channel; confirm layout, real counts, correct deltas across two days of data, and zero-3rd suppression. Use the `slack-blockkit-designer` agent.
---
## Phase 5 — Grafana (self-hosted EC2, VPN-only)
**5.1 `grafana_stack.py`:**
- Import an existing VPC (or a small dedicated one). EC2 `t4g.small` (ARM64), Amazon Linux 2023, Grafana OSS installed via user-data, EBS gp3 with `RETAIN`.
- **Security group: ingress only from VPN / office CIDRs** (see `office-ips` memory) on the ALB; ALB→instance on 3000.
- Internal/again-restricted **ALB** with HTTPS using the wildcard ACM cert `*.seahaven.com` (ARN from context, same pattern as the slack bot). Target group → instance:3000.
- Route53 A/alias `grafana.seahaven.com` → ALB in zone `Z06652411XKH89KTZD3XA`.
- **Instance role** with Athena (`StartQueryExecution`, `GetQueryResults`), Glue (`GetTable`/`GetPartitions`), and S3 read on the analytics + athena-results prefixes. No static keys.
**5.2 Provisioning (dashboards-as-code).** Ship via user-data / config-sync into `/etc/grafana/provisioning/`:
- `datasources/athena.yaml` — Athena datasource using the instance role (default auth provider), workgroup `apm-wo-analysis`, database `apm_wo_analysis`.
- `dashboards/apm.yaml` — provider pointing at the dashboards folder.
- `dashboards/apm-work-orders.json` — the committed dashboard. Use the `grafana-author` agent to build:
- category distribution (bar), escalation summary (stat/pie), action vs routine (donut), escalation-by-site (bar/table);
- trend time-series over `dt`;
- **filterable WO table** with `$site/$department/$category/$status/$hold_reason` template variables, WO-number cell data-links into APM, escalation-row coloring, CSV export;
- mismatch panel; chart-to-table drill via data links.
- Panel refresh hourly (data changes once/day). Kiosk URL for the wall display.
**5.3 Backup.** Snapshot/version the dashboard JSON in-repo (source of truth) and back up `grafana.db` (or use the SQLite on the retained EBS volume). Document restore.
**Validate:** dashboard loads from Athena over VPN; filters work live; the Slack 📊 button deep-links correctly.
---
## Phase 6 — Docs, Confluence, memory
- **README** (`sh-readme`): architecture, data flow, the classification model, ingestion (direct + drop-folder), and an ops **runbook** (`sh-runbook`): how the export gets uploaded, Grafana patching cadence, dashboard-JSON redeploy, EBS/config backup-restore, common failures.
- **Confluence** (`sh-confluence`): add a Mermaid subgraph for this stack to the "AWS Architecture Map" (page 1540098). If Confluence is unreachable, state the doc update as outstanding.
- **Memory** (`sh-distill`): update `apm-wo-comment-analysis` to "deployed" with final resource names; ensure the cross-links to `procurement-ingest`, `payments-dashboard`, `stampli-drop-folder`, and `macos-tcc-launchd` are present.
- **Pre-merge / pre-prod:** `sh-pr-check` on the branch, `sh-prod-ready` before the first real deploy.
---
## Deploy order (once code is in)
```bash
# 0. deploy role already exists (Phase 0.3)
cd cdk && pip install -r requirements.txt
cdk deploy apm-wo-analysis-pipeline # S3, Glue, Athena, Lambdas, IAM
# upload one export, confirm analytics/ partition + summary.json + Slack post
cdk deploy apm-wo-analysis-grafana # EC2, ALB, SG, Route53, datasource role
# load the dashboard, verify Athena queries and the kiosk view
```
## Cost
S3/Athena/Lambda ≈ pennies/month at ~350 rows/day; Grafana `t4g.small` ≈ $12/mo + small EBS + ALB hours. No per-user fees.
## Inputs still required before building
1. APM WO deep-link URL format (for the table cell links + Slack buttons).
2. Slack channel + which bot token (reuse `payments-dashboard` app vs new).
3. Confirm `grafana.seahaven.com` and VPN/office CIDRs for the SG.
4. Confirm the drop-folder path conventions if using the launchd uploader.

View file

@ -0,0 +1,18 @@
{
"uid": "apm-wo",
"title": "APM Work Orders",
"description": "Daily APM work-order analysis — breakdown, escalations, trend, filterable WO table, mismatches. Built in Phase 5 (docs/BUILD.md) by the grafana-author agent.",
"tags": ["apm", "work-orders"],
"timezone": "browser",
"schemaVersion": 39,
"version": 1,
"refresh": "1h",
"time": {
"from": "now-30d",
"to": "now"
},
"templating": {
"list": []
},
"panels": []
}

View file

@ -0,0 +1,14 @@
# Dashboard provider — loads the committed JSON from the repo (source of truth).
# Round-trip any UI edits back to grafana/dashboards/*.json (Phase 5).
apiVersion: 1
providers:
- name: apm
orgId: 1
folder: APM
type: file
disableDeletion: false
allowUiUpdates: false
updateIntervalSeconds: 60
options:
path: /var/lib/grafana/dashboards
foldersFromFilesStructure: false

View file

@ -0,0 +1,14 @@
# Athena datasource, authenticated via the EC2 instance IAM role (no static keys).
# Provisioned into /etc/grafana/provisioning/datasources/ via user-data (Phase 5).
apiVersion: 1
datasources:
- name: Athena
type: grafana-athena-datasource
isDefault: true
jsonData:
authType: ec2_iam_role
defaultRegion: us-east-1
catalog: AwsDataCatalog
database: apm_wo_analysis
workgroup: apm-wo-analysis
outputLocation: s3://apm-wo-analysis-exports-328440206208/athena-results/

View file

@ -0,0 +1,76 @@
"""Two-axis APM work-order classifier — the core domain logic.
Axis 1 — comment intent: regex over the HTML-stripped ``Last Comment``,
most-specific first.
Axis 2 — structured state: ``Hold Reason`` + ``WO Status``.
Resolution: comment intent wins when confident, else structured state, else
``Other``. A Claude Haiku fallback (Secrets Manager
``apm-wo-analysis/anthropic-api-key``) is reserved strictly for ambiguous
free-text with no structured signal. A mismatch detector flags when comment
intent contradicts structured state — surface, never suppress.
The authoritative spec is CLAUDE.md ("The classification model"). This module
is owned by the classifier-engineer agent; the rule ladder and Haiku fallback
are implemented in Phase 2 (docs/BUILD.md) and smoke-tested against a real
export before merge.
"""
from __future__ import annotations
# Axis 2 — Hold Reason → category.
HOLD_TO_CATEGORY = {
"SCHEDULING": "Awaiting Scheduling",
"REPORT": "Report / Docs Needed",
"VENDOR": "Awaiting Vendor / Parts",
"PARTS": "Awaiting Vendor / Parts",
"ORDER": "Awaiting Vendor / Parts",
"VERIFY": "Verification Needed",
"RESOURCE": "Resource Hold",
"NOEQUIP": "No Equipment",
}
# Axis 2 — WO Status signals.
WO_STATUS_CANCELLED = "RCAN" # → Cancelled
WO_STATUS_HOLD = "H" # corroborates On Hold
WO_STATUS_IN_FLIGHT = frozenset({"IP", "R", "RR"})
ESCALATION_CATEGORIES = frozenset(
{
"1st Escalation",
"2nd Escalation",
"3rd Escalation",
"SIM Ticket",
"Other Escalation",
}
)
ACTION_NEEDED_CATEGORIES = ESCALATION_CATEGORIES | frozenset(
{
"Awaiting Scheduling",
"Report / Docs Needed",
"Awaiting Report / Invoice",
"Awaiting Vendor / Parts",
"Status Inquiry",
"Vendor No-Show",
}
)
def strip_html(comment: str | None) -> str:
"""Unwrap the HTML-wrapped Last Comment and decode entities. (Phase 2)"""
raise NotImplementedError
def comment_intent(comment: str) -> str | None:
"""Ordered, most-specific-first rule ladder; returns a bucket or None. (Phase 2)"""
raise NotImplementedError
def classify(
wo_status: str | None,
hold_reason: str | None,
last_comment: str | None,
) -> tuple[str, str | None]:
"""Resolve to ``(final_category, mismatch_reason | None)``. (Phase 2)"""
raise NotImplementedError

View file

@ -0,0 +1,19 @@
"""S3-triggered classifier Lambda: parse export → two-axis classify → Parquet.
Triggered on ``s3:ObjectCreated`` under the ``raw/`` prefix regardless of how
the file arrives (direct upload or local drop-folder). For each non-blank-comment
row it strips HTML, runs ``classify()``, derives ``is_escalation`` / ``is_action``
/ ``mismatch``, writes a per-WO snapshot to ``analytics/dt=YYYY-MM-DD/`` as
Parquet (registering the Glue partition), emits a small ``summary.json`` for the
slack-post Lambda, then invokes it.
Runtime when wired up: Python 3.12, ARM64, 512 MB, 120 s, awswrangler layer.
Implemented in Phase 2 (docs/BUILD.md).
"""
from __future__ import annotations
def handler(event, context):
"""Lambda entry point. (Phase 2)"""
raise NotImplementedError

View file

@ -0,0 +1,3 @@
awswrangler>=3.9.0
openpyxl>=3.1.0
anthropic>=0.40.0

View file

@ -0,0 +1,23 @@
"""Block Kit builders for the daily summary post and the 3rd-escalation alert.
Two push surfaces, no App Home (see CLAUDE.md "Slack surfaces"):
- daily summary: header, vs-yesterday deltas, escalation + action/routine
fields, top sites, mismatch callout, and a 📊 Open dashboard link button to
Grafana. Long WO lists go in modals, never the channel (<100-block limit).
- 3rd-escalation alert: standalone, batched, one @here — suppressed entirely
on zero-3rd days (return None).
Owned by the slack-blockkit-designer agent; built in Phase 4 (docs/BUILD.md).
"""
from __future__ import annotations
def build_daily_summary(today: dict, yesterday: dict | None) -> list[dict]:
"""Build the daily summary blocks (with vs-yesterday deltas). (Phase 4)"""
raise NotImplementedError
def build_escalation_alert(today: dict) -> list[dict] | None:
"""Build the batched 3rd-escalation alert, or None when count is 0. (Phase 4)"""
raise NotImplementedError

View file

@ -0,0 +1,16 @@
"""Slack post + alert Lambda: daily summary and batched 3rd-escalation alert.
Invoked after the classifier finishes. Reads ``analytics/dt=today/summary.json``
and the yesterday partition (for deltas), posts the daily summary to the WO
channel via the reused bot token (SSM, same pattern as payments-dashboard), and
— only if the 3rd-escalation count > 0 — posts the standalone batched alert.
Implemented in Phase 4 (docs/BUILD.md).
"""
from __future__ import annotations
def handler(event, context):
"""Lambda entry point. (Phase 4)"""
raise NotImplementedError

View file

@ -0,0 +1 @@
slack_sdk>=3.33.0

View file

@ -0,0 +1,31 @@
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<!--
launchd agent for the apm-wo-analysis drop-folder uploader.
Install (paths must be OUTSIDE ~/Documents — macOS TCC sandbox):
cp scripts/drop_folder_upload.sh "$HOME/Library/Application Support/seahaven/apm-wo-drop/upload.sh"
cp scripts/com.seahaven.apm-wo-drop.plist "$HOME/Library/LaunchAgents/"
launchctl load "$HOME/Library/LaunchAgents/com.seahaven.apm-wo-drop.plist"
Replace <USER> below with the deploying account's home before loading.
-->
<plist version="1.0">
<dict>
<key>Label</key>
<string>com.seahaven.apm-wo-drop</string>
<key>ProgramArguments</key>
<array>
<string>/bin/bash</string>
<string>/Users/<USER>/Library/Application Support/seahaven/apm-wo-drop/upload.sh</string>
</array>
<key>WatchPaths</key>
<array>
<string>/Users/<USER>/APM-WO-Drop</string>
</array>
<key>StandardOutPath</key>
<string>/tmp/apm-wo-drop.out.log</string>
<key>StandardErrorPath</key>
<string>/tmp/apm-wo-drop.err.log</string>
</dict>
</plist>

29
scripts/drop_folder_upload.sh Executable file
View file

@ -0,0 +1,29 @@
#!/usr/bin/env bash
# apm-wo-analysis local drop-folder uploader (optional zero-touch ingestion).
#
# Mirrors the stampli-drop-folder pattern. When deployed, BOTH this script and
# the watched folder must live OUTSIDE ~/Documents (macOS TCC sandbox — see the
# macos-tcc-launchd memory). Suggested install locations:
# script: ~/Library/Application Support/seahaven/apm-wo-drop/upload.sh
# folder: ~/APM-WO-Drop
#
# Triggered by the launchd agent (scripts/com.seahaven.apm-wo-drop.plist) on a
# WatchPaths change. Uploads each new export to the raw/ prefix, then moves it to
# a local processed/ subfolder. Uses a least-privilege local AWS profile scoped
# to s3:PutObject on raw/ only. The classifier Lambda is S3-triggered from there.
set -euo pipefail
DROP_DIR="${APM_WO_DROP_DIR:-$HOME/APM-WO-Drop}"
PROCESSED_DIR="$DROP_DIR/processed"
BUCKET="apm-wo-analysis-exports-328440206208"
PROFILE="${APM_WO_AWS_PROFILE:-apm-wo-drop}"
mkdir -p "$PROCESSED_DIR"
shopt -s nullglob
for f in "$DROP_DIR"/*.xlsx "$DROP_DIR"/*.csv; do
[ -e "$f" ] || continue
name="$(basename "$f")"
aws --profile "$PROFILE" s3 cp "$f" "s3://${BUCKET}/raw/${name}"
mv "$f" "$PROCESSED_DIR/$name"
done

24
tests/test_classify.py Normal file
View file

@ -0,0 +1,24 @@
"""Smoke test for the two-axis classifier.
The full smoke test — run ``classify()`` over the sample export
(~/Downloads/_documents/Sheet1-1.xlsx) and assert "Other" <= ~10% with known
fixtures landing in the right buckets — is added in Phase 2 (docs/BUILD.md) and
owned by the classifier-engineer agent. This placeholder keeps the test layout
in place and guards the classification constants.
"""
import sys
from pathlib import Path
sys.path.insert(
0, str(Path(__file__).resolve().parent.parent / "lambdas" / "classifier")
)
import classify # noqa: E402
def test_classification_constants_well_formed():
assert "3rd Escalation" in classify.ESCALATION_CATEGORIES
assert classify.HOLD_TO_CATEGORY["SCHEDULING"] == "Awaiting Scheduling"
# Every escalation category is also action-needed.
assert classify.ESCALATION_CATEGORIES <= classify.ACTION_NEEDED_CATEGORIES