# apm-wo-analysis — Step-by-Step Build Guide Historical CDK build instructions. **PLAT-75 moved this stack to HCP Terraform** in seahaven-prod (`011934824531`). Authority is `terraform/` plus the README deployment section. Mgmt CDK stacks were deleted 2026-09-16; do not deploy from `cdk/`. End-to-end CDK build instructions below are the original mgmt path. Account `328440206208`, region `us-east-1`, all names kebab-case. > Convention gates that apply throughout: OIDC deploy role created **before** any CD; secrets in Secrets Manager; `cross_reviewer` on IAM/handler diffs (orchestrator since archived; use `cross_review.py` in `security-review`); `ruff` clean + `cdk synth` green before push; README + Confluence + memory updated as part of the work, not after. ## Target repo layout ``` apm-wo-analysis/ ├── CLAUDE.md ├── README.md ├── cdk/ │ ├── app.py │ ├── cdk.json │ ├── requirements.txt # aws-cdk-lib==2.253.1, constructs>=10.6.0 │ └── stacks/ │ ├── pipeline_stack.py # S3, Lambdas, Glue, Athena, IAM, schedule │ └── grafana_stack.py # VPC import, EC2, ALB, SG, Route53, datasource role ├── lambdas/ │ ├── classifier/ # S3-triggered: parse → classify → write parquet │ │ ├── handler.py │ │ ├── classify.py # the two-axis model (the core logic) │ │ └── requirements.txt # awswrangler, openpyxl, anthropic │ └── slack_post/ # builds + posts daily summary and alert │ ├── handler.py │ └── blockkit.py ├── grafana/ │ ├── provisioning/ │ │ ├── datasources/athena.yaml │ │ └── dashboards/apm.yaml │ └── dashboards/apm-work-orders.json ├── scripts/ │ ├── drop_folder_upload.sh # local launchd uploader (lives in repo, deployed outside ~/Documents) │ └── com.seahaven.apm-wo-drop.plist ├── tests/ │ └── test_classify.py # smoke test against a sample export └── .github/ ├── workflows/{ci.yaml,deploy.yaml} └── dependabot.yml ``` --- ## Phase 0 — Repo provisioning **0.1 Lock reuse patterns first.** Run the `Explore` agent over `payments-dashboard` (the `payments-slackAppHome` Lambda + how its Slack bot token is read from SSM) and the org `.github` reusable workflows (`ci-python-sam.yaml`, `cd-cdk.yaml`) so the new repo matches them exactly. Optionally scaffold with the `sh-bootstrap` skill. **0.2 Create the repo.** ```bash gh repo create Sea-Haven-Industries/apm-wo-analysis --private git init && git branch -M main git remote add origin git@github.com:Sea-Haven-Industries/apm-wo-analysis.git ``` **0.3 OIDC deploy role FIRST** (before any CD). Create `githubdeploy-apm-wo-analysis`, trust scoped to `repo:Sea-Haven-Industries/apm-wo-analysis:*`, with permissions to deploy the two stacks (CloudFormation, plus the resource services they create). Mirror the trust policy from `githubdeploy-procurement-ingest`. Store the ARN as the repo secret `AWS_DEPLOY_ROLE_ARN`. **0.4 CDK scaffold.** ```bash mkdir -p cdk/stacks lambdas tests grafana printf 'aws-cdk-lib==2.253.1\nconstructs>=10.6.0\n' > cdk/requirements.txt ``` `cdk/app.py` instantiates `PipelineStack` and `GrafanaStack`, both pinned to env `328440206208`/`us-east-1`. **0.5 CI/CD + Dependabot.** `.github/workflows/ci.yaml` calls the reusable `ci-python-sam.yaml@main` (lint `lambdas cdk` + `cdk synth`); `deploy.yaml` calls `cd-cdk.yaml@main` (Python 3.12, cdk dir `cdk`, OIDC, `cdk deploy --all`, concurrency single). Add `dependabot.yml` (pip for `cdk/` and each `lambdas/*`, github-actions). **0.6 Branch protection** on `main`: require the CI check + the Claude Code App review. **Validate:** open a throwaway PR with the empty scaffold; CI `cdk synth` is green. --- ## Phase 1 — Ingestion (direct S3 / local drop-folder, NO email) **1.1 Bucket** in `pipeline_stack.py`: ```python bucket = s3.Bucket(self, "Exports", bucket_name=f"apm-wo-analysis-exports-{self.account}", removal_policy=RETAIN, encryption=s3.BucketEncryption.S3_MANAGED, block_public_access=s3.BlockPublicAccess.BLOCK_ALL, lifecycle_rules=[s3.LifecycleRule(prefix="raw/", expiration=Duration.days(90))]) ``` Prefixes: `raw/` (incoming exports), `analytics/` (per-WO snapshots), `athena-results/`. **1.2 Direct upload path** (always available): ```bash aws s3 cp ./Sheet1-1.xlsx s3://apm-wo-analysis-exports-328440206208/raw/ ``` **1.3 Local drop-folder path** (optional zero-touch). Mirror `stampli-drop-folder`. **The script and the watched folder must live OUTSIDE `~/Documents`** (TCC sandbox — see `macos-tcc-launchd` memory). Suggested: folder `~/APM-WO-Drop`, script `~/Library/Application Support/seahaven/apm-wo-drop/upload.sh`, plist `~/Library/LaunchAgents/com.seahaven.apm-wo-drop.plist`. - `upload.sh`: on a new file in the drop folder, `aws s3 cp` it to `raw/`, then move it to a local `processed/` subfolder. - plist: `WatchPaths` = the drop folder; `ProgramArguments` = the script. `launchctl load` it. - Uses a least-privilege local IAM user/profile with `s3:PutObject` to `raw/` only. **Validate:** drop or upload a real export; confirm the object lands in `raw/`. --- ## Phase 2 — Classifier Lambda **2.1 `lambdas/classifier/classify.py`** — port the two-axis model exactly as specified in `CLAUDE.md`: - `strip_html(text)` — unwrap `…` and entities. - `comment_intent(text)` — ordered rule ladder (most-specific first), returns a bucket or `None`. - `HOLD_TO_CAT` map + `WO Status` rules (`RCAN`→Cancelled, etc.). - `classify(status, hold, comment)` → `(final_category, mismatch_reason | None)` with resolution policy comment-intent → structured → `Other`. - Haiku fallback (Secrets Manager `apm-wo-analysis/anthropic-api-key`) invoked **only** for `Other` rows with blank Hold Reason and a non-trivial comment. **2.2 `lambdas/classifier/handler.py`** — S3-triggered on `raw/`: 1. Read the object, parse with `openpyxl` (xlsx) / csv. 2. For each row: strip HTML, `classify()`, derive `is_escalation`, `is_action`, `mismatch`. 3. Write a per-WO snapshot to `analytics/dt=YYYY-MM-DD/` as Parquet via `awswrangler.s3.to_parquet(..., dataset=True, partition_cols=["dt"], database="apm_wo_analysis", table="apm_wo_snapshots")` (also registers the Glue partition). 4. Emit a small `analytics/dt=YYYY-MM-DD/summary.json` (counts per category, escalation total, action/routine, top sites, mismatch list) for the Slack Lambda to read cheaply. 5. On completion, async-invoke the slack-post Lambda (or fire an EventBridge event). **2.3 Runtime:** Python 3.12, ARM64, 512 MB, 120 s. Layer: `awswrangler` (AWS SDK for pandas) ARM64 managed layer. IAM: read `raw/`, write `analytics/`, read the Anthropic secret, Glue `CreatePartition`/`BatchCreatePartition`. **2.4 Cross-review (MANDATORY):** the handler signature + IAM policy diff go through `cross_reviewer` before merge (orchestrator archived; `cross_reviewer` now lives in `security-review`): ```bash python3 ~/Documents/repositories/seahaven/security-review/cross_review.py "Review this diff for breaking changes: " ``` **Validate (smoke test):** `pytest tests/test_classify.py` runs `classify()` over the sample export and asserts "Other" ≤ ~10% and that known fixtures land in the right buckets. Use the `classifier-engineer` agent for tuning. --- ## Phase 3 — Analytics dataset (Glue + Athena) **3.1 Glue database** `apm_wo_analysis` (CDK `glue.CfnDatabase`). **3.2 Table** `apm_wo_snapshots` over `s3://apm-wo-analysis-exports-328440206208/analytics/`, columns matching the snapshot (wo_number, description, equipment_code, site, due_date, department, wo_status, hold_reason, last_comment, last_comment_by, last_comment_date, contractor, category, is_escalation, is_action, mismatch), partitioned by `dt` with **partition projection** enabled: ``` projection.enabled = true projection.dt.type = date projection.dt.format = yyyy-MM-dd projection.dt.range = 2026-01-01,NOW storage.location.template = s3://.../analytics/dt=${dt}/ ``` Projection means no crawler and no `MSCK REPAIR`. **3.3 Athena workgroup** `apm-wo-analysis` with result location `s3://.../athena-results/` and result encryption. **Validate:** ```sql SELECT category, count(*) FROM apm_wo_analysis.apm_wo_snapshots WHERE dt = current_date GROUP BY 1 ORDER BY 2 DESC; ``` returns today's breakdown matching the smoke-test distribution. --- ## Phase 4 — Slack post + alert Lambda **4.1 `lambdas/slack_post/blockkit.py`** — build the daily summary and the batched 3rd-escalation alert (port the prototype). Daily post: header, context line with vs-yesterday deltas, escalation + action/routine fields, top sites, mismatch callout, action buttons incl. **📊 Open dashboard** (`url` → `https://grafana.seahaven.com/d/apm-wo/...?from=now-30d&to=now`), footer. Alert: batched, one `@here`, **return `None` / skip post when zero 3rd escalations**. **4.2 `lambdas/slack_post/handler.py`** — triggered after the classifier: 1. Read `analytics/dt=today/summary.json` and `dt=yesterday/summary.json` (deltas). 2. Post the daily summary to the WO channel via the reused bot token (SSM, same pattern as `payments-dashboard`). 3. If 3rd-escalation count > 0, post the standalone alert; else post nothing. 4. Drill-down `block_actions` (category/site buttons → `views.open` modal) handled by an API Gateway endpoint with Slack signature verification (or defer modals to v1.1 and rely on the Grafana table). **4.3 IAM:** read `analytics/`, read the Slack token secret/param. **Cross-review** the IAM diff. **Validate:** point at a test channel; confirm layout, real counts, correct deltas across two days of data, and zero-3rd suppression. Use the `slack-blockkit-designer` agent. --- ## Phase 5 — Grafana (self-hosted EC2, VPN-only) **5.1 `grafana_stack.py`:** - Import an existing VPC (or a small dedicated one). EC2 `t4g.small` (ARM64), Amazon Linux 2023, Grafana OSS installed via user-data, EBS gp3 with `RETAIN`. - **Security group: ingress only from VPN / office CIDRs** (see `office-ips` memory) on the ALB; ALB→instance on 3000. - Internal/again-restricted **ALB** with HTTPS using the wildcard ACM cert `*.seahaven.com` (ARN from context, same pattern as the slack bot). Target group → instance:3000. - Route53 A/alias `grafana.seahaven.com` → ALB in zone `Z06652411XKH89KTZD3XA`. - **Instance role** with Athena (`StartQueryExecution`, `GetQueryResults`), Glue (`GetTable`/`GetPartitions`), and S3 read on the analytics + athena-results prefixes. No static keys. **5.2 Provisioning (dashboards-as-code).** Ship via user-data / config-sync into `/etc/grafana/provisioning/`: - `datasources/athena.yaml` — Athena datasource using the instance role (default auth provider), workgroup `apm-wo-analysis`, database `apm_wo_analysis`. - `dashboards/apm.yaml` — provider pointing at the dashboards folder. - `dashboards/apm-work-orders.json` — the committed dashboard. Use the `grafana-author` agent to build: - category distribution (bar), escalation summary (stat/pie), action vs routine (donut), escalation-by-site (bar/table); - trend time-series over `dt`; - **filterable WO table** with `$site/$department/$category/$status/$hold_reason` template variables, WO-number cell data-links into APM, escalation-row coloring, CSV export; - mismatch panel; chart-to-table drill via data links. - Panel refresh hourly (data changes once/day). Kiosk URL for the wall display. **5.3 Backup.** Snapshot/version the dashboard JSON in-repo (source of truth) and back up `grafana.db` (or use the SQLite on the retained EBS volume). Document restore. **Validate:** dashboard loads from Athena over VPN; filters work live; the Slack 📊 button deep-links correctly. --- ## Phase 6 — Docs, Confluence, memory - **README** (`sh-readme`): architecture, data flow, the classification model, ingestion (direct + drop-folder), and an ops **runbook** (`sh-runbook`): how the export gets uploaded, Grafana patching cadence, dashboard-JSON redeploy, EBS/config backup-restore, common failures. - **Confluence** (`sh-confluence`): add a Mermaid subgraph for this stack to the "AWS Architecture Map" (page 1540098). If Confluence is unreachable, state the doc update as outstanding. - **Memory** (`sh-distill`): update `apm-wo-comment-analysis` to "deployed" with final resource names; ensure the cross-links to `procurement-ingest`, `payments-dashboard`, `stampli-drop-folder`, and `macos-tcc-launchd` are present. - **Pre-merge / pre-prod:** `sh-pr-check` on the branch, `sh-prod-ready` before the first real deploy. --- ## Deploy order (once code is in) ```bash # 0. deploy role already exists (Phase 0.3) cd cdk && pip install -r requirements.txt cdk deploy apm-wo-analysis-pipeline # S3, Glue, Athena, Lambdas, IAM # upload one export, confirm analytics/ partition + summary.json + Slack post cdk deploy apm-wo-analysis-grafana # EC2, ALB, SG, Route53, datasource role # load the dashboard, verify Athena queries and the kiosk view ``` ## Cost S3/Athena/Lambda ≈ pennies/month at ~350 rows/day; Grafana `t4g.small` ≈ $12/mo + small EBS + ALB hours. No per-user fees. ## Inputs still required before building 1. APM WO deep-link URL format (for the table cell links + Slack buttons). 2. Slack channel + which bot token (reuse `payments-dashboard` app vs new). 3. Confirm `grafana.seahaven.com` and VPN/office CIDRs for the SG. 4. Confirm the drop-folder path conventions if using the launchd uploader.