apm-wo-analysis/docs/BUILD.md

226 lines
13 KiB
Markdown
Raw Permalink Normal View History

# apm-wo-analysis — Step-by-Step Build Guide
Historical CDK build instructions. **PLAT-75 moved this stack to HCP Terraform**
in seahaven-prod (`011934824531`). Authority is `terraform/` plus the README
deployment section. Mgmt CDK stacks were deleted 2026-09-16; do not deploy from
`cdk/`.
End-to-end CDK build instructions below are the original mgmt path. Account
`328440206208`, region `us-east-1`, all names kebab-case.
> Convention gates that apply throughout: OIDC deploy role created **before** any CD; secrets in Secrets Manager; `cross_reviewer` on IAM/handler diffs (orchestrator since archived; use `cross_review.py` in `security-review`); `ruff` clean + `cdk synth` green before push; README + Confluence + memory updated as part of the work, not after.
## Target repo layout
```
apm-wo-analysis/
├── CLAUDE.md
├── README.md
├── cdk/
│ ├── app.py
│ ├── cdk.json
│ ├── requirements.txt # aws-cdk-lib==2.253.1, constructs>=10.6.0
│ └── stacks/
│ ├── pipeline_stack.py # S3, Lambdas, Glue, Athena, IAM, schedule
│ └── grafana_stack.py # VPC import, EC2, ALB, SG, Route53, datasource role
├── lambdas/
│ ├── classifier/ # S3-triggered: parse → classify → write parquet
│ │ ├── handler.py
│ │ ├── classify.py # the two-axis model (the core logic)
│ │ └── requirements.txt # awswrangler, openpyxl, anthropic
│ └── slack_post/ # builds + posts daily summary and alert
│ ├── handler.py
│ └── blockkit.py
├── grafana/
│ ├── provisioning/
│ │ ├── datasources/athena.yaml
│ │ └── dashboards/apm.yaml
│ └── dashboards/apm-work-orders.json
├── scripts/
│ ├── drop_folder_upload.sh # local launchd uploader (lives in repo, deployed outside ~/Documents)
│ └── com.seahaven.apm-wo-drop.plist
├── tests/
│ └── test_classify.py # smoke test against a sample export
└── .github/
├── workflows/{ci.yaml,deploy.yaml}
└── dependabot.yml
```
---
## Phase 0 — Repo provisioning
**0.1 Lock reuse patterns first.** Run the `Explore` agent over `payments-dashboard` (the `payments-slackAppHome` Lambda + how its Slack bot token is read from SSM) and the org `.github` reusable workflows (`ci-python-sam.yaml`, `cd-cdk.yaml`) so the new repo matches them exactly. Optionally scaffold with the `sh-bootstrap` skill.
**0.2 Create the repo.**
```bash
gh repo create Sea-Haven-Industries/apm-wo-analysis --private
git init && git branch -M main
git remote add origin git@github.com:Sea-Haven-Industries/apm-wo-analysis.git
```
**0.3 OIDC deploy role FIRST** (before any CD). Create `githubdeploy-apm-wo-analysis`, trust scoped to `repo:Sea-Haven-Industries/apm-wo-analysis:*`, with permissions to deploy the two stacks (CloudFormation, plus the resource services they create). Mirror the trust policy from `githubdeploy-procurement-ingest`. Store the ARN as the repo secret `AWS_DEPLOY_ROLE_ARN`.
**0.4 CDK scaffold.**
```bash
mkdir -p cdk/stacks lambdas tests grafana
printf 'aws-cdk-lib==2.253.1\nconstructs>=10.6.0\n' > cdk/requirements.txt
```
`cdk/app.py` instantiates `PipelineStack` and `GrafanaStack`, both pinned to env `328440206208`/`us-east-1`.
**0.5 CI/CD + Dependabot.** `.github/workflows/ci.yaml` calls the reusable `ci-python-sam.yaml@main` (lint `lambdas cdk` + `cdk synth`); `deploy.yaml` calls `cd-cdk.yaml@main` (Python 3.12, cdk dir `cdk`, OIDC, `cdk deploy --all`, concurrency single). Add `dependabot.yml` (pip for `cdk/` and each `lambdas/*`, github-actions).
**0.6 Branch protection** on `main`: require the CI check + the Claude Code App review.
**Validate:** open a throwaway PR with the empty scaffold; CI `cdk synth` is green.
---
## Phase 1 — Ingestion (direct S3 / local drop-folder, NO email)
**1.1 Bucket** in `pipeline_stack.py`:
```python
bucket = s3.Bucket(self, "Exports",
bucket_name=f"apm-wo-analysis-exports-{self.account}",
removal_policy=RETAIN, encryption=s3.BucketEncryption.S3_MANAGED,
block_public_access=s3.BlockPublicAccess.BLOCK_ALL,
lifecycle_rules=[s3.LifecycleRule(prefix="raw/", expiration=Duration.days(90))])
```
Prefixes: `raw/` (incoming exports), `analytics/` (per-WO snapshots), `athena-results/`.
**1.2 Direct upload path** (always available):
```bash
aws s3 cp ./Sheet1-1.xlsx s3://apm-wo-analysis-exports-328440206208/raw/
```
**1.3 Local drop-folder path** (optional zero-touch). Mirror `stampli-drop-folder`. **The script and the watched folder must live OUTSIDE `~/Documents`** (TCC sandbox — see `macos-tcc-launchd` memory). Suggested: folder `~/APM-WO-Drop`, script `~/Library/Application Support/seahaven/apm-wo-drop/upload.sh`, plist `~/Library/LaunchAgents/com.seahaven.apm-wo-drop.plist`.
- `upload.sh`: on a new file in the drop folder, `aws s3 cp` it to `raw/`, then move it to a local `processed/` subfolder.
- plist: `WatchPaths` = the drop folder; `ProgramArguments` = the script. `launchctl load` it.
- Uses a least-privilege local IAM user/profile with `s3:PutObject` to `raw/` only.
**Validate:** drop or upload a real export; confirm the object lands in `raw/`.
---
## Phase 2 — Classifier Lambda
**2.1 `lambdas/classifier/classify.py`** — port the two-axis model exactly as specified in `CLAUDE.md`:
- `strip_html(text)` — unwrap `<html>…</html>` and entities.
- `comment_intent(text)` — ordered rule ladder (most-specific first), returns a bucket or `None`.
- `HOLD_TO_CAT` map + `WO Status` rules (`RCAN`→Cancelled, etc.).
- `classify(status, hold, comment)` → `(final_category, mismatch_reason | None)` with resolution policy comment-intent → structured → `Other`.
- Haiku fallback (Secrets Manager `apm-wo-analysis/anthropic-api-key`) invoked **only** for `Other` rows with blank Hold Reason and a non-trivial comment.
**2.2 `lambdas/classifier/handler.py`** — S3-triggered on `raw/`:
1. Read the object, parse with `openpyxl` (xlsx) / csv.
2. For each row: strip HTML, `classify()`, derive `is_escalation`, `is_action`, `mismatch`.
3. Write a per-WO snapshot to `analytics/dt=YYYY-MM-DD/` as Parquet via `awswrangler.s3.to_parquet(..., dataset=True, partition_cols=["dt"], database="apm_wo_analysis", table="apm_wo_snapshots")` (also registers the Glue partition).
4. Emit a small `analytics/dt=YYYY-MM-DD/summary.json` (counts per category, escalation total, action/routine, top sites, mismatch list) for the Slack Lambda to read cheaply.
5. On completion, async-invoke the slack-post Lambda (or fire an EventBridge event).
**2.3 Runtime:** Python 3.12, ARM64, 512 MB, 120 s. Layer: `awswrangler` (AWS SDK for pandas) ARM64 managed layer. IAM: read `raw/`, write `analytics/`, read the Anthropic secret, Glue `CreatePartition`/`BatchCreatePartition`.
**2.4 Cross-review (MANDATORY):** the handler signature + IAM policy diff go through `cross_reviewer` before merge (orchestrator archived; `cross_reviewer` now lives in `security-review`):
```bash
python3 ~/Documents/repositories/seahaven/security-review/cross_review.py "Review this diff for breaking changes: <classifier IAM policy + handler contract>"
```
**Validate (smoke test):** `pytest tests/test_classify.py` runs `classify()` over the sample export and asserts "Other" ≤ ~10% and that known fixtures land in the right buckets. Use the `classifier-engineer` agent for tuning.
---
## Phase 3 — Analytics dataset (Glue + Athena)
**3.1 Glue database** `apm_wo_analysis` (CDK `glue.CfnDatabase`).
**3.2 Table** `apm_wo_snapshots` over `s3://apm-wo-analysis-exports-328440206208/analytics/`, columns matching the snapshot (wo_number, description, equipment_code, site, due_date, department, wo_status, hold_reason, last_comment, last_comment_by, last_comment_date, contractor, category, is_escalation, is_action, mismatch), partitioned by `dt` with **partition projection** enabled:
```
projection.enabled = true
projection.dt.type = date
projection.dt.format = yyyy-MM-dd
projection.dt.range = 2026-01-01,NOW
storage.location.template = s3://.../analytics/dt=${dt}/
```
Projection means no crawler and no `MSCK REPAIR`.
**3.3 Athena workgroup** `apm-wo-analysis` with result location `s3://.../athena-results/` and result encryption.
**Validate:**
```sql
SELECT category, count(*) FROM apm_wo_analysis.apm_wo_snapshots
WHERE dt = current_date GROUP BY 1 ORDER BY 2 DESC;
```
returns today's breakdown matching the smoke-test distribution.
---
## Phase 4 — Slack post + alert Lambda
**4.1 `lambdas/slack_post/blockkit.py`** — build the daily summary and the batched 3rd-escalation alert (port the prototype). Daily post: header, context line with vs-yesterday deltas, escalation + action/routine fields, top sites, mismatch callout, action buttons incl. **📊 Open dashboard** (`url` → `https://grafana.seahaven.com/d/apm-wo/...?from=now-30d&to=now`), footer. Alert: batched, one `@here`, **return `None` / skip post when zero 3rd escalations**.
**4.2 `lambdas/slack_post/handler.py`** — triggered after the classifier:
1. Read `analytics/dt=today/summary.json` and `dt=yesterday/summary.json` (deltas).
2. Post the daily summary to the WO channel via the reused bot token (SSM, same pattern as `payments-dashboard`).
3. If 3rd-escalation count > 0, post the standalone alert; else post nothing.
4. Drill-down `block_actions` (category/site buttons → `views.open` modal) handled by an API Gateway endpoint with Slack signature verification (or defer modals to v1.1 and rely on the Grafana table).
**4.3 IAM:** read `analytics/`, read the Slack token secret/param. **Cross-review** the IAM diff.
**Validate:** point at a test channel; confirm layout, real counts, correct deltas across two days of data, and zero-3rd suppression. Use the `slack-blockkit-designer` agent.
---
## Phase 5 — Grafana (self-hosted EC2, VPN-only)
**5.1 `grafana_stack.py`:**
- Import an existing VPC (or a small dedicated one). EC2 `t4g.small` (ARM64), Amazon Linux 2023, Grafana OSS installed via user-data, EBS gp3 with `RETAIN`.
- **Security group: ingress only from VPN / office CIDRs** (see `office-ips` memory) on the ALB; ALB→instance on 3000.
- Internal/again-restricted **ALB** with HTTPS using the wildcard ACM cert `*.seahaven.com` (ARN from context, same pattern as the slack bot). Target group → instance:3000.
- Route53 A/alias `grafana.seahaven.com` → ALB in zone `Z06652411XKH89KTZD3XA`.
- **Instance role** with Athena (`StartQueryExecution`, `GetQueryResults`), Glue (`GetTable`/`GetPartitions`), and S3 read on the analytics + athena-results prefixes. No static keys.
**5.2 Provisioning (dashboards-as-code).** Ship via user-data / config-sync into `/etc/grafana/provisioning/`:
- `datasources/athena.yaml` — Athena datasource using the instance role (default auth provider), workgroup `apm-wo-analysis`, database `apm_wo_analysis`.
- `dashboards/apm.yaml` — provider pointing at the dashboards folder.
- `dashboards/apm-work-orders.json` — the committed dashboard. Use the `grafana-author` agent to build:
- category distribution (bar), escalation summary (stat/pie), action vs routine (donut), escalation-by-site (bar/table);
- trend time-series over `dt`;
- **filterable WO table** with `$site/$department/$category/$status/$hold_reason` template variables, WO-number cell data-links into APM, escalation-row coloring, CSV export;
- mismatch panel; chart-to-table drill via data links.
- Panel refresh hourly (data changes once/day). Kiosk URL for the wall display.
**5.3 Backup.** Snapshot/version the dashboard JSON in-repo (source of truth) and back up `grafana.db` (or use the SQLite on the retained EBS volume). Document restore.
**Validate:** dashboard loads from Athena over VPN; filters work live; the Slack 📊 button deep-links correctly.
---
## Phase 6 — Docs, Confluence, memory
- **README** (`sh-readme`): architecture, data flow, the classification model, ingestion (direct + drop-folder), and an ops **runbook** (`sh-runbook`): how the export gets uploaded, Grafana patching cadence, dashboard-JSON redeploy, EBS/config backup-restore, common failures.
- **Confluence** (`sh-confluence`): add a Mermaid subgraph for this stack to the "AWS Architecture Map" (page 1540098). If Confluence is unreachable, state the doc update as outstanding.
- **Memory** (`sh-distill`): update `apm-wo-comment-analysis` to "deployed" with final resource names; ensure the cross-links to `procurement-ingest`, `payments-dashboard`, `stampli-drop-folder`, and `macos-tcc-launchd` are present.
- **Pre-merge / pre-prod:** `sh-pr-check` on the branch, `sh-prod-ready` before the first real deploy.
---
## Deploy order (once code is in)
```bash
# 0. deploy role already exists (Phase 0.3)
cd cdk && pip install -r requirements.txt
cdk deploy apm-wo-analysis-pipeline # S3, Glue, Athena, Lambdas, IAM
# upload one export, confirm analytics/ partition + summary.json + Slack post
cdk deploy apm-wo-analysis-grafana # EC2, ALB, SG, Route53, datasource role
# load the dashboard, verify Athena queries and the kiosk view
```
## Cost
S3/Athena/Lambda ≈ pennies/month at ~350 rows/day; Grafana `t4g.small` ≈ $12/mo + small EBS + ALB hours. No per-user fees.
## Inputs still required before building
1. APM WO deep-link URL format (for the table cell links + Slack buttons).
2. Slack channel + which bot token (reuse `payments-dashboard` app vs new).
3. Confirm `grafana.seahaven.com` and VPN/office CIDRs for the SG.
4. Confirm the drop-folder path conventions if using the launchd uploader.