mirror of
https://github.com/Sea-Haven-Industries/apm-wo-analysis.git
synced 2026-09-30 17:03:15 +00:00
219 lines
13 KiB
Markdown
219 lines
13 KiB
Markdown
# apm-wo-analysis — Step-by-Step Build Guide
|
|
|
|
End-to-end build instructions. Read `../CLAUDE.md` first for the domain model and locked decisions. Account `328440206208`, region `us-east-1`, all names kebab-case.
|
|
|
|
> Convention gates that apply throughout: OIDC deploy role created **before** any CD; secrets in Secrets Manager; `cross_reviewer` on IAM/handler diffs (orchestrator since archived; use `cross_review.py` in `security-review`); `ruff` clean + `cdk synth` green before push; README + Confluence + memory updated as part of the work, not after.
|
|
|
|
## Target repo layout
|
|
|
|
```
|
|
apm-wo-analysis/
|
|
├── CLAUDE.md
|
|
├── README.md
|
|
├── cdk/
|
|
│ ├── app.py
|
|
│ ├── cdk.json
|
|
│ ├── requirements.txt # aws-cdk-lib==2.253.1, constructs>=10.6.0
|
|
│ └── stacks/
|
|
│ ├── pipeline_stack.py # S3, Lambdas, Glue, Athena, IAM, schedule
|
|
│ └── grafana_stack.py # VPC import, EC2, ALB, SG, Route53, datasource role
|
|
├── lambdas/
|
|
│ ├── classifier/ # S3-triggered: parse → classify → write parquet
|
|
│ │ ├── handler.py
|
|
│ │ ├── classify.py # the two-axis model (the core logic)
|
|
│ │ └── requirements.txt # awswrangler, openpyxl, anthropic
|
|
│ └── slack_post/ # builds + posts daily summary and alert
|
|
│ ├── handler.py
|
|
│ └── blockkit.py
|
|
├── grafana/
|
|
│ ├── provisioning/
|
|
│ │ ├── datasources/athena.yaml
|
|
│ │ └── dashboards/apm.yaml
|
|
│ └── dashboards/apm-work-orders.json
|
|
├── scripts/
|
|
│ ├── drop_folder_upload.sh # local launchd uploader (lives in repo, deployed outside ~/Documents)
|
|
│ └── com.seahaven.apm-wo-drop.plist
|
|
├── tests/
|
|
│ └── test_classify.py # smoke test against a sample export
|
|
└── .github/
|
|
├── workflows/{ci.yaml,deploy.yaml}
|
|
└── dependabot.yml
|
|
```
|
|
|
|
---
|
|
|
|
## Phase 0 — Repo provisioning
|
|
|
|
**0.1 Lock reuse patterns first.** Run the `Explore` agent over `payments-dashboard` (the `payments-slackAppHome` Lambda + how its Slack bot token is read from SSM) and the org `.github` reusable workflows (`ci-python-sam.yaml`, `cd-cdk.yaml`) so the new repo matches them exactly. Optionally scaffold with the `sh-bootstrap` skill.
|
|
|
|
**0.2 Create the repo.**
|
|
```bash
|
|
gh repo create Sea-Haven-Industries/apm-wo-analysis --private
|
|
git init && git branch -M main
|
|
git remote add origin git@github.com:Sea-Haven-Industries/apm-wo-analysis.git
|
|
```
|
|
|
|
**0.3 OIDC deploy role FIRST** (before any CD). Create `githubdeploy-apm-wo-analysis`, trust scoped to `repo:Sea-Haven-Industries/apm-wo-analysis:*`, with permissions to deploy the two stacks (CloudFormation, plus the resource services they create). Mirror the trust policy from `githubdeploy-procurement-ingest`. Store the ARN as the repo secret `AWS_DEPLOY_ROLE_ARN`.
|
|
|
|
**0.4 CDK scaffold.**
|
|
```bash
|
|
mkdir -p cdk/stacks lambdas tests grafana
|
|
printf 'aws-cdk-lib==2.253.1\nconstructs>=10.6.0\n' > cdk/requirements.txt
|
|
```
|
|
`cdk/app.py` instantiates `PipelineStack` and `GrafanaStack`, both pinned to env `328440206208`/`us-east-1`.
|
|
|
|
**0.5 CI/CD + Dependabot.** `.github/workflows/ci.yaml` calls the reusable `ci-python-sam.yaml@main` (lint `lambdas cdk` + `cdk synth`); `deploy.yaml` calls `cd-cdk.yaml@main` (Python 3.12, cdk dir `cdk`, OIDC, `cdk deploy --all`, concurrency single). Add `dependabot.yml` (pip for `cdk/` and each `lambdas/*`, github-actions).
|
|
|
|
**0.6 Branch protection** on `main`: require the CI check + the Claude Code App review.
|
|
|
|
**Validate:** open a throwaway PR with the empty scaffold; CI `cdk synth` is green.
|
|
|
|
---
|
|
|
|
## Phase 1 — Ingestion (direct S3 / local drop-folder, NO email)
|
|
|
|
**1.1 Bucket** in `pipeline_stack.py`:
|
|
```python
|
|
bucket = s3.Bucket(self, "Exports",
|
|
bucket_name=f"apm-wo-analysis-exports-{self.account}",
|
|
removal_policy=RETAIN, encryption=s3.BucketEncryption.S3_MANAGED,
|
|
block_public_access=s3.BlockPublicAccess.BLOCK_ALL,
|
|
lifecycle_rules=[s3.LifecycleRule(prefix="raw/", expiration=Duration.days(90))])
|
|
```
|
|
Prefixes: `raw/` (incoming exports), `analytics/` (per-WO snapshots), `athena-results/`.
|
|
|
|
**1.2 Direct upload path** (always available):
|
|
```bash
|
|
aws s3 cp ./Sheet1-1.xlsx s3://apm-wo-analysis-exports-328440206208/raw/
|
|
```
|
|
|
|
**1.3 Local drop-folder path** (optional zero-touch). Mirror `stampli-drop-folder`. **The script and the watched folder must live OUTSIDE `~/Documents`** (TCC sandbox — see `macos-tcc-launchd` memory). Suggested: folder `~/APM-WO-Drop`, script `~/Library/Application Support/seahaven/apm-wo-drop/upload.sh`, plist `~/Library/LaunchAgents/com.seahaven.apm-wo-drop.plist`.
|
|
- `upload.sh`: on a new file in the drop folder, `aws s3 cp` it to `raw/`, then move it to a local `processed/` subfolder.
|
|
- plist: `WatchPaths` = the drop folder; `ProgramArguments` = the script. `launchctl load` it.
|
|
- Uses a least-privilege local IAM user/profile with `s3:PutObject` to `raw/` only.
|
|
|
|
**Validate:** drop or upload a real export; confirm the object lands in `raw/`.
|
|
|
|
---
|
|
|
|
## Phase 2 — Classifier Lambda
|
|
|
|
**2.1 `lambdas/classifier/classify.py`** — port the two-axis model exactly as specified in `CLAUDE.md`:
|
|
- `strip_html(text)` — unwrap `<html>…</html>` and entities.
|
|
- `comment_intent(text)` — ordered rule ladder (most-specific first), returns a bucket or `None`.
|
|
- `HOLD_TO_CAT` map + `WO Status` rules (`RCAN`→Cancelled, etc.).
|
|
- `classify(status, hold, comment)` → `(final_category, mismatch_reason | None)` with resolution policy comment-intent → structured → `Other`.
|
|
- Haiku fallback (Secrets Manager `apm-wo-analysis/anthropic-api-key`) invoked **only** for `Other` rows with blank Hold Reason and a non-trivial comment.
|
|
|
|
**2.2 `lambdas/classifier/handler.py`** — S3-triggered on `raw/`:
|
|
1. Read the object, parse with `openpyxl` (xlsx) / csv.
|
|
2. For each row: strip HTML, `classify()`, derive `is_escalation`, `is_action`, `mismatch`.
|
|
3. Write a per-WO snapshot to `analytics/dt=YYYY-MM-DD/` as Parquet via `awswrangler.s3.to_parquet(..., dataset=True, partition_cols=["dt"], database="apm_wo_analysis", table="apm_wo_snapshots")` (also registers the Glue partition).
|
|
4. Emit a small `analytics/dt=YYYY-MM-DD/summary.json` (counts per category, escalation total, action/routine, top sites, mismatch list) for the Slack Lambda to read cheaply.
|
|
5. On completion, async-invoke the slack-post Lambda (or fire an EventBridge event).
|
|
|
|
**2.3 Runtime:** Python 3.12, ARM64, 512 MB, 120 s. Layer: `awswrangler` (AWS SDK for pandas) ARM64 managed layer. IAM: read `raw/`, write `analytics/`, read the Anthropic secret, Glue `CreatePartition`/`BatchCreatePartition`.
|
|
|
|
**2.4 Cross-review (MANDATORY):** the handler signature + IAM policy diff go through `cross_reviewer` before merge (orchestrator archived; `cross_reviewer` now lives in `security-review`):
|
|
```bash
|
|
python3 ~/Documents/repositories/seahaven/security-review/cross_review.py "Review this diff for breaking changes: <classifier IAM policy + handler contract>"
|
|
```
|
|
|
|
**Validate (smoke test):** `pytest tests/test_classify.py` runs `classify()` over the sample export and asserts "Other" ≤ ~10% and that known fixtures land in the right buckets. Use the `classifier-engineer` agent for tuning.
|
|
|
|
---
|
|
|
|
## Phase 3 — Analytics dataset (Glue + Athena)
|
|
|
|
**3.1 Glue database** `apm_wo_analysis` (CDK `glue.CfnDatabase`).
|
|
|
|
**3.2 Table** `apm_wo_snapshots` over `s3://apm-wo-analysis-exports-328440206208/analytics/`, columns matching the snapshot (wo_number, description, equipment_code, site, due_date, department, wo_status, hold_reason, last_comment, last_comment_by, last_comment_date, contractor, category, is_escalation, is_action, mismatch), partitioned by `dt` with **partition projection** enabled:
|
|
```
|
|
projection.enabled = true
|
|
projection.dt.type = date
|
|
projection.dt.format = yyyy-MM-dd
|
|
projection.dt.range = 2026-01-01,NOW
|
|
storage.location.template = s3://.../analytics/dt=${dt}/
|
|
```
|
|
Projection means no crawler and no `MSCK REPAIR`.
|
|
|
|
**3.3 Athena workgroup** `apm-wo-analysis` with result location `s3://.../athena-results/` and result encryption.
|
|
|
|
**Validate:**
|
|
```sql
|
|
SELECT category, count(*) FROM apm_wo_analysis.apm_wo_snapshots
|
|
WHERE dt = current_date GROUP BY 1 ORDER BY 2 DESC;
|
|
```
|
|
returns today's breakdown matching the smoke-test distribution.
|
|
|
|
---
|
|
|
|
## Phase 4 — Slack post + alert Lambda
|
|
|
|
**4.1 `lambdas/slack_post/blockkit.py`** — build the daily summary and the batched 3rd-escalation alert (port the prototype). Daily post: header, context line with vs-yesterday deltas, escalation + action/routine fields, top sites, mismatch callout, action buttons incl. **📊 Open dashboard** (`url` → `https://grafana.seahaven.com/d/apm-wo/...?from=now-30d&to=now`), footer. Alert: batched, one `@here`, **return `None` / skip post when zero 3rd escalations**.
|
|
|
|
**4.2 `lambdas/slack_post/handler.py`** — triggered after the classifier:
|
|
1. Read `analytics/dt=today/summary.json` and `dt=yesterday/summary.json` (deltas).
|
|
2. Post the daily summary to the WO channel via the reused bot token (SSM, same pattern as `payments-dashboard`).
|
|
3. If 3rd-escalation count > 0, post the standalone alert; else post nothing.
|
|
4. Drill-down `block_actions` (category/site buttons → `views.open` modal) handled by an API Gateway endpoint with Slack signature verification (or defer modals to v1.1 and rely on the Grafana table).
|
|
|
|
**4.3 IAM:** read `analytics/`, read the Slack token secret/param. **Cross-review** the IAM diff.
|
|
|
|
**Validate:** point at a test channel; confirm layout, real counts, correct deltas across two days of data, and zero-3rd suppression. Use the `slack-blockkit-designer` agent.
|
|
|
|
---
|
|
|
|
## Phase 5 — Grafana (self-hosted EC2, VPN-only)
|
|
|
|
**5.1 `grafana_stack.py`:**
|
|
- Import an existing VPC (or a small dedicated one). EC2 `t4g.small` (ARM64), Amazon Linux 2023, Grafana OSS installed via user-data, EBS gp3 with `RETAIN`.
|
|
- **Security group: ingress only from VPN / office CIDRs** (see `office-ips` memory) on the ALB; ALB→instance on 3000.
|
|
- Internal/again-restricted **ALB** with HTTPS using the wildcard ACM cert `*.seahaven.com` (ARN from context, same pattern as the slack bot). Target group → instance:3000.
|
|
- Route53 A/alias `grafana.seahaven.com` → ALB in zone `Z06652411XKH89KTZD3XA`.
|
|
- **Instance role** with Athena (`StartQueryExecution`, `GetQueryResults`), Glue (`GetTable`/`GetPartitions`), and S3 read on the analytics + athena-results prefixes. No static keys.
|
|
|
|
**5.2 Provisioning (dashboards-as-code).** Ship via user-data / config-sync into `/etc/grafana/provisioning/`:
|
|
- `datasources/athena.yaml` — Athena datasource using the instance role (default auth provider), workgroup `apm-wo-analysis`, database `apm_wo_analysis`.
|
|
- `dashboards/apm.yaml` — provider pointing at the dashboards folder.
|
|
- `dashboards/apm-work-orders.json` — the committed dashboard. Use the `grafana-author` agent to build:
|
|
- category distribution (bar), escalation summary (stat/pie), action vs routine (donut), escalation-by-site (bar/table);
|
|
- trend time-series over `dt`;
|
|
- **filterable WO table** with `$site/$department/$category/$status/$hold_reason` template variables, WO-number cell data-links into APM, escalation-row coloring, CSV export;
|
|
- mismatch panel; chart-to-table drill via data links.
|
|
- Panel refresh hourly (data changes once/day). Kiosk URL for the wall display.
|
|
|
|
**5.3 Backup.** Snapshot/version the dashboard JSON in-repo (source of truth) and back up `grafana.db` (or use the SQLite on the retained EBS volume). Document restore.
|
|
|
|
**Validate:** dashboard loads from Athena over VPN; filters work live; the Slack 📊 button deep-links correctly.
|
|
|
|
---
|
|
|
|
## Phase 6 — Docs, Confluence, memory
|
|
|
|
- **README** (`sh-readme`): architecture, data flow, the classification model, ingestion (direct + drop-folder), and an ops **runbook** (`sh-runbook`): how the export gets uploaded, Grafana patching cadence, dashboard-JSON redeploy, EBS/config backup-restore, common failures.
|
|
- **Confluence** (`sh-confluence`): add a Mermaid subgraph for this stack to the "AWS Architecture Map" (page 1540098). If Confluence is unreachable, state the doc update as outstanding.
|
|
- **Memory** (`sh-distill`): update `apm-wo-comment-analysis` to "deployed" with final resource names; ensure the cross-links to `procurement-ingest`, `payments-dashboard`, `stampli-drop-folder`, and `macos-tcc-launchd` are present.
|
|
- **Pre-merge / pre-prod:** `sh-pr-check` on the branch, `sh-prod-ready` before the first real deploy.
|
|
|
|
---
|
|
|
|
## Deploy order (once code is in)
|
|
|
|
```bash
|
|
# 0. deploy role already exists (Phase 0.3)
|
|
cd cdk && pip install -r requirements.txt
|
|
cdk deploy apm-wo-analysis-pipeline # S3, Glue, Athena, Lambdas, IAM
|
|
# upload one export, confirm analytics/ partition + summary.json + Slack post
|
|
cdk deploy apm-wo-analysis-grafana # EC2, ALB, SG, Route53, datasource role
|
|
# load the dashboard, verify Athena queries and the kiosk view
|
|
```
|
|
|
|
## Cost
|
|
S3/Athena/Lambda ≈ pennies/month at ~350 rows/day; Grafana `t4g.small` ≈ $12/mo + small EBS + ALB hours. No per-user fees.
|
|
|
|
## Inputs still required before building
|
|
1. APM WO deep-link URL format (for the table cell links + Slack buttons).
|
|
2. Slack channel + which bot token (reuse `payments-dashboard` app vs new).
|
|
3. Confirm `grafana.seahaven.com` and VPN/office CIDRs for the SG.
|
|
4. Confirm the drop-folder path conventions if using the launchd uploader.
|