apm-wo-analysis/docs/BUILD.md
2026-07-14 19:24:08 -04:00

13 KiB

apm-wo-analysis — Step-by-Step Build Guide

End-to-end build instructions. Read ../CLAUDE.md first for the domain model and locked decisions. Account 328440206208, region us-east-1, all names kebab-case.

Convention gates that apply throughout: OIDC deploy role created before any CD; secrets in Secrets Manager; cross_reviewer on IAM/handler diffs (orchestrator since archived; use cross_review.py in security-review); ruff clean + cdk synth green before push; README + Confluence + memory updated as part of the work, not after.

Target repo layout

apm-wo-analysis/
├── CLAUDE.md
├── README.md
├── cdk/
│   ├── app.py
│   ├── cdk.json
│   ├── requirements.txt              # aws-cdk-lib==2.253.1, constructs>=10.6.0
│   └── stacks/
│       ├── pipeline_stack.py         # S3, Lambdas, Glue, Athena, IAM, schedule
│       └── grafana_stack.py          # VPC import, EC2, ALB, SG, Route53, datasource role
├── lambdas/
│   ├── classifier/                   # S3-triggered: parse → classify → write parquet
│   │   ├── handler.py
│   │   ├── classify.py               # the two-axis model (the core logic)
│   │   └── requirements.txt          # awswrangler, openpyxl, anthropic
│   └── slack_post/                   # builds + posts daily summary and alert
│       ├── handler.py
│       └── blockkit.py
├── grafana/
│   ├── provisioning/
│   │   ├── datasources/athena.yaml
│   │   └── dashboards/apm.yaml
│   └── dashboards/apm-work-orders.json
├── scripts/
│   ├── drop_folder_upload.sh         # local launchd uploader (lives in repo, deployed outside ~/Documents)
│   └── com.seahaven.apm-wo-drop.plist
├── tests/
│   └── test_classify.py              # smoke test against a sample export
└── .github/
    ├── workflows/{ci.yaml,deploy.yaml}
    └── dependabot.yml

Phase 0 — Repo provisioning

0.1 Lock reuse patterns first. Run the Explore agent over payments-dashboard (the payments-slackAppHome Lambda + how its Slack bot token is read from SSM) and the org .github reusable workflows (ci-python-sam.yaml, cd-cdk.yaml) so the new repo matches them exactly. Optionally scaffold with the sh-bootstrap skill.

0.2 Create the repo.

gh repo create Sea-Haven-Industries/apm-wo-analysis --private
git init && git branch -M main
git remote add origin git@github.com:Sea-Haven-Industries/apm-wo-analysis.git

0.3 OIDC deploy role FIRST (before any CD). Create githubdeploy-apm-wo-analysis, trust scoped to repo:Sea-Haven-Industries/apm-wo-analysis:*, with permissions to deploy the two stacks (CloudFormation, plus the resource services they create). Mirror the trust policy from githubdeploy-procurement-ingest. Store the ARN as the repo secret AWS_DEPLOY_ROLE_ARN.

0.4 CDK scaffold.

mkdir -p cdk/stacks lambdas tests grafana
printf 'aws-cdk-lib==2.253.1\nconstructs>=10.6.0\n' > cdk/requirements.txt

cdk/app.py instantiates PipelineStack and GrafanaStack, both pinned to env 328440206208/us-east-1.

0.5 CI/CD + Dependabot. .github/workflows/ci.yaml calls the reusable ci-python-sam.yaml@main (lint lambdas cdk + cdk synth); deploy.yaml calls cd-cdk.yaml@main (Python 3.12, cdk dir cdk, OIDC, cdk deploy --all, concurrency single). Add dependabot.yml (pip for cdk/ and each lambdas/*, github-actions).

0.6 Branch protection on main: require the CI check + the Claude Code App review.

Validate: open a throwaway PR with the empty scaffold; CI cdk synth is green.


Phase 1 — Ingestion (direct S3 / local drop-folder, NO email)

1.1 Bucket in pipeline_stack.py:

bucket = s3.Bucket(self, "Exports",
    bucket_name=f"apm-wo-analysis-exports-{self.account}",
    removal_policy=RETAIN, encryption=s3.BucketEncryption.S3_MANAGED,
    block_public_access=s3.BlockPublicAccess.BLOCK_ALL,
    lifecycle_rules=[s3.LifecycleRule(prefix="raw/", expiration=Duration.days(90))])

Prefixes: raw/ (incoming exports), analytics/ (per-WO snapshots), athena-results/.

1.2 Direct upload path (always available):

aws s3 cp ./Sheet1-1.xlsx s3://apm-wo-analysis-exports-328440206208/raw/

1.3 Local drop-folder path (optional zero-touch). Mirror stampli-drop-folder. The script and the watched folder must live OUTSIDE ~/Documents (TCC sandbox — see macos-tcc-launchd memory). Suggested: folder ~/APM-WO-Drop, script ~/Library/Application Support/seahaven/apm-wo-drop/upload.sh, plist ~/Library/LaunchAgents/com.seahaven.apm-wo-drop.plist.

  • upload.sh: on a new file in the drop folder, aws s3 cp it to raw/, then move it to a local processed/ subfolder.
  • plist: WatchPaths = the drop folder; ProgramArguments = the script. launchctl load it.
  • Uses a least-privilege local IAM user/profile with s3:PutObject to raw/ only.

Validate: drop or upload a real export; confirm the object lands in raw/.


Phase 2 — Classifier Lambda

2.1 lambdas/classifier/classify.py — port the two-axis model exactly as specified in CLAUDE.md:

  • strip_html(text) — unwrap <html>…</html> and entities.
  • comment_intent(text) — ordered rule ladder (most-specific first), returns a bucket or None.
  • HOLD_TO_CAT map + WO Status rules (RCAN→Cancelled, etc.).
  • classify(status, hold, comment) → (final_category, mismatch_reason | None) with resolution policy comment-intent → structured → Other.
  • Haiku fallback (Secrets Manager apm-wo-analysis/anthropic-api-key) invoked only for Other rows with blank Hold Reason and a non-trivial comment.

2.2 lambdas/classifier/handler.py — S3-triggered on raw/:

  1. Read the object, parse with openpyxl (xlsx) / csv.
  2. For each row: strip HTML, classify(), derive is_escalation, is_action, mismatch.
  3. Write a per-WO snapshot to analytics/dt=YYYY-MM-DD/ as Parquet via awswrangler.s3.to_parquet(..., dataset=True, partition_cols=["dt"], database="apm_wo_analysis", table="apm_wo_snapshots") (also registers the Glue partition).
  4. Emit a small analytics/dt=YYYY-MM-DD/summary.json (counts per category, escalation total, action/routine, top sites, mismatch list) for the Slack Lambda to read cheaply.
  5. On completion, async-invoke the slack-post Lambda (or fire an EventBridge event).

2.3 Runtime: Python 3.12, ARM64, 512 MB, 120 s. Layer: awswrangler (AWS SDK for pandas) ARM64 managed layer. IAM: read raw/, write analytics/, read the Anthropic secret, Glue CreatePartition/BatchCreatePartition.

2.4 Cross-review (MANDATORY): the handler signature + IAM policy diff go through cross_reviewer before merge (orchestrator archived; cross_reviewer now lives in security-review):

python3 ~/Documents/repositories/seahaven/security-review/cross_review.py "Review this diff for breaking changes: <classifier IAM policy + handler contract>"

Validate (smoke test): pytest tests/test_classify.py runs classify() over the sample export and asserts "Other" ≤ ~10% and that known fixtures land in the right buckets. Use the classifier-engineer agent for tuning.


Phase 3 — Analytics dataset (Glue + Athena)

3.1 Glue database apm_wo_analysis (CDK glue.CfnDatabase).

3.2 Table apm_wo_snapshots over s3://apm-wo-analysis-exports-328440206208/analytics/, columns matching the snapshot (wo_number, description, equipment_code, site, due_date, department, wo_status, hold_reason, last_comment, last_comment_by, last_comment_date, contractor, category, is_escalation, is_action, mismatch), partitioned by dt with partition projection enabled:

projection.enabled = true
projection.dt.type = date
projection.dt.format = yyyy-MM-dd
projection.dt.range = 2026-01-01,NOW
storage.location.template = s3://.../analytics/dt=${dt}/

Projection means no crawler and no MSCK REPAIR.

3.3 Athena workgroup apm-wo-analysis with result location s3://.../athena-results/ and result encryption.

Validate:

SELECT category, count(*) FROM apm_wo_analysis.apm_wo_snapshots
WHERE dt = current_date GROUP BY 1 ORDER BY 2 DESC;

returns today's breakdown matching the smoke-test distribution.


Phase 4 — Slack post + alert Lambda

4.1 lambdas/slack_post/blockkit.py — build the daily summary and the batched 3rd-escalation alert (port the prototype). Daily post: header, context line with vs-yesterday deltas, escalation + action/routine fields, top sites, mismatch callout, action buttons incl. 📊 Open dashboard (url → https://grafana.seahaven.com/d/apm-wo/...?from=now-30d&to=now), footer. Alert: batched, one @here, return None / skip post when zero 3rd escalations.

4.2 lambdas/slack_post/handler.py — triggered after the classifier:

  1. Read analytics/dt=today/summary.json and dt=yesterday/summary.json (deltas).
  2. Post the daily summary to the WO channel via the reused bot token (SSM, same pattern as payments-dashboard).
  3. If 3rd-escalation count > 0, post the standalone alert; else post nothing.
  4. Drill-down block_actions (category/site buttons → views.open modal) handled by an API Gateway endpoint with Slack signature verification (or defer modals to v1.1 and rely on the Grafana table).

4.3 IAM: read analytics/, read the Slack token secret/param. Cross-review the IAM diff.

Validate: point at a test channel; confirm layout, real counts, correct deltas across two days of data, and zero-3rd suppression. Use the slack-blockkit-designer agent.


Phase 5 — Grafana (self-hosted EC2, VPN-only)

5.1 grafana_stack.py:

  • Import an existing VPC (or a small dedicated one). EC2 t4g.small (ARM64), Amazon Linux 2023, Grafana OSS installed via user-data, EBS gp3 with RETAIN.
  • Security group: ingress only from VPN / office CIDRs (see office-ips memory) on the ALB; ALB→instance on 3000.
  • Internal/again-restricted ALB with HTTPS using the wildcard ACM cert *.seahaven.com (ARN from context, same pattern as the slack bot). Target group → instance:3000.
  • Route53 A/alias grafana.seahaven.com → ALB in zone Z06652411XKH89KTZD3XA.
  • Instance role with Athena (StartQueryExecution, GetQueryResults), Glue (GetTable/GetPartitions), and S3 read on the analytics + athena-results prefixes. No static keys.

5.2 Provisioning (dashboards-as-code). Ship via user-data / config-sync into /etc/grafana/provisioning/:

  • datasources/athena.yaml — Athena datasource using the instance role (default auth provider), workgroup apm-wo-analysis, database apm_wo_analysis.
  • dashboards/apm.yaml — provider pointing at the dashboards folder.
  • dashboards/apm-work-orders.json — the committed dashboard. Use the grafana-author agent to build:
    • category distribution (bar), escalation summary (stat/pie), action vs routine (donut), escalation-by-site (bar/table);
    • trend time-series over dt;
    • filterable WO table with $site/$department/$category/$status/$hold_reason template variables, WO-number cell data-links into APM, escalation-row coloring, CSV export;
    • mismatch panel; chart-to-table drill via data links.
  • Panel refresh hourly (data changes once/day). Kiosk URL for the wall display.

5.3 Backup. Snapshot/version the dashboard JSON in-repo (source of truth) and back up grafana.db (or use the SQLite on the retained EBS volume). Document restore.

Validate: dashboard loads from Athena over VPN; filters work live; the Slack 📊 button deep-links correctly.


Phase 6 — Docs, Confluence, memory

  • README (sh-readme): architecture, data flow, the classification model, ingestion (direct + drop-folder), and an ops runbook (sh-runbook): how the export gets uploaded, Grafana patching cadence, dashboard-JSON redeploy, EBS/config backup-restore, common failures.
  • Confluence (sh-confluence): add a Mermaid subgraph for this stack to the "AWS Architecture Map" (page 1540098). If Confluence is unreachable, state the doc update as outstanding.
  • Memory (sh-distill): update apm-wo-comment-analysis to "deployed" with final resource names; ensure the cross-links to procurement-ingest, payments-dashboard, stampli-drop-folder, and macos-tcc-launchd are present.
  • Pre-merge / pre-prod: sh-pr-check on the branch, sh-prod-ready before the first real deploy.

Deploy order (once code is in)

# 0. deploy role already exists (Phase 0.3)
cd cdk && pip install -r requirements.txt
cdk deploy apm-wo-analysis-pipeline      # S3, Glue, Athena, Lambdas, IAM
# upload one export, confirm analytics/ partition + summary.json + Slack post
cdk deploy apm-wo-analysis-grafana       # EC2, ALB, SG, Route53, datasource role
# load the dashboard, verify Athena queries and the kiosk view

Cost

S3/Athena/Lambda ≈ pennies/month at ~350 rows/day; Grafana t4g.small ≈ $12/mo + small EBS + ALB hours. No per-user fees.

Inputs still required before building

  1. APM WO deep-link URL format (for the table cell links + Slack buttons).
  2. Slack channel + which bot token (reuse payments-dashboard app vs new).
  3. Confirm grafana.seahaven.com and VPN/office CIDRs for the SG.
  4. Confirm the drop-folder path conventions if using the launchd uploader.