13 KiB
apm-wo-analysis — Step-by-Step Build Guide
End-to-end build instructions. Read ../CLAUDE.md first for the domain model and locked decisions. Account 328440206208, region us-east-1, all names kebab-case.
Convention gates that apply throughout: OIDC deploy role created before any CD; secrets in Secrets Manager;
cross_revieweron IAM/handler diffs (orchestrator since archived; usecross_review.pyinsecurity-review);ruffclean +cdk synthgreen before push; README + Confluence + memory updated as part of the work, not after.
Target repo layout
apm-wo-analysis/
├── CLAUDE.md
├── README.md
├── cdk/
│ ├── app.py
│ ├── cdk.json
│ ├── requirements.txt # aws-cdk-lib==2.253.1, constructs>=10.6.0
│ └── stacks/
│ ├── pipeline_stack.py # S3, Lambdas, Glue, Athena, IAM, schedule
│ └── grafana_stack.py # VPC import, EC2, ALB, SG, Route53, datasource role
├── lambdas/
│ ├── classifier/ # S3-triggered: parse → classify → write parquet
│ │ ├── handler.py
│ │ ├── classify.py # the two-axis model (the core logic)
│ │ └── requirements.txt # awswrangler, openpyxl, anthropic
│ └── slack_post/ # builds + posts daily summary and alert
│ ├── handler.py
│ └── blockkit.py
├── grafana/
│ ├── provisioning/
│ │ ├── datasources/athena.yaml
│ │ └── dashboards/apm.yaml
│ └── dashboards/apm-work-orders.json
├── scripts/
│ ├── drop_folder_upload.sh # local launchd uploader (lives in repo, deployed outside ~/Documents)
│ └── com.seahaven.apm-wo-drop.plist
├── tests/
│ └── test_classify.py # smoke test against a sample export
└── .github/
├── workflows/{ci.yaml,deploy.yaml}
└── dependabot.yml
Phase 0 — Repo provisioning
0.1 Lock reuse patterns first. Run the Explore agent over payments-dashboard (the payments-slackAppHome Lambda + how its Slack bot token is read from SSM) and the org .github reusable workflows (ci-python-sam.yaml, cd-cdk.yaml) so the new repo matches them exactly. Optionally scaffold with the sh-bootstrap skill.
0.2 Create the repo.
gh repo create Sea-Haven-Industries/apm-wo-analysis --private
git init && git branch -M main
git remote add origin git@github.com:Sea-Haven-Industries/apm-wo-analysis.git
0.3 OIDC deploy role FIRST (before any CD). Create githubdeploy-apm-wo-analysis, trust scoped to repo:Sea-Haven-Industries/apm-wo-analysis:*, with permissions to deploy the two stacks (CloudFormation, plus the resource services they create). Mirror the trust policy from githubdeploy-procurement-ingest. Store the ARN as the repo secret AWS_DEPLOY_ROLE_ARN.
0.4 CDK scaffold.
mkdir -p cdk/stacks lambdas tests grafana
printf 'aws-cdk-lib==2.253.1\nconstructs>=10.6.0\n' > cdk/requirements.txt
cdk/app.py instantiates PipelineStack and GrafanaStack, both pinned to env 328440206208/us-east-1.
0.5 CI/CD + Dependabot. .github/workflows/ci.yaml calls the reusable ci-python-sam.yaml@main (lint lambdas cdk + cdk synth); deploy.yaml calls cd-cdk.yaml@main (Python 3.12, cdk dir cdk, OIDC, cdk deploy --all, concurrency single). Add dependabot.yml (pip for cdk/ and each lambdas/*, github-actions).
0.6 Branch protection on main: require the CI check + the Claude Code App review.
Validate: open a throwaway PR with the empty scaffold; CI cdk synth is green.
Phase 1 — Ingestion (direct S3 / local drop-folder, NO email)
1.1 Bucket in pipeline_stack.py:
bucket = s3.Bucket(self, "Exports",
bucket_name=f"apm-wo-analysis-exports-{self.account}",
removal_policy=RETAIN, encryption=s3.BucketEncryption.S3_MANAGED,
block_public_access=s3.BlockPublicAccess.BLOCK_ALL,
lifecycle_rules=[s3.LifecycleRule(prefix="raw/", expiration=Duration.days(90))])
Prefixes: raw/ (incoming exports), analytics/ (per-WO snapshots), athena-results/.
1.2 Direct upload path (always available):
aws s3 cp ./Sheet1-1.xlsx s3://apm-wo-analysis-exports-328440206208/raw/
1.3 Local drop-folder path (optional zero-touch). Mirror stampli-drop-folder. The script and the watched folder must live OUTSIDE ~/Documents (TCC sandbox — see macos-tcc-launchd memory). Suggested: folder ~/APM-WO-Drop, script ~/Library/Application Support/seahaven/apm-wo-drop/upload.sh, plist ~/Library/LaunchAgents/com.seahaven.apm-wo-drop.plist.
upload.sh: on a new file in the drop folder,aws s3 cpit toraw/, then move it to a localprocessed/subfolder.- plist:
WatchPaths= the drop folder;ProgramArguments= the script.launchctl loadit. - Uses a least-privilege local IAM user/profile with
s3:PutObjecttoraw/only.
Validate: drop or upload a real export; confirm the object lands in raw/.
Phase 2 — Classifier Lambda
2.1 lambdas/classifier/classify.py — port the two-axis model exactly as specified in CLAUDE.md:
strip_html(text)— unwrap<html>…</html>and entities.comment_intent(text)— ordered rule ladder (most-specific first), returns a bucket orNone.HOLD_TO_CATmap +WO Statusrules (RCAN→Cancelled, etc.).classify(status, hold, comment)→(final_category, mismatch_reason | None)with resolution policy comment-intent → structured →Other.- Haiku fallback (Secrets Manager
apm-wo-analysis/anthropic-api-key) invoked only forOtherrows with blank Hold Reason and a non-trivial comment.
2.2 lambdas/classifier/handler.py — S3-triggered on raw/:
- Read the object, parse with
openpyxl(xlsx) / csv. - For each row: strip HTML,
classify(), deriveis_escalation,is_action,mismatch. - Write a per-WO snapshot to
analytics/dt=YYYY-MM-DD/as Parquet viaawswrangler.s3.to_parquet(..., dataset=True, partition_cols=["dt"], database="apm_wo_analysis", table="apm_wo_snapshots")(also registers the Glue partition). - Emit a small
analytics/dt=YYYY-MM-DD/summary.json(counts per category, escalation total, action/routine, top sites, mismatch list) for the Slack Lambda to read cheaply. - On completion, async-invoke the slack-post Lambda (or fire an EventBridge event).
2.3 Runtime: Python 3.12, ARM64, 512 MB, 120 s. Layer: awswrangler (AWS SDK for pandas) ARM64 managed layer. IAM: read raw/, write analytics/, read the Anthropic secret, Glue CreatePartition/BatchCreatePartition.
2.4 Cross-review (MANDATORY): the handler signature + IAM policy diff go through cross_reviewer before merge (orchestrator archived; cross_reviewer now lives in security-review):
python3 ~/Documents/repositories/seahaven/security-review/cross_review.py "Review this diff for breaking changes: <classifier IAM policy + handler contract>"
Validate (smoke test): pytest tests/test_classify.py runs classify() over the sample export and asserts "Other" ≤ ~10% and that known fixtures land in the right buckets. Use the classifier-engineer agent for tuning.
Phase 3 — Analytics dataset (Glue + Athena)
3.1 Glue database apm_wo_analysis (CDK glue.CfnDatabase).
3.2 Table apm_wo_snapshots over s3://apm-wo-analysis-exports-328440206208/analytics/, columns matching the snapshot (wo_number, description, equipment_code, site, due_date, department, wo_status, hold_reason, last_comment, last_comment_by, last_comment_date, contractor, category, is_escalation, is_action, mismatch), partitioned by dt with partition projection enabled:
projection.enabled = true
projection.dt.type = date
projection.dt.format = yyyy-MM-dd
projection.dt.range = 2026-01-01,NOW
storage.location.template = s3://.../analytics/dt=${dt}/
Projection means no crawler and no MSCK REPAIR.
3.3 Athena workgroup apm-wo-analysis with result location s3://.../athena-results/ and result encryption.
Validate:
SELECT category, count(*) FROM apm_wo_analysis.apm_wo_snapshots
WHERE dt = current_date GROUP BY 1 ORDER BY 2 DESC;
returns today's breakdown matching the smoke-test distribution.
Phase 4 — Slack post + alert Lambda
4.1 lambdas/slack_post/blockkit.py — build the daily summary and the batched 3rd-escalation alert (port the prototype). Daily post: header, context line with vs-yesterday deltas, escalation + action/routine fields, top sites, mismatch callout, action buttons incl. 📊 Open dashboard (url → https://grafana.seahaven.com/d/apm-wo/...?from=now-30d&to=now), footer. Alert: batched, one @here, return None / skip post when zero 3rd escalations.
4.2 lambdas/slack_post/handler.py — triggered after the classifier:
- Read
analytics/dt=today/summary.jsonanddt=yesterday/summary.json(deltas). - Post the daily summary to the WO channel via the reused bot token (SSM, same pattern as
payments-dashboard). - If 3rd-escalation count > 0, post the standalone alert; else post nothing.
- Drill-down
block_actions(category/site buttons →views.openmodal) handled by an API Gateway endpoint with Slack signature verification (or defer modals to v1.1 and rely on the Grafana table).
4.3 IAM: read analytics/, read the Slack token secret/param. Cross-review the IAM diff.
Validate: point at a test channel; confirm layout, real counts, correct deltas across two days of data, and zero-3rd suppression. Use the slack-blockkit-designer agent.
Phase 5 — Grafana (self-hosted EC2, VPN-only)
5.1 grafana_stack.py:
- Import an existing VPC (or a small dedicated one). EC2
t4g.small(ARM64), Amazon Linux 2023, Grafana OSS installed via user-data, EBS gp3 withRETAIN. - Security group: ingress only from VPN / office CIDRs (see
office-ipsmemory) on the ALB; ALB→instance on 3000. - Internal/again-restricted ALB with HTTPS using the wildcard ACM cert
*.seahaven.com(ARN from context, same pattern as the slack bot). Target group → instance:3000. - Route53 A/alias
grafana.seahaven.com→ ALB in zoneZ06652411XKH89KTZD3XA. - Instance role with Athena (
StartQueryExecution,GetQueryResults), Glue (GetTable/GetPartitions), and S3 read on the analytics + athena-results prefixes. No static keys.
5.2 Provisioning (dashboards-as-code). Ship via user-data / config-sync into /etc/grafana/provisioning/:
datasources/athena.yaml— Athena datasource using the instance role (default auth provider), workgroupapm-wo-analysis, databaseapm_wo_analysis.dashboards/apm.yaml— provider pointing at the dashboards folder.dashboards/apm-work-orders.json— the committed dashboard. Use thegrafana-authoragent to build:- category distribution (bar), escalation summary (stat/pie), action vs routine (donut), escalation-by-site (bar/table);
- trend time-series over
dt; - filterable WO table with
$site/$department/$category/$status/$hold_reasontemplate variables, WO-number cell data-links into APM, escalation-row coloring, CSV export; - mismatch panel; chart-to-table drill via data links.
- Panel refresh hourly (data changes once/day). Kiosk URL for the wall display.
5.3 Backup. Snapshot/version the dashboard JSON in-repo (source of truth) and back up grafana.db (or use the SQLite on the retained EBS volume). Document restore.
Validate: dashboard loads from Athena over VPN; filters work live; the Slack 📊 button deep-links correctly.
Phase 6 — Docs, Confluence, memory
- README (
sh-readme): architecture, data flow, the classification model, ingestion (direct + drop-folder), and an ops runbook (sh-runbook): how the export gets uploaded, Grafana patching cadence, dashboard-JSON redeploy, EBS/config backup-restore, common failures. - Confluence (
sh-confluence): add a Mermaid subgraph for this stack to the "AWS Architecture Map" (page 1540098). If Confluence is unreachable, state the doc update as outstanding. - Memory (
sh-distill): updateapm-wo-comment-analysisto "deployed" with final resource names; ensure the cross-links toprocurement-ingest,payments-dashboard,stampli-drop-folder, andmacos-tcc-launchdare present. - Pre-merge / pre-prod:
sh-pr-checkon the branch,sh-prod-readybefore the first real deploy.
Deploy order (once code is in)
# 0. deploy role already exists (Phase 0.3)
cd cdk && pip install -r requirements.txt
cdk deploy apm-wo-analysis-pipeline # S3, Glue, Athena, Lambdas, IAM
# upload one export, confirm analytics/ partition + summary.json + Slack post
cdk deploy apm-wo-analysis-grafana # EC2, ALB, SG, Route53, datasource role
# load the dashboard, verify Athena queries and the kiosk view
Cost
S3/Athena/Lambda ≈ pennies/month at ~350 rows/day; Grafana t4g.small ≈ $12/mo + small EBS + ALB hours. No per-user fees.
Inputs still required before building
- APM WO deep-link URL format (for the table cell links + Slack buttons).
- Slack channel + which bot token (reuse
payments-dashboardapp vs new). - Confirm
grafana.seahaven.comand VPN/office CIDRs for the SG. - Confirm the drop-folder path conventions if using the launchd uploader.