No description
Find a file
Adam Moussa f13d2e3ff1
Some checks are pending
Deploy / Deploy to prod (push) Waiting to run
feat(infra): migrate pipeline and Grafana to HCP Terraform (PLAT-75) (#57)
* feat(infra): migrate pipeline and Grafana to HCP Terraform (PLAT-75)

Move apm-wo-analysis into seahaven-prod under workspace apm-wo-analysis-prod
with in-repo hcptf/githubdeploy IAM, stub Lambdas, and GitHub Actions zip CD.

* chore(iam): add Checkov skip comments for HCP IAM documents

Pre-push HIGH findings are the DLM snapshot describe, tagged EC2 creates,
exec boundary DescribeLogGroups star, and the drop-uploader user policy.
2026-09-16 20:33:24 +00:00
.claude/agents docs: repoint cross-family review to security-review cross_review.py (orchestrator archived) (#31) 2026-07-14 19:24:08 -04:00
.github feat(infra): migrate pipeline and Grafana to HCP Terraform (PLAT-75) (#57) 2026-09-16 20:33:24 +00:00
cdk chore(deps): bump aws-cdk-lib in /cdk in the minor-and-patch group (#49) 2026-08-20 14:59:41 -04:00
docs feat(infra): migrate pipeline and Grafana to HCP Terraform (PLAT-75) (#57) 2026-09-16 20:33:24 +00:00
grafana feat(infra): migrate pipeline and Grafana to HCP Terraform (PLAT-75) (#57) 2026-09-16 20:33:24 +00:00
lambdas chore: resolve open code scanning alerts (URL host matching + workflow permissions) (#36) 2026-07-23 19:10:45 +00:00
scripts feat(infra): migrate pipeline and Grafana to HCP Terraform (PLAT-75) (#57) 2026-09-16 20:33:24 +00:00
slack Add Slack post + interactions Lambdas with drill-down modals (Phase 4) 2026-05-28 17:48:51 -04:00
terraform feat(infra): migrate pipeline and Grafana to HCP Terraform (PLAT-75) (#57) 2026-09-16 20:33:24 +00:00
tests feat(infra): migrate pipeline and Grafana to HCP Terraform (PLAT-75) (#57) 2026-09-16 20:33:24 +00:00
.gitignore feat(infra): migrate pipeline and Grafana to HCP Terraform (PLAT-75) (#57) 2026-09-16 20:33:24 +00:00
.mergify.yml chore(ci): switch auto-merge from seahaven-bot to Mergify (#53) 2026-08-24 13:53:44 -04:00
AGENTS.md ci: add org PR policy caller (PLAT-62) (#40) 2026-08-04 15:58:00 +00:00
CLAUDE.md feat(infra): migrate pipeline and Grafana to HCP Terraform (PLAT-75) (#57) 2026-09-16 20:33:24 +00:00
pyproject.toml feat(infra): migrate pipeline and Grafana to HCP Terraform (PLAT-75) (#57) 2026-09-16 20:33:24 +00:00
README.md feat(infra): migrate pipeline and Grafana to HCP Terraform (PLAT-75) (#57) 2026-09-16 20:33:24 +00:00
renovate.json feat(infra): migrate pipeline and Grafana to HCP Terraform (PLAT-75) (#57) 2026-09-16 20:33:24 +00:00

apm-wo-analysis

Python Terraform Slack CI

Daily analysis of Amazon APM work-order "Last Comment" data for Sea Haven facility ops. A curated daily filter-view export (~350 work orders) is classified on two axes (comment intent + structured WO Status/Hold Reason), pushed to Slack as a daily summary plus a batched 3rd-escalation alert, and surfaced in a self-hosted Grafana dashboard. Replaces a legacy Google Apps Script + versioned-Google-Sheet workflow.

This is a sibling concern to the apm@ email pipeline in procurement-ingest — it consumes a different feed (the manual export) and does not read those tables. There is deliberately no DynamoDB: this is an analytics workload backed by S3 + Athena (Grafana cannot query DynamoDB).

  • Account / region: seahaven-prod 011934824531 / us-east-1 (PLAT-75). Mgmt CDK stacks apm-wo-analysis-pipeline and apm-wo-analysis-grafana deleted 2026-09-16.
  • IaC: HCP Terraform workspace apm-wo-analysis-prod (working directory terraform/, file trigger terraform/**). Lambdas Python 3.12, ARM64. GitHub Actions owns zip + Grafana-config content.

Architecture

APM export (xlsx/csv)
  → S3 raw/  (direct `aws s3 cp` OR local launchd drop-folder)
      → classifier Lambda (HTML strip + two-axis classify, Haiku fallback)
          ├→ S3 analytics/dt=YYYY-MM-DD/  (per-WO snapshot, Parquet)
          │     → Glue table (partition projection) → Athena → Grafana (EC2, office-IP, kiosk)
          ├→ S3 meta/dt=YYYY-MM-DD/  (summary.json + details.json — NOT in the table prefix)
          └→ async-invoke slack-post Lambda
                → daily summary post  [📊 Open dashboard button]
                → standalone batched 3rd-escalation alert (suppressed when zero)
                → drill-down modals via apm-wo.seahaven.com (signature-verified)

One HCP workspace (apm-wo-analysis-prod) owns both the pipeline and Grafana:

Area Resources
Pipeline S3 exports + artifacts buckets, drop-uploader IAM user, classifier + slack-post + slack-interactions Lambdas (stub + ignore_changes), classifier DLQ, Glue DB + projection table, Athena workgroup, HTTP API (apm-wo.seahaven.com), SSM deploy contract, in-repo hcptf-* / githubdeploy-* IAM
Grafana Dedicated VPC (private instance, public ALB, one NAT), EC2 (Grafana OSS), office-IP-restricted ALB (grafana.seahaven.com), instance IAM role, daily DLM snapshot. Dashboards sync from S3 via GitHub Actions. Route53 aliases stay in the mgmt hosted zone and are flipped out of band.

AWS Resources

Resource Name Purpose
S3 bucket apm-wo-analysis-exports-011934824531 Single data bucket. Prefixes: raw/ (incoming, 90-day expiry), analytics/ (Parquet snapshots, kept), meta/ (summary/details JSON), athena-results/ (query output, 30-day expiry), grafana-config/ (dashboards-as-code). SSE-S3, BPA-all, enforce-SSL.
S3 bucket apm-wo-analysis-artifacts-011934824531 Lambda zip artifacts. GitHub Actions uploads functions/<name>/<sha>.zip.
IAM user apm-wo-drop-uploader Drop-folder identity; s3:PutObject on raw/* only. Access key created out-of-band, stored in local apm-wo-drop profile.
Glue database apm_wo_analysis Analytics catalog.
Glue table apm_wo_snapshots External Parquet table over analytics/, partition projection on dt (date, 2026-01-01..NOW) — no crawler, no MSCK. 17-column snapshot schema.
Athena workgroup apm-wo-analysis Enforced result location athena-results/, SSE-S3.
SQS queue apm-wo-analysis-classifier-dlq Dead-letter for failed classifier async invocations (14-day retention).
SSM parameter /apm-wo-analysis/grafana-base-url Grafana dashboard URL for the 📊 button / modal links (ops-editable, no redeploy).
HTTP API + domain apm-wo.seahaven.com → POST /slack/interactions Slack interactivity endpoint. Stage throttled 10 rps / 20 burst. ACM cert in seahaven-prod; Route53 alias in mgmt zone (OOB).
EC2 instance Grafana (t4g.small, AL2023, ARM64) Self-hosted Grafana OSS in a dedicated VPC private subnet, IMDSv2-only, SSM-managed. gp3 20 GB encrypted, DeleteOnTermination=false, tagged apm-grafana-backup.
ALB Grafana ALB (grafana.seahaven.com) Internet-facing, HTTPS 443, SG admits only office CIDRs (47.21.61.4/32, 96.250.164.146/32); forwards to instance:3000, health /api/health.
DLM policy Grafana volume backup Daily snapshot (07:00 UTC) of the tagged instance, 7 retained.
Route53 apm-wo.seahaven.com, grafana.seahaven.com Aliases in zone Z06652411XKH89KTZD3XA (seahaven.com).

Lambda Functions

All Python 3.12, ARM64, explicit LogGroup (/aws/lambda/<name>, 60-day retention).

Function Trigger Purpose
apm-wo-analysis-classifier S3 ObjectCreated on `raw/*.xlsx .csv`
apm-wo-analysis-slack-post Async-invoked by the classifier ({"dt": …}) Read today's + yesterday's meta/.../summary.json, post the daily summary, and (only when third_escalation_count > 0) the batched 3rd-escalation alert from details.json. 256 MB / 30 s.
apm-wo-analysis-slack-interactions HTTP API POST /slack/interactions Verify the Slack request signature, read meta/.../details.json, and views.open a filtered WO-list modal within Slack's 3 s trigger_id window. 256 MB / 30 s.

Configuration

Secrets Manager (names only — Terraform creates empty shells; values copied out of band)

Secret Purpose
apm-wo-analysis/anthropic-api-key Claude Haiku fallback for ambiguous free-text comments.
apm-wo-analysis/slack-credentials JSON { botToken, signingSecret, channelId } for the reused Slack app.

SSM Parameters

Parameter Purpose
/apm-wo-analysis/grafana-base-url Grafana dashboard deep-link base (https://grafana.seahaven.com/d/apm-wo/...).

Environment Variables (non-secret)

  • classifier: APM_HAIKU_FALLBACK (on/off), SLACK_POST_FUNCTION_NAME.
  • slack-post / slack-interactions: SLACK_SECRET_NAME, DASHBOARD_URL_PARAM, ANALYTICS_BUCKET.

GitHub Environment prod

  • DEPLOY_ROLE_ARN — the OIDC deploy role githubdeploy-apm-wo-analysis (/tf-managed/).

HCP workspace

  • apm-wo-analysis-prod in project seahaven-prod. VCS main, working directory terraform, file trigger terraform/**. TFC_AWS_APPLY_ROLE_ARN / TFC_AWS_PLAN_ROLE_ARN are workspace vars pointing at hcptf-apm-wo-analysis / hcptf-apm-wo-analysis-plan after the bootstrap window.

The classification model

Always two-axis, never comment-only. The legacy script's central flaw was reading only the comment while ignoring WO Status + Hold Reason, which left ~17% in "Other". The two-axis model cuts that to ~9% before any AI — and on the real 347-row export the implementation lands "Other" at 5.2% (18 rows) with 18 mismatches flagged.

  • Axis 1 — comment intent: regex over the HTML-stripped Last Comment, most-specific first (escalations → SIM ticket → vendor no-show → scheduling → reports → completion → … → other).
  • Axis 2 — structured state: Hold Reason → category, and WO Status signals (RCAN→Cancelled, H corroborates On Hold, IP/R/RR in-flight).
  • Resolution: comment intent wins when confident → else structured state → else Other. A Claude Haiku fallback (Secrets Manager key) is reserved for ambiguous free-text with no structured signal.
  • Mismatch detector (a feature): flags when comment intent contradicts structured state (e.g. "completed" while WO Status is IP). Surfaced, never suppressed.

Authoritative spec: CLAUDE.md. Implementation: lambdas/classifier/classify.py (logic) and lambdas/classifier/handler.py (S3 → Parquet + JSON + invoke).

Repository layout

terraform/               HCP Terraform: IAM, S3, Glue, Athena, Lambdas, API, VPC, Grafana
  bootstrap/stub/        committed Lambda stub; GHA replaces code via update-function-code
  templates/             Grafana user-data
cdk/                     frozen mgmt CDK source until the mgmt stacks are deleted
lambdas/
  classifier/            S3-triggered: parse → two-axis classify → Parquet + meta JSON
  slack_post/            blockkit.py (builders), handler.py (post), interactions.py
                         (modals), slackio.py (Secrets/SSM/S3/signature)
grafana/
  provisioning/          Athena datasource + dashboard provider (as code)
  dashboards/            apm-work-orders.json (uid apm-wo) — source of truth
slack/manifest.yaml      Slack app manifest (interactivity request URL)
scripts/                 drop-folder uploader, launchd plist, package_lambdas.py
tests/                   classifier smoke test + Block Kit / handler assertions
docs/BUILD.md            original CDK build guide (historical)
docs/RUNBOOK.md          operations

Ingestion (no email)

The export reaches S3 by direct upload or a local drop-folder, never SES/email.

  • Direct: aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-011934824531/raw/

  • Drop-folder (optional zero-touch): a launchd agent (scripts/apm-wo-uploader.sh

    • scripts/com.seahaven.apm-wo-uploader.plist) that watches ~/apm-wo-drop/, uploads new .xlsx/.csv files to raw/, and archives them locally. Uploads with the scoped apm-wo-drop profile (IAM user apm-wo-drop-uploader).

    Install — the runnable copy and watched folder must live outside ~/Documents (macOS TCC sandbox; a repo-path script fails silently with LastExitStatus=32256):

    install -d "$HOME/.local/bin" "$HOME/apm-wo-drop"
    cp scripts/apm-wo-uploader.sh "$HOME/.local/bin/apm-wo-uploader.sh"
    chmod +x "$HOME/.local/bin/apm-wo-uploader.sh"
    cp scripts/com.seahaven.apm-wo-uploader.plist "$HOME/Library/LaunchAgents/"
    launchctl load -w "$HOME/Library/LaunchAgents/com.seahaven.apm-wo-uploader.plist"
    aws configure --profile apm-wo-drop    # one-time, with the uploader's access key
    

Any .xlsx/.csv landing under raw/ invokes the classifier.

Local development

  • pyenv Python 3.12. ruff check + ruff format --check before pushing (hook-enforced).
  • Tests: python -m pytest tests/ -q (classifier rule ladder + handler transforms + Haiku fallback, Block Kit builders, Slack interactions, slack-post orchestration). No AWS needed — external boundaries are monkeypatched. Tests run in CI. Coverage via pytest-cov (tests/requirements.txt, non-gating).
  • The classification quality gate (deterministic "Other" share) runs in CI against a committed synthetic fixture tests/fixtures/sample_export.csv. Additionally, smoke-test against the real export before declaring any classification change done: ~/Downloads/_documents/Sheet1-1.xlsx (skips automatically when absent).
  • terraform fmt -check -recursive and terraform validate must pass in CI (terraform init -backend=false).

Deployment

Terraform owns containers. GitHub Actions owns Lambda zips and Grafana JSON.

  • CI (.github/workflows/ci.yaml): pytest + terraform fmt + terraform validate.
  • Content CD (.github/workflows/deploy.yaml): Environment prod, OIDC githubdeploy-apm-wo-analysis, update-function-code + aws s3 sync grafana/. paths-ignore for terraform/**. Never creates an HCP run.
  • Infra: HCP workspace apm-wo-analysis-prod, VCS on main, working directory terraform, file trigger terraform/**. First apply is Manual via the hcptf-bootstrap window; after seal, TFC_AWS_* point at hcptf-apm-wo-analysis / hcptf-apm-wo-analysis-plan.

Account/region: 011934824531 / us-east-1.

Lambda code seam: committed stub under terraform/bootstrap/ plus lifecycle.ignore_changes on filename/s3_key/source_code_hash. A post-deploy plan must be empty.

Operations

Verify it's working — drop a real export into raw/, then within ~1 min:

  • analytics/dt=<today>/*.parquet and meta/dt=<today>/{summary,details}.json appear,
  • the daily summary posts to the WO Slack channel (+ alert if any 3rd escalations),
  • the Grafana dashboard renders over the office network (grafana.seahaven.com).

Logs: CloudWatch /aws/lambda/apm-wo-analysis-{classifier,slack-post,slack-interactions} (60-day retention).

Failure modes:

  • Classifier failure (malformed export, transient error) → after Lambda retries, the event lands in apm-wo-analysis-classifier-dlq. Check the DLQ if a day's data is missing.
  • Slack post failure is best-effort and does not fail classification (data still lands in S3).
  • Grafana down → check the instance via SSM Session Manager (no SSH); systemctl status grafana-server; ALB target health.

Reprocess a day: re-upload the same export to raw/ — the classifier uses overwrite_partitions, so a same-day re-run replaces that dt partition idempotently.

Grafana admin: access is office-IP-restricted at the ALB; the instance is SSM-only. Dashboards are provisioned from grafana-config/ in S3 (synced on boot and by a 15-min systemd timer); edit dashboards in-repo, not in the UI (allowUiUpdates: false). grafana.db lives on the RETAIN'd gp3 volume and is snapshotted daily by DLM.

Documentation

  • Confluence "AWS Architecture Map" (IT space, page 1540098): a Mermaid subgraph for this stack — done (added under Serverless Applications, plus rows in CI/CD Pipelines, EC2 Inventory, and Key Data Stores).
  • Confluence stack page "APM WO Analysis (apm-wo-analysis)" (IT space, page 8749057, under AWS Cloud Infrastructure): full resource/Lambda/secrets detail — done.
  • Slack Apps Inventory (Confluence page 524569): "APM Work Orders" (new dedicated app, App ID A0B6P28V64B) added — done.
  • This README + docs/BUILD.md (phased build guide) + CLAUDE.md (domain spec).

Notes / Gotchas

  • Partition date comes from the S3 event time, not the Lambda wall-clock — stable across retries and the midnight boundary.
  • meta/ vs analytics/: summary/details JSON must stay out of analytics/ — Athena reads every object in the table prefix as Parquet and chokes on JSON.
  • Classifier deps (awswrangler/pandas/pyarrow/numpy) come from the AWS-managed SDK-for-pandas layer — bundling them blows Lambda's 250 MB unzipped limit. Verify the pinned layer ARN/version on region or runtime changes.
  • Grafana dashboard contract: dashboard uid must stay apm-wo (the Slack 📊 button deep-links to d/apm-wo). Athena plugin query key is rawSQL (capital). Datasource authType: default (the EC2 instance role; ec2_iam_role is rejected by the plugin). Config sync uses aws s3 sync --exact-timestamps (plain sync skips same-size edits). Template vars use refresh: 1 (on load).
  • No Client VPN exists — "VPN-only" Grafana is realized as office-IP SG restriction. Grafana sits in a dedicated VPC with a single NAT (one AZ) for instance egress.
  • Slack interactions endpoint is unauthenticated at the gateway by design; the Lambda verifies the Slack signature (replay window + HMAC). Stage-throttled.
  • Route53 for apm-wo.seahaven.com and grafana.seahaven.com stays in the mgmt hosted zone Z06652411XKH89KTZD3XA and is flipped out of band.

Known operational debt

  • HCP auto-apply stays off until the scoped plan is clean and the bootstrap trust window is closed.
  • Deferred review NITs: print()→logging, source-IP logging on signature failure, S3 versioning, the '${site:raw}' WO-table SQL tidy, and a CIDR-maintenance note in the runbook.

Status

Live in seahaven-prod under HCP workspace apm-wo-analysis-prod (PLAT-75, 2026-09-16). Classification logic is unchanged. Mgmt CDK stacks apm-wo-analysis-pipeline and apm-wo-analysis-grafana are deleted; the mgmt exports bucket is RETAIN cold archive. GitHub Actions deploy.yaml on main (Environment prod) owns Lambda zips and grafana-config/ sync. HCP auto-apply stays off until the scoped plan is clean.