Bumps the minor-and-patch group in /cdk with 1 update: [aws-cdk-lib](https://github.com/aws/aws-cdk). Updates `aws-cdk-lib` from 2.263.0 to 2.264.0 - [Release notes](https://github.com/aws/aws-cdk/releases) - [Changelog](https://github.com/aws/aws-cdk/blob/main/CHANGELOG.v2.alpha.md) - [Commits](https://github.com/aws/aws-cdk/compare/v2.263.0...v2.264.0) --- updated-dependencies: - dependency-name: aws-cdk-lib dependency-version: 2.264.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: minor-and-patch ... Signed-off-by: dependabot[bot] <support@github.com> |
||
|---|---|---|
| .claude/agents | ||
| .github | ||
| cdk | ||
| docs | ||
| grafana | ||
| lambdas | ||
| scripts | ||
| slack | ||
| tests | ||
| .gitignore | ||
| AGENTS.md | ||
| CLAUDE.md | ||
| pyproject.toml | ||
| README.md | ||
apm-wo-analysis
Daily analysis of Amazon APM work-order "Last Comment" data for Sea Haven
facility ops. A curated daily filter-view export (~350 work orders) is classified
on two axes (comment intent + structured WO Status/Hold Reason), pushed to
Slack as a daily summary plus a batched 3rd-escalation alert, and surfaced in a
self-hosted Grafana dashboard. Replaces a legacy Google Apps Script +
versioned-Google-Sheet workflow.
This is a sibling concern to the apm@ email pipeline in procurement-ingest
— it consumes a different feed (the manual export) and does not read those
tables. There is deliberately no DynamoDB: this is an analytics workload
backed by S3 + Athena (Grafana cannot query DynamoDB).
- Account / region: 328440206208 / us-east-1
- IaC: CDK (Python),
aws-cdk-lib==2.253.1. Lambdas Python 3.12, ARM64.
Architecture
APM export (xlsx/csv)
→ S3 raw/ (direct `aws s3 cp` OR local launchd drop-folder)
→ classifier Lambda (HTML strip + two-axis classify, Haiku fallback)
├→ S3 analytics/dt=YYYY-MM-DD/ (per-WO snapshot, Parquet)
│ → Glue table (partition projection) → Athena → Grafana (EC2, office-IP, kiosk)
├→ S3 meta/dt=YYYY-MM-DD/ (summary.json + details.json — NOT in the table prefix)
└→ async-invoke slack-post Lambda
→ daily summary post [📊 Open dashboard button]
→ standalone batched 3rd-escalation alert (suppressed when zero)
→ drill-down modals via apm-wo.seahaven.com (signature-verified)
Two CDK stacks (cdk/app.py instantiates both):
| Stack | Resources |
|---|---|
apm-wo-analysis-pipeline |
S3 exports bucket, drop-uploader IAM user, classifier + slack-post + slack-interactions Lambdas, classifier DLQ, Glue DB + projection table, Athena workgroup, HTTP API (apm-wo.seahaven.com), SSM param, scoped IAM |
apm-wo-analysis-grafana |
EC2 (Grafana OSS), internet-facing office-IP-restricted ALB (grafana.seahaven.com), security groups, instance IAM role, Route53 alias, daily DLM snapshot, dashboards-as-code S3 deployment |
AWS Resources
| Resource | Name | Purpose |
|---|---|---|
| S3 bucket | apm-wo-analysis-exports-328440206208 |
Single bucket. Prefixes: raw/ (incoming, 90-day expiry), analytics/ (Parquet snapshots, kept), meta/ (summary/details JSON), athena-results/ (query output, 30-day expiry), grafana-config/ (dashboards-as-code). SSE-S3, BPA-all, enforce-SSL, RETAIN. |
| IAM user | apm-wo-drop-uploader |
Drop-folder identity; s3:PutObject on raw/* only. Access key created out-of-band, stored in local apm-wo-drop profile. |
| Glue database | apm_wo_analysis |
Analytics catalog. |
| Glue table | apm_wo_snapshots |
External Parquet table over analytics/, partition projection on dt (date, 2026-01-01..NOW) — no crawler, no MSCK. 17-column snapshot schema. |
| Athena workgroup | apm-wo-analysis |
Enforced result location athena-results/, SSE-S3. |
| SQS queue | apm-wo-analysis-classifier-dlq |
Dead-letter for failed classifier async invocations (14-day retention). |
| SSM parameter | /apm-wo-analysis/grafana-base-url |
Grafana dashboard URL for the 📊 button / modal links (ops-editable, no redeploy). |
| HTTP API + domain | apm-wo.seahaven.com → POST /slack/interactions |
Slack interactivity endpoint. Stage throttled 10 rps / 20 burst. *.seahaven.com ACM cert; Route53 alias. |
| EC2 instance | Grafana (t4g.small, AL2023, ARM64) |
Self-hosted Grafana OSS in seahaven-vpc private subnets, IMDSv2-only, SSM-managed. gp3 20 GB encrypted, DeleteOnTermination=false, tagged apm-grafana-backup. |
| ALB | Grafana ALB (grafana.seahaven.com) |
Internet-facing, HTTPS 443, SG admits only office CIDRs (47.21.61.4/32, 96.250.164.146/32); forwards to instance:3000, health /api/health. |
| DLM policy | Grafana volume backup | Daily snapshot (07:00 UTC) of the tagged instance, 7 retained. |
| Route53 | apm-wo.seahaven.com, grafana.seahaven.com |
Aliases in zone Z06652411XKH89KTZD3XA (seahaven.com). |
Lambda Functions
All Python 3.12, ARM64, explicit LogGroup (/aws/lambda/<name>, 60-day retention).
| Function | Trigger | Purpose |
|---|---|---|
apm-wo-analysis-classifier |
S3 ObjectCreated on `raw/*.xlsx |
.csv` |
apm-wo-analysis-slack-post |
Async-invoked by the classifier ({"dt": …}) |
Read today's + yesterday's meta/.../summary.json, post the daily summary, and (only when third_escalation_count > 0) the batched 3rd-escalation alert from details.json. 256 MB / 30 s. |
apm-wo-analysis-slack-interactions |
HTTP API POST /slack/interactions |
Verify the Slack request signature, read meta/.../details.json, and views.open a filtered WO-list modal within Slack's 3 s trigger_id window. 256 MB / 30 s. |
Configuration
Secrets Manager (names only — created out-of-band, never in CloudFormation)
| Secret | Purpose |
|---|---|
apm-wo-analysis/anthropic-api-key |
Claude Haiku fallback for ambiguous free-text comments. |
apm-wo-analysis/slack-credentials |
JSON { botToken, signingSecret, channelId } for the reused Slack app. |
SSM Parameters
| Parameter | Purpose |
|---|---|
/apm-wo-analysis/grafana-base-url |
Grafana dashboard deep-link base (https://grafana.seahaven.com/d/apm-wo/...). |
Environment Variables (non-secret)
- classifier:
APM_HAIKU_FALLBACK(on/off),SLACK_POST_FUNCTION_NAME. - slack-post / slack-interactions:
SLACK_SECRET_NAME,DASHBOARD_URL_PARAM,ANALYTICS_BUCKET.
GitHub repo secret
AWS_DEPLOY_ROLE_ARN— the OIDC deploy rolegithubdeploy-apm-wo-analysis.
CDK context (cdk/cdk.json)
wildcardCertArn, hostedZoneId/hostedZoneName, slackInteractionsDomain, grafanaDomain, grafanaVpcId/grafanaAzs/grafana{Public,Private}SubnetIds, officeCidrs, athenaPluginVersion (3.2.0).
The classification model
Always two-axis, never comment-only. The legacy script's central flaw was
reading only the comment while ignoring WO Status + Hold Reason, which left
~17% in "Other". The two-axis model cuts that to ~9% before any AI — and on the
real 347-row export the implementation lands "Other" at 5.2% (18 rows) with
18 mismatches flagged.
- Axis 1 — comment intent: regex over the HTML-stripped
Last Comment, most-specific first (escalations → SIM ticket → vendor no-show → scheduling → reports → completion → … → other). - Axis 2 — structured state:
Hold Reason→ category, andWO Statussignals (RCAN→Cancelled,Hcorroborates On Hold,IP/R/RRin-flight). - Resolution: comment intent wins when confident → else structured state →
else
Other. A Claude Haiku fallback (Secrets Manager key) is reserved for ambiguous free-text with no structured signal. - Mismatch detector (a feature): flags when comment intent contradicts
structured state (e.g. "completed" while
WO StatusisIP). Surfaced, never suppressed.
Authoritative spec: CLAUDE.md. Implementation:
lambdas/classifier/classify.py (logic) and lambdas/classifier/handler.py
(S3 → Parquet + JSON + invoke).
Repository layout
cdk/
app.py CDK entry point — instantiates both stacks
cdk.json context: cert, zone, subnets, office CIDRs, plugin version
requirements.txt aws-cdk-lib==2.253.1, constructs>=10.6.0
assets/
grafana_userdata.sh EC2 bootstrap: install Grafana + Athena plugin, S3 config sync
stacks/
pipeline_stack.py S3, Lambdas, DLQ, Glue, Athena, HTTP API, IAM
grafana_stack.py VPC import, EC2, ALB, SG, Route53, instance role, DLM
lambdas/
classifier/ S3-triggered: parse → two-axis classify → Parquet + meta JSON
slack_post/ blockkit.py (builders), handler.py (post), interactions.py
(modals), slackio.py (Secrets/SSM/S3/signature)
grafana/
provisioning/ Athena datasource + dashboard provider (as code)
dashboards/ apm-work-orders.json (uid apm-wo) — source of truth
slack/manifest.yaml Slack app manifest (interactivity request URL)
scripts/ local drop-folder uploader + launchd plist
tests/ classifier smoke test + offline synth/blockkit assertions
docs/BUILD.md phased, end-to-end build guide
Ingestion (no email)
The export reaches S3 by direct upload or a local drop-folder, never SES/email.
-
Direct:
aws s3 cp ./export.xlsx s3://apm-wo-analysis-exports-328440206208/raw/ -
Drop-folder (optional zero-touch): a launchd agent (
scripts/apm-wo-uploader.shscripts/com.seahaven.apm-wo-uploader.plist) that watches~/apm-wo-drop/, uploads new.xlsx/.csvfiles toraw/, and archives them locally. Uploads with the scopedapm-wo-dropprofile (IAM userapm-wo-drop-uploader).
Install — the runnable copy and watched folder must live outside
~/Documents(macOS TCC sandbox; a repo-path script fails silently withLastExitStatus=32256):install -d "$HOME/.local/bin" "$HOME/apm-wo-drop" cp scripts/apm-wo-uploader.sh "$HOME/.local/bin/apm-wo-uploader.sh" chmod +x "$HOME/.local/bin/apm-wo-uploader.sh" cp scripts/com.seahaven.apm-wo-uploader.plist "$HOME/Library/LaunchAgents/" launchctl load -w "$HOME/Library/LaunchAgents/com.seahaven.apm-wo-uploader.plist" aws configure --profile apm-wo-drop # one-time, with the uploader's access key
Any .xlsx/.csv landing under raw/ invokes the classifier.
Local development
- pyenv Python 3.12.
ruff check+ruff format --checkbefore pushing (hook-enforced). - Tests:
python -m pytest tests/ -q(~158 tests — classifier rule ladder + handler transforms + Haiku fallback, Block Kit builders, the signature-verified Slack interactions endpoint, slack-post orchestration, and offlinecdk.assertionssynth checks). No AWS needed — external boundaries are monkeypatched. Tests run in CI (ci.yamlsetsrun-tests: true); coverage is reported viapytest-cov(tests/requirements.txt, ~94%, non-gating). - The classification quality gate (deterministic "Other" share) runs in CI against a
committed synthetic fixture
tests/fixtures/sample_export.csv. Additionally, smoke-test against the real export before declaring any classification change done:~/Downloads/_documents/Sheet1-1.xlsx(skips automatically when absent). cdk synthmust pass in CI before merge (Docker + QEMU — Lambda deps are bundled for ARM64;ci.yamlsetsenable-qemu: true).
Deployment
CI/CD via the org reusable workflows (no manual prod deploys in steady state):
- CI (
.github/workflows/ci.yaml) →ci-python-sam.yaml@main: ruff +pytest+cdk synth(QEMU-enabled). Runs on PRs intomain. - Deploy (
.github/workflows/deploy.yaml) →cd-cdk.yaml@main: OIDC assume-role,cdk deploy --all, single-flight concurrency. Runs on push tomain.
Stack name/region/account: apm-wo-analysis-{pipeline,grafana} / us-east-1 / 328440206208.
OIDC deploy role githubdeploy-apm-wo-analysis must exist before the first deploy.
Note: CD is temporarily disabled (deploy job gated
if: ${{ false }}on the phase-0 branch) while the Phase 0–5 stack is merged intomain, to avoid a deploy on every merge. Re-enable as the first Phase 6 step by reverting that commit. Both stacks are already deployed manually and validated in prod.
Manual deploy (emergency/reference; pipeline first so the bucket/table exist):
cd cdk && pip install -r requirements.txt
cdk deploy apm-wo-analysis-pipeline # S3, Glue, Athena, Lambdas, DLQ, API, IAM
cdk deploy apm-wo-analysis-grafana # EC2, ALB, SG, Route53, DLM, dashboards
Operations
Verify it's working — drop a real export into raw/, then within ~1 min:
analytics/dt=<today>/*.parquetandmeta/dt=<today>/{summary,details}.jsonappear,- the daily summary posts to the WO Slack channel (+ alert if any 3rd escalations),
- the Grafana dashboard renders over the office network (
grafana.seahaven.com).
Logs: CloudWatch /aws/lambda/apm-wo-analysis-{classifier,slack-post,slack-interactions} (60-day retention).
Failure modes:
- Classifier failure (malformed export, transient error) → after Lambda retries, the event lands in
apm-wo-analysis-classifier-dlq. Check the DLQ if a day's data is missing. - Slack post failure is best-effort and does not fail classification (data still lands in S3).
- Grafana down → check the instance via SSM Session Manager (no SSH);
systemctl status grafana-server; ALB target health.
Reprocess a day: re-upload the same export to raw/ — the classifier uses
overwrite_partitions, so a same-day re-run replaces that dt partition idempotently.
Grafana admin: access is office-IP-restricted at the ALB; the instance is
SSM-only. Dashboards are provisioned from grafana-config/ in S3 (synced on boot
and by a 15-min systemd timer); edit dashboards in-repo, not in the UI
(allowUiUpdates: false). grafana.db lives on the RETAIN'd gp3 volume and is
snapshotted daily by DLM.
Documentation
- Confluence "AWS Architecture Map" (IT space, page 1540098): a Mermaid subgraph for this stack — done (added under Serverless Applications, plus rows in CI/CD Pipelines, EC2 Inventory, and Key Data Stores).
- Confluence stack page "APM WO Analysis (apm-wo-analysis)" (IT space, page 8749057, under AWS Cloud Infrastructure): full resource/Lambda/secrets detail — done.
- Slack Apps Inventory (Confluence page 524569): "APM Work Orders" (new dedicated
app, App ID
A0B6P28V64B) added — done. - This README +
docs/BUILD.md(phased build guide) +CLAUDE.md(domain spec).
Notes / Gotchas
- Partition date comes from the S3 event time, not the Lambda wall-clock — stable across retries and the midnight boundary.
meta/vsanalytics/: summary/details JSON must stay out ofanalytics/— Athena reads every object in the table prefix as Parquet and chokes on JSON.- Classifier deps (awswrangler/pandas/pyarrow/numpy) come from the AWS-managed SDK-for-pandas layer — bundling them blows Lambda's 250 MB unzipped limit. Verify the pinned layer ARN/version on region or runtime changes.
- Grafana dashboard contract: dashboard
uidmust stayapm-wo(the Slack 📊 button deep-links tod/apm-wo). Athena plugin query key israwSQL(capital). DatasourceauthType: default(the EC2 instance role;ec2_iam_roleis rejected by the plugin). Config sync usesaws s3 sync --exact-timestamps(plain sync skips same-size edits). Template vars userefresh: 1(on load). - No Client VPN exists — "VPN-only" Grafana is realized as office-IP SG
restriction.
seahaven-vpchas a single NAT (one AZ) for instance egress. - Slack interactions endpoint is unauthenticated at the gateway by design; the Lambda verifies the Slack signature (replay window + HMAC). Stage-throttled.
Known operational debt
- Re-enable CD (revert the phase-0 disable) once the stack is merged.
- One clean instance replacement is owed to validate the committed user-data from a cold boot and to apply root-volume encryption (can't encrypt in place).
- Deferred review NITs:
print()→logging, source-IP logging on signature failure, S3 versioning, the'${site:raw}'WO-table SQL tidy, and a CIDR-maintenance note in the runbook.
Status
Phases 2–5 implemented, deployed to prod and validated end-to-end (classifier,
Slack post + alert + modal, Grafana dashboard). Stacked PRs #6→#11 are open and
unmerged; cross-review (#2/#4/#5) and /security-review of the two public endpoints
are cleared. Phase 6 (this docs pass + Confluence + runbook) is in progress on
feature/phase-6-docs.