Commit graph

10 commits

Author SHA1 Message Date
Adam Moussa
e4fc32599f
fix(pipeline): page when classifier has no daily invocation
Some checks failed
Deploy / deploy (push) Has been cancelled
2026-08-13 17:39:11 -04:00
Adam Moussa
44ad13d8cd
Add DLQ messages-present alarm for classifier (INFRA-57) (#26)
Some checks failed
Deploy / deploy (push) Has been cancelled
The classifier DLQ (apm-wo-analysis-classifier-dlq) had no alarm, so a
message landing there after Lambda exhausts async retries — a genuinely
dropped classifier run — surfaced nowhere.

Add a CloudWatch alarm on AWS/SQS ApproximateNumberOfMessagesVisible
(Maximum, threshold >0, 300s period, 1 evaluation period, notBreaching)
that pages the shared site-alerts SNS topic. Mirrors the
workorder-email-processor-dlq-messages alarm in procurement-ingest's
wo_stack and the payments-payroll-batch-dlq-messages alarm convention.
ALARM-only, no OK action, per the CloudWatch-alarm preference.
2026-06-17 15:02:10 -04:00
Adam Moussa
eae4d67e00 Apply cross-review findings (Phase 2/4/5 hardening)
From the cross_reviewer (GPT-4.1) per-PR passes, now that the orchestrator is
back up:

Phase 2 (classifier):
- Process ALL S3 records, not just event["Records"][0] — batched notifications
  no longer silently dropped (the review's only BLOCK).
- Derive the partition dt from the S3 event time, not the Lambda wall-clock —
  stable across retries / the midnight boundary.
- Add an SQS dead-letter queue so a failed run surfaces instead of dropping a
  day's data after Lambda's retries.

Phase 4 (Slack):
- Stage throttling (rate 10 / burst 20) on the public /slack/interactions HTTP
  API. (AWS WAF doesn't attach to apigwv2 HTTP APIs; stage throttling is the
  mechanism.)

Phase 5 (Grafana):
- Explicit encrypted=True on the gp3 root volume.

Tests: synth assertions for the DLQ, stage throttling, and the encrypted volume.
60/60 pass; cdk synth green for both stacks. Deferred NITs (print->logging, sig-
failure source-IP logging, S3 versioning, CIDR-maintenance runbook) -> Phase 6.

NOTE: like the earlier deploy fixes these sit on phase-5 but span phases — the
classifier/DLQ to #8, throttling to #10, encryption to #11 — reconcile at merge.
The encrypted-volume change needs the deferred clean instance replacement to
take effect (can't encrypt a live volume in place).
2026-05-29 11:29:14 -04:00
Adam Moussa
68f3acded8 Fix dashboard rendering: barchart panel type + metadata off the table prefix
Two issues found loading the deployed dashboard:

1. Panels used type "bar-chart" (hyphenated); Grafana's core panel is "barchart"
   — hence "plugin bar-chart required". Fixed both panels.

2. ALL panels showed "no data" because Athena failed with HIVE_BAD_DATA:
   the classifier wrote summary.json/details.json INTO analytics/dt=*/ — the
   same prefix the Glue table scans — so Athena tried to read the JSON as
   Parquet and every query failed. Move the metadata to a separate meta/dt=*/
   prefix: classifier writes there (grant_read_write meta/*), the Slack Lambdas
   read there (read_meta_json, grant_read meta/*), and analytics/ holds only
   Parquet. Verified: the category GROUP BY query now succeeds against Athena.
2026-05-28 18:49:41 -04:00
Adam Moussa
f193754c27 Fix deploy-time failures found in prod testing
Two issues only a real deploy/run surfaced (synth + offline tests passed):

1. Classifier exceeded Lambda's 250 MB unzipped limit (bundled awswrangler +
   pandas + pyarrow + numpy). Move them to the AWS-managed SDK-for-pandas layer
   (AWSSDKPandas-Python312-Arm64:27, awswrangler 3.16.1, pre-stripped to fit);
   bundle only openpyxl. Drop the unused anthropic SDK — _call_haiku uses stdlib
   urllib. Function package now ~890 KB.

2. Slack rejected the daily post with invalid_blocks: every category drill
   button shared action_id "drill_category". Qualify it as "drill_category:<cat>"
   for uniqueness; the interactions handler now matches on the prefix. Add a
   regression test asserting all daily-summary action_ids are unique.

Verified in prod: classifier writes Parquet + summary.json + details.json;
slack-post posts the daily summary + 3rd-escalation alert; the interactions
endpoint (apm-wo.seahaven.com) returns 401 on a bad signature. 58/58 tests pass.

NOTE: these fixes sit on the phase-5 branch but logically belong to earlier
phases — the layer fix to #8 (classifier), the Slack fix to #10 — and must be
moved/cherry-picked there before those PRs merge independently. See cleanup.
2026-05-28 18:29:16 -04:00
Adam Moussa
c3a2c936d1 Add Slack post + interactions Lambdas with drill-down modals (Phase 4)
Two push surfaces (no App Home) + interactive drill-down, per CLAUDE.md.

Block Kit (blockkit.py, pure/offline): build_daily_summary (header, vs-yesterday
deltas, escalation breakdown with 3rd highlighted, action/routine, top sites,
mismatch callout, category drill buttons + 📊 Open dashboard link, footer),
build_escalation_alert (one @here, returns None on zero-3rd — suppression), and
build_wo_modal (views.open payload, capped under Slack's 100-block limit).

Lambdas: slack_post/handler.py (classifier-invoked: read today/yesterday
summary.json, post daily summary, conditionally post the batched alert from
details.json) and slack_post/interactions.py (API Gateway: verify Slack
signature, filter details.json, views.open the WO modal within the 3s trigger_id
window). slackio.py centralizes Secrets Manager creds, the SSM dashboard URL,
signature verification, and analytics/ reads — keeping blockkit pure.

Classifier: emit analytics/dt=*/details.json (per-WO index for the modals) and
async-invoke slack-post after the snapshot write (best-effort; a Slack failure
never fails classification).

CDK: slack-post + interactions Lambdas (Docker-bundled slack_sdk), HTTP API on
apm-wo.seahaven.com (wildcard ACM cert + Route53 alias; signature-verified, so
the route is unauthenticated by design), SSM /apm-wo-analysis/grafana-base-url,
and scoped IAM (read analytics/, read the Slack secret + dashboard param;
classifier granted lambda:InvokeFunction on slack-post). Slack creds live in one
Secrets Manager secret apm-wo-analysis/slack-credentials {botToken, signingSecret,
channelId}; cdk.json gains cert/zone/domain context.

WO drill-downs link to Grafana only — no APM deep-links (per decision).

Deliverables for test time: slack/manifest.yaml (app manifest, interactivity
request_url = apm-wo.seahaven.com).

Tests: tests/test_blockkit.py (30 offline cases — deltas, zero-3rd None, <100
blocks under large inputs, modal truncation/overflow, dashboard URL) and Phase 4
assertions in test_pipeline_synth.py (both Lambdas, the API route/domain/alias,
and no broad/write IAM on the Slack roles). 49/49 tests pass; cdk synth green.
2026-05-28 17:48:51 -04:00
Adam Moussa
6ee218eea3 Add analytics dataset: projection table + Athena workgroup (Phase 3)
CDK-define the apm_wo_snapshots Glue table with partition projection over
analytics/ (dt as projected date partition, 2026-01-01..NOW). Projection means
no crawler, no MSCK REPAIR, and — critically — the classifier needs no Glue
catalog access at all.

pipeline_stack.py: glue.CfnTable (Parquet SerDe, 17-column schema mirroring the
classifier's snapshot incl. contractor_description and the two boolean flags) +
athena.CfnWorkGroup `apm-wo-analysis` (enforced result location, SSE-S3) + a
30-day lifecycle rule on athena-results/ (disposable query output in a RETAIN
bucket). Trim the classifier role: drop the entire Glue policy statement.

handler.py: stop registering the table at runtime — drop database=/table= from
to_parquet so the classifier writes pure Parquet; partition projection handles
the rest. Keeps overwrite_partitions for idempotent same-day re-uploads.

test_pipeline_synth.py: offline synth assertions (bundling skipped) — projection
properties, column schema/types, workgroup result enforcement, and that no IAM
policy grants glue:* to the classifier.

cdk synth green; 11/11 tests pass.
2026-05-28 17:24:36 -04:00
Adam Moussa
eb724580e0 Add two-axis classifier Lambda and CDK wiring (Phase 2)
Implement the core classification engine and wire it into the pipeline stack.

classify.py: two-axis classifier — HTML-strip, comment-intent regex buckets
(escalations → status inquiry, most-specific first), Hold Reason / WO Status
structured state, comment-vs-state mismatch detector, and a Claude Haiku
fallback (Secrets Manager key) reserved for ambiguous free-text. Exports
ESCALATION_CATEGORIES / ACTION_NEEDED_CATEGORIES.

handler.py: S3-triggered handler — parse xlsx/csv, classify each non-blank
row, write a per-WO Parquet snapshot to analytics/dt=YYYY-MM-DD/ (registers
the Glue partition via awswrangler) and a summary.json for slack-post (Phase 4).

pipeline_stack.py: Glue database, ARM64 Python 3.12 classifier Lambda
(Docker-bundled deps), S3 raw/ notification (.xlsx/.csv), and least-privilege
IAM (read raw/, read-write analytics/, scoped Glue catalog, read Anthropic key).

Smoke-tested against the real export: 347 rows, "Other" at 5.2% (target ~9%),
18 mismatches flagged. 7/7 unit + smoke tests pass; cdk synth green.
2026-05-28 17:09:52 -04:00
Adam Moussa
7befc8f0d7 Add drop-folder ingestion and scoped uploader IAM user
Complete Phase 1 ingestion. Add a least-privilege IAM user
(apm-wo-drop-uploader) to the pipeline stack, scoped to s3:PutObject
on the raw/ prefix only — the local launchd uploader authenticates as
this user via a dedicated profile, so a laptop credential leak cannot
read, list, or touch the analytics data.

Replace the scaffold uploader stub with the hardened stampli-pattern
script (lockfile, logging, timestamped archive, notifications, settle
delay) and align names to the convention (~/apm-wo-drop, ~/.local/bin,
com.seahaven.apm-wo-uploader). The plist sets PATH/HOME because launchd
runs with a stripped environment and otherwise cannot find aws.

The exports bucket already shipped in the Phase 0 scaffold, so the code
delta here is the uploader identity and tooling.
2026-05-28 16:32:56 -04:00
Adam Moussa
58b91bda70 Scaffold apm-wo-analysis repository
Stand up the Phase 0 CDK scaffold for the daily APM work-order
analysis pipeline: two-stack CDK app (pipeline + grafana), classifier
and slack-post Lambda packages, dashboards-as-code, the local
drop-folder uploader, and a classifier smoke-test placeholder.

Wire CI/CD to the org reusable workflows: ci.yaml -> ci-python-sam
(ruff + cdk synth) and deploy.yaml -> cd-cdk (OIDC, cdk deploy --all).
Pin aws-cdk-lib==2.253.1; Lambdas target Python 3.12 / arm64.

Rewrite .gitignore to the org Python-CDK standard so the source-of-
truth files (CLAUDE.md, docs/, .claude/agents) are tracked while build
artifacts (.venv, cdk.out, caches) stay ignored.

Domain logic, stack resources, and dashboards are stubbed and filled
in across Phases 1-5 (docs/BUILD.md). cdk synth is green for both
stacks; ruff check/format pass.
2026-05-28 16:13:13 -04:00