CDK-define the apm_wo_snapshots Glue table with partition projection over
analytics/ (dt as projected date partition, 2026-01-01..NOW). Projection means
no crawler, no MSCK REPAIR, and — critically — the classifier needs no Glue
catalog access at all.
pipeline_stack.py: glue.CfnTable (Parquet SerDe, 17-column schema mirroring the
classifier's snapshot incl. contractor_description and the two boolean flags) +
athena.CfnWorkGroup `apm-wo-analysis` (enforced result location, SSE-S3) + a
30-day lifecycle rule on athena-results/ (disposable query output in a RETAIN
bucket). Trim the classifier role: drop the entire Glue policy statement.
handler.py: stop registering the table at runtime — drop database=/table= from
to_parquet so the classifier writes pure Parquet; partition projection handles
the rest. Keeps overwrite_partitions for idempotent same-day re-uploads.
test_pipeline_synth.py: offline synth assertions (bundling skipped) — projection
properties, column schema/types, workgroup result enforcement, and that no IAM
policy grants glue:* to the classifier.
cdk synth green; 11/11 tests pass.
Implement the core classification engine and wire it into the pipeline stack.
classify.py: two-axis classifier — HTML-strip, comment-intent regex buckets
(escalations → status inquiry, most-specific first), Hold Reason / WO Status
structured state, comment-vs-state mismatch detector, and a Claude Haiku
fallback (Secrets Manager key) reserved for ambiguous free-text. Exports
ESCALATION_CATEGORIES / ACTION_NEEDED_CATEGORIES.
handler.py: S3-triggered handler — parse xlsx/csv, classify each non-blank
row, write a per-WO Parquet snapshot to analytics/dt=YYYY-MM-DD/ (registers
the Glue partition via awswrangler) and a summary.json for slack-post (Phase 4).
pipeline_stack.py: Glue database, ARM64 Python 3.12 classifier Lambda
(Docker-bundled deps), S3 raw/ notification (.xlsx/.csv), and least-privilege
IAM (read raw/, read-write analytics/, scoped Glue catalog, read Anthropic key).
Smoke-tested against the real export: 347 rows, "Other" at 5.2% (target ~9%),
18 mismatches flagged. 7/7 unit + smoke tests pass; cdk synth green.
Complete Phase 1 ingestion. Add a least-privilege IAM user
(apm-wo-drop-uploader) to the pipeline stack, scoped to s3:PutObject
on the raw/ prefix only — the local launchd uploader authenticates as
this user via a dedicated profile, so a laptop credential leak cannot
read, list, or touch the analytics data.
Replace the scaffold uploader stub with the hardened stampli-pattern
script (lockfile, logging, timestamped archive, notifications, settle
delay) and align names to the convention (~/apm-wo-drop, ~/.local/bin,
com.seahaven.apm-wo-uploader). The plist sets PATH/HOME because launchd
runs with a stripped environment and otherwise cannot find aws.
The exports bucket already shipped in the Phase 0 scaffold, so the code
delta here is the uploader identity and tooling.
Stand up the Phase 0 CDK scaffold for the daily APM work-order
analysis pipeline: two-stack CDK app (pipeline + grafana), classifier
and slack-post Lambda packages, dashboards-as-code, the local
drop-folder uploader, and a classifier smoke-test placeholder.
Wire CI/CD to the org reusable workflows: ci.yaml -> ci-python-sam
(ruff + cdk synth) and deploy.yaml -> cd-cdk (OIDC, cdk deploy --all).
Pin aws-cdk-lib==2.253.1; Lambdas target Python 3.12 / arm64.
Rewrite .gitignore to the org Python-CDK standard so the source-of-
truth files (CLAUDE.md, docs/, .claude/agents) are tracked while build
artifacts (.venv, cdk.out, caches) stay ignored.
Domain logic, stack resources, and dashboards are stubbed and filled
in across Phases 1-5 (docs/BUILD.md). cdk synth is green for both
stacks; ruff check/format pass.