mirror of
https://github.com/Sea-Haven-Industries/apm-wo-analysis.git
synced 2026-09-30 11:13:14 +00:00
Implement the core classification engine and wire it into the pipeline stack. classify.py: two-axis classifier — HTML-strip, comment-intent regex buckets (escalations → status inquiry, most-specific first), Hold Reason / WO Status structured state, comment-vs-state mismatch detector, and a Claude Haiku fallback (Secrets Manager key) reserved for ambiguous free-text. Exports ESCALATION_CATEGORIES / ACTION_NEEDED_CATEGORIES. handler.py: S3-triggered handler — parse xlsx/csv, classify each non-blank row, write a per-WO Parquet snapshot to analytics/dt=YYYY-MM-DD/ (registers the Glue partition via awswrangler) and a summary.json for slack-post (Phase 4). pipeline_stack.py: Glue database, ARM64 Python 3.12 classifier Lambda (Docker-bundled deps), S3 raw/ notification (.xlsx/.csv), and least-privilege IAM (read raw/, read-write analytics/, scoped Glue catalog, read Anthropic key). Smoke-tested against the real export: 347 rows, "Other" at 5.2% (target ~9%), 18 mismatches flagged. 7/7 unit + smoke tests pass; cdk synth green. |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| grafana_stack.py | ||
| pipeline_stack.py | ||