procurement-ingest/lambdas/wo/email_processor/telemetry.py

36 lines
1.5 KiB
Python
Raw Normal View History

feat: decompose email-processor handlers into flat siblings + lazy boto3 clients (refactor phase 5) (#113) Both email-processor God-handlers split along the seams that already work in the flat-sibling pattern established by lambdas/shared/, so bare-name imports keep working under the existing bundling glob. PO (5-way split): handler.py keeps only the event loop, fail-closed auth, and email_type routing. extraction.py holds extract_with_claude and _EMAIL_TAG_RE, importing EXTRACTION_PROMPT from prompts.py and parse_raw_email from shared/email_parsing.py rather than recreating a PO-local copy. enrichment.py is a pure code move of enrich_parsed and pad_zip (PO-only; WO has no enrichment stage) with zero behavior change. telemetry.py holds the EMF ParseMethod emit wrappers. persistence.py holds _write_fields/_merge_update/save_*, collapsing the byte-identical save_new_po/save_revision bodies into one _save_merge helper that both now call through, preserving the sticky Cancelled ConditionExpression guard for both callers; save_cancellation stays separate. WO (5 concerns, no enrichment stage): the handler loop keeps validate_ai_fallback and the re.fullmatch(r"[0-9]+", work_order_id) key guard ahead of both save_work_order and save_event, since the guard protects the DynamoDB partition key and the '#'-delimited comment_id range-key segment. _header_date_iso and comment_id determinism stay colocated with persistence.py's save_event for the retry-idempotent event_id key. EXTRACTION_PROMPT (PO) moves to prompts.py with cross-reference headers to derived_fields.py's authoritative trade/site/fiscal rule tables; handler.py re-exports it (from prompts import EXTRACTION_PROMPT) since four tests dereference handler.EXTRACTION_ PROMPT directly. WO's prompt moves the same way. I/O modules (extraction.py's bedrock client, persistence.py's dynamodb resource, handler.py's s3 client) get lazy cached boto3 accessors; pure modules (enrichment.py, prompts.py, telemetry.py) import no boto3. Test monkeypatch surfaces move to the module that now owns the client (e.g. persistence.dynamodb) everywhere tests patch it, and the moto-before-handler-import ordering in _po_parser_support.py is preserved so the moto-backed suites don't hit real AWS. Behavior-preservation pins, verified with tests: PO still emits ParseMethod=ai_fallback before the Bedrock call, with ai_fallback_rejected as the additive second datapoint on rejection. WO still emits after its gate with mutually-exclusive ai_fallback / ai_fallback_rejected. Shadow DerivedFieldAgreement telemetry stays ai_fallback-only. derived_fields.py is untouched (diff against feature/phase-3-shared-extraction is empty). handler(event, context) signatures and the save_* public contract are unchanged on both pipelines; goldens unchanged. PO_EXPECTED_TOP_LEVEL_MODULES and its WO equivalent in tests/test_bundle_consistency.py are updated for the new sibling modules so the AST bundle-consistency test still fails on an unshipped or uncommented-out sibling.
2026-07-20 15:34:53 -04:00
"""Parse-outcome telemetry for the work-order email processor.
Pure stdout EMF emitter -- no AWS client, no boto3. The extraction path is async
and the role already has logs:PutLogEvents, so parse outcomes are written as EMF
log lines rather than PutMetricData API calls.
"""
from emf import emit_parse_outcome
# CloudWatch EMF namespace + metric for parse-outcome observability.
METRIC_NAMESPACE = "Seahaven/WorkorderIngest"
def emit_parse_metric(method, template_id, reason_code, work_order_id):
"""Emit one CloudWatch EMF line recording the parse outcome.
Zero-latency (no PutMetricData API call): the extraction path is async and
the role already has logs:PutLogEvents. ParseMethod/TemplateId are the only
promoted (dimensioned) fields to keep cardinality low; ReasonCode and
work_order_id ride along as Logs-Insights-queryable properties.
Two dimension sets are published: ["ParseMethod"] (aggregated across all
template ids -- the series the fallback-rate alarm queries) AND
["ParseMethod", "TemplateId"] (per-template breakdown for Logs Insights /
dashboards). CloudWatch materializes only the exact dimension sets listed
here and does NOT auto-aggregate, so the alarm's single-dimension query
would receive no data unless ["ParseMethod"] is emitted explicitly."""
emit_parse_outcome(
METRIC_NAMESPACE,
method,
template_id,
reason_code,
"work_order_id",
work_order_id,
)