procurement-ingest/lambdas/shared/emf.py

83 lines
3.5 KiB
Python
Raw Permalink Normal View History

feat: extract lambdas/shared/ — single-source ses_auth, web_ui auth, email parsing, EMF emitter (refactor phase 3) (#111) Four modules move into the handbook-mandated lambdas/shared/ location, collapsing duplicated logic that had to be kept in sync by hand across the PO and WO pipelines: - ses_auth.py: the PO and WO copies were verified sha256-identical against the feature/phase-7-ops-recovery baseline before the move (no drift since the last audit). shared/ses_auth.py is the exact bytes of that one copy; both originals are git rm'd (the PO copy via rename, the WO copy as a straight delete). Bundling lands the module flat in /asset-output for both email processors, so the handlers keep `from ses_auth import authenticate_inbound_email` unchanged — zero handler diff for this move, which is what keeps fail-closed auth byte-identical through the change. - web_ui_auth.py: extracts the byte-identical _get_auth_token / _header / is_authenticated block plus the four token-cache globals out of both web_ui handlers. The per-stack INFRA-74 comments stay in each handler as-is (deliberately drifted wording, stack-specific) rather than being unified into the shared module. Fail-closed semantics (unset ARN or Secrets Manager exception -> deny) are unchanged. - email_parsing.py: parse_raw_email ships as the superset version that returns cc unconditionally. WO's output is bit-identical to before; PO simply ignores the cc field rather than being "cleaned up" to consume it. No second variant is kept. - emf.py: a generic emitter parameterized by namespace, dimension sets, and properties. Every call site's emitted EMF envelope is unchanged, including the load-bearing [["ParseMethod"],["ParseMethod","TemplateId"]] dimension-set shape the alarms and metric filters depend on. Emission ordering is untouched: PO still emits ai_fallback before the Bedrock call, WO still emits its mutually-exclusive ai_fallback/ai_fallback_rejected after its gate. The deliberate-double-count comments survive. _emit_derived_agreement_metric was found living inside derived_fields.py, so per the DERIVED-FIELDS exception it is left as a third, unconverted copy (derived_fields.py and the shadow DerivedFieldAgreement telemetry stay untouchable while that bake runs) — a comment there points at shared/emf.py for the eventual follow-up. Bundling: both email-processor cdk bundling commands gain a trailing `cp shared/*.py /asset-output/` (they were already cp-only post-Phase 7, so no pip step or manylinux pin is reintroduced). Both web_ui functions gain the same widened-root staging so web_ui_auth.py ships beside their handler; site_extractor's from_asset is untouched. Tests: PO_EXPECTED_TOP_LEVEL_MODULES gains the shared modules that now ship, the AST sibling-import check resolves imports whose source now lives under shared/, and the new shared cp line has its own revert/mutation detection. _SIBLING_MODULES resolution and _po_parser_support.py now load ses_auth/email_parsing/emf from shared/; the two-copy ses_auth byte-identity fixture-hygiene test is retired as obsolete now that there is one copy, and the ses_auth fixture parameterization over two identical copies is dropped. The sys.modules save/restore dance for template_parser (still duplicated per-pipeline) is left in place.
2026-07-20 13:38:23 -04:00
"""Shared CloudWatch EMF emitter for the procurement-ingest pipelines.
A single generic ``emit_metric`` builds the Embedded Metric Format envelope
(``_aws`` block + promoted dimension properties + the metric-value key) and
prints it to stdout, where the Lambda log subscription materializes the metric.
Both the PO and WO email processors emit their parse-outcome metric through
``emit_parse_outcome`` so the load-bearing dimension-set list
``[["ParseMethod"], ["ParseMethod", "TemplateId"]]`` is pinned in exactly ONE
place -- a one-sided dimension change to one pipeline is structurally
impossible.
Envelope byte-equivalence rests on CPython dict insertion order (preserved by
json.dumps with default separators): ``{**properties, metric_name: value}``
yields ``_aws`` first, then the caller's properties in order, then the metric
value last -- matching every inline emitter this module replaced.
"""
import json
from datetime import datetime, timezone
PARSE_METRIC_NAME = "ParseOutcome"
# Two dimension sets are published for every parse-outcome metric: ["ParseMethod"]
# (aggregated across all template ids -- the series the fallback-rate alarms
# query) AND ["ParseMethod", "TemplateId"] (per-template breakdown for Logs
# Insights / dashboards). CloudWatch materializes only the exact dimension sets
# listed here and does NOT auto-aggregate, so an alarm's single-dimension query
# would receive no data unless ["ParseMethod"] is emitted explicitly. Pinned
# ONCE here; both pipelines share it.
_PARSE_DIMENSION_SETS = [["ParseMethod"], ["ParseMethod", "TemplateId"]]
def emit_metric(
namespace, metric_name, dimension_sets, properties, *, value=1, unit="Count"
):
"""Print one CloudWatch EMF log line for ``metric_name`` in ``namespace``.
Zero-latency (no PutMetricData API call): the extraction path is async and
the role already has logs:PutLogEvents. ``properties`` are emitted in caller
order between the ``_aws`` block and the trailing metric-value key; the keys
named in ``dimension_sets`` are the promoted (dimensioned) fields, the rest
ride along as Logs-Insights-queryable properties.
"""
emf = {
"_aws": {
# EMF requires Timestamp (epoch ms); without it CloudWatch may not
# extract the metric datapoint from the log event.
"Timestamp": int(datetime.now(timezone.utc).timestamp() * 1000),
"CloudWatchMetrics": [
{
"Namespace": namespace,
"Dimensions": dimension_sets,
"Metrics": [{"Name": metric_name, "Unit": unit}],
}
],
},
**properties,
metric_name: value,
}
print(json.dumps(emf))
def emit_parse_outcome(namespace, method, template_id, reason_code, id_key, id_value):
"""Emit one parse-outcome EMF line shared by both email processors.
ParseMethod/TemplateId are the only promoted (dimensioned) fields to keep
cardinality low; ReasonCode and the pipeline id (``id_key``: ``po_number``
or ``work_order_id``) ride along as Logs-Insights-queryable properties. See
``_PARSE_DIMENSION_SETS`` for why both the single- and two-dimension sets
are published.
"""
emit_metric(
namespace,
PARSE_METRIC_NAME,
_PARSE_DIMENSION_SETS,
{
"ParseMethod": method,
"TemplateId": template_id or "unknown",
"ReasonCode": reason_code or "ok",
id_key: id_value or "",
},
)