procurement-ingest/lambdas/wo/email_processor/handler.py
Adam Moussa ff368beaa4
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)

tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.

Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.

Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.

New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
  'content' key, empty content list, non-JSON model text), asserting
  PO's pre-call ai_fallback metric survives with no partial write and
  the exception propagates; WO's no-datapoint-on-throttle behavior is
  pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
  monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
  zero writes, no raise — closing the hole where deleting the gate
  line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
  unset ARN and on a Secrets Manager exception, TTL cache refresh,
  Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
  no table scan, non-ASCII token, and a hostile-field-escaping
  regression lock. PO web_ui has no __init__.py, so these go through
  the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
  (table 'WorkOrders'): null-status never clobbers wo_status,
  created_at immutable via if_not_exists, status->wo_status mapping,
  None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
  per-pipeline multi-record failure-isolation (all-or-retry contract).
  The reprocess.py synthetic-event-shape contract test already landed
  in Phase 7, so it isn't duplicated here.

Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.

_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.

CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.

docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.

* test: lock attribute-context quote escaping in web_ui hostile-field test

The escaping regression lock asserted only the element-context vector
(raw <script> absent, &lt;script&gt; present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).

Add assertions that the onclick sink's JSON string renders its opening
quote as &quot; (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.

* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override

- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
  xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
  behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
  corrected, instead of requiring a passing test to be knowingly deleted. The
  module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
  defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
  reusable workflow's bare pytest) red-exits local subset runs by design, with the
  --cov-fail-under=0 override for iteration. Floor stays in addopts because the
  centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00

143 lines
6.5 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

"""
Email processor Lambda.
Triggered by S3 events when SES delivers an email.
Parses the raw email, sends it to Claude for structured extraction,
then writes the result to DynamoDB.
Phase 5: this handler is the thin event loop + fail-closed auth + routing. The
work has moved to flat sibling modules (bare-name imports resolve via the same
flat-landing bundling as ses_auth/template_parser):
extraction.py -- extract_with_bedrock + the <email>-tag neutralizer
telemetry.py -- emit_parse_metric (stdout EMF)
persistence.py -- save_work_order / save_event / comment_id determinism
prompts.py -- EXTRACTION_PROMPT (re-exported below for tests)
"""
import logging
import re
import boto3
from email_parsing import parse_raw_email
from extraction import extract_with_bedrock
from persistence import save_event, save_work_order
from prompts import EXTRACTION_PROMPT # noqa: F401 (re-export for tests)
from ses_auth import authenticate_inbound_email
from telemetry import emit_parse_metric
from template_parser import try_deterministic_parse, validate_ai_fallback
logger = logging.getLogger()
logger.setLevel(logging.INFO)
# Lazy cached S3 client. Keeps the public attribute name ``s3`` so the test
# monkeypatch target changes module only, not attribute name.
s3 = None
def _get_s3():
global s3
if s3 is None:
s3 = boto3.client("s3")
return s3
def handler(event, context):
"""Lambda entry point. Triggered by S3 ObjectCreated events."""
# Deploy-guard healthcheck (Phase 0): a top-level direct-invoke
# {"healthcheck": true} probe returns immediately, BEFORE any S3 fetch,
# SES sender-auth gate, or Records iteration. Real mail arrives as S3
# ObjectCreated events whose top-level keys ("Records") AWS controls, so
# email content can never set this key -- this creates no accept path for
# mail. It emits NO EMF and no log line matching the sender_auth_rejected
# metric-filter, so repeated post-deploy smoke invokes never page.
if isinstance(event, dict) and event.get("healthcheck") is True:
return {"healthcheck": "ok"}
for record in event.get("Records", []):
bucket = record["s3"]["bucket"]["name"]
key = record["s3"]["object"]["key"]
s3_key = f"s3://{bucket}/{key}"
logger.info(f"Processing email: {s3_key}")
# Fetch raw email from S3
response = _get_s3().get_object(Bucket=bucket, Key=key)
raw_email = response["Body"].read()
# Fail-closed sender authentication (INFRA-107): only mail with an
# SES-stamped dkim=pass verdict for an allowlisted domain may create
# or update work orders. Rejected mail is logged and skipped without
# erroring the invocation (no retries / DLQ spam).
if not authenticate_inbound_email(raw_email, s3_key):
continue
# Parse the raw email
email_data = parse_raw_email(raw_email)
logger.info(f"Subject: {email_data['subject']}")
# Deterministic template parse first; fall back to the AI extractor only
# on a miss or an invalid (fail-closed) result.
parsed, method, template_id, reason = try_deterministic_parse(email_data)
if parsed is None:
method = "ai_fallback"
# A Bedrock transport error (throttle, malformed response, etc.)
# previously emitted ZERO ParseOutcome datapoints -- the only emit
# sites are the post-gate success (below) and the ai_fallback_rejected
# branch. Wrap the call so a fallback attempt that dies in Bedrock
# records exactly one datapoint (ai_fallback / ReasonCode=bedrock_error)
# and then re-raises into the async-retry / DLQ path. The emit is in
# the except -- never pre-call -- so a gate-rejected email (Bedrock
# returned, validate_ai_fallback fails below) still emits ONLY
# ai_fallback_rejected, preserving the wo_stack "a rejected email
# emits nothing else" alarm contract (no double-count).
try:
parsed = extract_with_bedrock(email_data)
except Exception:
emit_parse_metric("ai_fallback", template_id, "bedrock_error", None)
raise
# Fail-closed validation gate on AI output: a prompt-injected
# email body could steer the model into returning arbitrary
# field values, so enforce the same structural contract on both
# parse paths BEFORE any DynamoDB write.
ok, val_reason = validate_ai_fallback(parsed)
if not ok:
logger.warning(
f"AI-fallback validation failed ({val_reason}), skipping: {key}"
)
emit_parse_metric(
"ai_fallback_rejected",
template_id,
val_reason,
parsed.get("work_order_id") if isinstance(parsed, dict) else None,
)
continue
logger.info(
f"Parsed ({method}/{template_id}/{reason}): "
f"type={parsed.get('email_type')}, wo={parsed.get('work_order_id')}"
)
emit_parse_metric(method, template_id, reason, parsed.get("work_order_id"))
# work_order_id becomes a DynamoDB partition key and the leading, '#'-
# delimited segment of the comment_id range key, so it must be digits
# only. The template path already guarantees this via validate(); the
# AI-fallback path returns raw model output, which a prompt-injected
# email body could steer into a non-numeric or '#'-bearing value that
# forges key segments or lands on an arbitrary WO. Enforce the same
# contract on both paths and skip (fail closed) on a violation.
# [0-9] not \d: \d is Unicode-aware and would admit fullwidth digits
# (e.g. "12345") as a distinct-but-lookalike partition key.
work_order_id = parsed.get("work_order_id")
if not work_order_id or not re.fullmatch(r"[0-9]+", str(work_order_id)):
logger.warning(f"Missing or non-numeric work order ID, skipping: {key}")
continue
# Always upsert the work order with any new info
save_work_order(parsed, s3_key)
# Save every email as an event for history tracking. The raw object key
# (not the s3:// URI) drives the retry-idempotent comment_id suffix;
# the parse method decides whether comment_time may enter the key (A1).
save_event(parsed, s3_key, key, method, email_data.get("date"))
return {"statusCode": 200, "body": "OK"}