procurement-ingest/lambdas/wo/email_processor/extraction.py

89 lines
3.3 KiB
Python
Raw Normal View History

feat: decompose email-processor handlers into flat siblings + lazy boto3 clients (refactor phase 5) (#113) Both email-processor God-handlers split along the seams that already work in the flat-sibling pattern established by lambdas/shared/, so bare-name imports keep working under the existing bundling glob. PO (5-way split): handler.py keeps only the event loop, fail-closed auth, and email_type routing. extraction.py holds extract_with_claude and _EMAIL_TAG_RE, importing EXTRACTION_PROMPT from prompts.py and parse_raw_email from shared/email_parsing.py rather than recreating a PO-local copy. enrichment.py is a pure code move of enrich_parsed and pad_zip (PO-only; WO has no enrichment stage) with zero behavior change. telemetry.py holds the EMF ParseMethod emit wrappers. persistence.py holds _write_fields/_merge_update/save_*, collapsing the byte-identical save_new_po/save_revision bodies into one _save_merge helper that both now call through, preserving the sticky Cancelled ConditionExpression guard for both callers; save_cancellation stays separate. WO (5 concerns, no enrichment stage): the handler loop keeps validate_ai_fallback and the re.fullmatch(r"[0-9]+", work_order_id) key guard ahead of both save_work_order and save_event, since the guard protects the DynamoDB partition key and the '#'-delimited comment_id range-key segment. _header_date_iso and comment_id determinism stay colocated with persistence.py's save_event for the retry-idempotent event_id key. EXTRACTION_PROMPT (PO) moves to prompts.py with cross-reference headers to derived_fields.py's authoritative trade/site/fiscal rule tables; handler.py re-exports it (from prompts import EXTRACTION_PROMPT) since four tests dereference handler.EXTRACTION_ PROMPT directly. WO's prompt moves the same way. I/O modules (extraction.py's bedrock client, persistence.py's dynamodb resource, handler.py's s3 client) get lazy cached boto3 accessors; pure modules (enrichment.py, prompts.py, telemetry.py) import no boto3. Test monkeypatch surfaces move to the module that now owns the client (e.g. persistence.dynamodb) everywhere tests patch it, and the moto-before-handler-import ordering in _po_parser_support.py is preserved so the moto-backed suites don't hit real AWS. Behavior-preservation pins, verified with tests: PO still emits ParseMethod=ai_fallback before the Bedrock call, with ai_fallback_rejected as the additive second datapoint on rejection. WO still emits after its gate with mutually-exclusive ai_fallback / ai_fallback_rejected. Shadow DerivedFieldAgreement telemetry stays ai_fallback-only. derived_fields.py is untouched (diff against feature/phase-3-shared-extraction is empty). handler(event, context) signatures and the save_* public contract are unchanged on both pipelines; goldens unchanged. PO_EXPECTED_TOP_LEVEL_MODULES and its WO equivalent in tests/test_bundle_consistency.py are updated for the new sibling modules so the AST bundle-consistency test still fails on an unshipped or uncommented-out sibling.
2026-07-20 15:34:53 -04:00
"""Bedrock AI-fallback extraction for the work-order email processor.
Owns the Bedrock model id, the <email>-tag neutralizer, and the
extract_with_bedrock call. The boto3 bedrock-runtime client is built lazily on
first use so tests can patch this module's ``bedrock`` attribute before any real
client is constructed (moto-before-handler invariant).
"""
import json
import os
import re
import boto3
from prompts import EXTRACTION_PROMPT
BEDROCK_MODEL_ID = os.environ.get(
"BEDROCK_MODEL_ID", "us.anthropic.claude-haiku-4-5-20251001-v1:0"
)
# An <email>/</email> (or whitespace-padded variant) appearing INSIDE the
# untrusted email text could forge the data-block boundary, so any such
# sequence is neutralized before wrapping. A single [\s/]* class (not two
# \s* around an optional /) keeps matching linear -- the two-quantifier form
# backtracks quadratically on "<" + a long whitespace run (attacker DoS).
_EMAIL_TAG_RE = re.compile(r"<[\s/]*email\b", re.IGNORECASE)
# Lazy cached Bedrock client. Keeps the public attribute name ``bedrock`` so the
# test monkeypatch target changes module only, not attribute name.
bedrock = None
def _get_bedrock():
global bedrock
if bedrock is None:
bedrock = boto3.client("bedrock-runtime")
return bedrock
def extract_with_bedrock(email_data: dict) -> dict:
"""Send parsed email to Claude on Bedrock for structured extraction.
The untrusted email body is wrapped in an explicit XML-tagged data block
(<email>) to delimit data from instructions; <email>-tag lookalikes inside
the untrusted text are neutralized so the boundary cannot be forged. The
system prompt instructs the model to treat the block as data only, which
(combined with the downstream validate_ai_fallback gate) defends against
prompt injection from DKIM-passing but attacker-controlled email bodies.
"""
email_text = (
f"Subject: {email_data['subject']}\n"
f"From: {email_data['sender']}\n"
f"To: {email_data['to']}\n"
f"CC: {email_data['cc']}\n"
f"Date: {email_data['date']}\n"
f"\n---\n\n"
f"{email_data['body']}"
)
email_text = _EMAIL_TAG_RE.sub("[email-tag]", email_text)
resp = _get_bedrock().invoke_model(
modelId=BEDROCK_MODEL_ID,
body=json.dumps(
{
"anthropic_version": "bedrock-2023-05-31",
"max_tokens": 1024,
# Greedy decoding: retries of the same email should get the
# same extraction back (advisory A1). Not a hard guarantee of
# determinism, so model output still never enters a table key.
"temperature": 0,
"messages": [
{
"role": "user",
"content": (
f"{EXTRACTION_PROMPT}\n\n<email>\n{email_text}\n</email>"
),
}
],
}
),
)
response_text = json.loads(resp["body"].read())["content"][0]["text"]
# Extract JSON from response (handle markdown code blocks)
json_match = re.search(r"```(?:json)?\s*(.*?)```", response_text, re.DOTALL)
if json_match:
response_text = json_match.group(1)
return json.loads(response_text.strip())