mirror of
https://github.com/Sea-Haven-Industries/procurement-ingest.git
synced 2026-09-30 08:23:14 +00:00
Some checks are pending
Deploy / deploy (push) Waiting to run
* Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * feature: Add PO template parser scaffold and design doc Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage: - coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands. - coupa_cancellation (2.9%): implemented. Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work. Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com> * Implement PO new_po extraction and value-level gate Replace the extract_new_po scaffold stub with the full section-windowed extractor (duplicate-label anchoring, sentinel ship-to, label-keyed U+2022 bullet split, Decimal money from three anchored contexts only) and add value-level gate rules V1-V13. Both new_po_not_implemented scaffold guards are removed; rules 6-8 (unrecognized_status, multiline_unsupported, non_usd) go live. The gate re-derives every byte proof from the email body so an extractor bug cannot vouch for itself: amount re-serialization with a digit/comma border check (the thousands-separator truncation kill switch), sum(lines)==total against both Total blocks, anchor/supplier identity proofs, USPS address shape on the raw pre-enrichment zip, bullet label discipline, and sentinel/artifact hygiene. Any failure falls closed to the LLM; a validation failure is never a parsed result. Refs: #99 * Wire template-first parse into PO handler with EMF metric Run try_deterministic_parse ahead of the Bedrock extractor and fall back only on a miss/invalid (fail-closed) result. The shared enrich_parsed post-stage and the save_cancellation/save_revision/ save_new_po routing are untouched, so both paths write identical DynamoDB shapes and the po-ingest-site-extractor stream contract is preserved. Each record emits one ParseMethod EMF line (Seahaven/PoIngest/ ParseOutcome, dimension sets [ParseMethod] and [ParseMethod,TemplateId], ReasonCode/po_number ride-alongs) mirroring the WO idiom. The metric fires before the Bedrock call so a Bedrock-side error still records the ai_fallback outcome. Refs: #99 * Add PO fallback-rate alarm retuned for ~57 emails/day The WO alarm's 15-min period and >=10-sample floor assume ~760/day and would be structurally dead at PO volume (a 15-min period holds ~0.6 emails, so the floor is never met). Retune: 6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...) volume floor so a single email can never breach a datapoint (1/8 = 12.5% < 20%), threshold >20% against a ~1% expected baseline, eval 4 / datapoints 2 (24h span) so noise self-clears while total template drift pages within ~12h. No element-wise MAX in the math expression (post-#102 rule); ALARM-only SnsAction to site-alerts, NOT_BREACHING. Gated with 'npx cdk synth po-ingest'. Also add template_parser.py to the bundling cp list -- without it every deployed invocation would ImportError (unit tests cannot catch an asset-bundling omission). Refs: #99, #102 * Add offline PO parser suite with scrubbed fixture corpus 132 tests: golden-file comparison for all 25 positive fixtures (17 single-line new-PO + 8 cancellations, Decimal-exact via parse_float=Decimal), every fail-closed gate reason code covered (body-level triggers via 17 synthetic adversarial .eml mutations, candidate-level via direct validate() unit tests), real multi-line and comment/non-Coupa fallback fixtures, dual line-ending parse identity, two-path enrich/save parity (site-extractor stream guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass + scrub-marker leak sweep), and Bedrock dispatch/EMF assertions. The suite loads handler/template_parser via importlib under unique module names and binds the handler's bare sibling imports around exec (tests/conftest.py load_handler gets the same treatment) -- the WO suite caches bare 'handler'/'template_parser' names in sys.modules, and bare imports here would silently bind to the wrong pipeline. moto is imported before the handler so its botocore stubber hook precedes boto3 session creation (the PO conftest chain now loads at pytest session start). Fixtures are scrubbed real S3 samples: transport/auth header values replaced with same-shape placeholders (structure kept so ses_auth still passes), per-file digit ciphers, amounts remapped with sum==total re-established. The .gitignore exception is scoped to the PO fixtures path only. Refs: #99 * Document PO template-first parser and retuned alarm README: PO flow is now template-first with Bedrock fallback; parser/gate section mirroring the WO writeup; Seahaven/PoIngest ParseOutcome namespace and the fallback-rate alarm numbers with their volume justification (deliberately not WO's settings); test-suite and repo-layout updates. Design doc: mark PR #1 complete in progress/checklist sections; document the six value-level gate reason codes and the scaffold guard removal; correct the stale data-access note (default CLI session is 328440206208) and note the ~90-day S3 lifecycle aging of the corpus; record the 2.3 layout addendum (leading Supplier bullet segment, EA evidence lines, summary unit-price tokens, decode-path line endings), the fixture-build pins (address join convention, quantity/unit/price source), the V10 sweep outcome, and resolutions for open questions Q3/Q6. Cross-family review and the Confluence architecture-map update are flagged outstanding for merge. Refs: #99 * Record cross-family review outcome for handler wiring GPT-4.1 cross_review.py run against the real handler diff returned no BLOCK and no security findings; both FIX items verified as no-change-needed (fallback logging already correct; non-dict AI output is the pre-existing issue #101 pattern this PR deliberately does not touch). Refs: #99 * Pin line-item currency to USD in the PO gate The non_usd rule only checked the Total-block top-level currency, so a new_po whose line item read 'for 55,206.00 CAD' under a USD Total block still template-parsed as ok -- a fail-open hole in the fail-closed gate. Every line item's captured currency and its re-derived body token must now byte-equal the proven-USD top-level currency; covered by a line-level CAD adversarial fixture (the existing adv-non-usd only exercised the Total-block variant) and a candidate-mutation unit test. * Scrub residual transport tokens from PO fixtures The first-pass harvest scrub sanitized only the primary SES/DKIM header blocks, leaving the real SES Feedback-ID sender-identity hash in 49 committed fixtures and, on the two non-Coupa fixtures, an embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the token classes the PR #99 fixture lesson requires placeholdered. Replace each with a same-shape ScrubbedFixture value (byte-safe, CRLF and folding preserved) so header structure and ses_auth behavior are unchanged. * Converge quantity/price to Decimal on both paths EXTRACTION_PROMPT declares quantity and price as JSON strings, so a prompt-obedient Bedrock response stores DynamoDB Strings where the template parser stores Numbers -- divergent attribute types for the same email on the purchase-orders stream. Coerce numeric strings to Decimal in the shared enrich_parsed post-stage (thousands-separator safe; non-numeric strings kept verbatim) so both paths converge; prompt rewording itself remains PR #2 scope. The two-path parity test was circular -- it replayed the parser- derived golden as 'the LLM output', so it could never see the type divergence. It now feeds a prompt-shaped payload (string quantity/ price, LLM-filled site_code) through enrich_parsed and save_new_po, and the fixture-hygiene test now asserts the scrubbed transport-token header classes so fixture regressions are caught. * Coerce bare-int quantity/price to Decimal in enrich_parsed GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged residual type drift: parse_float=Decimal rules out floats on the LLM path, but a bare JSON int survived as Python int. Coerce it so both parse paths emit one canonical Decimal type. * Scrub fixture-body PII and harden cancellation gate (sec review) /sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill verifier) confirmed two diff-introduced findings; both fixed here. F3 (medium, real PII in new fixtures): the harvest scrub replaced header tokens but left real third-party PII in message BODIES -- an Amazon contact's name/phone/personal email in non-coupa-02.eml and an internal t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn names recurring across the new_po corpus. Replaced every personal name, phone, personal email, and internal URL with synthetic placeholders (QP-soft-wrap aware) across both .eml bodies and expected goldens. Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs, the leaked tokens), closing the header-only gap that let this through. F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and matched with .search(), unlike the anchored new_po pattern -- a subject merely ending with the cancellation phrase could be routed to the sticky- Cancelled write. Fully anchored it and switched to .match, and added a body-corroboration gate (the real Coupa body independently restates 'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted subject whose body does not corroborate now fails closed to the LLM (new reason code cancellation_body_unconfirmed). Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po dispatch and undelimited extraction prompt (issue #101 family) are byte-identical to main and unchanged here. 401 tests pass; ruff/format clean; cdk synth po-ingest clean. --------- Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
990 lines
40 KiB
Python
990 lines
40 KiB
Python
"""Deterministic template parser for Coupa purchase order emails.
|
||
|
||
Pure module: no boto3, no network. Runs ahead of the AI extraction path in the
|
||
purchase-order email processor. Only returns a parsed result when it is proven
|
||
conformant to one of the known Coupa templates; otherwise it fails closed and
|
||
signals the caller to fall back to the Bedrock AI extractor.
|
||
|
||
Mirrors the WO parser idioms (lambdas/wo/email_processor/template_parser.py):
|
||
classify_template -> extract -> validate (FAIL CLOSED) -> try_deterministic_parse
|
||
returning (parsed|None, method, template_id, reason). A failure is NEVER a
|
||
parsed result.
|
||
|
||
Full-bucket triage of all 3,448 inbound emails (2026-07-16) fixed the scope:
|
||
|
||
T1 coupa_new_po -- 3,294 / 3,448 (95.5%). Subject
|
||
"***Copy for Reference*** New Purchase Order <PO#> has been issued".
|
||
multipart/alternative; the text/plain part is a stable label-delimited
|
||
layout (PO ID / Status / Order Date / Revision Date / Payment Term / Req # /
|
||
Submitted By / On Behalf Of / Supplier block / Shipping block w/ Location
|
||
Code + Attn / a Lines section whose per-line metadata is U+2022-delimited:
|
||
Need By / Category / Account / Period [/ optional Part Number]).
|
||
Maps to email_type "new_po". 99.8% single-line-item, 100% USD.
|
||
|
||
T2 coupa_cancellation -- 100 / 3,448 (2.9%). Subject
|
||
"<SiteName> Purchase Order #<PO#> has been cancelled". Minimal body; the
|
||
only field the handler's save_cancellation() needs is po_number.
|
||
Maps to email_type "cancellation".
|
||
|
||
DELIBERATELY OUT OF SCOPE -> always AI fallback (never template-parsed):
|
||
* "New Comment on Purchase Order for Amazon" (19 / 3,448) -- an email_type the
|
||
handler enum does not model; do not fabricate a PO record deterministically.
|
||
* revision -- 0 distinct emails in 3,448; no template to build.
|
||
* any multi-line-item new_po (6 / 3,448) -- Lines-array structure unobserved.
|
||
* any non-USD new_po (0 observed) -- non-USD path entirely unexercised.
|
||
* non-Coupa senders (35 / 3,448) -- already rejected at the ses_auth layer.
|
||
|
||
Derived fields (site_code, trade, fiscal_year) are NOT computed here: the parser
|
||
leaves them None and a shared post-stage (handler enrich_parsed(), the pad_zip
|
||
precedent) fills them identically on both the template and the LLM path, so the
|
||
gate judges extraction fidelity only. coupa_category is a verbatim label capture,
|
||
not a classifier, and IS extracted here.
|
||
|
||
Entry point: try_deterministic_parse(email_data) -> (parsed|None, method,
|
||
template_id, reason_code).
|
||
"""
|
||
|
||
import re
|
||
from decimal import Decimal, InvalidOperation
|
||
|
||
# ---------------------------------------------------------------------------
|
||
# IMPLEMENTATION STATUS
|
||
# [x] contract + recursive skeleton/normalize
|
||
# [x] classify_template (both templates)
|
||
# [x] coupa_cancellation extract + gate
|
||
# [x] coupa_new_po extract -- labeled fields, duplicate-label anchoring,
|
||
# U+2022 line split, Decimal amounts (built against the scrubbed real
|
||
# fixture corpus in tests/fixtures/)
|
||
# [x] coupa_new_po value-level gate rules V1-V13 (see validate())
|
||
# ---------------------------------------------------------------------------
|
||
|
||
# The contract keys (23), EXACTLY -- mirrors the AI EXTRACTION_PROMPT fields.
|
||
CONTRACT_KEYS = (
|
||
"email_type",
|
||
"po_number",
|
||
"po_status",
|
||
"source_system",
|
||
"submitted_by",
|
||
"on_behalf_of",
|
||
"order_date",
|
||
"revision_date",
|
||
"last_opened",
|
||
"acknowledged_at",
|
||
"payment_terms",
|
||
"requisition_number",
|
||
"department",
|
||
"view_order_url",
|
||
"supplier",
|
||
"site_code",
|
||
"ship_to",
|
||
"total_amount",
|
||
"currency",
|
||
"fiscal_year",
|
||
"trade",
|
||
"coupa_category",
|
||
"line_items",
|
||
)
|
||
|
||
SUPPLIER_KEYS = ("name",)
|
||
|
||
SHIP_TO_KEYS = (
|
||
"name",
|
||
"address",
|
||
"street",
|
||
"city",
|
||
"state",
|
||
"zip",
|
||
"location_code",
|
||
"attn",
|
||
)
|
||
|
||
LINE_ITEM_KEYS = (
|
||
"description",
|
||
"amount",
|
||
"currency",
|
||
"need_by",
|
||
"category",
|
||
"account_code",
|
||
"period",
|
||
"quantity",
|
||
"unit",
|
||
"price",
|
||
)
|
||
|
||
# Derived fields the parser must leave None (filled by the shared post-stage).
|
||
DERIVED_KEYS = ("site_code", "trade", "fiscal_year")
|
||
|
||
SOURCE_SYSTEM = "coupa"
|
||
|
||
# email_type per template.
|
||
_TEMPLATE_EMAIL_TYPE = {
|
||
"coupa_new_po": "new_po",
|
||
"coupa_cancellation": "cancellation",
|
||
}
|
||
VALID_EMAIL_TYPES = {"new_po", "revision", "cancellation"}
|
||
|
||
# Only these Status strings were observed on new_po (2,223 + 1,071 of 3,294).
|
||
# Anything else fails closed to the LLM -- NEVER default-to-new_po.
|
||
NEW_PO_SAFE_STATUSES = {"Issued - Created", "Issued - Scheduled for email"}
|
||
|
||
# Subject classifiers (Python unfolds header continuation lines before we see them).
|
||
_NEW_PO_SUBJECT = re.compile(
|
||
r"^\*\*\*Copy for Reference\*\*\* New Purchase Order\s+(?P<po>\S+)\s+has been issued$"
|
||
)
|
||
# Fully anchored, symmetric with _NEW_PO_SUBJECT: the whole subject must be
|
||
# "<SiteName> Purchase Order #<PO#> has been cancelled" -- the SiteName prefix
|
||
# is bounded (no '#', no newline, <=80 chars) so a subject that merely *ends*
|
||
# with the cancellation phrase (e.g. a forwarded/quoted thread, or arbitrary
|
||
# prefix text before the tail) is NOT misclassified as a cancellation and
|
||
# routed to the sticky-Cancelled write. Matched with .match (see
|
||
# classify_template), never .search.
|
||
_CANCELLATION_SUBJECT = re.compile(
|
||
r"^(?P<site>[^#\n]{0,80}?)Purchase Order\s+#(?P<po>[A-Z0-9-]+)\s+has been cancelled\s*$"
|
||
)
|
||
|
||
# PO number shape, e.g. 2D-21456967, FK-21920384, B187-17955555.
|
||
_PO_ID_RE = re.compile(r"^[A-Z0-9]{1,6}-\d+$")
|
||
|
||
# U+2022 bullet delimiting per-line metadata in the Lines section.
|
||
_BULLET = "•"
|
||
|
||
# --- new_po layout patterns (built against the scrubbed real fixture corpus) ---
|
||
# Money tokens are read ONLY from three anchored contexts: the Lines-section
|
||
# '<desc> for <amt> <CCY>' line, the Total-block standalone amount line, and the
|
||
# Items-summary '<qty> <UNIT> x <price>' line. NEVER free money-shaped scanning:
|
||
# the Items summary carries unit-price tokens distinct from line amounts.
|
||
_MONEY_RE = re.compile(r"\d{1,3}(?:,\d{3})*\.\d{2}")
|
||
_CURRENCY_RE = re.compile(r"[A-Z]{3}")
|
||
# Items-summary quantity line, e.g. '1.0 EACH x 55,206.00'.
|
||
_SUMMARY_ITEM_RE = re.compile(
|
||
r"^(?P<qty>\d+(?:\.\d+)?) (?P<unit>[A-Z]+) x (?P<price>\d{1,3}(?:,\d{3})*\.\d{2})$"
|
||
)
|
||
# Lines-block quantity evidence line, e.g. '1.0 EA' (gate rule V13 cross-check).
|
||
_LINE_QTY_RE = re.compile(r"^(?P<qty>\d+(?:\.\d+)?) (?P<unit>[A-Z]+)$")
|
||
# Lines-block description/amount line. GREEDY desc: '.+' binds the LAST ' for ',
|
||
# so a description containing the word 'for' can never shift the amount.
|
||
_LINE_DESC_AMT_RE = re.compile(
|
||
r"^(?P<desc>.+) for (?P<amt>\d{1,3}(?:,\d{3})*\.\d{2}) (?P<cur>[A-Z]{3})$"
|
||
)
|
||
_VIEW_ORDER_URL_RE = re.compile(r"^https://supplier\.coupahost\.com/orders/\S+$")
|
||
_ORDER_URL_ID_RE = re.compile(r"^https://supplier\.coupahost\.com/orders/(\d+)\b")
|
||
_LOCATION_CODE_RE = re.compile(r"^Location Code: (?P<lc>\d+)$")
|
||
_ATTN_RE = re.compile(r"^Attn: (?P<attn>.+)$")
|
||
# Ship-to city line immediately preceding 'United States'.
|
||
_CITY_STATE_ZIP_RE = re.compile(
|
||
r"^(?P<city>.+), (?P<state>[A-Z]{2}) (?P<zip>\d{5}(?:-\d{4})?)$"
|
||
)
|
||
_US_SENTINEL = "United States"
|
||
_SUPPLIER_MARKER = (
|
||
"SEA HAVEN" # case-sensitive; drift becomes fallback, never wrong data
|
||
)
|
||
|
||
# More Detail block: label lines, each expected exactly once (gate rule V9).
|
||
_MORE_DETAIL_FIELDS = {
|
||
"Department": "department",
|
||
"Status": "po_status",
|
||
"Last Opened": "last_opened",
|
||
"Order Date": "order_date",
|
||
"Acknowledged At": "acknowledged_at",
|
||
"Revision Date": "revision_date",
|
||
"Payment Term": "payment_terms",
|
||
"Req #": "requisition_number",
|
||
}
|
||
_MORE_DETAIL_LABELS = ("PO ID", *_MORE_DETAIL_FIELDS)
|
||
|
||
# Per-line bullet metadata: closed label set, assigned purely by leading label
|
||
# (longest label first), NEVER by ordinal position -- real data has an optional
|
||
# 'Part Number' segment between 'Category' and 'Account', and every run starts
|
||
# with a 'Supplier <name>' segment.
|
||
_BULLET_LABELS = ("Part Number", "Need By", "Category", "Account", "Period", "Supplier")
|
||
_BULLET_FIELDS = {
|
||
"Need By": "need_by",
|
||
"Category": "category",
|
||
"Account": "account_code",
|
||
"Period": "period",
|
||
}
|
||
|
||
# Frozen USPS state/territory codes (50 states + DC + territories).
|
||
_USPS_STATES = frozenset(
|
||
"""AL AK AZ AR CA CO CT DE FL GA HI ID IL IN IA KS KY LA ME MD MA MI MN MS
|
||
MO MT NE NV NH NJ NM NY NC ND OH OK OR PA RI SC SD TN TX UT VT VA WA WV WI
|
||
WY DC PR VI GU AS MP""".split()
|
||
)
|
||
|
||
# Sentinel: a label was present but its value did not parse. Must FAIL the gate
|
||
# (present-but-unparseable), distinct from an absent value (None).
|
||
_UNPARSEABLE = "__UNPARSEABLE__"
|
||
|
||
|
||
# ---------------------------------------------------------------------------
|
||
# Skeleton helpers (recursive -- unlike WO's flat contract)
|
||
# ---------------------------------------------------------------------------
|
||
def _empty_line_item():
|
||
return {k: None for k in LINE_ITEM_KEYS}
|
||
|
||
|
||
def _empty_candidate():
|
||
"""Full nested skeleton: every contract key present, None where absent."""
|
||
cand = {k: None for k in CONTRACT_KEYS}
|
||
cand["supplier"] = {k: None for k in SUPPLIER_KEYS}
|
||
cand["ship_to"] = {k: None for k in SHIP_TO_KEYS}
|
||
cand["line_items"] = [_empty_line_item()]
|
||
cand["source_system"] = SOURCE_SYSTEM
|
||
return cand
|
||
|
||
|
||
def _normalize(candidate):
|
||
"""Guarantee exact nested key presence before returning."""
|
||
out = _empty_candidate()
|
||
for k in CONTRACT_KEYS:
|
||
if k in candidate and k not in ("supplier", "ship_to", "line_items"):
|
||
out[k] = candidate[k]
|
||
supplier = candidate.get("supplier") or {}
|
||
out["supplier"] = {k: supplier.get(k) for k in SUPPLIER_KEYS}
|
||
ship_to = candidate.get("ship_to") or {}
|
||
out["ship_to"] = {k: ship_to.get(k) for k in SHIP_TO_KEYS}
|
||
items = candidate.get("line_items") or [{}]
|
||
out["line_items"] = [{k: (it or {}).get(k) for k in LINE_ITEM_KEYS} for it in items]
|
||
return out
|
||
|
||
|
||
def _clean(value):
|
||
"""Strip trailing CR and U+00A0 nbsp that every captured Coupa value carries;
|
||
collapse nothing else. Returns None for empty/placeholder 'None'."""
|
||
if value is None:
|
||
return None
|
||
v = value.replace("\r", "").replace(" ", " ").strip()
|
||
if v == "" or v == "None":
|
||
return None
|
||
return v
|
||
|
||
|
||
def _plain_lines(body):
|
||
return body.replace("\r\n", "\n").replace("\r", "\n").split("\n")
|
||
|
||
|
||
def _to_decimal(raw):
|
||
"""Parse a Coupa money token ('18,624.05') to Decimal, stripping thousands
|
||
separators. Returns _UNPARSEABLE if it does not parse (gate must reject)."""
|
||
if raw is None:
|
||
return None
|
||
token = raw.replace(",", "").strip()
|
||
try:
|
||
return Decimal(token)
|
||
except (InvalidOperation, ValueError):
|
||
return _UNPARSEABLE
|
||
|
||
|
||
# ---------------------------------------------------------------------------
|
||
# Line navigation helpers (shared by the extractor and the gate; both operate
|
||
# on email_data["body"] only -- the gate re-derives its own byte evidence and
|
||
# never trusts extractor-carried state it can re-derive)
|
||
# ---------------------------------------------------------------------------
|
||
def _visible(line):
|
||
"""True when the raw line carries visible content. The literal placeholder
|
||
'None' IS visible (unlike _clean, which maps it to None), so label values
|
||
of 'None' are found -- not skipped over into the next label line."""
|
||
return bool(line.replace("\r", "").replace("\xa0", "").strip())
|
||
|
||
|
||
def _indices(lines, label):
|
||
"""All indices whose _clean-ed content equals the label exactly."""
|
||
return [i for i, ln in enumerate(lines) if _clean(ln) == label]
|
||
|
||
|
||
def _find_after(lines, label, start):
|
||
"""First index >= start whose _clean-ed content equals label, or None."""
|
||
for i in range(start, len(lines)):
|
||
if _clean(lines[i]) == label:
|
||
return i
|
||
return None
|
||
|
||
|
||
def _next_visible(lines, idx, end=None):
|
||
"""Index of the first visible line strictly after idx (before end), or None."""
|
||
stop = len(lines) if end is None else min(end, len(lines))
|
||
for j in range(idx + 1, stop):
|
||
if _visible(lines[j]):
|
||
return j
|
||
return None
|
||
|
||
|
||
def _split_bullet_segments(raw_line):
|
||
"""Split a Lines-section metadata line on the bare U+2022 bullet; strip each
|
||
segment of spaces and nbsp; drop empty segments."""
|
||
segments = []
|
||
for seg in raw_line.replace("\r", "").split(_BULLET):
|
||
seg = seg.replace("\xa0", " ").strip()
|
||
if seg:
|
||
segments.append(seg)
|
||
return segments
|
||
|
||
|
||
def _match_bullet_label(segment):
|
||
"""(label, value) by leading-label prefix match (longest label first)
|
||
against the closed _BULLET_LABELS set, or (None, None) if unrecognized."""
|
||
for label in sorted(_BULLET_LABELS, key=len, reverse=True):
|
||
if segment == label:
|
||
return label, None
|
||
if segment.startswith(label + " "):
|
||
return label, segment[len(label) :].strip()
|
||
return None, None
|
||
|
||
|
||
def _walk_leaves(obj):
|
||
"""Yield every scalar leaf of a nested dict/list candidate."""
|
||
if isinstance(obj, dict):
|
||
for v in obj.values():
|
||
yield from _walk_leaves(v)
|
||
elif isinstance(obj, list):
|
||
for v in obj:
|
||
yield from _walk_leaves(v)
|
||
else:
|
||
yield obj
|
||
|
||
|
||
# ---------------------------------------------------------------------------
|
||
# Subject helpers
|
||
# ---------------------------------------------------------------------------
|
||
def classify_template(email_data):
|
||
"""Return (template_id, reason). template_id in
|
||
{coupa_new_po, coupa_cancellation, unknown}."""
|
||
subject = _clean(email_data.get("subject")) or ""
|
||
if _NEW_PO_SUBJECT.match(subject):
|
||
return "coupa_new_po", "ok"
|
||
if _CANCELLATION_SUBJECT.match(subject):
|
||
return "coupa_cancellation", "ok"
|
||
return "unknown", "subject_no_match"
|
||
|
||
|
||
def _subject_po_id(email_data):
|
||
"""PO number parsed from the subject, or None."""
|
||
subject = _clean(email_data.get("subject")) or ""
|
||
m = _NEW_PO_SUBJECT.match(subject)
|
||
if m:
|
||
return m.group("po")
|
||
m = _CANCELLATION_SUBJECT.match(subject)
|
||
if m:
|
||
return m.group("po")
|
||
return None
|
||
|
||
|
||
# ---------------------------------------------------------------------------
|
||
# T2: coupa_cancellation -> cancellation (simple + safe: po_number only)
|
||
# ---------------------------------------------------------------------------
|
||
def extract_cancellation(email_data):
|
||
"""Cancellation carries no PO detail we trust beyond the id; the handler's
|
||
save_cancellation() only needs po_number. Everything else stays None."""
|
||
candidate = _empty_candidate()
|
||
candidate["email_type"] = "cancellation"
|
||
candidate["po_number"] = _subject_po_id(email_data)
|
||
return _normalize(candidate)
|
||
|
||
|
||
# ---------------------------------------------------------------------------
|
||
# T1: coupa_new_po -> new_po
|
||
# ---------------------------------------------------------------------------
|
||
def _assign_bullet_metadata(item, raw_line):
|
||
"""Assign the U+2022 metadata segments to the item BY LEADING LABEL.
|
||
|
||
'Supplier' is recognized but not stored on the item (gate rule V5 proves it
|
||
byte-equals supplier.name from the body); 'Part Number' is recognized but
|
||
discarded (LINE_ITEM_KEYS has no slot -- inventing one would break the
|
||
key_set_mismatch rule and LLM-path shape parity). Unrecognized or duplicate
|
||
segments are the gate's job to reject (rule V8)."""
|
||
for segment in _split_bullet_segments(raw_line):
|
||
label, value = _match_bullet_label(segment)
|
||
field = _BULLET_FIELDS.get(label)
|
||
if field and item[field] is None:
|
||
item[field] = _clean(value)
|
||
|
||
|
||
def extract_new_po(email_data):
|
||
"""Extract the new_po contract from the text/plain body.
|
||
|
||
Permissive capture, section-windowed: each anchor line is located by exact
|
||
_clean-ed full-line equality, STRICTLY AFTER the previous anchor. Missing
|
||
anchors leave fields None -- the gate then fails closed. Representation-
|
||
agnostic: matches only on _clean-ed lines, never on '\\r'-suffixed literals
|
||
(body line endings are decode-path dependent).
|
||
|
||
Duplicate labels: 'Supplier', 'Shipping', 'Total' each appear TWICE (summary
|
||
placeholder + detail block; the first 'Shipping' value is literally 'None').
|
||
supplier.name anchors on the FIRST 'Supplier'; ship_to on the SECOND
|
||
'Shipping'; total on the SECOND 'Total'.
|
||
|
||
DERIVED_KEYS (site_code, trade, fiscal_year) stay None -- filled later by
|
||
the shared post-stage identically on both paths.
|
||
"""
|
||
candidate = _empty_candidate()
|
||
candidate["email_type"] = "new_po"
|
||
candidate["po_number"] = _subject_po_id(email_data)
|
||
|
||
lines = _plain_lines(email_data["body"])
|
||
|
||
# --- Summary section (start .. 'More Detail') ---
|
||
more_detail = _find_after(lines, "More Detail", 0)
|
||
summary_end = more_detail if more_detail is not None else len(lines)
|
||
for label, field in (
|
||
("Submitted By", "submitted_by"),
|
||
("On Behalf Of", "on_behalf_of"),
|
||
):
|
||
idx = _find_after(lines, label, 0)
|
||
if idx is not None and idx < summary_end:
|
||
j = _next_visible(lines, idx, summary_end)
|
||
if j is not None:
|
||
candidate[field] = _clean(lines[j])
|
||
sup1 = _find_after(lines, "Supplier", 0)
|
||
if sup1 is not None and sup1 < summary_end:
|
||
j = _next_visible(lines, sup1, summary_end)
|
||
if j is not None:
|
||
candidate["supplier"]["name"] = _clean(lines[j])
|
||
# view_order_url: the unique orders link in the summary (0 or >1 -> None).
|
||
url_lines = [
|
||
_clean(lines[i])
|
||
for i in range(summary_end)
|
||
if _VIEW_ORDER_URL_RE.match(_clean(lines[i]) or "")
|
||
]
|
||
if len(url_lines) == 1:
|
||
candidate["view_order_url"] = url_lines[0]
|
||
# Items-summary quantity lines ('1.0 EACH x 55,206.00'), collected in order.
|
||
summary_items = [
|
||
m
|
||
for i in range(summary_end)
|
||
if (m := _SUMMARY_ITEM_RE.fullmatch(_clean(lines[i]) or ""))
|
||
]
|
||
|
||
# --- More Detail block ('More Detail' .. second 'Supplier') ---
|
||
sup2 = None
|
||
if more_detail is not None:
|
||
sup2 = _find_after(lines, "Supplier", more_detail + 1)
|
||
md_end = sup2 if sup2 is not None else len(lines)
|
||
for label, field in _MORE_DETAIL_FIELDS.items():
|
||
idx = _find_after(lines, label, more_detail + 1)
|
||
if idx is not None and idx < md_end:
|
||
j = _next_visible(lines, idx, md_end)
|
||
if j is not None:
|
||
candidate[field] = _clean(lines[j])
|
||
# The FIRST 'Shipping' lives in this block; its value must be the literal
|
||
# 'None' placeholder (gate rule V4 tripwire) and is never used for ship_to.
|
||
|
||
# --- ship_to (second 'Shipping' .. 'Lines'), sentinel-anchored ---
|
||
ship2 = _find_after(lines, "Shipping", sup2 + 1) if sup2 is not None else None
|
||
lines_anchor = _find_after(lines, "Lines", ship2 + 1) if ship2 is not None else None
|
||
st_end = lines_anchor if lines_anchor is not None else len(lines)
|
||
ship_to = candidate["ship_to"]
|
||
if ship2 is not None:
|
||
name_idx = _next_visible(lines, ship2, st_end)
|
||
if name_idx is not None:
|
||
ship_to["name"] = _clean(lines[name_idx])
|
||
us_idx = _find_after(lines, _US_SENTINEL, name_idx + 1)
|
||
if us_idx is not None and us_idx < st_end:
|
||
city_idx = us_idx - 1
|
||
m = None
|
||
if city_idx > name_idx:
|
||
m = _CITY_STATE_ZIP_RE.fullmatch(_clean(lines[city_idx]) or "")
|
||
if m:
|
||
ship_to["city"] = m.group("city")
|
||
ship_to["state"] = m.group("state")
|
||
ship_to["zip"] = m.group("zip")
|
||
street = [
|
||
_clean(lines[j])
|
||
for j in range(name_idx + 1, city_idx)
|
||
if _visible(lines[j])
|
||
]
|
||
if street:
|
||
ship_to["street"] = "\n".join(street)
|
||
ship_to["address"] = "\n".join(
|
||
_clean(lines[j])
|
||
for j in range(name_idx, us_idx + 1)
|
||
if _visible(lines[j])
|
||
)
|
||
for j in range(us_idx + 1, st_end):
|
||
cl = _clean(lines[j]) or ""
|
||
lc = _LOCATION_CODE_RE.fullmatch(cl)
|
||
if lc and ship_to["location_code"] is None:
|
||
ship_to["location_code"] = lc.group("lc")
|
||
attn = _ATTN_RE.fullmatch(cl)
|
||
if attn and ship_to["attn"] is None:
|
||
ship_to["attn"] = _clean(attn.group("attn"))
|
||
|
||
# --- Lines section ('Lines' .. second 'Total') ---
|
||
# Item blocks are delimited by the lone U+00A0 line(s). EVERY block is
|
||
# extracted, even when >1, so gate rule 7 fires with honest
|
||
# multiline_unsupported data (never silently keep item 0).
|
||
total2 = (
|
||
_find_after(lines, "Total", lines_anchor + 1)
|
||
if lines_anchor is not None
|
||
else None
|
||
)
|
||
items = []
|
||
if lines_anchor is not None:
|
||
end = total2 if total2 is not None else len(lines)
|
||
blocks, block = [], []
|
||
for j in range(lines_anchor + 1, end):
|
||
if lines[j].replace("\r", "") == "\xa0":
|
||
blocks.append(block)
|
||
block = []
|
||
else:
|
||
block.append(lines[j])
|
||
blocks.append(block)
|
||
for block in blocks:
|
||
visible = [ln for ln in block if _visible(ln)]
|
||
if not visible:
|
||
continue
|
||
item = _empty_line_item()
|
||
for raw in visible:
|
||
dm = _LINE_DESC_AMT_RE.fullmatch(_clean(raw) or "")
|
||
if dm and item["description"] is None:
|
||
item["description"] = _clean(dm.group("desc"))
|
||
item["amount"] = _to_decimal(dm.group("amt"))
|
||
item["currency"] = dm.group("cur")
|
||
elif _BULLET in raw:
|
||
_assign_bullet_metadata(item, raw)
|
||
# The optional '<qty> EA' evidence line is not stored: quantity/
|
||
# unit/price come from the Items summary; gate rule V13
|
||
# cross-checks the EA line against it from the body.
|
||
items.append(item)
|
||
if items:
|
||
candidate["line_items"] = items
|
||
# coupa_category is a VERBATIM copy of item 0's Category bullet value.
|
||
candidate["coupa_category"] = items[0]["category"]
|
||
if len(summary_items) == 1 and len(items) == 1:
|
||
m = summary_items[0]
|
||
items[0]["quantity"] = _to_decimal(m.group("qty"))
|
||
items[0]["unit"] = m.group("unit")
|
||
items[0]["price"] = _to_decimal(m.group("price"))
|
||
|
||
# --- Total block (second 'Total' .. end) ---
|
||
if total2 is not None:
|
||
amt_idx = _next_visible(lines, total2)
|
||
if amt_idx is not None:
|
||
token = _clean(lines[amt_idx]) or ""
|
||
if _MONEY_RE.fullmatch(token):
|
||
candidate["total_amount"] = _to_decimal(token)
|
||
cur_idx = _next_visible(lines, amt_idx)
|
||
if cur_idx is not None:
|
||
cur_token = _clean(lines[cur_idx]) or ""
|
||
if _CURRENCY_RE.fullmatch(cur_token):
|
||
candidate["currency"] = cur_token
|
||
|
||
# DERIVED_KEYS intentionally left None (shared post-stage fills them).
|
||
return _normalize(candidate)
|
||
|
||
|
||
# ---------------------------------------------------------------------------
|
||
# Validation gate -- FAIL CLOSED
|
||
# ---------------------------------------------------------------------------
|
||
def _structural_keys_ok(candidate):
|
||
if set(candidate.keys()) != set(CONTRACT_KEYS):
|
||
return False
|
||
if set((candidate.get("supplier") or {}).keys()) != set(SUPPLIER_KEYS):
|
||
return False
|
||
if set((candidate.get("ship_to") or {}).keys()) != set(SHIP_TO_KEYS):
|
||
return False
|
||
items = candidate.get("line_items")
|
||
if not isinstance(items, list) or not items:
|
||
return False
|
||
return all(set((it or {}).keys()) == set(LINE_ITEM_KEYS) for it in items)
|
||
|
||
|
||
def validate(candidate, template_id, email_data):
|
||
"""Return (True, 'ok') only if provably conformant; else (False, reason).
|
||
Every rule must hold. See failClosedGateRules in the investigation report."""
|
||
# (1) known template
|
||
if template_id not in _TEMPLATE_EMAIL_TYPE:
|
||
return False, "subject_no_match"
|
||
|
||
# (2) exact nested key-set
|
||
if not _structural_keys_ok(candidate):
|
||
return False, "key_set_mismatch"
|
||
|
||
# (3) derived fields must NOT be populated by the parser
|
||
for k in DERIVED_KEYS:
|
||
if candidate.get(k) is not None:
|
||
return False, "derived_field_set"
|
||
|
||
# (4) email_type matches the template's expected type
|
||
expected_type = _TEMPLATE_EMAIL_TYPE[template_id]
|
||
et = candidate.get("email_type")
|
||
if et not in VALID_EMAIL_TYPES:
|
||
return False, "missing_required_field"
|
||
if et != expected_type:
|
||
return False, "email_type_mismatch"
|
||
|
||
# (5) po_number: valid shape AND byte-equals the subject id
|
||
subject_po = _subject_po_id(email_data)
|
||
po = candidate.get("po_number")
|
||
if not po or not _PO_ID_RE.match(str(po)):
|
||
return False, "missing_required_field"
|
||
if po != subject_po:
|
||
return False, "po_id_mismatch"
|
||
|
||
if template_id == "coupa_cancellation":
|
||
# Body corroboration: the subject SiteName prefix is free text, so the
|
||
# anchored subject alone cannot distinguish a genuine Coupa cancellation
|
||
# from an arbitrary "<prefix> Purchase Order #<po> has been cancelled"
|
||
# subject. The real Coupa body independently restates the id in a
|
||
# "Purchase Order #<po> ... has been cancelled" notice; require that
|
||
# (with the SAME po_number) before marking a PO sticky-Cancelled, so a
|
||
# misrouted/near-miss email fails closed to the LLM instead. (SEC review
|
||
# of PR #105, F1.)
|
||
body = email_data.get("body") or ""
|
||
restated = re.search(r"Purchase Order\s+#" + re.escape(str(po)) + r"\b", body)
|
||
if not restated or "cancelled" not in body.lower():
|
||
return False, "cancellation_body_unconfirmed"
|
||
return True, "ok"
|
||
|
||
# ---- coupa_new_po ----
|
||
# (6) status must be one of the confirmed-safe strings.
|
||
status = candidate.get("po_status")
|
||
if status not in NEW_PO_SAFE_STATUSES:
|
||
return False, "unrecognized_status"
|
||
|
||
# (7) single-line only -- multi-line Lines structure is unobserved.
|
||
if len(candidate.get("line_items") or []) != 1:
|
||
return False, "multiline_unsupported"
|
||
|
||
# (8) currency must be exactly USD (non-USD path entirely unexercised).
|
||
if candidate.get("currency") != "USD":
|
||
return False, "non_usd"
|
||
|
||
# (V1-V13) value-level rules: every byte proof is RE-DERIVED from
|
||
# email_data["body"] -- the gate never trusts extractor-carried state it
|
||
# can re-derive, so an extractor bug cannot vouch for itself.
|
||
return _validate_new_po_values(candidate, email_data)
|
||
|
||
|
||
def _money_border_ok(line_text, serialized):
|
||
"""True when `serialized` occurs in the source line with a character before
|
||
it that is not a digit or comma (the '18,624.05' -> '624.05' kill switch:
|
||
a truncated capture re-serializes as '624.05', but every occurrence of that
|
||
string in its source line is preceded by a comma or digit)."""
|
||
idx = line_text.find(serialized)
|
||
while idx != -1:
|
||
prev = line_text[idx - 1] if idx > 0 else ""
|
||
if prev not in "0123456789,":
|
||
return True
|
||
idx = line_text.find(serialized, idx + 1)
|
||
return False
|
||
|
||
|
||
def _validate_new_po_values(candidate, email_data): # noqa: PLR0911, PLR0912, PLR0915
|
||
"""Value-level gate rules V1-V13 for coupa_new_po. FAIL CLOSED.
|
||
|
||
Anchor integrity (V4) runs first because every later byte proof needs the
|
||
anchor frame; the remaining rules run in spec order. Each rule re-derives
|
||
its evidence from email_data["body"] -- duplicated regex work, negligible
|
||
at ~57 emails/day, in exchange for extractor-bug independence."""
|
||
lines = _plain_lines(email_data["body"])
|
||
items = candidate.get("line_items") or []
|
||
ship_to = candidate.get("ship_to") or {}
|
||
supplier_name = (candidate.get("supplier") or {}).get("name")
|
||
|
||
# --- V4 anchor integrity -> anchor_violation ---
|
||
more_detail_idxs = _indices(lines, "More Detail")
|
||
supplier_idxs = _indices(lines, "Supplier")
|
||
shipping_idxs = _indices(lines, "Shipping")
|
||
total_idxs = _indices(lines, "Total")
|
||
if len(more_detail_idxs) != 1:
|
||
return False, "anchor_violation"
|
||
if len(supplier_idxs) != 2 or len(shipping_idxs) != 2 or len(total_idxs) != 2:
|
||
return False, "anchor_violation"
|
||
more_detail = more_detail_idxs[0]
|
||
# The FIRST Shipping value must be the literal 'None' placeholder -- the
|
||
# dup-label-swap tripwire (a real address under the first label means the
|
||
# layout drifted; ship_to would have been read from the wrong block).
|
||
first_ship_val = ""
|
||
if shipping_idxs[0] + 1 < len(lines):
|
||
first_ship_val = (
|
||
lines[shipping_idxs[0] + 1].replace("\r", "").replace("\xa0", "").strip()
|
||
)
|
||
if first_ship_val != "None":
|
||
return False, "anchor_violation"
|
||
lines_anchor = _find_after(lines, "Lines", shipping_idxs[1] + 1)
|
||
if lines_anchor is None:
|
||
return False, "anchor_violation"
|
||
# Section ordering: summary Supplier < More Detail < detail Supplier <
|
||
# second Shipping < Lines < second Total; summary Total before More Detail.
|
||
if not (
|
||
supplier_idxs[0]
|
||
< more_detail
|
||
< supplier_idxs[1]
|
||
< shipping_idxs[1]
|
||
< lines_anchor
|
||
< total_idxs[1]
|
||
):
|
||
return False, "anchor_violation"
|
||
if not (total_idxs[0] < more_detail < shipping_idxs[0] < supplier_idxs[1]):
|
||
return False, "anchor_violation"
|
||
lines_end = total_idxs[1]
|
||
|
||
# --- V1 money fidelity -> amount_mismatch ---
|
||
# Every captured money value must re-locate its raw source token in the
|
||
# body: token fullmatches the grouped money shape, format(value, ',.2f')
|
||
# byte-equals it, and the character before it is not a digit/comma.
|
||
for_matches = []
|
||
for j in range(lines_anchor + 1, lines_end):
|
||
m = _LINE_DESC_AMT_RE.fullmatch(_clean(lines[j]) or "")
|
||
if m:
|
||
for_matches.append(m)
|
||
if len(for_matches) != len(items):
|
||
return False, "amount_mismatch"
|
||
for item, m in zip(items, for_matches):
|
||
amount = item.get("amount")
|
||
if not isinstance(amount, Decimal):
|
||
return False, "amount_mismatch"
|
||
serialized = format(amount, ",.2f")
|
||
if serialized != m.group("amt"):
|
||
return False, "amount_mismatch"
|
||
if not _money_border_ok(m.string, serialized):
|
||
return False, "amount_mismatch"
|
||
# Line-level currency pin (rule 8 companion): rule 8 already proved the
|
||
# top-level currency is exactly 'USD', but that only covers the Total
|
||
# block. Every line item's captured currency AND its re-derived body
|
||
# token must byte-equal it too -- a single non-USD line item fails
|
||
# closed even when the Total block reads USD (the non-USD path is
|
||
# entirely unexercised at every level, not just the total).
|
||
if item.get("currency") != candidate.get("currency"):
|
||
return False, "non_usd"
|
||
if m.group("cur") != candidate.get("currency"):
|
||
return False, "non_usd"
|
||
summary_matches = [
|
||
m
|
||
for i in range(more_detail)
|
||
if (m := _SUMMARY_ITEM_RE.fullmatch(_clean(lines[i]) or ""))
|
||
]
|
||
price = items[0].get("price") if items else None
|
||
if price is not None:
|
||
if not isinstance(price, Decimal) or len(summary_matches) != 1:
|
||
return False, "amount_mismatch"
|
||
serialized = format(price, ",.2f")
|
||
if serialized != summary_matches[0].group("price"):
|
||
return False, "amount_mismatch"
|
||
if not _money_border_ok(summary_matches[0].string, serialized):
|
||
return False, "amount_mismatch"
|
||
total = candidate.get("total_amount")
|
||
if not isinstance(total, Decimal):
|
||
return False, "amount_mismatch"
|
||
|
||
# --- V3 dual-Total proof -> amount_mismatch (V2 needs its token, so it
|
||
# derives here; sum proof follows immediately) ---
|
||
total_tokens = []
|
||
for t_idx in total_idxs:
|
||
a_idx = _next_visible(lines, t_idx)
|
||
if a_idx is None:
|
||
return False, "amount_mismatch"
|
||
token = _clean(lines[a_idx]) or ""
|
||
if not _MONEY_RE.fullmatch(token):
|
||
return False, "amount_mismatch"
|
||
c_idx = _next_visible(lines, a_idx)
|
||
cur_token = (_clean(lines[c_idx]) or "") if c_idx is not None else ""
|
||
if not _CURRENCY_RE.fullmatch(cur_token):
|
||
return False, "amount_mismatch"
|
||
total_tokens.append((token, cur_token))
|
||
if total_tokens[0] != total_tokens[1]:
|
||
return False, "amount_mismatch"
|
||
serialized = format(total, ",.2f")
|
||
if serialized != total_tokens[1][0]:
|
||
return False, "amount_mismatch"
|
||
if candidate.get("currency") != total_tokens[1][1]:
|
||
return False, "amount_mismatch"
|
||
|
||
# --- V2 sum proof -> amount_mismatch (exact Decimal equality; deliberately
|
||
# NO qty*price==amount rule -- corpus shows partial quantities) ---
|
||
if sum(item["amount"] for item in items) != total:
|
||
return False, "amount_mismatch"
|
||
|
||
# --- V5 supplier proof -> anchor_violation ---
|
||
if not supplier_name or _SUPPLIER_MARKER not in supplier_name:
|
||
return False, "anchor_violation"
|
||
det_idx = _next_visible(lines, supplier_idxs[1], shipping_idxs[1])
|
||
if det_idx is None or _clean(lines[det_idx]) != supplier_name:
|
||
return False, "anchor_violation"
|
||
sum_idx = _next_visible(lines, supplier_idxs[0], more_detail)
|
||
if sum_idx is None or _clean(lines[sum_idx]) != supplier_name:
|
||
return False, "anchor_violation"
|
||
if ship_to.get("name") == supplier_name:
|
||
return False, "anchor_violation"
|
||
|
||
# --- V8 bullet discipline -> bullet_label_unrecognized (V5's per-item
|
||
# Supplier-segment proof rides the same walk) ---
|
||
bullet_lines = [
|
||
lines[j] for j in range(lines_anchor + 1, lines_end) if _BULLET in lines[j]
|
||
]
|
||
if len(bullet_lines) != len(items):
|
||
return False, "bullet_label_unrecognized"
|
||
for raw in bullet_lines:
|
||
seen = {}
|
||
for segment in _split_bullet_segments(raw):
|
||
label, value = _match_bullet_label(segment)
|
||
if label is None or label in seen:
|
||
return False, "bullet_label_unrecognized"
|
||
seen[label] = value
|
||
if set(seen) - {
|
||
"Supplier",
|
||
"Need By",
|
||
"Category",
|
||
"Account",
|
||
"Period",
|
||
"Part Number",
|
||
}:
|
||
return False, "bullet_label_unrecognized"
|
||
if not {"Supplier", "Need By", "Category", "Account", "Period"} <= set(seen):
|
||
return False, "bullet_label_unrecognized"
|
||
if _clean(seen["Supplier"]) != supplier_name:
|
||
return False, "anchor_violation"
|
||
|
||
# --- V6 ship_to required -> missing_required_field ---
|
||
for field in ("name", "street", "city", "state", "zip", "location_code"):
|
||
if not ship_to.get(field):
|
||
return False, "missing_required_field"
|
||
if not re.fullmatch(r"\d+", ship_to["location_code"]):
|
||
return False, "missing_required_field"
|
||
lc_values = [
|
||
m.group("lc")
|
||
for j in range(shipping_idxs[1] + 1, lines_anchor)
|
||
if (m := _LOCATION_CODE_RE.fullmatch(_clean(lines[j]) or ""))
|
||
]
|
||
if ship_to["location_code"] not in lc_values:
|
||
return False, "missing_required_field"
|
||
attn_values = [
|
||
_clean(m.group("attn"))
|
||
for j in range(shipping_idxs[1] + 1, lines_anchor)
|
||
if (m := _ATTN_RE.fullmatch(_clean(lines[j]) or ""))
|
||
]
|
||
if attn_values:
|
||
if ship_to.get("attn") not in attn_values:
|
||
return False, "missing_required_field"
|
||
elif ship_to.get("attn") is not None:
|
||
return False, "missing_required_field"
|
||
|
||
# --- V7 address shape -> address_shape_invalid (gated on the RAW
|
||
# pre-enrichment zip: validate() runs BEFORE enrich_parsed, so a short zip
|
||
# like '2149' fails closed to the LLM path where pad_zip repairs it --
|
||
# both paths then get identical pad_zip treatment downstream) ---
|
||
us_idx = _find_after(lines, _US_SENTINEL, shipping_idxs[1] + 1)
|
||
if us_idx is None or us_idx >= lines_anchor:
|
||
return False, "address_shape_invalid"
|
||
m = _CITY_STATE_ZIP_RE.fullmatch(_clean(lines[us_idx - 1]) or "")
|
||
if not m:
|
||
return False, "address_shape_invalid"
|
||
if (
|
||
m.group("city") != ship_to["city"]
|
||
or m.group("state") != ship_to["state"]
|
||
or m.group("zip") != ship_to["zip"]
|
||
):
|
||
return False, "address_shape_invalid"
|
||
if ship_to["state"] not in _USPS_STATES:
|
||
return False, "address_shape_invalid"
|
||
if not re.fullmatch(r"\d{5}(-\d{4})?", ship_to["zip"]):
|
||
return False, "address_shape_invalid"
|
||
|
||
# --- V9 required labeled fields -> missing_required_field ---
|
||
for field in (
|
||
"po_status",
|
||
"order_date",
|
||
"payment_terms",
|
||
"requisition_number",
|
||
"submitted_by",
|
||
"view_order_url",
|
||
):
|
||
if candidate.get(field) is None:
|
||
return False, "missing_required_field"
|
||
for item in items:
|
||
for field in (
|
||
"description",
|
||
"amount",
|
||
"currency",
|
||
"need_by",
|
||
"category",
|
||
"account_code",
|
||
"period",
|
||
):
|
||
if item.get(field) is None:
|
||
return False, "missing_required_field"
|
||
for label in _MORE_DETAIL_LABELS:
|
||
if len(_indices(lines, label)) != 1:
|
||
return False, "missing_required_field"
|
||
if candidate.get("coupa_category") != items[0].get("category"):
|
||
return False, "missing_required_field"
|
||
# Nullable by design: on_behalf_of, department, last_opened, acknowledged_at,
|
||
# revision_date, attn, quantity, unit, price.
|
||
|
||
# --- V10 PO identity proofs -> po_id_mismatch ---
|
||
po = candidate["po_number"]
|
||
po_id_idx = _indices(lines, "PO ID")[0]
|
||
v_idx = _next_visible(lines, po_id_idx)
|
||
if v_idx is None or _clean(lines[v_idx]) != po:
|
||
return False, "po_id_mismatch"
|
||
if not _indices(lines, f"Amazon Purchase Order #{po}"):
|
||
return False, "po_id_mismatch"
|
||
um = _ORDER_URL_ID_RE.match(candidate.get("view_order_url") or "")
|
||
if not um or um.group(1) != po.split("-", 1)[1]:
|
||
return False, "po_id_mismatch"
|
||
|
||
# --- V11 sentinel discipline -> unparseable_value ---
|
||
for leaf in _walk_leaves(candidate):
|
||
if isinstance(leaf, str) and leaf == _UNPARSEABLE:
|
||
return False, "unparseable_value"
|
||
|
||
# --- V12 hygiene -> residual_artifact ---
|
||
for leaf in _walk_leaves(candidate):
|
||
if isinstance(leaf, str) and ("\r" in leaf or "\xa0" in leaf):
|
||
return False, "residual_artifact"
|
||
|
||
# --- V13 quantity/unit/price coherence -> amount_mismatch ---
|
||
quantity = items[0].get("quantity")
|
||
unit = items[0].get("unit")
|
||
if quantity is not None or unit is not None or price is not None:
|
||
if not isinstance(quantity, Decimal) or quantity <= 0:
|
||
return False, "amount_mismatch"
|
||
if not unit or not re.fullmatch(r"[A-Z]+", unit):
|
||
return False, "amount_mismatch"
|
||
if not isinstance(price, Decimal):
|
||
return False, "amount_mismatch"
|
||
if len(summary_matches) != 1:
|
||
return False, "amount_mismatch"
|
||
if _to_decimal(summary_matches[0].group("qty")) != quantity:
|
||
return False, "amount_mismatch"
|
||
if summary_matches[0].group("unit") != unit:
|
||
return False, "amount_mismatch"
|
||
# The Lines-block '<qty> EA' evidence line must numeric-equal the
|
||
# summary quantity when present.
|
||
ea_matches = [
|
||
m
|
||
for j in range(lines_anchor + 1, lines_end)
|
||
if (m := _LINE_QTY_RE.fullmatch(_clean(lines[j]) or ""))
|
||
]
|
||
if ea_matches:
|
||
if len(ea_matches) != 1:
|
||
return False, "amount_mismatch"
|
||
if _to_decimal(ea_matches[0].group("qty")) != quantity:
|
||
return False, "amount_mismatch"
|
||
|
||
return True, "ok"
|
||
|
||
|
||
def try_deterministic_parse(email_data):
|
||
"""Entry point. Returns (parsed|None, parse_method, template_id, reason).
|
||
|
||
On a proven-conformant parse returns (dict, 'template', template_id, 'ok').
|
||
On any miss/invalid/exception returns (None, 'ai_fallback', template_id,
|
||
reason) -- a failure is NEVER a parsed result."""
|
||
template_id = "unknown"
|
||
try:
|
||
template_id, reason = classify_template(email_data)
|
||
if template_id == "unknown":
|
||
return None, "ai_fallback", template_id, reason
|
||
|
||
if template_id == "coupa_new_po":
|
||
candidate = extract_new_po(email_data)
|
||
else:
|
||
candidate = extract_cancellation(email_data)
|
||
|
||
ok, reason = validate(candidate, template_id, email_data)
|
||
if not ok:
|
||
return None, "ai_fallback", template_id, reason
|
||
return candidate, "template", template_id, "ok"
|
||
except Exception: # noqa: BLE001 -- fail closed on ANY extractor error
|
||
return None, "ai_fallback", template_id, "extractor_raised"
|