* Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * feature: Add PO template parser scaffold and design doc Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage: - coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands. - coupa_cancellation (2.9%): implemented. Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work. Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com> * Implement PO new_po extraction and value-level gate Replace the extract_new_po scaffold stub with the full section-windowed extractor (duplicate-label anchoring, sentinel ship-to, label-keyed U+2022 bullet split, Decimal money from three anchored contexts only) and add value-level gate rules V1-V13. Both new_po_not_implemented scaffold guards are removed; rules 6-8 (unrecognized_status, multiline_unsupported, non_usd) go live. The gate re-derives every byte proof from the email body so an extractor bug cannot vouch for itself: amount re-serialization with a digit/comma border check (the thousands-separator truncation kill switch), sum(lines)==total against both Total blocks, anchor/supplier identity proofs, USPS address shape on the raw pre-enrichment zip, bullet label discipline, and sentinel/artifact hygiene. Any failure falls closed to the LLM; a validation failure is never a parsed result. Refs: #99 * Wire template-first parse into PO handler with EMF metric Run try_deterministic_parse ahead of the Bedrock extractor and fall back only on a miss/invalid (fail-closed) result. The shared enrich_parsed post-stage and the save_cancellation/save_revision/ save_new_po routing are untouched, so both paths write identical DynamoDB shapes and the po-ingest-site-extractor stream contract is preserved. Each record emits one ParseMethod EMF line (Seahaven/PoIngest/ ParseOutcome, dimension sets [ParseMethod] and [ParseMethod,TemplateId], ReasonCode/po_number ride-alongs) mirroring the WO idiom. The metric fires before the Bedrock call so a Bedrock-side error still records the ai_fallback outcome. Refs: #99 * Add PO fallback-rate alarm retuned for ~57 emails/day The WO alarm's 15-min period and >=10-sample floor assume ~760/day and would be structurally dead at PO volume (a 15-min period holds ~0.6 emails, so the floor is never met). Retune: 6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...) volume floor so a single email can never breach a datapoint (1/8 = 12.5% < 20%), threshold >20% against a ~1% expected baseline, eval 4 / datapoints 2 (24h span) so noise self-clears while total template drift pages within ~12h. No element-wise MAX in the math expression (post-#102 rule); ALARM-only SnsAction to site-alerts, NOT_BREACHING. Gated with 'npx cdk synth po-ingest'. Also add template_parser.py to the bundling cp list -- without it every deployed invocation would ImportError (unit tests cannot catch an asset-bundling omission). Refs: #99, #102 * Add offline PO parser suite with scrubbed fixture corpus 132 tests: golden-file comparison for all 25 positive fixtures (17 single-line new-PO + 8 cancellations, Decimal-exact via parse_float=Decimal), every fail-closed gate reason code covered (body-level triggers via 17 synthetic adversarial .eml mutations, candidate-level via direct validate() unit tests), real multi-line and comment/non-Coupa fallback fixtures, dual line-ending parse identity, two-path enrich/save parity (site-extractor stream guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass + scrub-marker leak sweep), and Bedrock dispatch/EMF assertions. The suite loads handler/template_parser via importlib under unique module names and binds the handler's bare sibling imports around exec (tests/conftest.py load_handler gets the same treatment) -- the WO suite caches bare 'handler'/'template_parser' names in sys.modules, and bare imports here would silently bind to the wrong pipeline. moto is imported before the handler so its botocore stubber hook precedes boto3 session creation (the PO conftest chain now loads at pytest session start). Fixtures are scrubbed real S3 samples: transport/auth header values replaced with same-shape placeholders (structure kept so ses_auth still passes), per-file digit ciphers, amounts remapped with sum==total re-established. The .gitignore exception is scoped to the PO fixtures path only. Refs: #99 * Document PO template-first parser and retuned alarm README: PO flow is now template-first with Bedrock fallback; parser/gate section mirroring the WO writeup; Seahaven/PoIngest ParseOutcome namespace and the fallback-rate alarm numbers with their volume justification (deliberately not WO's settings); test-suite and repo-layout updates. Design doc: mark PR #1 complete in progress/checklist sections; document the six value-level gate reason codes and the scaffold guard removal; correct the stale data-access note (default CLI session is 328440206208) and note the ~90-day S3 lifecycle aging of the corpus; record the 2.3 layout addendum (leading Supplier bullet segment, EA evidence lines, summary unit-price tokens, decode-path line endings), the fixture-build pins (address join convention, quantity/unit/price source), the V10 sweep outcome, and resolutions for open questions Q3/Q6. Cross-family review and the Confluence architecture-map update are flagged outstanding for merge. Refs: #99 * Record cross-family review outcome for handler wiring GPT-4.1 cross_review.py run against the real handler diff returned no BLOCK and no security findings; both FIX items verified as no-change-needed (fallback logging already correct; non-dict AI output is the pre-existing issue #101 pattern this PR deliberately does not touch). Refs: #99 * Pin line-item currency to USD in the PO gate The non_usd rule only checked the Total-block top-level currency, so a new_po whose line item read 'for 55,206.00 CAD' under a USD Total block still template-parsed as ok -- a fail-open hole in the fail-closed gate. Every line item's captured currency and its re-derived body token must now byte-equal the proven-USD top-level currency; covered by a line-level CAD adversarial fixture (the existing adv-non-usd only exercised the Total-block variant) and a candidate-mutation unit test. * Scrub residual transport tokens from PO fixtures The first-pass harvest scrub sanitized only the primary SES/DKIM header blocks, leaving the real SES Feedback-ID sender-identity hash in 49 committed fixtures and, on the two non-Coupa fixtures, an embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the token classes the PR #99 fixture lesson requires placeholdered. Replace each with a same-shape ScrubbedFixture value (byte-safe, CRLF and folding preserved) so header structure and ses_auth behavior are unchanged. * Converge quantity/price to Decimal on both paths EXTRACTION_PROMPT declares quantity and price as JSON strings, so a prompt-obedient Bedrock response stores DynamoDB Strings where the template parser stores Numbers -- divergent attribute types for the same email on the purchase-orders stream. Coerce numeric strings to Decimal in the shared enrich_parsed post-stage (thousands-separator safe; non-numeric strings kept verbatim) so both paths converge; prompt rewording itself remains PR #2 scope. The two-path parity test was circular -- it replayed the parser- derived golden as 'the LLM output', so it could never see the type divergence. It now feeds a prompt-shaped payload (string quantity/ price, LLM-filled site_code) through enrich_parsed and save_new_po, and the fixture-hygiene test now asserts the scrubbed transport-token header classes so fixture regressions are caught. * Coerce bare-int quantity/price to Decimal in enrich_parsed GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged residual type drift: parse_float=Decimal rules out floats on the LLM path, but a bare JSON int survived as Python int. Coerce it so both parse paths emit one canonical Decimal type. * Scrub fixture-body PII and harden cancellation gate (sec review) /sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill verifier) confirmed two diff-introduced findings; both fixed here. F3 (medium, real PII in new fixtures): the harvest scrub replaced header tokens but left real third-party PII in message BODIES -- an Amazon contact's name/phone/personal email in non-coupa-02.eml and an internal t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn names recurring across the new_po corpus. Replaced every personal name, phone, personal email, and internal URL with synthetic placeholders (QP-soft-wrap aware) across both .eml bodies and expected goldens. Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs, the leaked tokens), closing the header-only gap that let this through. F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and matched with .search(), unlike the anchored new_po pattern -- a subject merely ending with the cancellation phrase could be routed to the sticky- Cancelled write. Fully anchored it and switched to .match, and added a body-corroboration gate (the real Coupa body independently restates 'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted subject whose body does not corroborate now fails closed to the LLM (new reason code cancellation_body_unconfirmed). Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po dispatch and undelimited extraction prompt (issue #101 family) are byte-identical to main and unchanged here. 401 tests pass; ruff/format clean; cdk synth po-ingest clean. --------- Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
25 KiB
PO Template-First Parser — Design, Investigation & Progress
Status: PR #1 implemented (extraction + gate + wiring + alarm + tests); PR #2 (derived-classifier factoring) not started · Branch:
feat/po-template-parser(stacked on PR #99) · Last updated: 2026-07-16Living document for making the Coupa purchase-order email parser template-first with LLM fallback, mirroring the work-order (WO) processor's PR #99. Captures the investigation, the data we gathered, the decisions made, the current scaffold, and everything still to do.
1. Context & Goal
The PO email processor (lambdas/po/email_processor/handler.py) sends every inbound Coupa/Amazon purchase-order email to an LLM (Claude Haiku 4.5 on Bedrock) for structured extraction. WO PR #99 established the pattern we want to replicate here:
Run a deterministic template parser first. It returns a result only when the email provably conforms to a known template and passes a fail-closed validation gate; otherwise the handler falls back to the Bedrock LLM. This removes the LLM from the hot path for routine traffic while keeping full AI coverage for anything unexpected.
Goal: apply the same template-first + fail-closed-fallback approach to the PO processor, without regressing data quality or the downstream po-ingest-site-extractor pipeline (which reads site_code/ship_to off the purchase-orders DynamoDB stream).
Reference implementation: lambdas/wo/email_processor/template_parser.py (PR #99).
Why this is harder than WO
- The PO contract is ~40 fields with nested objects (
supplier{},ship_to{},line_items[]) vs WO's flat 16. - Much of the PO extraction prompt is derived classification (
trade,site_code,fiscal_year) already expressed as deterministic English rules. - Money fields require
Decimal(DynamoDB rejects floats). email_typeis not in the subject (unlike WO's two subject templates).
2. Investigation
Two phases: a read-only multi-agent (ultracode) feasibility study over a 120-email sample, then a full-bucket triage over all 3,448 inbound emails to get real distribution numbers.
2.1 Data access
- PO email bucket:
s3://po-ingest-emails-328440206208/inbound/(AWS account 328440206208, us-east-1). Reached via AWS profileStale (corrected 2026-07-16 during the fixture harvest): the default CLI session is now authenticated to 328440206208 directly — noamoussa-mgmt(SSO); default CLI creds are the personal account681986854588--profileflag needed.- ⚠️ Corpus is aging out:
inbound/objects carry an S3 lifecycle expiration (~90-day rolling window; oldest object 2026-04-17 at harvest time, 3,422 objects vs 3,448 at triage). Any further harvesting should not be deferred long.
2.2 Full-bucket triage results (all 3,448 emails, 2026-07-16)
Method: parallel ranged-GET (bytes=0-12000) of every object, then local classification (quoted-printable bodies, decodable offline).
| Email kind | Count | % | Notes |
|---|---|---|---|
new_po — ***Copy for Reference*** New Purchase Order <PO#> has been issued |
3,294 | 95.5% | Single uniform Coupa layout |
cancellation — <Site> Purchase Order #<PO#> has been cancelled |
100 | 2.9% | Distinct subject + minimal body |
comment — New Comment on Purchase Order for Amazon |
19 | 0.55% | Type NOT in handler enum (see §7) |
| non-Coupa senders (human replies, WO mail, an AWS SES setup notice) | 35 | 1.0% | Rejected at ses_auth before parse |
| revision | 0 | 0% | No distinct revision emails exist |
Within new_po:
- Status values:
Issued - Created(2,223) andIssued - Scheduled for email(1,071) — only these two. - Multi-line-item: 6 / 3,294 (0.18%), all ≤3 items.
- Non-USD: 0 (after excluding unit-token false positives like
EACH/HOUR). - So the single-line / USD scope covers ~99.8% of new_po traffic.
2.3 new_po layout (the one template that matters)
multipart/alternative; the text/plain part is a stable, label-delimited flattening:
- Top:
Amazon Purchase Order #<PO#>, thenSubmitted By,On Behalf Of,Supplier,Total,Items. More Detailblock:PO ID,Department,Status,Last Opened,Order Date,Acknowledged At,Revision Date,Payment Term,Req #,Shipping.Supplierdetail block (address).Shippingdetail block (ship-to address,Location Code:,Attn:).Linessection: one block per line item; per-line metadata is U+2022 (•) delimited:Need By·Category·Account·Period[· optionalPart Number].
Structural hazards found in real data:
Supplier,Shipping,Totallabels each appear twice (summary placeholder + detail block); the firstShippingvalue is literallyNone.- Optional
Part Numberbullet segment appears in a minority of emails, inserted betweenCategoryandAccount→ shifts any positional splitter. - Every captured value carries a trailing
\r; a lone\xa0(nbsp) line sits beforeTotal. - Real amounts include thousands separators (
18,624.05,32,405.00).
Addendum (2026-07-16 fixture harvest):
- Every per-line bullet run BEGINS with a
Supplier <name>segment beforeNeed By(the gate proves it byte-equalssupplier.nameon every item). - Multi-item Lines blocks carry a
<qty> EAevidence line before each description; the single-item Items summary carries an optional<qty> <UNIT> x <price>line (absent on some emails, e.g. new-po-14) which is the source forquantity/unit/price(the Lines-blockEAline is a gate cross-check only). - The Items summary contains unit-price money tokens distinct from line amounts (e.g.
50.0 EACH x 66.00with line amount8,356.00) — money must be anchored on thefor <amt> <CCY>/ Total-block / summary-xcontexts, never free money-shaped scanning; and lineamount ≠ qty×priceon partial quantities, so onlysum(lines) == totalis load-bearing. - Body line endings are decode-path dependent: Coupa QP encodes
=0D, sopolicy.defaultget_content()yields\r\neven after transport normalization, while other decode paths yield bare LF — the parser matches only_clean-ed lines and is tested against both representations.
2.4 cancellation layout
Subject: <SiteName> Purchase Order #<PO#> has been cancelled. Body is a short dashes-delimited notice — no Status label, no line items. The only field the handler's save_cancellation() needs is po_number.
3. Key Decisions
- Build two templates, not one and not three:
coupa_new_po(95.5%) — the main workload.coupa_cancellation(2.9%) — trivial and low-risk (justpo_number); we have 100 real examples.- Together ≈ 98.4% of all mail deterministically handled/routed.
- No revision template — 0 distinct revision emails in 3,448. The handler's
revisiontype essentially never arrives as its own email. email_typenever defaults to new_po. It is emitted only when the exact new_po subject regex matches andpo_status∈ the two confirmed-safe strings. Anything else → LLM. (Misrouting a cancellation to new_po would silently defeat the sticky-Cancelledguard.)- Fail closed, always. Any miss / invalid / exception returns
Noneand falls back to the LLM. A failure is never a parsed result. (Copies WO'stry/exceptposture.) Decimalfor all money, with thousands-separator stripping, to match the LLM path'sparse_float=Decimal.- Derived fields (
site_code,trade,fiscal_year) are NOT computed in the parser. They are leftNone; a shared post-stage (enrich_parsed(), thepad_zipprecedent) fills them identically on both the template and LLM paths, so the gate judges extraction fidelity only and both paths write byte-identical shapes downstream.coupa_categoryis a verbatim label capture, not a classifier, and IS extracted. - Structural nested contract with recursive key-set validation (unlike WO's flat tuple).
- Two-PR split: (1) extraction template + gate + handler wiring; (2) derived-classifier factoring, shadow-logged before it becomes authoritative.
4. Contract
Top-level CONTRACT_KEYS (23) plus three nested sub-tuples. Taken from the EXTRACTION_PROMPT in handler.py.
CONTRACT_KEYS = (
email_type, po_number, po_status, source_system, submitted_by, on_behalf_of,
order_date, revision_date, last_opened, acknowledged_at, payment_terms,
requisition_number, department, view_order_url, supplier, site_code, ship_to,
total_amount, currency, fiscal_year, trade, coupa_category, line_items,
)
SUPPLIER_KEYS = (name,)
SHIP_TO_KEYS = (name, address, street, city, state, zip, location_code, attn)
LINE_ITEM_KEYS = (description, amount, currency, need_by, category,
account_code, period, quantity, unit, price)
| Class | Keys |
|---|---|
| Verbatim / labeled extract | po_number, po_status, submitted_by, on_behalf_of, order_date, revision_date, last_opened, acknowledged_at, payment_terms, requisition_number, department, view_order_url, supplier.name, ship_to.location_code, ship_to.attn, coupa_category, all line_items[*] |
| Extract with care | total_amount, currency, ship_to.name/street/city/state/zip (Decimal; sentinel-anchored address) |
| Constant | source_system = "coupa" |
| Derived — post-stage, NOT parser, NOT gate-validated | site_code, trade, fiscal_year |
Handler enrichment metadata (raw_s3_key, processed_at, data_source, email_subject, ship_to_raw, top-level state) is added by enrich_parsed() and is not part of the parser contract.
A parity test (mirroring WO's test_contract_keys_match_extraction_prompt) should assert the flattened contract key set is a subset of the keys quoted in EXTRACTION_PROMPT.
5. Parser Design
File: lambdas/po/email_processor/template_parser.py — pure module (no boto3/network), mirroring WO idioms.
classify_template(email_data) -> (template_id, reason)— exact subject regex →coupa_new_po|coupa_cancellation|unknown.extract_new_po/extract_cancellation— build the nested candidate; leave derived fieldsNone._empty_candidate()/_normalize()— build & normalize the nested skeleton (supplier={},ship_to={},line_items=[one item])._clean()— strip trailing\r+\xa0, map"None"/empty →None._to_decimal()— thousands-separator-safeDecimal,_UNPARSEABLEsentinel on failure.validate(candidate, template_id, email_data) -> (bool, reason)— fail closed.try_deterministic_parse(email_data) -> (parsed|None, method, template_id, reason)— entry point;try/exceptfails closed on any error.
Handler integration (✅ wired in PR #1): after authenticate_inbound_email() + parse_raw_email(), try_deterministic_parse runs first; if None, extract_with_claude; the shared enrich_parsed() stage runs on parsed regardless of path, then the existing save_cancellation / save_revision / save_new_po routing (untouched). The ParseMethod EMF metric is emitted immediately after the deterministic attempt — before any Bedrock call — so a Bedrock-side error still records the ai_fallback outcome.
6. Fail-Closed Gate Rules
Reason codes emitted by validate() / try_deterministic_parse() (closed set):
subject_no_match · key_set_mismatch · derived_field_set · email_type_mismatch · missing_required_field · po_id_mismatch · unrecognized_status · multiline_unsupported · non_usd · amount_mismatch · anchor_violation · address_shape_invalid · bullet_label_unrecognized · unparseable_value · residual_artifact · extractor_raised · (ok / template on success).
(new_po_not_implemented disappeared with the scaffold guard — both copies removed; a test asserts the string no longer exists in the module.)
Implemented (both templates):
- Known template only (
subject_no_matchotherwise). - Exact nested key-set at every level (
key_set_mismatch). - Derived fields must be unset by the parser (
derived_field_set). email_typein enum and == template's expected type (email_type_mismatch).po_numbervalid shape^[A-Z0-9]{1,6}-\d+$and byte-equals the subject id (po_id_mismatch/missing_required_field).
Implemented (new_po): safe-status enum (unrecognized_status), single-line only (multiline_unsupported), USD only (non_usd) — enforced at both levels: the Total-block top-level currency must be exactly USD (rule 8) and every line item's captured currency plus its re-derived for <amt> <CCY> body token must byte-equal it (a single non-USD line item fails closed even when the Total block reads USD).
Implemented (new_po value-level rules V1–V13 — ✅ done, one reason code each): the gate re-derives every byte proof from email_data["body"] (never trusting extractor-carried state), with the six value-level reason codes:
amount_mismatch— money fidelity (each Decimal re-serializes byte-identically to its re-derived source token, non-digit/comma border — the18,624.05→624.05kill switch),sum(lines) == total(exact Decimal; deliberately noqty×price == amountrule — corpus shows partial quantities), dual-Totalbyte-identity, and quantity/unit/price coherence vs the summary andEAevidence lines.anchor_violation— exactly 2×Supplier/Shipping/Total, firstShippingis the literalNoneplaceholder, section ordering,supplier.namecontainsSEA HAVENand byte-equals the detail-block name + every item'sSupplierbullet segment,ship_to.name != supplier.name.address_shape_invalid— city line fullmatches^City, ST ZIP$immediately beforeUnited States, state in the frozen USPS set, zip^\d{5}(-\d{4})?$on the RAW pre-enrichment value (a short zip fails closed to the LLM path where the sharedpad_ziprepairs it — the gate never predicts enrichment).bullet_label_unrecognized— every U+2022 segment carries a recognized leading label from the closed set; core labels exactly once,Part Numberat most once.unparseable_value— recursive walk: no_UNPARSEABLEsentinel leaf.residual_artifact— recursive walk: no\r/\xa0in any string leaf. Plusmissing_required_fieldfor the required labeled fields (per-item metadata,Location Code:sourcing, no inventedAttn:) andpo_id_mismatchfor the bodyPO ID/ heading /orders/<id>-URL identity proofs.
Precondition (handled a layer earlier): authenticate_inbound_email() (SES dkim=pass for amazon.coupahost.com) must pass before the parser runs — the 35 non-Coupa emails are rejected there.
7. Scope — In / Out
IN (template-parsed):
- Single-line-item, USD,
new_powith the exact subject and a safeStatus. cancellation(subject match →po_number).
OUT (always LLM fallback / rejected):
- Comments (19) —
New Comment on Purchase Orderis anemail_typethe handler enum (new_po/revision/cancellation) does not model. The LLM currently shoehorns these. ⚠️ Pre-existing data-quality gap, independent of this work — decide separately whether comments should create/update PO records at all. - Revisions — no template (0 emails).
- Multi-line-item new_po (0.18%) —
Lines-array structure unobserved;multiline_unsupported. - Non-USD new_po (0 observed) — path unexercised;
non_usd. - Non-Coupa senders (1.0%) — rejected at ses_auth.
8. Progress
- Read-only ultracode feasibility investigation (8 agents, GO-with-conditions).
- Full-bucket triage over all 3,448 emails → distribution + scope numbers (§2.2).
- Confirmed cancellation & comment body layouts against real emails.
- Scaffold
template_parser.pyonfeat/po-template-parser:- Nested contract + recursive
_empty_candidate/_normalize. classify_templatefor both templates (real subject regexes).coupa_cancellationextract + gate — implemented & smoke-tested (parses a real cancellation →po_number,email_type=cancellation).coupa_new_poclassify + structural/status/single-line/USD gate.Scaffold guard— removed (both copies) with the PR #1 implementation below.ruff check+ruff format --checkpass.
- Nested contract + recursive
- PR #1 implementation (2026-07-16):
extract_new_po— section-windowed, anchor-disciplined, label-keyed bullet split,Decimalmoney, representation-agnostic line handling. All 17 harvested single-line new-POs parse totemplate/ok; the 3 real multi-line ones extract faithfully (sum==total) then gate-reject withmultiline_unsupported.- Value-level gate rules V1–V13 (§6) — every byte proof re-derived from the body.
- Handler wiring:
try_deterministic_parsefirst, Bedrock fallback onNone, sharedenrich_parsed()on both paths,save_*routing untouched. ParseMethodEMF metric set (Seahaven/PoIngest/ParseOutcome, dims[["ParseMethod"],["ParseMethod","TemplateId"]],ReasonCode/po_numberride-alongs), emitted before the Bedrock call.- Fallback-rate alarm in
cdk/po_stack.py, retuned for ~57/day (6h periods · ≥8 volume floor · >20% · 2-of-4) —npx cdk synth po-ingestgreen. cdk/po_stack.pybundlingcpnow includestemplate_parser.py(was a deploy-time ImportError waiting to happen).- Offline test suite
lambdas/po/email_processor/tests/(136 tests): goldens for all 25 positive fixtures, every §6 reason code covered, dual line-ending identity, two-path enrich/save parity, fixture hygiene (ses_auth + leak-sweep markers), Bedrock dispatch/EMF. Full-root pytest: 358 green. - Committed sanitized
.emlfixture corpus (header + body scrub, per-file digit cipher, narrow.gitignoreexception scoped tolambdas/po/email_processor/tests/fixtures/). - README updated (PO flow, parser section, alarm numbers + justification, tests, repo layout).
9. TODO / Path Forward
PR #1 — extraction template + wiring
- Implement
extract_new_po: labeled single-occurrence fields; duplicate-label anchoring (Supplier/Shipping/Total);ship_toby sentinel anchors;line_itemssplit on•by leading label;Decimalamounts;coupa_categoryverbatim. - Implement the new_po value-level gate rules (§6 V1–V13) and remove the scaffold guard (both copies).
- Handler integration:
try_deterministic_parsefirst, Bedrock fallback onNone; route through existingsave_*branches. - Emit
ParseMethod-only EMF set (namespaceSeahaven/PoIngest, dims[["ParseMethod"], ["ParseMethod","TemplateId"]]). - Fallback-rate CloudWatch alarm in
cdk/po_stack.py— retuned for ~57/day: 6h periods,IF((fb+tmpl)>=8, …)volume floor, >20% threshold, eval 4 / datapoints 2 (see the po_stack comment + README for the arithmetic;npx cdk synth po-ingestgated). - Offline tests + fixtures: goldens for every positive fixture; one adversarial per gate rule (thousands-sep, dup-label swap, Part-Number bullet shift pair, multiline, non-USD, bad status, PO-id/URL mismatch, unlabeled/duplicate bullets, prompt injection); comment + non-Coupa assertions; contract⊆prompt parity test.
- README updated. Confluence "AWS Architecture Map" update — ⚠️ outstanding (must land with the merge; PO subgraph gains the template-parser stage + fallback alarm).
cross_review.py(GPT-4.1) pass — run 2026-07-16 against the real handler diff: no BLOCK, no security findings, "safe to merge with minor fixes." Both FIX items verified as no-change-needed: the fallback log line already reportsai_fallbackcorrectly (a Claude-side exception propagates before it, unchanged posture), and the non-dict-from-Claude hazard is the pre-existing issue #101 pattern this PR deliberately leaves untouched.ruff check+ruff format+pytestgreen locally (354 tests; CI must confirm).
PR #2 — derived-classifier factoring (follow-up)
- Move
trade/site_code/fiscal_yearinto sharedenrich_parsed(). - Run in shadow mode: compute Python values, log Python-vs-LLM disagreement via EMF, keep LLM authoritative during a bake period.
- Harden
site_code(7-shape + skip-list, incl. multi-hop ATTN likeCBRE - RME - DLI6) andtradeagainst the full sample with per-shape fixtures. - Only after acceptable agreement: make Python authoritative and drop the fields from
EXTRACTION_PROMPT.
10. Open Questions
- Goal — cost, latency, or determinism? Determinism (same input → same output) is the strongest justification for the derived-classifier port regardless of per-email cost.
- Classifier factoring now or deferred? Recommendation: extraction template first (PR #1), factoring as shadow-logged PR #2.
- Strict single-line/USD-only v1 acceptable? ✅ Resolved (PR #1): yes. Covers ~99.8% of new_po; multi-line and non-USD fail closed with honest reason codes.
- Confirm the real cancellation
Statusspelling —handler.pyhardcodesCANCELLED_STATUS='Cancelled'; the cancellation body doesn't restate it. Verify against a known-cancelled PO in the table. (Doesn't affect parser output —save_cancellationhardcodes the string.) - Comments — should
New Comment on Purchase Orderemails create/update PO records at all, or be dropped? (Out of scope for the parser, but a real data-quality decision.) - Repo fixtures — ✅ Resolved (PR #1): committed sanitized corpus under
lambdas/po/email_processor/tests/fixtures/(35 scrubbed.eml: 17 single-line new-PO, 8 cancellations, 4 comments, 3 non-Coupa, 3 multi-line; plus 18 synthetic adversarial mutations). Transport/auth headers scrubbed to same-shape placeholders (structure preserved,ses_authstill passes), per-file digit cipher on identifiers, amounts remapped withsum==totalre-established, and a narrow.gitignoreexception scoped to the PO fixtures path (never blanket*.eml). Source S3 keys deliberately unrecorded (they ARE the SES receipt tokens the scrub replaced).
Fixture-build pins recorded during PR #1 (locked by goldens):
ship_to.street/ship_to.addressjoin convention: multi-line segments joined with'\n';address= the lines fromship_to.namethroughUnited Statesinclusive (this is whatenrich_parsedpromotes toship_to_rawon the site-extractor stream).quantity/unit/pricesource: the Items-summary<qty> <UNIT> x <price>line (e.g.1.0 EACH x 55,206.00); the Lines-block<qty> EAline is gate evidence only (V13 numeric cross-check). Emails without a summary line (e.g. new-po-14) leave all threenull— nullable by contract.- V10 URL-id sweep outcome: all 20 harvested new_po emails satisfy
orders/<id> == po_numberdigit suffix — the strict V10 check stays live (no demotion needed); a corpus-sweep test locks it. - Canonical
quantity/pricetype = Decimal (DynamoDB Number), converged in the sharedenrich_parsed:EXTRACTION_PROMPTdeclares both fields as JSON strings, so a prompt-obedient Bedrock response arrives asstrwhere the template path emitsDecimal. The shared post-stage coerces numeric strings (thousands-separator-safe) toDecimalso both paths write the same attribute type to the purchase-orders table stream; non-numeric strings are left verbatim. Prompt rewording itself stays PR #2 scope. The two-path parity test feeds a prompt-shaped payload (string quantity/price, LLM-filledsite_code) — never the parser-derived golden verbatim — so it cannot be circular. - Second-pass header scrub (post-review): the first-pass harvest scrub left the real SES
Feedback-IDsender-identity hash on every Coupa fixture and, on the two non-Coupa fixtures, an embedded second SES block'sX-Ses-Receipt, the Exchange cross-tenant UPN ciphertext, Gmail ARCfh=/X-Gm-*tokens, and related opaque routing blobs. All replaced with same-shapeScrubbedFixtureplaceholders (byte-safe, CRLF preserved); the fixture-hygiene test now asserts these token classes are scrubbed in every header block so regressions are caught.
11. Risk Register (high-severity)
| Risk | Mitigation |
|---|---|
Cancellation misrouted to new_po → sticky-Cancelled guard defeated, PO stays active |
email_type closed-set gate; never default-to-new_po; only emit on exact subject + safe status |
Thousands-separator truncation (18,624.05→624.05, ~30× too small, passes naive checks) |
Amount byte-equality: re-serialize Decimal to source grouped format; reject if token bordered by digit/comma; sum(lines)==total |
| Multi-line structure unobserved | Hard-fail len(line_items)!=1 → LLM; profile real multi-line before supporting |
Derived-field silent divergence (wrong trade/site_code still well-formed; site_code rules already incomplete) |
Factor into shared post-stage; keep OFF the gate; shadow-log before authoritative; per-shape fixtures |
Duplicate-label first-match (Supplier/Shipping/Total twice; first Shipping=None) |
Occurrence/sentinel anchoring; gate rejects Shipping=='None', requires Location Code:, supplier.name has SEA HAVEN, ship_to.name != supplier.name |
Appendix — References
- WO reference parser:
lambdas/wo/email_processor/template_parser.py(PR #99). - PO handler +
EXTRACTION_PROMPT:lambdas/po/email_processor/handler.py. - PO stack (alarms, IAM, Bedrock):
cdk/po_stack.py. - Bedrock model: inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0. - PO email bucket:
s3://po-ingest-emails-328440206208/inbound/(account 328440206208, profileamoussa-mgmt).