mirror of
https://github.com/Sea-Haven-Industries/procurement-ingest.git
synced 2026-09-30 09:33:15 +00:00
287 lines
30 KiB
Markdown
287 lines
30 KiB
Markdown
# PO Template-First Parser — Design, Investigation & Progress
|
||
|
||
> **Status:** PR #1 implemented (extraction + gate + wiring + alarm + tests); PR #2 (derived-classifier factoring — shadow mode) implemented · **Branch:** `feat/po-template-parser` (stacked on PR #99) · **Last updated:** 2026-07-16
|
||
>
|
||
> Living document for making the Coupa **purchase-order** email parser *template-first with LLM fallback*, mirroring the work-order (WO) processor's PR #99. Captures the investigation, the data we gathered, the decisions made, the current scaffold, and everything still to do.
|
||
|
||
---
|
||
|
||
## 1. Context & Goal
|
||
|
||
The PO email processor (`lambdas/po/email_processor/handler.py`) sends **every** inbound Coupa/Amazon purchase-order email to an LLM (Claude Haiku 4.5 on Bedrock) for structured extraction. WO PR #99 established the pattern we want to replicate here:
|
||
|
||
> Run a **deterministic template parser first**. It returns a result only when the email provably conforms to a known template and passes a **fail-closed validation gate**; otherwise the handler falls back to the Bedrock LLM. This removes the LLM from the hot path for routine traffic while keeping full AI coverage for anything unexpected.
|
||
|
||
**Goal:** apply the same template-first + fail-closed-fallback approach to the PO processor, without regressing data quality or the downstream `po-ingest-site-extractor` pipeline (which reads `site_code`/`ship_to` off the `purchase-orders` DynamoDB stream).
|
||
|
||
**Reference implementation:** `lambdas/wo/email_processor/template_parser.py` (PR #99).
|
||
|
||
### Why this is harder than WO
|
||
- The PO contract is ~40 fields with **nested** objects (`supplier{}`, `ship_to{}`, `line_items[]`) vs WO's flat 16.
|
||
- Much of the PO extraction prompt is **derived classification** (`trade`, `site_code`, `fiscal_year`) already expressed as deterministic English rules.
|
||
- Money fields require `Decimal` (DynamoDB rejects floats).
|
||
- `email_type` is **not** in the subject (unlike WO's two subject templates).
|
||
|
||
---
|
||
|
||
## 2. Investigation
|
||
|
||
Two phases: a read-only multi-agent (ultracode) feasibility study over a 120-email sample, then a full-bucket triage over **all 3,448** inbound emails to get real distribution numbers.
|
||
|
||
### 2.1 Data access
|
||
> **Migration note (2026-07 / PLAT-67 2026-08-05):** this section is historical — the harvest ran
|
||
> against the management account. The live bucket is
|
||
> `s3://po-ingest-emails-011934824531/inbound/` (seahaven-prod, profile
|
||
> `seahaven-prod`). The mgmt bucket `po-ingest-emails-328440206208` remains as
|
||
> PLAT-67 cold archive (tagged; wipe deferred).
|
||
- PO email bucket: `s3://po-ingest-emails-328440206208/inbound/` (AWS account **328440206208**, us-east-1).
|
||
- ~~Reached via AWS profile `amoussa-mgmt` (SSO); default CLI creds are the personal account `681986854588`~~ **Stale (corrected 2026-07-16 during the fixture harvest):** the default CLI session is now authenticated to **328440206208** directly — no `--profile` flag needed.
|
||
- ⚠️ **Corpus is aging out:** `inbound/` objects carry an S3 lifecycle expiration (~90-day rolling window; oldest object 2026-04-17 at harvest time, 3,422 objects vs 3,448 at triage). Any further harvesting should not be deferred long.
|
||
|
||
### 2.2 Full-bucket triage results (all 3,448 emails, 2026-07-16)
|
||
|
||
Method: parallel ranged-GET (`bytes=0-12000`) of every object, then local classification (quoted-printable bodies, decodable offline).
|
||
|
||
| Email kind | Count | % | Notes |
|
||
|---|---:|---:|---|
|
||
| **new_po** — `***Copy for Reference*** New Purchase Order <PO#> has been issued` | 3,294 | **95.5%** | Single uniform Coupa layout |
|
||
| **cancellation** — `<Site> Purchase Order #<PO#> has been cancelled` | 100 | **2.9%** | Distinct subject + minimal body |
|
||
| **comment** — `New Comment on Purchase Order for Amazon` | 19 | 0.55% | Type NOT in handler enum (see §7) |
|
||
| **non-Coupa senders** (human replies, WO mail, an AWS SES setup notice) | 35 | 1.0% | Rejected at ses_auth before parse |
|
||
| **revision** | **0** | 0% | No distinct revision emails exist |
|
||
|
||
Within **new_po**:
|
||
- **Status** values: `Issued - Created` (2,223) and `Issued - Scheduled for email` (1,071) — only these two.
|
||
- **Multi-line-item:** 6 / 3,294 (**0.18%**), all ≤3 items.
|
||
- **Non-USD:** 0 (after excluding unit-token false positives like `EACH`/`HOUR`).
|
||
- So the single-line / USD scope covers **~99.8%** of new_po traffic.
|
||
|
||
### 2.3 new_po layout (the one template that matters)
|
||
|
||
`multipart/alternative`; the **text/plain** part is a stable, label-delimited flattening:
|
||
- Top: `Amazon Purchase Order #<PO#>`, then `Submitted By`, `On Behalf Of`, `Supplier`, `Total`, `Items`.
|
||
- `More Detail` block: `PO ID`, `Department`, `Status`, `Last Opened`, `Order Date`, `Acknowledged At`, `Revision Date`, `Payment Term`, `Req #`, `Shipping`.
|
||
- `Supplier` detail block (address).
|
||
- `Shipping` detail block (ship-to address, `Location Code:`, `Attn:`).
|
||
- `Lines` section: one block per line item; per-line metadata is **U+2022 (`•`) delimited**: `Need By` · `Category` · `Account` · `Period` [· optional `Part Number`].
|
||
|
||
**Structural hazards found in real data:**
|
||
- `Supplier`, `Shipping`, `Total` labels **each appear twice** (summary placeholder + detail block); the *first* `Shipping` value is literally `None`.
|
||
- Optional `Part Number` bullet segment appears in a minority of emails, inserted between `Category` and `Account` → shifts any positional splitter.
|
||
- Every captured value carries a trailing `\r`; a lone `\xa0` (nbsp) line sits before `Total`.
|
||
- Real amounts include thousands separators (`18,624.05`, `32,405.00`).
|
||
|
||
**Addendum (2026-07-16 fixture harvest):**
|
||
- Every per-line bullet run BEGINS with a `Supplier <name>` segment before `Need By` (the gate proves it byte-equals `supplier.name` on every item).
|
||
- Multi-item Lines blocks carry a `<qty> EA` evidence line before each description; the single-item Items summary carries an optional `<qty> <UNIT> x <price>` line (absent on some emails, e.g. new-po-14) which is the **source** for `quantity`/`unit`/`price` (the Lines-block `EA` line is a gate cross-check only).
|
||
- The Items summary contains **unit-price money tokens distinct from line amounts** (e.g. `50.0 EACH x 66.00` with line amount `8,356.00`) — money must be anchored on the `for <amt> <CCY>` / Total-block / summary-`x` contexts, never free money-shaped scanning; and line `amount ≠ qty×price` on partial quantities, so only `sum(lines) == total` is load-bearing.
|
||
- Body line endings are **decode-path dependent**: Coupa QP encodes `=0D`, so `policy.default` `get_content()` yields `\r\n` even after transport normalization, while other decode paths yield bare LF — the parser matches only `_clean`-ed lines and is tested against both representations.
|
||
|
||
### 2.4 cancellation layout
|
||
|
||
Subject: `<SiteName> Purchase Order #<PO#> has been cancelled`. Body is a short dashes-delimited notice — **no `Status` label**, no line items. The only field the handler's `save_cancellation()` needs is `po_number`.
|
||
|
||
---
|
||
|
||
## 3. Key Decisions
|
||
|
||
1. **Build two templates**, not one and not three:
|
||
- `coupa_new_po` (95.5%) — the main workload.
|
||
- `coupa_cancellation` (2.9%) — trivial and low-risk (just `po_number`); we have 100 real examples.
|
||
- Together ≈ **98.4%** of all mail deterministically handled/routed.
|
||
2. **No revision template** — 0 distinct revision emails in 3,448. The handler's `revision` type essentially never arrives as its own email.
|
||
3. **`email_type` never defaults to new_po.** It is emitted only when the exact new_po subject regex matches **and** `po_status` ∈ the two confirmed-safe strings. Anything else → LLM. (Misrouting a cancellation to new_po would silently defeat the sticky-`Cancelled` guard.)
|
||
4. **Fail closed, always.** Any miss / invalid / exception returns `None` and falls back to the LLM. A failure is never a parsed result. (Copies WO's `try/except` posture.)
|
||
5. **`Decimal` for all money**, with thousands-separator stripping, to match the LLM path's `parse_float=Decimal`.
|
||
6. **Derived fields (`site_code`, `trade`, `fiscal_year`) are NOT computed in the parser.** They are left `None`; a shared post-stage (`enrich_parsed()`, the `pad_zip` precedent) fills them **identically** on both the template and LLM paths, so the gate judges *extraction fidelity only* and both paths write byte-identical shapes downstream. `coupa_category` is a **verbatim** label capture, not a classifier, and IS extracted.
|
||
7. **Structural nested contract** with recursive key-set validation (unlike WO's flat tuple).
|
||
8. **Two-PR split:** (1) extraction template + gate + handler wiring; (2) derived-classifier factoring, shadow-logged before it becomes authoritative.
|
||
|
||
---
|
||
|
||
## 4. Contract
|
||
|
||
Top-level `CONTRACT_KEYS` (23) plus three nested sub-tuples. Taken from the `EXTRACTION_PROMPT` in `handler.py`.
|
||
|
||
```
|
||
CONTRACT_KEYS = (
|
||
email_type, po_number, po_status, source_system, submitted_by, on_behalf_of,
|
||
order_date, revision_date, last_opened, acknowledged_at, payment_terms,
|
||
requisition_number, department, view_order_url, supplier, site_code, ship_to,
|
||
total_amount, currency, fiscal_year, trade, coupa_category, line_items,
|
||
)
|
||
SUPPLIER_KEYS = (name,)
|
||
SHIP_TO_KEYS = (name, address, street, city, state, zip, location_code, attn)
|
||
LINE_ITEM_KEYS = (description, amount, currency, need_by, category,
|
||
account_code, period, quantity, unit, price)
|
||
```
|
||
|
||
| Class | Keys |
|
||
|---|---|
|
||
| **Verbatim / labeled extract** | po_number, po_status, submitted_by, on_behalf_of, order_date, revision_date, last_opened, acknowledged_at, payment_terms, requisition_number, department, view_order_url, supplier.name, ship_to.location_code, ship_to.attn, coupa_category, all `line_items[*]` |
|
||
| **Extract with care** | total_amount, currency, ship_to.name/street/city/state/zip (`Decimal`; sentinel-anchored address) |
|
||
| **Constant** | source_system = `"coupa"` |
|
||
| **Derived — post-stage, NOT parser, NOT gate-validated** | site_code, trade, fiscal_year |
|
||
|
||
Handler enrichment metadata (`raw_s3_key`, `processed_at`, `data_source`, `email_subject`, `ship_to_raw`, top-level `state`) is added by `enrich_parsed()` and is **not** part of the parser contract.
|
||
|
||
A parity test (mirroring WO's `test_contract_keys_match_extraction_prompt`) should assert the flattened contract key set is a subset of the keys quoted in `EXTRACTION_PROMPT`.
|
||
|
||
---
|
||
|
||
## 5. Parser Design
|
||
|
||
File: `lambdas/po/email_processor/template_parser.py` — pure module (no boto3/network), mirroring WO idioms.
|
||
|
||
- `classify_template(email_data) -> (template_id, reason)` — exact subject regex → `coupa_new_po` | `coupa_cancellation` | `unknown`.
|
||
- `extract_new_po` / `extract_cancellation` — build the nested candidate; leave derived fields `None`.
|
||
- `_empty_candidate()` / `_normalize()` — build & normalize the **nested** skeleton (`supplier={}`, `ship_to={}`, `line_items=[one item]`).
|
||
- `_clean()` — strip trailing `\r` + `\xa0`, map `"None"`/empty → `None`.
|
||
- `_to_decimal()` — thousands-separator-safe `Decimal`, `_UNPARSEABLE` sentinel on failure.
|
||
- `validate(candidate, template_id, email_data) -> (bool, reason)` — fail closed.
|
||
- `try_deterministic_parse(email_data) -> (parsed|None, method, template_id, reason)` — entry point; `try/except` fails closed on any error.
|
||
|
||
**Handler integration (✅ wired in PR #1):** after `authenticate_inbound_email()` + `parse_raw_email()`, `try_deterministic_parse` runs first; if `None`, `extract_with_claude`; the shared `enrich_parsed()` stage runs on `parsed` regardless of path, then the existing `save_cancellation` / `save_revision` / `save_new_po` routing (untouched). The `ParseMethod` EMF metric is emitted immediately after the deterministic attempt — before any Bedrock call — so a Bedrock-side error still records the `ai_fallback` outcome.
|
||
|
||
---
|
||
|
||
## 6. Fail-Closed Gate Rules
|
||
|
||
Reason codes emitted by `validate()` / `try_deterministic_parse()` (closed set):
|
||
|
||
`subject_no_match` · `key_set_mismatch` · `derived_field_set` · `email_type_mismatch` · `missing_required_field` · `po_id_mismatch` · `unrecognized_status` · `multiline_unsupported` · `non_usd` · `amount_mismatch` · `anchor_violation` · `address_shape_invalid` · `bullet_label_unrecognized` · `unparseable_value` · `residual_artifact` · `extractor_raised` · (`ok` / `template` on success).
|
||
|
||
(`new_po_not_implemented` disappeared with the scaffold guard — both copies removed; a test asserts the string no longer exists in the module.)
|
||
|
||
**Implemented (both templates):**
|
||
- Known template only (`subject_no_match` otherwise).
|
||
- Exact nested key-set at every level (`key_set_mismatch`).
|
||
- Derived fields must be unset by the parser (`derived_field_set`).
|
||
- `email_type` in enum and == template's expected type (`email_type_mismatch`).
|
||
- `po_number` valid shape `^[A-Z0-9]{1,6}-\d+$` **and** byte-equals the subject id (`po_id_mismatch` / `missing_required_field`).
|
||
|
||
**Implemented (new_po):** safe-status enum (`unrecognized_status`), single-line only (`multiline_unsupported`), USD only (`non_usd`) — enforced at **both** levels: the Total-block top-level currency must be exactly `USD` (rule 8) **and** every line item's captured currency plus its re-derived `for <amt> <CCY>` body token must byte-equal it (a single non-USD line item fails closed even when the Total block reads USD).
|
||
|
||
**Implemented (new_po value-level rules V1–V13 — ✅ done, one reason code each):** the gate **re-derives every byte proof from `email_data["body"]`** (never trusting extractor-carried state), with the six value-level reason codes:
|
||
- `amount_mismatch` — money fidelity (each Decimal re-serializes byte-identically to its re-derived source token, non-digit/comma border — the `18,624.05`→`624.05` kill switch), `sum(lines) == total` (exact Decimal; deliberately **no** `qty×price == amount` rule — corpus shows partial quantities), dual-`Total` byte-identity, and quantity/unit/price coherence vs the summary and `EA` evidence lines.
|
||
- `anchor_violation` — exactly 2× `Supplier`/`Shipping`/`Total`, first `Shipping` is the literal `None` placeholder, section ordering, `supplier.name` contains `SEA HAVEN` and byte-equals the detail-block name + every item's `Supplier` bullet segment, `ship_to.name != supplier.name`.
|
||
- `address_shape_invalid` — city line fullmatches `^City, ST ZIP$` immediately before `United States`, state in the frozen USPS set, zip `^\d{5}(-\d{4})?$` on the RAW pre-enrichment value (a short zip fails closed to the LLM path where the shared `pad_zip` repairs it — the gate never predicts enrichment).
|
||
- `bullet_label_unrecognized` — every U+2022 segment carries a recognized leading label from the closed set; core labels exactly once, `Part Number` at most once.
|
||
- `unparseable_value` — recursive walk: no `_UNPARSEABLE` sentinel leaf.
|
||
- `residual_artifact` — recursive walk: no `\r`/`\xa0` in any string leaf.
|
||
Plus `missing_required_field` for the required labeled fields (per-item metadata, `Location Code:` sourcing, no invented `Attn:`) and `po_id_mismatch` for the body `PO ID` / heading / `orders/<id>`-URL identity proofs.
|
||
|
||
**Precondition (handled a layer earlier):** `authenticate_inbound_email()` (SES `dkim=pass` for `amazon.coupahost.com`) must pass before the parser runs — the 35 non-Coupa emails are rejected there.
|
||
|
||
---
|
||
|
||
## 7. Scope — In / Out
|
||
|
||
**IN (template-parsed):**
|
||
- Single-line-item, USD, `new_po` with the exact subject and a safe `Status`.
|
||
- `cancellation` (subject match → `po_number`).
|
||
|
||
**OUT (always LLM fallback / rejected):**
|
||
- **Comments** (19) — `New Comment on Purchase Order` is an `email_type` the handler enum (`new_po`/`revision`/`cancellation`) does not model. The LLM currently shoehorns these. ⚠️ **Pre-existing data-quality gap, independent of this work** — decide separately whether comments should create/update PO records at all.
|
||
- **Revisions** — no template (0 emails).
|
||
- **Multi-line-item new_po** (0.18%) — `Lines`-array structure unobserved; `multiline_unsupported`.
|
||
- **Non-USD new_po** (0 observed) — path unexercised; `non_usd`.
|
||
- **Non-Coupa senders** (1.0%) — rejected at ses_auth.
|
||
|
||
---
|
||
|
||
## 8. Progress
|
||
|
||
- [x] Read-only ultracode feasibility investigation (8 agents, GO-with-conditions).
|
||
- [x] Full-bucket triage over all 3,448 emails → distribution + scope numbers (§2.2).
|
||
- [x] Confirmed cancellation & comment body layouts against real emails.
|
||
- [x] **Scaffold** `template_parser.py` on `feat/po-template-parser`:
|
||
- [x] Nested contract + recursive `_empty_candidate`/`_normalize`.
|
||
- [x] `classify_template` for both templates (real subject regexes).
|
||
- [x] `coupa_cancellation` extract + gate — **implemented & smoke-tested** (parses a real cancellation → `po_number`, `email_type=cancellation`).
|
||
- [x] `coupa_new_po` classify + structural/status/single-line/USD gate.
|
||
- [x] ~~Scaffold guard~~ — removed (both copies) with the PR #1 implementation below.
|
||
- [x] `ruff check` + `ruff format --check` pass.
|
||
- [x] **PR #1 implementation (2026-07-16):**
|
||
- [x] `extract_new_po` — section-windowed, anchor-disciplined, label-keyed bullet split, `Decimal` money, representation-agnostic line handling. All 17 harvested single-line new-POs parse to `template/ok`; the 3 real multi-line ones extract faithfully (`sum==total`) then gate-reject with `multiline_unsupported`.
|
||
- [x] Value-level gate rules V1–V13 (§6) — every byte proof re-derived from the body.
|
||
- [x] Handler wiring: `try_deterministic_parse` first, Bedrock fallback on `None`, shared `enrich_parsed()` on both paths, `save_*` routing untouched.
|
||
- [x] `ParseMethod` EMF metric set (`Seahaven/PoIngest`/`ParseOutcome`, dims `[["ParseMethod"],["ParseMethod","TemplateId"]]`, `ReasonCode`/`po_number` ride-alongs), emitted before the Bedrock call.
|
||
- [x] Fallback-rate alarm in `cdk/po_stack.py`, retuned for ~57/day (6h periods · ≥8 volume floor · >20% · 2-of-4) — `npx cdk synth po-ingest` green.
|
||
- [x] `cdk/po_stack.py` bundling `cp` now includes `template_parser.py` (was a deploy-time ImportError waiting to happen).
|
||
- [x] Offline test suite `lambdas/po/email_processor/tests/` (136 tests): goldens for all 25 positive fixtures, every §6 reason code covered, dual line-ending identity, two-path enrich/save parity, fixture hygiene (ses_auth + leak-sweep markers), Bedrock dispatch/EMF. Full-root pytest: 358 green.
|
||
- [x] Committed sanitized `.eml` fixture corpus (header + body scrub, per-file digit cipher, narrow `.gitignore` exception scoped to `lambdas/po/email_processor/tests/fixtures/`).
|
||
- [x] README updated (PO flow, parser section, alarm numbers + justification, tests, repo layout).
|
||
|
||
### PR #2 implementation notes (2026-07-16)
|
||
|
||
- **Shared classifier factoring.** `trade`/`site_code`/`fiscal_year` are computed by the pure, total `derived_fields.derive_all(parsed) -> {site_code, trade, fiscal_year}` and applied inside the shared `enrich_parsed()` post-stage, so both parse paths run identical classification.
|
||
- **Shadow semantics — fill gaps, never overwrite.** For each field: if the incoming value is `null` and Python has a value, Python fills it (both paths); if a value is already present (only possible from the LLM on the ai_fallback path) it is **kept** — the LLM stays authoritative during the bake. The template path always arrives with all three `null` (the gate enforces the parser leaves them unset), so it is effectively Python-authoritative.
|
||
- **EMF metric shape.** On the **ai_fallback path only**, one `DerivedFieldAgreement` record per field is emitted (namespace `Seahaven/PoIngest`, value 1, dims `[["Field","Agreement"]]`). `Agreement` ∈ `agree` | `disagree` | `llm_null_python_filled` | `python_null`; skipped when both values are `null`. `po_number`/`PythonValue`/`LlmValue` are non-dimensioned ride-alongs (cardinality fixed at Field × Agreement). The template path emits **no** agreement metric. The whole block is wrapped defensively so no classification/telemetry error can fail this S3-async invocation (an uncaught exception → retry storm → DLQ).
|
||
- **`quantity`/`price` prompt tweak (deferred from PR #1) — done.** `EXTRACTION_PROMPT` now declares line-item `quantity`/`price` as `"number or null"` (was `"string or null"`); `parse_float=Decimal` already handles numerics and the `enrich_parsed` Decimal coercion stays as the safety net for a non-conforming model. The `site_code`/`fiscal_year`/`trade` **rule sections of the prompt are left intact** — the LLM stays authoritative on the fallback path until the post-bake follow-up.
|
||
- **Backtest agreement stats (final, 2026-07-16).** Harness: every harvested real inbound email (3,422 objects; corpus is the rolling ~90-day S3 window) → template parse → `derive_all()` → per-field compare against the LLM-written values in `purchase-orders` (17,605-record scan). Population: 3,190 template-parsed new_po / 100 cancellations / 132 fallbacks (matches PR #1 triage). The `llm_null_python_filled` buckets (478/577/564 per field) are all records processed pre-May 2026 with no `data_source` attribute — an older prompt/writer era; Python filling those is strictly additive. On the modern-era comparable set:
|
||
- `site_code` — 2,695/2,712 **99.4%**, and **zero Python-wrong**: all 17 misses are LLM errors (16× PO-prefix `2D` stored as a site code, 1× `4101` garbage vs Python's correct `DRN5`).
|
||
- `fiscal_year` — 2,626/2,626 **100%**.
|
||
- `trade` — 2,188/2,613 raw (83.7%); **96.8% ex-deliberate**. The 425 mismatches decompose: 316 deliberate prompt-faithful `General Building - General Building Project` → `General Building - Project` (the LLM ignored the prompt's own BBM rule — biggest intentional behavior shift, ~10% of POs); 14 bollard → `Fencing/Gates` (explicit prompt keyword); 7 `General Building Technician` → Handyman (prompt rule); 16 LLM off-menu labels (`Electrical - Emergency`, `Plumbing - Project`, `General Building - Locksmith` — closed label set wins); 18 BBM plumbing rows where the LLM is internally inconsistent (the ladder follows the per-row LLM majority, e.g. Technician→PM 116:9, fixture-repair→Reactive ~140:2); ~54 residual free-text one-offs (2.1%).
|
||
- Corpus-driven rules beyond the prompt (recorded here as the authoritative spec deltas): site-code shapes for code-first dash names (`WPT2 - Amazon.com Services LLC`), leading token (`QDE1 (Co-located inside WDE1)`), mid-name token with digit (`Amazon Fresh UVA5 Non-Inv (Prime)`), free description token 4-5 chars w/ digit (`... aqt DSF7`), rightmost-token parens/ATTN resolution, ≥2-alphabetic-char token floor (kills `B187`, `2D`, `4101`), skip-list additions LLC/INC/CORP/LTD/ATTN; trade BBM plumbing 5-rung ladder and `conveyance` keyword.
|
||
|
||
---
|
||
|
||
## 9. TODO / Path Forward
|
||
|
||
**PR #1 — extraction template + wiring**
|
||
- [x] Implement `extract_new_po`: labeled single-occurrence fields; duplicate-label anchoring (`Supplier`/`Shipping`/`Total`); `ship_to` by sentinel anchors; `line_items` split on `•` by leading label; `Decimal` amounts; `coupa_category` verbatim.
|
||
- [x] Implement the new_po value-level gate rules (§6 V1–V13) and remove the scaffold guard (both copies).
|
||
- [x] Handler integration: `try_deterministic_parse` first, Bedrock fallback on `None`; route through existing `save_*` branches.
|
||
- [x] Emit `ParseMethod`-only EMF set (namespace `Seahaven/PoIngest`, dims `[["ParseMethod"], ["ParseMethod","TemplateId"]]`).
|
||
- [x] Fallback-rate CloudWatch alarm in `cdk/po_stack.py` — retuned for ~57/day: 6h periods, `IF((fb+tmpl)>=8, …)` volume floor, >20% threshold, eval 4 / datapoints 2 (see the po_stack comment + README for the arithmetic; `npx cdk synth po-ingest` gated).
|
||
- [x] Offline tests + fixtures: goldens for every positive fixture; one adversarial per gate rule (thousands-sep, dup-label swap, Part-Number bullet shift pair, multiline, non-USD, bad status, PO-id/URL mismatch, unlabeled/duplicate bullets, prompt injection); comment + non-Coupa assertions; contract⊆prompt parity test.
|
||
- [x] README updated. Confluence "AWS Architecture Map" update — ⚠️ outstanding (must land with the merge; PO subgraph gains the template-parser stage + fallback alarm).
|
||
- [x] `cross_review.py` (GPT-4.1) pass — run 2026-07-16 against the real handler diff: **no BLOCK, no security findings, "safe to merge with minor fixes."** Both FIX items verified as no-change-needed: the fallback log line already reports `ai_fallback` correctly (a Claude-side exception propagates before it, unchanged posture), and the non-dict-from-Claude hazard is the pre-existing issue #101 pattern this PR deliberately leaves untouched.
|
||
- [x] `ruff check` + `ruff format` + `pytest` green locally (354 tests; CI must confirm).
|
||
|
||
**PR #2 — derived-classifier factoring (follow-up)**
|
||
- [x] Move `trade`/`site_code`/`fiscal_year` into shared `enrich_parsed()` (via `derived_fields.derive_all()`).
|
||
- [x] Run in **shadow mode**: compute Python values, log Python-vs-LLM disagreement via the `DerivedFieldAgreement` EMF metric, keep LLM authoritative during a bake period.
|
||
- [ ] Harden `site_code` (7-shape + skip-list, incl. multi-hop ATTN like `CBRE - RME - DLI6`) and `trade` against the full sample with per-shape fixtures.
|
||
- [ ] Only after acceptable agreement: make Python authoritative and drop the fields from `EXTRACTION_PROMPT`.
|
||
|
||
---
|
||
|
||
## 10. Open Questions
|
||
|
||
1. **Goal — cost, latency, or determinism?** Determinism (same input → same output) is the strongest justification for the derived-classifier port regardless of per-email cost.
|
||
2. **Classifier factoring now or deferred?** Recommendation: extraction template first (PR #1), factoring as shadow-logged PR #2.
|
||
3. **Strict single-line/USD-only v1 acceptable?** ✅ **Resolved (PR #1): yes.** Covers ~99.8% of new_po; multi-line and non-USD fail closed with honest reason codes.
|
||
4. **Confirm the real cancellation `Status` spelling** — `handler.py` hardcodes `CANCELLED_STATUS='Cancelled'`; the cancellation body doesn't restate it. Verify against a known-cancelled PO in the table. (Doesn't affect parser output — `save_cancellation` hardcodes the string.)
|
||
5. **Comments** — should `New Comment on Purchase Order` emails create/update PO records at all, or be dropped? (Out of scope for the parser, but a real data-quality decision.)
|
||
6. **Repo fixtures** — ✅ **Resolved (PR #1): committed sanitized corpus** under `lambdas/po/email_processor/tests/fixtures/` (35 scrubbed `.eml`: 17 single-line new-PO, 8 cancellations, 4 comments, 3 non-Coupa, 3 multi-line; plus 18 synthetic adversarial mutations). Transport/auth headers scrubbed to same-shape placeholders (structure preserved, `ses_auth` still passes), per-file digit cipher on identifiers, amounts remapped with `sum==total` re-established, and a narrow `.gitignore` exception scoped to the PO fixtures path (never blanket `*.eml`). Source S3 keys deliberately unrecorded (they ARE the SES receipt tokens the scrub replaced).
|
||
|
||
**Fixture-build pins recorded during PR #1 (locked by goldens):**
|
||
- **`ship_to.street`/`ship_to.address` join convention:** multi-line segments joined with `'\n'`; `address` = the lines from `ship_to.name` through `United States` inclusive (this is what `enrich_parsed` promotes to `ship_to_raw` on the site-extractor stream).
|
||
- **`quantity`/`unit`/`price` source:** the Items-summary `<qty> <UNIT> x <price>` line (e.g. `1.0 EACH x 55,206.00`); the Lines-block `<qty> EA` line is gate evidence only (V13 numeric cross-check). Emails without a summary line (e.g. new-po-14) leave all three `null` — nullable by contract.
|
||
- **V10 URL-id sweep outcome:** all 20 harvested new_po emails satisfy `orders/<id> == po_number` digit suffix — the strict V10 check stays live (no demotion needed); a corpus-sweep test locks it.
|
||
- **Canonical `quantity`/`price` type = Decimal (DynamoDB Number), converged in the shared `enrich_parsed`:** `EXTRACTION_PROMPT` declares both fields as JSON `"number or null"` (see the PR #2 prompt tweak above — `prompts.py:67-69`), not JSON strings; `parse_float=Decimal` already handles a conforming numeric response the same as the template path's `Decimal`. The shared post-stage's numeric-string-to-`Decimal` coercion (thousands-separator-safe) remains as a defensive net for a non-conforming model response that arrives as `str` anyway; non-numeric strings are left verbatim. `prompts.py` itself is frozen this phase — this is a doc-only correction. The two-path parity test feeds a prompt-shaped payload (string quantity/price, LLM-filled `site_code`) — never the parser-derived golden verbatim — so it cannot be circular.
|
||
- **Second-pass header scrub (post-review):** the first-pass harvest scrub left the real SES `Feedback-ID` sender-identity hash on every Coupa fixture and, on the two non-Coupa fixtures, an embedded second SES block's `X-Ses-Receipt`, the Exchange cross-tenant UPN ciphertext, Gmail ARC `fh=` / `X-Gm-*` tokens, and related opaque routing blobs. All replaced with same-shape `ScrubbedFixture` placeholders (byte-safe, CRLF preserved); the fixture-hygiene test now asserts these token classes are scrubbed in every header block so regressions are caught.
|
||
|
||
---
|
||
|
||
## 11. Risk Register (high-severity)
|
||
|
||
| Risk | Mitigation |
|
||
|---|---|
|
||
| **Cancellation misrouted to new_po** → sticky-`Cancelled` guard defeated, PO stays active | `email_type` closed-set gate; never default-to-new_po; only emit on exact subject + safe status |
|
||
| **Thousands-separator truncation** (`18,624.05`→`624.05`, ~30× too small, passes naive checks) | Amount byte-equality: re-serialize `Decimal` to source grouped format; reject if token bordered by digit/comma; `sum(lines)==total` |
|
||
| **Multi-line structure unobserved** | Hard-fail `len(line_items)!=1` → LLM; profile real multi-line before supporting |
|
||
| **Derived-field silent divergence** (wrong `trade`/`site_code` still well-formed; `site_code` rules already incomplete) | Factor into shared post-stage; keep OFF the gate; shadow-log before authoritative; per-shape fixtures |
|
||
| **Duplicate-label first-match** (`Supplier`/`Shipping`/`Total` twice; first `Shipping`=`None`) | Occurrence/sentinel anchoring; gate rejects `Shipping=='None'`, requires `Location Code:`, `supplier.name` has `SEA HAVEN`, `ship_to.name != supplier.name` |
|
||
|
||
---
|
||
|
||
## Appendix — References
|
||
|
||
- WO reference parser: `lambdas/wo/email_processor/template_parser.py` (PR #99).
|
||
- PO handler + `EXTRACTION_PROMPT`: `lambdas/po/email_processor/handler.py`.
|
||
- PO stack (alarms, IAM, Bedrock): `cdk/po_stack.py`.
|
||
- Bedrock model: inference profile `us.anthropic.claude-haiku-4-5-20251001-v1:0`.
|
||
- PO email bucket (historical, pre-migration): `s3://po-ingest-emails-328440206208/inbound/`; post-migration `s3://po-ingest-emails-011934824531/inbound/` (seahaven-prod).
|