feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
# PO Template-First Parser — Design, Investigation & Progress
2026-07-17 11:47:33 -04:00
> **Status:** PR #1 implemented (extraction + gate + wiring + alarm + tests); PR #2 (derived-classifier factoring — shadow mode) implemented · **Branch:** `feat/po-template-parser` (stacked on PR #99) · **Last updated:** 2026-07-16
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
>
> Living document for making the Coupa **purchase-order** email parser *template-first with LLM fallback*, mirroring the work-order (WO) processor's PR #99. Captures the investigation, the data we gathered, the decisions made, the current scaffold, and everything still to do.
---
## 1. Context & Goal
The PO email processor (`lambdas/po/email_processor/handler.py` ) sends **every** inbound Coupa/Amazon purchase-order email to an LLM (Claude Haiku 4.5 on Bedrock) for structured extraction. WO PR #99 established the pattern we want to replicate here:
> Run a **deterministic template parser first**. It returns a result only when the email provably conforms to a known template and passes a **fail-closed validation gate**; otherwise the handler falls back to the Bedrock LLM. This removes the LLM from the hot path for routine traffic while keeping full AI coverage for anything unexpected.
**Goal:** apply the same template-first + fail-closed-fallback approach to the PO processor, without regressing data quality or the downstream `po-ingest-site-extractor` pipeline (which reads `site_code` /`ship_to` off the `purchase-orders` DynamoDB stream).
**Reference implementation:** `lambdas/wo/email_processor/template_parser.py` (PR #99 ).
### Why this is harder than WO
- The PO contract is ~40 fields with **nested** objects (`supplier{}` , `ship_to{}` , `line_items[]` ) vs WO's flat 16.
- Much of the PO extraction prompt is **derived classification** (`trade` , `site_code` , `fiscal_year` ) already expressed as deterministic English rules.
- Money fields require `Decimal` (DynamoDB rejects floats).
- `email_type` is **not** in the subject (unlike WO's two subject templates).
---
## 2. Investigation
Two phases: a read-only multi-agent (ultracode) feasibility study over a 120-email sample, then a full-bucket triage over **all 3,448** inbound emails to get real distribution numbers.
### 2.1 Data access
Migrate to seahaven-prod: deploy role, backfill tooling, account-portability fixes (#125)
* feat(migration): prepare stacks and tooling for the seahaven-prod account move
Phase 1 of the mgmt (328440206208) -> seahaven-prod (011934824531)
migration. No behavior change in-account; everything here is additive or
account-portability hygiene:
- infra/deploy-role/: reviewed OIDC deploy-role artifacts for prod
(trust main-only, cdk-hnb659fds-* AssumeRole, smoke-invoke-lambda scoped
to exactly the two email-processor fn ARNs). Codifies the previously
out-of-band smoke-invoke grant.
- Table resource policies: make_slack_bot_read_policy in cdk/common.py,
applied to purchase-orders, verified-sites, WorkOrders,
WorkOrderComments (NOT pending-site-review; no bot consumer). Grants the
mgmt-resident seahaven-slack-bot roles read-only cross-account access
post-move (bot-side identity grants land in the slack-bot repo).
- scripts/migrate_tables.py: dry-run-default backfill tool implementing
the plan's per-table semantics (superset overwrite, ingested_at cutoff
for WorkOrderComments, backup-gated truncate-and-load for the two
site tables) plus a verify subcommand (count parity, spot checks,
sticky-Cancelled drift check).
- tests/test_resource_policy_helper.py: statement-shape unit tests +
static pins that exactly the four bot-read tables carry the policy.
- Account-literal fixes: account-agnostic fixture bucket in
test_reprocess_contract; runbook/README/po-template-parser account
references updated to prod with historical mgmt notes; README gains the
account-prerequisites list (imported-by-name dependencies).
deploy.yaml is deliberately unchanged (push-to-main auto-deploy kept).
Merge is held until migration Phase 0 completes; flipping the
AWS_DEPLOY_ROLE_ARN repo secret and merging this PR IS the first prod
deploy.
* fix(migration): verify backup AVAILABLE pre-truncate; document wildcard risk acceptance (cross-review FIX/NIT)
* refactor(migration): drop cross-account read grants (slack-bot decommissioned); harden backfill + deploy role
seahaven-slack-bot was decommissioned 2026-07-23 (stack DELETE_IN_PROGRESS,
consumer Lambdas gone); its successor sh-mcp is undeployed and uses
same-account DynamoDB access. So no live consumer reads these tables
cross-account. Per Adam's call, drop the cross-account grants entirely and
re-add correctly-scoped ones if/when sh-mcp deploys to a different account.
- Remove the four table resource policies + make_slack_bot_read_policy helper
+ its constants (cdk/common.py, po_stack.py, wo_stack.py) and the helper's
unit test. Both stacks synth with zero table ResourcePolicy.
- scripts/migrate_tables.py hardening (fixes from the sh-security-review
fan-out on the destructive backfill tool):
* validate --cutoff strictly (parse ISO-8601, require aware UTC, re-emit
canonical second-precision form) so a malformed cutoff can't silently
copy dual-window rows or drop history;
* reject `copy --all` up front (must run tables individually, in order,
with the stream-drain wait) instead of writing three tables then erroring;
* truncate backup gate now also checks recency (<1h) and TableId, not just
status+name;
* verify requires --cutoff whenever a cutoff table is in scope (else it
false-flags dual-window rows as MISSING);
* sticky-cancel is now PREVENTED copy-side (a non-Cancelled source item
never overwrites a dest-Cancelled PO), and the verify comment no longer
overstates what its source-side scan covers;
* spot-check all modes (truncate_load keys are verbatim, so key-existence
is sound there too).
- Deploy role: scope cloudformation:DescribeStacks to this repo's stacks +
CDKToolkit (was Resource:*, disclosed all tenant stacks in the shared prod
account); add a drift check warning on unexpected role policies and drop the
dead SMOKE_POLICY_NAME var; document the shared-account bootstrap-role
accepted risk in the deploy-role README.
* docs(deploy-role): fold in cross-review NITs (DescribeStacks maintenance note, warn-only drift rationale)
* ci: update workflow to use new workflow tag (ruff versioning fix)
* fix(migration): address Open SWE review findings on migrate_tables.py
- Validate the truncate backup on dry-run as well as --execute so a
missing/stale/wrong-incarnation --backup-arn surfaces on the rehearsal
run (finding f_24a48b8900).
- Assert configured keys match the live key schema of both tables before
any key projection, turning config/schema drift into a descriptive
abort instead of a mid-backfill KeyError (finding f_cb6b5a6c59).
- Clarify why key-existence spot-checks are sound for WorkOrderComments:
the copy Puts source items verbatim and the sample uses the same
cutoff filter, so per-account comment_id divergence never enters the
check (finding f_390b7d6c3b is a false positive; comment hardened).
2026-07-23 17:08:47 -04:00
> **Migration note (2026-07):** this section is historical — the harvest ran
> against the management account. Post-migration the live bucket is
> `s3://po-ingest-emails-011934824531/inbound/` (seahaven-prod, profile
> `seahaven-prod`); the mgmt bucket persists only until decommission.
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
- PO email bucket: `s3://po-ingest-emails-328440206208/inbound/` (AWS account **328440206208** , us-east-1).
- ~~Reached via AWS profile `amoussa-mgmt` (SSO); default CLI creds are the personal account `681986854588` ~~ **Stale (corrected 2026-07-16 during the fixture harvest):** the default CLI session is now authenticated to **328440206208** directly — no `--profile` flag needed.
- ⚠️ **Corpus is aging out:** `inbound/` objects carry an S3 lifecycle expiration (~90-day rolling window; oldest object 2026-04-17 at harvest time, 3,422 objects vs 3,448 at triage). Any further harvesting should not be deferred long.
### 2.2 Full-bucket triage results (all 3,448 emails, 2026-07-16)
Method: parallel ranged-GET (`bytes=0-12000` ) of every object, then local classification (quoted-printable bodies, decodable offline).
| Email kind | Count | % | Notes |
|---|---:|---:|---|
| **new_po** — `***Copy for Reference*** New Purchase Order <PO#> has been issued` | 3,294 | **95.5%** | Single uniform Coupa layout |
| **cancellation** — `<Site> Purchase Order #<PO#> has been cancelled` | 100 | **2.9%** | Distinct subject + minimal body |
| **comment** — `New Comment on Purchase Order for Amazon` | 19 | 0.55% | Type NOT in handler enum (see §7) |
| **non-Coupa senders** (human replies, WO mail, an AWS SES setup notice) | 35 | 1.0% | Rejected at ses_auth before parse |
| **revision** | **0** | 0% | No distinct revision emails exist |
Within **new_po** :
- **Status** values: `Issued - Created` (2,223) and `Issued - Scheduled for email` (1,071) — only these two.
- **Multi-line-item:** 6 / 3,294 (**0.18%**), all ≤3 items.
- **Non-USD:** 0 (after excluding unit-token false positives like `EACH` /`HOUR` ).
- So the single-line / USD scope covers ** ~99.8%** of new_po traffic.
### 2.3 new_po layout (the one template that matters)
`multipart/alternative` ; the **text/plain** part is a stable, label-delimited flattening:
- Top: `Amazon Purchase Order #<PO#>` , then `Submitted By` , `On Behalf Of` , `Supplier` , `Total` , `Items` .
- `More Detail` block: `PO ID` , `Department` , `Status` , `Last Opened` , `Order Date` , `Acknowledged At` , `Revision Date` , `Payment Term` , `Req #` , `Shipping` .
- `Supplier` detail block (address).
- `Shipping` detail block (ship-to address, `Location Code:` , `Attn:` ).
- `Lines` section: one block per line item; per-line metadata is **U+2022 (`•`) delimited** : `Need By` · `Category` · `Account` · `Period` [· optional `Part Number` ].
**Structural hazards found in real data:**
- `Supplier` , `Shipping` , `Total` labels **each appear twice** (summary placeholder + detail block); the *first* `Shipping` value is literally `None` .
- Optional `Part Number` bullet segment appears in a minority of emails, inserted between `Category` and `Account` → shifts any positional splitter.
- Every captured value carries a trailing `\r` ; a lone `\xa0` (nbsp) line sits before `Total` .
- Real amounts include thousands separators (`18,624.05` , `32,405.00` ).
**Addendum (2026-07-16 fixture harvest):**
- Every per-line bullet run BEGINS with a `Supplier <name>` segment before `Need By` (the gate proves it byte-equals `supplier.name` on every item).
- Multi-item Lines blocks carry a `<qty> EA` evidence line before each description; the single-item Items summary carries an optional `<qty> <UNIT> x <price>` line (absent on some emails, e.g. new-po-14) which is the **source** for `quantity` /`unit` /`price` (the Lines-block `EA` line is a gate cross-check only).
- The Items summary contains **unit-price money tokens distinct from line amounts** (e.g. `50.0 EACH x 66.00` with line amount `8,356.00` ) — money must be anchored on the `for <amt> <CCY>` / Total-block / summary-`x` contexts, never free money-shaped scanning; and line `amount ≠ qty× price` on partial quantities, so only `sum(lines) == total` is load-bearing.
- Body line endings are **decode-path dependent** : Coupa QP encodes `=0D` , so `policy.default` `get_content()` yields `\r\n` even after transport normalization, while other decode paths yield bare LF — the parser matches only `_clean` -ed lines and is tested against both representations.
### 2.4 cancellation layout
Subject: `<SiteName> Purchase Order #<PO#> has been cancelled` . Body is a short dashes-delimited notice — **no `Status` label** , no line items. The only field the handler's `save_cancellation()` needs is `po_number` .
---
## 3. Key Decisions
1. **Build two templates** , not one and not three:
- `coupa_new_po` (95.5%) — the main workload.
- `coupa_cancellation` (2.9%) — trivial and low-risk (just `po_number` ); we have 100 real examples.
- Together ≈ **98.4%** of all mail deterministically handled/routed.
2. **No revision template** — 0 distinct revision emails in 3,448. The handler's `revision` type essentially never arrives as its own email.
3. ** `email_type` never defaults to new_po.** It is emitted only when the exact new_po subject regex matches **and** `po_status` ∈ the two confirmed-safe strings. Anything else → LLM. (Misrouting a cancellation to new_po would silently defeat the sticky-`Cancelled` guard.)
4. **Fail closed, always.** Any miss / invalid / exception returns `None` and falls back to the LLM. A failure is never a parsed result. (Copies WO's `try/except` posture.)
5. ** `Decimal` for all money**, with thousands-separator stripping, to match the LLM path's `parse_float=Decimal` .
6. **Derived fields (`site_code`, `trade`, `fiscal_year`) are NOT computed in the parser.** They are left `None` ; a shared post-stage (`enrich_parsed()` , the `pad_zip` precedent) fills them **identically** on both the template and LLM paths, so the gate judges *extraction fidelity only* and both paths write byte-identical shapes downstream. `coupa_category` is a **verbatim** label capture, not a classifier, and IS extracted.
7. **Structural nested contract** with recursive key-set validation (unlike WO's flat tuple).
8. **Two-PR split:** (1) extraction template + gate + handler wiring; (2) derived-classifier factoring, shadow-logged before it becomes authoritative.
---
## 4. Contract
Top-level `CONTRACT_KEYS` (23) plus three nested sub-tuples. Taken from the `EXTRACTION_PROMPT` in `handler.py` .
```
CONTRACT_KEYS = (
email_type, po_number, po_status, source_system, submitted_by, on_behalf_of,
order_date, revision_date, last_opened, acknowledged_at, payment_terms,
requisition_number, department, view_order_url, supplier, site_code, ship_to,
total_amount, currency, fiscal_year, trade, coupa_category, line_items,
)
SUPPLIER_KEYS = (name,)
SHIP_TO_KEYS = (name, address, street, city, state, zip, location_code, attn)
LINE_ITEM_KEYS = (description, amount, currency, need_by, category,
account_code, period, quantity, unit, price)
```
| Class | Keys |
|---|---|
| **Verbatim / labeled extract** | po_number, po_status, submitted_by, on_behalf_of, order_date, revision_date, last_opened, acknowledged_at, payment_terms, requisition_number, department, view_order_url, supplier.name, ship_to.location_code, ship_to.attn, coupa_category, all `line_items[*]` |
| **Extract with care** | total_amount, currency, ship_to.name/street/city/state/zip (`Decimal` ; sentinel-anchored address) |
| **Constant** | source_system = `"coupa"` |
| **Derived — post-stage, NOT parser, NOT gate-validated** | site_code, trade, fiscal_year |
Handler enrichment metadata (`raw_s3_key` , `processed_at` , `data_source` , `email_subject` , `ship_to_raw` , top-level `state` ) is added by `enrich_parsed()` and is **not** part of the parser contract.
A parity test (mirroring WO's `test_contract_keys_match_extraction_prompt` ) should assert the flattened contract key set is a subset of the keys quoted in `EXTRACTION_PROMPT` .
---
## 5. Parser Design
File: `lambdas/po/email_processor/template_parser.py` — pure module (no boto3/network), mirroring WO idioms.
- `classify_template(email_data) -> (template_id, reason)` — exact subject regex → `coupa_new_po` | `coupa_cancellation` | `unknown` .
- `extract_new_po` / `extract_cancellation` — build the nested candidate; leave derived fields `None` .
- `_empty_candidate()` / `_normalize()` — build & normalize the **nested** skeleton (`supplier={}` , `ship_to={}` , `line_items=[one item]` ).
- `_clean()` — strip trailing `\r` + `\xa0` , map `"None"` /empty → `None` .
- `_to_decimal()` — thousands-separator-safe `Decimal` , `_UNPARSEABLE` sentinel on failure.
- `validate(candidate, template_id, email_data) -> (bool, reason)` — fail closed.
- `try_deterministic_parse(email_data) -> (parsed|None, method, template_id, reason)` — entry point; `try/except` fails closed on any error.
**Handler integration (✅ wired in PR #1 ):** after `authenticate_inbound_email()` + `parse_raw_email()` , `try_deterministic_parse` runs first; if `None` , `extract_with_claude` ; the shared `enrich_parsed()` stage runs on `parsed` regardless of path, then the existing `save_cancellation` / `save_revision` / `save_new_po` routing (untouched). The `ParseMethod` EMF metric is emitted immediately after the deterministic attempt — before any Bedrock call — so a Bedrock-side error still records the `ai_fallback` outcome.
---
## 6. Fail-Closed Gate Rules
Reason codes emitted by `validate()` / `try_deterministic_parse()` (closed set):
`subject_no_match` · `key_set_mismatch` · `derived_field_set` · `email_type_mismatch` · `missing_required_field` · `po_id_mismatch` · `unrecognized_status` · `multiline_unsupported` · `non_usd` · `amount_mismatch` · `anchor_violation` · `address_shape_invalid` · `bullet_label_unrecognized` · `unparseable_value` · `residual_artifact` · `extractor_raised` · (`ok` / `template` on success).
(`new_po_not_implemented` disappeared with the scaffold guard — both copies removed; a test asserts the string no longer exists in the module.)
**Implemented (both templates):**
- Known template only (`subject_no_match` otherwise).
- Exact nested key-set at every level (`key_set_mismatch` ).
- Derived fields must be unset by the parser (`derived_field_set` ).
- `email_type` in enum and == template's expected type (`email_type_mismatch` ).
- `po_number` valid shape `^[A-Z0-9]{1,6}-\d+$` **and** byte-equals the subject id (`po_id_mismatch` / `missing_required_field` ).
**Implemented (new_po):** safe-status enum (`unrecognized_status` ), single-line only (`multiline_unsupported` ), USD only (`non_usd` ) — enforced at **both** levels: the Total-block top-level currency must be exactly `USD` (rule 8) **and** every line item's captured currency plus its re-derived `for <amt> <CCY>` body token must byte-equal it (a single non-USD line item fails closed even when the Total block reads USD).
**Implemented (new_po value-level rules V1– V13 — ✅ done, one reason code each):** the gate **re-derives every byte proof from `email_data["body"]`** (never trusting extractor-carried state), with the six value-level reason codes:
- `amount_mismatch` — money fidelity (each Decimal re-serializes byte-identically to its re-derived source token, non-digit/comma border — the `18,624.05` →`624.05` kill switch), `sum(lines) == total` (exact Decimal; deliberately **no** `qty× price == amount` rule — corpus shows partial quantities), dual-`Total` byte-identity, and quantity/unit/price coherence vs the summary and `EA` evidence lines.
- `anchor_violation` — exactly 2× `Supplier` /`Shipping` /`Total` , first `Shipping` is the literal `None` placeholder, section ordering, `supplier.name` contains `SEA HAVEN` and byte-equals the detail-block name + every item's `Supplier` bullet segment, `ship_to.name != supplier.name` .
- `address_shape_invalid` — city line fullmatches `^City, ST ZIP$` immediately before `United States` , state in the frozen USPS set, zip `^\d{5}(-\d{4})?$` on the RAW pre-enrichment value (a short zip fails closed to the LLM path where the shared `pad_zip` repairs it — the gate never predicts enrichment).
- `bullet_label_unrecognized` — every U+2022 segment carries a recognized leading label from the closed set; core labels exactly once, `Part Number` at most once.
- `unparseable_value` — recursive walk: no `_UNPARSEABLE` sentinel leaf.
- `residual_artifact` — recursive walk: no `\r` /`\xa0` in any string leaf.
Plus `missing_required_field` for the required labeled fields (per-item metadata, `Location Code:` sourcing, no invented `Attn:` ) and `po_id_mismatch` for the body `PO ID` / heading / `orders/<id>` -URL identity proofs.
**Precondition (handled a layer earlier):** `authenticate_inbound_email()` (SES `dkim=pass` for `amazon.coupahost.com` ) must pass before the parser runs — the 35 non-Coupa emails are rejected there.
---
## 7. Scope — In / Out
**IN (template-parsed):**
- Single-line-item, USD, `new_po` with the exact subject and a safe `Status` .
- `cancellation` (subject match → `po_number` ).
**OUT (always LLM fallback / rejected):**
- **Comments** (19) — `New Comment on Purchase Order` is an `email_type` the handler enum (`new_po` /`revision` /`cancellation` ) does not model. The LLM currently shoehorns these. ⚠️ **Pre-existing data-quality gap, independent of this work** — decide separately whether comments should create/update PO records at all.
- **Revisions** — no template (0 emails).
- **Multi-line-item new_po** (0.18%) — `Lines` -array structure unobserved; `multiline_unsupported` .
- **Non-USD new_po** (0 observed) — path unexercised; `non_usd` .
- **Non-Coupa senders** (1.0%) — rejected at ses_auth.
---
## 8. Progress
- [x] Read-only ultracode feasibility investigation (8 agents, GO-with-conditions).
- [x] Full-bucket triage over all 3,448 emails → distribution + scope numbers (§2.2).
- [x] Confirmed cancellation & comment body layouts against real emails.
- [x] **Scaffold** `template_parser.py` on `feat/po-template-parser` :
- [x] Nested contract + recursive `_empty_candidate` /`_normalize` .
- [x] `classify_template` for both templates (real subject regexes).
- [x] `coupa_cancellation` extract + gate — **implemented & smoke-tested** (parses a real cancellation → `po_number` , `email_type=cancellation` ).
- [x] `coupa_new_po` classify + structural/status/single-line/USD gate.
- [x] ~~Scaffold guard~~ — removed (both copies) with the PR #1 implementation below.
- [x] `ruff check` + `ruff format --check` pass.
- [x] **PR #1 implementation (2026-07-16):**
- [x] `extract_new_po` — section-windowed, anchor-disciplined, label-keyed bullet split, `Decimal` money, representation-agnostic line handling. All 17 harvested single-line new-POs parse to `template/ok` ; the 3 real multi-line ones extract faithfully (`sum==total` ) then gate-reject with `multiline_unsupported` .
- [x] Value-level gate rules V1– V13 (§6) — every byte proof re-derived from the body.
- [x] Handler wiring: `try_deterministic_parse` first, Bedrock fallback on `None` , shared `enrich_parsed()` on both paths, `save_*` routing untouched.
- [x] `ParseMethod` EMF metric set (`Seahaven/PoIngest` /`ParseOutcome` , dims `[["ParseMethod"],["ParseMethod","TemplateId"]]` , `ReasonCode` /`po_number` ride-alongs), emitted before the Bedrock call.
- [x] Fallback-rate alarm in `cdk/po_stack.py` , retuned for ~57/day (6h periods · ≥8 volume floor · >20% · 2-of-4) — `npx cdk synth po-ingest` green.
- [x] `cdk/po_stack.py` bundling `cp` now includes `template_parser.py` (was a deploy-time ImportError waiting to happen).
- [x] Offline test suite `lambdas/po/email_processor/tests/` (136 tests): goldens for all 25 positive fixtures, every §6 reason code covered, dual line-ending identity, two-path enrich/save parity, fixture hygiene (ses_auth + leak-sweep markers), Bedrock dispatch/EMF. Full-root pytest: 358 green.
- [x] Committed sanitized `.eml` fixture corpus (header + body scrub, per-file digit cipher, narrow `.gitignore` exception scoped to `lambdas/po/email_processor/tests/fixtures/` ).
- [x] README updated (PO flow, parser section, alarm numbers + justification, tests, repo layout).
2026-07-17 11:47:33 -04:00
### PR #2 implementation notes (2026-07-16)
- **Shared classifier factoring.** `trade` /`site_code` /`fiscal_year` are computed by the pure, total `derived_fields.derive_all(parsed) -> {site_code, trade, fiscal_year}` and applied inside the shared `enrich_parsed()` post-stage, so both parse paths run identical classification.
- **Shadow semantics — fill gaps, never overwrite.** For each field: if the incoming value is `null` and Python has a value, Python fills it (both paths); if a value is already present (only possible from the LLM on the ai_fallback path) it is **kept** — the LLM stays authoritative during the bake. The template path always arrives with all three `null` (the gate enforces the parser leaves them unset), so it is effectively Python-authoritative.
- **EMF metric shape.** On the **ai_fallback path only** , one `DerivedFieldAgreement` record per field is emitted (namespace `Seahaven/PoIngest` , value 1, dims `[["Field","Agreement"]]` ). `Agreement` ∈ `agree` | `disagree` | `llm_null_python_filled` | `python_null` ; skipped when both values are `null` . `po_number` /`PythonValue` /`LlmValue` are non-dimensioned ride-alongs (cardinality fixed at Field × Agreement). The template path emits **no** agreement metric. The whole block is wrapped defensively so no classification/telemetry error can fail this S3-async invocation (an uncaught exception → retry storm → DLQ).
- **`quantity` /`price` prompt tweak (deferred from PR #1 ) — done.** `EXTRACTION_PROMPT` now declares line-item `quantity` /`price` as `"number or null"` (was `"string or null"` ); `parse_float=Decimal` already handles numerics and the `enrich_parsed` Decimal coercion stays as the safety net for a non-conforming model. The `site_code` /`fiscal_year` /`trade` **rule sections of the prompt are left intact** — the LLM stays authoritative on the fallback path until the post-bake follow-up.
- **Backtest agreement stats (final, 2026-07-16).** Harness: every harvested real inbound email (3,422 objects; corpus is the rolling ~90-day S3 window) → template parse → `derive_all()` → per-field compare against the LLM-written values in `purchase-orders` (17,605-record scan). Population: 3,190 template-parsed new_po / 100 cancellations / 132 fallbacks (matches PR #1 triage). The `llm_null_python_filled` buckets (478/577/564 per field) are all records processed pre-May 2026 with no `data_source` attribute — an older prompt/writer era; Python filling those is strictly additive. On the modern-era comparable set:
- `site_code` — 2,695/2,712 **99.4%** , and **zero Python-wrong** : all 17 misses are LLM errors (16× PO-prefix `2D` stored as a site code, 1× `4101` garbage vs Python's correct `DRN5` ).
- `fiscal_year` — 2,626/2,626 **100%** .
- `trade` — 2,188/2,613 raw (83.7%); **96.8% ex-deliberate** . The 425 mismatches decompose: 316 deliberate prompt-faithful `General Building - General Building Project` → `General Building - Project` (the LLM ignored the prompt's own BBM rule — biggest intentional behavior shift, ~10% of POs); 14 bollard → `Fencing/Gates` (explicit prompt keyword); 7 `General Building Technician` → Handyman (prompt rule); 16 LLM off-menu labels (`Electrical - Emergency` , `Plumbing - Project` , `General Building - Locksmith` — closed label set wins); 18 BBM plumbing rows where the LLM is internally inconsistent (the ladder follows the per-row LLM majority, e.g. Technician→PM 116:9, fixture-repair→Reactive ~140:2); ~54 residual free-text one-offs (2.1%).
- Corpus-driven rules beyond the prompt (recorded here as the authoritative spec deltas): site-code shapes for code-first dash names (`WPT2 - Amazon.com Services LLC` ), leading token (`QDE1 (Co-located inside WDE1)` ), mid-name token with digit (`Amazon Fresh UVA5 Non-Inv (Prime)` ), free description token 4-5 chars w/ digit (`... aqt DSF7` ), rightmost-token parens/ATTN resolution, ≥2-alphabetic-char token floor (kills `B187` , `2D` , `4101` ), skip-list additions LLC/INC/CORP/LTD/ATTN; trade BBM plumbing 5-rung ladder and `conveyance` keyword.
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
---
## 9. TODO / Path Forward
**PR #1 — extraction template + wiring**
- [x] Implement `extract_new_po` : labeled single-occurrence fields; duplicate-label anchoring (`Supplier` /`Shipping` /`Total` ); `ship_to` by sentinel anchors; `line_items` split on `•` by leading label; `Decimal` amounts; `coupa_category` verbatim.
- [x] Implement the new_po value-level gate rules (§6 V1– V13) and remove the scaffold guard (both copies).
- [x] Handler integration: `try_deterministic_parse` first, Bedrock fallback on `None` ; route through existing `save_*` branches.
- [x] Emit `ParseMethod` -only EMF set (namespace `Seahaven/PoIngest` , dims `[["ParseMethod"], ["ParseMethod","TemplateId"]]` ).
- [x] Fallback-rate CloudWatch alarm in `cdk/po_stack.py` — retuned for ~57/day: 6h periods, `IF((fb+tmpl)>=8, …)` volume floor, >20% threshold, eval 4 / datapoints 2 (see the po_stack comment + README for the arithmetic; `npx cdk synth po-ingest` gated).
- [x] Offline tests + fixtures: goldens for every positive fixture; one adversarial per gate rule (thousands-sep, dup-label swap, Part-Number bullet shift pair, multiline, non-USD, bad status, PO-id/URL mismatch, unlabeled/duplicate bullets, prompt injection); comment + non-Coupa assertions; contract⊆prompt parity test.
- [x] README updated. Confluence "AWS Architecture Map" update — ⚠️ outstanding (must land with the merge; PO subgraph gains the template-parser stage + fallback alarm).
- [x] `cross_review.py` (GPT-4.1) pass — run 2026-07-16 against the real handler diff: **no BLOCK, no security findings, "safe to merge with minor fixes."** Both FIX items verified as no-change-needed: the fallback log line already reports `ai_fallback` correctly (a Claude-side exception propagates before it, unchanged posture), and the non-dict-from-Claude hazard is the pre-existing issue #101 pattern this PR deliberately leaves untouched.
- [x] `ruff check` + `ruff format` + `pytest` green locally (354 tests; CI must confirm).
**PR #2 — derived-classifier factoring (follow-up)**
2026-07-17 11:47:33 -04:00
- [x] Move `trade` /`site_code` /`fiscal_year` into shared `enrich_parsed()` (via `derived_fields.derive_all()` ).
- [x] Run in **shadow mode** : compute Python values, log Python-vs-LLM disagreement via the `DerivedFieldAgreement` EMF metric, keep LLM authoritative during a bake period.
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
- [ ] Harden `site_code` (7-shape + skip-list, incl. multi-hop ATTN like `CBRE - RME - DLI6` ) and `trade` against the full sample with per-shape fixtures.
- [ ] Only after acceptable agreement: make Python authoritative and drop the fields from `EXTRACTION_PROMPT` .
---
## 10. Open Questions
1. **Goal — cost, latency, or determinism?** Determinism (same input → same output) is the strongest justification for the derived-classifier port regardless of per-email cost.
2. **Classifier factoring now or deferred?** Recommendation: extraction template first (PR #1 ), factoring as shadow-logged PR #2 .
3. **Strict single-line/USD-only v1 acceptable?** ✅ **Resolved (PR #1): yes.** Covers ~99.8% of new_po; multi-line and non-USD fail closed with honest reason codes.
4. **Confirm the real cancellation `Status` spelling** — `handler.py` hardcodes `CANCELLED_STATUS='Cancelled'` ; the cancellation body doesn't restate it. Verify against a known-cancelled PO in the table. (Doesn't affect parser output — `save_cancellation` hardcodes the string.)
5. **Comments** — should `New Comment on Purchase Order` emails create/update PO records at all, or be dropped? (Out of scope for the parser, but a real data-quality decision.)
6. **Repo fixtures** — ✅ **Resolved (PR #1): committed sanitized corpus** under `lambdas/po/email_processor/tests/fixtures/` (35 scrubbed `.eml` : 17 single-line new-PO, 8 cancellations, 4 comments, 3 non-Coupa, 3 multi-line; plus 18 synthetic adversarial mutations). Transport/auth headers scrubbed to same-shape placeholders (structure preserved, `ses_auth` still passes), per-file digit cipher on identifiers, amounts remapped with `sum==total` re-established, and a narrow `.gitignore` exception scoped to the PO fixtures path (never blanket `*.eml` ). Source S3 keys deliberately unrecorded (they ARE the SES receipt tokens the scrub replaced).
**Fixture-build pins recorded during PR #1 (locked by goldens):**
- **`ship_to.street` /`ship_to.address` join convention:** multi-line segments joined with `'\n'` ; `address` = the lines from `ship_to.name` through `United States` inclusive (this is what `enrich_parsed` promotes to `ship_to_raw` on the site-extractor stream).
- **`quantity` /`unit` /`price` source:** the Items-summary `<qty> <UNIT> x <price>` line (e.g. `1.0 EACH x 55,206.00` ); the Lines-block `<qty> EA` line is gate evidence only (V13 numeric cross-check). Emails without a summary line (e.g. new-po-14) leave all three `null` — nullable by contract.
- **V10 URL-id sweep outcome:** all 20 harvested new_po emails satisfy `orders/<id> == po_number` digit suffix — the strict V10 check stays live (no demotion needed); a corpus-sweep test locks it.
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
- **Canonical `quantity` /`price` type = Decimal (DynamoDB Number), converged in the shared `enrich_parsed` :** `EXTRACTION_PROMPT` declares both fields as JSON `"number or null"` (see the PR #2 prompt tweak above — `prompts.py:67-69` ), not JSON strings; `parse_float=Decimal` already handles a conforming numeric response the same as the template path's `Decimal` . The shared post-stage's numeric-string-to-`Decimal` coercion (thousands-separator-safe) remains as a defensive net for a non-conforming model response that arrives as `str` anyway; non-numeric strings are left verbatim. `prompts.py` itself is frozen this phase — this is a doc-only correction. The two-path parity test feeds a prompt-shaped payload (string quantity/price, LLM-filled `site_code` ) — never the parser-derived golden verbatim — so it cannot be circular.
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
- **Second-pass header scrub (post-review):** the first-pass harvest scrub left the real SES `Feedback-ID` sender-identity hash on every Coupa fixture and, on the two non-Coupa fixtures, an embedded second SES block's `X-Ses-Receipt` , the Exchange cross-tenant UPN ciphertext, Gmail ARC `fh=` / `X-Gm-*` tokens, and related opaque routing blobs. All replaced with same-shape `ScrubbedFixture` placeholders (byte-safe, CRLF preserved); the fixture-hygiene test now asserts these token classes are scrubbed in every header block so regressions are caught.
---
## 11. Risk Register (high-severity)
| Risk | Mitigation |
|---|---|
| **Cancellation misrouted to new_po** → sticky-`Cancelled` guard defeated, PO stays active | `email_type` closed-set gate; never default-to-new_po; only emit on exact subject + safe status |
| **Thousands-separator truncation** (`18,624.05` →`624.05` , ~30× too small, passes naive checks) | Amount byte-equality: re-serialize `Decimal` to source grouped format; reject if token bordered by digit/comma; `sum(lines)==total` |
| **Multi-line structure unobserved** | Hard-fail `len(line_items)!=1` → LLM; profile real multi-line before supporting |
| **Derived-field silent divergence** (wrong `trade` /`site_code` still well-formed; `site_code` rules already incomplete) | Factor into shared post-stage; keep OFF the gate; shadow-log before authoritative; per-shape fixtures |
| **Duplicate-label first-match** (`Supplier` /`Shipping` /`Total` twice; first `Shipping` =`None` ) | Occurrence/sentinel anchoring; gate rejects `Shipping=='None'` , requires `Location Code:` , `supplier.name` has `SEA HAVEN` , `ship_to.name != supplier.name` |
---
## Appendix — References
- WO reference parser: `lambdas/wo/email_processor/template_parser.py` (PR #99 ).
- PO handler + `EXTRACTION_PROMPT` : `lambdas/po/email_processor/handler.py` .
- PO stack (alarms, IAM, Bedrock): `cdk/po_stack.py` .
- Bedrock model: inference profile `us.anthropic.claude-haiku-4-5-20251001-v1:0` .
Migrate to seahaven-prod: deploy role, backfill tooling, account-portability fixes (#125)
* feat(migration): prepare stacks and tooling for the seahaven-prod account move
Phase 1 of the mgmt (328440206208) -> seahaven-prod (011934824531)
migration. No behavior change in-account; everything here is additive or
account-portability hygiene:
- infra/deploy-role/: reviewed OIDC deploy-role artifacts for prod
(trust main-only, cdk-hnb659fds-* AssumeRole, smoke-invoke-lambda scoped
to exactly the two email-processor fn ARNs). Codifies the previously
out-of-band smoke-invoke grant.
- Table resource policies: make_slack_bot_read_policy in cdk/common.py,
applied to purchase-orders, verified-sites, WorkOrders,
WorkOrderComments (NOT pending-site-review; no bot consumer). Grants the
mgmt-resident seahaven-slack-bot roles read-only cross-account access
post-move (bot-side identity grants land in the slack-bot repo).
- scripts/migrate_tables.py: dry-run-default backfill tool implementing
the plan's per-table semantics (superset overwrite, ingested_at cutoff
for WorkOrderComments, backup-gated truncate-and-load for the two
site tables) plus a verify subcommand (count parity, spot checks,
sticky-Cancelled drift check).
- tests/test_resource_policy_helper.py: statement-shape unit tests +
static pins that exactly the four bot-read tables carry the policy.
- Account-literal fixes: account-agnostic fixture bucket in
test_reprocess_contract; runbook/README/po-template-parser account
references updated to prod with historical mgmt notes; README gains the
account-prerequisites list (imported-by-name dependencies).
deploy.yaml is deliberately unchanged (push-to-main auto-deploy kept).
Merge is held until migration Phase 0 completes; flipping the
AWS_DEPLOY_ROLE_ARN repo secret and merging this PR IS the first prod
deploy.
* fix(migration): verify backup AVAILABLE pre-truncate; document wildcard risk acceptance (cross-review FIX/NIT)
* refactor(migration): drop cross-account read grants (slack-bot decommissioned); harden backfill + deploy role
seahaven-slack-bot was decommissioned 2026-07-23 (stack DELETE_IN_PROGRESS,
consumer Lambdas gone); its successor sh-mcp is undeployed and uses
same-account DynamoDB access. So no live consumer reads these tables
cross-account. Per Adam's call, drop the cross-account grants entirely and
re-add correctly-scoped ones if/when sh-mcp deploys to a different account.
- Remove the four table resource policies + make_slack_bot_read_policy helper
+ its constants (cdk/common.py, po_stack.py, wo_stack.py) and the helper's
unit test. Both stacks synth with zero table ResourcePolicy.
- scripts/migrate_tables.py hardening (fixes from the sh-security-review
fan-out on the destructive backfill tool):
* validate --cutoff strictly (parse ISO-8601, require aware UTC, re-emit
canonical second-precision form) so a malformed cutoff can't silently
copy dual-window rows or drop history;
* reject `copy --all` up front (must run tables individually, in order,
with the stream-drain wait) instead of writing three tables then erroring;
* truncate backup gate now also checks recency (<1h) and TableId, not just
status+name;
* verify requires --cutoff whenever a cutoff table is in scope (else it
false-flags dual-window rows as MISSING);
* sticky-cancel is now PREVENTED copy-side (a non-Cancelled source item
never overwrites a dest-Cancelled PO), and the verify comment no longer
overstates what its source-side scan covers;
* spot-check all modes (truncate_load keys are verbatim, so key-existence
is sound there too).
- Deploy role: scope cloudformation:DescribeStacks to this repo's stacks +
CDKToolkit (was Resource:*, disclosed all tenant stacks in the shared prod
account); add a drift check warning on unexpected role policies and drop the
dead SMOKE_POLICY_NAME var; document the shared-account bootstrap-role
accepted risk in the deploy-role README.
* docs(deploy-role): fold in cross-review NITs (DescribeStacks maintenance note, warn-only drift rationale)
* ci: update workflow to use new workflow tag (ruff versioning fix)
* fix(migration): address Open SWE review findings on migrate_tables.py
- Validate the truncate backup on dry-run as well as --execute so a
missing/stale/wrong-incarnation --backup-arn surfaces on the rehearsal
run (finding f_24a48b8900).
- Assert configured keys match the live key schema of both tables before
any key projection, turning config/schema drift into a descriptive
abort instead of a mid-backfill KeyError (finding f_cb6b5a6c59).
- Clarify why key-existence spot-checks are sound for WorkOrderComments:
the copy Puts source items verbatim and the sample uses the same
cutoff filter, so per-account comment_id divergence never enters the
check (finding f_390b7d6c3b is a false positive; comment hardened).
2026-07-23 17:08:47 -04:00
- PO email bucket (historical, pre-migration): `s3://po-ingest-emails-328440206208/inbound/` ; post-migration `s3://po-ingest-emails-011934824531/inbound/` (seahaven-prod).