feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
"""Deterministic template parser for Coupa purchase order emails.
|
|
|
|
|
|
|
|
|
|
|
|
Pure module: no boto3, no network. Runs ahead of the AI extraction path in the
|
|
|
|
|
|
purchase-order email processor. Only returns a parsed result when it is proven
|
|
|
|
|
|
conformant to one of the known Coupa templates; otherwise it fails closed and
|
|
|
|
|
|
signals the caller to fall back to the Bedrock AI extractor.
|
|
|
|
|
|
|
|
|
|
|
|
Mirrors the WO parser idioms (lambdas/wo/email_processor/template_parser.py):
|
|
|
|
|
|
classify_template -> extract -> validate (FAIL CLOSED) -> try_deterministic_parse
|
|
|
|
|
|
returning (parsed|None, method, template_id, reason). A failure is NEVER a
|
|
|
|
|
|
parsed result.
|
|
|
|
|
|
|
|
|
|
|
|
Full-bucket triage of all 3,448 inbound emails (2026-07-16) fixed the scope:
|
|
|
|
|
|
|
|
|
|
|
|
T1 coupa_new_po -- 3,294 / 3,448 (95.5%). Subject
|
|
|
|
|
|
"***Copy for Reference*** New Purchase Order <PO#> has been issued".
|
|
|
|
|
|
multipart/alternative; the text/plain part is a stable label-delimited
|
|
|
|
|
|
layout (PO ID / Status / Order Date / Revision Date / Payment Term / Req # /
|
|
|
|
|
|
Submitted By / On Behalf Of / Supplier block / Shipping block w/ Location
|
|
|
|
|
|
Code + Attn / a Lines section whose per-line metadata is U+2022-delimited:
|
|
|
|
|
|
Need By / Category / Account / Period [/ optional Part Number]).
|
|
|
|
|
|
Maps to email_type "new_po". 99.8% single-line-item, 100% USD.
|
|
|
|
|
|
|
|
|
|
|
|
T2 coupa_cancellation -- 100 / 3,448 (2.9%). Subject
|
|
|
|
|
|
"<SiteName> Purchase Order #<PO#> has been cancelled". Minimal body; the
|
|
|
|
|
|
only field the handler's save_cancellation() needs is po_number.
|
|
|
|
|
|
Maps to email_type "cancellation".
|
|
|
|
|
|
|
|
|
|
|
|
DELIBERATELY OUT OF SCOPE -> always AI fallback (never template-parsed):
|
|
|
|
|
|
* "New Comment on Purchase Order for Amazon" (19 / 3,448) -- an email_type the
|
|
|
|
|
|
handler enum does not model; do not fabricate a PO record deterministically.
|
|
|
|
|
|
* revision -- 0 distinct emails in 3,448; no template to build.
|
|
|
|
|
|
* any multi-line-item new_po (6 / 3,448) -- Lines-array structure unobserved.
|
|
|
|
|
|
* any non-USD new_po (0 observed) -- non-USD path entirely unexercised.
|
|
|
|
|
|
* non-Coupa senders (35 / 3,448) -- already rejected at the ses_auth layer.
|
|
|
|
|
|
|
|
|
|
|
|
Derived fields (site_code, trade, fiscal_year) are NOT computed here: the parser
|
|
|
|
|
|
leaves them None and a shared post-stage (handler enrich_parsed(), the pad_zip
|
|
|
|
|
|
precedent) fills them identically on both the template and the LLM path, so the
|
|
|
|
|
|
gate judges extraction fidelity only. coupa_category is a verbatim label capture,
|
|
|
|
|
|
not a classifier, and IS extracted here.
|
|
|
|
|
|
|
|
|
|
|
|
Entry point: try_deterministic_parse(email_data) -> (parsed|None, method,
|
|
|
|
|
|
template_id, reason_code).
|
|
|
|
|
|
"""
|
|
|
|
|
|
|
|
|
|
|
|
import re
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
from dataclasses import dataclass
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
from decimal import Decimal, InvalidOperation
|
|
|
|
|
|
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
# IMPLEMENTATION STATUS
|
|
|
|
|
|
# [x] contract + recursive skeleton/normalize
|
|
|
|
|
|
# [x] classify_template (both templates)
|
|
|
|
|
|
# [x] coupa_cancellation extract + gate
|
|
|
|
|
|
# [x] coupa_new_po extract -- labeled fields, duplicate-label anchoring,
|
|
|
|
|
|
# U+2022 line split, Decimal amounts (built against the scrubbed real
|
|
|
|
|
|
# fixture corpus in tests/fixtures/)
|
|
|
|
|
|
# [x] coupa_new_po value-level gate rules V1-V13 (see validate())
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
|
|
|
|
|
|
# The contract keys (23), EXACTLY -- mirrors the AI EXTRACTION_PROMPT fields.
|
|
|
|
|
|
CONTRACT_KEYS = (
|
|
|
|
|
|
"email_type",
|
|
|
|
|
|
"po_number",
|
|
|
|
|
|
"po_status",
|
|
|
|
|
|
"source_system",
|
|
|
|
|
|
"submitted_by",
|
|
|
|
|
|
"on_behalf_of",
|
|
|
|
|
|
"order_date",
|
|
|
|
|
|
"revision_date",
|
|
|
|
|
|
"last_opened",
|
|
|
|
|
|
"acknowledged_at",
|
|
|
|
|
|
"payment_terms",
|
|
|
|
|
|
"requisition_number",
|
|
|
|
|
|
"department",
|
|
|
|
|
|
"view_order_url",
|
|
|
|
|
|
"supplier",
|
|
|
|
|
|
"site_code",
|
|
|
|
|
|
"ship_to",
|
|
|
|
|
|
"total_amount",
|
|
|
|
|
|
"currency",
|
|
|
|
|
|
"fiscal_year",
|
|
|
|
|
|
"trade",
|
|
|
|
|
|
"coupa_category",
|
|
|
|
|
|
"line_items",
|
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
SUPPLIER_KEYS = ("name",)
|
|
|
|
|
|
|
|
|
|
|
|
SHIP_TO_KEYS = (
|
|
|
|
|
|
"name",
|
|
|
|
|
|
"address",
|
|
|
|
|
|
"street",
|
|
|
|
|
|
"city",
|
|
|
|
|
|
"state",
|
|
|
|
|
|
"zip",
|
|
|
|
|
|
"location_code",
|
|
|
|
|
|
"attn",
|
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
LINE_ITEM_KEYS = (
|
|
|
|
|
|
"description",
|
|
|
|
|
|
"amount",
|
|
|
|
|
|
"currency",
|
|
|
|
|
|
"need_by",
|
|
|
|
|
|
"category",
|
|
|
|
|
|
"account_code",
|
|
|
|
|
|
"period",
|
|
|
|
|
|
"quantity",
|
|
|
|
|
|
"unit",
|
|
|
|
|
|
"price",
|
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
# Derived fields the parser must leave None (filled by the shared post-stage).
|
|
|
|
|
|
DERIVED_KEYS = ("site_code", "trade", "fiscal_year")
|
|
|
|
|
|
|
|
|
|
|
|
SOURCE_SYSTEM = "coupa"
|
|
|
|
|
|
|
|
|
|
|
|
# email_type per template.
|
|
|
|
|
|
_TEMPLATE_EMAIL_TYPE = {
|
|
|
|
|
|
"coupa_new_po": "new_po",
|
|
|
|
|
|
"coupa_cancellation": "cancellation",
|
|
|
|
|
|
}
|
|
|
|
|
|
VALID_EMAIL_TYPES = {"new_po", "revision", "cancellation"}
|
|
|
|
|
|
|
|
|
|
|
|
# Only these Status strings were observed on new_po (2,223 + 1,071 of 3,294).
|
|
|
|
|
|
# Anything else fails closed to the LLM -- NEVER default-to-new_po.
|
|
|
|
|
|
NEW_PO_SAFE_STATUSES = {"Issued - Created", "Issued - Scheduled for email"}
|
|
|
|
|
|
|
PO ai-fallback fail-closed gate + prompt hardening (refactor phase 1) (#108)
* feat: PO ai-fallback fail-closed gate + prompt hardening, parity with #104 (refactor phase 1)
Ports WO's #104 AI-fallback security hardening to the PO email
processor, adapted for PO's nested contract instead of copying the
WO gate verbatim.
validate_ai_fallback() (template_parser.py) fail-closes raw Bedrock
output before it reaches enrich_parsed or any dispatch/save:
recursive key-set check with missing-key normalization (nested
contract: supplier{}, ship_to{}, line_items[]); po_number checked
against the same hardened prefix+hyphen+digits regex family that
guards the DynamoDB partition key the handler builds from it
(rejects fullwidth-digit and trailing-artifact injection); email_type
enforced against the {new_po, revision, cancellation} allow-list
before dispatch so a miss can never fall into the else -> save_new_po
branch; money fields accept Decimal/int/None only, matching PO's
parse_float=Decimal decode (a float-typed check would be wrong here).
A gate failure emits ParseMethod=ai_fallback_rejected and `continue`s
to the next record -- it never raises, so attacker-controlled input
can't churn the retry/DLQ path.
extract_with_claude() wraps the untrusted email in an <email> data
block and neutralizes forged <email>-tag lookalikes in the body with
the same linear-time regex approach as WO's _EMAIL_TAG_RE, and sets
temperature=0 on the Bedrock call.
Deliberate double-count: PO emits ParseMethod=ai_fallback before the
Bedrock call (so a Bedrock-side error still records the outcome), so
a rejected email always produces both an ai_fallback datapoint
(pre-call) and an ai_fallback_rejected datapoint (post-gate). This is
intentional, not a bug -- documented in handler.py, template_parser.py,
and the README.
cdk/po_stack.py: in-place property update to the existing
po-email-processor-template-fallback-rate alarm (same logical ID, no
rename/replacement) -- the fb/(fb+tmpl) expression is left
byte-identical to its pre-Phase-1 form and ai_fallback_rejected is
deliberately excluded from the numerator/denominator/volume floor,
since folding it in as WO does would double-count every rejection
(PO's pre-call emit already counts it once via fb). A net-new
EmailProcessorAiFallbackRejectedAlarm watches the rejected series on
its own, retuned for ~57 emails/day with the 6h/IF-floor/eval-4/
datapoints-2 idiom (not WO's 5-minute sparse idiom, which is
structurally dead at PO volume). Both alarms remain ALARM-only to
site-alerts, NOT_BREACHING, with no element-wise MAX in the math
(post-#102 rule).
* Block "Cancelled" po_status off the AI cancellation route
The AI-fallback gate type-checked po_status but let any string
through, unlike the template path which never emits "Cancelled" on a
new_po. Dispatch routes on email_type, so an AI-path new_po or revision
carrying po_status="Cancelled" would reach save_new_po/save_revision and
cancel a live PO via _merge_update's sticky-cancel write without ever
hitting save_cancellation. Reject the exact sticky marker on any
non-cancellation email_type so the AI path matches the template path's
guard; arbitrary non-marker status strings still pass.
email_type is already validated to the enum before this check, and a
cancellation reaches save_cancellation (which hardcodes the status), so
po_status stays irrelevant on that route.
2026-07-17 14:50:47 -04:00
|
|
|
|
# The sticky, authoritative cancellation status. MUST stay in sync with
|
|
|
|
|
|
# handler.CANCELLED_STATUS -- the marker the sticky-cancel ConditionExpression
|
|
|
|
|
|
# writes and compares against. The AI-fallback gate uses it to forbid a
|
|
|
|
|
|
# non-cancellation email_type from carrying "Cancelled" in po_status, so an
|
|
|
|
|
|
# AI-path new_po/revision cannot cancel a live PO off dispatch (parity with the
|
|
|
|
|
|
# template path, which never emits "Cancelled" on a new_po).
|
|
|
|
|
|
_CANCELLED_STATUS = "Cancelled"
|
|
|
|
|
|
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
# Subject classifiers (Python unfolds header continuation lines before we see them).
|
|
|
|
|
|
_NEW_PO_SUBJECT = re.compile(
|
|
|
|
|
|
r"^\*\*\*Copy for Reference\*\*\* New Purchase Order\s+(?P<po>\S+)\s+has been issued$"
|
|
|
|
|
|
)
|
|
|
|
|
|
# Fully anchored, symmetric with _NEW_PO_SUBJECT: the whole subject must be
|
|
|
|
|
|
# "<SiteName> Purchase Order #<PO#> has been cancelled" -- the SiteName prefix
|
|
|
|
|
|
# is bounded (no '#', no newline, <=80 chars) so a subject that merely *ends*
|
|
|
|
|
|
# with the cancellation phrase (e.g. a forwarded/quoted thread, or arbitrary
|
|
|
|
|
|
# prefix text before the tail) is NOT misclassified as a cancellation and
|
|
|
|
|
|
# routed to the sticky-Cancelled write. Matched with .match (see
|
|
|
|
|
|
# classify_template), never .search.
|
|
|
|
|
|
_CANCELLATION_SUBJECT = re.compile(
|
|
|
|
|
|
r"^(?P<site>[^#\n]{0,80}?)Purchase Order\s+#(?P<po>[A-Z0-9-]+)\s+has been cancelled\s*$"
|
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
# PO number shape, e.g. 2D-21456967, FK-21920384, B187-17955555.
|
|
|
|
|
|
_PO_ID_RE = re.compile(r"^[A-Z0-9]{1,6}-\d+$")
|
|
|
|
|
|
|
PO ai-fallback fail-closed gate + prompt hardening (refactor phase 1) (#108)
* feat: PO ai-fallback fail-closed gate + prompt hardening, parity with #104 (refactor phase 1)
Ports WO's #104 AI-fallback security hardening to the PO email
processor, adapted for PO's nested contract instead of copying the
WO gate verbatim.
validate_ai_fallback() (template_parser.py) fail-closes raw Bedrock
output before it reaches enrich_parsed or any dispatch/save:
recursive key-set check with missing-key normalization (nested
contract: supplier{}, ship_to{}, line_items[]); po_number checked
against the same hardened prefix+hyphen+digits regex family that
guards the DynamoDB partition key the handler builds from it
(rejects fullwidth-digit and trailing-artifact injection); email_type
enforced against the {new_po, revision, cancellation} allow-list
before dispatch so a miss can never fall into the else -> save_new_po
branch; money fields accept Decimal/int/None only, matching PO's
parse_float=Decimal decode (a float-typed check would be wrong here).
A gate failure emits ParseMethod=ai_fallback_rejected and `continue`s
to the next record -- it never raises, so attacker-controlled input
can't churn the retry/DLQ path.
extract_with_claude() wraps the untrusted email in an <email> data
block and neutralizes forged <email>-tag lookalikes in the body with
the same linear-time regex approach as WO's _EMAIL_TAG_RE, and sets
temperature=0 on the Bedrock call.
Deliberate double-count: PO emits ParseMethod=ai_fallback before the
Bedrock call (so a Bedrock-side error still records the outcome), so
a rejected email always produces both an ai_fallback datapoint
(pre-call) and an ai_fallback_rejected datapoint (post-gate). This is
intentional, not a bug -- documented in handler.py, template_parser.py,
and the README.
cdk/po_stack.py: in-place property update to the existing
po-email-processor-template-fallback-rate alarm (same logical ID, no
rename/replacement) -- the fb/(fb+tmpl) expression is left
byte-identical to its pre-Phase-1 form and ai_fallback_rejected is
deliberately excluded from the numerator/denominator/volume floor,
since folding it in as WO does would double-count every rejection
(PO's pre-call emit already counts it once via fb). A net-new
EmailProcessorAiFallbackRejectedAlarm watches the rejected series on
its own, retuned for ~57 emails/day with the 6h/IF-floor/eval-4/
datapoints-2 idiom (not WO's 5-minute sparse idiom, which is
structurally dead at PO volume). Both alarms remain ALARM-only to
site-alerts, NOT_BREACHING, with no element-wise MAX in the math
(post-#102 rule).
* Block "Cancelled" po_status off the AI cancellation route
The AI-fallback gate type-checked po_status but let any string
through, unlike the template path which never emits "Cancelled" on a
new_po. Dispatch routes on email_type, so an AI-path new_po or revision
carrying po_status="Cancelled" would reach save_new_po/save_revision and
cancel a live PO via _merge_update's sticky-cancel write without ever
hitting save_cancellation. Reject the exact sticky marker on any
non-cancellation email_type so the AI path matches the template path's
guard; arbitrary non-marker status strings still pass.
email_type is already validated to the enum before this check, and a
cancellation reaches save_cancellation (which hardcodes the status), so
po_status stays irrelevant on that route.
2026-07-17 14:50:47 -04:00
|
|
|
|
# AI-fallback PO id shape. Same prefix+hyphen+digits family as _PO_ID_RE, but
|
|
|
|
|
|
# HARDENED for the untrusted AI path exactly as WO hardened _WO_ID_RE: [0-9]
|
|
|
|
|
|
# not \d (rejects fullwidth Unicode digits like "2D-18206023" that render
|
|
|
|
|
|
# like ASCII but are a distinct DynamoDB partition key) and \A...\Z not ^...$
|
|
|
|
|
|
# (rejects trailing-newline lookalikes "2D-18206023\n"). The handler builds the
|
|
|
|
|
|
# purchase-orders partition key from po_number (handler _write_fields Key and
|
|
|
|
|
|
# save_cancellation), so an injected "123#x" ('#' not in the class) or bare
|
|
|
|
|
|
# "123" (no prefix-hyphen) must fail here. Distinct from _PO_ID_RE, which the
|
|
|
|
|
|
# template path additionally byte-equals against the subject id -- do NOT touch
|
|
|
|
|
|
# _PO_ID_RE or the template-path validate().
|
|
|
|
|
|
_AI_PO_ID_RE = re.compile(r"\A[A-Z0-9]{1,6}-[0-9]+\Z")
|
|
|
|
|
|
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
# U+2022 bullet delimiting per-line metadata in the Lines section.
|
|
|
|
|
|
_BULLET = "•"
|
|
|
|
|
|
|
|
|
|
|
|
# --- new_po layout patterns (built against the scrubbed real fixture corpus) ---
|
|
|
|
|
|
# Money tokens are read ONLY from three anchored contexts: the Lines-section
|
|
|
|
|
|
# '<desc> for <amt> <CCY>' line, the Total-block standalone amount line, and the
|
|
|
|
|
|
# Items-summary '<qty> <UNIT> x <price>' line. NEVER free money-shaped scanning:
|
|
|
|
|
|
# the Items summary carries unit-price tokens distinct from line amounts.
|
|
|
|
|
|
_MONEY_RE = re.compile(r"\d{1,3}(?:,\d{3})*\.\d{2}")
|
|
|
|
|
|
_CURRENCY_RE = re.compile(r"[A-Z]{3}")
|
|
|
|
|
|
# Items-summary quantity line, e.g. '1.0 EACH x 55,206.00'.
|
|
|
|
|
|
_SUMMARY_ITEM_RE = re.compile(
|
|
|
|
|
|
r"^(?P<qty>\d+(?:\.\d+)?) (?P<unit>[A-Z]+) x (?P<price>\d{1,3}(?:,\d{3})*\.\d{2})$"
|
|
|
|
|
|
)
|
|
|
|
|
|
# Lines-block quantity evidence line, e.g. '1.0 EA' (gate rule V13 cross-check).
|
|
|
|
|
|
_LINE_QTY_RE = re.compile(r"^(?P<qty>\d+(?:\.\d+)?) (?P<unit>[A-Z]+)$")
|
|
|
|
|
|
# Lines-block description/amount line. GREEDY desc: '.+' binds the LAST ' for ',
|
|
|
|
|
|
# so a description containing the word 'for' can never shift the amount.
|
|
|
|
|
|
_LINE_DESC_AMT_RE = re.compile(
|
|
|
|
|
|
r"^(?P<desc>.+) for (?P<amt>\d{1,3}(?:,\d{3})*\.\d{2}) (?P<cur>[A-Z]{3})$"
|
|
|
|
|
|
)
|
|
|
|
|
|
_VIEW_ORDER_URL_RE = re.compile(r"^https://supplier\.coupahost\.com/orders/\S+$")
|
|
|
|
|
|
_ORDER_URL_ID_RE = re.compile(r"^https://supplier\.coupahost\.com/orders/(\d+)\b")
|
|
|
|
|
|
_LOCATION_CODE_RE = re.compile(r"^Location Code: (?P<lc>\d+)$")
|
|
|
|
|
|
_ATTN_RE = re.compile(r"^Attn: (?P<attn>.+)$")
|
|
|
|
|
|
# Ship-to city line immediately preceding 'United States'.
|
|
|
|
|
|
_CITY_STATE_ZIP_RE = re.compile(
|
|
|
|
|
|
r"^(?P<city>.+), (?P<state>[A-Z]{2}) (?P<zip>\d{5}(?:-\d{4})?)$"
|
|
|
|
|
|
)
|
|
|
|
|
|
_US_SENTINEL = "United States"
|
|
|
|
|
|
_SUPPLIER_MARKER = (
|
|
|
|
|
|
"SEA HAVEN" # case-sensitive; drift becomes fallback, never wrong data
|
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
# More Detail block: label lines, each expected exactly once (gate rule V9).
|
|
|
|
|
|
_MORE_DETAIL_FIELDS = {
|
|
|
|
|
|
"Department": "department",
|
|
|
|
|
|
"Status": "po_status",
|
|
|
|
|
|
"Last Opened": "last_opened",
|
|
|
|
|
|
"Order Date": "order_date",
|
|
|
|
|
|
"Acknowledged At": "acknowledged_at",
|
|
|
|
|
|
"Revision Date": "revision_date",
|
|
|
|
|
|
"Payment Term": "payment_terms",
|
|
|
|
|
|
"Req #": "requisition_number",
|
|
|
|
|
|
}
|
|
|
|
|
|
_MORE_DETAIL_LABELS = ("PO ID", *_MORE_DETAIL_FIELDS)
|
|
|
|
|
|
|
|
|
|
|
|
# Per-line bullet metadata: closed label set, assigned purely by leading label
|
|
|
|
|
|
# (longest label first), NEVER by ordinal position -- real data has an optional
|
|
|
|
|
|
# 'Part Number' segment between 'Category' and 'Account', and every run starts
|
|
|
|
|
|
# with a 'Supplier <name>' segment.
|
|
|
|
|
|
_BULLET_LABELS = ("Part Number", "Need By", "Category", "Account", "Period", "Supplier")
|
|
|
|
|
|
_BULLET_FIELDS = {
|
|
|
|
|
|
"Need By": "need_by",
|
|
|
|
|
|
"Category": "category",
|
|
|
|
|
|
"Account": "account_code",
|
|
|
|
|
|
"Period": "period",
|
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
# Frozen USPS state/territory codes (50 states + DC + territories).
|
|
|
|
|
|
_USPS_STATES = frozenset(
|
|
|
|
|
|
"""AL AK AZ AR CA CO CT DE FL GA HI ID IL IN IA KS KY LA ME MD MA MI MN MS
|
|
|
|
|
|
MO MT NE NV NH NJ NM NY NC ND OH OK OR PA RI SC SD TN TX UT VT VA WA WV WI
|
|
|
|
|
|
WY DC PR VI GU AS MP""".split()
|
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
# Sentinel: a label was present but its value did not parse. Must FAIL the gate
|
|
|
|
|
|
# (present-but-unparseable), distinct from an absent value (None).
|
|
|
|
|
|
_UNPARSEABLE = "__UNPARSEABLE__"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
# Skeleton helpers (recursive -- unlike WO's flat contract)
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def _empty_line_item():
|
|
|
|
|
|
return {k: None for k in LINE_ITEM_KEYS}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _empty_candidate():
|
|
|
|
|
|
"""Full nested skeleton: every contract key present, None where absent."""
|
|
|
|
|
|
cand = {k: None for k in CONTRACT_KEYS}
|
|
|
|
|
|
cand["supplier"] = {k: None for k in SUPPLIER_KEYS}
|
|
|
|
|
|
cand["ship_to"] = {k: None for k in SHIP_TO_KEYS}
|
|
|
|
|
|
cand["line_items"] = [_empty_line_item()]
|
|
|
|
|
|
cand["source_system"] = SOURCE_SYSTEM
|
|
|
|
|
|
return cand
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _normalize(candidate):
|
|
|
|
|
|
"""Guarantee exact nested key presence before returning."""
|
|
|
|
|
|
out = _empty_candidate()
|
|
|
|
|
|
for k in CONTRACT_KEYS:
|
|
|
|
|
|
if k in candidate and k not in ("supplier", "ship_to", "line_items"):
|
|
|
|
|
|
out[k] = candidate[k]
|
|
|
|
|
|
supplier = candidate.get("supplier") or {}
|
|
|
|
|
|
out["supplier"] = {k: supplier.get(k) for k in SUPPLIER_KEYS}
|
|
|
|
|
|
ship_to = candidate.get("ship_to") or {}
|
|
|
|
|
|
out["ship_to"] = {k: ship_to.get(k) for k in SHIP_TO_KEYS}
|
|
|
|
|
|
items = candidate.get("line_items") or [{}]
|
|
|
|
|
|
out["line_items"] = [{k: (it or {}).get(k) for k in LINE_ITEM_KEYS} for it in items]
|
|
|
|
|
|
return out
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _clean(value):
|
|
|
|
|
|
"""Strip trailing CR and U+00A0 nbsp that every captured Coupa value carries;
|
|
|
|
|
|
collapse nothing else. Returns None for empty/placeholder 'None'."""
|
|
|
|
|
|
if value is None:
|
|
|
|
|
|
return None
|
|
|
|
|
|
v = value.replace("\r", "").replace(" ", " ").strip()
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
if v in ("", "None"):
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
return None
|
|
|
|
|
|
return v
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _plain_lines(body):
|
|
|
|
|
|
return body.replace("\r\n", "\n").replace("\r", "\n").split("\n")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _to_decimal(raw):
|
|
|
|
|
|
"""Parse a Coupa money token ('18,624.05') to Decimal, stripping thousands
|
|
|
|
|
|
separators. Returns _UNPARSEABLE if it does not parse (gate must reject)."""
|
|
|
|
|
|
if raw is None:
|
|
|
|
|
|
return None
|
|
|
|
|
|
token = raw.replace(",", "").strip()
|
|
|
|
|
|
try:
|
|
|
|
|
|
return Decimal(token)
|
|
|
|
|
|
except (InvalidOperation, ValueError):
|
|
|
|
|
|
return _UNPARSEABLE
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
# Line navigation helpers (shared by the extractor and the gate; both operate
|
|
|
|
|
|
# on email_data["body"] only -- the gate re-derives its own byte evidence and
|
|
|
|
|
|
# never trusts extractor-carried state it can re-derive)
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def _visible(line):
|
|
|
|
|
|
"""True when the raw line carries visible content. The literal placeholder
|
|
|
|
|
|
'None' IS visible (unlike _clean, which maps it to None), so label values
|
|
|
|
|
|
of 'None' are found -- not skipped over into the next label line."""
|
|
|
|
|
|
return bool(line.replace("\r", "").replace("\xa0", "").strip())
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _indices(lines, label):
|
|
|
|
|
|
"""All indices whose _clean-ed content equals the label exactly."""
|
|
|
|
|
|
return [i for i, ln in enumerate(lines) if _clean(ln) == label]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _find_after(lines, label, start):
|
|
|
|
|
|
"""First index >= start whose _clean-ed content equals label, or None."""
|
|
|
|
|
|
for i in range(start, len(lines)):
|
|
|
|
|
|
if _clean(lines[i]) == label:
|
|
|
|
|
|
return i
|
|
|
|
|
|
return None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _next_visible(lines, idx, end=None):
|
|
|
|
|
|
"""Index of the first visible line strictly after idx (before end), or None."""
|
|
|
|
|
|
stop = len(lines) if end is None else min(end, len(lines))
|
|
|
|
|
|
for j in range(idx + 1, stop):
|
|
|
|
|
|
if _visible(lines[j]):
|
|
|
|
|
|
return j
|
|
|
|
|
|
return None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _split_bullet_segments(raw_line):
|
|
|
|
|
|
"""Split a Lines-section metadata line on the bare U+2022 bullet; strip each
|
|
|
|
|
|
segment of spaces and nbsp; drop empty segments."""
|
|
|
|
|
|
segments = []
|
|
|
|
|
|
for seg in raw_line.replace("\r", "").split(_BULLET):
|
|
|
|
|
|
seg = seg.replace("\xa0", " ").strip()
|
|
|
|
|
|
if seg:
|
|
|
|
|
|
segments.append(seg)
|
|
|
|
|
|
return segments
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _match_bullet_label(segment):
|
|
|
|
|
|
"""(label, value) by leading-label prefix match (longest label first)
|
|
|
|
|
|
against the closed _BULLET_LABELS set, or (None, None) if unrecognized."""
|
|
|
|
|
|
for label in sorted(_BULLET_LABELS, key=len, reverse=True):
|
|
|
|
|
|
if segment == label:
|
|
|
|
|
|
return label, None
|
|
|
|
|
|
if segment.startswith(label + " "):
|
|
|
|
|
|
return label, segment[len(label) :].strip()
|
|
|
|
|
|
return None, None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _walk_leaves(obj):
|
|
|
|
|
|
"""Yield every scalar leaf of a nested dict/list candidate."""
|
|
|
|
|
|
if isinstance(obj, dict):
|
|
|
|
|
|
for v in obj.values():
|
|
|
|
|
|
yield from _walk_leaves(v)
|
|
|
|
|
|
elif isinstance(obj, list):
|
|
|
|
|
|
for v in obj:
|
|
|
|
|
|
yield from _walk_leaves(v)
|
|
|
|
|
|
else:
|
|
|
|
|
|
yield obj
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
# Subject helpers
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def classify_template(email_data):
|
|
|
|
|
|
"""Return (template_id, reason). template_id in
|
|
|
|
|
|
{coupa_new_po, coupa_cancellation, unknown}."""
|
|
|
|
|
|
subject = _clean(email_data.get("subject")) or ""
|
|
|
|
|
|
if _NEW_PO_SUBJECT.match(subject):
|
|
|
|
|
|
return "coupa_new_po", "ok"
|
|
|
|
|
|
if _CANCELLATION_SUBJECT.match(subject):
|
|
|
|
|
|
return "coupa_cancellation", "ok"
|
|
|
|
|
|
return "unknown", "subject_no_match"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _subject_po_id(email_data):
|
|
|
|
|
|
"""PO number parsed from the subject, or None."""
|
|
|
|
|
|
subject = _clean(email_data.get("subject")) or ""
|
|
|
|
|
|
m = _NEW_PO_SUBJECT.match(subject)
|
|
|
|
|
|
if m:
|
|
|
|
|
|
return m.group("po")
|
|
|
|
|
|
m = _CANCELLATION_SUBJECT.match(subject)
|
|
|
|
|
|
if m:
|
|
|
|
|
|
return m.group("po")
|
|
|
|
|
|
return None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
# T2: coupa_cancellation -> cancellation (simple + safe: po_number only)
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def extract_cancellation(email_data):
|
|
|
|
|
|
"""Cancellation carries no PO detail we trust beyond the id; the handler's
|
|
|
|
|
|
save_cancellation() only needs po_number. Everything else stays None."""
|
|
|
|
|
|
candidate = _empty_candidate()
|
|
|
|
|
|
candidate["email_type"] = "cancellation"
|
|
|
|
|
|
candidate["po_number"] = _subject_po_id(email_data)
|
|
|
|
|
|
return _normalize(candidate)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
# T1: coupa_new_po -> new_po
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def _assign_bullet_metadata(item, raw_line):
|
|
|
|
|
|
"""Assign the U+2022 metadata segments to the item BY LEADING LABEL.
|
|
|
|
|
|
|
|
|
|
|
|
'Supplier' is recognized but not stored on the item (gate rule V5 proves it
|
|
|
|
|
|
byte-equals supplier.name from the body); 'Part Number' is recognized but
|
|
|
|
|
|
discarded (LINE_ITEM_KEYS has no slot -- inventing one would break the
|
|
|
|
|
|
key_set_mismatch rule and LLM-path shape parity). Unrecognized or duplicate
|
|
|
|
|
|
segments are the gate's job to reject (rule V8)."""
|
|
|
|
|
|
for segment in _split_bullet_segments(raw_line):
|
|
|
|
|
|
label, value = _match_bullet_label(segment)
|
|
|
|
|
|
field = _BULLET_FIELDS.get(label)
|
|
|
|
|
|
if field and item[field] is None:
|
|
|
|
|
|
item[field] = _clean(value)
|
|
|
|
|
|
|
|
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
def _extract_summary_section(candidate, lines):
|
|
|
|
|
|
"""Summary block (start .. 'More Detail'): submitted_by, on_behalf_of, the
|
|
|
|
|
|
FIRST 'Supplier' name, the unique view_order_url, and the Items-summary
|
|
|
|
|
|
quantity lines. Returns (more_detail_index_or_None, summary_item_matches)."""
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
more_detail = _find_after(lines, "More Detail", 0)
|
|
|
|
|
|
summary_end = more_detail if more_detail is not None else len(lines)
|
|
|
|
|
|
for label, field in (
|
|
|
|
|
|
("Submitted By", "submitted_by"),
|
|
|
|
|
|
("On Behalf Of", "on_behalf_of"),
|
|
|
|
|
|
):
|
|
|
|
|
|
idx = _find_after(lines, label, 0)
|
|
|
|
|
|
if idx is not None and idx < summary_end:
|
|
|
|
|
|
j = _next_visible(lines, idx, summary_end)
|
|
|
|
|
|
if j is not None:
|
|
|
|
|
|
candidate[field] = _clean(lines[j])
|
|
|
|
|
|
sup1 = _find_after(lines, "Supplier", 0)
|
|
|
|
|
|
if sup1 is not None and sup1 < summary_end:
|
|
|
|
|
|
j = _next_visible(lines, sup1, summary_end)
|
|
|
|
|
|
if j is not None:
|
|
|
|
|
|
candidate["supplier"]["name"] = _clean(lines[j])
|
|
|
|
|
|
# view_order_url: the unique orders link in the summary (0 or >1 -> None).
|
|
|
|
|
|
url_lines = [
|
|
|
|
|
|
_clean(lines[i])
|
|
|
|
|
|
for i in range(summary_end)
|
|
|
|
|
|
if _VIEW_ORDER_URL_RE.match(_clean(lines[i]) or "")
|
|
|
|
|
|
]
|
|
|
|
|
|
if len(url_lines) == 1:
|
|
|
|
|
|
candidate["view_order_url"] = url_lines[0]
|
|
|
|
|
|
# Items-summary quantity lines ('1.0 EACH x 55,206.00'), collected in order.
|
|
|
|
|
|
summary_items = [
|
|
|
|
|
|
m
|
|
|
|
|
|
for i in range(summary_end)
|
|
|
|
|
|
if (m := _SUMMARY_ITEM_RE.fullmatch(_clean(lines[i]) or ""))
|
|
|
|
|
|
]
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return more_detail, summary_items
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
|
|
|
|
|
|
def _extract_more_detail_block(candidate, lines, more_detail):
|
|
|
|
|
|
"""More Detail block ('More Detail' .. second 'Supplier'): the labeled
|
|
|
|
|
|
header fields. Returns the second 'Supplier' index (or None).
|
|
|
|
|
|
|
|
|
|
|
|
The FIRST 'Shipping' lives in this block; its value must be the literal
|
|
|
|
|
|
'None' placeholder (gate rule V4 tripwire) and is never used for ship_to."""
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
sup2 = None
|
|
|
|
|
|
if more_detail is not None:
|
|
|
|
|
|
sup2 = _find_after(lines, "Supplier", more_detail + 1)
|
|
|
|
|
|
md_end = sup2 if sup2 is not None else len(lines)
|
|
|
|
|
|
for label, field in _MORE_DETAIL_FIELDS.items():
|
|
|
|
|
|
idx = _find_after(lines, label, more_detail + 1)
|
|
|
|
|
|
if idx is not None and idx < md_end:
|
|
|
|
|
|
j = _next_visible(lines, idx, md_end)
|
|
|
|
|
|
if j is not None:
|
|
|
|
|
|
candidate[field] = _clean(lines[j])
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return sup2
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _fill_location_attn(ship_to, lines, start, end):
|
|
|
|
|
|
"""Location Code + Attn lines after the 'United States' sentinel; the first
|
|
|
|
|
|
of each wins (mirrors the extractor's None-guarded assignment)."""
|
|
|
|
|
|
for j in range(start, end):
|
|
|
|
|
|
cl = _clean(lines[j]) or ""
|
|
|
|
|
|
lc = _LOCATION_CODE_RE.fullmatch(cl)
|
|
|
|
|
|
if lc and ship_to["location_code"] is None:
|
|
|
|
|
|
ship_to["location_code"] = lc.group("lc")
|
|
|
|
|
|
attn = _ATTN_RE.fullmatch(cl)
|
|
|
|
|
|
if attn and ship_to["attn"] is None:
|
|
|
|
|
|
ship_to["attn"] = _clean(attn.group("attn"))
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _fill_ship_to_address(ship_to, lines, ship2, st_end):
|
|
|
|
|
|
"""Populate ship_to from the second 'Shipping' block, sentinel-anchored on
|
|
|
|
|
|
the 'United States' line that terminates the address."""
|
|
|
|
|
|
name_idx = _next_visible(lines, ship2, st_end)
|
|
|
|
|
|
if name_idx is None:
|
|
|
|
|
|
return
|
|
|
|
|
|
ship_to["name"] = _clean(lines[name_idx])
|
|
|
|
|
|
us_idx = _find_after(lines, _US_SENTINEL, name_idx + 1)
|
|
|
|
|
|
if us_idx is None or us_idx >= st_end:
|
|
|
|
|
|
return
|
|
|
|
|
|
city_idx = us_idx - 1
|
|
|
|
|
|
m = None
|
|
|
|
|
|
if city_idx > name_idx:
|
|
|
|
|
|
m = _CITY_STATE_ZIP_RE.fullmatch(_clean(lines[city_idx]) or "")
|
|
|
|
|
|
if m:
|
|
|
|
|
|
ship_to["city"] = m.group("city")
|
|
|
|
|
|
ship_to["state"] = m.group("state")
|
|
|
|
|
|
ship_to["zip"] = m.group("zip")
|
|
|
|
|
|
street = [
|
|
|
|
|
|
_clean(lines[j])
|
|
|
|
|
|
for j in range(name_idx + 1, city_idx)
|
|
|
|
|
|
if _visible(lines[j])
|
|
|
|
|
|
]
|
|
|
|
|
|
if street:
|
|
|
|
|
|
ship_to["street"] = "\n".join(street)
|
|
|
|
|
|
ship_to["address"] = "\n".join(
|
|
|
|
|
|
_clean(lines[j]) for j in range(name_idx, us_idx + 1) if _visible(lines[j])
|
|
|
|
|
|
)
|
|
|
|
|
|
_fill_location_attn(ship_to, lines, us_idx + 1, st_end)
|
|
|
|
|
|
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
def _extract_ship_to(candidate, lines, sup2):
|
|
|
|
|
|
"""ship_to block (second 'Shipping' .. 'Lines'). Returns the 'Lines' anchor
|
|
|
|
|
|
index (or None)."""
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
ship2 = _find_after(lines, "Shipping", sup2 + 1) if sup2 is not None else None
|
|
|
|
|
|
lines_anchor = _find_after(lines, "Lines", ship2 + 1) if ship2 is not None else None
|
|
|
|
|
|
st_end = lines_anchor if lines_anchor is not None else len(lines)
|
|
|
|
|
|
if ship2 is not None:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
_fill_ship_to_address(candidate["ship_to"], lines, ship2, st_end)
|
|
|
|
|
|
return lines_anchor
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _parse_line_blocks(lines, start, end):
|
|
|
|
|
|
"""Split the Lines section into U+00A0-delimited blocks and build one line
|
|
|
|
|
|
item per non-empty block. EVERY block is extracted, even when >1, so gate
|
|
|
|
|
|
rule 7 fires with honest multiline_unsupported data (never silently keep
|
|
|
|
|
|
item 0)."""
|
|
|
|
|
|
blocks, block = [], []
|
|
|
|
|
|
for j in range(start, end):
|
|
|
|
|
|
if lines[j].replace("\r", "") == "\xa0":
|
|
|
|
|
|
blocks.append(block)
|
|
|
|
|
|
block = []
|
|
|
|
|
|
else:
|
|
|
|
|
|
block.append(lines[j])
|
|
|
|
|
|
blocks.append(block)
|
|
|
|
|
|
items = []
|
|
|
|
|
|
for block in blocks:
|
|
|
|
|
|
visible = [ln for ln in block if _visible(ln)]
|
|
|
|
|
|
if not visible:
|
|
|
|
|
|
continue
|
|
|
|
|
|
item = _empty_line_item()
|
|
|
|
|
|
for raw in visible:
|
|
|
|
|
|
dm = _LINE_DESC_AMT_RE.fullmatch(_clean(raw) or "")
|
|
|
|
|
|
if dm and item["description"] is None:
|
|
|
|
|
|
item["description"] = _clean(dm.group("desc"))
|
|
|
|
|
|
item["amount"] = _to_decimal(dm.group("amt"))
|
|
|
|
|
|
item["currency"] = dm.group("cur")
|
|
|
|
|
|
elif _BULLET in raw:
|
|
|
|
|
|
_assign_bullet_metadata(item, raw)
|
|
|
|
|
|
# The optional '<qty> EA' evidence line is not stored: quantity/
|
|
|
|
|
|
# unit/price come from the Items summary; gate rule V13 cross-checks
|
|
|
|
|
|
# the EA line against it from the body.
|
|
|
|
|
|
items.append(item)
|
|
|
|
|
|
return items
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _extract_line_items(candidate, lines, lines_anchor, summary_items):
|
|
|
|
|
|
"""Lines section ('Lines' .. second 'Total'). Populates line_items,
|
|
|
|
|
|
coupa_category, and (single-line only) quantity/unit/price from the Items
|
|
|
|
|
|
summary. Returns the second 'Total' index (or None)."""
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
total2 = (
|
|
|
|
|
|
_find_after(lines, "Total", lines_anchor + 1)
|
|
|
|
|
|
if lines_anchor is not None
|
|
|
|
|
|
else None
|
|
|
|
|
|
)
|
|
|
|
|
|
items = []
|
|
|
|
|
|
if lines_anchor is not None:
|
|
|
|
|
|
end = total2 if total2 is not None else len(lines)
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
items = _parse_line_blocks(lines, lines_anchor + 1, end)
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
if items:
|
|
|
|
|
|
candidate["line_items"] = items
|
|
|
|
|
|
# coupa_category is a VERBATIM copy of item 0's Category bullet value.
|
|
|
|
|
|
candidate["coupa_category"] = items[0]["category"]
|
|
|
|
|
|
if len(summary_items) == 1 and len(items) == 1:
|
|
|
|
|
|
m = summary_items[0]
|
|
|
|
|
|
items[0]["quantity"] = _to_decimal(m.group("qty"))
|
|
|
|
|
|
items[0]["unit"] = m.group("unit")
|
|
|
|
|
|
items[0]["price"] = _to_decimal(m.group("price"))
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return total2
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _extract_total_block(candidate, lines, total2):
|
|
|
|
|
|
"""Total block (second 'Total' .. end): total_amount + currency."""
|
|
|
|
|
|
if total2 is None:
|
|
|
|
|
|
return
|
|
|
|
|
|
amt_idx = _next_visible(lines, total2)
|
|
|
|
|
|
if amt_idx is None:
|
|
|
|
|
|
return
|
|
|
|
|
|
token = _clean(lines[amt_idx]) or ""
|
|
|
|
|
|
if _MONEY_RE.fullmatch(token):
|
|
|
|
|
|
candidate["total_amount"] = _to_decimal(token)
|
|
|
|
|
|
cur_idx = _next_visible(lines, amt_idx)
|
|
|
|
|
|
if cur_idx is not None:
|
|
|
|
|
|
cur_token = _clean(lines[cur_idx]) or ""
|
|
|
|
|
|
if _CURRENCY_RE.fullmatch(cur_token):
|
|
|
|
|
|
candidate["currency"] = cur_token
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def extract_new_po(email_data):
|
|
|
|
|
|
"""Extract the new_po contract from the text/plain body.
|
|
|
|
|
|
|
|
|
|
|
|
Permissive capture, section-windowed: each anchor line is located by exact
|
|
|
|
|
|
_clean-ed full-line equality, STRICTLY AFTER the previous anchor. Missing
|
|
|
|
|
|
anchors leave fields None -- the gate then fails closed. Representation-
|
|
|
|
|
|
agnostic: matches only on _clean-ed lines, never on '\\r'-suffixed literals
|
|
|
|
|
|
(body line endings are decode-path dependent).
|
|
|
|
|
|
|
|
|
|
|
|
Duplicate labels: 'Supplier', 'Shipping', 'Total' each appear TWICE (summary
|
|
|
|
|
|
placeholder + detail block; the first 'Shipping' value is literally 'None').
|
|
|
|
|
|
supplier.name anchors on the FIRST 'Supplier'; ship_to on the SECOND
|
|
|
|
|
|
'Shipping'; total on the SECOND 'Total'.
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
Delegated section by section to the _extract_* helpers, threading the
|
|
|
|
|
|
anchor indices each stage discovers into the next; DERIVED_KEYS (site_code,
|
|
|
|
|
|
trade, fiscal_year) stay None -- filled later by the shared post-stage
|
|
|
|
|
|
identically on both paths.
|
|
|
|
|
|
"""
|
|
|
|
|
|
candidate = _empty_candidate()
|
|
|
|
|
|
candidate["email_type"] = "new_po"
|
|
|
|
|
|
candidate["po_number"] = _subject_po_id(email_data)
|
|
|
|
|
|
|
|
|
|
|
|
lines = _plain_lines(email_data["body"])
|
|
|
|
|
|
|
|
|
|
|
|
more_detail, summary_items = _extract_summary_section(candidate, lines)
|
|
|
|
|
|
sup2 = _extract_more_detail_block(candidate, lines, more_detail)
|
|
|
|
|
|
lines_anchor = _extract_ship_to(candidate, lines, sup2)
|
|
|
|
|
|
total2 = _extract_line_items(candidate, lines, lines_anchor, summary_items)
|
|
|
|
|
|
_extract_total_block(candidate, lines, total2)
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
|
|
|
|
|
|
# DERIVED_KEYS intentionally left None (shared post-stage fills them).
|
|
|
|
|
|
return _normalize(candidate)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
# Validation gate -- FAIL CLOSED
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def _structural_keys_ok(candidate):
|
|
|
|
|
|
if set(candidate.keys()) != set(CONTRACT_KEYS):
|
|
|
|
|
|
return False
|
|
|
|
|
|
if set((candidate.get("supplier") or {}).keys()) != set(SUPPLIER_KEYS):
|
|
|
|
|
|
return False
|
|
|
|
|
|
if set((candidate.get("ship_to") or {}).keys()) != set(SHIP_TO_KEYS):
|
|
|
|
|
|
return False
|
|
|
|
|
|
items = candidate.get("line_items")
|
|
|
|
|
|
if not isinstance(items, list) or not items:
|
|
|
|
|
|
return False
|
|
|
|
|
|
return all(set((it or {}).keys()) == set(LINE_ITEM_KEYS) for it in items)
|
|
|
|
|
|
|
|
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
def validate(candidate, template_id, email_data): # noqa: C901, PLR0911, PLR0912
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
"""Return (True, 'ok') only if provably conformant; else (False, reason).
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
Every rule must hold. See failClosedGateRules in the investigation report.
|
|
|
|
|
|
|
|
|
|
|
|
A linear fail-closed rule ladder with one return per reason code -- the
|
|
|
|
|
|
branch count is the rule count. Splitting it further adds indirection, not
|
|
|
|
|
|
clarity, so the complexity/return/branch ceilings are suppressed here (the
|
|
|
|
|
|
value-level rules V1-V13 ARE decomposed, in _validate_new_po_values)."""
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
# (1) known template
|
|
|
|
|
|
if template_id not in _TEMPLATE_EMAIL_TYPE:
|
|
|
|
|
|
return False, "subject_no_match"
|
|
|
|
|
|
|
|
|
|
|
|
# (2) exact nested key-set
|
|
|
|
|
|
if not _structural_keys_ok(candidate):
|
|
|
|
|
|
return False, "key_set_mismatch"
|
|
|
|
|
|
|
|
|
|
|
|
# (3) derived fields must NOT be populated by the parser
|
|
|
|
|
|
for k in DERIVED_KEYS:
|
|
|
|
|
|
if candidate.get(k) is not None:
|
|
|
|
|
|
return False, "derived_field_set"
|
|
|
|
|
|
|
|
|
|
|
|
# (4) email_type matches the template's expected type
|
|
|
|
|
|
expected_type = _TEMPLATE_EMAIL_TYPE[template_id]
|
|
|
|
|
|
et = candidate.get("email_type")
|
|
|
|
|
|
if et not in VALID_EMAIL_TYPES:
|
|
|
|
|
|
return False, "missing_required_field"
|
|
|
|
|
|
if et != expected_type:
|
|
|
|
|
|
return False, "email_type_mismatch"
|
|
|
|
|
|
|
|
|
|
|
|
# (5) po_number: valid shape AND byte-equals the subject id
|
|
|
|
|
|
subject_po = _subject_po_id(email_data)
|
|
|
|
|
|
po = candidate.get("po_number")
|
|
|
|
|
|
if not po or not _PO_ID_RE.match(str(po)):
|
|
|
|
|
|
return False, "missing_required_field"
|
|
|
|
|
|
if po != subject_po:
|
|
|
|
|
|
return False, "po_id_mismatch"
|
|
|
|
|
|
|
|
|
|
|
|
if template_id == "coupa_cancellation":
|
|
|
|
|
|
# Body corroboration: the subject SiteName prefix is free text, so the
|
|
|
|
|
|
# anchored subject alone cannot distinguish a genuine Coupa cancellation
|
|
|
|
|
|
# from an arbitrary "<prefix> Purchase Order #<po> has been cancelled"
|
|
|
|
|
|
# subject. The real Coupa body independently restates the id in a
|
|
|
|
|
|
# "Purchase Order #<po> ... has been cancelled" notice; require that
|
|
|
|
|
|
# (with the SAME po_number) before marking a PO sticky-Cancelled, so a
|
|
|
|
|
|
# misrouted/near-miss email fails closed to the LLM instead. (SEC review
|
|
|
|
|
|
# of PR #105, F1.)
|
|
|
|
|
|
body = email_data.get("body") or ""
|
|
|
|
|
|
restated = re.search(r"Purchase Order\s+#" + re.escape(str(po)) + r"\b", body)
|
|
|
|
|
|
if not restated or "cancelled" not in body.lower():
|
|
|
|
|
|
return False, "cancellation_body_unconfirmed"
|
|
|
|
|
|
return True, "ok"
|
|
|
|
|
|
|
|
|
|
|
|
# ---- coupa_new_po ----
|
|
|
|
|
|
# (6) status must be one of the confirmed-safe strings.
|
|
|
|
|
|
status = candidate.get("po_status")
|
|
|
|
|
|
if status not in NEW_PO_SAFE_STATUSES:
|
|
|
|
|
|
return False, "unrecognized_status"
|
|
|
|
|
|
|
|
|
|
|
|
# (7) single-line only -- multi-line Lines structure is unobserved.
|
|
|
|
|
|
if len(candidate.get("line_items") or []) != 1:
|
|
|
|
|
|
return False, "multiline_unsupported"
|
|
|
|
|
|
|
|
|
|
|
|
# (8) currency must be exactly USD (non-USD path entirely unexercised).
|
|
|
|
|
|
if candidate.get("currency") != "USD":
|
|
|
|
|
|
return False, "non_usd"
|
|
|
|
|
|
|
|
|
|
|
|
# (V1-V13) value-level rules: every byte proof is RE-DERIVED from
|
|
|
|
|
|
# email_data["body"] -- the gate never trusts extractor-carried state it
|
|
|
|
|
|
# can re-derive, so an extractor bug cannot vouch for itself.
|
|
|
|
|
|
return _validate_new_po_values(candidate, email_data)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _money_border_ok(line_text, serialized):
|
|
|
|
|
|
"""True when `serialized` occurs in the source line with a character before
|
|
|
|
|
|
it that is not a digit or comma (the '18,624.05' -> '624.05' kill switch:
|
|
|
|
|
|
a truncated capture re-serializes as '624.05', but every occurrence of that
|
|
|
|
|
|
string in its source line is preceded by a comma or digit)."""
|
|
|
|
|
|
idx = line_text.find(serialized)
|
|
|
|
|
|
while idx != -1:
|
|
|
|
|
|
prev = line_text[idx - 1] if idx > 0 else ""
|
|
|
|
|
|
if prev not in "0123456789,":
|
|
|
|
|
|
return True
|
|
|
|
|
|
idx = line_text.find(serialized, idx + 1)
|
|
|
|
|
|
return False
|
|
|
|
|
|
|
|
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
@dataclass
|
|
|
|
|
|
class _NewPoAnchorFrame:
|
|
|
|
|
|
"""Anchor frame for the coupa_new_po value gate.
|
|
|
|
|
|
|
|
|
|
|
|
V4 (_build_anchor_frame) proves the duplicate-label anchor layout ONCE and
|
|
|
|
|
|
carries the shared byte evidence every later rule re-derives from the body:
|
|
|
|
|
|
the section indices, plus the Items-summary matches and the captured line
|
|
|
|
|
|
price -- so V13 (_check_qty_unit_price) can re-consume exactly what V1
|
|
|
|
|
|
(_check_money_fidelity) proved, without either rule trusting extractor-
|
|
|
|
|
|
carried state it cannot re-derive from email_data["body"]."""
|
|
|
|
|
|
|
|
|
|
|
|
lines: list
|
|
|
|
|
|
more_detail: int
|
|
|
|
|
|
supplier_idxs: list
|
|
|
|
|
|
shipping_idxs: list
|
|
|
|
|
|
total_idxs: list
|
|
|
|
|
|
lines_anchor: int
|
|
|
|
|
|
lines_end: int
|
|
|
|
|
|
summary_matches: list
|
|
|
|
|
|
price: object
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _anchor_order_ok(
|
|
|
|
|
|
more_detail, supplier_idxs, shipping_idxs, total_idxs, lines_anchor
|
|
|
|
|
|
):
|
|
|
|
|
|
"""Section ordering: summary Supplier < More Detail < detail Supplier <
|
|
|
|
|
|
second Shipping < Lines < second Total; summary Total before More Detail."""
|
|
|
|
|
|
return (
|
|
|
|
|
|
supplier_idxs[0]
|
|
|
|
|
|
< more_detail
|
|
|
|
|
|
< supplier_idxs[1]
|
|
|
|
|
|
< shipping_idxs[1]
|
|
|
|
|
|
< lines_anchor
|
|
|
|
|
|
< total_idxs[1]
|
|
|
|
|
|
) and (total_idxs[0] < more_detail < shipping_idxs[0] < supplier_idxs[1])
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
|
|
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
def _first_shipping_value(lines, first_shipping_idx):
|
|
|
|
|
|
"""Value line under the FIRST 'Shipping' (the dup-label-swap tripwire: a
|
|
|
|
|
|
real address here means the layout drifted and ship_to was read from the
|
|
|
|
|
|
wrong block; the genuine layout carries the literal 'None')."""
|
|
|
|
|
|
if first_shipping_idx + 1 >= len(lines):
|
|
|
|
|
|
return ""
|
|
|
|
|
|
return lines[first_shipping_idx + 1].replace("\r", "").replace("\xa0", "").strip()
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _build_anchor_frame(lines, items):
|
|
|
|
|
|
"""V4 anchor integrity. Returns (frame, None) when the duplicate-label
|
|
|
|
|
|
layout is proven, else (None, 'anchor_violation'). Also precomputes the
|
|
|
|
|
|
Items-summary matches and the captured line price onto the frame."""
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
more_detail_idxs = _indices(lines, "More Detail")
|
|
|
|
|
|
supplier_idxs = _indices(lines, "Supplier")
|
|
|
|
|
|
shipping_idxs = _indices(lines, "Shipping")
|
|
|
|
|
|
total_idxs = _indices(lines, "Total")
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
# The Coupa layout repeats each of Supplier/Shipping/Total exactly twice
|
|
|
|
|
|
# (summary placeholder + detail block) -- 2 is a structural constant of the
|
|
|
|
|
|
# template, not a tunable magic number.
|
|
|
|
|
|
if (
|
|
|
|
|
|
len(more_detail_idxs) != 1
|
|
|
|
|
|
or len(supplier_idxs) != 2 # noqa: PLR2004
|
|
|
|
|
|
or len(shipping_idxs) != 2 # noqa: PLR2004
|
|
|
|
|
|
or len(total_idxs) != 2 # noqa: PLR2004
|
|
|
|
|
|
):
|
|
|
|
|
|
return None, "anchor_violation"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
more_detail = more_detail_idxs[0]
|
|
|
|
|
|
lines_anchor = _find_after(lines, "Lines", shipping_idxs[1] + 1)
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
# lines_anchor is proven non-None before _anchor_order_ok consumes it (the
|
|
|
|
|
|
# `or` short-circuits); every index below is safe once the counts hold.
|
|
|
|
|
|
if (
|
|
|
|
|
|
lines_anchor is None
|
|
|
|
|
|
or _first_shipping_value(lines, shipping_idxs[0]) != "None"
|
|
|
|
|
|
or not _anchor_order_ok(
|
|
|
|
|
|
more_detail, supplier_idxs, shipping_idxs, total_idxs, lines_anchor
|
|
|
|
|
|
)
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
):
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return None, "anchor_violation"
|
|
|
|
|
|
summary_matches = [
|
|
|
|
|
|
m
|
|
|
|
|
|
for i in range(more_detail)
|
|
|
|
|
|
if (m := _SUMMARY_ITEM_RE.fullmatch(_clean(lines[i]) or ""))
|
|
|
|
|
|
]
|
|
|
|
|
|
frame = _NewPoAnchorFrame(
|
|
|
|
|
|
lines=lines,
|
|
|
|
|
|
more_detail=more_detail,
|
|
|
|
|
|
supplier_idxs=supplier_idxs,
|
|
|
|
|
|
shipping_idxs=shipping_idxs,
|
|
|
|
|
|
total_idxs=total_idxs,
|
|
|
|
|
|
lines_anchor=lines_anchor,
|
|
|
|
|
|
lines_end=total_idxs[1],
|
|
|
|
|
|
summary_matches=summary_matches,
|
|
|
|
|
|
price=items[0].get("price") if items else None,
|
|
|
|
|
|
)
|
|
|
|
|
|
return frame, None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _check_line_item_money(candidate, frame):
|
|
|
|
|
|
"""V1 line-level money fidelity -> amount_mismatch / non_usd.
|
|
|
|
|
|
|
|
|
|
|
|
Every captured line amount must re-locate its raw source token in the body:
|
|
|
|
|
|
the token fullmatches the grouped money shape, format(value, ',.2f') byte-
|
|
|
|
|
|
equals it, and the character before it is not a digit/comma. Line-level
|
|
|
|
|
|
currency is pinned too -- a single non-USD line item fails closed even when
|
|
|
|
|
|
the Total block reads USD (the non-USD path is entirely unexercised)."""
|
|
|
|
|
|
lines = frame.lines
|
|
|
|
|
|
items = candidate.get("line_items") or []
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
for_matches = []
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
for j in range(frame.lines_anchor + 1, frame.lines_end):
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
m = _LINE_DESC_AMT_RE.fullmatch(_clean(lines[j]) or "")
|
|
|
|
|
|
if m:
|
|
|
|
|
|
for_matches.append(m)
|
|
|
|
|
|
if len(for_matches) != len(items):
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "amount_mismatch"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
for item, m in zip(items, for_matches):
|
|
|
|
|
|
amount = item.get("amount")
|
|
|
|
|
|
if not isinstance(amount, Decimal):
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "amount_mismatch"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
serialized = format(amount, ",.2f")
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
if serialized != m.group("amt") or not _money_border_ok(m.string, serialized):
|
|
|
|
|
|
return "amount_mismatch"
|
|
|
|
|
|
currency = candidate.get("currency")
|
|
|
|
|
|
if item.get("currency") != currency or m.group("cur") != currency:
|
|
|
|
|
|
return "non_usd"
|
|
|
|
|
|
return None
|
|
|
|
|
|
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
def _check_summary_price(frame):
|
|
|
|
|
|
"""V1 summary-price fidelity -> amount_mismatch. The captured Items-summary
|
|
|
|
|
|
price must re-locate its raw token exactly as the line amounts do."""
|
|
|
|
|
|
price = frame.price
|
|
|
|
|
|
if price is None:
|
|
|
|
|
|
return None
|
|
|
|
|
|
if not isinstance(price, Decimal) or len(frame.summary_matches) != 1:
|
|
|
|
|
|
return "amount_mismatch"
|
|
|
|
|
|
serialized = format(price, ",.2f")
|
|
|
|
|
|
if serialized != frame.summary_matches[0].group("price"):
|
|
|
|
|
|
return "amount_mismatch"
|
|
|
|
|
|
if not _money_border_ok(frame.summary_matches[0].string, serialized):
|
|
|
|
|
|
return "amount_mismatch"
|
|
|
|
|
|
return None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _check_money_fidelity(candidate, frame):
|
|
|
|
|
|
"""V1 money fidelity: line-item amounts then the Items-summary price."""
|
|
|
|
|
|
return _check_line_item_money(candidate, frame) or _check_summary_price(frame)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _read_total_tokens(lines, total_idxs):
|
|
|
|
|
|
"""Read the (amount, currency) token pair under each 'Total' anchor.
|
|
|
|
|
|
Returns (tokens, True) or (None, False) on any missing/malformed token."""
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
total_tokens = []
|
|
|
|
|
|
for t_idx in total_idxs:
|
|
|
|
|
|
a_idx = _next_visible(lines, t_idx)
|
|
|
|
|
|
if a_idx is None:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return None, False
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
token = _clean(lines[a_idx]) or ""
|
|
|
|
|
|
if not _MONEY_RE.fullmatch(token):
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return None, False
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
c_idx = _next_visible(lines, a_idx)
|
|
|
|
|
|
cur_token = (_clean(lines[c_idx]) or "") if c_idx is not None else ""
|
|
|
|
|
|
if not _CURRENCY_RE.fullmatch(cur_token):
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return None, False
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
total_tokens.append((token, cur_token))
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return total_tokens, True
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _check_total_proof(candidate, frame):
|
|
|
|
|
|
"""V3 dual-Total proof + V2 sum proof -> amount_mismatch. Both Total blocks
|
|
|
|
|
|
must carry byte-identical tokens that byte-equal the captured total, and the
|
|
|
|
|
|
line amounts must sum to it exactly (no qty*price rule -- partial qtys)."""
|
|
|
|
|
|
items = candidate.get("line_items") or []
|
|
|
|
|
|
total = candidate.get("total_amount")
|
|
|
|
|
|
if not isinstance(total, Decimal):
|
|
|
|
|
|
return "amount_mismatch"
|
|
|
|
|
|
total_tokens, ok = _read_total_tokens(frame.lines, frame.total_idxs)
|
|
|
|
|
|
if (
|
|
|
|
|
|
not ok
|
|
|
|
|
|
or total_tokens[0] != total_tokens[1]
|
|
|
|
|
|
or format(total, ",.2f") != total_tokens[1][0]
|
|
|
|
|
|
or candidate.get("currency") != total_tokens[1][1]
|
|
|
|
|
|
or sum(item["amount"] for item in items) != total
|
|
|
|
|
|
):
|
|
|
|
|
|
return "amount_mismatch"
|
|
|
|
|
|
return None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _check_supplier_proof(candidate, frame):
|
|
|
|
|
|
"""V5 supplier proof -> anchor_violation. The supplier name must carry the
|
|
|
|
|
|
SEA HAVEN marker and byte-equal its restatements in both the summary and
|
|
|
|
|
|
the detail block, and must not collide with the ship_to name."""
|
|
|
|
|
|
lines = frame.lines
|
|
|
|
|
|
ship_to = candidate.get("ship_to") or {}
|
|
|
|
|
|
supplier_name = (candidate.get("supplier") or {}).get("name")
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
if not supplier_name or _SUPPLIER_MARKER not in supplier_name:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "anchor_violation"
|
|
|
|
|
|
det_idx = _next_visible(lines, frame.supplier_idxs[1], frame.shipping_idxs[1])
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
if det_idx is None or _clean(lines[det_idx]) != supplier_name:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "anchor_violation"
|
|
|
|
|
|
sum_idx = _next_visible(lines, frame.supplier_idxs[0], frame.more_detail)
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
if sum_idx is None or _clean(lines[sum_idx]) != supplier_name:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "anchor_violation"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
if ship_to.get("name") == supplier_name:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "anchor_violation"
|
|
|
|
|
|
return None
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
|
|
|
|
|
|
def _check_bullet_line(raw, supplier_name):
|
|
|
|
|
|
"""One Lines-section bullet metadata line: closed label set, no dupes, the
|
|
|
|
|
|
required labels present, and the per-item Supplier segment (V5) byte-equal
|
|
|
|
|
|
to supplier_name."""
|
|
|
|
|
|
seen = {}
|
|
|
|
|
|
for segment in _split_bullet_segments(raw):
|
|
|
|
|
|
label, value = _match_bullet_label(segment)
|
|
|
|
|
|
if label is None or label in seen:
|
|
|
|
|
|
return "bullet_label_unrecognized"
|
|
|
|
|
|
seen[label] = value
|
|
|
|
|
|
if set(seen) - {
|
|
|
|
|
|
"Supplier",
|
|
|
|
|
|
"Need By",
|
|
|
|
|
|
"Category",
|
|
|
|
|
|
"Account",
|
|
|
|
|
|
"Period",
|
|
|
|
|
|
"Part Number",
|
|
|
|
|
|
}:
|
|
|
|
|
|
return "bullet_label_unrecognized"
|
|
|
|
|
|
if not {"Supplier", "Need By", "Category", "Account", "Period"} <= set(seen):
|
|
|
|
|
|
return "bullet_label_unrecognized"
|
|
|
|
|
|
if _clean(seen["Supplier"]) != supplier_name:
|
|
|
|
|
|
return "anchor_violation"
|
|
|
|
|
|
return None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _check_bullet_discipline(candidate, frame):
|
|
|
|
|
|
"""V8 bullet discipline -> bullet_label_unrecognized (V5's per-item
|
|
|
|
|
|
Supplier-segment proof rides the same walk)."""
|
|
|
|
|
|
lines = frame.lines
|
|
|
|
|
|
items = candidate.get("line_items") or []
|
|
|
|
|
|
supplier_name = (candidate.get("supplier") or {}).get("name")
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
bullet_lines = [
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
lines[j]
|
|
|
|
|
|
for j in range(frame.lines_anchor + 1, frame.lines_end)
|
|
|
|
|
|
if _BULLET in lines[j]
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
]
|
|
|
|
|
|
if len(bullet_lines) != len(items):
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "bullet_label_unrecognized"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
for raw in bullet_lines:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
reason = _check_bullet_line(raw, supplier_name)
|
|
|
|
|
|
if reason:
|
|
|
|
|
|
return reason
|
|
|
|
|
|
return None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _check_ship_to_required(candidate, frame):
|
|
|
|
|
|
"""V6 ship_to required fields -> missing_required_field. Required fields
|
|
|
|
|
|
present, numeric location_code that re-derives from the body, and attn (if
|
|
|
|
|
|
present) matching a body Attn line."""
|
|
|
|
|
|
lines = frame.lines
|
|
|
|
|
|
ship_to = candidate.get("ship_to") or {}
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
for field in ("name", "street", "city", "state", "zip", "location_code"):
|
|
|
|
|
|
if not ship_to.get(field):
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "missing_required_field"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
if not re.fullmatch(r"\d+", ship_to["location_code"]):
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "missing_required_field"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
lc_values = [
|
|
|
|
|
|
m.group("lc")
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
for j in range(frame.shipping_idxs[1] + 1, frame.lines_anchor)
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
if (m := _LOCATION_CODE_RE.fullmatch(_clean(lines[j]) or ""))
|
|
|
|
|
|
]
|
|
|
|
|
|
if ship_to["location_code"] not in lc_values:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "missing_required_field"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
attn_values = [
|
|
|
|
|
|
_clean(m.group("attn"))
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
for j in range(frame.shipping_idxs[1] + 1, frame.lines_anchor)
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
if (m := _ATTN_RE.fullmatch(_clean(lines[j]) or ""))
|
|
|
|
|
|
]
|
|
|
|
|
|
if attn_values:
|
|
|
|
|
|
if ship_to.get("attn") not in attn_values:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "missing_required_field"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
elif ship_to.get("attn") is not None:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "missing_required_field"
|
|
|
|
|
|
return None
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
|
|
|
|
|
|
def _check_address_shape(candidate, frame):
|
|
|
|
|
|
"""V7 address shape -> address_shape_invalid. Gated on the RAW pre-
|
|
|
|
|
|
enrichment zip: validate() runs BEFORE enrich_parsed, so a short zip fails
|
|
|
|
|
|
closed to the LLM path where pad_zip repairs it (both paths then get
|
|
|
|
|
|
identical pad_zip treatment downstream)."""
|
|
|
|
|
|
lines = frame.lines
|
|
|
|
|
|
ship_to = candidate.get("ship_to") or {}
|
|
|
|
|
|
us_idx = _find_after(lines, _US_SENTINEL, frame.shipping_idxs[1] + 1)
|
|
|
|
|
|
if us_idx is None or us_idx >= frame.lines_anchor:
|
|
|
|
|
|
return "address_shape_invalid"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
m = _CITY_STATE_ZIP_RE.fullmatch(_clean(lines[us_idx - 1]) or "")
|
|
|
|
|
|
if not m:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "address_shape_invalid"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
if (
|
|
|
|
|
|
m.group("city") != ship_to["city"]
|
|
|
|
|
|
or m.group("state") != ship_to["state"]
|
|
|
|
|
|
or m.group("zip") != ship_to["zip"]
|
|
|
|
|
|
):
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "address_shape_invalid"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
if ship_to["state"] not in _USPS_STATES:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "address_shape_invalid"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
if not re.fullmatch(r"\d{5}(-\d{4})?", ship_to["zip"]):
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "address_shape_invalid"
|
|
|
|
|
|
return None
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
|
|
|
|
|
|
def _check_required_fields(candidate, frame):
|
|
|
|
|
|
"""V9 required labeled fields -> missing_required_field. Nullable by design:
|
|
|
|
|
|
on_behalf_of, department, last_opened, acknowledged_at, revision_date, attn,
|
|
|
|
|
|
quantity, unit, price."""
|
|
|
|
|
|
lines = frame.lines
|
|
|
|
|
|
items = candidate.get("line_items") or []
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
for field in (
|
|
|
|
|
|
"po_status",
|
|
|
|
|
|
"order_date",
|
|
|
|
|
|
"payment_terms",
|
|
|
|
|
|
"requisition_number",
|
|
|
|
|
|
"submitted_by",
|
|
|
|
|
|
"view_order_url",
|
|
|
|
|
|
):
|
|
|
|
|
|
if candidate.get(field) is None:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "missing_required_field"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
for item in items:
|
|
|
|
|
|
for field in (
|
|
|
|
|
|
"description",
|
|
|
|
|
|
"amount",
|
|
|
|
|
|
"currency",
|
|
|
|
|
|
"need_by",
|
|
|
|
|
|
"category",
|
|
|
|
|
|
"account_code",
|
|
|
|
|
|
"period",
|
|
|
|
|
|
):
|
|
|
|
|
|
if item.get(field) is None:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "missing_required_field"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
for label in _MORE_DETAIL_LABELS:
|
|
|
|
|
|
if len(_indices(lines, label)) != 1:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "missing_required_field"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
if candidate.get("coupa_category") != items[0].get("category"):
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "missing_required_field"
|
|
|
|
|
|
return None
|
|
|
|
|
|
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
def _check_po_identity(candidate, frame):
|
|
|
|
|
|
"""V10 PO identity proofs -> po_id_mismatch. The PO id must restate under
|
|
|
|
|
|
'PO ID', in an 'Amazon Purchase Order #<po>' line, and as the numeric tail
|
|
|
|
|
|
of the view_order_url."""
|
|
|
|
|
|
lines = frame.lines
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
po = candidate["po_number"]
|
|
|
|
|
|
po_id_idx = _indices(lines, "PO ID")[0]
|
|
|
|
|
|
v_idx = _next_visible(lines, po_id_idx)
|
|
|
|
|
|
if v_idx is None or _clean(lines[v_idx]) != po:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "po_id_mismatch"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
if not _indices(lines, f"Amazon Purchase Order #{po}"):
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "po_id_mismatch"
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
um = _ORDER_URL_ID_RE.match(candidate.get("view_order_url") or "")
|
|
|
|
|
|
if not um or um.group(1) != po.split("-", 1)[1]:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "po_id_mismatch"
|
|
|
|
|
|
return None
|
|
|
|
|
|
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
def _check_sentinels(candidate, frame):
|
|
|
|
|
|
"""V11 sentinel discipline -> unparseable_value. A present-but-unparseable
|
|
|
|
|
|
value (the _UNPARSEABLE sentinel) must fail closed."""
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
for leaf in _walk_leaves(candidate):
|
|
|
|
|
|
if isinstance(leaf, str) and leaf == _UNPARSEABLE:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "unparseable_value"
|
|
|
|
|
|
return None
|
|
|
|
|
|
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
def _check_hygiene(candidate, frame):
|
|
|
|
|
|
"""V12 hygiene -> residual_artifact. No captured value may retain a raw CR
|
|
|
|
|
|
or nbsp artifact."""
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
for leaf in _walk_leaves(candidate):
|
|
|
|
|
|
if isinstance(leaf, str) and ("\r" in leaf or "\xa0" in leaf):
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return "residual_artifact"
|
|
|
|
|
|
return None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _check_qty_unit_shape(quantity, unit, price):
|
|
|
|
|
|
"""V13 scalar shape: a positive Decimal quantity, an all-caps unit, and a
|
|
|
|
|
|
Decimal price."""
|
|
|
|
|
|
if (
|
|
|
|
|
|
not isinstance(quantity, Decimal)
|
|
|
|
|
|
or quantity <= 0
|
|
|
|
|
|
or not unit
|
|
|
|
|
|
or not re.fullmatch(r"[A-Z]+", unit)
|
|
|
|
|
|
or not isinstance(price, Decimal)
|
|
|
|
|
|
):
|
|
|
|
|
|
return "amount_mismatch"
|
|
|
|
|
|
return None
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
|
|
|
|
|
|
def _check_ea_line_evidence(lines, frame, quantity):
|
|
|
|
|
|
"""The Lines-block '<qty> EA' evidence line (when present) must numeric-
|
|
|
|
|
|
equal the summary quantity."""
|
|
|
|
|
|
ea_matches = [
|
|
|
|
|
|
m
|
|
|
|
|
|
for j in range(frame.lines_anchor + 1, frame.lines_end)
|
|
|
|
|
|
if (m := _LINE_QTY_RE.fullmatch(_clean(lines[j]) or ""))
|
|
|
|
|
|
]
|
|
|
|
|
|
if not ea_matches:
|
|
|
|
|
|
return None
|
|
|
|
|
|
if len(ea_matches) != 1:
|
|
|
|
|
|
return "amount_mismatch"
|
|
|
|
|
|
if _to_decimal(ea_matches[0].group("qty")) != quantity:
|
|
|
|
|
|
return "amount_mismatch"
|
|
|
|
|
|
return None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _check_qty_unit_price(candidate, frame):
|
|
|
|
|
|
"""V13 quantity/unit/price coherence -> amount_mismatch. Skipped entirely
|
|
|
|
|
|
when all three are absent; otherwise every piece must cohere with the
|
|
|
|
|
|
Items-summary line and the Lines-block EA evidence."""
|
|
|
|
|
|
items = candidate.get("line_items") or []
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
quantity = items[0].get("quantity")
|
|
|
|
|
|
unit = items[0].get("unit")
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
price = frame.price
|
|
|
|
|
|
if quantity is None and unit is None and price is None:
|
|
|
|
|
|
return None
|
|
|
|
|
|
reason = _check_qty_unit_shape(quantity, unit, price)
|
|
|
|
|
|
if reason:
|
|
|
|
|
|
return reason
|
|
|
|
|
|
if (
|
|
|
|
|
|
len(frame.summary_matches) != 1
|
|
|
|
|
|
or _to_decimal(frame.summary_matches[0].group("qty")) != quantity
|
|
|
|
|
|
or frame.summary_matches[0].group("unit") != unit
|
|
|
|
|
|
):
|
|
|
|
|
|
return "amount_mismatch"
|
|
|
|
|
|
return _check_ea_line_evidence(frame.lines, frame, quantity)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# Value-level rules in spec order, each returning a reason code or None. V4
|
|
|
|
|
|
# builds the shared anchor frame first (below); these consume it.
|
|
|
|
|
|
_NEW_PO_VALUE_CHECKS = (
|
|
|
|
|
|
_check_money_fidelity, # V1
|
|
|
|
|
|
_check_total_proof, # V3 + V2
|
|
|
|
|
|
_check_supplier_proof, # V5
|
|
|
|
|
|
_check_bullet_discipline, # V8
|
|
|
|
|
|
_check_ship_to_required, # V6
|
|
|
|
|
|
_check_address_shape, # V7
|
|
|
|
|
|
_check_required_fields, # V9
|
|
|
|
|
|
_check_po_identity, # V10
|
|
|
|
|
|
_check_sentinels, # V11
|
|
|
|
|
|
_check_hygiene, # V12
|
|
|
|
|
|
_check_qty_unit_price, # V13
|
|
|
|
|
|
)
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
|
|
|
|
|
|
def _validate_new_po_values(candidate, email_data):
|
|
|
|
|
|
"""Value-level gate rules V1-V13 for coupa_new_po. FAIL CLOSED.
|
|
|
|
|
|
|
|
|
|
|
|
V4 anchor integrity builds the shared anchor frame first (every later byte
|
|
|
|
|
|
proof needs it); the remaining rules then run in spec order via the per-rule
|
|
|
|
|
|
_check_* helpers, each re-deriving its evidence from email_data["body"] so
|
|
|
|
|
|
an extractor bug cannot vouch for itself. The first helper to return a
|
|
|
|
|
|
reason code short-circuits to (False, reason)."""
|
|
|
|
|
|
lines = _plain_lines(email_data["body"])
|
|
|
|
|
|
items = candidate.get("line_items") or []
|
|
|
|
|
|
frame, reason = _build_anchor_frame(lines, items)
|
|
|
|
|
|
if reason:
|
|
|
|
|
|
return False, reason
|
|
|
|
|
|
for check in _NEW_PO_VALUE_CHECKS:
|
|
|
|
|
|
reason = check(candidate, frame)
|
|
|
|
|
|
if reason:
|
|
|
|
|
|
return False, reason
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
return True, "ok"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def try_deterministic_parse(email_data):
|
|
|
|
|
|
"""Entry point. Returns (parsed|None, parse_method, template_id, reason).
|
|
|
|
|
|
|
|
|
|
|
|
On a proven-conformant parse returns (dict, 'template', template_id, 'ok').
|
|
|
|
|
|
On any miss/invalid/exception returns (None, 'ai_fallback', template_id,
|
|
|
|
|
|
reason) -- a failure is NEVER a parsed result."""
|
|
|
|
|
|
template_id = "unknown"
|
|
|
|
|
|
try:
|
|
|
|
|
|
template_id, reason = classify_template(email_data)
|
|
|
|
|
|
if template_id == "unknown":
|
|
|
|
|
|
return None, "ai_fallback", template_id, reason
|
|
|
|
|
|
|
|
|
|
|
|
if template_id == "coupa_new_po":
|
|
|
|
|
|
candidate = extract_new_po(email_data)
|
|
|
|
|
|
else:
|
|
|
|
|
|
candidate = extract_cancellation(email_data)
|
|
|
|
|
|
|
|
|
|
|
|
ok, reason = validate(candidate, template_id, email_data)
|
|
|
|
|
|
if not ok:
|
|
|
|
|
|
return None, "ai_fallback", template_id, reason
|
|
|
|
|
|
return candidate, "template", template_id, "ok"
|
|
|
|
|
|
except Exception: # noqa: BLE001 -- fail closed on ANY extractor error
|
|
|
|
|
|
return None, "ai_fallback", template_id, "extractor_raised"
|
PO ai-fallback fail-closed gate + prompt hardening (refactor phase 1) (#108)
* feat: PO ai-fallback fail-closed gate + prompt hardening, parity with #104 (refactor phase 1)
Ports WO's #104 AI-fallback security hardening to the PO email
processor, adapted for PO's nested contract instead of copying the
WO gate verbatim.
validate_ai_fallback() (template_parser.py) fail-closes raw Bedrock
output before it reaches enrich_parsed or any dispatch/save:
recursive key-set check with missing-key normalization (nested
contract: supplier{}, ship_to{}, line_items[]); po_number checked
against the same hardened prefix+hyphen+digits regex family that
guards the DynamoDB partition key the handler builds from it
(rejects fullwidth-digit and trailing-artifact injection); email_type
enforced against the {new_po, revision, cancellation} allow-list
before dispatch so a miss can never fall into the else -> save_new_po
branch; money fields accept Decimal/int/None only, matching PO's
parse_float=Decimal decode (a float-typed check would be wrong here).
A gate failure emits ParseMethod=ai_fallback_rejected and `continue`s
to the next record -- it never raises, so attacker-controlled input
can't churn the retry/DLQ path.
extract_with_claude() wraps the untrusted email in an <email> data
block and neutralizes forged <email>-tag lookalikes in the body with
the same linear-time regex approach as WO's _EMAIL_TAG_RE, and sets
temperature=0 on the Bedrock call.
Deliberate double-count: PO emits ParseMethod=ai_fallback before the
Bedrock call (so a Bedrock-side error still records the outcome), so
a rejected email always produces both an ai_fallback datapoint
(pre-call) and an ai_fallback_rejected datapoint (post-gate). This is
intentional, not a bug -- documented in handler.py, template_parser.py,
and the README.
cdk/po_stack.py: in-place property update to the existing
po-email-processor-template-fallback-rate alarm (same logical ID, no
rename/replacement) -- the fb/(fb+tmpl) expression is left
byte-identical to its pre-Phase-1 form and ai_fallback_rejected is
deliberately excluded from the numerator/denominator/volume floor,
since folding it in as WO does would double-count every rejection
(PO's pre-call emit already counts it once via fb). A net-new
EmailProcessorAiFallbackRejectedAlarm watches the rejected series on
its own, retuned for ~57 emails/day with the 6h/IF-floor/eval-4/
datapoints-2 idiom (not WO's 5-minute sparse idiom, which is
structurally dead at PO volume). Both alarms remain ALARM-only to
site-alerts, NOT_BREACHING, with no element-wise MAX in the math
(post-#102 rule).
* Block "Cancelled" po_status off the AI cancellation route
The AI-fallback gate type-checked po_status but let any string
through, unlike the template path which never emits "Cancelled" on a
new_po. Dispatch routes on email_type, so an AI-path new_po or revision
carrying po_status="Cancelled" would reach save_new_po/save_revision and
cancel a live PO via _merge_update's sticky-cancel write without ever
hitting save_cancellation. Reject the exact sticky marker on any
non-cancellation email_type so the AI path matches the template path's
guard; arbitrary non-marker status strings still pass.
email_type is already validated to the enum before this check, and a
cancellation reaches save_cancellation (which hardcodes the status), so
po_status stays irrelevant on that route.
2026-07-17 14:50:47 -04:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
# AI-fallback validation gate -- FAIL CLOSED
|
|
|
|
|
|
#
|
|
|
|
|
|
# Called on the raw Bedrock/Claude output BEFORE enrich_parsed and BEFORE any
|
|
|
|
|
|
# dispatch/save (handler.py). Mirrors WO's validate_ai_fallback, but PO's
|
|
|
|
|
|
# contract is NESTED and requires missing-key normalization, so the gate returns
|
|
|
|
|
|
# a THREE-tuple (ok, reason, normalized_candidate_or_None): on success the
|
|
|
|
|
|
# handler adopts `parsed = normalized` and never re-normalizes.
|
|
|
|
|
|
#
|
|
|
|
|
|
# Missing keys are TOLERATED (the LLM may omit null fields) and filled with None
|
|
|
|
|
|
# at every nesting level; EXTRA keys are REJECTED with "key_set_mismatch" at
|
|
|
|
|
|
# every nesting level. This deliberately does NOT reuse _normalize(), which
|
|
|
|
|
|
# silently drops extras and coerces line_items [] -> [one empty item] (that
|
|
|
|
|
|
# would change the downstream write shape -- the AI_PAYLOAD fixture ships
|
|
|
|
|
|
# line_items: [] and it must stay []).
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
|
|
|
|
|
|
# Top-level scalar fields that must be None or str (blocks LLM-emitted maps/
|
|
|
|
|
|
# lists from landing as DynamoDB Map/List attribute pollution). email_type,
|
|
|
|
|
|
# po_number, po_status are validated separately; supplier/ship_to/line_items are
|
|
|
|
|
|
# nested; total_amount is a money field.
|
|
|
|
|
|
_AI_TOP_STR_FIELDS = (
|
|
|
|
|
|
"source_system",
|
|
|
|
|
|
"submitted_by",
|
|
|
|
|
|
"on_behalf_of",
|
|
|
|
|
|
"order_date",
|
|
|
|
|
|
"revision_date",
|
|
|
|
|
|
"last_opened",
|
|
|
|
|
|
"acknowledged_at",
|
|
|
|
|
|
"payment_terms",
|
|
|
|
|
|
"requisition_number",
|
|
|
|
|
|
"department",
|
|
|
|
|
|
"view_order_url",
|
|
|
|
|
|
"site_code",
|
|
|
|
|
|
"currency",
|
|
|
|
|
|
"fiscal_year",
|
|
|
|
|
|
"trade",
|
|
|
|
|
|
"coupa_category",
|
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
# Line-item scalar fields that must be None or str. amount is a money field;
|
|
|
|
|
|
# quantity/price are money-or-str (enrich_parsed coerces numeric strings).
|
|
|
|
|
|
_AI_LINE_ITEM_STR_FIELDS = (
|
|
|
|
|
|
"description",
|
|
|
|
|
|
"currency",
|
|
|
|
|
|
"need_by",
|
|
|
|
|
|
"category",
|
|
|
|
|
|
"account_code",
|
|
|
|
|
|
"period",
|
|
|
|
|
|
"unit",
|
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _is_ai_money(value):
|
|
|
|
|
|
"""True for a valid strict money value: None | int | Decimal.
|
|
|
|
|
|
|
|
|
|
|
|
PO parses Bedrock output with parse_float=Decimal, so a float can never
|
|
|
|
|
|
legitimately occur and a float-typed check would be wrong. bool is an int
|
|
|
|
|
|
subclass and is EXPLICITLY rejected (a JSON true/false must not read as
|
|
|
|
|
|
1/0 into a money column)."""
|
|
|
|
|
|
if value is None:
|
|
|
|
|
|
return True
|
|
|
|
|
|
if isinstance(value, bool):
|
|
|
|
|
|
return False
|
|
|
|
|
|
return isinstance(value, (int, Decimal))
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _is_ai_money_or_str(value):
|
|
|
|
|
|
"""True for None | int | Decimal | str, bool rejected. str is tolerated for
|
|
|
|
|
|
quantity/price because enrich_parsed's shared coercion stage converts
|
|
|
|
|
|
numeric strings to Decimal and deliberately stores non-numeric strings
|
|
|
|
|
|
verbatim -- the gate must not break that documented contract."""
|
|
|
|
|
|
if isinstance(value, str):
|
|
|
|
|
|
return True
|
|
|
|
|
|
return _is_ai_money(value)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _normalize_nested_dict(value, keys):
|
|
|
|
|
|
"""Strict per-level normalize for a nested container (supplier/ship_to).
|
|
|
|
|
|
|
|
|
|
|
|
Returns (normalized_dict_or_None, ok):
|
|
|
|
|
|
* None -> ({k: None for k in keys}, True) (all-None dict)
|
|
|
|
|
|
* dict whose keys are a subset of `keys` -> (missing filled None, True)
|
|
|
|
|
|
* dict with any EXTRA key -> (None, False)
|
|
|
|
|
|
* any other type -> (None, False)
|
|
|
|
|
|
"""
|
|
|
|
|
|
if value is None:
|
|
|
|
|
|
return {k: None for k in keys}, True
|
|
|
|
|
|
if not isinstance(value, dict):
|
|
|
|
|
|
return None, False
|
|
|
|
|
|
if set(value.keys()) - set(keys):
|
|
|
|
|
|
return None, False
|
|
|
|
|
|
return {k: value.get(k) for k in keys}, True
|
|
|
|
|
|
|
|
|
|
|
|
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
def validate_ai_fallback(candidate): # noqa: C901, PLR0911, PLR0912
|
PO ai-fallback fail-closed gate + prompt hardening (refactor phase 1) (#108)
* feat: PO ai-fallback fail-closed gate + prompt hardening, parity with #104 (refactor phase 1)
Ports WO's #104 AI-fallback security hardening to the PO email
processor, adapted for PO's nested contract instead of copying the
WO gate verbatim.
validate_ai_fallback() (template_parser.py) fail-closes raw Bedrock
output before it reaches enrich_parsed or any dispatch/save:
recursive key-set check with missing-key normalization (nested
contract: supplier{}, ship_to{}, line_items[]); po_number checked
against the same hardened prefix+hyphen+digits regex family that
guards the DynamoDB partition key the handler builds from it
(rejects fullwidth-digit and trailing-artifact injection); email_type
enforced against the {new_po, revision, cancellation} allow-list
before dispatch so a miss can never fall into the else -> save_new_po
branch; money fields accept Decimal/int/None only, matching PO's
parse_float=Decimal decode (a float-typed check would be wrong here).
A gate failure emits ParseMethod=ai_fallback_rejected and `continue`s
to the next record -- it never raises, so attacker-controlled input
can't churn the retry/DLQ path.
extract_with_claude() wraps the untrusted email in an <email> data
block and neutralizes forged <email>-tag lookalikes in the body with
the same linear-time regex approach as WO's _EMAIL_TAG_RE, and sets
temperature=0 on the Bedrock call.
Deliberate double-count: PO emits ParseMethod=ai_fallback before the
Bedrock call (so a Bedrock-side error still records the outcome), so
a rejected email always produces both an ai_fallback datapoint
(pre-call) and an ai_fallback_rejected datapoint (post-gate). This is
intentional, not a bug -- documented in handler.py, template_parser.py,
and the README.
cdk/po_stack.py: in-place property update to the existing
po-email-processor-template-fallback-rate alarm (same logical ID, no
rename/replacement) -- the fb/(fb+tmpl) expression is left
byte-identical to its pre-Phase-1 form and ai_fallback_rejected is
deliberately excluded from the numerator/denominator/volume floor,
since folding it in as WO does would double-count every rejection
(PO's pre-call emit already counts it once via fb). A net-new
EmailProcessorAiFallbackRejectedAlarm watches the rejected series on
its own, retuned for ~57 emails/day with the 6h/IF-floor/eval-4/
datapoints-2 idiom (not WO's 5-minute sparse idiom, which is
structurally dead at PO volume). Both alarms remain ALARM-only to
site-alerts, NOT_BREACHING, with no element-wise MAX in the math
(post-#102 rule).
* Block "Cancelled" po_status off the AI cancellation route
The AI-fallback gate type-checked po_status but let any string
through, unlike the template path which never emits "Cancelled" on a
new_po. Dispatch routes on email_type, so an AI-path new_po or revision
carrying po_status="Cancelled" would reach save_new_po/save_revision and
cancel a live PO via _merge_update's sticky-cancel write without ever
hitting save_cancellation. Reject the exact sticky marker on any
non-cancellation email_type so the AI path matches the template path's
guard; arbitrary non-marker status strings still pass.
email_type is already validated to the enum before this check, and a
cancellation reaches save_cancellation (which hardcodes the status), so
po_status stays irrelevant on that route.
2026-07-17 14:50:47 -04:00
|
|
|
|
"""Fail-closed schema/type validation for the AI-fallback parse path.
|
|
|
|
|
|
|
|
|
|
|
|
Returns (ok, reason, normalized_candidate_or_None). On success the handler
|
|
|
|
|
|
adopts the returned normalized dict (`parsed = normalized`) and never
|
|
|
|
|
|
re-normalizes. Reason-code vocabulary: not_an_object, key_set_mismatch,
|
|
|
|
|
|
missing_required_field, invalid_status, invalid_money_type,
|
|
|
|
|
|
invalid_field_type, ok."""
|
|
|
|
|
|
# (1) json.loads on model output can yield list/str/int/None; only an object
|
|
|
|
|
|
# can satisfy the contract. Anything else must fail closed HERE rather than
|
|
|
|
|
|
# AttributeError at the handler's logger f-string into async retries / DLQ.
|
|
|
|
|
|
if not isinstance(candidate, dict):
|
|
|
|
|
|
return False, "not_an_object", None
|
|
|
|
|
|
|
|
|
|
|
|
# (2) key-set + missing-key normalization: extras rejected, missing -> None.
|
|
|
|
|
|
if set(candidate.keys()) - set(CONTRACT_KEYS):
|
|
|
|
|
|
return False, "key_set_mismatch", None
|
|
|
|
|
|
normalized = {k: candidate.get(k) for k in CONTRACT_KEYS}
|
|
|
|
|
|
|
|
|
|
|
|
supplier, ok = _normalize_nested_dict(normalized["supplier"], SUPPLIER_KEYS)
|
|
|
|
|
|
if not ok:
|
|
|
|
|
|
return False, "key_set_mismatch", None
|
|
|
|
|
|
normalized["supplier"] = supplier
|
|
|
|
|
|
|
|
|
|
|
|
ship_to, ok = _normalize_nested_dict(normalized["ship_to"], SHIP_TO_KEYS)
|
|
|
|
|
|
if not ok:
|
|
|
|
|
|
return False, "key_set_mismatch", None
|
|
|
|
|
|
normalized["ship_to"] = ship_to
|
|
|
|
|
|
|
|
|
|
|
|
# line_items: list or None. None -> []; [] stays [] (preserves the current
|
|
|
|
|
|
# downstream write shape). Every element must be a dict; each is normalized
|
|
|
|
|
|
# to exactly LINE_ITEM_KEYS with extras rejected.
|
|
|
|
|
|
items = normalized["line_items"]
|
|
|
|
|
|
if items is None:
|
|
|
|
|
|
items = []
|
|
|
|
|
|
elif not isinstance(items, list):
|
|
|
|
|
|
return False, "key_set_mismatch", None
|
|
|
|
|
|
norm_items = []
|
|
|
|
|
|
for it in items:
|
|
|
|
|
|
if not isinstance(it, dict):
|
|
|
|
|
|
return False, "key_set_mismatch", None
|
|
|
|
|
|
if set(it.keys()) - set(LINE_ITEM_KEYS):
|
|
|
|
|
|
return False, "key_set_mismatch", None
|
|
|
|
|
|
norm_items.append({k: it.get(k) for k in LINE_ITEM_KEYS})
|
|
|
|
|
|
normalized["line_items"] = norm_items
|
|
|
|
|
|
|
|
|
|
|
|
# (3) po_number: required non-empty, hardened prefix+hyphen+digits shape.
|
|
|
|
|
|
po = normalized["po_number"]
|
|
|
|
|
|
if not po or not _AI_PO_ID_RE.match(str(po)):
|
|
|
|
|
|
return False, "missing_required_field", None
|
|
|
|
|
|
|
|
|
|
|
|
# (4) email_type in the enum, enforced HERE (before dispatch) so a miss can
|
|
|
|
|
|
# never fall into the handler's else -> save_new_po branch. isinstance guard
|
|
|
|
|
|
# first: an unhashable JSON list/dict would raise TypeError on `in <set>`
|
|
|
|
|
|
# and escape the fail-closed gate.
|
|
|
|
|
|
et = normalized["email_type"]
|
|
|
|
|
|
if not isinstance(et, str) or et not in VALID_EMAIL_TYPES:
|
|
|
|
|
|
return False, "missing_required_field", None
|
|
|
|
|
|
|
|
|
|
|
|
# (5) po_status: None or str. PARTIAL DIVERGENCE from WO -- PO has NO closed
|
|
|
|
|
|
# AI-path status vocabulary (NEW_PO_SAFE_STATUSES is a template-path new_po
|
|
|
|
|
|
# allow-list; revision/cancellation statuses are uncharacterized), so
|
|
|
|
|
|
# arbitrary strings pass the type check -- with ONE exception: a
|
|
|
|
|
|
# non-cancellation email_type may not carry the sticky "Cancelled" status.
|
|
|
|
|
|
# Dispatch routes on email_type, so an AI-path new_po/revision carrying
|
|
|
|
|
|
# po_status="Cancelled" would reach save_new_po/save_revision and cancel a
|
|
|
|
|
|
# live PO via _merge_update while never hitting save_cancellation. The
|
|
|
|
|
|
# template path already forbids this (a cancellation misrouted as new_po
|
|
|
|
|
|
# defeats the sticky-Cancelled guard); mirror it here. email_type is already
|
|
|
|
|
|
# validated to the enum at step (4); a cancellation reaches save_cancellation,
|
|
|
|
|
|
# which hardcodes the status, so po_status is irrelevant on that route.
|
|
|
|
|
|
status = normalized["po_status"]
|
|
|
|
|
|
if status is not None and not isinstance(status, str):
|
|
|
|
|
|
return False, "invalid_status", None
|
|
|
|
|
|
if status == _CANCELLED_STATUS and normalized["email_type"] != "cancellation":
|
|
|
|
|
|
return False, "invalid_status", None
|
|
|
|
|
|
|
|
|
|
|
|
# (6) money fields, two tiers.
|
|
|
|
|
|
if not _is_ai_money(normalized["total_amount"]):
|
|
|
|
|
|
return False, "invalid_money_type", None
|
|
|
|
|
|
for it in normalized["line_items"]:
|
|
|
|
|
|
if not _is_ai_money(it["amount"]):
|
|
|
|
|
|
return False, "invalid_money_type", None
|
|
|
|
|
|
if not _is_ai_money_or_str(it["quantity"]):
|
|
|
|
|
|
return False, "invalid_money_type", None
|
|
|
|
|
|
if not _is_ai_money_or_str(it["price"]):
|
|
|
|
|
|
return False, "invalid_money_type", None
|
|
|
|
|
|
|
|
|
|
|
|
# (7) all remaining scalar fields must be None or str.
|
|
|
|
|
|
for field in _AI_TOP_STR_FIELDS:
|
|
|
|
|
|
val = normalized[field]
|
|
|
|
|
|
if val is not None and not isinstance(val, str):
|
|
|
|
|
|
return False, "invalid_field_type", None
|
|
|
|
|
|
if normalized["supplier"]["name"] is not None and not isinstance(
|
|
|
|
|
|
normalized["supplier"]["name"], str
|
|
|
|
|
|
):
|
|
|
|
|
|
return False, "invalid_field_type", None
|
|
|
|
|
|
for field in SHIP_TO_KEYS:
|
|
|
|
|
|
val = normalized["ship_to"][field]
|
|
|
|
|
|
if val is not None and not isinstance(val, str):
|
|
|
|
|
|
return False, "invalid_field_type", None
|
|
|
|
|
|
for it in normalized["line_items"]:
|
|
|
|
|
|
for field in _AI_LINE_ITEM_STR_FIELDS:
|
|
|
|
|
|
val = it[field]
|
|
|
|
|
|
if val is not None and not isinstance(val, str):
|
|
|
|
|
|
return False, "invalid_field_type", None
|
|
|
|
|
|
|
|
|
|
|
|
return True, "ok", normalized
|