procurement-ingest/lambdas/wo/email_processor/tests/test_parser.py
Adam Moussa ff368beaa4
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)

tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.

Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.

Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.

New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
  'content' key, empty content list, non-JSON model text), asserting
  PO's pre-call ai_fallback metric survives with no partial write and
  the exception propagates; WO's no-datapoint-on-throttle behavior is
  pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
  monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
  zero writes, no raise — closing the hole where deleting the gate
  line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
  unset ARN and on a Secrets Manager exception, TTL cache refresh,
  Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
  no table scan, non-ASCII token, and a hostile-field-escaping
  regression lock. PO web_ui has no __init__.py, so these go through
  the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
  (table 'WorkOrders'): null-status never clobbers wo_status,
  created_at immutable via if_not_exists, status->wo_status mapping,
  None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
  per-pipeline multi-record failure-isolation (all-or-retry contract).
  The reprocess.py synthetic-event-shape contract test already landed
  in Phase 7, so it isn't duplicated here.

Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.

_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.

CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.

docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.

* test: lock attribute-context quote escaping in web_ui hostile-field test

The escaping regression lock asserted only the element-context vector
(raw <script> absent, &lt;script&gt; present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).

Add assertions that the onclick sink's JSON string renders its opening
quote as &quot; (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.

* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override

- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
  xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
  behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
  corrected, instead of requiring a passing test to be knowingly deleted. The
  module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
  defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
  reusable workflow's bare pytest) red-exits local subset runs by design, with the
  --cov-fail-under=0 override for iteration. Floor stays in addopts because the
  centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00

174 lines
6.4 KiB
Python

"""Golden-file tests for the deterministic template parser.
Positives: every scrubbed real sample must parse via the deterministic path and
match its checked-in golden exactly. Negatives: every adversarial / drift sample
must fail closed to the AI fallback (parsed is None).
"""
import pytest
from _wo_parser_support import (
ADVERSARIAL_STEMS,
AI_FALLBACK_STEMS,
ASSIGN_STEMS,
UPDATE_STEMS,
load_email,
load_golden,
wo_template_parser as template_parser,
)
try_deterministic_parse = template_parser.try_deterministic_parse
@pytest.mark.parametrize("stem", UPDATE_STEMS)
def test_update_plaintext_matches_golden(stem):
email_data = load_email("update-plaintext", stem)
parsed, method, template_id, reason = try_deterministic_parse(email_data)
assert method == "template"
assert template_id == "update_plaintext"
assert reason == "ok"
assert parsed == load_golden(stem)
# Contract invariants for the comment shape.
assert set(parsed.keys()) == set(load_golden(stem).keys())
assert parsed["email_type"] == "comment"
assert parsed["status"] is None # never clobber a real wo_status on upsert
@pytest.mark.parametrize("stem", ASSIGN_STEMS)
def test_assign_html_matches_golden(stem):
email_data = load_email("assign-html", stem)
parsed, method, template_id, reason = try_deterministic_parse(email_data)
assert method == "template"
assert template_id == "assign_html"
assert reason == "ok"
assert parsed == load_golden(stem)
assert parsed["email_type"] == "new_work_order"
assert parsed["status"] == "assigned"
@pytest.mark.parametrize("stem", AI_FALLBACK_STEMS)
def test_ai_fallback_samples_return_none(stem):
email_data = load_email("ai-fallback", stem)
parsed, method, template_id, reason = try_deterministic_parse(email_data)
assert parsed is None, f"{stem} should have failed closed, got {parsed}"
assert method == "ai_fallback"
assert reason != "ok"
def test_positive_result_matches_ai_contract_keys():
# The deterministic result must carry EXACTLY the keys the AI fallback
# produces (the EXTRACTION_PROMPT JSON contract) -- no more, no less.
CONTRACT_KEYS = template_parser.CONTRACT_KEYS
parsed, *_ = try_deterministic_parse(
load_email("update-plaintext", UPDATE_STEMS[0])
)
assert set(parsed.keys()) == set(CONTRACT_KEYS)
def test_extractor_exception_fails_closed(monkeypatch):
# Any exception inside the extractor must yield None, never a partial parse.
monkeypatch.setattr(
template_parser,
"extract_update_plaintext",
lambda *_a, **_k: (_ for _ in ()).throw(RuntimeError("boom")),
)
parsed, method, _tid, reason = template_parser.try_deterministic_parse(
load_email("update-plaintext", UPDATE_STEMS[0])
)
assert parsed is None
assert method == "ai_fallback"
assert reason == "extractor_raised"
# --- Advisory A3: multi-paragraph comments must not be truncated ---
def test_multi_paragraph_comment_is_captured_in_full():
"""A blank line inside the New Comment block is a paragraph break, not the
end of the comment -- the block ends at the next label/separator."""
email_data = {
"subject": "AMAZON UPDATE WO DETAILS 11144580730",
"sender": "noreply@hxgnsmartcloud.com",
"to": "apm@int.seahaven.com",
"cc": "",
"date": "Mon, 27 Apr 2026 23:57:49 +0000 (UTC)",
"body": (
"New Comment:\n"
"first paragraph line one\n"
"first paragraph line two\n"
"\n"
"second paragraph after a blank line\n"
"\n"
"Creation Time(UTC): 2026-04-27 23:51:48\n"
"Submitted By: jdoe\n"
"\n"
"_________________\n"
"\n"
"Work Order: 11144580730 - fix dock door\n"
"\n"
"Building: WCO0.\n"
"\n"
"\n"
"Please review it in Amazon Portal\n"
),
}
parsed, method, template_id, reason = try_deterministic_parse(email_data)
assert method == "template"
assert template_id == "update_plaintext"
assert reason == "ok"
assert parsed["comment_text"] == (
"first paragraph line one\n"
"first paragraph line two\n"
"\n"
"second paragraph after a blank line"
)
# Fixed-position blocks still stop at blank lines: the portal footer after
# Building must not bleed into the address.
assert parsed["address"] is None
assert parsed["comment_time"] == "2026-04-27T23:51:48"
def test_huge_blank_run_in_comment_parses_in_linear_time():
"""A New Comment block padded with a massive blank-line run must not go
quadratic in the leading/trailing-blank trim (CWE-407 guard)."""
import time
_T1_LABELS = template_parser._T1_LABELS
_capture_block = template_parser._capture_block
lines = ["New Comment:"] + [""] * 500_000 + ["x"] + [""] * 500_000
start = time.monotonic()
out = _capture_block(lines, 0, _T1_LABELS, stop_on_blank=False)
elapsed = time.monotonic() - start
assert out == ["x"]
assert elapsed < 2.0 # quadratic trim took ~20s+ at this size
# --- Adversarial: prompt injection / near-miss subjects ---
def test_prompt_injection_in_comment_never_corrupts_other_fields():
"""Injection text in the comment body must either fall back OR parse safely
with the real work_order_id / site_code intact and the payload confined to
comment_text (it must never steer the structured fields)."""
email_data = load_email("adversarial", "prompt-injection-comment")
parsed, method, _tid, _reason = try_deterministic_parse(email_data)
if parsed is None:
assert method == "ai_fallback"
return
# Parsed safely: the subject-derived id wins, injected id/site are ignored.
assert parsed["work_order_id"] == "11144580730"
assert parsed["site_code"] == "WCO0"
assert parsed["email_type"] == "comment"
assert "00000000000" not in (parsed["work_order_id"] or "")
assert parsed["site_code"] != "HACK99"
# The injection payload is confined to the free-text comment field.
assert "Ignore all previous instructions" in parsed["comment_text"]
@pytest.mark.parametrize("stem", ADVERSARIAL_STEMS)
def test_adversarial_samples_do_not_raise(stem):
# The parser must be total: adversarial input returns a tuple, never raises.
result = try_deterministic_parse(load_email("adversarial", stem))
assert isinstance(result, tuple) and len(result) == 4