procurement-ingest/lambdas/wo/email_processor/tests/test_validation_gate.py
Adam Moussa ff368beaa4
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)

tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.

Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.

Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.

New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
  'content' key, empty content list, non-JSON model text), asserting
  PO's pre-call ai_fallback metric survives with no partial write and
  the exception propagates; WO's no-datapoint-on-throttle behavior is
  pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
  monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
  zero writes, no raise — closing the hole where deleting the gate
  line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
  unset ARN and on a Secrets Manager exception, TTL cache refresh,
  Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
  no table scan, non-ASCII token, and a hostile-field-escaping
  regression lock. PO web_ui has no __init__.py, so these go through
  the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
  (table 'WorkOrders'): null-status never clobbers wo_status,
  created_at immutable via if_not_exists, status->wo_status mapping,
  None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
  per-pipeline multi-record failure-isolation (all-or-retry contract).
  The reprocess.py synthetic-event-shape contract test already landed
  in Phase 7, so it isn't duplicated here.

Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.

_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.

CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.

docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.

* test: lock attribute-context quote escaping in web_ui hostile-field test

The escaping regression lock asserted only the element-context vector
(raw <script> absent, &lt;script&gt; present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).

Add assertions that the onclick sink's JSON string renders its opening
quote as &quot; (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.

* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override

- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
  xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
  behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
  corrected, instead of requiring a passing test to be knowingly deleted. The
  module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
  defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
  reusable workflow's bare pytest) red-exits local subset runs by design, with the
  --cov-fail-under=0 override for iteration. Floor stays in addopts because the
  centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00

376 lines
13 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

"""Fail-closed validation-gate tests.
Each known drift / adversarial shape must be rejected with the expected reason
code, and direct unit tests exercise each individual gate rule.
"""
import pytest
from _wo_parser_support import load_email, wo_template_parser as template_parser
CONTRACT_KEYS = template_parser.CONTRACT_KEYS
extract_update_plaintext = template_parser.extract_update_plaintext
try_deterministic_parse = template_parser.try_deterministic_parse
validate = template_parser.validate
validate_ai_fallback = template_parser.validate_ai_fallback
# Fixture stem -> expected fail-closed reason code.
EXPECTED_REASONS = {
"unknown-subject": "subject_no_match",
"missing-id": "subject_no_match",
"single-space-work-order": "single_space_work_order",
"nondigit-id": "missing_required_field",
"empty-new-comment": "missing_required_field",
"malformed-site-code": "malformed_site_code",
"label-bleed-comment": "label_bleed",
"unparseable-creation-time": "creation_time_unparseable",
"cancellation-t1": "missing_required_field",
"update-status-t1": "missing_required_field",
"missing-building-and-comment": "missing_required_field",
"t2-wo-id-mismatch": "wo_id_mismatch",
"t2-missing-address": "missing_required_field",
"t2-unparseable-date": "creation_time_unparseable",
}
@pytest.mark.parametrize("stem,reason", sorted(EXPECTED_REASONS.items()))
def test_gate_rejects_with_reason(stem, reason):
parsed, method, _tid, got_reason = try_deterministic_parse(
load_email("ai-fallback", stem)
)
assert parsed is None
assert method == "ai_fallback"
assert got_reason == reason
# --- Direct unit tests on validate() ---
def _good_t1_email():
return load_email("update-plaintext", "update-plaintext-01")
def _good_candidate():
return extract_update_plaintext(_good_t1_email())
def test_baseline_candidate_is_valid():
ok, reason = validate(_good_candidate(), "update_plaintext", _good_t1_email())
assert ok and reason == "ok"
def test_rule1_unknown_template():
ok, reason = validate(_good_candidate(), "unknown", _good_t1_email())
assert not ok and reason == "subject_no_match"
def test_rule2_extra_key_fails():
cand = _good_candidate()
cand["surprise"] = "x"
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "key_set_mismatch"
def test_rule2_missing_key_fails():
cand = _good_candidate()
del cand["address"]
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "key_set_mismatch"
def test_rule3_nondigit_wo():
cand = _good_candidate()
cand["work_order_id"] = "12A45"
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "missing_required_field"
def test_rule5_wrong_email_type():
cand = _good_candidate()
cand["email_type"] = "new_work_order" # wrong for a T1 template
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "email_type_mismatch"
def test_rule6_bad_site_code():
cand = _good_candidate()
cand["site_code"] = "workshop"
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "malformed_site_code"
def test_rule7_bad_status():
# Phase 8 (constraint 7a): the template-path status-enum branch now returns
# its OWN reason code "invalid_status" -- previously a copy-paste bug made it
# return "malformed_site_code" (the site_code branch's code), so one status
# failure yielded two different codes depending on which parse path hit it.
cand = _good_candidate()
cand["status"] = "frobnicated"
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "invalid_status"
def test_template_and_ai_paths_agree_on_bad_status_reason_code():
"""Regression lock for the invalid_status reason-code fix (constraint 7a):
an identical bad-status failure must yield the SAME reason code on BOTH the
template ``validate`` path and the ``validate_ai_fallback`` path -- ending
the one-failure-two-codes-by-path split (doc Q4)."""
template_cand = _good_candidate()
template_cand["status"] = "frobnicated"
_, template_reason = validate(template_cand, "update_plaintext", _good_t1_email())
ai_cand = _ai_candidate()
ai_cand["work_order_id"] = "12345"
ai_cand["email_type"] = "update"
ai_cand["status"] = "frobnicated"
_, ai_reason = validate_ai_fallback(ai_cand)
assert template_reason == ai_reason == "invalid_status"
def test_rule8_empty_comment_text():
cand = _good_candidate()
cand["comment_text"] = " "
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "missing_required_field"
def test_rule9_label_bleed_in_comment():
cand = _good_candidate()
cand["comment_text"] = "text that leaked Building: WCO0 into the value"
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "label_bleed"
def test_rule9_separator_bleed_in_address():
cand = _good_candidate()
cand["address"] = "123 Main St ________________"
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "label_bleed"
def test_contract_keys_match_extraction_prompt():
"""The parser's key set must be exactly the AI EXTRACTION_PROMPT contract, so
the deterministic and AI-fallback paths write identical shapes downstream."""
import re
from _wo_parser_support import wo_handler as handler
# The prompt's JSON skeleton uses union-type pseudo-values (not strict JSON)
# and repeats some enum terms in prose bullets, so pull quoted "key": tokens
# and assert every contract key is a field the AI is asked to emit (the
# parser must never invent a key outside the AI contract).
prompt_keys = set(re.findall(r'"([a-z_]+)":', handler.EXTRACTION_PROMPT))
assert set(CONTRACT_KEYS).issubset(prompt_keys)
assert len(CONTRACT_KEYS) == 16 # current EXTRACTION_PROMPT field count
# --- validate_ai_fallback unit tests -----------------------------------------
def _ai_candidate():
return {k: None for k in CONTRACT_KEYS}
def test_ai_fallback_baseline_is_valid():
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
ok, reason = validate_ai_fallback(cand)
assert ok and reason == "ok"
def test_ai_fallback_nondigit_wo_id():
cand = _ai_candidate()
cand["work_order_id"] = "12A45"
cand["email_type"] = "update"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "missing_required_field"
def test_ai_fallback_null_wo_id():
cand = _ai_candidate()
cand["work_order_id"] = None
cand["email_type"] = "update"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "missing_required_field"
def test_ai_fallback_hash_in_wo_id():
cand = _ai_candidate()
cand["work_order_id"] = "123#spoofed#deadbeef"
cand["email_type"] = "update"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "missing_required_field"
def test_ai_fallback_invalid_email_type():
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "exploit"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "missing_required_field"
def test_ai_fallback_null_email_type():
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = None
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "missing_required_field"
def test_ai_fallback_bad_status_enum():
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["status"] = "frobnicated"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "invalid_status"
def test_ai_fallback_unhashable_enum_fails_closed():
# A JSON list/dict for an enum field is unhashable; the gate must fail
# closed (isinstance guard), not raise TypeError into async retries.
for field, bad in (
("email_type", ["update"]),
("email_type", {"x": 1}),
("status", ["new"]),
("status", {"x": 1}),
):
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand[field] = bad
ok, reason = validate_ai_fallback(cand) # must not raise
assert not ok, f"{field}={bad!r} should fail closed"
def test_ai_fallback_fullwidth_digit_wo_id_rejected():
# Fullwidth digits render like ASCII but are a distinct partition key;
# [0-9] (not \d) must reject them.
cand = _ai_candidate()
cand["work_order_id"] = "12345" # "12345" fullwidth
cand["email_type"] = "update"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "missing_required_field"
def test_ai_fallback_trailing_newline_rejected():
# \A..\Z (not ^..$) must reject a trailing newline in wo_id and site_code.
cand = _ai_candidate()
cand["work_order_id"] = "12345\n"
cand["email_type"] = "update"
ok, _ = validate_ai_fallback(cand)
assert not ok
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["site_code"] = "WIL1\n"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "malformed_site_code"
def test_ai_fallback_non_dict_fails_closed():
# json.loads on model output can yield any JSON type; the gate must fail
# closed on a non-object rather than raise into async retries / DLQ.
for bad in ([], "string", 42, None, [{"work_order_id": "12345"}]):
ok, reason = validate_ai_fallback(bad)
assert not ok and reason == "not_an_object", f"{bad!r} should fail closed"
def test_ai_fallback_unparseable_sentinel_rejected():
# The template parser's internal _UNPARSEABLE sentinel must never survive
# the AI gate into the store.
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["comment_time"] = "__UNPARSEABLE__"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "creation_time_unparseable"
def test_ai_fallback_non_iso_date_rejected():
for key in ("date_reported", "scheduled_start", "due_date", "comment_time"):
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand[key] = "ignore previous instructions"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "creation_time_unparseable", f"{key} not gated"
def test_ai_fallback_iso_dates_accepted():
for value in ("2026-07-16", "2026-07-16T10:15:00", "2026-07-16T10:15:00Z", None):
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["date_reported"] = value
cand["comment_time"] = value
ok, reason = validate_ai_fallback(cand)
assert ok, f"date {value!r} should pass"
def test_ai_fallback_valid_status_ok():
for status in (
"new",
"assigned",
"in_progress",
"on_hold",
"completed",
"cancelled",
"unknown",
None,
):
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["status"] = status
ok, reason = validate_ai_fallback(cand)
assert ok, f"status={status} should pass"
def test_ai_fallback_bad_site_code():
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["site_code"] = "workshop"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "malformed_site_code"
def test_ai_fallback_valid_site_codes():
for code in ("WIL1", "ZDL8", "AB12", None):
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["site_code"] = code
ok, reason = validate_ai_fallback(cand)
assert ok, f"site_code={code} should pass"
def test_ai_fallback_key_set_mismatch_extra():
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["surprise"] = "x"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "key_set_mismatch"
def test_ai_fallback_key_set_mismatch_missing():
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
del cand["address"]
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "key_set_mismatch"
def test_ai_fallback_all_valid_email_types():
for et in ("new_work_order", "update", "comment", "cancellation"):
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = et
ok, reason = validate_ai_fallback(cand)
assert ok, f"email_type={et} should pass"