procurement-ingest/lambdas/wo/email_processor/tests/test_validation_gate.py

425 lines
14 KiB
Python
Raw Normal View History

feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
"""Fail-closed validation-gate tests.
Each known drift / adversarial shape must be rejected with the expected reason
code, and direct unit tests exercise each individual gate rule.
"""
import pytest
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118) * test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8) tests/conftest.py only loads for the tests/ root, not a standalone `pytest lambdas/po/email_processor/tests` run, so it could never carry session invariants like the dummy AWS env or the moto stubber registration. Add a single repo-root conftest.py (pytest.ini pins rootdir there, so it loads for every invocation) that sets the dummy AWS credentials/region, imports moto BEFORE any handler module so boto3 sessions pick up its stubber hook (carrying the explanatory comment verbatim from the old _po_parser_support.py), and exposes one load_lambda_module(pipeline, name) — the sys.modules save/restore dance stays, since template_parser is still a duplicated bare name across pipelines needing per-exec sibling binding. Add tests/support/ as the shared package both pipelines' local _*_parser_support.py modules delegate to: a superset FakeTable (PO's update_item recording + WO's put_item and keyed single-row store), FakeDynamoResource, load_email, and load_golden with parse_float=Decimal kept (load-bearing for exact money comparison at PO magnitudes — WO's prior load_golden had no parse_float and must not regress PO by losing it). Rewrite _wo_parser_support.py off the bare `import handler` / `from handler import parse_raw_email` strategy that was the source of the bare-name sys.modules collision the other two loaders defend against. Move test_po_merge.py and test_pad_zip.py into lambdas/po/email_processor/tests/ (PO-specific, belongs beside the code) via git mv so history follows; test_parse_raw_email.py and test_ses_auth.py stay at the repo root since they're genuinely cross-pipeline, parameterized over both handlers. Delete tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only, and imports a handler at collection time, bypassing the loader gate entirely — the golden suites already cover its role. Its pytest.ini exclusion comment goes with it. New scenario coverage, all built on the single loader + support package: - PO+WO Bedrock transport errors (ThrottlingException, missing 'content' key, empty content list, non-JSON model text), asserting PO's pre-call ai_fallback metric survives with no partial write and the exception propagates; WO's no-datapoint-on-throttle behavior is pinned with a documenting test rather than "fixed" by reordering. - Handler-level SES-auth reject seam per pipeline: no auth monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls, zero writes, no raise — closing the hole where deleting the gate line today still passes every test. - web_ui coverage for both PO and WO (0% before this): fail-closed on unset ARN and on a Secrets Manager exception, TTL cache refresh, Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with no table scan, non-ASCII token, and a hostile-field-escaping regression lock. PO web_ui has no __init__.py, so these go through the loader rather than package imports. - A moto-backed mirror of test_po_merge for WO merge semantics (table 'WorkOrders'): null-status never clobbers wo_status, created_at immutable via if_not_exists, status->wo_status mapping, None fields absent from SET, record_type only-when-present. - Small pins: the PO-DC-02 64-char EMF clamp regression and per-pipeline multi-record failure-isolation (all-or-retry contract). The reprocess.py synthetic-event-shape contract test already landed in Phase 7, so it isn't duplicated here. Two WO product-code fixes ride along, since this is the phase that exercises them: (a) the invalid_status reason-code fix in template_parser.py's status check, which previously returned malformed_site_code for the same failure validate_ai_fallback already labels invalid_status, making one failure surface two codes depending on path (grepped the dashboards/metric filters for malformed_site_code first — no external references found, safe to diverge the two codes); (b) wrapping the WO Bedrock call in handler.py so a transport failure emits ai_fallback/bedrock_error in an except-and-reraise. This is deliberately not a naive reorder: the emit sits in the except block, not pre-call, so a gate-rejected email still emits only ai_fallback_rejected and wo_stack's "a rejected email emits nothing else" alarm contract doesn't double-count. A test computes the emitted series by hand to pin the no-double-count behavior. Neither change touches the handler event/return contract. _validate_new_po_values in the PO template_parser.py is split into per-rule helpers, and the V4 anchor-frame dataclass now carries summary_matches/price so V13 can consume them; extract_new_po (C901=35) is included in the split. Add ruff.toml enabling C901/PLR so the mccabe/complexity suppressions scattered through the tree stop being decorative; derived_fields.py is under the shadow-bake freeze so its violations are silenced via a per-file ignore with a justification comment instead of an in-file edit, and the handful of other pre-existing violations surfaced by turning the config on get the same per-file-ignore treatment with a reason, or a fix where the file isn't frozen. scripts/ is added to the CI lint scope. CI gains an explicit --cov module list (lambdas/po and wo email_processor + web_ui, po/site_extractor, lambdas/shared) plus --cov-fail-under=80, since web_ui and site_extractor lack __init__.py markers and a bare --cov=lambdas silently skips them for the missing package marker; .coveragerc omits the test dirs themselves from the count. The Phase 0 AST bundle-consistency test stays in the standard pytest run. .gitignore picks up the resulting .coverage data file. docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT declares quantity/price as "number or null", not JSON strings, so parse_float=Decimal already handles a conforming Bedrock response — the doc previously implied the coercion path was the primary mechanism rather than a defensive net for non-conforming responses. * test: lock attribute-context quote escaping in web_ui hostile-field test The escaping regression lock asserted only the element-context vector (raw <script> absent, &lt;script&gt; present) while its docstring claimed quotes were covered -- the payload's " and ' were never asserted on, so a quote-escaping regression on the onclick row-link sink (attribute breakout -> event-handler injection) would have passed green. /sh-security-review finding WC-01 (confirmed medium, test-integrity). Add assertions that the onclick sink's JSON string renders its opening quote as &quot; (raw " after window.location= fails), that the payload's quote characters appear only entity-escaped, and that the raw payload never appears anywhere in the body. Mutation-verified: the test now fails when the sink's quote-escaping is dropped. * test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override - tests/test_web_ui_auth.py: replace the TypeError characterization pin with an xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False) behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is corrected, instead of requiring a passing test to be knowingly deleted. The module stays frozen this phase; the underlying hmac.compare_digest ASCII-only defect is tracked as a follow-up. - pytest.ini: document that the aggregate 80% floor (enforced in CI via the reusable workflow's bare pytest) red-exits local subset runs by design, with the --cov-fail-under=0 override for iteration. Floor stays in addopts because the centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
from _wo_parser_support import load_email, wo_template_parser as template_parser
CONTRACT_KEYS = template_parser.CONTRACT_KEYS
extract_update_plaintext = template_parser.extract_update_plaintext
try_deterministic_parse = template_parser.try_deterministic_parse
validate = template_parser.validate
validate_ai_fallback = template_parser.validate_ai_fallback
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
# Fixture stem -> expected fail-closed reason code.
EXPECTED_REASONS = {
"unknown-subject": "subject_no_match",
"missing-id": "subject_no_match",
"single-space-work-order": "single_space_work_order",
"nondigit-id": "missing_required_field",
"empty-new-comment": "missing_required_field",
"malformed-site-code": "malformed_site_code",
"label-bleed-comment": "label_bleed",
"unparseable-creation-time": "creation_time_unparseable",
"cancellation-t1": "missing_required_field",
"update-status-t1": "missing_required_field",
"missing-building-and-comment": "missing_required_field",
"t2-wo-id-mismatch": "wo_id_mismatch",
"t2-missing-address": "missing_required_field",
"t2-unparseable-date": "creation_time_unparseable",
}
@pytest.mark.parametrize("stem,reason", sorted(EXPECTED_REASONS.items()))
def test_gate_rejects_with_reason(stem, reason):
parsed, method, _tid, got_reason = try_deterministic_parse(
load_email("ai-fallback", stem)
)
assert parsed is None
assert method == "ai_fallback"
assert got_reason == reason
# --- Direct unit tests on validate() ---
def _good_t1_email():
return load_email("update-plaintext", "update-plaintext-01")
def _good_candidate():
return extract_update_plaintext(_good_t1_email())
def test_baseline_candidate_is_valid():
ok, reason = validate(_good_candidate(), "update_plaintext", _good_t1_email())
assert ok and reason == "ok"
def test_rule1_unknown_template():
ok, reason = validate(_good_candidate(), "unknown", _good_t1_email())
assert not ok and reason == "subject_no_match"
def test_rule2_extra_key_fails():
cand = _good_candidate()
cand["surprise"] = "x"
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "key_set_mismatch"
def test_rule2_missing_key_fails():
cand = _good_candidate()
del cand["address"]
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "key_set_mismatch"
def test_rule3_nondigit_wo():
cand = _good_candidate()
cand["work_order_id"] = "12A45"
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "missing_required_field"
def test_rule5_wrong_email_type():
cand = _good_candidate()
cand["email_type"] = "new_work_order" # wrong for a T1 template
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "email_type_mismatch"
def test_rule6_bad_site_code():
cand = _good_candidate()
cand["site_code"] = "workshop"
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "malformed_site_code"
def test_rule7_bad_status():
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118) * test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8) tests/conftest.py only loads for the tests/ root, not a standalone `pytest lambdas/po/email_processor/tests` run, so it could never carry session invariants like the dummy AWS env or the moto stubber registration. Add a single repo-root conftest.py (pytest.ini pins rootdir there, so it loads for every invocation) that sets the dummy AWS credentials/region, imports moto BEFORE any handler module so boto3 sessions pick up its stubber hook (carrying the explanatory comment verbatim from the old _po_parser_support.py), and exposes one load_lambda_module(pipeline, name) — the sys.modules save/restore dance stays, since template_parser is still a duplicated bare name across pipelines needing per-exec sibling binding. Add tests/support/ as the shared package both pipelines' local _*_parser_support.py modules delegate to: a superset FakeTable (PO's update_item recording + WO's put_item and keyed single-row store), FakeDynamoResource, load_email, and load_golden with parse_float=Decimal kept (load-bearing for exact money comparison at PO magnitudes — WO's prior load_golden had no parse_float and must not regress PO by losing it). Rewrite _wo_parser_support.py off the bare `import handler` / `from handler import parse_raw_email` strategy that was the source of the bare-name sys.modules collision the other two loaders defend against. Move test_po_merge.py and test_pad_zip.py into lambdas/po/email_processor/tests/ (PO-specific, belongs beside the code) via git mv so history follows; test_parse_raw_email.py and test_ses_auth.py stay at the repo root since they're genuinely cross-pipeline, parameterized over both handlers. Delete tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only, and imports a handler at collection time, bypassing the loader gate entirely — the golden suites already cover its role. Its pytest.ini exclusion comment goes with it. New scenario coverage, all built on the single loader + support package: - PO+WO Bedrock transport errors (ThrottlingException, missing 'content' key, empty content list, non-JSON model text), asserting PO's pre-call ai_fallback metric survives with no partial write and the exception propagates; WO's no-datapoint-on-throttle behavior is pinned with a documenting test rather than "fixed" by reordering. - Handler-level SES-auth reject seam per pipeline: no auth monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls, zero writes, no raise — closing the hole where deleting the gate line today still passes every test. - web_ui coverage for both PO and WO (0% before this): fail-closed on unset ARN and on a Secrets Manager exception, TTL cache refresh, Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with no table scan, non-ASCII token, and a hostile-field-escaping regression lock. PO web_ui has no __init__.py, so these go through the loader rather than package imports. - A moto-backed mirror of test_po_merge for WO merge semantics (table 'WorkOrders'): null-status never clobbers wo_status, created_at immutable via if_not_exists, status->wo_status mapping, None fields absent from SET, record_type only-when-present. - Small pins: the PO-DC-02 64-char EMF clamp regression and per-pipeline multi-record failure-isolation (all-or-retry contract). The reprocess.py synthetic-event-shape contract test already landed in Phase 7, so it isn't duplicated here. Two WO product-code fixes ride along, since this is the phase that exercises them: (a) the invalid_status reason-code fix in template_parser.py's status check, which previously returned malformed_site_code for the same failure validate_ai_fallback already labels invalid_status, making one failure surface two codes depending on path (grepped the dashboards/metric filters for malformed_site_code first — no external references found, safe to diverge the two codes); (b) wrapping the WO Bedrock call in handler.py so a transport failure emits ai_fallback/bedrock_error in an except-and-reraise. This is deliberately not a naive reorder: the emit sits in the except block, not pre-call, so a gate-rejected email still emits only ai_fallback_rejected and wo_stack's "a rejected email emits nothing else" alarm contract doesn't double-count. A test computes the emitted series by hand to pin the no-double-count behavior. Neither change touches the handler event/return contract. _validate_new_po_values in the PO template_parser.py is split into per-rule helpers, and the V4 anchor-frame dataclass now carries summary_matches/price so V13 can consume them; extract_new_po (C901=35) is included in the split. Add ruff.toml enabling C901/PLR so the mccabe/complexity suppressions scattered through the tree stop being decorative; derived_fields.py is under the shadow-bake freeze so its violations are silenced via a per-file ignore with a justification comment instead of an in-file edit, and the handful of other pre-existing violations surfaced by turning the config on get the same per-file-ignore treatment with a reason, or a fix where the file isn't frozen. scripts/ is added to the CI lint scope. CI gains an explicit --cov module list (lambdas/po and wo email_processor + web_ui, po/site_extractor, lambdas/shared) plus --cov-fail-under=80, since web_ui and site_extractor lack __init__.py markers and a bare --cov=lambdas silently skips them for the missing package marker; .coveragerc omits the test dirs themselves from the count. The Phase 0 AST bundle-consistency test stays in the standard pytest run. .gitignore picks up the resulting .coverage data file. docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT declares quantity/price as "number or null", not JSON strings, so parse_float=Decimal already handles a conforming Bedrock response — the doc previously implied the coercion path was the primary mechanism rather than a defensive net for non-conforming responses. * test: lock attribute-context quote escaping in web_ui hostile-field test The escaping regression lock asserted only the element-context vector (raw <script> absent, &lt;script&gt; present) while its docstring claimed quotes were covered -- the payload's " and ' were never asserted on, so a quote-escaping regression on the onclick row-link sink (attribute breakout -> event-handler injection) would have passed green. /sh-security-review finding WC-01 (confirmed medium, test-integrity). Add assertions that the onclick sink's JSON string renders its opening quote as &quot; (raw " after window.location= fails), that the payload's quote characters appear only entity-escaped, and that the raw payload never appears anywhere in the body. Mutation-verified: the test now fails when the sink's quote-escaping is dropped. * test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override - tests/test_web_ui_auth.py: replace the TypeError characterization pin with an xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False) behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is corrected, instead of requiring a passing test to be knowingly deleted. The module stays frozen this phase; the underlying hmac.compare_digest ASCII-only defect is tracked as a follow-up. - pytest.ini: document that the aggregate 80% floor (enforced in CI via the reusable workflow's bare pytest) red-exits local subset runs by design, with the --cov-fail-under=0 override for iteration. Floor stays in addopts because the centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
# Phase 8 (constraint 7a): the template-path status-enum branch now returns
# its OWN reason code "invalid_status" -- previously a copy-paste bug made it
# return "malformed_site_code" (the site_code branch's code), so one status
# failure yielded two different codes depending on which parse path hit it.
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
cand = _good_candidate()
cand["status"] = "frobnicated"
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118) * test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8) tests/conftest.py only loads for the tests/ root, not a standalone `pytest lambdas/po/email_processor/tests` run, so it could never carry session invariants like the dummy AWS env or the moto stubber registration. Add a single repo-root conftest.py (pytest.ini pins rootdir there, so it loads for every invocation) that sets the dummy AWS credentials/region, imports moto BEFORE any handler module so boto3 sessions pick up its stubber hook (carrying the explanatory comment verbatim from the old _po_parser_support.py), and exposes one load_lambda_module(pipeline, name) — the sys.modules save/restore dance stays, since template_parser is still a duplicated bare name across pipelines needing per-exec sibling binding. Add tests/support/ as the shared package both pipelines' local _*_parser_support.py modules delegate to: a superset FakeTable (PO's update_item recording + WO's put_item and keyed single-row store), FakeDynamoResource, load_email, and load_golden with parse_float=Decimal kept (load-bearing for exact money comparison at PO magnitudes — WO's prior load_golden had no parse_float and must not regress PO by losing it). Rewrite _wo_parser_support.py off the bare `import handler` / `from handler import parse_raw_email` strategy that was the source of the bare-name sys.modules collision the other two loaders defend against. Move test_po_merge.py and test_pad_zip.py into lambdas/po/email_processor/tests/ (PO-specific, belongs beside the code) via git mv so history follows; test_parse_raw_email.py and test_ses_auth.py stay at the repo root since they're genuinely cross-pipeline, parameterized over both handlers. Delete tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only, and imports a handler at collection time, bypassing the loader gate entirely — the golden suites already cover its role. Its pytest.ini exclusion comment goes with it. New scenario coverage, all built on the single loader + support package: - PO+WO Bedrock transport errors (ThrottlingException, missing 'content' key, empty content list, non-JSON model text), asserting PO's pre-call ai_fallback metric survives with no partial write and the exception propagates; WO's no-datapoint-on-throttle behavior is pinned with a documenting test rather than "fixed" by reordering. - Handler-level SES-auth reject seam per pipeline: no auth monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls, zero writes, no raise — closing the hole where deleting the gate line today still passes every test. - web_ui coverage for both PO and WO (0% before this): fail-closed on unset ARN and on a Secrets Manager exception, TTL cache refresh, Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with no table scan, non-ASCII token, and a hostile-field-escaping regression lock. PO web_ui has no __init__.py, so these go through the loader rather than package imports. - A moto-backed mirror of test_po_merge for WO merge semantics (table 'WorkOrders'): null-status never clobbers wo_status, created_at immutable via if_not_exists, status->wo_status mapping, None fields absent from SET, record_type only-when-present. - Small pins: the PO-DC-02 64-char EMF clamp regression and per-pipeline multi-record failure-isolation (all-or-retry contract). The reprocess.py synthetic-event-shape contract test already landed in Phase 7, so it isn't duplicated here. Two WO product-code fixes ride along, since this is the phase that exercises them: (a) the invalid_status reason-code fix in template_parser.py's status check, which previously returned malformed_site_code for the same failure validate_ai_fallback already labels invalid_status, making one failure surface two codes depending on path (grepped the dashboards/metric filters for malformed_site_code first — no external references found, safe to diverge the two codes); (b) wrapping the WO Bedrock call in handler.py so a transport failure emits ai_fallback/bedrock_error in an except-and-reraise. This is deliberately not a naive reorder: the emit sits in the except block, not pre-call, so a gate-rejected email still emits only ai_fallback_rejected and wo_stack's "a rejected email emits nothing else" alarm contract doesn't double-count. A test computes the emitted series by hand to pin the no-double-count behavior. Neither change touches the handler event/return contract. _validate_new_po_values in the PO template_parser.py is split into per-rule helpers, and the V4 anchor-frame dataclass now carries summary_matches/price so V13 can consume them; extract_new_po (C901=35) is included in the split. Add ruff.toml enabling C901/PLR so the mccabe/complexity suppressions scattered through the tree stop being decorative; derived_fields.py is under the shadow-bake freeze so its violations are silenced via a per-file ignore with a justification comment instead of an in-file edit, and the handful of other pre-existing violations surfaced by turning the config on get the same per-file-ignore treatment with a reason, or a fix where the file isn't frozen. scripts/ is added to the CI lint scope. CI gains an explicit --cov module list (lambdas/po and wo email_processor + web_ui, po/site_extractor, lambdas/shared) plus --cov-fail-under=80, since web_ui and site_extractor lack __init__.py markers and a bare --cov=lambdas silently skips them for the missing package marker; .coveragerc omits the test dirs themselves from the count. The Phase 0 AST bundle-consistency test stays in the standard pytest run. .gitignore picks up the resulting .coverage data file. docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT declares quantity/price as "number or null", not JSON strings, so parse_float=Decimal already handles a conforming Bedrock response — the doc previously implied the coercion path was the primary mechanism rather than a defensive net for non-conforming responses. * test: lock attribute-context quote escaping in web_ui hostile-field test The escaping regression lock asserted only the element-context vector (raw <script> absent, &lt;script&gt; present) while its docstring claimed quotes were covered -- the payload's " and ' were never asserted on, so a quote-escaping regression on the onclick row-link sink (attribute breakout -> event-handler injection) would have passed green. /sh-security-review finding WC-01 (confirmed medium, test-integrity). Add assertions that the onclick sink's JSON string renders its opening quote as &quot; (raw " after window.location= fails), that the payload's quote characters appear only entity-escaped, and that the raw payload never appears anywhere in the body. Mutation-verified: the test now fails when the sink's quote-escaping is dropped. * test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override - tests/test_web_ui_auth.py: replace the TypeError characterization pin with an xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False) behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is corrected, instead of requiring a passing test to be knowingly deleted. The module stays frozen this phase; the underlying hmac.compare_digest ASCII-only defect is tracked as a follow-up. - pytest.ini: document that the aggregate 80% floor (enforced in CI via the reusable workflow's bare pytest) red-exits local subset runs by design, with the --cov-fail-under=0 override for iteration. Floor stays in addopts because the centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
assert not ok and reason == "invalid_status"
def test_template_and_ai_paths_agree_on_bad_status_reason_code():
"""Regression lock for the invalid_status reason-code fix (constraint 7a):
an identical bad-status failure must yield the SAME reason code on BOTH the
template ``validate`` path and the ``validate_ai_fallback`` path -- ending
the one-failure-two-codes-by-path split (doc Q4)."""
template_cand = _good_candidate()
template_cand["status"] = "frobnicated"
_, template_reason = validate(template_cand, "update_plaintext", _good_t1_email())
ai_cand = _ai_candidate()
ai_cand["work_order_id"] = "12345"
ai_cand["email_type"] = "update"
ai_cand["status"] = "frobnicated"
_, ai_reason = validate_ai_fallback(ai_cand)
assert template_reason == ai_reason == "invalid_status"
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
def test_rule8_empty_comment_text():
cand = _good_candidate()
cand["comment_text"] = " "
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "missing_required_field"
def test_rule9_label_bleed_in_comment():
cand = _good_candidate()
cand["comment_text"] = "text that leaked Building: WCO0 into the value"
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "label_bleed"
def test_rule9_separator_bleed_in_address():
cand = _good_candidate()
cand["address"] = "123 Main St ________________"
ok, reason = validate(cand, "update_plaintext", _good_t1_email())
assert not ok and reason == "label_bleed"
def test_contract_keys_match_extraction_prompt():
"""The parser's key set must be exactly the AI EXTRACTION_PROMPT contract, so
the deterministic and AI-fallback paths write identical shapes downstream."""
import re
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118) * test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8) tests/conftest.py only loads for the tests/ root, not a standalone `pytest lambdas/po/email_processor/tests` run, so it could never carry session invariants like the dummy AWS env or the moto stubber registration. Add a single repo-root conftest.py (pytest.ini pins rootdir there, so it loads for every invocation) that sets the dummy AWS credentials/region, imports moto BEFORE any handler module so boto3 sessions pick up its stubber hook (carrying the explanatory comment verbatim from the old _po_parser_support.py), and exposes one load_lambda_module(pipeline, name) — the sys.modules save/restore dance stays, since template_parser is still a duplicated bare name across pipelines needing per-exec sibling binding. Add tests/support/ as the shared package both pipelines' local _*_parser_support.py modules delegate to: a superset FakeTable (PO's update_item recording + WO's put_item and keyed single-row store), FakeDynamoResource, load_email, and load_golden with parse_float=Decimal kept (load-bearing for exact money comparison at PO magnitudes — WO's prior load_golden had no parse_float and must not regress PO by losing it). Rewrite _wo_parser_support.py off the bare `import handler` / `from handler import parse_raw_email` strategy that was the source of the bare-name sys.modules collision the other two loaders defend against. Move test_po_merge.py and test_pad_zip.py into lambdas/po/email_processor/tests/ (PO-specific, belongs beside the code) via git mv so history follows; test_parse_raw_email.py and test_ses_auth.py stay at the repo root since they're genuinely cross-pipeline, parameterized over both handlers. Delete tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only, and imports a handler at collection time, bypassing the loader gate entirely — the golden suites already cover its role. Its pytest.ini exclusion comment goes with it. New scenario coverage, all built on the single loader + support package: - PO+WO Bedrock transport errors (ThrottlingException, missing 'content' key, empty content list, non-JSON model text), asserting PO's pre-call ai_fallback metric survives with no partial write and the exception propagates; WO's no-datapoint-on-throttle behavior is pinned with a documenting test rather than "fixed" by reordering. - Handler-level SES-auth reject seam per pipeline: no auth monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls, zero writes, no raise — closing the hole where deleting the gate line today still passes every test. - web_ui coverage for both PO and WO (0% before this): fail-closed on unset ARN and on a Secrets Manager exception, TTL cache refresh, Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with no table scan, non-ASCII token, and a hostile-field-escaping regression lock. PO web_ui has no __init__.py, so these go through the loader rather than package imports. - A moto-backed mirror of test_po_merge for WO merge semantics (table 'WorkOrders'): null-status never clobbers wo_status, created_at immutable via if_not_exists, status->wo_status mapping, None fields absent from SET, record_type only-when-present. - Small pins: the PO-DC-02 64-char EMF clamp regression and per-pipeline multi-record failure-isolation (all-or-retry contract). The reprocess.py synthetic-event-shape contract test already landed in Phase 7, so it isn't duplicated here. Two WO product-code fixes ride along, since this is the phase that exercises them: (a) the invalid_status reason-code fix in template_parser.py's status check, which previously returned malformed_site_code for the same failure validate_ai_fallback already labels invalid_status, making one failure surface two codes depending on path (grepped the dashboards/metric filters for malformed_site_code first — no external references found, safe to diverge the two codes); (b) wrapping the WO Bedrock call in handler.py so a transport failure emits ai_fallback/bedrock_error in an except-and-reraise. This is deliberately not a naive reorder: the emit sits in the except block, not pre-call, so a gate-rejected email still emits only ai_fallback_rejected and wo_stack's "a rejected email emits nothing else" alarm contract doesn't double-count. A test computes the emitted series by hand to pin the no-double-count behavior. Neither change touches the handler event/return contract. _validate_new_po_values in the PO template_parser.py is split into per-rule helpers, and the V4 anchor-frame dataclass now carries summary_matches/price so V13 can consume them; extract_new_po (C901=35) is included in the split. Add ruff.toml enabling C901/PLR so the mccabe/complexity suppressions scattered through the tree stop being decorative; derived_fields.py is under the shadow-bake freeze so its violations are silenced via a per-file ignore with a justification comment instead of an in-file edit, and the handful of other pre-existing violations surfaced by turning the config on get the same per-file-ignore treatment with a reason, or a fix where the file isn't frozen. scripts/ is added to the CI lint scope. CI gains an explicit --cov module list (lambdas/po and wo email_processor + web_ui, po/site_extractor, lambdas/shared) plus --cov-fail-under=80, since web_ui and site_extractor lack __init__.py markers and a bare --cov=lambdas silently skips them for the missing package marker; .coveragerc omits the test dirs themselves from the count. The Phase 0 AST bundle-consistency test stays in the standard pytest run. .gitignore picks up the resulting .coverage data file. docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT declares quantity/price as "number or null", not JSON strings, so parse_float=Decimal already handles a conforming Bedrock response — the doc previously implied the coercion path was the primary mechanism rather than a defensive net for non-conforming responses. * test: lock attribute-context quote escaping in web_ui hostile-field test The escaping regression lock asserted only the element-context vector (raw <script> absent, &lt;script&gt; present) while its docstring claimed quotes were covered -- the payload's " and ' were never asserted on, so a quote-escaping regression on the onclick row-link sink (attribute breakout -> event-handler injection) would have passed green. /sh-security-review finding WC-01 (confirmed medium, test-integrity). Add assertions that the onclick sink's JSON string renders its opening quote as &quot; (raw " after window.location= fails), that the payload's quote characters appear only entity-escaped, and that the raw payload never appears anywhere in the body. Mutation-verified: the test now fails when the sink's quote-escaping is dropped. * test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override - tests/test_web_ui_auth.py: replace the TypeError characterization pin with an xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False) behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is corrected, instead of requiring a passing test to be knowingly deleted. The module stays frozen this phase; the underlying hmac.compare_digest ASCII-only defect is tracked as a follow-up. - pytest.ini: document that the aggregate 80% floor (enforced in CI via the reusable workflow's bare pytest) red-exits local subset runs by design, with the --cov-fail-under=0 override for iteration. Floor stays in addopts because the centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
from _wo_parser_support import wo_handler as handler
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
# The prompt's JSON skeleton uses union-type pseudo-values (not strict JSON)
# and repeats some enum terms in prose bullets, so pull quoted "key": tokens
# and assert every contract key is a field the AI is asked to emit (the
# parser must never invent a key outside the AI contract).
prompt_keys = set(re.findall(r'"([a-z_]+)":', handler.EXTRACTION_PROMPT))
assert set(CONTRACT_KEYS).issubset(prompt_keys)
assert len(CONTRACT_KEYS) == 16 # current EXTRACTION_PROMPT field count
fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path (#104) * fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path The ai_fallback parse path applied no validation gate to raw Bedrock/LLM output before DynamoDB writes, and the extraction prompt concatenated the untrusted email body directly with no instructions-vs-data delimiter. A DKIM-passing attacker could prompt-inject arbitrary field values into the work-order store. Changes: - wrap untrusted email in \<email\> XML block with prompt instructing the model to treat its contents as data only - add validate_ai_fallback() in template_parser that enforces the same contract keys, enums, and patterns as the template path before any write - call validate_ai_fallback() in handler() dispatch; emit an ai_fallback_rejected EMF metric on failure and skip the record - add 17 unit tests covering every gate rule and two end-to-end dispatch tests (injected email_type, injected status) Refs #101 * style: apply ruff formatting to fix CI check * harden ai_fallback gate: review fixes + security-review findings Review follow-up on the ai_fallback validation gate (PR #104), plus findings from a fan-out /sh-security-review of the change surface. Reviewer FIX items: - Neutralize forged <email> delimiters in the untrusted body before wrapping, so an in-body </email> cannot escape the data block. - Fail closed on non-dict model output instead of crashing the handler into async retries; count ai_fallback_rejected parses in the fallback-rate alarm and add a dedicated rejected-parse alarm so a gate-rejection drift outage is not silent. - Return a distinct invalid_status reason (was malformed_site_code); validate ISO-8601 dates; README + docstring updates. Security-review findings (detector fan-out + proof-or-kill verifier): - ReDoS (confirmed, medium): the tag neutralizer used two \s* around an optional /, backtracking quadratically on "<" + a long whitespace run (~32s at 100k chars -- one email could time out the Lambda). Collapse to a single [\s/]* class: linear, same defanging. - Unhashable-type crash (confirmed): a JSON list/dict for email_type or status made `x in <set>` raise TypeError, escaping the gate into retries. Guard with isinstance(str) before membership. - Unicode/newline regex (confirmed): _WO_ID_RE/_SITE_CODE_RE used ^..$ with \d, admitting fullwidth digits ("12345" as a lookalike partition key) and trailing newlines. Switch to \A[0-9]+\Z (and the handler's inline recheck to [0-9]) so neither passes. - Alarm comment (confirmed, low): corrected the "slow trickle still pages" wording -- rejections >~25-30 min apart page on neither alarm, the same knowingly-accepted residual as sender-auth-rejected. Refuted: residual free-text prompt injection is inherent to trusting allowlisted senders, not a new primitive; no DynamoDB key-poisoning bypass survives both gates ('#' can never enter work_order_id). 7 new regression tests. All 260 tests pass; ruff clean; cdk synth OK. --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-16 16:23:28 -04:00
# --- validate_ai_fallback unit tests -----------------------------------------
def _ai_candidate():
return {k: None for k in CONTRACT_KEYS}
def test_ai_fallback_baseline_is_valid():
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
ok, reason = validate_ai_fallback(cand)
assert ok and reason == "ok"
def test_ai_fallback_nondigit_wo_id():
cand = _ai_candidate()
cand["work_order_id"] = "12A45"
cand["email_type"] = "update"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "missing_required_field"
def test_ai_fallback_null_wo_id():
cand = _ai_candidate()
cand["work_order_id"] = None
cand["email_type"] = "update"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "missing_required_field"
def test_ai_fallback_hash_in_wo_id():
cand = _ai_candidate()
cand["work_order_id"] = "123#spoofed#deadbeef"
cand["email_type"] = "update"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "missing_required_field"
def test_ai_fallback_invalid_email_type():
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "exploit"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "missing_required_field"
def test_ai_fallback_null_email_type():
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = None
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "missing_required_field"
def test_ai_fallback_bad_status_enum():
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["status"] = "frobnicated"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "invalid_status"
def test_ai_fallback_unhashable_enum_fails_closed():
# A JSON list/dict for an enum field is unhashable; the gate must fail
# closed (isinstance guard), not raise TypeError into async retries.
for field, bad in (
("email_type", ["update"]),
("email_type", {"x": 1}),
("status", ["new"]),
("status", {"x": 1}),
):
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand[field] = bad
ok, reason = validate_ai_fallback(cand) # must not raise
assert not ok, f"{field}={bad!r} should fail closed"
def test_ai_fallback_fullwidth_digit_wo_id_rejected():
# Fullwidth digits render like ASCII but are a distinct partition key;
# [0-9] (not \d) must reject them.
cand = _ai_candidate()
cand["work_order_id"] = "12345" # "12345" fullwidth
cand["email_type"] = "update"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "missing_required_field"
def test_ai_fallback_trailing_newline_rejected():
# \A..\Z (not ^..$) must reject a trailing newline in wo_id and site_code.
cand = _ai_candidate()
cand["work_order_id"] = "12345\n"
cand["email_type"] = "update"
ok, _ = validate_ai_fallback(cand)
assert not ok
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["site_code"] = "WIL1\n"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "malformed_site_code"
def test_ai_fallback_non_dict_fails_closed():
# json.loads on model output can yield any JSON type; the gate must fail
# closed on a non-object rather than raise into async retries / DLQ.
for bad in ([], "string", 42, None, [{"work_order_id": "12345"}]):
ok, reason = validate_ai_fallback(bad)
assert not ok and reason == "not_an_object", f"{bad!r} should fail closed"
def test_ai_fallback_unparseable_sentinel_rejected():
# The template parser's internal _UNPARSEABLE sentinel must never survive
# the AI gate into the store.
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["comment_time"] = "__UNPARSEABLE__"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "creation_time_unparseable"
def test_ai_fallback_non_iso_date_rejected():
for key in ("date_reported", "scheduled_start", "due_date", "comment_time"):
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand[key] = "ignore previous instructions"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "creation_time_unparseable", f"{key} not gated"
def test_ai_fallback_iso_dates_accepted():
for value in ("2026-07-16", "2026-07-16T10:15:00", "2026-07-16T10:15:00Z", None):
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["date_reported"] = value
cand["comment_time"] = value
ok, reason = validate_ai_fallback(cand)
assert ok, f"date {value!r} should pass"
def test_ai_fallback_valid_status_ok():
for status in (
"new",
"assigned",
"in_progress",
"on_hold",
"completed",
"cancelled",
"unknown",
None,
):
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["status"] = status
ok, reason = validate_ai_fallback(cand)
assert ok, f"status={status} should pass"
def test_ai_fallback_bad_site_code():
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["site_code"] = "workshop"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "malformed_site_code"
def test_ai_fallback_valid_site_codes():
for code in ("WIL1", "ZDL8", "AB12", None):
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["site_code"] = code
ok, reason = validate_ai_fallback(cand)
assert ok, f"site_code={code} should pass"
def test_ai_fallback_key_set_mismatch_extra():
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["surprise"] = "x"
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "key_set_mismatch"
def test_ai_fallback_key_set_mismatch_missing():
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
del cand["address"]
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "key_set_mismatch"
def test_ai_fallback_all_valid_email_types():
for et in ("new_work_order", "update", "comment", "cancellation"):
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = et
ok, reason = validate_ai_fallback(cand)
assert ok, f"email_type={et} should pass"
def test_ai_fallback_free_text_float_rejected():
# Prompt-injected JSON float in a free-text slot must fail closed before
# update_item (boto3 rejects Python floats -> async retries / DLQ).
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["severity"] = 1.5
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "invalid_field_type"
def test_ai_fallback_free_text_dict_rejected():
# LLM-emitted maps in scalar slots must be rejected (DynamoDB Map pollution).
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["assigned_to"] = {"a": 1}
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "invalid_field_type"
def test_ai_fallback_free_text_list_rejected():
# LLM-emitted lists in scalar slots must be rejected (DynamoDB List pollution).
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand["comment_text"] = ["x"]
ok, reason = validate_ai_fallback(cand)
assert not ok and reason == "invalid_field_type"
def test_ai_fallback_free_text_none_and_str_ok():
# None and str remain valid for every free-text field.
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
ok, reason = validate_ai_fallback(cand)
assert ok and reason == "ok"
for field in template_parser._AI_FREE_TEXT_STR_FIELDS:
cand = _ai_candidate()
cand["work_order_id"] = "12345"
cand["email_type"] = "update"
cand[field] = "ok"
ok, reason = validate_ai_fallback(cand)
assert ok and reason == "ok", field