mirror of
https://github.com/Sea-Haven-Industries/procurement-ingest.git
synced 2026-09-30 04:53:12 +00:00
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8) tests/conftest.py only loads for the tests/ root, not a standalone `pytest lambdas/po/email_processor/tests` run, so it could never carry session invariants like the dummy AWS env or the moto stubber registration. Add a single repo-root conftest.py (pytest.ini pins rootdir there, so it loads for every invocation) that sets the dummy AWS credentials/region, imports moto BEFORE any handler module so boto3 sessions pick up its stubber hook (carrying the explanatory comment verbatim from the old _po_parser_support.py), and exposes one load_lambda_module(pipeline, name) — the sys.modules save/restore dance stays, since template_parser is still a duplicated bare name across pipelines needing per-exec sibling binding. Add tests/support/ as the shared package both pipelines' local _*_parser_support.py modules delegate to: a superset FakeTable (PO's update_item recording + WO's put_item and keyed single-row store), FakeDynamoResource, load_email, and load_golden with parse_float=Decimal kept (load-bearing for exact money comparison at PO magnitudes — WO's prior load_golden had no parse_float and must not regress PO by losing it). Rewrite _wo_parser_support.py off the bare `import handler` / `from handler import parse_raw_email` strategy that was the source of the bare-name sys.modules collision the other two loaders defend against. Move test_po_merge.py and test_pad_zip.py into lambdas/po/email_processor/tests/ (PO-specific, belongs beside the code) via git mv so history follows; test_parse_raw_email.py and test_ses_auth.py stay at the repo root since they're genuinely cross-pipeline, parameterized over both handlers. Delete tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only, and imports a handler at collection time, bypassing the loader gate entirely — the golden suites already cover its role. Its pytest.ini exclusion comment goes with it. New scenario coverage, all built on the single loader + support package: - PO+WO Bedrock transport errors (ThrottlingException, missing 'content' key, empty content list, non-JSON model text), asserting PO's pre-call ai_fallback metric survives with no partial write and the exception propagates; WO's no-datapoint-on-throttle behavior is pinned with a documenting test rather than "fixed" by reordering. - Handler-level SES-auth reject seam per pipeline: no auth monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls, zero writes, no raise — closing the hole where deleting the gate line today still passes every test. - web_ui coverage for both PO and WO (0% before this): fail-closed on unset ARN and on a Secrets Manager exception, TTL cache refresh, Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with no table scan, non-ASCII token, and a hostile-field-escaping regression lock. PO web_ui has no __init__.py, so these go through the loader rather than package imports. - A moto-backed mirror of test_po_merge for WO merge semantics (table 'WorkOrders'): null-status never clobbers wo_status, created_at immutable via if_not_exists, status->wo_status mapping, None fields absent from SET, record_type only-when-present. - Small pins: the PO-DC-02 64-char EMF clamp regression and per-pipeline multi-record failure-isolation (all-or-retry contract). The reprocess.py synthetic-event-shape contract test already landed in Phase 7, so it isn't duplicated here. Two WO product-code fixes ride along, since this is the phase that exercises them: (a) the invalid_status reason-code fix in template_parser.py's status check, which previously returned malformed_site_code for the same failure validate_ai_fallback already labels invalid_status, making one failure surface two codes depending on path (grepped the dashboards/metric filters for malformed_site_code first — no external references found, safe to diverge the two codes); (b) wrapping the WO Bedrock call in handler.py so a transport failure emits ai_fallback/bedrock_error in an except-and-reraise. This is deliberately not a naive reorder: the emit sits in the except block, not pre-call, so a gate-rejected email still emits only ai_fallback_rejected and wo_stack's "a rejected email emits nothing else" alarm contract doesn't double-count. A test computes the emitted series by hand to pin the no-double-count behavior. Neither change touches the handler event/return contract. _validate_new_po_values in the PO template_parser.py is split into per-rule helpers, and the V4 anchor-frame dataclass now carries summary_matches/price so V13 can consume them; extract_new_po (C901=35) is included in the split. Add ruff.toml enabling C901/PLR so the mccabe/complexity suppressions scattered through the tree stop being decorative; derived_fields.py is under the shadow-bake freeze so its violations are silenced via a per-file ignore with a justification comment instead of an in-file edit, and the handful of other pre-existing violations surfaced by turning the config on get the same per-file-ignore treatment with a reason, or a fix where the file isn't frozen. scripts/ is added to the CI lint scope. CI gains an explicit --cov module list (lambdas/po and wo email_processor + web_ui, po/site_extractor, lambdas/shared) plus --cov-fail-under=80, since web_ui and site_extractor lack __init__.py markers and a bare --cov=lambdas silently skips them for the missing package marker; .coveragerc omits the test dirs themselves from the count. The Phase 0 AST bundle-consistency test stays in the standard pytest run. .gitignore picks up the resulting .coverage data file. docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT declares quantity/price as "number or null", not JSON strings, so parse_float=Decimal already handles a conforming Bedrock response — the doc previously implied the coercion path was the primary mechanism rather than a defensive net for non-conforming responses. * test: lock attribute-context quote escaping in web_ui hostile-field test The escaping regression lock asserted only the element-context vector (raw <script> absent, <script> present) while its docstring claimed quotes were covered -- the payload's " and ' were never asserted on, so a quote-escaping regression on the onclick row-link sink (attribute breakout -> event-handler injection) would have passed green. /sh-security-review finding WC-01 (confirmed medium, test-integrity). Add assertions that the onclick sink's JSON string renders its opening quote as " (raw " after window.location= fails), that the payload's quote characters appear only entity-escaped, and that the raw payload never appears anywhere in the body. Mutation-verified: the test now fails when the sink's quote-escaping is dropped. * test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override - tests/test_web_ui_auth.py: replace the TypeError characterization pin with an xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False) behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is corrected, instead of requiring a passing test to be knowingly deleted. The module stays frozen this phase; the underlying hmac.compare_digest ASCII-only defect is tracked as a follow-up. - pytest.ini: document that the aggregate 80% floor (enforced in CI via the reusable workflow's bare pytest) red-exits local subset runs by design, with the --cov-fail-under=0 override for iteration. Floor stays in addopts because the centralized ci-python-sam workflow exposes no per-run test command.
140 lines
6.6 KiB
Python
140 lines
6.6 KiB
Python
"""The ONE Lambda-module loader for the whole procurement-ingest test suite.
|
|
|
|
Every test root (the repo-root ``tests/`` suite AND the two per-pipeline
|
|
``lambdas/*/email_processor/tests`` suites) loads Lambda modules through this
|
|
single ``load_lambda_module`` -- there is exactly one copy of the sys.modules
|
|
save/restore dance in the repo, and the repo-root ``conftest.py`` re-exports it.
|
|
|
|
``from conftest import ...`` is deliberately NOT used: three ``conftest.py``
|
|
files exist across the roots and pytest's per-directory sys.path prepending
|
|
makes the bare name ``conftest`` resolve nondeterministically -- exactly the
|
|
bare-name-collision class this phase eliminates. The loader lives here instead,
|
|
imported the same way from every invocation directory as ``tests.support``.
|
|
"""
|
|
|
|
import importlib.util
|
|
import sys
|
|
from pathlib import Path
|
|
|
|
REPO_ROOT = Path(__file__).resolve().parents[2]
|
|
_SHARED_DIR = REPO_ROOT / "lambdas" / "shared"
|
|
|
|
|
|
# Sibling modules imported by bare name from the handlers (the Lambda runtime
|
|
# puts each function's own directory on sys.path; the CDK bundling then cp's the
|
|
# shared modules in flat beside handler.py so those bare imports resolve too).
|
|
# template_parser/derived_fields are duplicated PER PIPELINE, so their bare
|
|
# names MUST be bound to the right pipeline's file around each handler exec --
|
|
# relying on sys.path ordering (or on whatever a previously collected suite left
|
|
# in sys.modules) silently binds a handler to the OTHER pipeline's sibling.
|
|
# ses_auth/email_parsing/emf/web_ui_auth are now single-sourced under
|
|
# lambdas/shared/ (Phase 3); the loop below resolves them from there via a
|
|
# shared-dir fallback.
|
|
#
|
|
# Phase 5 decomposed each God-handler into flat siblings (prompts/telemetry/
|
|
# extraction/enrichment/persistence). The order below is DEPENDENCY-TOPOLOGICAL,
|
|
# not alphabetical: the loader binds each bare name in sys.modules right after
|
|
# exec'ing it, so a sibling whose module body does `from <x> import ...` must
|
|
# appear AFTER <x> here or its exec ImportErrors. The load-bearing edges are
|
|
# emf < telemetry, prompts < extraction, and derived_fields + telemetry <
|
|
# enrichment. WO has no enrichment/derived_fields sibling -- the loader's
|
|
# `if not sibling_path.exists(): continue` silently skips them there, so one
|
|
# unified tuple serves both pipelines.
|
|
#
|
|
# web_ui_auth was added (no dependencies; after emf) so
|
|
# load_lambda_module("po"|"wo", "web_ui/handler") can bind the web_ui handlers'
|
|
# bare `from web_ui_auth import is_authenticated` (lambdas/po/web_ui/handler.py:15
|
|
# and the wo equivalent) via the same shared-dir fallback.
|
|
_SIBLING_MODULES = (
|
|
"ses_auth",
|
|
"email_parsing",
|
|
"emf",
|
|
"web_ui_auth",
|
|
"prompts",
|
|
"template_parser",
|
|
"derived_fields",
|
|
"telemetry",
|
|
"extraction",
|
|
"enrichment",
|
|
"persistence",
|
|
)
|
|
|
|
|
|
def _load_module(path, module_name):
|
|
if module_name in sys.modules:
|
|
return sys.modules[module_name]
|
|
spec = importlib.util.spec_from_file_location(module_name, path)
|
|
module = importlib.util.module_from_spec(spec)
|
|
sys.modules[module_name] = module
|
|
spec.loader.exec_module(module)
|
|
return module
|
|
|
|
|
|
def load_lambda_module(pipeline, name):
|
|
"""Load a Lambda module by file path under a unique, deterministic name.
|
|
|
|
``pipeline`` is one of ``{"po", "wo", "shared"}`` and ``name`` is the path
|
|
under ``lambdas/<pipeline>/`` without the ``.py`` suffix (e.g.
|
|
``"email_processor/handler"``, ``"web_ui/handler"``, or ``"ses_auth"`` for
|
|
shared). The module name is ``f"{pipeline}_{name.replace('/', '_')}"`` --
|
|
this reproduces the existing unique names byte-for-byte
|
|
(``po_email_processor_handler``, ``wo_email_processor_handler``,
|
|
``shared_ses_auth``), so every existing ``sys.modules`` sibling key
|
|
(``po_email_processor_handler__persistence`` etc.) is unchanged.
|
|
|
|
The handler files all share the basename ``handler.py`` and are not
|
|
importable as packages, so a plain ``import handler`` would collide across
|
|
Lambdas. The same loader serves leaf modules (``ses_auth.py``,
|
|
``web_ui_auth.py``), which are likewise loaded by file path.
|
|
|
|
Handler modules import their siblings by bare name (e.g. ``from
|
|
template_parser import try_deterministic_parse``). Each sibling is loaded
|
|
from the handler's own directory (falling back to ``lambdas/shared/``) under
|
|
a unique module name and registered under its bare name only for the
|
|
duration of the handler exec, then the previous binding is restored -- so
|
|
this loader is deterministic regardless of collection order and of what the
|
|
per-Lambda test suites (which put their own module dir on sys.path) have
|
|
already cached in sys.modules.
|
|
|
|
The save/restore dance does NOT shrink to nothing: template_parser (and
|
|
derived_fields/prompts/telemetry/extraction/enrichment/persistence) remain
|
|
duplicated bare names ACROSS pipelines. One pytest session execs BOTH
|
|
handlers; without per-exec bare-name binding + restore, whichever pipeline
|
|
loads second silently binds to the first pipeline's sibling. Only
|
|
ses_auth/email_parsing/emf/web_ui_auth are single-sourced.
|
|
"""
|
|
path = REPO_ROOT / "lambdas" / pipeline / f"{name}.py"
|
|
module_name = f"{pipeline}_{name.replace('/', '_')}"
|
|
if module_name in sys.modules:
|
|
return sys.modules[module_name]
|
|
# Keep the handler dir on sys.path for parity with the Lambda runtime.
|
|
handler_dir = str(path.parent)
|
|
if handler_dir not in sys.path:
|
|
sys.path.insert(0, handler_dir)
|
|
if path.name != "handler.py":
|
|
# Leaf modules (e.g. ses_auth.py / web_ui_auth.py) have no sibling
|
|
# imports of the per-pipeline kind the dance guards.
|
|
return _load_module(path, module_name)
|
|
saved = {}
|
|
for sibling in _SIBLING_MODULES:
|
|
# Per-pipeline siblings (template_parser/derived_fields/...) resolve next
|
|
# to the handler; the shared, single-sourced siblings (ses_auth/
|
|
# email_parsing/emf/web_ui_auth) fall back to lambdas/shared/. No
|
|
# ambiguity: post Phase 3 the shared names exist ONLY under shared/, the
|
|
# per-pipeline names ONLY next to the handler.
|
|
sibling_path = path.parent / f"{sibling}.py"
|
|
if not sibling_path.exists():
|
|
sibling_path = _SHARED_DIR / f"{sibling}.py"
|
|
if not sibling_path.exists():
|
|
continue
|
|
saved[sibling] = sys.modules.get(sibling)
|
|
sys.modules[sibling] = _load_module(sibling_path, f"{module_name}__{sibling}")
|
|
try:
|
|
module = _load_module(path, module_name)
|
|
finally:
|
|
for sibling, previous in saved.items():
|
|
if previous is not None:
|
|
sys.modules[sibling] = previous
|
|
else:
|
|
sys.modules.pop(sibling, None)
|
|
return module
|