Commit graph

6 commits

Author SHA1 Message Date
Adam Moussa
c040050373
feat(webhook): SHOC WO webhook emitter - dark-ship streams + HMAC secret/rotation (PR-2) (#137)
* docs(webhook): revise SHOC webhook contract and plan for post-migration reality

Branch re-cut on main 2026-07-23 (old base carried stale PR #99 commits).

Contract Rev 2026-07-23:
- Producer account corrected: seahaven-prod (011934824531); mgmt frozen
- Reconciliation backstop is the new procurement read API, not SyncController
- wo_status "unknown" is real; SHOC must map it (checklist item added)
- write_origin forward-compat note for phase-2 write-back echo suppression
- SyncVendorReplies retirement flagged (dead table, no vendor_reply event)

Plan updates:
- Account gate: seahaven-prod only; never enable streams on mgmt tables
- Emitter ships DARK (ESMs enabled=False); activation is a deliberate flip
  after the SHOC receiver passes shared HMAC vectors
- Post-refactor conventions: common.py helpers, bundle-consistency AST pins,
  pytest.ini --cov additions, consolidated test roots
- Dedicated-CMK rationale, secret-ARN handooff step, consumer audit refreshed
  (slack-bot decommissioned), enum golden test, write_origin skip-branch test

* feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation

Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC
call-and-be-called effort). Everything ships DARK: both DynamoDB event
source mappings deploy enabled=False; activation is a deliberate
one-line follow-up PR gated on the SHOC receiver passing the shared
HMAC test vectors.

- Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments
  (in-place update, RETAIN + logical IDs untouched; no existing
  consumers — verified live, neither table had a stream).
- workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope ->
  HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard
  ordering (parallelization 1, bisect off, retry until 24h age,
  ReportBatchItemFailures); 429/5xx/timeout block the shard in order,
  other 4xx park to workorder-shoc-emitter-rejected; ESM failures ->
  workorder-shoc-emitter-failures (metadata; replay rebuilds from
  DynamoDB). Echo guard skips write_origin=shoc-write-api.
- Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK
  (alias workorder-ingest-shoc-webhook-kms); cross-account
  GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly
  arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy
  DESTROY deliberately (machine-generated material; avoids the
  fixed-name RETAIN-orphan deadlock).
- workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap,
  64-hex keys, kid = UTC %Y-%m-%dT%H.
- Alarms (ALARM-only -> site-alerts): emitter errors/throttles/
  duration + iterator-age (>=10 min) + failures/rejected queue
  depth; rotator standard trio.
- scripts/replay_shoc_webhooks.py: dry-run-default operator replay
  (rebuilds from tables, replay:true envelopes).
- Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared
  with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay
  signing pinned to identical vectors); bundle-consistency AST pins
  for both new bundles.
- README: WO stack + webhook feed section, alarm table, runbooks;
  removed stale seahaven-slack-bot consumer references.

* fix(webhook): kms:ViaService pins, https-only delivery, cross-account principal CI pin

GPT-4.1 cross-family review of the policy surface (no BLOCK): FIX applied
to the cross-account shoc-backend-dev Decrypt statement and both Lambda
role KMS grants (the key is only ever used via Secrets Manager); its
invariant-enforcement QUESTION answered durably with
tests/test_cross_account_principal_pin.py (any new foreign IAM principal
in cdk/ fails CI). Scanner mediums fixed: delivery.py and the replay
script now refuse non-https URLs (urllib follows file:// and http://).
SQS metadata-action and dynamodb:ListStreams NITs skipped: standard CDK
grant shapes; ListStreams has no resource-level scoping. The 4 gitleaks
HIGHs on docs/shoc-webhook-test-vectors.json are deliberate non-secrets
(shared receiver-verification vectors) suppressed machine-level with
justification.

* harden(webhook): resolve /sh-security-review findings (1 confirmed medium + cheap fixes)

High-recall detector fan-out (injection/authz/secrets-crypto/iac-iam/logic)
+ proof-or-kill verifier. Gate PASSES: 1 confirmed medium, 0 confirmed
critical/high. Confirmed finding fixed; several unverified-but-cheap
hardenings applied since the emitter ships dark and activation is weeks out.

- CONFIRMED medium (confused deputy): the rotation Lambda's generated
  invoke permission for secretsmanager.amazonaws.com carried no
  SourceAccount/SourceArn, so any account's Secrets Manager could invoke
  the rotator. Patched the generated CfnPermission in place (a second
  permission would be additive, not restrictive) to pin account + this
  secret ARN.
- delivery + replay: refuse to follow receiver 3xx redirects (no-redirect
  opener) so live X-SH-* auth headers can't be forwarded to a
  receiver-chosen Location and an http:// Location can't slip past the
  https guard. Fixed the "unfollowed 3xx" comment that was factually wrong.
- delivery: classify 401/403 as retryable (invalidate key cache + retry in
  order) instead of parking -- transient auth failures (rotation outran the
  TTL cache, clock skew) are availability events, not contract bugs.
- envelope: build_event now genuinely total (guarded eventID /
  ApproximateCreationDateTime subscripts) per its own never-raise contract.
- handler: catch-all so an unexpected per-record error (e.g. SQS park
  failure) reports only that record instead of failing the whole batch
  (which would re-deliver every earlier success for 24h); per-invocation
  emit/skip batch summary so a systemic silent drop is queryable/alarmable.
- rotator: narrow the AWSCURRENT-read except to ResourceNotFound/JSONDecode
  (transient SM/KMS errors re-raise so the overlap key isn't silently
  dropped); kid uniqueness checked against ALL retained kids with a random
  suffix on collision (never reissue a kid for a different secret).
- contract: skeleton-upsert required on ANY unknown work_order_id (not just
  comment-before-create) + monotonicity guard (ignore older updated_at), so
  a parked created or an out-of-order replay can't corrupt receiver state.

Unverified/refuted findings left as-is with rationale: the two "high" logic
claims (whole-batch crash triggers, ordering violation) were refuted on
reachability (real stream records carry required fields; persistence writes
strings only; full-state idempotent upsert absorbs the ordering gap). Signed
kid/version binding (AUTHZ-002) declined: coordinated contract change, not
cheap, no exploit with one algorithm/key.

* fix(webhook): drop kid from rotator test_ok log (CodeQL clear-text-logging FP)

GHAS CodeQL flagged py/clear-text-logging-sensitive-data (high) at
_test_secret's success log because head["kid"] is subscripted from the
same parsed-secret dict that holds head["secret"] — the taint tracker
can't tell the non-secret key id from the secret. The secret value is
never logged. Rather than dismiss the alert (fragile; re-alerts on line
moves), remove the flow: kid is already logged at stage time in
_create_secret and version_id correlates the steps, so the test_ok log
keeps only event + version_id. Also hardens against a future edit that
swaps the logged field.
2026-07-24 22:12:20 +00:00
Adam Moussa
f67d8b9907
feat(api): procurement-api read stack + OpenAPI docs (SHOC reconciliation path) (#127)
Some checks are pending
Deploy / deploy (push) Waiting to run
* feat(api): add procurement-api stack - read API + OpenAPI docs page

Third CDK stack: API Gateway REST API (IAM SigV4) over both pipelines'
tables, replacing SHOC's retired SyncController cross-account DynamoDB
scan as the reconciliation/backfill path.

- lambdas/api/: handler (healthcheck + docs-token gate + router dispatch),
  router (single route table), pagination (opaque cursor, hostile -> 400),
  Decimal-safe serialization, wo_repo/po_repo reads. No VendorReplies.
- OpenAPI 3.1 spec as source of truth incl. top-level webhooks section
  documenting the outbound SHOC feed; phase-2 write endpoints x-planned
  (router answers 501). Self-contained /docs page, no CDN.
- Auth: AWS_IAM on data routes + resource policy scoped to exactly
  arn:aws:iam::396287094661:role/shoc-backend-dev on GET/*; /docs and
  /openapi.json carve-out is token-gated in the Lambda via shared
  web_ui_auth (fail-closed, INFRA-74 posture).
- KMS: explicit Decrypt/DescribeKey on the DynamoDB CMK from SSM
  (name-imported table drops the key association - INFRA-104 class).
- Alarms: errors/throttles/duration(p99>=22.5s) + gateway 5xx, ALARM-only
  to site-alerts. No access logging in v1 (docs ?token= shim stays out of
  logs); cloud_watch_role=False.
- Tests: handler auth-seam + routing + Decimal round-trip; moto cursor
  pagination incl. hostile cursors; spec<->router drift gate; bundle
  AST pins for the api command; pytest.ini --cov + loader siblings.
- Deploy role: third stack DescribeStacks ARN + procurement-api smoke
  invoke ARN (re-run create-deploy-role.sh before merge).

* harden(api): apply sh-security-review findings to procurement-api

Fan-out (6 detectors) + review findings resolved:

Correctness / DoS:
- pagination: require EXACT key-set match (was subset) so a partial/foreign
  composite cursor can't reach DynamoDB as an inconsistent ExclusiveStartKey
  -> ValidationException -> 500; comments Query now pins the cursor's
  work_order_id to the path entity.
- handler: map botocore ValidationException to 400 (defense in depth) so a
  crafted cursor can't drive the zero-threshold 5xx alarm.
- web_ui_auth: compare tokens as bytes; a non-ASCII presented token now fails
  closed (401) instead of crashing hmac.compare_digest into a 500. Resolves the
  pre-existing xfail(strict) follow-up test; hardens the web UIs too.

Docs page:
- typeStr() now escapes the one spec-derived string that reached innerHTML.
- spec inlined into the docs <script> block escapes "<" -> < (</script>
  breakout guard); /openapi.json still served byte-faithful.
- Cache-Control: no-store + Referrer-Policy: no-referrer on docs responses so
  the ?token= URL stays out of caches/Referer.
- spec-drift test asserts the committed spec carries no "</" / "<!--".

IAM / IaC:
- resource policy enumerates the 7 data GET resources instead of GET/* so a
  future GET route can't silently inherit SHOC cross-account reach.
- kms:Decrypt grant gains a kms:ViaService=dynamodb condition.
- stage throttling (50 rps / 100 burst) bounds the unauthenticated /docs blast
  radius below the 10k account default.
- corrected the PATCH/POST comment (same-account callers aren't blocked by the
  resource policy; 501 handler + absent write grant are the gate).
- documented the RETAIN log-group first-deploy rollback trap and the
  resource-policy-needs-redeploy gotcha in-stack.

Mandatory GPT-4.1 cross-family review of the full policy surface: no BLOCK/FIX.
675 tests pass, ruff clean, cdk synth green.
2026-07-23 19:32:20 -04:00
Adam Moussa
ff368beaa4
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)

tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.

Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.

Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.

New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
  'content' key, empty content list, non-JSON model text), asserting
  PO's pre-call ai_fallback metric survives with no partial write and
  the exception propagates; WO's no-datapoint-on-throttle behavior is
  pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
  monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
  zero writes, no raise — closing the hole where deleting the gate
  line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
  unset ARN and on a Secrets Manager exception, TTL cache refresh,
  Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
  no table scan, non-ASCII token, and a hostile-field-escaping
  regression lock. PO web_ui has no __init__.py, so these go through
  the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
  (table 'WorkOrders'): null-status never clobbers wo_status,
  created_at immutable via if_not_exists, status->wo_status mapping,
  None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
  per-pipeline multi-record failure-isolation (all-or-retry contract).
  The reprocess.py synthetic-event-shape contract test already landed
  in Phase 7, so it isn't duplicated here.

Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.

_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.

CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.

docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.

* test: lock attribute-context quote escaping in web_ui hostile-field test

The escaping regression lock asserted only the element-context vector
(raw <script> absent, &lt;script&gt; present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).

Add assertions that the onclick sink's JSON string renders its opening
quote as &quot; (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.

* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override

- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
  xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
  behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
  corrected, instead of requiring a passing test to be knowingly deleted. The
  module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
  defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
  reusable workflow's bare pytest) red-exits local subset runs by design, with the
  --cov-fail-under=0 override for iteration. Floor stays in addopts because the
  centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
Adam Moussa
dc17de4026
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
Some checks are pending
Deploy / deploy (push) Waiting to run
* Add deterministic template parser for WO emails

The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.

The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order:  <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.

Refs: #23

* Migrate WO processor to Bedrock and fix comment_id collision

Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.

Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).

Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.

Refs: #23

* Migrate PO processor to Bedrock

Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.

* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm

Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.

Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.

Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.

* Add offline WO parser test suite

Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.

Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.

Refs: #23

* Document Bedrock migration and WO parse flow in README

Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.

Refs: #23

* Fix f-string lint and formatting in backfill scripts

Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.

* Emit ParseMethod-only EMF set so fallback alarm can fire

The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.

Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.

* Commit WO parser .eml fixtures for executable coverage

The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.

Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.

* feature: Add PO template parser scaffold and design doc

Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.

Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.

Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>

* Implement PO new_po extraction and value-level gate

Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.

The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.

Refs: #99

* Wire template-first parse into PO handler with EMF metric

Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.

Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.

Refs: #99

* Add PO fallback-rate alarm retuned for ~57 emails/day

The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.

Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).

Refs: #99, #102

* Add offline PO parser suite with scrubbed fixture corpus

132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.

The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).

Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.

Refs: #99

* Document PO template-first parser and retuned alarm

README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.

Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.

Refs: #99

* Record cross-family review outcome for handler wiring

GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).

Refs: #99

* Pin line-item currency to USD in the PO gate

The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.

* Scrub residual transport tokens from PO fixtures

The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.

* Converge quantity/price to Decimal on both paths

EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.

The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.

* Coerce bare-int quantity/price to Decimal in enrich_parsed

GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.

* Scrub fixture-body PII and harden cancellation gate (sec review)

/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.

F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.

F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).

Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.

401 tests pass; ruff/format clean; cdk synth po-ingest clean.

---------

Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
Adam Moussa
acc1961d21
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails

The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.

The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order:  <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.

Refs: #23

* Migrate WO processor to Bedrock and fix comment_id collision

Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.

Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).

Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.

Refs: #23

* Migrate PO processor to Bedrock

Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.

* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm

Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.

Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.

Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.

* Add offline WO parser test suite

Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.

Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.

Refs: #23

* Document Bedrock migration and WO parse flow in README

Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.

Refs: #23

* Fix f-string lint and formatting in backfill scripts

Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.

* Emit ParseMethod-only EMF set so fallback alarm can fire

The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.

Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.

* Commit WO parser .eml fixtures for executable coverage

The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.

Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.

* Fix WO parser advisories A1-A3 (PR #99 follow-ups)

A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.

A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.

A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.

Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
  header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
  path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
  used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
  range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
Adam Moussa
941a0ec0fe
Add pytest suite and enable tests in CI (#95)
Some checks are pending
Deploy / deploy (push) Waiting to run
The repo had no automated tests, so regressions in the email
parsing pipeline could only be caught manually. The reusable
ci-python-sam workflow already supports a run-tests input at the
pinned SHA; enable it and add a first suite covering the pure
functions pad_zip and parse_raw_email.

tests/conftest.py sets AWS_DEFAULT_REGION and dummy credentials
before any handler import because the handlers create boto3
clients at module import time, and loads each handler.py via
importlib under a unique module name since the files share a
basename. pytest.ini scopes discovery to tests/ so the manual
test_local.py script at the repo root is not collected.
2026-07-15 19:06:20 -04:00