2026-05-12 15:21:06 -04:00
|
|
|
|
"""
|
|
|
|
|
|
Email processor Lambda.
|
|
|
|
|
|
|
|
|
|
|
|
Triggered by S3 events when SES delivers an email.
|
|
|
|
|
|
Parses the raw email, sends it to Claude for structured extraction,
|
|
|
|
|
|
then writes the result to DynamoDB.
|
feat: decompose email-processor handlers into flat siblings + lazy boto3 clients (refactor phase 5) (#113)
Both email-processor God-handlers split along the seams that already
work in the flat-sibling pattern established by lambdas/shared/, so
bare-name imports keep working under the existing bundling glob.
PO (5-way split): handler.py keeps only the event loop, fail-closed
auth, and email_type routing. extraction.py holds extract_with_claude
and _EMAIL_TAG_RE, importing EXTRACTION_PROMPT from prompts.py and
parse_raw_email from shared/email_parsing.py rather than recreating a
PO-local copy. enrichment.py is a pure code move of enrich_parsed and
pad_zip (PO-only; WO has no enrichment stage) with zero behavior
change. telemetry.py holds the EMF ParseMethod emit wrappers.
persistence.py holds _write_fields/_merge_update/save_*, collapsing
the byte-identical save_new_po/save_revision bodies into one
_save_merge helper that both now call through, preserving the sticky
Cancelled ConditionExpression guard for both callers; save_cancellation
stays separate.
WO (5 concerns, no enrichment stage): the handler loop keeps
validate_ai_fallback and the re.fullmatch(r"[0-9]+", work_order_id)
key guard ahead of both save_work_order and save_event, since the
guard protects the DynamoDB partition key and the '#'-delimited
comment_id range-key segment. _header_date_iso and comment_id
determinism stay colocated with persistence.py's save_event for the
retry-idempotent event_id key.
EXTRACTION_PROMPT (PO) moves to prompts.py with cross-reference
headers to derived_fields.py's authoritative trade/site/fiscal rule
tables; handler.py re-exports it (from prompts import
EXTRACTION_PROMPT) since four tests dereference handler.EXTRACTION_
PROMPT directly. WO's prompt moves the same way.
I/O modules (extraction.py's bedrock client, persistence.py's
dynamodb resource, handler.py's s3 client) get lazy cached boto3
accessors; pure modules (enrichment.py, prompts.py, telemetry.py)
import no boto3. Test monkeypatch surfaces move to the module that
now owns the client (e.g. persistence.dynamodb) everywhere tests
patch it, and the moto-before-handler-import ordering in
_po_parser_support.py is preserved so the moto-backed suites don't
hit real AWS.
Behavior-preservation pins, verified with tests: PO still emits
ParseMethod=ai_fallback before the Bedrock call, with
ai_fallback_rejected as the additive second datapoint on rejection.
WO still emits after its gate with mutually-exclusive ai_fallback /
ai_fallback_rejected. Shadow DerivedFieldAgreement telemetry stays
ai_fallback-only. derived_fields.py is untouched (diff against
feature/phase-3-shared-extraction is empty). handler(event, context)
signatures and the save_* public contract are unchanged on both
pipelines; goldens unchanged.
PO_EXPECTED_TOP_LEVEL_MODULES and its WO equivalent in
tests/test_bundle_consistency.py are updated for the new sibling
modules so the AST bundle-consistency test still fails on an
unshipped or uncommented-out sibling.
2026-07-20 15:34:53 -04:00
|
|
|
|
|
|
|
|
|
|
Phase 5: this handler is the thin event loop + fail-closed auth + routing. The
|
|
|
|
|
|
work has moved to flat sibling modules (bare-name imports resolve via the same
|
|
|
|
|
|
flat-landing bundling as ses_auth/template_parser):
|
|
|
|
|
|
extraction.py -- extract_with_bedrock + the <email>-tag neutralizer
|
|
|
|
|
|
telemetry.py -- emit_parse_metric (stdout EMF)
|
|
|
|
|
|
persistence.py -- save_work_order / save_event / comment_id determinism
|
|
|
|
|
|
prompts.py -- EXTRACTION_PROMPT (re-exported below for tests)
|
2026-05-12 15:21:06 -04:00
|
|
|
|
"""
|
|
|
|
|
|
|
|
|
|
|
|
import logging
|
|
|
|
|
|
import re
|
|
|
|
|
|
|
|
|
|
|
|
import boto3
|
feat: extract lambdas/shared/ — single-source ses_auth, web_ui auth, email parsing, EMF emitter (refactor phase 3) (#111)
Four modules move into the handbook-mandated lambdas/shared/ location,
collapsing duplicated logic that had to be kept in sync by hand across
the PO and WO pipelines:
- ses_auth.py: the PO and WO copies were verified sha256-identical
against the feature/phase-7-ops-recovery baseline before the move
(no drift since the last audit). shared/ses_auth.py is the exact
bytes of that one copy; both originals are git rm'd (the PO copy
via rename, the WO copy as a straight delete). Bundling lands the
module flat in /asset-output for both email processors, so the
handlers keep `from ses_auth import authenticate_inbound_email`
unchanged — zero handler diff for this move, which is what keeps
fail-closed auth byte-identical through the change.
- web_ui_auth.py: extracts the byte-identical _get_auth_token /
_header / is_authenticated block plus the four token-cache globals
out of both web_ui handlers. The per-stack INFRA-74 comments stay
in each handler as-is (deliberately drifted wording, stack-specific)
rather than being unified into the shared module. Fail-closed
semantics (unset ARN or Secrets Manager exception -> deny) are
unchanged.
- email_parsing.py: parse_raw_email ships as the superset version that
returns cc unconditionally. WO's output is bit-identical to before;
PO simply ignores the cc field rather than being "cleaned up" to
consume it. No second variant is kept.
- emf.py: a generic emitter parameterized by namespace, dimension
sets, and properties. Every call site's emitted EMF envelope is
unchanged, including the load-bearing
[["ParseMethod"],["ParseMethod","TemplateId"]] dimension-set shape
the alarms and metric filters depend on. Emission ordering is
untouched: PO still emits ai_fallback before the Bedrock call, WO
still emits its mutually-exclusive ai_fallback/ai_fallback_rejected
after its gate. The deliberate-double-count comments survive.
_emit_derived_agreement_metric was found living inside
derived_fields.py, so per the DERIVED-FIELDS exception it is left
as a third, unconverted copy (derived_fields.py and the shadow
DerivedFieldAgreement telemetry stay untouchable while that bake
runs) — a comment there points at shared/emf.py for the eventual
follow-up.
Bundling: both email-processor cdk bundling commands gain a trailing
`cp shared/*.py /asset-output/` (they were already cp-only post-Phase
7, so no pip step or manylinux pin is reintroduced). Both web_ui
functions gain the same widened-root staging so web_ui_auth.py ships
beside their handler; site_extractor's from_asset is untouched.
Tests: PO_EXPECTED_TOP_LEVEL_MODULES gains the shared modules that now
ship, the AST sibling-import check resolves imports whose source now
lives under shared/, and the new shared cp line has its own
revert/mutation detection. _SIBLING_MODULES resolution and
_po_parser_support.py now load ses_auth/email_parsing/emf from
shared/; the two-copy ses_auth byte-identity fixture-hygiene test is
retired as obsolete now that there is one copy, and the ses_auth
fixture parameterization over two identical copies is dropped. The
sys.modules save/restore dance for template_parser (still duplicated
per-pipeline) is left in place.
2026-07-20 13:38:23 -04:00
|
|
|
|
from email_parsing import parse_raw_email
|
feat: decompose email-processor handlers into flat siblings + lazy boto3 clients (refactor phase 5) (#113)
Both email-processor God-handlers split along the seams that already
work in the flat-sibling pattern established by lambdas/shared/, so
bare-name imports keep working under the existing bundling glob.
PO (5-way split): handler.py keeps only the event loop, fail-closed
auth, and email_type routing. extraction.py holds extract_with_claude
and _EMAIL_TAG_RE, importing EXTRACTION_PROMPT from prompts.py and
parse_raw_email from shared/email_parsing.py rather than recreating a
PO-local copy. enrichment.py is a pure code move of enrich_parsed and
pad_zip (PO-only; WO has no enrichment stage) with zero behavior
change. telemetry.py holds the EMF ParseMethod emit wrappers.
persistence.py holds _write_fields/_merge_update/save_*, collapsing
the byte-identical save_new_po/save_revision bodies into one
_save_merge helper that both now call through, preserving the sticky
Cancelled ConditionExpression guard for both callers; save_cancellation
stays separate.
WO (5 concerns, no enrichment stage): the handler loop keeps
validate_ai_fallback and the re.fullmatch(r"[0-9]+", work_order_id)
key guard ahead of both save_work_order and save_event, since the
guard protects the DynamoDB partition key and the '#'-delimited
comment_id range-key segment. _header_date_iso and comment_id
determinism stay colocated with persistence.py's save_event for the
retry-idempotent event_id key.
EXTRACTION_PROMPT (PO) moves to prompts.py with cross-reference
headers to derived_fields.py's authoritative trade/site/fiscal rule
tables; handler.py re-exports it (from prompts import
EXTRACTION_PROMPT) since four tests dereference handler.EXTRACTION_
PROMPT directly. WO's prompt moves the same way.
I/O modules (extraction.py's bedrock client, persistence.py's
dynamodb resource, handler.py's s3 client) get lazy cached boto3
accessors; pure modules (enrichment.py, prompts.py, telemetry.py)
import no boto3. Test monkeypatch surfaces move to the module that
now owns the client (e.g. persistence.dynamodb) everywhere tests
patch it, and the moto-before-handler-import ordering in
_po_parser_support.py is preserved so the moto-backed suites don't
hit real AWS.
Behavior-preservation pins, verified with tests: PO still emits
ParseMethod=ai_fallback before the Bedrock call, with
ai_fallback_rejected as the additive second datapoint on rejection.
WO still emits after its gate with mutually-exclusive ai_fallback /
ai_fallback_rejected. Shadow DerivedFieldAgreement telemetry stays
ai_fallback-only. derived_fields.py is untouched (diff against
feature/phase-3-shared-extraction is empty). handler(event, context)
signatures and the save_* public contract are unchanged on both
pipelines; goldens unchanged.
PO_EXPECTED_TOP_LEVEL_MODULES and its WO equivalent in
tests/test_bundle_consistency.py are updated for the new sibling
modules so the AST bundle-consistency test still fails on an
unshipped or uncommented-out sibling.
2026-07-20 15:34:53 -04:00
|
|
|
|
from extraction import extract_with_bedrock
|
|
|
|
|
|
from persistence import save_event, save_work_order
|
|
|
|
|
|
from prompts import EXTRACTION_PROMPT # noqa: F401 (re-export for tests)
|
Add fail-closed SES sender authentication (INFRA-107) (#98)
* Add fail-closed SES sender authentication
The From header and any raw-MIME Authentication-Results copies are
attacker-forgeable, so a forged email to apm@int.seahaven.com or
amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107,
CRITICAL). Both S3-triggered email processors now authenticate the
sender against the Authentication-Results header SES itself prepends
at delivery: only the topmost header is consulted, its authserv-id
must be amazonses.com, and it must carry dkim=pass for a domain in
the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set
in CDK so ops can adjust without code changes).
Allowlists come from live traffic observed 2026-07-15 on both ingest
buckets: WO mail arrives via the apm@ Google Groups forward, which
re-signs as seahaven.com (the hxgnsmartcloud.com signature does not
survive the forward); PO mail passes for amazon.coupahost.com.
amazonses.com also passes on PO mail but is deliberately excluded --
every SES customer's outbound mail passes for it.
Every failure path (env var unset, header missing or unparseable,
verdict fail, unaligned domain) rejects the email: a structured
warning with the reason and S3 key is logged and the record skipped
without erroring the invocation, so rejected mail causes no Lambda
retries or DLQ messages. Handler signatures and event sources are
unchanged.
Refs: INFRA-107
* Harden AR parser per cross-family review
Cross-family (GPT-4.1) review findings: terminate the dkim result
token at end-of-clause, whitespace, or a comment so a value like
"dkim=pass-fake" can never be read as a pass; normalize trailing
dots off allowlist entries so "seahaven.com." matches; make the
compat32 parser policy explicit. Adds tests for result-token
boundaries, comments after the result, quoted domain values, and
folding inside a dkim clause.
Refs: INFRA-107
* Harden AR parsing and alarm on sender-auth rejects
The SES-stamped Authentication-Results value echoes attacker-controlled
SMTP-session tokens (envelope-from, helo, header.from) as their own
semicolon-delimited property clauses. A naive split(";") tore an RFC 5321
quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass
clause, so a fully spoofed email was accepted on the genuinely
SES-stamped topmost header. Tokenise comment- and quoted-string-aware
(RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on
semicolons outside a quoted-string, and fail closed on unbalanced
quotes/comments so a ';' inside a quoted pvalue can never start a clause.
Rejected mail returns normally (no error, no retry, no DLQ message), so a
signing-domain drift or a wrong allowlist would silently discard 100% of
legitimate mail while every alarm stayed green. Add a CloudWatch Logs
metric filter + alarm on the sender_auth_rejected warning to both stacks
so a false-reject storm pages instead of vanishing. This is also the
safety net for the WO seahaven.com allowlist assumption, which must be
validated against a live SES-stamped header (a plain Gmail auto-forward
re-signs under the sending Workspace domain, not seahaven.com).
Refs: INFRA-107
* chore: retrigger CI (no run recorded for 7c74ac1)
* Fix quoted-AUID DKIM domain spoof in sender auth
Resolve three confirmed /sh-security-review findings on the fail-closed
SES sender-authentication control.
HIGH: header.i/header.d domain extraction was not quoted-string aware.
An attacker with a valid DKIM key for their own domain could set an
RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive
extractor stopped at the closing quote and returned seahaven.com,
accepting forged mail. Extraction now tokenises the clause with the same
quoted-string discipline already used for clause splitting: header.d
(the plain signing domain) is authoritative when present, otherwise the
header.i domain is the part after the AUID's LAST top-level "@", so a "@"
inside a quoted local-part is treated as signer-controlled label text and
yields the true signer (attacker.com), not seahaven.com.
LOW: the topmost-header parse ran outside evaluate_sender_authentication's
try/except, so an unexpected parser exception on crafted input could
propagate into the handler and Lambda async retries/DLQ. The parse now
fails CLOSED with an authentication_results_unparseable reason.
MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to
a low-volume total-reject outage (a trickle that never sums to 3). Both
stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3
/ datapoints_to_alarm=2, so a sustained reject condition pages even at one
reject per period while a lone stray probe self-clears.
Refs: INFRA-107
* Load Lambda function dir on sys.path in tests
Rebasing INFRA-107 onto main folded #95's pytest suite into this
branch's tests. The unified conftest loads the PO/WO handlers by file
path, and handler.py now does `from ses_auth import
authenticate_inbound_email` -- a bare sibling import that resolves in
the Lambda only because the runtime puts each function's own directory
on sys.path. The shared load_handler now adds that directory so the
handler tests import correctly alongside the sender-auth tests.
Refs: INFRA-107
* Note #97 test files in README directory tree
The rebase onto main brought in #97's tests/requirements.txt and
tests/test_po_merge.py. List both in the directory tree so it matches
the tree on disk.
Refs: INFRA-107
* Document INFRA-107 forwarder-binding risk acceptance
Record the accepted risk that WO sender auth binds to the apm@ forward's
re-signing domain (seahaven.com) rather than the Hexagon originator; the
apm@ Google Group's restricted posting policy is the load-bearing control
(escalates to HIGH if the group is opened to external posting). Also
correct the sender-auth-rejected alarm docs to match the shipped config
(>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and
note the SES-AR-01/02 parser hardening follow-ups.
Refs: INFRA-107
2026-07-15 20:58:47 -04:00
|
|
|
|
from ses_auth import authenticate_inbound_email
|
feat: decompose email-processor handlers into flat siblings + lazy boto3 clients (refactor phase 5) (#113)
Both email-processor God-handlers split along the seams that already
work in the flat-sibling pattern established by lambdas/shared/, so
bare-name imports keep working under the existing bundling glob.
PO (5-way split): handler.py keeps only the event loop, fail-closed
auth, and email_type routing. extraction.py holds extract_with_claude
and _EMAIL_TAG_RE, importing EXTRACTION_PROMPT from prompts.py and
parse_raw_email from shared/email_parsing.py rather than recreating a
PO-local copy. enrichment.py is a pure code move of enrich_parsed and
pad_zip (PO-only; WO has no enrichment stage) with zero behavior
change. telemetry.py holds the EMF ParseMethod emit wrappers.
persistence.py holds _write_fields/_merge_update/save_*, collapsing
the byte-identical save_new_po/save_revision bodies into one
_save_merge helper that both now call through, preserving the sticky
Cancelled ConditionExpression guard for both callers; save_cancellation
stays separate.
WO (5 concerns, no enrichment stage): the handler loop keeps
validate_ai_fallback and the re.fullmatch(r"[0-9]+", work_order_id)
key guard ahead of both save_work_order and save_event, since the
guard protects the DynamoDB partition key and the '#'-delimited
comment_id range-key segment. _header_date_iso and comment_id
determinism stay colocated with persistence.py's save_event for the
retry-idempotent event_id key.
EXTRACTION_PROMPT (PO) moves to prompts.py with cross-reference
headers to derived_fields.py's authoritative trade/site/fiscal rule
tables; handler.py re-exports it (from prompts import
EXTRACTION_PROMPT) since four tests dereference handler.EXTRACTION_
PROMPT directly. WO's prompt moves the same way.
I/O modules (extraction.py's bedrock client, persistence.py's
dynamodb resource, handler.py's s3 client) get lazy cached boto3
accessors; pure modules (enrichment.py, prompts.py, telemetry.py)
import no boto3. Test monkeypatch surfaces move to the module that
now owns the client (e.g. persistence.dynamodb) everywhere tests
patch it, and the moto-before-handler-import ordering in
_po_parser_support.py is preserved so the moto-backed suites don't
hit real AWS.
Behavior-preservation pins, verified with tests: PO still emits
ParseMethod=ai_fallback before the Bedrock call, with
ai_fallback_rejected as the additive second datapoint on rejection.
WO still emits after its gate with mutually-exclusive ai_fallback /
ai_fallback_rejected. Shadow DerivedFieldAgreement telemetry stays
ai_fallback-only. derived_fields.py is untouched (diff against
feature/phase-3-shared-extraction is empty). handler(event, context)
signatures and the save_* public contract are unchanged on both
pipelines; goldens unchanged.
PO_EXPECTED_TOP_LEVEL_MODULES and its WO equivalent in
tests/test_bundle_consistency.py are updated for the new sibling
modules so the AST bundle-consistency test still fails on an
unshipped or uncommented-out sibling.
2026-07-20 15:34:53 -04:00
|
|
|
|
from telemetry import emit_parse_metric
|
fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path (#104)
* fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path
The ai_fallback parse path applied no validation gate to raw Bedrock/LLM
output before DynamoDB writes, and the extraction prompt concatenated the
untrusted email body directly with no instructions-vs-data delimiter. A
DKIM-passing attacker could prompt-inject arbitrary field values into the
work-order store.
Changes:
- wrap untrusted email in \<email\> XML block with prompt instructing the
model to treat its contents as data only
- add validate_ai_fallback() in template_parser that enforces the same
contract keys, enums, and patterns as the template path before any write
- call validate_ai_fallback() in handler() dispatch; emit an
ai_fallback_rejected EMF metric on failure and skip the record
- add 17 unit tests covering every gate rule and two end-to-end dispatch
tests (injected email_type, injected status)
Refs #101
* style: apply ruff formatting to fix CI check
* harden ai_fallback gate: review fixes + security-review findings
Review follow-up on the ai_fallback validation gate (PR #104), plus
findings from a fan-out /sh-security-review of the change surface.
Reviewer FIX items:
- Neutralize forged <email> delimiters in the untrusted body before
wrapping, so an in-body </email> cannot escape the data block.
- Fail closed on non-dict model output instead of crashing the handler
into async retries; count ai_fallback_rejected parses in the
fallback-rate alarm and add a dedicated rejected-parse alarm so a
gate-rejection drift outage is not silent.
- Return a distinct invalid_status reason (was malformed_site_code);
validate ISO-8601 dates; README + docstring updates.
Security-review findings (detector fan-out + proof-or-kill verifier):
- ReDoS (confirmed, medium): the tag neutralizer used two \s* around an
optional /, backtracking quadratically on "<" + a long whitespace run
(~32s at 100k chars -- one email could time out the Lambda). Collapse
to a single [\s/]* class: linear, same defanging.
- Unhashable-type crash (confirmed): a JSON list/dict for email_type or
status made `x in <set>` raise TypeError, escaping the gate into
retries. Guard with isinstance(str) before membership.
- Unicode/newline regex (confirmed): _WO_ID_RE/_SITE_CODE_RE used ^..$
with \d, admitting fullwidth digits ("12345" as a lookalike
partition key) and trailing newlines. Switch to \A[0-9]+\Z (and the
handler's inline recheck to [0-9]) so neither passes.
- Alarm comment (confirmed, low): corrected the "slow trickle still
pages" wording -- rejections >~25-30 min apart page on neither alarm,
the same knowingly-accepted residual as sender-auth-rejected.
Refuted: residual free-text prompt injection is inherent to trusting
allowlisted senders, not a new primitive; no DynamoDB key-poisoning
bypass survives both gates ('#' can never enter work_order_id).
7 new regression tests. All 260 tests pass; ruff clean; cdk synth OK.
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-16 16:23:28 -04:00
|
|
|
|
from template_parser import try_deterministic_parse, validate_ai_fallback
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
|
|
2026-05-12 15:21:06 -04:00
|
|
|
|
logger = logging.getLogger()
|
|
|
|
|
|
logger.setLevel(logging.INFO)
|
|
|
|
|
|
|
feat: decompose email-processor handlers into flat siblings + lazy boto3 clients (refactor phase 5) (#113)
Both email-processor God-handlers split along the seams that already
work in the flat-sibling pattern established by lambdas/shared/, so
bare-name imports keep working under the existing bundling glob.
PO (5-way split): handler.py keeps only the event loop, fail-closed
auth, and email_type routing. extraction.py holds extract_with_claude
and _EMAIL_TAG_RE, importing EXTRACTION_PROMPT from prompts.py and
parse_raw_email from shared/email_parsing.py rather than recreating a
PO-local copy. enrichment.py is a pure code move of enrich_parsed and
pad_zip (PO-only; WO has no enrichment stage) with zero behavior
change. telemetry.py holds the EMF ParseMethod emit wrappers.
persistence.py holds _write_fields/_merge_update/save_*, collapsing
the byte-identical save_new_po/save_revision bodies into one
_save_merge helper that both now call through, preserving the sticky
Cancelled ConditionExpression guard for both callers; save_cancellation
stays separate.
WO (5 concerns, no enrichment stage): the handler loop keeps
validate_ai_fallback and the re.fullmatch(r"[0-9]+", work_order_id)
key guard ahead of both save_work_order and save_event, since the
guard protects the DynamoDB partition key and the '#'-delimited
comment_id range-key segment. _header_date_iso and comment_id
determinism stay colocated with persistence.py's save_event for the
retry-idempotent event_id key.
EXTRACTION_PROMPT (PO) moves to prompts.py with cross-reference
headers to derived_fields.py's authoritative trade/site/fiscal rule
tables; handler.py re-exports it (from prompts import
EXTRACTION_PROMPT) since four tests dereference handler.EXTRACTION_
PROMPT directly. WO's prompt moves the same way.
I/O modules (extraction.py's bedrock client, persistence.py's
dynamodb resource, handler.py's s3 client) get lazy cached boto3
accessors; pure modules (enrichment.py, prompts.py, telemetry.py)
import no boto3. Test monkeypatch surfaces move to the module that
now owns the client (e.g. persistence.dynamodb) everywhere tests
patch it, and the moto-before-handler-import ordering in
_po_parser_support.py is preserved so the moto-backed suites don't
hit real AWS.
Behavior-preservation pins, verified with tests: PO still emits
ParseMethod=ai_fallback before the Bedrock call, with
ai_fallback_rejected as the additive second datapoint on rejection.
WO still emits after its gate with mutually-exclusive ai_fallback /
ai_fallback_rejected. Shadow DerivedFieldAgreement telemetry stays
ai_fallback-only. derived_fields.py is untouched (diff against
feature/phase-3-shared-extraction is empty). handler(event, context)
signatures and the save_* public contract are unchanged on both
pipelines; goldens unchanged.
PO_EXPECTED_TOP_LEVEL_MODULES and its WO equivalent in
tests/test_bundle_consistency.py are updated for the new sibling
modules so the AST bundle-consistency test still fails on an
unshipped or uncommented-out sibling.
2026-07-20 15:34:53 -04:00
|
|
|
|
# Lazy cached S3 client. Keeps the public attribute name ``s3`` so the test
|
|
|
|
|
|
# monkeypatch target changes module only, not attribute name.
|
|
|
|
|
|
s3 = None
|
2026-05-12 15:21:06 -04:00
|
|
|
|
|
|
|
|
|
|
|
feat: decompose email-processor handlers into flat siblings + lazy boto3 clients (refactor phase 5) (#113)
Both email-processor God-handlers split along the seams that already
work in the flat-sibling pattern established by lambdas/shared/, so
bare-name imports keep working under the existing bundling glob.
PO (5-way split): handler.py keeps only the event loop, fail-closed
auth, and email_type routing. extraction.py holds extract_with_claude
and _EMAIL_TAG_RE, importing EXTRACTION_PROMPT from prompts.py and
parse_raw_email from shared/email_parsing.py rather than recreating a
PO-local copy. enrichment.py is a pure code move of enrich_parsed and
pad_zip (PO-only; WO has no enrichment stage) with zero behavior
change. telemetry.py holds the EMF ParseMethod emit wrappers.
persistence.py holds _write_fields/_merge_update/save_*, collapsing
the byte-identical save_new_po/save_revision bodies into one
_save_merge helper that both now call through, preserving the sticky
Cancelled ConditionExpression guard for both callers; save_cancellation
stays separate.
WO (5 concerns, no enrichment stage): the handler loop keeps
validate_ai_fallback and the re.fullmatch(r"[0-9]+", work_order_id)
key guard ahead of both save_work_order and save_event, since the
guard protects the DynamoDB partition key and the '#'-delimited
comment_id range-key segment. _header_date_iso and comment_id
determinism stay colocated with persistence.py's save_event for the
retry-idempotent event_id key.
EXTRACTION_PROMPT (PO) moves to prompts.py with cross-reference
headers to derived_fields.py's authoritative trade/site/fiscal rule
tables; handler.py re-exports it (from prompts import
EXTRACTION_PROMPT) since four tests dereference handler.EXTRACTION_
PROMPT directly. WO's prompt moves the same way.
I/O modules (extraction.py's bedrock client, persistence.py's
dynamodb resource, handler.py's s3 client) get lazy cached boto3
accessors; pure modules (enrichment.py, prompts.py, telemetry.py)
import no boto3. Test monkeypatch surfaces move to the module that
now owns the client (e.g. persistence.dynamodb) everywhere tests
patch it, and the moto-before-handler-import ordering in
_po_parser_support.py is preserved so the moto-backed suites don't
hit real AWS.
Behavior-preservation pins, verified with tests: PO still emits
ParseMethod=ai_fallback before the Bedrock call, with
ai_fallback_rejected as the additive second datapoint on rejection.
WO still emits after its gate with mutually-exclusive ai_fallback /
ai_fallback_rejected. Shadow DerivedFieldAgreement telemetry stays
ai_fallback-only. derived_fields.py is untouched (diff against
feature/phase-3-shared-extraction is empty). handler(event, context)
signatures and the save_* public contract are unchanged on both
pipelines; goldens unchanged.
PO_EXPECTED_TOP_LEVEL_MODULES and its WO equivalent in
tests/test_bundle_consistency.py are updated for the new sibling
modules so the AST bundle-consistency test still fails on an
unshipped or uncommented-out sibling.
2026-07-20 15:34:53 -04:00
|
|
|
|
def _get_s3():
|
|
|
|
|
|
global s3
|
|
|
|
|
|
if s3 is None:
|
|
|
|
|
|
s3 = boto3.client("s3")
|
|
|
|
|
|
return s3
|
2026-05-12 15:21:06 -04:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def handler(event, context):
|
|
|
|
|
|
"""Lambda entry point. Triggered by S3 ObjectCreated events."""
|
feat: deploy-pipeline guards — healthcheck, smoke gate, bundle glob + AST test (refactor phase 0) (#107)
* feat: deploy-pipeline guards — healthcheck, smoke gate, bundle glob + AST test (refactor phase 0)
Deploys of po-email-processor and workorder-email-processor had no
verification step, so an init-time ImportError in the bundled zip
could ship silently and only surface on the next real S3 event. This
adds a synchronous post-deploy smoke gate wired into the deploy
workflow: both Lambdas are invoked with {"healthcheck": true} and the
FunctionError field is checked, since an Unhandled init error still
returns HTTP 200 on RequestResponse invokes and would false-pass a
plain exit-code check.
The healthcheck branch is the first statement in each handler, before
any boto3/S3 use or ses_auth, and only fires on a top-level direct
invoke ("healthcheck" is not a key AWS ever sets on a real S3
ObjectCreated event, so mail content can't reach this path). It emits
no EMF metrics and no log text that could match the
sender-auth-rejected metric filter, so two deploys in one window
won't trip the alarm.
Separately, the PO stack's asset bundling copied a hand-maintained
four-file allowlist into the zip, so every new sibling module
handler.py imports had to be added by hand or the deploy shipped a
Lambda that ImportErrors at cold start (bit us for template_parser in
PR #105 and nearly for derived_fields in PR #2). Replaced it with a
non-recursive ./*.py glob so top-level source files ship
automatically while tests/ and the stale package/ dir still cannot,
and added an AST-based bundle-consistency test that parses each
handler's first-party imports and fails CI if the bundling command
would omit any of them (a revert to an incomplete allowlist, or code
moved into a subdirectory the glob doesn't cover).
Includes the refactor-evaluation report that scoped this phase.
* fix: review nits — unambiguous bundling-command extraction, smoke payload-parse message, dead asserts
- tests/test_bundle_consistency.py: _extract_bundling_command now collects
all command=[...] matches and demands exactly one per stack file, instead
of silently returning whichever ast.walk visits first if a second bundled
function is ever added.
- scripts/post-deploy-smoke.sh: distinguish an unparseable response payload
from a payload mismatch so the failure message says what actually happened
(the previous "could not parse" branch was unreachable — the inline python
always exited 0).
- test_po_healthcheck.py: drop the substring assertions on stdout that were
dead behind the stricter `captured.out == ""` assertion; keep the stderr
filter-pattern check.
Review follow-up on PR #107; no behavior change to any shipped code path.
2026-07-17 13:18:45 -04:00
|
|
|
|
# Deploy-guard healthcheck (Phase 0): a top-level direct-invoke
|
|
|
|
|
|
# {"healthcheck": true} probe returns immediately, BEFORE any S3 fetch,
|
|
|
|
|
|
# SES sender-auth gate, or Records iteration. Real mail arrives as S3
|
|
|
|
|
|
# ObjectCreated events whose top-level keys ("Records") AWS controls, so
|
|
|
|
|
|
# email content can never set this key -- this creates no accept path for
|
|
|
|
|
|
# mail. It emits NO EMF and no log line matching the sender_auth_rejected
|
|
|
|
|
|
# metric-filter, so repeated post-deploy smoke invokes never page.
|
|
|
|
|
|
if isinstance(event, dict) and event.get("healthcheck") is True:
|
|
|
|
|
|
return {"healthcheck": "ok"}
|
|
|
|
|
|
|
2026-05-12 15:21:06 -04:00
|
|
|
|
for record in event.get("Records", []):
|
|
|
|
|
|
bucket = record["s3"]["bucket"]["name"]
|
|
|
|
|
|
key = record["s3"]["object"]["key"]
|
Add fail-closed SES sender authentication (INFRA-107) (#98)
* Add fail-closed SES sender authentication
The From header and any raw-MIME Authentication-Results copies are
attacker-forgeable, so a forged email to apm@int.seahaven.com or
amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107,
CRITICAL). Both S3-triggered email processors now authenticate the
sender against the Authentication-Results header SES itself prepends
at delivery: only the topmost header is consulted, its authserv-id
must be amazonses.com, and it must carry dkim=pass for a domain in
the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set
in CDK so ops can adjust without code changes).
Allowlists come from live traffic observed 2026-07-15 on both ingest
buckets: WO mail arrives via the apm@ Google Groups forward, which
re-signs as seahaven.com (the hxgnsmartcloud.com signature does not
survive the forward); PO mail passes for amazon.coupahost.com.
amazonses.com also passes on PO mail but is deliberately excluded --
every SES customer's outbound mail passes for it.
Every failure path (env var unset, header missing or unparseable,
verdict fail, unaligned domain) rejects the email: a structured
warning with the reason and S3 key is logged and the record skipped
without erroring the invocation, so rejected mail causes no Lambda
retries or DLQ messages. Handler signatures and event sources are
unchanged.
Refs: INFRA-107
* Harden AR parser per cross-family review
Cross-family (GPT-4.1) review findings: terminate the dkim result
token at end-of-clause, whitespace, or a comment so a value like
"dkim=pass-fake" can never be read as a pass; normalize trailing
dots off allowlist entries so "seahaven.com." matches; make the
compat32 parser policy explicit. Adds tests for result-token
boundaries, comments after the result, quoted domain values, and
folding inside a dkim clause.
Refs: INFRA-107
* Harden AR parsing and alarm on sender-auth rejects
The SES-stamped Authentication-Results value echoes attacker-controlled
SMTP-session tokens (envelope-from, helo, header.from) as their own
semicolon-delimited property clauses. A naive split(";") tore an RFC 5321
quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass
clause, so a fully spoofed email was accepted on the genuinely
SES-stamped topmost header. Tokenise comment- and quoted-string-aware
(RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on
semicolons outside a quoted-string, and fail closed on unbalanced
quotes/comments so a ';' inside a quoted pvalue can never start a clause.
Rejected mail returns normally (no error, no retry, no DLQ message), so a
signing-domain drift or a wrong allowlist would silently discard 100% of
legitimate mail while every alarm stayed green. Add a CloudWatch Logs
metric filter + alarm on the sender_auth_rejected warning to both stacks
so a false-reject storm pages instead of vanishing. This is also the
safety net for the WO seahaven.com allowlist assumption, which must be
validated against a live SES-stamped header (a plain Gmail auto-forward
re-signs under the sending Workspace domain, not seahaven.com).
Refs: INFRA-107
* chore: retrigger CI (no run recorded for 7c74ac1)
* Fix quoted-AUID DKIM domain spoof in sender auth
Resolve three confirmed /sh-security-review findings on the fail-closed
SES sender-authentication control.
HIGH: header.i/header.d domain extraction was not quoted-string aware.
An attacker with a valid DKIM key for their own domain could set an
RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive
extractor stopped at the closing quote and returned seahaven.com,
accepting forged mail. Extraction now tokenises the clause with the same
quoted-string discipline already used for clause splitting: header.d
(the plain signing domain) is authoritative when present, otherwise the
header.i domain is the part after the AUID's LAST top-level "@", so a "@"
inside a quoted local-part is treated as signer-controlled label text and
yields the true signer (attacker.com), not seahaven.com.
LOW: the topmost-header parse ran outside evaluate_sender_authentication's
try/except, so an unexpected parser exception on crafted input could
propagate into the handler and Lambda async retries/DLQ. The parse now
fails CLOSED with an authentication_results_unparseable reason.
MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to
a low-volume total-reject outage (a trickle that never sums to 3). Both
stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3
/ datapoints_to_alarm=2, so a sustained reject condition pages even at one
reject per period while a lone stray probe self-clears.
Refs: INFRA-107
* Load Lambda function dir on sys.path in tests
Rebasing INFRA-107 onto main folded #95's pytest suite into this
branch's tests. The unified conftest loads the PO/WO handlers by file
path, and handler.py now does `from ses_auth import
authenticate_inbound_email` -- a bare sibling import that resolves in
the Lambda only because the runtime puts each function's own directory
on sys.path. The shared load_handler now adds that directory so the
handler tests import correctly alongside the sender-auth tests.
Refs: INFRA-107
* Note #97 test files in README directory tree
The rebase onto main brought in #97's tests/requirements.txt and
tests/test_po_merge.py. List both in the directory tree so it matches
the tree on disk.
Refs: INFRA-107
* Document INFRA-107 forwarder-binding risk acceptance
Record the accepted risk that WO sender auth binds to the apm@ forward's
re-signing domain (seahaven.com) rather than the Hexagon originator; the
apm@ Google Group's restricted posting policy is the load-bearing control
(escalates to HIGH if the group is opened to external posting). Also
correct the sender-auth-rejected alarm docs to match the shipped config
(>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and
note the SES-AR-01/02 parser hardening follow-ups.
Refs: INFRA-107
2026-07-15 20:58:47 -04:00
|
|
|
|
s3_key = f"s3://{bucket}/{key}"
|
2026-05-12 15:21:06 -04:00
|
|
|
|
|
Add fail-closed SES sender authentication (INFRA-107) (#98)
* Add fail-closed SES sender authentication
The From header and any raw-MIME Authentication-Results copies are
attacker-forgeable, so a forged email to apm@int.seahaven.com or
amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107,
CRITICAL). Both S3-triggered email processors now authenticate the
sender against the Authentication-Results header SES itself prepends
at delivery: only the topmost header is consulted, its authserv-id
must be amazonses.com, and it must carry dkim=pass for a domain in
the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set
in CDK so ops can adjust without code changes).
Allowlists come from live traffic observed 2026-07-15 on both ingest
buckets: WO mail arrives via the apm@ Google Groups forward, which
re-signs as seahaven.com (the hxgnsmartcloud.com signature does not
survive the forward); PO mail passes for amazon.coupahost.com.
amazonses.com also passes on PO mail but is deliberately excluded --
every SES customer's outbound mail passes for it.
Every failure path (env var unset, header missing or unparseable,
verdict fail, unaligned domain) rejects the email: a structured
warning with the reason and S3 key is logged and the record skipped
without erroring the invocation, so rejected mail causes no Lambda
retries or DLQ messages. Handler signatures and event sources are
unchanged.
Refs: INFRA-107
* Harden AR parser per cross-family review
Cross-family (GPT-4.1) review findings: terminate the dkim result
token at end-of-clause, whitespace, or a comment so a value like
"dkim=pass-fake" can never be read as a pass; normalize trailing
dots off allowlist entries so "seahaven.com." matches; make the
compat32 parser policy explicit. Adds tests for result-token
boundaries, comments after the result, quoted domain values, and
folding inside a dkim clause.
Refs: INFRA-107
* Harden AR parsing and alarm on sender-auth rejects
The SES-stamped Authentication-Results value echoes attacker-controlled
SMTP-session tokens (envelope-from, helo, header.from) as their own
semicolon-delimited property clauses. A naive split(";") tore an RFC 5321
quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass
clause, so a fully spoofed email was accepted on the genuinely
SES-stamped topmost header. Tokenise comment- and quoted-string-aware
(RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on
semicolons outside a quoted-string, and fail closed on unbalanced
quotes/comments so a ';' inside a quoted pvalue can never start a clause.
Rejected mail returns normally (no error, no retry, no DLQ message), so a
signing-domain drift or a wrong allowlist would silently discard 100% of
legitimate mail while every alarm stayed green. Add a CloudWatch Logs
metric filter + alarm on the sender_auth_rejected warning to both stacks
so a false-reject storm pages instead of vanishing. This is also the
safety net for the WO seahaven.com allowlist assumption, which must be
validated against a live SES-stamped header (a plain Gmail auto-forward
re-signs under the sending Workspace domain, not seahaven.com).
Refs: INFRA-107
* chore: retrigger CI (no run recorded for 7c74ac1)
* Fix quoted-AUID DKIM domain spoof in sender auth
Resolve three confirmed /sh-security-review findings on the fail-closed
SES sender-authentication control.
HIGH: header.i/header.d domain extraction was not quoted-string aware.
An attacker with a valid DKIM key for their own domain could set an
RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive
extractor stopped at the closing quote and returned seahaven.com,
accepting forged mail. Extraction now tokenises the clause with the same
quoted-string discipline already used for clause splitting: header.d
(the plain signing domain) is authoritative when present, otherwise the
header.i domain is the part after the AUID's LAST top-level "@", so a "@"
inside a quoted local-part is treated as signer-controlled label text and
yields the true signer (attacker.com), not seahaven.com.
LOW: the topmost-header parse ran outside evaluate_sender_authentication's
try/except, so an unexpected parser exception on crafted input could
propagate into the handler and Lambda async retries/DLQ. The parse now
fails CLOSED with an authentication_results_unparseable reason.
MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to
a low-volume total-reject outage (a trickle that never sums to 3). Both
stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3
/ datapoints_to_alarm=2, so a sustained reject condition pages even at one
reject per period while a lone stray probe self-clears.
Refs: INFRA-107
* Load Lambda function dir on sys.path in tests
Rebasing INFRA-107 onto main folded #95's pytest suite into this
branch's tests. The unified conftest loads the PO/WO handlers by file
path, and handler.py now does `from ses_auth import
authenticate_inbound_email` -- a bare sibling import that resolves in
the Lambda only because the runtime puts each function's own directory
on sys.path. The shared load_handler now adds that directory so the
handler tests import correctly alongside the sender-auth tests.
Refs: INFRA-107
* Note #97 test files in README directory tree
The rebase onto main brought in #97's tests/requirements.txt and
tests/test_po_merge.py. List both in the directory tree so it matches
the tree on disk.
Refs: INFRA-107
* Document INFRA-107 forwarder-binding risk acceptance
Record the accepted risk that WO sender auth binds to the apm@ forward's
re-signing domain (seahaven.com) rather than the Hexagon originator; the
apm@ Google Group's restricted posting policy is the load-bearing control
(escalates to HIGH if the group is opened to external posting). Also
correct the sender-auth-rejected alarm docs to match the shipped config
(>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and
note the SES-AR-01/02 parser hardening follow-ups.
Refs: INFRA-107
2026-07-15 20:58:47 -04:00
|
|
|
|
logger.info(f"Processing email: {s3_key}")
|
2026-05-12 15:21:06 -04:00
|
|
|
|
|
|
|
|
|
|
# Fetch raw email from S3
|
feat: decompose email-processor handlers into flat siblings + lazy boto3 clients (refactor phase 5) (#113)
Both email-processor God-handlers split along the seams that already
work in the flat-sibling pattern established by lambdas/shared/, so
bare-name imports keep working under the existing bundling glob.
PO (5-way split): handler.py keeps only the event loop, fail-closed
auth, and email_type routing. extraction.py holds extract_with_claude
and _EMAIL_TAG_RE, importing EXTRACTION_PROMPT from prompts.py and
parse_raw_email from shared/email_parsing.py rather than recreating a
PO-local copy. enrichment.py is a pure code move of enrich_parsed and
pad_zip (PO-only; WO has no enrichment stage) with zero behavior
change. telemetry.py holds the EMF ParseMethod emit wrappers.
persistence.py holds _write_fields/_merge_update/save_*, collapsing
the byte-identical save_new_po/save_revision bodies into one
_save_merge helper that both now call through, preserving the sticky
Cancelled ConditionExpression guard for both callers; save_cancellation
stays separate.
WO (5 concerns, no enrichment stage): the handler loop keeps
validate_ai_fallback and the re.fullmatch(r"[0-9]+", work_order_id)
key guard ahead of both save_work_order and save_event, since the
guard protects the DynamoDB partition key and the '#'-delimited
comment_id range-key segment. _header_date_iso and comment_id
determinism stay colocated with persistence.py's save_event for the
retry-idempotent event_id key.
EXTRACTION_PROMPT (PO) moves to prompts.py with cross-reference
headers to derived_fields.py's authoritative trade/site/fiscal rule
tables; handler.py re-exports it (from prompts import
EXTRACTION_PROMPT) since four tests dereference handler.EXTRACTION_
PROMPT directly. WO's prompt moves the same way.
I/O modules (extraction.py's bedrock client, persistence.py's
dynamodb resource, handler.py's s3 client) get lazy cached boto3
accessors; pure modules (enrichment.py, prompts.py, telemetry.py)
import no boto3. Test monkeypatch surfaces move to the module that
now owns the client (e.g. persistence.dynamodb) everywhere tests
patch it, and the moto-before-handler-import ordering in
_po_parser_support.py is preserved so the moto-backed suites don't
hit real AWS.
Behavior-preservation pins, verified with tests: PO still emits
ParseMethod=ai_fallback before the Bedrock call, with
ai_fallback_rejected as the additive second datapoint on rejection.
WO still emits after its gate with mutually-exclusive ai_fallback /
ai_fallback_rejected. Shadow DerivedFieldAgreement telemetry stays
ai_fallback-only. derived_fields.py is untouched (diff against
feature/phase-3-shared-extraction is empty). handler(event, context)
signatures and the save_* public contract are unchanged on both
pipelines; goldens unchanged.
PO_EXPECTED_TOP_LEVEL_MODULES and its WO equivalent in
tests/test_bundle_consistency.py are updated for the new sibling
modules so the AST bundle-consistency test still fails on an
unshipped or uncommented-out sibling.
2026-07-20 15:34:53 -04:00
|
|
|
|
response = _get_s3().get_object(Bucket=bucket, Key=key)
|
2026-05-12 15:21:06 -04:00
|
|
|
|
raw_email = response["Body"].read()
|
|
|
|
|
|
|
Add fail-closed SES sender authentication (INFRA-107) (#98)
* Add fail-closed SES sender authentication
The From header and any raw-MIME Authentication-Results copies are
attacker-forgeable, so a forged email to apm@int.seahaven.com or
amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107,
CRITICAL). Both S3-triggered email processors now authenticate the
sender against the Authentication-Results header SES itself prepends
at delivery: only the topmost header is consulted, its authserv-id
must be amazonses.com, and it must carry dkim=pass for a domain in
the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set
in CDK so ops can adjust without code changes).
Allowlists come from live traffic observed 2026-07-15 on both ingest
buckets: WO mail arrives via the apm@ Google Groups forward, which
re-signs as seahaven.com (the hxgnsmartcloud.com signature does not
survive the forward); PO mail passes for amazon.coupahost.com.
amazonses.com also passes on PO mail but is deliberately excluded --
every SES customer's outbound mail passes for it.
Every failure path (env var unset, header missing or unparseable,
verdict fail, unaligned domain) rejects the email: a structured
warning with the reason and S3 key is logged and the record skipped
without erroring the invocation, so rejected mail causes no Lambda
retries or DLQ messages. Handler signatures and event sources are
unchanged.
Refs: INFRA-107
* Harden AR parser per cross-family review
Cross-family (GPT-4.1) review findings: terminate the dkim result
token at end-of-clause, whitespace, or a comment so a value like
"dkim=pass-fake" can never be read as a pass; normalize trailing
dots off allowlist entries so "seahaven.com." matches; make the
compat32 parser policy explicit. Adds tests for result-token
boundaries, comments after the result, quoted domain values, and
folding inside a dkim clause.
Refs: INFRA-107
* Harden AR parsing and alarm on sender-auth rejects
The SES-stamped Authentication-Results value echoes attacker-controlled
SMTP-session tokens (envelope-from, helo, header.from) as their own
semicolon-delimited property clauses. A naive split(";") tore an RFC 5321
quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass
clause, so a fully spoofed email was accepted on the genuinely
SES-stamped topmost header. Tokenise comment- and quoted-string-aware
(RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on
semicolons outside a quoted-string, and fail closed on unbalanced
quotes/comments so a ';' inside a quoted pvalue can never start a clause.
Rejected mail returns normally (no error, no retry, no DLQ message), so a
signing-domain drift or a wrong allowlist would silently discard 100% of
legitimate mail while every alarm stayed green. Add a CloudWatch Logs
metric filter + alarm on the sender_auth_rejected warning to both stacks
so a false-reject storm pages instead of vanishing. This is also the
safety net for the WO seahaven.com allowlist assumption, which must be
validated against a live SES-stamped header (a plain Gmail auto-forward
re-signs under the sending Workspace domain, not seahaven.com).
Refs: INFRA-107
* chore: retrigger CI (no run recorded for 7c74ac1)
* Fix quoted-AUID DKIM domain spoof in sender auth
Resolve three confirmed /sh-security-review findings on the fail-closed
SES sender-authentication control.
HIGH: header.i/header.d domain extraction was not quoted-string aware.
An attacker with a valid DKIM key for their own domain could set an
RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive
extractor stopped at the closing quote and returned seahaven.com,
accepting forged mail. Extraction now tokenises the clause with the same
quoted-string discipline already used for clause splitting: header.d
(the plain signing domain) is authoritative when present, otherwise the
header.i domain is the part after the AUID's LAST top-level "@", so a "@"
inside a quoted local-part is treated as signer-controlled label text and
yields the true signer (attacker.com), not seahaven.com.
LOW: the topmost-header parse ran outside evaluate_sender_authentication's
try/except, so an unexpected parser exception on crafted input could
propagate into the handler and Lambda async retries/DLQ. The parse now
fails CLOSED with an authentication_results_unparseable reason.
MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to
a low-volume total-reject outage (a trickle that never sums to 3). Both
stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3
/ datapoints_to_alarm=2, so a sustained reject condition pages even at one
reject per period while a lone stray probe self-clears.
Refs: INFRA-107
* Load Lambda function dir on sys.path in tests
Rebasing INFRA-107 onto main folded #95's pytest suite into this
branch's tests. The unified conftest loads the PO/WO handlers by file
path, and handler.py now does `from ses_auth import
authenticate_inbound_email` -- a bare sibling import that resolves in
the Lambda only because the runtime puts each function's own directory
on sys.path. The shared load_handler now adds that directory so the
handler tests import correctly alongside the sender-auth tests.
Refs: INFRA-107
* Note #97 test files in README directory tree
The rebase onto main brought in #97's tests/requirements.txt and
tests/test_po_merge.py. List both in the directory tree so it matches
the tree on disk.
Refs: INFRA-107
* Document INFRA-107 forwarder-binding risk acceptance
Record the accepted risk that WO sender auth binds to the apm@ forward's
re-signing domain (seahaven.com) rather than the Hexagon originator; the
apm@ Google Group's restricted posting policy is the load-bearing control
(escalates to HIGH if the group is opened to external posting). Also
correct the sender-auth-rejected alarm docs to match the shipped config
(>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and
note the SES-AR-01/02 parser hardening follow-ups.
Refs: INFRA-107
2026-07-15 20:58:47 -04:00
|
|
|
|
# Fail-closed sender authentication (INFRA-107): only mail with an
|
|
|
|
|
|
# SES-stamped dkim=pass verdict for an allowlisted domain may create
|
|
|
|
|
|
# or update work orders. Rejected mail is logged and skipped without
|
|
|
|
|
|
# erroring the invocation (no retries / DLQ spam).
|
|
|
|
|
|
if not authenticate_inbound_email(raw_email, s3_key):
|
|
|
|
|
|
continue
|
|
|
|
|
|
|
2026-05-12 15:21:06 -04:00
|
|
|
|
# Parse the raw email
|
|
|
|
|
|
email_data = parse_raw_email(raw_email)
|
|
|
|
|
|
logger.info(f"Subject: {email_data['subject']}")
|
|
|
|
|
|
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
|
# Deterministic template parse first; fall back to the AI extractor only
|
|
|
|
|
|
# on a miss or an invalid (fail-closed) result.
|
|
|
|
|
|
parsed, method, template_id, reason = try_deterministic_parse(email_data)
|
|
|
|
|
|
if parsed is None:
|
|
|
|
|
|
parsed = extract_with_bedrock(email_data)
|
|
|
|
|
|
method = "ai_fallback"
|
fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path (#104)
* fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path
The ai_fallback parse path applied no validation gate to raw Bedrock/LLM
output before DynamoDB writes, and the extraction prompt concatenated the
untrusted email body directly with no instructions-vs-data delimiter. A
DKIM-passing attacker could prompt-inject arbitrary field values into the
work-order store.
Changes:
- wrap untrusted email in \<email\> XML block with prompt instructing the
model to treat its contents as data only
- add validate_ai_fallback() in template_parser that enforces the same
contract keys, enums, and patterns as the template path before any write
- call validate_ai_fallback() in handler() dispatch; emit an
ai_fallback_rejected EMF metric on failure and skip the record
- add 17 unit tests covering every gate rule and two end-to-end dispatch
tests (injected email_type, injected status)
Refs #101
* style: apply ruff formatting to fix CI check
* harden ai_fallback gate: review fixes + security-review findings
Review follow-up on the ai_fallback validation gate (PR #104), plus
findings from a fan-out /sh-security-review of the change surface.
Reviewer FIX items:
- Neutralize forged <email> delimiters in the untrusted body before
wrapping, so an in-body </email> cannot escape the data block.
- Fail closed on non-dict model output instead of crashing the handler
into async retries; count ai_fallback_rejected parses in the
fallback-rate alarm and add a dedicated rejected-parse alarm so a
gate-rejection drift outage is not silent.
- Return a distinct invalid_status reason (was malformed_site_code);
validate ISO-8601 dates; README + docstring updates.
Security-review findings (detector fan-out + proof-or-kill verifier):
- ReDoS (confirmed, medium): the tag neutralizer used two \s* around an
optional /, backtracking quadratically on "<" + a long whitespace run
(~32s at 100k chars -- one email could time out the Lambda). Collapse
to a single [\s/]* class: linear, same defanging.
- Unhashable-type crash (confirmed): a JSON list/dict for email_type or
status made `x in <set>` raise TypeError, escaping the gate into
retries. Guard with isinstance(str) before membership.
- Unicode/newline regex (confirmed): _WO_ID_RE/_SITE_CODE_RE used ^..$
with \d, admitting fullwidth digits ("12345" as a lookalike
partition key) and trailing newlines. Switch to \A[0-9]+\Z (and the
handler's inline recheck to [0-9]) so neither passes.
- Alarm comment (confirmed, low): corrected the "slow trickle still
pages" wording -- rejections >~25-30 min apart page on neither alarm,
the same knowingly-accepted residual as sender-auth-rejected.
Refuted: residual free-text prompt injection is inherent to trusting
allowlisted senders, not a new primitive; no DynamoDB key-poisoning
bypass survives both gates ('#' can never enter work_order_id).
7 new regression tests. All 260 tests pass; ruff clean; cdk synth OK.
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-16 16:23:28 -04:00
|
|
|
|
# Fail-closed validation gate on AI output: a prompt-injected
|
|
|
|
|
|
# email body could steer the model into returning arbitrary
|
|
|
|
|
|
# field values, so enforce the same structural contract on both
|
|
|
|
|
|
# parse paths BEFORE any DynamoDB write.
|
|
|
|
|
|
ok, val_reason = validate_ai_fallback(parsed)
|
|
|
|
|
|
if not ok:
|
|
|
|
|
|
logger.warning(
|
|
|
|
|
|
f"AI-fallback validation failed ({val_reason}), skipping: {key}"
|
|
|
|
|
|
)
|
|
|
|
|
|
emit_parse_metric(
|
|
|
|
|
|
"ai_fallback_rejected",
|
|
|
|
|
|
template_id,
|
|
|
|
|
|
val_reason,
|
|
|
|
|
|
parsed.get("work_order_id") if isinstance(parsed, dict) else None,
|
|
|
|
|
|
)
|
|
|
|
|
|
continue
|
2026-05-12 15:21:06 -04:00
|
|
|
|
logger.info(
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
|
f"Parsed ({method}/{template_id}/{reason}): "
|
|
|
|
|
|
f"type={parsed.get('email_type')}, wo={parsed.get('work_order_id')}"
|
2026-05-12 15:21:06 -04:00
|
|
|
|
)
|
|
|
|
|
|
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
|
emit_parse_metric(method, template_id, reason, parsed.get("work_order_id"))
|
|
|
|
|
|
|
|
|
|
|
|
# work_order_id becomes a DynamoDB partition key and the leading, '#'-
|
|
|
|
|
|
# delimited segment of the comment_id range key, so it must be digits
|
|
|
|
|
|
# only. The template path already guarantees this via validate(); the
|
|
|
|
|
|
# AI-fallback path returns raw model output, which a prompt-injected
|
|
|
|
|
|
# email body could steer into a non-numeric or '#'-bearing value that
|
|
|
|
|
|
# forges key segments or lands on an arbitrary WO. Enforce the same
|
|
|
|
|
|
# contract on both paths and skip (fail closed) on a violation.
|
fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path (#104)
* fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path
The ai_fallback parse path applied no validation gate to raw Bedrock/LLM
output before DynamoDB writes, and the extraction prompt concatenated the
untrusted email body directly with no instructions-vs-data delimiter. A
DKIM-passing attacker could prompt-inject arbitrary field values into the
work-order store.
Changes:
- wrap untrusted email in \<email\> XML block with prompt instructing the
model to treat its contents as data only
- add validate_ai_fallback() in template_parser that enforces the same
contract keys, enums, and patterns as the template path before any write
- call validate_ai_fallback() in handler() dispatch; emit an
ai_fallback_rejected EMF metric on failure and skip the record
- add 17 unit tests covering every gate rule and two end-to-end dispatch
tests (injected email_type, injected status)
Refs #101
* style: apply ruff formatting to fix CI check
* harden ai_fallback gate: review fixes + security-review findings
Review follow-up on the ai_fallback validation gate (PR #104), plus
findings from a fan-out /sh-security-review of the change surface.
Reviewer FIX items:
- Neutralize forged <email> delimiters in the untrusted body before
wrapping, so an in-body </email> cannot escape the data block.
- Fail closed on non-dict model output instead of crashing the handler
into async retries; count ai_fallback_rejected parses in the
fallback-rate alarm and add a dedicated rejected-parse alarm so a
gate-rejection drift outage is not silent.
- Return a distinct invalid_status reason (was malformed_site_code);
validate ISO-8601 dates; README + docstring updates.
Security-review findings (detector fan-out + proof-or-kill verifier):
- ReDoS (confirmed, medium): the tag neutralizer used two \s* around an
optional /, backtracking quadratically on "<" + a long whitespace run
(~32s at 100k chars -- one email could time out the Lambda). Collapse
to a single [\s/]* class: linear, same defanging.
- Unhashable-type crash (confirmed): a JSON list/dict for email_type or
status made `x in <set>` raise TypeError, escaping the gate into
retries. Guard with isinstance(str) before membership.
- Unicode/newline regex (confirmed): _WO_ID_RE/_SITE_CODE_RE used ^..$
with \d, admitting fullwidth digits ("12345" as a lookalike
partition key) and trailing newlines. Switch to \A[0-9]+\Z (and the
handler's inline recheck to [0-9]) so neither passes.
- Alarm comment (confirmed, low): corrected the "slow trickle still
pages" wording -- rejections >~25-30 min apart page on neither alarm,
the same knowingly-accepted residual as sender-auth-rejected.
Refuted: residual free-text prompt injection is inherent to trusting
allowlisted senders, not a new primitive; no DynamoDB key-poisoning
bypass survives both gates ('#' can never enter work_order_id).
7 new regression tests. All 260 tests pass; ruff clean; cdk synth OK.
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-16 16:23:28 -04:00
|
|
|
|
# [0-9] not \d: \d is Unicode-aware and would admit fullwidth digits
|
|
|
|
|
|
# (e.g. "12345") as a distinct-but-lookalike partition key.
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
|
work_order_id = parsed.get("work_order_id")
|
fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path (#104)
* fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path
The ai_fallback parse path applied no validation gate to raw Bedrock/LLM
output before DynamoDB writes, and the extraction prompt concatenated the
untrusted email body directly with no instructions-vs-data delimiter. A
DKIM-passing attacker could prompt-inject arbitrary field values into the
work-order store.
Changes:
- wrap untrusted email in \<email\> XML block with prompt instructing the
model to treat its contents as data only
- add validate_ai_fallback() in template_parser that enforces the same
contract keys, enums, and patterns as the template path before any write
- call validate_ai_fallback() in handler() dispatch; emit an
ai_fallback_rejected EMF metric on failure and skip the record
- add 17 unit tests covering every gate rule and two end-to-end dispatch
tests (injected email_type, injected status)
Refs #101
* style: apply ruff formatting to fix CI check
* harden ai_fallback gate: review fixes + security-review findings
Review follow-up on the ai_fallback validation gate (PR #104), plus
findings from a fan-out /sh-security-review of the change surface.
Reviewer FIX items:
- Neutralize forged <email> delimiters in the untrusted body before
wrapping, so an in-body </email> cannot escape the data block.
- Fail closed on non-dict model output instead of crashing the handler
into async retries; count ai_fallback_rejected parses in the
fallback-rate alarm and add a dedicated rejected-parse alarm so a
gate-rejection drift outage is not silent.
- Return a distinct invalid_status reason (was malformed_site_code);
validate ISO-8601 dates; README + docstring updates.
Security-review findings (detector fan-out + proof-or-kill verifier):
- ReDoS (confirmed, medium): the tag neutralizer used two \s* around an
optional /, backtracking quadratically on "<" + a long whitespace run
(~32s at 100k chars -- one email could time out the Lambda). Collapse
to a single [\s/]* class: linear, same defanging.
- Unhashable-type crash (confirmed): a JSON list/dict for email_type or
status made `x in <set>` raise TypeError, escaping the gate into
retries. Guard with isinstance(str) before membership.
- Unicode/newline regex (confirmed): _WO_ID_RE/_SITE_CODE_RE used ^..$
with \d, admitting fullwidth digits ("12345" as a lookalike
partition key) and trailing newlines. Switch to \A[0-9]+\Z (and the
handler's inline recheck to [0-9]) so neither passes.
- Alarm comment (confirmed, low): corrected the "slow trickle still
pages" wording -- rejections >~25-30 min apart page on neither alarm,
the same knowingly-accepted residual as sender-auth-rejected.
Refuted: residual free-text prompt injection is inherent to trusting
allowlisted senders, not a new primitive; no DynamoDB key-poisoning
bypass survives both gates ('#' can never enter work_order_id).
7 new regression tests. All 260 tests pass; ruff clean; cdk synth OK.
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-16 16:23:28 -04:00
|
|
|
|
if not work_order_id or not re.fullmatch(r"[0-9]+", str(work_order_id)):
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
|
logger.warning(f"Missing or non-numeric work order ID, skipping: {key}")
|
2026-05-12 15:21:06 -04:00
|
|
|
|
continue
|
|
|
|
|
|
|
|
|
|
|
|
# Always upsert the work order with any new info
|
|
|
|
|
|
save_work_order(parsed, s3_key)
|
|
|
|
|
|
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
|
# Save every email as an event for history tracking. The raw object key
|
|
|
|
|
|
# (not the s3:// URI) drives the retry-idempotent comment_id suffix;
|
|
|
|
|
|
# the parse method decides whether comment_time may enter the key (A1).
|
|
|
|
|
|
save_event(parsed, s3_key, key, method, email_data.get("date"))
|
2026-05-12 15:21:06 -04:00
|
|
|
|
|
|
|
|
|
|
return {"statusCode": 200, "body": "OK"}
|