feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
|
"""Deterministic template parser for Hexagon EAM work-order emails.
|
|
|
|
|
|
|
|
|
|
|
|
Pure module: no boto3, no network. Runs ahead of the AI extraction path in the
|
|
|
|
|
|
work-order email processor. Only returns a parsed result when it is proven
|
|
|
|
|
|
conformant to one of the two known Hexagon templates; otherwise it fails closed
|
|
|
|
|
|
and signals the caller to fall back to the AI extractor.
|
|
|
|
|
|
|
|
|
|
|
|
Two templates (see BUILD SPEC / recon):
|
|
|
|
|
|
T1 update_plaintext -- Subject "AMAZON UPDATE WO DETAILS <id>", text/plain,
|
|
|
|
|
|
labels New Comment: / Creation Time(UTC): / Submitted By: /
|
|
|
|
|
|
"Work Order: <id> - <desc>" (double space) / "Building: <SITE>." + address.
|
|
|
|
|
|
Maps to email_type "comment".
|
|
|
|
|
|
T2 assign_html -- Subject "AMAZON assign Work Order <id> on building <SITE>",
|
|
|
|
|
|
simple HTML, <br>-delimited WO Description: / Severity: / Date Reported: /
|
|
|
|
|
|
Scheduled Start Date: / Address:. Maps to email_type "new_work_order".
|
|
|
|
|
|
|
|
|
|
|
|
Entry point: try_deterministic_parse(email_data) -> (parsed|None, method,
|
|
|
|
|
|
template_id, reason_code).
|
|
|
|
|
|
"""
|
|
|
|
|
|
|
|
|
|
|
|
import html as html_module
|
|
|
|
|
|
import re
|
|
|
|
|
|
|
|
|
|
|
|
# The contract keys (16), EXACTLY -- mirrors the AI EXTRACTION_PROMPT fields. A conformant parse is a dict with these keys
|
|
|
|
|
|
# and no others (validation rule 2).
|
|
|
|
|
|
CONTRACT_KEYS = (
|
|
|
|
|
|
"email_type",
|
|
|
|
|
|
"work_order_id",
|
|
|
|
|
|
"description",
|
|
|
|
|
|
"status",
|
|
|
|
|
|
"site_code",
|
|
|
|
|
|
"building",
|
|
|
|
|
|
"address",
|
|
|
|
|
|
"severity",
|
|
|
|
|
|
"priority",
|
|
|
|
|
|
"date_reported",
|
|
|
|
|
|
"scheduled_start",
|
|
|
|
|
|
"due_date",
|
|
|
|
|
|
"assigned_to",
|
|
|
|
|
|
"commenter",
|
|
|
|
|
|
"comment_text",
|
|
|
|
|
|
"comment_time",
|
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
VALID_EMAIL_TYPES = {"new_work_order", "update", "comment", "cancellation"}
|
|
|
|
|
|
VALID_STATUSES = {
|
|
|
|
|
|
"new",
|
|
|
|
|
|
"assigned",
|
|
|
|
|
|
"in_progress",
|
|
|
|
|
|
"on_hold",
|
|
|
|
|
|
"completed",
|
|
|
|
|
|
"cancelled",
|
|
|
|
|
|
"unknown",
|
|
|
|
|
|
}
|
|
|
|
|
|
|
2026-08-04 19:39:20 -04:00
|
|
|
|
# Free-text scalar fields that must be None or str (blocks LLM-emitted maps/
|
|
|
|
|
|
# lists from landing as DynamoDB Map/List attribute pollution, and floats that
|
|
|
|
|
|
# would crash update_item). email_type, status, work_order_id, site_code, and
|
|
|
|
|
|
# date fields are validated separately.
|
|
|
|
|
|
_AI_FREE_TEXT_STR_FIELDS = (
|
|
|
|
|
|
"description",
|
|
|
|
|
|
"building",
|
|
|
|
|
|
"address",
|
|
|
|
|
|
"severity",
|
|
|
|
|
|
"priority",
|
|
|
|
|
|
"assigned_to",
|
|
|
|
|
|
"commenter",
|
|
|
|
|
|
"comment_text",
|
|
|
|
|
|
)
|
|
|
|
|
|
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
|
# Subject classifiers.
|
|
|
|
|
|
_T1_SUBJECT = re.compile(r"^AMAZON UPDATE WO DETAILS\s+(?P<wo>\S+)\s*$")
|
|
|
|
|
|
_T2_SUBJECT = re.compile(
|
|
|
|
|
|
r"^AMAZON assign Work Order\s+(?P<wo>\S+)\s+on building\s+(?P<site>\S+)\s*$"
|
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
# Known label tokens per template (lower-cased, used for scanning + bleed guard).
|
|
|
|
|
|
_T1_LABELS = (
|
|
|
|
|
|
"new comment:",
|
|
|
|
|
|
"creation time(utc):",
|
|
|
|
|
|
"submitted by:",
|
|
|
|
|
|
"work order:",
|
|
|
|
|
|
"building:",
|
|
|
|
|
|
)
|
|
|
|
|
|
_T2_LABELS = (
|
|
|
|
|
|
"wo description:",
|
|
|
|
|
|
"severity:",
|
|
|
|
|
|
"date reported:",
|
|
|
|
|
|
"scheduled start date:",
|
|
|
|
|
|
"address:",
|
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
_SEPARATOR_RE = re.compile(r"_{4,}")
|
fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path (#104)
* fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path
The ai_fallback parse path applied no validation gate to raw Bedrock/LLM
output before DynamoDB writes, and the extraction prompt concatenated the
untrusted email body directly with no instructions-vs-data delimiter. A
DKIM-passing attacker could prompt-inject arbitrary field values into the
work-order store.
Changes:
- wrap untrusted email in \<email\> XML block with prompt instructing the
model to treat its contents as data only
- add validate_ai_fallback() in template_parser that enforces the same
contract keys, enums, and patterns as the template path before any write
- call validate_ai_fallback() in handler() dispatch; emit an
ai_fallback_rejected EMF metric on failure and skip the record
- add 17 unit tests covering every gate rule and two end-to-end dispatch
tests (injected email_type, injected status)
Refs #101
* style: apply ruff formatting to fix CI check
* harden ai_fallback gate: review fixes + security-review findings
Review follow-up on the ai_fallback validation gate (PR #104), plus
findings from a fan-out /sh-security-review of the change surface.
Reviewer FIX items:
- Neutralize forged <email> delimiters in the untrusted body before
wrapping, so an in-body </email> cannot escape the data block.
- Fail closed on non-dict model output instead of crashing the handler
into async retries; count ai_fallback_rejected parses in the
fallback-rate alarm and add a dedicated rejected-parse alarm so a
gate-rejection drift outage is not silent.
- Return a distinct invalid_status reason (was malformed_site_code);
validate ISO-8601 dates; README + docstring updates.
Security-review findings (detector fan-out + proof-or-kill verifier):
- ReDoS (confirmed, medium): the tag neutralizer used two \s* around an
optional /, backtracking quadratically on "<" + a long whitespace run
(~32s at 100k chars -- one email could time out the Lambda). Collapse
to a single [\s/]* class: linear, same defanging.
- Unhashable-type crash (confirmed): a JSON list/dict for email_type or
status made `x in <set>` raise TypeError, escaping the gate into
retries. Guard with isinstance(str) before membership.
- Unicode/newline regex (confirmed): _WO_ID_RE/_SITE_CODE_RE used ^..$
with \d, admitting fullwidth digits ("12345" as a lookalike
partition key) and trailing newlines. Switch to \A[0-9]+\Z (and the
handler's inline recheck to [0-9]) so neither passes.
- Alarm comment (confirmed, low): corrected the "slow trickle still
pages" wording -- rejections >~25-30 min apart page on neither alarm,
the same knowingly-accepted residual as sender-auth-rejected.
Refuted: residual free-text prompt injection is inherent to trusting
allowlisted senders, not a new primitive; no DynamoDB key-poisoning
bypass survives both gates ('#' can never enter work_order_id).
7 new regression tests. All 260 tests pass; ruff clean; cdk synth OK.
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-16 16:23:28 -04:00
|
|
|
|
# \A...\Z (not ^...$, whose $ also matches just before a trailing newline) and
|
|
|
|
|
|
# explicit [0-9] (not \d, which is Unicode-aware and would admit fullwidth
|
|
|
|
|
|
# digits) so a value like "WIL1\n" or "12345" cannot pass as well-formed.
|
|
|
|
|
|
_SITE_CODE_RE = re.compile(r"\A[A-Z]{2,4}[0-9]{1,2}\Z")
|
|
|
|
|
|
_WO_ID_RE = re.compile(r"\A[0-9]+\Z")
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _empty_candidate():
|
|
|
|
|
|
return {k: None for k in CONTRACT_KEYS}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _normalize(candidate):
|
|
|
|
|
|
"""Guarantee all contract keys exist (None for absent) before returning."""
|
|
|
|
|
|
out = _empty_candidate()
|
|
|
|
|
|
for k in CONTRACT_KEYS:
|
|
|
|
|
|
if k in candidate:
|
|
|
|
|
|
out[k] = candidate[k]
|
|
|
|
|
|
return out
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def classify_template(email_data):
|
|
|
|
|
|
"""Return (template_id, reason). template_id in {update_plaintext,
|
|
|
|
|
|
assign_html, unknown}."""
|
|
|
|
|
|
subject = (email_data.get("subject") or "").strip()
|
|
|
|
|
|
if _T1_SUBJECT.match(subject):
|
|
|
|
|
|
return "update_plaintext", "ok"
|
|
|
|
|
|
if _T2_SUBJECT.match(subject):
|
|
|
|
|
|
return "assign_html", "ok"
|
|
|
|
|
|
return "unknown", "subject_no_match"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _subject_ids(email_data):
|
|
|
|
|
|
"""Return (wo_id, site_code) parsed from the subject, or (None, None)."""
|
|
|
|
|
|
subject = (email_data.get("subject") or "").strip()
|
|
|
|
|
|
m = _T1_SUBJECT.match(subject)
|
|
|
|
|
|
if m:
|
|
|
|
|
|
return m.group("wo"), None
|
|
|
|
|
|
m = _T2_SUBJECT.match(subject)
|
|
|
|
|
|
if m:
|
|
|
|
|
|
return m.group("wo"), m.group("site")
|
|
|
|
|
|
return None, None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _matches_any_label(line, labels):
|
|
|
|
|
|
low = line.strip().lower()
|
|
|
|
|
|
return any(low.startswith(lab) for lab in labels)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _contains_label_or_separator(text, labels):
|
|
|
|
|
|
"""Label-bleed guard: True if text carries a known label token or a
|
|
|
|
|
|
separator run (indicates the value over-ran into the next field)."""
|
|
|
|
|
|
if text is None:
|
|
|
|
|
|
return False
|
|
|
|
|
|
if _SEPARATOR_RE.search(text):
|
|
|
|
|
|
return True
|
|
|
|
|
|
low = text.lower()
|
|
|
|
|
|
return any(lab in low for lab in labels)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _html_to_lines(body):
|
|
|
|
|
|
"""Tiny HTML->text: turn <br>/</p>/</tr> and real CRLFs into line breaks,
|
|
|
|
|
|
strip remaining tags, unescape entities. Returns a list of raw lines."""
|
|
|
|
|
|
text = re.sub(r"(?i)<br\s*/?>", "\n", body)
|
|
|
|
|
|
text = re.sub(r"(?i)</p\s*>", "\n", text)
|
|
|
|
|
|
text = re.sub(r"(?i)</tr\s*>", "\n", text)
|
|
|
|
|
|
text = re.sub(r"<[^>]+>", "", text)
|
|
|
|
|
|
text = html_module.unescape(text)
|
|
|
|
|
|
text = text.replace("\r\n", "\n").replace("\r", "\n")
|
|
|
|
|
|
return text.split("\n")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _plain_lines(body):
|
|
|
|
|
|
return body.replace("\r\n", "\n").replace("\r", "\n").split("\n")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _find_label_index(lines, label):
|
|
|
|
|
|
ll = label.lower()
|
|
|
|
|
|
for i, line in enumerate(lines):
|
|
|
|
|
|
if line.strip().lower().startswith(ll):
|
|
|
|
|
|
return i
|
|
|
|
|
|
return -1
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _inline_value(line, label):
|
|
|
|
|
|
"""Value = remainder of the label line after the label token."""
|
|
|
|
|
|
idx = line.lower().find(label.lower())
|
|
|
|
|
|
return line[idx + len(label) :].strip()
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _capture_block(lines, start_index, labels, stop_on_blank=True):
|
|
|
|
|
|
"""Collect lines after start_index until a blank line (unless
|
|
|
|
|
|
stop_on_blank=False), a separator run, or a known label. Returns a list of
|
|
|
|
|
|
stripped non-consumed lines (may be empty).
|
|
|
|
|
|
|
|
|
|
|
|
stop_on_blank=False is for free-text blocks that legitimately contain blank
|
|
|
|
|
|
lines (multi-paragraph comments, advisory A3): internal blanks are kept as
|
|
|
|
|
|
empty strings, leading/trailing blanks are trimmed."""
|
|
|
|
|
|
collected = []
|
|
|
|
|
|
j = start_index + 1
|
|
|
|
|
|
while j < len(lines):
|
|
|
|
|
|
s = lines[j].strip()
|
|
|
|
|
|
if s == "":
|
|
|
|
|
|
if stop_on_blank:
|
|
|
|
|
|
break
|
|
|
|
|
|
# Leading blanks are never collected, so a long blank run cannot
|
|
|
|
|
|
# accumulate ahead of the O(1)-per-line trim below.
|
|
|
|
|
|
if collected:
|
|
|
|
|
|
collected.append("")
|
|
|
|
|
|
j += 1
|
|
|
|
|
|
continue
|
|
|
|
|
|
if _SEPARATOR_RE.fullmatch(s) or _SEPARATOR_RE.search(s):
|
|
|
|
|
|
break
|
|
|
|
|
|
if _matches_any_label(s, labels):
|
|
|
|
|
|
break
|
|
|
|
|
|
collected.append(s)
|
|
|
|
|
|
j += 1
|
|
|
|
|
|
while collected and collected[-1] == "":
|
|
|
|
|
|
collected.pop()
|
|
|
|
|
|
return collected
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
# T1: update_plaintext -> comment
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def extract_update_plaintext(email_data):
|
|
|
|
|
|
"""Extract all contract keys from a T1 plaintext update email. Returns a full
|
|
|
|
|
|
dict (values None where absent). Correctness is enforced by validate()."""
|
|
|
|
|
|
subject_wo, _ = _subject_ids(email_data)
|
|
|
|
|
|
lines = _plain_lines(email_data.get("body") or "")
|
|
|
|
|
|
candidate = _empty_candidate()
|
|
|
|
|
|
candidate["email_type"] = "comment"
|
|
|
|
|
|
candidate["work_order_id"] = subject_wo
|
|
|
|
|
|
# status stays None on a comment upsert -- never clobber a real wo_status.
|
|
|
|
|
|
|
|
|
|
|
|
# --- New Comment: block up to separator / next label. Blank lines do NOT
|
|
|
|
|
|
# end the block (multi-paragraph comments, advisory A3); the next label
|
|
|
|
|
|
# (normally Creation Time(UTC):) is the terminator. ---
|
|
|
|
|
|
nc_idx = _find_label_index(lines, "New Comment:")
|
|
|
|
|
|
if nc_idx >= 0:
|
|
|
|
|
|
pieces = []
|
|
|
|
|
|
inline = _inline_value(lines[nc_idx], "New Comment:")
|
|
|
|
|
|
if inline:
|
|
|
|
|
|
pieces.append(inline)
|
|
|
|
|
|
pieces.extend(_capture_block(lines, nc_idx, _T1_LABELS, stop_on_blank=False))
|
|
|
|
|
|
candidate["comment_text"] = "\n".join(pieces) if pieces else None
|
|
|
|
|
|
|
|
|
|
|
|
# --- Creation Time(UTC): -> ISO ---
|
|
|
|
|
|
ct_idx = _find_label_index(lines, "Creation Time(UTC):")
|
|
|
|
|
|
if ct_idx >= 0:
|
|
|
|
|
|
raw = _inline_value(lines[ct_idx], "Creation Time(UTC):")
|
|
|
|
|
|
candidate["comment_time"] = _parse_dt(raw, "%Y-%m-%d %H:%M:%S")
|
|
|
|
|
|
|
|
|
|
|
|
# --- Submitted By: -> commenter (username or joined ARN) ---
|
|
|
|
|
|
sb_idx = _find_label_index(lines, "Submitted By:")
|
|
|
|
|
|
if sb_idx >= 0:
|
|
|
|
|
|
pieces = []
|
|
|
|
|
|
inline = _inline_value(lines[sb_idx], "Submitted By:")
|
|
|
|
|
|
if inline:
|
|
|
|
|
|
pieces.append(inline)
|
|
|
|
|
|
# ARN continuation lines (rare) join with no separator.
|
|
|
|
|
|
pieces.extend(_capture_block(lines, sb_idx, _T1_LABELS))
|
|
|
|
|
|
candidate["commenter"] = "".join(pieces) if pieces else None
|
|
|
|
|
|
|
|
|
|
|
|
# --- Work Order: <id> - <desc> ---
|
|
|
|
|
|
wo_idx = _find_label_index(lines, "Work Order:")
|
|
|
|
|
|
if wo_idx >= 0:
|
|
|
|
|
|
m = re.match(r"\s*Work Order:\s+(\S+)\s+-\s+(.*)$", lines[wo_idx])
|
|
|
|
|
|
if m:
|
|
|
|
|
|
desc = m.group(2).strip()
|
|
|
|
|
|
candidate["description"] = desc or None
|
|
|
|
|
|
|
|
|
|
|
|
# --- Building: <SITE>. + optional address block ---
|
|
|
|
|
|
b_idx = _find_label_index(lines, "Building:")
|
|
|
|
|
|
if b_idx >= 0:
|
|
|
|
|
|
site = _inline_value(lines[b_idx], "Building:").rstrip(".").strip()
|
|
|
|
|
|
if site:
|
|
|
|
|
|
candidate["site_code"] = site
|
|
|
|
|
|
candidate["building"] = site
|
|
|
|
|
|
addr = _capture_block(lines, b_idx, _T1_LABELS)
|
|
|
|
|
|
candidate["address"] = "\n".join(addr) if addr else None
|
|
|
|
|
|
|
|
|
|
|
|
return _normalize(candidate)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
# T2: assign_html -> new_work_order
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def extract_assign_html(email_data):
|
|
|
|
|
|
"""Extract all contract keys from a T2 HTML assign email."""
|
|
|
|
|
|
subject_wo, subject_site = _subject_ids(email_data)
|
|
|
|
|
|
lines = _html_to_lines(email_data.get("body") or "")
|
|
|
|
|
|
candidate = _empty_candidate()
|
|
|
|
|
|
candidate["email_type"] = "new_work_order"
|
|
|
|
|
|
candidate["status"] = "assigned"
|
|
|
|
|
|
candidate["work_order_id"] = subject_wo
|
|
|
|
|
|
candidate["site_code"] = subject_site
|
|
|
|
|
|
candidate["building"] = subject_site
|
|
|
|
|
|
|
|
|
|
|
|
d_idx = _find_label_index(lines, "WO Description:")
|
|
|
|
|
|
if d_idx >= 0:
|
|
|
|
|
|
desc = _inline_value(lines[d_idx], "WO Description:")
|
|
|
|
|
|
candidate["description"] = desc or None
|
|
|
|
|
|
|
|
|
|
|
|
s_idx = _find_label_index(lines, "Severity:")
|
|
|
|
|
|
if s_idx >= 0:
|
|
|
|
|
|
sev = _inline_value(lines[s_idx], "Severity:")
|
|
|
|
|
|
candidate["severity"] = sev or None
|
|
|
|
|
|
|
|
|
|
|
|
dr_idx = _find_label_index(lines, "Date Reported:")
|
|
|
|
|
|
if dr_idx >= 0:
|
|
|
|
|
|
raw = _inline_value(lines[dr_idx], "Date Reported:")
|
|
|
|
|
|
candidate["date_reported"] = _parse_dt(raw, "%Y-%m-%d %H:%M")
|
|
|
|
|
|
|
|
|
|
|
|
ss_idx = _find_label_index(lines, "Scheduled Start Date:")
|
|
|
|
|
|
if ss_idx >= 0:
|
|
|
|
|
|
raw = _inline_value(lines[ss_idx], "Scheduled Start Date:")
|
|
|
|
|
|
candidate["scheduled_start"] = _parse_date(raw, "%Y-%m-%d")
|
|
|
|
|
|
|
|
|
|
|
|
a_idx = _find_label_index(lines, "Address:")
|
|
|
|
|
|
if a_idx >= 0:
|
|
|
|
|
|
inline = _inline_value(lines[a_idx], "Address:")
|
|
|
|
|
|
pieces = []
|
|
|
|
|
|
if inline:
|
|
|
|
|
|
pieces.append(inline)
|
|
|
|
|
|
pieces.extend(_capture_block(lines, a_idx, _T2_LABELS))
|
|
|
|
|
|
candidate["address"] = "\n".join(pieces) if pieces else None
|
|
|
|
|
|
|
|
|
|
|
|
return _normalize(candidate)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _parse_dt(raw, fmt):
|
|
|
|
|
|
from datetime import datetime
|
|
|
|
|
|
|
|
|
|
|
|
try:
|
|
|
|
|
|
return datetime.strptime(raw.strip(), fmt).isoformat()
|
|
|
|
|
|
except (ValueError, AttributeError):
|
|
|
|
|
|
return _UNPARSEABLE
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _parse_date(raw, fmt):
|
|
|
|
|
|
from datetime import datetime
|
|
|
|
|
|
|
|
|
|
|
|
try:
|
|
|
|
|
|
return datetime.strptime(raw.strip(), fmt).date().isoformat()
|
|
|
|
|
|
except (ValueError, AttributeError):
|
|
|
|
|
|
return _UNPARSEABLE
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# Sentinel: a label was present but its date/time value did not parse. This must
|
|
|
|
|
|
# FAIL the gate (present-but-unparseable), distinct from an absent value (None).
|
|
|
|
|
|
_UNPARSEABLE = "__UNPARSEABLE__"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
# Validation gate -- FAIL CLOSED
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def validate(candidate, template_id, email_data):
|
|
|
|
|
|
"""Return (True, 'ok') only if the candidate is provably conformant; else
|
|
|
|
|
|
(False, reason). Every rule must hold."""
|
|
|
|
|
|
# (1) known template
|
|
|
|
|
|
if template_id not in ("update_plaintext", "assign_html"):
|
|
|
|
|
|
return False, "subject_no_match"
|
|
|
|
|
|
|
|
|
|
|
|
# (2) keys EXACTLY the contract set
|
|
|
|
|
|
if set(candidate.keys()) != set(CONTRACT_KEYS):
|
|
|
|
|
|
return False, "key_set_mismatch"
|
|
|
|
|
|
|
|
|
|
|
|
# Unparseable date sentinels never survive.
|
|
|
|
|
|
for key in ("comment_time", "date_reported", "scheduled_start"):
|
|
|
|
|
|
if candidate.get(key) == _UNPARSEABLE:
|
|
|
|
|
|
return False, "creation_time_unparseable"
|
|
|
|
|
|
|
|
|
|
|
|
subject_wo, subject_site = _subject_ids(email_data)
|
|
|
|
|
|
body = email_data.get("body") or ""
|
|
|
|
|
|
|
|
|
|
|
|
# (3) work_order_id non-empty digits AND == subject id
|
|
|
|
|
|
wo = candidate.get("work_order_id")
|
|
|
|
|
|
if not wo or not _WO_ID_RE.match(str(wo)):
|
|
|
|
|
|
return False, "missing_required_field"
|
|
|
|
|
|
if wo != subject_wo:
|
|
|
|
|
|
return False, "wo_id_mismatch"
|
|
|
|
|
|
|
|
|
|
|
|
# (5) email_type in enum AND == template's expected type
|
|
|
|
|
|
expected_type = "comment" if template_id == "update_plaintext" else "new_work_order"
|
|
|
|
|
|
et = candidate.get("email_type")
|
|
|
|
|
|
if et not in VALID_EMAIL_TYPES:
|
|
|
|
|
|
return False, "missing_required_field"
|
|
|
|
|
|
if et != expected_type:
|
|
|
|
|
|
return False, "email_type_mismatch"
|
|
|
|
|
|
|
|
|
|
|
|
# (6) site_code if set matches the code pattern
|
|
|
|
|
|
site = candidate.get("site_code")
|
|
|
|
|
|
if site is not None and not _SITE_CODE_RE.match(str(site)):
|
|
|
|
|
|
return False, "malformed_site_code"
|
|
|
|
|
|
|
|
|
|
|
|
# (7) status if non-null in the enum
|
|
|
|
|
|
status = candidate.get("status")
|
|
|
|
|
|
if status is not None and status not in VALID_STATUSES:
|
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)
tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.
Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.
Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.
New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
'content' key, empty content list, non-JSON model text), asserting
PO's pre-call ai_fallback metric survives with no partial write and
the exception propagates; WO's no-datapoint-on-throttle behavior is
pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
zero writes, no raise — closing the hole where deleting the gate
line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
unset ARN and on a Secrets Manager exception, TTL cache refresh,
Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
no table scan, non-ASCII token, and a hostile-field-escaping
regression lock. PO web_ui has no __init__.py, so these go through
the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
(table 'WorkOrders'): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping,
None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
per-pipeline multi-record failure-isolation (all-or-retry contract).
The reprocess.py synthetic-event-shape contract test already landed
in Phase 7, so it isn't duplicated here.
Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.
_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.
CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.
docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.
* test: lock attribute-context quote escaping in web_ui hostile-field test
The escaping regression lock asserted only the element-context vector
(raw <script> absent, <script> present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).
Add assertions that the onclick sink's JSON string renders its opening
quote as " (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.
* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override
- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
corrected, instead of requiring a passing test to be knowingly deleted. The
module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
reusable workflow's bare pytest) red-exits local subset runs by design, with the
--cov-fail-under=0 override for iteration. Floor stays in addopts because the
centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00
|
|
|
|
return False, "invalid_status"
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
|
|
|
|
|
|
|
if template_id == "update_plaintext":
|
|
|
|
|
|
# (4) body must contain "Work Order: <id>" with LITERAL double space.
|
|
|
|
|
|
if f"Work Order: {subject_wo}" not in body:
|
|
|
|
|
|
return False, "single_space_work_order"
|
|
|
|
|
|
# (10) if Creation Time present it must have parsed.
|
|
|
|
|
|
if "Creation Time(UTC):" in body and candidate.get("comment_time") is None:
|
|
|
|
|
|
return False, "creation_time_unparseable"
|
|
|
|
|
|
# (8) comment needs work_order_id + non-empty comment_text
|
|
|
|
|
|
if not (candidate.get("comment_text") or "").strip():
|
|
|
|
|
|
return False, "missing_required_field"
|
|
|
|
|
|
# (9) label-bleed guard on comment_text
|
|
|
|
|
|
if _contains_label_or_separator(candidate.get("comment_text"), _T1_LABELS):
|
|
|
|
|
|
return False, "label_bleed"
|
|
|
|
|
|
if _contains_label_or_separator(candidate.get("address"), _T1_LABELS):
|
|
|
|
|
|
return False, "label_bleed"
|
|
|
|
|
|
else: # assign_html
|
|
|
|
|
|
# (4) body id if present must == subject
|
|
|
|
|
|
body_ids = re.findall(r"Work Order\D*(\d+)", body)
|
|
|
|
|
|
for bid in body_ids:
|
|
|
|
|
|
if bid != subject_wo:
|
|
|
|
|
|
return False, "wo_id_mismatch"
|
|
|
|
|
|
# (10) all five labels present and both dates parse.
|
|
|
|
|
|
for label in _T2_LABELS:
|
|
|
|
|
|
if label.lower() not in body.lower():
|
|
|
|
|
|
return False, "missing_required_field"
|
|
|
|
|
|
if candidate.get("date_reported") is None:
|
|
|
|
|
|
return False, "creation_time_unparseable"
|
|
|
|
|
|
if candidate.get("scheduled_start") is None:
|
|
|
|
|
|
return False, "creation_time_unparseable"
|
|
|
|
|
|
# (8) new_work_order needs work_order_id + site_code + non-empty description
|
|
|
|
|
|
if not candidate.get("site_code"):
|
|
|
|
|
|
return False, "missing_required_field"
|
|
|
|
|
|
if not (candidate.get("description") or "").strip():
|
|
|
|
|
|
return False, "missing_required_field"
|
|
|
|
|
|
# (9) label-bleed guard on description + address
|
|
|
|
|
|
if _contains_label_or_separator(candidate.get("description"), _T2_LABELS):
|
|
|
|
|
|
return False, "label_bleed"
|
|
|
|
|
|
if _contains_label_or_separator(candidate.get("address"), _T2_LABELS):
|
|
|
|
|
|
return False, "label_bleed"
|
|
|
|
|
|
|
|
|
|
|
|
return True, "ok"
|
|
|
|
|
|
|
|
|
|
|
|
|
fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path (#104)
* fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path
The ai_fallback parse path applied no validation gate to raw Bedrock/LLM
output before DynamoDB writes, and the extraction prompt concatenated the
untrusted email body directly with no instructions-vs-data delimiter. A
DKIM-passing attacker could prompt-inject arbitrary field values into the
work-order store.
Changes:
- wrap untrusted email in \<email\> XML block with prompt instructing the
model to treat its contents as data only
- add validate_ai_fallback() in template_parser that enforces the same
contract keys, enums, and patterns as the template path before any write
- call validate_ai_fallback() in handler() dispatch; emit an
ai_fallback_rejected EMF metric on failure and skip the record
- add 17 unit tests covering every gate rule and two end-to-end dispatch
tests (injected email_type, injected status)
Refs #101
* style: apply ruff formatting to fix CI check
* harden ai_fallback gate: review fixes + security-review findings
Review follow-up on the ai_fallback validation gate (PR #104), plus
findings from a fan-out /sh-security-review of the change surface.
Reviewer FIX items:
- Neutralize forged <email> delimiters in the untrusted body before
wrapping, so an in-body </email> cannot escape the data block.
- Fail closed on non-dict model output instead of crashing the handler
into async retries; count ai_fallback_rejected parses in the
fallback-rate alarm and add a dedicated rejected-parse alarm so a
gate-rejection drift outage is not silent.
- Return a distinct invalid_status reason (was malformed_site_code);
validate ISO-8601 dates; README + docstring updates.
Security-review findings (detector fan-out + proof-or-kill verifier):
- ReDoS (confirmed, medium): the tag neutralizer used two \s* around an
optional /, backtracking quadratically on "<" + a long whitespace run
(~32s at 100k chars -- one email could time out the Lambda). Collapse
to a single [\s/]* class: linear, same defanging.
- Unhashable-type crash (confirmed): a JSON list/dict for email_type or
status made `x in <set>` raise TypeError, escaping the gate into
retries. Guard with isinstance(str) before membership.
- Unicode/newline regex (confirmed): _WO_ID_RE/_SITE_CODE_RE used ^..$
with \d, admitting fullwidth digits ("12345" as a lookalike
partition key) and trailing newlines. Switch to \A[0-9]+\Z (and the
handler's inline recheck to [0-9]) so neither passes.
- Alarm comment (confirmed, low): corrected the "slow trickle still
pages" wording -- rejections >~25-30 min apart page on neither alarm,
the same knowingly-accepted residual as sender-auth-rejected.
Refuted: residual free-text prompt injection is inherent to trusting
allowlisted senders, not a new primitive; no DynamoDB key-poisoning
bypass survives both gates ('#' can never enter work_order_id).
7 new regression tests. All 260 tests pass; ruff clean; cdk synth OK.
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-16 16:23:28 -04:00
|
|
|
|
def validate_ai_fallback(candidate):
|
|
|
|
|
|
"""Fail-closed schema/enum validation for the AI-fallback parse path.
|
|
|
|
|
|
|
|
|
|
|
|
Called on the raw Bedrock/Claude output BEFORE any DynamoDB write. The
|
|
|
|
|
|
contract is a subset of the template-path validate(): no template-specific
|
|
|
|
|
|
rules (subject-id matching, label-bleed, required fields per template),
|
|
|
|
|
|
but ALL structural/enum guards apply so the AI path converges on the same
|
|
|
|
|
|
structural contract as the template path.
|
|
|
|
|
|
|
|
|
|
|
|
Returns (True, 'ok') or (False, reason)."""
|
|
|
|
|
|
from datetime import datetime
|
|
|
|
|
|
|
|
|
|
|
|
# json.loads on model output can yield any JSON type; only an object can
|
|
|
|
|
|
# satisfy the contract, and anything else must fail closed here rather
|
|
|
|
|
|
# than crash the handler into async S3 retries / DLQ.
|
|
|
|
|
|
if not isinstance(candidate, dict):
|
|
|
|
|
|
return False, "not_an_object"
|
|
|
|
|
|
|
|
|
|
|
|
# keys EXACTLY the contract set
|
|
|
|
|
|
if set(candidate.keys()) != set(CONTRACT_KEYS):
|
|
|
|
|
|
return False, "key_set_mismatch"
|
|
|
|
|
|
|
|
|
|
|
|
# work_order_id non-empty digits-only. The handler keeps its own check
|
|
|
|
|
|
# where the DynamoDB key is built (defense in depth); this gate is the
|
|
|
|
|
|
# single contract callers rely on.
|
|
|
|
|
|
wo = candidate.get("work_order_id")
|
|
|
|
|
|
if not wo or not _WO_ID_RE.match(str(wo)):
|
|
|
|
|
|
return False, "missing_required_field"
|
|
|
|
|
|
|
|
|
|
|
|
# email_type in enum. isinstance guard first: a non-str model value (JSON
|
|
|
|
|
|
# list/dict) is unhashable and `x in <set>` would raise, escaping the gate
|
|
|
|
|
|
# into async retries -- the opposite of fail-closed.
|
|
|
|
|
|
et = candidate.get("email_type")
|
|
|
|
|
|
if not isinstance(et, str) or et not in VALID_EMAIL_TYPES:
|
|
|
|
|
|
return False, "missing_required_field"
|
|
|
|
|
|
|
|
|
|
|
|
# status if non-null in the enum (same unhashable-type guard).
|
|
|
|
|
|
status = candidate.get("status")
|
|
|
|
|
|
if status is not None and (
|
|
|
|
|
|
not isinstance(status, str) or status not in VALID_STATUSES
|
|
|
|
|
|
):
|
|
|
|
|
|
return False, "invalid_status"
|
|
|
|
|
|
|
|
|
|
|
|
# site_code if non-null matches the code pattern
|
|
|
|
|
|
site = candidate.get("site_code")
|
|
|
|
|
|
if site is not None and not _SITE_CODE_RE.match(str(site)):
|
|
|
|
|
|
return False, "malformed_site_code"
|
|
|
|
|
|
|
|
|
|
|
|
# Date fields if non-null must be ISO-8601 strings (the prompt's declared
|
|
|
|
|
|
# format), converging with the template path's parsed-date guarantee and
|
|
|
|
|
|
# keeping the _UNPARSEABLE sentinel / arbitrary model prose out of the
|
|
|
|
|
|
# store.
|
|
|
|
|
|
for key in ("date_reported", "scheduled_start", "due_date", "comment_time"):
|
|
|
|
|
|
val = candidate.get(key)
|
|
|
|
|
|
if val is None:
|
|
|
|
|
|
continue
|
|
|
|
|
|
try:
|
|
|
|
|
|
datetime.fromisoformat(str(val))
|
|
|
|
|
|
except ValueError:
|
|
|
|
|
|
return False, "creation_time_unparseable"
|
|
|
|
|
|
|
2026-08-04 19:39:20 -04:00
|
|
|
|
# Free-text scalars must be None or str. A prompt-injected float reaches
|
|
|
|
|
|
# DynamoDB as Python float and crashes update_item; a dict/list lands as
|
|
|
|
|
|
# a Map/List attribute and pollutes downstream readers (PO parity).
|
|
|
|
|
|
for field in _AI_FREE_TEXT_STR_FIELDS:
|
|
|
|
|
|
val = candidate.get(field)
|
|
|
|
|
|
if val is not None and not isinstance(val, str):
|
|
|
|
|
|
return False, "invalid_field_type"
|
|
|
|
|
|
|
fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path (#104)
* fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path
The ai_fallback parse path applied no validation gate to raw Bedrock/LLM
output before DynamoDB writes, and the extraction prompt concatenated the
untrusted email body directly with no instructions-vs-data delimiter. A
DKIM-passing attacker could prompt-inject arbitrary field values into the
work-order store.
Changes:
- wrap untrusted email in \<email\> XML block with prompt instructing the
model to treat its contents as data only
- add validate_ai_fallback() in template_parser that enforces the same
contract keys, enums, and patterns as the template path before any write
- call validate_ai_fallback() in handler() dispatch; emit an
ai_fallback_rejected EMF metric on failure and skip the record
- add 17 unit tests covering every gate rule and two end-to-end dispatch
tests (injected email_type, injected status)
Refs #101
* style: apply ruff formatting to fix CI check
* harden ai_fallback gate: review fixes + security-review findings
Review follow-up on the ai_fallback validation gate (PR #104), plus
findings from a fan-out /sh-security-review of the change surface.
Reviewer FIX items:
- Neutralize forged <email> delimiters in the untrusted body before
wrapping, so an in-body </email> cannot escape the data block.
- Fail closed on non-dict model output instead of crashing the handler
into async retries; count ai_fallback_rejected parses in the
fallback-rate alarm and add a dedicated rejected-parse alarm so a
gate-rejection drift outage is not silent.
- Return a distinct invalid_status reason (was malformed_site_code);
validate ISO-8601 dates; README + docstring updates.
Security-review findings (detector fan-out + proof-or-kill verifier):
- ReDoS (confirmed, medium): the tag neutralizer used two \s* around an
optional /, backtracking quadratically on "<" + a long whitespace run
(~32s at 100k chars -- one email could time out the Lambda). Collapse
to a single [\s/]* class: linear, same defanging.
- Unhashable-type crash (confirmed): a JSON list/dict for email_type or
status made `x in <set>` raise TypeError, escaping the gate into
retries. Guard with isinstance(str) before membership.
- Unicode/newline regex (confirmed): _WO_ID_RE/_SITE_CODE_RE used ^..$
with \d, admitting fullwidth digits ("12345" as a lookalike
partition key) and trailing newlines. Switch to \A[0-9]+\Z (and the
handler's inline recheck to [0-9]) so neither passes.
- Alarm comment (confirmed, low): corrected the "slow trickle still
pages" wording -- rejections >~25-30 min apart page on neither alarm,
the same knowingly-accepted residual as sender-auth-rejected.
Refuted: residual free-text prompt injection is inherent to trusting
allowlisted senders, not a new primitive; no DynamoDB key-poisoning
bypass survives both gates ('#' can never enter work_order_id).
7 new regression tests. All 260 tests pass; ruff clean; cdk synth OK.
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-16 16:23:28 -04:00
|
|
|
|
return True, "ok"
|
|
|
|
|
|
|
|
|
|
|
|
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
|
def try_deterministic_parse(email_data):
|
|
|
|
|
|
"""Entry point. Returns (parsed|None, parse_method, template_id, reason).
|
|
|
|
|
|
|
|
|
|
|
|
On a proven-conformant parse returns (dict, 'template', template_id, 'ok').
|
|
|
|
|
|
On any miss/invalid/exception returns (None, 'ai_fallback', template_id,
|
|
|
|
|
|
reason) -- a failure is NEVER a parsed result."""
|
|
|
|
|
|
template_id = "unknown"
|
|
|
|
|
|
try:
|
|
|
|
|
|
template_id, reason = classify_template(email_data)
|
|
|
|
|
|
if template_id == "unknown":
|
|
|
|
|
|
return None, "ai_fallback", template_id, reason
|
|
|
|
|
|
|
|
|
|
|
|
if template_id == "update_plaintext":
|
|
|
|
|
|
candidate = extract_update_plaintext(email_data)
|
|
|
|
|
|
else:
|
|
|
|
|
|
candidate = extract_assign_html(email_data)
|
|
|
|
|
|
|
|
|
|
|
|
ok, reason = validate(candidate, template_id, email_data)
|
|
|
|
|
|
if not ok:
|
|
|
|
|
|
return None, "ai_fallback", template_id, reason
|
|
|
|
|
|
return candidate, "template", template_id, "ok"
|
|
|
|
|
|
except Exception: # noqa: BLE001 -- fail closed on ANY extractor error
|
|
|
|
|
|
return None, "ai_fallback", template_id, "extractor_raised"
|