procurement-ingest/cdk/po_stack.py

702 lines
32 KiB
Python
Raw Normal View History

"""CDK stack for the Coupa PO email ingestion pipeline."""
import aws_cdk as cdk
from aws_cdk import (
Duration,
RemovalPolicy,
Stack,
aws_cloudwatch as cloudwatch,
aws_cloudwatch_actions as cw_actions,
aws_dynamodb as dynamodb,
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
aws_iam as iam,
aws_kms as kms,
aws_lambda as lambda_,
aws_lambda_event_sources as lambda_event_sources,
aws_logs as logs,
aws_s3 as s3,
aws_s3_notifications as s3n,
aws_ses as ses,
aws_ses_actions as ses_actions,
aws_secretsmanager as secretsmanager,
aws_sns as sns,
aws_sqs as sqs,
aws_ssm as ssm,
)
from constructs import Construct
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70) * Add CloudWatch alarm coverage for po-ingest and workorder-ingest Expands alarm coverage across both CDK stacks. All alarms are ALARM-only (no OK action) to the shared site-alerts SNS topic, with TreatMissingData NOT_BREACHING. The site-alerts topic is now imported once near the top of each stack so every alarm reuses one Topic instance. po-ingest (cdk/po_stack.py): - Errors: po-ingest-site-extractor - Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui - Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption), po-ingest-site-extractor, po-web-ui - DynamoDB throttle + system-error: purchase-orders, verified-sites, pending-site-review workorder-ingest (cdk/wo_stack.py): - Throttles: workorder-email-processor - Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption) - DynamoDB throttle + system-error: WorkOrders, WorkOrderComments DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation dimension set, so each table alarm is a Sum math expression across operations via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.259.0). Refs INFRA-41 / audit H-8. * Drop NEEDS ADAM SIGN-OFF wording from alarm comments Duration alarm thresholds are owner-approved; remove the sign-off flag from po_stack.py and wo_stack.py comments. Threshold values, eval config, and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
# Operations these tables actually issue (PutItem/UpdateItem/DeleteItem writes,
# GetItem/Query/BatchGetItem reads). DynamoDB emits ThrottledRequests/SystemErrors
# keyed by TableName + Operation only, so the CDK *_for_operations helpers (which
# render a SUM MathExpression across these per-operation metrics) are the correct,
# non-deprecated way to roll a table up to a single alarmable series.
_DDB_ALARM_OPERATIONS = [
dynamodb.Operation.GET_ITEM,
dynamodb.Operation.BATCH_GET_ITEM,
dynamodb.Operation.QUERY,
dynamodb.Operation.SCAN,
dynamodb.Operation.PUT_ITEM,
dynamodb.Operation.UPDATE_ITEM,
dynamodb.Operation.DELETE_ITEM,
dynamodb.Operation.BATCH_WRITE_ITEM,
]
def _add_ddb_alarms(scope, id_prefix, table, alarm_name_prefix, alarm_topic):
"""Add throttle + system-error alarms for a DynamoDB table.
Both fire on any non-zero datapoint in a 5-min window. ALARM-only SnsAction
to site-alerts (no OK action); TreatMissingData NOT_BREACHING.
"""
table.metric_throttled_requests_for_operations(
operations=_DDB_ALARM_OPERATIONS,
period=Duration.minutes(5),
statistic="Sum",
).create_alarm(
scope,
f"{id_prefix}ThrottlesAlarm",
alarm_name=f"{alarm_name_prefix}-throttles",
alarm_description=f"{alarm_name_prefix} DynamoDB throttled requests",
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
table.metric_system_errors_for_operations(
operations=_DDB_ALARM_OPERATIONS,
period=Duration.minutes(5),
statistic="Sum",
).create_alarm(
scope,
f"{id_prefix}SystemErrorsAlarm",
alarm_name=f"{alarm_name_prefix}-system-errors",
alarm_description=f"{alarm_name_prefix} DynamoDB server-side (5xx) errors",
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
Add fail-closed SES sender authentication (INFRA-107) (#98) * Add fail-closed SES sender authentication The From header and any raw-MIME Authentication-Results copies are attacker-forgeable, so a forged email to apm@int.seahaven.com or amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107, CRITICAL). Both S3-triggered email processors now authenticate the sender against the Authentication-Results header SES itself prepends at delivery: only the topmost header is consulted, its authserv-id must be amazonses.com, and it must carry dkim=pass for a domain in the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set in CDK so ops can adjust without code changes). Allowlists come from live traffic observed 2026-07-15 on both ingest buckets: WO mail arrives via the apm@ Google Groups forward, which re-signs as seahaven.com (the hxgnsmartcloud.com signature does not survive the forward); PO mail passes for amazon.coupahost.com. amazonses.com also passes on PO mail but is deliberately excluded -- every SES customer's outbound mail passes for it. Every failure path (env var unset, header missing or unparseable, verdict fail, unaligned domain) rejects the email: a structured warning with the reason and S3 key is logged and the record skipped without erroring the invocation, so rejected mail causes no Lambda retries or DLQ messages. Handler signatures and event sources are unchanged. Refs: INFRA-107 * Harden AR parser per cross-family review Cross-family (GPT-4.1) review findings: terminate the dkim result token at end-of-clause, whitespace, or a comment so a value like "dkim=pass-fake" can never be read as a pass; normalize trailing dots off allowlist entries so "seahaven.com." matches; make the compat32 parser policy explicit. Adds tests for result-token boundaries, comments after the result, quoted domain values, and folding inside a dkim clause. Refs: INFRA-107 * Harden AR parsing and alarm on sender-auth rejects The SES-stamped Authentication-Results value echoes attacker-controlled SMTP-session tokens (envelope-from, helo, header.from) as their own semicolon-delimited property clauses. A naive split(";") tore an RFC 5321 quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass clause, so a fully spoofed email was accepted on the genuinely SES-stamped topmost header. Tokenise comment- and quoted-string-aware (RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on semicolons outside a quoted-string, and fail closed on unbalanced quotes/comments so a ';' inside a quoted pvalue can never start a clause. Rejected mail returns normally (no error, no retry, no DLQ message), so a signing-domain drift or a wrong allowlist would silently discard 100% of legitimate mail while every alarm stayed green. Add a CloudWatch Logs metric filter + alarm on the sender_auth_rejected warning to both stacks so a false-reject storm pages instead of vanishing. This is also the safety net for the WO seahaven.com allowlist assumption, which must be validated against a live SES-stamped header (a plain Gmail auto-forward re-signs under the sending Workspace domain, not seahaven.com). Refs: INFRA-107 * chore: retrigger CI (no run recorded for 7c74ac1) * Fix quoted-AUID DKIM domain spoof in sender auth Resolve three confirmed /sh-security-review findings on the fail-closed SES sender-authentication control. HIGH: header.i/header.d domain extraction was not quoted-string aware. An attacker with a valid DKIM key for their own domain could set an RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive extractor stopped at the closing quote and returned seahaven.com, accepting forged mail. Extraction now tokenises the clause with the same quoted-string discipline already used for clause splitting: header.d (the plain signing domain) is authoritative when present, otherwise the header.i domain is the part after the AUID's LAST top-level "@", so a "@" inside a quoted local-part is treated as signer-controlled label text and yields the true signer (attacker.com), not seahaven.com. LOW: the topmost-header parse ran outside evaluate_sender_authentication's try/except, so an unexpected parser exception on crafted input could propagate into the handler and Lambda async retries/DLQ. The parse now fails CLOSED with an authentication_results_unparseable reason. MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to a low-volume total-reject outage (a trickle that never sums to 3). Both stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3 / datapoints_to_alarm=2, so a sustained reject condition pages even at one reject per period while a lone stray probe self-clears. Refs: INFRA-107 * Load Lambda function dir on sys.path in tests Rebasing INFRA-107 onto main folded #95's pytest suite into this branch's tests. The unified conftest loads the PO/WO handlers by file path, and handler.py now does `from ses_auth import authenticate_inbound_email` -- a bare sibling import that resolves in the Lambda only because the runtime puts each function's own directory on sys.path. The shared load_handler now adds that directory so the handler tests import correctly alongside the sender-auth tests. Refs: INFRA-107 * Note #97 test files in README directory tree The rebase onto main brought in #97's tests/requirements.txt and tests/test_po_merge.py. List both in the directory tree so it matches the tree on disk. Refs: INFRA-107 * Document INFRA-107 forwarder-binding risk acceptance Record the accepted risk that WO sender auth binds to the apm@ forward's re-signing domain (seahaven.com) rather than the Hexagon originator; the apm@ Google Group's restricted posting policy is the load-bearing control (escalates to HIGH if the group is opened to external posting). Also correct the sender-auth-rejected alarm docs to match the shipped config (>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and note the SES-AR-01/02 parser hardening follow-ups. Refs: INFRA-107
2026-07-15 20:58:47 -04:00
# CloudWatch namespace for the log-derived sender-authentication metrics.
_SENDER_AUTH_METRIC_NAMESPACE = "Seahaven/ProcurementIngest"
def _add_sender_auth_rejected_alarm(scope, id_prefix, function_name, alarm_topic):
"""Metric-filter + alarm on ``sender_auth_rejected`` warnings (INFRA-107).
A rejected inbound email is skipped without erroring the invocation, so it
is invisible to the Errors/Throttles/DLQ alarms. This turns the structured
warning log into a CloudWatch metric and pages when rejections spike --
catching a silent false-reject storm (allowlist wrong, signing-domain
drift, SES header-format change) that would otherwise discard legitimate
mail while the pipeline reports healthy.
ALARM-only SnsAction to site-alerts; no OK action. The metric filter reads
the function's own log group (imported by the deterministic
``/aws/lambda/<fn>`` name, created by the function's log_retention). A plain
substring pattern is used because Lambda prefixes each line with its own
level/timestamp/request-id, so the JSON payload is not a standalone JSON
log event a `{$.event=...}` pattern could match.
"""
metric_name = f"{function_name}-sender-auth-rejected"
logs.MetricFilter(
scope,
f"{id_prefix}SenderAuthRejectedFilter",
log_group=logs.LogGroup.from_log_group_name(
scope,
f"{id_prefix}LogGroup",
f"/aws/lambda/{function_name}",
),
filter_pattern=logs.FilterPattern.literal('"sender_auth_rejected"'),
metric_namespace=_SENDER_AUTH_METRIC_NAMESPACE,
metric_name=metric_name,
metric_value="1",
default_value=0,
)
# Fire on a *sustained* reject condition rather than a volume spike. The
# earlier Sum>=3-over-15-min threshold had a blind spot that is exactly the
# failure this alarm exists to catch: a low-traffic pipeline in total
# drift outage (allowlist wrong / signing-domain changed) may only produce
# a trickle of rejects -- one every few minutes -- that never sums to 3 in
# any window, so the outage never pages. Instead: >=1 reject per 5-min
# period, alarming when 2 of the last 6 periods breach (evaluation_periods=6
# / datapoints_to_alarm=2). Six periods (30 min) with only 2 required
# datapoints closes the sparse-outage residual: even rejections >10-15 min
# apart can still place two breaching datapoints in a single 30-min
# evaluation window. A single stray spoof probe (one lone period) is
# tolerated and self-clears, but a sustained reject condition trips even
# at very low arrival rates. default_value=0 on the metric filter keeps the
# series continuous so NOT_BREACHING only applies before the first datapoint
# ever arrives.
Add fail-closed SES sender authentication (INFRA-107) (#98) * Add fail-closed SES sender authentication The From header and any raw-MIME Authentication-Results copies are attacker-forgeable, so a forged email to apm@int.seahaven.com or amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107, CRITICAL). Both S3-triggered email processors now authenticate the sender against the Authentication-Results header SES itself prepends at delivery: only the topmost header is consulted, its authserv-id must be amazonses.com, and it must carry dkim=pass for a domain in the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set in CDK so ops can adjust without code changes). Allowlists come from live traffic observed 2026-07-15 on both ingest buckets: WO mail arrives via the apm@ Google Groups forward, which re-signs as seahaven.com (the hxgnsmartcloud.com signature does not survive the forward); PO mail passes for amazon.coupahost.com. amazonses.com also passes on PO mail but is deliberately excluded -- every SES customer's outbound mail passes for it. Every failure path (env var unset, header missing or unparseable, verdict fail, unaligned domain) rejects the email: a structured warning with the reason and S3 key is logged and the record skipped without erroring the invocation, so rejected mail causes no Lambda retries or DLQ messages. Handler signatures and event sources are unchanged. Refs: INFRA-107 * Harden AR parser per cross-family review Cross-family (GPT-4.1) review findings: terminate the dkim result token at end-of-clause, whitespace, or a comment so a value like "dkim=pass-fake" can never be read as a pass; normalize trailing dots off allowlist entries so "seahaven.com." matches; make the compat32 parser policy explicit. Adds tests for result-token boundaries, comments after the result, quoted domain values, and folding inside a dkim clause. Refs: INFRA-107 * Harden AR parsing and alarm on sender-auth rejects The SES-stamped Authentication-Results value echoes attacker-controlled SMTP-session tokens (envelope-from, helo, header.from) as their own semicolon-delimited property clauses. A naive split(";") tore an RFC 5321 quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass clause, so a fully spoofed email was accepted on the genuinely SES-stamped topmost header. Tokenise comment- and quoted-string-aware (RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on semicolons outside a quoted-string, and fail closed on unbalanced quotes/comments so a ';' inside a quoted pvalue can never start a clause. Rejected mail returns normally (no error, no retry, no DLQ message), so a signing-domain drift or a wrong allowlist would silently discard 100% of legitimate mail while every alarm stayed green. Add a CloudWatch Logs metric filter + alarm on the sender_auth_rejected warning to both stacks so a false-reject storm pages instead of vanishing. This is also the safety net for the WO seahaven.com allowlist assumption, which must be validated against a live SES-stamped header (a plain Gmail auto-forward re-signs under the sending Workspace domain, not seahaven.com). Refs: INFRA-107 * chore: retrigger CI (no run recorded for 7c74ac1) * Fix quoted-AUID DKIM domain spoof in sender auth Resolve three confirmed /sh-security-review findings on the fail-closed SES sender-authentication control. HIGH: header.i/header.d domain extraction was not quoted-string aware. An attacker with a valid DKIM key for their own domain could set an RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive extractor stopped at the closing quote and returned seahaven.com, accepting forged mail. Extraction now tokenises the clause with the same quoted-string discipline already used for clause splitting: header.d (the plain signing domain) is authoritative when present, otherwise the header.i domain is the part after the AUID's LAST top-level "@", so a "@" inside a quoted local-part is treated as signer-controlled label text and yields the true signer (attacker.com), not seahaven.com. LOW: the topmost-header parse ran outside evaluate_sender_authentication's try/except, so an unexpected parser exception on crafted input could propagate into the handler and Lambda async retries/DLQ. The parse now fails CLOSED with an authentication_results_unparseable reason. MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to a low-volume total-reject outage (a trickle that never sums to 3). Both stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3 / datapoints_to_alarm=2, so a sustained reject condition pages even at one reject per period while a lone stray probe self-clears. Refs: INFRA-107 * Load Lambda function dir on sys.path in tests Rebasing INFRA-107 onto main folded #95's pytest suite into this branch's tests. The unified conftest loads the PO/WO handlers by file path, and handler.py now does `from ses_auth import authenticate_inbound_email` -- a bare sibling import that resolves in the Lambda only because the runtime puts each function's own directory on sys.path. The shared load_handler now adds that directory so the handler tests import correctly alongside the sender-auth tests. Refs: INFRA-107 * Note #97 test files in README directory tree The rebase onto main brought in #97's tests/requirements.txt and tests/test_po_merge.py. List both in the directory tree so it matches the tree on disk. Refs: INFRA-107 * Document INFRA-107 forwarder-binding risk acceptance Record the accepted risk that WO sender auth binds to the apm@ forward's re-signing domain (seahaven.com) rather than the Hexagon originator; the apm@ Google Group's restricted posting policy is the load-bearing control (escalates to HIGH if the group is opened to external posting). Also correct the sender-auth-rejected alarm docs to match the shipped config (>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and note the SES-AR-01/02 parser hardening follow-ups. Refs: INFRA-107
2026-07-15 20:58:47 -04:00
cloudwatch.Metric(
namespace=_SENDER_AUTH_METRIC_NAMESPACE,
metric_name=metric_name,
period=Duration.minutes(5),
statistic="Sum",
).create_alarm(
scope,
f"{id_prefix}SenderAuthRejectedAlarm",
alarm_name=f"{function_name}-sender-auth-rejected",
alarm_description=(
f"{function_name} rejected inbound mail on sender authentication "
"(possible allowlist/DKIM-domain drift silently dropping real mail)"
),
threshold=1,
evaluation_periods=6,
Add fail-closed SES sender authentication (INFRA-107) (#98) * Add fail-closed SES sender authentication The From header and any raw-MIME Authentication-Results copies are attacker-forgeable, so a forged email to apm@int.seahaven.com or amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107, CRITICAL). Both S3-triggered email processors now authenticate the sender against the Authentication-Results header SES itself prepends at delivery: only the topmost header is consulted, its authserv-id must be amazonses.com, and it must carry dkim=pass for a domain in the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set in CDK so ops can adjust without code changes). Allowlists come from live traffic observed 2026-07-15 on both ingest buckets: WO mail arrives via the apm@ Google Groups forward, which re-signs as seahaven.com (the hxgnsmartcloud.com signature does not survive the forward); PO mail passes for amazon.coupahost.com. amazonses.com also passes on PO mail but is deliberately excluded -- every SES customer's outbound mail passes for it. Every failure path (env var unset, header missing or unparseable, verdict fail, unaligned domain) rejects the email: a structured warning with the reason and S3 key is logged and the record skipped without erroring the invocation, so rejected mail causes no Lambda retries or DLQ messages. Handler signatures and event sources are unchanged. Refs: INFRA-107 * Harden AR parser per cross-family review Cross-family (GPT-4.1) review findings: terminate the dkim result token at end-of-clause, whitespace, or a comment so a value like "dkim=pass-fake" can never be read as a pass; normalize trailing dots off allowlist entries so "seahaven.com." matches; make the compat32 parser policy explicit. Adds tests for result-token boundaries, comments after the result, quoted domain values, and folding inside a dkim clause. Refs: INFRA-107 * Harden AR parsing and alarm on sender-auth rejects The SES-stamped Authentication-Results value echoes attacker-controlled SMTP-session tokens (envelope-from, helo, header.from) as their own semicolon-delimited property clauses. A naive split(";") tore an RFC 5321 quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass clause, so a fully spoofed email was accepted on the genuinely SES-stamped topmost header. Tokenise comment- and quoted-string-aware (RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on semicolons outside a quoted-string, and fail closed on unbalanced quotes/comments so a ';' inside a quoted pvalue can never start a clause. Rejected mail returns normally (no error, no retry, no DLQ message), so a signing-domain drift or a wrong allowlist would silently discard 100% of legitimate mail while every alarm stayed green. Add a CloudWatch Logs metric filter + alarm on the sender_auth_rejected warning to both stacks so a false-reject storm pages instead of vanishing. This is also the safety net for the WO seahaven.com allowlist assumption, which must be validated against a live SES-stamped header (a plain Gmail auto-forward re-signs under the sending Workspace domain, not seahaven.com). Refs: INFRA-107 * chore: retrigger CI (no run recorded for 7c74ac1) * Fix quoted-AUID DKIM domain spoof in sender auth Resolve three confirmed /sh-security-review findings on the fail-closed SES sender-authentication control. HIGH: header.i/header.d domain extraction was not quoted-string aware. An attacker with a valid DKIM key for their own domain could set an RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive extractor stopped at the closing quote and returned seahaven.com, accepting forged mail. Extraction now tokenises the clause with the same quoted-string discipline already used for clause splitting: header.d (the plain signing domain) is authoritative when present, otherwise the header.i domain is the part after the AUID's LAST top-level "@", so a "@" inside a quoted local-part is treated as signer-controlled label text and yields the true signer (attacker.com), not seahaven.com. LOW: the topmost-header parse ran outside evaluate_sender_authentication's try/except, so an unexpected parser exception on crafted input could propagate into the handler and Lambda async retries/DLQ. The parse now fails CLOSED with an authentication_results_unparseable reason. MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to a low-volume total-reject outage (a trickle that never sums to 3). Both stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3 / datapoints_to_alarm=2, so a sustained reject condition pages even at one reject per period while a lone stray probe self-clears. Refs: INFRA-107 * Load Lambda function dir on sys.path in tests Rebasing INFRA-107 onto main folded #95's pytest suite into this branch's tests. The unified conftest loads the PO/WO handlers by file path, and handler.py now does `from ses_auth import authenticate_inbound_email` -- a bare sibling import that resolves in the Lambda only because the runtime puts each function's own directory on sys.path. The shared load_handler now adds that directory so the handler tests import correctly alongside the sender-auth tests. Refs: INFRA-107 * Note #97 test files in README directory tree The rebase onto main brought in #97's tests/requirements.txt and tests/test_po_merge.py. List both in the directory tree so it matches the tree on disk. Refs: INFRA-107 * Document INFRA-107 forwarder-binding risk acceptance Record the accepted risk that WO sender auth binds to the apm@ forward's re-signing domain (seahaven.com) rather than the Hexagon originator; the apm@ Google Group's restricted posting policy is the load-bearing control (escalates to HIGH if the group is opened to external posting). Also correct the sender-auth-rejected alarm docs to match the shipped config (>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and note the SES-AR-01/02 parser hardening follow-ups. Refs: INFRA-107
2026-07-15 20:58:47 -04:00
datapoints_to_alarm=2,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
class PoIngestStack(Stack):
def __init__(self, scope: Construct, construct_id: str, **kwargs):
super().__init__(scope, construct_id, **kwargs)
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70) * Add CloudWatch alarm coverage for po-ingest and workorder-ingest Expands alarm coverage across both CDK stacks. All alarms are ALARM-only (no OK action) to the shared site-alerts SNS topic, with TreatMissingData NOT_BREACHING. The site-alerts topic is now imported once near the top of each stack so every alarm reuses one Topic instance. po-ingest (cdk/po_stack.py): - Errors: po-ingest-site-extractor - Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui - Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption), po-ingest-site-extractor, po-web-ui - DynamoDB throttle + system-error: purchase-orders, verified-sites, pending-site-review workorder-ingest (cdk/wo_stack.py): - Throttles: workorder-email-processor - Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption) - DynamoDB throttle + system-error: WorkOrders, WorkOrderComments DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation dimension set, so each table alarm is a Sum math expression across operations via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.259.0). Refs INFRA-41 / audit H-8. * Drop NEEDS ADAM SIGN-OFF wording from alarm comments Duration alarm thresholds are owner-approved; remove the sign-off flag from po_stack.py and wo_stack.py comments. Threshold values, eval config, and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
# --- Shared alarm SNS topic (site-alerts) ---
# Imported once near the top so every alarm in this stack reuses the same
# Topic construct instance (avoids duplicate logical IDs). ALARM-only
# SnsAction; no OK action, per the CloudWatch-alarm preference. The
# topic's CMK (alias/seahaven-alarm-topics) lives on the topic itself.
alarm_topic = sns.Topic.from_topic_arn(
self,
"SiteAlertsTopic",
f"arn:aws:sns:{self.region}:{self.account}:site-alerts",
)
# --- S3 bucket for raw emails ---
email_bucket = s3.Bucket(
self,
"EmailBucket",
bucket_name=f"po-ingest-emails-{self.account}",
block_public_access=s3.BlockPublicAccess.BLOCK_ALL,
removal_policy=RemovalPolicy.RETAIN,
lifecycle_rules=[
s3.LifecycleRule(expiration=Duration.days(90)),
],
)
# --- Shared customer-managed CMK for sensitive DynamoDB tables ---
# Owned by the account-baseline app (alias/seahaven-dynamodb, INFRA-95 /
# M-3); ARN published to SSM. The purchase-orders table was migrated to
# SSE-KMS out-of-band, so declaring encryption_key here reconciles the
# drift and — via grant_read_write_data below — propagates the required
# kms:Decrypt/GenerateDataKey/DescribeKey to the consumer roles.
dynamodb_cmk = kms.Key.from_key_arn(
self,
"DynamoDbCmk",
ssm.StringParameter.value_for_string_parameter(
self, "/seahaven/dynamodb/cmk-arn"
),
)
# --- Purchase-orders DynamoDB table ---
# Owned by this stack. Streams enabled for the site-extractor pipeline.
# Other stacks (seahaven-slack-bot) reference this table via fromTableName().
po_table = dynamodb.Table(
self,
"PurchaseOrdersTable",
table_name="purchase-orders",
partition_key=dynamodb.Attribute(
name="po_number",
type=dynamodb.AttributeType.STRING,
),
billing_mode=dynamodb.BillingMode.PAY_PER_REQUEST,
removal_policy=RemovalPolicy.RETAIN,
stream=dynamodb.StreamViewType.NEW_IMAGE,
encryption=dynamodb.TableEncryption.CUSTOMER_MANAGED,
encryption_key=dynamodb_cmk,
)
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
# --- Anthropic API key secret removed (Bedrock migration) ---
# PO parsing stays fully AI but moved from the Anthropic API to the
# Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0,
# so no provider API key is needed. The old secret
# "po-ingest/anthropic-api-key" had RemovalPolicy.RETAIN, so it is
# ORPHANED (not deleted) by this change: delete it manually post-deploy
# and revoke the stored key at Anthropic.
# --- DLQ for failed async invocations (INFRA-41 / audit H-8) ---
# SES → S3 → Lambda is async; without an OnFailure destination a failed
# parse (bad email, transient error) is silently dropped after Lambda's
# retries. CDK generates the queue name to avoid colliding with the
# interim CLI-created po-email-processor-dlq (removed post-deploy).
email_processor_dlq = sqs.Queue(
self,
"EmailProcessorDlq",
retention_period=Duration.days(14),
enforce_ssl=True,
)
# --- Lambda function ---
email_processor = lambda_.Function(
self,
"EmailProcessor",
function_name="po-email-processor",
runtime=lambda_.Runtime.PYTHON_3_12,
architecture=lambda_.Architecture.ARM_64,
handler="handler.handler",
code=lambda_.Code.from_asset(
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
"../lambdas/po/email_processor",
bundling=cdk.BundlingOptions(
image=lambda_.Runtime.PYTHON_3_12.bundling_image,
command=[
"bash",
"-c",
"pip install --platform manylinux2014_aarch64 --only-binary=:all: "
"-r requirements.txt -t /asset-output && "
feat: Python derived-field classifier with shadow telemetry (PO PR 2) (#106) * feat: Python derived-field classifier with shadow telemetry for PO ingest Port the site_code/trade/fiscal_year rules from EXTRACTION_PROMPT into a pure, total derived_fields module applied in the shared enrich_parsed() post-stage. Python fills gaps on both parse paths (the template path has no LLM values, closing the derived-field gap opened by the PR #105 two-PR split) and never overwrites a non-null LLM value; on ai_fallback a DerivedFieldAgreement EMF record per field shadows Python against the LLM during the bake. Rules hardened against a full-corpus backtest (3,422 real emails vs the LLM-written baseline): site_code 99.4% with zero Python-wrong cases, fiscal_year 100%, trade 96.8% ex-deliberate. Also: quantity/price now declared numeric in the prompt, and derived_fields.py added to the po_stack bundling copy (deploy-time ImportError otherwise). * fix: security-review hardening — EMF value length clamp, aggregate trade CPU budget sh-security-review (4 detectors + proof-or-kill verifier): PASS, 0 confirmed critical/high. Fixes the one confirmed low (unbounded LLM-value str() into the DerivedFieldAgreement EMF log line, clamped to 64 chars) and adds the verifier-recommended defense-in-depth aggregate character budget across line items in derive_trade (per-item caps alone allowed ~10s full-core on a pathological direct-call input; unreachable through the deployed handler but cheap to bound). Bundling cp list now carries a warning comment (GPT-4.1 cross-review FIX).
2026-07-17 11:47:33 -04:00
# NOTE: every module handler.py imports as a sibling
# MUST be listed here or the deploy ships a Lambda that
# ImportErrors at runtime (bit us for template_parser
# in PR #105 and nearly for derived_fields in PR #2).
"cp handler.py ses_auth.py template_parser.py "
"derived_fields.py /asset-output/",
],
),
),
timeout=Duration.seconds(60),
memory_size=256,
log_retention=logs.RetentionDays.TWO_MONTHS,
dead_letter_queue=email_processor_dlq,
environment={
"PO_TABLE": "purchase-orders",
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
"BEDROCK_MODEL_ID": "us.anthropic.claude-haiku-4-5-20251001-v1:0",
Add fail-closed SES sender authentication (INFRA-107) (#98) * Add fail-closed SES sender authentication The From header and any raw-MIME Authentication-Results copies are attacker-forgeable, so a forged email to apm@int.seahaven.com or amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107, CRITICAL). Both S3-triggered email processors now authenticate the sender against the Authentication-Results header SES itself prepends at delivery: only the topmost header is consulted, its authserv-id must be amazonses.com, and it must carry dkim=pass for a domain in the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set in CDK so ops can adjust without code changes). Allowlists come from live traffic observed 2026-07-15 on both ingest buckets: WO mail arrives via the apm@ Google Groups forward, which re-signs as seahaven.com (the hxgnsmartcloud.com signature does not survive the forward); PO mail passes for amazon.coupahost.com. amazonses.com also passes on PO mail but is deliberately excluded -- every SES customer's outbound mail passes for it. Every failure path (env var unset, header missing or unparseable, verdict fail, unaligned domain) rejects the email: a structured warning with the reason and S3 key is logged and the record skipped without erroring the invocation, so rejected mail causes no Lambda retries or DLQ messages. Handler signatures and event sources are unchanged. Refs: INFRA-107 * Harden AR parser per cross-family review Cross-family (GPT-4.1) review findings: terminate the dkim result token at end-of-clause, whitespace, or a comment so a value like "dkim=pass-fake" can never be read as a pass; normalize trailing dots off allowlist entries so "seahaven.com." matches; make the compat32 parser policy explicit. Adds tests for result-token boundaries, comments after the result, quoted domain values, and folding inside a dkim clause. Refs: INFRA-107 * Harden AR parsing and alarm on sender-auth rejects The SES-stamped Authentication-Results value echoes attacker-controlled SMTP-session tokens (envelope-from, helo, header.from) as their own semicolon-delimited property clauses. A naive split(";") tore an RFC 5321 quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass clause, so a fully spoofed email was accepted on the genuinely SES-stamped topmost header. Tokenise comment- and quoted-string-aware (RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on semicolons outside a quoted-string, and fail closed on unbalanced quotes/comments so a ';' inside a quoted pvalue can never start a clause. Rejected mail returns normally (no error, no retry, no DLQ message), so a signing-domain drift or a wrong allowlist would silently discard 100% of legitimate mail while every alarm stayed green. Add a CloudWatch Logs metric filter + alarm on the sender_auth_rejected warning to both stacks so a false-reject storm pages instead of vanishing. This is also the safety net for the WO seahaven.com allowlist assumption, which must be validated against a live SES-stamped header (a plain Gmail auto-forward re-signs under the sending Workspace domain, not seahaven.com). Refs: INFRA-107 * chore: retrigger CI (no run recorded for 7c74ac1) * Fix quoted-AUID DKIM domain spoof in sender auth Resolve three confirmed /sh-security-review findings on the fail-closed SES sender-authentication control. HIGH: header.i/header.d domain extraction was not quoted-string aware. An attacker with a valid DKIM key for their own domain could set an RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive extractor stopped at the closing quote and returned seahaven.com, accepting forged mail. Extraction now tokenises the clause with the same quoted-string discipline already used for clause splitting: header.d (the plain signing domain) is authoritative when present, otherwise the header.i domain is the part after the AUID's LAST top-level "@", so a "@" inside a quoted local-part is treated as signer-controlled label text and yields the true signer (attacker.com), not seahaven.com. LOW: the topmost-header parse ran outside evaluate_sender_authentication's try/except, so an unexpected parser exception on crafted input could propagate into the handler and Lambda async retries/DLQ. The parse now fails CLOSED with an authentication_results_unparseable reason. MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to a low-volume total-reject outage (a trickle that never sums to 3). Both stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3 / datapoints_to_alarm=2, so a sustained reject condition pages even at one reject per period while a lone stray probe self-clears. Refs: INFRA-107 * Load Lambda function dir on sys.path in tests Rebasing INFRA-107 onto main folded #95's pytest suite into this branch's tests. The unified conftest loads the PO/WO handlers by file path, and handler.py now does `from ses_auth import authenticate_inbound_email` -- a bare sibling import that resolves in the Lambda only because the runtime puts each function's own directory on sys.path. The shared load_handler now adds that directory so the handler tests import correctly alongside the sender-auth tests. Refs: INFRA-107 * Note #97 test files in README directory tree The rebase onto main brought in #97's tests/requirements.txt and tests/test_po_merge.py. List both in the directory tree so it matches the tree on disk. Refs: INFRA-107 * Document INFRA-107 forwarder-binding risk acceptance Record the accepted risk that WO sender auth binds to the apm@ forward's re-signing domain (seahaven.com) rather than the Hexagon originator; the apm@ Google Group's restricted posting policy is the load-bearing control (escalates to HIGH if the group is opened to external posting). Also correct the sender-auth-rejected alarm docs to match the shipped config (>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and note the SES-AR-01/02 parser hardening follow-ups. Refs: INFRA-107
2026-07-15 20:58:47 -04:00
# Fail-closed sender auth (INFRA-107): the handler only
# accepts mail whose SES-stamped Authentication-Results
# header carries dkim=pass for one of these domains.
# Observed on live traffic 2026-07-15: Coupa PO mail passes
# DKIM for amazon.coupahost.com (and amazonses.com, which is
# deliberately NOT allowlisted — every SES customer's mail
# passes that). Unset/empty ⇒ the handler rejects all mail.
"ALLOWED_DKIM_DOMAINS": "amazon.coupahost.com",
},
)
# Grant permissions
email_bucket.grant_read(email_processor)
po_table.grant_read_write_data(email_processor)
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
# --- Bedrock InvokeModel grant ---
# The us.* inference profile can route cross-region, so the grant MUST
# cover both the inference-profile ARN AND the per-region foundation-model
# ARNs (empty account field) for every region the profile can reach
# (us-east-1/us-east-2/us-west-2). A profile-only grant AccessDenies at
# runtime whenever the profile routes to a region whose foundation-model
# ARN is not allowed.
email_processor.add_to_role_policy(
iam.PolicyStatement(
actions=[
"bedrock:InvokeModel",
"bedrock:InvokeModelWithResponseStream",
],
resources=[
"arn:aws:bedrock:us-east-1:328440206208:inference-profile/us.anthropic.claude-haiku-4-5-20251001-v1:0",
"arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-haiku-4-5-20251001-v1:0",
"arn:aws:bedrock:us-east-2::foundation-model/anthropic.claude-haiku-4-5-20251001-v1:0",
"arn:aws:bedrock:us-west-2::foundation-model/anthropic.claude-haiku-4-5-20251001-v1:0",
],
)
)
# --- Errors alarm (INFRA-41 / audit H-8) ---
# ALARM-only (no OK action, per the CloudWatch-alarm preference) to the
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70) * Add CloudWatch alarm coverage for po-ingest and workorder-ingest Expands alarm coverage across both CDK stacks. All alarms are ALARM-only (no OK action) to the shared site-alerts SNS topic, with TreatMissingData NOT_BREACHING. The site-alerts topic is now imported once near the top of each stack so every alarm reuses one Topic instance. po-ingest (cdk/po_stack.py): - Errors: po-ingest-site-extractor - Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui - Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption), po-ingest-site-extractor, po-web-ui - DynamoDB throttle + system-error: purchase-orders, verified-sites, pending-site-review workorder-ingest (cdk/wo_stack.py): - Throttles: workorder-email-processor - Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption) - DynamoDB throttle + system-error: WorkOrders, WorkOrderComments DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation dimension set, so each table alarm is a Sum math expression across operations via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.259.0). Refs INFRA-41 / audit H-8. * Drop NEEDS ADAM SIGN-OFF wording from alarm comments Duration alarm thresholds are owner-approved; remove the sign-off flag from po_stack.py and wo_stack.py comments. Threshold values, eval config, and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
# shared site-alerts topic. Any errored invocation in a 5-min window pages.
email_processor.metric_errors(
period=Duration.minutes(5),
statistic="Sum",
).create_alarm(
self,
"EmailProcessorErrorsAlarm",
alarm_name="po-email-processor-errors",
alarm_description="po-email-processor async invocation errors",
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
Add fail-closed SES sender authentication (INFRA-107) (#98) * Add fail-closed SES sender authentication The From header and any raw-MIME Authentication-Results copies are attacker-forgeable, so a forged email to apm@int.seahaven.com or amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107, CRITICAL). Both S3-triggered email processors now authenticate the sender against the Authentication-Results header SES itself prepends at delivery: only the topmost header is consulted, its authserv-id must be amazonses.com, and it must carry dkim=pass for a domain in the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set in CDK so ops can adjust without code changes). Allowlists come from live traffic observed 2026-07-15 on both ingest buckets: WO mail arrives via the apm@ Google Groups forward, which re-signs as seahaven.com (the hxgnsmartcloud.com signature does not survive the forward); PO mail passes for amazon.coupahost.com. amazonses.com also passes on PO mail but is deliberately excluded -- every SES customer's outbound mail passes for it. Every failure path (env var unset, header missing or unparseable, verdict fail, unaligned domain) rejects the email: a structured warning with the reason and S3 key is logged and the record skipped without erroring the invocation, so rejected mail causes no Lambda retries or DLQ messages. Handler signatures and event sources are unchanged. Refs: INFRA-107 * Harden AR parser per cross-family review Cross-family (GPT-4.1) review findings: terminate the dkim result token at end-of-clause, whitespace, or a comment so a value like "dkim=pass-fake" can never be read as a pass; normalize trailing dots off allowlist entries so "seahaven.com." matches; make the compat32 parser policy explicit. Adds tests for result-token boundaries, comments after the result, quoted domain values, and folding inside a dkim clause. Refs: INFRA-107 * Harden AR parsing and alarm on sender-auth rejects The SES-stamped Authentication-Results value echoes attacker-controlled SMTP-session tokens (envelope-from, helo, header.from) as their own semicolon-delimited property clauses. A naive split(";") tore an RFC 5321 quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass clause, so a fully spoofed email was accepted on the genuinely SES-stamped topmost header. Tokenise comment- and quoted-string-aware (RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on semicolons outside a quoted-string, and fail closed on unbalanced quotes/comments so a ';' inside a quoted pvalue can never start a clause. Rejected mail returns normally (no error, no retry, no DLQ message), so a signing-domain drift or a wrong allowlist would silently discard 100% of legitimate mail while every alarm stayed green. Add a CloudWatch Logs metric filter + alarm on the sender_auth_rejected warning to both stacks so a false-reject storm pages instead of vanishing. This is also the safety net for the WO seahaven.com allowlist assumption, which must be validated against a live SES-stamped header (a plain Gmail auto-forward re-signs under the sending Workspace domain, not seahaven.com). Refs: INFRA-107 * chore: retrigger CI (no run recorded for 7c74ac1) * Fix quoted-AUID DKIM domain spoof in sender auth Resolve three confirmed /sh-security-review findings on the fail-closed SES sender-authentication control. HIGH: header.i/header.d domain extraction was not quoted-string aware. An attacker with a valid DKIM key for their own domain could set an RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive extractor stopped at the closing quote and returned seahaven.com, accepting forged mail. Extraction now tokenises the clause with the same quoted-string discipline already used for clause splitting: header.d (the plain signing domain) is authoritative when present, otherwise the header.i domain is the part after the AUID's LAST top-level "@", so a "@" inside a quoted local-part is treated as signer-controlled label text and yields the true signer (attacker.com), not seahaven.com. LOW: the topmost-header parse ran outside evaluate_sender_authentication's try/except, so an unexpected parser exception on crafted input could propagate into the handler and Lambda async retries/DLQ. The parse now fails CLOSED with an authentication_results_unparseable reason. MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to a low-volume total-reject outage (a trickle that never sums to 3). Both stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3 / datapoints_to_alarm=2, so a sustained reject condition pages even at one reject per period while a lone stray probe self-clears. Refs: INFRA-107 * Load Lambda function dir on sys.path in tests Rebasing INFRA-107 onto main folded #95's pytest suite into this branch's tests. The unified conftest loads the PO/WO handlers by file path, and handler.py now does `from ses_auth import authenticate_inbound_email` -- a bare sibling import that resolves in the Lambda only because the runtime puts each function's own directory on sys.path. The shared load_handler now adds that directory so the handler tests import correctly alongside the sender-auth tests. Refs: INFRA-107 * Note #97 test files in README directory tree The rebase onto main brought in #97's tests/requirements.txt and tests/test_po_merge.py. List both in the directory tree so it matches the tree on disk. Refs: INFRA-107 * Document INFRA-107 forwarder-binding risk acceptance Record the accepted risk that WO sender auth binds to the apm@ forward's re-signing domain (seahaven.com) rather than the Hexagon originator; the apm@ Google Group's restricted posting policy is the load-bearing control (escalates to HIGH if the group is opened to external posting). Also correct the sender-auth-rejected alarm docs to match the shipped config (>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and note the SES-AR-01/02 parser hardening follow-ups. Refs: INFRA-107
2026-07-15 20:58:47 -04:00
# --- Sender-auth rejection alarm (INFRA-107) ---
# A rejected email (bad/unaligned DKIM verdict) returns normally, so it
# produces NO Lambda error, NO DLQ message and NO retry -- only a
# `sender_auth_rejected` warning log. Without this metric filter + alarm a
# domain drift (Coupa rotates its signing subdomain, SES changes its
# Authentication-Results format, the allowlist is wrong) would silently
# discard 100% of legitimate PO mail while every other alarm stays green.
# A CloudWatch Logs metric filter turns those warnings into a metric so a
# false-reject storm pages instead of vanishing. default_value=0 keeps the
# series populated (alarm stays OK, never INSUFFICIENT_DATA) between events.
_add_sender_auth_rejected_alarm(
self, "EmailProcessor", "po-email-processor", alarm_topic
)
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70) * Add CloudWatch alarm coverage for po-ingest and workorder-ingest Expands alarm coverage across both CDK stacks. All alarms are ALARM-only (no OK action) to the shared site-alerts SNS topic, with TreatMissingData NOT_BREACHING. The site-alerts topic is now imported once near the top of each stack so every alarm reuses one Topic instance. po-ingest (cdk/po_stack.py): - Errors: po-ingest-site-extractor - Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui - Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption), po-ingest-site-extractor, po-web-ui - DynamoDB throttle + system-error: purchase-orders, verified-sites, pending-site-review workorder-ingest (cdk/wo_stack.py): - Throttles: workorder-email-processor - Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption) - DynamoDB throttle + system-error: WorkOrders, WorkOrderComments DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation dimension set, so each table alarm is a Sum math expression across operations via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.259.0). Refs INFRA-41 / audit H-8. * Drop NEEDS ADAM SIGN-OFF wording from alarm comments Duration alarm thresholds are owner-approved; remove the sign-off flag from po_stack.py and wo_stack.py comments. Threshold values, eval config, and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
# --- Throttles alarm: po-email-processor ---
# Any throttled invocation (concurrency cap hit) in a 5-min window pages.
# ALARM-only to site-alerts; no OK action; NOT_BREACHING when no data.
email_processor.metric_throttles(
period=Duration.minutes(5),
statistic="Sum",
).create_alarm(
self,
"EmailProcessorThrottlesAlarm",
alarm_name="po-email-processor-throttles",
alarm_description="po-email-processor invocation throttles",
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
# --- DLQ messages-present alarm ---
# Pages when any message lands in the EmailProcessorDlq: a message here
# means a PO email was permanently dropped after Lambda exhausted its
# async retries. Maximum over a single 5-min window > 0 fires; missing
# data (no messages metric emitted) is not breaching. Reuses the shared
# site-alerts topic, ALARM-only, like the errors alarm above. The metric
# helper derives the QueueName dimension from the queue construct, so the
# alarm tracks the CDK-generated queue name without hardcoding it.
email_processor_dlq.metric_approximate_number_of_messages_visible(
period=Duration.minutes(5),
statistic="Maximum",
).create_alarm(
self,
"EmailProcessorDlqMessagesAlarm",
alarm_name="po-email-processor-dlq-messages",
alarm_description="po-email-processor DLQ has messages (dropped PO emails)",
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70) * Add CloudWatch alarm coverage for po-ingest and workorder-ingest Expands alarm coverage across both CDK stacks. All alarms are ALARM-only (no OK action) to the shared site-alerts SNS topic, with TreatMissingData NOT_BREACHING. The site-alerts topic is now imported once near the top of each stack so every alarm reuses one Topic instance. po-ingest (cdk/po_stack.py): - Errors: po-ingest-site-extractor - Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui - Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption), po-ingest-site-extractor, po-web-ui - DynamoDB throttle + system-error: purchase-orders, verified-sites, pending-site-review workorder-ingest (cdk/wo_stack.py): - Throttles: workorder-email-processor - Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption) - DynamoDB throttle + system-error: WorkOrders, WorkOrderComments DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation dimension set, so each table alarm is a Sum math expression across operations via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.259.0). Refs INFRA-41 / audit H-8. * Drop NEEDS ADAM SIGN-OFF wording from alarm comments Duration alarm thresholds are owner-approved; remove the sign-off flag from po_stack.py and wo_stack.py comments. Threshold values, eval config, and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
# --- Duration alarm: po-email-processor (orphan adoption) ---
# Adopts the orphaned CLI alarm Lambda-Duration-po-email-processor under
# the repo's <fn>-duration naming (NEW logical name → no deploy collision;
# delete the orphan post-deploy). p99 / 45000 ms
# (75% of the 60s timeout) / eval 3 of 3 — tighter than the orphan's
# Maximum>=48000 / 1-of-1.
email_processor.metric_duration(
period=Duration.minutes(5),
statistic="p99",
).create_alarm(
self,
"EmailProcessorDurationAlarm",
alarm_name="po-email-processor-duration",
alarm_description="po-email-processor p99 duration approaching the 60s timeout",
threshold=45000,
evaluation_periods=3,
datapoints_to_alarm=2,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * feature: Add PO template parser scaffold and design doc Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage: - coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands. - coupa_cancellation (2.9%): implemented. Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work. Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com> * Implement PO new_po extraction and value-level gate Replace the extract_new_po scaffold stub with the full section-windowed extractor (duplicate-label anchoring, sentinel ship-to, label-keyed U+2022 bullet split, Decimal money from three anchored contexts only) and add value-level gate rules V1-V13. Both new_po_not_implemented scaffold guards are removed; rules 6-8 (unrecognized_status, multiline_unsupported, non_usd) go live. The gate re-derives every byte proof from the email body so an extractor bug cannot vouch for itself: amount re-serialization with a digit/comma border check (the thousands-separator truncation kill switch), sum(lines)==total against both Total blocks, anchor/supplier identity proofs, USPS address shape on the raw pre-enrichment zip, bullet label discipline, and sentinel/artifact hygiene. Any failure falls closed to the LLM; a validation failure is never a parsed result. Refs: #99 * Wire template-first parse into PO handler with EMF metric Run try_deterministic_parse ahead of the Bedrock extractor and fall back only on a miss/invalid (fail-closed) result. The shared enrich_parsed post-stage and the save_cancellation/save_revision/ save_new_po routing are untouched, so both paths write identical DynamoDB shapes and the po-ingest-site-extractor stream contract is preserved. Each record emits one ParseMethod EMF line (Seahaven/PoIngest/ ParseOutcome, dimension sets [ParseMethod] and [ParseMethod,TemplateId], ReasonCode/po_number ride-alongs) mirroring the WO idiom. The metric fires before the Bedrock call so a Bedrock-side error still records the ai_fallback outcome. Refs: #99 * Add PO fallback-rate alarm retuned for ~57 emails/day The WO alarm's 15-min period and >=10-sample floor assume ~760/day and would be structurally dead at PO volume (a 15-min period holds ~0.6 emails, so the floor is never met). Retune: 6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...) volume floor so a single email can never breach a datapoint (1/8 = 12.5% < 20%), threshold >20% against a ~1% expected baseline, eval 4 / datapoints 2 (24h span) so noise self-clears while total template drift pages within ~12h. No element-wise MAX in the math expression (post-#102 rule); ALARM-only SnsAction to site-alerts, NOT_BREACHING. Gated with 'npx cdk synth po-ingest'. Also add template_parser.py to the bundling cp list -- without it every deployed invocation would ImportError (unit tests cannot catch an asset-bundling omission). Refs: #99, #102 * Add offline PO parser suite with scrubbed fixture corpus 132 tests: golden-file comparison for all 25 positive fixtures (17 single-line new-PO + 8 cancellations, Decimal-exact via parse_float=Decimal), every fail-closed gate reason code covered (body-level triggers via 17 synthetic adversarial .eml mutations, candidate-level via direct validate() unit tests), real multi-line and comment/non-Coupa fallback fixtures, dual line-ending parse identity, two-path enrich/save parity (site-extractor stream guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass + scrub-marker leak sweep), and Bedrock dispatch/EMF assertions. The suite loads handler/template_parser via importlib under unique module names and binds the handler's bare sibling imports around exec (tests/conftest.py load_handler gets the same treatment) -- the WO suite caches bare 'handler'/'template_parser' names in sys.modules, and bare imports here would silently bind to the wrong pipeline. moto is imported before the handler so its botocore stubber hook precedes boto3 session creation (the PO conftest chain now loads at pytest session start). Fixtures are scrubbed real S3 samples: transport/auth header values replaced with same-shape placeholders (structure kept so ses_auth still passes), per-file digit ciphers, amounts remapped with sum==total re-established. The .gitignore exception is scoped to the PO fixtures path only. Refs: #99 * Document PO template-first parser and retuned alarm README: PO flow is now template-first with Bedrock fallback; parser/gate section mirroring the WO writeup; Seahaven/PoIngest ParseOutcome namespace and the fallback-rate alarm numbers with their volume justification (deliberately not WO's settings); test-suite and repo-layout updates. Design doc: mark PR #1 complete in progress/checklist sections; document the six value-level gate reason codes and the scaffold guard removal; correct the stale data-access note (default CLI session is 328440206208) and note the ~90-day S3 lifecycle aging of the corpus; record the 2.3 layout addendum (leading Supplier bullet segment, EA evidence lines, summary unit-price tokens, decode-path line endings), the fixture-build pins (address join convention, quantity/unit/price source), the V10 sweep outcome, and resolutions for open questions Q3/Q6. Cross-family review and the Confluence architecture-map update are flagged outstanding for merge. Refs: #99 * Record cross-family review outcome for handler wiring GPT-4.1 cross_review.py run against the real handler diff returned no BLOCK and no security findings; both FIX items verified as no-change-needed (fallback logging already correct; non-dict AI output is the pre-existing issue #101 pattern this PR deliberately does not touch). Refs: #99 * Pin line-item currency to USD in the PO gate The non_usd rule only checked the Total-block top-level currency, so a new_po whose line item read 'for 55,206.00 CAD' under a USD Total block still template-parsed as ok -- a fail-open hole in the fail-closed gate. Every line item's captured currency and its re-derived body token must now byte-equal the proven-USD top-level currency; covered by a line-level CAD adversarial fixture (the existing adv-non-usd only exercised the Total-block variant) and a candidate-mutation unit test. * Scrub residual transport tokens from PO fixtures The first-pass harvest scrub sanitized only the primary SES/DKIM header blocks, leaving the real SES Feedback-ID sender-identity hash in 49 committed fixtures and, on the two non-Coupa fixtures, an embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the token classes the PR #99 fixture lesson requires placeholdered. Replace each with a same-shape ScrubbedFixture value (byte-safe, CRLF and folding preserved) so header structure and ses_auth behavior are unchanged. * Converge quantity/price to Decimal on both paths EXTRACTION_PROMPT declares quantity and price as JSON strings, so a prompt-obedient Bedrock response stores DynamoDB Strings where the template parser stores Numbers -- divergent attribute types for the same email on the purchase-orders stream. Coerce numeric strings to Decimal in the shared enrich_parsed post-stage (thousands-separator safe; non-numeric strings kept verbatim) so both paths converge; prompt rewording itself remains PR #2 scope. The two-path parity test was circular -- it replayed the parser- derived golden as 'the LLM output', so it could never see the type divergence. It now feeds a prompt-shaped payload (string quantity/ price, LLM-filled site_code) through enrich_parsed and save_new_po, and the fixture-hygiene test now asserts the scrubbed transport-token header classes so fixture regressions are caught. * Coerce bare-int quantity/price to Decimal in enrich_parsed GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged residual type drift: parse_float=Decimal rules out floats on the LLM path, but a bare JSON int survived as Python int. Coerce it so both parse paths emit one canonical Decimal type. * Scrub fixture-body PII and harden cancellation gate (sec review) /sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill verifier) confirmed two diff-introduced findings; both fixed here. F3 (medium, real PII in new fixtures): the harvest scrub replaced header tokens but left real third-party PII in message BODIES -- an Amazon contact's name/phone/personal email in non-coupa-02.eml and an internal t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn names recurring across the new_po corpus. Replaced every personal name, phone, personal email, and internal URL with synthetic placeholders (QP-soft-wrap aware) across both .eml bodies and expected goldens. Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs, the leaked tokens), closing the header-only gap that let this through. F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and matched with .search(), unlike the anchored new_po pattern -- a subject merely ending with the cancellation phrase could be routed to the sticky- Cancelled write. Fully anchored it and switched to .match, and added a body-corroboration gate (the real Coupa body independently restates 'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted subject whose body does not corroborate now fails closed to the LLM (new reason code cancellation_body_unconfirmed). Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po dispatch and undelimited extraction prompt (issue #101 family) are byte-identical to main and unchanged here. 401 tests pass; ruff/format clean; cdk synth po-ingest clean. --------- Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
# --- Template fallback-rate alarm: po-email-processor ---
# The processor tries a deterministic template parse first and only calls
# the Bedrock AI extractor on a miss/invalid. A sustained rise in the
# ai_fallback share signals Coupa template drift (coverage collapse).
# EMF metric Seahaven/PoIngest/ParseOutcome, dimensioned by ParseMethod
# (template|ai_fallback).
#
# RETUNED for PO volume (~57 emails/day ≈ 14.25 per 6h period) -- the WO
# alarm's 15-min period / >=10-sample floor assume ~760/day and would be
# structurally DEAD here (a 15-min period holds ~0.6 PO emails, so the
# floor is never met and the IF always takes the 0 branch):
# * period 6h: a stable ~14-email denominator per datapoint.
# * volume floor >=8: at the floor, one fallback email = 12.5% < 20%,
# so a single email can NEVER breach a datapoint; a breach needs >=2
# fallbacks in one 6h window (2/8 = 25%) or >=3 at typical volume
# (3/14 ≈ 21%). Sparse overnight/weekend windows (<8 emails) take
# the 0 branch -- non-breaching by design (accepted trade: a Friday-
# evening drift may not page until weekend volume accrues).
# * threshold >20%: expected baseline fallback ≈1% (comments 0.55% +
# multi-line 0.18% + non-USD 0) -- far below the threshold.
# * 2 of 4 datapoints (24h span): isolated noise self-clears, while
# total template drift (100% fallback) pages within ~12h.
# Post-#102 rule: NO element-wise MAX(timeseries, scalar) in alarm math;
# the IF volume floor guarantees the non-zero denominator. Any change to
# this expression must be gated by `npx cdk synth po-ingest`.
fb_metric = cloudwatch.Metric(
namespace="Seahaven/PoIngest",
metric_name="ParseOutcome",
dimensions_map={"ParseMethod": "ai_fallback"},
statistic="Sum",
period=Duration.hours(6),
)
tmpl_metric = cloudwatch.Metric(
namespace="Seahaven/PoIngest",
metric_name="ParseOutcome",
dimensions_map={"ParseMethod": "template"},
statistic="Sum",
period=Duration.hours(6),
)
fallback_rate = cloudwatch.MathExpression(
expression=(
"IF((FILL(fb,0)+FILL(tmpl,0))>=8, "
"100*FILL(fb,0)/(FILL(fb,0)+FILL(tmpl,0)), 0)"
),
using_metrics={"fb": fb_metric, "tmpl": tmpl_metric},
period=Duration.hours(6),
label="TemplateFallbackRatePct",
)
fallback_rate.create_alarm(
self,
"EmailProcessorTemplateFallbackRateAlarm",
alarm_name="po-email-processor-template-fallback-rate",
alarm_description=(
"po-email-processor deterministic-template coverage collapse: "
">20% of parses fell back to the Bedrock AI extractor"
),
threshold=20,
evaluation_periods=4,
datapoints_to_alarm=2,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
# S3 event notification → Lambda
email_bucket.add_event_notification(
s3.EventType.OBJECT_CREATED,
s3n.LambdaDestination(email_processor),
s3.NotificationKeyFilter(prefix="inbound/"),
)
# --- SES Receipt Rule ---
# Reuse the existing INBOUND_MAIL rule set (shared with workorder-ingest)
rule_set = ses.ReceiptRuleSet.from_receipt_rule_set_name(
self,
"ExistingRuleSet",
"INBOUND_MAIL",
)
rule_set.add_rule(
"PoEmailRule",
recipients=["amazon_po@int.seahaven.com"],
actions=[
ses_actions.S3(
bucket=email_bucket,
object_key_prefix="inbound/",
),
],
)
Land safe fixes from 2026-06-17 security sweep (#97) * Remove gratuitous KMS grant on shared DynamoDB CMK wo-email-processor held grant_encrypt_decrypt on the shared seahaven-dynamodb CMK, but the WorkOrders/WorkOrderComments tables are not encrypted with that CMK. The grant was dead weight that extended the WO processor's decrypt reach to the CMK protecting the purchase-orders table (cross-stack decrypt). Drop it to restore least privilege; re-add as part of the table CMK migration (INFRA-6). Refs: INFRA-6 * Require Secrets Manager key for Anthropic client Remove the silent fallback to a plaintext ANTHROPIC_API_KEY env var in both email processors; require ANTHROPIC_API_KEY_SECRET_ARN and raise if absent so a misconfigured deploy fails loudly instead of using an unmanaged key. Adapted from f175323 on security/sweep-2026-06-17. The From-header sender-domain allowlist from that commit is intentionally dropped: the From header is spoofable (INFRA-107, confirmed critical) and sender authentication is being reworked in a separate PR. Refs: INFRA-107 * Merge PO revisions and handle out-of-order events save_revision did a full put_item overwrite, so a revision omitting line_items/supplier permanently deleted them. save_new_po used a conditional put that silently dropped the PO when an out-of-order cancellation had already created a skeleton row. Switch both to field-level merge update_items: a revision now SETs only the fields it carries, and a new_po backfills data into a pre-existing Cancelled skeleton while preserving the Cancelled status. No email can now delete data established by an earlier one. * Gate web UIs behind auth and escape currency XSS The po-web-ui and workorder-web-ui handlers had no auth: any invocation path returned the full PO/WO DB. Add a fail-closed shared-secret gate (X-Auth-Token / Bearer, constant-time compared to WEB_UI_AUTH_TOKEN) so a future re-attached Function URL cannot re-expose the data (URLs removed under INFRA-74). Wire the token from the SSM String param /procurement-ingest/web-ui-auth-token. Also fix stored XSS in po-web-ui fmt_currency: the non-numeric fallback returned str(val) unescaped, so a prompt-injected email could make Claude emit total_amount as <script>. Escape it. Refs: INFRA-74 * Document sweep security fixes and merge semantics Update the README for the 2026-06-17 security sweep: required Secrets Manager key (no plaintext env fallback), web UI auth gate + SSM token setup step, output-escaping note, and the new PO revision/cancellation merge behavior. Adapted from d91f45e on security/sweep-2026-06-17; the sender allowlist documentation is dropped along with the allowlist itself (deferred to the INFRA-107 sender-authentication rework). Refs: INFRA-107 * fix: resolve web UI auth token from Secrets Manager at runtime Replace the plaintext SSM String parameter with a Secrets Manager secret referenced by ARN only. The token is fetched and cached at module level on first invocation, keeping shared secrets out of CloudFormation templates and Lambda environment variables. Refs: PR-97 * Add TTL to web UI auth token cache for rotation The web-ui handlers cached the Secrets Manager auth token at module level with no expiry, so a rotated secret was only picked up when the warm container recycled — an emergency rotation could take hours to take effect. Cache the fetched value for a 5-minute TTL instead, so a rotated token propagates within the TTL while still avoiding a Secrets Manager call on every request. Still fails closed when the secret is unset or unreadable. Refs: INFRA-74 * Log Secrets Manager failures in web UI auth token fetch The web UI auth gate correctly fails closed when the shared token cannot be read, but _get_auth_token() swallowed every exception silently. A Secrets Manager permission or config error then made every request 401 with no operational signal, leaving an outage indistinguishable from ordinary unauthenticated traffic. Add a module-level logger to both web_ui handlers and log the fetch failure with logger.exception() in the except block before returning None. Behavior is unchanged (still fails closed); the failure is now visible in CloudWatch. The secret value is never logged. The two handlers stay byte-consistent in the mirrored _get_auth_token() region. The companion finding on the CDK import of the shared procurement-ingest/web-ui-auth-token secret was evaluated and left as-is: the token is a single secret shared by both the PO and WO stacks, so from_secret_name_v2 (which scopes grant_read via the standard 6-char suffix wildcard) is correct; making it a CDK-managed Secret in both stacks would collide the two stacks on the same explicit secret name at deploy time. Refs: INFRA-74 * Make Cancelled PO status sticky via atomic write The PO merge path read status with a get_item (_is_cancelled) and then wrote with an unconditional update_item. Two defects followed from this: - Race (Issue A): a cancellation landing between the read and the write was silently un-cancelled by a revision carrying a non-cancelled po_status — a TOCTOU on a table with concurrent email processing. - Over-broad strip (Issue B): save_revision dropped po_status whenever the PO was Cancelled, so legitimate status updates on non-cancelled POs and status-less revisions were affected rather than only the true un-cancel transition. Enforce the invariant server-side instead. "Cancelled" is a sticky, authoritative status: once set, later new_po/revision emails may enrich other fields but must never move it to a non-cancelled status. When the payload carries a non-cancelled po_status, _merge_update issues the update_item guarded by ConditionExpression "attribute_not_exists(po_status) OR po_status <> :marker", evaluated atomically at write time, so a cancellation that lands first always wins. On ConditionalCheckFailedException the same fields are re-written without po_status/cancelled_at, enriching the record while Cancelled sticks. Payloads with no status change, or an already -Cancelled status, take a plain merge — the status is only ever suppressed on a real un-cancel. This removes the non-atomic get_item from the write path; _is_cancelled is deleted. Key schema and attribute names are unchanged, so the cross-stack purchase-orders contract (read-only by seahaven-slack-bot) holds. Add moto-backed tests covering un-cancel suppression with field enrichment, status-less merge onto a Cancelled PO, legitimate status updates on non-cancelled POs, new_po backfill of a Cancelled skeleton, fresh create/merge, and authoritative save_cancellation. Refs: #97
2026-07-15 20:17:46 -04:00
# --- Web UI auth token secret ---
# Shared secret for the web UI auth gate, stored in Secrets Manager and
# resolved at runtime so the token never appears in CloudFormation templates
# or Lambda environment variables. Create this secret before deploying
# either stack; both PO and WO stacks reference it by name.
web_ui_auth_secret = secretsmanager.Secret.from_secret_name_v2(
self,
"WebUiAuthToken",
"procurement-ingest/web-ui-auth-token",
)
# --- Web UI Lambda ---
web_ui = lambda_.Function(
self,
"WebUI",
function_name="po-web-ui",
runtime=lambda_.Runtime.PYTHON_3_12,
architecture=lambda_.Architecture.ARM_64,
handler="handler.handler",
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
code=lambda_.Code.from_asset("../lambdas/po/web_ui"),
timeout=Duration.seconds(60),
memory_size=256,
log_retention=logs.RetentionDays.TWO_MONTHS,
environment={
"PO_TABLE": "purchase-orders",
Land safe fixes from 2026-06-17 security sweep (#97) * Remove gratuitous KMS grant on shared DynamoDB CMK wo-email-processor held grant_encrypt_decrypt on the shared seahaven-dynamodb CMK, but the WorkOrders/WorkOrderComments tables are not encrypted with that CMK. The grant was dead weight that extended the WO processor's decrypt reach to the CMK protecting the purchase-orders table (cross-stack decrypt). Drop it to restore least privilege; re-add as part of the table CMK migration (INFRA-6). Refs: INFRA-6 * Require Secrets Manager key for Anthropic client Remove the silent fallback to a plaintext ANTHROPIC_API_KEY env var in both email processors; require ANTHROPIC_API_KEY_SECRET_ARN and raise if absent so a misconfigured deploy fails loudly instead of using an unmanaged key. Adapted from f175323 on security/sweep-2026-06-17. The From-header sender-domain allowlist from that commit is intentionally dropped: the From header is spoofable (INFRA-107, confirmed critical) and sender authentication is being reworked in a separate PR. Refs: INFRA-107 * Merge PO revisions and handle out-of-order events save_revision did a full put_item overwrite, so a revision omitting line_items/supplier permanently deleted them. save_new_po used a conditional put that silently dropped the PO when an out-of-order cancellation had already created a skeleton row. Switch both to field-level merge update_items: a revision now SETs only the fields it carries, and a new_po backfills data into a pre-existing Cancelled skeleton while preserving the Cancelled status. No email can now delete data established by an earlier one. * Gate web UIs behind auth and escape currency XSS The po-web-ui and workorder-web-ui handlers had no auth: any invocation path returned the full PO/WO DB. Add a fail-closed shared-secret gate (X-Auth-Token / Bearer, constant-time compared to WEB_UI_AUTH_TOKEN) so a future re-attached Function URL cannot re-expose the data (URLs removed under INFRA-74). Wire the token from the SSM String param /procurement-ingest/web-ui-auth-token. Also fix stored XSS in po-web-ui fmt_currency: the non-numeric fallback returned str(val) unescaped, so a prompt-injected email could make Claude emit total_amount as <script>. Escape it. Refs: INFRA-74 * Document sweep security fixes and merge semantics Update the README for the 2026-06-17 security sweep: required Secrets Manager key (no plaintext env fallback), web UI auth gate + SSM token setup step, output-escaping note, and the new PO revision/cancellation merge behavior. Adapted from d91f45e on security/sweep-2026-06-17; the sender allowlist documentation is dropped along with the allowlist itself (deferred to the INFRA-107 sender-authentication rework). Refs: INFRA-107 * fix: resolve web UI auth token from Secrets Manager at runtime Replace the plaintext SSM String parameter with a Secrets Manager secret referenced by ARN only. The token is fetched and cached at module level on first invocation, keeping shared secrets out of CloudFormation templates and Lambda environment variables. Refs: PR-97 * Add TTL to web UI auth token cache for rotation The web-ui handlers cached the Secrets Manager auth token at module level with no expiry, so a rotated secret was only picked up when the warm container recycled — an emergency rotation could take hours to take effect. Cache the fetched value for a 5-minute TTL instead, so a rotated token propagates within the TTL while still avoiding a Secrets Manager call on every request. Still fails closed when the secret is unset or unreadable. Refs: INFRA-74 * Log Secrets Manager failures in web UI auth token fetch The web UI auth gate correctly fails closed when the shared token cannot be read, but _get_auth_token() swallowed every exception silently. A Secrets Manager permission or config error then made every request 401 with no operational signal, leaving an outage indistinguishable from ordinary unauthenticated traffic. Add a module-level logger to both web_ui handlers and log the fetch failure with logger.exception() in the except block before returning None. Behavior is unchanged (still fails closed); the failure is now visible in CloudWatch. The secret value is never logged. The two handlers stay byte-consistent in the mirrored _get_auth_token() region. The companion finding on the CDK import of the shared procurement-ingest/web-ui-auth-token secret was evaluated and left as-is: the token is a single secret shared by both the PO and WO stacks, so from_secret_name_v2 (which scopes grant_read via the standard 6-char suffix wildcard) is correct; making it a CDK-managed Secret in both stacks would collide the two stacks on the same explicit secret name at deploy time. Refs: INFRA-74 * Make Cancelled PO status sticky via atomic write The PO merge path read status with a get_item (_is_cancelled) and then wrote with an unconditional update_item. Two defects followed from this: - Race (Issue A): a cancellation landing between the read and the write was silently un-cancelled by a revision carrying a non-cancelled po_status — a TOCTOU on a table with concurrent email processing. - Over-broad strip (Issue B): save_revision dropped po_status whenever the PO was Cancelled, so legitimate status updates on non-cancelled POs and status-less revisions were affected rather than only the true un-cancel transition. Enforce the invariant server-side instead. "Cancelled" is a sticky, authoritative status: once set, later new_po/revision emails may enrich other fields but must never move it to a non-cancelled status. When the payload carries a non-cancelled po_status, _merge_update issues the update_item guarded by ConditionExpression "attribute_not_exists(po_status) OR po_status <> :marker", evaluated atomically at write time, so a cancellation that lands first always wins. On ConditionalCheckFailedException the same fields are re-written without po_status/cancelled_at, enriching the record while Cancelled sticks. Payloads with no status change, or an already -Cancelled status, take a plain merge — the status is only ever suppressed on a real un-cancel. This removes the non-atomic get_item from the write path; _is_cancelled is deleted. Key schema and attribute names are unchanged, so the cross-stack purchase-orders contract (read-only by seahaven-slack-bot) holds. Add moto-backed tests covering un-cancel suppression with field enrichment, status-less merge onto a Cancelled PO, legitimate status updates on non-cancelled POs, new_po backfill of a Cancelled skeleton, fresh create/merge, and authoritative save_cancellation. Refs: #97
2026-07-15 20:17:46 -04:00
# Defense-in-depth shared secret for the web UI handler. The
# handler fails closed if this ARN is unset or the secret is
# missing, so any future invocation path cannot re-expose the
# PO DB unauthenticated. The secret value is fetched at runtime
# from Secrets Manager (not embedded in env vars or template).
"WEB_UI_AUTH_TOKEN_SECRET_ARN": web_ui_auth_secret.secret_arn,
},
)
po_table.grant_read_data(web_ui)
Land safe fixes from 2026-06-17 security sweep (#97) * Remove gratuitous KMS grant on shared DynamoDB CMK wo-email-processor held grant_encrypt_decrypt on the shared seahaven-dynamodb CMK, but the WorkOrders/WorkOrderComments tables are not encrypted with that CMK. The grant was dead weight that extended the WO processor's decrypt reach to the CMK protecting the purchase-orders table (cross-stack decrypt). Drop it to restore least privilege; re-add as part of the table CMK migration (INFRA-6). Refs: INFRA-6 * Require Secrets Manager key for Anthropic client Remove the silent fallback to a plaintext ANTHROPIC_API_KEY env var in both email processors; require ANTHROPIC_API_KEY_SECRET_ARN and raise if absent so a misconfigured deploy fails loudly instead of using an unmanaged key. Adapted from f175323 on security/sweep-2026-06-17. The From-header sender-domain allowlist from that commit is intentionally dropped: the From header is spoofable (INFRA-107, confirmed critical) and sender authentication is being reworked in a separate PR. Refs: INFRA-107 * Merge PO revisions and handle out-of-order events save_revision did a full put_item overwrite, so a revision omitting line_items/supplier permanently deleted them. save_new_po used a conditional put that silently dropped the PO when an out-of-order cancellation had already created a skeleton row. Switch both to field-level merge update_items: a revision now SETs only the fields it carries, and a new_po backfills data into a pre-existing Cancelled skeleton while preserving the Cancelled status. No email can now delete data established by an earlier one. * Gate web UIs behind auth and escape currency XSS The po-web-ui and workorder-web-ui handlers had no auth: any invocation path returned the full PO/WO DB. Add a fail-closed shared-secret gate (X-Auth-Token / Bearer, constant-time compared to WEB_UI_AUTH_TOKEN) so a future re-attached Function URL cannot re-expose the data (URLs removed under INFRA-74). Wire the token from the SSM String param /procurement-ingest/web-ui-auth-token. Also fix stored XSS in po-web-ui fmt_currency: the non-numeric fallback returned str(val) unescaped, so a prompt-injected email could make Claude emit total_amount as <script>. Escape it. Refs: INFRA-74 * Document sweep security fixes and merge semantics Update the README for the 2026-06-17 security sweep: required Secrets Manager key (no plaintext env fallback), web UI auth gate + SSM token setup step, output-escaping note, and the new PO revision/cancellation merge behavior. Adapted from d91f45e on security/sweep-2026-06-17; the sender allowlist documentation is dropped along with the allowlist itself (deferred to the INFRA-107 sender-authentication rework). Refs: INFRA-107 * fix: resolve web UI auth token from Secrets Manager at runtime Replace the plaintext SSM String parameter with a Secrets Manager secret referenced by ARN only. The token is fetched and cached at module level on first invocation, keeping shared secrets out of CloudFormation templates and Lambda environment variables. Refs: PR-97 * Add TTL to web UI auth token cache for rotation The web-ui handlers cached the Secrets Manager auth token at module level with no expiry, so a rotated secret was only picked up when the warm container recycled — an emergency rotation could take hours to take effect. Cache the fetched value for a 5-minute TTL instead, so a rotated token propagates within the TTL while still avoiding a Secrets Manager call on every request. Still fails closed when the secret is unset or unreadable. Refs: INFRA-74 * Log Secrets Manager failures in web UI auth token fetch The web UI auth gate correctly fails closed when the shared token cannot be read, but _get_auth_token() swallowed every exception silently. A Secrets Manager permission or config error then made every request 401 with no operational signal, leaving an outage indistinguishable from ordinary unauthenticated traffic. Add a module-level logger to both web_ui handlers and log the fetch failure with logger.exception() in the except block before returning None. Behavior is unchanged (still fails closed); the failure is now visible in CloudWatch. The secret value is never logged. The two handlers stay byte-consistent in the mirrored _get_auth_token() region. The companion finding on the CDK import of the shared procurement-ingest/web-ui-auth-token secret was evaluated and left as-is: the token is a single secret shared by both the PO and WO stacks, so from_secret_name_v2 (which scopes grant_read via the standard 6-char suffix wildcard) is correct; making it a CDK-managed Secret in both stacks would collide the two stacks on the same explicit secret name at deploy time. Refs: INFRA-74 * Make Cancelled PO status sticky via atomic write The PO merge path read status with a get_item (_is_cancelled) and then wrote with an unconditional update_item. Two defects followed from this: - Race (Issue A): a cancellation landing between the read and the write was silently un-cancelled by a revision carrying a non-cancelled po_status — a TOCTOU on a table with concurrent email processing. - Over-broad strip (Issue B): save_revision dropped po_status whenever the PO was Cancelled, so legitimate status updates on non-cancelled POs and status-less revisions were affected rather than only the true un-cancel transition. Enforce the invariant server-side instead. "Cancelled" is a sticky, authoritative status: once set, later new_po/revision emails may enrich other fields but must never move it to a non-cancelled status. When the payload carries a non-cancelled po_status, _merge_update issues the update_item guarded by ConditionExpression "attribute_not_exists(po_status) OR po_status <> :marker", evaluated atomically at write time, so a cancellation that lands first always wins. On ConditionalCheckFailedException the same fields are re-written without po_status/cancelled_at, enriching the record while Cancelled sticks. Payloads with no status change, or an already -Cancelled status, take a plain merge — the status is only ever suppressed on a real un-cancel. This removes the non-atomic get_item from the write path; _is_cancelled is deleted. Key schema and attribute names are unchanged, so the cross-stack purchase-orders contract (read-only by seahaven-slack-bot) holds. Add moto-backed tests covering un-cancel suppression with field enrichment, status-less merge onto a Cancelled PO, legitimate status updates on non-cancelled POs, new_po backfill of a Cancelled skeleton, fresh create/merge, and authoritative save_cancellation. Refs: #97
2026-07-15 20:17:46 -04:00
web_ui_auth_secret.grant_read(web_ui)
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70) * Add CloudWatch alarm coverage for po-ingest and workorder-ingest Expands alarm coverage across both CDK stacks. All alarms are ALARM-only (no OK action) to the shared site-alerts SNS topic, with TreatMissingData NOT_BREACHING. The site-alerts topic is now imported once near the top of each stack so every alarm reuses one Topic instance. po-ingest (cdk/po_stack.py): - Errors: po-ingest-site-extractor - Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui - Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption), po-ingest-site-extractor, po-web-ui - DynamoDB throttle + system-error: purchase-orders, verified-sites, pending-site-review workorder-ingest (cdk/wo_stack.py): - Throttles: workorder-email-processor - Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption) - DynamoDB throttle + system-error: WorkOrders, WorkOrderComments DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation dimension set, so each table alarm is a Sum math expression across operations via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.259.0). Refs INFRA-41 / audit H-8. * Drop NEEDS ADAM SIGN-OFF wording from alarm comments Duration alarm thresholds are owner-approved; remove the sign-off flag from po_stack.py and wo_stack.py comments. Threshold values, eval config, and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
# --- Throttles alarm: po-web-ui ---
web_ui.metric_throttles(
period=Duration.minutes(5),
statistic="Sum",
).create_alarm(
self,
"WebUiThrottlesAlarm",
alarm_name="po-web-ui-throttles",
alarm_description="po-web-ui invocation throttles",
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
# --- Duration alarm: po-web-ui ---
# Net-new (no orphan exists for this function).
# p99 / 45000 ms (75% of the 60s timeout) / eval 3, datapoints 2.
web_ui.metric_duration(
period=Duration.minutes(5),
statistic="p99",
).create_alarm(
self,
"WebUiDurationAlarm",
alarm_name="po-web-ui-duration",
alarm_description="po-web-ui p99 duration approaching the 60s timeout",
threshold=45000,
evaluation_periods=3,
datapoints_to_alarm=2,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
# Public Function URL removed 2026-06-08 (INFRA-74 / audit C-5): the
# unauthenticated FunctionUrlAuthType.NONE URL was deleted out-of-band
# via CLI. Removing the construct (and its auto-generated Principal:*
# invoke permission) reconciles IaC with the live state.
# --- Verified sites table (extracted from PO ship-to addresses) ---
verified_sites_table = dynamodb.Table(
self,
"VerifiedSitesTable",
table_name="verified-sites",
partition_key=dynamodb.Attribute(
name="siteCode",
type=dynamodb.AttributeType.STRING,
),
billing_mode=dynamodb.BillingMode.PAY_PER_REQUEST,
removal_policy=RemovalPolicy.RETAIN,
)
# by-state GSI removed 2026-06-03 (audit M-20): 0 reads in 30d against
# 518 WCU of write amplification. Re-add if a state-level query path ships.
# --- Site extractor Lambda (DynamoDB Streams → verified-sites) ---
site_extractor = lambda_.Function(
self,
"SiteExtractor",
function_name="po-ingest-site-extractor",
runtime=lambda_.Runtime.PYTHON_3_12,
architecture=lambda_.Architecture.ARM_64,
handler="handler.handler",
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
code=lambda_.Code.from_asset("../lambdas/po/site_extractor"),
timeout=Duration.seconds(60),
memory_size=256,
log_retention=logs.RetentionDays.TWO_MONTHS,
environment={
"VERIFIED_SITES_TABLE": verified_sites_table.table_name,
"PENDING_REVIEW_TABLE": "pending-site-review",
},
)
verified_sites_table.grant_read_write_data(site_extractor)
site_extractor.add_event_source(
lambda_event_sources.DynamoEventSource(
po_table,
starting_position=lambda_.StartingPosition.TRIM_HORIZON,
batch_size=10,
max_batching_window=Duration.seconds(30),
bisect_batch_on_error=True,
retry_attempts=3,
)
)
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70) * Add CloudWatch alarm coverage for po-ingest and workorder-ingest Expands alarm coverage across both CDK stacks. All alarms are ALARM-only (no OK action) to the shared site-alerts SNS topic, with TreatMissingData NOT_BREACHING. The site-alerts topic is now imported once near the top of each stack so every alarm reuses one Topic instance. po-ingest (cdk/po_stack.py): - Errors: po-ingest-site-extractor - Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui - Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption), po-ingest-site-extractor, po-web-ui - DynamoDB throttle + system-error: purchase-orders, verified-sites, pending-site-review workorder-ingest (cdk/wo_stack.py): - Throttles: workorder-email-processor - Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption) - DynamoDB throttle + system-error: WorkOrders, WorkOrderComments DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation dimension set, so each table alarm is a Sum math expression across operations via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.259.0). Refs INFRA-41 / audit H-8. * Drop NEEDS ADAM SIGN-OFF wording from alarm comments Duration alarm thresholds are owner-approved; remove the sign-off flag from po_stack.py and wo_stack.py comments. Threshold values, eval config, and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
# --- Errors alarm: po-ingest-site-extractor ---
# Stream-consumer errors retry per the event-source config, but a
# persistent failure stalls the verified-sites pipeline. ALARM-only to
# site-alerts; no OK action; NOT_BREACHING when no data.
site_extractor.metric_errors(
period=Duration.minutes(5),
statistic="Sum",
).create_alarm(
self,
"SiteExtractorErrorsAlarm",
alarm_name="po-ingest-site-extractor-errors",
alarm_description="po-ingest-site-extractor invocation errors",
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
# --- Throttles alarm: po-ingest-site-extractor ---
site_extractor.metric_throttles(
period=Duration.minutes(5),
statistic="Sum",
).create_alarm(
self,
"SiteExtractorThrottlesAlarm",
alarm_name="po-ingest-site-extractor-throttles",
alarm_description="po-ingest-site-extractor invocation throttles",
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
# --- Duration alarm: po-ingest-site-extractor ---
# Net-new (no orphan exists for this function).
# p99 / 45000 ms (75% of the 60s timeout) / eval 3, datapoints 2.
site_extractor.metric_duration(
period=Duration.minutes(5),
statistic="p99",
).create_alarm(
self,
"SiteExtractorDurationAlarm",
alarm_name="po-ingest-site-extractor-duration",
alarm_description="po-ingest-site-extractor p99 duration approaching the 60s timeout",
threshold=45000,
evaluation_periods=3,
datapoints_to_alarm=2,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
cdk.CfnOutput(
self,
"VerifiedSitesTableName",
value=verified_sites_table.table_name,
description="Verified site addresses extracted from POs",
)
# --- Pending site review table (POs with no extractable site code) ---
pending_review_table = dynamodb.Table(
self,
"PendingSiteReviewTable",
table_name="pending-site-review",
partition_key=dynamodb.Attribute(
name="po_number",
type=dynamodb.AttributeType.STRING,
),
billing_mode=dynamodb.BillingMode.PAY_PER_REQUEST,
removal_policy=RemovalPolicy.RETAIN,
)
pending_review_table.grant_read_write_data(site_extractor)
verified_sites_table.grant_read_data(site_extractor)
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70) * Add CloudWatch alarm coverage for po-ingest and workorder-ingest Expands alarm coverage across both CDK stacks. All alarms are ALARM-only (no OK action) to the shared site-alerts SNS topic, with TreatMissingData NOT_BREACHING. The site-alerts topic is now imported once near the top of each stack so every alarm reuses one Topic instance. po-ingest (cdk/po_stack.py): - Errors: po-ingest-site-extractor - Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui - Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption), po-ingest-site-extractor, po-web-ui - DynamoDB throttle + system-error: purchase-orders, verified-sites, pending-site-review workorder-ingest (cdk/wo_stack.py): - Throttles: workorder-email-processor - Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption) - DynamoDB throttle + system-error: WorkOrders, WorkOrderComments DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation dimension set, so each table alarm is a Sum math expression across operations via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.259.0). Refs INFRA-41 / audit H-8. * Drop NEEDS ADAM SIGN-OFF wording from alarm comments Duration alarm thresholds are owner-approved; remove the sign-off flag from po_stack.py and wo_stack.py comments. Threshold values, eval config, and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
# --- DynamoDB throttle + system-error alarms ---
# ThrottledRequests / SystemErrors emit at TableName + Operation only
# (verified against live CloudWatch: no TableName-only rollup exists, and
# metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.259.0).
# Each table currently has zero throttle/error datapoints, so the series
# only materialise on first occurrence — NOT_BREACHING keeps them OK until
# then.
_add_ddb_alarms(
self, "PurchaseOrdersTable", po_table, "purchase-orders", alarm_topic
)
_add_ddb_alarms(
self,
"VerifiedSitesTable",
verified_sites_table,
"verified-sites",
alarm_topic,
)
_add_ddb_alarms(
self,
"PendingSiteReviewTable",
pending_review_table,
"pending-site-review",
alarm_topic,
)