mirror of
https://github.com/Sea-Haven-Industries/procurement-ingest.git
synced 2026-09-30 09:33:15 +00:00
* Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
630 lines
28 KiB
Python
630 lines
28 KiB
Python
"""CDK stack for the Coupa PO email ingestion pipeline."""
|
|
|
|
import aws_cdk as cdk
|
|
from aws_cdk import (
|
|
Duration,
|
|
RemovalPolicy,
|
|
Stack,
|
|
aws_cloudwatch as cloudwatch,
|
|
aws_cloudwatch_actions as cw_actions,
|
|
aws_dynamodb as dynamodb,
|
|
aws_iam as iam,
|
|
aws_kms as kms,
|
|
aws_lambda as lambda_,
|
|
aws_lambda_event_sources as lambda_event_sources,
|
|
aws_logs as logs,
|
|
aws_s3 as s3,
|
|
aws_s3_notifications as s3n,
|
|
aws_ses as ses,
|
|
aws_ses_actions as ses_actions,
|
|
aws_secretsmanager as secretsmanager,
|
|
aws_sns as sns,
|
|
aws_sqs as sqs,
|
|
aws_ssm as ssm,
|
|
)
|
|
from constructs import Construct
|
|
|
|
# Operations these tables actually issue (PutItem/UpdateItem/DeleteItem writes,
|
|
# GetItem/Query/BatchGetItem reads). DynamoDB emits ThrottledRequests/SystemErrors
|
|
# keyed by TableName + Operation only, so the CDK *_for_operations helpers (which
|
|
# render a SUM MathExpression across these per-operation metrics) are the correct,
|
|
# non-deprecated way to roll a table up to a single alarmable series.
|
|
_DDB_ALARM_OPERATIONS = [
|
|
dynamodb.Operation.GET_ITEM,
|
|
dynamodb.Operation.BATCH_GET_ITEM,
|
|
dynamodb.Operation.QUERY,
|
|
dynamodb.Operation.SCAN,
|
|
dynamodb.Operation.PUT_ITEM,
|
|
dynamodb.Operation.UPDATE_ITEM,
|
|
dynamodb.Operation.DELETE_ITEM,
|
|
dynamodb.Operation.BATCH_WRITE_ITEM,
|
|
]
|
|
|
|
|
|
def _add_ddb_alarms(scope, id_prefix, table, alarm_name_prefix, alarm_topic):
|
|
"""Add throttle + system-error alarms for a DynamoDB table.
|
|
|
|
Both fire on any non-zero datapoint in a 5-min window. ALARM-only SnsAction
|
|
to site-alerts (no OK action); TreatMissingData NOT_BREACHING.
|
|
"""
|
|
table.metric_throttled_requests_for_operations(
|
|
operations=_DDB_ALARM_OPERATIONS,
|
|
period=Duration.minutes(5),
|
|
statistic="Sum",
|
|
).create_alarm(
|
|
scope,
|
|
f"{id_prefix}ThrottlesAlarm",
|
|
alarm_name=f"{alarm_name_prefix}-throttles",
|
|
alarm_description=f"{alarm_name_prefix} DynamoDB throttled requests",
|
|
threshold=0,
|
|
evaluation_periods=1,
|
|
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
|
|
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
|
|
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
|
|
|
|
table.metric_system_errors_for_operations(
|
|
operations=_DDB_ALARM_OPERATIONS,
|
|
period=Duration.minutes(5),
|
|
statistic="Sum",
|
|
).create_alarm(
|
|
scope,
|
|
f"{id_prefix}SystemErrorsAlarm",
|
|
alarm_name=f"{alarm_name_prefix}-system-errors",
|
|
alarm_description=f"{alarm_name_prefix} DynamoDB server-side (5xx) errors",
|
|
threshold=0,
|
|
evaluation_periods=1,
|
|
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
|
|
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
|
|
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
|
|
|
|
|
|
# CloudWatch namespace for the log-derived sender-authentication metrics.
|
|
_SENDER_AUTH_METRIC_NAMESPACE = "Seahaven/ProcurementIngest"
|
|
|
|
|
|
def _add_sender_auth_rejected_alarm(scope, id_prefix, function_name, alarm_topic):
|
|
"""Metric-filter + alarm on ``sender_auth_rejected`` warnings (INFRA-107).
|
|
|
|
A rejected inbound email is skipped without erroring the invocation, so it
|
|
is invisible to the Errors/Throttles/DLQ alarms. This turns the structured
|
|
warning log into a CloudWatch metric and pages when rejections spike --
|
|
catching a silent false-reject storm (allowlist wrong, signing-domain
|
|
drift, SES header-format change) that would otherwise discard legitimate
|
|
mail while the pipeline reports healthy.
|
|
|
|
ALARM-only SnsAction to site-alerts; no OK action. The metric filter reads
|
|
the function's own log group (imported by the deterministic
|
|
``/aws/lambda/<fn>`` name, created by the function's log_retention). A plain
|
|
substring pattern is used because Lambda prefixes each line with its own
|
|
level/timestamp/request-id, so the JSON payload is not a standalone JSON
|
|
log event a `{$.event=...}` pattern could match.
|
|
"""
|
|
metric_name = f"{function_name}-sender-auth-rejected"
|
|
logs.MetricFilter(
|
|
scope,
|
|
f"{id_prefix}SenderAuthRejectedFilter",
|
|
log_group=logs.LogGroup.from_log_group_name(
|
|
scope,
|
|
f"{id_prefix}LogGroup",
|
|
f"/aws/lambda/{function_name}",
|
|
),
|
|
filter_pattern=logs.FilterPattern.literal('"sender_auth_rejected"'),
|
|
metric_namespace=_SENDER_AUTH_METRIC_NAMESPACE,
|
|
metric_name=metric_name,
|
|
metric_value="1",
|
|
default_value=0,
|
|
)
|
|
|
|
# Fire on a *sustained* reject condition rather than a volume spike. The
|
|
# earlier Sum>=3-over-15-min threshold had a blind spot that is exactly the
|
|
# failure this alarm exists to catch: a low-traffic pipeline in total
|
|
# drift outage (allowlist wrong / signing-domain changed) may only produce
|
|
# a trickle of rejects -- one every few minutes -- that never sums to 3 in
|
|
# any window, so the outage never pages. Instead: >=1 reject per 5-min
|
|
# period, alarming when 2 of the last 3 periods breach (evaluation_periods=3
|
|
# / datapoints_to_alarm=2, the same idiom as the duration alarm). A single
|
|
# stray spoof probe (one lone period) is tolerated and self-clears, but a
|
|
# sustained reject condition trips within ~10-15 min even at one reject per
|
|
# period. default_value=0 on the metric filter keeps the series continuous
|
|
# so NOT_BREACHING only applies before the first datapoint ever arrives.
|
|
cloudwatch.Metric(
|
|
namespace=_SENDER_AUTH_METRIC_NAMESPACE,
|
|
metric_name=metric_name,
|
|
period=Duration.minutes(5),
|
|
statistic="Sum",
|
|
).create_alarm(
|
|
scope,
|
|
f"{id_prefix}SenderAuthRejectedAlarm",
|
|
alarm_name=f"{function_name}-sender-auth-rejected",
|
|
alarm_description=(
|
|
f"{function_name} rejected inbound mail on sender authentication "
|
|
"(possible allowlist/DKIM-domain drift silently dropping real mail)"
|
|
),
|
|
threshold=1,
|
|
evaluation_periods=3,
|
|
datapoints_to_alarm=2,
|
|
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
|
|
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
|
|
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
|
|
|
|
|
|
class PoIngestStack(Stack):
|
|
def __init__(self, scope: Construct, construct_id: str, **kwargs):
|
|
super().__init__(scope, construct_id, **kwargs)
|
|
|
|
# --- Shared alarm SNS topic (site-alerts) ---
|
|
# Imported once near the top so every alarm in this stack reuses the same
|
|
# Topic construct instance (avoids duplicate logical IDs). ALARM-only
|
|
# SnsAction; no OK action, per the CloudWatch-alarm preference. The
|
|
# topic's CMK (alias/seahaven-alarm-topics) lives on the topic itself.
|
|
alarm_topic = sns.Topic.from_topic_arn(
|
|
self,
|
|
"SiteAlertsTopic",
|
|
f"arn:aws:sns:{self.region}:{self.account}:site-alerts",
|
|
)
|
|
|
|
# --- S3 bucket for raw emails ---
|
|
email_bucket = s3.Bucket(
|
|
self,
|
|
"EmailBucket",
|
|
bucket_name=f"po-ingest-emails-{self.account}",
|
|
block_public_access=s3.BlockPublicAccess.BLOCK_ALL,
|
|
removal_policy=RemovalPolicy.RETAIN,
|
|
lifecycle_rules=[
|
|
s3.LifecycleRule(expiration=Duration.days(90)),
|
|
],
|
|
)
|
|
|
|
# --- Shared customer-managed CMK for sensitive DynamoDB tables ---
|
|
# Owned by the account-baseline app (alias/seahaven-dynamodb, INFRA-95 /
|
|
# M-3); ARN published to SSM. The purchase-orders table was migrated to
|
|
# SSE-KMS out-of-band, so declaring encryption_key here reconciles the
|
|
# drift and — via grant_read_write_data below — propagates the required
|
|
# kms:Decrypt/GenerateDataKey/DescribeKey to the consumer roles.
|
|
dynamodb_cmk = kms.Key.from_key_arn(
|
|
self,
|
|
"DynamoDbCmk",
|
|
ssm.StringParameter.value_for_string_parameter(
|
|
self, "/seahaven/dynamodb/cmk-arn"
|
|
),
|
|
)
|
|
|
|
# --- Purchase-orders DynamoDB table ---
|
|
# Owned by this stack. Streams enabled for the site-extractor pipeline.
|
|
# Other stacks (seahaven-slack-bot) reference this table via fromTableName().
|
|
po_table = dynamodb.Table(
|
|
self,
|
|
"PurchaseOrdersTable",
|
|
table_name="purchase-orders",
|
|
partition_key=dynamodb.Attribute(
|
|
name="po_number",
|
|
type=dynamodb.AttributeType.STRING,
|
|
),
|
|
billing_mode=dynamodb.BillingMode.PAY_PER_REQUEST,
|
|
removal_policy=RemovalPolicy.RETAIN,
|
|
stream=dynamodb.StreamViewType.NEW_IMAGE,
|
|
encryption=dynamodb.TableEncryption.CUSTOMER_MANAGED,
|
|
encryption_key=dynamodb_cmk,
|
|
)
|
|
|
|
# --- Anthropic API key secret removed (Bedrock migration) ---
|
|
# PO parsing stays fully AI but moved from the Anthropic API to the
|
|
# Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0,
|
|
# so no provider API key is needed. The old secret
|
|
# "po-ingest/anthropic-api-key" had RemovalPolicy.RETAIN, so it is
|
|
# ORPHANED (not deleted) by this change: delete it manually post-deploy
|
|
# and revoke the stored key at Anthropic.
|
|
|
|
# --- DLQ for failed async invocations (INFRA-41 / audit H-8) ---
|
|
# SES → S3 → Lambda is async; without an OnFailure destination a failed
|
|
# parse (bad email, transient error) is silently dropped after Lambda's
|
|
# retries. CDK generates the queue name to avoid colliding with the
|
|
# interim CLI-created po-email-processor-dlq (removed post-deploy).
|
|
email_processor_dlq = sqs.Queue(
|
|
self,
|
|
"EmailProcessorDlq",
|
|
retention_period=Duration.days(14),
|
|
enforce_ssl=True,
|
|
)
|
|
|
|
# --- Lambda function ---
|
|
email_processor = lambda_.Function(
|
|
self,
|
|
"EmailProcessor",
|
|
function_name="po-email-processor",
|
|
runtime=lambda_.Runtime.PYTHON_3_12,
|
|
architecture=lambda_.Architecture.ARM_64,
|
|
handler="handler.handler",
|
|
code=lambda_.Code.from_asset(
|
|
"../lambdas/po/email_processor",
|
|
bundling=cdk.BundlingOptions(
|
|
image=lambda_.Runtime.PYTHON_3_12.bundling_image,
|
|
command=[
|
|
"bash",
|
|
"-c",
|
|
"pip install --platform manylinux2014_aarch64 --only-binary=:all: "
|
|
"-r requirements.txt -t /asset-output && "
|
|
"cp handler.py ses_auth.py /asset-output/",
|
|
],
|
|
),
|
|
),
|
|
timeout=Duration.seconds(60),
|
|
memory_size=256,
|
|
log_retention=logs.RetentionDays.TWO_MONTHS,
|
|
dead_letter_queue=email_processor_dlq,
|
|
environment={
|
|
"PO_TABLE": "purchase-orders",
|
|
"BEDROCK_MODEL_ID": "us.anthropic.claude-haiku-4-5-20251001-v1:0",
|
|
# Fail-closed sender auth (INFRA-107): the handler only
|
|
# accepts mail whose SES-stamped Authentication-Results
|
|
# header carries dkim=pass for one of these domains.
|
|
# Observed on live traffic 2026-07-15: Coupa PO mail passes
|
|
# DKIM for amazon.coupahost.com (and amazonses.com, which is
|
|
# deliberately NOT allowlisted — every SES customer's mail
|
|
# passes that). Unset/empty ⇒ the handler rejects all mail.
|
|
"ALLOWED_DKIM_DOMAINS": "amazon.coupahost.com",
|
|
},
|
|
)
|
|
|
|
# Grant permissions
|
|
email_bucket.grant_read(email_processor)
|
|
po_table.grant_read_write_data(email_processor)
|
|
|
|
# --- Bedrock InvokeModel grant ---
|
|
# The us.* inference profile can route cross-region, so the grant MUST
|
|
# cover both the inference-profile ARN AND the per-region foundation-model
|
|
# ARNs (empty account field) for every region the profile can reach
|
|
# (us-east-1/us-east-2/us-west-2). A profile-only grant AccessDenies at
|
|
# runtime whenever the profile routes to a region whose foundation-model
|
|
# ARN is not allowed.
|
|
email_processor.add_to_role_policy(
|
|
iam.PolicyStatement(
|
|
actions=[
|
|
"bedrock:InvokeModel",
|
|
"bedrock:InvokeModelWithResponseStream",
|
|
],
|
|
resources=[
|
|
"arn:aws:bedrock:us-east-1:328440206208:inference-profile/us.anthropic.claude-haiku-4-5-20251001-v1:0",
|
|
"arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-haiku-4-5-20251001-v1:0",
|
|
"arn:aws:bedrock:us-east-2::foundation-model/anthropic.claude-haiku-4-5-20251001-v1:0",
|
|
"arn:aws:bedrock:us-west-2::foundation-model/anthropic.claude-haiku-4-5-20251001-v1:0",
|
|
],
|
|
)
|
|
)
|
|
|
|
# --- Errors alarm (INFRA-41 / audit H-8) ---
|
|
# ALARM-only (no OK action, per the CloudWatch-alarm preference) to the
|
|
# shared site-alerts topic. Any errored invocation in a 5-min window pages.
|
|
email_processor.metric_errors(
|
|
period=Duration.minutes(5),
|
|
statistic="Sum",
|
|
).create_alarm(
|
|
self,
|
|
"EmailProcessorErrorsAlarm",
|
|
alarm_name="po-email-processor-errors",
|
|
alarm_description="po-email-processor async invocation errors",
|
|
threshold=0,
|
|
evaluation_periods=1,
|
|
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
|
|
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
|
|
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
|
|
|
|
# --- Sender-auth rejection alarm (INFRA-107) ---
|
|
# A rejected email (bad/unaligned DKIM verdict) returns normally, so it
|
|
# produces NO Lambda error, NO DLQ message and NO retry -- only a
|
|
# `sender_auth_rejected` warning log. Without this metric filter + alarm a
|
|
# domain drift (Coupa rotates its signing subdomain, SES changes its
|
|
# Authentication-Results format, the allowlist is wrong) would silently
|
|
# discard 100% of legitimate PO mail while every other alarm stays green.
|
|
# A CloudWatch Logs metric filter turns those warnings into a metric so a
|
|
# false-reject storm pages instead of vanishing. default_value=0 keeps the
|
|
# series populated (alarm stays OK, never INSUFFICIENT_DATA) between events.
|
|
_add_sender_auth_rejected_alarm(
|
|
self, "EmailProcessor", "po-email-processor", alarm_topic
|
|
)
|
|
|
|
# --- Throttles alarm: po-email-processor ---
|
|
# Any throttled invocation (concurrency cap hit) in a 5-min window pages.
|
|
# ALARM-only to site-alerts; no OK action; NOT_BREACHING when no data.
|
|
email_processor.metric_throttles(
|
|
period=Duration.minutes(5),
|
|
statistic="Sum",
|
|
).create_alarm(
|
|
self,
|
|
"EmailProcessorThrottlesAlarm",
|
|
alarm_name="po-email-processor-throttles",
|
|
alarm_description="po-email-processor invocation throttles",
|
|
threshold=0,
|
|
evaluation_periods=1,
|
|
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
|
|
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
|
|
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
|
|
|
|
# --- DLQ messages-present alarm ---
|
|
# Pages when any message lands in the EmailProcessorDlq: a message here
|
|
# means a PO email was permanently dropped after Lambda exhausted its
|
|
# async retries. Maximum over a single 5-min window > 0 fires; missing
|
|
# data (no messages metric emitted) is not breaching. Reuses the shared
|
|
# site-alerts topic, ALARM-only, like the errors alarm above. The metric
|
|
# helper derives the QueueName dimension from the queue construct, so the
|
|
# alarm tracks the CDK-generated queue name without hardcoding it.
|
|
email_processor_dlq.metric_approximate_number_of_messages_visible(
|
|
period=Duration.minutes(5),
|
|
statistic="Maximum",
|
|
).create_alarm(
|
|
self,
|
|
"EmailProcessorDlqMessagesAlarm",
|
|
alarm_name="po-email-processor-dlq-messages",
|
|
alarm_description="po-email-processor DLQ has messages (dropped PO emails)",
|
|
threshold=0,
|
|
evaluation_periods=1,
|
|
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
|
|
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
|
|
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
|
|
|
|
# --- Duration alarm: po-email-processor (orphan adoption) ---
|
|
# Adopts the orphaned CLI alarm Lambda-Duration-po-email-processor under
|
|
# the repo's <fn>-duration naming (NEW logical name → no deploy collision;
|
|
# delete the orphan post-deploy). p99 / 45000 ms
|
|
# (75% of the 60s timeout) / eval 3 of 3 — tighter than the orphan's
|
|
# Maximum>=48000 / 1-of-1.
|
|
email_processor.metric_duration(
|
|
period=Duration.minutes(5),
|
|
statistic="p99",
|
|
).create_alarm(
|
|
self,
|
|
"EmailProcessorDurationAlarm",
|
|
alarm_name="po-email-processor-duration",
|
|
alarm_description="po-email-processor p99 duration approaching the 60s timeout",
|
|
threshold=45000,
|
|
evaluation_periods=3,
|
|
datapoints_to_alarm=2,
|
|
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
|
|
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
|
|
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
|
|
|
|
# S3 event notification → Lambda
|
|
email_bucket.add_event_notification(
|
|
s3.EventType.OBJECT_CREATED,
|
|
s3n.LambdaDestination(email_processor),
|
|
s3.NotificationKeyFilter(prefix="inbound/"),
|
|
)
|
|
|
|
# --- SES Receipt Rule ---
|
|
# Reuse the existing INBOUND_MAIL rule set (shared with workorder-ingest)
|
|
rule_set = ses.ReceiptRuleSet.from_receipt_rule_set_name(
|
|
self,
|
|
"ExistingRuleSet",
|
|
"INBOUND_MAIL",
|
|
)
|
|
|
|
rule_set.add_rule(
|
|
"PoEmailRule",
|
|
recipients=["amazon_po@int.seahaven.com"],
|
|
actions=[
|
|
ses_actions.S3(
|
|
bucket=email_bucket,
|
|
object_key_prefix="inbound/",
|
|
),
|
|
],
|
|
)
|
|
|
|
# --- Web UI auth token secret ---
|
|
# Shared secret for the web UI auth gate, stored in Secrets Manager and
|
|
# resolved at runtime so the token never appears in CloudFormation templates
|
|
# or Lambda environment variables. Create this secret before deploying
|
|
# either stack; both PO and WO stacks reference it by name.
|
|
web_ui_auth_secret = secretsmanager.Secret.from_secret_name_v2(
|
|
self,
|
|
"WebUiAuthToken",
|
|
"procurement-ingest/web-ui-auth-token",
|
|
)
|
|
|
|
# --- Web UI Lambda ---
|
|
web_ui = lambda_.Function(
|
|
self,
|
|
"WebUI",
|
|
function_name="po-web-ui",
|
|
runtime=lambda_.Runtime.PYTHON_3_12,
|
|
architecture=lambda_.Architecture.ARM_64,
|
|
handler="handler.handler",
|
|
code=lambda_.Code.from_asset("../lambdas/po/web_ui"),
|
|
timeout=Duration.seconds(60),
|
|
memory_size=256,
|
|
log_retention=logs.RetentionDays.TWO_MONTHS,
|
|
environment={
|
|
"PO_TABLE": "purchase-orders",
|
|
# Defense-in-depth shared secret for the web UI handler. The
|
|
# handler fails closed if this ARN is unset or the secret is
|
|
# missing, so any future invocation path cannot re-expose the
|
|
# PO DB unauthenticated. The secret value is fetched at runtime
|
|
# from Secrets Manager (not embedded in env vars or template).
|
|
"WEB_UI_AUTH_TOKEN_SECRET_ARN": web_ui_auth_secret.secret_arn,
|
|
},
|
|
)
|
|
|
|
po_table.grant_read_data(web_ui)
|
|
web_ui_auth_secret.grant_read(web_ui)
|
|
|
|
# --- Throttles alarm: po-web-ui ---
|
|
web_ui.metric_throttles(
|
|
period=Duration.minutes(5),
|
|
statistic="Sum",
|
|
).create_alarm(
|
|
self,
|
|
"WebUiThrottlesAlarm",
|
|
alarm_name="po-web-ui-throttles",
|
|
alarm_description="po-web-ui invocation throttles",
|
|
threshold=0,
|
|
evaluation_periods=1,
|
|
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
|
|
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
|
|
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
|
|
|
|
# --- Duration alarm: po-web-ui ---
|
|
# Net-new (no orphan exists for this function).
|
|
# p99 / 45000 ms (75% of the 60s timeout) / eval 3, datapoints 2.
|
|
web_ui.metric_duration(
|
|
period=Duration.minutes(5),
|
|
statistic="p99",
|
|
).create_alarm(
|
|
self,
|
|
"WebUiDurationAlarm",
|
|
alarm_name="po-web-ui-duration",
|
|
alarm_description="po-web-ui p99 duration approaching the 60s timeout",
|
|
threshold=45000,
|
|
evaluation_periods=3,
|
|
datapoints_to_alarm=2,
|
|
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
|
|
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
|
|
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
|
|
|
|
# Public Function URL removed 2026-06-08 (INFRA-74 / audit C-5): the
|
|
# unauthenticated FunctionUrlAuthType.NONE URL was deleted out-of-band
|
|
# via CLI. Removing the construct (and its auto-generated Principal:*
|
|
# invoke permission) reconciles IaC with the live state.
|
|
|
|
# --- Verified sites table (extracted from PO ship-to addresses) ---
|
|
verified_sites_table = dynamodb.Table(
|
|
self,
|
|
"VerifiedSitesTable",
|
|
table_name="verified-sites",
|
|
partition_key=dynamodb.Attribute(
|
|
name="siteCode",
|
|
type=dynamodb.AttributeType.STRING,
|
|
),
|
|
billing_mode=dynamodb.BillingMode.PAY_PER_REQUEST,
|
|
removal_policy=RemovalPolicy.RETAIN,
|
|
)
|
|
# by-state GSI removed 2026-06-03 (audit M-20): 0 reads in 30d against
|
|
# 518 WCU of write amplification. Re-add if a state-level query path ships.
|
|
|
|
# --- Site extractor Lambda (DynamoDB Streams → verified-sites) ---
|
|
site_extractor = lambda_.Function(
|
|
self,
|
|
"SiteExtractor",
|
|
function_name="po-ingest-site-extractor",
|
|
runtime=lambda_.Runtime.PYTHON_3_12,
|
|
architecture=lambda_.Architecture.ARM_64,
|
|
handler="handler.handler",
|
|
code=lambda_.Code.from_asset("../lambdas/po/site_extractor"),
|
|
timeout=Duration.seconds(60),
|
|
memory_size=256,
|
|
log_retention=logs.RetentionDays.TWO_MONTHS,
|
|
environment={
|
|
"VERIFIED_SITES_TABLE": verified_sites_table.table_name,
|
|
"PENDING_REVIEW_TABLE": "pending-site-review",
|
|
},
|
|
)
|
|
|
|
verified_sites_table.grant_read_write_data(site_extractor)
|
|
|
|
site_extractor.add_event_source(
|
|
lambda_event_sources.DynamoEventSource(
|
|
po_table,
|
|
starting_position=lambda_.StartingPosition.TRIM_HORIZON,
|
|
batch_size=10,
|
|
max_batching_window=Duration.seconds(30),
|
|
bisect_batch_on_error=True,
|
|
retry_attempts=3,
|
|
)
|
|
)
|
|
|
|
# --- Errors alarm: po-ingest-site-extractor ---
|
|
# Stream-consumer errors retry per the event-source config, but a
|
|
# persistent failure stalls the verified-sites pipeline. ALARM-only to
|
|
# site-alerts; no OK action; NOT_BREACHING when no data.
|
|
site_extractor.metric_errors(
|
|
period=Duration.minutes(5),
|
|
statistic="Sum",
|
|
).create_alarm(
|
|
self,
|
|
"SiteExtractorErrorsAlarm",
|
|
alarm_name="po-ingest-site-extractor-errors",
|
|
alarm_description="po-ingest-site-extractor invocation errors",
|
|
threshold=0,
|
|
evaluation_periods=1,
|
|
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
|
|
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
|
|
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
|
|
|
|
# --- Throttles alarm: po-ingest-site-extractor ---
|
|
site_extractor.metric_throttles(
|
|
period=Duration.minutes(5),
|
|
statistic="Sum",
|
|
).create_alarm(
|
|
self,
|
|
"SiteExtractorThrottlesAlarm",
|
|
alarm_name="po-ingest-site-extractor-throttles",
|
|
alarm_description="po-ingest-site-extractor invocation throttles",
|
|
threshold=0,
|
|
evaluation_periods=1,
|
|
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
|
|
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
|
|
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
|
|
|
|
# --- Duration alarm: po-ingest-site-extractor ---
|
|
# Net-new (no orphan exists for this function).
|
|
# p99 / 45000 ms (75% of the 60s timeout) / eval 3, datapoints 2.
|
|
site_extractor.metric_duration(
|
|
period=Duration.minutes(5),
|
|
statistic="p99",
|
|
).create_alarm(
|
|
self,
|
|
"SiteExtractorDurationAlarm",
|
|
alarm_name="po-ingest-site-extractor-duration",
|
|
alarm_description="po-ingest-site-extractor p99 duration approaching the 60s timeout",
|
|
threshold=45000,
|
|
evaluation_periods=3,
|
|
datapoints_to_alarm=2,
|
|
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
|
|
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
|
|
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
|
|
|
|
cdk.CfnOutput(
|
|
self,
|
|
"VerifiedSitesTableName",
|
|
value=verified_sites_table.table_name,
|
|
description="Verified site addresses extracted from POs",
|
|
)
|
|
|
|
# --- Pending site review table (POs with no extractable site code) ---
|
|
pending_review_table = dynamodb.Table(
|
|
self,
|
|
"PendingSiteReviewTable",
|
|
table_name="pending-site-review",
|
|
partition_key=dynamodb.Attribute(
|
|
name="po_number",
|
|
type=dynamodb.AttributeType.STRING,
|
|
),
|
|
billing_mode=dynamodb.BillingMode.PAY_PER_REQUEST,
|
|
removal_policy=RemovalPolicy.RETAIN,
|
|
)
|
|
|
|
pending_review_table.grant_read_write_data(site_extractor)
|
|
verified_sites_table.grant_read_data(site_extractor)
|
|
|
|
# --- DynamoDB throttle + system-error alarms ---
|
|
# ThrottledRequests / SystemErrors emit at TableName + Operation only
|
|
# (verified against live CloudWatch: no TableName-only rollup exists, and
|
|
# metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.259.0).
|
|
# Each table currently has zero throttle/error datapoints, so the series
|
|
# only materialise on first occurrence — NOT_BREACHING keeps them OK until
|
|
# then.
|
|
_add_ddb_alarms(
|
|
self, "PurchaseOrdersTable", po_table, "purchase-orders", alarm_topic
|
|
)
|
|
_add_ddb_alarms(
|
|
self,
|
|
"VerifiedSitesTable",
|
|
verified_sites_table,
|
|
"verified-sites",
|
|
alarm_topic,
|
|
)
|
|
_add_ddb_alarms(
|
|
self,
|
|
"PendingSiteReviewTable",
|
|
pending_review_table,
|
|
"pending-site-review",
|
|
alarm_topic,
|
|
)
|