2026-04-07 12:12:30 -04:00
|
|
|
"""CDK stack for the Coupa PO email ingestion pipeline."""
|
|
|
|
|
|
|
|
|
|
import aws_cdk as cdk
|
|
|
|
|
from aws_cdk import (
|
|
|
|
|
Duration,
|
|
|
|
|
RemovalPolicy,
|
|
|
|
|
Stack,
|
Reconcile IaC with out-of-band DLQ + Function URL changes (INFRA-74, INFRA-41) (#50)
Make CDK the source of truth for two sets of changes applied out-of-band
via CLI to the po-ingest and WorkorderIngestStack stacks.
INFRA-74 (audit C-5): remove the public FunctionUrlAuthType.NONE Function
URL construct (and its auto-generated Principal:* invoke permission +
output) from both po-web-ui and workorder-web-ui. The URLs were already
deleted live via CLI; CFN's delete is idempotent.
INFRA-41 (audit H-8): add a CDK-managed SQS dead-letter queue
(dead_letter_queue=, 14d retention, SSL-enforced, CDK-generated name) and
an ALARM-only Errors alarm (Sum, threshold>0, site-alerts topic) for both
po-email-processor and workorder-email-processor, mirroring the
apm-wo-analysis-classifier DLQ and payments-payroll-batch alarm patterns.
Interim CLI resources (per-fn -dlq queues, -errors alarms, dlq-send inline
policies, OnFailure event-invoke-configs) removed post-deploy.
2026-06-08 16:02:29 -04:00
|
|
|
aws_cloudwatch as cloudwatch,
|
|
|
|
|
aws_cloudwatch_actions as cw_actions,
|
2026-04-07 12:12:30 -04:00
|
|
|
aws_dynamodb as dynamodb,
|
2026-06-10 19:31:55 -04:00
|
|
|
aws_kms as kms,
|
2026-04-07 12:12:30 -04:00
|
|
|
aws_lambda as lambda_,
|
2026-04-30 14:26:53 -04:00
|
|
|
aws_lambda_event_sources as lambda_event_sources,
|
|
|
|
|
aws_logs as logs,
|
2026-04-07 12:12:30 -04:00
|
|
|
aws_s3 as s3,
|
|
|
|
|
aws_s3_notifications as s3n,
|
|
|
|
|
aws_ses as ses,
|
|
|
|
|
aws_ses_actions as ses_actions,
|
|
|
|
|
aws_secretsmanager as secretsmanager,
|
Reconcile IaC with out-of-band DLQ + Function URL changes (INFRA-74, INFRA-41) (#50)
Make CDK the source of truth for two sets of changes applied out-of-band
via CLI to the po-ingest and WorkorderIngestStack stacks.
INFRA-74 (audit C-5): remove the public FunctionUrlAuthType.NONE Function
URL construct (and its auto-generated Principal:* invoke permission +
output) from both po-web-ui and workorder-web-ui. The URLs were already
deleted live via CLI; CFN's delete is idempotent.
INFRA-41 (audit H-8): add a CDK-managed SQS dead-letter queue
(dead_letter_queue=, 14d retention, SSL-enforced, CDK-generated name) and
an ALARM-only Errors alarm (Sum, threshold>0, site-alerts topic) for both
po-email-processor and workorder-email-processor, mirroring the
apm-wo-analysis-classifier DLQ and payments-payroll-batch alarm patterns.
Interim CLI resources (per-fn -dlq queues, -errors alarms, dlq-send inline
policies, OnFailure event-invoke-configs) removed post-deploy.
2026-06-08 16:02:29 -04:00
|
|
|
aws_sns as sns,
|
2026-06-10 19:31:55 -04:00
|
|
|
aws_ssm as ssm,
|
2026-04-07 12:12:30 -04:00
|
|
|
)
|
|
|
|
|
from constructs import Construct
|
|
|
|
|
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
import common
|
Add fail-closed SES sender authentication (INFRA-107) (#98)
* Add fail-closed SES sender authentication
The From header and any raw-MIME Authentication-Results copies are
attacker-forgeable, so a forged email to apm@int.seahaven.com or
amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107,
CRITICAL). Both S3-triggered email processors now authenticate the
sender against the Authentication-Results header SES itself prepends
at delivery: only the topmost header is consulted, its authserv-id
must be amazonses.com, and it must carry dkim=pass for a domain in
the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set
in CDK so ops can adjust without code changes).
Allowlists come from live traffic observed 2026-07-15 on both ingest
buckets: WO mail arrives via the apm@ Google Groups forward, which
re-signs as seahaven.com (the hxgnsmartcloud.com signature does not
survive the forward); PO mail passes for amazon.coupahost.com.
amazonses.com also passes on PO mail but is deliberately excluded --
every SES customer's outbound mail passes for it.
Every failure path (env var unset, header missing or unparseable,
verdict fail, unaligned domain) rejects the email: a structured
warning with the reason and S3 key is logged and the record skipped
without erroring the invocation, so rejected mail causes no Lambda
retries or DLQ messages. Handler signatures and event sources are
unchanged.
Refs: INFRA-107
* Harden AR parser per cross-family review
Cross-family (GPT-4.1) review findings: terminate the dkim result
token at end-of-clause, whitespace, or a comment so a value like
"dkim=pass-fake" can never be read as a pass; normalize trailing
dots off allowlist entries so "seahaven.com." matches; make the
compat32 parser policy explicit. Adds tests for result-token
boundaries, comments after the result, quoted domain values, and
folding inside a dkim clause.
Refs: INFRA-107
* Harden AR parsing and alarm on sender-auth rejects
The SES-stamped Authentication-Results value echoes attacker-controlled
SMTP-session tokens (envelope-from, helo, header.from) as their own
semicolon-delimited property clauses. A naive split(";") tore an RFC 5321
quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass
clause, so a fully spoofed email was accepted on the genuinely
SES-stamped topmost header. Tokenise comment- and quoted-string-aware
(RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on
semicolons outside a quoted-string, and fail closed on unbalanced
quotes/comments so a ';' inside a quoted pvalue can never start a clause.
Rejected mail returns normally (no error, no retry, no DLQ message), so a
signing-domain drift or a wrong allowlist would silently discard 100% of
legitimate mail while every alarm stayed green. Add a CloudWatch Logs
metric filter + alarm on the sender_auth_rejected warning to both stacks
so a false-reject storm pages instead of vanishing. This is also the
safety net for the WO seahaven.com allowlist assumption, which must be
validated against a live SES-stamped header (a plain Gmail auto-forward
re-signs under the sending Workspace domain, not seahaven.com).
Refs: INFRA-107
* chore: retrigger CI (no run recorded for 7c74ac1)
* Fix quoted-AUID DKIM domain spoof in sender auth
Resolve three confirmed /sh-security-review findings on the fail-closed
SES sender-authentication control.
HIGH: header.i/header.d domain extraction was not quoted-string aware.
An attacker with a valid DKIM key for their own domain could set an
RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive
extractor stopped at the closing quote and returned seahaven.com,
accepting forged mail. Extraction now tokenises the clause with the same
quoted-string discipline already used for clause splitting: header.d
(the plain signing domain) is authoritative when present, otherwise the
header.i domain is the part after the AUID's LAST top-level "@", so a "@"
inside a quoted local-part is treated as signer-controlled label text and
yields the true signer (attacker.com), not seahaven.com.
LOW: the topmost-header parse ran outside evaluate_sender_authentication's
try/except, so an unexpected parser exception on crafted input could
propagate into the handler and Lambda async retries/DLQ. The parse now
fails CLOSED with an authentication_results_unparseable reason.
MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to
a low-volume total-reject outage (a trickle that never sums to 3). Both
stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3
/ datapoints_to_alarm=2, so a sustained reject condition pages even at one
reject per period while a lone stray probe self-clears.
Refs: INFRA-107
* Load Lambda function dir on sys.path in tests
Rebasing INFRA-107 onto main folded #95's pytest suite into this
branch's tests. The unified conftest loads the PO/WO handlers by file
path, and handler.py now does `from ses_auth import
authenticate_inbound_email` -- a bare sibling import that resolves in
the Lambda only because the runtime puts each function's own directory
on sys.path. The shared load_handler now adds that directory so the
handler tests import correctly alongside the sender-auth tests.
Refs: INFRA-107
* Note #97 test files in README directory tree
The rebase onto main brought in #97's tests/requirements.txt and
tests/test_po_merge.py. List both in the directory tree so it matches
the tree on disk.
Refs: INFRA-107
* Document INFRA-107 forwarder-binding risk acceptance
Record the accepted risk that WO sender auth binds to the apm@ forward's
re-signing domain (seahaven.com) rather than the Hexagon originator; the
apm@ Google Group's restricted posting policy is the load-bearing control
(escalates to HIGH if the group is opened to external posting). Also
correct the sender-auth-rejected alarm docs to match the shipped config
(>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and
note the SES-AR-01/02 parser hardening follow-ups.
Refs: INFRA-107
2026-07-15 20:58:47 -04:00
|
|
|
|
|
|
|
|
|
2026-04-07 12:12:30 -04:00
|
|
|
class PoIngestStack(Stack):
|
|
|
|
|
def __init__(self, scope: Construct, construct_id: str, **kwargs):
|
|
|
|
|
super().__init__(scope, construct_id, **kwargs)
|
|
|
|
|
|
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70)
* Add CloudWatch alarm coverage for po-ingest and workorder-ingest
Expands alarm coverage across both CDK stacks. All alarms are ALARM-only
(no OK action) to the shared site-alerts SNS topic, with TreatMissingData
NOT_BREACHING. The site-alerts topic is now imported once near the top of
each stack so every alarm reuses one Topic instance.
po-ingest (cdk/po_stack.py):
- Errors: po-ingest-site-extractor
- Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui
- Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption),
po-ingest-site-extractor, po-web-ui
- DynamoDB throttle + system-error: purchase-orders, verified-sites,
pending-site-review
workorder-ingest (cdk/wo_stack.py):
- Throttles: workorder-email-processor
- Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption)
- DynamoDB throttle + system-error: WorkOrders, WorkOrderComments
DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation
dimension set, so each table alarm is a Sum math expression across operations
via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests
is deprecated/invalid in aws-cdk-lib 2.259.0).
Refs INFRA-41 / audit H-8.
* Drop NEEDS ADAM SIGN-OFF wording from alarm comments
Duration alarm thresholds are owner-approved; remove the sign-off flag
from po_stack.py and wo_stack.py comments. Threshold values, eval config,
and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
|
|
|
# --- Shared alarm SNS topic (site-alerts) ---
|
|
|
|
|
# Imported once near the top so every alarm in this stack reuses the same
|
|
|
|
|
# Topic construct instance (avoids duplicate logical IDs). ALARM-only
|
|
|
|
|
# SnsAction; no OK action, per the CloudWatch-alarm preference. The
|
|
|
|
|
# topic's CMK (alias/seahaven-alarm-topics) lives on the topic itself.
|
|
|
|
|
alarm_topic = sns.Topic.from_topic_arn(
|
|
|
|
|
self,
|
|
|
|
|
"SiteAlertsTopic",
|
|
|
|
|
f"arn:aws:sns:{self.region}:{self.account}:site-alerts",
|
|
|
|
|
)
|
|
|
|
|
|
2026-04-07 12:12:30 -04:00
|
|
|
# --- S3 bucket for raw emails ---
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
email_bucket = common.make_email_bucket(self, "EmailBucket", "po-ingest-emails")
|
2026-04-07 12:12:30 -04:00
|
|
|
|
2026-06-10 19:31:55 -04:00
|
|
|
# --- Shared customer-managed CMK for sensitive DynamoDB tables ---
|
|
|
|
|
# Owned by the account-baseline app (alias/seahaven-dynamodb, INFRA-95 /
|
|
|
|
|
# M-3); ARN published to SSM. The purchase-orders table was migrated to
|
|
|
|
|
# SSE-KMS out-of-band, so declaring encryption_key here reconciles the
|
|
|
|
|
# drift and — via grant_read_write_data below — propagates the required
|
|
|
|
|
# kms:Decrypt/GenerateDataKey/DescribeKey to the consumer roles.
|
|
|
|
|
dynamodb_cmk = kms.Key.from_key_arn(
|
|
|
|
|
self,
|
|
|
|
|
"DynamoDbCmk",
|
|
|
|
|
ssm.StringParameter.value_for_string_parameter(
|
|
|
|
|
self, "/seahaven/dynamodb/cmk-arn"
|
|
|
|
|
),
|
|
|
|
|
)
|
|
|
|
|
|
2026-04-30 14:26:53 -04:00
|
|
|
# --- Purchase-orders DynamoDB table ---
|
|
|
|
|
# Owned by this stack. Streams enabled for the site-extractor pipeline.
|
|
|
|
|
# Other stacks (seahaven-slack-bot) reference this table via fromTableName().
|
|
|
|
|
po_table = dynamodb.Table(
|
2026-05-08 16:01:21 -04:00
|
|
|
self,
|
|
|
|
|
"PurchaseOrdersTable",
|
2026-04-30 14:26:53 -04:00
|
|
|
table_name="purchase-orders",
|
|
|
|
|
partition_key=dynamodb.Attribute(
|
|
|
|
|
name="po_number",
|
|
|
|
|
type=dynamodb.AttributeType.STRING,
|
|
|
|
|
),
|
|
|
|
|
billing_mode=dynamodb.BillingMode.PAY_PER_REQUEST,
|
|
|
|
|
removal_policy=RemovalPolicy.RETAIN,
|
Align PO schema with enriched records and improve extraction prompt
Replaces extraction prompt with domain-specific rules: trade
classification taxonomy (23 categories), site_code skip list,
zip padding, revision email type, and structured extraction for
fiscal_year, trade, and coupa_category.
Handler changes:
- New "revision" email type overwrites existing PO via put_item
- enrich_parsed() adds top-level state, ship_to_raw, data_source
- pad_zip() zero-pads short zip codes (e.g., "7001" → "07001")
- Removed invoice_total/invoice_count (Payee Central only)
Web UI: added revision badge, new detail fields (site code, state,
trade, fiscal year, coupa category, data source), line item table
now shows Qty/Unit/Price columns, list view shows Site and Trade.
CDK: fixed StreamViewType to match deployed table (NEW_IMAGE).
README: documented PO record schema and revision flow.
2026-05-01 19:40:18 -04:00
|
|
|
stream=dynamodb.StreamViewType.NEW_IMAGE,
|
2026-06-10 19:31:55 -04:00
|
|
|
encryption=dynamodb.TableEncryption.CUSTOMER_MANAGED,
|
|
|
|
|
encryption_key=dynamodb_cmk,
|
2026-04-07 12:12:30 -04:00
|
|
|
)
|
|
|
|
|
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
# --- Anthropic API key secret removed (Bedrock migration) ---
|
|
|
|
|
# PO parsing stays fully AI but moved from the Anthropic API to the
|
|
|
|
|
# Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0,
|
|
|
|
|
# so no provider API key is needed. The old secret
|
|
|
|
|
# "po-ingest/anthropic-api-key" had RemovalPolicy.RETAIN, so it is
|
|
|
|
|
# ORPHANED (not deleted) by this change: delete it manually post-deploy
|
|
|
|
|
# and revoke the stored key at Anthropic.
|
2026-04-07 12:12:30 -04:00
|
|
|
|
Reconcile IaC with out-of-band DLQ + Function URL changes (INFRA-74, INFRA-41) (#50)
Make CDK the source of truth for two sets of changes applied out-of-band
via CLI to the po-ingest and WorkorderIngestStack stacks.
INFRA-74 (audit C-5): remove the public FunctionUrlAuthType.NONE Function
URL construct (and its auto-generated Principal:* invoke permission +
output) from both po-web-ui and workorder-web-ui. The URLs were already
deleted live via CLI; CFN's delete is idempotent.
INFRA-41 (audit H-8): add a CDK-managed SQS dead-letter queue
(dead_letter_queue=, 14d retention, SSL-enforced, CDK-generated name) and
an ALARM-only Errors alarm (Sum, threshold>0, site-alerts topic) for both
po-email-processor and workorder-email-processor, mirroring the
apm-wo-analysis-classifier DLQ and payments-payroll-batch alarm patterns.
Interim CLI resources (per-fn -dlq queues, -errors alarms, dlq-send inline
policies, OnFailure event-invoke-configs) removed post-deploy.
2026-06-08 16:02:29 -04:00
|
|
|
# --- DLQ for failed async invocations (INFRA-41 / audit H-8) ---
|
|
|
|
|
# SES → S3 → Lambda is async; without an OnFailure destination a failed
|
|
|
|
|
# parse (bad email, transient error) is silently dropped after Lambda's
|
|
|
|
|
# retries. CDK generates the queue name to avoid colliding with the
|
|
|
|
|
# interim CLI-created po-email-processor-dlq (removed post-deploy).
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
email_processor_dlq = common.make_processor_dlq(self, "EmailProcessorDlq")
|
Reconcile IaC with out-of-band DLQ + Function URL changes (INFRA-74, INFRA-41) (#50)
Make CDK the source of truth for two sets of changes applied out-of-band
via CLI to the po-ingest and WorkorderIngestStack stacks.
INFRA-74 (audit C-5): remove the public FunctionUrlAuthType.NONE Function
URL construct (and its auto-generated Principal:* invoke permission +
output) from both po-web-ui and workorder-web-ui. The URLs were already
deleted live via CLI; CFN's delete is idempotent.
INFRA-41 (audit H-8): add a CDK-managed SQS dead-letter queue
(dead_letter_queue=, 14d retention, SSL-enforced, CDK-generated name) and
an ALARM-only Errors alarm (Sum, threshold>0, site-alerts topic) for both
po-email-processor and workorder-email-processor, mirroring the
apm-wo-analysis-classifier DLQ and payments-payroll-batch alarm patterns.
Interim CLI resources (per-fn -dlq queues, -errors alarms, dlq-send inline
policies, OnFailure event-invoke-configs) removed post-deploy.
2026-06-08 16:02:29 -04:00
|
|
|
|
2026-04-07 12:12:30 -04:00
|
|
|
# --- Lambda function ---
|
|
|
|
|
email_processor = lambda_.Function(
|
2026-05-08 16:01:21 -04:00
|
|
|
self,
|
|
|
|
|
"EmailProcessor",
|
2026-04-07 12:12:30 -04:00
|
|
|
function_name="po-email-processor",
|
|
|
|
|
runtime=lambda_.Runtime.PYTHON_3_12,
|
2026-04-30 14:26:53 -04:00
|
|
|
architecture=lambda_.Architecture.ARM_64,
|
2026-04-07 12:12:30 -04:00
|
|
|
handler="handler.handler",
|
2026-05-08 16:01:21 -04:00
|
|
|
code=lambda_.Code.from_asset(
|
feat: widen email-processor asset roots to lambdas/ with scoped globs + excludes (refactor phase 2) (#109)
Both email-processor Code.from_asset calls now bundle from lambdas/
instead of their per-function subdirectory, so Phase 3's shared/
module is reachable from the asset root once it lands. The bundling
commands were rewritten for the new cwd (pip install -r <po|wo>/
email_processor/requirements.txt -t /asset-output && cp <po|wo>/
email_processor/*.py /asset-output/), preserving the ARM64
--platform manylinux2014_aarch64 --only-binary=:all: pin exactly —
its removal shipped x86 wheels into the ARM64 function and caused a
100% outage (PR #34).
All five from_asset calls (both email processors, po web_ui, po
site_extractor, wo web_ui) now exclude **/__pycache__/**; the two
widened ones also exclude **/tests/** and **/package/**. Without the
package/ exclude, the stale untracked 44 MB
lambdas/po/email_processor/package/ dir (local-only, never present
in CI) would diverge local vs CI asset hashes and force spurious
redeploys — from_asset doesn't honor .gitignore. That dir is left in
place; deleting it is Adam's call.
WO's prod zip shrinks as deliberate cleanup, not a byte-identical
match to PO: the old `cp -r .` shipped tests/ (real scrubbed .eml
fixtures), __pycache__/, and requirements.txt into production. The
acceptance bar for WO is runtime-imported module set unchanged +
smoke, not a byte-identical zip; PO keeps the byte-identical
first-party file set guarantee. tests/test_bundle_consistency.py is
updated in the same change to recognize the scoped
`cp po/email_processor/*.py` (resp. wo) glob as the new
unconditionally-safe shape, without loosening the allowlist-revert
detection, the detection-logic mutation test, or the
PO_EXPECTED_TOP_LEVEL_MODULES exact-set pin.
No code moved under lambdas/ in this change (git diff main...HEAD --
lambdas/ is empty); only CDK asset wiring and its tests changed.
2026-07-17 15:47:01 -04:00
|
|
|
"../lambdas",
|
|
|
|
|
exclude=["**/__pycache__/**", "**/tests/**", "**/package/**"],
|
2026-05-08 16:01:21 -04:00
|
|
|
bundling=cdk.BundlingOptions(
|
|
|
|
|
image=lambda_.Runtime.PYTHON_3_12.bundling_image,
|
|
|
|
|
command=[
|
|
|
|
|
"bash",
|
|
|
|
|
"-c",
|
feat: deploy-pipeline guards — healthcheck, smoke gate, bundle glob + AST test (refactor phase 0) (#107)
* feat: deploy-pipeline guards — healthcheck, smoke gate, bundle glob + AST test (refactor phase 0)
Deploys of po-email-processor and workorder-email-processor had no
verification step, so an init-time ImportError in the bundled zip
could ship silently and only surface on the next real S3 event. This
adds a synchronous post-deploy smoke gate wired into the deploy
workflow: both Lambdas are invoked with {"healthcheck": true} and the
FunctionError field is checked, since an Unhandled init error still
returns HTTP 200 on RequestResponse invokes and would false-pass a
plain exit-code check.
The healthcheck branch is the first statement in each handler, before
any boto3/S3 use or ses_auth, and only fires on a top-level direct
invoke ("healthcheck" is not a key AWS ever sets on a real S3
ObjectCreated event, so mail content can't reach this path). It emits
no EMF metrics and no log text that could match the
sender-auth-rejected metric filter, so two deploys in one window
won't trip the alarm.
Separately, the PO stack's asset bundling copied a hand-maintained
four-file allowlist into the zip, so every new sibling module
handler.py imports had to be added by hand or the deploy shipped a
Lambda that ImportErrors at cold start (bit us for template_parser in
PR #105 and nearly for derived_fields in PR #2). Replaced it with a
non-recursive ./*.py glob so top-level source files ship
automatically while tests/ and the stale package/ dir still cannot,
and added an AST-based bundle-consistency test that parses each
handler's first-party imports and fails CI if the bundling command
would omit any of them (a revert to an incomplete allowlist, or code
moved into a subdirectory the glob doesn't cover).
Includes the refactor-evaluation report that scoped this phase.
* fix: review nits — unambiguous bundling-command extraction, smoke payload-parse message, dead asserts
- tests/test_bundle_consistency.py: _extract_bundling_command now collects
all command=[...] matches and demands exactly one per stack file, instead
of silently returning whichever ast.walk visits first if a second bundled
function is ever added.
- scripts/post-deploy-smoke.sh: distinguish an unparseable response payload
from a payload mismatch so the failure message says what actually happened
(the previous "could not parse" branch was unreachable — the inline python
always exited 0).
- test_po_healthcheck.py: drop the substring assertions on stdout that were
dead behind the stricter `captured.out == ""` assertion; keep the stderr
filter-pattern check.
Review follow-up on PR #107; no behavior change to any shipped code path.
2026-07-17 13:18:45 -04:00
|
|
|
# NOTE: non-recursive glob (not `cp -r`) so tests/ and
|
|
|
|
|
# the stale package/ dir are never shipped -- only
|
|
|
|
|
# top-level .py siblings of handler.py. This replaces
|
|
|
|
|
# a hand-maintained four-file allowlist that twice
|
|
|
|
|
# nearly shipped a broken Lambda (missing
|
|
|
|
|
# template_parser in PR #105, nearly missing
|
|
|
|
|
# derived_fields in PR #2) because a new sibling
|
|
|
|
|
# import wasn't added to the list. The glob makes
|
|
|
|
|
# that class of bug structurally impossible;
|
|
|
|
|
# tests/test_bundle_consistency.py ast-parses
|
|
|
|
|
# handler.py's first-party imports and asserts this
|
|
|
|
|
# command ships all of them, so a future revert back
|
|
|
|
|
# to an allowlist that omits a sibling fails CI.
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
# pip step removed in Phase 7: requirements.txt is now
|
|
|
|
|
# empty (boto3 comes from the Lambda runtime), so nothing
|
|
|
|
|
# is installed and the manylinux pin has nothing to pin.
|
feat: extract lambdas/shared/ — single-source ses_auth, web_ui auth, email parsing, EMF emitter (refactor phase 3) (#111)
Four modules move into the handbook-mandated lambdas/shared/ location,
collapsing duplicated logic that had to be kept in sync by hand across
the PO and WO pipelines:
- ses_auth.py: the PO and WO copies were verified sha256-identical
against the feature/phase-7-ops-recovery baseline before the move
(no drift since the last audit). shared/ses_auth.py is the exact
bytes of that one copy; both originals are git rm'd (the PO copy
via rename, the WO copy as a straight delete). Bundling lands the
module flat in /asset-output for both email processors, so the
handlers keep `from ses_auth import authenticate_inbound_email`
unchanged — zero handler diff for this move, which is what keeps
fail-closed auth byte-identical through the change.
- web_ui_auth.py: extracts the byte-identical _get_auth_token /
_header / is_authenticated block plus the four token-cache globals
out of both web_ui handlers. The per-stack INFRA-74 comments stay
in each handler as-is (deliberately drifted wording, stack-specific)
rather than being unified into the shared module. Fail-closed
semantics (unset ARN or Secrets Manager exception -> deny) are
unchanged.
- email_parsing.py: parse_raw_email ships as the superset version that
returns cc unconditionally. WO's output is bit-identical to before;
PO simply ignores the cc field rather than being "cleaned up" to
consume it. No second variant is kept.
- emf.py: a generic emitter parameterized by namespace, dimension
sets, and properties. Every call site's emitted EMF envelope is
unchanged, including the load-bearing
[["ParseMethod"],["ParseMethod","TemplateId"]] dimension-set shape
the alarms and metric filters depend on. Emission ordering is
untouched: PO still emits ai_fallback before the Bedrock call, WO
still emits its mutually-exclusive ai_fallback/ai_fallback_rejected
after its gate. The deliberate-double-count comments survive.
_emit_derived_agreement_metric was found living inside
derived_fields.py, so per the DERIVED-FIELDS exception it is left
as a third, unconverted copy (derived_fields.py and the shadow
DerivedFieldAgreement telemetry stay untouchable while that bake
runs) — a comment there points at shared/emf.py for the eventual
follow-up.
Bundling: both email-processor cdk bundling commands gain a trailing
`cp shared/*.py /asset-output/` (they were already cp-only post-Phase
7, so no pip step or manylinux pin is reintroduced). Both web_ui
functions gain the same widened-root staging so web_ui_auth.py ships
beside their handler; site_extractor's from_asset is untouched.
Tests: PO_EXPECTED_TOP_LEVEL_MODULES gains the shared modules that now
ship, the AST sibling-import check resolves imports whose source now
lives under shared/, and the new shared cp line has its own
revert/mutation detection. _SIBLING_MODULES resolution and
_po_parser_support.py now load ses_auth/email_parsing/emf from
shared/; the two-copy ses_auth byte-identity fixture-hygiene test is
retired as obsolete now that there is one copy, and the ses_auth
fixture parameterization over two identical copies is dropped. The
sys.modules save/restore dance for template_parser (still duplicated
per-pipeline) is left in place.
2026-07-20 13:38:23 -04:00
|
|
|
# shared/*.py ships the four modules extracted to
|
|
|
|
|
# lambdas/shared/ (Phase 3): ses_auth, web_ui_auth,
|
|
|
|
|
# email_parsing, emf. Flat cp keeps the bare-name
|
|
|
|
|
# imports (e.g. `from ses_auth import ...`) resolving
|
|
|
|
|
# unchanged in /asset-output.
|
|
|
|
|
"cp po/email_processor/*.py /asset-output/ && "
|
|
|
|
|
"cp shared/*.py /asset-output/",
|
2026-05-08 16:01:21 -04:00
|
|
|
],
|
|
|
|
|
),
|
|
|
|
|
),
|
2026-04-07 12:12:30 -04:00
|
|
|
timeout=Duration.seconds(60),
|
|
|
|
|
memory_size=256,
|
2026-04-30 14:26:53 -04:00
|
|
|
log_retention=logs.RetentionDays.TWO_MONTHS,
|
Reconcile IaC with out-of-band DLQ + Function URL changes (INFRA-74, INFRA-41) (#50)
Make CDK the source of truth for two sets of changes applied out-of-band
via CLI to the po-ingest and WorkorderIngestStack stacks.
INFRA-74 (audit C-5): remove the public FunctionUrlAuthType.NONE Function
URL construct (and its auto-generated Principal:* invoke permission +
output) from both po-web-ui and workorder-web-ui. The URLs were already
deleted live via CLI; CFN's delete is idempotent.
INFRA-41 (audit H-8): add a CDK-managed SQS dead-letter queue
(dead_letter_queue=, 14d retention, SSL-enforced, CDK-generated name) and
an ALARM-only Errors alarm (Sum, threshold>0, site-alerts topic) for both
po-email-processor and workorder-email-processor, mirroring the
apm-wo-analysis-classifier DLQ and payments-payroll-batch alarm patterns.
Interim CLI resources (per-fn -dlq queues, -errors alarms, dlq-send inline
policies, OnFailure event-invoke-configs) removed post-deploy.
2026-06-08 16:02:29 -04:00
|
|
|
dead_letter_queue=email_processor_dlq,
|
2026-04-07 12:12:30 -04:00
|
|
|
environment={
|
|
|
|
|
"PO_TABLE": "purchase-orders",
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
"BEDROCK_MODEL_ID": "us.anthropic.claude-haiku-4-5-20251001-v1:0",
|
Add fail-closed SES sender authentication (INFRA-107) (#98)
* Add fail-closed SES sender authentication
The From header and any raw-MIME Authentication-Results copies are
attacker-forgeable, so a forged email to apm@int.seahaven.com or
amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107,
CRITICAL). Both S3-triggered email processors now authenticate the
sender against the Authentication-Results header SES itself prepends
at delivery: only the topmost header is consulted, its authserv-id
must be amazonses.com, and it must carry dkim=pass for a domain in
the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set
in CDK so ops can adjust without code changes).
Allowlists come from live traffic observed 2026-07-15 on both ingest
buckets: WO mail arrives via the apm@ Google Groups forward, which
re-signs as seahaven.com (the hxgnsmartcloud.com signature does not
survive the forward); PO mail passes for amazon.coupahost.com.
amazonses.com also passes on PO mail but is deliberately excluded --
every SES customer's outbound mail passes for it.
Every failure path (env var unset, header missing or unparseable,
verdict fail, unaligned domain) rejects the email: a structured
warning with the reason and S3 key is logged and the record skipped
without erroring the invocation, so rejected mail causes no Lambda
retries or DLQ messages. Handler signatures and event sources are
unchanged.
Refs: INFRA-107
* Harden AR parser per cross-family review
Cross-family (GPT-4.1) review findings: terminate the dkim result
token at end-of-clause, whitespace, or a comment so a value like
"dkim=pass-fake" can never be read as a pass; normalize trailing
dots off allowlist entries so "seahaven.com." matches; make the
compat32 parser policy explicit. Adds tests for result-token
boundaries, comments after the result, quoted domain values, and
folding inside a dkim clause.
Refs: INFRA-107
* Harden AR parsing and alarm on sender-auth rejects
The SES-stamped Authentication-Results value echoes attacker-controlled
SMTP-session tokens (envelope-from, helo, header.from) as their own
semicolon-delimited property clauses. A naive split(";") tore an RFC 5321
quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass
clause, so a fully spoofed email was accepted on the genuinely
SES-stamped topmost header. Tokenise comment- and quoted-string-aware
(RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on
semicolons outside a quoted-string, and fail closed on unbalanced
quotes/comments so a ';' inside a quoted pvalue can never start a clause.
Rejected mail returns normally (no error, no retry, no DLQ message), so a
signing-domain drift or a wrong allowlist would silently discard 100% of
legitimate mail while every alarm stayed green. Add a CloudWatch Logs
metric filter + alarm on the sender_auth_rejected warning to both stacks
so a false-reject storm pages instead of vanishing. This is also the
safety net for the WO seahaven.com allowlist assumption, which must be
validated against a live SES-stamped header (a plain Gmail auto-forward
re-signs under the sending Workspace domain, not seahaven.com).
Refs: INFRA-107
* chore: retrigger CI (no run recorded for 7c74ac1)
* Fix quoted-AUID DKIM domain spoof in sender auth
Resolve three confirmed /sh-security-review findings on the fail-closed
SES sender-authentication control.
HIGH: header.i/header.d domain extraction was not quoted-string aware.
An attacker with a valid DKIM key for their own domain could set an
RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive
extractor stopped at the closing quote and returned seahaven.com,
accepting forged mail. Extraction now tokenises the clause with the same
quoted-string discipline already used for clause splitting: header.d
(the plain signing domain) is authoritative when present, otherwise the
header.i domain is the part after the AUID's LAST top-level "@", so a "@"
inside a quoted local-part is treated as signer-controlled label text and
yields the true signer (attacker.com), not seahaven.com.
LOW: the topmost-header parse ran outside evaluate_sender_authentication's
try/except, so an unexpected parser exception on crafted input could
propagate into the handler and Lambda async retries/DLQ. The parse now
fails CLOSED with an authentication_results_unparseable reason.
MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to
a low-volume total-reject outage (a trickle that never sums to 3). Both
stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3
/ datapoints_to_alarm=2, so a sustained reject condition pages even at one
reject per period while a lone stray probe self-clears.
Refs: INFRA-107
* Load Lambda function dir on sys.path in tests
Rebasing INFRA-107 onto main folded #95's pytest suite into this
branch's tests. The unified conftest loads the PO/WO handlers by file
path, and handler.py now does `from ses_auth import
authenticate_inbound_email` -- a bare sibling import that resolves in
the Lambda only because the runtime puts each function's own directory
on sys.path. The shared load_handler now adds that directory so the
handler tests import correctly alongside the sender-auth tests.
Refs: INFRA-107
* Note #97 test files in README directory tree
The rebase onto main brought in #97's tests/requirements.txt and
tests/test_po_merge.py. List both in the directory tree so it matches
the tree on disk.
Refs: INFRA-107
* Document INFRA-107 forwarder-binding risk acceptance
Record the accepted risk that WO sender auth binds to the apm@ forward's
re-signing domain (seahaven.com) rather than the Hexagon originator; the
apm@ Google Group's restricted posting policy is the load-bearing control
(escalates to HIGH if the group is opened to external posting). Also
correct the sender-auth-rejected alarm docs to match the shipped config
(>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and
note the SES-AR-01/02 parser hardening follow-ups.
Refs: INFRA-107
2026-07-15 20:58:47 -04:00
|
|
|
# Fail-closed sender auth (INFRA-107): the handler only
|
|
|
|
|
# accepts mail whose SES-stamped Authentication-Results
|
|
|
|
|
# header carries dkim=pass for one of these domains.
|
|
|
|
|
# Observed on live traffic 2026-07-15: Coupa PO mail passes
|
|
|
|
|
# DKIM for amazon.coupahost.com (and amazonses.com, which is
|
|
|
|
|
# deliberately NOT allowlisted — every SES customer's mail
|
|
|
|
|
# passes that). Unset/empty ⇒ the handler rejects all mail.
|
|
|
|
|
"ALLOWED_DKIM_DOMAINS": "amazon.coupahost.com",
|
2026-04-07 12:12:30 -04:00
|
|
|
},
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
# Grant permissions
|
|
|
|
|
email_bucket.grant_read(email_processor)
|
|
|
|
|
po_table.grant_read_write_data(email_processor)
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
|
|
|
|
|
# --- Bedrock InvokeModel grant ---
|
|
|
|
|
# The us.* inference profile can route cross-region, so the grant MUST
|
|
|
|
|
# cover both the inference-profile ARN AND the per-region foundation-model
|
|
|
|
|
# ARNs (empty account field) for every region the profile can reach
|
|
|
|
|
# (us-east-1/us-east-2/us-west-2). A profile-only grant AccessDenies at
|
|
|
|
|
# runtime whenever the profile routes to a region whose foundation-model
|
|
|
|
|
# ARN is not allowed.
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
email_processor.add_to_role_policy(common.make_bedrock_invoke_statement(self))
|
|
|
|
|
|
|
|
|
|
# --- Standard per-Lambda alarms: po-email-processor ---
|
|
|
|
|
# errors (INFRA-41 / audit H-8), throttles, DLQ-messages (dropped PO
|
|
|
|
|
# emails), and a p99 duration alarm (orphan adoption of the CLI
|
|
|
|
|
# Lambda-Duration-po-email-processor under <fn>-duration naming, 45000 ms
|
|
|
|
|
# = 75% of the 60s timeout, eval 3 / dp 2). All ALARM-only to site-alerts.
|
|
|
|
|
common.add_standard_lambda_alarms(
|
Reconcile IaC with out-of-band DLQ + Function URL changes (INFRA-74, INFRA-41) (#50)
Make CDK the source of truth for two sets of changes applied out-of-band
via CLI to the po-ingest and WorkorderIngestStack stacks.
INFRA-74 (audit C-5): remove the public FunctionUrlAuthType.NONE Function
URL construct (and its auto-generated Principal:* invoke permission +
output) from both po-web-ui and workorder-web-ui. The URLs were already
deleted live via CLI; CFN's delete is idempotent.
INFRA-41 (audit H-8): add a CDK-managed SQS dead-letter queue
(dead_letter_queue=, 14d retention, SSL-enforced, CDK-generated name) and
an ALARM-only Errors alarm (Sum, threshold>0, site-alerts topic) for both
po-email-processor and workorder-email-processor, mirroring the
apm-wo-analysis-classifier DLQ and payments-payroll-batch alarm patterns.
Interim CLI resources (per-fn -dlq queues, -errors alarms, dlq-send inline
policies, OnFailure event-invoke-configs) removed post-deploy.
2026-06-08 16:02:29 -04:00
|
|
|
self,
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
"EmailProcessor",
|
|
|
|
|
email_processor,
|
|
|
|
|
"po-email-processor",
|
|
|
|
|
alarm_topic,
|
|
|
|
|
duration_statistic="p99",
|
|
|
|
|
errors=True,
|
|
|
|
|
dlq=email_processor_dlq,
|
|
|
|
|
descriptions={
|
|
|
|
|
"errors": "po-email-processor async invocation errors",
|
|
|
|
|
"throttles": "po-email-processor invocation throttles",
|
|
|
|
|
"dlq": "po-email-processor DLQ has messages (dropped PO emails)",
|
|
|
|
|
"duration": "po-email-processor p99 duration approaching the 60s timeout",
|
|
|
|
|
},
|
|
|
|
|
)
|
Reconcile IaC with out-of-band DLQ + Function URL changes (INFRA-74, INFRA-41) (#50)
Make CDK the source of truth for two sets of changes applied out-of-band
via CLI to the po-ingest and WorkorderIngestStack stacks.
INFRA-74 (audit C-5): remove the public FunctionUrlAuthType.NONE Function
URL construct (and its auto-generated Principal:* invoke permission +
output) from both po-web-ui and workorder-web-ui. The URLs were already
deleted live via CLI; CFN's delete is idempotent.
INFRA-41 (audit H-8): add a CDK-managed SQS dead-letter queue
(dead_letter_queue=, 14d retention, SSL-enforced, CDK-generated name) and
an ALARM-only Errors alarm (Sum, threshold>0, site-alerts topic) for both
po-email-processor and workorder-email-processor, mirroring the
apm-wo-analysis-classifier DLQ and payments-payroll-batch alarm patterns.
Interim CLI resources (per-fn -dlq queues, -errors alarms, dlq-send inline
policies, OnFailure event-invoke-configs) removed post-deploy.
2026-06-08 16:02:29 -04:00
|
|
|
|
Add fail-closed SES sender authentication (INFRA-107) (#98)
* Add fail-closed SES sender authentication
The From header and any raw-MIME Authentication-Results copies are
attacker-forgeable, so a forged email to apm@int.seahaven.com or
amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107,
CRITICAL). Both S3-triggered email processors now authenticate the
sender against the Authentication-Results header SES itself prepends
at delivery: only the topmost header is consulted, its authserv-id
must be amazonses.com, and it must carry dkim=pass for a domain in
the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set
in CDK so ops can adjust without code changes).
Allowlists come from live traffic observed 2026-07-15 on both ingest
buckets: WO mail arrives via the apm@ Google Groups forward, which
re-signs as seahaven.com (the hxgnsmartcloud.com signature does not
survive the forward); PO mail passes for amazon.coupahost.com.
amazonses.com also passes on PO mail but is deliberately excluded --
every SES customer's outbound mail passes for it.
Every failure path (env var unset, header missing or unparseable,
verdict fail, unaligned domain) rejects the email: a structured
warning with the reason and S3 key is logged and the record skipped
without erroring the invocation, so rejected mail causes no Lambda
retries or DLQ messages. Handler signatures and event sources are
unchanged.
Refs: INFRA-107
* Harden AR parser per cross-family review
Cross-family (GPT-4.1) review findings: terminate the dkim result
token at end-of-clause, whitespace, or a comment so a value like
"dkim=pass-fake" can never be read as a pass; normalize trailing
dots off allowlist entries so "seahaven.com." matches; make the
compat32 parser policy explicit. Adds tests for result-token
boundaries, comments after the result, quoted domain values, and
folding inside a dkim clause.
Refs: INFRA-107
* Harden AR parsing and alarm on sender-auth rejects
The SES-stamped Authentication-Results value echoes attacker-controlled
SMTP-session tokens (envelope-from, helo, header.from) as their own
semicolon-delimited property clauses. A naive split(";") tore an RFC 5321
quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass
clause, so a fully spoofed email was accepted on the genuinely
SES-stamped topmost header. Tokenise comment- and quoted-string-aware
(RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on
semicolons outside a quoted-string, and fail closed on unbalanced
quotes/comments so a ';' inside a quoted pvalue can never start a clause.
Rejected mail returns normally (no error, no retry, no DLQ message), so a
signing-domain drift or a wrong allowlist would silently discard 100% of
legitimate mail while every alarm stayed green. Add a CloudWatch Logs
metric filter + alarm on the sender_auth_rejected warning to both stacks
so a false-reject storm pages instead of vanishing. This is also the
safety net for the WO seahaven.com allowlist assumption, which must be
validated against a live SES-stamped header (a plain Gmail auto-forward
re-signs under the sending Workspace domain, not seahaven.com).
Refs: INFRA-107
* chore: retrigger CI (no run recorded for 7c74ac1)
* Fix quoted-AUID DKIM domain spoof in sender auth
Resolve three confirmed /sh-security-review findings on the fail-closed
SES sender-authentication control.
HIGH: header.i/header.d domain extraction was not quoted-string aware.
An attacker with a valid DKIM key for their own domain could set an
RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive
extractor stopped at the closing quote and returned seahaven.com,
accepting forged mail. Extraction now tokenises the clause with the same
quoted-string discipline already used for clause splitting: header.d
(the plain signing domain) is authoritative when present, otherwise the
header.i domain is the part after the AUID's LAST top-level "@", so a "@"
inside a quoted local-part is treated as signer-controlled label text and
yields the true signer (attacker.com), not seahaven.com.
LOW: the topmost-header parse ran outside evaluate_sender_authentication's
try/except, so an unexpected parser exception on crafted input could
propagate into the handler and Lambda async retries/DLQ. The parse now
fails CLOSED with an authentication_results_unparseable reason.
MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to
a low-volume total-reject outage (a trickle that never sums to 3). Both
stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3
/ datapoints_to_alarm=2, so a sustained reject condition pages even at one
reject per period while a lone stray probe self-clears.
Refs: INFRA-107
* Load Lambda function dir on sys.path in tests
Rebasing INFRA-107 onto main folded #95's pytest suite into this
branch's tests. The unified conftest loads the PO/WO handlers by file
path, and handler.py now does `from ses_auth import
authenticate_inbound_email` -- a bare sibling import that resolves in
the Lambda only because the runtime puts each function's own directory
on sys.path. The shared load_handler now adds that directory so the
handler tests import correctly alongside the sender-auth tests.
Refs: INFRA-107
* Note #97 test files in README directory tree
The rebase onto main brought in #97's tests/requirements.txt and
tests/test_po_merge.py. List both in the directory tree so it matches
the tree on disk.
Refs: INFRA-107
* Document INFRA-107 forwarder-binding risk acceptance
Record the accepted risk that WO sender auth binds to the apm@ forward's
re-signing domain (seahaven.com) rather than the Hexagon originator; the
apm@ Google Group's restricted posting policy is the load-bearing control
(escalates to HIGH if the group is opened to external posting). Also
correct the sender-auth-rejected alarm docs to match the shipped config
(>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and
note the SES-AR-01/02 parser hardening follow-ups.
Refs: INFRA-107
2026-07-15 20:58:47 -04:00
|
|
|
# --- Sender-auth rejection alarm (INFRA-107) ---
|
|
|
|
|
# A rejected email (bad/unaligned DKIM verdict) returns normally, so it
|
|
|
|
|
# produces NO Lambda error, NO DLQ message and NO retry -- only a
|
|
|
|
|
# `sender_auth_rejected` warning log. Without this metric filter + alarm a
|
|
|
|
|
# domain drift (Coupa rotates its signing subdomain, SES changes its
|
|
|
|
|
# Authentication-Results format, the allowlist is wrong) would silently
|
|
|
|
|
# discard 100% of legitimate PO mail while every other alarm stays green.
|
|
|
|
|
# A CloudWatch Logs metric filter turns those warnings into a metric so a
|
|
|
|
|
# false-reject storm pages instead of vanishing. default_value=0 keeps the
|
|
|
|
|
# series populated (alarm stays OK, never INSUFFICIENT_DATA) between events.
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
common.add_sender_auth_rejected_alarm(
|
Add fail-closed SES sender authentication (INFRA-107) (#98)
* Add fail-closed SES sender authentication
The From header and any raw-MIME Authentication-Results copies are
attacker-forgeable, so a forged email to apm@int.seahaven.com or
amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107,
CRITICAL). Both S3-triggered email processors now authenticate the
sender against the Authentication-Results header SES itself prepends
at delivery: only the topmost header is consulted, its authserv-id
must be amazonses.com, and it must carry dkim=pass for a domain in
the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set
in CDK so ops can adjust without code changes).
Allowlists come from live traffic observed 2026-07-15 on both ingest
buckets: WO mail arrives via the apm@ Google Groups forward, which
re-signs as seahaven.com (the hxgnsmartcloud.com signature does not
survive the forward); PO mail passes for amazon.coupahost.com.
amazonses.com also passes on PO mail but is deliberately excluded --
every SES customer's outbound mail passes for it.
Every failure path (env var unset, header missing or unparseable,
verdict fail, unaligned domain) rejects the email: a structured
warning with the reason and S3 key is logged and the record skipped
without erroring the invocation, so rejected mail causes no Lambda
retries or DLQ messages. Handler signatures and event sources are
unchanged.
Refs: INFRA-107
* Harden AR parser per cross-family review
Cross-family (GPT-4.1) review findings: terminate the dkim result
token at end-of-clause, whitespace, or a comment so a value like
"dkim=pass-fake" can never be read as a pass; normalize trailing
dots off allowlist entries so "seahaven.com." matches; make the
compat32 parser policy explicit. Adds tests for result-token
boundaries, comments after the result, quoted domain values, and
folding inside a dkim clause.
Refs: INFRA-107
* Harden AR parsing and alarm on sender-auth rejects
The SES-stamped Authentication-Results value echoes attacker-controlled
SMTP-session tokens (envelope-from, helo, header.from) as their own
semicolon-delimited property clauses. A naive split(";") tore an RFC 5321
quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass
clause, so a fully spoofed email was accepted on the genuinely
SES-stamped topmost header. Tokenise comment- and quoted-string-aware
(RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on
semicolons outside a quoted-string, and fail closed on unbalanced
quotes/comments so a ';' inside a quoted pvalue can never start a clause.
Rejected mail returns normally (no error, no retry, no DLQ message), so a
signing-domain drift or a wrong allowlist would silently discard 100% of
legitimate mail while every alarm stayed green. Add a CloudWatch Logs
metric filter + alarm on the sender_auth_rejected warning to both stacks
so a false-reject storm pages instead of vanishing. This is also the
safety net for the WO seahaven.com allowlist assumption, which must be
validated against a live SES-stamped header (a plain Gmail auto-forward
re-signs under the sending Workspace domain, not seahaven.com).
Refs: INFRA-107
* chore: retrigger CI (no run recorded for 7c74ac1)
* Fix quoted-AUID DKIM domain spoof in sender auth
Resolve three confirmed /sh-security-review findings on the fail-closed
SES sender-authentication control.
HIGH: header.i/header.d domain extraction was not quoted-string aware.
An attacker with a valid DKIM key for their own domain could set an
RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive
extractor stopped at the closing quote and returned seahaven.com,
accepting forged mail. Extraction now tokenises the clause with the same
quoted-string discipline already used for clause splitting: header.d
(the plain signing domain) is authoritative when present, otherwise the
header.i domain is the part after the AUID's LAST top-level "@", so a "@"
inside a quoted local-part is treated as signer-controlled label text and
yields the true signer (attacker.com), not seahaven.com.
LOW: the topmost-header parse ran outside evaluate_sender_authentication's
try/except, so an unexpected parser exception on crafted input could
propagate into the handler and Lambda async retries/DLQ. The parse now
fails CLOSED with an authentication_results_unparseable reason.
MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to
a low-volume total-reject outage (a trickle that never sums to 3). Both
stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3
/ datapoints_to_alarm=2, so a sustained reject condition pages even at one
reject per period while a lone stray probe self-clears.
Refs: INFRA-107
* Load Lambda function dir on sys.path in tests
Rebasing INFRA-107 onto main folded #95's pytest suite into this
branch's tests. The unified conftest loads the PO/WO handlers by file
path, and handler.py now does `from ses_auth import
authenticate_inbound_email` -- a bare sibling import that resolves in
the Lambda only because the runtime puts each function's own directory
on sys.path. The shared load_handler now adds that directory so the
handler tests import correctly alongside the sender-auth tests.
Refs: INFRA-107
* Note #97 test files in README directory tree
The rebase onto main brought in #97's tests/requirements.txt and
tests/test_po_merge.py. List both in the directory tree so it matches
the tree on disk.
Refs: INFRA-107
* Document INFRA-107 forwarder-binding risk acceptance
Record the accepted risk that WO sender auth binds to the apm@ forward's
re-signing domain (seahaven.com) rather than the Hexagon originator; the
apm@ Google Group's restricted posting policy is the load-bearing control
(escalates to HIGH if the group is opened to external posting). Also
correct the sender-auth-rejected alarm docs to match the shipped config
(>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and
note the SES-AR-01/02 parser hardening follow-ups.
Refs: INFRA-107
2026-07-15 20:58:47 -04:00
|
|
|
self, "EmailProcessor", "po-email-processor", alarm_topic
|
|
|
|
|
)
|
|
|
|
|
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
# --- Template fallback-rate alarm: po-email-processor ---
|
|
|
|
|
# The processor tries a deterministic template parse first and only calls
|
|
|
|
|
# the Bedrock AI extractor on a miss/invalid. A sustained rise in the
|
|
|
|
|
# ai_fallback share signals Coupa template drift (coverage collapse).
|
|
|
|
|
# EMF metric Seahaven/PoIngest/ParseOutcome, dimensioned by ParseMethod
|
|
|
|
|
# (template|ai_fallback).
|
|
|
|
|
#
|
|
|
|
|
# RETUNED for PO volume (~57 emails/day ≈ 14.25 per 6h period) -- the WO
|
|
|
|
|
# alarm's 15-min period / >=10-sample floor assume ~760/day and would be
|
|
|
|
|
# structurally DEAD here (a 15-min period holds ~0.6 PO emails, so the
|
|
|
|
|
# floor is never met and the IF always takes the 0 branch):
|
|
|
|
|
# * period 6h: a stable ~14-email denominator per datapoint.
|
|
|
|
|
# * volume floor >=8: at the floor, one fallback email = 12.5% < 20%,
|
|
|
|
|
# so a single email can NEVER breach a datapoint; a breach needs >=2
|
|
|
|
|
# fallbacks in one 6h window (2/8 = 25%) or >=3 at typical volume
|
|
|
|
|
# (3/14 ≈ 21%). Sparse overnight/weekend windows (<8 emails) take
|
|
|
|
|
# the 0 branch -- non-breaching by design (accepted trade: a Friday-
|
|
|
|
|
# evening drift may not page until weekend volume accrues).
|
|
|
|
|
# * threshold >20%: expected baseline fallback ≈1% (comments 0.55% +
|
|
|
|
|
# multi-line 0.18% + non-USD 0) -- far below the threshold.
|
|
|
|
|
# * 2 of 4 datapoints (24h span): isolated noise self-clears, while
|
|
|
|
|
# total template drift (100% fallback) pages within ~12h.
|
|
|
|
|
# Post-#102 rule: NO element-wise MAX(timeseries, scalar) in alarm math;
|
|
|
|
|
# the IF volume floor guarantees the non-zero denominator. Any change to
|
|
|
|
|
# this expression must be gated by `npx cdk synth po-ingest`.
|
PO ai-fallback fail-closed gate + prompt hardening (refactor phase 1) (#108)
* feat: PO ai-fallback fail-closed gate + prompt hardening, parity with #104 (refactor phase 1)
Ports WO's #104 AI-fallback security hardening to the PO email
processor, adapted for PO's nested contract instead of copying the
WO gate verbatim.
validate_ai_fallback() (template_parser.py) fail-closes raw Bedrock
output before it reaches enrich_parsed or any dispatch/save:
recursive key-set check with missing-key normalization (nested
contract: supplier{}, ship_to{}, line_items[]); po_number checked
against the same hardened prefix+hyphen+digits regex family that
guards the DynamoDB partition key the handler builds from it
(rejects fullwidth-digit and trailing-artifact injection); email_type
enforced against the {new_po, revision, cancellation} allow-list
before dispatch so a miss can never fall into the else -> save_new_po
branch; money fields accept Decimal/int/None only, matching PO's
parse_float=Decimal decode (a float-typed check would be wrong here).
A gate failure emits ParseMethod=ai_fallback_rejected and `continue`s
to the next record -- it never raises, so attacker-controlled input
can't churn the retry/DLQ path.
extract_with_claude() wraps the untrusted email in an <email> data
block and neutralizes forged <email>-tag lookalikes in the body with
the same linear-time regex approach as WO's _EMAIL_TAG_RE, and sets
temperature=0 on the Bedrock call.
Deliberate double-count: PO emits ParseMethod=ai_fallback before the
Bedrock call (so a Bedrock-side error still records the outcome), so
a rejected email always produces both an ai_fallback datapoint
(pre-call) and an ai_fallback_rejected datapoint (post-gate). This is
intentional, not a bug -- documented in handler.py, template_parser.py,
and the README.
cdk/po_stack.py: in-place property update to the existing
po-email-processor-template-fallback-rate alarm (same logical ID, no
rename/replacement) -- the fb/(fb+tmpl) expression is left
byte-identical to its pre-Phase-1 form and ai_fallback_rejected is
deliberately excluded from the numerator/denominator/volume floor,
since folding it in as WO does would double-count every rejection
(PO's pre-call emit already counts it once via fb). A net-new
EmailProcessorAiFallbackRejectedAlarm watches the rejected series on
its own, retuned for ~57 emails/day with the 6h/IF-floor/eval-4/
datapoints-2 idiom (not WO's 5-minute sparse idiom, which is
structurally dead at PO volume). Both alarms remain ALARM-only to
site-alerts, NOT_BREACHING, with no element-wise MAX in the math
(post-#102 rule).
* Block "Cancelled" po_status off the AI cancellation route
The AI-fallback gate type-checked po_status but let any string
through, unlike the template path which never emits "Cancelled" on a
new_po. Dispatch routes on email_type, so an AI-path new_po or revision
carrying po_status="Cancelled" would reach save_new_po/save_revision and
cancel a live PO via _merge_update's sticky-cancel write without ever
hitting save_cancellation. Reject the exact sticky marker on any
non-cancellation email_type so the AI path matches the template path's
guard; arbitrary non-marker status strings still pass.
email_type is already validated to the enum before this check, and a
cancellation reaches save_cancellation (which hardcodes the status), so
po_status stays irrelevant on that route.
2026-07-17 14:50:47 -04:00
|
|
|
#
|
|
|
|
|
# DOUBLE-COUNT ACCOUNTING (Phase 1 / PO AI-fallback gate): PO emits
|
|
|
|
|
# ParseMethod=ai_fallback BEFORE the Bedrock call for EVERY AI-path
|
|
|
|
|
# email (handler pre-call emit; try_deterministic_parse returns
|
|
|
|
|
# "ai_fallback" on every template miss), so a gate-rejected email
|
|
|
|
|
# already appears exactly once in `fb`. Therefore fb = ALL fallback
|
|
|
|
|
# attempts (accepted + rejected), fb + tmpl = ALL emails, and
|
|
|
|
|
# rate = fb/(fb+tmpl) is exact -- the expression below is deliberately
|
|
|
|
|
# left BYTE-IDENTICAL to the pre-Phase-1 form, and `rej`
|
|
|
|
|
# (ai_fallback_rejected) is deliberately EXCLUDED from this
|
|
|
|
|
# expression's numerator, denominator, and volume floor, and is never
|
|
|
|
|
# added to using_metrics. This is NOT an oversight: folding `rej` in
|
|
|
|
|
# here as WO does (fb+rej numerator / fb+rej+tmpl denominator) would
|
|
|
|
|
# double-count every rejected email in both numerator and
|
|
|
|
|
# denominator (PO's pre-call emit already counts it once via `fb`),
|
|
|
|
|
# inflating the observed rate toward 100% and double-counting toward
|
|
|
|
|
# the >=8 volume floor -- a prompt-injection probing burst would then
|
|
|
|
|
# falsely page this template-drift alarm on top of the dedicated
|
|
|
|
|
# rejected alarm below. The rejected series gets its own alarm
|
|
|
|
|
# instead (EmailProcessorAiFallbackRejectedAlarm, below).
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
common.make_fallback_rate_alarm(
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
self,
|
|
|
|
|
"EmailProcessorTemplateFallbackRateAlarm",
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
namespace="Seahaven/PoIngest",
|
|
|
|
|
alarm_topic=alarm_topic,
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
alarm_name="po-email-processor-template-fallback-rate",
|
|
|
|
|
alarm_description=(
|
|
|
|
|
"po-email-processor deterministic-template coverage collapse: "
|
|
|
|
|
">20% of parses fell back to the Bedrock AI extractor"
|
|
|
|
|
),
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
rejected_included=False,
|
|
|
|
|
period=Duration.hours(6),
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
threshold=20,
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
floor=8,
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
evaluation_periods=4,
|
|
|
|
|
datapoints_to_alarm=2,
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
)
|
feat: template-first PO parser with fail-closed gate and Bedrock fallback (#105)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* feature: Add PO template parser scaffold and design doc
Mirror WO PR #99's template-first approach for the Coupa PO processor. Two templates identified from a full 3,448-email triage:
- coupa_new_po (95.5%): scaffolded; fails closed to the LLM until extract_new_po lands.
- coupa_cancellation (2.9%): implemented.
Nested contract with recursive validation, Decimal money, and a fail-closed gate. Derived fields (site_code/trade/fiscal_year) are deferred to a shared post-stage. Comments, revisions, multi-line, and non-USD emails fall back to Bedrock. docs/po-template-parser.md records the investigation, decisions, and remaining work.
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
* Implement PO new_po extraction and value-level gate
Replace the extract_new_po scaffold stub with the full
section-windowed extractor (duplicate-label anchoring, sentinel
ship-to, label-keyed U+2022 bullet split, Decimal money from three
anchored contexts only) and add value-level gate rules V1-V13.
Both new_po_not_implemented scaffold guards are removed; rules 6-8
(unrecognized_status, multiline_unsupported, non_usd) go live.
The gate re-derives every byte proof from the email body so an
extractor bug cannot vouch for itself: amount re-serialization
with a digit/comma border check (the thousands-separator
truncation kill switch), sum(lines)==total against both Total
blocks, anchor/supplier identity proofs, USPS address shape on
the raw pre-enrichment zip, bullet label discipline, and
sentinel/artifact hygiene. Any failure falls closed to the LLM;
a validation failure is never a parsed result.
Refs: #99
* Wire template-first parse into PO handler with EMF metric
Run try_deterministic_parse ahead of the Bedrock extractor and
fall back only on a miss/invalid (fail-closed) result. The shared
enrich_parsed post-stage and the save_cancellation/save_revision/
save_new_po routing are untouched, so both paths write identical
DynamoDB shapes and the po-ingest-site-extractor stream contract
is preserved.
Each record emits one ParseMethod EMF line (Seahaven/PoIngest/
ParseOutcome, dimension sets [ParseMethod] and
[ParseMethod,TemplateId], ReasonCode/po_number ride-alongs)
mirroring the WO idiom. The metric fires before the Bedrock call
so a Bedrock-side error still records the ai_fallback outcome.
Refs: #99
* Add PO fallback-rate alarm retuned for ~57 emails/day
The WO alarm's 15-min period and >=10-sample floor assume
~760/day and would be structurally dead at PO volume (a 15-min
period holds ~0.6 emails, so the floor is never met). Retune:
6-hour periods (~14.25 expected emails), IF((fb+tmpl)>=8,...)
volume floor so a single email can never breach a datapoint
(1/8 = 12.5% < 20%), threshold >20% against a ~1% expected
baseline, eval 4 / datapoints 2 (24h span) so noise self-clears
while total template drift pages within ~12h. No element-wise
MAX in the math expression (post-#102 rule); ALARM-only
SnsAction to site-alerts, NOT_BREACHING. Gated with
'npx cdk synth po-ingest'.
Also add template_parser.py to the bundling cp list -- without it
every deployed invocation would ImportError (unit tests cannot
catch an asset-bundling omission).
Refs: #99, #102
* Add offline PO parser suite with scrubbed fixture corpus
132 tests: golden-file comparison for all 25 positive fixtures
(17 single-line new-PO + 8 cancellations, Decimal-exact via
parse_float=Decimal), every fail-closed gate reason code covered
(body-level triggers via 17 synthetic adversarial .eml mutations,
candidate-level via direct validate() unit tests), real multi-line
and comment/non-Coupa fallback fixtures, dual line-ending parse
identity, two-path enrich/save parity (site-extractor stream
guard), V10 URL-id corpus sweep, fixture hygiene (ses_auth pass +
scrub-marker leak sweep), and Bedrock dispatch/EMF assertions.
The suite loads handler/template_parser via importlib under
unique module names and binds the handler's bare sibling imports
around exec (tests/conftest.py load_handler gets the same
treatment) -- the WO suite caches bare 'handler'/'template_parser'
names in sys.modules, and bare imports here would silently bind
to the wrong pipeline. moto is imported before the handler so its
botocore stubber hook precedes boto3 session creation (the PO
conftest chain now loads at pytest session start).
Fixtures are scrubbed real S3 samples: transport/auth header
values replaced with same-shape placeholders (structure kept so
ses_auth still passes), per-file digit ciphers, amounts remapped
with sum==total re-established. The .gitignore exception is
scoped to the PO fixtures path only.
Refs: #99
* Document PO template-first parser and retuned alarm
README: PO flow is now template-first with Bedrock fallback;
parser/gate section mirroring the WO writeup; Seahaven/PoIngest
ParseOutcome namespace and the fallback-rate alarm numbers with
their volume justification (deliberately not WO's settings);
test-suite and repo-layout updates.
Design doc: mark PR #1 complete in progress/checklist sections;
document the six value-level gate reason codes and the scaffold
guard removal; correct the stale data-access note (default CLI
session is 328440206208) and note the ~90-day S3 lifecycle aging
of the corpus; record the 2.3 layout addendum (leading Supplier
bullet segment, EA evidence lines, summary unit-price tokens,
decode-path line endings), the fixture-build pins (address join
convention, quantity/unit/price source), the V10 sweep outcome,
and resolutions for open questions Q3/Q6. Cross-family review and
the Confluence architecture-map update are flagged outstanding
for merge.
Refs: #99
* Record cross-family review outcome for handler wiring
GPT-4.1 cross_review.py run against the real handler diff
returned no BLOCK and no security findings; both FIX items
verified as no-change-needed (fallback logging already correct;
non-dict AI output is the pre-existing issue #101 pattern this
PR deliberately does not touch).
Refs: #99
* Pin line-item currency to USD in the PO gate
The non_usd rule only checked the Total-block top-level currency, so a
new_po whose line item read 'for 55,206.00 CAD' under a USD Total block
still template-parsed as ok -- a fail-open hole in the fail-closed
gate. Every line item's captured currency and its re-derived body token
must now byte-equal the proven-USD top-level currency; covered by a
line-level CAD adversarial fixture (the existing adv-non-usd only
exercised the Total-block variant) and a candidate-mutation unit test.
* Scrub residual transport tokens from PO fixtures
The first-pass harvest scrub sanitized only the primary SES/DKIM
header blocks, leaving the real SES Feedback-ID sender-identity hash
in 49 committed fixtures and, on the two non-Coupa fixtures, an
embedded second SES block's X-Ses-Receipt, the Exchange cross-tenant
UPN ciphertext, and Gmail ARC fh= / X-Gm-* tokens -- exactly the
token classes the PR #99 fixture lesson requires placeholdered.
Replace each with a same-shape ScrubbedFixture value (byte-safe,
CRLF and folding preserved) so header structure and ses_auth
behavior are unchanged.
* Converge quantity/price to Decimal on both paths
EXTRACTION_PROMPT declares quantity and price as JSON strings, so a
prompt-obedient Bedrock response stores DynamoDB Strings where the
template parser stores Numbers -- divergent attribute types for the
same email on the purchase-orders stream. Coerce numeric strings to
Decimal in the shared enrich_parsed post-stage (thousands-separator
safe; non-numeric strings kept verbatim) so both paths converge;
prompt rewording itself remains PR #2 scope.
The two-path parity test was circular -- it replayed the parser-
derived golden as 'the LLM output', so it could never see the type
divergence. It now feeds a prompt-shaped payload (string quantity/
price, LLM-filled site_code) through enrich_parsed and save_new_po,
and the fixture-hygiene test now asserts the scrubbed transport-token
header classes so fixture regressions are caught.
* Coerce bare-int quantity/price to Decimal in enrich_parsed
GPT-4.1 cross-family review of the final PR diff (no BLOCK) flagged
residual type drift: parse_float=Decimal rules out floats on the LLM
path, but a bare JSON int survived as Python int. Coerce it so both
parse paths emit one canonical Decimal type.
* Scrub fixture-body PII and harden cancellation gate (sec review)
/sh-security-review of PR #105 (5 fresh-context detectors + proof-or-kill
verifier) confirmed two diff-introduced findings; both fixed here.
F3 (medium, real PII in new fixtures): the harvest scrub replaced header
tokens but left real third-party PII in message BODIES -- an Amazon
contact's name/phone/personal email in non-coupa-02.eml and an internal
t.corp.amazon.com ticket URL in comment-02.eml, plus real submitter/attn
names recurring across the new_po corpus. Replaced every personal name,
phone, personal email, and internal URL with synthetic placeholders
(QP-soft-wrap aware) across both .eml bodies and expected goldens.
Extended test_fixture_hygiene to scan BODIES (phone shapes, corp URLs,
the leaked tokens), closing the header-only gap that let this through.
F1 (medium, cancellation gate): _CANCELLATION_SUBJECT was unanchored and
matched with .search(), unlike the anchored new_po pattern -- a subject
merely ending with the cancellation phrase could be routed to the sticky-
Cancelled write. Fully anchored it and switched to .match, and added a
body-corroboration gate (the real Coupa body independently restates
'Purchase Order #<po> ... has been cancelled'); a near-miss/misrouted
subject whose body does not corroborate now fails closed to the LLM
(new reason code cancellation_body_unconfirmed).
Pre-existing (advisory, not this PR): the LLM-fallback else->save_new_po
dispatch and undelimited extraction prompt (issue #101 family) are
byte-identical to main and unchanged here.
401 tests pass; ruff/format clean; cdk synth po-ingest clean.
---------
Signed-off-by: Adam Moussa <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 17:50:59 -04:00
|
|
|
|
PO ai-fallback fail-closed gate + prompt hardening (refactor phase 1) (#108)
* feat: PO ai-fallback fail-closed gate + prompt hardening, parity with #104 (refactor phase 1)
Ports WO's #104 AI-fallback security hardening to the PO email
processor, adapted for PO's nested contract instead of copying the
WO gate verbatim.
validate_ai_fallback() (template_parser.py) fail-closes raw Bedrock
output before it reaches enrich_parsed or any dispatch/save:
recursive key-set check with missing-key normalization (nested
contract: supplier{}, ship_to{}, line_items[]); po_number checked
against the same hardened prefix+hyphen+digits regex family that
guards the DynamoDB partition key the handler builds from it
(rejects fullwidth-digit and trailing-artifact injection); email_type
enforced against the {new_po, revision, cancellation} allow-list
before dispatch so a miss can never fall into the else -> save_new_po
branch; money fields accept Decimal/int/None only, matching PO's
parse_float=Decimal decode (a float-typed check would be wrong here).
A gate failure emits ParseMethod=ai_fallback_rejected and `continue`s
to the next record -- it never raises, so attacker-controlled input
can't churn the retry/DLQ path.
extract_with_claude() wraps the untrusted email in an <email> data
block and neutralizes forged <email>-tag lookalikes in the body with
the same linear-time regex approach as WO's _EMAIL_TAG_RE, and sets
temperature=0 on the Bedrock call.
Deliberate double-count: PO emits ParseMethod=ai_fallback before the
Bedrock call (so a Bedrock-side error still records the outcome), so
a rejected email always produces both an ai_fallback datapoint
(pre-call) and an ai_fallback_rejected datapoint (post-gate). This is
intentional, not a bug -- documented in handler.py, template_parser.py,
and the README.
cdk/po_stack.py: in-place property update to the existing
po-email-processor-template-fallback-rate alarm (same logical ID, no
rename/replacement) -- the fb/(fb+tmpl) expression is left
byte-identical to its pre-Phase-1 form and ai_fallback_rejected is
deliberately excluded from the numerator/denominator/volume floor,
since folding it in as WO does would double-count every rejection
(PO's pre-call emit already counts it once via fb). A net-new
EmailProcessorAiFallbackRejectedAlarm watches the rejected series on
its own, retuned for ~57 emails/day with the 6h/IF-floor/eval-4/
datapoints-2 idiom (not WO's 5-minute sparse idiom, which is
structurally dead at PO volume). Both alarms remain ALARM-only to
site-alerts, NOT_BREACHING, with no element-wise MAX in the math
(post-#102 rule).
* Block "Cancelled" po_status off the AI cancellation route
The AI-fallback gate type-checked po_status but let any string
through, unlike the template path which never emits "Cancelled" on a
new_po. Dispatch routes on email_type, so an AI-path new_po or revision
carrying po_status="Cancelled" would reach save_new_po/save_revision and
cancel a live PO via _merge_update's sticky-cancel write without ever
hitting save_cancellation. Reject the exact sticky marker on any
non-cancellation email_type so the AI path matches the template path's
guard; arbitrary non-marker status strings still pass.
email_type is already validated to the enum before this check, and a
cancellation reaches save_cancellation (which hardcodes the status), so
po_status stays irrelevant on that route.
2026-07-17 14:50:47 -04:00
|
|
|
# --- AI-fallback rejected alarm: po-email-processor (Phase 1) ---
|
|
|
|
|
# The validate_ai_fallback gate (template_parser.py) fail-closes Bedrock
|
|
|
|
|
# output that doesn't match PO's contract (structurally wrong shape,
|
|
|
|
|
# injected po_number/email_type, wrong field types) and emits
|
|
|
|
|
# ParseMethod=ai_fallback_rejected instead of writing it. That is a
|
|
|
|
|
# SILENT skip (`continue`, never raise) by design -- attacker-controlled
|
|
|
|
|
# input must not churn the retry/DLQ path -- so without a dedicated
|
|
|
|
|
# alarm a sustained rejection run (prompt-injection probing, or a
|
|
|
|
|
# template-drift outage whose AI output also happens to fail the gate)
|
|
|
|
|
# is invisible everywhere except this metric and the ReasonCode log
|
|
|
|
|
# line.
|
|
|
|
|
#
|
|
|
|
|
# RETUNED for PO volume (~57 emails/day, baseline ai_fallback rate
|
|
|
|
|
# ~1% => ~0.6 AI-fallback emails/day, expected rejections ~= 0) -- NOT
|
|
|
|
|
# WO's 5-min/2-of-6 sparse idiom (wo_stack.py), which needs two
|
|
|
|
|
# rejections inside a single 30-min window and is structurally dead at
|
|
|
|
|
# this volume. Mirrors the PO fallback-rate alarm's 6h/eval-4/dp-2
|
|
|
|
|
# retune idiom above, but with a COUNT floor on the rejected series
|
|
|
|
|
# itself rather than an email-volume floor: an email-volume floor
|
|
|
|
|
# (fb+tmpl>=N) would suppress paging in exactly the sparse
|
|
|
|
|
# overnight/weekend windows where a silently-dropped email matters
|
|
|
|
|
# most, and there is no denominator here, so there is nothing else to
|
|
|
|
|
# guard against divide-by-zero. FILL(rej,0) turns the sparse EMF
|
|
|
|
|
# series (no datapoint in quiet periods -- no metric-filter
|
|
|
|
|
# default_value exists for EMF) into a dense 0-series so every
|
|
|
|
|
# evaluation window has data. Post-#102 rule still holds: NO
|
|
|
|
|
# element-wise MAX(timeseries, scalar) anywhere in this expression.
|
|
|
|
|
#
|
|
|
|
|
# Tuning: rejections self-clear unless >=2 breaching datapoints land in
|
|
|
|
|
# >=2 distinct 6h windows within 24h (sustained probing, or template
|
|
|
|
|
# drift whose AI output also fails the gate), which pages within
|
|
|
|
|
# ~12-24h. Accepted residual (matches WO's accepted residual): because
|
|
|
|
|
# the breach is measured per 6h window, ANY burst of rejections
|
|
|
|
|
# confined to a single 6h window -- whether one stray email or dozens
|
|
|
|
|
# in a 20-minute spike -- is one breaching datapoint and never pages
|
|
|
|
|
# this alarm by itself. This is deliberate anti-flap tuning at ~0
|
|
|
|
|
# expected rejections/day, not a coverage gap in the fail-closed gate:
|
|
|
|
|
# every burst email is still rejected before any DynamoDB write, and
|
|
|
|
|
# the burst stays fully visible as ai_fallback_rejected datapoints and
|
|
|
|
|
# ReasonCode log lines, with the pre-call ai_fallback emit also raising
|
|
|
|
|
# the fallback-rate numerator above. A same-window burst detector
|
|
|
|
|
# (1-of-1 at a higher threshold) is a tracked follow-up if faster
|
|
|
|
|
# single-window paging is wanted.
|
|
|
|
|
rejected_metric = cloudwatch.Metric(
|
|
|
|
|
namespace="Seahaven/PoIngest",
|
|
|
|
|
metric_name="ParseOutcome",
|
|
|
|
|
dimensions_map={"ParseMethod": "ai_fallback_rejected"},
|
|
|
|
|
statistic="Sum",
|
|
|
|
|
period=Duration.hours(6),
|
|
|
|
|
)
|
|
|
|
|
rejected_floor = cloudwatch.MathExpression(
|
|
|
|
|
expression="IF(FILL(rej,0)>=1, FILL(rej,0), 0)",
|
|
|
|
|
using_metrics={"rej": rejected_metric},
|
|
|
|
|
period=Duration.hours(6),
|
|
|
|
|
label="AiFallbackRejectedCount",
|
|
|
|
|
)
|
|
|
|
|
rejected_floor.create_alarm(
|
|
|
|
|
self,
|
|
|
|
|
"EmailProcessorAiFallbackRejectedAlarm",
|
|
|
|
|
alarm_name="po-email-processor-ai-fallback-rejected",
|
|
|
|
|
alarm_description=(
|
|
|
|
|
"po-email-processor is rejecting Bedrock AI-fallback output at "
|
|
|
|
|
"the validation gate (possible prompt-injection probing or "
|
|
|
|
|
"template drift silently dropping real mail)"
|
|
|
|
|
),
|
|
|
|
|
threshold=1,
|
|
|
|
|
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
|
|
|
|
|
evaluation_periods=4,
|
|
|
|
|
datapoints_to_alarm=2,
|
|
|
|
|
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
|
|
|
|
|
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
|
|
|
|
|
|
2026-04-07 12:12:30 -04:00
|
|
|
# S3 event notification → Lambda
|
|
|
|
|
email_bucket.add_event_notification(
|
|
|
|
|
s3.EventType.OBJECT_CREATED,
|
|
|
|
|
s3n.LambdaDestination(email_processor),
|
|
|
|
|
s3.NotificationKeyFilter(prefix="inbound/"),
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
# --- SES Receipt Rule ---
|
|
|
|
|
# Reuse the existing INBOUND_MAIL rule set (shared with workorder-ingest)
|
|
|
|
|
rule_set = ses.ReceiptRuleSet.from_receipt_rule_set_name(
|
2026-05-08 16:01:21 -04:00
|
|
|
self,
|
|
|
|
|
"ExistingRuleSet",
|
|
|
|
|
"INBOUND_MAIL",
|
2026-04-07 12:12:30 -04:00
|
|
|
)
|
|
|
|
|
|
|
|
|
|
rule_set.add_rule(
|
|
|
|
|
"PoEmailRule",
|
|
|
|
|
recipients=["amazon_po@int.seahaven.com"],
|
|
|
|
|
actions=[
|
|
|
|
|
ses_actions.S3(
|
|
|
|
|
bucket=email_bucket,
|
|
|
|
|
object_key_prefix="inbound/",
|
|
|
|
|
),
|
|
|
|
|
],
|
|
|
|
|
)
|
|
|
|
|
|
Land safe fixes from 2026-06-17 security sweep (#97)
* Remove gratuitous KMS grant on shared DynamoDB CMK
wo-email-processor held grant_encrypt_decrypt on the shared
seahaven-dynamodb CMK, but the WorkOrders/WorkOrderComments tables
are not encrypted with that CMK. The grant was dead weight that
extended the WO processor's decrypt reach to the CMK protecting the
purchase-orders table (cross-stack decrypt). Drop it to restore
least privilege; re-add as part of the table CMK migration (INFRA-6).
Refs: INFRA-6
* Require Secrets Manager key for Anthropic client
Remove the silent fallback to a plaintext ANTHROPIC_API_KEY env var
in both email processors; require ANTHROPIC_API_KEY_SECRET_ARN and
raise if absent so a misconfigured deploy fails loudly instead of
using an unmanaged key.
Adapted from f175323 on security/sweep-2026-06-17. The From-header
sender-domain allowlist from that commit is intentionally dropped:
the From header is spoofable (INFRA-107, confirmed critical) and
sender authentication is being reworked in a separate PR.
Refs: INFRA-107
* Merge PO revisions and handle out-of-order events
save_revision did a full put_item overwrite, so a revision omitting
line_items/supplier permanently deleted them. save_new_po used a
conditional put that silently dropped the PO when an out-of-order
cancellation had already created a skeleton row.
Switch both to field-level merge update_items: a revision now SETs
only the fields it carries, and a new_po backfills data into a
pre-existing Cancelled skeleton while preserving the Cancelled
status. No email can now delete data established by an earlier one.
* Gate web UIs behind auth and escape currency XSS
The po-web-ui and workorder-web-ui handlers had no auth: any
invocation path returned the full PO/WO DB. Add a fail-closed
shared-secret gate (X-Auth-Token / Bearer, constant-time compared to
WEB_UI_AUTH_TOKEN) so a future re-attached Function URL cannot
re-expose the data (URLs removed under INFRA-74). Wire the token from
the SSM String param /procurement-ingest/web-ui-auth-token.
Also fix stored XSS in po-web-ui fmt_currency: the non-numeric
fallback returned str(val) unescaped, so a prompt-injected email
could make Claude emit total_amount as <script>. Escape it.
Refs: INFRA-74
* Document sweep security fixes and merge semantics
Update the README for the 2026-06-17 security sweep: required
Secrets Manager key (no plaintext env fallback), web UI auth gate +
SSM token setup step, output-escaping note, and the new PO
revision/cancellation merge behavior.
Adapted from d91f45e on security/sweep-2026-06-17; the sender
allowlist documentation is dropped along with the allowlist itself
(deferred to the INFRA-107 sender-authentication rework).
Refs: INFRA-107
* fix: resolve web UI auth token from Secrets Manager at runtime
Replace the plaintext SSM String parameter with a Secrets Manager secret
referenced by ARN only. The token is fetched and cached at module level on
first invocation, keeping shared secrets out of CloudFormation templates and
Lambda environment variables.
Refs: PR-97
* Add TTL to web UI auth token cache for rotation
The web-ui handlers cached the Secrets Manager auth token at module
level with no expiry, so a rotated secret was only picked up when the
warm container recycled — an emergency rotation could take hours to
take effect. Cache the fetched value for a 5-minute TTL instead, so a
rotated token propagates within the TTL while still avoiding a Secrets
Manager call on every request. Still fails closed when the secret is
unset or unreadable.
Refs: INFRA-74
* Log Secrets Manager failures in web UI auth token fetch
The web UI auth gate correctly fails closed when the shared token
cannot be read, but _get_auth_token() swallowed every exception
silently. A Secrets Manager permission or config error then made
every request 401 with no operational signal, leaving an outage
indistinguishable from ordinary unauthenticated traffic.
Add a module-level logger to both web_ui handlers and log the
fetch failure with logger.exception() in the except block before
returning None. Behavior is unchanged (still fails closed); the
failure is now visible in CloudWatch. The secret value is never
logged. The two handlers stay byte-consistent in the mirrored
_get_auth_token() region.
The companion finding on the CDK import of the shared
procurement-ingest/web-ui-auth-token secret was evaluated and left
as-is: the token is a single secret shared by both the PO and WO
stacks, so from_secret_name_v2 (which scopes grant_read via the
standard 6-char suffix wildcard) is correct; making it a CDK-managed
Secret in both stacks would collide the two stacks on the same
explicit secret name at deploy time.
Refs: INFRA-74
* Make Cancelled PO status sticky via atomic write
The PO merge path read status with a get_item (_is_cancelled) and then
wrote with an unconditional update_item. Two defects followed from this:
- Race (Issue A): a cancellation landing between the read and the write
was silently un-cancelled by a revision carrying a non-cancelled
po_status — a TOCTOU on a table with concurrent email processing.
- Over-broad strip (Issue B): save_revision dropped po_status whenever
the PO was Cancelled, so legitimate status updates on non-cancelled
POs and status-less revisions were affected rather than only the true
un-cancel transition.
Enforce the invariant server-side instead. "Cancelled" is a sticky,
authoritative status: once set, later new_po/revision emails may enrich
other fields but must never move it to a non-cancelled status. When the
payload carries a non-cancelled po_status, _merge_update issues the
update_item guarded by ConditionExpression "attribute_not_exists(po_status)
OR po_status <> :marker", evaluated atomically at write time, so a
cancellation that lands first always wins. On ConditionalCheckFailedException
the same fields are re-written without po_status/cancelled_at, enriching the
record while Cancelled sticks. Payloads with no status change, or an already
-Cancelled status, take a plain merge — the status is only ever suppressed on
a real un-cancel. This removes the non-atomic get_item from the write path;
_is_cancelled is deleted. Key schema and attribute names are unchanged, so the
cross-stack purchase-orders contract (read-only by seahaven-slack-bot) holds.
Add moto-backed tests covering un-cancel suppression with field enrichment,
status-less merge onto a Cancelled PO, legitimate status updates on
non-cancelled POs, new_po backfill of a Cancelled skeleton, fresh
create/merge, and authoritative save_cancellation.
Refs: #97
2026-07-15 20:17:46 -04:00
|
|
|
# --- Web UI auth token secret ---
|
|
|
|
|
# Shared secret for the web UI auth gate, stored in Secrets Manager and
|
|
|
|
|
# resolved at runtime so the token never appears in CloudFormation templates
|
|
|
|
|
# or Lambda environment variables. Create this secret before deploying
|
|
|
|
|
# either stack; both PO and WO stacks reference it by name.
|
|
|
|
|
web_ui_auth_secret = secretsmanager.Secret.from_secret_name_v2(
|
|
|
|
|
self,
|
|
|
|
|
"WebUiAuthToken",
|
|
|
|
|
"procurement-ingest/web-ui-auth-token",
|
|
|
|
|
)
|
|
|
|
|
|
2026-04-07 12:12:30 -04:00
|
|
|
# --- Web UI Lambda ---
|
|
|
|
|
web_ui = lambda_.Function(
|
2026-05-08 16:01:21 -04:00
|
|
|
self,
|
|
|
|
|
"WebUI",
|
2026-04-07 12:12:30 -04:00
|
|
|
function_name="po-web-ui",
|
|
|
|
|
runtime=lambda_.Runtime.PYTHON_3_12,
|
2026-04-30 14:26:53 -04:00
|
|
|
architecture=lambda_.Architecture.ARM_64,
|
2026-04-07 12:12:30 -04:00
|
|
|
handler="handler.handler",
|
feat: widen email-processor asset roots to lambdas/ with scoped globs + excludes (refactor phase 2) (#109)
Both email-processor Code.from_asset calls now bundle from lambdas/
instead of their per-function subdirectory, so Phase 3's shared/
module is reachable from the asset root once it lands. The bundling
commands were rewritten for the new cwd (pip install -r <po|wo>/
email_processor/requirements.txt -t /asset-output && cp <po|wo>/
email_processor/*.py /asset-output/), preserving the ARM64
--platform manylinux2014_aarch64 --only-binary=:all: pin exactly —
its removal shipped x86 wheels into the ARM64 function and caused a
100% outage (PR #34).
All five from_asset calls (both email processors, po web_ui, po
site_extractor, wo web_ui) now exclude **/__pycache__/**; the two
widened ones also exclude **/tests/** and **/package/**. Without the
package/ exclude, the stale untracked 44 MB
lambdas/po/email_processor/package/ dir (local-only, never present
in CI) would diverge local vs CI asset hashes and force spurious
redeploys — from_asset doesn't honor .gitignore. That dir is left in
place; deleting it is Adam's call.
WO's prod zip shrinks as deliberate cleanup, not a byte-identical
match to PO: the old `cp -r .` shipped tests/ (real scrubbed .eml
fixtures), __pycache__/, and requirements.txt into production. The
acceptance bar for WO is runtime-imported module set unchanged +
smoke, not a byte-identical zip; PO keeps the byte-identical
first-party file set guarantee. tests/test_bundle_consistency.py is
updated in the same change to recognize the scoped
`cp po/email_processor/*.py` (resp. wo) glob as the new
unconditionally-safe shape, without loosening the allowlist-revert
detection, the detection-logic mutation test, or the
PO_EXPECTED_TOP_LEVEL_MODULES exact-set pin.
No code moved under lambdas/ in this change (git diff main...HEAD --
lambdas/ is empty); only CDK asset wiring and its tests changed.
2026-07-17 15:47:01 -04:00
|
|
|
code=lambda_.Code.from_asset(
|
feat: extract lambdas/shared/ — single-source ses_auth, web_ui auth, email parsing, EMF emitter (refactor phase 3) (#111)
Four modules move into the handbook-mandated lambdas/shared/ location,
collapsing duplicated logic that had to be kept in sync by hand across
the PO and WO pipelines:
- ses_auth.py: the PO and WO copies were verified sha256-identical
against the feature/phase-7-ops-recovery baseline before the move
(no drift since the last audit). shared/ses_auth.py is the exact
bytes of that one copy; both originals are git rm'd (the PO copy
via rename, the WO copy as a straight delete). Bundling lands the
module flat in /asset-output for both email processors, so the
handlers keep `from ses_auth import authenticate_inbound_email`
unchanged — zero handler diff for this move, which is what keeps
fail-closed auth byte-identical through the change.
- web_ui_auth.py: extracts the byte-identical _get_auth_token /
_header / is_authenticated block plus the four token-cache globals
out of both web_ui handlers. The per-stack INFRA-74 comments stay
in each handler as-is (deliberately drifted wording, stack-specific)
rather than being unified into the shared module. Fail-closed
semantics (unset ARN or Secrets Manager exception -> deny) are
unchanged.
- email_parsing.py: parse_raw_email ships as the superset version that
returns cc unconditionally. WO's output is bit-identical to before;
PO simply ignores the cc field rather than being "cleaned up" to
consume it. No second variant is kept.
- emf.py: a generic emitter parameterized by namespace, dimension
sets, and properties. Every call site's emitted EMF envelope is
unchanged, including the load-bearing
[["ParseMethod"],["ParseMethod","TemplateId"]] dimension-set shape
the alarms and metric filters depend on. Emission ordering is
untouched: PO still emits ai_fallback before the Bedrock call, WO
still emits its mutually-exclusive ai_fallback/ai_fallback_rejected
after its gate. The deliberate-double-count comments survive.
_emit_derived_agreement_metric was found living inside
derived_fields.py, so per the DERIVED-FIELDS exception it is left
as a third, unconverted copy (derived_fields.py and the shadow
DerivedFieldAgreement telemetry stay untouchable while that bake
runs) — a comment there points at shared/emf.py for the eventual
follow-up.
Bundling: both email-processor cdk bundling commands gain a trailing
`cp shared/*.py /asset-output/` (they were already cp-only post-Phase
7, so no pip step or manylinux pin is reintroduced). Both web_ui
functions gain the same widened-root staging so web_ui_auth.py ships
beside their handler; site_extractor's from_asset is untouched.
Tests: PO_EXPECTED_TOP_LEVEL_MODULES gains the shared modules that now
ship, the AST sibling-import check resolves imports whose source now
lives under shared/, and the new shared cp line has its own
revert/mutation detection. _SIBLING_MODULES resolution and
_po_parser_support.py now load ses_auth/email_parsing/emf from
shared/; the two-copy ses_auth byte-identity fixture-hygiene test is
retired as obsolete now that there is one copy, and the ses_auth
fixture parameterization over two identical copies is dropped. The
sys.modules save/restore dance for template_parser (still duplicated
per-pipeline) is left in place.
2026-07-20 13:38:23 -04:00
|
|
|
"../lambdas",
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
exclude=["**/__pycache__/**", "requirements.txt"],
|
feat: extract lambdas/shared/ — single-source ses_auth, web_ui auth, email parsing, EMF emitter (refactor phase 3) (#111)
Four modules move into the handbook-mandated lambdas/shared/ location,
collapsing duplicated logic that had to be kept in sync by hand across
the PO and WO pipelines:
- ses_auth.py: the PO and WO copies were verified sha256-identical
against the feature/phase-7-ops-recovery baseline before the move
(no drift since the last audit). shared/ses_auth.py is the exact
bytes of that one copy; both originals are git rm'd (the PO copy
via rename, the WO copy as a straight delete). Bundling lands the
module flat in /asset-output for both email processors, so the
handlers keep `from ses_auth import authenticate_inbound_email`
unchanged — zero handler diff for this move, which is what keeps
fail-closed auth byte-identical through the change.
- web_ui_auth.py: extracts the byte-identical _get_auth_token /
_header / is_authenticated block plus the four token-cache globals
out of both web_ui handlers. The per-stack INFRA-74 comments stay
in each handler as-is (deliberately drifted wording, stack-specific)
rather than being unified into the shared module. Fail-closed
semantics (unset ARN or Secrets Manager exception -> deny) are
unchanged.
- email_parsing.py: parse_raw_email ships as the superset version that
returns cc unconditionally. WO's output is bit-identical to before;
PO simply ignores the cc field rather than being "cleaned up" to
consume it. No second variant is kept.
- emf.py: a generic emitter parameterized by namespace, dimension
sets, and properties. Every call site's emitted EMF envelope is
unchanged, including the load-bearing
[["ParseMethod"],["ParseMethod","TemplateId"]] dimension-set shape
the alarms and metric filters depend on. Emission ordering is
untouched: PO still emits ai_fallback before the Bedrock call, WO
still emits its mutually-exclusive ai_fallback/ai_fallback_rejected
after its gate. The deliberate-double-count comments survive.
_emit_derived_agreement_metric was found living inside
derived_fields.py, so per the DERIVED-FIELDS exception it is left
as a third, unconverted copy (derived_fields.py and the shadow
DerivedFieldAgreement telemetry stay untouchable while that bake
runs) — a comment there points at shared/emf.py for the eventual
follow-up.
Bundling: both email-processor cdk bundling commands gain a trailing
`cp shared/*.py /asset-output/` (they were already cp-only post-Phase
7, so no pip step or manylinux pin is reintroduced). Both web_ui
functions gain the same widened-root staging so web_ui_auth.py ships
beside their handler; site_extractor's from_asset is untouched.
Tests: PO_EXPECTED_TOP_LEVEL_MODULES gains the shared modules that now
ship, the AST sibling-import check resolves imports whose source now
lives under shared/, and the new shared cp line has its own
revert/mutation detection. _SIBLING_MODULES resolution and
_po_parser_support.py now load ses_auth/email_parsing/emf from
shared/; the two-copy ses_auth byte-identity fixture-hygiene test is
retired as obsolete now that there is one copy, and the ses_auth
fixture parameterization over two identical copies is dropped. The
sys.modules save/restore dance for template_parser (still duplicated
per-pipeline) is left in place.
2026-07-20 13:38:23 -04:00
|
|
|
bundling=cdk.BundlingOptions(
|
|
|
|
|
image=lambda_.Runtime.PYTHON_3_12.bundling_image,
|
|
|
|
|
command=[
|
|
|
|
|
"bash",
|
|
|
|
|
"-c",
|
|
|
|
|
# web_ui_auth.py is shared (lambdas/shared) and must land
|
|
|
|
|
# FLAT beside handler.py so the bare
|
|
|
|
|
# `from web_ui_auth import is_authenticated` resolves at
|
|
|
|
|
# runtime. Only web_ui_auth is copied from shared/ -- the
|
|
|
|
|
# other shared modules (ses_auth/email_parsing/emf) are
|
|
|
|
|
# email-processor-only and must not bloat the web UI zip.
|
|
|
|
|
# requirements.txt stays excluded (Dependabot anchor only,
|
|
|
|
|
# never runtime): keeps the deployed file list = {handler,
|
|
|
|
|
# web_ui_auth}. NOTE: the top-level `exclude=` on
|
|
|
|
|
# from_asset only filters the asset-hash fingerprint, NOT
|
|
|
|
|
# the directory Docker bundling actually mounts, so a
|
|
|
|
|
# local __pycache__/requirements.txt on disk at synth
|
|
|
|
|
# time WOULD otherwise leak into the bundled zip -- strip
|
|
|
|
|
# them explicitly post-cp instead of relying on exclude.
|
|
|
|
|
"cp -r po/web_ui/. /asset-output/ && "
|
|
|
|
|
"cp shared/web_ui_auth.py /asset-output/ && "
|
|
|
|
|
"rm -rf /asset-output/__pycache__ /asset-output/requirements.txt",
|
|
|
|
|
],
|
|
|
|
|
),
|
feat: widen email-processor asset roots to lambdas/ with scoped globs + excludes (refactor phase 2) (#109)
Both email-processor Code.from_asset calls now bundle from lambdas/
instead of their per-function subdirectory, so Phase 3's shared/
module is reachable from the asset root once it lands. The bundling
commands were rewritten for the new cwd (pip install -r <po|wo>/
email_processor/requirements.txt -t /asset-output && cp <po|wo>/
email_processor/*.py /asset-output/), preserving the ARM64
--platform manylinux2014_aarch64 --only-binary=:all: pin exactly —
its removal shipped x86 wheels into the ARM64 function and caused a
100% outage (PR #34).
All five from_asset calls (both email processors, po web_ui, po
site_extractor, wo web_ui) now exclude **/__pycache__/**; the two
widened ones also exclude **/tests/** and **/package/**. Without the
package/ exclude, the stale untracked 44 MB
lambdas/po/email_processor/package/ dir (local-only, never present
in CI) would diverge local vs CI asset hashes and force spurious
redeploys — from_asset doesn't honor .gitignore. That dir is left in
place; deleting it is Adam's call.
WO's prod zip shrinks as deliberate cleanup, not a byte-identical
match to PO: the old `cp -r .` shipped tests/ (real scrubbed .eml
fixtures), __pycache__/, and requirements.txt into production. The
acceptance bar for WO is runtime-imported module set unchanged +
smoke, not a byte-identical zip; PO keeps the byte-identical
first-party file set guarantee. tests/test_bundle_consistency.py is
updated in the same change to recognize the scoped
`cp po/email_processor/*.py` (resp. wo) glob as the new
unconditionally-safe shape, without loosening the allowlist-revert
detection, the detection-logic mutation test, or the
PO_EXPECTED_TOP_LEVEL_MODULES exact-set pin.
No code moved under lambdas/ in this change (git diff main...HEAD --
lambdas/ is empty); only CDK asset wiring and its tests changed.
2026-07-17 15:47:01 -04:00
|
|
|
),
|
2026-04-07 12:12:30 -04:00
|
|
|
timeout=Duration.seconds(60),
|
|
|
|
|
memory_size=256,
|
2026-04-30 14:26:53 -04:00
|
|
|
log_retention=logs.RetentionDays.TWO_MONTHS,
|
2026-04-07 12:12:30 -04:00
|
|
|
environment={
|
|
|
|
|
"PO_TABLE": "purchase-orders",
|
Land safe fixes from 2026-06-17 security sweep (#97)
* Remove gratuitous KMS grant on shared DynamoDB CMK
wo-email-processor held grant_encrypt_decrypt on the shared
seahaven-dynamodb CMK, but the WorkOrders/WorkOrderComments tables
are not encrypted with that CMK. The grant was dead weight that
extended the WO processor's decrypt reach to the CMK protecting the
purchase-orders table (cross-stack decrypt). Drop it to restore
least privilege; re-add as part of the table CMK migration (INFRA-6).
Refs: INFRA-6
* Require Secrets Manager key for Anthropic client
Remove the silent fallback to a plaintext ANTHROPIC_API_KEY env var
in both email processors; require ANTHROPIC_API_KEY_SECRET_ARN and
raise if absent so a misconfigured deploy fails loudly instead of
using an unmanaged key.
Adapted from f175323 on security/sweep-2026-06-17. The From-header
sender-domain allowlist from that commit is intentionally dropped:
the From header is spoofable (INFRA-107, confirmed critical) and
sender authentication is being reworked in a separate PR.
Refs: INFRA-107
* Merge PO revisions and handle out-of-order events
save_revision did a full put_item overwrite, so a revision omitting
line_items/supplier permanently deleted them. save_new_po used a
conditional put that silently dropped the PO when an out-of-order
cancellation had already created a skeleton row.
Switch both to field-level merge update_items: a revision now SETs
only the fields it carries, and a new_po backfills data into a
pre-existing Cancelled skeleton while preserving the Cancelled
status. No email can now delete data established by an earlier one.
* Gate web UIs behind auth and escape currency XSS
The po-web-ui and workorder-web-ui handlers had no auth: any
invocation path returned the full PO/WO DB. Add a fail-closed
shared-secret gate (X-Auth-Token / Bearer, constant-time compared to
WEB_UI_AUTH_TOKEN) so a future re-attached Function URL cannot
re-expose the data (URLs removed under INFRA-74). Wire the token from
the SSM String param /procurement-ingest/web-ui-auth-token.
Also fix stored XSS in po-web-ui fmt_currency: the non-numeric
fallback returned str(val) unescaped, so a prompt-injected email
could make Claude emit total_amount as <script>. Escape it.
Refs: INFRA-74
* Document sweep security fixes and merge semantics
Update the README for the 2026-06-17 security sweep: required
Secrets Manager key (no plaintext env fallback), web UI auth gate +
SSM token setup step, output-escaping note, and the new PO
revision/cancellation merge behavior.
Adapted from d91f45e on security/sweep-2026-06-17; the sender
allowlist documentation is dropped along with the allowlist itself
(deferred to the INFRA-107 sender-authentication rework).
Refs: INFRA-107
* fix: resolve web UI auth token from Secrets Manager at runtime
Replace the plaintext SSM String parameter with a Secrets Manager secret
referenced by ARN only. The token is fetched and cached at module level on
first invocation, keeping shared secrets out of CloudFormation templates and
Lambda environment variables.
Refs: PR-97
* Add TTL to web UI auth token cache for rotation
The web-ui handlers cached the Secrets Manager auth token at module
level with no expiry, so a rotated secret was only picked up when the
warm container recycled — an emergency rotation could take hours to
take effect. Cache the fetched value for a 5-minute TTL instead, so a
rotated token propagates within the TTL while still avoiding a Secrets
Manager call on every request. Still fails closed when the secret is
unset or unreadable.
Refs: INFRA-74
* Log Secrets Manager failures in web UI auth token fetch
The web UI auth gate correctly fails closed when the shared token
cannot be read, but _get_auth_token() swallowed every exception
silently. A Secrets Manager permission or config error then made
every request 401 with no operational signal, leaving an outage
indistinguishable from ordinary unauthenticated traffic.
Add a module-level logger to both web_ui handlers and log the
fetch failure with logger.exception() in the except block before
returning None. Behavior is unchanged (still fails closed); the
failure is now visible in CloudWatch. The secret value is never
logged. The two handlers stay byte-consistent in the mirrored
_get_auth_token() region.
The companion finding on the CDK import of the shared
procurement-ingest/web-ui-auth-token secret was evaluated and left
as-is: the token is a single secret shared by both the PO and WO
stacks, so from_secret_name_v2 (which scopes grant_read via the
standard 6-char suffix wildcard) is correct; making it a CDK-managed
Secret in both stacks would collide the two stacks on the same
explicit secret name at deploy time.
Refs: INFRA-74
* Make Cancelled PO status sticky via atomic write
The PO merge path read status with a get_item (_is_cancelled) and then
wrote with an unconditional update_item. Two defects followed from this:
- Race (Issue A): a cancellation landing between the read and the write
was silently un-cancelled by a revision carrying a non-cancelled
po_status — a TOCTOU on a table with concurrent email processing.
- Over-broad strip (Issue B): save_revision dropped po_status whenever
the PO was Cancelled, so legitimate status updates on non-cancelled
POs and status-less revisions were affected rather than only the true
un-cancel transition.
Enforce the invariant server-side instead. "Cancelled" is a sticky,
authoritative status: once set, later new_po/revision emails may enrich
other fields but must never move it to a non-cancelled status. When the
payload carries a non-cancelled po_status, _merge_update issues the
update_item guarded by ConditionExpression "attribute_not_exists(po_status)
OR po_status <> :marker", evaluated atomically at write time, so a
cancellation that lands first always wins. On ConditionalCheckFailedException
the same fields are re-written without po_status/cancelled_at, enriching the
record while Cancelled sticks. Payloads with no status change, or an already
-Cancelled status, take a plain merge — the status is only ever suppressed on
a real un-cancel. This removes the non-atomic get_item from the write path;
_is_cancelled is deleted. Key schema and attribute names are unchanged, so the
cross-stack purchase-orders contract (read-only by seahaven-slack-bot) holds.
Add moto-backed tests covering un-cancel suppression with field enrichment,
status-less merge onto a Cancelled PO, legitimate status updates on
non-cancelled POs, new_po backfill of a Cancelled skeleton, fresh
create/merge, and authoritative save_cancellation.
Refs: #97
2026-07-15 20:17:46 -04:00
|
|
|
# Defense-in-depth shared secret for the web UI handler. The
|
|
|
|
|
# handler fails closed if this ARN is unset or the secret is
|
|
|
|
|
# missing, so any future invocation path cannot re-expose the
|
|
|
|
|
# PO DB unauthenticated. The secret value is fetched at runtime
|
|
|
|
|
# from Secrets Manager (not embedded in env vars or template).
|
|
|
|
|
"WEB_UI_AUTH_TOKEN_SECRET_ARN": web_ui_auth_secret.secret_arn,
|
2026-04-07 12:12:30 -04:00
|
|
|
},
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
po_table.grant_read_data(web_ui)
|
Land safe fixes from 2026-06-17 security sweep (#97)
* Remove gratuitous KMS grant on shared DynamoDB CMK
wo-email-processor held grant_encrypt_decrypt on the shared
seahaven-dynamodb CMK, but the WorkOrders/WorkOrderComments tables
are not encrypted with that CMK. The grant was dead weight that
extended the WO processor's decrypt reach to the CMK protecting the
purchase-orders table (cross-stack decrypt). Drop it to restore
least privilege; re-add as part of the table CMK migration (INFRA-6).
Refs: INFRA-6
* Require Secrets Manager key for Anthropic client
Remove the silent fallback to a plaintext ANTHROPIC_API_KEY env var
in both email processors; require ANTHROPIC_API_KEY_SECRET_ARN and
raise if absent so a misconfigured deploy fails loudly instead of
using an unmanaged key.
Adapted from f175323 on security/sweep-2026-06-17. The From-header
sender-domain allowlist from that commit is intentionally dropped:
the From header is spoofable (INFRA-107, confirmed critical) and
sender authentication is being reworked in a separate PR.
Refs: INFRA-107
* Merge PO revisions and handle out-of-order events
save_revision did a full put_item overwrite, so a revision omitting
line_items/supplier permanently deleted them. save_new_po used a
conditional put that silently dropped the PO when an out-of-order
cancellation had already created a skeleton row.
Switch both to field-level merge update_items: a revision now SETs
only the fields it carries, and a new_po backfills data into a
pre-existing Cancelled skeleton while preserving the Cancelled
status. No email can now delete data established by an earlier one.
* Gate web UIs behind auth and escape currency XSS
The po-web-ui and workorder-web-ui handlers had no auth: any
invocation path returned the full PO/WO DB. Add a fail-closed
shared-secret gate (X-Auth-Token / Bearer, constant-time compared to
WEB_UI_AUTH_TOKEN) so a future re-attached Function URL cannot
re-expose the data (URLs removed under INFRA-74). Wire the token from
the SSM String param /procurement-ingest/web-ui-auth-token.
Also fix stored XSS in po-web-ui fmt_currency: the non-numeric
fallback returned str(val) unescaped, so a prompt-injected email
could make Claude emit total_amount as <script>. Escape it.
Refs: INFRA-74
* Document sweep security fixes and merge semantics
Update the README for the 2026-06-17 security sweep: required
Secrets Manager key (no plaintext env fallback), web UI auth gate +
SSM token setup step, output-escaping note, and the new PO
revision/cancellation merge behavior.
Adapted from d91f45e on security/sweep-2026-06-17; the sender
allowlist documentation is dropped along with the allowlist itself
(deferred to the INFRA-107 sender-authentication rework).
Refs: INFRA-107
* fix: resolve web UI auth token from Secrets Manager at runtime
Replace the plaintext SSM String parameter with a Secrets Manager secret
referenced by ARN only. The token is fetched and cached at module level on
first invocation, keeping shared secrets out of CloudFormation templates and
Lambda environment variables.
Refs: PR-97
* Add TTL to web UI auth token cache for rotation
The web-ui handlers cached the Secrets Manager auth token at module
level with no expiry, so a rotated secret was only picked up when the
warm container recycled — an emergency rotation could take hours to
take effect. Cache the fetched value for a 5-minute TTL instead, so a
rotated token propagates within the TTL while still avoiding a Secrets
Manager call on every request. Still fails closed when the secret is
unset or unreadable.
Refs: INFRA-74
* Log Secrets Manager failures in web UI auth token fetch
The web UI auth gate correctly fails closed when the shared token
cannot be read, but _get_auth_token() swallowed every exception
silently. A Secrets Manager permission or config error then made
every request 401 with no operational signal, leaving an outage
indistinguishable from ordinary unauthenticated traffic.
Add a module-level logger to both web_ui handlers and log the
fetch failure with logger.exception() in the except block before
returning None. Behavior is unchanged (still fails closed); the
failure is now visible in CloudWatch. The secret value is never
logged. The two handlers stay byte-consistent in the mirrored
_get_auth_token() region.
The companion finding on the CDK import of the shared
procurement-ingest/web-ui-auth-token secret was evaluated and left
as-is: the token is a single secret shared by both the PO and WO
stacks, so from_secret_name_v2 (which scopes grant_read via the
standard 6-char suffix wildcard) is correct; making it a CDK-managed
Secret in both stacks would collide the two stacks on the same
explicit secret name at deploy time.
Refs: INFRA-74
* Make Cancelled PO status sticky via atomic write
The PO merge path read status with a get_item (_is_cancelled) and then
wrote with an unconditional update_item. Two defects followed from this:
- Race (Issue A): a cancellation landing between the read and the write
was silently un-cancelled by a revision carrying a non-cancelled
po_status — a TOCTOU on a table with concurrent email processing.
- Over-broad strip (Issue B): save_revision dropped po_status whenever
the PO was Cancelled, so legitimate status updates on non-cancelled
POs and status-less revisions were affected rather than only the true
un-cancel transition.
Enforce the invariant server-side instead. "Cancelled" is a sticky,
authoritative status: once set, later new_po/revision emails may enrich
other fields but must never move it to a non-cancelled status. When the
payload carries a non-cancelled po_status, _merge_update issues the
update_item guarded by ConditionExpression "attribute_not_exists(po_status)
OR po_status <> :marker", evaluated atomically at write time, so a
cancellation that lands first always wins. On ConditionalCheckFailedException
the same fields are re-written without po_status/cancelled_at, enriching the
record while Cancelled sticks. Payloads with no status change, or an already
-Cancelled status, take a plain merge — the status is only ever suppressed on
a real un-cancel. This removes the non-atomic get_item from the write path;
_is_cancelled is deleted. Key schema and attribute names are unchanged, so the
cross-stack purchase-orders contract (read-only by seahaven-slack-bot) holds.
Add moto-backed tests covering un-cancel suppression with field enrichment,
status-less merge onto a Cancelled PO, legitimate status updates on
non-cancelled POs, new_po backfill of a Cancelled skeleton, fresh
create/merge, and authoritative save_cancellation.
Refs: #97
2026-07-15 20:17:46 -04:00
|
|
|
web_ui_auth_secret.grant_read(web_ui)
|
2026-04-07 12:12:30 -04:00
|
|
|
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
# --- Standard per-Lambda alarms: po-web-ui ---
|
|
|
|
|
# Throttles + p99 duration only (no errors alarm, no DLQ -- web_ui is a
|
|
|
|
|
# synchronous read path with no async DLQ). p99 / 45000 ms (75% of the
|
|
|
|
|
# 60s timeout) / eval 3, datapoints 2.
|
|
|
|
|
common.add_standard_lambda_alarms(
|
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70)
* Add CloudWatch alarm coverage for po-ingest and workorder-ingest
Expands alarm coverage across both CDK stacks. All alarms are ALARM-only
(no OK action) to the shared site-alerts SNS topic, with TreatMissingData
NOT_BREACHING. The site-alerts topic is now imported once near the top of
each stack so every alarm reuses one Topic instance.
po-ingest (cdk/po_stack.py):
- Errors: po-ingest-site-extractor
- Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui
- Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption),
po-ingest-site-extractor, po-web-ui
- DynamoDB throttle + system-error: purchase-orders, verified-sites,
pending-site-review
workorder-ingest (cdk/wo_stack.py):
- Throttles: workorder-email-processor
- Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption)
- DynamoDB throttle + system-error: WorkOrders, WorkOrderComments
DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation
dimension set, so each table alarm is a Sum math expression across operations
via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests
is deprecated/invalid in aws-cdk-lib 2.259.0).
Refs INFRA-41 / audit H-8.
* Drop NEEDS ADAM SIGN-OFF wording from alarm comments
Duration alarm thresholds are owner-approved; remove the sign-off flag
from po_stack.py and wo_stack.py comments. Threshold values, eval config,
and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
|
|
|
self,
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
"WebUi",
|
|
|
|
|
web_ui,
|
|
|
|
|
"po-web-ui",
|
|
|
|
|
alarm_topic,
|
|
|
|
|
duration_statistic="p99",
|
|
|
|
|
errors=False,
|
|
|
|
|
dlq=None,
|
|
|
|
|
descriptions={
|
|
|
|
|
"throttles": "po-web-ui invocation throttles",
|
|
|
|
|
"duration": "po-web-ui p99 duration approaching the 60s timeout",
|
|
|
|
|
},
|
|
|
|
|
)
|
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70)
* Add CloudWatch alarm coverage for po-ingest and workorder-ingest
Expands alarm coverage across both CDK stacks. All alarms are ALARM-only
(no OK action) to the shared site-alerts SNS topic, with TreatMissingData
NOT_BREACHING. The site-alerts topic is now imported once near the top of
each stack so every alarm reuses one Topic instance.
po-ingest (cdk/po_stack.py):
- Errors: po-ingest-site-extractor
- Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui
- Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption),
po-ingest-site-extractor, po-web-ui
- DynamoDB throttle + system-error: purchase-orders, verified-sites,
pending-site-review
workorder-ingest (cdk/wo_stack.py):
- Throttles: workorder-email-processor
- Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption)
- DynamoDB throttle + system-error: WorkOrders, WorkOrderComments
DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation
dimension set, so each table alarm is a Sum math expression across operations
via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests
is deprecated/invalid in aws-cdk-lib 2.259.0).
Refs INFRA-41 / audit H-8.
* Drop NEEDS ADAM SIGN-OFF wording from alarm comments
Duration alarm thresholds are owner-approved; remove the sign-off flag
from po_stack.py and wo_stack.py comments. Threshold values, eval config,
and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
|
|
|
|
Reconcile IaC with out-of-band DLQ + Function URL changes (INFRA-74, INFRA-41) (#50)
Make CDK the source of truth for two sets of changes applied out-of-band
via CLI to the po-ingest and WorkorderIngestStack stacks.
INFRA-74 (audit C-5): remove the public FunctionUrlAuthType.NONE Function
URL construct (and its auto-generated Principal:* invoke permission +
output) from both po-web-ui and workorder-web-ui. The URLs were already
deleted live via CLI; CFN's delete is idempotent.
INFRA-41 (audit H-8): add a CDK-managed SQS dead-letter queue
(dead_letter_queue=, 14d retention, SSL-enforced, CDK-generated name) and
an ALARM-only Errors alarm (Sum, threshold>0, site-alerts topic) for both
po-email-processor and workorder-email-processor, mirroring the
apm-wo-analysis-classifier DLQ and payments-payroll-batch alarm patterns.
Interim CLI resources (per-fn -dlq queues, -errors alarms, dlq-send inline
policies, OnFailure event-invoke-configs) removed post-deploy.
2026-06-08 16:02:29 -04:00
|
|
|
# Public Function URL removed 2026-06-08 (INFRA-74 / audit C-5): the
|
|
|
|
|
# unauthenticated FunctionUrlAuthType.NONE URL was deleted out-of-band
|
|
|
|
|
# via CLI. Removing the construct (and its auto-generated Principal:*
|
|
|
|
|
# invoke permission) reconciles IaC with the live state.
|
2026-04-30 14:26:53 -04:00
|
|
|
|
|
|
|
|
# --- Verified sites table (extracted from PO ship-to addresses) ---
|
|
|
|
|
verified_sites_table = dynamodb.Table(
|
2026-05-08 16:01:21 -04:00
|
|
|
self,
|
|
|
|
|
"VerifiedSitesTable",
|
2026-04-30 14:26:53 -04:00
|
|
|
table_name="verified-sites",
|
|
|
|
|
partition_key=dynamodb.Attribute(
|
|
|
|
|
name="siteCode",
|
|
|
|
|
type=dynamodb.AttributeType.STRING,
|
|
|
|
|
),
|
|
|
|
|
billing_mode=dynamodb.BillingMode.PAY_PER_REQUEST,
|
|
|
|
|
removal_policy=RemovalPolicy.RETAIN,
|
|
|
|
|
)
|
2026-06-03 15:32:23 -04:00
|
|
|
# by-state GSI removed 2026-06-03 (audit M-20): 0 reads in 30d against
|
|
|
|
|
# 518 WCU of write amplification. Re-add if a state-level query path ships.
|
2026-04-30 14:26:53 -04:00
|
|
|
|
|
|
|
|
# --- Site extractor Lambda (DynamoDB Streams → verified-sites) ---
|
|
|
|
|
site_extractor = lambda_.Function(
|
2026-05-08 16:01:21 -04:00
|
|
|
self,
|
|
|
|
|
"SiteExtractor",
|
2026-04-30 14:26:53 -04:00
|
|
|
function_name="po-ingest-site-extractor",
|
|
|
|
|
runtime=lambda_.Runtime.PYTHON_3_12,
|
|
|
|
|
architecture=lambda_.Architecture.ARM_64,
|
|
|
|
|
handler="handler.handler",
|
feat: widen email-processor asset roots to lambdas/ with scoped globs + excludes (refactor phase 2) (#109)
Both email-processor Code.from_asset calls now bundle from lambdas/
instead of their per-function subdirectory, so Phase 3's shared/
module is reachable from the asset root once it lands. The bundling
commands were rewritten for the new cwd (pip install -r <po|wo>/
email_processor/requirements.txt -t /asset-output && cp <po|wo>/
email_processor/*.py /asset-output/), preserving the ARM64
--platform manylinux2014_aarch64 --only-binary=:all: pin exactly —
its removal shipped x86 wheels into the ARM64 function and caused a
100% outage (PR #34).
All five from_asset calls (both email processors, po web_ui, po
site_extractor, wo web_ui) now exclude **/__pycache__/**; the two
widened ones also exclude **/tests/** and **/package/**. Without the
package/ exclude, the stale untracked 44 MB
lambdas/po/email_processor/package/ dir (local-only, never present
in CI) would diverge local vs CI asset hashes and force spurious
redeploys — from_asset doesn't honor .gitignore. That dir is left in
place; deleting it is Adam's call.
WO's prod zip shrinks as deliberate cleanup, not a byte-identical
match to PO: the old `cp -r .` shipped tests/ (real scrubbed .eml
fixtures), __pycache__/, and requirements.txt into production. The
acceptance bar for WO is runtime-imported module set unchanged +
smoke, not a byte-identical zip; PO keeps the byte-identical
first-party file set guarantee. tests/test_bundle_consistency.py is
updated in the same change to recognize the scoped
`cp po/email_processor/*.py` (resp. wo) glob as the new
unconditionally-safe shape, without loosening the allowlist-revert
detection, the detection-logic mutation test, or the
PO_EXPECTED_TOP_LEVEL_MODULES exact-set pin.
No code moved under lambdas/ in this change (git diff main...HEAD --
lambdas/ is empty); only CDK asset wiring and its tests changed.
2026-07-17 15:47:01 -04:00
|
|
|
code=lambda_.Code.from_asset(
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
# requirements.txt is excluded from the bundle: it exists only
|
|
|
|
|
# as a Dependabot anchor (git-based scan sees it), never pip-
|
|
|
|
|
# installed (this is a plain non-bundled asset) and never needed
|
|
|
|
|
# at runtime (boto3 comes from the Lambda runtime). Excluding it
|
|
|
|
|
# keeps the deployed asset hash neutral vs base while the manifest
|
|
|
|
|
# still lands in git for Dependabot.
|
|
|
|
|
"../lambdas/po/site_extractor",
|
|
|
|
|
exclude=["**/__pycache__/**", "requirements.txt"],
|
feat: widen email-processor asset roots to lambdas/ with scoped globs + excludes (refactor phase 2) (#109)
Both email-processor Code.from_asset calls now bundle from lambdas/
instead of their per-function subdirectory, so Phase 3's shared/
module is reachable from the asset root once it lands. The bundling
commands were rewritten for the new cwd (pip install -r <po|wo>/
email_processor/requirements.txt -t /asset-output && cp <po|wo>/
email_processor/*.py /asset-output/), preserving the ARM64
--platform manylinux2014_aarch64 --only-binary=:all: pin exactly —
its removal shipped x86 wheels into the ARM64 function and caused a
100% outage (PR #34).
All five from_asset calls (both email processors, po web_ui, po
site_extractor, wo web_ui) now exclude **/__pycache__/**; the two
widened ones also exclude **/tests/** and **/package/**. Without the
package/ exclude, the stale untracked 44 MB
lambdas/po/email_processor/package/ dir (local-only, never present
in CI) would diverge local vs CI asset hashes and force spurious
redeploys — from_asset doesn't honor .gitignore. That dir is left in
place; deleting it is Adam's call.
WO's prod zip shrinks as deliberate cleanup, not a byte-identical
match to PO: the old `cp -r .` shipped tests/ (real scrubbed .eml
fixtures), __pycache__/, and requirements.txt into production. The
acceptance bar for WO is runtime-imported module set unchanged +
smoke, not a byte-identical zip; PO keeps the byte-identical
first-party file set guarantee. tests/test_bundle_consistency.py is
updated in the same change to recognize the scoped
`cp po/email_processor/*.py` (resp. wo) glob as the new
unconditionally-safe shape, without loosening the allowlist-revert
detection, the detection-logic mutation test, or the
PO_EXPECTED_TOP_LEVEL_MODULES exact-set pin.
No code moved under lambdas/ in this change (git diff main...HEAD --
lambdas/ is empty); only CDK asset wiring and its tests changed.
2026-07-17 15:47:01 -04:00
|
|
|
),
|
2026-04-30 14:26:53 -04:00
|
|
|
timeout=Duration.seconds(60),
|
|
|
|
|
memory_size=256,
|
|
|
|
|
log_retention=logs.RetentionDays.TWO_MONTHS,
|
|
|
|
|
environment={
|
|
|
|
|
"VERIFIED_SITES_TABLE": verified_sites_table.table_name,
|
2026-04-30 15:01:09 -04:00
|
|
|
"PENDING_REVIEW_TABLE": "pending-site-review",
|
2026-04-30 14:26:53 -04:00
|
|
|
},
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
verified_sites_table.grant_read_write_data(site_extractor)
|
|
|
|
|
|
|
|
|
|
site_extractor.add_event_source(
|
|
|
|
|
lambda_event_sources.DynamoEventSource(
|
|
|
|
|
po_table,
|
|
|
|
|
starting_position=lambda_.StartingPosition.TRIM_HORIZON,
|
|
|
|
|
batch_size=10,
|
|
|
|
|
max_batching_window=Duration.seconds(30),
|
|
|
|
|
bisect_batch_on_error=True,
|
|
|
|
|
retry_attempts=3,
|
|
|
|
|
)
|
|
|
|
|
)
|
|
|
|
|
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
# --- Standard per-Lambda alarms: po-ingest-site-extractor ---
|
|
|
|
|
# errors + throttles + p99 duration. NO DLQ alarm: site_extractor is a
|
|
|
|
|
# DynamoEventSource stream consumer with no async DLQ attached (dlq=None).
|
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70)
* Add CloudWatch alarm coverage for po-ingest and workorder-ingest
Expands alarm coverage across both CDK stacks. All alarms are ALARM-only
(no OK action) to the shared site-alerts SNS topic, with TreatMissingData
NOT_BREACHING. The site-alerts topic is now imported once near the top of
each stack so every alarm reuses one Topic instance.
po-ingest (cdk/po_stack.py):
- Errors: po-ingest-site-extractor
- Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui
- Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption),
po-ingest-site-extractor, po-web-ui
- DynamoDB throttle + system-error: purchase-orders, verified-sites,
pending-site-review
workorder-ingest (cdk/wo_stack.py):
- Throttles: workorder-email-processor
- Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption)
- DynamoDB throttle + system-error: WorkOrders, WorkOrderComments
DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation
dimension set, so each table alarm is a Sum math expression across operations
via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests
is deprecated/invalid in aws-cdk-lib 2.259.0).
Refs INFRA-41 / audit H-8.
* Drop NEEDS ADAM SIGN-OFF wording from alarm comments
Duration alarm thresholds are owner-approved; remove the sign-off flag
from po_stack.py and wo_stack.py comments. Threshold values, eval config,
and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
|
|
|
# Stream-consumer errors retry per the event-source config, but a
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
# persistent failure stalls the verified-sites pipeline. p99 / 45000 ms
|
|
|
|
|
# (75% of the 60s timeout) / eval 3, datapoints 2.
|
|
|
|
|
common.add_standard_lambda_alarms(
|
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70)
* Add CloudWatch alarm coverage for po-ingest and workorder-ingest
Expands alarm coverage across both CDK stacks. All alarms are ALARM-only
(no OK action) to the shared site-alerts SNS topic, with TreatMissingData
NOT_BREACHING. The site-alerts topic is now imported once near the top of
each stack so every alarm reuses one Topic instance.
po-ingest (cdk/po_stack.py):
- Errors: po-ingest-site-extractor
- Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui
- Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption),
po-ingest-site-extractor, po-web-ui
- DynamoDB throttle + system-error: purchase-orders, verified-sites,
pending-site-review
workorder-ingest (cdk/wo_stack.py):
- Throttles: workorder-email-processor
- Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption)
- DynamoDB throttle + system-error: WorkOrders, WorkOrderComments
DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation
dimension set, so each table alarm is a Sum math expression across operations
via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests
is deprecated/invalid in aws-cdk-lib 2.259.0).
Refs INFRA-41 / audit H-8.
* Drop NEEDS ADAM SIGN-OFF wording from alarm comments
Duration alarm thresholds are owner-approved; remove the sign-off flag
from po_stack.py and wo_stack.py comments. Threshold values, eval config,
and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
|
|
|
self,
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
"SiteExtractor",
|
|
|
|
|
site_extractor,
|
|
|
|
|
"po-ingest-site-extractor",
|
|
|
|
|
alarm_topic,
|
|
|
|
|
duration_statistic="p99",
|
|
|
|
|
errors=True,
|
|
|
|
|
dlq=None,
|
|
|
|
|
descriptions={
|
|
|
|
|
"errors": "po-ingest-site-extractor invocation errors",
|
|
|
|
|
"throttles": "po-ingest-site-extractor invocation throttles",
|
|
|
|
|
"duration": "po-ingest-site-extractor p99 duration approaching the 60s timeout",
|
|
|
|
|
},
|
|
|
|
|
)
|
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70)
* Add CloudWatch alarm coverage for po-ingest and workorder-ingest
Expands alarm coverage across both CDK stacks. All alarms are ALARM-only
(no OK action) to the shared site-alerts SNS topic, with TreatMissingData
NOT_BREACHING. The site-alerts topic is now imported once near the top of
each stack so every alarm reuses one Topic instance.
po-ingest (cdk/po_stack.py):
- Errors: po-ingest-site-extractor
- Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui
- Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption),
po-ingest-site-extractor, po-web-ui
- DynamoDB throttle + system-error: purchase-orders, verified-sites,
pending-site-review
workorder-ingest (cdk/wo_stack.py):
- Throttles: workorder-email-processor
- Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption)
- DynamoDB throttle + system-error: WorkOrders, WorkOrderComments
DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation
dimension set, so each table alarm is a Sum math expression across operations
via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests
is deprecated/invalid in aws-cdk-lib 2.259.0).
Refs INFRA-41 / audit H-8.
* Drop NEEDS ADAM SIGN-OFF wording from alarm comments
Duration alarm thresholds are owner-approved; remove the sign-off flag
from po_stack.py and wo_stack.py comments. Threshold values, eval config,
and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
|
|
|
|
2026-05-08 16:01:21 -04:00
|
|
|
cdk.CfnOutput(
|
|
|
|
|
self,
|
|
|
|
|
"VerifiedSitesTableName",
|
2026-04-30 14:26:53 -04:00
|
|
|
value=verified_sites_table.table_name,
|
|
|
|
|
description="Verified site addresses extracted from POs",
|
|
|
|
|
)
|
2026-04-30 15:01:09 -04:00
|
|
|
|
|
|
|
|
# --- Pending site review table (POs with no extractable site code) ---
|
|
|
|
|
pending_review_table = dynamodb.Table(
|
2026-05-08 16:01:21 -04:00
|
|
|
self,
|
|
|
|
|
"PendingSiteReviewTable",
|
2026-04-30 15:01:09 -04:00
|
|
|
table_name="pending-site-review",
|
|
|
|
|
partition_key=dynamodb.Attribute(
|
|
|
|
|
name="po_number",
|
|
|
|
|
type=dynamodb.AttributeType.STRING,
|
|
|
|
|
),
|
|
|
|
|
billing_mode=dynamodb.BillingMode.PAY_PER_REQUEST,
|
|
|
|
|
removal_policy=RemovalPolicy.RETAIN,
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
pending_review_table.grant_read_write_data(site_extractor)
|
|
|
|
|
verified_sites_table.grant_read_data(site_extractor)
|
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70)
* Add CloudWatch alarm coverage for po-ingest and workorder-ingest
Expands alarm coverage across both CDK stacks. All alarms are ALARM-only
(no OK action) to the shared site-alerts SNS topic, with TreatMissingData
NOT_BREACHING. The site-alerts topic is now imported once near the top of
each stack so every alarm reuses one Topic instance.
po-ingest (cdk/po_stack.py):
- Errors: po-ingest-site-extractor
- Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui
- Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption),
po-ingest-site-extractor, po-web-ui
- DynamoDB throttle + system-error: purchase-orders, verified-sites,
pending-site-review
workorder-ingest (cdk/wo_stack.py):
- Throttles: workorder-email-processor
- Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption)
- DynamoDB throttle + system-error: WorkOrders, WorkOrderComments
DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation
dimension set, so each table alarm is a Sum math expression across operations
via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests
is deprecated/invalid in aws-cdk-lib 2.259.0).
Refs INFRA-41 / audit H-8.
* Drop NEEDS ADAM SIGN-OFF wording from alarm comments
Duration alarm thresholds are owner-approved; remove the sign-off flag
from po_stack.py and wo_stack.py comments. Threshold values, eval config,
and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
|
|
|
|
|
|
|
|
# --- DynamoDB throttle + system-error alarms ---
|
|
|
|
|
# ThrottledRequests / SystemErrors emit at TableName + Operation only
|
|
|
|
|
# (verified against live CloudWatch: no TableName-only rollup exists, and
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
# metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.261.0).
|
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70)
* Add CloudWatch alarm coverage for po-ingest and workorder-ingest
Expands alarm coverage across both CDK stacks. All alarms are ALARM-only
(no OK action) to the shared site-alerts SNS topic, with TreatMissingData
NOT_BREACHING. The site-alerts topic is now imported once near the top of
each stack so every alarm reuses one Topic instance.
po-ingest (cdk/po_stack.py):
- Errors: po-ingest-site-extractor
- Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui
- Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption),
po-ingest-site-extractor, po-web-ui
- DynamoDB throttle + system-error: purchase-orders, verified-sites,
pending-site-review
workorder-ingest (cdk/wo_stack.py):
- Throttles: workorder-email-processor
- Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption)
- DynamoDB throttle + system-error: WorkOrders, WorkOrderComments
DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation
dimension set, so each table alarm is a Sum math expression across operations
via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests
is deprecated/invalid in aws-cdk-lib 2.259.0).
Refs INFRA-41 / audit H-8.
* Drop NEEDS ADAM SIGN-OFF wording from alarm comments
Duration alarm thresholds are owner-approved; remove the sign-off flag
from po_stack.py and wo_stack.py comments. Threshold values, eval config,
and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
|
|
|
# Each table currently has zero throttle/error datapoints, so the series
|
|
|
|
|
# only materialise on first occurrence — NOT_BREACHING keeps them OK until
|
|
|
|
|
# then.
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
common.add_ddb_alarms(
|
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70)
* Add CloudWatch alarm coverage for po-ingest and workorder-ingest
Expands alarm coverage across both CDK stacks. All alarms are ALARM-only
(no OK action) to the shared site-alerts SNS topic, with TreatMissingData
NOT_BREACHING. The site-alerts topic is now imported once near the top of
each stack so every alarm reuses one Topic instance.
po-ingest (cdk/po_stack.py):
- Errors: po-ingest-site-extractor
- Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui
- Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption),
po-ingest-site-extractor, po-web-ui
- DynamoDB throttle + system-error: purchase-orders, verified-sites,
pending-site-review
workorder-ingest (cdk/wo_stack.py):
- Throttles: workorder-email-processor
- Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption)
- DynamoDB throttle + system-error: WorkOrders, WorkOrderComments
DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation
dimension set, so each table alarm is a Sum math expression across operations
via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests
is deprecated/invalid in aws-cdk-lib 2.259.0).
Refs INFRA-41 / audit H-8.
* Drop NEEDS ADAM SIGN-OFF wording from alarm comments
Duration alarm thresholds are owner-approved; remove the sign-off flag
from po_stack.py and wo_stack.py comments. Threshold values, eval config,
and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
|
|
|
self, "PurchaseOrdersTable", po_table, "purchase-orders", alarm_topic
|
|
|
|
|
)
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
common.add_ddb_alarms(
|
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70)
* Add CloudWatch alarm coverage for po-ingest and workorder-ingest
Expands alarm coverage across both CDK stacks. All alarms are ALARM-only
(no OK action) to the shared site-alerts SNS topic, with TreatMissingData
NOT_BREACHING. The site-alerts topic is now imported once near the top of
each stack so every alarm reuses one Topic instance.
po-ingest (cdk/po_stack.py):
- Errors: po-ingest-site-extractor
- Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui
- Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption),
po-ingest-site-extractor, po-web-ui
- DynamoDB throttle + system-error: purchase-orders, verified-sites,
pending-site-review
workorder-ingest (cdk/wo_stack.py):
- Throttles: workorder-email-processor
- Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption)
- DynamoDB throttle + system-error: WorkOrders, WorkOrderComments
DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation
dimension set, so each table alarm is a Sum math expression across operations
via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests
is deprecated/invalid in aws-cdk-lib 2.259.0).
Refs INFRA-41 / audit H-8.
* Drop NEEDS ADAM SIGN-OFF wording from alarm comments
Duration alarm thresholds are owner-approved; remove the sign-off flag
from po_stack.py and wo_stack.py comments. Threshold values, eval config,
and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
|
|
|
self,
|
|
|
|
|
"VerifiedSitesTable",
|
|
|
|
|
verified_sites_table,
|
|
|
|
|
"verified-sites",
|
|
|
|
|
alarm_topic,
|
|
|
|
|
)
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
common.add_ddb_alarms(
|
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70)
* Add CloudWatch alarm coverage for po-ingest and workorder-ingest
Expands alarm coverage across both CDK stacks. All alarms are ALARM-only
(no OK action) to the shared site-alerts SNS topic, with TreatMissingData
NOT_BREACHING. The site-alerts topic is now imported once near the top of
each stack so every alarm reuses one Topic instance.
po-ingest (cdk/po_stack.py):
- Errors: po-ingest-site-extractor
- Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui
- Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption),
po-ingest-site-extractor, po-web-ui
- DynamoDB throttle + system-error: purchase-orders, verified-sites,
pending-site-review
workorder-ingest (cdk/wo_stack.py):
- Throttles: workorder-email-processor
- Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption)
- DynamoDB throttle + system-error: WorkOrders, WorkOrderComments
DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation
dimension set, so each table alarm is a Sum math expression across operations
via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests
is deprecated/invalid in aws-cdk-lib 2.259.0).
Refs INFRA-41 / audit H-8.
* Drop NEEDS ADAM SIGN-OFF wording from alarm comments
Duration alarm thresholds are owner-approved; remove the sign-off flag
from po_stack.py and wo_stack.py comments. Threshold values, eval config,
and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
|
|
|
self,
|
|
|
|
|
"PendingSiteReviewTable",
|
|
|
|
|
pending_review_table,
|
|
|
|
|
"pending-site-review",
|
|
|
|
|
alarm_topic,
|
|
|
|
|
)
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB
alarms, the sender-auth-rejected metric filter + alarm, the standard
per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email
bucket, the async DLQ, the template-fallback-rate math alarm) move into
cdk/common.py.
Every helper is a PLAIN function taking (scope, id, ...), called with each
stack's own Stack as scope and the exact literal construct ids used inline
before, so every synthesized logical ID is byte-stable. A Construct
subclass would reparent the tree and make CloudFormation attempt to
replace the RETAIN-protected purchase-orders/WorkOrders tables and named
buckets -- data loss -- so it is forbidden. Per-function alarm variance
(PO p99 vs WO p95 duration, po-web-ui throttles+duration only,
site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved
through call-site arguments, not baked into the helpers.
make_bedrock_invoke_statement derives the inference-profile and us-east-1
foundation-model ARNs from Stack.of(scope).account/.region instead of the
hardcoded 328440206208/us-east-1 literals. The environment stays
account-agnostic (region-only), so the account resolves to the
AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the
same ARN the literal named in-account (a benign in-place IAM policy
update, never a replacement) and is account-portable rather than pinned to
the frozen management account.
The account= pin evaluated for cdk.Environment was deliberately NOT added:
resolving every account-derived value (bucket names, Lambda::Permission
source account, SNS action ARN) to literals makes CloudFormation flag the
RETAIN email buckets as requiring replacement against the deployed
account-agnostic templates -- a data-loss risk that outranks the pin, which
buys nothing (the resolved values are unchanged).
Also: net-new CfnOutputs for the five Lambda function ARNs and the
owned/consumed table names, exact-pin constructs==10.6.0, and fix the
stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment.
The common.py extraction is zero-cdk-diff on both stacks (byte-stable
logical IDs, no asset/property change); the only deltas versus deployed
are the intended benign Bedrock IAM in-place update and the additive
CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM
move; its BLOCK was a verified false positive (it read AWS::AccountId as a
wildcard -- it is a deploy-time-resolved concrete value naming one account
and one inference-profile, region is pinned us-east-1, and the grant is
strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
|
|
|
|
|
|
|
|
# --- Function ARN + consumed-table-name outputs (Phase 4, additive) ---
|
|
|
|
|
cdk.CfnOutput(
|
|
|
|
|
self,
|
|
|
|
|
"EmailProcessorFunctionArn",
|
|
|
|
|
value=email_processor.function_arn,
|
|
|
|
|
description="ARN of the po-email-processor Lambda",
|
|
|
|
|
)
|
|
|
|
|
cdk.CfnOutput(
|
|
|
|
|
self,
|
|
|
|
|
"WebUiFunctionArn",
|
|
|
|
|
value=web_ui.function_arn,
|
|
|
|
|
description="ARN of the po-web-ui Lambda",
|
|
|
|
|
)
|
|
|
|
|
cdk.CfnOutput(
|
|
|
|
|
self,
|
|
|
|
|
"SiteExtractorFunctionArn",
|
|
|
|
|
value=site_extractor.function_arn,
|
|
|
|
|
description="ARN of the po-ingest-site-extractor Lambda",
|
|
|
|
|
)
|
|
|
|
|
cdk.CfnOutput(
|
|
|
|
|
self,
|
|
|
|
|
"PurchaseOrdersTableName",
|
|
|
|
|
value=po_table.table_name,
|
|
|
|
|
description="purchase-orders DynamoDB table consumed by this stack",
|
|
|
|
|
)
|
|
|
|
|
cdk.CfnOutput(
|
|
|
|
|
self,
|
|
|
|
|
"PendingSiteReviewTableName",
|
|
|
|
|
value=pending_review_table.table_name,
|
|
|
|
|
description="pending-site-review DynamoDB table",
|
|
|
|
|
)
|