The From header and any raw-MIME Authentication-Results copies are
attacker-forgeable, so a forged email to apm@int.seahaven.com or
amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107,
CRITICAL). Both S3-triggered email processors now authenticate the
sender against the Authentication-Results header SES itself prepends
at delivery: only the topmost header is consulted, its authserv-id
must be amazonses.com, and it must carry dkim=pass for a domain in
the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set
in CDK so ops can adjust without code changes).
Allowlists come from live traffic observed 2026-07-15 on both ingest
buckets: WO mail arrives via the apm@ Google Groups forward, which
re-signs as seahaven.com (the hxgnsmartcloud.com signature does not
survive the forward); PO mail passes for amazon.coupahost.com.
amazonses.com also passes on PO mail but is deliberately excluded --
every SES customer's outbound mail passes for it.
Every failure path (env var unset, header missing or unparseable,
verdict fail, unaligned domain) rejects the email: a structured
warning with the reason and S3 key is logged and the record skipped
without erroring the invocation, so rejected mail causes no Lambda
retries or DLQ messages. Handler signatures and event sources are
unchanged.
Refs: INFRA-107
The existing workorder-email-processor-errors alarm fires on any errored
async invocation, but a message only reaches the DLQ after Lambda exhausts
its async retries and gives up — a genuinely dropped work-order email that
the errors alarm alone does not distinguish.
Add an ALARM-only CloudWatch alarm (workorder-email-processor-dlq-messages)
on the EmailProcessorDlq ApproximateNumberOfMessagesVisible metric
(Statistic MAXIMUM, period 5m, evaluationPeriods 1, threshold > 0,
treatMissingData NOT_BREACHING). Routes to the same shared site-alerts SNS
topic via SnsAction, mirroring the errors-alarm construct style.
Refs INFRA-41 / audit H-8.
Add block_public_access=BlockPublicAccess.BLOCK_ALL to the po-ingest and
workorder-ingest EmailBucket constructs. The buckets are already private
at runtime via account-level and AWS-default BPA, so this is a no-op for
behavior; it closes the codification gap that left CKV_AWS_53-56 firing
on the synthesized templates and blocking the security pre-push gate
(and the DLQ-alarm PRs that ride on it).
* Add CloudWatch alarm coverage for po-ingest and workorder-ingest
Expands alarm coverage across both CDK stacks. All alarms are ALARM-only
(no OK action) to the shared site-alerts SNS topic, with TreatMissingData
NOT_BREACHING. The site-alerts topic is now imported once near the top of
each stack so every alarm reuses one Topic instance.
po-ingest (cdk/po_stack.py):
- Errors: po-ingest-site-extractor
- Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui
- Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption),
po-ingest-site-extractor, po-web-ui
- DynamoDB throttle + system-error: purchase-orders, verified-sites,
pending-site-review
workorder-ingest (cdk/wo_stack.py):
- Throttles: workorder-email-processor
- Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption)
- DynamoDB throttle + system-error: WorkOrders, WorkOrderComments
DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation
dimension set, so each table alarm is a Sum math expression across operations
via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests
is deprecated/invalid in aws-cdk-lib 2.259.0).
Refs INFRA-41 / audit H-8.
* Drop NEEDS ADAM SIGN-OFF wording from alarm comments
Duration alarm thresholds are owner-approved; remove the sign-off flag
from po_stack.py and wo_stack.py comments. Threshold values, eval config,
and orphan-delete notes are unchanged.
The purchase-orders table was migrated to SSE-KMS (alias/seahaven-dynamodb,
INFRA-95/M-3) out-of-band, but po_stack never declared the key, so
grant_read_write_data did not propagate kms perms. po-email-processor failed
~99.6% of invocations with kms:Decrypt AccessDeniedException, a data-loss
outage on the PO ingestion write path.
- po_stack: declare encryption_key on purchase-orders (reconciles SSE drift;
no-op against the already-encrypted live table) so the existing grants add
kms:Decrypt/GenerateDataKey/DescribeKey to EmailProcessor, WebUI, SiteExtractor.
- wo_stack: pre-emptive grant_encrypt_decrypt on the WO processor role ahead of
the WorkOrders CMK migration (INFRA-6); tables left unencrypted, no table change.
GPT-4.1 cross-review: no blockers.
Make CDK the source of truth for two sets of changes applied out-of-band
via CLI to the po-ingest and WorkorderIngestStack stacks.
INFRA-74 (audit C-5): remove the public FunctionUrlAuthType.NONE Function
URL construct (and its auto-generated Principal:* invoke permission +
output) from both po-web-ui and workorder-web-ui. The URLs were already
deleted live via CLI; CFN's delete is idempotent.
INFRA-41 (audit H-8): add a CDK-managed SQS dead-letter queue
(dead_letter_queue=, 14d retention, SSL-enforced, CDK-generated name) and
an ALARM-only Errors alarm (Sum, threshold>0, site-alerts topic) for both
po-email-processor and workorder-email-processor, mirroring the
apm-wo-analysis-classifier DLQ and payments-payroll-batch alarm patterns.
Interim CLI resources (per-fn -dlq queues, -errors alarms, dlq-send inline
policies, OnFailure event-invoke-configs) removed post-deploy.
* Drop read-idle GSIs: by-state and status-index
Audit M-20: 30 days of CloudWatch metrics show 0 reads on both
indexes against 518 (by-state) and ~50k (status-index) WCU of write
amplification. No code path queries either index.
site-code-index follows in the next commit - CloudFormation allows
only one GSI change per table per deploy.
* Drop read-idle site-code-index GSI
Second half of the M-20 cleanup - deployed separately because
CloudFormation allows one GSI change per table per update.
* Merge workorder-ingest pipeline into unified repo
Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/.
Two independent CloudFormation stacks in one CDK app. Fix WO stack
compliance: ARM64 architecture, 60-day log retention, aarch64 bundling,
RETAIN on Anthropic secret. Remove stale CodePipeline buildspec.
* Fix test_local.py import path and remove dead shared/models.py
test_local.py referenced the old lambdas/email_processor path. Updated
to lambdas/wo/email_processor. Removed shared/ directory entirely as
nothing imports from it.
* Escape HTML in both web UI dashboards to prevent XSS
Both Function URLs are public (auth_type=NONE) and render
email-derived content via f-strings. Attacker-crafted emails
could inject scripts. Added html.escape() on all interpolated
values in both PO and WO dashboards.
* Add pagination to WO web UI scan
get_work_orders() only fetched the first 1MB page from DynamoDB.
Loop on LastEvaluatedKey to match the PO web UI pattern.
* Fix esc(None) TypeError and javascript: scheme in PO web UI
Coerce supplier name through `or ""` before escaping to handle
nested None from DynamoDB. Add scheme allowlist on view_order_url
to block javascript:/data: hrefs from LLM-extracted URLs.
* Fix WO render_badge None guard, updated_at slice, and backfill path
Add null guard to WO render_badge matching the PO version. Use
`or ""` before slicing updated_at to handle explicit None values.
Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor.
* Harden WO web UI and fix JS-context XSS in both dashboards
- Use json.dumps for onclick URLs to prevent JS string breakout
- Add .lower() to WO render_badge color lookup matching PO pattern
- Add pagination to get_comments query
- Cap get_work_orders to 500 results matching PO pattern
* Apply ruff formatting to web UI handlers