procurement-ingest/cdk/wo_stack.py

848 lines
38 KiB
Python
Raw Normal View History

Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
"""CDK stack for the work order email ingestion pipeline."""
import aws_cdk as cdk
from aws_cdk import (
Duration,
RemovalPolicy,
Stack,
aws_cloudwatch as cloudwatch,
aws_cloudwatch_actions as cw_actions,
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
aws_dynamodb as dynamodb,
feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC call-and-be-called effort). Everything ships DARK: both DynamoDB event source mappings deploy enabled=False; activation is a deliberate one-line follow-up PR gated on the SHOC receiver passing the shared HMAC test vectors. - Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments (in-place update, RETAIN + logical IDs untouched; no existing consumers — verified live, neither table had a stream). - workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope -> HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard ordering (parallelization 1, bisect off, retry until 24h age, ReportBatchItemFailures); 429/5xx/timeout block the shard in order, other 4xx park to workorder-shoc-emitter-rejected; ESM failures -> workorder-shoc-emitter-failures (metadata; replay rebuilds from DynamoDB). Echo guard skips write_origin=shoc-write-api. - Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK (alias workorder-ingest-shoc-webhook-kms); cross-account GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy DESTROY deliberately (machine-generated material; avoids the fixed-name RETAIN-orphan deadlock). - workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap, 64-hex keys, kid = UTC %Y-%m-%dT%H. - Alarms (ALARM-only -> site-alerts): emitter errors/throttles/ duration + iterator-age (>=10 min) + failures/rejected queue depth; rotator standard trio. - scripts/replay_shoc_webhooks.py: dry-run-default operator replay (rebuilds from tables, replay:true envelopes). - Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay signing pinned to identical vectors); bundle-consistency AST pins for both new bundles. - README: WO stack + webhook feed section, alarm table, runbooks; removed stale seahaven-slack-bot consumer references.
2026-07-24 14:54:07 -04:00
aws_iam as iam,
aws_kms as kms,
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
aws_lambda as lambda_,
feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC call-and-be-called effort). Everything ships DARK: both DynamoDB event source mappings deploy enabled=False; activation is a deliberate one-line follow-up PR gated on the SHOC receiver passing the shared HMAC test vectors. - Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments (in-place update, RETAIN + logical IDs untouched; no existing consumers — verified live, neither table had a stream). - workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope -> HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard ordering (parallelization 1, bisect off, retry until 24h age, ReportBatchItemFailures); 429/5xx/timeout block the shard in order, other 4xx park to workorder-shoc-emitter-rejected; ESM failures -> workorder-shoc-emitter-failures (metadata; replay rebuilds from DynamoDB). Echo guard skips write_origin=shoc-write-api. - Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK (alias workorder-ingest-shoc-webhook-kms); cross-account GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy DESTROY deliberately (machine-generated material; avoids the fixed-name RETAIN-orphan deadlock). - workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap, 64-hex keys, kid = UTC %Y-%m-%dT%H. - Alarms (ALARM-only -> site-alerts): emitter errors/throttles/ duration + iterator-age (>=10 min) + failures/rejected queue depth; rotator standard trio. - scripts/replay_shoc_webhooks.py: dry-run-default operator replay (rebuilds from tables, replay:true envelopes). - Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay signing pinned to identical vectors); bundle-consistency AST pins for both new bundles. - README: WO stack + webhook feed section, alarm table, runbooks; removed stale seahaven-slack-bot consumer references.
2026-07-24 14:54:07 -04:00
aws_lambda_event_sources as lambda_event_sources,
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
aws_s3 as s3,
aws_s3_notifications as s3n,
aws_ses as ses,
aws_ses_actions as ses_actions,
aws_secretsmanager as secretsmanager,
aws_sns as sns,
feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC call-and-be-called effort). Everything ships DARK: both DynamoDB event source mappings deploy enabled=False; activation is a deliberate one-line follow-up PR gated on the SHOC receiver passing the shared HMAC test vectors. - Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments (in-place update, RETAIN + logical IDs untouched; no existing consumers — verified live, neither table had a stream). - workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope -> HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard ordering (parallelization 1, bisect off, retry until 24h age, ReportBatchItemFailures); 429/5xx/timeout block the shard in order, other 4xx park to workorder-shoc-emitter-rejected; ESM failures -> workorder-shoc-emitter-failures (metadata; replay rebuilds from DynamoDB). Echo guard skips write_origin=shoc-write-api. - Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK (alias workorder-ingest-shoc-webhook-kms); cross-account GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy DESTROY deliberately (machine-generated material; avoids the fixed-name RETAIN-orphan deadlock). - workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap, 64-hex keys, kid = UTC %Y-%m-%dT%H. - Alarms (ALARM-only -> site-alerts): emitter errors/throttles/ duration + iterator-age (>=10 min) + failures/rejected queue depth; rotator standard trio. - scripts/replay_shoc_webhooks.py: dry-run-default operator replay (rebuilds from tables, replay:true envelopes). - Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay signing pinned to identical vectors); bundle-consistency AST pins for both new bundles. - README: WO stack + webhook feed section, alarm table, runbooks; removed stale seahaven-slack-bot consumer references.
2026-07-24 14:54:07 -04:00
aws_sqs as sqs,
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
)
from constructs import Construct
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112) The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
import common
Add fail-closed SES sender authentication (INFRA-107) (#98) * Add fail-closed SES sender authentication The From header and any raw-MIME Authentication-Results copies are attacker-forgeable, so a forged email to apm@int.seahaven.com or amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107, CRITICAL). Both S3-triggered email processors now authenticate the sender against the Authentication-Results header SES itself prepends at delivery: only the topmost header is consulted, its authserv-id must be amazonses.com, and it must carry dkim=pass for a domain in the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set in CDK so ops can adjust without code changes). Allowlists come from live traffic observed 2026-07-15 on both ingest buckets: WO mail arrives via the apm@ Google Groups forward, which re-signs as seahaven.com (the hxgnsmartcloud.com signature does not survive the forward); PO mail passes for amazon.coupahost.com. amazonses.com also passes on PO mail but is deliberately excluded -- every SES customer's outbound mail passes for it. Every failure path (env var unset, header missing or unparseable, verdict fail, unaligned domain) rejects the email: a structured warning with the reason and S3 key is logged and the record skipped without erroring the invocation, so rejected mail causes no Lambda retries or DLQ messages. Handler signatures and event sources are unchanged. Refs: INFRA-107 * Harden AR parser per cross-family review Cross-family (GPT-4.1) review findings: terminate the dkim result token at end-of-clause, whitespace, or a comment so a value like "dkim=pass-fake" can never be read as a pass; normalize trailing dots off allowlist entries so "seahaven.com." matches; make the compat32 parser policy explicit. Adds tests for result-token boundaries, comments after the result, quoted domain values, and folding inside a dkim clause. Refs: INFRA-107 * Harden AR parsing and alarm on sender-auth rejects The SES-stamped Authentication-Results value echoes attacker-controlled SMTP-session tokens (envelope-from, helo, header.from) as their own semicolon-delimited property clauses. A naive split(";") tore an RFC 5321 quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass clause, so a fully spoofed email was accepted on the genuinely SES-stamped topmost header. Tokenise comment- and quoted-string-aware (RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on semicolons outside a quoted-string, and fail closed on unbalanced quotes/comments so a ';' inside a quoted pvalue can never start a clause. Rejected mail returns normally (no error, no retry, no DLQ message), so a signing-domain drift or a wrong allowlist would silently discard 100% of legitimate mail while every alarm stayed green. Add a CloudWatch Logs metric filter + alarm on the sender_auth_rejected warning to both stacks so a false-reject storm pages instead of vanishing. This is also the safety net for the WO seahaven.com allowlist assumption, which must be validated against a live SES-stamped header (a plain Gmail auto-forward re-signs under the sending Workspace domain, not seahaven.com). Refs: INFRA-107 * chore: retrigger CI (no run recorded for 7c74ac1) * Fix quoted-AUID DKIM domain spoof in sender auth Resolve three confirmed /sh-security-review findings on the fail-closed SES sender-authentication control. HIGH: header.i/header.d domain extraction was not quoted-string aware. An attacker with a valid DKIM key for their own domain could set an RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive extractor stopped at the closing quote and returned seahaven.com, accepting forged mail. Extraction now tokenises the clause with the same quoted-string discipline already used for clause splitting: header.d (the plain signing domain) is authoritative when present, otherwise the header.i domain is the part after the AUID's LAST top-level "@", so a "@" inside a quoted local-part is treated as signer-controlled label text and yields the true signer (attacker.com), not seahaven.com. LOW: the topmost-header parse ran outside evaluate_sender_authentication's try/except, so an unexpected parser exception on crafted input could propagate into the handler and Lambda async retries/DLQ. The parse now fails CLOSED with an authentication_results_unparseable reason. MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to a low-volume total-reject outage (a trickle that never sums to 3). Both stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3 / datapoints_to_alarm=2, so a sustained reject condition pages even at one reject per period while a lone stray probe self-clears. Refs: INFRA-107 * Load Lambda function dir on sys.path in tests Rebasing INFRA-107 onto main folded #95's pytest suite into this branch's tests. The unified conftest loads the PO/WO handlers by file path, and handler.py now does `from ses_auth import authenticate_inbound_email` -- a bare sibling import that resolves in the Lambda only because the runtime puts each function's own directory on sys.path. The shared load_handler now adds that directory so the handler tests import correctly alongside the sender-auth tests. Refs: INFRA-107 * Note #97 test files in README directory tree The rebase onto main brought in #97's tests/requirements.txt and tests/test_po_merge.py. List both in the directory tree so it matches the tree on disk. Refs: INFRA-107 * Document INFRA-107 forwarder-binding risk acceptance Record the accepted risk that WO sender auth binds to the apm@ forward's re-signing domain (seahaven.com) rather than the Hexagon originator; the apm@ Google Group's restricted posting policy is the load-bearing control (escalates to HIGH if the group is opened to external posting). Also correct the sender-auth-rejected alarm docs to match the shipped config (>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and note the SES-AR-01/02 parser hardening follow-ups. Refs: INFRA-107
2026-07-15 20:58:47 -04:00
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
class WorkorderIngestStack(Stack):
def __init__(self, scope: Construct, construct_id: str, **kwargs):
super().__init__(scope, construct_id, **kwargs)
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70) * Add CloudWatch alarm coverage for po-ingest and workorder-ingest Expands alarm coverage across both CDK stacks. All alarms are ALARM-only (no OK action) to the shared site-alerts SNS topic, with TreatMissingData NOT_BREACHING. The site-alerts topic is now imported once near the top of each stack so every alarm reuses one Topic instance. po-ingest (cdk/po_stack.py): - Errors: po-ingest-site-extractor - Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui - Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption), po-ingest-site-extractor, po-web-ui - DynamoDB throttle + system-error: purchase-orders, verified-sites, pending-site-review workorder-ingest (cdk/wo_stack.py): - Throttles: workorder-email-processor - Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption) - DynamoDB throttle + system-error: WorkOrders, WorkOrderComments DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation dimension set, so each table alarm is a Sum math expression across operations via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.259.0). Refs INFRA-41 / audit H-8. * Drop NEEDS ADAM SIGN-OFF wording from alarm comments Duration alarm thresholds are owner-approved; remove the sign-off flag from po_stack.py and wo_stack.py comments. Threshold values, eval config, and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
# --- Shared alarm SNS topic (site-alerts) ---
# Imported once near the top so every alarm in this stack reuses the same
# Topic construct instance (avoids duplicate logical IDs). ALARM-only
# SnsAction; no OK action, per the CloudWatch-alarm preference. The
# topic's CMK (alias/seahaven-alarm-topics) lives on the topic itself.
alarm_topic = sns.Topic.from_topic_arn(
self,
"SiteAlertsTopic",
f"arn:aws:sns:{self.region}:{self.account}:site-alerts",
)
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
# --- S3 bucket for raw emails ---
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112) The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
email_bucket = common.make_email_bucket(
self, "EmailBucket", "workorder-ingest-emails"
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
)
# --- DynamoDB tables ---
work_orders_table = dynamodb.Table(
self,
"WorkOrdersTable",
table_name="WorkOrders",
partition_key=dynamodb.Attribute(
name="work_order_id",
type=dynamodb.AttributeType.STRING,
),
billing_mode=dynamodb.BillingMode.PAY_PER_REQUEST,
feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC call-and-be-called effort). Everything ships DARK: both DynamoDB event source mappings deploy enabled=False; activation is a deliberate one-line follow-up PR gated on the SHOC receiver passing the shared HMAC test vectors. - Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments (in-place update, RETAIN + logical IDs untouched; no existing consumers — verified live, neither table had a stream). - workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope -> HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard ordering (parallelization 1, bisect off, retry until 24h age, ReportBatchItemFailures); 429/5xx/timeout block the shard in order, other 4xx park to workorder-shoc-emitter-rejected; ESM failures -> workorder-shoc-emitter-failures (metadata; replay rebuilds from DynamoDB). Echo guard skips write_origin=shoc-write-api. - Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK (alias workorder-ingest-shoc-webhook-kms); cross-account GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy DESTROY deliberately (machine-generated material; avoids the fixed-name RETAIN-orphan deadlock). - workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap, 64-hex keys, kid = UTC %Y-%m-%dT%H. - Alarms (ALARM-only -> site-alerts): emitter errors/throttles/ duration + iterator-age (>=10 min) + failures/rejected queue depth; rotator standard trio. - scripts/replay_shoc_webhooks.py: dry-run-default operator replay (rebuilds from tables, replay:true envelopes). - Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay signing pinned to identical vectors); bundle-consistency AST pins for both new bundles. - README: WO stack + webhook feed section, alarm table, runbooks; removed stale seahaven-slack-bot consumer references.
2026-07-24 14:54:07 -04:00
# NEW_AND_OLD_IMAGES: the SHOC emitter needs OLD.wo_status to
# classify the cancelled transition (docs/shoc-webhook-plan.md
# Phase 2). In-place CFN update -- no table replacement.
stream=dynamodb.StreamViewType.NEW_AND_OLD_IMAGES,
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
removal_policy=RemovalPolicy.RETAIN,
)
# site-code-index and status-index GSIs removed 2026-06-03 (audit M-20):
# 0 reads in 30d against ~50k WCU each of write amplification. Re-add if
# a site-code or status query path ships.
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
comments_table = dynamodb.Table(
self,
"CommentsTable",
table_name="WorkOrderComments",
partition_key=dynamodb.Attribute(
name="work_order_id",
type=dynamodb.AttributeType.STRING,
),
sort_key=dynamodb.Attribute(
name="comment_id",
type=dynamodb.AttributeType.STRING,
),
billing_mode=dynamodb.BillingMode.PAY_PER_REQUEST,
feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC call-and-be-called effort). Everything ships DARK: both DynamoDB event source mappings deploy enabled=False; activation is a deliberate one-line follow-up PR gated on the SHOC receiver passing the shared HMAC test vectors. - Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments (in-place update, RETAIN + logical IDs untouched; no existing consumers — verified live, neither table had a stream). - workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope -> HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard ordering (parallelization 1, bisect off, retry until 24h age, ReportBatchItemFailures); 429/5xx/timeout block the shard in order, other 4xx park to workorder-shoc-emitter-rejected; ESM failures -> workorder-shoc-emitter-failures (metadata; replay rebuilds from DynamoDB). Echo guard skips write_origin=shoc-write-api. - Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK (alias workorder-ingest-shoc-webhook-kms); cross-account GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy DESTROY deliberately (machine-generated material; avoids the fixed-name RETAIN-orphan deadlock). - workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap, 64-hex keys, kid = UTC %Y-%m-%dT%H. - Alarms (ALARM-only -> site-alerts): emitter errors/throttles/ duration + iterator-age (>=10 min) + failures/rejected queue depth; rotator standard trio. - scripts/replay_shoc_webhooks.py: dry-run-default operator replay (rebuilds from tables, replay:true envelopes). - Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay signing pinned to identical vectors); bundle-consistency AST pins for both new bundles. - README: WO stack + webhook feed section, alarm table, runbooks; removed stale seahaven-slack-bot consumer references.
2026-07-24 14:54:07 -04:00
# Streamed for the SHOC emitter (docs/shoc-webhook-plan.md Phase 2).
stream=dynamodb.StreamViewType.NEW_AND_OLD_IMAGES,
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
removal_policy=RemovalPolicy.RETAIN,
)
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
# --- Anthropic API key secret removed (Bedrock migration) ---
# Parsing moved from the Anthropic API to the Bedrock inference profile
# us.anthropic.claude-haiku-4-5-20251001-v1:0, so no provider API key is
# needed. The old secret "workorder-ingest/anthropic-api-key" had
# RemovalPolicy.RETAIN, so it is ORPHANED (not deleted) by this change:
# delete it manually post-deploy and revoke the stored key at Anthropic.
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
# --- DLQ for failed async invocations (INFRA-41 / audit H-8) ---
# SES → S3 → Lambda is async; without an OnFailure destination a failed
# parse (bad email, transient error) is silently dropped after Lambda's
# retries. CDK generates the queue name to avoid colliding with the
# interim CLI-created workorder-email-processor-dlq (removed post-deploy).
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112) The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
email_processor_dlq = common.make_processor_dlq(self, "EmailProcessorDlq")
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
# --- Lambda function ---
email_processor_log_group = common.make_function_log_group(
self, "EmailProcessor", "workorder-email-processor"
)
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
email_processor = lambda_.Function(
self,
"EmailProcessor",
function_name="workorder-email-processor",
runtime=lambda_.Runtime.PYTHON_3_12,
architecture=lambda_.Architecture.ARM_64,
handler="handler.handler",
code=lambda_.Code.from_asset(
feat: widen email-processor asset roots to lambdas/ with scoped globs + excludes (refactor phase 2) (#109) Both email-processor Code.from_asset calls now bundle from lambdas/ instead of their per-function subdirectory, so Phase 3's shared/ module is reachable from the asset root once it lands. The bundling commands were rewritten for the new cwd (pip install -r <po|wo>/ email_processor/requirements.txt -t /asset-output && cp <po|wo>/ email_processor/*.py /asset-output/), preserving the ARM64 --platform manylinux2014_aarch64 --only-binary=:all: pin exactly — its removal shipped x86 wheels into the ARM64 function and caused a 100% outage (PR #34). All five from_asset calls (both email processors, po web_ui, po site_extractor, wo web_ui) now exclude **/__pycache__/**; the two widened ones also exclude **/tests/** and **/package/**. Without the package/ exclude, the stale untracked 44 MB lambdas/po/email_processor/package/ dir (local-only, never present in CI) would diverge local vs CI asset hashes and force spurious redeploys — from_asset doesn't honor .gitignore. That dir is left in place; deleting it is Adam's call. WO's prod zip shrinks as deliberate cleanup, not a byte-identical match to PO: the old `cp -r .` shipped tests/ (real scrubbed .eml fixtures), __pycache__/, and requirements.txt into production. The acceptance bar for WO is runtime-imported module set unchanged + smoke, not a byte-identical zip; PO keeps the byte-identical first-party file set guarantee. tests/test_bundle_consistency.py is updated in the same change to recognize the scoped `cp po/email_processor/*.py` (resp. wo) glob as the new unconditionally-safe shape, without loosening the allowlist-revert detection, the detection-logic mutation test, or the PO_EXPECTED_TOP_LEVEL_MODULES exact-set pin. No code moved under lambdas/ in this change (git diff main...HEAD -- lambdas/ is empty); only CDK asset wiring and its tests changed.
2026-07-17 15:47:01 -04:00
"../lambdas",
exclude=["**/__pycache__/**", "**/tests/**", "**/package/**"],
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
bundling=cdk.BundlingOptions(
image=lambda_.Runtime.PYTHON_3_12.bundling_image,
command=[
"bash",
"-c",
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110) * feat: ops/recovery tooling + dependency hygiene (refactor phase 7) Generalize scripts/reprocess.py from a PO-only full-sweep script into a pipeline-general recovery tool. Targeted replay (--key/--prefix/--since) is now the default, and the full inbound/ sweep is demoted behind an explicit --all that documents its five hazards (async concurrency does not serialize, use RequestResponse if order matters, metric double-count, Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves the correct function + bucket; dry-run-by-default / --execute is preserved. A new tests/test_reprocess_contract.py pins the synthetic S3 event shape and asserts the raw list_objects_v2 key is emitted untransformed (the handler is the single decode point; a pre-decoded key would corrupt keys containing spaces or '+'). Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console redrive-to-source, so it documents the receive -> extract key -> targeted reprocess --key -> verify -> purge procedure, the real recovery windows (14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table RETAIN policy and is the true replay floor), and that sender-auth and ai_fallback_rejected drops are fail-closed skips that never reach the DLQ. Linked from the README alarms and scripts sections. Drop the vendored boto3 floor pin from both email-processor requirements (the Lambda runtime provides boto3; lambda-template.md empty-with-comment form). With nothing left to install, the email-processor bundling becomes cp-only -- the whole pip step is removed, which is the only acceptable way the manylinux2014_aarch64 pin disappears (removing the pin while keeping a pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and add pinned po/web_ui + po/site_extractor manifests (excluded from their bundles, so hash-neutral) so their new Dependabot entries have something to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and /lambdas/po/site_extractor. cdk diff is confined to exactly the two email processors' asset hashes on both stacks. The wo/web_ui dead-manifest reduction was deliberately left out: that manifest already ships inside the plain (non-bundled) WebUI asset on main, so reducing or excluding it would redeploy workorder-web-ui for no functional change -- deferred to keep the blast radius to the two intended targets. The untracked 44 MB lambdas/po/email_processor/package/ dir was removed from the filesystem (asset-hash-neutral given Phase 2's package/ exclude); it is untracked, so there is nothing to commit for it. * Reject --all combined with --prefix/--since in reprocess.py --all is a distinct mode (the demoted full-prefix sweep), but the args.all branch unconditionally set prefix=inbound/ and since=None, so passing it alongside a narrower selector silently discarded that selector. `--all --since 2026-07-01` swept the entire corpus instead of the bounded window, triggering every documented --all hazard (Bedrock re-bill, metric double- count, merged-field regression) on objects the operator never targeted -- contradicting the tool's safety goal. Add the missing mutual-exclusion guard alongside the existing --key one, and pin --all+--prefix, --all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
# pip step removed in Phase 7: requirements.txt is now empty
# (boto3 comes from the Lambda runtime), so nothing is installed
# and the manylinux pin has nothing to pin. cp-only is safe.
feat: extract lambdas/shared/ — single-source ses_auth, web_ui auth, email parsing, EMF emitter (refactor phase 3) (#111) Four modules move into the handbook-mandated lambdas/shared/ location, collapsing duplicated logic that had to be kept in sync by hand across the PO and WO pipelines: - ses_auth.py: the PO and WO copies were verified sha256-identical against the feature/phase-7-ops-recovery baseline before the move (no drift since the last audit). shared/ses_auth.py is the exact bytes of that one copy; both originals are git rm'd (the PO copy via rename, the WO copy as a straight delete). Bundling lands the module flat in /asset-output for both email processors, so the handlers keep `from ses_auth import authenticate_inbound_email` unchanged — zero handler diff for this move, which is what keeps fail-closed auth byte-identical through the change. - web_ui_auth.py: extracts the byte-identical _get_auth_token / _header / is_authenticated block plus the four token-cache globals out of both web_ui handlers. The per-stack INFRA-74 comments stay in each handler as-is (deliberately drifted wording, stack-specific) rather than being unified into the shared module. Fail-closed semantics (unset ARN or Secrets Manager exception -> deny) are unchanged. - email_parsing.py: parse_raw_email ships as the superset version that returns cc unconditionally. WO's output is bit-identical to before; PO simply ignores the cc field rather than being "cleaned up" to consume it. No second variant is kept. - emf.py: a generic emitter parameterized by namespace, dimension sets, and properties. Every call site's emitted EMF envelope is unchanged, including the load-bearing [["ParseMethod"],["ParseMethod","TemplateId"]] dimension-set shape the alarms and metric filters depend on. Emission ordering is untouched: PO still emits ai_fallback before the Bedrock call, WO still emits its mutually-exclusive ai_fallback/ai_fallback_rejected after its gate. The deliberate-double-count comments survive. _emit_derived_agreement_metric was found living inside derived_fields.py, so per the DERIVED-FIELDS exception it is left as a third, unconverted copy (derived_fields.py and the shadow DerivedFieldAgreement telemetry stay untouchable while that bake runs) — a comment there points at shared/emf.py for the eventual follow-up. Bundling: both email-processor cdk bundling commands gain a trailing `cp shared/*.py /asset-output/` (they were already cp-only post-Phase 7, so no pip step or manylinux pin is reintroduced). Both web_ui functions gain the same widened-root staging so web_ui_auth.py ships beside their handler; site_extractor's from_asset is untouched. Tests: PO_EXPECTED_TOP_LEVEL_MODULES gains the shared modules that now ship, the AST sibling-import check resolves imports whose source now lives under shared/, and the new shared cp line has its own revert/mutation detection. _SIBLING_MODULES resolution and _po_parser_support.py now load ses_auth/email_parsing/emf from shared/; the two-copy ses_auth byte-identity fixture-hygiene test is retired as obsolete now that there is one copy, and the ses_auth fixture parameterization over two identical copies is dropped. The sys.modules save/restore dance for template_parser (still duplicated per-pipeline) is left in place.
2026-07-20 13:38:23 -04:00
# shared/*.py ships the four modules extracted to
# lambdas/shared/ (Phase 3): ses_auth, web_ui_auth,
# email_parsing, emf. Flat cp keeps the bare-name
# imports (e.g. `from ses_auth import ...`) resolving
# unchanged in /asset-output.
"cp wo/email_processor/*.py /asset-output/ && "
"cp shared/*.py /asset-output/",
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
],
),
),
timeout=Duration.seconds(60),
memory_size=256,
log_group=email_processor_log_group,
dead_letter_queue=email_processor_dlq,
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
environment={
"WORK_ORDERS_TABLE": work_orders_table.table_name,
"COMMENTS_TABLE": comments_table.table_name,
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
"BEDROCK_MODEL_ID": "us.anthropic.claude-haiku-4-5-20251001-v1:0",
Add fail-closed SES sender authentication (INFRA-107) (#98) * Add fail-closed SES sender authentication The From header and any raw-MIME Authentication-Results copies are attacker-forgeable, so a forged email to apm@int.seahaven.com or amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107, CRITICAL). Both S3-triggered email processors now authenticate the sender against the Authentication-Results header SES itself prepends at delivery: only the topmost header is consulted, its authserv-id must be amazonses.com, and it must carry dkim=pass for a domain in the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set in CDK so ops can adjust without code changes). Allowlists come from live traffic observed 2026-07-15 on both ingest buckets: WO mail arrives via the apm@ Google Groups forward, which re-signs as seahaven.com (the hxgnsmartcloud.com signature does not survive the forward); PO mail passes for amazon.coupahost.com. amazonses.com also passes on PO mail but is deliberately excluded -- every SES customer's outbound mail passes for it. Every failure path (env var unset, header missing or unparseable, verdict fail, unaligned domain) rejects the email: a structured warning with the reason and S3 key is logged and the record skipped without erroring the invocation, so rejected mail causes no Lambda retries or DLQ messages. Handler signatures and event sources are unchanged. Refs: INFRA-107 * Harden AR parser per cross-family review Cross-family (GPT-4.1) review findings: terminate the dkim result token at end-of-clause, whitespace, or a comment so a value like "dkim=pass-fake" can never be read as a pass; normalize trailing dots off allowlist entries so "seahaven.com." matches; make the compat32 parser policy explicit. Adds tests for result-token boundaries, comments after the result, quoted domain values, and folding inside a dkim clause. Refs: INFRA-107 * Harden AR parsing and alarm on sender-auth rejects The SES-stamped Authentication-Results value echoes attacker-controlled SMTP-session tokens (envelope-from, helo, header.from) as their own semicolon-delimited property clauses. A naive split(";") tore an RFC 5321 quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass clause, so a fully spoofed email was accepted on the genuinely SES-stamped topmost header. Tokenise comment- and quoted-string-aware (RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on semicolons outside a quoted-string, and fail closed on unbalanced quotes/comments so a ';' inside a quoted pvalue can never start a clause. Rejected mail returns normally (no error, no retry, no DLQ message), so a signing-domain drift or a wrong allowlist would silently discard 100% of legitimate mail while every alarm stayed green. Add a CloudWatch Logs metric filter + alarm on the sender_auth_rejected warning to both stacks so a false-reject storm pages instead of vanishing. This is also the safety net for the WO seahaven.com allowlist assumption, which must be validated against a live SES-stamped header (a plain Gmail auto-forward re-signs under the sending Workspace domain, not seahaven.com). Refs: INFRA-107 * chore: retrigger CI (no run recorded for 7c74ac1) * Fix quoted-AUID DKIM domain spoof in sender auth Resolve three confirmed /sh-security-review findings on the fail-closed SES sender-authentication control. HIGH: header.i/header.d domain extraction was not quoted-string aware. An attacker with a valid DKIM key for their own domain could set an RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive extractor stopped at the closing quote and returned seahaven.com, accepting forged mail. Extraction now tokenises the clause with the same quoted-string discipline already used for clause splitting: header.d (the plain signing domain) is authoritative when present, otherwise the header.i domain is the part after the AUID's LAST top-level "@", so a "@" inside a quoted local-part is treated as signer-controlled label text and yields the true signer (attacker.com), not seahaven.com. LOW: the topmost-header parse ran outside evaluate_sender_authentication's try/except, so an unexpected parser exception on crafted input could propagate into the handler and Lambda async retries/DLQ. The parse now fails CLOSED with an authentication_results_unparseable reason. MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to a low-volume total-reject outage (a trickle that never sums to 3). Both stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3 / datapoints_to_alarm=2, so a sustained reject condition pages even at one reject per period while a lone stray probe self-clears. Refs: INFRA-107 * Load Lambda function dir on sys.path in tests Rebasing INFRA-107 onto main folded #95's pytest suite into this branch's tests. The unified conftest loads the PO/WO handlers by file path, and handler.py now does `from ses_auth import authenticate_inbound_email` -- a bare sibling import that resolves in the Lambda only because the runtime puts each function's own directory on sys.path. The shared load_handler now adds that directory so the handler tests import correctly alongside the sender-auth tests. Refs: INFRA-107 * Note #97 test files in README directory tree The rebase onto main brought in #97's tests/requirements.txt and tests/test_po_merge.py. List both in the directory tree so it matches the tree on disk. Refs: INFRA-107 * Document INFRA-107 forwarder-binding risk acceptance Record the accepted risk that WO sender auth binds to the apm@ forward's re-signing domain (seahaven.com) rather than the Hexagon originator; the apm@ Google Group's restricted posting policy is the load-bearing control (escalates to HIGH if the group is opened to external posting). Also correct the sender-auth-rejected alarm docs to match the shipped config (>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and note the SES-AR-01/02 parser hardening follow-ups. Refs: INFRA-107
2026-07-15 20:58:47 -04:00
# Fail-closed sender auth (INFRA-107): the handler only
# accepts mail whose SES-stamped Authentication-Results
# header carries dkim=pass for one of these domains. APM
# mail arrives via the apm@ Google Groups forward, which
# re-signs as seahaven.com (observed on live traffic
# 2026-07-15: "dkim=pass header.i=@seahaven.com"; the
# original hxgnsmartcloud.com signature does not survive
# the forward). Unset/empty ⇒ the handler rejects all mail.
"ALLOWED_DKIM_DOMAINS": "seahaven.com",
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
},
)
# Grant permissions
email_bucket.grant_read(email_processor)
work_orders_table.grant_read_write_data(email_processor)
comments_table.grant_read_write_data(email_processor)
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
# --- Bedrock InvokeModel grant ---
# The us.* inference profile can route cross-region, so the grant MUST
# cover both the inference-profile ARN AND the per-region foundation-model
# ARNs (empty account field) for every region the profile can reach
# (us-east-1/us-east-2/us-west-2). A profile-only grant AccessDenies at
# runtime whenever the profile routes to a region whose foundation-model
# ARN is not allowed.
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112) The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
email_processor.add_to_role_policy(common.make_bedrock_invoke_statement(self))
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
Land safe fixes from 2026-06-17 security sweep (#97) * Remove gratuitous KMS grant on shared DynamoDB CMK wo-email-processor held grant_encrypt_decrypt on the shared seahaven-dynamodb CMK, but the WorkOrders/WorkOrderComments tables are not encrypted with that CMK. The grant was dead weight that extended the WO processor's decrypt reach to the CMK protecting the purchase-orders table (cross-stack decrypt). Drop it to restore least privilege; re-add as part of the table CMK migration (INFRA-6). Refs: INFRA-6 * Require Secrets Manager key for Anthropic client Remove the silent fallback to a plaintext ANTHROPIC_API_KEY env var in both email processors; require ANTHROPIC_API_KEY_SECRET_ARN and raise if absent so a misconfigured deploy fails loudly instead of using an unmanaged key. Adapted from f175323 on security/sweep-2026-06-17. The From-header sender-domain allowlist from that commit is intentionally dropped: the From header is spoofable (INFRA-107, confirmed critical) and sender authentication is being reworked in a separate PR. Refs: INFRA-107 * Merge PO revisions and handle out-of-order events save_revision did a full put_item overwrite, so a revision omitting line_items/supplier permanently deleted them. save_new_po used a conditional put that silently dropped the PO when an out-of-order cancellation had already created a skeleton row. Switch both to field-level merge update_items: a revision now SETs only the fields it carries, and a new_po backfills data into a pre-existing Cancelled skeleton while preserving the Cancelled status. No email can now delete data established by an earlier one. * Gate web UIs behind auth and escape currency XSS The po-web-ui and workorder-web-ui handlers had no auth: any invocation path returned the full PO/WO DB. Add a fail-closed shared-secret gate (X-Auth-Token / Bearer, constant-time compared to WEB_UI_AUTH_TOKEN) so a future re-attached Function URL cannot re-expose the data (URLs removed under INFRA-74). Wire the token from the SSM String param /procurement-ingest/web-ui-auth-token. Also fix stored XSS in po-web-ui fmt_currency: the non-numeric fallback returned str(val) unescaped, so a prompt-injected email could make Claude emit total_amount as <script>. Escape it. Refs: INFRA-74 * Document sweep security fixes and merge semantics Update the README for the 2026-06-17 security sweep: required Secrets Manager key (no plaintext env fallback), web UI auth gate + SSM token setup step, output-escaping note, and the new PO revision/cancellation merge behavior. Adapted from d91f45e on security/sweep-2026-06-17; the sender allowlist documentation is dropped along with the allowlist itself (deferred to the INFRA-107 sender-authentication rework). Refs: INFRA-107 * fix: resolve web UI auth token from Secrets Manager at runtime Replace the plaintext SSM String parameter with a Secrets Manager secret referenced by ARN only. The token is fetched and cached at module level on first invocation, keeping shared secrets out of CloudFormation templates and Lambda environment variables. Refs: PR-97 * Add TTL to web UI auth token cache for rotation The web-ui handlers cached the Secrets Manager auth token at module level with no expiry, so a rotated secret was only picked up when the warm container recycled — an emergency rotation could take hours to take effect. Cache the fetched value for a 5-minute TTL instead, so a rotated token propagates within the TTL while still avoiding a Secrets Manager call on every request. Still fails closed when the secret is unset or unreadable. Refs: INFRA-74 * Log Secrets Manager failures in web UI auth token fetch The web UI auth gate correctly fails closed when the shared token cannot be read, but _get_auth_token() swallowed every exception silently. A Secrets Manager permission or config error then made every request 401 with no operational signal, leaving an outage indistinguishable from ordinary unauthenticated traffic. Add a module-level logger to both web_ui handlers and log the fetch failure with logger.exception() in the except block before returning None. Behavior is unchanged (still fails closed); the failure is now visible in CloudWatch. The secret value is never logged. The two handlers stay byte-consistent in the mirrored _get_auth_token() region. The companion finding on the CDK import of the shared procurement-ingest/web-ui-auth-token secret was evaluated and left as-is: the token is a single secret shared by both the PO and WO stacks, so from_secret_name_v2 (which scopes grant_read via the standard 6-char suffix wildcard) is correct; making it a CDK-managed Secret in both stacks would collide the two stacks on the same explicit secret name at deploy time. Refs: INFRA-74 * Make Cancelled PO status sticky via atomic write The PO merge path read status with a get_item (_is_cancelled) and then wrote with an unconditional update_item. Two defects followed from this: - Race (Issue A): a cancellation landing between the read and the write was silently un-cancelled by a revision carrying a non-cancelled po_status — a TOCTOU on a table with concurrent email processing. - Over-broad strip (Issue B): save_revision dropped po_status whenever the PO was Cancelled, so legitimate status updates on non-cancelled POs and status-less revisions were affected rather than only the true un-cancel transition. Enforce the invariant server-side instead. "Cancelled" is a sticky, authoritative status: once set, later new_po/revision emails may enrich other fields but must never move it to a non-cancelled status. When the payload carries a non-cancelled po_status, _merge_update issues the update_item guarded by ConditionExpression "attribute_not_exists(po_status) OR po_status <> :marker", evaluated atomically at write time, so a cancellation that lands first always wins. On ConditionalCheckFailedException the same fields are re-written without po_status/cancelled_at, enriching the record while Cancelled sticks. Payloads with no status change, or an already -Cancelled status, take a plain merge — the status is only ever suppressed on a real un-cancel. This removes the non-atomic get_item from the write path; _is_cancelled is deleted. Key schema and attribute names are unchanged, so the cross-stack purchase-orders contract (read-only by seahaven-slack-bot) holds. Add moto-backed tests covering un-cancel suppression with field enrichment, status-less merge onto a Cancelled PO, legitimate status updates on non-cancelled POs, new_po backfill of a Cancelled skeleton, fresh create/merge, and authoritative save_cancellation. Refs: #97
2026-07-15 20:17:46 -04:00
# NOTE: The pre-emptive grant_encrypt_decrypt on the shared DynamoDB CMK
# (alias/seahaven-dynamodb) was removed (security sweep 2026-06-17). The
# WorkOrders/WorkOrderComments tables are NOT SSE-KMS encrypted with that
# CMK, so the grant was unused for these tables yet handed
# wo-email-processor kms:Decrypt on the CMK that also protects the
# purchase-orders table (cross-stack decrypt reach). Re-add this grant only
# as part of the actual CMK migration of these tables (INFRA-6), at which
# point grant_read_write_data on the (then encrypted) tables would propagate
# the needed key permissions automatically.
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112) The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
# --- Standard per-Lambda alarms: workorder-email-processor ---
# errors (INFRA-41 / audit H-8), throttles, DLQ-visible-messages (dropped
# emails), and a p95 duration alarm (orphan adoption of the CLI
# Lambda-Duration-workorder-email-processor under <fn>-duration naming,
# 45000 ms = 75% of the 60s timeout, eval 3 / dp 2). p95 (NOT p99) is the
# WO-specific duration statistic. All ALARM-only to site-alerts.
common.add_standard_lambda_alarms(
self,
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112) The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
"EmailProcessor",
email_processor,
"workorder-email-processor",
alarm_topic,
duration_statistic="p95",
errors=True,
dlq=email_processor_dlq,
descriptions={
"errors": "workorder-email-processor async invocation errors",
"throttles": "workorder-email-processor invocation throttles",
"dlq": "workorder-email-processor DLQ has visible messages (dropped emails)",
"duration": "workorder-email-processor p95 duration approaching the 60s timeout",
},
)
Add fail-closed SES sender authentication (INFRA-107) (#98) * Add fail-closed SES sender authentication The From header and any raw-MIME Authentication-Results copies are attacker-forgeable, so a forged email to apm@int.seahaven.com or amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107, CRITICAL). Both S3-triggered email processors now authenticate the sender against the Authentication-Results header SES itself prepends at delivery: only the topmost header is consulted, its authserv-id must be amazonses.com, and it must carry dkim=pass for a domain in the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set in CDK so ops can adjust without code changes). Allowlists come from live traffic observed 2026-07-15 on both ingest buckets: WO mail arrives via the apm@ Google Groups forward, which re-signs as seahaven.com (the hxgnsmartcloud.com signature does not survive the forward); PO mail passes for amazon.coupahost.com. amazonses.com also passes on PO mail but is deliberately excluded -- every SES customer's outbound mail passes for it. Every failure path (env var unset, header missing or unparseable, verdict fail, unaligned domain) rejects the email: a structured warning with the reason and S3 key is logged and the record skipped without erroring the invocation, so rejected mail causes no Lambda retries or DLQ messages. Handler signatures and event sources are unchanged. Refs: INFRA-107 * Harden AR parser per cross-family review Cross-family (GPT-4.1) review findings: terminate the dkim result token at end-of-clause, whitespace, or a comment so a value like "dkim=pass-fake" can never be read as a pass; normalize trailing dots off allowlist entries so "seahaven.com." matches; make the compat32 parser policy explicit. Adds tests for result-token boundaries, comments after the result, quoted domain values, and folding inside a dkim clause. Refs: INFRA-107 * Harden AR parsing and alarm on sender-auth rejects The SES-stamped Authentication-Results value echoes attacker-controlled SMTP-session tokens (envelope-from, helo, header.from) as their own semicolon-delimited property clauses. A naive split(";") tore an RFC 5321 quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass clause, so a fully spoofed email was accepted on the genuinely SES-stamped topmost header. Tokenise comment- and quoted-string-aware (RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on semicolons outside a quoted-string, and fail closed on unbalanced quotes/comments so a ';' inside a quoted pvalue can never start a clause. Rejected mail returns normally (no error, no retry, no DLQ message), so a signing-domain drift or a wrong allowlist would silently discard 100% of legitimate mail while every alarm stayed green. Add a CloudWatch Logs metric filter + alarm on the sender_auth_rejected warning to both stacks so a false-reject storm pages instead of vanishing. This is also the safety net for the WO seahaven.com allowlist assumption, which must be validated against a live SES-stamped header (a plain Gmail auto-forward re-signs under the sending Workspace domain, not seahaven.com). Refs: INFRA-107 * chore: retrigger CI (no run recorded for 7c74ac1) * Fix quoted-AUID DKIM domain spoof in sender auth Resolve three confirmed /sh-security-review findings on the fail-closed SES sender-authentication control. HIGH: header.i/header.d domain extraction was not quoted-string aware. An attacker with a valid DKIM key for their own domain could set an RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive extractor stopped at the closing quote and returned seahaven.com, accepting forged mail. Extraction now tokenises the clause with the same quoted-string discipline already used for clause splitting: header.d (the plain signing domain) is authoritative when present, otherwise the header.i domain is the part after the AUID's LAST top-level "@", so a "@" inside a quoted local-part is treated as signer-controlled label text and yields the true signer (attacker.com), not seahaven.com. LOW: the topmost-header parse ran outside evaluate_sender_authentication's try/except, so an unexpected parser exception on crafted input could propagate into the handler and Lambda async retries/DLQ. The parse now fails CLOSED with an authentication_results_unparseable reason. MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to a low-volume total-reject outage (a trickle that never sums to 3). Both stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3 / datapoints_to_alarm=2, so a sustained reject condition pages even at one reject per period while a lone stray probe self-clears. Refs: INFRA-107 * Load Lambda function dir on sys.path in tests Rebasing INFRA-107 onto main folded #95's pytest suite into this branch's tests. The unified conftest loads the PO/WO handlers by file path, and handler.py now does `from ses_auth import authenticate_inbound_email` -- a bare sibling import that resolves in the Lambda only because the runtime puts each function's own directory on sys.path. The shared load_handler now adds that directory so the handler tests import correctly alongside the sender-auth tests. Refs: INFRA-107 * Note #97 test files in README directory tree The rebase onto main brought in #97's tests/requirements.txt and tests/test_po_merge.py. List both in the directory tree so it matches the tree on disk. Refs: INFRA-107 * Document INFRA-107 forwarder-binding risk acceptance Record the accepted risk that WO sender auth binds to the apm@ forward's re-signing domain (seahaven.com) rather than the Hexagon originator; the apm@ Google Group's restricted posting policy is the load-bearing control (escalates to HIGH if the group is opened to external posting). Also correct the sender-auth-rejected alarm docs to match the shipped config (>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and note the SES-AR-01/02 parser hardening follow-ups. Refs: INFRA-107
2026-07-15 20:58:47 -04:00
# --- Sender-auth rejection alarm (INFRA-107) ---
# A rejected email (bad/unaligned DKIM verdict) returns normally, so it
# produces NO Lambda error, NO DLQ message and NO retry -- only a
# `sender_auth_rejected` warning log. The WO allowlist trusts dkim=pass
# for seahaven.com on the assumption the apm@ forward re-signs there; if
# that assumption is wrong (e.g. a Gmail auto-forward re-signs under a
# different domain), 100% of legitimate work-order mail is silently
# dropped. This metric filter + alarm turns those warnings into a paging
# signal so a false-reject storm surfaces instead of a silent outage.
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112) The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
common.add_sender_auth_rejected_alarm(
self,
"EmailProcessor",
"workorder-email-processor",
alarm_topic,
email_processor_log_group,
Add fail-closed SES sender authentication (INFRA-107) (#98) * Add fail-closed SES sender authentication The From header and any raw-MIME Authentication-Results copies are attacker-forgeable, so a forged email to apm@int.seahaven.com or amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107, CRITICAL). Both S3-triggered email processors now authenticate the sender against the Authentication-Results header SES itself prepends at delivery: only the topmost header is consulted, its authserv-id must be amazonses.com, and it must carry dkim=pass for a domain in the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set in CDK so ops can adjust without code changes). Allowlists come from live traffic observed 2026-07-15 on both ingest buckets: WO mail arrives via the apm@ Google Groups forward, which re-signs as seahaven.com (the hxgnsmartcloud.com signature does not survive the forward); PO mail passes for amazon.coupahost.com. amazonses.com also passes on PO mail but is deliberately excluded -- every SES customer's outbound mail passes for it. Every failure path (env var unset, header missing or unparseable, verdict fail, unaligned domain) rejects the email: a structured warning with the reason and S3 key is logged and the record skipped without erroring the invocation, so rejected mail causes no Lambda retries or DLQ messages. Handler signatures and event sources are unchanged. Refs: INFRA-107 * Harden AR parser per cross-family review Cross-family (GPT-4.1) review findings: terminate the dkim result token at end-of-clause, whitespace, or a comment so a value like "dkim=pass-fake" can never be read as a pass; normalize trailing dots off allowlist entries so "seahaven.com." matches; make the compat32 parser policy explicit. Adds tests for result-token boundaries, comments after the result, quoted domain values, and folding inside a dkim clause. Refs: INFRA-107 * Harden AR parsing and alarm on sender-auth rejects The SES-stamped Authentication-Results value echoes attacker-controlled SMTP-session tokens (envelope-from, helo, header.from) as their own semicolon-delimited property clauses. A naive split(";") tore an RFC 5321 quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass clause, so a fully spoofed email was accepted on the genuinely SES-stamped topmost header. Tokenise comment- and quoted-string-aware (RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on semicolons outside a quoted-string, and fail closed on unbalanced quotes/comments so a ';' inside a quoted pvalue can never start a clause. Rejected mail returns normally (no error, no retry, no DLQ message), so a signing-domain drift or a wrong allowlist would silently discard 100% of legitimate mail while every alarm stayed green. Add a CloudWatch Logs metric filter + alarm on the sender_auth_rejected warning to both stacks so a false-reject storm pages instead of vanishing. This is also the safety net for the WO seahaven.com allowlist assumption, which must be validated against a live SES-stamped header (a plain Gmail auto-forward re-signs under the sending Workspace domain, not seahaven.com). Refs: INFRA-107 * chore: retrigger CI (no run recorded for 7c74ac1) * Fix quoted-AUID DKIM domain spoof in sender auth Resolve three confirmed /sh-security-review findings on the fail-closed SES sender-authentication control. HIGH: header.i/header.d domain extraction was not quoted-string aware. An attacker with a valid DKIM key for their own domain could set an RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive extractor stopped at the closing quote and returned seahaven.com, accepting forged mail. Extraction now tokenises the clause with the same quoted-string discipline already used for clause splitting: header.d (the plain signing domain) is authoritative when present, otherwise the header.i domain is the part after the AUID's LAST top-level "@", so a "@" inside a quoted local-part is treated as signer-controlled label text and yields the true signer (attacker.com), not seahaven.com. LOW: the topmost-header parse ran outside evaluate_sender_authentication's try/except, so an unexpected parser exception on crafted input could propagate into the handler and Lambda async retries/DLQ. The parse now fails CLOSED with an authentication_results_unparseable reason. MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to a low-volume total-reject outage (a trickle that never sums to 3). Both stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3 / datapoints_to_alarm=2, so a sustained reject condition pages even at one reject per period while a lone stray probe self-clears. Refs: INFRA-107 * Load Lambda function dir on sys.path in tests Rebasing INFRA-107 onto main folded #95's pytest suite into this branch's tests. The unified conftest loads the PO/WO handlers by file path, and handler.py now does `from ses_auth import authenticate_inbound_email` -- a bare sibling import that resolves in the Lambda only because the runtime puts each function's own directory on sys.path. The shared load_handler now adds that directory so the handler tests import correctly alongside the sender-auth tests. Refs: INFRA-107 * Note #97 test files in README directory tree The rebase onto main brought in #97's tests/requirements.txt and tests/test_po_merge.py. List both in the directory tree so it matches the tree on disk. Refs: INFRA-107 * Document INFRA-107 forwarder-binding risk acceptance Record the accepted risk that WO sender auth binds to the apm@ forward's re-signing domain (seahaven.com) rather than the Hexagon originator; the apm@ Google Group's restricted posting policy is the load-bearing control (escalates to HIGH if the group is opened to external posting). Also correct the sender-auth-rejected alarm docs to match the shipped config (>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and note the SES-AR-01/02 parser hardening follow-ups. Refs: INFRA-107
2026-07-15 20:58:47 -04:00
)
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
# --- Template fallback-rate alarm: workorder-email-processor ---
# The processor tries a deterministic template parse first and only calls
# the Bedrock AI extractor on a miss/invalid. A sustained rise in the
# ai_fallback share signals Hexagon template drift (coverage collapse).
# EMF metric Seahaven/WorkorderIngest/ParseOutcome, dimensioned by
# ParseMethod (template|ai_fallback). 15-min periods (deliberate deviation
# from the 5-min house style) accumulate a stable denominator at the low
# ~760/day volume; FILL(0) + a >=10-sample volume floor prevent
# low-volume false pages and INSUFFICIENT_DATA. ALARM-only SnsAction to
# site-alerts, no OK action, NOT_BREACHING -- matching the stack idiom.
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112) The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
# rejected_included=True: the WO expression folds ai_fallback_rejected
# (rej) into BOTH numerator and denominator -- a drift outage whose AI
# output also fails the gate must still count as fallback, otherwise it
# would LOWER the observed rate while silently dropping mail. (Contrast
# PO, which excludes rej to avoid a pre-call double-count.)
common.make_fallback_rate_alarm(
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
self,
"EmailProcessorTemplateFallbackRateAlarm",
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112) The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
namespace="Seahaven/WorkorderIngest",
alarm_topic=alarm_topic,
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
alarm_name="workorder-email-processor-template-fallback-rate",
alarm_description=(
"workorder-email-processor deterministic-template coverage "
"collapse: >15% of parses fell back to the Bedrock AI extractor"
),
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112) The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
rejected_included=True,
period=Duration.minutes(15),
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
threshold=15,
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112) The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
floor=10,
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
evaluation_periods=3,
datapoints_to_alarm=2,
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112) The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
)
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99) * Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path (#104) * fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path The ai_fallback parse path applied no validation gate to raw Bedrock/LLM output before DynamoDB writes, and the extraction prompt concatenated the untrusted email body directly with no instructions-vs-data delimiter. A DKIM-passing attacker could prompt-inject arbitrary field values into the work-order store. Changes: - wrap untrusted email in \<email\> XML block with prompt instructing the model to treat its contents as data only - add validate_ai_fallback() in template_parser that enforces the same contract keys, enums, and patterns as the template path before any write - call validate_ai_fallback() in handler() dispatch; emit an ai_fallback_rejected EMF metric on failure and skip the record - add 17 unit tests covering every gate rule and two end-to-end dispatch tests (injected email_type, injected status) Refs #101 * style: apply ruff formatting to fix CI check * harden ai_fallback gate: review fixes + security-review findings Review follow-up on the ai_fallback validation gate (PR #104), plus findings from a fan-out /sh-security-review of the change surface. Reviewer FIX items: - Neutralize forged <email> delimiters in the untrusted body before wrapping, so an in-body </email> cannot escape the data block. - Fail closed on non-dict model output instead of crashing the handler into async retries; count ai_fallback_rejected parses in the fallback-rate alarm and add a dedicated rejected-parse alarm so a gate-rejection drift outage is not silent. - Return a distinct invalid_status reason (was malformed_site_code); validate ISO-8601 dates; README + docstring updates. Security-review findings (detector fan-out + proof-or-kill verifier): - ReDoS (confirmed, medium): the tag neutralizer used two \s* around an optional /, backtracking quadratically on "<" + a long whitespace run (~32s at 100k chars -- one email could time out the Lambda). Collapse to a single [\s/]* class: linear, same defanging. - Unhashable-type crash (confirmed): a JSON list/dict for email_type or status made `x in <set>` raise TypeError, escaping the gate into retries. Guard with isinstance(str) before membership. - Unicode/newline regex (confirmed): _WO_ID_RE/_SITE_CODE_RE used ^..$ with \d, admitting fullwidth digits ("12345" as a lookalike partition key) and trailing newlines. Switch to \A[0-9]+\Z (and the handler's inline recheck to [0-9]) so neither passes. - Alarm comment (confirmed, low): corrected the "slow trickle still pages" wording -- rejections >~25-30 min apart page on neither alarm, the same knowingly-accepted residual as sender-auth-rejected. Refuted: residual free-text prompt injection is inherent to trusting allowlisted senders, not a new primitive; no DynamoDB key-poisoning bypass survives both gates ('#' can never enter work_order_id). 7 new regression tests. All 260 tests pass; ruff clean; cdk synth OK. --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-16 16:23:28 -04:00
# --- AI-fallback rejected alarm: workorder-email-processor ---
# A parse rejected by the validate_ai_fallback gate is dropped without
# error/retry/DLQ (fail closed), so like sender-auth rejections it
# needs its own pager or a sustained rejection condition (prompt-
# injection probing, or template drift whose AI output fails the gate)
# stays silent. Same sparse-arrival idiom as the sender-auth-rejected
# alarm: >=1 rejection per 5-min period, 2 of the last 6 periods (30
# min), so a lone probe self-clears but a burst pages within ~10 min.
# Coverage residual (matching the sender-auth-rejected sibling and
# knowingly accepted): rejections spaced >~25-30 min apart never place
# two breaching datapoints in one 30-min window, and the fallback-rate
# alarm dilutes them below 15% against normal template volume, so a
# *very* sparse silent-drop trickle is not paged by either alarm.
# EMF emits no datapoint in quiet periods (no metric-filter
# default_value here); NOT_BREACHING treats those gaps as OK.
cloudwatch.Metric(
namespace="Seahaven/WorkorderIngest",
metric_name="ParseOutcome",
dimensions_map={"ParseMethod": "ai_fallback_rejected"},
statistic="Sum",
period=Duration.minutes(5),
).create_alarm(
self,
"EmailProcessorAiFallbackRejectedAlarm",
alarm_name="workorder-email-processor-ai-fallback-rejected",
alarm_description=(
"workorder-email-processor is rejecting Bedrock AI-fallback "
"output at the validation gate (possible prompt-injection "
"probing or template drift silently dropping real mail)"
),
threshold=1,
evaluation_periods=6,
datapoints_to_alarm=2,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
# S3 event notification -> Lambda
email_bucket.add_event_notification(
s3.EventType.OBJECT_CREATED,
s3n.LambdaDestination(email_processor),
s3.NotificationKeyFilter(prefix="inbound/"),
)
# --- SES Receipt Rule ---
rule_set = ses.ReceiptRuleSet.from_receipt_rule_set_name(
self,
"ExistingRuleSet",
"INBOUND_MAIL",
)
rule_set.add_rule(
"WorkorderEmailRule",
recipients=["apm@int.seahaven.com"],
actions=[
ses_actions.S3(
bucket=email_bucket,
object_key_prefix="inbound/",
),
],
)
Land safe fixes from 2026-06-17 security sweep (#97) * Remove gratuitous KMS grant on shared DynamoDB CMK wo-email-processor held grant_encrypt_decrypt on the shared seahaven-dynamodb CMK, but the WorkOrders/WorkOrderComments tables are not encrypted with that CMK. The grant was dead weight that extended the WO processor's decrypt reach to the CMK protecting the purchase-orders table (cross-stack decrypt). Drop it to restore least privilege; re-add as part of the table CMK migration (INFRA-6). Refs: INFRA-6 * Require Secrets Manager key for Anthropic client Remove the silent fallback to a plaintext ANTHROPIC_API_KEY env var in both email processors; require ANTHROPIC_API_KEY_SECRET_ARN and raise if absent so a misconfigured deploy fails loudly instead of using an unmanaged key. Adapted from f175323 on security/sweep-2026-06-17. The From-header sender-domain allowlist from that commit is intentionally dropped: the From header is spoofable (INFRA-107, confirmed critical) and sender authentication is being reworked in a separate PR. Refs: INFRA-107 * Merge PO revisions and handle out-of-order events save_revision did a full put_item overwrite, so a revision omitting line_items/supplier permanently deleted them. save_new_po used a conditional put that silently dropped the PO when an out-of-order cancellation had already created a skeleton row. Switch both to field-level merge update_items: a revision now SETs only the fields it carries, and a new_po backfills data into a pre-existing Cancelled skeleton while preserving the Cancelled status. No email can now delete data established by an earlier one. * Gate web UIs behind auth and escape currency XSS The po-web-ui and workorder-web-ui handlers had no auth: any invocation path returned the full PO/WO DB. Add a fail-closed shared-secret gate (X-Auth-Token / Bearer, constant-time compared to WEB_UI_AUTH_TOKEN) so a future re-attached Function URL cannot re-expose the data (URLs removed under INFRA-74). Wire the token from the SSM String param /procurement-ingest/web-ui-auth-token. Also fix stored XSS in po-web-ui fmt_currency: the non-numeric fallback returned str(val) unescaped, so a prompt-injected email could make Claude emit total_amount as <script>. Escape it. Refs: INFRA-74 * Document sweep security fixes and merge semantics Update the README for the 2026-06-17 security sweep: required Secrets Manager key (no plaintext env fallback), web UI auth gate + SSM token setup step, output-escaping note, and the new PO revision/cancellation merge behavior. Adapted from d91f45e on security/sweep-2026-06-17; the sender allowlist documentation is dropped along with the allowlist itself (deferred to the INFRA-107 sender-authentication rework). Refs: INFRA-107 * fix: resolve web UI auth token from Secrets Manager at runtime Replace the plaintext SSM String parameter with a Secrets Manager secret referenced by ARN only. The token is fetched and cached at module level on first invocation, keeping shared secrets out of CloudFormation templates and Lambda environment variables. Refs: PR-97 * Add TTL to web UI auth token cache for rotation The web-ui handlers cached the Secrets Manager auth token at module level with no expiry, so a rotated secret was only picked up when the warm container recycled — an emergency rotation could take hours to take effect. Cache the fetched value for a 5-minute TTL instead, so a rotated token propagates within the TTL while still avoiding a Secrets Manager call on every request. Still fails closed when the secret is unset or unreadable. Refs: INFRA-74 * Log Secrets Manager failures in web UI auth token fetch The web UI auth gate correctly fails closed when the shared token cannot be read, but _get_auth_token() swallowed every exception silently. A Secrets Manager permission or config error then made every request 401 with no operational signal, leaving an outage indistinguishable from ordinary unauthenticated traffic. Add a module-level logger to both web_ui handlers and log the fetch failure with logger.exception() in the except block before returning None. Behavior is unchanged (still fails closed); the failure is now visible in CloudWatch. The secret value is never logged. The two handlers stay byte-consistent in the mirrored _get_auth_token() region. The companion finding on the CDK import of the shared procurement-ingest/web-ui-auth-token secret was evaluated and left as-is: the token is a single secret shared by both the PO and WO stacks, so from_secret_name_v2 (which scopes grant_read via the standard 6-char suffix wildcard) is correct; making it a CDK-managed Secret in both stacks would collide the two stacks on the same explicit secret name at deploy time. Refs: INFRA-74 * Make Cancelled PO status sticky via atomic write The PO merge path read status with a get_item (_is_cancelled) and then wrote with an unconditional update_item. Two defects followed from this: - Race (Issue A): a cancellation landing between the read and the write was silently un-cancelled by a revision carrying a non-cancelled po_status — a TOCTOU on a table with concurrent email processing. - Over-broad strip (Issue B): save_revision dropped po_status whenever the PO was Cancelled, so legitimate status updates on non-cancelled POs and status-less revisions were affected rather than only the true un-cancel transition. Enforce the invariant server-side instead. "Cancelled" is a sticky, authoritative status: once set, later new_po/revision emails may enrich other fields but must never move it to a non-cancelled status. When the payload carries a non-cancelled po_status, _merge_update issues the update_item guarded by ConditionExpression "attribute_not_exists(po_status) OR po_status <> :marker", evaluated atomically at write time, so a cancellation that lands first always wins. On ConditionalCheckFailedException the same fields are re-written without po_status/cancelled_at, enriching the record while Cancelled sticks. Payloads with no status change, or an already -Cancelled status, take a plain merge — the status is only ever suppressed on a real un-cancel. This removes the non-atomic get_item from the write path; _is_cancelled is deleted. Key schema and attribute names are unchanged, so the cross-stack purchase-orders contract (read-only by seahaven-slack-bot) holds. Add moto-backed tests covering un-cancel suppression with field enrichment, status-less merge onto a Cancelled PO, legitimate status updates on non-cancelled POs, new_po backfill of a Cancelled skeleton, fresh create/merge, and authoritative save_cancellation. Refs: #97
2026-07-15 20:17:46 -04:00
# --- Web UI auth token secret ---
# Shared secret for the web UI auth gate, stored in Secrets Manager and
# resolved at runtime so the token never appears in CloudFormation templates
# or Lambda environment variables. Create this secret before deploying
# either stack; both PO and WO stacks reference it by name.
web_ui_auth_secret = secretsmanager.Secret.from_secret_name_v2(
self,
"WebUiAuthToken",
"procurement-ingest/web-ui-auth-token",
)
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
# --- Web UI Lambda ---
web_ui_log_group = common.make_function_log_group(
self, "WebUI", "workorder-web-ui"
)
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
web_ui = lambda_.Function(
self,
"WebUI",
function_name="workorder-web-ui",
runtime=lambda_.Runtime.PYTHON_3_12,
architecture=lambda_.Architecture.ARM_64,
handler="handler.handler",
feat: widen email-processor asset roots to lambdas/ with scoped globs + excludes (refactor phase 2) (#109) Both email-processor Code.from_asset calls now bundle from lambdas/ instead of their per-function subdirectory, so Phase 3's shared/ module is reachable from the asset root once it lands. The bundling commands were rewritten for the new cwd (pip install -r <po|wo>/ email_processor/requirements.txt -t /asset-output && cp <po|wo>/ email_processor/*.py /asset-output/), preserving the ARM64 --platform manylinux2014_aarch64 --only-binary=:all: pin exactly — its removal shipped x86 wheels into the ARM64 function and caused a 100% outage (PR #34). All five from_asset calls (both email processors, po web_ui, po site_extractor, wo web_ui) now exclude **/__pycache__/**; the two widened ones also exclude **/tests/** and **/package/**. Without the package/ exclude, the stale untracked 44 MB lambdas/po/email_processor/package/ dir (local-only, never present in CI) would diverge local vs CI asset hashes and force spurious redeploys — from_asset doesn't honor .gitignore. That dir is left in place; deleting it is Adam's call. WO's prod zip shrinks as deliberate cleanup, not a byte-identical match to PO: the old `cp -r .` shipped tests/ (real scrubbed .eml fixtures), __pycache__/, and requirements.txt into production. The acceptance bar for WO is runtime-imported module set unchanged + smoke, not a byte-identical zip; PO keeps the byte-identical first-party file set guarantee. tests/test_bundle_consistency.py is updated in the same change to recognize the scoped `cp po/email_processor/*.py` (resp. wo) glob as the new unconditionally-safe shape, without loosening the allowlist-revert detection, the detection-logic mutation test, or the PO_EXPECTED_TOP_LEVEL_MODULES exact-set pin. No code moved under lambdas/ in this change (git diff main...HEAD -- lambdas/ is empty); only CDK asset wiring and its tests changed.
2026-07-17 15:47:01 -04:00
code=lambda_.Code.from_asset(
feat: extract lambdas/shared/ — single-source ses_auth, web_ui auth, email parsing, EMF emitter (refactor phase 3) (#111) Four modules move into the handbook-mandated lambdas/shared/ location, collapsing duplicated logic that had to be kept in sync by hand across the PO and WO pipelines: - ses_auth.py: the PO and WO copies were verified sha256-identical against the feature/phase-7-ops-recovery baseline before the move (no drift since the last audit). shared/ses_auth.py is the exact bytes of that one copy; both originals are git rm'd (the PO copy via rename, the WO copy as a straight delete). Bundling lands the module flat in /asset-output for both email processors, so the handlers keep `from ses_auth import authenticate_inbound_email` unchanged — zero handler diff for this move, which is what keeps fail-closed auth byte-identical through the change. - web_ui_auth.py: extracts the byte-identical _get_auth_token / _header / is_authenticated block plus the four token-cache globals out of both web_ui handlers. The per-stack INFRA-74 comments stay in each handler as-is (deliberately drifted wording, stack-specific) rather than being unified into the shared module. Fail-closed semantics (unset ARN or Secrets Manager exception -> deny) are unchanged. - email_parsing.py: parse_raw_email ships as the superset version that returns cc unconditionally. WO's output is bit-identical to before; PO simply ignores the cc field rather than being "cleaned up" to consume it. No second variant is kept. - emf.py: a generic emitter parameterized by namespace, dimension sets, and properties. Every call site's emitted EMF envelope is unchanged, including the load-bearing [["ParseMethod"],["ParseMethod","TemplateId"]] dimension-set shape the alarms and metric filters depend on. Emission ordering is untouched: PO still emits ai_fallback before the Bedrock call, WO still emits its mutually-exclusive ai_fallback/ai_fallback_rejected after its gate. The deliberate-double-count comments survive. _emit_derived_agreement_metric was found living inside derived_fields.py, so per the DERIVED-FIELDS exception it is left as a third, unconverted copy (derived_fields.py and the shadow DerivedFieldAgreement telemetry stay untouchable while that bake runs) — a comment there points at shared/emf.py for the eventual follow-up. Bundling: both email-processor cdk bundling commands gain a trailing `cp shared/*.py /asset-output/` (they were already cp-only post-Phase 7, so no pip step or manylinux pin is reintroduced). Both web_ui functions gain the same widened-root staging so web_ui_auth.py ships beside their handler; site_extractor's from_asset is untouched. Tests: PO_EXPECTED_TOP_LEVEL_MODULES gains the shared modules that now ship, the AST sibling-import check resolves imports whose source now lives under shared/, and the new shared cp line has its own revert/mutation detection. _SIBLING_MODULES resolution and _po_parser_support.py now load ses_auth/email_parsing/emf from shared/; the two-copy ses_auth byte-identity fixture-hygiene test is retired as obsolete now that there is one copy, and the ses_auth fixture parameterization over two identical copies is dropped. The sys.modules save/restore dance for template_parser (still duplicated per-pipeline) is left in place.
2026-07-20 13:38:23 -04:00
"../lambdas",
exclude=["**/__pycache__/**"],
bundling=cdk.BundlingOptions(
image=lambda_.Runtime.PYTHON_3_12.bundling_image,
command=[
"bash",
"-c",
# web_ui_auth.py is shared (lambdas/shared) and must land
# FLAT beside handler.py so
# `from web_ui_auth import is_authenticated` resolves at
# runtime. Only web_ui_auth is copied from shared/. WO
# baseline keeps __init__.py and requirements.txt, so the
# whole web_ui dir is copied; deployed file list becomes
# {__init__, handler, requirements.txt, web_ui_auth}.
# NOTE: the top-level `exclude=` on from_asset only
# filters the asset-hash fingerprint, NOT the directory
# Docker bundling actually mounts, so a local
# __pycache__ on disk at synth time WOULD otherwise leak
# into the bundled zip -- strip it explicitly post-cp
# instead of relying on exclude.
"cp -r wo/web_ui/. /asset-output/ && "
"cp shared/web_ui_auth.py /asset-output/ && "
"rm -rf /asset-output/__pycache__",
],
),
feat: widen email-processor asset roots to lambdas/ with scoped globs + excludes (refactor phase 2) (#109) Both email-processor Code.from_asset calls now bundle from lambdas/ instead of their per-function subdirectory, so Phase 3's shared/ module is reachable from the asset root once it lands. The bundling commands were rewritten for the new cwd (pip install -r <po|wo>/ email_processor/requirements.txt -t /asset-output && cp <po|wo>/ email_processor/*.py /asset-output/), preserving the ARM64 --platform manylinux2014_aarch64 --only-binary=:all: pin exactly — its removal shipped x86 wheels into the ARM64 function and caused a 100% outage (PR #34). All five from_asset calls (both email processors, po web_ui, po site_extractor, wo web_ui) now exclude **/__pycache__/**; the two widened ones also exclude **/tests/** and **/package/**. Without the package/ exclude, the stale untracked 44 MB lambdas/po/email_processor/package/ dir (local-only, never present in CI) would diverge local vs CI asset hashes and force spurious redeploys — from_asset doesn't honor .gitignore. That dir is left in place; deleting it is Adam's call. WO's prod zip shrinks as deliberate cleanup, not a byte-identical match to PO: the old `cp -r .` shipped tests/ (real scrubbed .eml fixtures), __pycache__/, and requirements.txt into production. The acceptance bar for WO is runtime-imported module set unchanged + smoke, not a byte-identical zip; PO keeps the byte-identical first-party file set guarantee. tests/test_bundle_consistency.py is updated in the same change to recognize the scoped `cp po/email_processor/*.py` (resp. wo) glob as the new unconditionally-safe shape, without loosening the allowlist-revert detection, the detection-logic mutation test, or the PO_EXPECTED_TOP_LEVEL_MODULES exact-set pin. No code moved under lambdas/ in this change (git diff main...HEAD -- lambdas/ is empty); only CDK asset wiring and its tests changed.
2026-07-17 15:47:01 -04:00
),
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
timeout=Duration.seconds(15),
memory_size=128,
log_group=web_ui_log_group,
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
environment={
"WORK_ORDERS_TABLE": work_orders_table.table_name,
"COMMENTS_TABLE": comments_table.table_name,
Land safe fixes from 2026-06-17 security sweep (#97) * Remove gratuitous KMS grant on shared DynamoDB CMK wo-email-processor held grant_encrypt_decrypt on the shared seahaven-dynamodb CMK, but the WorkOrders/WorkOrderComments tables are not encrypted with that CMK. The grant was dead weight that extended the WO processor's decrypt reach to the CMK protecting the purchase-orders table (cross-stack decrypt). Drop it to restore least privilege; re-add as part of the table CMK migration (INFRA-6). Refs: INFRA-6 * Require Secrets Manager key for Anthropic client Remove the silent fallback to a plaintext ANTHROPIC_API_KEY env var in both email processors; require ANTHROPIC_API_KEY_SECRET_ARN and raise if absent so a misconfigured deploy fails loudly instead of using an unmanaged key. Adapted from f175323 on security/sweep-2026-06-17. The From-header sender-domain allowlist from that commit is intentionally dropped: the From header is spoofable (INFRA-107, confirmed critical) and sender authentication is being reworked in a separate PR. Refs: INFRA-107 * Merge PO revisions and handle out-of-order events save_revision did a full put_item overwrite, so a revision omitting line_items/supplier permanently deleted them. save_new_po used a conditional put that silently dropped the PO when an out-of-order cancellation had already created a skeleton row. Switch both to field-level merge update_items: a revision now SETs only the fields it carries, and a new_po backfills data into a pre-existing Cancelled skeleton while preserving the Cancelled status. No email can now delete data established by an earlier one. * Gate web UIs behind auth and escape currency XSS The po-web-ui and workorder-web-ui handlers had no auth: any invocation path returned the full PO/WO DB. Add a fail-closed shared-secret gate (X-Auth-Token / Bearer, constant-time compared to WEB_UI_AUTH_TOKEN) so a future re-attached Function URL cannot re-expose the data (URLs removed under INFRA-74). Wire the token from the SSM String param /procurement-ingest/web-ui-auth-token. Also fix stored XSS in po-web-ui fmt_currency: the non-numeric fallback returned str(val) unescaped, so a prompt-injected email could make Claude emit total_amount as <script>. Escape it. Refs: INFRA-74 * Document sweep security fixes and merge semantics Update the README for the 2026-06-17 security sweep: required Secrets Manager key (no plaintext env fallback), web UI auth gate + SSM token setup step, output-escaping note, and the new PO revision/cancellation merge behavior. Adapted from d91f45e on security/sweep-2026-06-17; the sender allowlist documentation is dropped along with the allowlist itself (deferred to the INFRA-107 sender-authentication rework). Refs: INFRA-107 * fix: resolve web UI auth token from Secrets Manager at runtime Replace the plaintext SSM String parameter with a Secrets Manager secret referenced by ARN only. The token is fetched and cached at module level on first invocation, keeping shared secrets out of CloudFormation templates and Lambda environment variables. Refs: PR-97 * Add TTL to web UI auth token cache for rotation The web-ui handlers cached the Secrets Manager auth token at module level with no expiry, so a rotated secret was only picked up when the warm container recycled — an emergency rotation could take hours to take effect. Cache the fetched value for a 5-minute TTL instead, so a rotated token propagates within the TTL while still avoiding a Secrets Manager call on every request. Still fails closed when the secret is unset or unreadable. Refs: INFRA-74 * Log Secrets Manager failures in web UI auth token fetch The web UI auth gate correctly fails closed when the shared token cannot be read, but _get_auth_token() swallowed every exception silently. A Secrets Manager permission or config error then made every request 401 with no operational signal, leaving an outage indistinguishable from ordinary unauthenticated traffic. Add a module-level logger to both web_ui handlers and log the fetch failure with logger.exception() in the except block before returning None. Behavior is unchanged (still fails closed); the failure is now visible in CloudWatch. The secret value is never logged. The two handlers stay byte-consistent in the mirrored _get_auth_token() region. The companion finding on the CDK import of the shared procurement-ingest/web-ui-auth-token secret was evaluated and left as-is: the token is a single secret shared by both the PO and WO stacks, so from_secret_name_v2 (which scopes grant_read via the standard 6-char suffix wildcard) is correct; making it a CDK-managed Secret in both stacks would collide the two stacks on the same explicit secret name at deploy time. Refs: INFRA-74 * Make Cancelled PO status sticky via atomic write The PO merge path read status with a get_item (_is_cancelled) and then wrote with an unconditional update_item. Two defects followed from this: - Race (Issue A): a cancellation landing between the read and the write was silently un-cancelled by a revision carrying a non-cancelled po_status — a TOCTOU on a table with concurrent email processing. - Over-broad strip (Issue B): save_revision dropped po_status whenever the PO was Cancelled, so legitimate status updates on non-cancelled POs and status-less revisions were affected rather than only the true un-cancel transition. Enforce the invariant server-side instead. "Cancelled" is a sticky, authoritative status: once set, later new_po/revision emails may enrich other fields but must never move it to a non-cancelled status. When the payload carries a non-cancelled po_status, _merge_update issues the update_item guarded by ConditionExpression "attribute_not_exists(po_status) OR po_status <> :marker", evaluated atomically at write time, so a cancellation that lands first always wins. On ConditionalCheckFailedException the same fields are re-written without po_status/cancelled_at, enriching the record while Cancelled sticks. Payloads with no status change, or an already -Cancelled status, take a plain merge — the status is only ever suppressed on a real un-cancel. This removes the non-atomic get_item from the write path; _is_cancelled is deleted. Key schema and attribute names are unchanged, so the cross-stack purchase-orders contract (read-only by seahaven-slack-bot) holds. Add moto-backed tests covering un-cancel suppression with field enrichment, status-less merge onto a Cancelled PO, legitimate status updates on non-cancelled POs, new_po backfill of a Cancelled skeleton, fresh create/merge, and authoritative save_cancellation. Refs: #97
2026-07-15 20:17:46 -04:00
# Defense-in-depth shared secret for the web UI handler. The
# handler fails closed if this ARN is unset or the secret is
# missing, so any future invocation path cannot re-expose the
# WO DB unauthenticated. The secret value is fetched at runtime
# from Secrets Manager (not embedded in env vars or template).
"WEB_UI_AUTH_TOKEN_SECRET_ARN": web_ui_auth_secret.secret_arn,
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
},
)
work_orders_table.grant_read_data(web_ui)
comments_table.grant_read_data(web_ui)
Land safe fixes from 2026-06-17 security sweep (#97) * Remove gratuitous KMS grant on shared DynamoDB CMK wo-email-processor held grant_encrypt_decrypt on the shared seahaven-dynamodb CMK, but the WorkOrders/WorkOrderComments tables are not encrypted with that CMK. The grant was dead weight that extended the WO processor's decrypt reach to the CMK protecting the purchase-orders table (cross-stack decrypt). Drop it to restore least privilege; re-add as part of the table CMK migration (INFRA-6). Refs: INFRA-6 * Require Secrets Manager key for Anthropic client Remove the silent fallback to a plaintext ANTHROPIC_API_KEY env var in both email processors; require ANTHROPIC_API_KEY_SECRET_ARN and raise if absent so a misconfigured deploy fails loudly instead of using an unmanaged key. Adapted from f175323 on security/sweep-2026-06-17. The From-header sender-domain allowlist from that commit is intentionally dropped: the From header is spoofable (INFRA-107, confirmed critical) and sender authentication is being reworked in a separate PR. Refs: INFRA-107 * Merge PO revisions and handle out-of-order events save_revision did a full put_item overwrite, so a revision omitting line_items/supplier permanently deleted them. save_new_po used a conditional put that silently dropped the PO when an out-of-order cancellation had already created a skeleton row. Switch both to field-level merge update_items: a revision now SETs only the fields it carries, and a new_po backfills data into a pre-existing Cancelled skeleton while preserving the Cancelled status. No email can now delete data established by an earlier one. * Gate web UIs behind auth and escape currency XSS The po-web-ui and workorder-web-ui handlers had no auth: any invocation path returned the full PO/WO DB. Add a fail-closed shared-secret gate (X-Auth-Token / Bearer, constant-time compared to WEB_UI_AUTH_TOKEN) so a future re-attached Function URL cannot re-expose the data (URLs removed under INFRA-74). Wire the token from the SSM String param /procurement-ingest/web-ui-auth-token. Also fix stored XSS in po-web-ui fmt_currency: the non-numeric fallback returned str(val) unescaped, so a prompt-injected email could make Claude emit total_amount as <script>. Escape it. Refs: INFRA-74 * Document sweep security fixes and merge semantics Update the README for the 2026-06-17 security sweep: required Secrets Manager key (no plaintext env fallback), web UI auth gate + SSM token setup step, output-escaping note, and the new PO revision/cancellation merge behavior. Adapted from d91f45e on security/sweep-2026-06-17; the sender allowlist documentation is dropped along with the allowlist itself (deferred to the INFRA-107 sender-authentication rework). Refs: INFRA-107 * fix: resolve web UI auth token from Secrets Manager at runtime Replace the plaintext SSM String parameter with a Secrets Manager secret referenced by ARN only. The token is fetched and cached at module level on first invocation, keeping shared secrets out of CloudFormation templates and Lambda environment variables. Refs: PR-97 * Add TTL to web UI auth token cache for rotation The web-ui handlers cached the Secrets Manager auth token at module level with no expiry, so a rotated secret was only picked up when the warm container recycled — an emergency rotation could take hours to take effect. Cache the fetched value for a 5-minute TTL instead, so a rotated token propagates within the TTL while still avoiding a Secrets Manager call on every request. Still fails closed when the secret is unset or unreadable. Refs: INFRA-74 * Log Secrets Manager failures in web UI auth token fetch The web UI auth gate correctly fails closed when the shared token cannot be read, but _get_auth_token() swallowed every exception silently. A Secrets Manager permission or config error then made every request 401 with no operational signal, leaving an outage indistinguishable from ordinary unauthenticated traffic. Add a module-level logger to both web_ui handlers and log the fetch failure with logger.exception() in the except block before returning None. Behavior is unchanged (still fails closed); the failure is now visible in CloudWatch. The secret value is never logged. The two handlers stay byte-consistent in the mirrored _get_auth_token() region. The companion finding on the CDK import of the shared procurement-ingest/web-ui-auth-token secret was evaluated and left as-is: the token is a single secret shared by both the PO and WO stacks, so from_secret_name_v2 (which scopes grant_read via the standard 6-char suffix wildcard) is correct; making it a CDK-managed Secret in both stacks would collide the two stacks on the same explicit secret name at deploy time. Refs: INFRA-74 * Make Cancelled PO status sticky via atomic write The PO merge path read status with a get_item (_is_cancelled) and then wrote with an unconditional update_item. Two defects followed from this: - Race (Issue A): a cancellation landing between the read and the write was silently un-cancelled by a revision carrying a non-cancelled po_status — a TOCTOU on a table with concurrent email processing. - Over-broad strip (Issue B): save_revision dropped po_status whenever the PO was Cancelled, so legitimate status updates on non-cancelled POs and status-less revisions were affected rather than only the true un-cancel transition. Enforce the invariant server-side instead. "Cancelled" is a sticky, authoritative status: once set, later new_po/revision emails may enrich other fields but must never move it to a non-cancelled status. When the payload carries a non-cancelled po_status, _merge_update issues the update_item guarded by ConditionExpression "attribute_not_exists(po_status) OR po_status <> :marker", evaluated atomically at write time, so a cancellation that lands first always wins. On ConditionalCheckFailedException the same fields are re-written without po_status/cancelled_at, enriching the record while Cancelled sticks. Payloads with no status change, or an already -Cancelled status, take a plain merge — the status is only ever suppressed on a real un-cancel. This removes the non-atomic get_item from the write path; _is_cancelled is deleted. Key schema and attribute names are unchanged, so the cross-stack purchase-orders contract (read-only by seahaven-slack-bot) holds. Add moto-backed tests covering un-cancel suppression with field enrichment, status-less merge onto a Cancelled PO, legitimate status updates on non-cancelled POs, new_po backfill of a Cancelled skeleton, fresh create/merge, and authoritative save_cancellation. Refs: #97
2026-07-15 20:17:46 -04:00
web_ui_auth_secret.grant_read(web_ui)
Merge workorder-ingest into unified procurement repo (#22) * Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70) * Add CloudWatch alarm coverage for po-ingest and workorder-ingest Expands alarm coverage across both CDK stacks. All alarms are ALARM-only (no OK action) to the shared site-alerts SNS topic, with TreatMissingData NOT_BREACHING. The site-alerts topic is now imported once near the top of each stack so every alarm reuses one Topic instance. po-ingest (cdk/po_stack.py): - Errors: po-ingest-site-extractor - Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui - Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption), po-ingest-site-extractor, po-web-ui - DynamoDB throttle + system-error: purchase-orders, verified-sites, pending-site-review workorder-ingest (cdk/wo_stack.py): - Throttles: workorder-email-processor - Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption) - DynamoDB throttle + system-error: WorkOrders, WorkOrderComments DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation dimension set, so each table alarm is a Sum math expression across operations via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.259.0). Refs INFRA-41 / audit H-8. * Drop NEEDS ADAM SIGN-OFF wording from alarm comments Duration alarm thresholds are owner-approved; remove the sign-off flag from po_stack.py and wo_stack.py comments. Threshold values, eval config, and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
# --- DynamoDB throttle + system-error alarms ---
# ThrottledRequests / SystemErrors emit at TableName + Operation only
# (verified against live CloudWatch: no TableName-only rollup exists, and
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112) The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
# metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.261.0).
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70) * Add CloudWatch alarm coverage for po-ingest and workorder-ingest Expands alarm coverage across both CDK stacks. All alarms are ALARM-only (no OK action) to the shared site-alerts SNS topic, with TreatMissingData NOT_BREACHING. The site-alerts topic is now imported once near the top of each stack so every alarm reuses one Topic instance. po-ingest (cdk/po_stack.py): - Errors: po-ingest-site-extractor - Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui - Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption), po-ingest-site-extractor, po-web-ui - DynamoDB throttle + system-error: purchase-orders, verified-sites, pending-site-review workorder-ingest (cdk/wo_stack.py): - Throttles: workorder-email-processor - Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption) - DynamoDB throttle + system-error: WorkOrders, WorkOrderComments DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation dimension set, so each table alarm is a Sum math expression across operations via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.259.0). Refs INFRA-41 / audit H-8. * Drop NEEDS ADAM SIGN-OFF wording from alarm comments Duration alarm thresholds are owner-approved; remove the sign-off flag from po_stack.py and wo_stack.py comments. Threshold values, eval config, and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
# Each table currently has zero throttle/error datapoints, so the series
# only materialise on first occurrence — NOT_BREACHING keeps them OK until
# then.
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112) The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
common.add_ddb_alarms(
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70) * Add CloudWatch alarm coverage for po-ingest and workorder-ingest Expands alarm coverage across both CDK stacks. All alarms are ALARM-only (no OK action) to the shared site-alerts SNS topic, with TreatMissingData NOT_BREACHING. The site-alerts topic is now imported once near the top of each stack so every alarm reuses one Topic instance. po-ingest (cdk/po_stack.py): - Errors: po-ingest-site-extractor - Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui - Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption), po-ingest-site-extractor, po-web-ui - DynamoDB throttle + system-error: purchase-orders, verified-sites, pending-site-review workorder-ingest (cdk/wo_stack.py): - Throttles: workorder-email-processor - Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption) - DynamoDB throttle + system-error: WorkOrders, WorkOrderComments DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation dimension set, so each table alarm is a Sum math expression across operations via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.259.0). Refs INFRA-41 / audit H-8. * Drop NEEDS ADAM SIGN-OFF wording from alarm comments Duration alarm thresholds are owner-approved; remove the sign-off flag from po_stack.py and wo_stack.py comments. Threshold values, eval config, and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
self, "WorkOrdersTable", work_orders_table, "WorkOrders", alarm_topic
)
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112) The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
common.add_ddb_alarms(
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70) * Add CloudWatch alarm coverage for po-ingest and workorder-ingest Expands alarm coverage across both CDK stacks. All alarms are ALARM-only (no OK action) to the shared site-alerts SNS topic, with TreatMissingData NOT_BREACHING. The site-alerts topic is now imported once near the top of each stack so every alarm reuses one Topic instance. po-ingest (cdk/po_stack.py): - Errors: po-ingest-site-extractor - Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui - Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption), po-ingest-site-extractor, po-web-ui - DynamoDB throttle + system-error: purchase-orders, verified-sites, pending-site-review workorder-ingest (cdk/wo_stack.py): - Throttles: workorder-email-processor - Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption) - DynamoDB throttle + system-error: WorkOrders, WorkOrderComments DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation dimension set, so each table alarm is a Sum math expression across operations via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.259.0). Refs INFRA-41 / audit H-8. * Drop NEEDS ADAM SIGN-OFF wording from alarm comments Duration alarm thresholds are owner-approved; remove the sign-off flag from po_stack.py and wo_stack.py comments. Threshold values, eval config, and orphan-delete notes are unchanged.
2026-06-17 14:46:03 -04:00
self, "WorkOrderComments", comments_table, "WorkOrderComments", alarm_topic
)
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112) The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal).
2026-07-20 14:14:57 -04:00
# --- Function ARN + consumed-table-name outputs (Phase 4, additive) ---
cdk.CfnOutput(
self,
"EmailProcessorFunctionArn",
value=email_processor.function_arn,
description="ARN of the workorder-email-processor Lambda",
)
cdk.CfnOutput(
self,
"WebUiFunctionArn",
value=web_ui.function_arn,
description="ARN of the workorder-web-ui Lambda",
)
cdk.CfnOutput(
self,
"WorkOrdersTableName",
value=work_orders_table.table_name,
description="WorkOrders DynamoDB table",
)
cdk.CfnOutput(
self,
"WorkOrderCommentsTableName",
value=comments_table.table_name,
description="WorkOrderComments DynamoDB table",
)
# Public Function URL removed 2026-06-08 (INFRA-74 / audit C-5): the
# unauthenticated FunctionUrlAuthType.NONE URL was deleted out-of-band
# via CLI. Removing the construct (and its auto-generated Principal:*
# invoke permission) reconciles IaC with the live state.
feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC call-and-be-called effort). Everything ships DARK: both DynamoDB event source mappings deploy enabled=False; activation is a deliberate one-line follow-up PR gated on the SHOC receiver passing the shared HMAC test vectors. - Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments (in-place update, RETAIN + logical IDs untouched; no existing consumers — verified live, neither table had a stream). - workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope -> HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard ordering (parallelization 1, bisect off, retry until 24h age, ReportBatchItemFailures); 429/5xx/timeout block the shard in order, other 4xx park to workorder-shoc-emitter-rejected; ESM failures -> workorder-shoc-emitter-failures (metadata; replay rebuilds from DynamoDB). Echo guard skips write_origin=shoc-write-api. - Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK (alias workorder-ingest-shoc-webhook-kms); cross-account GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy DESTROY deliberately (machine-generated material; avoids the fixed-name RETAIN-orphan deadlock). - workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap, 64-hex keys, kid = UTC %Y-%m-%dT%H. - Alarms (ALARM-only -> site-alerts): emitter errors/throttles/ duration + iterator-age (>=10 min) + failures/rejected queue depth; rotator standard trio. - scripts/replay_shoc_webhooks.py: dry-run-default operator replay (rebuilds from tables, replay:true envelopes). - Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay signing pinned to identical vectors); bundle-consistency AST pins for both new bundles. - README: WO stack + webhook feed section, alarm table, runbooks; removed stale seahaven-slack-bot consumer references.
2026-07-24 14:54:07 -04:00
# =====================================================================
# --- SHOC webhook emitter (docs/shoc-webhook-plan.md) ---
# Realtime work-order feed to the SHOC backend: DynamoDB Streams on the
# two WO tables -> workorder-shoc-emitter -> HMAC-signed HTTPS POST.
# Contract: docs/shoc-webhook-contract.md (Rev 2026-07-23). Built in
# _add_shoc_webhook_emitter (module helper, common.py plain-helper
# style) to keep __init__ under the PLR0915 statement ceiling; the
# stack stays the construct scope, so extraction does not move any
# logical ID.
# =====================================================================
_add_shoc_webhook_emitter(self, work_orders_table, comments_table, alarm_topic)
def _add_shoc_webhook_emitter(stack, work_orders_table, comments_table, alarm_topic):
"""SHOC webhook emitter (docs/shoc-webhook-plan.md Phases 1-4).
Dedicated HMAC CMK + secret with cross-account SHOC read grants, the
30-day rotation Lambda, the two failure queues, the stream-driven emitter
Lambda with its two (dark, enabled=False) event source mappings, and the
full alarm set. `stack` is the construct scope for every child, exactly as
if this code were inline in __init__.
"""
# SHOC consumer principal for the cross-account read grants. EXACT role
# ARN only -- future shoc-backend-staging/-prod roles are each a
# deliberate, individually-reviewed policy addition (no wildcard or
# prefix trust).
shoc_consumer_principal = iam.ArnPrincipal(
"arn:aws:iam::396287094661:role/shoc-backend-dev"
)
# --- Dedicated CMK for the HMAC secret ---
# Dedicated key, NOT alias/seahaven-dynamodb: reusing the DynamoDB CMK
# would hand the SHOC cross-account grant decrypt reach over the PO
# table's encryption key -- the dedicated key scopes the grant to
# exactly this secret (plan Phase 1).
shoc_webhook_key = kms.Key(
stack,
"ShocWebhookHmacKey",
alias="workorder-ingest-shoc-webhook-kms",
description=(
"Dedicated CMK for the workorder-ingest/shoc-webhook-hmac "
"secret (cross-account readable by the SHOC backend)"
),
enable_key_rotation=True,
# GOTCHA: DESTROY is deliberate -- do not "harden" this to RETAIN.
# The key protects only machine-generated HMAC material that is
# fully regenerable by one rotation, and DESTROY avoids the
# fixed-name RETAIN-orphan deadlock on the alias (mirrors the
# secret's rationale below).
removal_policy=RemovalPolicy.DESTROY,
pending_window=Duration.days(7),
)
# Key-policy half of the cross-account read grant (the secret resource
# policy below is the other half; either one alone fails silently at
# the receiver). resources=["*"] is key-scoped, not account-wide --
# KMS key policies only ever apply to this key.
# kms:ViaService pins the grant to Secrets Manager decrypt paths only
# (GPT-4.1 cross-review FIX): a compromised shoc-backend-dev cannot use
# this key for arbitrary KMS operations outside the secret fetch.
feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC call-and-be-called effort). Everything ships DARK: both DynamoDB event source mappings deploy enabled=False; activation is a deliberate one-line follow-up PR gated on the SHOC receiver passing the shared HMAC test vectors. - Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments (in-place update, RETAIN + logical IDs untouched; no existing consumers — verified live, neither table had a stream). - workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope -> HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard ordering (parallelization 1, bisect off, retry until 24h age, ReportBatchItemFailures); 429/5xx/timeout block the shard in order, other 4xx park to workorder-shoc-emitter-rejected; ESM failures -> workorder-shoc-emitter-failures (metadata; replay rebuilds from DynamoDB). Echo guard skips write_origin=shoc-write-api. - Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK (alias workorder-ingest-shoc-webhook-kms); cross-account GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy DESTROY deliberately (machine-generated material; avoids the fixed-name RETAIN-orphan deadlock). - workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap, 64-hex keys, kid = UTC %Y-%m-%dT%H. - Alarms (ALARM-only -> site-alerts): emitter errors/throttles/ duration + iterator-age (>=10 min) + failures/rejected queue depth; rotator standard trio. - scripts/replay_shoc_webhooks.py: dry-run-default operator replay (rebuilds from tables, replay:true envelopes). - Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay signing pinned to identical vectors); bundle-consistency AST pins for both new bundles. - README: WO stack + webhook feed section, alarm table, runbooks; removed stale seahaven-slack-bot consumer references.
2026-07-24 14:54:07 -04:00
shoc_webhook_key.add_to_resource_policy(
iam.PolicyStatement(
actions=["kms:Decrypt"],
principals=[shoc_consumer_principal],
resources=["*"],
conditions={
"StringEquals": {
"kms:ViaService": (f"secretsmanager.{stack.region}.amazonaws.com")
}
},
feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC call-and-be-called effort). Everything ships DARK: both DynamoDB event source mappings deploy enabled=False; activation is a deliberate one-line follow-up PR gated on the SHOC receiver passing the shared HMAC test vectors. - Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments (in-place update, RETAIN + logical IDs untouched; no existing consumers — verified live, neither table had a stream). - workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope -> HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard ordering (parallelization 1, bisect off, retry until 24h age, ReportBatchItemFailures); 429/5xx/timeout block the shard in order, other 4xx park to workorder-shoc-emitter-rejected; ESM failures -> workorder-shoc-emitter-failures (metadata; replay rebuilds from DynamoDB). Echo guard skips write_origin=shoc-write-api. - Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK (alias workorder-ingest-shoc-webhook-kms); cross-account GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy DESTROY deliberately (machine-generated material; avoids the fixed-name RETAIN-orphan deadlock). - workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap, 64-hex keys, kid = UTC %Y-%m-%dT%H. - Alarms (ALARM-only -> site-alerts): emitter errors/throttles/ duration + iterator-age (>=10 min) + failures/rejected queue depth; rotator standard trio. - scripts/replay_shoc_webhooks.py: dry-run-default operator replay (rebuilds from tables, replay:true envelopes). - Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay signing pinned to identical vectors); bundle-consistency AST pins for both new bundles. - README: WO stack + webhook feed section, alarm table, runbooks; removed stale seahaven-slack-bot consumer references.
2026-07-24 14:54:07 -04:00
)
)
# --- HMAC signing secret ---
# Value shape (contract section 6.1):
# {"keys": [{"kid": "<YYYY-MM-DDTHH>", "secret": "<64 hex>"}, ...]},
# newest first, max 2; the producer signs with keys[0]. The
# generate_secret_string below is BOOTSTRAP shape only ({"keys": []}
# plus throwaway entropy the rotator ignores); the first rotation
# (rotate_immediately default) populates the real keys.
shoc_hmac_secret = secretsmanager.Secret(
stack,
"ShocWebhookHmacSecret",
secret_name="workorder-ingest/shoc-webhook-hmac",
encryption_key=shoc_webhook_key,
description=(
"HMAC signing keys for the SHOC work-order webhook "
"(docs/shoc-webhook-contract.md section 6)"
),
# GOTCHA: DESTROY is deliberate -- do not "harden" this to RETAIN.
# The value is machine-generated HMAC material with no operator-set
# content, fully regenerable by one rotation, so RETAIN buys
# nothing and would expose the fixed-name RETAIN orphan deadlock
# (a failed first create orphans an empty shell holding the global
# name; see reference_secret_retain_orphan_deadlock).
removal_policy=RemovalPolicy.DESTROY,
generate_secret_string=secretsmanager.SecretStringGenerator(
secret_string_template='{"keys": []}',
generate_string_key="bootstrap_entropy",
password_length=32,
exclude_punctuation=True,
),
)
cdk.Tags.of(shoc_hmac_secret).add("Purpose", "shoc-webhook-hmac")
cdk.Tags.of(shoc_hmac_secret).add("ManagedBy", "procurement-ingest-cdk")
# Secret-resource-policy half of the cross-account read grant (the key
# policy above is the other half). DescribeSecret lets the receiver
# resolve secret metadata without any broader list permission.
shoc_hmac_secret.add_to_resource_policy(
iam.PolicyStatement(
actions=[
"secretsmanager:GetSecretValue",
"secretsmanager:DescribeSecret",
],
principals=[shoc_consumer_principal],
resources=["*"],
)
)
# --- HMAC rotation Lambda ---
# 30-day schedule: generates a new key, prepends as keys[0], truncates
# to 2 entries. Single-user rotation (receivers re-fetch on a <=5-min
# TTL), so the standard 4-step rotation collapses to
# createSecret/finishSecret.
shoc_hmac_rotator_log_group = common.make_function_log_group(
stack, "ShocHmacRotator", "workorder-shoc-hmac-rotator"
)
shoc_hmac_rotator = lambda_.Function(
stack,
"ShocHmacRotator",
function_name="workorder-shoc-hmac-rotator",
runtime=lambda_.Runtime.PYTHON_3_12,
architecture=lambda_.Architecture.ARM_64,
handler="handler.handler",
code=lambda_.Code.from_asset(
"../lambdas",
exclude=["**/__pycache__/**", "**/tests/**", "**/package/**"],
bundling=cdk.BundlingOptions(
image=lambda_.Runtime.PYTHON_3_12.bundling_image,
command=[
"bash",
"-c",
# Stdlib + boto3-from-runtime only; nothing installed.
"cp wo/shoc_hmac_rotator/*.py /asset-output/",
],
),
),
timeout=Duration.seconds(60),
memory_size=128,
log_group=shoc_hmac_rotator_log_group,
)
# Rotation permissions, scoped to the one secret. CDK's secret_arn
# token resolves to the full ARN including the -?????? suffix wildcard,
# so no separate "*"-suffixed resource variant is needed.
shoc_hmac_rotator.add_to_role_policy(
iam.PolicyStatement(
actions=[
"secretsmanager:DescribeSecret",
"secretsmanager:GetSecretValue",
"secretsmanager:PutSecretValue",
"secretsmanager:UpdateSecretVersionStage",
],
resources=[shoc_hmac_secret.secret_arn],
)
)
# Explicit statement instead of grant_encrypt_decrypt so the grant can
# carry kms:ViaService (cross-review FIX): the rotator only ever touches
# this key through Secrets Manager put/get, never the KMS API directly.
shoc_hmac_rotator.add_to_role_policy(
iam.PolicyStatement(
actions=[
"kms:Decrypt",
"kms:Encrypt",
"kms:GenerateDataKey*",
"kms:ReEncrypt*",
],
resources=[shoc_webhook_key.key_arn],
conditions={
"StringEquals": {
"kms:ViaService": (f"secretsmanager.{stack.region}.amazonaws.com")
}
},
)
)
feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC call-and-be-called effort). Everything ships DARK: both DynamoDB event source mappings deploy enabled=False; activation is a deliberate one-line follow-up PR gated on the SHOC receiver passing the shared HMAC test vectors. - Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments (in-place update, RETAIN + logical IDs untouched; no existing consumers — verified live, neither table had a stream). - workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope -> HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard ordering (parallelization 1, bisect off, retry until 24h age, ReportBatchItemFailures); 429/5xx/timeout block the shard in order, other 4xx park to workorder-shoc-emitter-rejected; ESM failures -> workorder-shoc-emitter-failures (metadata; replay rebuilds from DynamoDB). Echo guard skips write_origin=shoc-write-api. - Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK (alias workorder-ingest-shoc-webhook-kms); cross-account GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy DESTROY deliberately (machine-generated material; avoids the fixed-name RETAIN-orphan deadlock). - workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap, 64-hex keys, kid = UTC %Y-%m-%dT%H. - Alarms (ALARM-only -> site-alerts): emitter errors/throttles/ duration + iterator-age (>=10 min) + failures/rejected queue depth; rotator standard trio. - scripts/replay_shoc_webhooks.py: dry-run-default operator replay (rebuilds from tables, replay:true envelopes). - Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay signing pinned to identical vectors); bundle-consistency AST pins for both new bundles. - README: WO stack + webhook feed section, alarm table, runbooks; removed stale seahaven-slack-bot consumer references.
2026-07-24 14:54:07 -04:00
shoc_hmac_secret.add_rotation_schedule(
"Rotation",
rotation_lambda=shoc_hmac_rotator,
automatically_after=Duration.days(30),
)
# --- Standard per-Lambda alarms: workorder-shoc-hmac-rotator ---
# errors + throttles + p99 duration. No DLQ alarm: rotation is invoked
# synchronously by Secrets Manager (dlq=None); a failed rotation
# surfaces as an invocation error.
common.add_standard_lambda_alarms(
stack,
"ShocHmacRotator",
shoc_hmac_rotator,
"workorder-shoc-hmac-rotator",
alarm_topic,
duration_statistic="p99",
errors=True,
dlq=None,
descriptions={
"errors": "workorder-shoc-hmac-rotator invocation errors",
"throttles": "workorder-shoc-hmac-rotator invocation throttles",
"duration": (
"workorder-shoc-hmac-rotator p99 duration approaching the 60s timeout"
),
},
)
# --- Emitter failure queues ---
# Failures queue: ESM on_failure destination. It receives ESM failure
# METADATA (shard/sequence pointers), not full payloads -- replay
# rebuilds events from DynamoDB (contract section 8).
shoc_emitter_failures_queue = sqs.Queue(
stack,
"ShocEmitterFailuresQueue",
queue_name="workorder-shoc-emitter-failures",
retention_period=Duration.days(14),
enforce_ssl=True,
)
# Rejected queue: full {envelope, response_status} payloads parked by
# the handler on non-retryable 4xx responses (contract section 7).
shoc_emitter_rejected_queue = sqs.Queue(
stack,
"ShocEmitterRejectedQueue",
queue_name="workorder-shoc-emitter-rejected",
retention_period=Duration.days(14),
enforce_ssl=True,
)
# --- Emitter Lambda ---
shoc_emitter_log_group = common.make_function_log_group(
stack, "ShocEmitter", "workorder-shoc-emitter"
)
shoc_emitter = lambda_.Function(
stack,
"ShocEmitter",
function_name="workorder-shoc-emitter",
runtime=lambda_.Runtime.PYTHON_3_12,
architecture=lambda_.Architecture.ARM_64,
handler="handler.handler",
code=lambda_.Code.from_asset(
"../lambdas",
exclude=["**/__pycache__/**", "**/tests/**", "**/package/**"],
bundling=cdk.BundlingOptions(
image=lambda_.Runtime.PYTHON_3_12.bundling_image,
command=[
"bash",
"-c",
# Stdlib HTTP (urllib.request) + boto3-from-runtime
# only; nothing installed.
"cp wo/shoc_emitter/*.py /asset-output/",
],
),
),
timeout=Duration.seconds(60),
memory_size=256,
log_group=shoc_emitter_log_group,
environment={
# Non-sensitive endpoint URL (HMAC is the auth, the URL is
# not). path TBD by SHOC -- confirmed in the activation PR.
"SHOC_WEBHOOK_URL": (
"https://api.dev.seahaven.com/api/webhooks/work-orders"
),
"HMAC_SECRET_ARN": shoc_hmac_secret.secret_arn,
"REJECTED_QUEUE_URL": shoc_emitter_rejected_queue.queue_url,
},
)
work_orders_table.grant_stream_read(shoc_emitter)
comments_table.grant_stream_read(shoc_emitter)
shoc_hmac_secret.grant_read(shoc_emitter)
# Explicit statement instead of grant_decrypt so the grant carries
# kms:ViaService (cross-review FIX): the emitter only decrypts this key
# through Secrets Manager GetSecretValue.
shoc_emitter.add_to_role_policy(
iam.PolicyStatement(
actions=["kms:Decrypt"],
resources=[shoc_webhook_key.key_arn],
conditions={
"StringEquals": {
"kms:ViaService": (f"secretsmanager.{stack.region}.amazonaws.com")
}
},
)
)
feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC call-and-be-called effort). Everything ships DARK: both DynamoDB event source mappings deploy enabled=False; activation is a deliberate one-line follow-up PR gated on the SHOC receiver passing the shared HMAC test vectors. - Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments (in-place update, RETAIN + logical IDs untouched; no existing consumers — verified live, neither table had a stream). - workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope -> HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard ordering (parallelization 1, bisect off, retry until 24h age, ReportBatchItemFailures); 429/5xx/timeout block the shard in order, other 4xx park to workorder-shoc-emitter-rejected; ESM failures -> workorder-shoc-emitter-failures (metadata; replay rebuilds from DynamoDB). Echo guard skips write_origin=shoc-write-api. - Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK (alias workorder-ingest-shoc-webhook-kms); cross-account GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy DESTROY deliberately (machine-generated material; avoids the fixed-name RETAIN-orphan deadlock). - workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap, 64-hex keys, kid = UTC %Y-%m-%dT%H. - Alarms (ALARM-only -> site-alerts): emitter errors/throttles/ duration + iterator-age (>=10 min) + failures/rejected queue depth; rotator standard trio. - scripts/replay_shoc_webhooks.py: dry-run-default operator replay (rebuilds from tables, replay:true envelopes). - Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay signing pinned to identical vectors); bundle-consistency AST pins for both new bundles. - README: WO stack + webhook feed section, alarm table, runbooks; removed stale seahaven-slack-bot consumer references.
2026-07-24 14:54:07 -04:00
shoc_emitter_rejected_queue.grant_send_messages(shoc_emitter)
# ------------------------------------------------------------------
# SHIPS DARK: both event source mappings deploy with enabled=False,
# deliberately. The full stack (Lambda, queues, alarms, secret,
# rotation) deploys and is testable with zero deliveries while SHOC
# has no receiver, so our merge cadence never depends on Luby's.
# Activation is a one-line enabled=True PR gated on the SHOC receiver
# passing the shared HMAC test vectors (plan Phase 3). LATEST start
# position => no historical flood at activation; SHOC loads history
# via the procurement read API instead.
# Ordering knobs: parallelization_factor=1, bisect_batch_on_error=
# False and retry_attempts=-1 (retry until the 24h record age) are
# REQUIRED for strict per-work-order in-order delivery -- a retryable
# failure blocks the shard rather than skipping ahead, and
# report_batch_item_failures keeps earlier in-batch successes from
# being re-delivered.
# ------------------------------------------------------------------
shoc_emitter.add_event_source(
lambda_event_sources.DynamoEventSource(
work_orders_table,
starting_position=lambda_.StartingPosition.LATEST,
batch_size=10,
bisect_batch_on_error=False,
retry_attempts=-1,
max_record_age=Duration.hours(24),
parallelization_factor=1,
report_batch_item_failures=True,
enabled=False,
on_failure=lambda_event_sources.SqsDlq(shoc_emitter_failures_queue),
)
)
shoc_emitter.add_event_source(
lambda_event_sources.DynamoEventSource(
comments_table,
starting_position=lambda_.StartingPosition.LATEST,
batch_size=10,
bisect_batch_on_error=False,
retry_attempts=-1,
max_record_age=Duration.hours(24),
parallelization_factor=1,
report_batch_item_failures=True,
enabled=False,
on_failure=lambda_event_sources.SqsDlq(shoc_emitter_failures_queue),
)
)
# --- Standard per-Lambda alarms: workorder-shoc-emitter ---
# errors + throttles + p99 duration. No DLQ alarm here: the emitter is
# a stream consumer with no async DLQ (dlq=None); its failure surfaces
# are the two SQS queues alarmed bespoke below.
common.add_standard_lambda_alarms(
stack,
"ShocEmitter",
shoc_emitter,
"workorder-shoc-emitter",
alarm_topic,
duration_statistic="p99",
errors=True,
dlq=None,
descriptions={
"errors": "workorder-shoc-emitter invocation errors",
"throttles": "workorder-shoc-emitter invocation throttles",
"duration": (
"workorder-shoc-emitter p99 duration approaching the 60s timeout"
),
},
)
# --- Bespoke emitter alarms (plan Phase 4) ---
# These don't fit add_standard_lambda_alarms' shape and stay bespoke.
# Iterator age >= 10 min sustained means SHOC is likely down and the
# shard is blocking (exactly the ordered-backpressure design working);
# the two SQS-visible alarms page the operator replay runbook.
shoc_emitter.metric(
"IteratorAge",
statistic="Maximum",
period=Duration.minutes(5),
).create_alarm(
stack,
"ShocEmitterIteratorAgeAlarm",
alarm_name="workorder-shoc-emitter-iterator-age",
alarm_description=(
"workorder-shoc-emitter stream lag >= 10 min "
"(SHOC receiver likely down; shard blocking on retries)"
),
threshold=600000,
evaluation_periods=3,
datapoints_to_alarm=2,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
shoc_emitter_failures_queue.metric_approximate_number_of_messages_visible(
period=Duration.minutes(5),
statistic="Maximum",
).create_alarm(
stack,
"ShocEmitterFailuresMessagesAlarm",
alarm_name="workorder-shoc-emitter-failures-messages",
alarm_description=(
"workorder-shoc-emitter retry-exhausted stream records parked "
"(ESM failure metadata; replay rebuilds from DynamoDB)"
),
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
shoc_emitter_rejected_queue.metric_approximate_number_of_messages_visible(
period=Duration.minutes(5),
statistic="Maximum",
).create_alarm(
stack,
"ShocEmitterRejectedMessagesAlarm",
alarm_name="workorder-shoc-emitter-rejected-messages",
alarm_description=(
"workorder-shoc-emitter parked non-retryable 4xx deliveries "
"(contract bug; inspect payloads and replay)"
),
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
cdk.CfnOutput(
stack,
"ShocWebhookHmacSecretArn",
value=shoc_hmac_secret.secret_arn,
description=(
"SHOC webhook HMAC secret ARN -- hand off to Luby (SHOC team) "
"for the cross-account receiver fetch"
),
)