chore(infra): remove cdk tree after hcp cutover (PLAT-89) (#164)

* chore(infra): remove cdk tree after hcp cutover

Delete retired CDK sources, retarget bundle/principal contract tests to
Terraform packaging, disable CDK synth in CI, and scrub deploy-adjacent docs.

* fix(test): restore exact SHOC principal pin in terraform

Pin shoc_consumer_role_arn's Terraform default and example to the trusted
ARN, and require grant sites to consume local.shoc_consumer_role_arn only.
This commit is contained in:
Adam Moussa 2026-08-07 12:20:29 -04:00 • committed by GitHub
parent e54e53d0e8
commit 5a3729c583
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
18 changed files with 467 additions and 3108 deletions

View file

@ -1,4 +1,3 @@
[run]
omit =
*/tests/*
*/cdk.out/*

View file

@ -1,16 +1,5 @@
version: 2
updates:
- package-ecosystem: "pip"
directory: "/cdk"
schedule:
interval: "weekly"
groups:
minor-and-patch:
update-types:
- "minor"
- "patch"
commit-message:
prefix: "chore(deps)"
- package-ecosystem: "pip"
directory: "/lambdas/po/email_processor"
schedule:
@ -99,3 +88,14 @@ updates:
- "patch"
commit-message:
prefix: "chore(deps)"
- package-ecosystem: "terraform"
directory: "/terraform"
schedule:
interval: "weekly"
groups:
minor-and-patch:
update-types:
- "minor"
- "patch"
commit-message:
prefix: "chore(deps)"

View file

@ -10,9 +10,10 @@ jobs:
ci:
uses: Sea-Haven-Industries/.github/.github/workflows/ci-python-sam.yaml@3f746774229d41770727e2e4fd63ed5f5555a8b3 # v1.0.3
with:
source-dirs: "lambdas cdk tests scripts"
source-dirs: "lambdas tests scripts"
run-tests: true
run-cdk-synth: true
# Deploy path is HCP Terraform; reusable name still covers Python lint/tests.
run-cdk-synth: false
run-sam-validate: false
# The published OpenAPI contract must stay compliant with redocly.yaml;
# this is the CI mirror of the local `npm run lint:api`.

1
.gitignore vendored
View file

@ -6,7 +6,6 @@ build/
.venv/
venv/
node_modules/
cdk.out/
.env
*.eml
# Test fixtures: scrubbed Hexagon sample emails are committed so the WO template

View file

@ -111,7 +111,7 @@ A read-only REST API (API Gateway + the `procurement-api` Lambda, `lambdas/api/`
- **Spec is source of truth:** `lambdas/api/openapi.json` (OpenAPI 3.1). Its top-level `webhooks` section documents the outbound SHOC work-order feed, so one page describes both directions (call + be-called). `tests/test_api_spec_drift.py` pins the spec's paths to the router table, so spec and implementation cannot drift.
- **Auth:** data routes use API Gateway `AWS_IAM` (SigV4) plus a resource policy allowing exactly `arn:aws:iam::396287094661:role/shoc-backend-dev` on `GET/*`; same-account admin callers authorize via identity policy (Postman signs SigV4 natively). Docs routes are auth `NONE` at the gateway (resource-policy carve-out for exactly those two GETs) but the handler fails closed on the shared token (`lambdas/shared/web_ui_auth.py`, secret `procurement-ingest/web-ui-auth-token`) — not an unauthenticated data path (INFRA-74 posture).
- **Custom domain:** `https://procurement-api.seahaven.com` (REGIONAL API Gateway domain, TLS 1.2, empty base-path mapping to the `prod` stage, so callers hit `/work-orders` with no `/prod` segment). The stable SHOC-facing endpoint; the raw `*.execute-api.us-east-1.amazonaws.com/prod` URL still works. **Cross-account DNS:** the `seahaven.com` public zone is in the mgmt account (`328440206208`), so the ACM cert's validation record and the A-alias are added there out of band — `scripts/setup_procurement_api_domain.sh cert` issues the cert (DNS-validated against the mgmt zone) and writes its ARN to prod SSM `/procurement-api/custom-domain/certificate-arn`, which the stack reads (`value_for_string_parameter`, CFN-resolved at deploy); after `cdk deploy procurement-api`, `scripts/setup_procurement_api_domain.sh alias` adds the A-alias from the stack's `ProcurementApiAliasTarget`/`ProcurementApiAliasHostedZoneId` outputs. SigV4 is unaffected (same underlying API id + resource policy); the docs page injects whichever host served the request into `servers[0].url`.
- **Custom domain:** `https://procurement-api.seahaven.com` (REGIONAL API Gateway domain, TLS 1.2, empty base-path mapping to the `prod` stage, so callers hit `/work-orders` with no `/prod` segment). The stable SHOC-facing endpoint; the raw `*.execute-api.us-east-1.amazonaws.com/prod` URL still works. **Cross-account DNS:** the `seahaven.com` public zone is in the mgmt account (`328440206208`), so the ACM cert's validation record and the A-alias are added there out of band — `scripts/setup_procurement_api_domain.sh cert` issues the cert (DNS-validated against the mgmt zone) and writes its ARN to prod SSM `/procurement-api/custom-domain/certificate-arn`, which Terraform reads (`data.aws_ssm_parameter` in `terraform/api.tf`); after an HCP apply that owns the API Gateway DomainName, `scripts/setup_procurement_api_domain.sh alias` adds the A-alias from API Gateway `get-domain-name` (`regionalDomainName` / `regionalHostedZoneId`). SigV4 is unaffected (same underlying API id + resource policy); the docs page injects whichever host served the request into `servers[0].url`.
- **KMS:** `purchase-orders` is CMK-encrypted; the imported-by-name table doesn't carry the key association, so the stack grants `kms:Decrypt`/`DescribeKey` on the CMK from SSM `/seahaven/dynamodb/cmk-arn` explicitly (the INFRA-104 failure class).
- **No access logging in v1** (keeps the `?token=` shim out of any log and avoids the account-level API Gateway CloudWatch role); rotate the docs token before ever enabling it. No CORS (server-to-server + Postman callers).
- **Alarms:** `procurement-api-errors`/`-throttles`/`-duration` (p99 ≥ 22.5 s) + gateway `procurement-api-5xx`, all ALARM-only → `site-alerts`. No 4XX alarm (401/403/404 are expected traffic).
@ -119,37 +119,31 @@ A read-only REST API (API Gateway + the `procurement-api` Lambda, `lambdas/api/`
## Architecture
**IaC:** Terraform under `terraform/` (HCP remote apply, Manual until sealed). Workspace `procurement-ingest-prod` in project `seahaven-prod` is the sole mutate path for live resources in seahaven-prod `011934824531` / `us-east-1` (PLAT-86 import-in-place). Former CDK stacks (`po-ingest`, `WorkorderIngestStack`, `procurement-api`) are historical; `cdk/` remains for reference/synth until a follow-up hygiene pass. Never apply against mgmt `328440206208` — PLAT-67 left cold-archive RETAIN leftovers there only.
**IaC:** Terraform under `terraform/` (HCP remote apply, Manual until sealed). Workspace `procurement-ingest-prod` in project `seahaven-prod` is the sole mutate path for live resources in seahaven-prod `011934824531` / `us-east-1` (PLAT-86 import-in-place). Former CDK stacks (`po-ingest`, `WorkorderIngestStack`, `procurement-api`) are historical and removed from the tree (PLAT-89). Never apply against mgmt `328440206208` — PLAT-67 left cold-archive RETAIN leftovers there only.
All Lambdas: Python 3.12, ARM64, 60-day log retention.
**LLM provider — Amazon Bedrock.** Both email processors call Claude Haiku 4.5 through the Bedrock inference profile `us.anthropic.claude-haiku-4-5-20251001-v1:0` (`bedrock-runtime` `InvokeModel`), configured via the `BEDROCK_MODEL_ID` env var. There is **no Anthropic API key and no Secrets Manager secret** any more. Each processor role is granted `bedrock:InvokeModel` + `bedrock:InvokeModelWithResponseStream` on **both** the inference-profile ARN **and** the per-region foundation-model ARNs for `us-east-1` / `us-east-2` / `us-west-2` (empty-account foundation-model ARNs) — the `us.*` profile can route cross-region, so a profile-only grant would `AccessDenied` at runtime under load.
> **Retired secrets (manual cleanup outstanding):** the former secrets `po-ingest/anthropic-api-key` and `workorder-ingest/anthropic-api-key` had `RemovalPolicy.RETAIN`, so removing them from the CDK stacks **orphans** them rather than deleting them. Delete both by hand post-deploy and revoke the stored keys at the provider. The processors no longer read any `ANTHROPIC_API_KEY_SECRET_ARN` — authentication to Bedrock is via the Lambda execution-role IAM grant, so there is no provider API key or Secrets Manager fetch on the parse path.
> **Retired secrets (manual cleanup outstanding):** the former secrets `po-ingest/anthropic-api-key` and `workorder-ingest/anthropic-api-key` were retained when removed from IaC, so they were **orphaned** rather than deleted. Delete both by hand and revoke the stored keys at the provider. The processors no longer read any `ANTHROPIC_API_KEY_SECRET_ARN` — authentication to Bedrock is via the Lambda execution-role IAM grant, so there is no provider API key or Secrets Manager fetch on the parse path.
**SES:** Both stacks add rules to the shared `INBOUND_MAIL` receipt rule set on `int.seahaven.com`.
**SES:** Both pipelines add rules to the shared `INBOUND_MAIL` receipt rule set on `int.seahaven.com`.
### Shared CDK helpers (`cdk/common.py`, Phase 4)
### Shared alarms and packaging helpers
The ~379 lines that `po_stack.py` and `wo_stack.py` both defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock `InvokeModel` grant, the raw-email bucket, the async-invoke DLQ, and the template-fallback-rate math alarm) are collapsed into `cdk/common.py`.
DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock `InvokeModel` grant, the raw-email bucket, the async-invoke DLQ, and the template-fallback-rate math alarm live under `terraform/` (`po_*.tf`, `wo_*.tf`, `api*.tf`). Lambda zip packaging is `terraform/build_packages.sh` (non-recursive `copy_py_dir` + `copy_shared_all` / selective `copy_src_file`), gated by `tests/test_bundle_consistency.py`.
**Logical-ID-safety rule (load-bearing).** Every helper is a **plain function** taking `(scope, id, ...)` — never a `Construct` subclass. Each stack calls a helper with its **own Stack instance as `scope`** and the **exact same literal construct id** it used inline before the extraction, so every synthesized logical ID is byte-stable. A `Construct` subclass would insert an extra tree node, reparent every child's logical ID, and CloudFormation would attempt to **replace** the `RemovalPolicy.RETAIN`-protected `purchase-orders` / `WorkOrders` / `WorkOrderComments` tables and the named S3 buckets — a data-loss event. Nothing in `common.py` subclasses `Construct`, and it imports only the CDK constructs its helpers touch (`dynamodb`, `cloudwatch`/`cw_actions`, `iam`, `logs`, `s3`, `sqs` — no `kms`, `ssm`, or Lambda event-source imports).
Extracted helpers: `add_ddb_alarms`, `add_sender_auth_rejected_alarm`, `add_standard_lambda_alarms` (bespoke `descriptions=` dict passed through verbatim per call site — no alarm text is generated or homogenized), `make_bedrock_invoke_statement`, `make_email_bucket`, `make_processor_dlq`, `make_fallback_rate_alarm`. Per-function alarm variance is preserved exactly through call-site arguments, not baked into the helpers: PO's email-processor duration alarm uses `p99`, WO's uses `p95`; `po-web-ui` gets throttles + duration only (no errors, no DLQ); `po-ingest-site-extractor` has no DLQ alarm (it's a DynamoDB-stream consumer, not async-invoked); `workorder-web-ui` gets **zero** alarms — the helper is simply never called for it, so it cannot silently add any. The two "AI-fallback-rejected" alarms (PO's 6-hour count-floor `IF(FILL(rej,0)>=1,...)`, WO's 5-minute/2-of-6 sparse idiom) are a different shape from `make_fallback_rate_alarm` and stay as distinct inline call sites in each stack rather than being forced through the shared helper.
**Bedrock ARN, now account/region-derived.** `make_bedrock_invoke_statement(scope)` builds the inference-profile ARN and the `us-east-1` foundation-model ARN from `Stack.of(scope).account` / `.region` instead of the previous hardcoded `328440206208`/`us-east-1` literals (the `us-east-2`/`us-west-2` foundation-model ARNs stay literal — they're fixed cross-region reach targets, not the stack's own region). Because the environment is account-agnostic (above), `stack.account` is the `AWS::AccountId` pseudo-parameter, so the derived ARN synthesizes as an `Fn::Sub`/`Ref` that **resolves at deploy time to the same ARN** the hardcoded literal named in-account. Versus the deployed stack this is a benign **in-place** IAM policy update (IAM policies never require replacement) — and it makes the grant account-portable instead of pinned to the management account. This IAM PolicyStatement change is the surface the mandatory GPT-4.1 cross-family review covers.
**New `CfnOutput`s.** Each stack now exports its Lambda function ARNs and the DynamoDB table names it owns/consumes, all net-new/additive (no existing output changes): `po-ingest` — `EmailProcessorFunctionArn`, `WebUiFunctionArn`, `SiteExtractorFunctionArn`, `PurchaseOrdersTableName`, `PendingSiteReviewTableName` (the existing `VerifiedSitesTableName` output is unchanged); `workorder-ingest` — `EmailProcessorFunctionArn`, `WebUiFunctionArn`, `WorkOrdersTableName`, `WorkOrderCommentsTableName`.
Per-function alarm variance is preserved: PO's email-processor duration alarm uses `p99`, WO's uses `p95`; `po-web-ui` gets throttles + duration only (no errors, no DLQ); `po-ingest-site-extractor` has no DLQ alarm (it's a DynamoDB-stream consumer, not async-invoked); `workorder-web-ui` gets **zero** alarms.
### Sender authentication (INFRA-107)
The `From` header and any `Authentication-Results` header inside the raw MIME are attacker-forgeable, so neither is trusted. Instead, both email processors authenticate the sender against the verdicts SES itself stamps at delivery time, failing closed. The authenticator is **single-sourced** at `lambdas/shared/ses_auth.py` (Phase 3) — previously duplicated byte-for-byte in each pipeline's `email_processor/` dir and kept in sync by a fixture-hygiene test; now one copy, so a future hardening fix to the fail-closed logic lands **once** instead of needing two identical edits. The CDK bundling `cp`s it flat beside each handler so the handlers' unchanged `from ses_auth import authenticate_inbound_email` still resolves at runtime (see [`lambdas/shared/`](#deploy-pipeline-guards-phase-0)). The gate:
The `From` header and any `Authentication-Results` header inside the raw MIME are attacker-forgeable, so neither is trusted. Instead, both email processors authenticate the sender against the verdicts SES itself stamps at delivery time, failing closed. The authenticator is **single-sourced** at `lambdas/shared/ses_auth.py` (Phase 3) — previously duplicated byte-for-byte in each pipeline's `email_processor/` dir and kept in sync by a fixture-hygiene test; now one copy, so a future hardening fix to the fail-closed logic lands **once** instead of needing two identical edits. Packaging copies it flat beside each handler (`copy_shared_all` in `terraform/build_packages.sh`) so the handlers' unchanged `from ses_auth import authenticate_inbound_email` still resolves at runtime (see [`lambdas/shared/`](#deploy-pipeline-guards-phase-0)). The gate:
1. Take **only the topmost** `Authentication-Results` header (SES prepends its trace headers; any lower copies arrived inside the message and are ignored).
2. Require its authserv-id to be `amazonses.com`.
3. Require a `dkim=pass` clause whose `header.d=`/`header.i=` domain is in the pipeline's allowlist.
The allowlist is the `ALLOWED_DKIM_DOMAINS` Lambda environment variable (comma-separated, set per stack in CDK — no code change needed to adjust):
The allowlist is the `ALLOWED_DKIM_DOMAINS` Lambda environment variable (comma-separated, set per pipeline in Terraform — no code change needed to adjust):
| Pipeline | `ALLOWED_DKIM_DOMAINS` | Why |
|---|---|---|
@ -162,11 +156,11 @@ The allowlist is the `ALLOWED_DKIM_DOMAINS` Lambda environment variable (comma-s
On any failure (env var unset, header missing/unparseable, verdict fail, unaligned domain) the processor logs a structured `sender_auth_rejected` warning with the reason and S3 key, skips the email, and returns normally — rejected mail never triggers Lambda retries or DLQ messages, but the `<fn>-sender-auth-rejected` CloudWatch alarm (see [CloudWatch alarms](#cloudwatch-alarms)) pages on a rejection spike so a drift-induced outage is not silent. Unit tests live in `tests/test_ses_auth.py`.
**Failure handling (INFRA-41):** Each email-processor is async-invoked (S3 → Lambda). Both have a CDK-managed SQS dead-letter queue (`dead_letter_queue=`, 14-day retention, SSL-enforced) so a failed parse is captured rather than silently dropped after Lambda's retries.
**Failure handling (INFRA-41):** Each email-processor is async-invoked (S3 → Lambda). Both have a Terraform-managed SQS dead-letter queue (14-day retention, SSL-enforced) so a failed parse is captured rather than silently dropped after Lambda's retries.
### CloudWatch alarms
Every alarm is **ALARM-only** (no OK action), sends to the shared `site-alerts` SNS topic (imported once per stack via `Topic.from_topic_arn`), and uses `TreatMissingData.NOT_BREACHING`.
Every alarm is **ALARM-only** (no OK action), sends to the shared `site-alerts` SNS topic, and uses treat-missing-data not-breaching.
**Lambda alarms** (`AWS/Lambda`, `FunctionName` dimension):
@ -177,7 +171,7 @@ Every alarm is **ALARM-only** (no OK action), sends to the shared `site-alerts`
| `<fn>-duration` | `po-email-processor`, `po-ingest-site-extractor`, `po-web-ui`, `workorder-shoc-emitter`, `workorder-shoc-hmac-rotator` (p99); `workorder-email-processor` (p95) | `Duration` percentile, 5 min, `>= 45000` ms (75% of the 60s timeout), eval 3 / datapoints 2 |
| `<fn>-sender-auth-rejected` | `po-email-processor`, `workorder-email-processor` | Log-metric-filter count (namespace `Seahaven/ProcurementIngest`, `default_value=0`) on `sender_auth_rejected` warnings, `Sum` 5 min, `>= 1`, eval 3 / datapoints 2 |
The `<fn>-sender-auth-rejected` alarm closes the silent-drop gap in INFRA-107: a rejected email returns normally (no error, no retry, no DLQ message), so without a log-metric filter a signing-domain drift or a wrong allowlist would discard 100% of legitimate mail while every other alarm stayed green. It counts `sender_auth_rejected` warnings per 5-minute period (`default_value=0` keeps the series continuous) and pages when 2 of the last 3 periods each see at least one rejection — a lone stray spoof probe to the internal ingest address self-clears, but a sustained false-reject storm pages within ~10–15 minutes even at low mail volume; the config is easy to tune in the CDK helper. (A residual gap remains for a *very* sparse total-reject outage — see the SES-AR-01/02 hardening issue.)
The `<fn>-sender-auth-rejected` alarm closes the silent-drop gap in INFRA-107: a rejected email returns normally (no error, no retry, no DLQ message), so without a log-metric filter a signing-domain drift or a wrong allowlist would discard 100% of legitimate mail while every other alarm stayed green. It counts `sender_auth_rejected` warnings per 5-minute period (`default_value=0` keeps the series continuous) and pages when 2 of the last 3 periods each see at least one rejection — a lone stray spoof probe to the internal ingest address self-clears, but a sustained false-reject storm pages within ~10–15 minutes even at low mail volume; the config is easy to tune in Terraform. (A residual gap remains for a *very* sparse total-reject outage — see the SES-AR-01/02 hardening issue.)
The `<fn>-duration` and `<fn>-throttles` alarms for `po-email-processor` and `workorder-email-processor` supersede the orphaned, CLI-created `Lambda-Duration-*` / `Lambda-Throttles-*` alarms (deleted post-deploy).
@ -219,22 +213,7 @@ This placement is deliberate, not incidental: real mail always arrives as an S3
The script runs `set -euo pipefail` and exits non-zero on any invoke failure, any `FunctionError`, or a payload mismatch, failing the apply gate.
**PO bundling: glob replaces the hand-maintained allowlist.** `cdk/po_stack.py`'s asset bundling command ships PO's Lambda source with a non-recursive glob instead of a hand-maintained list of filenames (`cp handler.py ses_auth.py template_parser.py derived_fields.py /asset-output/`). The glob is functionally identical for today's file set — non-recursive, so `tests/` and other subdirectories are still excluded — but structurally eliminates the failure mode that shipped a broken bundle twice (PR #105 omitted `template_parser.py`; PR #2 nearly omitted `derived_fields.py`): a new sibling module the handler imports now ships automatically instead of requiring someone to remember to add it to the list.
**Phase 2: widened asset root, both processors on the glob.** Both `cdk/po_stack.py` and `cdk/wo_stack.py` widen their bundled email-processor's `Code.from_asset` root from the per-pipeline dir (`../lambdas/po/email_processor`, `../lambdas/wo/email_processor`) to the shared parent, `../lambdas` — the prerequisite for the Phase 3 `lambdas/shared/` extraction, which needs a bundling root able to reach a sibling `shared/` package outside either pipeline's own dir (this move ships **zero handler code changes** — `git diff -- lambdas/` is empty for this PR). With bundling present, CDK mounts the asset root as the container's working directory, so both the `pip install -r` path and the `cp` source operand became repo-relative to `lambdas/`: `-r po/email_processor/requirements.txt` (resp. `wo/email_processor/requirements.txt`) and `cp po/email_processor/*.py /asset-output/` (resp. `cp wo/email_processor/*.py /asset-output/`). The pip `--platform manylinux2014_aarch64 --only-binary=:all:` pin — removing it once shipped x86 wheels into the ARM64 function and caused a total outage (PR #34) — is preserved byte-for-byte on both.
Both bundled `from_asset` calls also gain `exclude=['**/__pycache__/**', '**/tests/**', '**/package/**']`. This is load-bearing, not cosmetic: `Code.from_asset` does not honor `.gitignore`, and widening the root to `../lambdas` means the untracked, 44 MB `lambdas/po/email_processor/package/` dir (a stale vendored dependency tree; deletion is a separate, deliberate call — not part of this change) would otherwise be staged into the *source fingerprint* `from_asset` hashes to decide whether to re-bundle. Because CI never has that local-only directory, an un-excluded root would diverge the local vs. CI asset hash on every synth/deploy and force spurious redeploys; the `**/tests/**` and `__pycache__` excludes keep the hash stable for the same reason. Note the exclude does **not** decouple the two pipelines' asset hashes: `from_asset` hashes with its default `AssetHashType.SOURCE`, so the fingerprint is computed over *all* of `../lambdas` minus only the excluded `__pycache__`/`tests`/`package` paths — PO's and WO's first-party source (both `email_processor` trees, plus the two `web_ui`s and the `site_extractor`) therefore both feed **both** email-processors' hash. Editing any non-excluded file under `lambdas/` changes both email-processors' source fingerprint and redeploys both functions with byte-identical bundles. That coupling is an accepted cost of the shared-root design (the bundling `cp` glob still copies only each pipeline's own `*.py` into the zip); the excludes exist solely to strip local-only/irrelevant cruft that would diverge local vs. CI, not to isolate PO's tree from WO's — which `SOURCE` hashing cannot do here.
**WO bundling: glob replaces the whole-dir copy — deliberate prod-zip shrinkage.** WO's bundling command changes from a recursive `cp -r . /asset-output/` (the entire `wo/email_processor/` source dir, copied into the deployed zip) to the same scoped, non-recursive glob PO uses: `cp wo/email_processor/*.py /asset-output/`. This intentionally drops from the production zip:
- `requirements.txt` — needed only at bundle time (`pip install -r ...`), never at runtime;
- the entire `tests/` tree (`lambdas/wo/email_processor/tests/`) — real scrubbed `.eml` fixtures, golden JSON, and test modules;
- any first-party `__pycache__/*.pyc` a local `cp -r .` would have picked up (the currently-deployed zip carries none, but the exclude keeps future local builds equally clean).
This is cleanup, not a regression: none of those file classes are imported at runtime by `handler.handler`, so the acceptance bar for this change on WO is "the runtime-imported module set is unchanged, plus a post-deploy smoke pass" — not a byte-identical zip diff (that stricter bar applies to PO only, whose deployed zip was already this tight before this change). All four first-party top-level `.py` files WO's handler needs — `__init__.py`, `handler.py`, `ses_auth.py`, `template_parser.py` — are still shipped; the glob retains `__init__.py` because it is itself a top-level `.py` file, not a special case requiring a separate copy rule.
**Plain (non-bundled) `from_asset` calls gain `exclude` too.** The three non-bundled Lambda assets — `po-web-ui`, `po-ingest-site-extractor`, `workorder-web-ui` — each add `exclude=['**/__pycache__/**']`. Nothing else about these three changes: each keeps its own scoped asset path (`../lambdas/po/web_ui`, etc.) rather than widening to `../lambdas`, and none gains bundling. Without the exclude, a developer's local `__pycache__` — again invisible to `from_asset`'s `.gitignore`-blind staging — makes that function's asset hash nondeterministic across machines and forces spurious redeploys.
`tests/test_bundle_consistency.py` guards all of the above with a pure-AST check (no synth, no boto3, no handler import): it parses each handler.py's top-level first-party sibling imports, extracts the bundling `command=[...]` string from the corresponding CDK stack file, and asserts every required sibling module is guaranteed to ship. It recognizes both the scoped glob (`cp po/email_processor/*.py` / `cp wo/email_processor/*.py`, with or without a path prefix) and a whole-dir recursive copy (`cp -r . /asset-output/`) as unconditionally-safe shapes, and falls back to literal filename matching for any other (allowlist-style) shape. It pins each stack's command to the scoped-glob form specifically — a future revert to a narrowed single-file copy, a commented-out glob, or a filename allowlist missing a sibling all fail CI loudly instead of silently shipping a broken bundle. Runs in the existing pytest step, before synth.
**Lambda packaging (`terraform/build_packages.sh`).** HCP plan/apply packages each function via non-recursive `copy_py_dir` (top-level `*.py` only, so `tests/` is excluded) plus `copy_shared_all` for the email processors and selective `copy_src_file` for web UI auth / API static assets. That eliminates the failure mode that shipped a broken bundle twice (PR #105 omitted `template_parser.py`; PR #2 nearly omitted `derived_fields.py`): a new sibling module the handler imports ships automatically when added under the pipeline dir, and `tests/test_bundle_consistency.py` pins the exact-set + ships-all contracts against the packaging script (no Terraform plan required). Pip installs for third-party deps still use `--platform manylinux2014_aarch64 --only-binary=:all:` — removing that pin once shipped x86 wheels into the ARM64 function and caused a total outage (PR #34).
### Phase 3: shared module extraction (`lambdas/shared/`)
@ -247,17 +226,17 @@ Four first-party modules that were previously duplicated per pipeline (or inline
| `email_parsing.py` | `parse_raw_email` (the WO superset that returns `cc` unconditionally; PO simply ignores `cc`) | both email processors |
| `emf.py` | generic CloudWatch EMF emitter (`emit_metric`, `emit_parse_outcome`) parameterized by namespace / dimension-sets / properties | both email processors (PO also uses `emit_metric` for `DerivedFieldAgreement`) |
**Flat-landing import rule.** The shared dir has **no `__init__.py`** — the modules are consumed by bare name (`from ses_auth import ...`, `from emf import emit_parse_outcome`), exactly as when they were siblings. This works because the bundling `cp` lands them **flat in `/asset-output/`** beside `handler.py`, so at runtime each shared module sits on the function's own `sys.path` under its bare name — the handler import lines are unchanged, which is what keeps the byte-identical fail-closed `ses_auth` behavior through the move. The load-bearing `emf` dimension-set list `[["ParseMethod"], ["ParseMethod", "TemplateId"]]` is now pinned **once** in `emf.py` (one-sided dimension drift between the two pipelines becomes structurally impossible), while the deliberate per-pipeline **emission-ordering** differences stay in the handlers (PO emits `ai_fallback` *before* the Bedrock call with an intentional double-count; WO emits mutually-exclusive `ai_fallback`/`ai_fallback_rejected` after its gate).
**Flat-landing import rule.** The shared dir has **no `__init__.py`** — the modules are consumed by bare name (`from ses_auth import ...`, `from emf import emit_parse_outcome`), exactly as when they were siblings. This works because packaging lands them **flat in the zip root** beside `handler.py`, so at runtime each shared module sits on the function's own `sys.path` under its bare name — the handler import lines are unchanged, which is what keeps the byte-identical fail-closed `ses_auth` behavior through the move. The load-bearing `emf` dimension-set list `[["ParseMethod"], ["ParseMethod", "TemplateId"]]` is now pinned **once** in `emf.py` (one-sided dimension drift between the two pipelines becomes structurally impossible), while the deliberate per-pipeline **emission-ordering** differences stay in the handlers (PO emits `ai_fallback` *before* the Bedrock call with an intentional double-count; WO emits mutually-exclusive `ai_fallback`/`ai_fallback_rejected` after its gate).
**Bundling — email processors.** Both email-processor commands append a second glob, `cp shared/*.py /asset-output/`, after their own `cp <pipeline>/email_processor/*.py`. This ships all four shared modules flat into each email-processor zip. `web_ui_auth.py` therefore rides along into both email-processor bundles even though the email handlers never import it — a harmless, deliberate consequence of the all-of-`shared/` glob (`PO_EXPECTED_TOP_LEVEL_MODULES` and the bundle-parity expectations account for it). The base is cp-only (Phase 7 removed the pip install / `manylinux` pin from both email-processor commands, and this phase moves only pure first-party modules with no new dependencies, so it **stays** cp-only — no pip step is reintroduced).
**Packaging — email processors.** Both email-processor packages run `copy_py_dir` for the pipeline dir plus `copy_shared_all`. This ships all four shared modules flat into each email-processor zip. `web_ui_auth.py` therefore rides along into both email-processor bundles even though the email handlers never import it — a harmless, deliberate consequence of copying all of `shared/` (`PO_EXPECTED_TOP_LEVEL_MODULES` and the bundle-parity expectations account for it).
**Bundling — web UIs.** Both `po-web-ui` and `workorder-web-ui` gain the same widened-root Docker bundling mechanism: their `Code.from_asset` root widens to `../lambdas` and their command copies the function's own dir contents plus **only** `shared/web_ui_auth.py` (`cp shared/web_ui_auth.py`, *not* `cp shared/*.py`) — the web UIs need only the auth module, and shipping the email-processor-only modules would break the "deployed set + `web_ui_auth`, nothing else" parity. The per-stack INFRA-74 comments stay in each handler (their wording is deliberately pipeline-specific and is not unified). `site_extractor`'s asset is untouched.
**Packaging — web UIs.** Both `po-web-ui` and `workorder-web-ui` copy their own dir plus **only** `shared/web_ui_auth.py` (not all of `shared/`) — the web UIs need only the auth module. `site_extractor` packages its own dir only.
`tests/test_bundle_consistency.py` is updated in lockstep without losing teeth: `_first_party_sibling_imports` resolves shared-sourced imports under `lambdas/shared/`; the command extractor selects the email-processor command now that each stack has two bundled functions; `_bundling_ships_all` accumulates shipped module stems across **both** globs; `PO_EXPECTED_TOP_LEVEL_MODULES` gains the four shared modules; and a new pin + mutation test require the `cp shared/*.py` line to be actually executed (a commented-out or removed shared `cp` fails CI).
`tests/test_bundle_consistency.py` is updated in lockstep without losing teeth: `_first_party_sibling_imports` resolves shared-sourced imports under `lambdas/shared/`; packaging recipes are parsed from `terraform/build_packages.sh`; ships-all accumulates across `copy_py_dir` + `copy_shared_all` / `copy_src_file`; and pins require shared / web_ui_auth copies to be actually present (a commented-out or removed copy fails CI).
### Phase 5: handler decomposition + lazy boto3 clients
Both email-processor God-handlers are decomposed along the seams that already work into **flat sibling modules** in the same directory (bare-name imports, exactly like the `ses_auth`/`template_parser`/`derived_fields` pattern), so the Phase 0/2/3 `cp <pipeline>/email_processor/*.py` glob ships every new sibling automatically — no bundling change beyond the exact-set pin. Every move is a pure delete-here/add-there; `derived_fields.py`, both `template_parser.py`, the `validate_ai_fallback` gates, `lambdas/shared/`, and `cdk/` are byte-untouched.
Both email-processor God-handlers are decomposed along the seams that already work into **flat sibling modules** in the same directory (bare-name imports, exactly like the `ses_auth`/`template_parser`/`derived_fields` pattern), so the packaging `copy_py_dir` ships every new sibling automatically — no packaging change beyond the exact-set pin. Every move is a pure delete-here/add-there; `derived_fields.py`, both `template_parser.py`, the `validate_ai_fallback` gates, and `lambdas/shared/` are byte-untouched.
**PO** (`handler.py` → 5 siblings + `prompts.py`): `handler.py` keeps the event loop, fail-closed SES auth, and `email_type` routing; `extraction.py` owns `extract_with_claude` + `_EMAIL_TAG_RE`; `enrichment.py` owns `enrich_parsed` + `pad_zip` (PO-only — WO has no enrichment stage) as a byte-identical move including the derived-field shadow block; `telemetry.py` owns the EMF `ParseMethod`/`DerivedFieldAgreement` emit wrappers; `persistence.py` owns `_write_fields`/`_merge_update`/`save_*` — with the two byte-identical `save_new_po`/`save_revision` **collapsed into one `_save_merge`** plus two thin wrappers differing only in the log verb (behavior-identical to both originals, sticky-cancel `ConditionExpression` guard intact; `save_cancellation` stays its own function). **WO** splits into ~5 concerns (no `enrichment`), keeping `validate_ai_fallback` **and** the `re.fullmatch(r"[0-9]+", work_order_id)` key guard in the handler loop AHEAD of both `save_work_order` and `save_event` (it protects the partition key and the `#`-delimited `comment_id` range-key segment), and keeps `_header_date_iso`/`comment_id` determinism together with `save_event` in `persistence.py`.
@ -277,7 +256,7 @@ Both email-processor God-handlers are decomposed along the seams that already wo
### `purchase-orders` table (owned here)
The `purchase-orders` DynamoDB table is **owned by this repo's `po-ingest` stack** (defined in `cdk/po_stack.py` with `RemovalPolicy.RETAIN`, `StreamViewType.NEW_IMAGE`, and SSE-KMS encryption with the shared customer-managed CMK `alias/seahaven-dynamodb`, INFRA-95 / M-3). The `po-email-processor` Lambda is the authoritative writer — it performs the merge inserts, field-level revision merges, and cancellation updates described above.
The `purchase-orders` DynamoDB table is **owned by this repo's Terraform config** (defined in `terraform/po_ddb.tf` with retain lifecycle, `StreamViewType.NEW_IMAGE`, and SSE-KMS encryption with the shared customer-managed CMK `alias/seahaven-dynamodb`, INFRA-95 / M-3). The `po-email-processor` Lambda is the authoritative writer — it performs the merge inserts, field-level revision merges, and cancellation updates described above.
**Consumers (read-only):**
@ -293,7 +272,7 @@ The `purchase-orders` DynamoDB table is **owned by this repo's `po-ingest` stack
### `WorkOrders` and `WorkOrderComments` tables (owned here)
Both tables are **owned by this repo's `WorkorderIngestStack`** (`cdk/wo_stack.py`, `RemovalPolicy.RETAIN`):
Both tables are **owned by this repo's Terraform config** (`terraform/wo_ddb.tf`, retain lifecycle):
- `WorkOrders` — PK `work_order_id` (S).
- `WorkOrderComments` — PK `work_order_id` (S), SK `comment_id` (S).
@ -323,7 +302,7 @@ Both currently use default DynamoDB encryption — they are **not** yet on the s
### `verified-sites` table (owned here)
Owned by this repo's `po-ingest` stack (`cdk/po_stack.py`). PK `siteCode` (S); default DynamoDB encryption (NOT the shared CMK).
Owned by this repo's Terraform config (`terraform/po_ddb.tf`). PK `siteCode` (S); default DynamoDB encryption (NOT the shared CMK).
**Consumer (read-only) — data contract:** `seahaven-slack-bot` (former reader, incl. the INFRA-138/INFRA-180 `by-state` GSI drift saga) was decommissioned 2026-07-23 — the GSI-drift issue died with it. Current consumer: the `procurement-api` Lambda (`GET /verified-sites*`). The API's `VerifiedSite` schema deliberately promises only pipeline-written fields (`siteCode`, `address`, `city`, `state`, `zip`, `fullAddress`, `locationCode`, `poCount`, `sourcePOs`) — `latitude`/`longitude`/`notes` were manually curated legacy attributes the extractor never writes, and are not part of the contract.
@ -337,7 +316,7 @@ The canonical map of Sea Haven's AWS infrastructure lives in Confluence. This pr
## CI/CD
GitHub Actions with reusable workflows from `Sea-Haven-Industries/.github` (all pinned to a commit SHA of `main`):
- **CI** (`ci.yaml`, PR to `main`): linting + `cdk synth` via `ci-python-sam.yaml` (historical CDK tree still synthesizes until a follow-up hygiene pass removes it).
- **CI** (`ci.yaml`, PR to `main`): linting + pytest via `ci-python-sam.yaml` (`run-cdk-synth: false`; HCP Terraform is the deploy path).
- **Terraform CI** (`ci-terraform.yaml`, PR to `main` when `terraform/**` changes): `fmt -check`, `init -backend=false`, `validate`.
- **CD:** HCP Terraform workspace `procurement-ingest-prod` (VCS-bound to `main`, working directory `terraform/`, Manual apply until sealed). There is no push-to-main CDK/SAM deploy workflow. Post-apply smoke: `scripts/post-deploy-smoke.sh`.
- Plus dependency review and PR labeler workflows on every PR
@ -386,13 +365,13 @@ pytest
**Test-root consolidation (Phase 8).** Both test roots are kept (`tests/` for genuinely cross-pipeline suites, plus each pipeline's own `lambdas/*/email_processor/tests/`), but loading is now single-sourced. A new repo-root `conftest.py` — not `tests/conftest.py`, which is never an ancestor of the pipeline test roots and so cannot load for a standalone `pytest lambdas/po/email_processor/tests` run — sets the dummy AWS env, imports `moto` before any handler import (registers moto's botocore stubber hook so boto3 sessions created afterwards are stubbable), and exposes the one `load_lambda_module(pipeline, name)` loader (implemented in `tests/support/loader.py`) that every handler exec in the suite goes through, including its `sys.modules` save/restore dance for the per-pipeline duplicated bare names (`template_parser`, `persistence`, etc. — `derived_fields`/`prompts`/`telemetry`/`extraction`/`enrichment` too). `tests/support/` is a small shared package (`FakeTable`, `FakeDynamoResource`, `FakeS3`, `load_email`, `load_golden`) used by both pipelines' `_po_parser_support.py`/`_wo_parser_support.py` shims — `load_golden` decodes JSON money with `parse_float=Decimal` unconditionally (load-bearing for PO's exact-money goldens; proven safe for WO, whose 55 golden fixtures contain zero float-typed JSON numbers). `test_local.py` (a root-level manual script that imported a handler at collection time, bypassing the loader gate, and globbed a nonexistent `samples/` dir) is deleted — the golden suites cover its role.
Coverage:
- `tests/` — shared handler + cross-pipeline tests: `parse_raw_email` (`test_parse_raw_email.py`), the fail-closed sender-authentication parser (`test_ses_auth.py`, INFRA-107), the Phase 0 CDK-bundling/handler-import AST consistency check (`test_bundle_consistency.py` — see [Deploy-Pipeline Guards](#deploy-pipeline-guards-phase-0)), the `scripts/reprocess.py` synthetic S3-event-shape contract (`test_reprocess_contract.py`, Phase 7), and, new in Phase 8: the handler-level SES-auth reject seam per pipeline (`test_handler_auth_seam.py` — no auth monkeypatch + an empty `ALLOWED_DKIM_DOMAINS` must yield zero Bedrock calls, zero writes, no raise), and both web_ui functions' first coverage — fail-closed auth (`test_web_ui_auth.py`) and the handlers themselves, including a 401-without-a-table-scan assertion and a hostile-field escaping regression lock (`test_web_ui_handlers.py`).
- `tests/` — shared handler + cross-pipeline tests: `parse_raw_email` (`test_parse_raw_email.py`), the fail-closed sender-authentication parser (`test_ses_auth.py`, INFRA-107), the packaging/handler-import consistency check (`test_bundle_consistency.py` — see [Deploy-Pipeline Guards](#deploy-pipeline-guards-phase-0)), the `scripts/reprocess.py` synthetic S3-event-shape contract (`test_reprocess_contract.py`, Phase 7), and, new in Phase 8: the handler-level SES-auth reject seam per pipeline (`test_handler_auth_seam.py` — no auth monkeypatch + an empty `ALLOWED_DKIM_DOMAINS` must yield zero Bedrock calls, zero writes, no raise), and both web_ui functions' first coverage — fail-closed auth (`test_web_ui_auth.py`) and the handlers themselves, including a 401-without-a-table-scan assertion and a hostile-field escaping regression lock (`test_web_ui_handlers.py`).
- `lambdas/wo/email_processor/tests/` — the deterministic WO parser suite: golden-file tests over 55 real scrubbed `.eml` fixtures (`test_parser.py`), fail-closed validation-gate rules and adversarial/injection cases (`test_validation_gate.py`), the issue #23 `comment_id` idempotency invariants (`test_comment_id.py`), the Bedrock-fallback dispatch/EMF-metric behavior with a mocked `invoke_model` (`test_bedrock_fallback.py`, which also carries the Phase 5 behavior pins — WO's mutually-exclusive emit and the `[0-9]+` key guard ahead of both saves), and the Phase 0 direct-invoke healthcheck contract (`test_healthcheck.py`). Golden JSON lives under `tests/fixtures/expected/`. New in Phase 8: `test_wo_bedrock_transport.py` (transport errors — throttle, missing `content`, empty content, non-JSON model text — plus the hand-computed no-double-count pin for the `handler.py` except-and-reraise metric wrap, and the multi-record failure-isolation pin) and `test_wo_merge.py` (a moto-backed mirror of `test_po_merge.py` against the `WorkOrders` table — null-status never clobbers `wo_status`, `created_at` immutable, `status`→`wo_status` mapping, `None` fields absent from `SET`, `record_type` only-when-present). After Phase 5 the suite patches accessors on the owning siblings (`extraction.bedrock`, `persistence.dynamodb`, `persistence.datetime`, `persistence.save_*`) rather than on `handler`.
- `lambdas/po/email_processor/tests/` — the deterministic PO parser suite: golden-file tests over real scrubbed `.eml` fixtures (17 single-line new-PO + 8 cancellations, exact `Decimal`-aware golden comparison via `parse_float=Decimal`), fail-closed validation-gate coverage for **every** gate reason code (fixture-driven for body-level triggers under `fixtures/adversarial/`, direct `validate()` unit tests for candidate-level mutations), real multi-line and comment/non-Coupa fallback fixtures under `fixtures/ai-fallback/`, dual line-ending (CRLF/LF) parse-identity, two-path `enrich_parsed`/`save_new_po` parity (the site-extractor stream-contract guard), fixture hygiene (`ses_auth` pass + scrub-marker leak sweep), the Bedrock-fallback dispatch/EMF-metric behavior plus the Phase 5 behavior pins (`test_po_bedrock_fallback.py` — the pre-Bedrock `ai_fallback` emit and the additive rejected double-count), the `_save_merge` collapse parity to both `save_new_po`/`save_revision` (`test_po_save_merge_parity.py`), the `ai_fallback`-only shadow telemetry after the `enrich_parsed` move (`test_po_derived_wiring.py`, which gains the PO-DC-02 64-char `po_number` clamp regression pin in Phase 8), and the Phase 0 direct-invoke healthcheck contract (`test_po_healthcheck.py`). New in Phase 8: `test_po_bedrock_transport.py` (the same four transport-error cases as WO, asserting the pre-call `ai_fallback` metric survives and no partial write occurs, plus the multi-record failure-isolation pin) and, moved here from `tests/`: `test_po_merge.py` (#97, moto-backed merge-write semantics) and `test_pad_zip.py` (zip-code padding). After Phase 5 the suite patches accessors/constants on the owning siblings (`po_extraction.bedrock`/`.BEDROCK_MODEL_ID`/`._EMAIL_TAG_RE`, `po_persistence.dynamodb`/`.PO_TABLE`, `po_enrichment.derive_all`/`.pad_zip`, `po_telemetry.DERIVED_METRIC_NAME`) rather than on `po_handler`; re-exported names it calls (`save_*`, `extract_with_claude`, `enrich_parsed`, `_emit_parse_method_metric`, `EXTRACTION_PROMPT`) stay on `po_handler`.
All three roots are discovered by `pytest.ini` (`testpaths`).
**Coverage floor (Phase 8).** `pytest.ini`'s `addopts` runs `pytest-cov` with an explicit `--cov` path per first-party package (`lambdas/po/email_processor`, `lambdas/wo/email_processor`, `lambdas/po/web_ui`, `lambdas/wo/web_ui`, `lambdas/po/site_extractor`, `lambdas/shared`) rather than relying on an `__init__.py` package marker — none of these dirs have one, and adding one would perturb the CDK bundling asset-hash fingerprint for zero runtime benefit; `pytest-cov`'s path form measures by source file regardless of package markers. `site_extractor` is deliberately included even though it measures 0% until Phase 6 lands — coverage honesty, not a silent skip. `.coveragerc` omits `*/tests/*` and `*/cdk.out/*` so the pipeline test dirs (which sit inside their own `--cov` path) don't dilute the number. `--cov-fail-under` is the new permanent CI floor (constraint 10 — never ratcheted down).
**Coverage floor (Phase 8).** `pytest.ini`'s `addopts` runs `pytest-cov` with an explicit `--cov` path per first-party package (`lambdas/po/email_processor`, `lambdas/wo/email_processor`, `lambdas/po/web_ui`, `lambdas/wo/web_ui`, `lambdas/po/site_extractor`, `lambdas/shared`) rather than relying on an `__init__.py` package marker — none of these dirs have one, and adding one would perturb flat packaging for zero runtime benefit; `pytest-cov`'s path form measures by source file regardless of package markers. `site_extractor` is deliberately included even though it measures 0% until Phase 6 lands — coverage honesty, not a silent skip. `.coveragerc` omits `*/tests/*` so the pipeline test dirs (which sit inside their own `--cov` path) don't dilute the number. `--cov-fail-under` is the new permanent CI floor (constraint 10 — never ratcheted down).
**Ruff C901/PLR floor (Phase 8).** `ruff.toml` adds `extend-select = ["C901", "PLR"]` with `max-complexity = 12`, making the tree's pre-existing `# noqa: PLR09xx` suppressions load-bearing instead of inert (no config previously enabled the rules they suppressed). `scripts/` is in the lint scope. Findings that could not be split or were out of this phase's file-ownership were resolved with a per-file `[lint.per-file-ignores]` entry carrying a written justification — most notably `derived_fields.py` (shadow-bake freeze: even an in-file `noqa` comment is a barred edit) and the other Phase 3/5/6/7 frozen modules (`enrichment.py`, WO `template_parser.py`, both `ses_auth.py`/`emf.py`, `site_extractor/handler.py`, `scripts/reprocess.py`). PO's `template_parser.py` — this phase's one in-scope split target — clears the ceiling by decomposing `extract_new_po` and `_validate_new_po_values` into per-rule helpers instead of an ignore.
@ -420,24 +399,15 @@ python scripts/replay_shoc_webhooks.py --url https://... --since 2026-07-24T02:0
## Directory Structure
```
cdk/
app.py # Three stacks: po-ingest + WorkorderIngestStack + procurement-api (region-only env)
common.py # Phase 4: shared plain-function CDK helpers (alarms, Bedrock grant,
# email bucket, processor DLQ, fallback-rate alarm) -- called with
# each stack's own scope + literal construct ids, logical-ID-safe
po_stack.py # Purchase order pipeline resources
wo_stack.py # Work order pipeline resources + the SHOC webhook feed (secret/CMK,
# rotator, emitter, disabled ESMs, queues, alarms)
procurement_api_stack.py # REST API (IAM SigV4 + resource policy) over both pipelines' tables + token-gated /docs
lambdas/ # Phase 2: shared Code.from_asset("../lambdas") bundling root for
# BOTH po-email-processor and workorder-email-processor (and, since
# Phase 3, both web_ui functions) -- each email-processor command
# `cp`s its own po/email_processor/*.py (or wo/...) subset PLUS
# `cp shared/*.py`; each web_ui command `cp`s its own dir PLUS only
# `shared/web_ui_auth.py`; site_extractor keeps its narrower, non-
# bundled asset root
terraform/ # Sole IaC path (HCP workspace procurement-ingest-prod)
build_packages.sh # Lambda zip packaging (copy_py_dir / copy_shared_all / copy_src_file)
*.tf # PO / WO / API / SHOC resources imported under PLAT-86
lambdas/ # Runtime source; packaged by terraform/build_packages.sh
# email processors: copy_py_dir + copy_shared_all;
# web_ui: copy_py_dir + shared/web_ui_auth.py only;
# site_extractor / SHOC: copy_py_dir of their own dirs
shared/ # Phase 3: single-sourced first-party modules, landed FLAT (no
# __init__.py -- bare-name imports) into each bundle by the cp above
# __init__.py -- bare-name imports) into each email-processor zip
ses_auth.py # fail-closed SES sender-auth (INFRA-107) -- one copy, one fix
web_ui_auth.py # fail-closed X-Auth-Token gate + token cache (INFRA-74; also gates the API /docs routes)
email_parsing.py # parse_raw_email (WO superset; returns cc unconditionally)
@ -455,8 +425,7 @@ lambdas/ # Phase 2: shared Code.from_asset("../lambdas") bundling
fonts.css # SHOC fonts (DM Sans/Montserrat/JetBrains Mono, @fontsource latin subsets) as data URIs, inlined into /docs
po/ # PO pipeline Lambdas
email_processor/ # Phase 5: God-handler decomposed into flat siblings (bare-name
# imports; the Phase 0/2/3 `cp <pipeline>/email_processor/*.py`
# glob ships every new sibling automatically)
# imports; copy_py_dir ships every new sibling automatically)
handler.py # event loop + fail-closed SES auth + email_type routing + {"healthcheck": true} early-return; lazy s3 accessor; re-exports EXTRACTION_PROMPT (4 tests deref handler.EXTRACTION_PROMPT)
extraction.py # extract_with_claude + _EMAIL_TAG_RE + EXTRACTION_PROMPT import; lazy bedrock accessor
enrichment.py # enrich_parsed + pad_zip (PO-only; no boto3); derived-field shadow block

View file

@ -1,30 +0,0 @@
#!/usr/bin/env python3
import aws_cdk as cdk
from po_stack import PoIngestStack
from procurement_api_stack import ProcurementApiStack
from wo_stack import WorkorderIngestStack
app = cdk.App()
PoIngestStack(
app,
"po-ingest",
stack_name="po-ingest",
env=cdk.Environment(region="us-east-1"),
)
WorkorderIngestStack(
app,
"workorder-ingest",
stack_name="WorkorderIngestStack",
env=cdk.Environment(region="us-east-1"),
)
ProcurementApiStack(
app,
"procurement-api",
stack_name="procurement-api",
env=cdk.Environment(region="us-east-1"),
)
app.synth()

View file

@ -1,6 +0,0 @@
{
"app": "python3 app.py",
"context": {
"@aws-cdk/core:bootstrapQualifier": "hnb659fds"
}
}

View file

@ -1,357 +0,0 @@
"""Shared CDK helpers for the procurement-ingest stacks (Phase 4).
Plain free functions extracted from po_stack.py / wo_stack.py. Each takes the
same Stack `scope` and the same literal construct `id` the stacks passed inline,
so every synthesized logical ID is byte-stable. NOTHING here is a Construct
subclass: a subclass would insert a tree node and reparent/replace the
RETAIN-protected tables and named buckets.
"""
from aws_cdk import (
Duration,
RemovalPolicy,
Stack,
aws_cloudwatch as cloudwatch,
aws_cloudwatch_actions as cw_actions,
aws_dynamodb as dynamodb,
aws_iam as iam,
aws_logs as logs,
aws_s3 as s3,
aws_sqs as sqs,
)
# --- DynamoDB alarm operations (moved verbatim from both stacks) ---
_DDB_ALARM_OPERATIONS = [
dynamodb.Operation.GET_ITEM,
dynamodb.Operation.BATCH_GET_ITEM,
dynamodb.Operation.QUERY,
dynamodb.Operation.SCAN,
dynamodb.Operation.PUT_ITEM,
dynamodb.Operation.UPDATE_ITEM,
dynamodb.Operation.DELETE_ITEM,
dynamodb.Operation.BATCH_WRITE_ITEM,
]
# CloudWatch namespace for the log-derived sender-authentication metrics.
_SENDER_AUTH_METRIC_NAMESPACE = "Seahaven/ProcurementIngest"
def add_ddb_alarms(scope, id_prefix, table, alarm_name_prefix, alarm_topic):
"""Add throttle + system-error alarms for a DynamoDB table.
Both fire on any non-zero datapoint in a 5-min window. ALARM-only SnsAction
to site-alerts (no OK action); TreatMissingData NOT_BREACHING.
"""
table.metric_throttled_requests_for_operations(
operations=_DDB_ALARM_OPERATIONS,
period=Duration.minutes(5),
statistic="Sum",
).create_alarm(
scope,
f"{id_prefix}ThrottlesAlarm",
alarm_name=f"{alarm_name_prefix}-throttles",
alarm_description=f"{alarm_name_prefix} DynamoDB throttled requests",
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
table.metric_system_errors_for_operations(
operations=_DDB_ALARM_OPERATIONS,
period=Duration.minutes(5),
statistic="Sum",
).create_alarm(
scope,
f"{id_prefix}SystemErrorsAlarm",
alarm_name=f"{alarm_name_prefix}-system-errors",
alarm_description=f"{alarm_name_prefix} DynamoDB server-side (5xx) errors",
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
def make_function_log_group(scope, id_prefix, function_name):
"""Explicit log group for a Lambda, replacing the deprecated
``log_retention`` prop (INFRA-114).
Making the group a real stack resource (a) drops the LogRetention custom
resource whose role carried wildcard ``logs:PutRetentionPolicy`` (checkov
CKV_AWS_111), and (b) makes the group an orderable CFN dependency:
consumers like the sender-auth metric filter now deploy AFTER the group
exists. The previous by-name import raced group creation in a fresh
account and failed the first seahaven-prod deploy (the mgmt account
masked this because its groups predated the filter).
RETAIN matches the repo convention for stateful resources and mirrors the
old behavior (LogRetention never deleted groups on stack delete). NOTE:
never ``cdk deploy`` this app to mgmt (328440206208) — PLAT-67 deleted the
mgmt stacks and left RETAIN cold-archive log groups; a redeploy would
collide with those leftovers.
"""
return logs.LogGroup(
scope,
f"{id_prefix}LogGroup",
log_group_name=f"/aws/lambda/{function_name}",
retention=logs.RetentionDays.TWO_MONTHS,
removal_policy=RemovalPolicy.RETAIN,
)
def add_sender_auth_rejected_alarm(
scope, id_prefix, function_name, alarm_topic, log_group
):
"""Metric-filter + alarm on ``sender_auth_rejected`` warnings (INFRA-107).
A rejected inbound email is skipped without erroring the invocation, so it
is invisible to the Errors/Throttles/DLQ alarms. This turns the structured
warning log into a CloudWatch metric and pages when rejections spike --
catching a silent false-reject storm (allowlist wrong, signing-domain
drift, SES header-format change) that would otherwise discard legitimate
mail while the pipeline reports healthy.
ALARM-only SnsAction to site-alerts; no OK action. The metric filter reads
the function's own log group, passed in as the EXPLICIT LogGroup construct
(from ``make_function_log_group``) so CFN orders the filter after the
group exists -- a by-name import here failed the first fresh-account
deploy. A plain substring pattern is used because Lambda prefixes each
line with its own level/timestamp/request-id, so the JSON payload is not
a standalone JSON log event a `{$.event=...}` pattern could match.
"""
metric_name = f"{function_name}-sender-auth-rejected"
logs.MetricFilter(
scope,
f"{id_prefix}SenderAuthRejectedFilter",
log_group=log_group,
filter_pattern=logs.FilterPattern.literal('"sender_auth_rejected"'),
metric_namespace=_SENDER_AUTH_METRIC_NAMESPACE,
metric_name=metric_name,
metric_value="1",
default_value=0,
)
cloudwatch.Metric(
namespace=_SENDER_AUTH_METRIC_NAMESPACE,
metric_name=metric_name,
period=Duration.minutes(5),
statistic="Sum",
).create_alarm(
scope,
f"{id_prefix}SenderAuthRejectedAlarm",
alarm_name=f"{function_name}-sender-auth-rejected",
alarm_description=(
f"{function_name} rejected inbound mail on sender authentication "
"(possible allowlist/DKIM-domain drift silently dropping real mail)"
),
threshold=1,
evaluation_periods=6,
datapoints_to_alarm=2,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
def add_standard_lambda_alarms(
scope,
id_prefix,
fn,
name_prefix,
topic,
*,
duration_statistic,
errors=True,
dlq=None,
descriptions,
):
"""Standard per-Lambda alarm set: errors (optional), throttles, dlq
(optional), duration. Every alarm bespoke-described via `descriptions`
(keys: errors/throttles/dlq/duration passed through VERBATIM). Construct
ids are f"{id_prefix}<Kind>Alarm", alarm names f"{name_prefix}-<kind>",
exactly the inline literals. Order of creation is irrelevant to logical IDs
(ids are explicit) so the fixed errors->throttles->dlq->duration order here
reproduces both PO (dlq before duration) and WO (duration before dlq)
templates identically.
"""
if errors:
fn.metric_errors(period=Duration.minutes(5), statistic="Sum").create_alarm(
scope,
f"{id_prefix}ErrorsAlarm",
alarm_name=f"{name_prefix}-errors",
alarm_description=descriptions["errors"],
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(topic))
fn.metric_throttles(period=Duration.minutes(5), statistic="Sum").create_alarm(
scope,
f"{id_prefix}ThrottlesAlarm",
alarm_name=f"{name_prefix}-throttles",
alarm_description=descriptions["throttles"],
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(topic))
if dlq is not None:
dlq.metric_approximate_number_of_messages_visible(
period=Duration.minutes(5),
statistic="Maximum",
).create_alarm(
scope,
f"{id_prefix}DlqMessagesAlarm",
alarm_name=f"{name_prefix}-dlq-messages",
alarm_description=descriptions["dlq"],
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(topic))
fn.metric_duration(
period=Duration.minutes(5),
statistic=duration_statistic,
).create_alarm(
scope,
f"{id_prefix}DurationAlarm",
alarm_name=f"{name_prefix}-duration",
alarm_description=descriptions["duration"],
threshold=45000,
evaluation_periods=3,
datapoints_to_alarm=2,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(topic))
def make_bedrock_invoke_statement(scope):
"""Return the Bedrock InvokeModel PolicyStatement for the email processor.
The us.* inference profile can route cross-region, so the grant covers the
inference-profile ARN plus the per-region foundation-model ARNs
(us-east-1/us-east-2/us-west-2). ARN #1 account and ARN #1/#2 region are
DERIVED from Stack.of(scope).account/.region (not hardcoded 328440206208);
ARN #3/#4 regions (us-east-2/us-west-2) are cross-region reach targets, NOT
the stack's own region, so they stay literal. Caller attaches via
fn.add_to_role_policy(...) on the SAME function -> logical-ID-safe.
"""
stack = Stack.of(scope)
account = stack.account
region = stack.region
return iam.PolicyStatement(
actions=[
"bedrock:InvokeModel",
"bedrock:InvokeModelWithResponseStream",
],
resources=[
f"arn:aws:bedrock:{region}:{account}:inference-profile/us.anthropic.claude-haiku-4-5-20251001-v1:0",
f"arn:aws:bedrock:{region}::foundation-model/anthropic.claude-haiku-4-5-20251001-v1:0",
"arn:aws:bedrock:us-east-2::foundation-model/anthropic.claude-haiku-4-5-20251001-v1:0",
"arn:aws:bedrock:us-west-2::foundation-model/anthropic.claude-haiku-4-5-20251001-v1:0",
],
)
def make_email_bucket(scope, id, name_prefix):
"""Raw-email S3 bucket: BLOCK_ALL public, RETAIN, 90-day expiration.
bucket_name = f"{name_prefix}-{account}" (account derived from the stack),
reproducing f"po-ingest-emails-{self.account}" /
f"workorder-ingest-emails-{self.account}" exactly.
"""
return s3.Bucket(
scope,
id,
bucket_name=f"{name_prefix}-{Stack.of(scope).account}",
block_public_access=s3.BlockPublicAccess.BLOCK_ALL,
removal_policy=RemovalPolicy.RETAIN,
lifecycle_rules=[
s3.LifecycleRule(expiration=Duration.days(90)),
],
)
def make_processor_dlq(scope, id):
"""Async-invoke DLQ: 14-day retention, enforce_ssl. CDK-generated name."""
return sqs.Queue(
scope,
id,
retention_period=Duration.days(14),
enforce_ssl=True,
)
def make_fallback_rate_alarm(
scope,
id,
*,
namespace,
alarm_topic,
alarm_name,
alarm_description,
rejected_included,
period,
threshold,
floor,
evaluation_periods,
datapoints_to_alarm,
):
"""Template-fallback-rate MathExpression alarm.
rejected_included=False (PO): numerator FILL(fb,0), denom fb+tmpl.
rejected_included=True (WO): numerator (fb+rej), denom fb+rej+tmpl.
Expression string, FILL, label reproduced BYTE-FOR-BYTE. GREATER_THAN,
NOT_BREACHING. NO element-wise MAX (post-#102). The two DISTINCT rejected
alarms are NOT built here.
"""
fb_metric = cloudwatch.Metric(
namespace=namespace,
metric_name="ParseOutcome",
dimensions_map={"ParseMethod": "ai_fallback"},
statistic="Sum",
period=period,
)
tmpl_metric = cloudwatch.Metric(
namespace=namespace,
metric_name="ParseOutcome",
dimensions_map={"ParseMethod": "template"},
statistic="Sum",
period=period,
)
if rejected_included:
fb_rej_metric = cloudwatch.Metric(
namespace=namespace,
metric_name="ParseOutcome",
dimensions_map={"ParseMethod": "ai_fallback_rejected"},
statistic="Sum",
period=period,
)
using_metrics = {"fb": fb_metric, "rej": fb_rej_metric, "tmpl": tmpl_metric}
sum_terms = "FILL(fb,0)+FILL(rej,0)+FILL(tmpl,0)"
numerator = "(FILL(fb,0)+FILL(rej,0))"
else:
using_metrics = {"fb": fb_metric, "tmpl": tmpl_metric}
sum_terms = "FILL(fb,0)+FILL(tmpl,0)"
numerator = "FILL(fb,0)"
expression = f"IF(({sum_terms})>={floor}, 100*{numerator}/({sum_terms}), 0)"
fallback_rate = cloudwatch.MathExpression(
expression=expression,
using_metrics=using_metrics,
period=period,
label="TemplateFallbackRatePct",
)
fallback_rate.create_alarm(
scope,
id,
alarm_name=alarm_name,
alarm_description=alarm_description,
threshold=threshold,
evaluation_periods=evaluation_periods,
datapoints_to_alarm=datapoints_to_alarm,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))

View file

@ -1,616 +0,0 @@
"""CDK stack for the Coupa PO email ingestion pipeline."""
import aws_cdk as cdk
from aws_cdk import (
Duration,
RemovalPolicy,
Stack,
aws_cloudwatch as cloudwatch,
aws_cloudwatch_actions as cw_actions,
aws_dynamodb as dynamodb,
aws_kms as kms,
aws_lambda as lambda_,
aws_lambda_event_sources as lambda_event_sources,
aws_s3 as s3,
aws_s3_notifications as s3n,
aws_ses as ses,
aws_ses_actions as ses_actions,
aws_secretsmanager as secretsmanager,
aws_sns as sns,
aws_ssm as ssm,
)
from constructs import Construct
import common
class PoIngestStack(Stack):
def __init__(self, scope: Construct, construct_id: str, **kwargs):
super().__init__(scope, construct_id, **kwargs)
# --- Shared alarm SNS topic (site-alerts) ---
# Imported once near the top so every alarm in this stack reuses the same
# Topic construct instance (avoids duplicate logical IDs). ALARM-only
# SnsAction; no OK action, per the CloudWatch-alarm preference. The
# topic's CMK (alias/seahaven-alarm-topics) lives on the topic itself.
alarm_topic = sns.Topic.from_topic_arn(
self,
"SiteAlertsTopic",
f"arn:aws:sns:{self.region}:{self.account}:site-alerts",
)
# --- S3 bucket for raw emails ---
email_bucket = common.make_email_bucket(self, "EmailBucket", "po-ingest-emails")
# --- Shared customer-managed CMK for sensitive DynamoDB tables ---
# Owned by the account-baseline app (alias/seahaven-dynamodb, INFRA-95 /
# M-3); ARN published to SSM. The purchase-orders table was migrated to
# SSE-KMS out-of-band, so declaring encryption_key here reconciles the
# drift and — via grant_read_write_data below — propagates the required
# kms:Decrypt/GenerateDataKey/DescribeKey to the consumer roles.
dynamodb_cmk = kms.Key.from_key_arn(
self,
"DynamoDbCmk",
ssm.StringParameter.value_for_string_parameter(
self, "/seahaven/dynamodb/cmk-arn"
),
)
# --- Purchase-orders DynamoDB table ---
# Owned by this stack. Streams enabled for the site-extractor pipeline.
# Other stacks (seahaven-slack-bot) reference this table via fromTableName().
po_table = dynamodb.Table(
self,
"PurchaseOrdersTable",
table_name="purchase-orders",
partition_key=dynamodb.Attribute(
name="po_number",
type=dynamodb.AttributeType.STRING,
),
billing_mode=dynamodb.BillingMode.PAY_PER_REQUEST,
removal_policy=RemovalPolicy.RETAIN,
stream=dynamodb.StreamViewType.NEW_IMAGE,
encryption=dynamodb.TableEncryption.CUSTOMER_MANAGED,
encryption_key=dynamodb_cmk,
)
# --- Anthropic API key secret removed (Bedrock migration) ---
# PO parsing stays fully AI but moved from the Anthropic API to the
# Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0,
# so no provider API key is needed. The old secret
# "po-ingest/anthropic-api-key" had RemovalPolicy.RETAIN, so it is
# ORPHANED (not deleted) by this change: delete it manually post-deploy
# and revoke the stored key at Anthropic.
# --- DLQ for failed async invocations (INFRA-41 / audit H-8) ---
# SES → S3 → Lambda is async; without an OnFailure destination a failed
# parse (bad email, transient error) is silently dropped after Lambda's
# retries. CDK generates the queue name to avoid colliding with the
# interim CLI-created po-email-processor-dlq (removed post-deploy).
email_processor_dlq = common.make_processor_dlq(self, "EmailProcessorDlq")
# --- Lambda function ---
email_processor_log_group = common.make_function_log_group(
self, "EmailProcessor", "po-email-processor"
)
email_processor = lambda_.Function(
self,
"EmailProcessor",
function_name="po-email-processor",
runtime=lambda_.Runtime.PYTHON_3_12,
architecture=lambda_.Architecture.ARM_64,
handler="handler.handler",
code=lambda_.Code.from_asset(
"../lambdas",
exclude=["**/__pycache__/**", "**/tests/**", "**/package/**"],
bundling=cdk.BundlingOptions(
image=lambda_.Runtime.PYTHON_3_12.bundling_image,
command=[
"bash",
"-c",
# NOTE: non-recursive glob (not `cp -r`) so tests/ and
# the stale package/ dir are never shipped -- only
# top-level .py siblings of handler.py. This replaces
# a hand-maintained four-file allowlist that twice
# nearly shipped a broken Lambda (missing
# template_parser in PR #105, nearly missing
# derived_fields in PR #2) because a new sibling
# import wasn't added to the list. The glob makes
# that class of bug structurally impossible;
# tests/test_bundle_consistency.py ast-parses
# handler.py's first-party imports and asserts this
# command ships all of them, so a future revert back
# to an allowlist that omits a sibling fails CI.
# pip step removed in Phase 7: requirements.txt is now
# empty (boto3 comes from the Lambda runtime), so nothing
# is installed and the manylinux pin has nothing to pin.
# shared/*.py ships the four modules extracted to
# lambdas/shared/ (Phase 3): ses_auth, web_ui_auth,
# email_parsing, emf. Flat cp keeps the bare-name
# imports (e.g. `from ses_auth import ...`) resolving
# unchanged in /asset-output.
"cp po/email_processor/*.py /asset-output/ && "
"cp shared/*.py /asset-output/",
],
),
),
timeout=Duration.seconds(60),
memory_size=256,
log_group=email_processor_log_group,
dead_letter_queue=email_processor_dlq,
environment={
"PO_TABLE": "purchase-orders",
"BEDROCK_MODEL_ID": "us.anthropic.claude-haiku-4-5-20251001-v1:0",
# Fail-closed sender auth (INFRA-107): the handler only
# accepts mail whose SES-stamped Authentication-Results
# header carries dkim=pass for one of these domains.
# Observed on live traffic 2026-07-15: Coupa PO mail passes
# DKIM for amazon.coupahost.com (and amazonses.com, which is
# deliberately NOT allowlisted — every SES customer's mail
# passes that). Unset/empty ⇒ the handler rejects all mail.
"ALLOWED_DKIM_DOMAINS": "amazon.coupahost.com",
},
)
# Grant permissions
email_bucket.grant_read(email_processor)
po_table.grant_read_write_data(email_processor)
# --- Bedrock InvokeModel grant ---
# The us.* inference profile can route cross-region, so the grant MUST
# cover both the inference-profile ARN AND the per-region foundation-model
# ARNs (empty account field) for every region the profile can reach
# (us-east-1/us-east-2/us-west-2). A profile-only grant AccessDenies at
# runtime whenever the profile routes to a region whose foundation-model
# ARN is not allowed.
email_processor.add_to_role_policy(common.make_bedrock_invoke_statement(self))
# --- Standard per-Lambda alarms: po-email-processor ---
# errors (INFRA-41 / audit H-8), throttles, DLQ-messages (dropped PO
# emails), and a p99 duration alarm (orphan adoption of the CLI
# Lambda-Duration-po-email-processor under <fn>-duration naming, 45000 ms
# = 75% of the 60s timeout, eval 3 / dp 2). All ALARM-only to site-alerts.
common.add_standard_lambda_alarms(
self,
"EmailProcessor",
email_processor,
"po-email-processor",
alarm_topic,
duration_statistic="p99",
errors=True,
dlq=email_processor_dlq,
descriptions={
"errors": "po-email-processor async invocation errors",
"throttles": "po-email-processor invocation throttles",
"dlq": "po-email-processor DLQ has messages (dropped PO emails)",
"duration": "po-email-processor p99 duration approaching the 60s timeout",
},
)
# --- Sender-auth rejection alarm (INFRA-107) ---
# A rejected email (bad/unaligned DKIM verdict) returns normally, so it
# produces NO Lambda error, NO DLQ message and NO retry -- only a
# `sender_auth_rejected` warning log. Without this metric filter + alarm a
# domain drift (Coupa rotates its signing subdomain, SES changes its
# Authentication-Results format, the allowlist is wrong) would silently
# discard 100% of legitimate PO mail while every other alarm stays green.
# A CloudWatch Logs metric filter turns those warnings into a metric so a
# false-reject storm pages instead of vanishing. default_value=0 keeps the
# series populated (alarm stays OK, never INSUFFICIENT_DATA) between events.
common.add_sender_auth_rejected_alarm(
self,
"EmailProcessor",
"po-email-processor",
alarm_topic,
email_processor_log_group,
)
# --- Template fallback-rate alarm: po-email-processor ---
# The processor tries a deterministic template parse first and only calls
# the Bedrock AI extractor on a miss/invalid. A sustained rise in the
# ai_fallback share signals Coupa template drift (coverage collapse).
# EMF metric Seahaven/PoIngest/ParseOutcome, dimensioned by ParseMethod
# (template|ai_fallback).
#
# RETUNED for PO volume (~57 emails/day ≈ 14.25 per 6h period) -- the WO
# alarm's 15-min period / >=10-sample floor assume ~760/day and would be
# structurally DEAD here (a 15-min period holds ~0.6 PO emails, so the
# floor is never met and the IF always takes the 0 branch):
# * period 6h: a stable ~14-email denominator per datapoint.
# * volume floor >=8: at the floor, one fallback email = 12.5% < 20%,
# so a single email can NEVER breach a datapoint; a breach needs >=2
# fallbacks in one 6h window (2/8 = 25%) or >=3 at typical volume
# (3/14 ≈ 21%). Sparse overnight/weekend windows (<8 emails) take
# the 0 branch -- non-breaching by design (accepted trade: a Friday-
# evening drift may not page until weekend volume accrues).
# * threshold >20%: expected baseline fallback ≈1% (comments 0.55% +
# multi-line 0.18% + non-USD 0) -- far below the threshold.
# * 2 of 4 datapoints (24h span): isolated noise self-clears, while
# total template drift (100% fallback) pages within ~12h.
# Post-#102 rule: NO element-wise MAX(timeseries, scalar) in alarm math;
# the IF volume floor guarantees the non-zero denominator. Any change to
# this expression must be gated by `npx cdk synth po-ingest`.
#
# DOUBLE-COUNT ACCOUNTING (Phase 1 / PO AI-fallback gate): PO emits
# ParseMethod=ai_fallback BEFORE the Bedrock call for EVERY AI-path
# email (handler pre-call emit; try_deterministic_parse returns
# "ai_fallback" on every template miss), so a gate-rejected email
# already appears exactly once in `fb`. Therefore fb = ALL fallback
# attempts (accepted + rejected), fb + tmpl = ALL emails, and
# rate = fb/(fb+tmpl) is exact -- the expression below is deliberately
# left BYTE-IDENTICAL to the pre-Phase-1 form, and `rej`
# (ai_fallback_rejected) is deliberately EXCLUDED from this
# expression's numerator, denominator, and volume floor, and is never
# added to using_metrics. This is NOT an oversight: folding `rej` in
# here as WO does (fb+rej numerator / fb+rej+tmpl denominator) would
# double-count every rejected email in both numerator and
# denominator (PO's pre-call emit already counts it once via `fb`),
# inflating the observed rate toward 100% and double-counting toward
# the >=8 volume floor -- a prompt-injection probing burst would then
# falsely page this template-drift alarm on top of the dedicated
# rejected alarm below. The rejected series gets its own alarm
# instead (EmailProcessorAiFallbackRejectedAlarm, below).
common.make_fallback_rate_alarm(
self,
"EmailProcessorTemplateFallbackRateAlarm",
namespace="Seahaven/PoIngest",
alarm_topic=alarm_topic,
alarm_name="po-email-processor-template-fallback-rate",
alarm_description=(
"po-email-processor deterministic-template coverage collapse: "
">20% of parses fell back to the Bedrock AI extractor"
),
rejected_included=False,
period=Duration.hours(6),
threshold=20,
floor=8,
evaluation_periods=4,
datapoints_to_alarm=2,
)
# --- AI-fallback rejected alarm: po-email-processor (Phase 1) ---
# The validate_ai_fallback gate (template_parser.py) fail-closes Bedrock
# output that doesn't match PO's contract (structurally wrong shape,
# injected po_number/email_type, wrong field types) and emits
# ParseMethod=ai_fallback_rejected instead of writing it. That is a
# SILENT skip (`continue`, never raise) by design -- attacker-controlled
# input must not churn the retry/DLQ path -- so without a dedicated
# alarm a sustained rejection run (prompt-injection probing, or a
# template-drift outage whose AI output also happens to fail the gate)
# is invisible everywhere except this metric and the ReasonCode log
# line.
#
# RETUNED for PO volume (~57 emails/day, baseline ai_fallback rate
# ~1% => ~0.6 AI-fallback emails/day, expected rejections ~= 0) -- NOT
# WO's 5-min/2-of-6 sparse idiom (wo_stack.py), which needs two
# rejections inside a single 30-min window and is structurally dead at
# this volume. Mirrors the PO fallback-rate alarm's 6h/eval-4/dp-2
# retune idiom above, but with a COUNT floor on the rejected series
# itself rather than an email-volume floor: an email-volume floor
# (fb+tmpl>=N) would suppress paging in exactly the sparse
# overnight/weekend windows where a silently-dropped email matters
# most, and there is no denominator here, so there is nothing else to
# guard against divide-by-zero. FILL(rej,0) turns the sparse EMF
# series (no datapoint in quiet periods -- no metric-filter
# default_value exists for EMF) into a dense 0-series so every
# evaluation window has data. Post-#102 rule still holds: NO
# element-wise MAX(timeseries, scalar) anywhere in this expression.
#
# Tuning: rejections self-clear unless >=2 breaching datapoints land in
# >=2 distinct 6h windows within 24h (sustained probing, or template
# drift whose AI output also fails the gate), which pages within
# ~12-24h. Accepted residual (matches WO's accepted residual): because
# the breach is measured per 6h window, ANY burst of rejections
# confined to a single 6h window -- whether one stray email or dozens
# in a 20-minute spike -- is one breaching datapoint and never pages
# this alarm by itself. This is deliberate anti-flap tuning at ~0
# expected rejections/day, not a coverage gap in the fail-closed gate:
# every burst email is still rejected before any DynamoDB write, and
# the burst stays fully visible as ai_fallback_rejected datapoints and
# ReasonCode log lines, with the pre-call ai_fallback emit also raising
# the fallback-rate numerator above. A same-window burst detector
# (1-of-1 at a higher threshold) is a tracked follow-up if faster
# single-window paging is wanted.
rejected_metric = cloudwatch.Metric(
namespace="Seahaven/PoIngest",
metric_name="ParseOutcome",
dimensions_map={"ParseMethod": "ai_fallback_rejected"},
statistic="Sum",
period=Duration.hours(6),
)
rejected_floor = cloudwatch.MathExpression(
expression="IF(FILL(rej,0)>=1, FILL(rej,0), 0)",
using_metrics={"rej": rejected_metric},
period=Duration.hours(6),
label="AiFallbackRejectedCount",
)
rejected_floor.create_alarm(
self,
"EmailProcessorAiFallbackRejectedAlarm",
alarm_name="po-email-processor-ai-fallback-rejected",
alarm_description=(
"po-email-processor is rejecting Bedrock AI-fallback output at "
"the validation gate (possible prompt-injection probing or "
"template drift silently dropping real mail)"
),
threshold=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
evaluation_periods=4,
datapoints_to_alarm=2,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
# S3 event notification → Lambda
email_bucket.add_event_notification(
s3.EventType.OBJECT_CREATED,
s3n.LambdaDestination(email_processor),
s3.NotificationKeyFilter(prefix="inbound/"),
)
# --- SES Receipt Rule ---
# Reuse the existing INBOUND_MAIL rule set (shared with workorder-ingest)
rule_set = ses.ReceiptRuleSet.from_receipt_rule_set_name(
self,
"ExistingRuleSet",
"INBOUND_MAIL",
)
rule_set.add_rule(
"PoEmailRule",
recipients=["amazon_po@int.seahaven.com"],
actions=[
ses_actions.S3(
bucket=email_bucket,
object_key_prefix="inbound/",
),
],
)
# --- Web UI auth token secret ---
# Shared secret for the web UI auth gate, stored in Secrets Manager and
# resolved at runtime so the token never appears in CloudFormation templates
# or Lambda environment variables. Create this secret before deploying
# either stack; both PO and WO stacks reference it by name.
web_ui_auth_secret = secretsmanager.Secret.from_secret_name_v2(
self,
"WebUiAuthToken",
"procurement-ingest/web-ui-auth-token",
)
# --- Web UI Lambda ---
web_ui_log_group = common.make_function_log_group(self, "WebUI", "po-web-ui")
web_ui = lambda_.Function(
self,
"WebUI",
function_name="po-web-ui",
runtime=lambda_.Runtime.PYTHON_3_12,
architecture=lambda_.Architecture.ARM_64,
handler="handler.handler",
code=lambda_.Code.from_asset(
"../lambdas",
exclude=["**/__pycache__/**", "requirements.txt"],
bundling=cdk.BundlingOptions(
image=lambda_.Runtime.PYTHON_3_12.bundling_image,
command=[
"bash",
"-c",
# web_ui_auth.py is shared (lambdas/shared) and must land
# FLAT beside handler.py so the bare
# `from web_ui_auth import is_authenticated` resolves at
# runtime. Only web_ui_auth is copied from shared/ -- the
# other shared modules (ses_auth/email_parsing/emf) are
# email-processor-only and must not bloat the web UI zip.
# requirements.txt stays excluded (Dependabot anchor only,
# never runtime): keeps the deployed file list = {handler,
# web_ui_auth}. NOTE: the top-level `exclude=` on
# from_asset only filters the asset-hash fingerprint, NOT
# the directory Docker bundling actually mounts, so a
# local __pycache__/requirements.txt on disk at synth
# time WOULD otherwise leak into the bundled zip -- strip
# them explicitly post-cp instead of relying on exclude.
"cp -r po/web_ui/. /asset-output/ && "
"cp shared/web_ui_auth.py /asset-output/ && "
"rm -rf /asset-output/__pycache__ /asset-output/requirements.txt",
],
),
),
timeout=Duration.seconds(60),
memory_size=256,
log_group=web_ui_log_group,
environment={
"PO_TABLE": "purchase-orders",
# Defense-in-depth shared secret for the web UI handler. The
# handler fails closed if this ARN is unset or the secret is
# missing, so any future invocation path cannot re-expose the
# PO DB unauthenticated. The secret value is fetched at runtime
# from Secrets Manager (not embedded in env vars or template).
"WEB_UI_AUTH_TOKEN_SECRET_ARN": web_ui_auth_secret.secret_arn,
},
)
po_table.grant_read_data(web_ui)
web_ui_auth_secret.grant_read(web_ui)
# --- Standard per-Lambda alarms: po-web-ui ---
# Throttles + p99 duration only (no errors alarm, no DLQ -- web_ui is a
# synchronous read path with no async DLQ). p99 / 45000 ms (75% of the
# 60s timeout) / eval 3, datapoints 2.
common.add_standard_lambda_alarms(
self,
"WebUi",
web_ui,
"po-web-ui",
alarm_topic,
duration_statistic="p99",
errors=False,
dlq=None,
descriptions={
"throttles": "po-web-ui invocation throttles",
"duration": "po-web-ui p99 duration approaching the 60s timeout",
},
)
# Public Function URL removed 2026-06-08 (INFRA-74 / audit C-5): the
# unauthenticated FunctionUrlAuthType.NONE URL was deleted out-of-band
# via CLI. Removing the construct (and its auto-generated Principal:*
# invoke permission) reconciles IaC with the live state.
# --- Verified sites table (extracted from PO ship-to addresses) ---
verified_sites_table = dynamodb.Table(
self,
"VerifiedSitesTable",
table_name="verified-sites",
partition_key=dynamodb.Attribute(
name="siteCode",
type=dynamodb.AttributeType.STRING,
),
billing_mode=dynamodb.BillingMode.PAY_PER_REQUEST,
removal_policy=RemovalPolicy.RETAIN,
)
# by-state GSI removed 2026-06-03 (audit M-20): 0 reads in 30d against
# 518 WCU of write amplification. Re-add if a state-level query path ships.
# --- Site extractor Lambda (DynamoDB Streams → verified-sites) ---
site_extractor_log_group = common.make_function_log_group(
self, "SiteExtractor", "po-ingest-site-extractor"
)
site_extractor = lambda_.Function(
self,
"SiteExtractor",
function_name="po-ingest-site-extractor",
runtime=lambda_.Runtime.PYTHON_3_12,
architecture=lambda_.Architecture.ARM_64,
handler="handler.handler",
code=lambda_.Code.from_asset(
# requirements.txt is excluded from the bundle: it exists only
# as a Dependabot anchor (git-based scan sees it), never pip-
# installed (this is a plain non-bundled asset) and never needed
# at runtime (boto3 comes from the Lambda runtime). Excluding it
# keeps the deployed asset hash neutral vs base while the manifest
# still lands in git for Dependabot.
"../lambdas/po/site_extractor",
exclude=["**/__pycache__/**", "requirements.txt"],
),
timeout=Duration.seconds(60),
memory_size=256,
log_group=site_extractor_log_group,
environment={
"VERIFIED_SITES_TABLE": verified_sites_table.table_name,
"PENDING_REVIEW_TABLE": "pending-site-review",
},
)
verified_sites_table.grant_read_write_data(site_extractor)
site_extractor.add_event_source(
lambda_event_sources.DynamoEventSource(
po_table,
starting_position=lambda_.StartingPosition.TRIM_HORIZON,
batch_size=10,
max_batching_window=Duration.seconds(30),
bisect_batch_on_error=True,
retry_attempts=3,
)
)
# --- Standard per-Lambda alarms: po-ingest-site-extractor ---
# errors + throttles + p99 duration. NO DLQ alarm: site_extractor is a
# DynamoEventSource stream consumer with no async DLQ attached (dlq=None).
# Stream-consumer errors retry per the event-source config, but a
# persistent failure stalls the verified-sites pipeline. p99 / 45000 ms
# (75% of the 60s timeout) / eval 3, datapoints 2.
common.add_standard_lambda_alarms(
self,
"SiteExtractor",
site_extractor,
"po-ingest-site-extractor",
alarm_topic,
duration_statistic="p99",
errors=True,
dlq=None,
descriptions={
"errors": "po-ingest-site-extractor invocation errors",
"throttles": "po-ingest-site-extractor invocation throttles",
"duration": "po-ingest-site-extractor p99 duration approaching the 60s timeout",
},
)
cdk.CfnOutput(
self,
"VerifiedSitesTableName",
value=verified_sites_table.table_name,
description="Verified site addresses extracted from POs",
)
# --- Pending site review table (POs with no extractable site code) ---
pending_review_table = dynamodb.Table(
self,
"PendingSiteReviewTable",
table_name="pending-site-review",
partition_key=dynamodb.Attribute(
name="po_number",
type=dynamodb.AttributeType.STRING,
),
billing_mode=dynamodb.BillingMode.PAY_PER_REQUEST,
removal_policy=RemovalPolicy.RETAIN,
)
pending_review_table.grant_read_write_data(site_extractor)
verified_sites_table.grant_read_data(site_extractor)
# --- DynamoDB throttle + system-error alarms ---
# ThrottledRequests / SystemErrors emit at TableName + Operation only
# (verified against live CloudWatch: no TableName-only rollup exists, and
# metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.261.0).
# Each table currently has zero throttle/error datapoints, so the series
# only materialise on first occurrence — NOT_BREACHING keeps them OK until
# then.
common.add_ddb_alarms(
self, "PurchaseOrdersTable", po_table, "purchase-orders", alarm_topic
)
common.add_ddb_alarms(
self,
"VerifiedSitesTable",
verified_sites_table,
"verified-sites",
alarm_topic,
)
common.add_ddb_alarms(
self,
"PendingSiteReviewTable",
pending_review_table,
"pending-site-review",
alarm_topic,
)
# --- Function ARN + consumed-table-name outputs (Phase 4, additive) ---
cdk.CfnOutput(
self,
"EmailProcessorFunctionArn",
value=email_processor.function_arn,
description="ARN of the po-email-processor Lambda",
)
cdk.CfnOutput(
self,
"WebUiFunctionArn",
value=web_ui.function_arn,
description="ARN of the po-web-ui Lambda",
)
cdk.CfnOutput(
self,
"SiteExtractorFunctionArn",
value=site_extractor.function_arn,
description="ARN of the po-ingest-site-extractor Lambda",
)
cdk.CfnOutput(
self,
"PurchaseOrdersTableName",
value=po_table.table_name,
description="purchase-orders DynamoDB table consumed by this stack",
)
cdk.CfnOutput(
self,
"PendingSiteReviewTableName",
value=pending_review_table.table_name,
description="pending-site-review DynamoDB table",
)

View file

@ -1,427 +0,0 @@
"""procurement-api stack: read-only REST API over both pipelines' tables.
Third stack in the app. Serves work orders + comments (WO stack tables) and
purchase orders + verified sites (PO stack tables) to SigV4 callers -- the
primary consumer is the SHOC backend, for which this API replaces the retired
SyncController cross-account DynamoDB scan as the reconciliation/backfill
path. Also hosts the token-gated OpenAPI docs page (/docs, /openapi.json).
Tables are imported by fixed physical name (Table.from_table_name), NOT
passed as cross-stack objects: object passing would synthesize CFN Exports
from the owning stacks and lock them against future changes to the tables.
The one thing name-import does NOT carry is the purchase-orders CMK
association -- see the explicit KMS grant below.
"""
import aws_cdk as cdk
from aws_cdk import (
Duration,
Stack,
aws_apigateway as apigateway,
aws_certificatemanager as acm,
aws_cloudwatch as cloudwatch,
aws_cloudwatch_actions as cw_actions,
aws_dynamodb as dynamodb,
aws_iam as iam,
aws_kms as kms,
aws_lambda as lambda_,
aws_secretsmanager as secretsmanager,
aws_sns as sns,
aws_ssm as ssm,
)
from constructs import Construct
import common
# The only cross-account caller. Grants are to this exact role ARN -- future
# shoc-backend-staging/-prod roles are each a deliberate, individually
# cross-reviewed addition (no wildcard/prefix trust).
SHOC_BACKEND_DEV_ROLE_ARN = "arn:aws:iam::396287094661:role/shoc-backend-dev"
STAGE_NAME = "prod"
# Custom domain for the SHOC-facing read API. The seahaven.com zone is in the
# mgmt account, so the cert is issued out of band and its ARN handed in via
# this SSM parameter (see the custom-domain block below).
CUSTOM_DOMAIN_NAME = "procurement-api.seahaven.com"
CUSTOM_DOMAIN_CERT_ARN_SSM_PARAM = "/procurement-api/custom-domain/certificate-arn"
class ProcurementApiStack(Stack):
def __init__(self, scope: Construct, construct_id: str, **kwargs) -> None:
super().__init__(scope, construct_id, **kwargs)
alarm_topic = sns.Topic.from_topic_arn(
self,
"SiteAlertsTopic",
f"arn:aws:sns:{self.region}:{self.account}:site-alerts",
)
work_orders_table = dynamodb.Table.from_table_name(
self, "WorkOrdersTable", "WorkOrders"
)
comments_table = dynamodb.Table.from_table_name(
self, "CommentsTable", "WorkOrderComments"
)
po_table = dynamodb.Table.from_table_name(
self, "PurchaseOrdersTable", "purchase-orders"
)
verified_sites_table = dynamodb.Table.from_table_name(
self, "VerifiedSitesTable", "verified-sites"
)
web_ui_auth_secret = secretsmanager.Secret.from_secret_name_v2(
self,
"WebUiAuthToken",
"procurement-ingest/web-ui-auth-token",
)
# purchase-orders is SSE-KMS encrypted with the org DynamoDB CMK. A
# name-imported Table has no encryption-key association, so
# grant_read_data alone leaves the reader without kms:Decrypt and every
# purchase-orders read AccessDenies at runtime (the INFRA-104 failure
# class). Import the key from the same SSM parameter po_stack uses and
# grant it explicitly.
dynamodb_cmk = kms.Key.from_key_arn(
self,
"DynamoDbCmk",
ssm.StringParameter.value_for_string_parameter(
self, "/seahaven/dynamodb/cmk-arn"
),
)
# RETAIN log group (INFRA-114). This is the only stateful resource in
# this otherwise-stateless stack, so it carries the fresh-deploy
# rollback trap: if the FIRST create fails after the group exists,
# rollback deletes everything else but retains the group, and the retry
# CREATE then collides on "/aws/lambda/procurement-api already exists".
# Recovery: delete that log group before re-running a failed first
# deploy (same class as the RETAIN-orphan recovery in the deploy-role
# runbook / reference_cfn_deploy_role_gotchas).
api_log_group = common.make_function_log_group(
self, "ProcurementApi", "procurement-api"
)
api_fn = lambda_.Function(
self,
"ProcurementApi",
function_name="procurement-api",
runtime=lambda_.Runtime.PYTHON_3_12,
architecture=lambda_.Architecture.ARM_64,
handler="handler.handler",
code=lambda_.Code.from_asset(
"../lambdas",
exclude=["**/__pycache__/**"],
bundling=cdk.BundlingOptions(
image=lambda_.Runtime.PYTHON_3_12.bundling_image,
command=[
"bash",
"-c",
# Non-recursive glob ships every api/ sibling (the
# allowlist-omission trap from PR #105/PR #2); the spec,
# docs page, and the vendored Redoc bundle ride along
# because the handler serves them from its own package
# dir. web_ui_auth.py must land FLAT beside handler.py
# for the bare import.
"cp api/*.py /asset-output/ && "
"cp api/openapi.json /asset-output/ && "
"cp api/docs.html /asset-output/ && "
"cp api/redoc.standalone.js /asset-output/ && "
"cp api/fonts.css /asset-output/ && "
"cp shared/web_ui_auth.py /asset-output/ && "
"rm -rf /asset-output/__pycache__",
],
),
),
timeout=Duration.seconds(30),
memory_size=256,
log_group=api_log_group,
environment={
"WORK_ORDERS_TABLE": work_orders_table.table_name,
"COMMENTS_TABLE": comments_table.table_name,
"PO_TABLE": po_table.table_name,
"VERIFIED_SITES_TABLE": verified_sites_table.table_name,
"WEB_UI_AUTH_TOKEN_SECRET_ARN": web_ui_auth_secret.secret_arn,
},
)
work_orders_table.grant_read_data(api_fn)
comments_table.grant_read_data(api_fn)
po_table.grant_read_data(api_fn)
verified_sites_table.grant_read_data(api_fn)
web_ui_auth_secret.grant_read(api_fn)
api_fn.add_to_role_policy(
iam.PolicyStatement(
actions=["kms:Decrypt", "kms:DescribeKey"],
resources=[dynamodb_cmk.key_arn],
# Scope the grant to the DynamoDB data path only: the role can
# decrypt purchase-orders items via DynamoDB, never call
# kms:Decrypt directly on arbitrary ciphertext under the shared
# org CMK.
conditions={
"StringEquals": {
"kms:ViaService": f"dynamodb.{self.region}.amazonaws.com"
}
},
)
)
# Resource policy: once a REST API has one, anything not explicitly
# allowed is denied -- so the docs routes need their own Allow or the
# NONE-auth methods would be black-holed. The docs routes are still
# token-gated inside the Lambda (fail-closed), so this is not an
# unauthenticated data path (INFRA-74 posture).
#
# The SHOC data grant enumerates the exact GET resources rather than
# GET/*: adding a RESOURCE must be as reviewable as adding a PRINCIPAL,
# so a future GET route can't silently inherit cross-account reach
# without a policy diff. (Note: this resource policy binds CROSS-account
# callers only; a same-account principal holding execute-api:Invoke is
# authorized by its own identity policy under AWS union semantics -- it
# is NOT constrained here, including on the planned PATCH/POST methods,
# which is why those also rely on the 501 handler + absent write grant,
# not on this policy, until phase 2.)
shoc_data_resources = [
f"execute-api:/{STAGE_NAME}/GET/work-orders",
f"execute-api:/{STAGE_NAME}/GET/work-orders/*",
f"execute-api:/{STAGE_NAME}/GET/purchase-orders",
f"execute-api:/{STAGE_NAME}/GET/purchase-orders/*",
f"execute-api:/{STAGE_NAME}/GET/verified-sites",
f"execute-api:/{STAGE_NAME}/GET/verified-sites/*",
]
api_policy = iam.PolicyDocument(
statements=[
iam.PolicyStatement(
sid="ShocBackendDevDataRead",
principals=[iam.ArnPrincipal(SHOC_BACKEND_DEV_ROLE_ARN)],
actions=["execute-api:Invoke"],
resources=shoc_data_resources,
),
# SECURITY INVARIANT: this AnyPrincipal allow is safe only
# while /docs//openapi.json serve static docs (handler
# enforces the shared token, fail-closed). Widening these
# routes to dynamic data, or enabling access logging (which
# would record the docs ?token= shim), requires a security
# re-review + docs-token rotation.
iam.PolicyStatement(
sid="DocsTokenGatedRoutes",
principals=[iam.AnyPrincipal()],
actions=["execute-api:Invoke"],
resources=[
f"execute-api:/{STAGE_NAME}/GET/docs",
f"execute-api:/{STAGE_NAME}/GET/openapi.json",
],
),
]
)
# OPERATIONAL NOTE: API Gateway serves the resource policy from the
# deployed stage snapshot, and CDK's Deployment hash is computed from
# resources/methods, not the RestApi Policy. A later policy-ONLY change
# (e.g. revoking the SHOC role) will UPDATE the RestApi but keep serving
# the old policy until a new Deployment is forced (any method/resource
# change, or a salted deployment). When tightening this policy, force a
# redeploy and verify the effective policy post-deploy.
api = apigateway.RestApi(
self,
"ProcurementRestApi",
rest_api_name="procurement-api",
description=(
"Read API over procurement-ingest work orders + purchase "
"orders; token-gated OpenAPI docs at /docs"
),
endpoint_types=[apigateway.EndpointType.REGIONAL],
policy=api_policy,
deploy_options=apigateway.StageOptions(
stage_name=STAGE_NAME,
# Bound the blast radius of the unauthenticated /docs routes
# (and the whole API) below the 10k rps account default -- this
# is a low-volume reconciliation/backfill API, not a hot path.
throttling_rate_limit=50,
throttling_burst_limit=100,
),
# No access/execution logging in v1: avoids the account-level API
# Gateway CloudWatch role prerequisite AND keeps the docs ?token=
# query shim out of any log. Rotate the docs token before ever
# enabling access logging here.
cloud_watch_role=False,
)
integration = apigateway.LambdaIntegration(api_fn)
iam_auth = {"authorization_type": apigateway.AuthorizationType.IAM}
work_orders = api.root.add_resource("work-orders")
work_orders.add_method("GET", integration, **iam_auth)
wo_by_id = work_orders.add_resource("{workOrderId}")
wo_by_id.add_method("GET", integration, **iam_auth)
# Phase-2 planned write endpoints: deployed but the handler answers 501
# and the role holds no DynamoDB write grant. The resource policy denies
# these to the cross-account SHOC role (GET-only enumeration above); a
# same-account caller is NOT blocked by the resource policy, so the 501
# handler + absent write grant are the real gate until phase 2 lands the
# deliberate policy + handler + write-grant change with its own review.
wo_by_id.add_method("PATCH", integration, **iam_auth)
wo_comments = wo_by_id.add_resource("comments")
wo_comments.add_method("GET", integration, **iam_auth)
wo_comments.add_method("POST", integration, **iam_auth)
purchase_orders = api.root.add_resource("purchase-orders")
purchase_orders.add_method("GET", integration, **iam_auth)
purchase_orders.add_resource("{poNumber}").add_method(
"GET", integration, **iam_auth
)
verified_sites = api.root.add_resource("verified-sites")
verified_sites.add_method("GET", integration, **iam_auth)
verified_sites.add_resource("{siteCode}").add_method(
"GET", integration, **iam_auth
)
api.root.add_resource("docs").add_method(
"GET",
integration,
authorization_type=apigateway.AuthorizationType.NONE,
)
api.root.add_resource("openapi.json").add_method(
"GET",
integration,
authorization_type=apigateway.AuthorizationType.NONE,
)
# --- Custom domain: procurement-api.seahaven.com ---
# Stable, brandable endpoint for the SHOC backend to sign SigV4 against
# (replaces the opaque execute-api URL). REGIONAL to match the API, so
# the ACM cert lives in this account+region (us-east-1).
#
# CROSS-ACCOUNT DNS: the seahaven.com public zone
# (Z06652411XKH89KTZD3XA) is in the mgmt account (328440206208), not
# here. So the cert's DNS-validation CNAME and the final A-alias record
# are added to that zone OUT OF BAND (see the README runbook /
# scripts/setup_procurement_api_domain.sh), and the issued cert ARN is
# handed to this stack via SSM. Reading it with value_for_string_
# parameter keeps the reference CFN-resolved at deploy -- no synth-time
# AWS creds, unlike a cross-account HostedZone.from_lookup. The
# A-alias record itself is created out of band too (the zone is not in
# this account, so CDK cannot own it); the outputs below give the exact
# alias target.
cert_arn = ssm.StringParameter.value_for_string_parameter(
self, CUSTOM_DOMAIN_CERT_ARN_SSM_PARAM
)
certificate = acm.Certificate.from_certificate_arn(
self, "ProcurementApiCertificate", cert_arn
)
custom_domain = apigateway.DomainName(
self,
"ProcurementApiDomain",
domain_name=CUSTOM_DOMAIN_NAME,
certificate=certificate,
endpoint_type=apigateway.EndpointType.REGIONAL,
security_policy=apigateway.SecurityPolicy.TLS_1_2,
)
# Empty base path: the custom domain root maps to the prod stage, so
# callers hit https://procurement-api.seahaven.com/work-orders (no
# /prod segment -- the mapping strips it).
custom_domain.add_base_path_mapping(api, stage=api.deployment_stage)
# --- Alarms (ALARM-only -> site-alerts, NOT_BREACHING) ---
# common.add_standard_lambda_alarms is NOT used here: its duration
# threshold is a fixed 45000 ms (75% of the processors' 60 s timeout),
# which this function's 30 s timeout can never reach. Same idiom,
# right-sized thresholds.
api_fn.metric_errors(period=Duration.minutes(5), statistic="Sum").create_alarm(
self,
"ProcurementApiErrorsAlarm",
alarm_name="procurement-api-errors",
alarm_description=(
"procurement-api Lambda raised (bundle/init failures; the "
"handler catches request errors, so any signal here is "
"structural)"
),
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
api_fn.metric_throttles(
period=Duration.minutes(5), statistic="Sum"
).create_alarm(
self,
"ProcurementApiThrottlesAlarm",
alarm_name="procurement-api-throttles",
alarm_description="procurement-api Lambda throttled",
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
api_fn.metric_duration(
period=Duration.minutes(5), statistic="p99"
).create_alarm(
self,
"ProcurementApiDurationAlarm",
alarm_name="procurement-api-duration",
alarm_description=(
"procurement-api p99 duration >= 22.5s (75% of the 30s "
"timeout; scans degrading toward timeout)"
),
threshold=22500,
evaluation_periods=3,
datapoints_to_alarm=2,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
# Gateway-side 5xx: catches what the Lambda's own Errors metric can't
# (the handler returns clean 500s; integration faults surface here).
# No 4XX alarm -- 401/403/404 are expected traffic.
cloudwatch.Metric(
namespace="AWS/ApiGateway",
metric_name="5XXError",
dimensions_map={"ApiName": "procurement-api", "Stage": STAGE_NAME},
period=Duration.minutes(5),
statistic="Sum",
).create_alarm(
self,
"ProcurementApi5xxAlarm",
alarm_name="procurement-api-5xx",
alarm_description="procurement-api gateway 5XX responses",
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
cdk.CfnOutput(
self,
"ApiEndpointUrl",
value=api.url,
description="procurement-api invoke URL (stage prod)",
)
cdk.CfnOutput(
self,
"ProcurementApiFunctionArn",
value=api_fn.function_arn,
description="ARN of the procurement-api Lambda",
)
cdk.CfnOutput(
self,
"ProcurementApiCustomDomainUrl",
value=f"https://{CUSTOM_DOMAIN_NAME}/",
description="procurement-api custom domain (SHOC-facing endpoint)",
)
# A-alias target for the mgmt-account Route53 record. Create an ALIAS A
# record: procurement-api.seahaven.com -> this regional domain name,
# with this hosted-zone id, EvaluateTargetHealth=false.
cdk.CfnOutput(
self,
"ProcurementApiAliasTarget",
value=custom_domain.domain_name_alias_domain_name,
description="Route53 A-alias target (add in the mgmt seahaven.com zone)",
)
cdk.CfnOutput(
self,
"ProcurementApiAliasHostedZoneId",
value=custom_domain.domain_name_alias_hosted_zone_id,
description="Route53 A-alias target hosted-zone id (mgmt zone record)",
)

View file

@ -1,2 +0,0 @@
aws-cdk-lib==2.263.0
constructs==10.8.0

View file

@ -1,910 +0,0 @@
"""CDK stack for the work order email ingestion pipeline."""
import aws_cdk as cdk
import common
from aws_cdk import (
Duration,
RemovalPolicy,
Stack,
)
from aws_cdk import (
aws_cloudwatch as cloudwatch,
)
from aws_cdk import (
aws_cloudwatch_actions as cw_actions,
)
from aws_cdk import (
aws_dynamodb as dynamodb,
)
from aws_cdk import (
aws_iam as iam,
)
from aws_cdk import (
aws_kms as kms,
)
from aws_cdk import (
aws_lambda as lambda_,
)
from aws_cdk import (
aws_lambda_event_sources as lambda_event_sources,
)
from aws_cdk import (
aws_s3 as s3,
)
from aws_cdk import (
aws_s3_notifications as s3n,
)
from aws_cdk import (
aws_secretsmanager as secretsmanager,
)
from aws_cdk import (
aws_ses as ses,
)
from aws_cdk import (
aws_ses_actions as ses_actions,
)
from aws_cdk import (
aws_sns as sns,
)
from aws_cdk import (
aws_sqs as sqs,
)
from constructs import Construct
class WorkorderIngestStack(Stack):
def __init__(self, scope: Construct, construct_id: str, **kwargs):
super().__init__(scope, construct_id, **kwargs)
# --- Shared alarm SNS topic (site-alerts) ---
# Imported once near the top so every alarm in this stack reuses the same
# Topic construct instance (avoids duplicate logical IDs). ALARM-only
# SnsAction; no OK action, per the CloudWatch-alarm preference. The
# topic's CMK (alias/seahaven-alarm-topics) lives on the topic itself.
alarm_topic = sns.Topic.from_topic_arn(
self,
"SiteAlertsTopic",
f"arn:aws:sns:{self.region}:{self.account}:site-alerts",
)
# --- S3 bucket for raw emails ---
email_bucket = common.make_email_bucket(
self, "EmailBucket", "workorder-ingest-emails"
)
# --- DynamoDB tables ---
work_orders_table = dynamodb.Table(
self,
"WorkOrdersTable",
table_name="WorkOrders",
partition_key=dynamodb.Attribute(
name="work_order_id",
type=dynamodb.AttributeType.STRING,
),
billing_mode=dynamodb.BillingMode.PAY_PER_REQUEST,
# NEW_AND_OLD_IMAGES: the SHOC emitter needs OLD.wo_status to
# classify the cancelled transition (docs/shoc-webhook-plan.md
# Phase 2). In-place CFN update -- no table replacement.
stream=dynamodb.StreamViewType.NEW_AND_OLD_IMAGES,
removal_policy=RemovalPolicy.RETAIN,
)
# site-code-index and status-index GSIs removed 2026-06-03 (audit M-20):
# 0 reads in 30d against ~50k WCU each of write amplification. Re-add if
# a site-code or status query path ships.
comments_table = dynamodb.Table(
self,
"CommentsTable",
table_name="WorkOrderComments",
partition_key=dynamodb.Attribute(
name="work_order_id",
type=dynamodb.AttributeType.STRING,
),
sort_key=dynamodb.Attribute(
name="comment_id",
type=dynamodb.AttributeType.STRING,
),
billing_mode=dynamodb.BillingMode.PAY_PER_REQUEST,
# Streamed for the SHOC emitter (docs/shoc-webhook-plan.md Phase 2).
stream=dynamodb.StreamViewType.NEW_AND_OLD_IMAGES,
removal_policy=RemovalPolicy.RETAIN,
)
# --- Anthropic API key secret removed (Bedrock migration) ---
# Parsing moved from the Anthropic API to the Bedrock inference profile
# us.anthropic.claude-haiku-4-5-20251001-v1:0, so no provider API key is
# needed. The old secret "workorder-ingest/anthropic-api-key" had
# RemovalPolicy.RETAIN, so it is ORPHANED (not deleted) by this change:
# delete it manually post-deploy and revoke the stored key at Anthropic.
# --- DLQ for failed async invocations (INFRA-41 / audit H-8) ---
# SES → S3 → Lambda is async; without an OnFailure destination a failed
# parse (bad email, transient error) is silently dropped after Lambda's
# retries. CDK generates the queue name to avoid colliding with the
# interim CLI-created workorder-email-processor-dlq (removed post-deploy).
email_processor_dlq = common.make_processor_dlq(self, "EmailProcessorDlq")
# --- Lambda function ---
email_processor_log_group = common.make_function_log_group(
self, "EmailProcessor", "workorder-email-processor"
)
email_processor = lambda_.Function(
self,
"EmailProcessor",
function_name="workorder-email-processor",
runtime=lambda_.Runtime.PYTHON_3_12,
architecture=lambda_.Architecture.ARM_64,
handler="handler.handler",
code=lambda_.Code.from_asset(
"../lambdas",
exclude=["**/__pycache__/**", "**/tests/**", "**/package/**"],
bundling=cdk.BundlingOptions(
image=lambda_.Runtime.PYTHON_3_12.bundling_image,
command=[
"bash",
"-c",
# pip step removed in Phase 7: requirements.txt is now empty
# (boto3 comes from the Lambda runtime), so nothing is installed
# and the manylinux pin has nothing to pin. cp-only is safe.
# shared/*.py ships the four modules extracted to
# lambdas/shared/ (Phase 3): ses_auth, web_ui_auth,
# email_parsing, emf. Flat cp keeps the bare-name
# imports (e.g. `from ses_auth import ...`) resolving
# unchanged in /asset-output.
"cp wo/email_processor/*.py /asset-output/ && "
"cp shared/*.py /asset-output/",
],
),
),
timeout=Duration.seconds(60),
memory_size=256,
log_group=email_processor_log_group,
dead_letter_queue=email_processor_dlq,
environment={
"WORK_ORDERS_TABLE": work_orders_table.table_name,
"COMMENTS_TABLE": comments_table.table_name,
"BEDROCK_MODEL_ID": "us.anthropic.claude-haiku-4-5-20251001-v1:0",
# Fail-closed sender auth (INFRA-107): the handler only
# accepts mail whose SES-stamped Authentication-Results
# header carries dkim=pass for one of these domains. APM
# mail arrives via the apm@ Google Groups forward, which
# re-signs as seahaven.com (observed on live traffic
# 2026-07-15: "dkim=pass header.i=@seahaven.com"; the
# original hxgnsmartcloud.com signature does not survive
# the forward). Unset/empty ⇒ the handler rejects all mail.
"ALLOWED_DKIM_DOMAINS": "seahaven.com",
},
)
# Grant permissions
email_bucket.grant_read(email_processor)
work_orders_table.grant_read_write_data(email_processor)
comments_table.grant_read_write_data(email_processor)
# --- Bedrock InvokeModel grant ---
# The us.* inference profile can route cross-region, so the grant MUST
# cover both the inference-profile ARN AND the per-region foundation-model
# ARNs (empty account field) for every region the profile can reach
# (us-east-1/us-east-2/us-west-2). A profile-only grant AccessDenies at
# runtime whenever the profile routes to a region whose foundation-model
# ARN is not allowed.
email_processor.add_to_role_policy(common.make_bedrock_invoke_statement(self))
# NOTE: The pre-emptive grant_encrypt_decrypt on the shared DynamoDB CMK
# (alias/seahaven-dynamodb) was removed (security sweep 2026-06-17). The
# WorkOrders/WorkOrderComments tables are NOT SSE-KMS encrypted with that
# CMK, so the grant was unused for these tables yet handed
# wo-email-processor kms:Decrypt on the CMK that also protects the
# purchase-orders table (cross-stack decrypt reach). Re-add this grant only
# as part of the actual CMK migration of these tables (INFRA-6), at which
# point grant_read_write_data on the (then encrypted) tables would propagate
# the needed key permissions automatically.
# --- Standard per-Lambda alarms: workorder-email-processor ---
# errors (INFRA-41 / audit H-8), throttles, DLQ-visible-messages (dropped
# emails), and a p95 duration alarm (orphan adoption of the CLI
# Lambda-Duration-workorder-email-processor under <fn>-duration naming,
# 45000 ms = 75% of the 60s timeout, eval 3 / dp 2). p95 (NOT p99) is the
# WO-specific duration statistic. All ALARM-only to site-alerts.
common.add_standard_lambda_alarms(
self,
"EmailProcessor",
email_processor,
"workorder-email-processor",
alarm_topic,
duration_statistic="p95",
errors=True,
dlq=email_processor_dlq,
descriptions={
"errors": "workorder-email-processor async invocation errors",
"throttles": "workorder-email-processor invocation throttles",
"dlq": "workorder-email-processor DLQ has visible messages (dropped emails)",
"duration": "workorder-email-processor p95 duration approaching the 60s timeout",
},
)
# --- Sender-auth rejection alarm (INFRA-107) ---
# A rejected email (bad/unaligned DKIM verdict) returns normally, so it
# produces NO Lambda error, NO DLQ message and NO retry -- only a
# `sender_auth_rejected` warning log. The WO allowlist trusts dkim=pass
# for seahaven.com on the assumption the apm@ forward re-signs there; if
# that assumption is wrong (e.g. a Gmail auto-forward re-signs under a
# different domain), 100% of legitimate work-order mail is silently
# dropped. This metric filter + alarm turns those warnings into a paging
# signal so a false-reject storm surfaces instead of a silent outage.
common.add_sender_auth_rejected_alarm(
self,
"EmailProcessor",
"workorder-email-processor",
alarm_topic,
email_processor_log_group,
)
# --- Template fallback-rate alarm: workorder-email-processor ---
# The processor tries a deterministic template parse first and only calls
# the Bedrock AI extractor on a miss/invalid. A sustained rise in the
# ai_fallback share signals Hexagon template drift (coverage collapse).
# EMF metric Seahaven/WorkorderIngest/ParseOutcome, dimensioned by
# ParseMethod (template|ai_fallback). 15-min periods (deliberate deviation
# from the 5-min house style) accumulate a stable denominator at the low
# ~760/day volume; FILL(0) + a >=10-sample volume floor prevent
# low-volume false pages and INSUFFICIENT_DATA. ALARM-only SnsAction to
# site-alerts, no OK action, NOT_BREACHING -- matching the stack idiom.
# rejected_included=True: the WO expression folds ai_fallback_rejected
# (rej) into BOTH numerator and denominator -- a drift outage whose AI
# output also fails the gate must still count as fallback, otherwise it
# would LOWER the observed rate while silently dropping mail. (Contrast
# PO, which excludes rej to avoid a pre-call double-count.)
common.make_fallback_rate_alarm(
self,
"EmailProcessorTemplateFallbackRateAlarm",
namespace="Seahaven/WorkorderIngest",
alarm_topic=alarm_topic,
alarm_name="workorder-email-processor-template-fallback-rate",
alarm_description=(
"workorder-email-processor deterministic-template coverage "
"collapse: >15% of parses fell back to the Bedrock AI extractor"
),
rejected_included=True,
period=Duration.minutes(15),
threshold=15,
floor=10,
evaluation_periods=3,
datapoints_to_alarm=2,
)
# --- AI-fallback rejected alarm: workorder-email-processor ---
# A parse rejected by the validate_ai_fallback gate is dropped without
# error/retry/DLQ (fail closed), so like sender-auth rejections it
# needs its own pager or a sustained rejection condition (prompt-
# injection probing, or template drift whose AI output fails the gate)
# stays silent. Same sparse-arrival idiom as the sender-auth-rejected
# alarm: >=1 rejection per 5-min period, 2 of the last 6 periods (30
# min), so a lone probe self-clears but a burst pages within ~10 min.
# Coverage residual (matching the sender-auth-rejected sibling and
# knowingly accepted): rejections spaced >~25-30 min apart never place
# two breaching datapoints in one 30-min window, and the fallback-rate
# alarm dilutes them below 15% against normal template volume, so a
# *very* sparse silent-drop trickle is not paged by either alarm.
# EMF emits no datapoint in quiet periods (no metric-filter
# default_value here); NOT_BREACHING treats those gaps as OK.
cloudwatch.Metric(
namespace="Seahaven/WorkorderIngest",
metric_name="ParseOutcome",
dimensions_map={"ParseMethod": "ai_fallback_rejected"},
statistic="Sum",
period=Duration.minutes(5),
).create_alarm(
self,
"EmailProcessorAiFallbackRejectedAlarm",
alarm_name="workorder-email-processor-ai-fallback-rejected",
alarm_description=(
"workorder-email-processor is rejecting Bedrock AI-fallback "
"output at the validation gate (possible prompt-injection "
"probing or template drift silently dropping real mail)"
),
threshold=1,
evaluation_periods=6,
datapoints_to_alarm=2,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
# S3 event notification -> Lambda
email_bucket.add_event_notification(
s3.EventType.OBJECT_CREATED,
s3n.LambdaDestination(email_processor),
s3.NotificationKeyFilter(prefix="inbound/"),
)
# --- SES Receipt Rule ---
rule_set = ses.ReceiptRuleSet.from_receipt_rule_set_name(
self,
"ExistingRuleSet",
"INBOUND_MAIL",
)
rule_set.add_rule(
"WorkorderEmailRule",
recipients=["apm@int.seahaven.com"],
actions=[
ses_actions.S3(
bucket=email_bucket,
object_key_prefix="inbound/",
),
],
)
# --- Web UI auth token secret ---
# Shared secret for the web UI auth gate, stored in Secrets Manager and
# resolved at runtime so the token never appears in CloudFormation templates
# or Lambda environment variables. Create this secret before deploying
# either stack; both PO and WO stacks reference it by name.
web_ui_auth_secret = secretsmanager.Secret.from_secret_name_v2(
self,
"WebUiAuthToken",
"procurement-ingest/web-ui-auth-token",
)
# --- Web UI Lambda ---
web_ui_log_group = common.make_function_log_group(
self, "WebUI", "workorder-web-ui"
)
web_ui = lambda_.Function(
self,
"WebUI",
function_name="workorder-web-ui",
runtime=lambda_.Runtime.PYTHON_3_12,
architecture=lambda_.Architecture.ARM_64,
handler="handler.handler",
code=lambda_.Code.from_asset(
"../lambdas",
exclude=["**/__pycache__/**"],
bundling=cdk.BundlingOptions(
image=lambda_.Runtime.PYTHON_3_12.bundling_image,
command=[
"bash",
"-c",
# web_ui_auth.py is shared (lambdas/shared) and must land
# FLAT beside handler.py so
# `from web_ui_auth import is_authenticated` resolves at
# runtime. Only web_ui_auth is copied from shared/. WO
# baseline keeps __init__.py and requirements.txt, so the
# whole web_ui dir is copied; deployed file list becomes
# {__init__, handler, requirements.txt, web_ui_auth}.
# NOTE: the top-level `exclude=` on from_asset only
# filters the asset-hash fingerprint, NOT the directory
# Docker bundling actually mounts, so a local
# __pycache__ on disk at synth time WOULD otherwise leak
# into the bundled zip -- strip it explicitly post-cp
# instead of relying on exclude.
"cp -r wo/web_ui/. /asset-output/ && "
"cp shared/web_ui_auth.py /asset-output/ && "
"rm -rf /asset-output/__pycache__",
],
),
),
timeout=Duration.seconds(15),
memory_size=128,
log_group=web_ui_log_group,
environment={
"WORK_ORDERS_TABLE": work_orders_table.table_name,
"COMMENTS_TABLE": comments_table.table_name,
# Defense-in-depth shared secret for the web UI handler. The
# handler fails closed if this ARN is unset or the secret is
# missing, so any future invocation path cannot re-expose the
# WO DB unauthenticated. The secret value is fetched at runtime
# from Secrets Manager (not embedded in env vars or template).
"WEB_UI_AUTH_TOKEN_SECRET_ARN": web_ui_auth_secret.secret_arn,
},
)
work_orders_table.grant_read_data(web_ui)
comments_table.grant_read_data(web_ui)
web_ui_auth_secret.grant_read(web_ui)
# --- DynamoDB throttle + system-error alarms ---
# ThrottledRequests / SystemErrors emit at TableName + Operation only
# (verified against live CloudWatch: no TableName-only rollup exists, and
# metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.261.0).
# Each table currently has zero throttle/error datapoints, so the series
# only materialise on first occurrence — NOT_BREACHING keeps them OK until
# then.
common.add_ddb_alarms(
self, "WorkOrdersTable", work_orders_table, "WorkOrders", alarm_topic
)
common.add_ddb_alarms(
self, "WorkOrderComments", comments_table, "WorkOrderComments", alarm_topic
)
# --- Function ARN + consumed-table-name outputs (Phase 4, additive) ---
cdk.CfnOutput(
self,
"EmailProcessorFunctionArn",
value=email_processor.function_arn,
description="ARN of the workorder-email-processor Lambda",
)
cdk.CfnOutput(
self,
"WebUiFunctionArn",
value=web_ui.function_arn,
description="ARN of the workorder-web-ui Lambda",
)
cdk.CfnOutput(
self,
"WorkOrdersTableName",
value=work_orders_table.table_name,
description="WorkOrders DynamoDB table",
)
cdk.CfnOutput(
self,
"WorkOrderCommentsTableName",
value=comments_table.table_name,
description="WorkOrderComments DynamoDB table",
)
# Public Function URL removed 2026-06-08 (INFRA-74 / audit C-5): the
# unauthenticated FunctionUrlAuthType.NONE URL was deleted out-of-band
# via CLI. Removing the construct (and its auto-generated Principal:*
# invoke permission) reconciles IaC with the live state.
# =====================================================================
# --- SHOC webhook emitter (docs/shoc-webhook-plan.md) ---
# Realtime work-order feed to the SHOC backend: DynamoDB Streams on the
# two WO tables -> workorder-shoc-emitter -> HMAC-signed HTTPS POST.
# Contract: docs/shoc-webhook-contract.md (Rev 2026-07-23). Built in
# _add_shoc_webhook_emitter (module helper, common.py plain-helper
# style) to keep __init__ under the PLR0915 statement ceiling; the
# stack stays the construct scope, so extraction does not move any
# logical ID.
# =====================================================================
_add_shoc_webhook_emitter(self, work_orders_table, comments_table, alarm_topic)
def _scope_rotation_invoke_permission(stack, rotator_fn, secret):
"""Add SourceAccount/SourceArn to the generated rotation invoke permission.
``add_rotation_schedule`` emits an ``AWS::Lambda::Permission`` for the
``secretsmanager.amazonaws.com`` service principal with no source
conditions. Rather than add a second (additive) permission, find that
generated ``CfnPermission`` and pin it to this account + secret so only
this secret's Secrets Manager can invoke the rotator.
"""
patched = False
for child in stack.node.find_all():
if (
isinstance(child, lambda_.CfnPermission)
and child.principal == "secretsmanager.amazonaws.com"
and stack.resolve(child.function_name)
== stack.resolve(rotator_fn.function_arn)
):
child.source_account = stack.account
child.source_arn = secret.secret_arn
patched = True
if not patched: # fail loud if CDK changes the generated shape on upgrade
raise RuntimeError(
"rotation invoke CfnPermission not found; cannot scope source conditions"
)
def _add_shoc_webhook_emitter(stack, work_orders_table, comments_table, alarm_topic):
"""SHOC webhook emitter (docs/shoc-webhook-plan.md Phases 1-4).
Dedicated HMAC CMK + secret with cross-account SHOC read grants, the
30-day rotation Lambda, the two failure queues, the stream-driven emitter
Lambda with its two (dark, enabled=False) event source mappings, and the
full alarm set. `stack` is the construct scope for every child, exactly as
if this code were inline in __init__.
"""
# SHOC consumer principal for the cross-account read grants. EXACT role
# ARN only -- future shoc-backend-staging/-prod roles are each a
# deliberate, individually-reviewed policy addition (no wildcard or
# prefix trust).
shoc_consumer_principal = iam.ArnPrincipal(
"arn:aws:iam::396287094661:role/shoc-backend-dev"
)
# --- Dedicated CMK for the HMAC secret ---
# Dedicated key, NOT alias/seahaven-dynamodb: reusing the DynamoDB CMK
# would hand the SHOC cross-account grant decrypt reach over the PO
# table's encryption key -- the dedicated key scopes the grant to
# exactly this secret (plan Phase 1).
shoc_webhook_key = kms.Key(
stack,
"ShocWebhookHmacKey",
alias="workorder-ingest-shoc-webhook-kms",
description=(
"Dedicated CMK for the workorder-ingest/shoc-webhook-hmac "
"secret (cross-account readable by the SHOC backend)"
),
enable_key_rotation=True,
# GOTCHA: DESTROY is deliberate -- do not "harden" this to RETAIN.
# The key protects only machine-generated HMAC material that is
# fully regenerable by one rotation, and DESTROY avoids the
# fixed-name RETAIN-orphan deadlock on the alias (mirrors the
# secret's rationale below).
removal_policy=RemovalPolicy.DESTROY,
pending_window=Duration.days(7),
)
# Key-policy half of the cross-account read grant (the secret resource
# policy below is the other half; either one alone fails silently at
# the receiver). resources=["*"] is key-scoped, not account-wide --
# KMS key policies only ever apply to this key.
# kms:ViaService pins the grant to Secrets Manager decrypt paths only
# (GPT-4.1 cross-review FIX): a compromised shoc-backend-dev cannot use
# this key for arbitrary KMS operations outside the secret fetch.
shoc_webhook_key.add_to_resource_policy(
iam.PolicyStatement(
actions=["kms:Decrypt"],
principals=[shoc_consumer_principal],
resources=["*"],
conditions={
"StringEquals": {
"kms:ViaService": (f"secretsmanager.{stack.region}.amazonaws.com")
}
},
)
)
# --- HMAC signing secret ---
# Value shape (contract section 6.1):
# {"keys": [{"kid": "<YYYY-MM-DDTHH>", "secret": "<64 hex>"}, ...]},
# newest first, max 2; the producer signs with keys[0]. The
# generate_secret_string below is BOOTSTRAP shape only ({"keys": []}
# plus throwaway entropy the rotator ignores); the first rotation
# (rotate_immediately default) populates the real keys.
shoc_hmac_secret = secretsmanager.Secret(
stack,
"ShocWebhookHmacSecret",
secret_name="workorder-ingest/shoc-webhook-hmac",
encryption_key=shoc_webhook_key,
description=(
"HMAC signing keys for the SHOC work-order webhook "
"(docs/shoc-webhook-contract.md section 6)"
),
# GOTCHA: DESTROY is deliberate -- do not "harden" this to RETAIN.
# The value is machine-generated HMAC material with no operator-set
# content, fully regenerable by one rotation, so RETAIN buys
# nothing and would expose the fixed-name RETAIN orphan deadlock
# (a failed first create orphans an empty shell holding the global
# name; see reference_secret_retain_orphan_deadlock).
removal_policy=RemovalPolicy.DESTROY,
generate_secret_string=secretsmanager.SecretStringGenerator(
secret_string_template='{"keys": []}',
generate_string_key="bootstrap_entropy",
password_length=32,
exclude_punctuation=True,
),
)
cdk.Tags.of(shoc_hmac_secret).add("Purpose", "shoc-webhook-hmac")
cdk.Tags.of(shoc_hmac_secret).add("ManagedBy", "procurement-ingest-cdk")
# Secret-resource-policy half of the cross-account read grant (the key
# policy above is the other half). DescribeSecret lets the receiver
# resolve secret metadata without any broader list permission.
shoc_hmac_secret.add_to_resource_policy(
iam.PolicyStatement(
actions=[
"secretsmanager:GetSecretValue",
"secretsmanager:DescribeSecret",
],
principals=[shoc_consumer_principal],
resources=["*"],
)
)
# --- HMAC rotation Lambda ---
# 30-day schedule: generates a new key, prepends as keys[0], truncates
# to 2 entries. Single-user rotation (receivers re-fetch on a <=5-min
# TTL), so the standard 4-step rotation collapses to
# createSecret/finishSecret.
shoc_hmac_rotator_log_group = common.make_function_log_group(
stack, "ShocHmacRotator", "workorder-shoc-hmac-rotator"
)
shoc_hmac_rotator = lambda_.Function(
stack,
"ShocHmacRotator",
function_name="workorder-shoc-hmac-rotator",
runtime=lambda_.Runtime.PYTHON_3_12,
architecture=lambda_.Architecture.ARM_64,
handler="handler.handler",
code=lambda_.Code.from_asset(
"../lambdas",
exclude=["**/__pycache__/**", "**/tests/**", "**/package/**"],
bundling=cdk.BundlingOptions(
image=lambda_.Runtime.PYTHON_3_12.bundling_image,
command=[
"bash",
"-c",
# Stdlib + boto3-from-runtime only; nothing installed.
"cp wo/shoc_hmac_rotator/*.py /asset-output/",
],
),
),
timeout=Duration.seconds(60),
memory_size=128,
log_group=shoc_hmac_rotator_log_group,
)
# Rotation permissions, scoped to the one secret. CDK's secret_arn
# token resolves to the full ARN including the -?????? suffix wildcard,
# so no separate "*"-suffixed resource variant is needed.
shoc_hmac_rotator.add_to_role_policy(
iam.PolicyStatement(
actions=[
"secretsmanager:DescribeSecret",
"secretsmanager:GetSecretValue",
"secretsmanager:PutSecretValue",
"secretsmanager:UpdateSecretVersionStage",
],
resources=[shoc_hmac_secret.secret_arn],
)
)
# Explicit statement instead of grant_encrypt_decrypt so the grant can
# carry kms:ViaService (cross-review FIX): the rotator only ever touches
# this key through Secrets Manager put/get, never the KMS API directly.
shoc_hmac_rotator.add_to_role_policy(
iam.PolicyStatement(
actions=[
"kms:Decrypt",
"kms:Encrypt",
"kms:GenerateDataKey*",
"kms:ReEncrypt*",
],
resources=[shoc_webhook_key.key_arn],
conditions={
"StringEquals": {
"kms:ViaService": (f"secretsmanager.{stack.region}.amazonaws.com")
}
},
)
)
shoc_hmac_secret.add_rotation_schedule(
"Rotation",
rotation_lambda=shoc_hmac_rotator,
automatically_after=Duration.days(30),
)
# Scope the Secrets-Manager-service invoke permission to THIS secret
# (cross-review FIX / confused-deputy): add_rotation_schedule emits an
# AWS::Lambda::Permission for secretsmanager.amazonaws.com with no
# SourceAccount/SourceArn, so any account's Secrets Manager could invoke
# the rotator by pointing a foreign secret's RotationLambdaARN at it.
# Lambda permissions are additive (OR), so a second scoped permission
# would NOT revoke the unscoped one -- patch the generated permission in
# place. source_arn pins the invoker to this secret; source_account is the
# belt-and-braces account bound. (Blast radius was already contained by
# the rotator role being resource-scoped to this secret, but this closes
# the unauthenticated invoke primitive per AWS rotation guidance.)
_scope_rotation_invoke_permission(stack, shoc_hmac_rotator, shoc_hmac_secret)
# --- Standard per-Lambda alarms: workorder-shoc-hmac-rotator ---
# errors + throttles + p99 duration. No DLQ alarm: rotation is invoked
# synchronously by Secrets Manager (dlq=None); a failed rotation
# surfaces as an invocation error.
common.add_standard_lambda_alarms(
stack,
"ShocHmacRotator",
shoc_hmac_rotator,
"workorder-shoc-hmac-rotator",
alarm_topic,
duration_statistic="p99",
errors=True,
dlq=None,
descriptions={
"errors": "workorder-shoc-hmac-rotator invocation errors",
"throttles": "workorder-shoc-hmac-rotator invocation throttles",
"duration": (
"workorder-shoc-hmac-rotator p99 duration approaching the 60s timeout"
),
},
)
# --- Emitter failure queues ---
# Failures queue: ESM on_failure destination. It receives ESM failure
# METADATA (shard/sequence pointers), not full payloads -- replay
# rebuilds events from DynamoDB (contract section 8).
shoc_emitter_failures_queue = sqs.Queue(
stack,
"ShocEmitterFailuresQueue",
queue_name="workorder-shoc-emitter-failures",
retention_period=Duration.days(14),
enforce_ssl=True,
)
# Rejected queue: full {envelope, response_status} payloads parked by
# the handler on non-retryable 4xx responses (contract section 7).
shoc_emitter_rejected_queue = sqs.Queue(
stack,
"ShocEmitterRejectedQueue",
queue_name="workorder-shoc-emitter-rejected",
retention_period=Duration.days(14),
enforce_ssl=True,
)
# --- Emitter Lambda ---
shoc_emitter_log_group = common.make_function_log_group(
stack, "ShocEmitter", "workorder-shoc-emitter"
)
shoc_emitter = lambda_.Function(
stack,
"ShocEmitter",
function_name="workorder-shoc-emitter",
runtime=lambda_.Runtime.PYTHON_3_12,
architecture=lambda_.Architecture.ARM_64,
handler="handler.handler",
code=lambda_.Code.from_asset(
"../lambdas",
exclude=["**/__pycache__/**", "**/tests/**", "**/package/**"],
bundling=cdk.BundlingOptions(
image=lambda_.Runtime.PYTHON_3_12.bundling_image,
command=[
"bash",
"-c",
# Stdlib HTTP (urllib.request) + boto3-from-runtime
# only; nothing installed.
"cp wo/shoc_emitter/*.py /asset-output/",
],
),
),
timeout=Duration.seconds(60),
memory_size=256,
log_group=shoc_emitter_log_group,
environment={
# Non-sensitive endpoint URL (HMAC is the auth, the URL is
# not). path TBD by SHOC -- confirmed in the activation PR.
"SHOC_WEBHOOK_URL": (
"https://api.dev.seahaven.com/api/webhooks/work-orders"
),
"HMAC_SECRET_ARN": shoc_hmac_secret.secret_arn,
"REJECTED_QUEUE_URL": shoc_emitter_rejected_queue.queue_url,
},
)
work_orders_table.grant_stream_read(shoc_emitter)
comments_table.grant_stream_read(shoc_emitter)
shoc_hmac_secret.grant_read(shoc_emitter)
# Explicit statement instead of grant_decrypt so the grant carries
# kms:ViaService (cross-review FIX): the emitter only decrypts this key
# through Secrets Manager GetSecretValue.
shoc_emitter.add_to_role_policy(
iam.PolicyStatement(
actions=["kms:Decrypt"],
resources=[shoc_webhook_key.key_arn],
conditions={
"StringEquals": {
"kms:ViaService": (f"secretsmanager.{stack.region}.amazonaws.com")
}
},
)
)
shoc_emitter_rejected_queue.grant_send_messages(shoc_emitter)
# ------------------------------------------------------------------
# ACTIVATED (enabled=True) 2026-07-30 after SHOC's receiver passed the
# shared HMAC test vectors. Shipped DARK originally (enabled=False) so the
# stack could deploy and be tested with zero deliveries while SHOC had no
# receiver; this activation PR is the deliberate one-line flip (plan Phase
# 3). LATEST start position => the feed begins now, no historical flood;
# SHOC backfills history via the procurement read API, not the stream.
# Ordering knobs: parallelization_factor=1, bisect_batch_on_error=
# False and retry_attempts=-1 (retry until the 24h record age) are
# REQUIRED for strict per-work-order in-order delivery -- a retryable
# failure blocks the shard rather than skipping ahead, and
# report_batch_item_failures keeps earlier in-batch successes from
# being re-delivered.
# ------------------------------------------------------------------
shoc_emitter.add_event_source(
lambda_event_sources.DynamoEventSource(
work_orders_table,
starting_position=lambda_.StartingPosition.LATEST,
batch_size=10,
bisect_batch_on_error=False,
retry_attempts=-1,
max_record_age=Duration.hours(24),
parallelization_factor=1,
report_batch_item_failures=True,
enabled=True,
on_failure=lambda_event_sources.SqsDlq(shoc_emitter_failures_queue),
)
)
shoc_emitter.add_event_source(
lambda_event_sources.DynamoEventSource(
comments_table,
starting_position=lambda_.StartingPosition.LATEST,
batch_size=10,
bisect_batch_on_error=False,
retry_attempts=-1,
max_record_age=Duration.hours(24),
parallelization_factor=1,
report_batch_item_failures=True,
enabled=True,
on_failure=lambda_event_sources.SqsDlq(shoc_emitter_failures_queue),
)
)
# --- Standard per-Lambda alarms: workorder-shoc-emitter ---
# errors + throttles + p99 duration. No DLQ alarm here: the emitter is
# a stream consumer with no async DLQ (dlq=None); its failure surfaces
# are the two SQS queues alarmed bespoke below.
common.add_standard_lambda_alarms(
stack,
"ShocEmitter",
shoc_emitter,
"workorder-shoc-emitter",
alarm_topic,
duration_statistic="p99",
errors=True,
dlq=None,
descriptions={
"errors": "workorder-shoc-emitter invocation errors",
"throttles": "workorder-shoc-emitter invocation throttles",
"duration": (
"workorder-shoc-emitter p99 duration approaching the 60s timeout"
),
},
)
# --- Bespoke emitter alarms (plan Phase 4) ---
# These don't fit add_standard_lambda_alarms' shape and stay bespoke.
# Iterator age >= 10 min sustained means SHOC is likely down and the
# shard is blocking (exactly the ordered-backpressure design working);
# the two SQS-visible alarms page the operator replay runbook.
shoc_emitter.metric(
"IteratorAge",
statistic="Maximum",
period=Duration.minutes(5),
).create_alarm(
stack,
"ShocEmitterIteratorAgeAlarm",
alarm_name="workorder-shoc-emitter-iterator-age",
alarm_description=(
"workorder-shoc-emitter stream lag >= 10 min "
"(SHOC receiver likely down; shard blocking on retries)"
),
threshold=600000,
evaluation_periods=3,
datapoints_to_alarm=2,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_OR_EQUAL_TO_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
shoc_emitter_failures_queue.metric_approximate_number_of_messages_visible(
period=Duration.minutes(5),
statistic="Maximum",
).create_alarm(
stack,
"ShocEmitterFailuresMessagesAlarm",
alarm_name="workorder-shoc-emitter-failures-messages",
alarm_description=(
"workorder-shoc-emitter retry-exhausted stream records parked "
"(ESM failure metadata; replay rebuilds from DynamoDB)"
),
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
shoc_emitter_rejected_queue.metric_approximate_number_of_messages_visible(
period=Duration.minutes(5),
statistic="Maximum",
).create_alarm(
stack,
"ShocEmitterRejectedMessagesAlarm",
alarm_name="workorder-shoc-emitter-rejected-messages",
alarm_description=(
"workorder-shoc-emitter parked non-retryable 4xx deliveries "
"(contract bug; inspect payloads and replay)"
),
threshold=0,
evaluation_periods=1,
comparison_operator=cloudwatch.ComparisonOperator.GREATER_THAN_THRESHOLD,
treat_missing_data=cloudwatch.TreatMissingData.NOT_BREACHING,
).add_alarm_action(cw_actions.SnsAction(alarm_topic))
cdk.CfnOutput(
stack,
"ShocWebhookHmacSecretArn",
value=shoc_hmac_secret.secret_arn,
description=(
"SHOC webhook HMAC secret ARN -- hand off to Luby (SHOC team) "
"for the cross-account receiver fetch"
),
)

View file

@ -1,6 +1,5 @@
# Phase 8: make the C901/PLR complexity ceilings ENFORCED, not decorative —
# the pre-existing `noqa: PLR09xx` suppressions were inert with no config.
extend-exclude = ["cdk.out"]
[lint]
extend-select = ["C901", "PLR"]
@ -26,8 +25,5 @@ max-complexity = 12
"lambdas/wo/email_processor/template_parser.py" = ["C901", "PLR0911", "PLR0912"]
"lambdas/shared/ses_auth.py" = ["C901", "PLR0911"]
"lambdas/shared/emf.py" = ["PLR0913"] # 6-arg emitter IS the Phase 3 envelope contract
# Phase 4 plain helpers take (scope, id, ...) positionally by design to keep
# construct logical IDs byte-stable; the arg count IS the contract (like emf.py).
"cdk/common.py" = ["PLR0913"]
"lambdas/po/site_extractor/handler.py" = ["PLR0911"] # Phase 6 owns this file
"scripts/reprocess.py" = ["PLR0912"] # argparse validation ladder; Phase 7 file

View file

@ -14,14 +14,14 @@
# - adds its DNS-validation CNAME to the mgmt seahaven.com zone
# - waits for ISSUED, then writes the cert ARN to prod SSM
# (/procurement-api/custom-domain/certificate-arn)
# 2. deploy the procurement-api stack (cdk deploy procurement-api) -- it
# reads the SSM param and creates the API Gateway DomainName + mapping
# 2. HCP Terraform apply on workspace procurement-ingest-prod -- it reads
# the SSM param and owns the API Gateway DomainName + mapping
# 3. ./setup_procurement_api_domain.sh alias
# - reads the stack's regional alias target from the outputs
# - reads the regional alias target from API Gateway get-domain-name
# - adds the A-alias (procurement-api.seahaven.com -> API GW) to the
# mgmt zone
#
# Requires SSO sessions for BOTH profiles (prod for ACM/SSM, mgmt for Route53).
# Requires SSO sessions for BOTH profiles (prod for ACM/SSM/API GW, mgmt for Route53).
###############################################################################
set -euo pipefail
@ -33,7 +33,6 @@ PROD_ACCOUNT="011934824531"
MGMT_ACCOUNT="328440206208"
ZONE_ID="Z06652411XKH89KTZD3XA" # seahaven.com public zone, in the mgmt account
SSM_PARAM="/procurement-api/custom-domain/certificate-arn"
STACK_NAME="procurement-api"
_verify_account() {
local profile="$1" expected="$2"
@ -89,23 +88,23 @@ JSON
aws ssm put-parameter --profile "${PROD_PROFILE}" --region "${REGION}" \
--name "${SSM_PARAM}" --type String --overwrite --value "${cert_arn}" >/dev/null
echo "OK: cert ISSUED and SSM param set. Now: cdk deploy ${STACK_NAME}, then '$0 alias'."
echo "OK: cert ISSUED and SSM param set. Next: HCP apply on procurement-ingest-prod, then '$0 alias'."
}
cmd_alias() {
_verify_account "${PROD_PROFILE}" "${PROD_ACCOUNT}"
_verify_account "${MGMT_PROFILE}" "${MGMT_ACCOUNT}"
echo "==> Reading regional alias target from the ${STACK_NAME} stack outputs"
echo "==> Reading regional alias target from API Gateway domain ${DOMAIN}"
local target zone
target="$(aws cloudformation describe-stacks --profile "${PROD_PROFILE}" --region "${REGION}" \
--stack-name "${STACK_NAME}" \
--query "Stacks[0].Outputs[?OutputKey=='ProcurementApiAliasTarget'].OutputValue | [0]" --output text)"
zone="$(aws cloudformation describe-stacks --profile "${PROD_PROFILE}" --region "${REGION}" \
--stack-name "${STACK_NAME}" \
--query "Stacks[0].Outputs[?OutputKey=='ProcurementApiAliasHostedZoneId'].OutputValue | [0]" --output text)"
target="$(aws apigateway get-domain-name --profile "${PROD_PROFILE}" --region "${REGION}" \
--domain-name "${DOMAIN}" \
--query "regionalDomainName" --output text)"
zone="$(aws apigateway get-domain-name --profile "${PROD_PROFILE}" --region "${REGION}" \
--domain-name "${DOMAIN}" \
--query "regionalHostedZoneId" --output text)"
if [[ -z "${target}" || "${target}" == "None" ]]; then
echo "ERROR: no alias target output; deploy the stack first." >&2
echo "ERROR: no regionalDomainName for ${DOMAIN}; HCP-apply the API custom domain first." >&2
exit 1
fi

View file

@ -4,5 +4,6 @@ web_ui_auth_token_secret_arn = "arn:aws:secretsmanager:us-east-1:011934824531:se
shoc_hmac_secret_arn = "arn:aws:secretsmanager:us-east-1:011934824531:secret:workorder-ingest/shoc-webhook-hmac-puYTcB"
# Required: no Terraform default. Live emitter target today (contract Rev 2026-07-23).
shoc_webhook_url = "https://api.dev.seahaven.com/api/webhooks/work-orders"
# Required exact-ARN pin for cross-account HMAC read (docs/shoc-webhook-contract.md).
# Exact-ARN pin for cross-account HMAC/API trust (matches variable default;
# docs/shoc-webhook-contract.md). Changing this is a deliberate IAM review.
shoc_consumer_role_arn = "arn:aws:iam::396287094661:role/shoc-backend-dev"

View file

@ -21,5 +21,6 @@ variable "shoc_webhook_url" {
variable "shoc_consumer_role_arn" {
type = string
description = "Exact IAM role ARN allowed to GetSecretValue / kms:Decrypt the SHOC HMAC secret (cross-account consumer). Live pin today is arn:aws:iam::396287094661:role/shoc-backend-dev per docs/shoc-webhook-contract.md."
description = "Exact IAM role ARN allowed to GetSecretValue / kms:Decrypt the SHOC HMAC secret and invoke the read API (cross-account consumer). Default is the live pin per docs/shoc-webhook-contract.md; changing it is a deliberate cross-family IAM review."
default = "arn:aws:iam::396287094661:role/shoc-backend-dev"
}

File diff suppressed because it is too large Load diff

View file

@ -1,12 +1,15 @@
"""Pin the cross-account principal surface of the CDK app.
"""Pin the cross-account principal surface of the Terraform app.
The SHOC integration deliberately trusts EXACTLY ONE foreign principal:
``arn:aws:iam::396287094661:role/shoc-backend-dev`` (read API resource
policy in procurement_api_stack.py, HMAC secret + KMS grants in
wo_stack.py). Future shoc-backend-staging/-prod roles are each a
deliberate, individually-reviewed policy addition — so any new foreign
account id or role ARN appearing in cdk/ must consciously update this
pin (and go through the mandatory GPT-4.1 cross-family IAM review).
``arn:aws:iam::396287094661:role/shoc-backend-dev`` (API resource policy in
api.tf, HMAC secret + KMS grants in wo_shoc.tf). The exact ARN is pinned as
the ``shoc_consumer_role_arn`` variable default and in
``terraform.tfvars.example``; grant sites consume ``local.shoc_consumer_role_arn``
only. Future shoc-backend-staging/-prod roles are each a deliberate,
individually-reviewed policy addition — so any new foreign account id or role
ARN appearing under terraform/, or any change to the default/example pin, must
consciously update this test (and go through the mandatory GPT-4.1 cross-family
IAM review).
Raised as a QUESTION in the 2026-07-24 cross-family review of the
webhook emitter policy surface: "how is the exact-one-principal
@ -17,26 +20,77 @@ import re
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
CDK_DIR = REPO_ROOT / "cdk"
TF_DIR = REPO_ROOT / "terraform"
VARIABLES_TF = TF_DIR / "variables.tf"
TFVARS_EXAMPLE = TF_DIR / "terraform.tfvars.example"
# The one foreign principal the app may reference, and the only files
# allowed to reference it.
# The one foreign principal the app may trust.
ALLOWED_FOREIGN_PRINCIPAL = "arn:aws:iam::396287094661:role/shoc-backend-dev"
ALLOWED_FILES = {"procurement_api_stack.py", "wo_stack.py"}
# Files allowed to embed that ARN as a literal (default + example). Grant
# sites must use local.shoc_consumer_role_arn instead.
ALLOWED_LITERAL_FILES = {"variables.tf", "terraform.tfvars.example"}
# Grant sites that must reference the local; dropping one half fails the pin.
GRANT_FILES = ("api.tf", "wo_shoc.tf")
# Accounts that are not "foreign": seahaven-prod (the deploy target).
HOME_ACCOUNTS = {"011934824531"}
_IAM_ARN_RE = re.compile(r"arn:aws:iam::(\d{12}):\S*?(?=[\"'\s])")
_VAR_DEFAULT_RE = re.compile(
r'variable\s+"shoc_consumer_role_arn"\s*\{(.*?)^\}',
re.DOTALL | re.MULTILINE,
)
_DEFAULT_VALUE_RE = re.compile(r'default\s*=\s*"([^"]+)"')
_TFVARS_VALUE_RE = re.compile(
r'^shoc_consumer_role_arn\s*=\s*"([^"]+)"\s*$',
re.MULTILINE,
)
_LOCAL_REF = "local.shoc_consumer_role_arn"
def _cdk_sources():
return sorted(CDK_DIR.glob("*.py"))
def _terraform_sources():
paths = sorted(TF_DIR.glob("*.tf"))
if TFVARS_EXAMPLE.exists():
paths.append(TFVARS_EXAMPLE)
return paths
def test_only_the_pinned_foreign_principal_appears_in_cdk_sources():
def test_variable_default_is_the_pinned_foreign_principal():
"""The Terraform default must equal the exact trusted ARN.
Grant sites consume the variable via local.shoc_consumer_role_arn. Pinning
the default restores the exact-principal invariant the former CDK scan
enforced: widening trust requires editing this default (and this test).
"""
block = _VAR_DEFAULT_RE.search(VARIABLES_TF.read_text())
assert block, "variables.tf must declare variable shoc_consumer_role_arn"
default = _DEFAULT_VALUE_RE.search(block.group(1))
assert default, (
"variable shoc_consumer_role_arn must set default = "
f'"{ALLOWED_FOREIGN_PRINCIPAL}" so the exact principal is pinned in-repo'
)
assert default.group(1) == ALLOWED_FOREIGN_PRINCIPAL, (
f"shoc_consumer_role_arn default is {default.group(1)!r}, expected "
f"{ALLOWED_FOREIGN_PRINCIPAL!r}"
)
def test_tfvars_example_is_the_pinned_foreign_principal():
match = _TFVARS_VALUE_RE.search(TFVARS_EXAMPLE.read_text())
assert match, (
"terraform.tfvars.example must set shoc_consumer_role_arn to the pinned ARN"
)
assert match.group(1) == ALLOWED_FOREIGN_PRINCIPAL, (
f"terraform.tfvars.example shoc_consumer_role_arn is {match.group(1)!r}, "
f"expected {ALLOWED_FOREIGN_PRINCIPAL!r}"
)
def test_only_the_pinned_foreign_principal_appears_in_terraform_sources():
findings = []
for path in _cdk_sources():
for path in _terraform_sources():
for match in _IAM_ARN_RE.finditer(path.read_text()):
account = match.group(1)
if account in HOME_ACCOUNTS:
@ -46,12 +100,37 @@ def test_only_the_pinned_foreign_principal_appears_in_cdk_sources():
unexpected = [
(name, arn)
for name, arn in findings
if arn != ALLOWED_FOREIGN_PRINCIPAL or name not in ALLOWED_FILES
if arn != ALLOWED_FOREIGN_PRINCIPAL or name not in ALLOWED_LITERAL_FILES
]
assert not unexpected, (
"Unexpected foreign IAM principal(s) in cdk/ — every cross-account "
f"trust addition must update this pin deliberately: {unexpected}"
"Unexpected foreign IAM principal(s) under terraform/ — every "
"cross-account trust addition must update this pin deliberately: "
f"{unexpected}"
)
# Both grant sites must still reference the pinned role (deleting one
# half of the secret/KMS grant pair fails silently at the receiver).
assert {name for name, _ in findings} == ALLOWED_FILES
# Both default/example sites must still name the pinned role as a literal.
assert {name for name, _ in findings} == ALLOWED_LITERAL_FILES
def test_grant_files_use_local_shoc_consumer_role_arn_only():
"""Grant sites must consume the local — never a hardcoded foreign ARN.
api.tf (API resource policy) and wo_shoc.tf (KMS + secret policy) are the
two halves of the trust surface. Each must reference
local.shoc_consumer_role_arn so dropping one half fails CI, and neither
may embed a raw foreign IAM ARN (that would bypass the variable default pin).
"""
for filename in GRANT_FILES:
text = (TF_DIR / filename).read_text()
assert _LOCAL_REF in text, (
f"{filename} must reference {_LOCAL_REF} so the SHOC cross-account "
"grant surface cannot silently drop one half of the trust pair"
)
foreign = [
m.group(0)
for m in _IAM_ARN_RE.finditer(text)
if m.group(1) not in HOME_ACCOUNTS
]
assert not foreign, (
f"{filename} must not hardcode foreign IAM ARNs "
f"(use {_LOCAL_REF}): {foreign}"
)