mirror of
https://github.com/Sea-Haven-Industries/procurement-ingest.git
synced 2026-09-30 11:53:13 +00:00
20 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
06e96d8513
|
feat(webhook): activate SHOC work-order webhook emitter (#138)
Some checks failed
Deploy / deploy (push) Has been cancelled
* feat(webhook): ACTIVATE SHOC WO webhook emitter (enabled=True) The deliberate one-line activation flip (plan Phase 3). Turns on both DynamoDB stream event-source mappings for workorder-shoc-emitter, which shipped dark (enabled=False) in PR-2. LATEST start position => the live feed begins at deploy, no historical flood; SHOC loads history via the procurement read API first. DO NOT MERGE until: (1) PR-2 (feat/shoc-wo-webhook) is merged to main; (2) SHOC's receiver passes the shared HMAC test vectors (docs/shoc-webhook-test-vectors.json) at the target endpoint; (3) the cross-account secret read from shoc-backend-dev is confirmed working. Draft, gated on Luby. * docs(webhook): stamp activation date and correct backfill wording * docs(webhook): adjust activation comment date to 2026-07-30. Signed-off-by: Adam Moussa <adam@seahavenind.com> * docs(readme): mark SHOC webhook emitter active as of 2026-07-30 --------- Signed-off-by: Adam Moussa <adam@seahavenind.com> |
||
|
|
c040050373
|
feat(webhook): SHOC WO webhook emitter - dark-ship streams + HMAC secret/rotation (PR-2) (#137)
* docs(webhook): revise SHOC webhook contract and plan for post-migration reality Branch re-cut on main 2026-07-23 (old base carried stale PR #99 commits). Contract Rev 2026-07-23: - Producer account corrected: seahaven-prod (011934824531); mgmt frozen - Reconciliation backstop is the new procurement read API, not SyncController - wo_status "unknown" is real; SHOC must map it (checklist item added) - write_origin forward-compat note for phase-2 write-back echo suppression - SyncVendorReplies retirement flagged (dead table, no vendor_reply event) Plan updates: - Account gate: seahaven-prod only; never enable streams on mgmt tables - Emitter ships DARK (ESMs enabled=False); activation is a deliberate flip after the SHOC receiver passes shared HMAC vectors - Post-refactor conventions: common.py helpers, bundle-consistency AST pins, pytest.ini --cov additions, consolidated test roots - Dedicated-CMK rationale, secret-ARN handooff step, consumer audit refreshed (slack-bot decommissioned), enum golden test, write_origin skip-branch test * feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC call-and-be-called effort). Everything ships DARK: both DynamoDB event source mappings deploy enabled=False; activation is a deliberate one-line follow-up PR gated on the SHOC receiver passing the shared HMAC test vectors. - Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments (in-place update, RETAIN + logical IDs untouched; no existing consumers — verified live, neither table had a stream). - workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope -> HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard ordering (parallelization 1, bisect off, retry until 24h age, ReportBatchItemFailures); 429/5xx/timeout block the shard in order, other 4xx park to workorder-shoc-emitter-rejected; ESM failures -> workorder-shoc-emitter-failures (metadata; replay rebuilds from DynamoDB). Echo guard skips write_origin=shoc-write-api. - Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK (alias workorder-ingest-shoc-webhook-kms); cross-account GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy DESTROY deliberately (machine-generated material; avoids the fixed-name RETAIN-orphan deadlock). - workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap, 64-hex keys, kid = UTC %Y-%m-%dT%H. - Alarms (ALARM-only -> site-alerts): emitter errors/throttles/ duration + iterator-age (>=10 min) + failures/rejected queue depth; rotator standard trio. - scripts/replay_shoc_webhooks.py: dry-run-default operator replay (rebuilds from tables, replay:true envelopes). - Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay signing pinned to identical vectors); bundle-consistency AST pins for both new bundles. - README: WO stack + webhook feed section, alarm table, runbooks; removed stale seahaven-slack-bot consumer references. * fix(webhook): kms:ViaService pins, https-only delivery, cross-account principal CI pin GPT-4.1 cross-family review of the policy surface (no BLOCK): FIX applied to the cross-account shoc-backend-dev Decrypt statement and both Lambda role KMS grants (the key is only ever used via Secrets Manager); its invariant-enforcement QUESTION answered durably with tests/test_cross_account_principal_pin.py (any new foreign IAM principal in cdk/ fails CI). Scanner mediums fixed: delivery.py and the replay script now refuse non-https URLs (urllib follows file:// and http://). SQS metadata-action and dynamodb:ListStreams NITs skipped: standard CDK grant shapes; ListStreams has no resource-level scoping. The 4 gitleaks HIGHs on docs/shoc-webhook-test-vectors.json are deliberate non-secrets (shared receiver-verification vectors) suppressed machine-level with justification. * harden(webhook): resolve /sh-security-review findings (1 confirmed medium + cheap fixes) High-recall detector fan-out (injection/authz/secrets-crypto/iac-iam/logic) + proof-or-kill verifier. Gate PASSES: 1 confirmed medium, 0 confirmed critical/high. Confirmed finding fixed; several unverified-but-cheap hardenings applied since the emitter ships dark and activation is weeks out. - CONFIRMED medium (confused deputy): the rotation Lambda's generated invoke permission for secretsmanager.amazonaws.com carried no SourceAccount/SourceArn, so any account's Secrets Manager could invoke the rotator. Patched the generated CfnPermission in place (a second permission would be additive, not restrictive) to pin account + this secret ARN. - delivery + replay: refuse to follow receiver 3xx redirects (no-redirect opener) so live X-SH-* auth headers can't be forwarded to a receiver-chosen Location and an http:// Location can't slip past the https guard. Fixed the "unfollowed 3xx" comment that was factually wrong. - delivery: classify 401/403 as retryable (invalidate key cache + retry in order) instead of parking -- transient auth failures (rotation outran the TTL cache, clock skew) are availability events, not contract bugs. - envelope: build_event now genuinely total (guarded eventID / ApproximateCreationDateTime subscripts) per its own never-raise contract. - handler: catch-all so an unexpected per-record error (e.g. SQS park failure) reports only that record instead of failing the whole batch (which would re-deliver every earlier success for 24h); per-invocation emit/skip batch summary so a systemic silent drop is queryable/alarmable. - rotator: narrow the AWSCURRENT-read except to ResourceNotFound/JSONDecode (transient SM/KMS errors re-raise so the overlap key isn't silently dropped); kid uniqueness checked against ALL retained kids with a random suffix on collision (never reissue a kid for a different secret). - contract: skeleton-upsert required on ANY unknown work_order_id (not just comment-before-create) + monotonicity guard (ignore older updated_at), so a parked created or an out-of-order replay can't corrupt receiver state. Unverified/refuted findings left as-is with rationale: the two "high" logic claims (whole-batch crash triggers, ordering violation) were refuted on reachability (real stream records carry required fields; persistence writes strings only; full-state idempotent upsert absorbs the ordering gap). Signed kid/version binding (AUTHZ-002) declined: coordinated contract change, not cheap, no exploit with one algorithm/key. * fix(webhook): drop kid from rotator test_ok log (CodeQL clear-text-logging FP) GHAS CodeQL flagged py/clear-text-logging-sensitive-data (high) at _test_secret's success log because head["kid"] is subscripted from the same parsed-secret dict that holds head["secret"] — the taint tracker can't tell the non-secret key id from the secret. The secret value is never logged. Rather than dismiss the alert (fragile; re-alerts on line moves), remove the flow: kid is already logged at stage time in _create_secret and version_id correlates the steps, so the test_ok log keeps only event + version_id. Also hardens against a future edit that swaps the logged field. |
||
|
|
00d0d32337
|
fix(cdk): explicit Lambda LogGroups replace log_retention (INFRA-114) (#126)
Some checks are pending
Deploy / deploy (push) Waiting to run
First seahaven-prod deploy failed CREATE on the sender-auth MetricFilter: it imported /aws/lambda/<fn> by name, which pre-existed in mgmt but not in a fresh account. All 5 functions now get an explicit logs.LogGroup (TWO_MONTHS, RETAIN) via common.make_function_log_group, and the metric filter takes the construct so CFN orders it after the group exists. Also removes the deprecated LogRetention custom resource and its wildcard logs:PutRetentionPolicy role (CKV_AWS_111). Function roles keep AWSLambdaBasicExecutionRole (verified in the synthesized template), so log-write permissions are unchanged; cross-review's grant_write FIX was a false positive on that basis. Supersedes PR #84, which hardcoded the mgmt logs-CMK ARN and predates the common.py refactor. mgmt collision note: these CREATEs would collide with the pre-existing groups in mgmt; acceptable because the deploy secret now targets prod and mgmt is frozen pending decommission. |
||
|
|
f8eb18f02b
|
feat: collapse duplicated CDK into cdk/common.py plain helpers (refactor phase 4) (#112)
Some checks are pending
Deploy / deploy (push) Waiting to run
The ~379 lines po_stack.py and wo_stack.py defined identically (DynamoDB alarms, the sender-auth-rejected metric filter + alarm, the standard per-Lambda alarm set, the Bedrock InvokeModel grant, the raw-email bucket, the async DLQ, the template-fallback-rate math alarm) move into cdk/common.py. Every helper is a PLAIN function taking (scope, id, ...), called with each stack's own Stack as scope and the exact literal construct ids used inline before, so every synthesized logical ID is byte-stable. A Construct subclass would reparent the tree and make CloudFormation attempt to replace the RETAIN-protected purchase-orders/WorkOrders tables and named buckets -- data loss -- so it is forbidden. Per-function alarm variance (PO p99 vs WO p95 duration, po-web-ui throttles+duration only, site-extractor no DLQ alarm, workorder-web-ui zero alarms) is preserved through call-site arguments, not baked into the helpers. make_bedrock_invoke_statement derives the inference-profile and us-east-1 foundation-model ARNs from Stack.of(scope).account/.region instead of the hardcoded 328440206208/us-east-1 literals. The environment stays account-agnostic (region-only), so the account resolves to the AWS::AccountId pseudo-parameter: the derived ARN resolves at deploy to the same ARN the literal named in-account (a benign in-place IAM policy update, never a replacement) and is account-portable rather than pinned to the frozen management account. The account= pin evaluated for cdk.Environment was deliberately NOT added: resolving every account-derived value (bucket names, Lambda::Permission source account, SNS action ARN) to literals makes CloudFormation flag the RETAIN email buckets as requiring replacement against the deployed account-agnostic templates -- a data-loss risk that outranks the pin, which buys nothing (the resolved values are unchanged). Also: net-new CfnOutputs for the five Lambda function ARNs and the owned/consumed table names, exact-pin constructs==10.6.0, and fix the stale aws-cdk-lib 2.259.0 -> 2.261.0 version comment. The common.py extraction is zero-cdk-diff on both stacks (byte-stable logical IDs, no asset/property change); the only deltas versus deployed are the intended benign Bedrock IAM in-place update and the additive CfnOutputs. Mandatory GPT-4.1 cross-family review ran on the Bedrock IAM move; its BLOCK was a verified false positive (it read AWS::AccountId as a wildcard -- it is a deploy-time-resolved concrete value naming one account and one inference-profile, region is pinned us-east-1, and the grant is strictly more least-privilege-correct than the hardcoded literal). |
||
|
|
30112cc680
|
feat: extract lambdas/shared/ — single-source ses_auth, web_ui auth, email parsing, EMF emitter (refactor phase 3) (#111)
Some checks are pending
Deploy / deploy (push) Waiting to run
Four modules move into the handbook-mandated lambdas/shared/ location, collapsing duplicated logic that had to be kept in sync by hand across the PO and WO pipelines: - ses_auth.py: the PO and WO copies were verified sha256-identical against the feature/phase-7-ops-recovery baseline before the move (no drift since the last audit). shared/ses_auth.py is the exact bytes of that one copy; both originals are git rm'd (the PO copy via rename, the WO copy as a straight delete). Bundling lands the module flat in /asset-output for both email processors, so the handlers keep `from ses_auth import authenticate_inbound_email` unchanged — zero handler diff for this move, which is what keeps fail-closed auth byte-identical through the change. - web_ui_auth.py: extracts the byte-identical _get_auth_token / _header / is_authenticated block plus the four token-cache globals out of both web_ui handlers. The per-stack INFRA-74 comments stay in each handler as-is (deliberately drifted wording, stack-specific) rather than being unified into the shared module. Fail-closed semantics (unset ARN or Secrets Manager exception -> deny) are unchanged. - email_parsing.py: parse_raw_email ships as the superset version that returns cc unconditionally. WO's output is bit-identical to before; PO simply ignores the cc field rather than being "cleaned up" to consume it. No second variant is kept. - emf.py: a generic emitter parameterized by namespace, dimension sets, and properties. Every call site's emitted EMF envelope is unchanged, including the load-bearing [["ParseMethod"],["ParseMethod","TemplateId"]] dimension-set shape the alarms and metric filters depend on. Emission ordering is untouched: PO still emits ai_fallback before the Bedrock call, WO still emits its mutually-exclusive ai_fallback/ai_fallback_rejected after its gate. The deliberate-double-count comments survive. _emit_derived_agreement_metric was found living inside derived_fields.py, so per the DERIVED-FIELDS exception it is left as a third, unconverted copy (derived_fields.py and the shadow DerivedFieldAgreement telemetry stay untouchable while that bake runs) — a comment there points at shared/emf.py for the eventual follow-up. Bundling: both email-processor cdk bundling commands gain a trailing `cp shared/*.py /asset-output/` (they were already cp-only post-Phase 7, so no pip step or manylinux pin is reintroduced). Both web_ui functions gain the same widened-root staging so web_ui_auth.py ships beside their handler; site_extractor's from_asset is untouched. Tests: PO_EXPECTED_TOP_LEVEL_MODULES gains the shared modules that now ship, the AST sibling-import check resolves imports whose source now lives under shared/, and the new shared cp line has its own revert/mutation detection. _SIBLING_MODULES resolution and _po_parser_support.py now load ses_auth/email_parsing/emf from shared/; the two-copy ses_auth byte-identity fixture-hygiene test is retired as obsolete now that there is one copy, and the ses_auth fixture parameterization over two identical copies is dropped. The sys.modules save/restore dance for template_parser (still duplicated per-pipeline) is left in place. |
||
|
|
75fe91c198
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7) Generalize scripts/reprocess.py from a PO-only full-sweep script into a pipeline-general recovery tool. Targeted replay (--key/--prefix/--since) is now the default, and the full inbound/ sweep is demoted behind an explicit --all that documents its five hazards (async concurrency does not serialize, use RequestResponse if order matters, metric double-count, Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves the correct function + bucket; dry-run-by-default / --execute is preserved. A new tests/test_reprocess_contract.py pins the synthetic S3 event shape and asserts the raw list_objects_v2 key is emitted untransformed (the handler is the single decode point; a pre-decoded key would corrupt keys containing spaces or '+'). Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console redrive-to-source, so it documents the receive -> extract key -> targeted reprocess --key -> verify -> purge procedure, the real recovery windows (14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table RETAIN policy and is the true replay floor), and that sender-auth and ai_fallback_rejected drops are fail-closed skips that never reach the DLQ. Linked from the README alarms and scripts sections. Drop the vendored boto3 floor pin from both email-processor requirements (the Lambda runtime provides boto3; lambda-template.md empty-with-comment form). With nothing left to install, the email-processor bundling becomes cp-only -- the whole pip step is removed, which is the only acceptable way the manylinux2014_aarch64 pin disappears (removing the pin while keeping a pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and add pinned po/web_ui + po/site_extractor manifests (excluded from their bundles, so hash-neutral) so their new Dependabot entries have something to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and /lambdas/po/site_extractor. cdk diff is confined to exactly the two email processors' asset hashes on both stacks. The wo/web_ui dead-manifest reduction was deliberately left out: that manifest already ships inside the plain (non-bundled) WebUI asset on main, so reducing or excluding it would redeploy workorder-web-ui for no functional change -- deferred to keep the blast radius to the two intended targets. The untracked 44 MB lambdas/po/email_processor/package/ dir was removed from the filesystem (asset-hash-neutral given Phase 2's package/ exclude); it is untracked, so there is nothing to commit for it. * Reject --all combined with --prefix/--since in reprocess.py --all is a distinct mode (the demoted full-prefix sweep), but the args.all branch unconditionally set prefix=inbound/ and since=None, so passing it alongside a narrower selector silently discarded that selector. `--all --since 2026-07-01` swept the entire corpus instead of the bounded window, triggering every documented --all hazard (Bedrock re-bill, metric double- count, merged-field regression) on objects the operator never targeted -- contradicting the tool's safety goal. Add the missing mutual-exclusion guard alongside the existing --key one, and pin --all+--prefix, --all+--since, and all three together as argparse rejections. |
||
|
|
df03f3497f
|
feat: widen email-processor asset roots to lambdas/ with scoped globs + excludes (refactor phase 2) (#109)
Some checks failed
Deploy / deploy (push) Has been cancelled
Both email-processor Code.from_asset calls now bundle from lambdas/ instead of their per-function subdirectory, so Phase 3's shared/ module is reachable from the asset root once it lands. The bundling commands were rewritten for the new cwd (pip install -r <po|wo>/ email_processor/requirements.txt -t /asset-output && cp <po|wo>/ email_processor/*.py /asset-output/), preserving the ARM64 --platform manylinux2014_aarch64 --only-binary=:all: pin exactly — its removal shipped x86 wheels into the ARM64 function and caused a 100% outage (PR #34). All five from_asset calls (both email processors, po web_ui, po site_extractor, wo web_ui) now exclude **/__pycache__/**; the two widened ones also exclude **/tests/** and **/package/**. Without the package/ exclude, the stale untracked 44 MB lambdas/po/email_processor/package/ dir (local-only, never present in CI) would diverge local vs CI asset hashes and force spurious redeploys — from_asset doesn't honor .gitignore. That dir is left in place; deleting it is Adam's call. WO's prod zip shrinks as deliberate cleanup, not a byte-identical match to PO: the old `cp -r .` shipped tests/ (real scrubbed .eml fixtures), __pycache__/, and requirements.txt into production. The acceptance bar for WO is runtime-imported module set unchanged + smoke, not a byte-identical zip; PO keeps the byte-identical first-party file set guarantee. tests/test_bundle_consistency.py is updated in the same change to recognize the scoped `cp po/email_processor/*.py` (resp. wo) glob as the new unconditionally-safe shape, without loosening the allowlist-revert detection, the detection-logic mutation test, or the PO_EXPECTED_TOP_LEVEL_MODULES exact-set pin. No code moved under lambdas/ in this change (git diff main...HEAD -- lambdas/ is empty); only CDK asset wiring and its tests changed. |
||
|
|
e86ea970a8
|
fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path (#104)
Some checks are pending
Deploy / deploy (push) Waiting to run
* fix: add fail-closed validation gate and XML-delimited prompt on ai_fallback path The ai_fallback parse path applied no validation gate to raw Bedrock/LLM output before DynamoDB writes, and the extraction prompt concatenated the untrusted email body directly with no instructions-vs-data delimiter. A DKIM-passing attacker could prompt-inject arbitrary field values into the work-order store. Changes: - wrap untrusted email in \<email\> XML block with prompt instructing the model to treat its contents as data only - add validate_ai_fallback() in template_parser that enforces the same contract keys, enums, and patterns as the template path before any write - call validate_ai_fallback() in handler() dispatch; emit an ai_fallback_rejected EMF metric on failure and skip the record - add 17 unit tests covering every gate rule and two end-to-end dispatch tests (injected email_type, injected status) Refs #101 * style: apply ruff formatting to fix CI check * harden ai_fallback gate: review fixes + security-review findings Review follow-up on the ai_fallback validation gate (PR #104), plus findings from a fan-out /sh-security-review of the change surface. Reviewer FIX items: - Neutralize forged <email> delimiters in the untrusted body before wrapping, so an in-body </email> cannot escape the data block. - Fail closed on non-dict model output instead of crashing the handler into async retries; count ai_fallback_rejected parses in the fallback-rate alarm and add a dedicated rejected-parse alarm so a gate-rejection drift outage is not silent. - Return a distinct invalid_status reason (was malformed_site_code); validate ISO-8601 dates; README + docstring updates. Security-review findings (detector fan-out + proof-or-kill verifier): - ReDoS (confirmed, medium): the tag neutralizer used two \s* around an optional /, backtracking quadratically on "<" + a long whitespace run (~32s at 100k chars -- one email could time out the Lambda). Collapse to a single [\s/]* class: linear, same defanging. - Unhashable-type crash (confirmed): a JSON list/dict for email_type or status made `x in <set>` raise TypeError, escaping the gate into retries. Guard with isinstance(str) before membership. - Unicode/newline regex (confirmed): _WO_ID_RE/_SITE_CODE_RE used ^..$ with \d, admitting fullwidth digits ("12345" as a lookalike partition key) and trailing newlines. Switch to \A[0-9]+\Z (and the handler's inline recheck to [0-9]) so neither passes. - Alarm comment (confirmed, low): corrected the "slow trickle still pages" wording -- rejections >~25-30 min apart page on neither alarm, the same knowingly-accepted residual as sender-auth-rejected. Refuted: residual free-text prompt injection is inherent to trusting allowlisted senders, not a new primitive; no DynamoDB key-poisoning bypass survives both gates ('#' can never enter work_order_id). 7 new regression tests. All 260 tests pass; ruff clean; cdk synth OK. --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> Co-authored-by: Adam Moussa <adam@seahavenind.com> |
||
|
|
11bf0f5d12
|
fix(ses_auth): harden comment stripping and alarm evaluation window (#103)
Some checks are pending
Deploy / deploy (push) Waiting to run
SES-AR-01: treat ")" at depth 0 as an unmatched close, rejecting the value as not well-formed so a ")(...)" pair cannot manufacture a depth-0 gap where a smuggled dkim=pass clause gets parsed. SES-AR-02: emit a space when a comment is removed so comments act as CFWS folding whitespace (RFC 5322). Without this, "dk(z)im=pass" would become "dkim=pass" and an attacker comment could glue unrelated tokens. Alarm: widen evaluation_periods from 3→6 (30-min window) with datapoints_to_alarm still at 2, closing the sparse-outage residual where rejections >10-15 min apart fail to place breaching datapoints in 2 of 3 consecutive periods. Both ses_auth.py copies stay byte-identical. All 86 tests pass. Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> |
||
|
|
8e16d34dc7
|
Fix WO fallback-rate alarm math expression (deploy hotfix) (#102)
Some checks are pending
Deploy / deploy (push) Waiting to run
The EmailProcessorTemplateFallbackRateAlarm added in #99 used MAX([FILL(fb,0)+FILL(tmpl,0),1]) as a divide-by-zero guard, but CloudWatch metric math has no element-wise MAX over an array; it rejects the array operand at deploy time with 'Unsupported operand type(s) for MAX', failing the WorkorderIngestStack update (it rolled back cleanly). The outer IF((...)>=10, ...) volume floor already guarantees a non-zero denominator in the true branch, so divide by (FILL(fb,0)+FILL(tmpl,0)) directly and drop the invalid guard. Unblocks the #99 deploy (PO stack already applied; WO rolled back). |
||
|
|
acc1961d21
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO. |
||
|
|
7b9e26d79d
|
Add fail-closed SES sender authentication (INFRA-107) (#98)
Some checks are pending
Deploy / deploy (push) Waiting to run
* Add fail-closed SES sender authentication
The From header and any raw-MIME Authentication-Results copies are
attacker-forgeable, so a forged email to apm@int.seahaven.com or
amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107,
CRITICAL). Both S3-triggered email processors now authenticate the
sender against the Authentication-Results header SES itself prepends
at delivery: only the topmost header is consulted, its authserv-id
must be amazonses.com, and it must carry dkim=pass for a domain in
the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set
in CDK so ops can adjust without code changes).
Allowlists come from live traffic observed 2026-07-15 on both ingest
buckets: WO mail arrives via the apm@ Google Groups forward, which
re-signs as seahaven.com (the hxgnsmartcloud.com signature does not
survive the forward); PO mail passes for amazon.coupahost.com.
amazonses.com also passes on PO mail but is deliberately excluded --
every SES customer's outbound mail passes for it.
Every failure path (env var unset, header missing or unparseable,
verdict fail, unaligned domain) rejects the email: a structured
warning with the reason and S3 key is logged and the record skipped
without erroring the invocation, so rejected mail causes no Lambda
retries or DLQ messages. Handler signatures and event sources are
unchanged.
Refs: INFRA-107
* Harden AR parser per cross-family review
Cross-family (GPT-4.1) review findings: terminate the dkim result
token at end-of-clause, whitespace, or a comment so a value like
"dkim=pass-fake" can never be read as a pass; normalize trailing
dots off allowlist entries so "seahaven.com." matches; make the
compat32 parser policy explicit. Adds tests for result-token
boundaries, comments after the result, quoted domain values, and
folding inside a dkim clause.
Refs: INFRA-107
* Harden AR parsing and alarm on sender-auth rejects
The SES-stamped Authentication-Results value echoes attacker-controlled
SMTP-session tokens (envelope-from, helo, header.from) as their own
semicolon-delimited property clauses. A naive split(";") tore an RFC 5321
quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass
clause, so a fully spoofed email was accepted on the genuinely
SES-stamped topmost header. Tokenise comment- and quoted-string-aware
(RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on
semicolons outside a quoted-string, and fail closed on unbalanced
quotes/comments so a ';' inside a quoted pvalue can never start a clause.
Rejected mail returns normally (no error, no retry, no DLQ message), so a
signing-domain drift or a wrong allowlist would silently discard 100% of
legitimate mail while every alarm stayed green. Add a CloudWatch Logs
metric filter + alarm on the sender_auth_rejected warning to both stacks
so a false-reject storm pages instead of vanishing. This is also the
safety net for the WO seahaven.com allowlist assumption, which must be
validated against a live SES-stamped header (a plain Gmail auto-forward
re-signs under the sending Workspace domain, not seahaven.com).
Refs: INFRA-107
* chore: retrigger CI (no run recorded for
|
||
|
|
97a5886139
|
Land safe fixes from 2026-06-17 security sweep (#97)
Some checks are pending
Deploy / deploy (push) Waiting to run
* Remove gratuitous KMS grant on shared DynamoDB CMK wo-email-processor held grant_encrypt_decrypt on the shared seahaven-dynamodb CMK, but the WorkOrders/WorkOrderComments tables are not encrypted with that CMK. The grant was dead weight that extended the WO processor's decrypt reach to the CMK protecting the purchase-orders table (cross-stack decrypt). Drop it to restore least privilege; re-add as part of the table CMK migration (INFRA-6). Refs: INFRA-6 * Require Secrets Manager key for Anthropic client Remove the silent fallback to a plaintext ANTHROPIC_API_KEY env var in both email processors; require ANTHROPIC_API_KEY_SECRET_ARN and raise if absent so a misconfigured deploy fails loudly instead of using an unmanaged key. Adapted from |
||
|
|
35dc32390c
|
Add DLQ messages-present alarm for workorder-email-processor (#68)
Some checks failed
Deploy / deploy (push) Has been cancelled
The existing workorder-email-processor-errors alarm fires on any errored async invocation, but a message only reaches the DLQ after Lambda exhausts its async retries and gives up — a genuinely dropped work-order email that the errors alarm alone does not distinguish. Add an ALARM-only CloudWatch alarm (workorder-email-processor-dlq-messages) on the EmailProcessorDlq ApproximateNumberOfMessagesVisible metric (Statistic MAXIMUM, period 5m, evaluationPeriods 1, threshold > 0, treatMissingData NOT_BREACHING). Routes to the same shared site-alerts SNS topic via SnsAction, mirroring the errors-alarm construct style. Refs INFRA-41 / audit H-8. |
||
|
|
86f2ebb42d
|
Codify S3 Block Public Access on ingest buckets (#71)
Add block_public_access=BlockPublicAccess.BLOCK_ALL to the po-ingest and workorder-ingest EmailBucket constructs. The buckets are already private at runtime via account-level and AWS-default BPA, so this is a no-op for behavior; it closes the codification gap that left CKV_AWS_53-56 firing on the synthesized templates and blocking the security pre-push gate (and the DLQ-alarm PRs that ride on it). |
||
|
|
0fdf407e1d
|
Add CloudWatch alarm coverage for po-ingest and workorder-ingest (#70)
Some checks are pending
Deploy / deploy (push) Waiting to run
* Add CloudWatch alarm coverage for po-ingest and workorder-ingest Expands alarm coverage across both CDK stacks. All alarms are ALARM-only (no OK action) to the shared site-alerts SNS topic, with TreatMissingData NOT_BREACHING. The site-alerts topic is now imported once near the top of each stack so every alarm reuses one Topic instance. po-ingest (cdk/po_stack.py): - Errors: po-ingest-site-extractor - Throttles: po-email-processor, po-ingest-site-extractor, po-web-ui - Duration (p99, >=45000ms, eval3/dp2): po-email-processor (orphan adoption), po-ingest-site-extractor, po-web-ui - DynamoDB throttle + system-error: purchase-orders, verified-sites, pending-site-review workorder-ingest (cdk/wo_stack.py): - Throttles: workorder-email-processor - Duration (p95, >=45000ms, eval3/dp2): workorder-email-processor (orphan adoption) - DynamoDB throttle + system-error: WorkOrders, WorkOrderComments DynamoDB ThrottledRequests/SystemErrors emit only at the TableName+Operation dimension set, so each table alarm is a Sum math expression across operations via the non-deprecated metric_*_for_operations helpers (metric_throttled_requests is deprecated/invalid in aws-cdk-lib 2.259.0). Refs INFRA-41 / audit H-8. * Drop NEEDS ADAM SIGN-OFF wording from alarm comments Duration alarm thresholds are owner-approved; remove the sign-off flag from po_stack.py and wo_stack.py comments. Threshold values, eval config, and orphan-delete notes are unchanged. |
||
|
|
47fa35688d
|
fix(po): grant KMS on seahaven-dynamodb CMK to purchase-orders consumers (INFRA-104) (#56)
Some checks are pending
Deploy / deploy (push) Waiting to run
The purchase-orders table was migrated to SSE-KMS (alias/seahaven-dynamodb, INFRA-95/M-3) out-of-band, but po_stack never declared the key, so grant_read_write_data did not propagate kms perms. po-email-processor failed ~99.6% of invocations with kms:Decrypt AccessDeniedException, a data-loss outage on the PO ingestion write path. - po_stack: declare encryption_key on purchase-orders (reconciles SSE drift; no-op against the already-encrypted live table) so the existing grants add kms:Decrypt/GenerateDataKey/DescribeKey to EmailProcessor, WebUI, SiteExtractor. - wo_stack: pre-emptive grant_encrypt_decrypt on the WO processor role ahead of the WorkOrders CMK migration (INFRA-6); tables left unencrypted, no table change. GPT-4.1 cross-review: no blockers. |
||
|
|
109c565cbf
|
Reconcile IaC with out-of-band DLQ + Function URL changes (INFRA-74, INFRA-41) (#50)
Some checks are pending
Deploy / deploy (push) Waiting to run
Make CDK the source of truth for two sets of changes applied out-of-band via CLI to the po-ingest and WorkorderIngestStack stacks. INFRA-74 (audit C-5): remove the public FunctionUrlAuthType.NONE Function URL construct (and its auto-generated Principal:* invoke permission + output) from both po-web-ui and workorder-web-ui. The URLs were already deleted live via CLI; CFN's delete is idempotent. INFRA-41 (audit H-8): add a CDK-managed SQS dead-letter queue (dead_letter_queue=, 14d retention, SSL-enforced, CDK-generated name) and an ALARM-only Errors alarm (Sum, threshold>0, site-alerts topic) for both po-email-processor and workorder-email-processor, mirroring the apm-wo-analysis-classifier DLQ and payments-payroll-batch alarm patterns. Interim CLI resources (per-fn -dlq queues, -errors alarms, dlq-send inline policies, OnFailure event-invoke-configs) removed post-deploy. |
||
|
|
9370abf9ab
|
Drop three read-idle GSIs (audit M-20) (#35)
Some checks are pending
Deploy / deploy (push) Waiting to run
* Drop read-idle GSIs: by-state and status-index Audit M-20: 30 days of CloudWatch metrics show 0 reads on both indexes against 518 (by-state) and ~50k (status-index) WCU of write amplification. No code path queries either index. site-code-index follows in the next commit - CloudFormation allows only one GSI change per table per deploy. * Drop read-idle site-code-index GSI Second half of the M-20 cleanup - deployed separately because CloudFormation allows one GSI change per table per update. |
||
|
|
5112c1345b
|
Merge workorder-ingest into unified procurement repo (#22)
* Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers |