procurement-ingest/docs/shoc-webhook-plan.md
Adam Moussa c040050373
feat(webhook): SHOC WO webhook emitter - dark-ship streams + HMAC secret/rotation (PR-2) (#137)
* docs(webhook): revise SHOC webhook contract and plan for post-migration reality

Branch re-cut on main 2026-07-23 (old base carried stale PR #99 commits).

Contract Rev 2026-07-23:
- Producer account corrected: seahaven-prod (011934824531); mgmt frozen
- Reconciliation backstop is the new procurement read API, not SyncController
- wo_status "unknown" is real; SHOC must map it (checklist item added)
- write_origin forward-compat note for phase-2 write-back echo suppression
- SyncVendorReplies retirement flagged (dead table, no vendor_reply event)

Plan updates:
- Account gate: seahaven-prod only; never enable streams on mgmt tables
- Emitter ships DARK (ESMs enabled=False); activation is a deliberate flip
  after the SHOC receiver passes shared HMAC vectors
- Post-refactor conventions: common.py helpers, bundle-consistency AST pins,
  pytest.ini --cov additions, consolidated test roots
- Dedicated-CMK rationale, secret-ARN handooff step, consumer audit refreshed
  (slack-bot decommissioned), enum golden test, write_origin skip-branch test

* feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation

Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC
call-and-be-called effort). Everything ships DARK: both DynamoDB event
source mappings deploy enabled=False; activation is a deliberate
one-line follow-up PR gated on the SHOC receiver passing the shared
HMAC test vectors.

- Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments
  (in-place update, RETAIN + logical IDs untouched; no existing
  consumers — verified live, neither table had a stream).
- workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope ->
  HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard
  ordering (parallelization 1, bisect off, retry until 24h age,
  ReportBatchItemFailures); 429/5xx/timeout block the shard in order,
  other 4xx park to workorder-shoc-emitter-rejected; ESM failures ->
  workorder-shoc-emitter-failures (metadata; replay rebuilds from
  DynamoDB). Echo guard skips write_origin=shoc-write-api.
- Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK
  (alias workorder-ingest-shoc-webhook-kms); cross-account
  GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly
  arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy
  DESTROY deliberately (machine-generated material; avoids the
  fixed-name RETAIN-orphan deadlock).
- workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap,
  64-hex keys, kid = UTC %Y-%m-%dT%H.
- Alarms (ALARM-only -> site-alerts): emitter errors/throttles/
  duration + iterator-age (>=10 min) + failures/rejected queue
  depth; rotator standard trio.
- scripts/replay_shoc_webhooks.py: dry-run-default operator replay
  (rebuilds from tables, replay:true envelopes).
- Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared
  with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay
  signing pinned to identical vectors); bundle-consistency AST pins
  for both new bundles.
- README: WO stack + webhook feed section, alarm table, runbooks;
  removed stale seahaven-slack-bot consumer references.

* fix(webhook): kms:ViaService pins, https-only delivery, cross-account principal CI pin

GPT-4.1 cross-family review of the policy surface (no BLOCK): FIX applied
to the cross-account shoc-backend-dev Decrypt statement and both Lambda
role KMS grants (the key is only ever used via Secrets Manager); its
invariant-enforcement QUESTION answered durably with
tests/test_cross_account_principal_pin.py (any new foreign IAM principal
in cdk/ fails CI). Scanner mediums fixed: delivery.py and the replay
script now refuse non-https URLs (urllib follows file:// and http://).
SQS metadata-action and dynamodb:ListStreams NITs skipped: standard CDK
grant shapes; ListStreams has no resource-level scoping. The 4 gitleaks
HIGHs on docs/shoc-webhook-test-vectors.json are deliberate non-secrets
(shared receiver-verification vectors) suppressed machine-level with
justification.

* harden(webhook): resolve /sh-security-review findings (1 confirmed medium + cheap fixes)

High-recall detector fan-out (injection/authz/secrets-crypto/iac-iam/logic)
+ proof-or-kill verifier. Gate PASSES: 1 confirmed medium, 0 confirmed
critical/high. Confirmed finding fixed; several unverified-but-cheap
hardenings applied since the emitter ships dark and activation is weeks out.

- CONFIRMED medium (confused deputy): the rotation Lambda's generated
  invoke permission for secretsmanager.amazonaws.com carried no
  SourceAccount/SourceArn, so any account's Secrets Manager could invoke
  the rotator. Patched the generated CfnPermission in place (a second
  permission would be additive, not restrictive) to pin account + this
  secret ARN.
- delivery + replay: refuse to follow receiver 3xx redirects (no-redirect
  opener) so live X-SH-* auth headers can't be forwarded to a
  receiver-chosen Location and an http:// Location can't slip past the
  https guard. Fixed the "unfollowed 3xx" comment that was factually wrong.
- delivery: classify 401/403 as retryable (invalidate key cache + retry in
  order) instead of parking -- transient auth failures (rotation outran the
  TTL cache, clock skew) are availability events, not contract bugs.
- envelope: build_event now genuinely total (guarded eventID /
  ApproximateCreationDateTime subscripts) per its own never-raise contract.
- handler: catch-all so an unexpected per-record error (e.g. SQS park
  failure) reports only that record instead of failing the whole batch
  (which would re-deliver every earlier success for 24h); per-invocation
  emit/skip batch summary so a systemic silent drop is queryable/alarmable.
- rotator: narrow the AWSCURRENT-read except to ResourceNotFound/JSONDecode
  (transient SM/KMS errors re-raise so the overlap key isn't silently
  dropped); kid uniqueness checked against ALL retained kids with a random
  suffix on collision (never reissue a kid for a different secret).
- contract: skeleton-upsert required on ANY unknown work_order_id (not just
  comment-before-create) + monotonicity guard (ignore older updated_at), so
  a parked created or an out-of-order replay can't corrupt receiver state.

Unverified/refuted findings left as-is with rationale: the two "high" logic
claims (whole-batch crash triggers, ordering violation) were refuted on
reachability (real stream records carry required fields; persistence writes
strings only; full-state idempotent upsert absorbs the ordering gap). Signed
kid/version binding (AUTHZ-002) declined: coordinated contract change, not
cheap, no exploit with one algorithm/key.

* fix(webhook): drop kid from rotator test_ok log (CodeQL clear-text-logging FP)

GHAS CodeQL flagged py/clear-text-logging-sensitive-data (high) at
_test_secret's success log because head["kid"] is subscripted from the
same parsed-secret dict that holds head["secret"] — the taint tracker
can't tell the non-secret key id from the secret. The secret value is
never logged. Rather than dismiss the alert (fragile; re-alerts on line
moves), remove the flow: kid is already logged at stage time in
_create_secret and version_id correlates the steps, so the test_ok log
keeps only event + version_id. Also hardens against a future edit that
swaps the logged field.
2026-07-24 22:12:20 +00:00

17 KiB

Implementation Plan — SHOC Realtime Work-Order Webhook

Branch: feat/shoc-wo-webhook (re-cut on main 2026-07-23; all build work targets the post-refactor layout — cdk/common.py helpers, explicit LogGroups, decomposed handlers, consolidated test roots — not the pre-refactor style earlier drafts implied). Companion docs: docs/shoc-webhook-contract.md (Rev 2026-07-23 — the producer/consumer contract; Luby builds the receiver against it) and the procurement read API OpenAPI spec (lambdas/api/openapi.json, built in the sibling procurement-api effort — the reconciliation/backfill path). Requestor: Luby (SHOC team). Decisions captured 2026-07-16: all WO event types, seconds-latency, full-payload push, HMAC + automated rotation, in-order per WO, prod-data-in-dev accepted, dual-write retained (DynamoDB retirement deferred for scope discipline — the original blocker, seahaven-slack-bot, was decommissioned 2026-07-23; retirement is now sequenced after SHOC prod + mgmt decommission, not this PR). Account (2026-07-23): everything here builds in seahaven-prod (011934824531). The old management account (328440206208) is frozen/rollback-only.

Architecture

SES → S3 → workorder-email-processor ──writes──▶ WorkOrders ─stream─▶
                                     └─writes──▶ WorkOrderComments ─stream─▶ workorder-shoc-emitter ──HMAC POST──▶ SHOC API
                                                                                    │ (retry-exhausted / non-retryable)
                                                                                    ▼
                                                                        shoc-emitter failure queues → alarms → replay script

Rationale (vs. inline call from the processor): DynamoDB Streams gives per-partition-key (per-work_order_id) commit-order delivery with shard-blocking retries — the ordering guarantee Luby requires — and keeps the email-ingestion hot path completely decoupled from SHOC availability. No change to PR #99's handler.

Phase 0 — Prerequisites (blocking)

  1. Branch re-cut on main (done 2026-07-23).
  2. Luby signs off docs/shoc-webhook-contract.md Rev 2026-07-23 (endpoint path, PR #10 entity mapping, the unknown-status mapping, SyncVendorReplies retirement).
  3. Confirm the SHOC receiver role ARN for the resource policies (arn:aws:iam::396287094661:role/shoc-backend-dev) — the role lives in SHOC's account, outside this repo's control; confirm it still exists before the CDK references it (the first cross-account get-secret-value test is the hard proof).
  4. ⚠️ ACCOUNT GATE: every deploy in this plan targets seahaven-prod (011934824531) ONLY. Never deploy this stack to mgmt (328440206208), and never "temporarily" enable streams on mgmt's copies of the WO tables — that is the only path by which the SES dual-delivery bake could ever double-send to SHOC. Verify aws sts get-caller-identity resolves to 011934824531 before cdk deploy.
  5. Confirm with Luby which SHOC environment should receive the feed now — "prod-data-in-dev" was accepted 2026-07-16, pre-migration; re-confirm the target URL.

Phase 1 — Secret + rotation (CDK, wo_stack.py)

This phase is an incremental deploy onto the already-migrated workorder-ingest stack in seahaven-prod — it is not part of the account-migration work.

  • New customer-managed KMS key workorder-ingest-shoc-webhook-kms (kebab-case per naming-conventions.md). Required because the default aws/secretsmanager key cannot serve cross-account reads. Deliberately a dedicated key, not alias/seahaven-dynamodb: reusing the DynamoDB CMK would hand the SHOC cross-account grant decrypt reach over the PO table's encryption key — the dedicated key scopes the grant to exactly this secret.
  • New Secrets Manager secret workorder-ingest/shoc-webhook-hmac (naming per secrets-and-config.md stack-name/value-name), CDK-managed, encrypted with the CMK. Value shape: {"keys": [{"kid", "secret"}, ...]} (dual-key, contract §6.1).
  • RemovalPolicy: DESTROY, deliberately. The value is machine-generated HMAC material with no operator-set content — fully regenerable by one rotation — so RETAIN buys nothing and would expose the fixed-name RETAIN orphan deadlock (failed first create orphans an empty shell holding the global name → every later create fails AlreadyExists; see reference_secret_retain_orphan_deadlock). Add a GOTCHA comment at the removal-policy line so nobody "hardens" it to RETAIN later. If a first deploy does fail, check for and force-delete any empty orphaned shell before retrying. Accidental-deletion recovery: recreate + rotate (new keys), notify Luby (receivers refresh within their 5-min TTL), watch the -failures queue and replay the gap — documented in the README runbook. Secret gets a description and Purpose/ManagedBy tags for discoverability.
  • Rotation Lambda workorder-shoc-hmac-rotator (Python 3.12, ARM64, explicit LogGroup via common.make_function_log_group, alarms via the house common.add_standard_lambda_alarms idiom) on a 30-day rotation_schedule: generates a new 32-byte key, prepends as keys[0], truncates to 2 entries. Single-user rotation (no external system holds the value — receivers re-fetch), so the standard 4-step rotation collapses to createSecret/finishSecret. Rotation-outage math: producer signs keys[0] with a ≤5-min cache; receiver keeps the previous key and refreshes on unknown kid (contract §6.1, receiver-side) — no delivery window where signatures can't verify.
  • Cross-account read grants — both halves required, to the exact role ARN (arn:aws:iam::396287094661:role/shoc-backend-dev), no wildcards: the secret resource policy grants secretsmanager:GetSecretValue AND the KMS key policy grants kms:Decrypt; either one alone fails silently at the receiver. Future staging/prod receiver roles are each an explicit policy addition with its own cross-family review — no prefix/wildcard trust. IAM/resource-policy change → mandatory GPT-4.1 cross-family review (Phase 6).
  • Post-deploy verification: assume shoc-backend-dev (or have Luby run it) and perform a real cross-account get-secret-value before declaring Phase 1 done (synth ≠ authz — the live call is the proof).
  • Handoff: send Luby the secret ARN, region, and the contract Rev pointer as an explicit step — the grant is useless if she doesn't have the ARN.

Phase 2 — Streams on the WO tables (CDK)

  • Enable stream=NEW_AND_OLD_IMAGES on WorkOrders and WorkOrderComments. In-place CFN update, no replacement, additive. Consumer audit (refreshed 2026-07-23): neither table has a stream, so no stream consumers can exist, and the tables have NO cross-stack consumers at all — seahaven-slack-bot (the former sole reader) was decommissioned 2026-07-23 and its grants dropped from CDK. Re-run this audit before deploy anyway; the reasoning pattern (re-audit before assuming safety) is the point. Note the change in the README data-contract section. Pre-deploy checks: (a) confirm the prod tables still carry RemovalPolicy.RETAIN and their pre-migration logical IDs; (b) cdk diff must show the table changes as in-place updates — if either table shows [Replacement] (e.g. from an accidental logical-ID change), abort; these are RETAIN production tables and a replacement would orphan them.
  • OLD_IMAGE is needed to detect the wo_status → cancelled transition (work_order.cancelled classification).

Phase 3 — Emitter Lambda workorder-shoc-emitter

  • Python 3.12, ARM64, explicit LogGroup via common.make_function_log_group, kebab-case name, memory_size=256, timeout=60s (POST timeout 10s per contract §2). Stdlib HTTP (urllib.request) — no new deps. Source at lambdas/wo/shoc_emitter/.
  • Ships DARK: both event-source mappings deploy with enabled=False — the full stack (Lambda, queues, alarms, secret, rotation) deploys and is testable with zero deliveries while SHOC has no receiver, so our merge cadence never depends on Luby's. Activation is a deliberate one-line enabled=True PR after her receiver passes the shared HMAC test vectors. While disabled, stream records simply age out unprocessed (24 h retention) — that's fine; SHOC loads history via the read API at activation anyway.
  • Two event-source mappings (one per stream): parallelization_factor=1 and bisect_batch_on_function_error=False (both required for strict in-order), batch_size=10, maximum_record_age = 24h (stream retention), retry_attempts=-1 (retry until expiry — blocking the shard is what preserves order), on_failure → SQS destination workorder-shoc-emitter-failures (14-day, SSL-enforced, house DLQ style).
  • Handler logic:
    1. Deserialize the Dynamo stream image → plain JSON (Decimal-safe).
    2. Classify: WorkOrders INSERT → work_order.created; MODIFY → work_order.updated, or work_order.cancelled when OLD.wo_status ≠ cancelled ∧ NEW.wo_status = cancelled; WorkOrderComments INSERT → work_order.comment_added. REMOVE events: skip (no deletes in this pipeline). Echo guard (forward-compat): if NEW.write_origin == "shoc-write-api", skip — a no-op branch today (nothing writes that attribute yet) that keeps the phase-2 write-back API from echoing SHOC's own writes back at it without re-touching this loop.
    3. Build envelope (contract §3); delivery_id = the stream record eventID (unique, stable across ESM retries); occurred_at = ApproximateCreationDateTime.
    4. Sign (contract §6) with keys[0] from the secret — cached per handbook Lambda pattern with a 5-minute TTL so rotation propagates.
    5. POST with a User-Agent: workorder-shoc-emitter/1 header; classify response per contract §7: 2xx → done; 429/5xx/timeout/conn-error → raise (ESM blocks + retries in order); other 4xx → send record to workorder-shoc-emitter-rejected SQS queue, log structured warning, return success (a contract bug must not block the shard for 24h).
    6. Structured JSON logging for every delivery attempt/outcome (delivery_id, event_type, work_order_id, status code, latency) — Logs-Insights-queryable, same style as the processor's sender_auth_rejected records.
  • Throughput note (parallelization_factor=1): ordering serializes per shard, but volume is ~760 events/day with bursts of a few events/second against a ~10s worst-case POST — a shard-blocking bottleneck only matters when SHOC is degraded, which is exactly when we want ordered blocking + the iterator-age alarm rather than parallel hammering. Accepted.
  • Env config: SHOC_WEBHOOK_URL (non-sensitive endpoint URL — env var/CDK context per secrets-and-config.md; HMAC is the auth, the URL is not), HMAC_SECRET_ARN.
  • IAM: stream read (grant_stream_read on both tables), GetSecretValue scoped to the one secret, kms:Decrypt on the CMK, SQS send to the two failure queues. → cross-family review.
  • CI plumbing (post-refactor, easy to miss): the emitter's and rotator's bundling commands need new AST-pin coverage in tests/test_bundle_consistency.py (the test currently only recognizes email_processor/web_ui command shapes — an unrecognized bundle ships unchecked); add lambdas/wo/shoc_emitter (and the rotator dir) to pytest.ini's explicit --cov= list (unlisted dirs are invisible to the 80% floor); touch tests/support/loader.py only if a real sibling-name collision exists (per-directory resolution usually isolates it).

Phase 4 — Monitoring (house alarm style: ALARM-only → site-alerts, NOT_BREACHING)

Alarm Metric Threshold
workorder-shoc-emitter-errors Lambda Errors Sum 5 min > 0, eval 1
workorder-shoc-emitter-throttles Lambda Throttles Sum 5 min > 0, eval 1
workorder-shoc-emitter-iterator-age IteratorAge Max 5 min ≥ 600,000 ms (10 min lag = SHOC likely down), eval 3/2
workorder-shoc-emitter-failures-messages SQS visible Max 5 min > 0, eval 1
workorder-shoc-emitter-rejected-messages SQS visible Max 5 min > 0, eval 1

House rules made explicit: every alarm sets only AlarmActions (no OKActions, no InsufficientDataActions — feedback_cloudwatch_alarms), and every alarm must actually carry the site-alerts action (no action = monitoring theater). Errors/throttles/duration come from common.add_standard_lambda_alarms; iterator-age and the two SQS-visible alarms don't fit that helper's shape and stay bespoke — don't misread the helper as covering everything. site-alerts exists in seahaven-prod (live-verified 2026-07-23) via Topic.from_topic_arn, correctly on the alias/seahaven-alarm-topics CMK — never alias/aws/sns, whose key policy silently blocks CloudWatch publishes (feedback_sns_alarm_topic_kms). Post-deploy, verify delivery on one new alarm with set-alarm-state + describe-alarm-history --history-item-type Action.

Optional EMF DeliveryOutcome metric (delivered/rejected, mirroring the ParseMethod pattern) for a delivery-rate dashboard — cheap, same zero-latency EMF approach.

Phase 5 — Replay tool + tests

  • scripts/replay_shoc_webhooks.py: rebuild events from WorkOrders/WorkOrderComments for a --work-order-id list or --since window and re-POST with "replay": true (contract §8). Dry-run default, house script style. Documented with usage examples in the README Scripts section (operator runbook for the -failures/-rejected alarms).
  • Unit tests (offline, no AWS — extend the pytest suite under the consolidated tests/ conventions): signature generation golden vectors (shareable with Luby so both sides verify against identical vectors); stream-record → envelope mapping incl. Decimal handling and null fields; cancelled transition classification incl. already-cancelled MODIFYs (no false cancelled events); the write_origin == "shoc-write-api" skip branch; an enum golden test pinning emitted wo_status/record_type values to lambdas/wo/email_processor/prompts.py (enum drift fails CI here, not in SHOC prod); response-classification matrix (2xx/429/5xx/timeout/4xx); secret-cache TTL + unknown-kid refresh.
  • Interaction note: PR #99 advisory A1 (AI-fallback comment_time nondeterminism) can duplicate comment rows on async retries → duplicate comment_added webhooks. SHOC dedupe on comment_id/delivery_id absorbs it; fixing A1 stays tracked separately.

Phase 6 — Gates (before merge, in order)

  1. ruff check + ruff format --check + pytest locally (pre-push hook enforced). Pre-flight: deploy credentials resolve to seahaven-prod (Phase 0 gate 4).
  2. GPT-4.1 cross-family review (cross_review.py) of the complete policy surface as a single diff — Lambda execution-role policy + KMS key policy + secret resource policy + rotation-Lambda role together, never piecemeal (cross-policy gaps hide between separately-reviewed fragments). Mandatory. All policy changes ship in this one PR so the review sees the complete surface.
  3. /sh-security-review — mandatory (auth surface: HMAC design + rotation; untrusted-input: email-derived data serialized into outbound requests; cross-account exposure). Resolve every confirmed critical/high.
  4. PR against main, CI green (gh pr checks), no admin-merge.

Phase 7 — Docs + memory (same conversation as ship, per global instructions)

  • README: new emitter Lambda, streams, secret + CMK, alarms, data-flow diagram, SHOC consumer contract section, replay-script runbook, the DESTROY-not-RETAIN rationale on the secret — and remove the stale "Consumer (read-only) — data contract" paragraph naming seahaven-slack-bot (decommissioned).
  • Confluence AWS Architecture Map (id 1540098): add emitter, streams, the workorder-ingest/shoc-webhook-hmac secret + CMK, the two failure SQS queues, and the cross-account SHOC edge (both the webhook POST and the secret-read trust) to the WorkorderIngestStack subgraph. Read the page first — the 2026-07-23 migration already edited it (v35); don't clobber those edits.
  • Memory: update project_procurement_ingest.md and project_shoc.md (cross-referenced) — webhook shipped, contract Rev, secret/CMK names + exact cross-account grant ARNs, rotation design, DESTROY-policy rationale, dark-ship/activation state, read API replaces SyncController, echo-guard design.

Deploy/cutover order

  1. cdk deploy (streams + secret + emitter ship in one deploy, ESMs disabled — zero deliveries, zero alarm noise; our merge cadence never depends on SHOC's readiness).
  2. Luby builds the receiver against the contract + shared test vectors and loads history via the read API; receiver passes the vectors at the target env.
  3. Activation: one-line enabled=True PR. The ESM starts at LATEST — no historical flood.
  4. Verify end-to-end with a live APM email; watch iterator-age + rejected alarms for 48h.
  5. SyncController's Dynamo scan (incl. SyncVendorReplies) retires with the mgmt decommission; the read API is the standing reconciliation path.

Explicit non-goals (this PR)

  • No DynamoDB retirement / write-path switch (kept out for scope discipline — the former blocker, seahaven-slack-bot, is decommissioned; revisit after SHOC prod + mgmt decommission).
  • No PO-pipeline webhook.
  • No change to workorder-email-processor or its parser.
  • No write-back endpoints (phase 2 of the read API, own PR + own IAM cross-review).