From c8f779a60b286359594ab3928bb451e6c3245c67 Mon Sep 17 00:00:00 2001 From: Adam Moussa Date: Thu, 23 Jul 2026 18:34:42 -0400 Subject: [PATCH] docs(webhook): revise SHOC webhook contract and plan for post-migration reality Branch re-cut on main 2026-07-23 (old base carried stale PR #99 commits). Contract Rev 2026-07-23: - Producer account corrected: seahaven-prod (011934824531); mgmt frozen - Reconciliation backstop is the new procurement read API, not SyncController - wo_status "unknown" is real; SHOC must map it (checklist item added) - write_origin forward-compat note for phase-2 write-back echo suppression - SyncVendorReplies retirement flagged (dead table, no vendor_reply event) Plan updates: - Account gate: seahaven-prod only; never enable streams on mgmt tables - Emitter ships DARK (ESMs enabled=False); activation is a deliberate flip after the SHOC receiver passes shared HMAC vectors - Post-refactor conventions: common.py helpers, bundle-consistency AST pins, pytest.ini --cov additions, consolidated test roots - Dedicated-CMK rationale, secret-ARN handooff step, consumer audit refreshed (slack-bot decommissioned), enum golden test, write_origin skip-branch test --- docs/shoc-webhook-contract.md | 164 ++++++++++++++++++++++++++++++++++ docs/shoc-webhook-plan.md | 108 ++++++++++++++++++++++ 2 files changed, 272 insertions(+) create mode 100644 docs/shoc-webhook-contract.md create mode 100644 docs/shoc-webhook-plan.md diff --git a/docs/shoc-webhook-contract.md b/docs/shoc-webhook-contract.md new file mode 100644 index 0000000..d83af83 --- /dev/null +++ b/docs/shoc-webhook-contract.md @@ -0,0 +1,164 @@ +# SHOC Work-Order Webhook — Delivery Contract (v1) + +**Status:** DRAFT — for SHOC team (Luby) review. **Rev 2026-07-23** (supersedes the 2026-07-16 draft: producer account corrected to seahaven-prod; reconciliation backstop changed to the new read API; §4.1 `unknown` status note; §3 `write_origin` forward-compat note. Sections 2, 5, 6, 7, and 10 are unchanged from the 07-16 draft). +**Producer:** `workorder-shoc-emitter` Lambda, Sea Haven **seahaven-prod** AWS account (**011934824531**, us-east-1). The former management account (328440206208) is frozen/rollback-only and will never host this feed, its secret, or its streams. +**Consumer:** SHOC backend (.NET 8), initially `https://api.dev.seahaven.com`. Endpoint path is SHOC's choice — suggested `POST /api/webhooks/work-orders`; map entities onto the work-orders domain from shoc-backend PR #10. + +## 1. Overview + +Every work-order mutation the procurement-ingest pipeline writes to DynamoDB is pushed to SHOC as an HTTPS POST within seconds. The feed is driven by DynamoDB Streams, so events are emitted **in the exact order the pipeline committed them**, per work order. DynamoDB remains the source of truth; this webhook is a realtime feed. The reconciliation backstop (and the initial-history load — the feed starts at activation time, not from history) is the **procurement read API** (`GET /work-orders`, `GET /work-orders/{id}/comments`, IAM SigV4 — see its OpenAPI doc). SHOC's existing SyncController DynamoDB scan is transitional: it points at the old management account and retires when those stacks are decommissioned. `SyncVendorReplies` should be dropped on the SHOC side — the `VendorReplies` table is dead (its writer was deleted), and no `vendor_reply` webhook event exists or is planned. + +## 2. Transport + +- HTTPS POST, `Content-Type: application/json; charset=utf-8`, body ≤ 256 KB. +- Producer timeout is **10 seconds**. Respond `2xx` as fast as possible; if your processing is slow, accept-and-enqueue internally rather than processing inline. + +## 3. Event envelope + +```json +{ + "schema_version": 1, + "delivery_id": "4b7c2f0e-...", + "event_type": "work_order.updated", + "occurred_at": "2026-07-16T14:03:22.114208+00:00", + "source": "procurement-ingest/workorder-shoc-emitter", + "replay": false, + "data": { ... } +} +``` + +| Field | Meaning | +|---|---| +| `schema_version` | Integer. Breaking payload changes increment it; SHOC should reject versions it doesn't know. | +| `delivery_id` | Unique per source event, **stable across producer retries** — your idempotency key. Persist it; ignore any delivery whose id you've already processed. | +| `event_type` | See §4. | +| `occurred_at` | ISO 8601 **with UTC offset** (`+00:00`), when the pipeline committed the write. | +| `replay` | `true` when re-sent by the operator replay tool after an outage. Same idempotency rules apply. | +| `data` | Event-type-specific body, §4. | + +> **Forward-compat (phase-2 write-back):** when SHOC later gains write endpoints on the read API, records SHOC itself wrote will carry a `write_origin` attribute (e.g. `"shoc-write-api"`), and the emitter will skip them so SHOC never receives an echo of its own write. Nothing to build now — just don't reject envelopes if a `data.write_origin` field appears later. + +## 4. Event types + +### 4.1 `work_order.created` / `work_order.updated` / `work_order.cancelled` + +Emitted from `WorkOrders` table writes: `created` on first insert, `updated` on any subsequent change, `cancelled` when `wo_status` transitions to `cancelled` (a `cancelled` event is a specialization of `updated` — same body). `data` is the **full current work-order state** (not a diff); fields absent from the source email are `null`: + +```json +{ + "work_order_id": "11144580730", + "wo_status": "assigned", + "description": "Dock door 14 won't close", + "customer": "AMAZON", + "site_code": "JFK8", + "building": "JFK8", + "address": "546 Gulf Ave, Staten Island, NY 10314", + "severity": "3-Normal", + "priority": "Medium", + "assigned_to": "Sea Haven Industries", + "date_reported": "2026-07-14T09:12:00", + "scheduled_start": "2026-07-17T08:00:00", + "due_date": "2026-07-21T17:00:00", + "record_type": "new_work_order", + "created_at": "2026-07-16T14:03:22.114208+00:00", + "updated_at": "2026-07-16T14:03:22.114208+00:00" +} +``` + +`wo_status` ∈ `new | assigned | in_progress | on_hold | completed | cancelled | unknown`. `record_type` (the triggering email's type) ∈ `new_work_order | update | comment | cancellation`. + +> **`unknown` is a real value** the extractor emits when the source email doesn't state a status — SHOC must define an explicit mapping for it (and defensively for any unrecognized future value) rather than letting it fall through a switch. Field names and enums in this section mirror the producer's write path, `lambdas/wo/email_processor/persistence.py` (+ `prompts.py` for enums) on `main` — that code is the authoritative source if this doc ever drifts. + +### 4.2 `work_order.comment_added` + +Emitted from `WorkOrderComments` inserts — one per source email (comments, but also update/cancellation event records): + +```json +{ + "work_order_id": "11144580730", + "comment_id": "11144580730#2026-04-27T23:51:48#a1b2c3d4e5f6", + "record_type": "comment", + "commenter": "APM Technician", + "text": "Vendor dispatched, ETA tomorrow AM.", + "created_at": "2026-04-27T23:51:48", + "ingested_at": "2026-07-16T14:03:22.114208+00:00" +} +``` + +`comment_id` is unique per source email and stable across retries — use it (or `delivery_id`) for dedupe. + +## 5. Ordering and delivery semantics + +- **At-least-once.** Duplicates are possible on retry; dedupe on `delivery_id`. +- **Per-work-order, per-event-family ordering is guaranteed:** all `work_order.*` state events for a given `work_order_id` arrive in commit order (a `cancelled` never precedes its `created`). Comment events are likewise ordered among themselves per work order. +- **Cross-family ordering is NOT guaranteed:** a `comment_added` may occasionally arrive before the `created` for its work order (separate streams). **Requirement:** on a comment for an unknown `work_order_id`, upsert a skeleton work order and let the state event backfill it — the same semantics the pipeline itself uses for out-of-order source emails. +- Expected volume ≈ 22,900 events/month (~760/day); >90% are comment events. Bursts of a few events/second are possible. + +## 6. Authentication — HMAC signature + +Every request carries: + +``` +X-SH-Timestamp: 1784642602 (unix seconds, producer clock) +X-SH-Key-Id: 2026-07-20T00 +X-SH-Signature: v1=hex(HMAC_SHA256(secret, "{timestamp}.{raw_body}")) +``` + +Verification requirements (fail closed — reject with `401` on any failure): + +1. Reject if `|now - X-SH-Timestamp| > 300s` (replay window). +2. Look up the secret for `X-SH-Key-Id` from the shared secret material (§6.1). Reject unknown key ids. +3. Recompute `HMAC-SHA256(secret, timestamp + "." + raw_body)` over the **raw request bytes** (before any JSON parsing/re-serialization) and compare **constant-time** against the `v1=` value. + +### 6.1 Key material and rotation + +The secret lives in **seahaven-prod (011934824531)** Secrets Manager: **`workorder-ingest/shoc-webhook-hmac`**, encrypted with a dedicated customer-managed KMS key so it is readable cross-account. The SHOC backend role (`arn:aws:iam::396287094661:role/shoc-backend-dev`) is granted `secretsmanager:GetSecretValue` + `kms:Decrypt` via resource policies — exact role ARN only; future `shoc-backend-staging`/`-prod` roles are each a deliberate, individually-reviewed policy addition (no wildcard/prefix trust). Secret value shape: + +```json +{ + "keys": [ + { "kid": "2026-07-20T00", "secret": "<64 hex chars>" }, + { "kid": "2026-06-20T00", "secret": "<64 hex chars>" } + ] +} +``` + +- Rotation is automated (30-day schedule): a new key is prepended, the previous key is kept for one overlap cycle. +- The producer always signs with `keys[0]`. The receiver must accept **any** listed `kid`. +- **Receiver must re-fetch the secret at most every 5 minutes** (cache TTL ≤ 300s) so rotation propagates with zero downtime. Also re-fetch immediately on an unknown-`kid` before rejecting. + +## 7. Response contract + +| Receiver response | Producer behavior | +|---|---| +| `2xx` | Delivered. Body ignored. | +| `429`, `5xx`, timeout, connection error | Retried with in-order blocking (that work order's events queue behind it) for up to **24 hours**; then parked for operator replay. | +| Any other `4xx` | Not retried — parked immediately and alarmed on our side. A `4xx` means a contract bug; expect us to call you. | + +## 8. Backstop / replay + +If SHOC is down longer than the retry window or deliveries are parked, an operator replay tool re-sends events rebuilt from DynamoDB (marked `"replay": true`). Because DynamoDB stays the source of truth, the procurement read API can also fully reconcile at any time (paginated `GET /work-orders` + per-WO comments). Missed webhooks are therefore an availability inconvenience, never data loss. + +## 9. SHOC-side checklist + +- [ ] Endpoint path confirmed + implemented against PR #10 entity model +- [ ] Raw-body HMAC verification, constant-time compare, ±300s timestamp window, fail-closed +- [ ] Cross-account secret fetch with ≤5-min cache + refresh-on-unknown-kid +- [ ] Dedupe store on `delivery_id` (and/or `comment_id`) +- [ ] Skeleton-upsert on comment-before-create +- [ ] Fast `2xx` ack (<10s), internal queue if processing is slow +- [ ] Timestamp parsing accepts `+00:00` offsets +- [ ] Reject unknown `schema_version` +- [ ] Explicit mapping for `wo_status: unknown` (and a safe default for unrecognized future values) +- [ ] Initial history loaded via the read API before relying on the feed (feed starts at activation, not from history) +- [ ] `SyncVendorReplies` dropped (dead table; no vendor-reply event exists) + +## 10. Environments + +| Env | URL | Status | +|---|---|---| +| dev | `https://api.dev.seahaven.com/` | first target (prod data in dev accepted 2026-07-16) | +| staging | `https://api.staging.seahaven.com/` | when staging deploys | +| prod | TBD | when SHOC prod exists | + +The target URL is producer-side config (per-env), so promotion is a config change, not a code change. diff --git a/docs/shoc-webhook-plan.md b/docs/shoc-webhook-plan.md new file mode 100644 index 0000000..4b47635 --- /dev/null +++ b/docs/shoc-webhook-plan.md @@ -0,0 +1,108 @@ +# Implementation Plan — SHOC Realtime Work-Order Webhook + +**Branch:** `feat/shoc-wo-webhook` (re-cut on `main` 2026-07-23; all build work targets the post-refactor layout — `cdk/common.py` helpers, explicit LogGroups, decomposed handlers, consolidated test roots — not the pre-refactor style earlier drafts implied). +**Companion docs:** `docs/shoc-webhook-contract.md` (Rev 2026-07-23 — the producer/consumer contract; Luby builds the receiver against it) and the procurement read API OpenAPI spec (`lambdas/api/openapi.json`, built in the sibling `procurement-api` effort — the reconciliation/backfill path). +**Requestor:** Luby (SHOC team). Decisions captured 2026-07-16: all WO event types, seconds-latency, full-payload push, HMAC + automated rotation, in-order per WO, prod-data-in-dev accepted, dual-write retained (DynamoDB retirement deferred for scope discipline — the original blocker, seahaven-slack-bot, was decommissioned 2026-07-23; retirement is now sequenced after SHOC prod + mgmt decommission, not this PR). +**Account (2026-07-23):** everything here builds in **seahaven-prod (011934824531)**. The old management account (328440206208) is frozen/rollback-only. + +## Architecture + +``` +SES → S3 → workorder-email-processor ──writes──▶ WorkOrders ─stream─▶ + └─writes──▶ WorkOrderComments ─stream─▶ workorder-shoc-emitter ──HMAC POST──▶ SHOC API + │ (retry-exhausted / non-retryable) + ▼ + shoc-emitter failure queues → alarms → replay script +``` + +Rationale (vs. inline call from the processor): DynamoDB Streams gives per-partition-key (per-`work_order_id`) commit-order delivery with shard-blocking retries — the ordering guarantee Luby requires — and keeps the email-ingestion hot path completely decoupled from SHOC availability. No change to PR #99's handler. + +## Phase 0 — Prerequisites (blocking) + +1. Branch re-cut on `main` (done 2026-07-23). +2. Luby signs off `docs/shoc-webhook-contract.md` Rev 2026-07-23 (endpoint path, PR #10 entity mapping, the `unknown`-status mapping, `SyncVendorReplies` retirement). +3. Confirm the SHOC receiver role ARN for the resource policies (`arn:aws:iam::396287094661:role/shoc-backend-dev`) — the role lives in SHOC's account, outside this repo's control; confirm it still exists before the CDK references it (the first cross-account `get-secret-value` test is the hard proof). +4. **⚠️ ACCOUNT GATE: every deploy in this plan targets seahaven-prod (011934824531) ONLY.** Never deploy this stack to mgmt (328440206208), and never "temporarily" enable streams on mgmt's copies of the WO tables — that is the only path by which the SES dual-delivery bake could ever double-send to SHOC. Verify `aws sts get-caller-identity` resolves to 011934824531 before `cdk deploy`. +5. Confirm with Luby which SHOC environment should receive the feed now — "prod-data-in-dev" was accepted 2026-07-16, pre-migration; re-confirm the target URL. + +## Phase 1 — Secret + rotation (CDK, `wo_stack.py`) + +This phase is an *incremental* deploy onto the already-migrated `workorder-ingest` stack in seahaven-prod — it is not part of the account-migration work. + +- New customer-managed KMS key `workorder-ingest-shoc-webhook-kms` (kebab-case per naming-conventions.md). Required because the default `aws/secretsmanager` key cannot serve cross-account reads. **Deliberately a dedicated key, not `alias/seahaven-dynamodb`:** reusing the DynamoDB CMK would hand the SHOC cross-account grant decrypt reach over the PO table's encryption key — the dedicated key scopes the grant to exactly this secret. +- New Secrets Manager secret **`workorder-ingest/shoc-webhook-hmac`** (naming per secrets-and-config.md `stack-name/value-name`), CDK-managed, encrypted with the CMK. Value shape: `{"keys": [{"kid", "secret"}, ...]}` (dual-key, contract §6.1). +- **RemovalPolicy: `DESTROY`, deliberately.** The value is machine-generated HMAC material with no operator-set content — fully regenerable by one rotation — so `RETAIN` buys nothing and would expose the fixed-name RETAIN orphan deadlock (failed first create orphans an empty shell holding the global name → every later create fails `AlreadyExists`; see `reference_secret_retain_orphan_deadlock`). Add a GOTCHA comment at the removal-policy line so nobody "hardens" it to RETAIN later. If a first deploy does fail, check for and force-delete any **empty** orphaned shell before retrying. Accidental-deletion recovery: recreate + rotate (new keys), notify Luby (receivers refresh within their 5-min TTL), watch the `-failures` queue and replay the gap — documented in the README runbook. Secret gets a `description` and `Purpose`/`ManagedBy` tags for discoverability. +- Rotation Lambda `workorder-shoc-hmac-rotator` (Python 3.12, ARM64, explicit LogGroup via `common.make_function_log_group`, alarms via the house `common.add_standard_lambda_alarms` idiom) on a 30-day `rotation_schedule`: generates a new 32-byte key, prepends as `keys[0]`, truncates to 2 entries. Single-user rotation (no external system holds the value — receivers re-fetch), so the standard 4-step rotation collapses to createSecret/finishSecret. Rotation-outage math: producer signs `keys[0]` with a ≤5-min cache; receiver keeps the previous key and refreshes on unknown `kid` (contract §6.1, receiver-side) — no delivery window where signatures can't verify. +- Cross-account read grants — **both halves required, to the exact role ARN** (`arn:aws:iam::396287094661:role/shoc-backend-dev`), no wildcards: the secret **resource policy** grants `secretsmanager:GetSecretValue` AND the **KMS key policy** grants `kms:Decrypt`; either one alone fails silently at the receiver. Future staging/prod receiver roles are each an explicit policy addition with its own cross-family review — no prefix/wildcard trust. **IAM/resource-policy change → mandatory GPT-4.1 cross-family review (Phase 6).** +- **Post-deploy verification:** assume `shoc-backend-dev` (or have Luby run it) and perform a real cross-account `get-secret-value` before declaring Phase 1 done (synth ≠ authz — the live call is the proof). +- **Handoff:** send Luby the secret ARN, region, and the contract Rev pointer as an explicit step — the grant is useless if she doesn't have the ARN. + +## Phase 2 — Streams on the WO tables (CDK) + +- Enable `stream=NEW_AND_OLD_IMAGES` on `WorkOrders` and `WorkOrderComments`. In-place CFN update, no replacement, additive. Consumer audit (refreshed 2026-07-23): **neither table has a stream, so no stream consumers can exist, and the tables have NO cross-stack consumers at all** — seahaven-slack-bot (the former sole reader) was decommissioned 2026-07-23 and its grants dropped from CDK. Re-run this audit before deploy anyway; the reasoning pattern (re-audit before assuming safety) is the point. Note the change in the README data-contract section. **Pre-deploy checks:** (a) confirm the prod tables still carry `RemovalPolicy.RETAIN` and their pre-migration logical IDs; (b) `cdk diff` must show the table changes as in-place updates — if either table shows `[Replacement]` (e.g. from an accidental logical-ID change), abort; these are RETAIN production tables and a replacement would orphan them. +- OLD_IMAGE is needed to detect the `wo_status → cancelled` transition (`work_order.cancelled` classification). + +## Phase 3 — Emitter Lambda `workorder-shoc-emitter` + +- Python 3.12, ARM64, explicit LogGroup via `common.make_function_log_group`, kebab-case name, `memory_size=256`, `timeout=60s` (POST timeout 10s per contract §2). Stdlib HTTP (`urllib.request`) — no new deps. Source at `lambdas/wo/shoc_emitter/`. +- **Ships DARK:** both event-source mappings deploy with `enabled=False` — the full stack (Lambda, queues, alarms, secret, rotation) deploys and is testable with zero deliveries while SHOC has no receiver, so our merge cadence never depends on Luby's. Activation is a deliberate one-line `enabled=True` PR after her receiver passes the shared HMAC test vectors. While disabled, stream records simply age out unprocessed (24 h retention) — that's fine; SHOC loads history via the read API at activation anyway. +- Two event-source mappings (one per stream): `parallelization_factor=1` and `bisect_batch_on_function_error=False` (both required for strict in-order), `batch_size=10`, `maximum_record_age` = 24h (stream retention), `retry_attempts=-1` (retry until expiry — blocking the shard is what preserves order), `on_failure` → SQS destination `workorder-shoc-emitter-failures` (14-day, SSL-enforced, house DLQ style). +- Handler logic: + 1. Deserialize the Dynamo stream image → plain JSON (Decimal-safe). + 2. Classify: `WorkOrders` INSERT → `work_order.created`; MODIFY → `work_order.updated`, or `work_order.cancelled` when OLD.wo_status ≠ `cancelled` ∧ NEW.wo_status = `cancelled`; `WorkOrderComments` INSERT → `work_order.comment_added`. REMOVE events: skip (no deletes in this pipeline). **Echo guard (forward-compat):** if NEW.`write_origin == "shoc-write-api"`, skip — a no-op branch today (nothing writes that attribute yet) that keeps the phase-2 write-back API from echoing SHOC's own writes back at it without re-touching this loop. + 3. Build envelope (contract §3); `delivery_id` = the stream record `eventID` (unique, stable across ESM retries); `occurred_at` = `ApproximateCreationDateTime`. + 4. Sign (contract §6) with `keys[0]` from the secret — cached per handbook Lambda pattern with a **5-minute TTL** so rotation propagates. + 5. POST with a `User-Agent: workorder-shoc-emitter/1` header; classify response per contract §7: `2xx` → done; `429/5xx/timeout/conn-error` → raise (ESM blocks + retries in order); other `4xx` → send record to `workorder-shoc-emitter-rejected` SQS queue, log structured warning, return success (a contract bug must not block the shard for 24h). + 6. Structured JSON logging for every delivery attempt/outcome (delivery_id, event_type, work_order_id, status code, latency) — Logs-Insights-queryable, same style as the processor's `sender_auth_rejected` records. +- Throughput note (`parallelization_factor=1`): ordering serializes per shard, but volume is ~760 events/day with bursts of a few events/second against a ~10s worst-case POST — a shard-blocking bottleneck only matters when SHOC is degraded, which is exactly when we *want* ordered blocking + the iterator-age alarm rather than parallel hammering. Accepted. +- Env config: `SHOC_WEBHOOK_URL` (non-sensitive endpoint URL — env var/CDK context per secrets-and-config.md; HMAC is the auth, the URL is not), `HMAC_SECRET_ARN`. +- IAM: stream read (`grant_stream_read` on both tables), `GetSecretValue` scoped to the one secret, `kms:Decrypt` on the CMK, SQS send to the two failure queues. **→ cross-family review.** +- CI plumbing (post-refactor, easy to miss): the emitter's and rotator's bundling commands need **new AST-pin coverage in `tests/test_bundle_consistency.py`** (the test currently only recognizes email_processor/web_ui command shapes — an unrecognized bundle ships unchecked); add `lambdas/wo/shoc_emitter` (and the rotator dir) to `pytest.ini`'s explicit `--cov=` list (unlisted dirs are invisible to the 80% floor); touch `tests/support/loader.py` only if a real sibling-name collision exists (per-directory resolution usually isolates it). + +## Phase 4 — Monitoring (house alarm style: ALARM-only → `site-alerts`, `NOT_BREACHING`) + +| Alarm | Metric | Threshold | +|---|---|---| +| `workorder-shoc-emitter-errors` | Lambda Errors Sum 5 min | > 0, eval 1 | +| `workorder-shoc-emitter-throttles` | Lambda Throttles Sum 5 min | > 0, eval 1 | +| `workorder-shoc-emitter-iterator-age` | IteratorAge Max 5 min | ≥ 600,000 ms (10 min lag = SHOC likely down), eval 3/2 | +| `workorder-shoc-emitter-failures-messages` | SQS visible Max 5 min | > 0, eval 1 | +| `workorder-shoc-emitter-rejected-messages` | SQS visible Max 5 min | > 0, eval 1 | + +House rules made explicit: every alarm sets **only** `AlarmActions` (no `OKActions`, no `InsufficientDataActions` — `feedback_cloudwatch_alarms`), and every alarm must actually carry the `site-alerts` action (no action = monitoring theater). Errors/throttles/duration come from `common.add_standard_lambda_alarms`; iterator-age and the two SQS-visible alarms don't fit that helper's shape and stay bespoke — don't misread the helper as covering everything. `site-alerts` exists in seahaven-prod (live-verified 2026-07-23) via `Topic.from_topic_arn`, correctly on the `alias/seahaven-alarm-topics` CMK — never `alias/aws/sns`, whose key policy silently blocks CloudWatch publishes (`feedback_sns_alarm_topic_kms`). Post-deploy, verify delivery on one new alarm with `set-alarm-state` + `describe-alarm-history --history-item-type Action`. + +Optional EMF `DeliveryOutcome` metric (delivered/rejected, mirroring the ParseMethod pattern) for a delivery-rate dashboard — cheap, same zero-latency EMF approach. + +## Phase 5 — Replay tool + tests + +- `scripts/replay_shoc_webhooks.py`: rebuild events from `WorkOrders`/`WorkOrderComments` for a `--work-order-id` list or `--since` window and re-POST with `"replay": true` (contract §8). Dry-run default, house script style. Documented with usage examples in the README Scripts section (operator runbook for the `-failures`/`-rejected` alarms). +- Unit tests (offline, no AWS — extend the pytest suite under the consolidated `tests/` conventions): signature generation golden vectors (shareable with Luby so both sides verify against identical vectors); stream-record → envelope mapping incl. Decimal handling and `null` fields; `cancelled` transition classification incl. already-cancelled MODIFYs (no false `cancelled` events); the `write_origin == "shoc-write-api"` skip branch; an enum golden test pinning emitted `wo_status`/`record_type` values to `lambdas/wo/email_processor/prompts.py` (enum drift fails CI here, not in SHOC prod); response-classification matrix (2xx/429/5xx/timeout/4xx); secret-cache TTL + unknown-kid refresh. +- Interaction note: PR #99 advisory A1 (AI-fallback `comment_time` nondeterminism) can duplicate comment rows on async retries → duplicate `comment_added` webhooks. SHOC dedupe on `comment_id`/`delivery_id` absorbs it; fixing A1 stays tracked separately. + +## Phase 6 — Gates (before merge, in order) + +1. `ruff check` + `ruff format --check` + `pytest` locally (pre-push hook enforced). Pre-flight: deploy credentials resolve to seahaven-prod (Phase 0 gate 4). +2. **GPT-4.1 cross-family review** (`cross_review.py`) of the complete policy surface **as a single diff** — Lambda execution-role policy + KMS key policy + secret resource policy + rotation-Lambda role together, never piecemeal (cross-policy gaps hide between separately-reviewed fragments). Mandatory. All policy changes ship in this one PR so the review sees the complete surface. +3. **`/sh-security-review`** — mandatory (auth surface: HMAC design + rotation; untrusted-input: email-derived data serialized into outbound requests; cross-account exposure). Resolve every confirmed critical/high. +4. PR against `main`, CI green (`gh pr checks`), no admin-merge. + +## Phase 7 — Docs + memory (same conversation as ship, per global instructions) + +- README: new emitter Lambda, streams, secret + CMK, alarms, data-flow diagram, SHOC consumer contract section, replay-script runbook, the DESTROY-not-RETAIN rationale on the secret — and **remove the stale "Consumer (read-only) — data contract" paragraph naming seahaven-slack-bot** (decommissioned). +- Confluence **AWS Architecture Map** (id 1540098): add emitter, streams, the `workorder-ingest/shoc-webhook-hmac` secret + CMK, the two failure SQS queues, and the cross-account SHOC edge (both the webhook POST and the secret-read trust) to the WorkorderIngestStack subgraph. Read the page first — the 2026-07-23 migration already edited it (v35); don't clobber those edits. +- Memory: update `project_procurement_ingest.md` and `project_shoc.md` (cross-referenced) — webhook shipped, contract Rev, secret/CMK names + exact cross-account grant ARNs, rotation design, DESTROY-policy rationale, dark-ship/activation state, read API replaces SyncController, echo-guard design. + +## Deploy/cutover order + +1. `cdk deploy` (streams + secret + emitter ship in one deploy, **ESMs disabled** — zero deliveries, zero alarm noise; our merge cadence never depends on SHOC's readiness). +2. Luby builds the receiver against the contract + shared test vectors and loads history via the read API; receiver passes the vectors at the target env. +3. Activation: one-line `enabled=True` PR. The ESM starts at `LATEST` — no historical flood. +4. Verify end-to-end with a live APM email; watch `iterator-age` + `rejected` alarms for 48h. +5. SyncController's Dynamo scan (incl. `SyncVendorReplies`) retires with the mgmt decommission; the read API is the standing reconciliation path. + +## Explicit non-goals (this PR) + +- No DynamoDB retirement / write-path switch (kept out for scope discipline — the former blocker, seahaven-slack-bot, is decommissioned; revisit after SHOC prod + mgmt decommission). +- No PO-pipeline webhook. +- No change to `workorder-email-processor` or its parser. +- No write-back endpoints (phase 2 of the read API, own PR + own IAM cross-review).