17 KiB
Implementation Plan — SHOC Realtime Work-Order Webhook
Branch: feat/shoc-wo-webhook (re-cut on main 2026-07-23; all build work targets the post-refactor layout — cdk/common.py helpers, explicit LogGroups, decomposed handlers, consolidated test roots — not the pre-refactor style earlier drafts implied).
Companion docs: docs/shoc-webhook-contract.md (Rev 2026-07-23 — the producer/consumer contract; Luby builds the receiver against it) and the procurement read API OpenAPI spec (lambdas/api/openapi.json, built in the sibling procurement-api effort — the reconciliation/backfill path).
Requestor: Luby (SHOC team). Decisions captured 2026-07-16: all WO event types, seconds-latency, full-payload push, HMAC + automated rotation, in-order per WO, prod-data-in-dev accepted, dual-write retained (DynamoDB retirement deferred for scope discipline — the original blocker, seahaven-slack-bot, was decommissioned 2026-07-23; revisit after SHOC prod).
Account: everything here builds in seahaven-prod (011934824531). Mgmt procurement-ingest stacks were deleted in PLAT-67 (2026-08-05); RETAIN leftovers there are cold archive only.
Architecture
SES → S3 → workorder-email-processor ──writes──▶ WorkOrders ─stream─▶
└─writes──▶ WorkOrderComments ─stream─▶ workorder-shoc-emitter ──HMAC POST──▶ SHOC API
│ (retry-exhausted / non-retryable)
▼
shoc-emitter failure queues → alarms → replay script
Rationale (vs. inline call from the processor): DynamoDB Streams gives per-partition-key (per-work_order_id) commit-order delivery with shard-blocking retries — the ordering guarantee Luby requires — and keeps the email-ingestion hot path completely decoupled from SHOC availability. No change to PR #99's handler.
Phase 0 — Prerequisites (blocking)
- Branch re-cut on
main(done 2026-07-23). - Luby signs off
docs/shoc-webhook-contract.mdRev 2026-07-23 (endpoint path, PR #10 entity mapping, theunknown-status mapping,SyncVendorRepliesretirement). - Confirm the SHOC receiver role ARN for the resource policies (
arn:aws:iam::396287094661:role/shoc-backend-dev) — the role lives in SHOC's account, outside this repo's control; confirm it still exists before the CDK references it (the first cross-accountget-secret-valuetest is the hard proof). - ⚠️ ACCOUNT GATE: every deploy in this plan targets seahaven-prod (011934824531) ONLY. Never deploy this app to mgmt (328440206208) — templates are account-agnostic and would recreate stacks against cold-archive RETAIN data (PLAT-67). Verify
aws sts get-caller-identityresolves to 011934824531 beforecdk deploy. - Confirm with Luby which SHOC environment should receive the feed now — "prod-data-in-dev" was accepted 2026-07-16, pre-migration; re-confirm the target URL.
Phase 1 — Secret + rotation (CDK, wo_stack.py)
This phase is an incremental deploy onto the already-migrated workorder-ingest stack in seahaven-prod — it is not part of the account-migration work.
- New customer-managed KMS key
workorder-ingest-shoc-webhook-kms(kebab-case per naming-conventions.md). Required because the defaultaws/secretsmanagerkey cannot serve cross-account reads. Deliberately a dedicated key, notalias/seahaven-dynamodb: reusing the DynamoDB CMK would hand the SHOC cross-account grant decrypt reach over the PO table's encryption key — the dedicated key scopes the grant to exactly this secret. - New Secrets Manager secret
workorder-ingest/shoc-webhook-hmac(naming per secrets-and-config.mdstack-name/value-name), CDK-managed, encrypted with the CMK. Value shape:{"keys": [{"kid", "secret"}, ...]}(dual-key, contract §6.1). - RemovalPolicy:
DESTROY, deliberately. The value is machine-generated HMAC material with no operator-set content — fully regenerable by one rotation — soRETAINbuys nothing and would expose the fixed-name RETAIN orphan deadlock (failed first create orphans an empty shell holding the global name → every later create failsAlreadyExists; seereference_secret_retain_orphan_deadlock). Add a GOTCHA comment at the removal-policy line so nobody "hardens" it to RETAIN later. If a first deploy does fail, check for and force-delete any empty orphaned shell before retrying. Accidental-deletion recovery: recreate + rotate (new keys), notify Luby (receivers refresh within their 5-min TTL), watch the-failuresqueue and replay the gap — documented in the README runbook. Secret gets adescriptionandPurpose/ManagedBytags for discoverability. - Rotation Lambda
workorder-shoc-hmac-rotator(Python 3.12, ARM64, explicit LogGroup viacommon.make_function_log_group, alarms via the housecommon.add_standard_lambda_alarmsidiom) on a 30-dayrotation_schedule: generates a new 32-byte key, prepends askeys[0], truncates to 2 entries. Single-user rotation (no external system holds the value — receivers re-fetch), so the standard 4-step rotation collapses to createSecret/finishSecret. Rotation-outage math: producer signskeys[0]with a ≤5-min cache; receiver keeps the previous key and refreshes on unknownkid(contract §6.1, receiver-side) — no delivery window where signatures can't verify. - Cross-account read grants — both halves required, to the exact role ARN (
arn:aws:iam::396287094661:role/shoc-backend-dev), no wildcards: the secret resource policy grantssecretsmanager:GetSecretValueAND the KMS key policy grantskms:Decrypt; either one alone fails silently at the receiver. Future staging/prod receiver roles are each an explicit policy addition with its own cross-family review — no prefix/wildcard trust. IAM/resource-policy change → mandatory GPT-4.1 cross-family review (Phase 6). - Post-deploy verification: assume
shoc-backend-dev(or have Luby run it) and perform a real cross-accountget-secret-valuebefore declaring Phase 1 done (synth ≠ authz — the live call is the proof). - Handoff: send Luby the secret ARN, region, and the contract Rev pointer as an explicit step — the grant is useless if she doesn't have the ARN.
Phase 2 — Streams on the WO tables (CDK)
- Enable
stream=NEW_AND_OLD_IMAGESonWorkOrdersandWorkOrderComments. In-place CFN update, no replacement, additive. Consumer audit (refreshed 2026-07-23): neither table has a stream, so no stream consumers can exist, and the tables have NO cross-stack consumers at all — seahaven-slack-bot (the former sole reader) was decommissioned 2026-07-23 and its grants dropped from CDK. Re-run this audit before deploy anyway; the reasoning pattern (re-audit before assuming safety) is the point. Note the change in the README data-contract section. Pre-deploy checks: (a) confirm the prod tables still carryRemovalPolicy.RETAINand their pre-migration logical IDs; (b)cdk diffmust show the table changes as in-place updates — if either table shows[Replacement](e.g. from an accidental logical-ID change), abort; these are RETAIN production tables and a replacement would orphan them. - OLD_IMAGE is needed to detect the
wo_status → cancelledtransition (work_order.cancelledclassification).
Phase 3 — Emitter Lambda workorder-shoc-emitter
- Python 3.12, ARM64, explicit LogGroup via
common.make_function_log_group, kebab-case name,memory_size=256,timeout=60s(POST timeout 10s per contract §2). Stdlib HTTP (urllib.request) — no new deps. Source atlambdas/wo/shoc_emitter/. - Ships DARK: both event-source mappings deploy with
enabled=False— the full stack (Lambda, queues, alarms, secret, rotation) deploys and is testable with zero deliveries while SHOC has no receiver, so our merge cadence never depends on Luby's. Activation is a deliberate one-lineenabled=TruePR after her receiver passes the shared HMAC test vectors. While disabled, stream records simply age out unprocessed (24 h retention) — that's fine; SHOC loads history via the read API at activation anyway. - Two event-source mappings (one per stream):
parallelization_factor=1andbisect_batch_on_function_error=False(both required for strict in-order),batch_size=10,maximum_record_age= 24h (stream retention),retry_attempts=-1(retry until expiry — blocking the shard is what preserves order),on_failure→ SQS destinationworkorder-shoc-emitter-failures(14-day, SSL-enforced, house DLQ style). - Handler logic:
- Deserialize the Dynamo stream image → plain JSON (Decimal-safe).
- Classify:
WorkOrdersINSERT →work_order.created; MODIFY →work_order.updated, orwork_order.cancelledwhen OLD.wo_status ≠cancelled∧ NEW.wo_status =cancelled;WorkOrderCommentsINSERT →work_order.comment_added. REMOVE events: skip (no deletes in this pipeline). Echo guard (forward-compat): if NEW.write_origin == "shoc-write-api", skip — a no-op branch today (nothing writes that attribute yet) that keeps the phase-2 write-back API from echoing SHOC's own writes back at it without re-touching this loop. - Build envelope (contract §3);
delivery_id= the stream recordeventID(unique, stable across ESM retries);occurred_at=ApproximateCreationDateTime. - Sign (contract §6) with
keys[0]from the secret — cached per handbook Lambda pattern with a 5-minute TTL so rotation propagates. - POST with a
User-Agent: workorder-shoc-emitter/1header; classify response per contract §7:2xx→ done;429/5xx/timeout/conn-error→ raise (ESM blocks + retries in order); other4xx→ send record toworkorder-shoc-emitter-rejectedSQS queue, log structured warning, return success (a contract bug must not block the shard for 24h). - Structured JSON logging for every delivery attempt/outcome (delivery_id, event_type, work_order_id, status code, latency) — Logs-Insights-queryable, same style as the processor's
sender_auth_rejectedrecords.
- Throughput note (
parallelization_factor=1): ordering serializes per shard, but volume is ~760 events/day with bursts of a few events/second against a ~10s worst-case POST — a shard-blocking bottleneck only matters when SHOC is degraded, which is exactly when we want ordered blocking + the iterator-age alarm rather than parallel hammering. Accepted. - Env config:
SHOC_WEBHOOK_URL(non-sensitive endpoint URL — env var/CDK context per secrets-and-config.md; HMAC is the auth, the URL is not),HMAC_SECRET_ARN. - IAM: stream read (
grant_stream_readon both tables),GetSecretValuescoped to the one secret,kms:Decrypton the CMK, SQS send to the two failure queues. → cross-family review. - CI plumbing (post-refactor, easy to miss): the emitter's and rotator's bundling commands need new AST-pin coverage in
tests/test_bundle_consistency.py(the test currently only recognizes email_processor/web_ui command shapes — an unrecognized bundle ships unchecked); addlambdas/wo/shoc_emitter(and the rotator dir) topytest.ini's explicit--cov=list (unlisted dirs are invisible to the 80% floor); touchtests/support/loader.pyonly if a real sibling-name collision exists (per-directory resolution usually isolates it).
Phase 4 — Monitoring (house alarm style: ALARM-only → site-alerts, NOT_BREACHING)
| Alarm | Metric | Threshold |
|---|---|---|
workorder-shoc-emitter-errors |
Lambda Errors Sum 5 min | > 0, eval 1 |
workorder-shoc-emitter-throttles |
Lambda Throttles Sum 5 min | > 0, eval 1 |
workorder-shoc-emitter-iterator-age |
IteratorAge Max 5 min | ≥ 600,000 ms (10 min lag = SHOC likely down), eval 3/2 |
workorder-shoc-emitter-failures-messages |
SQS visible Max 5 min | > 0, eval 1 |
workorder-shoc-emitter-rejected-messages |
SQS visible Max 5 min | > 0, eval 1 |
House rules made explicit: every alarm sets only AlarmActions (no OKActions, no InsufficientDataActions — feedback_cloudwatch_alarms), and every alarm must actually carry the site-alerts action (no action = monitoring theater). Errors/throttles/duration come from common.add_standard_lambda_alarms; iterator-age and the two SQS-visible alarms don't fit that helper's shape and stay bespoke — don't misread the helper as covering everything. site-alerts exists in seahaven-prod (live-verified 2026-07-23) via Topic.from_topic_arn, correctly on the alias/seahaven-alarm-topics CMK — never alias/aws/sns, whose key policy silently blocks CloudWatch publishes (feedback_sns_alarm_topic_kms). Post-deploy, verify delivery on one new alarm with set-alarm-state + describe-alarm-history --history-item-type Action.
Optional EMF DeliveryOutcome metric (delivered/rejected, mirroring the ParseMethod pattern) for a delivery-rate dashboard — cheap, same zero-latency EMF approach.
Phase 5 — Replay tool + tests
scripts/replay_shoc_webhooks.py: rebuild events fromWorkOrders/WorkOrderCommentsfor a--work-order-idlist or--sincewindow and re-POST with"replay": true(contract §8). Dry-run default, house script style. Documented with usage examples in the README Scripts section (operator runbook for the-failures/-rejectedalarms).- Unit tests (offline, no AWS — extend the pytest suite under the consolidated
tests/conventions): signature generation golden vectors (shareable with Luby so both sides verify against identical vectors); stream-record → envelope mapping incl. Decimal handling andnullfields;cancelledtransition classification incl. already-cancelled MODIFYs (no falsecancelledevents); thewrite_origin == "shoc-write-api"skip branch; an enum golden test pinning emittedwo_status/record_typevalues tolambdas/wo/email_processor/prompts.py(enum drift fails CI here, not in SHOC prod); response-classification matrix (2xx/429/5xx/timeout/4xx); secret-cache TTL + unknown-kid refresh. - Interaction note: PR #99 advisory A1 (AI-fallback
comment_timenondeterminism) can duplicate comment rows on async retries → duplicatecomment_addedwebhooks. SHOC dedupe oncomment_id/delivery_idabsorbs it; fixing A1 stays tracked separately.
Phase 6 — Gates (before merge, in order)
ruff check+ruff format --check+pytestlocally (pre-push hook enforced). Pre-flight: deploy credentials resolve to seahaven-prod (Phase 0 gate 4).- GPT-4.1 cross-family review (
cross_review.py) of the complete policy surface as a single diff — Lambda execution-role policy + KMS key policy + secret resource policy + rotation-Lambda role together, never piecemeal (cross-policy gaps hide between separately-reviewed fragments). Mandatory. All policy changes ship in this one PR so the review sees the complete surface. /sh-security-review— mandatory (auth surface: HMAC design + rotation; untrusted-input: email-derived data serialized into outbound requests; cross-account exposure). Resolve every confirmed critical/high.- PR against
main, CI green (gh pr checks), no admin-merge.
Phase 7 — Docs + memory (same conversation as ship, per global instructions)
- README: new emitter Lambda, streams, secret + CMK, alarms, data-flow diagram, SHOC consumer contract section, replay-script runbook, the DESTROY-not-RETAIN rationale on the secret — and remove the stale "Consumer (read-only) — data contract" paragraph naming seahaven-slack-bot (decommissioned).
- Confluence AWS Architecture Map (id 1540098): add emitter, streams, the
workorder-ingest/shoc-webhook-hmacsecret + CMK, the two failure SQS queues, and the cross-account SHOC edge (both the webhook POST and the secret-read trust) to the WorkorderIngestStack subgraph. Read the page first — the 2026-07-23 migration already edited it (v35); don't clobber those edits. - Memory: update
project_procurement_ingest.mdandproject_shoc.md(cross-referenced) — webhook shipped, contract Rev, secret/CMK names + exact cross-account grant ARNs, rotation design, DESTROY-policy rationale, dark-ship/activation state, read API replaces SyncController, echo-guard design.
Deploy/cutover order
cdk deploy(streams + secret + emitter ship in one deploy, ESMs disabled — zero deliveries, zero alarm noise; our merge cadence never depends on SHOC's readiness).- Luby builds the receiver against the contract + shared test vectors and loads history via the read API; receiver passes the vectors at the target env.
- Activation: one-line
enabled=TruePR. The ESM starts atLATEST— no historical flood. - Verify end-to-end with a live APM email; watch
iterator-age+rejectedalarms for 48h. - SyncController's Dynamo scan (incl.
SyncVendorReplies) is SHOC-owned retirement work now that mgmt stacks are gone (PLAT-67); the read API is the standing reconciliation path.
Explicit non-goals (this PR)
- No DynamoDB retirement / write-path switch (kept out for scope discipline — the former blocker, seahaven-slack-bot, is decommissioned; revisit after SHOC prod).
- No PO-pipeline webhook.
- No change to
workorder-email-processoror its parser. - No write-back endpoints (phase 2 of the read API, own PR + own IAM cross-review).