feat(webhook): SHOC WO webhook emitter - dark-ship streams + HMAC secret/rotation (PR-2) (#137)
* docs(webhook): revise SHOC webhook contract and plan for post-migration reality
Branch re-cut on main 2026-07-23 (old base carried stale PR #99 commits).
Contract Rev 2026-07-23:
- Producer account corrected: seahaven-prod (011934824531); mgmt frozen
- Reconciliation backstop is the new procurement read API, not SyncController
- wo_status "unknown" is real; SHOC must map it (checklist item added)
- write_origin forward-compat note for phase-2 write-back echo suppression
- SyncVendorReplies retirement flagged (dead table, no vendor_reply event)
Plan updates:
- Account gate: seahaven-prod only; never enable streams on mgmt tables
- Emitter ships DARK (ESMs enabled=False); activation is a deliberate flip
after the SHOC receiver passes shared HMAC vectors
- Post-refactor conventions: common.py helpers, bundle-consistency AST pins,
pytest.ini --cov additions, consolidated test roots
- Dedicated-CMK rationale, secret-ARN handooff step, consumer audit refreshed
(slack-bot decommissioned), enum golden test, write_origin skip-branch test
* feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation
Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC
call-and-be-called effort). Everything ships DARK: both DynamoDB event
source mappings deploy enabled=False; activation is a deliberate
one-line follow-up PR gated on the SHOC receiver passing the shared
HMAC test vectors.
- Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments
(in-place update, RETAIN + logical IDs untouched; no existing
consumers — verified live, neither table had a stream).
- workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope ->
HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard
ordering (parallelization 1, bisect off, retry until 24h age,
ReportBatchItemFailures); 429/5xx/timeout block the shard in order,
other 4xx park to workorder-shoc-emitter-rejected; ESM failures ->
workorder-shoc-emitter-failures (metadata; replay rebuilds from
DynamoDB). Echo guard skips write_origin=shoc-write-api.
- Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK
(alias workorder-ingest-shoc-webhook-kms); cross-account
GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly
arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy
DESTROY deliberately (machine-generated material; avoids the
fixed-name RETAIN-orphan deadlock).
- workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap,
64-hex keys, kid = UTC %Y-%m-%dT%H.
- Alarms (ALARM-only -> site-alerts): emitter errors/throttles/
duration + iterator-age (>=10 min) + failures/rejected queue
depth; rotator standard trio.
- scripts/replay_shoc_webhooks.py: dry-run-default operator replay
(rebuilds from tables, replay:true envelopes).
- Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared
with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay
signing pinned to identical vectors); bundle-consistency AST pins
for both new bundles.
- README: WO stack + webhook feed section, alarm table, runbooks;
removed stale seahaven-slack-bot consumer references.
* fix(webhook): kms:ViaService pins, https-only delivery, cross-account principal CI pin
GPT-4.1 cross-family review of the policy surface (no BLOCK): FIX applied
to the cross-account shoc-backend-dev Decrypt statement and both Lambda
role KMS grants (the key is only ever used via Secrets Manager); its
invariant-enforcement QUESTION answered durably with
tests/test_cross_account_principal_pin.py (any new foreign IAM principal
in cdk/ fails CI). Scanner mediums fixed: delivery.py and the replay
script now refuse non-https URLs (urllib follows file:// and http://).
SQS metadata-action and dynamodb:ListStreams NITs skipped: standard CDK
grant shapes; ListStreams has no resource-level scoping. The 4 gitleaks
HIGHs on docs/shoc-webhook-test-vectors.json are deliberate non-secrets
(shared receiver-verification vectors) suppressed machine-level with
justification.
* harden(webhook): resolve /sh-security-review findings (1 confirmed medium + cheap fixes)
High-recall detector fan-out (injection/authz/secrets-crypto/iac-iam/logic)
+ proof-or-kill verifier. Gate PASSES: 1 confirmed medium, 0 confirmed
critical/high. Confirmed finding fixed; several unverified-but-cheap
hardenings applied since the emitter ships dark and activation is weeks out.
- CONFIRMED medium (confused deputy): the rotation Lambda's generated
invoke permission for secretsmanager.amazonaws.com carried no
SourceAccount/SourceArn, so any account's Secrets Manager could invoke
the rotator. Patched the generated CfnPermission in place (a second
permission would be additive, not restrictive) to pin account + this
secret ARN.
- delivery + replay: refuse to follow receiver 3xx redirects (no-redirect
opener) so live X-SH-* auth headers can't be forwarded to a
receiver-chosen Location and an http:// Location can't slip past the
https guard. Fixed the "unfollowed 3xx" comment that was factually wrong.
- delivery: classify 401/403 as retryable (invalidate key cache + retry in
order) instead of parking -- transient auth failures (rotation outran the
TTL cache, clock skew) are availability events, not contract bugs.
- envelope: build_event now genuinely total (guarded eventID /
ApproximateCreationDateTime subscripts) per its own never-raise contract.
- handler: catch-all so an unexpected per-record error (e.g. SQS park
failure) reports only that record instead of failing the whole batch
(which would re-deliver every earlier success for 24h); per-invocation
emit/skip batch summary so a systemic silent drop is queryable/alarmable.
- rotator: narrow the AWSCURRENT-read except to ResourceNotFound/JSONDecode
(transient SM/KMS errors re-raise so the overlap key isn't silently
dropped); kid uniqueness checked against ALL retained kids with a random
suffix on collision (never reissue a kid for a different secret).
- contract: skeleton-upsert required on ANY unknown work_order_id (not just
comment-before-create) + monotonicity guard (ignore older updated_at), so
a parked created or an out-of-order replay can't corrupt receiver state.
Unverified/refuted findings left as-is with rationale: the two "high" logic
claims (whole-batch crash triggers, ordering violation) were refuted on
reachability (real stream records carry required fields; persistence writes
strings only; full-state idempotent upsert absorbs the ordering gap). Signed
kid/version binding (AUTHZ-002) declined: coordinated contract change, not
cheap, no exploit with one algorithm/key.
* fix(webhook): drop kid from rotator test_ok log (CodeQL clear-text-logging FP)
GHAS CodeQL flagged py/clear-text-logging-sensitive-data (high) at
_test_secret's success log because head["kid"] is subscripted from the
same parsed-secret dict that holds head["secret"] — the taint tracker
can't tell the non-secret key id from the secret. The secret value is
never logged. Rather than dismiss the alert (fragile; re-alerts on line
moves), remove the flow: kid is already logged at stage time in
_create_secret and version_id correlates the steps, so the test_ok log
keeps only event + version_id. Also hardens against a future edit that
swaps the logged field.
2026-07-24 18:12:20 -04:00
|
|
|
"""
|
|
|
|
|
SHOC work-order webhook emitter Lambda.
|
|
|
|
|
|
|
|
|
|
Consumes the WorkOrders/WorkOrderComments DynamoDB streams and pushes each
|
|
|
|
|
mutation to SHOC as an HMAC-signed HTTPS POST (docs/shoc-webhook-contract.md).
|
|
|
|
|
This handler is the thin event loop; the work lives in flat sibling modules
|
|
|
|
|
(bare-name imports resolve via the same flat-landing bundling as the email
|
|
|
|
|
processor's siblings):
|
|
|
|
|
envelope.py -- stream record -> envelope mapping + event classification
|
|
|
|
|
delivery.py -- secret cache, signing, POST, response classification
|
|
|
|
|
|
|
|
|
|
Ships DARK: both event-source mappings deploy with enabled=False, so this code
|
|
|
|
|
runs zero deliveries until the activation PR flips them on after SHOC's
|
|
|
|
|
receiver passes the shared HMAC test vectors.
|
|
|
|
|
|
|
|
|
|
Ordering semantics: the ESMs run parallelization_factor=1 with bisect-on-error
|
|
|
|
|
off, and this loop processes records strictly in order. A retryable failure
|
|
|
|
|
(429/5xx/timeout/connection error) stops the batch immediately and reports
|
|
|
|
|
that record via report_batch_item_failures -- earlier successes are not
|
|
|
|
|
re-delivered, and the ESM blocks the shard and retries from the failed record,
|
|
|
|
|
preserving per-work-order commit order. Non-retryable 4xx responses are a
|
|
|
|
|
contract bug, not an availability blip: the full envelope is parked on the
|
|
|
|
|
rejected queue (alarmed) and the loop continues, so a bad payload can never
|
|
|
|
|
block the shard for 24 hours.
|
|
|
|
|
|
|
|
|
|
No healthcheck branch on purpose: stream consumers are not smoke-gated
|
|
|
|
|
(po-ingest-site-extractor precedent) -- there is no direct-invoke path to
|
|
|
|
|
probe, and a synthetic stream record would be a real delivery.
|
|
|
|
|
"""
|
|
|
|
|
|
|
|
|
|
import json
|
|
|
|
|
import logging
|
|
|
|
|
import os
|
|
|
|
|
import time
|
|
|
|
|
|
|
|
|
|
import boto3
|
|
|
|
|
import delivery
|
|
|
|
|
import envelope
|
|
|
|
|
from botocore.config import Config
|
|
|
|
|
|
|
|
|
|
logger = logging.getLogger()
|
|
|
|
|
logger.setLevel(logging.INFO)
|
|
|
|
|
|
|
|
|
|
REJECTED_QUEUE_URL = os.environ.get("REJECTED_QUEUE_URL")
|
|
|
|
|
|
|
|
|
|
# Bound the SQS client's timeouts (Open SWE #17): a slow SQS response while
|
|
|
|
|
# parking a rejected envelope must not hang the invocation toward its 60s
|
|
|
|
|
# timeout and re-deliver the whole batch.
|
|
|
|
|
_BOTO_CONFIG = Config(
|
|
|
|
|
connect_timeout=3, read_timeout=5, retries={"max_attempts": 2, "mode": "standard"}
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
# Lazy cached SQS client. Keeps the public attribute name ``sqs`` so the test
|
|
|
|
|
# monkeypatch target changes module only, not attribute name.
|
|
|
|
|
sqs = None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _get_sqs():
|
|
|
|
|
global sqs
|
|
|
|
|
if sqs is None:
|
|
|
|
|
sqs = boto3.client("sqs", config=_BOTO_CONFIG)
|
|
|
|
|
return sqs
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _delivery_log(event: dict, status_code, latency_ms: int, outcome: str) -> str:
|
|
|
|
|
"""One structured record per delivery attempt (Logs-Insights-queryable)."""
|
|
|
|
|
return json.dumps(
|
|
|
|
|
{
|
|
|
|
|
"event": "shoc_delivery",
|
|
|
|
|
"delivery_id": event["delivery_id"],
|
|
|
|
|
"event_type": event["event_type"],
|
|
|
|
|
"work_order_id": event["data"].get("work_order_id"),
|
|
|
|
|
"status_code": status_code,
|
|
|
|
|
"latency_ms": latency_ms,
|
|
|
|
|
"outcome": outcome,
|
|
|
|
|
}
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _park_rejected(event: dict, status_code: int):
|
|
|
|
|
"""Send the full envelope to the rejected queue for operator replay."""
|
|
|
|
|
_get_sqs().send_message(
|
|
|
|
|
QueueUrl=REJECTED_QUEUE_URL,
|
|
|
|
|
MessageBody=json.dumps({"envelope": event, "response_status": status_code}),
|
|
|
|
|
)
|
|
|
|
|
logger.warning(
|
|
|
|
|
json.dumps(
|
|
|
|
|
{
|
|
|
|
|
"event": "shoc_delivery_rejected",
|
|
|
|
|
"delivery_id": event["delivery_id"],
|
|
|
|
|
"event_type": event["event_type"],
|
|
|
|
|
"work_order_id": event["data"].get("work_order_id"),
|
|
|
|
|
"response_status": status_code,
|
|
|
|
|
}
|
|
|
|
|
)
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _batch_summary(records: int, emitted: int, skipped: int) -> str:
|
|
|
|
|
"""Per-invocation emit/skip tally.
|
|
|
|
|
|
2026-08-04 19:23:50 -04:00
|
|
|
Skips (REMOVE, echo guard, blank-text comments, comment MODIFYs,
|
|
|
|
|
unmappable records) return no delivery and would otherwise be invisible:
|
|
|
|
|
a systemic classification break -- e.g. a table rename desyncing the
|
|
|
|
|
eventSourceARN parse -- would drop 100%% of events while every batch
|
|
|
|
|
still reports success. Logging the ratio makes that queryable in Logs
|
|
|
|
|
Insights and alarmable. Blank-text comment skips also emit a distinct
|
|
|
|
|
``shoc_skipped_empty_comment`` line so their volume stays observable
|
|
|
|
|
apart from other skip reasons.
|
feat(webhook): SHOC WO webhook emitter - dark-ship streams + HMAC secret/rotation (PR-2) (#137)
* docs(webhook): revise SHOC webhook contract and plan for post-migration reality
Branch re-cut on main 2026-07-23 (old base carried stale PR #99 commits).
Contract Rev 2026-07-23:
- Producer account corrected: seahaven-prod (011934824531); mgmt frozen
- Reconciliation backstop is the new procurement read API, not SyncController
- wo_status "unknown" is real; SHOC must map it (checklist item added)
- write_origin forward-compat note for phase-2 write-back echo suppression
- SyncVendorReplies retirement flagged (dead table, no vendor_reply event)
Plan updates:
- Account gate: seahaven-prod only; never enable streams on mgmt tables
- Emitter ships DARK (ESMs enabled=False); activation is a deliberate flip
after the SHOC receiver passes shared HMAC vectors
- Post-refactor conventions: common.py helpers, bundle-consistency AST pins,
pytest.ini --cov additions, consolidated test roots
- Dedicated-CMK rationale, secret-ARN handooff step, consumer audit refreshed
(slack-bot decommissioned), enum golden test, write_origin skip-branch test
* feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation
Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC
call-and-be-called effort). Everything ships DARK: both DynamoDB event
source mappings deploy enabled=False; activation is a deliberate
one-line follow-up PR gated on the SHOC receiver passing the shared
HMAC test vectors.
- Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments
(in-place update, RETAIN + logical IDs untouched; no existing
consumers — verified live, neither table had a stream).
- workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope ->
HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard
ordering (parallelization 1, bisect off, retry until 24h age,
ReportBatchItemFailures); 429/5xx/timeout block the shard in order,
other 4xx park to workorder-shoc-emitter-rejected; ESM failures ->
workorder-shoc-emitter-failures (metadata; replay rebuilds from
DynamoDB). Echo guard skips write_origin=shoc-write-api.
- Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK
(alias workorder-ingest-shoc-webhook-kms); cross-account
GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly
arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy
DESTROY deliberately (machine-generated material; avoids the
fixed-name RETAIN-orphan deadlock).
- workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap,
64-hex keys, kid = UTC %Y-%m-%dT%H.
- Alarms (ALARM-only -> site-alerts): emitter errors/throttles/
duration + iterator-age (>=10 min) + failures/rejected queue
depth; rotator standard trio.
- scripts/replay_shoc_webhooks.py: dry-run-default operator replay
(rebuilds from tables, replay:true envelopes).
- Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared
with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay
signing pinned to identical vectors); bundle-consistency AST pins
for both new bundles.
- README: WO stack + webhook feed section, alarm table, runbooks;
removed stale seahaven-slack-bot consumer references.
* fix(webhook): kms:ViaService pins, https-only delivery, cross-account principal CI pin
GPT-4.1 cross-family review of the policy surface (no BLOCK): FIX applied
to the cross-account shoc-backend-dev Decrypt statement and both Lambda
role KMS grants (the key is only ever used via Secrets Manager); its
invariant-enforcement QUESTION answered durably with
tests/test_cross_account_principal_pin.py (any new foreign IAM principal
in cdk/ fails CI). Scanner mediums fixed: delivery.py and the replay
script now refuse non-https URLs (urllib follows file:// and http://).
SQS metadata-action and dynamodb:ListStreams NITs skipped: standard CDK
grant shapes; ListStreams has no resource-level scoping. The 4 gitleaks
HIGHs on docs/shoc-webhook-test-vectors.json are deliberate non-secrets
(shared receiver-verification vectors) suppressed machine-level with
justification.
* harden(webhook): resolve /sh-security-review findings (1 confirmed medium + cheap fixes)
High-recall detector fan-out (injection/authz/secrets-crypto/iac-iam/logic)
+ proof-or-kill verifier. Gate PASSES: 1 confirmed medium, 0 confirmed
critical/high. Confirmed finding fixed; several unverified-but-cheap
hardenings applied since the emitter ships dark and activation is weeks out.
- CONFIRMED medium (confused deputy): the rotation Lambda's generated
invoke permission for secretsmanager.amazonaws.com carried no
SourceAccount/SourceArn, so any account's Secrets Manager could invoke
the rotator. Patched the generated CfnPermission in place (a second
permission would be additive, not restrictive) to pin account + this
secret ARN.
- delivery + replay: refuse to follow receiver 3xx redirects (no-redirect
opener) so live X-SH-* auth headers can't be forwarded to a
receiver-chosen Location and an http:// Location can't slip past the
https guard. Fixed the "unfollowed 3xx" comment that was factually wrong.
- delivery: classify 401/403 as retryable (invalidate key cache + retry in
order) instead of parking -- transient auth failures (rotation outran the
TTL cache, clock skew) are availability events, not contract bugs.
- envelope: build_event now genuinely total (guarded eventID /
ApproximateCreationDateTime subscripts) per its own never-raise contract.
- handler: catch-all so an unexpected per-record error (e.g. SQS park
failure) reports only that record instead of failing the whole batch
(which would re-deliver every earlier success for 24h); per-invocation
emit/skip batch summary so a systemic silent drop is queryable/alarmable.
- rotator: narrow the AWSCURRENT-read except to ResourceNotFound/JSONDecode
(transient SM/KMS errors re-raise so the overlap key isn't silently
dropped); kid uniqueness checked against ALL retained kids with a random
suffix on collision (never reissue a kid for a different secret).
- contract: skeleton-upsert required on ANY unknown work_order_id (not just
comment-before-create) + monotonicity guard (ignore older updated_at), so
a parked created or an out-of-order replay can't corrupt receiver state.
Unverified/refuted findings left as-is with rationale: the two "high" logic
claims (whole-batch crash triggers, ordering violation) were refuted on
reachability (real stream records carry required fields; persistence writes
strings only; full-state idempotent upsert absorbs the ordering gap). Signed
kid/version binding (AUTHZ-002) declined: coordinated contract change, not
cheap, no exploit with one algorithm/key.
* fix(webhook): drop kid from rotator test_ok log (CodeQL clear-text-logging FP)
GHAS CodeQL flagged py/clear-text-logging-sensitive-data (high) at
_test_secret's success log because head["kid"] is subscripted from the
same parsed-secret dict that holds head["secret"] — the taint tracker
can't tell the non-secret key id from the secret. The secret value is
never logged. Rather than dismiss the alert (fragile; re-alerts on line
moves), remove the flow: kid is already logged at stage time in
_create_secret and version_id correlates the steps, so the test_ok log
keeps only event + version_id. Also hardens against a future edit that
swaps the logged field.
2026-07-24 18:12:20 -04:00
|
|
|
"""
|
|
|
|
|
return json.dumps(
|
|
|
|
|
{
|
|
|
|
|
"event": "shoc_batch_summary",
|
|
|
|
|
"records": records,
|
|
|
|
|
"emitted": emitted,
|
|
|
|
|
"skipped": skipped,
|
|
|
|
|
}
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
2026-08-04 19:23:50 -04:00
|
|
|
def _log_empty_comment_skip(record: dict) -> None:
|
|
|
|
|
"""Structured log for a blank-text comment_added skip (Logs Insights)."""
|
|
|
|
|
stream = record.get("dynamodb") or {}
|
|
|
|
|
# Attribute-value encoded NewImage; only the S forms are read for the
|
|
|
|
|
# log fields -- classification already decided this is a blank skip.
|
|
|
|
|
new_image = stream.get("NewImage") or {}
|
|
|
|
|
work_order_id = (new_image.get("work_order_id") or {}).get("S")
|
|
|
|
|
comment_id = (new_image.get("comment_id") or {}).get("S")
|
|
|
|
|
logger.info(
|
|
|
|
|
json.dumps(
|
|
|
|
|
{
|
|
|
|
|
"event": "shoc_skipped_empty_comment",
|
|
|
|
|
"work_order_id": work_order_id,
|
|
|
|
|
"comment_id": comment_id,
|
|
|
|
|
}
|
|
|
|
|
)
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
feat(webhook): SHOC WO webhook emitter - dark-ship streams + HMAC secret/rotation (PR-2) (#137)
* docs(webhook): revise SHOC webhook contract and plan for post-migration reality
Branch re-cut on main 2026-07-23 (old base carried stale PR #99 commits).
Contract Rev 2026-07-23:
- Producer account corrected: seahaven-prod (011934824531); mgmt frozen
- Reconciliation backstop is the new procurement read API, not SyncController
- wo_status "unknown" is real; SHOC must map it (checklist item added)
- write_origin forward-compat note for phase-2 write-back echo suppression
- SyncVendorReplies retirement flagged (dead table, no vendor_reply event)
Plan updates:
- Account gate: seahaven-prod only; never enable streams on mgmt tables
- Emitter ships DARK (ESMs enabled=False); activation is a deliberate flip
after the SHOC receiver passes shared HMAC vectors
- Post-refactor conventions: common.py helpers, bundle-consistency AST pins,
pytest.ini --cov additions, consolidated test roots
- Dedicated-CMK rationale, secret-ARN handooff step, consumer audit refreshed
(slack-bot decommissioned), enum golden test, write_origin skip-branch test
* feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation
Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC
call-and-be-called effort). Everything ships DARK: both DynamoDB event
source mappings deploy enabled=False; activation is a deliberate
one-line follow-up PR gated on the SHOC receiver passing the shared
HMAC test vectors.
- Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments
(in-place update, RETAIN + logical IDs untouched; no existing
consumers — verified live, neither table had a stream).
- workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope ->
HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard
ordering (parallelization 1, bisect off, retry until 24h age,
ReportBatchItemFailures); 429/5xx/timeout block the shard in order,
other 4xx park to workorder-shoc-emitter-rejected; ESM failures ->
workorder-shoc-emitter-failures (metadata; replay rebuilds from
DynamoDB). Echo guard skips write_origin=shoc-write-api.
- Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK
(alias workorder-ingest-shoc-webhook-kms); cross-account
GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly
arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy
DESTROY deliberately (machine-generated material; avoids the
fixed-name RETAIN-orphan deadlock).
- workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap,
64-hex keys, kid = UTC %Y-%m-%dT%H.
- Alarms (ALARM-only -> site-alerts): emitter errors/throttles/
duration + iterator-age (>=10 min) + failures/rejected queue
depth; rotator standard trio.
- scripts/replay_shoc_webhooks.py: dry-run-default operator replay
(rebuilds from tables, replay:true envelopes).
- Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared
with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay
signing pinned to identical vectors); bundle-consistency AST pins
for both new bundles.
- README: WO stack + webhook feed section, alarm table, runbooks;
removed stale seahaven-slack-bot consumer references.
* fix(webhook): kms:ViaService pins, https-only delivery, cross-account principal CI pin
GPT-4.1 cross-family review of the policy surface (no BLOCK): FIX applied
to the cross-account shoc-backend-dev Decrypt statement and both Lambda
role KMS grants (the key is only ever used via Secrets Manager); its
invariant-enforcement QUESTION answered durably with
tests/test_cross_account_principal_pin.py (any new foreign IAM principal
in cdk/ fails CI). Scanner mediums fixed: delivery.py and the replay
script now refuse non-https URLs (urllib follows file:// and http://).
SQS metadata-action and dynamodb:ListStreams NITs skipped: standard CDK
grant shapes; ListStreams has no resource-level scoping. The 4 gitleaks
HIGHs on docs/shoc-webhook-test-vectors.json are deliberate non-secrets
(shared receiver-verification vectors) suppressed machine-level with
justification.
* harden(webhook): resolve /sh-security-review findings (1 confirmed medium + cheap fixes)
High-recall detector fan-out (injection/authz/secrets-crypto/iac-iam/logic)
+ proof-or-kill verifier. Gate PASSES: 1 confirmed medium, 0 confirmed
critical/high. Confirmed finding fixed; several unverified-but-cheap
hardenings applied since the emitter ships dark and activation is weeks out.
- CONFIRMED medium (confused deputy): the rotation Lambda's generated
invoke permission for secretsmanager.amazonaws.com carried no
SourceAccount/SourceArn, so any account's Secrets Manager could invoke
the rotator. Patched the generated CfnPermission in place (a second
permission would be additive, not restrictive) to pin account + this
secret ARN.
- delivery + replay: refuse to follow receiver 3xx redirects (no-redirect
opener) so live X-SH-* auth headers can't be forwarded to a
receiver-chosen Location and an http:// Location can't slip past the
https guard. Fixed the "unfollowed 3xx" comment that was factually wrong.
- delivery: classify 401/403 as retryable (invalidate key cache + retry in
order) instead of parking -- transient auth failures (rotation outran the
TTL cache, clock skew) are availability events, not contract bugs.
- envelope: build_event now genuinely total (guarded eventID /
ApproximateCreationDateTime subscripts) per its own never-raise contract.
- handler: catch-all so an unexpected per-record error (e.g. SQS park
failure) reports only that record instead of failing the whole batch
(which would re-deliver every earlier success for 24h); per-invocation
emit/skip batch summary so a systemic silent drop is queryable/alarmable.
- rotator: narrow the AWSCURRENT-read except to ResourceNotFound/JSONDecode
(transient SM/KMS errors re-raise so the overlap key isn't silently
dropped); kid uniqueness checked against ALL retained kids with a random
suffix on collision (never reissue a kid for a different secret).
- contract: skeleton-upsert required on ANY unknown work_order_id (not just
comment-before-create) + monotonicity guard (ignore older updated_at), so
a parked created or an out-of-order replay can't corrupt receiver state.
Unverified/refuted findings left as-is with rationale: the two "high" logic
claims (whole-batch crash triggers, ordering violation) were refuted on
reachability (real stream records carry required fields; persistence writes
strings only; full-state idempotent upsert absorbs the ordering gap). Signed
kid/version binding (AUTHZ-002) declined: coordinated contract change, not
cheap, no exploit with one algorithm/key.
* fix(webhook): drop kid from rotator test_ok log (CodeQL clear-text-logging FP)
GHAS CodeQL flagged py/clear-text-logging-sensitive-data (high) at
_test_secret's success log because head["kid"] is subscripted from the
same parsed-secret dict that holds head["secret"] — the taint tracker
can't tell the non-secret key id from the secret. The secret value is
never logged. Rather than dismiss the alert (fragile; re-alerts on line
moves), remove the flow: kid is already logged at stage time in
_create_secret and version_id correlates the steps, so the test_ok log
keeps only event + version_id. Also hardens against a future edit that
swaps the logged field.
2026-07-24 18:12:20 -04:00
|
|
|
def handler(event, context):
|
|
|
|
|
"""Lambda entry point. Triggered by the two WO-table stream ESMs."""
|
|
|
|
|
records = event.get("Records", [])
|
|
|
|
|
emitted = 0
|
|
|
|
|
skipped = 0
|
|
|
|
|
for record in records:
|
|
|
|
|
webhook_event = envelope.build_event(record)
|
|
|
|
|
if webhook_event is None:
|
|
|
|
|
skipped += 1
|
2026-08-04 19:23:50 -04:00
|
|
|
if envelope.is_blank_comment_skip(record):
|
|
|
|
|
_log_empty_comment_skip(record)
|
feat(webhook): SHOC WO webhook emitter - dark-ship streams + HMAC secret/rotation (PR-2) (#137)
* docs(webhook): revise SHOC webhook contract and plan for post-migration reality
Branch re-cut on main 2026-07-23 (old base carried stale PR #99 commits).
Contract Rev 2026-07-23:
- Producer account corrected: seahaven-prod (011934824531); mgmt frozen
- Reconciliation backstop is the new procurement read API, not SyncController
- wo_status "unknown" is real; SHOC must map it (checklist item added)
- write_origin forward-compat note for phase-2 write-back echo suppression
- SyncVendorReplies retirement flagged (dead table, no vendor_reply event)
Plan updates:
- Account gate: seahaven-prod only; never enable streams on mgmt tables
- Emitter ships DARK (ESMs enabled=False); activation is a deliberate flip
after the SHOC receiver passes shared HMAC vectors
- Post-refactor conventions: common.py helpers, bundle-consistency AST pins,
pytest.ini --cov additions, consolidated test roots
- Dedicated-CMK rationale, secret-ARN handooff step, consumer audit refreshed
(slack-bot decommissioned), enum golden test, write_origin skip-branch test
* feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation
Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC
call-and-be-called effort). Everything ships DARK: both DynamoDB event
source mappings deploy enabled=False; activation is a deliberate
one-line follow-up PR gated on the SHOC receiver passing the shared
HMAC test vectors.
- Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments
(in-place update, RETAIN + logical IDs untouched; no existing
consumers — verified live, neither table had a stream).
- workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope ->
HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard
ordering (parallelization 1, bisect off, retry until 24h age,
ReportBatchItemFailures); 429/5xx/timeout block the shard in order,
other 4xx park to workorder-shoc-emitter-rejected; ESM failures ->
workorder-shoc-emitter-failures (metadata; replay rebuilds from
DynamoDB). Echo guard skips write_origin=shoc-write-api.
- Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK
(alias workorder-ingest-shoc-webhook-kms); cross-account
GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly
arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy
DESTROY deliberately (machine-generated material; avoids the
fixed-name RETAIN-orphan deadlock).
- workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap,
64-hex keys, kid = UTC %Y-%m-%dT%H.
- Alarms (ALARM-only -> site-alerts): emitter errors/throttles/
duration + iterator-age (>=10 min) + failures/rejected queue
depth; rotator standard trio.
- scripts/replay_shoc_webhooks.py: dry-run-default operator replay
(rebuilds from tables, replay:true envelopes).
- Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared
with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay
signing pinned to identical vectors); bundle-consistency AST pins
for both new bundles.
- README: WO stack + webhook feed section, alarm table, runbooks;
removed stale seahaven-slack-bot consumer references.
* fix(webhook): kms:ViaService pins, https-only delivery, cross-account principal CI pin
GPT-4.1 cross-family review of the policy surface (no BLOCK): FIX applied
to the cross-account shoc-backend-dev Decrypt statement and both Lambda
role KMS grants (the key is only ever used via Secrets Manager); its
invariant-enforcement QUESTION answered durably with
tests/test_cross_account_principal_pin.py (any new foreign IAM principal
in cdk/ fails CI). Scanner mediums fixed: delivery.py and the replay
script now refuse non-https URLs (urllib follows file:// and http://).
SQS metadata-action and dynamodb:ListStreams NITs skipped: standard CDK
grant shapes; ListStreams has no resource-level scoping. The 4 gitleaks
HIGHs on docs/shoc-webhook-test-vectors.json are deliberate non-secrets
(shared receiver-verification vectors) suppressed machine-level with
justification.
* harden(webhook): resolve /sh-security-review findings (1 confirmed medium + cheap fixes)
High-recall detector fan-out (injection/authz/secrets-crypto/iac-iam/logic)
+ proof-or-kill verifier. Gate PASSES: 1 confirmed medium, 0 confirmed
critical/high. Confirmed finding fixed; several unverified-but-cheap
hardenings applied since the emitter ships dark and activation is weeks out.
- CONFIRMED medium (confused deputy): the rotation Lambda's generated
invoke permission for secretsmanager.amazonaws.com carried no
SourceAccount/SourceArn, so any account's Secrets Manager could invoke
the rotator. Patched the generated CfnPermission in place (a second
permission would be additive, not restrictive) to pin account + this
secret ARN.
- delivery + replay: refuse to follow receiver 3xx redirects (no-redirect
opener) so live X-SH-* auth headers can't be forwarded to a
receiver-chosen Location and an http:// Location can't slip past the
https guard. Fixed the "unfollowed 3xx" comment that was factually wrong.
- delivery: classify 401/403 as retryable (invalidate key cache + retry in
order) instead of parking -- transient auth failures (rotation outran the
TTL cache, clock skew) are availability events, not contract bugs.
- envelope: build_event now genuinely total (guarded eventID /
ApproximateCreationDateTime subscripts) per its own never-raise contract.
- handler: catch-all so an unexpected per-record error (e.g. SQS park
failure) reports only that record instead of failing the whole batch
(which would re-deliver every earlier success for 24h); per-invocation
emit/skip batch summary so a systemic silent drop is queryable/alarmable.
- rotator: narrow the AWSCURRENT-read except to ResourceNotFound/JSONDecode
(transient SM/KMS errors re-raise so the overlap key isn't silently
dropped); kid uniqueness checked against ALL retained kids with a random
suffix on collision (never reissue a kid for a different secret).
- contract: skeleton-upsert required on ANY unknown work_order_id (not just
comment-before-create) + monotonicity guard (ignore older updated_at), so
a parked created or an out-of-order replay can't corrupt receiver state.
Unverified/refuted findings left as-is with rationale: the two "high" logic
claims (whole-batch crash triggers, ordering violation) were refuted on
reachability (real stream records carry required fields; persistence writes
strings only; full-state idempotent upsert absorbs the ordering gap). Signed
kid/version binding (AUTHZ-002) declined: coordinated contract change, not
cheap, no exploit with one algorithm/key.
* fix(webhook): drop kid from rotator test_ok log (CodeQL clear-text-logging FP)
GHAS CodeQL flagged py/clear-text-logging-sensitive-data (high) at
_test_secret's success log because head["kid"] is subscripted from the
same parsed-secret dict that holds head["secret"] — the taint tracker
can't tell the non-secret key id from the secret. The secret value is
never logged. Rather than dismiss the alert (fragile; re-alerts on line
moves), remove the flow: kid is already logged at stage time in
_create_secret and version_id correlates the steps, so the test_ok log
keeps only event + version_id. Also hardens against a future edit that
swaps the logged field.
2026-07-24 18:12:20 -04:00
|
|
|
continue
|
|
|
|
|
# Capture the sequence number BEFORE any delivery attempt so the
|
|
|
|
|
# failure paths below can never KeyError on the subscript (Open SWE
|
|
|
|
|
# #9/#26 -- the catch-all exists to keep an exception from re-delivering
|
|
|
|
|
# the whole batch, so it must not itself raise). SequenceNumber is
|
|
|
|
|
# present on every real DynamoDB stream record; a record lacking it is
|
|
|
|
|
# unmappable-to-a-checkpoint, so log and skip rather than crash.
|
|
|
|
|
seq = record.get("dynamodb", {}).get("SequenceNumber")
|
|
|
|
|
if seq is None:
|
|
|
|
|
logger.warning(
|
|
|
|
|
json.dumps(
|
|
|
|
|
{
|
|
|
|
|
"event": "shoc_missing_sequence_number",
|
|
|
|
|
"delivery_id": webhook_event["delivery_id"],
|
|
|
|
|
}
|
|
|
|
|
)
|
|
|
|
|
)
|
|
|
|
|
skipped += 1
|
|
|
|
|
continue
|
|
|
|
|
emitted += 1
|
|
|
|
|
start = time.monotonic()
|
|
|
|
|
try:
|
|
|
|
|
outcome, status_code = delivery.deliver(webhook_event)
|
|
|
|
|
latency_ms = int((time.monotonic() - start) * 1000)
|
|
|
|
|
logger.info(_delivery_log(webhook_event, status_code, latency_ms, outcome))
|
|
|
|
|
if outcome == "rejected":
|
|
|
|
|
_park_rejected(webhook_event, status_code)
|
|
|
|
|
except delivery.RetryableDeliveryError as exc:
|
|
|
|
|
latency_ms = int((time.monotonic() - start) * 1000)
|
|
|
|
|
logger.warning(
|
|
|
|
|
_delivery_log(webhook_event, exc.status_code, latency_ms, "retryable")
|
|
|
|
|
)
|
|
|
|
|
logger.info(_batch_summary(len(records), emitted, skipped))
|
|
|
|
|
# Stop here: reporting this record's sequence number makes the ESM
|
|
|
|
|
# retry from it in order; earlier successes are not re-delivered.
|
|
|
|
|
return {"batchItemFailures": [{"itemIdentifier": seq}]}
|
|
|
|
|
except Exception:
|
|
|
|
|
# Catch-all so an unexpected error (e.g. an SQS park failure)
|
|
|
|
|
# cannot escape and fail the WHOLE invocation -- that would make
|
|
|
|
|
# the ESM re-deliver every earlier success in the batch for up to
|
|
|
|
|
# 24h. Report only THIS record so the ESM retries from it in order.
|
|
|
|
|
logger.exception(
|
|
|
|
|
json.dumps(
|
|
|
|
|
{
|
|
|
|
|
"event": "shoc_delivery_error",
|
|
|
|
|
"delivery_id": webhook_event["delivery_id"],
|
|
|
|
|
"event_type": webhook_event["event_type"],
|
|
|
|
|
}
|
|
|
|
|
)
|
|
|
|
|
)
|
|
|
|
|
logger.info(_batch_summary(len(records), emitted, skipped))
|
|
|
|
|
return {"batchItemFailures": [{"itemIdentifier": seq}]}
|
|
|
|
|
logger.info(_batch_summary(len(records), emitted, skipped))
|
|
|
|
|
return {"batchItemFailures": []}
|