procurement-ingest/docs/runbook-dlq-recovery.md
Adam Moussa 073201f633
Migrate to seahaven-prod: deploy role, backfill tooling, account-portability fixes (#125)
* feat(migration): prepare stacks and tooling for the seahaven-prod account move

Phase 1 of the mgmt (328440206208) -> seahaven-prod (011934824531)
migration. No behavior change in-account; everything here is additive or
account-portability hygiene:

- infra/deploy-role/: reviewed OIDC deploy-role artifacts for prod
  (trust main-only, cdk-hnb659fds-* AssumeRole, smoke-invoke-lambda scoped
  to exactly the two email-processor fn ARNs). Codifies the previously
  out-of-band smoke-invoke grant.
- Table resource policies: make_slack_bot_read_policy in cdk/common.py,
  applied to purchase-orders, verified-sites, WorkOrders,
  WorkOrderComments (NOT pending-site-review; no bot consumer). Grants the
  mgmt-resident seahaven-slack-bot roles read-only cross-account access
  post-move (bot-side identity grants land in the slack-bot repo).
- scripts/migrate_tables.py: dry-run-default backfill tool implementing
  the plan's per-table semantics (superset overwrite, ingested_at cutoff
  for WorkOrderComments, backup-gated truncate-and-load for the two
  site tables) plus a verify subcommand (count parity, spot checks,
  sticky-Cancelled drift check).
- tests/test_resource_policy_helper.py: statement-shape unit tests +
  static pins that exactly the four bot-read tables carry the policy.
- Account-literal fixes: account-agnostic fixture bucket in
  test_reprocess_contract; runbook/README/po-template-parser account
  references updated to prod with historical mgmt notes; README gains the
  account-prerequisites list (imported-by-name dependencies).

deploy.yaml is deliberately unchanged (push-to-main auto-deploy kept).
Merge is held until migration Phase 0 completes; flipping the
AWS_DEPLOY_ROLE_ARN repo secret and merging this PR IS the first prod
deploy.

* fix(migration): verify backup AVAILABLE pre-truncate; document wildcard risk acceptance (cross-review FIX/NIT)

* refactor(migration): drop cross-account read grants (slack-bot decommissioned); harden backfill + deploy role

seahaven-slack-bot was decommissioned 2026-07-23 (stack DELETE_IN_PROGRESS,
consumer Lambdas gone); its successor sh-mcp is undeployed and uses
same-account DynamoDB access. So no live consumer reads these tables
cross-account. Per Adam's call, drop the cross-account grants entirely and
re-add correctly-scoped ones if/when sh-mcp deploys to a different account.

- Remove the four table resource policies + make_slack_bot_read_policy helper
  + its constants (cdk/common.py, po_stack.py, wo_stack.py) and the helper's
  unit test. Both stacks synth with zero table ResourcePolicy.
- scripts/migrate_tables.py hardening (fixes from the sh-security-review
  fan-out on the destructive backfill tool):
  * validate --cutoff strictly (parse ISO-8601, require aware UTC, re-emit
    canonical second-precision form) so a malformed cutoff can't silently
    copy dual-window rows or drop history;
  * reject `copy --all` up front (must run tables individually, in order,
    with the stream-drain wait) instead of writing three tables then erroring;
  * truncate backup gate now also checks recency (<1h) and TableId, not just
    status+name;
  * verify requires --cutoff whenever a cutoff table is in scope (else it
    false-flags dual-window rows as MISSING);
  * sticky-cancel is now PREVENTED copy-side (a non-Cancelled source item
    never overwrites a dest-Cancelled PO), and the verify comment no longer
    overstates what its source-side scan covers;
  * spot-check all modes (truncate_load keys are verbatim, so key-existence
    is sound there too).
- Deploy role: scope cloudformation:DescribeStacks to this repo's stacks +
  CDKToolkit (was Resource:*, disclosed all tenant stacks in the shared prod
  account); add a drift check warning on unexpected role policies and drop the
  dead SMOKE_POLICY_NAME var; document the shared-account bootstrap-role
  accepted risk in the deploy-role README.

* docs(deploy-role): fold in cross-review NITs (DescribeStacks maintenance note, warn-only drift rationale)

* ci: update workflow to use new workflow tag (ruff versioning fix)

* fix(migration): address Open SWE review findings on migrate_tables.py

- Validate the truncate backup on dry-run as well as --execute so a
  missing/stale/wrong-incarnation --backup-arn surfaces on the rehearsal
  run (finding f_24a48b8900).
- Assert configured keys match the live key schema of both tables before
  any key projection, turning config/schema drift into a descriptive
  abort instead of a mid-backfill KeyError (finding f_cb6b5a6c59).
- Clarify why key-existence spot-checks are sound for WorkOrderComments:
  the copy Puts source items verbatim and the sample uses the same
  cutoff filter, so per-account comment_id divergence never enters the
  check (finding f_390b7d6c3b is a false positive; comment hardened).
2026-07-23 17:08:47 -04:00

5.6 KiB

DLQ Recovery Runbook — Email-Processor Dead-Letter Queues

Operational procedure for draining an email-processor dead-letter queue (DLQ) after a failed async parse. There is no console redrive-to-source for these queues; recovery is a manual, targeted re-invoke via scripts/reprocess.py.

Scope / mechanism

Both email processors set dead_letter_queue= on the lambda_.Function construct. This is the legacy per-function Lambda DeadLetterConfig (asynchronous-invocation DLQ), not an EventInvokeConfig on-failure Destination — aws lambda get-function-event-invoke-config returns ResourceNotFoundException for both functions (no destination config exists). A failed async invocation lands on the DLQ only after Lambda exhausts its automatic retries.

There is no console redrive-to-source: the SQS console's "redrive to source" applies only to SQS-to-SQS DLQ relationships, not to a Lambda DeadLetterConfig target. Recovery is manual, via targeted re-invoke.

Resource inventory (acct 011934824531 seahaven-prod, us-east-1)

Pipeline Function DLQ queue name Raw-email bucket DLQ alarm
PO po-email-processor CDK-generated; fill from stack resources after the first prod deploy po-ingest-emails-011934824531 po-email-processor-dlq-messages
WO workorder-email-processor CDK-generated; fill from stack resources after the first prod deploy workorder-ingest-emails-011934824531 workorder-email-processor-dlq-messages

DLQ URLs are https://sqs.us-east-1.amazonaws.com/011934824531/<queue-name>. (Until mgmt decommission completes, the pre-migration mgmt-account queues po-ingest-EmailProcessorDlqA753DED5-az8LUZE3ubtz and WorkorderIngestStack-EmailProcessorDlqA753DED5-Q8H555LrqSU1 still exist in 328440206208 with any pre-cutover dead letters.) Both DLQs: 14-day retention, SSE, TLS-enforced, VisibilityTimeout 30s.

Recovery procedure (no redrive — receive → extract key → targeted re-invoke → verify → purge)

  1. Trigger. The <fn>-dlq-messages alarm fires (ApproximateNumberOfMessagesVisible Maximum, 5 min, > 0, eval 1).

  2. RECEIVE the message (do not purge yet):

    aws sqs receive-message \
      --queue-url <DLQ_URL> \
      --max-number-of-messages 1 \
      --visibility-timeout 120 \
      --wait-time-seconds 5
    

    Capture the ReceiptHandle from the response.

  3. EXTRACT the S3 key. The DeadLetterConfig message Body is the original async invocation payload — the S3 event JSON. Read Records[0].s3.bucket.name and Records[0].s3.object.key from the Body. This is the raw object key that reprocess/S3 emitted (no URL-decoding applied).

  4. RE-INVOKE (targeted; dry-run first). Confirm the key with a dry-run, then execute:

    # PO queue:
    python scripts/reprocess.py --pipeline po --key '<key>'            # dry-run
    python scripts/reprocess.py --pipeline po --key '<key>' --execute  # re-invoke
    
    # WO queue:
    python scripts/reprocess.py --pipeline wo --key '<key>' --execute
    

    This re-invokes the same function with the same raw-key synthetic S3 event (a single object — not --all).

  5. VERIFY the write. Confirm the downstream effect landed before proceeding: the DynamoDB item exists / was updated (purchase-orders for PO, WorkOrders for WO), and the function's log group shows a clean parse (no new error, no new DLQ message). Do not proceed until verified.

  6. PURGE the one message. Delete only the processed message by its ReceiptHandle:

    aws sqs delete-message --queue-url <DLQ_URL> --receipt-handle '<ReceiptHandle>'
    

    Do not purge-queue — that would drop unexamined breadcrumbs. AWS access is otherwise read-only; delete-message on a DLQ you are actively draining is the one write this runbook performs.

Recovery windows

  • DLQ breadcrumb retention: 14 days (MessageRetentionPeriod=1209600s, confirmed live on both queues). After 14 days the breadcrumb is gone.
  • Raw-email S3 retention: 90 days. Both buckets have RemovalPolicy.RETAIN, but a single Enabled lifecycle rule (Expiration.Days=90, Filter.Prefix='') expires objects bucket-wide. The lifecycle rule OVERRIDES the RETAIN policy — S3 is the real replay floor: a raw email is gone at ~90 days regardless of the table RETAIN policy.

Because 14 days (DLQ) < 90 days (S3), any object referenced by a live DLQ breadcrumb is always still in S3, so DLQ replay within its 14-day window is never blocked by S3 expiry. The 90-day floor binds only for replays reconstructed from other sources (e.g. logs) after the breadcrumb has expired.

Intentional non-DLQ drops (do NOT hunt for these in the DLQ)

Two drop classes never produce a DLQ message because they are fail-closed skips, not errors — the handler returns normally (no raise, no retry, no DLQ message):

  • Sender-auth rejections. Logged as a structured sender_auth_rejected warning and skipped; covered by the <fn>-sender-auth-rejected alarm (log-metric filter), not the DLQ.
  • ai_fallback_rejected drops. AI-fallback output that failed the fail-closed validate_ai_fallback() gate, emitted as a ParseMethod=ai_fallback_rejected EMF datapoint and dropped without a DynamoDB write; covered by the <fn>-ai-fallback-rejected alarm, not the DLQ.

If mail is missing but the DLQ is empty, check those two alarms / log filters — the email was intentionally rejected. Re-invoking it via reprocess will just be rejected again; fix the sender-auth config or the upstream email, not the DLQ.