procurement-ingest/docs/runbook-dlq-recovery.md
Adam Moussa 75fe91c198
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)

Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').

Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.

Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.

cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.

The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.

* Reject --all combined with --prefix/--since in reprocess.py

--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00

5.3 KiB

DLQ Recovery Runbook — Email-Processor Dead-Letter Queues

Operational procedure for draining an email-processor dead-letter queue (DLQ) after a failed async parse. There is no console redrive-to-source for these queues; recovery is a manual, targeted re-invoke via scripts/reprocess.py.

Scope / mechanism

Both email processors set dead_letter_queue= on the lambda_.Function construct. This is the legacy per-function Lambda DeadLetterConfig (asynchronous-invocation DLQ), not an EventInvokeConfig on-failure Destination — aws lambda get-function-event-invoke-config returns ResourceNotFoundException for both functions (no destination config exists). A failed async invocation lands on the DLQ only after Lambda exhausts its automatic retries.

There is no console redrive-to-source: the SQS console's "redrive to source" applies only to SQS-to-SQS DLQ relationships, not to a Lambda DeadLetterConfig target. Recovery is manual, via targeted re-invoke.

Resource inventory (acct 328440206208, us-east-1)

Pipeline Function DLQ queue name Raw-email bucket DLQ alarm
PO po-email-processor po-ingest-EmailProcessorDlqA753DED5-az8LUZE3ubtz po-ingest-emails-328440206208 po-email-processor-dlq-messages
WO workorder-email-processor WorkorderIngestStack-EmailProcessorDlqA753DED5-Q8H555LrqSU1 workorder-ingest-emails-328440206208 workorder-email-processor-dlq-messages

DLQ URLs are https://sqs.us-east-1.amazonaws.com/328440206208/<queue-name>. Both DLQs: 14-day retention, SSE, TLS-enforced, VisibilityTimeout 30s.

Recovery procedure (no redrive — receive → extract key → targeted re-invoke → verify → purge)

  1. Trigger. The <fn>-dlq-messages alarm fires (ApproximateNumberOfMessagesVisible Maximum, 5 min, > 0, eval 1).

  2. RECEIVE the message (do not purge yet):

    aws sqs receive-message \
      --queue-url <DLQ_URL> \
      --max-number-of-messages 1 \
      --visibility-timeout 120 \
      --wait-time-seconds 5
    

    Capture the ReceiptHandle from the response.

  3. EXTRACT the S3 key. The DeadLetterConfig message Body is the original async invocation payload — the S3 event JSON. Read Records[0].s3.bucket.name and Records[0].s3.object.key from the Body. This is the raw object key that reprocess/S3 emitted (no URL-decoding applied).

  4. RE-INVOKE (targeted; dry-run first). Confirm the key with a dry-run, then execute:

    # PO queue:
    python scripts/reprocess.py --pipeline po --key '<key>'            # dry-run
    python scripts/reprocess.py --pipeline po --key '<key>' --execute  # re-invoke
    
    # WO queue:
    python scripts/reprocess.py --pipeline wo --key '<key>' --execute
    

    This re-invokes the same function with the same raw-key synthetic S3 event (a single object — not --all).

  5. VERIFY the write. Confirm the downstream effect landed before proceeding: the DynamoDB item exists / was updated (purchase-orders for PO, WorkOrders for WO), and the function's log group shows a clean parse (no new error, no new DLQ message). Do not proceed until verified.

  6. PURGE the one message. Delete only the processed message by its ReceiptHandle:

    aws sqs delete-message --queue-url <DLQ_URL> --receipt-handle '<ReceiptHandle>'
    

    Do not purge-queue — that would drop unexamined breadcrumbs. AWS access is otherwise read-only; delete-message on a DLQ you are actively draining is the one write this runbook performs.

Recovery windows

  • DLQ breadcrumb retention: 14 days (MessageRetentionPeriod=1209600s, confirmed live on both queues). After 14 days the breadcrumb is gone.
  • Raw-email S3 retention: 90 days. Both buckets have RemovalPolicy.RETAIN, but a single Enabled lifecycle rule (Expiration.Days=90, Filter.Prefix='') expires objects bucket-wide. The lifecycle rule OVERRIDES the RETAIN policy — S3 is the real replay floor: a raw email is gone at ~90 days regardless of the table RETAIN policy.

Because 14 days (DLQ) < 90 days (S3), any object referenced by a live DLQ breadcrumb is always still in S3, so DLQ replay within its 14-day window is never blocked by S3 expiry. The 90-day floor binds only for replays reconstructed from other sources (e.g. logs) after the breadcrumb has expired.

Intentional non-DLQ drops (do NOT hunt for these in the DLQ)

Two drop classes never produce a DLQ message because they are fail-closed skips, not errors — the handler returns normally (no raise, no retry, no DLQ message):

  • Sender-auth rejections. Logged as a structured sender_auth_rejected warning and skipped; covered by the <fn>-sender-auth-rejected alarm (log-metric filter), not the DLQ.
  • ai_fallback_rejected drops. AI-fallback output that failed the fail-closed validate_ai_fallback() gate, emitted as a ParseMethod=ai_fallback_rejected EMF datapoint and dropped without a DynamoDB write; covered by the <fn>-ai-fallback-rejected alarm, not the DLQ.

If mail is missing but the DLQ is empty, check those two alarms / log filters — the email was intentionally rejected. Re-invoking it via reprocess will just be rejected again; fix the sender-auth config or the upstream email, not the DLQ.