Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
# DLQ Recovery Runbook — Email-Processor Dead-Letter Queues
|
|
|
|
|
|
|
|
|
|
Operational procedure for draining an email-processor dead-letter queue (DLQ)
|
|
|
|
|
after a failed async parse. There is **no console redrive-to-source** for these
|
|
|
|
|
queues; recovery is a manual, targeted re-invoke via `scripts/reprocess.py`.
|
|
|
|
|
|
|
|
|
|
## Scope / mechanism
|
|
|
|
|
|
|
|
|
|
Both email processors set `dead_letter_queue=` on the `lambda_.Function`
|
|
|
|
|
construct. This is the legacy per-function **Lambda `DeadLetterConfig`**
|
|
|
|
|
(asynchronous-invocation DLQ), **not** an EventInvokeConfig on-failure
|
|
|
|
|
Destination — `aws lambda get-function-event-invoke-config` returns
|
|
|
|
|
`ResourceNotFoundException` for both functions (no destination config exists). A
|
|
|
|
|
failed async invocation lands on the DLQ only **after** Lambda exhausts its
|
|
|
|
|
automatic retries.
|
|
|
|
|
|
|
|
|
|
There is **no console redrive-to-source**: the SQS console's "redrive to source"
|
|
|
|
|
applies only to SQS-to-SQS DLQ relationships, not to a Lambda `DeadLetterConfig`
|
|
|
|
|
target. Recovery is manual, via targeted re-invoke.
|
|
|
|
|
|
Migrate to seahaven-prod: deploy role, backfill tooling, account-portability fixes (#125)
* feat(migration): prepare stacks and tooling for the seahaven-prod account move
Phase 1 of the mgmt (328440206208) -> seahaven-prod (011934824531)
migration. No behavior change in-account; everything here is additive or
account-portability hygiene:
- infra/deploy-role/: reviewed OIDC deploy-role artifacts for prod
(trust main-only, cdk-hnb659fds-* AssumeRole, smoke-invoke-lambda scoped
to exactly the two email-processor fn ARNs). Codifies the previously
out-of-band smoke-invoke grant.
- Table resource policies: make_slack_bot_read_policy in cdk/common.py,
applied to purchase-orders, verified-sites, WorkOrders,
WorkOrderComments (NOT pending-site-review; no bot consumer). Grants the
mgmt-resident seahaven-slack-bot roles read-only cross-account access
post-move (bot-side identity grants land in the slack-bot repo).
- scripts/migrate_tables.py: dry-run-default backfill tool implementing
the plan's per-table semantics (superset overwrite, ingested_at cutoff
for WorkOrderComments, backup-gated truncate-and-load for the two
site tables) plus a verify subcommand (count parity, spot checks,
sticky-Cancelled drift check).
- tests/test_resource_policy_helper.py: statement-shape unit tests +
static pins that exactly the four bot-read tables carry the policy.
- Account-literal fixes: account-agnostic fixture bucket in
test_reprocess_contract; runbook/README/po-template-parser account
references updated to prod with historical mgmt notes; README gains the
account-prerequisites list (imported-by-name dependencies).
deploy.yaml is deliberately unchanged (push-to-main auto-deploy kept).
Merge is held until migration Phase 0 completes; flipping the
AWS_DEPLOY_ROLE_ARN repo secret and merging this PR IS the first prod
deploy.
* fix(migration): verify backup AVAILABLE pre-truncate; document wildcard risk acceptance (cross-review FIX/NIT)
* refactor(migration): drop cross-account read grants (slack-bot decommissioned); harden backfill + deploy role
seahaven-slack-bot was decommissioned 2026-07-23 (stack DELETE_IN_PROGRESS,
consumer Lambdas gone); its successor sh-mcp is undeployed and uses
same-account DynamoDB access. So no live consumer reads these tables
cross-account. Per Adam's call, drop the cross-account grants entirely and
re-add correctly-scoped ones if/when sh-mcp deploys to a different account.
- Remove the four table resource policies + make_slack_bot_read_policy helper
+ its constants (cdk/common.py, po_stack.py, wo_stack.py) and the helper's
unit test. Both stacks synth with zero table ResourcePolicy.
- scripts/migrate_tables.py hardening (fixes from the sh-security-review
fan-out on the destructive backfill tool):
* validate --cutoff strictly (parse ISO-8601, require aware UTC, re-emit
canonical second-precision form) so a malformed cutoff can't silently
copy dual-window rows or drop history;
* reject `copy --all` up front (must run tables individually, in order,
with the stream-drain wait) instead of writing three tables then erroring;
* truncate backup gate now also checks recency (<1h) and TableId, not just
status+name;
* verify requires --cutoff whenever a cutoff table is in scope (else it
false-flags dual-window rows as MISSING);
* sticky-cancel is now PREVENTED copy-side (a non-Cancelled source item
never overwrites a dest-Cancelled PO), and the verify comment no longer
overstates what its source-side scan covers;
* spot-check all modes (truncate_load keys are verbatim, so key-existence
is sound there too).
- Deploy role: scope cloudformation:DescribeStacks to this repo's stacks +
CDKToolkit (was Resource:*, disclosed all tenant stacks in the shared prod
account); add a drift check warning on unexpected role policies and drop the
dead SMOKE_POLICY_NAME var; document the shared-account bootstrap-role
accepted risk in the deploy-role README.
* docs(deploy-role): fold in cross-review NITs (DescribeStacks maintenance note, warn-only drift rationale)
* ci: update workflow to use new workflow tag (ruff versioning fix)
* fix(migration): address Open SWE review findings on migrate_tables.py
- Validate the truncate backup on dry-run as well as --execute so a
missing/stale/wrong-incarnation --backup-arn surfaces on the rehearsal
run (finding f_24a48b8900).
- Assert configured keys match the live key schema of both tables before
any key projection, turning config/schema drift into a descriptive
abort instead of a mid-backfill KeyError (finding f_cb6b5a6c59).
- Clarify why key-existence spot-checks are sound for WorkOrderComments:
the copy Puts source items verbatim and the sample uses the same
cutoff filter, so per-account comment_id divergence never enters the
check (finding f_390b7d6c3b is a false positive; comment hardened).
2026-07-23 17:08:47 -04:00
|
|
|
## Resource inventory (acct 011934824531 seahaven-prod, us-east-1)
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
|
2026-08-13 17:39:06 -04:00
|
|
|
| Pipeline | Function | DLQ queue name | Raw-email bucket | DLQ alarms |
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
|---|---|---|---|---|
|
2026-08-13 17:39:06 -04:00
|
|
|
| PO | `po-email-processor` | `po-ingest-EmailProcessorDlqA753DED5-Mn33HvhsDPEu` | `po-ingest-emails-011934824531` | `po-email-processor-dlq-messages` |
|
|
|
|
|
| WO | `workorder-email-processor` | `WorkorderIngestStack-EmailProcessorDlqA753DED5-qCTHrsoEucas` | `workorder-ingest-emails-011934824531` | `workorder-email-processor-dlq-messages` (depth, visible `> 0`); `workorder-email-processor-dlq-age` (oldest message `>= 86400` s) |
|
Migrate to seahaven-prod: deploy role, backfill tooling, account-portability fixes (#125)
* feat(migration): prepare stacks and tooling for the seahaven-prod account move
Phase 1 of the mgmt (328440206208) -> seahaven-prod (011934824531)
migration. No behavior change in-account; everything here is additive or
account-portability hygiene:
- infra/deploy-role/: reviewed OIDC deploy-role artifacts for prod
(trust main-only, cdk-hnb659fds-* AssumeRole, smoke-invoke-lambda scoped
to exactly the two email-processor fn ARNs). Codifies the previously
out-of-band smoke-invoke grant.
- Table resource policies: make_slack_bot_read_policy in cdk/common.py,
applied to purchase-orders, verified-sites, WorkOrders,
WorkOrderComments (NOT pending-site-review; no bot consumer). Grants the
mgmt-resident seahaven-slack-bot roles read-only cross-account access
post-move (bot-side identity grants land in the slack-bot repo).
- scripts/migrate_tables.py: dry-run-default backfill tool implementing
the plan's per-table semantics (superset overwrite, ingested_at cutoff
for WorkOrderComments, backup-gated truncate-and-load for the two
site tables) plus a verify subcommand (count parity, spot checks,
sticky-Cancelled drift check).
- tests/test_resource_policy_helper.py: statement-shape unit tests +
static pins that exactly the four bot-read tables carry the policy.
- Account-literal fixes: account-agnostic fixture bucket in
test_reprocess_contract; runbook/README/po-template-parser account
references updated to prod with historical mgmt notes; README gains the
account-prerequisites list (imported-by-name dependencies).
deploy.yaml is deliberately unchanged (push-to-main auto-deploy kept).
Merge is held until migration Phase 0 completes; flipping the
AWS_DEPLOY_ROLE_ARN repo secret and merging this PR IS the first prod
deploy.
* fix(migration): verify backup AVAILABLE pre-truncate; document wildcard risk acceptance (cross-review FIX/NIT)
* refactor(migration): drop cross-account read grants (slack-bot decommissioned); harden backfill + deploy role
seahaven-slack-bot was decommissioned 2026-07-23 (stack DELETE_IN_PROGRESS,
consumer Lambdas gone); its successor sh-mcp is undeployed and uses
same-account DynamoDB access. So no live consumer reads these tables
cross-account. Per Adam's call, drop the cross-account grants entirely and
re-add correctly-scoped ones if/when sh-mcp deploys to a different account.
- Remove the four table resource policies + make_slack_bot_read_policy helper
+ its constants (cdk/common.py, po_stack.py, wo_stack.py) and the helper's
unit test. Both stacks synth with zero table ResourcePolicy.
- scripts/migrate_tables.py hardening (fixes from the sh-security-review
fan-out on the destructive backfill tool):
* validate --cutoff strictly (parse ISO-8601, require aware UTC, re-emit
canonical second-precision form) so a malformed cutoff can't silently
copy dual-window rows or drop history;
* reject `copy --all` up front (must run tables individually, in order,
with the stream-drain wait) instead of writing three tables then erroring;
* truncate backup gate now also checks recency (<1h) and TableId, not just
status+name;
* verify requires --cutoff whenever a cutoff table is in scope (else it
false-flags dual-window rows as MISSING);
* sticky-cancel is now PREVENTED copy-side (a non-Cancelled source item
never overwrites a dest-Cancelled PO), and the verify comment no longer
overstates what its source-side scan covers;
* spot-check all modes (truncate_load keys are verbatim, so key-existence
is sound there too).
- Deploy role: scope cloudformation:DescribeStacks to this repo's stacks +
CDKToolkit (was Resource:*, disclosed all tenant stacks in the shared prod
account); add a drift check warning on unexpected role policies and drop the
dead SMOKE_POLICY_NAME var; document the shared-account bootstrap-role
accepted risk in the deploy-role README.
* docs(deploy-role): fold in cross-review NITs (DescribeStacks maintenance note, warn-only drift rationale)
* ci: update workflow to use new workflow tag (ruff versioning fix)
* fix(migration): address Open SWE review findings on migrate_tables.py
- Validate the truncate backup on dry-run as well as --execute so a
missing/stale/wrong-incarnation --backup-arn surfaces on the rehearsal
run (finding f_24a48b8900).
- Assert configured keys match the live key schema of both tables before
any key projection, turning config/schema drift into a descriptive
abort instead of a mid-backfill KeyError (finding f_cb6b5a6c59).
- Clarify why key-existence spot-checks are sound for WorkOrderComments:
the copy Puts source items verbatim and the sample uses the same
cutoff filter, so per-account comment_id divergence never enters the
check (finding f_390b7d6c3b is a false positive; comment hardened).
2026-07-23 17:08:47 -04:00
|
|
|
|
|
|
|
|
DLQ URLs are `https://sqs.us-east-1.amazonaws.com/011934824531/<queue-name>`.
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
Both DLQs: 14-day retention, SSE, TLS-enforced, `VisibilityTimeout` 30s.
|
2026-08-04 20:24:38 -04:00
|
|
|
(Mgmt-account DLQs were removed with the PLAT-67 stack teardown on 2026-08-05.)
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
|
|
|
|
|
## Recovery procedure (no redrive — receive → extract key → targeted re-invoke → verify → purge)
|
|
|
|
|
|
|
|
|
|
1. **Trigger.** The `<fn>-dlq-messages` alarm fires
|
|
|
|
|
(`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1).
|
2026-08-13 17:39:06 -04:00
|
|
|
The WO queue also has `workorder-email-processor-dlq-age`, which fires when
|
|
|
|
|
`ApproximateAgeOfOldestMessage` is `>= 86400` seconds (1 day). That is the
|
|
|
|
|
ignored-page guard: drain before the 14-day retention window (`1209600` s)
|
|
|
|
|
expires.
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
|
|
|
|
|
2. **RECEIVE** the message (do not purge yet):
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
aws sqs receive-message \
|
|
|
|
|
--queue-url <DLQ_URL> \
|
|
|
|
|
--max-number-of-messages 1 \
|
|
|
|
|
--visibility-timeout 120 \
|
|
|
|
|
--wait-time-seconds 5
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Capture the `ReceiptHandle` from the response.
|
|
|
|
|
|
|
|
|
|
3. **EXTRACT the S3 key.** The `DeadLetterConfig` message `Body` is the original
|
|
|
|
|
async invocation payload — the S3 event JSON. Read
|
|
|
|
|
`Records[0].s3.bucket.name` and `Records[0].s3.object.key` from the `Body`.
|
|
|
|
|
This is the **raw** object key that reprocess/S3 emitted (no URL-decoding
|
|
|
|
|
applied).
|
|
|
|
|
|
|
|
|
|
4. **RE-INVOKE (targeted; dry-run first).** Confirm the key with a dry-run, then
|
|
|
|
|
execute:
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
# PO queue:
|
|
|
|
|
python scripts/reprocess.py --pipeline po --key '<key>' # dry-run
|
|
|
|
|
python scripts/reprocess.py --pipeline po --key '<key>' --execute # re-invoke
|
|
|
|
|
|
|
|
|
|
# WO queue:
|
|
|
|
|
python scripts/reprocess.py --pipeline wo --key '<key>' --execute
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
This re-invokes the **same** function with the same raw-key synthetic S3
|
|
|
|
|
event (a single object — **not** `--all`).
|
|
|
|
|
|
|
|
|
|
5. **VERIFY the write.** Confirm the downstream effect landed before proceeding:
|
|
|
|
|
the DynamoDB item exists / was updated (`purchase-orders` for PO,
|
2026-08-13 17:39:06 -04:00
|
|
|
`work-orders` for WO), and the function's log group shows a clean parse (no new
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
error, no new DLQ message). Do not proceed until verified.
|
|
|
|
|
|
|
|
|
|
6. **PURGE the one message.** Delete only the processed message by its
|
|
|
|
|
`ReceiptHandle`:
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
aws sqs delete-message --queue-url <DLQ_URL> --receipt-handle '<ReceiptHandle>'
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Do **not** `purge-queue` — that would drop unexamined breadcrumbs. AWS access
|
|
|
|
|
is otherwise read-only; `delete-message` on a DLQ you are actively draining is
|
|
|
|
|
the one write this runbook performs.
|
|
|
|
|
|
|
|
|
|
## Recovery windows
|
|
|
|
|
|
|
|
|
|
- **DLQ breadcrumb retention: 14 days** (`MessageRetentionPeriod=1209600s`,
|
|
|
|
|
confirmed live on both queues). After 14 days the breadcrumb is gone.
|
|
|
|
|
- **Raw-email S3 retention: 90 days.** Both buckets have
|
|
|
|
|
`RemovalPolicy.RETAIN`, **but** a single Enabled lifecycle rule
|
|
|
|
|
(`Expiration.Days=90`, `Filter.Prefix=''`) expires objects bucket-wide. **The
|
|
|
|
|
lifecycle rule OVERRIDES the RETAIN policy — S3 is the real replay floor:** a
|
|
|
|
|
raw email is gone at ~90 days regardless of the table RETAIN policy.
|
|
|
|
|
|
|
|
|
|
Because 14 days (DLQ) < 90 days (S3), any object referenced by a live DLQ
|
|
|
|
|
breadcrumb is always still in S3, so DLQ replay within its 14-day window is never
|
|
|
|
|
blocked by S3 expiry. The 90-day floor binds only for replays reconstructed from
|
|
|
|
|
other sources (e.g. logs) after the breadcrumb has expired.
|
|
|
|
|
|
|
|
|
|
## Intentional non-DLQ drops (do NOT hunt for these in the DLQ)
|
|
|
|
|
|
|
|
|
|
Two drop classes **never** produce a DLQ message because they are fail-closed
|
|
|
|
|
**skips, not errors** — the handler returns normally (no raise, no retry, no DLQ
|
|
|
|
|
message):
|
|
|
|
|
|
|
|
|
|
- **Sender-auth rejections.** Logged as a structured `sender_auth_rejected`
|
|
|
|
|
warning and skipped; covered by the `<fn>-sender-auth-rejected` alarm
|
|
|
|
|
(log-metric filter), **not** the DLQ.
|
|
|
|
|
- **`ai_fallback_rejected` drops.** AI-fallback output that failed the
|
|
|
|
|
fail-closed `validate_ai_fallback()` gate, emitted as a
|
|
|
|
|
`ParseMethod=ai_fallback_rejected` EMF datapoint and dropped without a
|
|
|
|
|
DynamoDB write; covered by the `<fn>-ai-fallback-rejected` alarm, **not** the
|
|
|
|
|
DLQ.
|
|
|
|
|
|
|
|
|
|
If mail is missing but the DLQ is empty, check those two alarms / log filters —
|
|
|
|
|
the email was intentionally rejected. Re-invoking it via reprocess will just be
|
|
|
|
|
rejected again; fix the sender-auth config or the upstream email, not the DLQ.
|