2026-04-07 12:12:30 -04:00
|
|
|
"""
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
Reprocess missed procurement-ingest emails by re-invoking an email-processor
|
|
|
|
|
Lambda with a synthetic S3 event.
|
|
|
|
|
|
|
|
|
|
TARGETED replay is the default and preferred mode: point it at exactly the
|
|
|
|
|
object(s) you need to replay.
|
|
|
|
|
|
|
|
|
|
# dry-run a single object (default is dry-run — lists, does not invoke)
|
|
|
|
|
python scripts/reprocess.py --pipeline po --key inbound/2026/msg.eml
|
|
|
|
|
# actually re-invoke that one object
|
|
|
|
|
python scripts/reprocess.py --pipeline po --key inbound/2026/msg.eml --execute
|
|
|
|
|
# replay a prefix, or everything modified since a timestamp
|
|
|
|
|
python scripts/reprocess.py --pipeline wo --prefix inbound/2026/07/
|
|
|
|
|
python scripts/reprocess.py --pipeline po --since 2026-07-16T00:00:00Z --execute
|
|
|
|
|
|
|
|
|
|
FULL-PREFIX sweep is DEMOTED behind an explicit --all flag and carries five
|
|
|
|
|
load-bearing caveats (see the --all help text and the in-code comment above the
|
|
|
|
|
sweep branch):
|
|
|
|
|
|
|
|
|
|
python scripts/reprocess.py --pipeline po --all --execute # replays ALL of inbound/
|
|
|
|
|
|
|
|
|
|
1. Event-type (async) invocation is CONCURRENT, so sorting does NOT
|
|
|
|
|
serialize the replay.
|
|
|
|
|
2. Switch InvocationType to "RequestResponse" if ordering matters.
|
|
|
|
|
3. Metrics get double-counted on replay.
|
|
|
|
|
4. Bedrock is re-billed for every ai_fallback email replayed.
|
|
|
|
|
5. Out-of-order replay can REGRESS already-merged fields.
|
|
|
|
|
|
|
|
|
|
Dry-run is the default in ALL modes (targeted and --all alike); nothing is
|
|
|
|
|
invoked unless --execute is passed.
|
2026-04-07 12:12:30 -04:00
|
|
|
"""
|
|
|
|
|
|
|
|
|
|
import argparse
|
|
|
|
|
import json
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
from datetime import datetime, timezone
|
2026-04-07 12:12:30 -04:00
|
|
|
|
|
|
|
|
import boto3
|
|
|
|
|
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
# Per-pipeline resolution: the function to invoke and the raw-email bucket to
|
|
|
|
|
# list, keyed by --pipeline. Both are derived at runtime from --pipeline +
|
|
|
|
|
# the caller's account id; there is no PO-hardcoded default.
|
|
|
|
|
PIPELINES = {
|
|
|
|
|
"po": {"function": "po-email-processor", "bucket": "po-ingest-emails-{acct}"},
|
|
|
|
|
"wo": {
|
|
|
|
|
"function": "workorder-email-processor",
|
|
|
|
|
"bucket": "workorder-ingest-emails-{acct}",
|
|
|
|
|
},
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
DEFAULT_PREFIX = "inbound/"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def parse_since(value: str) -> datetime:
|
|
|
|
|
"""Parse an ISO-8601 UTC timestamp (accepts a trailing 'Z') into an
|
|
|
|
|
aware UTC datetime so it can be compared against S3 LastModified."""
|
|
|
|
|
normalized = value.replace("Z", "+00:00")
|
|
|
|
|
dt = datetime.fromisoformat(normalized)
|
|
|
|
|
if dt.tzinfo is None:
|
|
|
|
|
dt = dt.replace(tzinfo=timezone.utc)
|
|
|
|
|
return dt.astimezone(timezone.utc)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def list_keys(s3, bucket: str, prefix: str, since: datetime | None) -> list[str]:
|
|
|
|
|
"""Paginate list_objects_v2 under prefix, returning raw object keys.
|
|
|
|
|
|
|
|
|
|
When since is set, only objects with LastModified >= since are kept.
|
|
|
|
|
The keys returned are the RAW list_objects_v2 obj["Key"] — no
|
|
|
|
|
URL-encoding or decoding is applied (see build_s3_event).
|
|
|
|
|
"""
|
|
|
|
|
keys: list[str] = []
|
2026-04-07 12:12:30 -04:00
|
|
|
paginator = s3.get_paginator("list_objects_v2")
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
for page in paginator.paginate(Bucket=bucket, Prefix=prefix):
|
2026-04-07 12:12:30 -04:00
|
|
|
for obj in page.get("Contents", []):
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
if since is not None and obj["LastModified"] < since:
|
|
|
|
|
continue
|
2026-04-07 12:12:30 -04:00
|
|
|
keys.append(obj["Key"])
|
|
|
|
|
return keys
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def build_s3_event(bucket: str, key: str) -> dict:
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
"""Build the synthetic S3 notification event the handler consumes.
|
|
|
|
|
|
|
|
|
|
CONTRACT (pinned by tests/test_reprocess_contract.py): the shape is
|
|
|
|
|
exactly Records[0].s3.bucket.name / Records[0].s3.object.key, and the
|
|
|
|
|
key is emitted RAW — reprocess applies NO urllib.parse.quote/unquote.
|
|
|
|
|
A real S3 notification URL-encodes the object key; reprocess builds from
|
|
|
|
|
the raw list_objects_v2 key (or the raw --key arg), and the handler is
|
|
|
|
|
what would decode. Replaying a pre-decoded/pre-encoded key would not
|
|
|
|
|
match what s3.get_object(Bucket, Key) expects.
|
|
|
|
|
"""
|
2026-04-07 12:12:30 -04:00
|
|
|
return {
|
|
|
|
|
"Records": [
|
|
|
|
|
{
|
|
|
|
|
"s3": {
|
|
|
|
|
"bucket": {"name": bucket},
|
|
|
|
|
"object": {"key": key},
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
]
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def main():
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
parser = argparse.ArgumentParser(
|
|
|
|
|
description=(
|
|
|
|
|
"Reprocess missed procurement-ingest emails by re-invoking an "
|
|
|
|
|
"email-processor Lambda with a synthetic S3 event. TARGETED replay "
|
|
|
|
|
"(--key/--prefix/--since) is the default and preferred mode; "
|
|
|
|
|
"full-prefix sweep requires the explicit --all flag."
|
|
|
|
|
)
|
|
|
|
|
)
|
|
|
|
|
parser.add_argument(
|
|
|
|
|
"--pipeline",
|
|
|
|
|
required=True,
|
|
|
|
|
choices=["po", "wo"],
|
|
|
|
|
help=(
|
|
|
|
|
"Which pipeline to replay: 'po' (po-email-processor / "
|
|
|
|
|
"po-ingest-emails-<acct>) or 'wo' (workorder-email-processor / "
|
|
|
|
|
"workorder-ingest-emails-<acct>)."
|
|
|
|
|
),
|
|
|
|
|
)
|
|
|
|
|
# Targeting selectors — at least one of {--key, --prefix, --since, --all}
|
|
|
|
|
# is required (enforced below); --key is exclusive.
|
|
|
|
|
parser.add_argument(
|
|
|
|
|
"--key",
|
|
|
|
|
help="Replay exactly one raw object key (e.g. inbound/2026/msg.eml).",
|
|
|
|
|
)
|
|
|
|
|
parser.add_argument(
|
|
|
|
|
"--prefix",
|
|
|
|
|
help=(
|
|
|
|
|
"Replay every object under this S3 prefix "
|
|
|
|
|
"(default targeting prefix is inbound/)."
|
|
|
|
|
),
|
|
|
|
|
)
|
|
|
|
|
parser.add_argument(
|
|
|
|
|
"--since",
|
|
|
|
|
help=(
|
|
|
|
|
"Only objects with LastModified >= this ISO-8601 UTC timestamp "
|
|
|
|
|
"(e.g. 2026-07-16T00:00:00Z). Combines with --prefix."
|
|
|
|
|
),
|
|
|
|
|
)
|
|
|
|
|
parser.add_argument(
|
|
|
|
|
"--all",
|
|
|
|
|
action="store_true",
|
|
|
|
|
help=(
|
|
|
|
|
"DEMOTED full-prefix sweep: replay EVERY object under inbound/ "
|
|
|
|
|
"(not a targeted subset). "
|
|
|
|
|
"CAVEATS: (1) Event-type (async) invocation is CONCURRENT so "
|
|
|
|
|
"sorting does NOT serialize the replay; "
|
|
|
|
|
"(2) switch to RequestResponse if order matters; "
|
|
|
|
|
"(3) metrics get double-counted on replay; "
|
|
|
|
|
"(4) Bedrock is re-billed; "
|
|
|
|
|
"(5) out-of-order replay can REGRESS already-merged fields. "
|
|
|
|
|
"Prefer --key/--prefix/--since. Still dry-run unless --execute."
|
|
|
|
|
),
|
|
|
|
|
)
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
parser.add_argument(
|
|
|
|
|
"--execute",
|
|
|
|
|
action="store_true",
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
help=(
|
|
|
|
|
"Actually invoke the Lambda. Default is dry-run (list only). "
|
|
|
|
|
"Applies to EVERY mode, including --all."
|
|
|
|
|
),
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
)
|
2026-04-07 12:12:30 -04:00
|
|
|
args = parser.parse_args()
|
|
|
|
|
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
# Exactly-one target gate: full sweep is ONLY reachable via --all, so an
|
|
|
|
|
# accidental whole-prefix sweep from running with no selector is impossible.
|
|
|
|
|
if not any([args.key, args.prefix, args.since, args.all]):
|
|
|
|
|
parser.error("choose a target: --key, --prefix, --since, or --all (full sweep)")
|
|
|
|
|
# --key is mutually exclusive with the listing selectors and with --all.
|
|
|
|
|
if args.key and (args.prefix or args.since or args.all):
|
|
|
|
|
parser.error("--key cannot be combined with --prefix, --since, or --all")
|
|
|
|
|
# --all is a distinct MODE (the demoted full-prefix sweep), not a modifier:
|
|
|
|
|
# reject it alongside the targeting selectors. Without this, the args.all
|
|
|
|
|
# clobber branch below would silently discard a narrower --prefix/--since
|
|
|
|
|
# (e.g. `--all --since 2026-07-01` would sweep the ENTIRE corpus, not the
|
|
|
|
|
# bounded window) and trigger every --all hazard on unintended objects.
|
|
|
|
|
if args.all and (args.prefix or args.since):
|
|
|
|
|
parser.error("--all cannot be combined with --prefix or --since")
|
|
|
|
|
|
|
|
|
|
account_id = boto3.client("sts").get_caller_identity()["Account"]
|
|
|
|
|
function_name = PIPELINES[args.pipeline]["function"]
|
|
|
|
|
bucket = PIPELINES[args.pipeline]["bucket"].format(acct=account_id)
|
|
|
|
|
|
|
|
|
|
if args.key:
|
|
|
|
|
# Single raw key — skip listing entirely. The key is passed through
|
|
|
|
|
# to build_s3_event untransformed (raw-key contract).
|
|
|
|
|
keys = [args.key]
|
|
|
|
|
else:
|
|
|
|
|
s3 = boto3.client("s3")
|
|
|
|
|
# --all is the DEMOTED full-prefix sweep. Five caveats, all load-bearing:
|
|
|
|
|
# 1. Event-type (async) invocation is CONCURRENT, so sorting the key
|
|
|
|
|
# list does NOT serialize the replay.
|
|
|
|
|
# 2. If ordering matters, switch InvocationType to "RequestResponse"
|
|
|
|
|
# (serial, waits per call).
|
|
|
|
|
# 3. Metrics (ParseOutcome EMF, alarms) get double-counted on replay.
|
|
|
|
|
# 4. Bedrock is re-billed for every ai_fallback email replayed.
|
|
|
|
|
# 5. Out-of-order replay can REGRESS already-merged fields (a stale
|
|
|
|
|
# email overwriting a newer merge).
|
|
|
|
|
# Targeted replay (--key / --prefix / --since) avoids all five at
|
|
|
|
|
# whole-corpus scale; prefer it.
|
|
|
|
|
if args.all:
|
|
|
|
|
prefix = DEFAULT_PREFIX
|
|
|
|
|
since = None
|
|
|
|
|
else:
|
|
|
|
|
# --prefix / --since target. --since without --prefix defaults the
|
|
|
|
|
# prefix to inbound/ so it filters the raw-email tree, not the bucket.
|
|
|
|
|
prefix = args.prefix if args.prefix else DEFAULT_PREFIX
|
|
|
|
|
since = parse_since(args.since) if args.since else None
|
|
|
|
|
keys = list_keys(s3, bucket, prefix, since)
|
2026-04-07 12:12:30 -04:00
|
|
|
|
|
|
|
|
if not keys:
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
print(f"No matching objects in s3://{bucket}/ — nothing to reprocess.")
|
2026-04-07 12:12:30 -04:00
|
|
|
return
|
|
|
|
|
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
if args.all:
|
|
|
|
|
print(
|
|
|
|
|
f"[--all full sweep] {len(keys)} email(s) under "
|
|
|
|
|
f"s3://{bucket}/{DEFAULT_PREFIX}\n"
|
|
|
|
|
)
|
|
|
|
|
else:
|
|
|
|
|
print(f"{len(keys)} email(s) selected in s3://{bucket}/\n")
|
|
|
|
|
|
|
|
|
|
lambda_client = boto3.client("lambda") if args.execute else None
|
2026-04-07 12:12:30 -04:00
|
|
|
|
|
|
|
|
for key in keys:
|
|
|
|
|
if args.execute:
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
print(f" Invoking {function_name} for {key} ... ", end="", flush=True)
|
2026-04-07 12:12:30 -04:00
|
|
|
resp = lambda_client.invoke(
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)
Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').
Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.
Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.
cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.
The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.
* Reject --all combined with --prefix/--since in reprocess.py
--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
|
|
|
FunctionName=function_name,
|
2026-04-07 12:12:30 -04:00
|
|
|
InvocationType="Event", # async — don't wait for each one
|
|
|
|
|
Payload=json.dumps(build_s3_event(bucket, key)),
|
|
|
|
|
)
|
|
|
|
|
print(f"status {resp['StatusCode']}")
|
|
|
|
|
else:
|
|
|
|
|
print(f" [dry-run] {key}")
|
|
|
|
|
|
|
|
|
|
if not args.execute:
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails
The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.
The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order: <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.
Refs: #23
* Migrate WO processor to Bedrock and fix comment_id collision
Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.
Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).
Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.
Refs: #23
* Migrate PO processor to Bedrock
Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.
* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm
Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.
Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.
Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.
* Add offline WO parser test suite
Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.
Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.
Refs: #23
* Document Bedrock migration and WO parse flow in README
Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.
Refs: #23
* Fix f-string lint and formatting in backfill scripts
Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.
* Emit ParseMethod-only EMF set so fallback alarm can fire
The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.
Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.
* Commit WO parser .eml fixtures for executable coverage
The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.
Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.
* Fix WO parser advisories A1-A3 (PR #99 follow-ups)
A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.
A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.
A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.
Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
|
|
|
print("\nDry run complete. Re-run with --execute to invoke the Lambda.")
|
2026-04-07 12:12:30 -04:00
|
|
|
|
|
|
|
|
|
|
|
|
|
if __name__ == "__main__":
|
|
|
|
|
main()
|