Commit graph

7 commits

Author SHA1 Message Date
Adam Moussa
073201f633
Migrate to seahaven-prod: deploy role, backfill tooling, account-portability fixes (#125)
* feat(migration): prepare stacks and tooling for the seahaven-prod account move

Phase 1 of the mgmt (328440206208) -> seahaven-prod (011934824531)
migration. No behavior change in-account; everything here is additive or
account-portability hygiene:

- infra/deploy-role/: reviewed OIDC deploy-role artifacts for prod
  (trust main-only, cdk-hnb659fds-* AssumeRole, smoke-invoke-lambda scoped
  to exactly the two email-processor fn ARNs). Codifies the previously
  out-of-band smoke-invoke grant.
- Table resource policies: make_slack_bot_read_policy in cdk/common.py,
  applied to purchase-orders, verified-sites, WorkOrders,
  WorkOrderComments (NOT pending-site-review; no bot consumer). Grants the
  mgmt-resident seahaven-slack-bot roles read-only cross-account access
  post-move (bot-side identity grants land in the slack-bot repo).
- scripts/migrate_tables.py: dry-run-default backfill tool implementing
  the plan's per-table semantics (superset overwrite, ingested_at cutoff
  for WorkOrderComments, backup-gated truncate-and-load for the two
  site tables) plus a verify subcommand (count parity, spot checks,
  sticky-Cancelled drift check).
- tests/test_resource_policy_helper.py: statement-shape unit tests +
  static pins that exactly the four bot-read tables carry the policy.
- Account-literal fixes: account-agnostic fixture bucket in
  test_reprocess_contract; runbook/README/po-template-parser account
  references updated to prod with historical mgmt notes; README gains the
  account-prerequisites list (imported-by-name dependencies).

deploy.yaml is deliberately unchanged (push-to-main auto-deploy kept).
Merge is held until migration Phase 0 completes; flipping the
AWS_DEPLOY_ROLE_ARN repo secret and merging this PR IS the first prod
deploy.

* fix(migration): verify backup AVAILABLE pre-truncate; document wildcard risk acceptance (cross-review FIX/NIT)

* refactor(migration): drop cross-account read grants (slack-bot decommissioned); harden backfill + deploy role

seahaven-slack-bot was decommissioned 2026-07-23 (stack DELETE_IN_PROGRESS,
consumer Lambdas gone); its successor sh-mcp is undeployed and uses
same-account DynamoDB access. So no live consumer reads these tables
cross-account. Per Adam's call, drop the cross-account grants entirely and
re-add correctly-scoped ones if/when sh-mcp deploys to a different account.

- Remove the four table resource policies + make_slack_bot_read_policy helper
  + its constants (cdk/common.py, po_stack.py, wo_stack.py) and the helper's
  unit test. Both stacks synth with zero table ResourcePolicy.
- scripts/migrate_tables.py hardening (fixes from the sh-security-review
  fan-out on the destructive backfill tool):
  * validate --cutoff strictly (parse ISO-8601, require aware UTC, re-emit
    canonical second-precision form) so a malformed cutoff can't silently
    copy dual-window rows or drop history;
  * reject `copy --all` up front (must run tables individually, in order,
    with the stream-drain wait) instead of writing three tables then erroring;
  * truncate backup gate now also checks recency (<1h) and TableId, not just
    status+name;
  * verify requires --cutoff whenever a cutoff table is in scope (else it
    false-flags dual-window rows as MISSING);
  * sticky-cancel is now PREVENTED copy-side (a non-Cancelled source item
    never overwrites a dest-Cancelled PO), and the verify comment no longer
    overstates what its source-side scan covers;
  * spot-check all modes (truncate_load keys are verbatim, so key-existence
    is sound there too).
- Deploy role: scope cloudformation:DescribeStacks to this repo's stacks +
  CDKToolkit (was Resource:*, disclosed all tenant stacks in the shared prod
  account); add a drift check warning on unexpected role policies and drop the
  dead SMOKE_POLICY_NAME var; document the shared-account bootstrap-role
  accepted risk in the deploy-role README.

* docs(deploy-role): fold in cross-review NITs (DescribeStacks maintenance note, warn-only drift rationale)

* ci: update workflow to use new workflow tag (ruff versioning fix)

* fix(migration): address Open SWE review findings on migrate_tables.py

- Validate the truncate backup on dry-run as well as --execute so a
  missing/stale/wrong-incarnation --backup-arn surfaces on the rehearsal
  run (finding f_24a48b8900).
- Assert configured keys match the live key schema of both tables before
  any key projection, turning config/schema drift into a descriptive
  abort instead of a mid-backfill KeyError (finding f_cb6b5a6c59).
- Clarify why key-existence spot-checks are sound for WorkOrderComments:
  the copy Puts source items verbatim and the sample uses the same
  cutoff filter, so per-account comment_id divergence never enters the
  check (finding f_390b7d6c3b is a false positive; comment hardened).
2026-07-23 17:08:47 -04:00
Adam Moussa
75fe91c198
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7)

Generalize scripts/reprocess.py from a PO-only full-sweep script into a
pipeline-general recovery tool. Targeted replay (--key/--prefix/--since)
is now the default, and the full inbound/ sweep is demoted behind an
explicit --all that documents its five hazards (async concurrency does
not serialize, use RequestResponse if order matters, metric double-count,
Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves
the correct function + bucket; dry-run-by-default / --execute is preserved.
A new tests/test_reprocess_contract.py pins the synthetic S3 event shape
and asserts the raw list_objects_v2 key is emitted untransformed (the
handler is the single decode point; a pre-decoded key would corrupt keys
containing spaces or '+').

Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console
redrive-to-source, so it documents the receive -> extract key -> targeted
reprocess --key -> verify -> purge procedure, the real recovery windows
(14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table
RETAIN policy and is the true replay floor), and that sender-auth and
ai_fallback_rejected drops are fail-closed skips that never reach the DLQ.
Linked from the README alarms and scripts sections.

Drop the vendored boto3 floor pin from both email-processor requirements
(the Lambda runtime provides boto3; lambda-template.md empty-with-comment
form). With nothing left to install, the email-processor bundling becomes
cp-only -- the whole pip step is removed, which is the only acceptable way
the manylinux2014_aarch64 pin disappears (removing the pin while keeping a
pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and
add pinned po/web_ui + po/site_extractor manifests (excluded from their
bundles, so hash-neutral) so their new Dependabot entries have something
to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and
/lambdas/po/site_extractor.

cdk diff is confined to exactly the two email processors' asset hashes on
both stacks. The wo/web_ui dead-manifest reduction was deliberately left
out: that manifest already ships inside the plain (non-bundled) WebUI
asset on main, so reducing or excluding it would redeploy workorder-web-ui
for no functional change -- deferred to keep the blast radius to the two
intended targets.

The untracked 44 MB lambdas/po/email_processor/package/ dir was removed
from the filesystem (asset-hash-neutral given Phase 2's package/ exclude);
it is untracked, so there is nothing to commit for it.

* Reject --all combined with --prefix/--since in reprocess.py

--all is a distinct mode (the demoted full-prefix sweep), but the args.all
branch unconditionally set prefix=inbound/ and since=None, so passing it
alongside a narrower selector silently discarded that selector. `--all
--since 2026-07-01` swept the entire corpus instead of the bounded window,
triggering every documented --all hazard (Bedrock re-bill, metric double-
count, merged-field regression) on objects the operator never targeted --
contradicting the tool's safety goal. Add the missing mutual-exclusion
guard alongside the existing --key one, and pin --all+--prefix,
--all+--since, and all three together as argparse rejections.
2026-07-20 12:53:34 -04:00
Adam Moussa
cb5539bd68
feat: deploy-pipeline guards — healthcheck, smoke gate, bundle glob + AST test (refactor phase 0) (#107)
Some checks are pending
Deploy / deploy (push) Waiting to run
* feat: deploy-pipeline guards — healthcheck, smoke gate, bundle glob + AST test (refactor phase 0)

Deploys of po-email-processor and workorder-email-processor had no
verification step, so an init-time ImportError in the bundled zip
could ship silently and only surface on the next real S3 event. This
adds a synchronous post-deploy smoke gate wired into the deploy
workflow: both Lambdas are invoked with {"healthcheck": true} and the
FunctionError field is checked, since an Unhandled init error still
returns HTTP 200 on RequestResponse invokes and would false-pass a
plain exit-code check.

The healthcheck branch is the first statement in each handler, before
any boto3/S3 use or ses_auth, and only fires on a top-level direct
invoke ("healthcheck" is not a key AWS ever sets on a real S3
ObjectCreated event, so mail content can't reach this path). It emits
no EMF metrics and no log text that could match the
sender-auth-rejected metric filter, so two deploys in one window
won't trip the alarm.

Separately, the PO stack's asset bundling copied a hand-maintained
four-file allowlist into the zip, so every new sibling module
handler.py imports had to be added by hand or the deploy shipped a
Lambda that ImportErrors at cold start (bit us for template_parser in
PR #105 and nearly for derived_fields in PR #2). Replaced it with a
non-recursive ./*.py glob so top-level source files ship
automatically while tests/ and the stale package/ dir still cannot,
and added an AST-based bundle-consistency test that parses each
handler's first-party imports and fails CI if the bundling command
would omit any of them (a revert to an incomplete allowlist, or code
moved into a subdirectory the glob doesn't cover).

Includes the refactor-evaluation report that scoped this phase.

* fix: review nits — unambiguous bundling-command extraction, smoke payload-parse message, dead asserts

- tests/test_bundle_consistency.py: _extract_bundling_command now collects
  all command=[...] matches and demands exactly one per stack file, instead
  of silently returning whichever ast.walk visits first if a second bundled
  function is ever added.
- scripts/post-deploy-smoke.sh: distinguish an unparseable response payload
  from a payload mismatch so the failure message says what actually happened
  (the previous "could not parse" branch was unreachable — the inline python
  always exited 0).
- test_po_healthcheck.py: drop the substring assertions on stdout that were
  dead behind the stricter `captured.out == ""` assertion; keep the stderr
  filter-pattern check.

Review follow-up on PR #107; no behavior change to any shipped code path.
2026-07-17 13:18:45 -04:00
Adam Moussa
acc1961d21
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails

The workorder-email-processor sends every one of ~22.9k emails/month to
an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment
template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse
those two shapes deterministically, offline, so the AI call is reserved
for the long tail.

The module is pure (no boto3, no network). try_deterministic_parse
classifies by subject, extracts the shared contract fields, and returns
a result ONLY when it passes a strict fail-closed validation gate: exact
contract-key set, subject/id agreement, the literal "Work Order:  <id>"
double space, per-type required fields, site-code shape, and a
label-bleed guard so a value that over-ran into the next field fails.
Any miss, drift, or extractor exception yields None so the caller falls
back to the AI extractor -- data is never corrupted, only the fallback
rate rises.

Refs: #23

* Migrate WO processor to Bedrock and fix comment_id collision

Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel
on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0
(BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or
Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept
byte-identical, so the AI-fallback output is unchanged. Try the new
deterministic template parser first and only call Bedrock on a
miss/invalid result.

Fix issue #23: the WorkOrderComments range key was
work_order_id#<comment_time>, so two emails on one WO with an identical
or absent comment time collided and overwrote each other. Derive a
12-hex suffix from the S3 object key alone -- deterministic, so an async
retry of the same object is byte-identical (idempotent) while distinct
emails get distinct keys -- and keep wall-clock now() out of the key
(literal 'nocomment' segment when comment_time is absent).

Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest
ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome
observability, replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc), and drop the anthropic dependency.

Refs: #23

* Migrate PO processor to Bedrock

Switch the PO email processor's AI extraction from the Anthropic SDK to
bedrock-runtime InvokeModel on the inference profile
us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so
it no longer needs a provider API key or Secrets Manager secret. PO
parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT
is kept byte-identical and the Bedrock text output is still decoded with
json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects
floats). Replace the deprecated datetime.utcnow() with
datetime.now(timezone.utc) and drop the anthropic dependency.

* Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm

Both stacks moved their processors from the Anthropic API to the Bedrock
inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each
processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream
on BOTH the inference-profile ARN AND the per-region foundation-model
ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile
routes cross-region, so a profile-only grant AccessDenies at runtime.

Remove both anthropic-api-key Secret constructs, their grant_read, and
the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets
had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in
the README for manual post-deploy deletion and key revocation.

Add the workorder-email-processor-template-fallback-rate alarm: a
FILL(0) + >=10-sample volume-floor MathExpression over the EMF
ParseOutcome metric (15-min periods) that pages when the AI-fallback
share exceeds 15% sustained, catching Hexagon template drift. ALARM-only
SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the
existing stack idiom.

* Add offline WO parser test suite

Cover the deterministic parser with golden-file tests over 55 real
scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By,
address present/absent, br+CRLF assign addresses), fail-closed
validation-gate rules, adversarial and prompt-injection cases that must
route to ai_fallback or parse without corrupting other fields, the issue
#23 comment_id idempotency invariants, and the Bedrock-fallback dispatch
plus EMF-metric emission with a mocked invoke_model.

Extend pytest.ini testpaths to discover the co-located suite, and update
tests/conftest.load_handler to put a handler's own directory on sys.path
so the WO handler's new `from template_parser import ...` resolves under
the existing shared handler tests. Point test_local.py at the new
template-first + Bedrock flow.

Refs: #23

* Document Bedrock migration and WO parse flow in README

Record the provider switch to the Bedrock inference profile (no Anthropic
API key or Secrets Manager secret, with the retired secrets flagged for
manual deletion), the WO deterministic-template-first + AI-fallback flow,
the new ParseOutcome EMF metric and template-fallback-rate alarm, the
issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift,
and offline test instructions.

Refs: #23

* Fix f-string lint and formatting in backfill scripts

Drop the f prefix from two f-strings that carry no placeholders
(F541) and apply ruff format, so `ruff check` / `ruff format --check`
pass in CI.

* Emit ParseMethod-only EMF set so fallback alarm can fire

The fallback-rate alarm queries the ParseOutcome series keyed on
ParseMethod alone, but the emitter published only the joint
(ParseMethod, TemplateId) dimension set. CloudWatch materializes
exactly the listed dimension sets and does not auto-aggregate, so the
alarm's series never received data: it evaluated a constant 0 and
could never page on template-drift coverage collapse.

Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and
update the EMF regression test to assert both sets are present.

* Commit WO parser .eml fixtures for executable coverage

The parser test suite globbed for input .eml fixtures that the repo's
`*.eml` ignore rule kept uncommitted, so every parametrized golden and
fail-closed test collected zero cases and CI could not exercise the
deterministic parser that handles 100% of WO email volume.

Add a fixtures-only negation to .gitignore and commit the 55 scrubbed
positive samples (50 update-plaintext, 5 assign-html) plus 14
ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each
fail-closed reason code (subject_no_match, single_space_work_order,
malformed_site_code, label_bleed, creation_time_unparseable,
wo_id_mismatch, missing_required_field) and the adversarial set proves
the parser is total and confines prompt-injection payloads to
comment_text without steering the structured fields.

* Fix WO parser advisories A1-A3 (PR #99 follow-ups)

A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model
output and not stable across Lambda async retries, so on the ai_fallback
path the comment_id range-key time segment now derives from the email Date
header (deterministic per S3 object) instead of the model's comment_time.
The template path is unchanged (its comment_time is a pure function of the
raw email). Bedrock invoke pins temperature 0 so retries reproduce the same
extraction. Closes the #23 reopening on the AI path.

A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so
CloudWatch reliably extracts the ParseOutcome datapoint that the
fallback-rate alarm depends on.

A3 — T1 New Comment capture no longer truncates at the first blank line;
multi-paragraph comments are captured through internal blanks and terminate
at the next label/separator. 17 golden files regenerated from the real
fixtures accordingly.

Hardening from the sh-security-review pass on this diff:
- _header_date_iso is total: OverflowError/OSError from an extreme Date
  header fall back to 'nocomment' instead of failing the invocation.
- _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic
  path on a crafted large blank run.
- work_order_id is enforced digits-only on BOTH parse paths before it is
  used as a DynamoDB key, so prompt-injected AI output cannot forge '#'
  range-key segments or land on an arbitrary WO.
2026-07-16 12:45:11 -04:00
Adam Moussa
5112c1345b
Merge workorder-ingest into unified procurement repo (#22)
* Merge workorder-ingest pipeline into unified repo

Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/.
Two independent CloudFormation stacks in one CDK app. Fix WO stack
compliance: ARM64 architecture, 60-day log retention, aarch64 bundling,
RETAIN on Anthropic secret. Remove stale CodePipeline buildspec.

* Fix test_local.py import path and remove dead shared/models.py

test_local.py referenced the old lambdas/email_processor path. Updated
to lambdas/wo/email_processor. Removed shared/ directory entirely as
nothing imports from it.

* Escape HTML in both web UI dashboards to prevent XSS

Both Function URLs are public (auth_type=NONE) and render
email-derived content via f-strings. Attacker-crafted emails
could inject scripts. Added html.escape() on all interpolated
values in both PO and WO dashboards.

* Add pagination to WO web UI scan

get_work_orders() only fetched the first 1MB page from DynamoDB.
Loop on LastEvaluatedKey to match the PO web UI pattern.

* Fix esc(None) TypeError and javascript: scheme in PO web UI

Coerce supplier name through `or ""` before escaping to handle
nested None from DynamoDB. Add scheme allowlist on view_order_url
to block javascript:/data: hrefs from LLM-extracted URLs.

* Fix WO render_badge None guard, updated_at slice, and backfill path

Add null guard to WO render_badge matching the PO version. Use
`or ""` before slicing updated_at to handle explicit None values.
Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor.

* Harden WO web UI and fix JS-context XSS in both dashboards

- Use json.dumps for onclick URLs to prevent JS string breakout
- Add .lower() to WO render_badge color lookup matching PO pattern
- Add pagination to get_comments query
- Cap get_work_orders to 500 results matching PO pattern

* Apply ruff formatting to web UI handlers
2026-05-12 15:21:06 -04:00
Adam Moussa
ec416079f5 Add verified-sites pipeline via DynamoDB Streams
Enable DynamoDB Streams on purchase-orders table and add a site-extractor
Lambda that extracts Amazon facility codes and addresses from PO ship-to
data, upserting them into a new verified-sites table. Includes a backfill
script for existing POs and upgrades existing Lambdas to arm64 + 60-day
log retention.
2026-04-30 14:26:53 -04:00
Adam Moussa
fc690dd958 Initial commit: PO email ingestion pipeline
CDK stack with SES receipt rule, S3 bucket, email processor Lambda
(Claude-powered extraction), web UI Lambda with Function URL, and
DynamoDB for storage. Includes reprocessing script for missed emails.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-07 12:12:30 -04:00