mirror of
https://github.com/Sea-Haven-Industries/procurement-ingest.git
synced 2026-09-30 22:23:14 +00:00
11 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
cf8ca2eeb6
|
feat(api): custom domain procurement-api.seahaven.com (stacked on PR-2) (#140)
* feat(api): custom domain procurement-api.seahaven.com for the read API Stacked on feat/shoc-wo-webhook. Gives the SHOC-facing read API a stable, brandable endpoint instead of the opaque execute-api URL. - procurement_api_stack.py: REGIONAL API Gateway DomainName (TLS 1.2) + empty base-path mapping to the prod stage, so callers hit https://procurement-api.seahaven.com/work-orders (no /prod segment). The ACM cert ARN is read from SSM (/procurement-api/custom-domain/certificate-arn) via value_for_string_parameter, because the seahaven.com zone is in the mgmt account (cross-account DNS) and the cert is issued out of band. Outputs expose the regional alias target + hosted-zone id for the mgmt A-record. - scripts/setup_procurement_api_domain.sh: idempotent two-step runbook (cert: request + mgmt-zone validation + wait + SSM; alias: post-deploy A-record from stack outputs). Verifies both account identities. - handler._base_url: omit the /{stage} segment for a custom-domain request (the base-path mapping serves the stage at the root) so the docs never advertise a broken server URL; execute-api hosts keep /{stage}. - openapi.json: custom domain added as servers[0] (recommended), execute-api kept as the direct fallback + the per-request injection target. No IAM/auth/policy change (same API id + resource policy), so the SigV4 surface and the mandatory cross-family gates are unaffected. 751 pytest, ruff, cdk synth, redocly lint all green. * fix(api): use .endswith('.amazonaws.com') instead of substring check for execute-api detection The prior '.execute-api.' in domain substring check is fragile and triggers CodeQL incomplete-sanitization warnings. All API Gateway default domains end with .amazonaws.com, so a suffix check is more precise and also silences the false-positive alert. Refs: https://github.com/Sea-Haven-Industries/procurement-ingest/security/code-scanning/6 |
||
|
|
c040050373
|
feat(webhook): SHOC WO webhook emitter - dark-ship streams + HMAC secret/rotation (PR-2) (#137)
* docs(webhook): revise SHOC webhook contract and plan for post-migration reality Branch re-cut on main 2026-07-23 (old base carried stale PR #99 commits). Contract Rev 2026-07-23: - Producer account corrected: seahaven-prod (011934824531); mgmt frozen - Reconciliation backstop is the new procurement read API, not SyncController - wo_status "unknown" is real; SHOC must map it (checklist item added) - write_origin forward-compat note for phase-2 write-back echo suppression - SyncVendorReplies retirement flagged (dead table, no vendor_reply event) Plan updates: - Account gate: seahaven-prod only; never enable streams on mgmt tables - Emitter ships DARK (ESMs enabled=False); activation is a deliberate flip after the SHOC receiver passes shared HMAC vectors - Post-refactor conventions: common.py helpers, bundle-consistency AST pins, pytest.ini --cov additions, consolidated test roots - Dedicated-CMK rationale, secret-ARN handooff step, consumer audit refreshed (slack-bot decommissioned), enum golden test, write_origin skip-branch test * feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC call-and-be-called effort). Everything ships DARK: both DynamoDB event source mappings deploy enabled=False; activation is a deliberate one-line follow-up PR gated on the SHOC receiver passing the shared HMAC test vectors. - Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments (in-place update, RETAIN + logical IDs untouched; no existing consumers — verified live, neither table had a stream). - workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope -> HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard ordering (parallelization 1, bisect off, retry until 24h age, ReportBatchItemFailures); 429/5xx/timeout block the shard in order, other 4xx park to workorder-shoc-emitter-rejected; ESM failures -> workorder-shoc-emitter-failures (metadata; replay rebuilds from DynamoDB). Echo guard skips write_origin=shoc-write-api. - Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK (alias workorder-ingest-shoc-webhook-kms); cross-account GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy DESTROY deliberately (machine-generated material; avoids the fixed-name RETAIN-orphan deadlock). - workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap, 64-hex keys, kid = UTC %Y-%m-%dT%H. - Alarms (ALARM-only -> site-alerts): emitter errors/throttles/ duration + iterator-age (>=10 min) + failures/rejected queue depth; rotator standard trio. - scripts/replay_shoc_webhooks.py: dry-run-default operator replay (rebuilds from tables, replay:true envelopes). - Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay signing pinned to identical vectors); bundle-consistency AST pins for both new bundles. - README: WO stack + webhook feed section, alarm table, runbooks; removed stale seahaven-slack-bot consumer references. * fix(webhook): kms:ViaService pins, https-only delivery, cross-account principal CI pin GPT-4.1 cross-family review of the policy surface (no BLOCK): FIX applied to the cross-account shoc-backend-dev Decrypt statement and both Lambda role KMS grants (the key is only ever used via Secrets Manager); its invariant-enforcement QUESTION answered durably with tests/test_cross_account_principal_pin.py (any new foreign IAM principal in cdk/ fails CI). Scanner mediums fixed: delivery.py and the replay script now refuse non-https URLs (urllib follows file:// and http://). SQS metadata-action and dynamodb:ListStreams NITs skipped: standard CDK grant shapes; ListStreams has no resource-level scoping. The 4 gitleaks HIGHs on docs/shoc-webhook-test-vectors.json are deliberate non-secrets (shared receiver-verification vectors) suppressed machine-level with justification. * harden(webhook): resolve /sh-security-review findings (1 confirmed medium + cheap fixes) High-recall detector fan-out (injection/authz/secrets-crypto/iac-iam/logic) + proof-or-kill verifier. Gate PASSES: 1 confirmed medium, 0 confirmed critical/high. Confirmed finding fixed; several unverified-but-cheap hardenings applied since the emitter ships dark and activation is weeks out. - CONFIRMED medium (confused deputy): the rotation Lambda's generated invoke permission for secretsmanager.amazonaws.com carried no SourceAccount/SourceArn, so any account's Secrets Manager could invoke the rotator. Patched the generated CfnPermission in place (a second permission would be additive, not restrictive) to pin account + this secret ARN. - delivery + replay: refuse to follow receiver 3xx redirects (no-redirect opener) so live X-SH-* auth headers can't be forwarded to a receiver-chosen Location and an http:// Location can't slip past the https guard. Fixed the "unfollowed 3xx" comment that was factually wrong. - delivery: classify 401/403 as retryable (invalidate key cache + retry in order) instead of parking -- transient auth failures (rotation outran the TTL cache, clock skew) are availability events, not contract bugs. - envelope: build_event now genuinely total (guarded eventID / ApproximateCreationDateTime subscripts) per its own never-raise contract. - handler: catch-all so an unexpected per-record error (e.g. SQS park failure) reports only that record instead of failing the whole batch (which would re-deliver every earlier success for 24h); per-invocation emit/skip batch summary so a systemic silent drop is queryable/alarmable. - rotator: narrow the AWSCURRENT-read except to ResourceNotFound/JSONDecode (transient SM/KMS errors re-raise so the overlap key isn't silently dropped); kid uniqueness checked against ALL retained kids with a random suffix on collision (never reissue a kid for a different secret). - contract: skeleton-upsert required on ANY unknown work_order_id (not just comment-before-create) + monotonicity guard (ignore older updated_at), so a parked created or an out-of-order replay can't corrupt receiver state. Unverified/refuted findings left as-is with rationale: the two "high" logic claims (whole-batch crash triggers, ordering violation) were refuted on reachability (real stream records carry required fields; persistence writes strings only; full-state idempotent upsert absorbs the ordering gap). Signed kid/version binding (AUTHZ-002) declined: coordinated contract change, not cheap, no exploit with one algorithm/key. * fix(webhook): drop kid from rotator test_ok log (CodeQL clear-text-logging FP) GHAS CodeQL flagged py/clear-text-logging-sensitive-data (high) at _test_secret's success log because head["kid"] is subscripted from the same parsed-secret dict that holds head["secret"] — the taint tracker can't tell the non-secret key id from the secret. The secret value is never logged. Rather than dismiss the alert (fragile; re-alerts on line moves), remove the flow: kid is already logged at stage time in _create_secret and version_id correlates the steps, so the test_ok log keeps only event + version_id. Also hardens against a future edit that swaps the logged field. |
||
|
|
b89e98a879
|
feat(api): Redocly lint gate + SHOC-themed /docs (Redoc theming, topbar, collapsible samples) (#130)
Some checks are pending
Deploy / deploy (push) Waiting to run
* feat(api): Add @redocly/cli as a dev dependency
Signed-off-by: Adam Moussa <adam@seahavenind.com>
* feat(api): Add Redocly configuration file with custom rules
Signed-off-by: Adam Moussa <adam@seahavenind.com>
* chore(api): Redocly lint config + bring openapi.json into compliance
redocly.yaml from the Redocly guidelines builder, with three generated
rules corrected: response-contains-property had the status codes as the
required body fields (intent was the Error schema's top-level 'error';
403 exempt since API Gateway emits AWS's {message} shape, 501 not 503);
operation-4xx-problem-details-rfc7807 off (adopting RFC 7807 would be a
runtime + SHOC-contract change, decided against); the two inert casing
rules (parameter names, schema properties) removed because both name
sets are contract-pinned (gateway resource paths, DynamoDB items).
Spec changes, no runtime impact: operationIds renamed to method-prefixed
kebab-case (get-work-orders, post-work-order-comment, ...); tags added to
all 15 operations + root tags object (groups the Redoc sidebar); examples
on all six parameters; license field; server description punctuation; two
descriptions reworded to start capitalized. Real linter catches fixed:
the two x-planned ops were missing their {workOrderId} path parameter
and any 4xx response (403 added - true today, gateway rejects unsigned).
.redocly.lint-ignore.yaml pins the six deliberate exceptions: webhook
keys are the shipped SHOC contract event names (not renameable), and the
x-planned ops answer only 501 (no 2xx to document).
package.json: npm run lint:api. Verified: lint 0 errors, 675 pytest,
headless-Chrome render of the tagged docs page.
* feat(api): SHOC design-system theme for /docs (vendored fonts)
Themes the Redoc page with the canonical SHOC token set: Montserrat 600
headings / DM Sans body / JetBrains Mono code, primary #1c75bc, navy
#262262 sidebar text + right panel, #f9fafb background, 244px sidebar.
sortRequiredPropsFirst on; 200 responses pre-expanded.
Fonts ship as lambdas/api/fonts.css (latin woff2 subsets from
@fontsource 5.3.0, embedded as data URIs, ~90KB) and inline via a new
__FONTS_CSS__ placeholder with the same </style breakout guard --
the offline single-response invariant holds, nothing fetches Google
Fonts (test-pinned). Bundling cp + bundle-consistency pin + spec-drift
asset checks extended.
Verified: headless-Chrome render (theme + fonts applied), ruff, 675
pytest, cdk synth + staged-asset check.
* feat(api): SHOC gradient topbar on /docs
64px fixed header with the SHOC shell gradient token (#1b1f52 ->
#1c4f8f -> #1c75bc), Sea Haven wordmark in Montserrat 600, page name
right-aligned in DM Sans. Redoc's scrollYOffset: 64 keeps the sticky
sidebar and anchor scrolling clear of the fixed bar. Verified via
headless-Chrome render.
* style(api): normalize /docs header and right-panel blues
The right panel's #262262 is a purple-leaning navy that clashed with
the cyan-leaning gradient, and the bar's brightest point sat directly
over the dark panel. Right panel now uses #1b1f52 (the gradient's own
dark endpoint) and the gradient runs bright-to-dark so its dark end
lands flush on the panel -- no seam, one blue family. Verified via
headless-Chrome render.
* style(api): right-panel gradient on /docs via bundle-pinned override
Redoc's theme only takes solid colors (it derives shades from
rightPanel.backgroundColor), so the gradient (#1b3d79 -> #1b3068 ->
#1b1f52, continuing the topbar blend) rides as a CSS override on the
styled-components classes of the per-section right-panel divs
(.sc-iGgWBj.sc-gsFSXq + the .sc-dExYaf stub). Those names are
deterministic for the vendored 2.5.3 bundle (verified across loads) but
change on any Redoc bump: re-derive via headless probe (find elements
whose computed background equals the rightPanel color). If they stop
matching, the panel falls back to the solid #1b1f52 theme color --
cosmetic only. Verified via headless-Chrome render.
* feat(api): collapsible samples column on /docs
Redoc CE has no built-in panel toggle, so the topbar gains a Hide/Show
samples button that flips .samples-collapsed on <html>: the right-panel
divs hide (same bundle-pinned styled-components classes as the gradient
override) and each section's content half takes the full width. Choice
persists in localStorage; aria-pressed tracks state. If the pinned
classes stop matching after a Redoc bump the toggle goes inert --
cosmetic only. Both states verified via headless-Chrome render.
* ci(api): spec-lint CI gate + npm Dependabot coverage
New spec-lint job mirrors the local npm run lint:api so openapi.json
cannot drift from redocly.yaml with green CI. Dependabot gains the npm
ecosystem (package.json is new; nothing watched @redocly/cli).
* feat(api): docs finishing touches - x-tagGroups, favicon, docs:preview
x-tagGroups sections the Redoc sidebar (Read API / Meta / SHOC Feed);
inline data-URI SVG favicon (SHOC blue) stops the browser's follow-up
/favicon.ico request 403ing at the gateway; npm run docs:preview wraps
the real-handler local render (scripts/preview_docs.py); README gains a
docs-page architecture section covering the inline pattern, theme,
pinned-selector caveat, and tooling. Lint 0 errors, 675 pytest,
headless render verified.
---------
Signed-off-by: Adam Moussa <adam@seahavenind.com>
|
||
|
|
f67d8b9907
|
feat(api): procurement-api read stack + OpenAPI docs (SHOC reconciliation path) (#127)
Some checks are pending
Deploy / deploy (push) Waiting to run
* feat(api): add procurement-api stack - read API + OpenAPI docs page Third CDK stack: API Gateway REST API (IAM SigV4) over both pipelines' tables, replacing SHOC's retired SyncController cross-account DynamoDB scan as the reconciliation/backfill path. - lambdas/api/: handler (healthcheck + docs-token gate + router dispatch), router (single route table), pagination (opaque cursor, hostile -> 400), Decimal-safe serialization, wo_repo/po_repo reads. No VendorReplies. - OpenAPI 3.1 spec as source of truth incl. top-level webhooks section documenting the outbound SHOC feed; phase-2 write endpoints x-planned (router answers 501). Self-contained /docs page, no CDN. - Auth: AWS_IAM on data routes + resource policy scoped to exactly arn:aws:iam::396287094661:role/shoc-backend-dev on GET/*; /docs and /openapi.json carve-out is token-gated in the Lambda via shared web_ui_auth (fail-closed, INFRA-74 posture). - KMS: explicit Decrypt/DescribeKey on the DynamoDB CMK from SSM (name-imported table drops the key association - INFRA-104 class). - Alarms: errors/throttles/duration(p99>=22.5s) + gateway 5xx, ALARM-only to site-alerts. No access logging in v1 (docs ?token= shim stays out of logs); cloud_watch_role=False. - Tests: handler auth-seam + routing + Decimal round-trip; moto cursor pagination incl. hostile cursors; spec<->router drift gate; bundle AST pins for the api command; pytest.ini --cov + loader siblings. - Deploy role: third stack DescribeStacks ARN + procurement-api smoke invoke ARN (re-run create-deploy-role.sh before merge). * harden(api): apply sh-security-review findings to procurement-api Fan-out (6 detectors) + review findings resolved: Correctness / DoS: - pagination: require EXACT key-set match (was subset) so a partial/foreign composite cursor can't reach DynamoDB as an inconsistent ExclusiveStartKey -> ValidationException -> 500; comments Query now pins the cursor's work_order_id to the path entity. - handler: map botocore ValidationException to 400 (defense in depth) so a crafted cursor can't drive the zero-threshold 5xx alarm. - web_ui_auth: compare tokens as bytes; a non-ASCII presented token now fails closed (401) instead of crashing hmac.compare_digest into a 500. Resolves the pre-existing xfail(strict) follow-up test; hardens the web UIs too. Docs page: - typeStr() now escapes the one spec-derived string that reached innerHTML. - spec inlined into the docs <script> block escapes "<" -> < (</script> breakout guard); /openapi.json still served byte-faithful. - Cache-Control: no-store + Referrer-Policy: no-referrer on docs responses so the ?token= URL stays out of caches/Referer. - spec-drift test asserts the committed spec carries no "</" / "<!--". IAM / IaC: - resource policy enumerates the 7 data GET resources instead of GET/* so a future GET route can't silently inherit SHOC cross-account reach. - kms:Decrypt grant gains a kms:ViaService=dynamodb condition. - stage throttling (50 rps / 100 burst) bounds the unauthenticated /docs blast radius below the 10k account default. - corrected the PATCH/POST comment (same-account callers aren't blocked by the resource policy; 501 handler + absent write grant are the gate). - documented the RETAIN log-group first-deploy rollback trap and the resource-policy-needs-redeploy gotcha in-stack. Mandatory GPT-4.1 cross-family review of the full policy surface: no BLOCK/FIX. 675 tests pass, ruff clean, cdk synth green. |
||
|
|
073201f633
|
Migrate to seahaven-prod: deploy role, backfill tooling, account-portability fixes (#125)
* feat(migration): prepare stacks and tooling for the seahaven-prod account move
Phase 1 of the mgmt (328440206208) -> seahaven-prod (011934824531)
migration. No behavior change in-account; everything here is additive or
account-portability hygiene:
- infra/deploy-role/: reviewed OIDC deploy-role artifacts for prod
(trust main-only, cdk-hnb659fds-* AssumeRole, smoke-invoke-lambda scoped
to exactly the two email-processor fn ARNs). Codifies the previously
out-of-band smoke-invoke grant.
- Table resource policies: make_slack_bot_read_policy in cdk/common.py,
applied to purchase-orders, verified-sites, WorkOrders,
WorkOrderComments (NOT pending-site-review; no bot consumer). Grants the
mgmt-resident seahaven-slack-bot roles read-only cross-account access
post-move (bot-side identity grants land in the slack-bot repo).
- scripts/migrate_tables.py: dry-run-default backfill tool implementing
the plan's per-table semantics (superset overwrite, ingested_at cutoff
for WorkOrderComments, backup-gated truncate-and-load for the two
site tables) plus a verify subcommand (count parity, spot checks,
sticky-Cancelled drift check).
- tests/test_resource_policy_helper.py: statement-shape unit tests +
static pins that exactly the four bot-read tables carry the policy.
- Account-literal fixes: account-agnostic fixture bucket in
test_reprocess_contract; runbook/README/po-template-parser account
references updated to prod with historical mgmt notes; README gains the
account-prerequisites list (imported-by-name dependencies).
deploy.yaml is deliberately unchanged (push-to-main auto-deploy kept).
Merge is held until migration Phase 0 completes; flipping the
AWS_DEPLOY_ROLE_ARN repo secret and merging this PR IS the first prod
deploy.
* fix(migration): verify backup AVAILABLE pre-truncate; document wildcard risk acceptance (cross-review FIX/NIT)
* refactor(migration): drop cross-account read grants (slack-bot decommissioned); harden backfill + deploy role
seahaven-slack-bot was decommissioned 2026-07-23 (stack DELETE_IN_PROGRESS,
consumer Lambdas gone); its successor sh-mcp is undeployed and uses
same-account DynamoDB access. So no live consumer reads these tables
cross-account. Per Adam's call, drop the cross-account grants entirely and
re-add correctly-scoped ones if/when sh-mcp deploys to a different account.
- Remove the four table resource policies + make_slack_bot_read_policy helper
+ its constants (cdk/common.py, po_stack.py, wo_stack.py) and the helper's
unit test. Both stacks synth with zero table ResourcePolicy.
- scripts/migrate_tables.py hardening (fixes from the sh-security-review
fan-out on the destructive backfill tool):
* validate --cutoff strictly (parse ISO-8601, require aware UTC, re-emit
canonical second-precision form) so a malformed cutoff can't silently
copy dual-window rows or drop history;
* reject `copy --all` up front (must run tables individually, in order,
with the stream-drain wait) instead of writing three tables then erroring;
* truncate backup gate now also checks recency (<1h) and TableId, not just
status+name;
* verify requires --cutoff whenever a cutoff table is in scope (else it
false-flags dual-window rows as MISSING);
* sticky-cancel is now PREVENTED copy-side (a non-Cancelled source item
never overwrites a dest-Cancelled PO), and the verify comment no longer
overstates what its source-side scan covers;
* spot-check all modes (truncate_load keys are verbatim, so key-existence
is sound there too).
- Deploy role: scope cloudformation:DescribeStacks to this repo's stacks +
CDKToolkit (was Resource:*, disclosed all tenant stacks in the shared prod
account); add a drift check warning on unexpected role policies and drop the
dead SMOKE_POLICY_NAME var; document the shared-account bootstrap-role
accepted risk in the deploy-role README.
* docs(deploy-role): fold in cross-review NITs (DescribeStacks maintenance note, warn-only drift rationale)
* ci: update workflow to use new workflow tag (ruff versioning fix)
* fix(migration): address Open SWE review findings on migrate_tables.py
- Validate the truncate backup on dry-run as well as --execute so a
missing/stale/wrong-incarnation --backup-arn surfaces on the rehearsal
run (finding f_24a48b8900).
- Assert configured keys match the live key schema of both tables before
any key projection, turning config/schema drift into a descriptive
abort instead of a mid-backfill KeyError (finding f_cb6b5a6c59).
- Clarify why key-existence spot-checks are sound for WorkOrderComments:
the copy Puts source items verbatim and the sample uses the same
cutoff filter, so per-account comment_id divergence never enters the
check (finding f_390b7d6c3b is a false positive; comment hardened).
|
||
|
|
75fe91c198
|
Ops/recovery tooling + dependency hygiene (refactor phase 7) (#110)
* feat: ops/recovery tooling + dependency hygiene (refactor phase 7) Generalize scripts/reprocess.py from a PO-only full-sweep script into a pipeline-general recovery tool. Targeted replay (--key/--prefix/--since) is now the default, and the full inbound/ sweep is demoted behind an explicit --all that documents its five hazards (async concurrency does not serialize, use RequestResponse if order matters, metric double-count, Bedrock re-bill, out-of-order field regression). --pipeline po|wo resolves the correct function + bucket; dry-run-by-default / --execute is preserved. A new tests/test_reprocess_contract.py pins the synthetic S3 event shape and asserts the raw list_objects_v2 key is emitted untransformed (the handler is the single decode point; a pre-decoded key would corrupt keys containing spaces or '+'). Add docs/runbook-dlq-recovery.md: the async on-failure DLQ has no console redrive-to-source, so it documents the receive -> extract key -> targeted reprocess --key -> verify -> purge procedure, the real recovery windows (14-day DLQ breadcrumb, 90-day raw-email S3 that overrides the table RETAIN policy and is the true replay floor), and that sender-auth and ai_fallback_rejected drops are fail-closed skips that never reach the DLQ. Linked from the README alarms and scripts sections. Drop the vendored boto3 floor pin from both email-processor requirements (the Lambda runtime provides boto3; lambda-template.md empty-with-comment form). With nothing left to install, the email-processor bundling becomes cp-only -- the whole pip step is removed, which is the only acceptable way the manylinux2014_aarch64 pin disappears (removing the pin while keeping a pip install caused the PR #34 x86-wheel outage). Exact-pin moto==5.2.2 and add pinned po/web_ui + po/site_extractor manifests (excluded from their bundles, so hash-neutral) so their new Dependabot entries have something to act on; add Dependabot entries for /tests, /lambdas/po/web_ui, and /lambdas/po/site_extractor. cdk diff is confined to exactly the two email processors' asset hashes on both stacks. The wo/web_ui dead-manifest reduction was deliberately left out: that manifest already ships inside the plain (non-bundled) WebUI asset on main, so reducing or excluding it would redeploy workorder-web-ui for no functional change -- deferred to keep the blast radius to the two intended targets. The untracked 44 MB lambdas/po/email_processor/package/ dir was removed from the filesystem (asset-hash-neutral given Phase 2's package/ exclude); it is untracked, so there is nothing to commit for it. * Reject --all combined with --prefix/--since in reprocess.py --all is a distinct mode (the demoted full-prefix sweep), but the args.all branch unconditionally set prefix=inbound/ and since=None, so passing it alongside a narrower selector silently discarded that selector. `--all --since 2026-07-01` swept the entire corpus instead of the bounded window, triggering every documented --all hazard (Bedrock re-bill, metric double- count, merged-field regression) on objects the operator never targeted -- contradicting the tool's safety goal. Add the missing mutual-exclusion guard alongside the existing --key one, and pin --all+--prefix, --all+--since, and all three together as argparse rejections. |
||
|
|
cb5539bd68
|
feat: deploy-pipeline guards — healthcheck, smoke gate, bundle glob + AST test (refactor phase 0) (#107)
Some checks are pending
Deploy / deploy (push) Waiting to run
* feat: deploy-pipeline guards — healthcheck, smoke gate, bundle glob + AST test (refactor phase 0)
Deploys of po-email-processor and workorder-email-processor had no
verification step, so an init-time ImportError in the bundled zip
could ship silently and only surface on the next real S3 event. This
adds a synchronous post-deploy smoke gate wired into the deploy
workflow: both Lambdas are invoked with {"healthcheck": true} and the
FunctionError field is checked, since an Unhandled init error still
returns HTTP 200 on RequestResponse invokes and would false-pass a
plain exit-code check.
The healthcheck branch is the first statement in each handler, before
any boto3/S3 use or ses_auth, and only fires on a top-level direct
invoke ("healthcheck" is not a key AWS ever sets on a real S3
ObjectCreated event, so mail content can't reach this path). It emits
no EMF metrics and no log text that could match the
sender-auth-rejected metric filter, so two deploys in one window
won't trip the alarm.
Separately, the PO stack's asset bundling copied a hand-maintained
four-file allowlist into the zip, so every new sibling module
handler.py imports had to be added by hand or the deploy shipped a
Lambda that ImportErrors at cold start (bit us for template_parser in
PR #105 and nearly for derived_fields in PR #2). Replaced it with a
non-recursive ./*.py glob so top-level source files ship
automatically while tests/ and the stale package/ dir still cannot,
and added an AST-based bundle-consistency test that parses each
handler's first-party imports and fails CI if the bundling command
would omit any of them (a revert to an incomplete allowlist, or code
moved into a subdirectory the glob doesn't cover).
Includes the refactor-evaluation report that scoped this phase.
* fix: review nits — unambiguous bundling-command extraction, smoke payload-parse message, dead asserts
- tests/test_bundle_consistency.py: _extract_bundling_command now collects
all command=[...] matches and demands exactly one per stack file, instead
of silently returning whichever ast.walk visits first if a second bundled
function is ever added.
- scripts/post-deploy-smoke.sh: distinguish an unparseable response payload
from a payload mismatch so the failure message says what actually happened
(the previous "could not parse" branch was unreachable — the inline python
always exited 0).
- test_po_healthcheck.py: drop the substring assertions on stdout that were
dead behind the stricter `captured.out == ""` assertion; keep the stderr
filter-pattern check.
Review follow-up on PR #107; no behavior change to any shipped code path.
|
||
|
|
acc1961d21
|
feat: template-first WO parser + Bedrock fallback, PO Bedrock switch (#99)
* Add deterministic template parser for WO emails The workorder-email-processor sends every one of ~22.9k emails/month to an LLM, but ~93.6% are the plain-text "AMAZON UPDATE WO DETAILS" comment template and ~6.4% the HTML "AMAZON assign Work Order" template. Parse those two shapes deterministically, offline, so the AI call is reserved for the long tail. The module is pure (no boto3, no network). try_deterministic_parse classifies by subject, extracts the shared contract fields, and returns a result ONLY when it passes a strict fail-closed validation gate: exact contract-key set, subject/id agreement, the literal "Work Order: <id>" double space, per-type required fields, site-code shape, and a label-bleed guard so a value that over-ran into the next field fails. Any miss, drift, or extractor exception yields None so the caller falls back to the AI extractor -- data is never corrupted, only the fallback rate rises. Refs: #23 * Migrate WO processor to Bedrock and fix comment_id collision Switch the AI path from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so parsing no longer needs a provider API key or Secrets Manager secret. The EXTRACTION_PROMPT and JSON contract are kept byte-identical, so the AI-fallback output is unchanged. Try the new deterministic template parser first and only call Bedrock on a miss/invalid result. Fix issue #23: the WorkOrderComments range key was work_order_id#<comment_time>, so two emails on one WO with an identical or absent comment time collided and overwrote each other. Derive a 12-hex suffix from the S3 object key alone -- deterministic, so an async retry of the same object is byte-identical (idempotent) while distinct emails get distinct keys -- and keep wall-clock now() out of the key (literal 'nocomment' segment when comment_time is absent). Also emit one CloudWatch EMF line per record (Seahaven/WorkorderIngest ParseOutcome, dimensioned by ParseMethod/TemplateId) for parse-outcome observability, replace the deprecated datetime.utcnow() with datetime.now(timezone.utc), and drop the anthropic dependency. Refs: #23 * Migrate PO processor to Bedrock Switch the PO email processor's AI extraction from the Anthropic SDK to bedrock-runtime InvokeModel on the inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0 (BEDROCK_MODEL_ID env), so it no longer needs a provider API key or Secrets Manager secret. PO parsing stays fully AI -- only the provider changes. The EXTRACTION_PROMPT is kept byte-identical and the Bedrock text output is still decoded with json.loads(..., parse_float=Decimal), which DynamoDB requires (it rejects floats). Replace the deprecated datetime.utcnow() with datetime.now(timezone.utc) and drop the anthropic dependency. * Grant Bedrock IAM, drop Anthropic secrets, add fallback alarm Both stacks moved their processors from the Anthropic API to the Bedrock inference profile us.anthropic.claude-haiku-4-5-20251001-v1:0. Grant each processor role bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream on BOTH the inference-profile ARN AND the per-region foundation-model ARNs for us-east-1/us-east-2/us-west-2 (empty-account) -- the us.* profile routes cross-region, so a profile-only grant AccessDenies at runtime. Remove both anthropic-api-key Secret constructs, their grant_read, and the ANTHROPIC_API_KEY_SECRET_ARN env; add BEDROCK_MODEL_ID. The secrets had RemovalPolicy.RETAIN so they are orphaned, not deleted -- flagged in the README for manual post-deploy deletion and key revocation. Add the workorder-email-processor-template-fallback-rate alarm: a FILL(0) + >=10-sample volume-floor MathExpression over the EMF ParseOutcome metric (15-min periods) that pages when the AI-fallback share exceeds 15% sustained, catching Hexagon template drift. ALARM-only SnsAction to site-alerts, no OK action, NOT_BREACHING, matching the existing stack idiom. * Add offline WO parser test suite Cover the deterministic parser with golden-file tests over 55 real scrubbed .eml fixtures (both comment sub-shapes, username Submitted-By, address present/absent, br+CRLF assign addresses), fail-closed validation-gate rules, adversarial and prompt-injection cases that must route to ai_fallback or parse without corrupting other fields, the issue #23 comment_id idempotency invariants, and the Bedrock-fallback dispatch plus EMF-metric emission with a mocked invoke_model. Extend pytest.ini testpaths to discover the co-located suite, and update tests/conftest.load_handler to put a handler's own directory on sys.path so the WO handler's new `from template_parser import ...` resolves under the existing shared handler tests. Point test_local.py at the new template-first + Bedrock flow. Refs: #23 * Document Bedrock migration and WO parse flow in README Record the provider switch to the Bedrock inference profile (no Anthropic API key or Secrets Manager secret, with the retired secrets flagged for manual deletion), the WO deterministic-template-first + AI-fallback flow, the new ParseOutcome EMF metric and template-fallback-rate alarm, the issue #23 comment_id format change, the +00:00 aware-UTC timestamp shift, and offline test instructions. Refs: #23 * Fix f-string lint and formatting in backfill scripts Drop the f prefix from two f-strings that carry no placeholders (F541) and apply ruff format, so `ruff check` / `ruff format --check` pass in CI. * Emit ParseMethod-only EMF set so fallback alarm can fire The fallback-rate alarm queries the ParseOutcome series keyed on ParseMethod alone, but the emitter published only the joint (ParseMethod, TemplateId) dimension set. CloudWatch materializes exactly the listed dimension sets and does not auto-aggregate, so the alarm's series never received data: it evaluated a constant 0 and could never page on template-drift coverage collapse. Publish both ["ParseMethod"] and ["ParseMethod","TemplateId"] and update the EMF regression test to assert both sets are present. * Commit WO parser .eml fixtures for executable coverage The parser test suite globbed for input .eml fixtures that the repo's `*.eml` ignore rule kept uncommitted, so every parametrized golden and fail-closed test collected zero cases and CI could not exercise the deterministic parser that handles 100% of WO email volume. Add a fixtures-only negation to .gitignore and commit the 55 scrubbed positive samples (50 update-plaintext, 5 assign-html) plus 14 ai-fallback and 3 adversarial fixtures. The ai-fallback set covers each fail-closed reason code (subject_no_match, single_space_work_order, malformed_site_code, label_bleed, creation_time_unparseable, wo_id_mismatch, missing_required_field) and the adversarial set proves the parser is total and confines prompt-injection payloads to comment_text without steering the structured fields. * Fix WO parser advisories A1-A3 (PR #99 follow-ups) A1 — AI-fallback comment_id nondeterminism: parsed comment_time is model output and not stable across Lambda async retries, so on the ai_fallback path the comment_id range-key time segment now derives from the email Date header (deterministic per S3 object) instead of the model's comment_time. The template path is unchanged (its comment_time is a pure function of the raw email). Bedrock invoke pins temperature 0 so retries reproduce the same extraction. Closes the #23 reopening on the AI path. A2 — EMF record now carries the spec-required _aws.Timestamp (epoch ms) so CloudWatch reliably extracts the ParseOutcome datapoint that the fallback-rate alarm depends on. A3 — T1 New Comment capture no longer truncates at the first blank line; multi-paragraph comments are captured through internal blanks and terminate at the next label/separator. 17 golden files regenerated from the real fixtures accordingly. Hardening from the sh-security-review pass on this diff: - _header_date_iso is total: OverflowError/OSError from an extreme Date header fall back to 'nocomment' instead of failing the invocation. - _capture_block trims blanks in O(n) (no pop(0)) — removes a quadratic path on a crafted large blank run. - work_order_id is enforced digits-only on BOTH parse paths before it is used as a DynamoDB key, so prompt-injected AI output cannot forge '#' range-key segments or land on an arbitrary WO. |
||
|
|
5112c1345b
|
Merge workorder-ingest into unified procurement repo (#22)
* Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers |
||
|
|
ec416079f5 |
Add verified-sites pipeline via DynamoDB Streams
Enable DynamoDB Streams on purchase-orders table and add a site-extractor Lambda that extracts Amazon facility codes and addresses from PO ship-to data, upserting them into a new verified-sites table. Includes a backfill script for existing POs and upgrades existing Lambdas to arm64 + 60-day log retention. |
||
|
|
fc690dd958 |
Initial commit: PO email ingestion pipeline
CDK stack with SES receipt rule, S3 bucket, email processor Lambda (Claude-powered extraction), web UI Lambda with Function URL, and DynamoDB for storage. Includes reprocessing script for missed emails. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> |