* Merge workorder-ingest pipeline into unified repo
Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/.
Two independent CloudFormation stacks in one CDK app. Fix WO stack
compliance: ARM64 architecture, 60-day log retention, aarch64 bundling,
RETAIN on Anthropic secret. Remove stale CodePipeline buildspec.
* Fix test_local.py import path and remove dead shared/models.py
test_local.py referenced the old lambdas/email_processor path. Updated
to lambdas/wo/email_processor. Removed shared/ directory entirely as
nothing imports from it.
* Escape HTML in both web UI dashboards to prevent XSS
Both Function URLs are public (auth_type=NONE) and render
email-derived content via f-strings. Attacker-crafted emails
could inject scripts. Added html.escape() on all interpolated
values in both PO and WO dashboards.
* Add pagination to WO web UI scan
get_work_orders() only fetched the first 1MB page from DynamoDB.
Loop on LastEvaluatedKey to match the PO web UI pattern.
* Fix esc(None) TypeError and javascript: scheme in PO web UI
Coerce supplier name through `or ""` before escaping to handle
nested None from DynamoDB. Add scheme allowlist on view_order_url
to block javascript:/data: hrefs from LLM-extracted URLs.
* Fix WO render_badge None guard, updated_at slice, and backfill path
Add null guard to WO render_badge matching the PO version. Use
`or ""` before slicing updated_at to handle explicit None values.
Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor.
* Harden WO web UI and fix JS-context XSS in both dashboards
- Use json.dumps for onclick URLs to prevent JS string breakout
- Add .lower() to WO render_badge color lookup matching PO pattern
- Add pagination to get_comments query
- Cap get_work_orders to 500 results matching PO pattern
* Apply ruff formatting to web UI handlers
* Add CI workflow and apply ruff formatting
* Disable cdk synth — email_processor uses pre-built package dir
The email_processor Lambda bundles deps into a gitignored package/
directory. cdk synth fails in CI without a build step to recreate it.
Disabling until packaging is standardized.
* Use CDK BundlingOptions for email_processor Lambda packaging
Replaces the pre-built gitignored package/ directory with CDK's
built-in bundling. Deps are now installed inside a Docker container
during cdk synth, so the build works identically locally and in CI.
Re-enables run-cdk-synth in the CI workflow.
Replaces extraction prompt with domain-specific rules: trade
classification taxonomy (23 categories), site_code skip list,
zip padding, revision email type, and structured extraction for
fiscal_year, trade, and coupa_category.
Handler changes:
- New "revision" email type overwrites existing PO via put_item
- enrich_parsed() adds top-level state, ship_to_raw, data_source
- pad_zip() zero-pads short zip codes (e.g., "7001" → "07001")
- Removed invoice_total/invoice_count (Payee Central only)
Web UI: added revision badge, new detail fields (site code, state,
trade, fiscal year, coupa category, data source), line item table
now shows Qty/Unit/Price columns, list view shows Site and Trade.
CDK: fixed StreamViewType to match deployed table (NEW_IMAGE).
README: documented PO record schema and revision flow.
The Claude extraction prompt now explicitly asks for site_code (the
Amazon facility code) and structured ship_to address fields (street,
city, state, zip). The site-extractor Lambda prefers these direct
fields when available, falling back to regex for older PO records.
CDK stack with SES receipt rule, S3 bucket, email processor Lambda
(Claude-powered extraction), web UI Lambda with Function URL, and
DynamoDB for storage. Includes reprocessing script for missed emails.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>