mirror of
https://github.com/Sea-Haven-Industries/procurement-ingest.git
synced 2026-09-30 17:43:14 +00:00
* Merge workorder-ingest pipeline into unified repo Move PO lambdas under lambdas/po/, add WO pipeline under lambdas/wo/. Two independent CloudFormation stacks in one CDK app. Fix WO stack compliance: ARM64 architecture, 60-day log retention, aarch64 bundling, RETAIN on Anthropic secret. Remove stale CodePipeline buildspec. * Fix test_local.py import path and remove dead shared/models.py test_local.py referenced the old lambdas/email_processor path. Updated to lambdas/wo/email_processor. Removed shared/ directory entirely as nothing imports from it. * Escape HTML in both web UI dashboards to prevent XSS Both Function URLs are public (auth_type=NONE) and render email-derived content via f-strings. Attacker-crafted emails could inject scripts. Added html.escape() on all interpolated values in both PO and WO dashboards. * Add pagination to WO web UI scan get_work_orders() only fetched the first 1MB page from DynamoDB. Loop on LastEvaluatedKey to match the PO web UI pattern. * Fix esc(None) TypeError and javascript: scheme in PO web UI Coerce supplier name through `or ""` before escaping to handle nested None from DynamoDB. Add scheme allowlist on view_order_url to block javascript:/data: hrefs from LLM-extracted URLs. * Fix WO render_badge None guard, updated_at slice, and backfill path Add null guard to WO render_badge matching the PO version. Use `or ""` before slicing updated_at to handle explicit None values. Fix backfill_sites.py sys.path to use new lambdas/po/site_extractor. * Harden WO web UI and fix JS-context XSS in both dashboards - Use json.dumps for onclick URLs to prevent JS string breakout - Add .lower() to WO render_badge color lookup matching PO pattern - Add pagination to get_comments query - Cap get_work_orders to 500 results matching PO pattern * Apply ruff formatting to web UI handlers
47 lines
1.2 KiB
Python
47 lines
1.2 KiB
Python
#!/usr/bin/env python3
|
|
"""
|
|
Local test script - parses sample emails through Claude without AWS.
|
|
|
|
Usage:
|
|
export ANTHROPIC_API_KEY=sk-ant-...
|
|
python test_local.py
|
|
"""
|
|
|
|
import json
|
|
import sys
|
|
from pathlib import Path
|
|
|
|
sys.path.insert(0, str(Path(__file__).parent / "lambdas" / "wo" / "email_processor"))
|
|
|
|
from handler import parse_raw_email, extract_with_claude
|
|
|
|
|
|
def main():
|
|
samples_dir = Path(__file__).parent / "samples"
|
|
eml_files = list(samples_dir.glob("*.eml"))
|
|
|
|
if not eml_files:
|
|
print("No .eml files found in samples/")
|
|
return
|
|
|
|
for eml_path in eml_files:
|
|
print(f"\n{'='*80}")
|
|
print(f"FILE: {eml_path.name}")
|
|
print(f"{'='*80}")
|
|
|
|
raw = eml_path.read_bytes()
|
|
email_data = parse_raw_email(raw)
|
|
|
|
print(f"Subject: {email_data['subject']}")
|
|
print(f"From: {email_data['sender']}")
|
|
print(f"Date: {email_data['date']}")
|
|
print(f"Body preview: {email_data['body'][:200]}...")
|
|
print()
|
|
|
|
print("Sending to Claude for extraction...")
|
|
parsed = extract_with_claude(email_data)
|
|
print(json.dumps(parsed, indent=2))
|
|
|
|
|
|
if __name__ == "__main__":
|
|
main()
|