Coupa PO email ingestion pipeline
Find a file
dependabot[bot] 636177ecfd
Update boto3 requirement in /lambdas/email_processor
Updates the requirements on [boto3](https://github.com/boto/boto3) to permit the latest version.
- [Release notes](https://github.com/boto/boto3/releases)
- [Commits](https://github.com/boto/boto3/compare/1.35.0...1.43.2)

---
updated-dependencies:
- dependency-name: boto3
  dependency-version: 1.43.2
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-05-02 21:17:26 +00:00
.github Add Dependabot version update configuration 2026-05-02 17:14:45 -04:00
cdk Align PO schema with enriched records and improve extraction prompt 2026-05-01 19:53:26 -04:00
lambdas Update boto3 requirement in /lambdas/email_processor 2026-05-02 21:17:26 +00:00
scripts Add verified-sites pipeline via DynamoDB Streams 2026-04-30 14:26:53 -04:00
.gitignore Initial commit: PO email ingestion pipeline 2026-04-07 12:12:30 -04:00
buildspec.yml Add CI/CD pipeline and fix stack name to kebab-case (#3) 2026-05-01 19:17:19 -04:00
README.md Align PO schema with enriched records and improve extraction prompt 2026-05-01 19:53:26 -04:00

PO Ingest

Coupa purchase-order email ingestion pipeline. SES receives Amazon PO emails, Claude extracts structured data, and the result lands in the shared purchase-orders DynamoDB table.

Flow

  1. Coupa sends a PO email to amazon_po@int.seahaven.com.
  2. SES (using the shared INBOUND_MAIL rule set) drops the raw MIME into s3://po-ingest-emails-{AccountId}/inbound/.
  3. S3 ObjectCreated fires the po-email-processor Lambda.
  4. The Lambda parses the email, sends it to Claude Haiku 4.5 for structured JSON extraction, and writes to DynamoDB.
    • email_type: new_po — conditional PutItem on purchase-orders (idempotent on po_number).
    • email_type: revision — unconditional PutItem overwriting the existing record with updated data.
    • email_type: cancellation — UpdateItem marking the existing row Cancelled.
  5. DynamoDB Streams (NEW_IMAGE) on purchase-orders feeds two downstream consumers:
    • LedgerFlow (seahaven-slack-bot/po-sync) — daily KB sync.
    • Verified-sites pipeline (po-ingest-site-extractor) — real-time site address extraction (see below).

The extraction prompt includes domain-specific rules for site code identification (with a skip list for false positives like RME, BBM, JLL), trade classification across 23 categories (Plumbing PM/Reactive, Electrical, HVAC, Dock Doors, etc.), fiscal year derivation, and ship-to address parsing with zip code zero-padding.

A separate po-web-ui Lambda (Function URL, unauthenticated) renders a simple HTML dashboard scanning the table.

PO record schema

Each record in purchase-orders includes:

  • Core: po_number (PK), email_type (new_po/revision/cancellation), po_status, source_system
  • People/dates: submitted_by, on_behalf_of, order_date, revision_date, payment_terms, requisition_number, department
  • Site: site_code, state (top-level), ship_to (structured), ship_to_raw (original text)
  • Classification: trade, fiscal_year, coupa_category
  • Financials: total_amount, currency, line_items[] (with description, amount, quantity, unit, price, need_by)
  • Metadata: data_source ("email" or "email+payee_scrape"), email_subject, processed_at, raw_s3_key

Verified-sites pipeline

The po-ingest-site-extractor Lambda is triggered by the DynamoDB Stream on every PO INSERT/MODIFY. It:

  1. Extracts an Amazon facility site code from ship_to.name using a regex cascade (parentheses, LLC - CODE, Station CODE, DS - CODE) with a fallback to the first line_items description.
  2. Parses ship_to.address into structured fields (street, city, state, zip).
  3. Upserts to the verified-sites DynamoDB table — atomically increments poCount and appends the PO number to sourcePOs.

POs with no extractable site code fall through to an address reverse-lookup against the verified-sites cache (normalized street + zip). If still unresolved, the PO is written to the pending-site-review table for manual verification against Payee Central.

Backfill stats (initial run): 14,825 POs scanned → 9,900 with extractable site codes → 1,100 unique sites.

Architecture

  • IaC: AWS CDK (Python), stack name po-ingest, region us-east-1.
  • Lambdas (all Python 3.12, arm64, 60-day log retention):
    • po-email-processor — S3-triggered, parses PO emails via Claude Haiku.
    • po-web-ui — Function URL, HTML dashboard.
    • po-ingest-site-extractor — DynamoDB Streams-triggered, extracts site addresses.
  • Storage:
    • S3 po-ingest-emails-{AccountId} — 90-day lifecycle expiry.
    • DynamoDB purchase-orders — owned by this stack, Streams enabled (NEW_AND_OLD_IMAGES).
    • DynamoDB verified-sites — PK siteCode, GSI by-state on state.
    • DynamoDB pending-site-review — PK po_number. POs with no extractable site code and no address match, awaiting manual Payee Central verification.
  • Secrets: Anthropic API key in Secrets Manager at po-ingest/anthropic-api-key.
  • SES: adds the PoEmailRule to the existing INBOUND_MAIL receipt rule set (shared with workorder-ingest).
  • CI/CD: CodePipeline V2 (po-ingest-pipeline) → CodeBuild (po-ingest-build). Pushes to main auto-deploy via buildspec.yml.

CI/CD

Merges to main trigger the po-ingest-pipeline (CodePipeline V2) which runs CodeBuild to cdk deploy. The pipeline uses the existing CodeStar connection to the Sea-Haven-Industries GitHub org.

Branch protection: main requires a PR (no direct push), no deletion, no force push.

Setup

  1. Bootstrap CDK in the account if you haven't already: cdk bootstrap aws://{AccountId}/us-east-1.
  2. Store the Anthropic API key:
    aws secretsmanager create-secret \
      --name po-ingest/anthropic-api-key \
      --secret-string "sk-ant-..."
    
  3. Install Lambda dependencies into the deployable package directory (gitignored):
    pip install -r lambdas/email_processor/requirements.txt -t lambdas/email_processor/package/
    
  4. Deploy:
    cd cdk
    pip install -r requirements.txt
    cdk deploy
    
  5. The WebUIUrl CloudFormation output is the dashboard URL.

Reprocessing

To re-run the processor against every email still sitting in inbound/ (useful after a parser change):

python scripts/reprocess.py            # dry-run — lists keys
python scripts/reprocess.py --execute  # invokes po-email-processor for each

Inserts are conditional on po_number, so re-processing existing POs is a no-op.

Backfilling verified sites

The stream Lambda handles all future POs automatically. To backfill from historical PO data (one-time):

python scripts/backfill_sites.py

Uses the same extraction logic as the Lambda. Idempotent — safe to re-run.