2026-04-20 19:31:22 -04:00
# PO Ingest
Coupa purchase-order email ingestion pipeline. SES receives Amazon PO emails, Claude extracts structured data, and the result lands in the shared `purchase-orders` DynamoDB table.
## Flow
1. Coupa sends a PO email to `amazon_po@int.seahaven.com` .
2. SES (using the shared `INBOUND_MAIL` rule set) drops the raw MIME into `s3://po-ingest-emails-{AccountId}/inbound/` .
3. S3 `ObjectCreated` fires the `po-email-processor` Lambda.
4. The Lambda parses the email, sends it to Claude Haiku 4.5 for structured JSON extraction, and writes to DynamoDB.
- `email_type: new_po` — conditional `PutItem` on `purchase-orders` (idempotent on `po_number` ).
Align PO schema with enriched records and improve extraction prompt
Replaces extraction prompt with domain-specific rules: trade
classification taxonomy (23 categories), site_code skip list,
zip padding, revision email type, and structured extraction for
fiscal_year, trade, and coupa_category.
Handler changes:
- New "revision" email type overwrites existing PO via put_item
- enrich_parsed() adds top-level state, ship_to_raw, data_source
- pad_zip() zero-pads short zip codes (e.g., "7001" → "07001")
- Removed invoice_total/invoice_count (Payee Central only)
Web UI: added revision badge, new detail fields (site code, state,
trade, fiscal year, coupa category, data source), line item table
now shows Qty/Unit/Price columns, list view shows Site and Trade.
CDK: fixed StreamViewType to match deployed table (NEW_IMAGE).
README: documented PO record schema and revision flow.
2026-05-01 19:40:18 -04:00
- `email_type: revision` — unconditional `PutItem` overwriting the existing record with updated data.
2026-04-20 19:31:22 -04:00
- `email_type: cancellation` — `UpdateItem` marking the existing row `Cancelled` .
Align PO schema with enriched records and improve extraction prompt
Replaces extraction prompt with domain-specific rules: trade
classification taxonomy (23 categories), site_code skip list,
zip padding, revision email type, and structured extraction for
fiscal_year, trade, and coupa_category.
Handler changes:
- New "revision" email type overwrites existing PO via put_item
- enrich_parsed() adds top-level state, ship_to_raw, data_source
- pad_zip() zero-pads short zip codes (e.g., "7001" → "07001")
- Removed invoice_total/invoice_count (Payee Central only)
Web UI: added revision badge, new detail fields (site code, state,
trade, fiscal year, coupa category, data source), line item table
now shows Qty/Unit/Price columns, list view shows Site and Trade.
CDK: fixed StreamViewType to match deployed table (NEW_IMAGE).
README: documented PO record schema and revision flow.
2026-05-01 19:40:18 -04:00
5. DynamoDB Streams (NEW_IMAGE) on `purchase-orders` feeds two downstream consumers:
2026-04-30 14:30:54 -04:00
- **LedgerFlow** (`seahaven-slack-bot/po-sync` ) — daily KB sync.
- **Verified-sites pipeline** (`po-ingest-site-extractor` ) — real-time site address extraction (see below).
2026-04-20 19:31:22 -04:00
Align PO schema with enriched records and improve extraction prompt
Replaces extraction prompt with domain-specific rules: trade
classification taxonomy (23 categories), site_code skip list,
zip padding, revision email type, and structured extraction for
fiscal_year, trade, and coupa_category.
Handler changes:
- New "revision" email type overwrites existing PO via put_item
- enrich_parsed() adds top-level state, ship_to_raw, data_source
- pad_zip() zero-pads short zip codes (e.g., "7001" → "07001")
- Removed invoice_total/invoice_count (Payee Central only)
Web UI: added revision badge, new detail fields (site code, state,
trade, fiscal year, coupa category, data source), line item table
now shows Qty/Unit/Price columns, list view shows Site and Trade.
CDK: fixed StreamViewType to match deployed table (NEW_IMAGE).
README: documented PO record schema and revision flow.
2026-05-01 19:40:18 -04:00
The extraction prompt includes domain-specific rules for site code identification (with a skip list for false positives like RME, BBM, JLL), trade classification across 23 categories (Plumbing PM/Reactive, Electrical, HVAC, Dock Doors, etc.), fiscal year derivation, and ship-to address parsing with zip code zero-padding.
2026-04-20 19:31:22 -04:00
A separate `po-web-ui` Lambda (Function URL, unauthenticated) renders a simple HTML dashboard scanning the table.
Align PO schema with enriched records and improve extraction prompt
Replaces extraction prompt with domain-specific rules: trade
classification taxonomy (23 categories), site_code skip list,
zip padding, revision email type, and structured extraction for
fiscal_year, trade, and coupa_category.
Handler changes:
- New "revision" email type overwrites existing PO via put_item
- enrich_parsed() adds top-level state, ship_to_raw, data_source
- pad_zip() zero-pads short zip codes (e.g., "7001" → "07001")
- Removed invoice_total/invoice_count (Payee Central only)
Web UI: added revision badge, new detail fields (site code, state,
trade, fiscal year, coupa category, data source), line item table
now shows Qty/Unit/Price columns, list view shows Site and Trade.
CDK: fixed StreamViewType to match deployed table (NEW_IMAGE).
README: documented PO record schema and revision flow.
2026-05-01 19:40:18 -04:00
### PO record schema
Each record in `purchase-orders` includes:
- **Core**: `po_number` (PK), `email_type` (new_po/revision/cancellation), `po_status` , `source_system`
- **People/dates**: `submitted_by` , `on_behalf_of` , `order_date` , `revision_date` , `payment_terms` , `requisition_number` , `department`
- **Site**: `site_code` , `state` (top-level), `ship_to` (structured), `ship_to_raw` (original text)
- **Classification**: `trade` , `fiscal_year` , `coupa_category`
- **Financials**: `total_amount` , `currency` , `line_items[]` (with `description` , `amount` , `quantity` , `unit` , `price` , `need_by` )
- **Metadata**: `data_source` ("email" or "email+payee_scrape"), `email_subject` , `processed_at` , `raw_s3_key`
2026-04-30 14:30:54 -04:00
### Verified-sites pipeline
The `po-ingest-site-extractor` Lambda is triggered by the DynamoDB Stream on every PO INSERT/MODIFY. It:
1. Extracts an Amazon facility site code from `ship_to.name` using a regex cascade (parentheses, `LLC - CODE` , `Station CODE` , `DS - CODE` ) with a fallback to the first `line_items` description.
2. Parses `ship_to.address` into structured fields (street, city, state, zip).
3. Upserts to the `verified-sites` DynamoDB table — atomically increments `poCount` and appends the PO number to `sourcePOs` .
2026-04-30 15:03:43 -04:00
POs with no extractable site code fall through to an address reverse-lookup against the verified-sites cache (normalized street + zip). If still unresolved, the PO is written to the `pending-site-review` table for manual verification against Payee Central.
2026-04-30 14:30:54 -04:00
Backfill stats (initial run): 14,825 POs scanned → 9,900 with extractable site codes → 1,100 unique sites.
2026-04-20 19:31:22 -04:00
## Architecture
2026-05-01 19:17:19 -04:00
- **IaC:** AWS CDK (Python), stack name `po-ingest` , region `us-east-1` .
2026-04-30 14:30:54 -04:00
- **Lambdas** (all Python 3.12, arm64, 60-day log retention):
- `po-email-processor` — S3-triggered, parses PO emails via Claude Haiku.
- `po-web-ui` — Function URL, HTML dashboard.
- `po-ingest-site-extractor` — DynamoDB Streams-triggered, extracts site addresses.
- **Storage:**
- S3 `po-ingest-emails-{AccountId}` — 90-day lifecycle expiry.
- DynamoDB `purchase-orders` — owned by this stack, Streams enabled (NEW_AND_OLD_IMAGES).
- DynamoDB `verified-sites` — PK `siteCode` , GSI `by-state` on `state` .
2026-04-30 15:03:43 -04:00
- DynamoDB `pending-site-review` — PK `po_number` . POs with no extractable site code and no address match, awaiting manual Payee Central verification.
2026-04-20 19:31:22 -04:00
- **Secrets:** Anthropic API key in Secrets Manager at `po-ingest/anthropic-api-key` .
- **SES:** adds the `PoEmailRule` to the existing `INBOUND_MAIL` receipt rule set (shared with `workorder-ingest` ).
2026-05-01 19:17:19 -04:00
- **CI/CD:** CodePipeline V2 (`po-ingest-pipeline` ) → CodeBuild (`po-ingest-build` ). Pushes to `main` auto-deploy via `buildspec.yml` .
## CI/CD
Merges to `main` trigger the `po-ingest-pipeline` (CodePipeline V2) which runs CodeBuild to `cdk deploy` . The pipeline uses the existing CodeStar connection to the Sea-Haven-Industries GitHub org.
**Branch protection:** `main` requires a PR (no direct push), no deletion, no force push.
2026-04-20 19:31:22 -04:00
## Setup
1. Bootstrap CDK in the account if you haven't already: `cdk bootstrap aws://{AccountId}/us-east-1` .
2. Store the Anthropic API key:
```bash
aws secretsmanager create-secret \
--name po-ingest/anthropic-api-key \
--secret-string "sk-ant-..."
```
3. Install Lambda dependencies into the deployable package directory (gitignored):
```bash
pip install -r lambdas/email_processor/requirements.txt -t lambdas/email_processor/package/
```
4. Deploy:
```bash
cd cdk
pip install -r requirements.txt
cdk deploy
```
5. The `WebUIUrl` CloudFormation output is the dashboard URL.
## Reprocessing
To re-run the processor against every email still sitting in `inbound/` (useful after a parser change):
```bash
python scripts/reprocess.py # dry-run — lists keys
python scripts/reprocess.py --execute # invokes po-email-processor for each
```
Inserts are conditional on `po_number` , so re-processing existing POs is a no-op.
2026-04-30 14:30:54 -04:00
## Backfilling verified sites
The stream Lambda handles all future POs automatically. To backfill from historical PO data (one-time):
```bash
python scripts/backfill_sites.py
```
Uses the same extraction logic as the Lambda. Idempotent — safe to re-run.