procurement-ingest/README.md
Adam Moussa e66062c934 Align PO schema with enriched records and improve extraction prompt
Replaces extraction prompt with domain-specific rules: trade
classification taxonomy (23 categories), site_code skip list,
zip padding, revision email type, and structured extraction for
fiscal_year, trade, and coupa_category.

Handler changes:
- New "revision" email type overwrites existing PO via put_item
- enrich_parsed() adds top-level state, ship_to_raw, data_source
- pad_zip() zero-pads short zip codes (e.g., "7001" → "07001")
- Removed invoice_total/invoice_count (Payee Central only)

Web UI: added revision badge, new detail fields (site code, state,
trade, fiscal year, coupa category, data source), line item table
now shows Qty/Unit/Price columns, list view shows Site and Trade.

CDK: fixed StreamViewType to match deployed table (NEW_IMAGE).
README: documented PO record schema and revision flow.
2026-05-01 19:53:26 -04:00

106 lines
5.7 KiB
Markdown

# PO Ingest
Coupa purchase-order email ingestion pipeline. SES receives Amazon PO emails, Claude extracts structured data, and the result lands in the shared `purchase-orders` DynamoDB table.
## Flow
1. Coupa sends a PO email to `amazon_po@int.seahaven.com`.
2. SES (using the shared `INBOUND_MAIL` rule set) drops the raw MIME into `s3://po-ingest-emails-{AccountId}/inbound/`.
3. S3 `ObjectCreated` fires the `po-email-processor` Lambda.
4. The Lambda parses the email, sends it to Claude Haiku 4.5 for structured JSON extraction, and writes to DynamoDB.
- `email_type: new_po` — conditional `PutItem` on `purchase-orders` (idempotent on `po_number`).
- `email_type: revision` — unconditional `PutItem` overwriting the existing record with updated data.
- `email_type: cancellation` — `UpdateItem` marking the existing row `Cancelled`.
5. DynamoDB Streams (NEW_IMAGE) on `purchase-orders` feeds two downstream consumers:
- **LedgerFlow** (`seahaven-slack-bot/po-sync`) — daily KB sync.
- **Verified-sites pipeline** (`po-ingest-site-extractor`) — real-time site address extraction (see below).
The extraction prompt includes domain-specific rules for site code identification (with a skip list for false positives like RME, BBM, JLL), trade classification across 23 categories (Plumbing PM/Reactive, Electrical, HVAC, Dock Doors, etc.), fiscal year derivation, and ship-to address parsing with zip code zero-padding.
A separate `po-web-ui` Lambda (Function URL, unauthenticated) renders a simple HTML dashboard scanning the table.
### PO record schema
Each record in `purchase-orders` includes:
- **Core**: `po_number` (PK), `email_type` (new_po/revision/cancellation), `po_status`, `source_system`
- **People/dates**: `submitted_by`, `on_behalf_of`, `order_date`, `revision_date`, `payment_terms`, `requisition_number`, `department`
- **Site**: `site_code`, `state` (top-level), `ship_to` (structured), `ship_to_raw` (original text)
- **Classification**: `trade`, `fiscal_year`, `coupa_category`
- **Financials**: `total_amount`, `currency`, `line_items[]` (with `description`, `amount`, `quantity`, `unit`, `price`, `need_by`)
- **Metadata**: `data_source` ("email" or "email+payee_scrape"), `email_subject`, `processed_at`, `raw_s3_key`
### Verified-sites pipeline
The `po-ingest-site-extractor` Lambda is triggered by the DynamoDB Stream on every PO INSERT/MODIFY. It:
1. Extracts an Amazon facility site code from `ship_to.name` using a regex cascade (parentheses, `LLC - CODE`, `Station CODE`, `DS - CODE`) with a fallback to the first `line_items` description.
2. Parses `ship_to.address` into structured fields (street, city, state, zip).
3. Upserts to the `verified-sites` DynamoDB table — atomically increments `poCount` and appends the PO number to `sourcePOs`.
POs with no extractable site code fall through to an address reverse-lookup against the verified-sites cache (normalized street + zip). If still unresolved, the PO is written to the `pending-site-review` table for manual verification against Payee Central.
Backfill stats (initial run): 14,825 POs scanned → 9,900 with extractable site codes → 1,100 unique sites.
## Architecture
- **IaC:** AWS CDK (Python), stack name `po-ingest`, region `us-east-1`.
- **Lambdas** (all Python 3.12, arm64, 60-day log retention):
- `po-email-processor` — S3-triggered, parses PO emails via Claude Haiku.
- `po-web-ui` — Function URL, HTML dashboard.
- `po-ingest-site-extractor` — DynamoDB Streams-triggered, extracts site addresses.
- **Storage:**
- S3 `po-ingest-emails-{AccountId}` — 90-day lifecycle expiry.
- DynamoDB `purchase-orders` — owned by this stack, Streams enabled (NEW_AND_OLD_IMAGES).
- DynamoDB `verified-sites` — PK `siteCode`, GSI `by-state` on `state`.
- DynamoDB `pending-site-review` — PK `po_number`. POs with no extractable site code and no address match, awaiting manual Payee Central verification.
- **Secrets:** Anthropic API key in Secrets Manager at `po-ingest/anthropic-api-key`.
- **SES:** adds the `PoEmailRule` to the existing `INBOUND_MAIL` receipt rule set (shared with `workorder-ingest`).
- **CI/CD:** CodePipeline V2 (`po-ingest-pipeline`) → CodeBuild (`po-ingest-build`). Pushes to `main` auto-deploy via `buildspec.yml`.
## CI/CD
Merges to `main` trigger the `po-ingest-pipeline` (CodePipeline V2) which runs CodeBuild to `cdk deploy`. The pipeline uses the existing CodeStar connection to the Sea-Haven-Industries GitHub org.
**Branch protection:** `main` requires a PR (no direct push), no deletion, no force push.
## Setup
1. Bootstrap CDK in the account if you haven't already: `cdk bootstrap aws://{AccountId}/us-east-1`.
2. Store the Anthropic API key:
```bash
aws secretsmanager create-secret \
--name po-ingest/anthropic-api-key \
--secret-string "sk-ant-..."
```
3. Install Lambda dependencies into the deployable package directory (gitignored):
```bash
pip install -r lambdas/email_processor/requirements.txt -t lambdas/email_processor/package/
```
4. Deploy:
```bash
cd cdk
pip install -r requirements.txt
cdk deploy
```
5. The `WebUIUrl` CloudFormation output is the dashboard URL.
## Reprocessing
To re-run the processor against every email still sitting in `inbound/` (useful after a parser change):
```bash
python scripts/reprocess.py # dry-run — lists keys
python scripts/reprocess.py --execute # invokes po-email-processor for each
```
Inserts are conditional on `po_number`, so re-processing existing POs is a no-op.
## Backfilling verified sites
The stream Lambda handles all future POs automatically. To backfill from historical PO data (one-time):
```bash
python scripts/backfill_sites.py
```
Uses the same extraction logic as the Lambda. Idempotent — safe to re-run.