mirror of
https://github.com/Sea-Haven-Industries/procurement-ingest.git
synced 2026-09-30 18:53:14 +00:00
Replaces extraction prompt with domain-specific rules: trade classification taxonomy (23 categories), site_code skip list, zip padding, revision email type, and structured extraction for fiscal_year, trade, and coupa_category. Handler changes: - New "revision" email type overwrites existing PO via put_item - enrich_parsed() adds top-level state, ship_to_raw, data_source - pad_zip() zero-pads short zip codes (e.g., "7001" → "07001") - Removed invoice_total/invoice_count (Payee Central only) Web UI: added revision badge, new detail fields (site code, state, trade, fiscal year, coupa category, data source), line item table now shows Qty/Unit/Price columns, list view shows Site and Trade. CDK: fixed StreamViewType to match deployed table (NEW_IMAGE). README: documented PO record schema and revision flow.
106 lines
5.7 KiB
Markdown
106 lines
5.7 KiB
Markdown
# PO Ingest
|
|
|
|
Coupa purchase-order email ingestion pipeline. SES receives Amazon PO emails, Claude extracts structured data, and the result lands in the shared `purchase-orders` DynamoDB table.
|
|
|
|
## Flow
|
|
|
|
1. Coupa sends a PO email to `amazon_po@int.seahaven.com`.
|
|
2. SES (using the shared `INBOUND_MAIL` rule set) drops the raw MIME into `s3://po-ingest-emails-{AccountId}/inbound/`.
|
|
3. S3 `ObjectCreated` fires the `po-email-processor` Lambda.
|
|
4. The Lambda parses the email, sends it to Claude Haiku 4.5 for structured JSON extraction, and writes to DynamoDB.
|
|
- `email_type: new_po` — conditional `PutItem` on `purchase-orders` (idempotent on `po_number`).
|
|
- `email_type: revision` — unconditional `PutItem` overwriting the existing record with updated data.
|
|
- `email_type: cancellation` — `UpdateItem` marking the existing row `Cancelled`.
|
|
5. DynamoDB Streams (NEW_IMAGE) on `purchase-orders` feeds two downstream consumers:
|
|
- **LedgerFlow** (`seahaven-slack-bot/po-sync`) — daily KB sync.
|
|
- **Verified-sites pipeline** (`po-ingest-site-extractor`) — real-time site address extraction (see below).
|
|
|
|
The extraction prompt includes domain-specific rules for site code identification (with a skip list for false positives like RME, BBM, JLL), trade classification across 23 categories (Plumbing PM/Reactive, Electrical, HVAC, Dock Doors, etc.), fiscal year derivation, and ship-to address parsing with zip code zero-padding.
|
|
|
|
A separate `po-web-ui` Lambda (Function URL, unauthenticated) renders a simple HTML dashboard scanning the table.
|
|
|
|
### PO record schema
|
|
|
|
Each record in `purchase-orders` includes:
|
|
- **Core**: `po_number` (PK), `email_type` (new_po/revision/cancellation), `po_status`, `source_system`
|
|
- **People/dates**: `submitted_by`, `on_behalf_of`, `order_date`, `revision_date`, `payment_terms`, `requisition_number`, `department`
|
|
- **Site**: `site_code`, `state` (top-level), `ship_to` (structured), `ship_to_raw` (original text)
|
|
- **Classification**: `trade`, `fiscal_year`, `coupa_category`
|
|
- **Financials**: `total_amount`, `currency`, `line_items[]` (with `description`, `amount`, `quantity`, `unit`, `price`, `need_by`)
|
|
- **Metadata**: `data_source` ("email" or "email+payee_scrape"), `email_subject`, `processed_at`, `raw_s3_key`
|
|
|
|
### Verified-sites pipeline
|
|
|
|
The `po-ingest-site-extractor` Lambda is triggered by the DynamoDB Stream on every PO INSERT/MODIFY. It:
|
|
|
|
1. Extracts an Amazon facility site code from `ship_to.name` using a regex cascade (parentheses, `LLC - CODE`, `Station CODE`, `DS - CODE`) with a fallback to the first `line_items` description.
|
|
2. Parses `ship_to.address` into structured fields (street, city, state, zip).
|
|
3. Upserts to the `verified-sites` DynamoDB table — atomically increments `poCount` and appends the PO number to `sourcePOs`.
|
|
|
|
POs with no extractable site code fall through to an address reverse-lookup against the verified-sites cache (normalized street + zip). If still unresolved, the PO is written to the `pending-site-review` table for manual verification against Payee Central.
|
|
|
|
Backfill stats (initial run): 14,825 POs scanned → 9,900 with extractable site codes → 1,100 unique sites.
|
|
|
|
## Architecture
|
|
|
|
- **IaC:** AWS CDK (Python), stack name `po-ingest`, region `us-east-1`.
|
|
- **Lambdas** (all Python 3.12, arm64, 60-day log retention):
|
|
- `po-email-processor` — S3-triggered, parses PO emails via Claude Haiku.
|
|
- `po-web-ui` — Function URL, HTML dashboard.
|
|
- `po-ingest-site-extractor` — DynamoDB Streams-triggered, extracts site addresses.
|
|
- **Storage:**
|
|
- S3 `po-ingest-emails-{AccountId}` — 90-day lifecycle expiry.
|
|
- DynamoDB `purchase-orders` — owned by this stack, Streams enabled (NEW_AND_OLD_IMAGES).
|
|
- DynamoDB `verified-sites` — PK `siteCode`, GSI `by-state` on `state`.
|
|
- DynamoDB `pending-site-review` — PK `po_number`. POs with no extractable site code and no address match, awaiting manual Payee Central verification.
|
|
- **Secrets:** Anthropic API key in Secrets Manager at `po-ingest/anthropic-api-key`.
|
|
- **SES:** adds the `PoEmailRule` to the existing `INBOUND_MAIL` receipt rule set (shared with `workorder-ingest`).
|
|
- **CI/CD:** CodePipeline V2 (`po-ingest-pipeline`) → CodeBuild (`po-ingest-build`). Pushes to `main` auto-deploy via `buildspec.yml`.
|
|
|
|
## CI/CD
|
|
|
|
Merges to `main` trigger the `po-ingest-pipeline` (CodePipeline V2) which runs CodeBuild to `cdk deploy`. The pipeline uses the existing CodeStar connection to the Sea-Haven-Industries GitHub org.
|
|
|
|
**Branch protection:** `main` requires a PR (no direct push), no deletion, no force push.
|
|
|
|
## Setup
|
|
|
|
1. Bootstrap CDK in the account if you haven't already: `cdk bootstrap aws://{AccountId}/us-east-1`.
|
|
2. Store the Anthropic API key:
|
|
```bash
|
|
aws secretsmanager create-secret \
|
|
--name po-ingest/anthropic-api-key \
|
|
--secret-string "sk-ant-..."
|
|
```
|
|
3. Install Lambda dependencies into the deployable package directory (gitignored):
|
|
```bash
|
|
pip install -r lambdas/email_processor/requirements.txt -t lambdas/email_processor/package/
|
|
```
|
|
4. Deploy:
|
|
```bash
|
|
cd cdk
|
|
pip install -r requirements.txt
|
|
cdk deploy
|
|
```
|
|
5. The `WebUIUrl` CloudFormation output is the dashboard URL.
|
|
|
|
## Reprocessing
|
|
|
|
To re-run the processor against every email still sitting in `inbound/` (useful after a parser change):
|
|
|
|
```bash
|
|
python scripts/reprocess.py # dry-run — lists keys
|
|
python scripts/reprocess.py --execute # invokes po-email-processor for each
|
|
```
|
|
|
|
Inserts are conditional on `po_number`, so re-processing existing POs is a no-op.
|
|
|
|
## Backfilling verified sites
|
|
|
|
The stream Lambda handles all future POs automatically. To backfill from historical PO data (one-time):
|
|
|
|
```bash
|
|
python scripts/backfill_sites.py
|
|
```
|
|
|
|
Uses the same extraction logic as the Lambda. Idempotent — safe to re-run.
|