export const meta = { name: 'phase-5-handler-decomposition', description: 'Phase 5 of the procurement-ingest refactor (docs/refactor-evaluation.md): decompose both email-processor God-handlers into flat siblings along the seams that already work. PO handler.py splits into handler.py (event loop + auth + routing) / extraction.py (extract_with_claude + prompt import; parse_raw_email already lives in shared/email_parsing.py) / enrichment.py (enrich_parsed, pad_zip — PO-only) / telemetry.py (the EMF ParseMethod emit wrappers) / persistence.py (_write_fields/_merge_update/save_* with save_new_po+save_revision collapsed into one behavior-identical _save_merge). WO splits into ~5 concerns, keeping validate_ai_fallback + the [0-9]+ work_order_id key guard in the handler loop ahead of both saves and _header_date_iso/comment_id determinism with persistence. EXTRACTION_PROMPT moves to prompts.py (re-exported into handler for 4 tests). Lazy cached boto3 accessors in I/O modules; pure modules import no boto3. Requires Phases 0-3 on the base branch (relies on the flat-sibling lambdas/shared/ pattern + the Phase 0/2/3 bundling glob that auto-ships new siblings). Untrusted-input parse/extract/gate/auth-routing code moves, so push is gated on /sh-security-review in the main loop. Committed locally, never pushed.', phases: [ { title: 'Setup', detail: 'verify Phases 0+1+2+3 on base (lambdas/shared/ + ses_auth/web_ui_auth/email_parsing/emf present), branch feature/phase-5-handler-decomposition', model: 'haiku' }, { title: 'Recon', detail: '4 mappers: PO handler seams + save_new_po/save_revision byte-compare, WO handler seams + key-guard/date-determinism, bundling glob + AST test + monkeypatch/loader patch surface, deployed-zip baseline (read-only AWS)' }, { title: 'Spec', detail: 'serial fable spec: pin per-pipeline module split, _save_merge collapse proof, prompts.py move + re-export, lazy boto3 accessor shape, patch-surface + PO_EXPECTED_TOP_LEVEL_MODULES edits, behavior-pin tests' }, { title: 'Implement', detail: 'opus: PO siblings + handler; opus: WO siblings + handler; opus: all test plumbing + bundle-test + README — disjoint files', model: 'opus' }, { title: 'Verify', detail: 'mechanical CHECKS (suite/goldens/ruff/synth/AST/zero cdk+shared+derived diff/ordering greps) + 3 fable lenses (behavior-preservation, boto3-laziness, save-merge-collapse)' }, { title: 'Fix', detail: 'opus fixer, full re-verify, max 3 rounds', model: 'opus' }, { title: 'Package', detail: 'single commit via -F (no push)', model: 'sonnet' }, ], } // ---------------------------------------------------------------- constants const REPO = '/Users/adammoussa/Documents/repositories/seahaven/procurement-ingest' const BRANCH = 'feature/phase-5-handler-decomposition' let _args = args if (typeof _args === 'string') { try { _args = JSON.parse(_args) } catch (e) { _args = null } } const BASE = (_args && _args.base) || 'main' const CONSTRAINTS = ` PINNED BEHAVIORAL CONSTRAINTS (docs/refactor-evaluation.md Phase 5 — violating any is a build failure): 1. PO SPLIT INTO FLAT SIBLINGS under lambdas/po/email_processor/ (same dir, flat-landing so bare-name imports keep working — the ses_auth/ template_parser/derive_all pattern): handler.py -- event loop + fail-closed auth + email_type routing ONLY. extraction.py -- extract_with_claude + _EMAIL_TAG_RE; imports EXTRACTION_PROMPT from prompts. parse_raw_email is ALREADY in shared/email_parsing.py after Phase 3 -- import it, do NOT recreate a PO copy. enrichment.py -- enrich_parsed + pad_zip (PO-ONLY; WO has NO enrichment stage). The derived-field block moves here as a PURE code move. telemetry.py -- the EMF ParseMethod emit wrappers (_emit_parse_method_metric and the derived-agreement emit call surface used by enrichment) -- but see constraint 4: derived_fields.py and its shadow telemetry BEHAVIOR are untouchable. persistence.py -- _write_fields / _merge_update / save_*. COLLAPSE the byte-identical save_new_po/save_revision into ONE _save_merge, behavior-identical to BOTH originals INCLUDING the sticky-cancel ConditionExpression guard; save_cancellation stays its own function. 2. WO SPLIT is ~5 concerns, NOT 7 (WO has no enrichment stage). Its split MUST keep validate_ai_fallback AND the re.fullmatch(r"[0-9]+", work_order_id) key guard in the handler loop AHEAD of BOTH save_work_order and save_event (the guard protects the DynamoDB partition key + the '#'-delimited comment_id range-key segment), and MUST keep _header_date_iso / comment_id determinism WITH persistence (the retry-idempotent event_id key). Do not move the guard below either save; do not fork comment_id determinism from save_event. 3. Move EXTRACTION_PROMPT (the large ~181-line prompt) to prompts.py with cross-reference headers to derived_fields (the trade/site/fiscal rule tables have a second authoritative copy there). KEEP a re-export in handler ('from prompts import EXTRACTION_PROMPT') since 4 tests dereference handler.EXTRACTION_PROMPT. Do the same shape for WO if its prompt moves. 4. BEHAVIOR-PRESERVATION (immovable, pin with tests): * PO emits ParseMethod=ai_fallback BEFORE the Bedrock call (handler loop: _emit_parse_method_metric fires before extract_with_claude); the ai_fallback_rejected emit is the additive second datapoint on rejection. * WO emits AFTER its gate, with mutually-exclusive ai_fallback / ai_fallback_rejected (a rejected WO email emits ONLY ai_fallback_rejected). * Shadow DerivedFieldAgreement telemetry stays ai_fallback-only. * derived_fields.py and the shadow-telemetry BEHAVIOR are UNTOUCHABLE (the bake is in progress). Moving enrich_parsed into enrichment.py must be a PURE code move -- byte-identical logic, ZERO behavior change. git diff ${BASE}...HEAD -- lambdas/po/email_processor/derived_fields.py MUST be empty. 5. SIGNATURES PRESERVED: handler(event, context) unchanged (event shape + return contract identical) on BOTH pipelines; the save_* public contract unchanged (the _save_merge collapse must be behavior-identical to both save_new_po and save_revision -- callers route new_po/revision through it). Goldens UNCHANGED. 6. BUNDLING: the Phase 0 top-level cp glob (cp ./*.py or cp /*.py) plus the AST bundle-consistency test already ship & verify new flat siblings automatically -- but the exact-set pin PO_EXPECTED_TOP_LEVEL_MODULES (and the WO equivalent) must be updated for the new sibling modules IN THE SAME CHANGE, keeping the test's teeth (a new sibling that is NOT added to the expected set, or a commented-out cp, must still fail the AST test). 7. LAZY CACHED boto3 accessors in the I/O modules (extraction.py's bedrock, persistence.py's dynamodb, handler.py's s3 -- whatever each module actually uses); PURE modules (enrichment.py logic, prompts.py, telemetry.py if it makes no AWS call) import NO boto3. Update the monkeypatch.setattr(handler, "dynamodb", fake) patch surface in the SAME change -- tests must patch the NEW module's accessor (e.g. the persistence module), everywhere the tests patch it. Preserve the moto-before-handler import ordering (_po_parser_support.py:26-33 -- verify current lines) or the moto-backed suites hit real AWS. 8. ZERO cdk/ diff (git diff ${BASE}...HEAD -- cdk/ must be empty) and ZERO lambdas/shared/ diff (Phase 3's shared modules are untouchable here). 9. UNTOUCHABLE FILES (git diff ${BASE}...HEAD must be empty for each): lambdas/po/email_processor/derived_fields.py, both template_parser.py, both validate_ai_fallback gate modules, lambdas/shared/*, cdk/*, and lambdas/po/site_extractor/ + both web_ui/ (Phase 5 touches only the two email processors). No opportunistic refactors -- every handler diff is a pure move (delete-here/add-there), an import line, or the _save_merge collapse per the binding spec. 10. AWS access is READ-ONLY (get-function, downloading deployed zips via presigned URLs). NEVER cdk deploy, never invoke, never mutate. ` const PREAMBLE = ` You are one of several agents building refactor Phase 5 in the git repo at ${REPO} on branch ${BRANCH} (already checked out — do NOT switch branches, do NOT create branches, do NOT commit, NEVER push, do NOT run cdk deploy or touch AWS resources beyond read-only calls). Authoritative spec: docs/refactor-evaluation.md, section "Phase 5 — Handler decomposition + lazy boto3 clients". Work ONLY in the files you are told you own; other agents are concurrently editing other files in this same working tree. ${CONSTRAINTS} Your final message is consumed by an orchestrator script, not a human — return only the structured data requested. ` // ------------------------------------------------------------------ schemas const RECON = { type: 'object', required: ['summary', 'facts'], properties: { summary: { type: 'string' }, facts: { type: 'array', items: { type: 'string' } }, blockers: { type: 'array', items: { type: 'string' } }, }, } const SPEC = { type: 'object', required: ['poSplit', 'woSplit', 'promptMove', 'botoLaziness', 'saveMergeCollapse', 'testPlumbing', 'behaviorPins', 'notes'], properties: { poSplit: { type: 'string', description: 'per PO sibling (handler/extraction/enrichment/telemetry/persistence): exact functions/constants it owns, its imports, and the source file:line ranges moved into it — nothing else may change; state explicitly that parse_raw_email is imported from shared/email_parsing.py and enrich_parsed is a byte-identical move' }, woSplit: { type: 'string', description: 'per WO sibling (~5 concerns): exact ownership, and the pinned proof that validate_ai_fallback + the [0-9]+ key guard stay in the handler loop ahead of both saves and _header_date_iso/comment_id determinism stays with persistence' }, promptMove: { type: 'string', description: 'prompts.py contents (which prompt(s) move), the cross-reference headers to derived_fields, and the exact re-export line kept in each handler so handler.EXTRACTION_PROMPT still resolves (name the 4 tests that dereference it)' }, botoLaziness: { type: 'string', description: 'per module: which boto3 client/resource it needs, the lazy cached accessor shape (module-level None + get_x() cache), which modules import NO boto3, and the exact new monkeypatch target(s) that replace monkeypatch.setattr(handler, "dynamodb", fake) everywhere tests patch it' }, saveMergeCollapse: { type: 'string', description: 'the single _save_merge signature + body, with a line-by-line proof it is behavior-identical to BOTH save_new_po and save_revision including the sticky-cancel ConditionExpression path; how the router calls it for new_po vs revision' }, testPlumbing: { type: 'string', description: 'every test/support file change: loader/_SIBLING_MODULES resolution for the new siblings, _po_parser_support.py/_wo_parser_support.py edits (moto-before-handler ordering preserved), the monkeypatch patch-surface rewrites, PO_EXPECTED_TOP_LEVEL_MODULES + WO equivalent edits with mutation-detection shapes, and the new behavior-pin tests' }, behaviorPins: { type: 'string', description: 'the exact tests/assertions pinning: PO ai_fallback emit precedes the Bedrock invoke; WO emits mutually-exclusive; shadow telemetry ai_fallback-only; _save_merge parity; the WO key-guard ordering' }, notes: { type: 'string' }, }, } const IMPL = { type: 'object', required: ['filesChanged', 'summary', 'checksRun'], properties: { filesChanged: { type: 'array', items: { type: 'string' } }, summary: { type: 'string' }, checksRun: { type: 'string' }, blockers: { type: 'array', items: { type: 'string' } }, }, } const CHECKS = { type: 'object', required: ['passed', 'details'], properties: { passed: { type: 'boolean' }, details: { type: 'string' }, scopeViolations: { type: 'array', items: { type: 'string' } }, }, } const FINDINGS = { type: 'object', required: ['findings'], properties: { findings: { type: 'array', items: { type: 'object', required: ['title', 'severity', 'confirmed', 'evidence', 'fix'], properties: { title: { type: 'string' }, severity: { enum: ['critical', 'high', 'medium', 'low'] }, confirmed: { type: 'boolean' }, evidence: { type: 'string' }, fix: { type: 'string' }, }, }, }, }, } // ------------------------------------------------------------------- setup phase('Setup') const setup = await agent(` In ${REPO}: 1. SEQUENCING GATE — Phases 0, 1, 2 AND 3 must all be on ${BASE}. Phase 5 relies on the flat-sibling lambdas/shared/ pattern + the Phase 0/2/3 bundling glob that auto-ships new siblings, and parse_raw_email now lives in shared/email_parsing.py after Phase 3. git fetch origin, then pick the base ref: origin/${BASE} if that remote ref exists, otherwise the local branch ${BASE} (a stacked local-only base is expected and fine). Verify on the base ref: (a) Phase 0: tests/test_bundle_consistency.py exists (git show :tests/test_bundle_consistency.py | head -3); (b) Phase 1: validate_ai_fallback exists in the PO pipeline (git grep validate_ai_fallback -- lambdas/po); (c) Phase 2: BOTH cdk/po_stack.py and cdk/wo_stack.py on the base ref contain Code.from_asset("../lambdas") for the email processors (git show :cdk/po_stack.py | grep -n '\\.\\./lambdas', same for wo_stack.py); (d) Phase 3: lambdas/shared/ EXISTS on the base ref and the four shared modules are present — ses_auth.py, web_ui_auth.py, email_parsing.py, emf.py (git show :lambdas/shared/email_parsing.py | head -3 for each; and confirm parse_raw_email lives in shared/email_parsing.py, NOT in the PO email_processor dir, on the base ref). If ANY is missing — ESPECIALLY Phase 3 (no lambdas/shared/ or no email_parsing.py on base) — STOP with a blocker naming the unmet phase and do nothing else. 2. Verify clean working tree (untracked .coverage / .claude/ / the local 44 MB lambdas/po/email_processor/package/ dir are fine; any OTHER dirt = blocker, never stash or discard). 3. git checkout ${BASE}; then git pull --ff-only ONLY if the branch has an upstream (a local-only base skips the pull — not a blocker); then git checkout -b ${BRANCH} 4. gh pr list --state open --json number,title,headRefName (overlap check). Return facts: HEAD sha, per-phase gate evidence (especially the Phase 3 shared/ + email_parsing.py proof), open PRs, blockers. `, { label: 'setup:branch', model: 'haiku', schema: RECON }) if (!setup || (setup.blockers && setup.blockers.length)) { return { aborted: 'setup blockers', blockers: setup ? setup.blockers : ['setup agent died'], facts: setup ? setup.facts : [] } } log(`Branch ${BRANCH} ready off ${BASE}. ${setup.summary}`) // ------------------------------------------------------------------- recon phase('Recon') const recon = await parallel([ () => agent(`${PREAMBLE} Read-only recon of the PO email-processor God-handler (lambdas/po/email_processor/handler.py) and its decomposition seams: 1. Quote with file:line the exact boundaries of every function/constant that moves and where it lands per constraint 1: EXTRACTION_PROMPT + _EMAIL_TAG_RE + extract_with_claude (-> extraction.py/prompts.py); _emit_parse_method_metric + _derived_agreement + _emit_derived_agreement_metric (-> telemetry.py, but note which of these the derived-agreement path needs and whether any lives in derived_fields.py per constraint 4); enrich_parsed + pad_zip (-> enrichment.py); _write_fields + _merge_update + save_new_po + save_revision + save_cancellation (-> persistence.py). 2. save_new_po vs save_revision: diff their BODIES line by line and prove they are byte-identical modulo the docstring/log string (the D9 collapse target). Quote both verbatim so the spec can pin _save_merge. 3. The handler loop (handler.py:662-758 currently): quote the exact ordering — healthcheck early-return, Records loop, s3.get_object, authenticate_inbound_email, parse_raw_email, try_deterministic_parse, _emit_parse_method_metric (PRE-Bedrock), extract_with_claude, validate_ai_fallback, ai_fallback_rejected emit + continue, po_number guard, enrich_parsed, email_type routing to save_*. 4. Which boto3 clients each moving chunk uses (s3, dynamodb resource, bedrock client) and the current module-level client lines (handler.py:27-29). 5. The current 'from derived_fields import derive_all', 'from ses_auth import authenticate_inbound_email', 'from template_parser import ...' import lines and what parse_raw_email's import will become (shared/email_parsing.py after Phase 3 -- confirm on the working tree whether it is already imported from there). 6. Which tests dereference handler.EXTRACTION_PROMPT (grep) and which patch handler.dynamodb / handler.bedrock / handler.s3. 25-35 precise facts.`, { label: 'recon:po-handler', model: 'sonnet', phase: 'Recon', schema: RECON }), () => agent(`${PREAMBLE} Read-only recon of the WO email-processor God-handler (lambdas/wo/email_processor/handler.py) and its ~5-concern split: 1. Quote with file:line every function/constant and the concern it belongs to: EXTRACTION_PROMPT + _EMAIL_TAG_RE + extract_with_bedrock; emit_parse_metric; save_work_order; _header_date_iso + save_event; the handler loop. 2. The handler loop (handler.py:350-435 currently): quote the exact ordering — healthcheck, Records loop, s3.get_object, authenticate_inbound_email, parse_raw_email, try_deterministic_parse, extract_with_bedrock, validate_ai_fallback, ai_fallback_rejected emit + continue, emit_parse_metric (POST-gate), the re.fullmatch(r"[0-9]+", work_order_id) key guard, save_work_order, save_event. CRITICAL per constraint 2: pin that the gate + the [0-9]+ key guard both precede BOTH saves, and that comment_id determinism (_header_date_iso + the object-key sha suffix, handler.py:277-347) is coupled to save_event. 3. Prove WO has NO enrichment stage (no enrich_parsed equivalent) so its split is 5 concerns not 7. 4. WO's EMF emit is mutually-exclusive (rejected email emits ONLY ai_fallback_rejected then continues; accepted email emits emit_parse_metric once AFTER the gate) -- quote the exact lines proving this, contrasted with PO's pre-Bedrock double-count. 5. boto3 clients WO uses (handler.py:27-29) and which WO tests patch handler.dynamodb / handler.bedrock and dereference handler.EXTRACTION_PROMPT. 6. _wo_parser_support.py bare-sys.path 'import handler' strategy and how it will resolve the new siblings. 20-30 precise facts.`, { label: 'recon:wo-handler', model: 'sonnet', phase: 'Recon', schema: RECON }), () => agent(`${PREAMBLE} Read-only recon of the bundling glob, the AST bundle-consistency test, and the test patch/loader surface this phase must rewire: 1. tests/test_bundle_consistency.py: quote PO_EXPECTED_TOP_LEVEL_MODULES (currently {"handler","ses_auth","template_parser","derived_fields"} on pre-Phase-3 — read the value ON THIS BRANCH after Phase 3 has added the shared modules), whether a WO_EXPECTED_TOP_LEVEL_MODULES (or equivalent) exists, the _first_party_sibling_imports AST logic, the exact-set assertion, and every mutation-detection assertion (a commented-out cp / a missing module in the expected set must fail). Confirm the Phase 0 cp glob form in both stacks (cp ./*.py or cp /*.py) auto-ships NEW flat siblings with no cdk edit needed this phase. 2. tests/conftest.py importlib loader: how it resolves module paths, the sys.modules save/restore, _SIBLING_MODULES exact current contents (which bare names it lists per pipeline). 3. lambdas/po/email_processor/tests/_po_parser_support.py: the independent importlib reimplementation, the load-bearing moto-before-handler import order (~lines 26-33 — quote it), and how handler + its new siblings get loaded. 4. lambdas/wo/email_processor/tests/_wo_parser_support.py: same. 5. Every place tests do monkeypatch.setattr(handler, "dynamodb"/"bedrock"/"s3", fake) — grep both pipelines' tests and quote each file:line (this is the patch surface constraint 7 says must move to the new module accessors). 6. Which tests reference handler.save_new_po / handler.save_revision / handler.enrich_parsed / handler.extract_with_claude / handler.EXTRACTION_PROMPT directly (they may need repointing after the split). 20-30 facts.`, { label: 'recon:bundling-tests', model: 'sonnet', phase: 'Recon', schema: RECON }), () => agent(`${PREAMBLE} Read-only AWS recon (region us-east-1, READ-ONLY): for po-email-processor and workorder-email-processor (find exact function names via the stacks/aws lambda list-functions) run aws lambda get-function; download each Code.Location presigned zip to a scratch dir and produce the COMPLETE top-level file list (unzip -l) separated into (a) first-party top-level .py, (b) dependency dirs, (c) anything else; record each function's CodeSha256. This deployed first-party module set is the baseline the post-split staged bundle is judged against: after the split the staged bundle's first-party .py set must be exactly the deployed set MINUS nothing PLUS the new sibling modules (extraction/enrichment/telemetry/ persistence/prompts for PO, the WO equivalents) and with the collapsed save_new_po/save_revision still resolvable via imports (no NEW top-level module should vanish that a handler still imports). Confirm parse_raw_email is NOT a top-level module in either deployed email-processor zip if Phase 3 already shipped it from shared/ (it lands flat as email_parsing.py). Return the categorized lists as facts.`, { label: 'recon:deployed-baseline', model: 'sonnet', phase: 'Recon', schema: RECON }), ]) const reconOk = recon.filter(Boolean) const pack = reconOk.map(r => `## ${r.summary}\n${r.facts.join('\n')}`).join('\n\n') const reconBlockers = reconOk.flatMap(r => r.blockers || []) .filter(b => b && !/^\s*(none|n\/a)\b/i.test(b)) log(`Recon complete: ${reconOk.length}/4 mappers, ${reconBlockers.length} advisory notes`) // Recon "blockers" here are downstream guidance (verify the WO dynamodb // monkeypatch surface per constraint 7 before persistence.py's split; // compare the post-split staged bundle against the local pre-Phase-5 bundle, // NOT the stale pre-Phase-3 deployed AWS zip), NOT stop conditions — main-loop // verified Phase 3 is on base and the seams move cleanly. The real gates are // pytest-green (catches any missed monkeypatch surface) and the bundle-parity // verifier. Fold the notes into the spec context instead of aborting. const reconAdvisories = reconBlockers.length ? `\n\nRECON ADVISORIES (recon flagged these; they are downstream verification guidance — the impl must update the full dynamodb/bedrock/s3 monkeypatch surface (constraint 7) so no test breaks, and the parity check compares against the local pre-Phase-5 staged bundle not the pre-Phase-3 deployed zip — NOT reasons to stop; pytest-green + bundle-parity are the real gates):\n- ${reconBlockers.join('\n- ')}` : '' // -------------------------------------------------------------------- spec phase('Spec') const spec = await agent(`${PREAMBLE} You are the SPEC agent — the single authority that pins every contested decision BEFORE parallel implementation (parallel leaves cannot see each other's choices). Using the recon pack below plus your own reads of the actual files, produce the binding implementation spec: - poSplit: per constraint 1's five PO siblings — exact functions/constants each owns, the imports each needs, and the source line ranges moved. State explicitly: parse_raw_email is imported from shared/email_parsing.py (NOT recreated); enrich_parsed + its derived-field block move into enrichment.py as a BYTE-IDENTICAL logic move (constraint 4). Resolve where the derived-agreement emit surface lives (telemetry.py vs left in enrichment) WITHOUT touching derived_fields.py. - woSplit: per constraint 2's ~5 WO concerns — ownership, plus the explicit pin that validate_ai_fallback + the [0-9]+ key guard stay in the handler loop AHEAD of both saves and _header_date_iso/comment_id determinism stays coupled to persistence (save_event). WO has NO enrichment stage. - promptMove: prompts.py contents (PO prompt for sure; WO prompt too if it moves), the cross-reference headers to derived_fields, and the EXACT re-export line each handler keeps so handler.EXTRACTION_PROMPT still resolves (name the 4 PO tests + any WO tests that dereference it). - botoLaziness: per module, the lazy cached boto3 accessor shape (module-level None cache + get_x()); which modules import NO boto3 (pure: enrichment logic, prompts, telemetry if it emits via print only); and the EXACT new monkeypatch target(s) replacing monkeypatch.setattr(handler, "dynamodb", fake) at every test site recon found — pick ONE consistent rule (patch the owning module's accessor/attribute) and state why. Preserve the moto-before-handler ordering. - saveMergeCollapse: the single _save_merge signature + body with a line-by-line proof it is behavior-identical to BOTH save_new_po and save_revision including the sticky-cancel ConditionExpression path; how the router dispatches new_po vs revision through it; save_cancellation stays separate. - testPlumbing: every file, every edit — _SIBLING_MODULES additions for the new bare names, _po_parser_support.py / _wo_parser_support.py loader edits (moto ordering preserved), the monkeypatch patch-surface rewrites, repointing any test that referenced handler.save_new_po/enrich_parsed/etc. to the new module, PO_EXPECTED_TOP_LEVEL_MODULES + the WO equivalent updated to the new sibling set (with mutation shapes: a missing sibling in the set, or a commented-out cp, must FAIL the AST test), and the NEW behavior-pin tests. - behaviorPins: the exact new tests/assertions pinning constraint 4 & 5 — PO ai_fallback emit precedes the Bedrock invoke (a throttle test is the cleanest proof), WO emits mutually-exclusive, shadow telemetry ai_fallback-only, _save_merge parity to both originals, WO key-guard ordering ahead of both saves. Recon pack:\n${pack}${reconAdvisories}`, { label: 'spec:pin-decomposition', phase: 'Spec', schema: SPEC }) if (!spec) return { aborted: 'spec agent died — rerun workflow', reconBlockers } const specBlock = `BINDING SPEC (from the spec agent — implement EXACTLY this):\n${JSON.stringify(spec, null, 2)}` log('Spec pinned: PO split, WO split, prompt move, boto3 laziness, _save_merge collapse, test plumbing, behavior pins') // --------------------------------------------------------------- implement phase('Implement') const impl = await parallel([ () => agent(`${PREAMBLE} YOU OWN: lambdas/po/email_processor/ EXCEPT its tests/ directory (another agent owns all test and support files). Do not touch cdk/, lambdas/wo/, lambdas/shared/, or README. Task: execute the PO split per spec.poSplit + spec.promptMove + spec.botoLaziness + spec.saveMergeCollapse: 1. Create prompts.py (EXTRACTION_PROMPT + cross-ref headers); keep the re-export line in handler.py. 2. Create extraction.py (extract_with_claude + _EMAIL_TAG_RE, lazy bedrock accessor, imports EXTRACTION_PROMPT from prompts and parse_raw_email from the shared email_parsing module). 3. Create enrichment.py (enrich_parsed + pad_zip + the derived-field block as a BYTE-IDENTICAL logic move — derived_fields.py stays untouched; no boto3). 4. Create telemetry.py (the EMF ParseMethod emit wrappers per spec) — pure (print-only, no boto3) unless spec says otherwise. 5. Create persistence.py (_write_fields, _merge_update, the collapsed _save_merge behavior-identical to BOTH save_new_po and save_revision incl. the sticky- cancel guard, save_cancellation; lazy dynamodb accessor). 6. Rewrite handler.py down to the event loop + auth + routing, importing from the new siblings; keep handler(event, context) signature + the healthcheck early-return + the PRE-Bedrock ai_fallback emit ordering EXACTLY. Every change is a pure move / import line / the _save_merge collapse — no opportunistic refactors (constraint 9). derived_fields.py diff MUST stay empty. Run before returning: ruff check lambdas/po && ruff format lambdas/po --check, plus an import smoke: python3 -c with sys.path prepended for lambdas/shared + the PO function dir, importing every new module + handler (tests are NOT yours to run — the plumbing agent lands them in parallel). ${specBlock}`, { label: 'impl:po-siblings', model: 'opus', phase: 'Implement', schema: IMPL }), () => agent(`${PREAMBLE} YOU OWN: lambdas/wo/email_processor/ EXCEPT its tests/ directory (another agent owns all test and support files). Do not touch cdk/, lambdas/po/, lambdas/shared/, or README. Task: execute the WO ~5-concern split per spec.woSplit + spec.promptMove + spec.botoLaziness: 1. Create prompts.py if the WO prompt moves (re-export kept in handler). 2. Create extraction.py (extract_with_bedrock + _EMAIL_TAG_RE, lazy bedrock). 3. Create telemetry.py (emit_parse_metric) — pure unless spec says otherwise. 4. Create persistence.py (save_work_order, _header_date_iso, save_event — comment_id determinism stays coupled here; lazy dynamodb). 5. Rewrite handler.py to the event loop + auth + routing, KEEPING validate_ai_fallback + the re.fullmatch(r"[0-9]+", work_order_id) key guard in the loop AHEAD of BOTH saves (constraint 2), the POST-gate mutually- exclusive emit ordering, the healthcheck early-return, and handler(event, context) signature EXACTLY. Pure moves / imports only (constraint 9). No enrichment stage exists — do not invent one. Run before returning: ruff check lambdas/wo && ruff format lambdas/wo --check, plus an import smoke: python3 -c with sys.path prepended for lambdas/shared + the WO function dir, importing every new module + handler. ${specBlock}`, { label: 'impl:wo-siblings', model: 'opus', phase: 'Implement', schema: IMPL }), () => agent(`${PREAMBLE} YOU OWN: all test and support files (tests/ at repo root including tests/test_bundle_consistency.py and tests/conftest.py, lambdas/po/email_processor/tests/, lambdas/wo/email_processor/tests/) and README.md ONLY. Task A: apply spec.testPlumbing in full — _SIBLING_MODULES additions for the new bare names, _po_parser_support.py / _wo_parser_support.py loader edits (PRESERVING the moto-before-handler ordering comment + import), rewrite the monkeypatch patch surface everywhere tests patch handler.dynamodb/bedrock/s3 to the new module accessors per spec.botoLaziness, repoint any test that referenced handler.save_new_po / handler.save_revision / handler.enrich_parsed / handler.extract_with_claude to the new modules (KEEP handler.EXTRACTION_PROMPT resolving via the re-export — do NOT repoint those 4 tests), and update PO_EXPECTED_TOP_LEVEL_MODULES + the WO equivalent to the new sibling set WITHOUT losing teeth (constraint 6 — a missing sibling in the set or a commented-out cp must FAIL the AST test; keep existing mutation tests green). Task B: add the NEW behavior-pin tests per spec.behaviorPins — PO ai_fallback emit precedes the Bedrock invoke (throttle-based proof), WO emits mutually- exclusive, shadow telemetry ai_fallback-only, _save_merge parity to both originals, WO key-guard ordering ahead of both saves. Task C: README — update the repo-layout / module map for the new PO+WO sibling modules (the flat-landing import rule already ships them via the Phase 0 glob), note the save_new_po/save_revision collapse into _save_merge and prompts.py. Match existing README style; goldens stay UNCHANGED. Run before returning: pytest -q --no-cov at repo root (must be green against the OTHER agents' edits — they land in parallel; poll by re-running up to ~10 min before reporting a blocker) and ruff check on every file you touched. ${specBlock}`, { label: 'impl:test-plumbing-readme', model: 'opus', phase: 'Implement', schema: IMPL }), ]) const implOk = impl.filter(Boolean) const implBlockers = implOk.flatMap(r => r.blockers || []) log(`Implement complete: ${implOk.length}/3 agents, blockers: ${implBlockers.length}`) // ---------------------------------------------------- verify + fix loop const EXPECTED_SCOPE = [ 'lambdas/po/email_processor/', 'lambdas/wo/email_processor/', 'tests/', 'README.md', ] const mechanicalPrompt = `${PREAMBLE} Independent re-verification — trust nothing self-reported. Run ALL gates, quoting failures verbatim: 1. pytest -q --no-cov (repo root, all three roots green; GOLDENS UNCHANGED — git diff ${BASE}...HEAD on any golden .json/.eml fixture must be empty). 2. ruff check . && ruff format --check . 3. cd cdk && npx cdk synth po-ingest -q && npx cdk synth workorder-ingest -q (artifact-id selectors, NOT stack_name). The Docker bundling runs — confirm each staged email-processor asset now contains the new sibling modules. 4. AST BUNDLE TEST: the test_bundle_consistency.py suite is GREEN and its teeth hold — on a scratch copy, comment out the cp glob line and drop a sibling from PO_EXPECTED_TOP_LEVEL_MODULES (and the WO equivalent); EACH mutation must FAIL the test. Restore. 5. ZERO-DIFF UNTOUCHABLES (each must output NOTHING): git diff ${BASE}...HEAD -- cdk/ git diff ${BASE}...HEAD -- lambdas/shared/ git diff ${BASE}...HEAD -- lambdas/po/email_processor/derived_fields.py git diff ${BASE}...HEAD -- lambdas/po/email_processor/template_parser.py git diff ${BASE}...HEAD -- lambdas/wo/email_processor/template_parser.py plus both validate_ai_fallback gate modules (resolve their filenames first) and lambdas/po/site_extractor/ + both web_ui/. 6. SIGNATURES: grep proves handler(event, context) unchanged on both pipelines and the healthcheck early-return + return {"statusCode": 200, ...} contract intact. handler.EXTRACTION_PROMPT still resolves (the re-export line is present in both handlers). 7. ORDERING PINS (grep with line numbers): - PO: _emit_parse_method_metric / the ai_fallback emit still PRECEDES the extract_with_claude (Bedrock) invoke in the handler loop. - WO: the emit stays POST-gate and mutually-exclusive (rejected path emits ONLY ai_fallback_rejected then continue); the re.fullmatch(r"[0-9]+", ...) key guard still sits AHEAD of BOTH save_work_order and save_event. 8. BOTO3 LAZINESS: grep proves the pure modules (PO enrichment.py, prompts.py, telemetry if print-only; WO equivalents) import NO boto3, and the I/O modules use lazy cached accessors (no module-level eager client that a test cannot patch). No monkeypatch.setattr(handler, "dynamodb"/... ) remains pointing at a now-empty handler attribute. 9. git status --porcelain scope check: every modified/added/deleted path under ${EXPECTED_SCOPE.join(', ')} (untracked .coverage/.claude/package/ tolerated). passed=true only if all green. YOU MAY NOT edit files.` const lenses = [ { key: 'behavior-preservation', prompt: `${PREAMBLE} ADVERSARIAL REVIEW — behavior-preservation lens. The decomposition must not change one observable byte of metric or telemetry behavior. (1) METRIC ORDERING: read the post-split PO handler loop and prove _emit_parse_method_metric fires BEFORE extract_with_claude (so a Bedrock-side throttle still records ai_fallback) and that ai_fallback_rejected is still the additive second datapoint; then the WO loop and prove the emit is POST-gate and MUTUALLY-EXCLUSIVE (a rejected email emits ONLY ai_fallback_rejected, an accepted one emits once after the gate). Construct the emitted EMF envelope from each moved emitter and diff it against ${BASE}'s output — namespace, dimension-set list [["ParseMethod"],["ParseMethod", "TemplateId"]], property keys/types identical. (2) SHADOW TELEMETRY: prove DerivedFieldAgreement is still emitted ONLY on the ai_fallback path after enrich_parsed moved to enrichment.py, and that git diff ${BASE}...HEAD -- derived_fields.py is EMPTY (the enrich_parsed move is byte-identical logic). (3) ROUTING: prove email_type routing (cancellation/ revision/new_po for PO; the WO save_work_order + save_event pair) still reaches the same saves in the same order. confirmed=true only with a concrete file:line proof or a sample-emission diff.` }, { key: 'boto3-laziness', prompt: `${PREAMBLE} ADVERSARIAL REVIEW — boto3-laziness lens. (1) PURE MODULES: prove the modules constraint 7 calls pure (PO enrichment.py, prompts.py, telemetry if print-only; WO equivalents) import NO boto3 at all — grep + read. A stray 'import boto3' or eager client in a "pure" module is a finding. (2) LAZY ACCESSORS: in the I/O modules (extraction bedrock, persistence dynamodb, handler s3) prove the client is created lazily and cached (module-level None + get_x()), not eagerly at import — an eager module-level client that tests cannot patch is a finding. (3) PATCH-SURFACE COMPLETENESS: enumerate EVERY test that previously did monkeypatch.setattr(handler, "dynamodb"/"bedrock"/"s3", fake) and prove each now patches the correct NEW module accessor — a test that still patches a now-dead handler attribute silently exercises the REAL client (or moto-misses) and is a critical finding. (4) MOTO ORDERING: prove the moto-before-handler import order survived in _po_parser_support.py and _wo_parser_support.py — reorder-and-run to confirm the suites would hit real AWS without it. confirmed=true only with file:line or reproduced-failure evidence.` }, { key: 'save-merge-collapse', prompt: `${PREAMBLE} ADVERSARIAL REVIEW — save-merge-collapse lens. The one _save_merge that replaces save_new_po AND save_revision must be byte-behavior-identical to BOTH. (1) Read _save_merge and BOTH ${BASE} originals; prove line-by-line the field-building, the _merge_update call, and the sticky-cancel path are identical (incl. the ConditionExpression "attribute_not_exists(po_status) OR po_status <> :__cancelled_marker" and the ConditionalCheckFailedException fallback that re-writes WITHOUT po_status/cancelled_at). (2) Prove the router dispatches email_type=="revision" and the new_po else-branch both through _save_merge, and that save_cancellation stayed SEPARATE and unchanged. (3) Construct a scenario where the collapse could diverge: a revision landing on a Cancelled PO, a new_po with po_status=None, a new_po carrying a non-cancelled status onto a Cancelled skeleton — and prove _save_merge behaves exactly as the original would have for each. (4) Confirm the save_* public contract callers rely on is unchanged and the goldens still pass. confirmed=true only with a concrete divergence sketch or a line-by-line equivalence proof.` }, ] let round = 0 let checks = null let confirmed = [] while (round < 3) { phase('Verify') const results = await parallel([ () => agent(mechanicalPrompt, { label: `verify:mechanical-r${round}`, model: 'sonnet', phase: 'Verify', schema: CHECKS }), ...lenses.map(l => () => agent(l.prompt, { label: `verify:${l.key}-r${round}`, phase: 'Verify', schema: FINDINGS })), ]) checks = results[0] confirmed = results.slice(1).filter(Boolean) .flatMap(r => r.findings || []) .filter(f => f.confirmed && f.severity !== 'low') const green = checks && checks.passed log(`Verify round ${round}: mechanical ${checks && checks.passed ? 'GREEN' : 'RED'}, confirmed findings: ${confirmed.length}`) if (green && confirmed.length === 0) break round += 1 if (round >= 3) break phase('Fix') await agent(`${PREAMBLE} You are the fix agent — you may edit files under: ${EXPECTED_SCOPE.join(', ')}. Fix EVERY item below minimally; the binding spec and 10 pinned constraints still hold (a finding that conflicts with a constraint is reported, not "fixed" — the constraint wins, esp. constraint 4's derived_fields untouchability + byte-identical enrich move, constraint 5's signature/goldens preservation, and constraint 8's zero cdk/shared diff). Re-run the specific failing gate/test per fix. MECHANICAL:\n${checks ? checks.details : '(agent died — rerun all gates)'} CONFIRMED FINDINGS:\n${JSON.stringify(confirmed, null, 2)} ${specBlock}`, { label: `fix:round-${round}`, model: 'opus', phase: 'Fix', schema: IMPL }) } const verifyClean = checks && checks.passed && confirmed.length === 0 if (!verifyClean) { return { status: 'NEEDS ATTENTION — verify not clean after 3 rounds; branch left uncommitted', branch: BRANCH, mechanical: checks, unresolvedFindings: confirmed, implBlockers, reconBlockers, spec, } } // ----------------------------------------------------------------- package phase('Package') const commit = await agent(`${PREAMBLE.replace('do NOT commit, ', '')} YOU are the commit agent: 1. Read ~/Documents/repositories/seahaven/engineering-handbook/commit-messages.md and follow it exactly. 2. git add only paths under: ${EXPECTED_SCOPE.join(', ')} and .claude/workflows/phase-5-handler-decomposition.js. NOT .coverage, NOT package/. Verify the staged set with git status — the new sibling files AND the deletions of the moved-out code from the handlers MUST be staged. 3. ONE commit; write the message to /tmp/phase5-commit-msg.txt and use git commit -F /tmp/phase5-commit-msg.txt (backticks in -m get eaten by zsh). Suggested subject: "feat: decompose email-processor handlers into flat siblings + lazy boto3 clients (refactor phase 5)" Body: the PO 5-way split (handler/extraction/enrichment/telemetry/persistence) with the save_new_po/save_revision -> _save_merge collapse one-liner, the WO ~5-concern split keeping the gate+key-guard in the loop, prompts.py + the handler.EXTRACTION_PROMPT re-export, the lazy-boto3 + patch-surface move, and the behavior-preservation pins (PO pre-Bedrock emit, WO mutually-exclusive, shadow telemetry ai_fallback-only, derived_fields untouched). NO AI attribution / Co-Authored-By lines. 4. Do NOT push. Return commit sha + shortstat in summary.`, { label: 'package:commit', model: 'sonnet', phase: 'Package', schema: IMPL }) return { status: 'BUILT — committed locally, NOT pushed', branch: BRANCH, base: BASE, commit: commit ? commit.summary : 'commit agent died — commit manually', spec: { poSplit: spec.poSplit, woSplit: spec.woSplit, saveMergeCollapse: spec.saveMergeCollapse, botoLaziness: spec.botoLaziness, notes: spec.notes }, implementation: implOk.map(r => r.summary), filesChanged: implOk.flatMap(r => r.filesChanged), verifyRounds: round + 1, blockers: implBlockers.concat(reconBlockers), outstandingGates: [ '/sh-security-review (MANDATORY before push — the untrusted-input parse/extract/validate-gate/auth-routing code is being MOVED across modules; CLAUDE.md gates untrusted-input handling changes and a move still touches the surface. Run on the committed diff, re-run after any post-review fix)', 'cross-family cross_review.py NOT required (handler(event, context) event/return signatures unchanged and no IAM change — internal module moves only)', 'push + PR + gh pr checks green', 'deploy-then-merge: full suite goldens unchanged, smoke green, ONE live email per pipeline (PO + WO) showing the correct ParseMethod, both the fallback + rejected alarm series behaving; the post-merge CI redeploy is a fresh asset (new sibling files) so confirm the smoke gate + AST bundle test caught any missing-module regression BEFORE merge', ], }