procurement-ingest/.claude/workflows/phase-5-handler-decomposition.js
Adam Moussa ca4f43a2cc
Some checks are pending
Deploy / deploy (push) Waiting to run
feat: decompose email-processor handlers into flat siblings + lazy boto3 clients (refactor phase 5) (#113)
Both email-processor God-handlers split along the seams that already
work in the flat-sibling pattern established by lambdas/shared/, so
bare-name imports keep working under the existing bundling glob.

PO (5-way split): handler.py keeps only the event loop, fail-closed
auth, and email_type routing. extraction.py holds extract_with_claude
and _EMAIL_TAG_RE, importing EXTRACTION_PROMPT from prompts.py and
parse_raw_email from shared/email_parsing.py rather than recreating a
PO-local copy. enrichment.py is a pure code move of enrich_parsed and
pad_zip (PO-only; WO has no enrichment stage) with zero behavior
change. telemetry.py holds the EMF ParseMethod emit wrappers.
persistence.py holds _write_fields/_merge_update/save_*, collapsing
the byte-identical save_new_po/save_revision bodies into one
_save_merge helper that both now call through, preserving the sticky
Cancelled ConditionExpression guard for both callers; save_cancellation
stays separate.

WO (5 concerns, no enrichment stage): the handler loop keeps
validate_ai_fallback and the re.fullmatch(r"[0-9]+", work_order_id)
key guard ahead of both save_work_order and save_event, since the
guard protects the DynamoDB partition key and the '#'-delimited
comment_id range-key segment. _header_date_iso and comment_id
determinism stay colocated with persistence.py's save_event for the
retry-idempotent event_id key.

EXTRACTION_PROMPT (PO) moves to prompts.py with cross-reference
headers to derived_fields.py's authoritative trade/site/fiscal rule
tables; handler.py re-exports it (from prompts import
EXTRACTION_PROMPT) since four tests dereference handler.EXTRACTION_
PROMPT directly. WO's prompt moves the same way.

I/O modules (extraction.py's bedrock client, persistence.py's
dynamodb resource, handler.py's s3 client) get lazy cached boto3
accessors; pure modules (enrichment.py, prompts.py, telemetry.py)
import no boto3. Test monkeypatch surfaces move to the module that
now owns the client (e.g. persistence.dynamodb) everywhere tests
patch it, and the moto-before-handler-import ordering in
_po_parser_support.py is preserved so the moto-backed suites don't
hit real AWS.

Behavior-preservation pins, verified with tests: PO still emits
ParseMethod=ai_fallback before the Bedrock call, with
ai_fallback_rejected as the additive second datapoint on rejection.
WO still emits after its gate with mutually-exclusive ai_fallback /
ai_fallback_rejected. Shadow DerivedFieldAgreement telemetry stays
ai_fallback-only. derived_fields.py is untouched (diff against
feature/phase-3-shared-extraction is empty). handler(event, context)
signatures and the save_* public contract are unchanged on both
pipelines; goldens unchanged.

PO_EXPECTED_TOP_LEVEL_MODULES and its WO equivalent in
tests/test_bundle_consistency.py are updated for the new sibling
modules so the AST bundle-consistency test still fails on an
unshipped or uncommented-out sibling.
2026-07-20 15:34:53 -04:00

692 lines
43 KiB
JavaScript

export const meta = {
name: 'phase-5-handler-decomposition',
description: 'Phase 5 of the procurement-ingest refactor (docs/refactor-evaluation.md): decompose both email-processor God-handlers into flat siblings along the seams that already work. PO handler.py splits into handler.py (event loop + auth + routing) / extraction.py (extract_with_claude + prompt import; parse_raw_email already lives in shared/email_parsing.py) / enrichment.py (enrich_parsed, pad_zip — PO-only) / telemetry.py (the EMF ParseMethod emit wrappers) / persistence.py (_write_fields/_merge_update/save_* with save_new_po+save_revision collapsed into one behavior-identical _save_merge). WO splits into ~5 concerns, keeping validate_ai_fallback + the [0-9]+ work_order_id key guard in the handler loop ahead of both saves and _header_date_iso/comment_id determinism with persistence. EXTRACTION_PROMPT moves to prompts.py (re-exported into handler for 4 tests). Lazy cached boto3 accessors in I/O modules; pure modules import no boto3. Requires Phases 0-3 on the base branch (relies on the flat-sibling lambdas/shared/ pattern + the Phase 0/2/3 bundling glob that auto-ships new siblings). Untrusted-input parse/extract/gate/auth-routing code moves, so push is gated on /sh-security-review in the main loop. Committed locally, never pushed.',
phases: [
{ title: 'Setup', detail: 'verify Phases 0+1+2+3 on base (lambdas/shared/ + ses_auth/web_ui_auth/email_parsing/emf present), branch feature/phase-5-handler-decomposition', model: 'haiku' },
{ title: 'Recon', detail: '4 mappers: PO handler seams + save_new_po/save_revision byte-compare, WO handler seams + key-guard/date-determinism, bundling glob + AST test + monkeypatch/loader patch surface, deployed-zip baseline (read-only AWS)' },
{ title: 'Spec', detail: 'serial fable spec: pin per-pipeline module split, _save_merge collapse proof, prompts.py move + re-export, lazy boto3 accessor shape, patch-surface + PO_EXPECTED_TOP_LEVEL_MODULES edits, behavior-pin tests' },
{ title: 'Implement', detail: 'opus: PO siblings + handler; opus: WO siblings + handler; opus: all test plumbing + bundle-test + README — disjoint files', model: 'opus' },
{ title: 'Verify', detail: 'mechanical CHECKS (suite/goldens/ruff/synth/AST/zero cdk+shared+derived diff/ordering greps) + 3 fable lenses (behavior-preservation, boto3-laziness, save-merge-collapse)' },
{ title: 'Fix', detail: 'opus fixer, full re-verify, max 3 rounds', model: 'opus' },
{ title: 'Package', detail: 'single commit via -F (no push)', model: 'sonnet' },
],
}
// ---------------------------------------------------------------- constants
const REPO = '/Users/adammoussa/Documents/repositories/seahaven/procurement-ingest'
const BRANCH = 'feature/phase-5-handler-decomposition'
let _args = args
if (typeof _args === 'string') {
try { _args = JSON.parse(_args) } catch (e) { _args = null }
}
const BASE = (_args && _args.base) || 'main'
const CONSTRAINTS = `
PINNED BEHAVIORAL CONSTRAINTS (docs/refactor-evaluation.md Phase 5 — violating any is a build failure):
1. PO SPLIT INTO FLAT SIBLINGS under lambdas/po/email_processor/ (same dir,
flat-landing so bare-name imports keep working — the ses_auth/
template_parser/derive_all pattern):
handler.py -- event loop + fail-closed auth + email_type routing ONLY.
extraction.py -- extract_with_claude + _EMAIL_TAG_RE; imports
EXTRACTION_PROMPT from prompts. parse_raw_email is
ALREADY in shared/email_parsing.py after Phase 3 --
import it, do NOT recreate a PO copy.
enrichment.py -- enrich_parsed + pad_zip (PO-ONLY; WO has NO enrichment
stage). The derived-field block moves here as a PURE
code move.
telemetry.py -- the EMF ParseMethod emit wrappers (_emit_parse_method_metric
and the derived-agreement emit call surface used by
enrichment) -- but see constraint 4: derived_fields.py
and its shadow telemetry BEHAVIOR are untouchable.
persistence.py -- _write_fields / _merge_update / save_*. COLLAPSE the
byte-identical save_new_po/save_revision into ONE
_save_merge, behavior-identical to BOTH originals
INCLUDING the sticky-cancel ConditionExpression guard;
save_cancellation stays its own function.
2. WO SPLIT is ~5 concerns, NOT 7 (WO has no enrichment stage). Its split MUST
keep validate_ai_fallback AND the re.fullmatch(r"[0-9]+", work_order_id) key
guard in the handler loop AHEAD of BOTH save_work_order and save_event
(the guard protects the DynamoDB partition key + the '#'-delimited comment_id
range-key segment), and MUST keep _header_date_iso / comment_id determinism
WITH persistence (the retry-idempotent event_id key). Do not move the guard
below either save; do not fork comment_id determinism from save_event.
3. Move EXTRACTION_PROMPT (the large ~181-line prompt) to prompts.py with
cross-reference headers to derived_fields (the trade/site/fiscal rule tables
have a second authoritative copy there). KEEP a re-export in handler
('from prompts import EXTRACTION_PROMPT') since 4 tests dereference
handler.EXTRACTION_PROMPT. Do the same shape for WO if its prompt moves.
4. BEHAVIOR-PRESERVATION (immovable, pin with tests):
* PO emits ParseMethod=ai_fallback BEFORE the Bedrock call (handler loop:
_emit_parse_method_metric fires before extract_with_claude); the
ai_fallback_rejected emit is the additive second datapoint on rejection.
* WO emits AFTER its gate, with mutually-exclusive ai_fallback /
ai_fallback_rejected (a rejected WO email emits ONLY ai_fallback_rejected).
* Shadow DerivedFieldAgreement telemetry stays ai_fallback-only.
* derived_fields.py and the shadow-telemetry BEHAVIOR are UNTOUCHABLE (the
bake is in progress). Moving enrich_parsed into enrichment.py must be a
PURE code move -- byte-identical logic, ZERO behavior change.
git diff ${BASE}...HEAD -- lambdas/po/email_processor/derived_fields.py
MUST be empty.
5. SIGNATURES PRESERVED: handler(event, context) unchanged (event shape +
return contract identical) on BOTH pipelines; the save_* public contract
unchanged (the _save_merge collapse must be behavior-identical to both
save_new_po and save_revision -- callers route new_po/revision through it).
Goldens UNCHANGED.
6. BUNDLING: the Phase 0 top-level cp glob (cp ./*.py or cp <fn>/*.py) plus the
AST bundle-consistency test already ship & verify new flat siblings
automatically -- but the exact-set pin PO_EXPECTED_TOP_LEVEL_MODULES (and the
WO equivalent) must be updated for the new sibling modules IN THE SAME CHANGE,
keeping the test's teeth (a new sibling that is NOT added to the expected set,
or a commented-out cp, must still fail the AST test).
7. LAZY CACHED boto3 accessors in the I/O modules (extraction.py's bedrock,
persistence.py's dynamodb, handler.py's s3 -- whatever each module actually
uses); PURE modules (enrichment.py logic, prompts.py, telemetry.py if it
makes no AWS call) import NO boto3. Update the
monkeypatch.setattr(handler, "dynamodb", fake) patch surface in the SAME
change -- tests must patch the NEW module's accessor (e.g. the persistence
module), everywhere the tests patch it. Preserve the moto-before-handler
import ordering (_po_parser_support.py:26-33 -- verify current lines) or the
moto-backed suites hit real AWS.
8. ZERO cdk/ diff (git diff ${BASE}...HEAD -- cdk/ must be empty) and ZERO
lambdas/shared/ diff (Phase 3's shared modules are untouchable here).
9. UNTOUCHABLE FILES (git diff ${BASE}...HEAD must be empty for each):
lambdas/po/email_processor/derived_fields.py, both template_parser.py, both
validate_ai_fallback gate modules, lambdas/shared/*, cdk/*, and
lambdas/po/site_extractor/ + both web_ui/ (Phase 5 touches only the two
email processors). No opportunistic refactors -- every handler diff is a
pure move (delete-here/add-there), an import line, or the _save_merge
collapse per the binding spec.
10. AWS access is READ-ONLY (get-function, downloading deployed zips via
presigned URLs). NEVER cdk deploy, never invoke, never mutate.
`
const PREAMBLE = `
You are one of several agents building refactor Phase 5 in the git repo at
${REPO} on branch ${BRANCH} (already checked out — do NOT switch branches,
do NOT create branches, do NOT commit, NEVER push, do NOT run cdk deploy or
touch AWS resources beyond read-only calls).
Authoritative spec: docs/refactor-evaluation.md, section "Phase 5 — Handler
decomposition + lazy boto3 clients".
Work ONLY in the files you are told you own; other agents are concurrently
editing other files in this same working tree.
${CONSTRAINTS}
Your final message is consumed by an orchestrator script, not a human —
return only the structured data requested.
`
// ------------------------------------------------------------------ schemas
const RECON = {
type: 'object',
required: ['summary', 'facts'],
properties: {
summary: { type: 'string' },
facts: { type: 'array', items: { type: 'string' } },
blockers: { type: 'array', items: { type: 'string' } },
},
}
const SPEC = {
type: 'object',
required: ['poSplit', 'woSplit', 'promptMove', 'botoLaziness', 'saveMergeCollapse', 'testPlumbing', 'behaviorPins', 'notes'],
properties: {
poSplit: { type: 'string', description: 'per PO sibling (handler/extraction/enrichment/telemetry/persistence): exact functions/constants it owns, its imports, and the source file:line ranges moved into it — nothing else may change; state explicitly that parse_raw_email is imported from shared/email_parsing.py and enrich_parsed is a byte-identical move' },
woSplit: { type: 'string', description: 'per WO sibling (~5 concerns): exact ownership, and the pinned proof that validate_ai_fallback + the [0-9]+ key guard stay in the handler loop ahead of both saves and _header_date_iso/comment_id determinism stays with persistence' },
promptMove: { type: 'string', description: 'prompts.py contents (which prompt(s) move), the cross-reference headers to derived_fields, and the exact re-export line kept in each handler so handler.EXTRACTION_PROMPT still resolves (name the 4 tests that dereference it)' },
botoLaziness: { type: 'string', description: 'per module: which boto3 client/resource it needs, the lazy cached accessor shape (module-level None + get_x() cache), which modules import NO boto3, and the exact new monkeypatch target(s) that replace monkeypatch.setattr(handler, "dynamodb", fake) everywhere tests patch it' },
saveMergeCollapse: { type: 'string', description: 'the single _save_merge signature + body, with a line-by-line proof it is behavior-identical to BOTH save_new_po and save_revision including the sticky-cancel ConditionExpression path; how the router calls it for new_po vs revision' },
testPlumbing: { type: 'string', description: 'every test/support file change: loader/_SIBLING_MODULES resolution for the new siblings, _po_parser_support.py/_wo_parser_support.py edits (moto-before-handler ordering preserved), the monkeypatch patch-surface rewrites, PO_EXPECTED_TOP_LEVEL_MODULES + WO equivalent edits with mutation-detection shapes, and the new behavior-pin tests' },
behaviorPins: { type: 'string', description: 'the exact tests/assertions pinning: PO ai_fallback emit precedes the Bedrock invoke; WO emits mutually-exclusive; shadow telemetry ai_fallback-only; _save_merge parity; the WO key-guard ordering' },
notes: { type: 'string' },
},
}
const IMPL = {
type: 'object',
required: ['filesChanged', 'summary', 'checksRun'],
properties: {
filesChanged: { type: 'array', items: { type: 'string' } },
summary: { type: 'string' },
checksRun: { type: 'string' },
blockers: { type: 'array', items: { type: 'string' } },
},
}
const CHECKS = {
type: 'object',
required: ['passed', 'details'],
properties: {
passed: { type: 'boolean' },
details: { type: 'string' },
scopeViolations: { type: 'array', items: { type: 'string' } },
},
}
const FINDINGS = {
type: 'object',
required: ['findings'],
properties: {
findings: {
type: 'array',
items: {
type: 'object',
required: ['title', 'severity', 'confirmed', 'evidence', 'fix'],
properties: {
title: { type: 'string' },
severity: { enum: ['critical', 'high', 'medium', 'low'] },
confirmed: { type: 'boolean' },
evidence: { type: 'string' },
fix: { type: 'string' },
},
},
},
},
}
// ------------------------------------------------------------------- setup
phase('Setup')
const setup = await agent(`
In ${REPO}:
1. SEQUENCING GATE — Phases 0, 1, 2 AND 3 must all be on ${BASE}. Phase 5
relies on the flat-sibling lambdas/shared/ pattern + the Phase 0/2/3
bundling glob that auto-ships new siblings, and parse_raw_email now lives in
shared/email_parsing.py after Phase 3. git fetch origin, then pick the base
ref: origin/${BASE} if that remote ref exists, otherwise the local branch
${BASE} (a stacked local-only base is expected and fine). Verify on the base
ref:
(a) Phase 0: tests/test_bundle_consistency.py exists
(git show <baseref>:tests/test_bundle_consistency.py | head -3);
(b) Phase 1: validate_ai_fallback exists in the PO pipeline
(git grep validate_ai_fallback <baseref> -- lambdas/po);
(c) Phase 2: BOTH cdk/po_stack.py and cdk/wo_stack.py on the base ref
contain Code.from_asset("../lambdas") for the email processors
(git show <baseref>:cdk/po_stack.py | grep -n '\\.\\./lambdas', same for
wo_stack.py);
(d) Phase 3: lambdas/shared/ EXISTS on the base ref and the four shared
modules are present — ses_auth.py, web_ui_auth.py, email_parsing.py,
emf.py (git show <baseref>:lambdas/shared/email_parsing.py | head -3 for
each; and confirm parse_raw_email lives in shared/email_parsing.py, NOT
in the PO email_processor dir, on the base ref).
If ANY is missing — ESPECIALLY Phase 3 (no lambdas/shared/ or no
email_parsing.py on base) — STOP with a blocker naming the unmet phase and
do nothing else.
2. Verify clean working tree (untracked .coverage / .claude/ / the local
44 MB lambdas/po/email_processor/package/ dir are fine; any OTHER dirt =
blocker, never stash or discard).
3. git checkout ${BASE}; then git pull --ff-only ONLY if the branch has an
upstream (a local-only base skips the pull — not a blocker); then
git checkout -b ${BRANCH}
4. gh pr list --state open --json number,title,headRefName (overlap check).
Return facts: HEAD sha, per-phase gate evidence (especially the Phase 3
shared/ + email_parsing.py proof), open PRs, blockers.
`, { label: 'setup:branch', model: 'haiku', schema: RECON })
if (!setup || (setup.blockers && setup.blockers.length)) {
return { aborted: 'setup blockers', blockers: setup ? setup.blockers : ['setup agent died'], facts: setup ? setup.facts : [] }
}
log(`Branch ${BRANCH} ready off ${BASE}. ${setup.summary}`)
// ------------------------------------------------------------------- recon
phase('Recon')
const recon = await parallel([
() => agent(`${PREAMBLE}
Read-only recon of the PO email-processor God-handler
(lambdas/po/email_processor/handler.py) and its decomposition seams:
1. Quote with file:line the exact boundaries of every function/constant that
moves and where it lands per constraint 1: EXTRACTION_PROMPT + _EMAIL_TAG_RE
+ extract_with_claude (-> extraction.py/prompts.py); _emit_parse_method_metric
+ _derived_agreement + _emit_derived_agreement_metric (-> telemetry.py, but
note which of these the derived-agreement path needs and whether any lives in
derived_fields.py per constraint 4); enrich_parsed + pad_zip (-> enrichment.py);
_write_fields + _merge_update + save_new_po + save_revision + save_cancellation
(-> persistence.py).
2. save_new_po vs save_revision: diff their BODIES line by line and prove they
are byte-identical modulo the docstring/log string (the D9 collapse target).
Quote both verbatim so the spec can pin _save_merge.
3. The handler loop (handler.py:662-758 currently): quote the exact ordering —
healthcheck early-return, Records loop, s3.get_object, authenticate_inbound_email,
parse_raw_email, try_deterministic_parse, _emit_parse_method_metric (PRE-Bedrock),
extract_with_claude, validate_ai_fallback, ai_fallback_rejected emit + continue,
po_number guard, enrich_parsed, email_type routing to save_*.
4. Which boto3 clients each moving chunk uses (s3, dynamodb resource, bedrock
client) and the current module-level client lines (handler.py:27-29).
5. The current 'from derived_fields import derive_all',
'from ses_auth import authenticate_inbound_email',
'from template_parser import ...' import lines and what parse_raw_email's
import will become (shared/email_parsing.py after Phase 3 -- confirm on the
working tree whether it is already imported from there).
6. Which tests dereference handler.EXTRACTION_PROMPT (grep) and which patch
handler.dynamodb / handler.bedrock / handler.s3.
25-35 precise facts.`,
{ label: 'recon:po-handler', model: 'sonnet', phase: 'Recon', schema: RECON }),
() => agent(`${PREAMBLE}
Read-only recon of the WO email-processor God-handler
(lambdas/wo/email_processor/handler.py) and its ~5-concern split:
1. Quote with file:line every function/constant and the concern it belongs to:
EXTRACTION_PROMPT + _EMAIL_TAG_RE + extract_with_bedrock; emit_parse_metric;
save_work_order; _header_date_iso + save_event; the handler loop.
2. The handler loop (handler.py:350-435 currently): quote the exact ordering —
healthcheck, Records loop, s3.get_object, authenticate_inbound_email,
parse_raw_email, try_deterministic_parse, extract_with_bedrock,
validate_ai_fallback, ai_fallback_rejected emit + continue, emit_parse_metric
(POST-gate), the re.fullmatch(r"[0-9]+", work_order_id) key guard,
save_work_order, save_event. CRITICAL per constraint 2: pin that the gate +
the [0-9]+ key guard both precede BOTH saves, and that comment_id determinism
(_header_date_iso + the object-key sha suffix, handler.py:277-347) is coupled
to save_event.
3. Prove WO has NO enrichment stage (no enrich_parsed equivalent) so its split
is 5 concerns not 7.
4. WO's EMF emit is mutually-exclusive (rejected email emits ONLY
ai_fallback_rejected then continues; accepted email emits emit_parse_metric
once AFTER the gate) -- quote the exact lines proving this, contrasted with
PO's pre-Bedrock double-count.
5. boto3 clients WO uses (handler.py:27-29) and which WO tests patch
handler.dynamodb / handler.bedrock and dereference handler.EXTRACTION_PROMPT.
6. _wo_parser_support.py bare-sys.path 'import handler' strategy and how it
will resolve the new siblings.
20-30 precise facts.`,
{ label: 'recon:wo-handler', model: 'sonnet', phase: 'Recon', schema: RECON }),
() => agent(`${PREAMBLE}
Read-only recon of the bundling glob, the AST bundle-consistency test, and the
test patch/loader surface this phase must rewire:
1. tests/test_bundle_consistency.py: quote PO_EXPECTED_TOP_LEVEL_MODULES
(currently {"handler","ses_auth","template_parser","derived_fields"} on
pre-Phase-3 — read the value ON THIS BRANCH after Phase 3 has added the
shared modules), whether a WO_EXPECTED_TOP_LEVEL_MODULES (or equivalent)
exists, the _first_party_sibling_imports AST logic, the exact-set assertion,
and every mutation-detection assertion (a commented-out cp / a missing module
in the expected set must fail). Confirm the Phase 0 cp glob form in both
stacks (cp ./*.py or cp <fn>/*.py) auto-ships NEW flat siblings with no cdk
edit needed this phase.
2. tests/conftest.py importlib loader: how it resolves module paths, the
sys.modules save/restore, _SIBLING_MODULES exact current contents (which
bare names it lists per pipeline).
3. lambdas/po/email_processor/tests/_po_parser_support.py: the independent
importlib reimplementation, the load-bearing moto-before-handler import order
(~lines 26-33 — quote it), and how handler + its new siblings get loaded.
4. lambdas/wo/email_processor/tests/_wo_parser_support.py: same.
5. Every place tests do monkeypatch.setattr(handler, "dynamodb"/"bedrock"/"s3",
fake) — grep both pipelines' tests and quote each file:line (this is the
patch surface constraint 7 says must move to the new module accessors).
6. Which tests reference handler.save_new_po / handler.save_revision /
handler.enrich_parsed / handler.extract_with_claude / handler.EXTRACTION_PROMPT
directly (they may need repointing after the split).
20-30 facts.`,
{ label: 'recon:bundling-tests', model: 'sonnet', phase: 'Recon', schema: RECON }),
() => agent(`${PREAMBLE}
Read-only AWS recon (region us-east-1, READ-ONLY): for po-email-processor and
workorder-email-processor (find exact function names via the stacks/aws lambda
list-functions) run aws lambda get-function; download each Code.Location
presigned zip to a scratch dir and produce the COMPLETE top-level file list
(unzip -l) separated into (a) first-party top-level .py, (b) dependency dirs,
(c) anything else; record each function's CodeSha256. This deployed first-party
module set is the baseline the post-split staged bundle is judged against: after
the split the staged bundle's first-party .py set must be exactly the deployed
set MINUS nothing PLUS the new sibling modules (extraction/enrichment/telemetry/
persistence/prompts for PO, the WO equivalents) and with the collapsed
save_new_po/save_revision still resolvable via imports (no NEW top-level module
should vanish that a handler still imports). Confirm parse_raw_email is NOT a
top-level module in either deployed email-processor zip if Phase 3 already
shipped it from shared/ (it lands flat as email_parsing.py). Return the
categorized lists as facts.`,
{ label: 'recon:deployed-baseline', model: 'sonnet', phase: 'Recon', schema: RECON }),
])
const reconOk = recon.filter(Boolean)
const pack = reconOk.map(r => `## ${r.summary}\n${r.facts.join('\n')}`).join('\n\n')
const reconBlockers = reconOk.flatMap(r => r.blockers || [])
.filter(b => b && !/^\s*(none|n\/a)\b/i.test(b))
log(`Recon complete: ${reconOk.length}/4 mappers, ${reconBlockers.length} advisory notes`)
// Recon "blockers" here are downstream guidance (verify the WO dynamodb
// monkeypatch surface per constraint 7 before persistence.py's split;
// compare the post-split staged bundle against the local pre-Phase-5 bundle,
// NOT the stale pre-Phase-3 deployed AWS zip), NOT stop conditions — main-loop
// verified Phase 3 is on base and the seams move cleanly. The real gates are
// pytest-green (catches any missed monkeypatch surface) and the bundle-parity
// verifier. Fold the notes into the spec context instead of aborting.
const reconAdvisories = reconBlockers.length
? `\n\nRECON ADVISORIES (recon flagged these; they are downstream verification guidance — the impl must update the full dynamodb/bedrock/s3 monkeypatch surface (constraint 7) so no test breaks, and the parity check compares against the local pre-Phase-5 staged bundle not the pre-Phase-3 deployed zip — NOT reasons to stop; pytest-green + bundle-parity are the real gates):\n- ${reconBlockers.join('\n- ')}`
: ''
// -------------------------------------------------------------------- spec
phase('Spec')
const spec = await agent(`${PREAMBLE}
You are the SPEC agent — the single authority that pins every contested
decision BEFORE parallel implementation (parallel leaves cannot see each
other's choices). Using the recon pack below plus your own reads of the actual
files, produce the binding implementation spec:
- poSplit: per constraint 1's five PO siblings — exact functions/constants each
owns, the imports each needs, and the source line ranges moved. State
explicitly: parse_raw_email is imported from shared/email_parsing.py (NOT
recreated); enrich_parsed + its derived-field block move into enrichment.py as
a BYTE-IDENTICAL logic move (constraint 4). Resolve where the derived-agreement
emit surface lives (telemetry.py vs left in enrichment) WITHOUT touching
derived_fields.py.
- woSplit: per constraint 2's ~5 WO concerns — ownership, plus the explicit pin
that validate_ai_fallback + the [0-9]+ key guard stay in the handler loop
AHEAD of both saves and _header_date_iso/comment_id determinism stays coupled
to persistence (save_event). WO has NO enrichment stage.
- promptMove: prompts.py contents (PO prompt for sure; WO prompt too if it moves),
the cross-reference headers to derived_fields, and the EXACT re-export line
each handler keeps so handler.EXTRACTION_PROMPT still resolves (name the 4 PO
tests + any WO tests that dereference it).
- botoLaziness: per module, the lazy cached boto3 accessor shape (module-level
None cache + get_x()); which modules import NO boto3 (pure: enrichment logic,
prompts, telemetry if it emits via print only); and the EXACT new monkeypatch
target(s) replacing monkeypatch.setattr(handler, "dynamodb", fake) at every
test site recon found — pick ONE consistent rule (patch the owning module's
accessor/attribute) and state why. Preserve the moto-before-handler ordering.
- saveMergeCollapse: the single _save_merge signature + body with a line-by-line
proof it is behavior-identical to BOTH save_new_po and save_revision including
the sticky-cancel ConditionExpression path; how the router dispatches new_po
vs revision through it; save_cancellation stays separate.
- testPlumbing: every file, every edit — _SIBLING_MODULES additions for the new
bare names, _po_parser_support.py / _wo_parser_support.py loader edits (moto
ordering preserved), the monkeypatch patch-surface rewrites, repointing any
test that referenced handler.save_new_po/enrich_parsed/etc. to the new module,
PO_EXPECTED_TOP_LEVEL_MODULES + the WO equivalent updated to the new sibling
set (with mutation shapes: a missing sibling in the set, or a commented-out cp,
must FAIL the AST test), and the NEW behavior-pin tests.
- behaviorPins: the exact new tests/assertions pinning constraint 4 & 5 — PO
ai_fallback emit precedes the Bedrock invoke (a throttle test is the cleanest
proof), WO emits mutually-exclusive, shadow telemetry ai_fallback-only,
_save_merge parity to both originals, WO key-guard ordering ahead of both saves.
Recon pack:\n${pack}${reconAdvisories}`,
{ label: 'spec:pin-decomposition', phase: 'Spec', schema: SPEC })
if (!spec) return { aborted: 'spec agent died — rerun workflow', reconBlockers }
const specBlock = `BINDING SPEC (from the spec agent — implement EXACTLY this):\n${JSON.stringify(spec, null, 2)}`
log('Spec pinned: PO split, WO split, prompt move, boto3 laziness, _save_merge collapse, test plumbing, behavior pins')
// --------------------------------------------------------------- implement
phase('Implement')
const impl = await parallel([
() => agent(`${PREAMBLE}
YOU OWN: lambdas/po/email_processor/ EXCEPT its tests/ directory (another agent
owns all test and support files). Do not touch cdk/, lambdas/wo/,
lambdas/shared/, or README.
Task: execute the PO split per spec.poSplit + spec.promptMove +
spec.botoLaziness + spec.saveMergeCollapse:
1. Create prompts.py (EXTRACTION_PROMPT + cross-ref headers); keep the
re-export line in handler.py.
2. Create extraction.py (extract_with_claude + _EMAIL_TAG_RE, lazy bedrock
accessor, imports EXTRACTION_PROMPT from prompts and parse_raw_email from
the shared email_parsing module).
3. Create enrichment.py (enrich_parsed + pad_zip + the derived-field block as a
BYTE-IDENTICAL logic move — derived_fields.py stays untouched; no boto3).
4. Create telemetry.py (the EMF ParseMethod emit wrappers per spec) — pure
(print-only, no boto3) unless spec says otherwise.
5. Create persistence.py (_write_fields, _merge_update, the collapsed _save_merge
behavior-identical to BOTH save_new_po and save_revision incl. the sticky-
cancel guard, save_cancellation; lazy dynamodb accessor).
6. Rewrite handler.py down to the event loop + auth + routing, importing from
the new siblings; keep handler(event, context) signature + the healthcheck
early-return + the PRE-Bedrock ai_fallback emit ordering EXACTLY.
Every change is a pure move / import line / the _save_merge collapse — no
opportunistic refactors (constraint 9). derived_fields.py diff MUST stay empty.
Run before returning: ruff check lambdas/po && ruff format lambdas/po --check,
plus an import smoke: python3 -c with sys.path prepended for lambdas/shared +
the PO function dir, importing every new module + handler (tests are NOT yours
to run — the plumbing agent lands them in parallel).
${specBlock}`,
{ label: 'impl:po-siblings', model: 'opus', phase: 'Implement', schema: IMPL }),
() => agent(`${PREAMBLE}
YOU OWN: lambdas/wo/email_processor/ EXCEPT its tests/ directory (another agent
owns all test and support files). Do not touch cdk/, lambdas/po/,
lambdas/shared/, or README.
Task: execute the WO ~5-concern split per spec.woSplit + spec.promptMove +
spec.botoLaziness:
1. Create prompts.py if the WO prompt moves (re-export kept in handler).
2. Create extraction.py (extract_with_bedrock + _EMAIL_TAG_RE, lazy bedrock).
3. Create telemetry.py (emit_parse_metric) — pure unless spec says otherwise.
4. Create persistence.py (save_work_order, _header_date_iso, save_event —
comment_id determinism stays coupled here; lazy dynamodb).
5. Rewrite handler.py to the event loop + auth + routing, KEEPING
validate_ai_fallback + the re.fullmatch(r"[0-9]+", work_order_id) key guard
in the loop AHEAD of BOTH saves (constraint 2), the POST-gate mutually-
exclusive emit ordering, the healthcheck early-return, and handler(event,
context) signature EXACTLY.
Pure moves / imports only (constraint 9). No enrichment stage exists — do not
invent one.
Run before returning: ruff check lambdas/wo && ruff format lambdas/wo --check,
plus an import smoke: python3 -c with sys.path prepended for lambdas/shared +
the WO function dir, importing every new module + handler.
${specBlock}`,
{ label: 'impl:wo-siblings', model: 'opus', phase: 'Implement', schema: IMPL }),
() => agent(`${PREAMBLE}
YOU OWN: all test and support files (tests/ at repo root including
tests/test_bundle_consistency.py and tests/conftest.py,
lambdas/po/email_processor/tests/, lambdas/wo/email_processor/tests/) and
README.md ONLY.
Task A: apply spec.testPlumbing in full — _SIBLING_MODULES additions for the new
bare names, _po_parser_support.py / _wo_parser_support.py loader edits
(PRESERVING the moto-before-handler ordering comment + import), rewrite the
monkeypatch patch surface everywhere tests patch handler.dynamodb/bedrock/s3 to
the new module accessors per spec.botoLaziness, repoint any test that referenced
handler.save_new_po / handler.save_revision / handler.enrich_parsed /
handler.extract_with_claude to the new modules (KEEP handler.EXTRACTION_PROMPT
resolving via the re-export — do NOT repoint those 4 tests), and update
PO_EXPECTED_TOP_LEVEL_MODULES + the WO equivalent to the new sibling set WITHOUT
losing teeth (constraint 6 — a missing sibling in the set or a commented-out cp
must FAIL the AST test; keep existing mutation tests green).
Task B: add the NEW behavior-pin tests per spec.behaviorPins — PO ai_fallback
emit precedes the Bedrock invoke (throttle-based proof), WO emits mutually-
exclusive, shadow telemetry ai_fallback-only, _save_merge parity to both
originals, WO key-guard ordering ahead of both saves.
Task C: README — update the repo-layout / module map for the new PO+WO sibling
modules (the flat-landing import rule already ships them via the Phase 0 glob),
note the save_new_po/save_revision collapse into _save_merge and prompts.py.
Match existing README style; goldens stay UNCHANGED.
Run before returning: pytest -q --no-cov at repo root (must be green against the
OTHER agents' edits — they land in parallel; poll by re-running up to ~10 min
before reporting a blocker) and ruff check on every file you touched.
${specBlock}`,
{ label: 'impl:test-plumbing-readme', model: 'opus', phase: 'Implement', schema: IMPL }),
])
const implOk = impl.filter(Boolean)
const implBlockers = implOk.flatMap(r => r.blockers || [])
log(`Implement complete: ${implOk.length}/3 agents, blockers: ${implBlockers.length}`)
// ---------------------------------------------------- verify + fix loop
const EXPECTED_SCOPE = [
'lambdas/po/email_processor/',
'lambdas/wo/email_processor/',
'tests/',
'README.md',
]
const mechanicalPrompt = `${PREAMBLE}
Independent re-verification — trust nothing self-reported. Run ALL gates,
quoting failures verbatim:
1. pytest -q --no-cov (repo root, all three roots green; GOLDENS UNCHANGED —
git diff ${BASE}...HEAD on any golden .json/.eml fixture must be empty).
2. ruff check . && ruff format --check .
3. cd cdk && npx cdk synth po-ingest -q && npx cdk synth workorder-ingest -q
(artifact-id selectors, NOT stack_name). The Docker bundling runs — confirm
each staged email-processor asset now contains the new sibling modules.
4. AST BUNDLE TEST: the test_bundle_consistency.py suite is GREEN and its teeth
hold — on a scratch copy, comment out the cp glob line and drop a sibling
from PO_EXPECTED_TOP_LEVEL_MODULES (and the WO equivalent); EACH mutation
must FAIL the test. Restore.
5. ZERO-DIFF UNTOUCHABLES (each must output NOTHING):
git diff ${BASE}...HEAD -- cdk/
git diff ${BASE}...HEAD -- lambdas/shared/
git diff ${BASE}...HEAD -- lambdas/po/email_processor/derived_fields.py
git diff ${BASE}...HEAD -- lambdas/po/email_processor/template_parser.py
git diff ${BASE}...HEAD -- lambdas/wo/email_processor/template_parser.py
plus both validate_ai_fallback gate modules (resolve their filenames first)
and lambdas/po/site_extractor/ + both web_ui/.
6. SIGNATURES: grep proves handler(event, context) unchanged on both pipelines
and the healthcheck early-return + return {"statusCode": 200, ...} contract
intact. handler.EXTRACTION_PROMPT still resolves (the re-export line is
present in both handlers).
7. ORDERING PINS (grep with line numbers):
- PO: _emit_parse_method_metric / the ai_fallback emit still PRECEDES the
extract_with_claude (Bedrock) invoke in the handler loop.
- WO: the emit stays POST-gate and mutually-exclusive (rejected path emits
ONLY ai_fallback_rejected then continue); the re.fullmatch(r"[0-9]+", ...)
key guard still sits AHEAD of BOTH save_work_order and save_event.
8. BOTO3 LAZINESS: grep proves the pure modules (PO enrichment.py, prompts.py,
telemetry if print-only; WO equivalents) import NO boto3, and the I/O modules
use lazy cached accessors (no module-level eager client that a test cannot
patch). No monkeypatch.setattr(handler, "dynamodb"/... ) remains pointing at
a now-empty handler attribute.
9. git status --porcelain scope check: every modified/added/deleted path under
${EXPECTED_SCOPE.join(', ')} (untracked .coverage/.claude/package/ tolerated).
passed=true only if all green. YOU MAY NOT edit files.`
const lenses = [
{ key: 'behavior-preservation', prompt: `${PREAMBLE}
ADVERSARIAL REVIEW — behavior-preservation lens. The decomposition must not
change one observable byte of metric or telemetry behavior. (1) METRIC ORDERING:
read the post-split PO handler loop and prove _emit_parse_method_metric fires
BEFORE extract_with_claude (so a Bedrock-side throttle still records ai_fallback)
and that ai_fallback_rejected is still the additive second datapoint; then the
WO loop and prove the emit is POST-gate and MUTUALLY-EXCLUSIVE (a rejected email
emits ONLY ai_fallback_rejected, an accepted one emits once after the gate).
Construct the emitted EMF envelope from each moved emitter and diff it against
${BASE}'s output — namespace, dimension-set list [["ParseMethod"],["ParseMethod",
"TemplateId"]], property keys/types identical. (2) SHADOW TELEMETRY: prove
DerivedFieldAgreement is still emitted ONLY on the ai_fallback path after
enrich_parsed moved to enrichment.py, and that
git diff ${BASE}...HEAD -- derived_fields.py is EMPTY (the enrich_parsed move is
byte-identical logic). (3) ROUTING: prove email_type routing (cancellation/
revision/new_po for PO; the WO save_work_order + save_event pair) still reaches
the same saves in the same order. confirmed=true only with a concrete
file:line proof or a sample-emission diff.` },
{ key: 'boto3-laziness', prompt: `${PREAMBLE}
ADVERSARIAL REVIEW — boto3-laziness lens. (1) PURE MODULES: prove the modules
constraint 7 calls pure (PO enrichment.py, prompts.py, telemetry if print-only;
WO equivalents) import NO boto3 at all — grep + read. A stray 'import boto3' or
eager client in a "pure" module is a finding. (2) LAZY ACCESSORS: in the I/O
modules (extraction bedrock, persistence dynamodb, handler s3) prove the client
is created lazily and cached (module-level None + get_x()), not eagerly at import
— an eager module-level client that tests cannot patch is a finding. (3)
PATCH-SURFACE COMPLETENESS: enumerate EVERY test that previously did
monkeypatch.setattr(handler, "dynamodb"/"bedrock"/"s3", fake) and prove each now
patches the correct NEW module accessor — a test that still patches a now-dead
handler attribute silently exercises the REAL client (or moto-misses) and is a
critical finding. (4) MOTO ORDERING: prove the moto-before-handler import order
survived in _po_parser_support.py and _wo_parser_support.py — reorder-and-run to
confirm the suites would hit real AWS without it. confirmed=true only with
file:line or reproduced-failure evidence.` },
{ key: 'save-merge-collapse', prompt: `${PREAMBLE}
ADVERSARIAL REVIEW — save-merge-collapse lens. The one _save_merge that replaces
save_new_po AND save_revision must be byte-behavior-identical to BOTH. (1) Read
_save_merge and BOTH ${BASE} originals; prove line-by-line the field-building,
the _merge_update call, and the sticky-cancel path are identical (incl. the
ConditionExpression "attribute_not_exists(po_status) OR po_status <>
:__cancelled_marker" and the ConditionalCheckFailedException fallback that
re-writes WITHOUT po_status/cancelled_at). (2) Prove the router dispatches
email_type=="revision" and the new_po else-branch both through _save_merge, and
that save_cancellation stayed SEPARATE and unchanged. (3) Construct a scenario
where the collapse could diverge: a revision landing on a Cancelled PO, a new_po
with po_status=None, a new_po carrying a non-cancelled status onto a Cancelled
skeleton — and prove _save_merge behaves exactly as the original would have for
each. (4) Confirm the save_* public contract callers rely on is unchanged and
the goldens still pass. confirmed=true only with a concrete divergence sketch
or a line-by-line equivalence proof.` },
]
let round = 0
let checks = null
let confirmed = []
while (round < 3) {
phase('Verify')
const results = await parallel([
() => agent(mechanicalPrompt, { label: `verify:mechanical-r${round}`, model: 'sonnet', phase: 'Verify', schema: CHECKS }),
...lenses.map(l => () =>
agent(l.prompt, { label: `verify:${l.key}-r${round}`, phase: 'Verify', schema: FINDINGS })),
])
checks = results[0]
confirmed = results.slice(1).filter(Boolean)
.flatMap(r => r.findings || [])
.filter(f => f.confirmed && f.severity !== 'low')
const green = checks && checks.passed
log(`Verify round ${round}: mechanical ${checks && checks.passed ? 'GREEN' : 'RED'}, confirmed findings: ${confirmed.length}`)
if (green && confirmed.length === 0) break
round += 1
if (round >= 3) break
phase('Fix')
await agent(`${PREAMBLE}
You are the fix agent — you may edit files under: ${EXPECTED_SCOPE.join(', ')}.
Fix EVERY item below minimally; the binding spec and 10 pinned constraints
still hold (a finding that conflicts with a constraint is reported, not "fixed"
— the constraint wins, esp. constraint 4's derived_fields untouchability +
byte-identical enrich move, constraint 5's signature/goldens preservation, and
constraint 8's zero cdk/shared diff). Re-run the specific failing gate/test per
fix.
MECHANICAL:\n${checks ? checks.details : '(agent died — rerun all gates)'}
CONFIRMED FINDINGS:\n${JSON.stringify(confirmed, null, 2)}
${specBlock}`,
{ label: `fix:round-${round}`, model: 'opus', phase: 'Fix', schema: IMPL })
}
const verifyClean = checks && checks.passed && confirmed.length === 0
if (!verifyClean) {
return {
status: 'NEEDS ATTENTION — verify not clean after 3 rounds; branch left uncommitted',
branch: BRANCH,
mechanical: checks,
unresolvedFindings: confirmed,
implBlockers,
reconBlockers,
spec,
}
}
// ----------------------------------------------------------------- package
phase('Package')
const commit = await agent(`${PREAMBLE.replace('do NOT commit, ', '')}
YOU are the commit agent:
1. Read ~/Documents/repositories/seahaven/engineering-handbook/commit-messages.md
and follow it exactly.
2. git add only paths under: ${EXPECTED_SCOPE.join(', ')} and
.claude/workflows/phase-5-handler-decomposition.js. NOT .coverage, NOT
package/. Verify the staged set with git status — the new sibling files AND
the deletions of the moved-out code from the handlers MUST be staged.
3. ONE commit; write the message to /tmp/phase5-commit-msg.txt and use
git commit -F /tmp/phase5-commit-msg.txt (backticks in -m get eaten by zsh).
Suggested subject:
"feat: decompose email-processor handlers into flat siblings + lazy boto3 clients (refactor phase 5)"
Body: the PO 5-way split (handler/extraction/enrichment/telemetry/persistence)
with the save_new_po/save_revision -> _save_merge collapse one-liner, the WO
~5-concern split keeping the gate+key-guard in the loop, prompts.py + the
handler.EXTRACTION_PROMPT re-export, the lazy-boto3 + patch-surface move, and
the behavior-preservation pins (PO pre-Bedrock emit, WO mutually-exclusive,
shadow telemetry ai_fallback-only, derived_fields untouched). NO AI
attribution / Co-Authored-By lines.
4. Do NOT push. Return commit sha + shortstat in summary.`,
{ label: 'package:commit', model: 'sonnet', phase: 'Package', schema: IMPL })
return {
status: 'BUILT — committed locally, NOT pushed',
branch: BRANCH,
base: BASE,
commit: commit ? commit.summary : 'commit agent died — commit manually',
spec: { poSplit: spec.poSplit, woSplit: spec.woSplit, saveMergeCollapse: spec.saveMergeCollapse, botoLaziness: spec.botoLaziness, notes: spec.notes },
implementation: implOk.map(r => r.summary),
filesChanged: implOk.flatMap(r => r.filesChanged),
verifyRounds: round + 1,
blockers: implBlockers.concat(reconBlockers),
outstandingGates: [
'/sh-security-review (MANDATORY before push — the untrusted-input parse/extract/validate-gate/auth-routing code is being MOVED across modules; CLAUDE.md gates untrusted-input handling changes and a move still touches the surface. Run on the committed diff, re-run after any post-review fix)',
'cross-family cross_review.py NOT required (handler(event, context) event/return signatures unchanged and no IAM change — internal module moves only)',
'push + PR + gh pr checks green',
'deploy-then-merge: full suite goldens unchanged, smoke green, ONE live email per pipeline (PO + WO) showing the correct ParseMethod, both the fallback + rejected alarm series behaving; the post-merge CI redeploy is a fresh asset (new sibling files) so confirm the smoke gate + AST bundle test caught any missing-module regression BEFORE merge',
],
}