procurement-ingest/.claude/workflows/phase-8-test-consolidation.js
Adam Moussa ff368beaa4
test: consolidate test roots — one loader, shared support, enforced CI floor (phase 8) (#118)
* test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)

tests/conftest.py only loads for the tests/ root, not a standalone
`pytest lambdas/po/email_processor/tests` run, so it could never carry
session invariants like the dummy AWS env or the moto stubber
registration. Add a single repo-root conftest.py (pytest.ini pins
rootdir there, so it loads for every invocation) that sets the dummy
AWS credentials/region, imports moto BEFORE any handler module so
boto3 sessions pick up its stubber hook (carrying the explanatory
comment verbatim from the old _po_parser_support.py), and exposes one
load_lambda_module(pipeline, name) — the sys.modules save/restore
dance stays, since template_parser is still a duplicated bare name
across pipelines needing per-exec sibling binding.

Add tests/support/ as the shared package both pipelines' local
_*_parser_support.py modules delegate to: a superset FakeTable (PO's
update_item recording + WO's put_item and keyed single-row store),
FakeDynamoResource, load_email, and load_golden with parse_float=Decimal
kept (load-bearing for exact money comparison at PO magnitudes — WO's
prior load_golden had no parse_float and must not regress PO by losing
it). Rewrite _wo_parser_support.py off the bare `import handler` /
`from handler import parse_raw_email` strategy that was the source of
the bare-name sys.modules collision the other two loaders defend
against.

Move test_po_merge.py and test_pad_zip.py into
lambdas/po/email_processor/tests/ (PO-specific, belongs beside the
code) via git mv so history follows; test_parse_raw_email.py and
test_ses_auth.py stay at the repo root since they're genuinely
cross-pipeline, parameterized over both handlers. Delete
tests/test_local.py: it globs a nonexistent samples/ dir, is WO-only,
and imports a handler at collection time, bypassing the loader gate
entirely — the golden suites already cover its role. Its pytest.ini
exclusion comment goes with it.

New scenario coverage, all built on the single loader + support
package:
- PO+WO Bedrock transport errors (ThrottlingException, missing
  'content' key, empty content list, non-JSON model text), asserting
  PO's pre-call ai_fallback metric survives with no partial write and
  the exception propagates; WO's no-datapoint-on-throttle behavior is
  pinned with a documenting test rather than "fixed" by reordering.
- Handler-level SES-auth reject seam per pipeline: no auth
  monkeypatch + empty ALLOWED_DKIM_DOMAINS asserts zero Bedrock calls,
  zero writes, no raise — closing the hole where deleting the gate
  line today still passes every test.
- web_ui coverage for both PO and WO (0% before this): fail-closed on
  unset ARN and on a Secrets Manager exception, TTL cache refresh,
  Bearer/X-Auth-Token/header-case-insensitivity, wrong-token 401 with
  no table scan, non-ASCII token, and a hostile-field-escaping
  regression lock. PO web_ui has no __init__.py, so these go through
  the loader rather than package imports.
- A moto-backed mirror of test_po_merge for WO merge semantics
  (table 'WorkOrders'): null-status never clobbers wo_status,
  created_at immutable via if_not_exists, status->wo_status mapping,
  None fields absent from SET, record_type only-when-present.
- Small pins: the PO-DC-02 64-char EMF clamp regression and
  per-pipeline multi-record failure-isolation (all-or-retry contract).
  The reprocess.py synthetic-event-shape contract test already landed
  in Phase 7, so it isn't duplicated here.

Two WO product-code fixes ride along, since this is the phase that
exercises them: (a) the invalid_status reason-code fix in
template_parser.py's status check, which previously returned
malformed_site_code for the same failure validate_ai_fallback already
labels invalid_status, making one failure surface two codes depending
on path (grepped the dashboards/metric filters for
malformed_site_code first — no external references found, safe to
diverge the two codes); (b) wrapping the WO Bedrock call in
handler.py so a transport failure emits ai_fallback/bedrock_error in
an except-and-reraise. This is deliberately not a naive reorder: the
emit sits in the except block, not pre-call, so a gate-rejected email
still emits only ai_fallback_rejected and wo_stack's "a rejected
email emits nothing else" alarm contract doesn't double-count. A test
computes the emitted series by hand to pin the no-double-count
behavior. Neither change touches the handler event/return contract.

_validate_new_po_values in the PO template_parser.py is split into
per-rule helpers, and the V4 anchor-frame dataclass now carries
summary_matches/price so V13 can consume them; extract_new_po
(C901=35) is included in the split. Add ruff.toml enabling C901/PLR
so the mccabe/complexity suppressions scattered through the tree stop
being decorative; derived_fields.py is under the shadow-bake freeze
so its violations are silenced via a per-file ignore with a
justification comment instead of an in-file edit, and the handful of
other pre-existing violations surfaced by turning the config on get
the same per-file-ignore treatment with a reason, or a fix where the
file isn't frozen. scripts/ is added to the CI lint scope.

CI gains an explicit --cov module list (lambdas/po and wo
email_processor + web_ui, po/site_extractor, lambdas/shared) plus
--cov-fail-under=80, since web_ui and site_extractor lack __init__.py
markers and a bare --cov=lambdas silently skips them for the missing
package marker; .coveragerc omits the test dirs themselves from the
count. The Phase 0 AST bundle-consistency test stays in the standard
pytest run. .gitignore picks up the resulting .coverage data file.

docs/po-template-parser.md gets a small correction: the EXTRACTION_PROMPT
declares quantity/price as "number or null", not JSON strings, so
parse_float=Decimal already handles a conforming Bedrock response —
the doc previously implied the coercion path was the primary
mechanism rather than a defensive net for non-conforming responses.

* test: lock attribute-context quote escaping in web_ui hostile-field test

The escaping regression lock asserted only the element-context vector
(raw <script> absent, &lt;script&gt; present) while its docstring claimed
quotes were covered -- the payload's " and ' were never asserted on, so
a quote-escaping regression on the onclick row-link sink (attribute
breakout -> event-handler injection) would have passed green.
/sh-security-review finding WC-01 (confirmed medium, test-integrity).

Add assertions that the onclick sink's JSON string renders its opening
quote as &quot; (raw " after window.location= fails), that the
payload's quote characters appear only entity-escaped, and that the
raw payload never appears anywhere in the body. Mutation-verified: the
test now fails when the sink's quote-escaping is dropped.

* test: address Open SWE review — xfail the web_ui non-ASCII auth pin, document subset coverage-floor override

- tests/test_web_ui_auth.py: replace the TypeError characterization pin with an
  xfail(strict, raises=TypeError) asserting the DESIRED fail-closed (False)
  behavior. Documents the intended fix and auto-fails (xpass) once web_ui_auth is
  corrected, instead of requiring a passing test to be knowingly deleted. The
  module stays frozen this phase; the underlying hmac.compare_digest ASCII-only
  defect is tracked as a follow-up.
- pytest.ini: document that the aggregate 80% floor (enforced in CI via the
  reusable workflow's bare pytest) red-exits local subset runs by design, with the
  --cov-fail-under=0 override for iteration. Floor stays in addopts because the
  centralized ci-python-sam workflow exposes no per-run test command.
2026-07-20 16:19:15 -04:00

712 lines
46 KiB
JavaScript

export const meta = {
name: 'phase-8-test-consolidation',
description: 'Phase 8 of the procurement-ingest refactor (docs/refactor-evaluation.md §4): test-root consolidation — keep BOTH roots but fix the LOADING via ONE new repo-root conftest.py (dummy AWS env + moto BUILTIN_HANDLERS before any handler import + single load_lambda_module keeping one sys.modules save/restore); a shared tests/support/ package (superset FakeTable, FakeDynamoResource, load_email, load_golden with parse_float=Decimal); rewrite _wo_parser_support.py off the bare import strategy; move test_po_merge/test_pad_zip into the PO tests dir (test_parse_raw_email/test_ses_auth stay root); delete test_local.py; add the missing scenarios (PO+WO Bedrock transport errors, handler SES-auth reject seam, web_ui both, WO merge semantics, small pins); CI gains --cov with an explicit module list (or __init__.py), a fail-under, the Phase 0 AST bundle test, an enforced ruff C901/PLR config incl scripts/. The ONLY product-code change is the WO handler Bedrock-error metric-wrap (except-and-reraise, no double-count) + invalid_status reason-code fix, and the PO _validate_new_po_values per-rule split. LAST PR — validates the new module boundaries, so it hard-gates on Phases 3 AND 5 merged. Committed locally, never pushed (push gated on /sh-security-review in the main loop for the WO metric change).',
phases: [
{ title: 'Setup', detail: 'verify Phases 3 (lambdas/shared/) AND 5 (decomposed handler siblings) on base, branch feature/phase-8-test-consolidation', model: 'haiku' },
{ title: 'Recon', detail: '4 mappers: the three loaders + FakeTable/support duplication, test inventory + coverage gaps, WO metric/reason-code + PO validation-split surface, CI/lint/ruff/README drift' },
{ title: 'Spec', detail: 'serial fable spec: pin root conftest, support package, loader rewrites, file moves/deletions, new scenarios, WO metric-wrap + reason-code, PO validation split, ruff config, CI changes, README, disjoint ownership' },
{ title: 'Implement', detail: 'opus: all tests + support + root conftest; opus: WO handler + PO template_parser product-code + newly-surfaced lint fixes; sonnet: ruff/pytest/CI config + README — disjoint files', model: 'opus' },
{ title: 'Verify', detail: 'mechanical gates (full suite green under the new loader, goldens UNCHANGED, ruff green incl scripts/, both cdk synth, --cov covers web_ui/site_extractor, WO no-double-count pin passes) + 3 fable lenses (loader-integrity, WO-metric-contract, coverage-honesty)' },
{ title: 'Fix', detail: 'opus fixer, full re-verify, max 3 rounds', model: 'opus' },
{ title: 'Package', detail: 'single commit via -F (no push)', model: 'sonnet' },
],
}
// ---------------------------------------------------------------- constants
const REPO = '/Users/adammoussa/Documents/repositories/seahaven/procurement-ingest'
const BRANCH = 'feature/phase-8-test-consolidation'
let _args = args
if (typeof _args === 'string') {
try { _args = JSON.parse(_args) } catch (e) { _args = null }
}
const BASE = (_args && _args.base) || 'main'
const CONSTRAINTS = `
PINNED BEHAVIORAL CONSTRAINTS (docs/refactor-evaluation.md Phase 8 + §4 — violating any is a build failure):
1. KEEP BOTH TEST ROOTS; fix the LOADING, not the split. ONE loader: a NEW
repo-root conftest.py (tests/conftest.py does NOT load for a standalone
'pytest lambdas/po' run, so it cannot carry session invariants) that:
sets the dummy AWS env; registers moto BUILTIN_HANDLERS BEFORE any handler
import (carry the explanatory comment currently buried at
_po_parser_support.py:33 verbatim — boto3 sessions only pick up moto's
stubber hook if created AFTER registration); and exposes a single
load_lambda_module(pipeline, name) keeping ONE copy of the sys.modules
save/restore dance (tests/conftest.py:70-85). That dance does NOT shrink to
nothing — template_parser is still a duplicated bare name across pipelines
and still needs per-exec sibling binding.
2. Rewrite _wo_parser_support.py OFF the bare 'import handler' /
'from handler import parse_raw_email' strategy (it is the source of the
bare-name sys.modules collision the other two loaders defend against — see
_wo_parser_support.py:17-18,39). Add a shared tests/support/ package with:
a SUPERSET FakeTable (PO's update_item recording + WO's put_item and the
keyed single-row store from _wo_parser_support.py:51-67), FakeDynamoResource,
load_email, load_golden — KEEPING PO's parse_float=Decimal in load_golden
(load-bearing: exact money comparison at PO magnitudes;
_po_parser_support.py load_golden). PO-only helpers keep Decimal; WO's
load_golden has no parse_float today — the superset must not regress PO.
3. FILE MOVES: tests/test_po_merge.py and tests/test_pad_zip.py ->
lambdas/po/email_processor/tests/ (PO-specific, belong beside the code).
tests/test_parse_raw_email.py and tests/test_ses_auth.py STAY at root
(genuinely cross-pipeline, parameterized over BOTH handlers). git mv so
history follows; fix their imports to the new support package.
4. DELETE tests/../test_local.py (repo root: globs a nonexistent samples/,
WO-only, imports a handler at collection so it bypasses the loader gate;
the golden suites cover its role). Deletion is PREFERRED over a --pipeline
rewrite. Its pytest.ini exclusion comment goes with it.
5. ADD the missing scenarios by module (§4 priority order), using the new
single loader + support package for EVERY new test:
(a) PO+WO Bedrock transport errors: ThrottlingException, missing 'content'
key, empty content list, non-JSON model text. Assert PO's pre-call
ai_fallback metric survived + no partial write + exception propagates.
PIN WO's no-datapoint-on-throttle behavior with a DOCUMENTING test —
do NOT "fix" it by reordering (constraint 7 owns the real fix).
(b) Handler-level SES-auth reject seam, per pipeline: NO auth monkeypatch +
empty ALLOWED_DKIM_DOMAINS -> assert ZERO Bedrock calls, ZERO writes,
NO raise (env read at call time; FakeS3 still needed, gate sits after
get_object). Today deleting the gate line passes all tests — this closes
that hole.
(c) web_ui BOTH functions (0% today; auth is the mandatory-review surface):
fail-closed on unset ARN (module reload — ARN read at import),
fail-closed on a Secrets-Manager exception, TTL cache refresh,
Bearer / X-Auth-Token / header case-insensitivity, wrong-token 401,
401 WITHOUT a table scan, non-ASCII token, hostile-field escaping
regression lock. PO web_ui lacks __init__.py — use the loader, not
package imports.
(d) WO merge semantics: a moto-backed mirror of test_po_merge (table name
'WorkOrders', NOT kebab): null-status never clobbers wo_status,
created_at immutable via if_not_exists, status->wo_status mapping, None
fields absent from SET, record_type only-when-present.
(e) Small pins: PO-DC-02 64-char EMF clamp regression, per-pipeline
multi-record failure-isolation (all-or-retry contract), reprocess.py
synthetic-event-shape contract IF Phase 7 has not already added it
(recon confirms — do not duplicate).
6. CI CHANGES (in the same PR): add --cov with an EXPLICIT module list (OR add
__init__.py so --cov=lambdas stops silently skipping web_ui/site_extractor
for the missing package marker) — pick ONE, state why; add a fail-under
ONCE the web_ui/site_extractor suites exist; keep the Phase 0 AST
bundle-consistency test in the standard pytest run; ADD a ruff config
(pyproject.toml or ruff.toml) that ENABLES C901/PLR so the complexity
ceilings are ENFORCED not decorative; include scripts/ in the lint scope;
keep the strict 'ci / ci' required check; NO admin-bypass pushes.
7. WO HANDLER PRODUCT-CODE CHANGES (the ONLY product-code deltas besides the PO
split): (a) invalid_status reason-code fix — GREP the dashboards/metric
filters for 'malformed_site_code' FIRST and report before renaming (one
status failure currently yields two codes by path). (b) WO Bedrock-error
metric fix: WRAP the Bedrock call so ai_fallback + bedrock_error are emitted
in an except-and-RERAISE. This is NOT a naive reorder — a reorder emits the
metric unconditionally pre-call and DOUBLE-COUNTS gate-rejected emails
against wo_stack's "a rejected email emits nothing else" alarm contract.
The handler EVENT/RETURN CONTRACT is unchanged (no signature change). PIN
the no-double-count with a test that computes the emitted series by hand.
8. PO _validate_new_po_values PER-RULE SPLIT (template_parser.py): break the
monolith into per-rule helpers; the V4 anchor-frame dataclass must CARRY
summary_matches / price so V13 can consume them; include extract_new_po
(C901=35) in the split scope. The NEW ruff PLR/C901 config makes the
previously-INERT 'noqa: PLR09xx' suppressions LIVE — every newly-surfaced
violation across the tree must be FIXED or noqa'd-with-written-justification
IN THIS SAME PR. EXCEPTION: files under the shadow-bake freeze
(derived_fields.py) are UNTOUCHABLE — surface their violations via a
per-file ignore in the ruff config (with a justification comment), NEVER an
in-file edit. Docstring/README drift batch (findings 33-38) lands here too.
9. UNTOUCHABLE (git diff ${BASE}...HEAD must be empty for each): every product
module EXCEPT lambdas/wo/email_processor/handler.py (constraint 7) and
lambdas/po/email_processor/template_parser.py (constraint 8). Specifically
derived_fields.py (shadow bake), both ses_auth, both validate_ai_fallback
gates, lambdas/shared/*, all CDK stacks, site_extractor. GOLDEN fixtures
(tests/**/expected/*.json and the .eml corpora) are byte-frozen — a test
that only passes after a golden edit is a FAIL. New tests must NOT edit
existing goldens; WO's one SES-stamped fixture (synthesized header block)
is the only new fixture allowed.
10. This is the LAST phase: the resulting suite / coverage fail-under / ruff
config become the NEW CI floor. Do not weaken any existing gate to make a
new one pass.
`
const PREAMBLE = `
You are one of several agents building refactor Phase 8 in the git repo at
${REPO} on branch ${BRANCH} (already checked out — do NOT switch branches,
do NOT create branches, do NOT commit, NEVER push, do NOT run cdk deploy or
touch AWS resources beyond read-only calls).
Authoritative spec: docs/refactor-evaluation.md, section "Phase 8 —
Test-root consolidation" and the whole of "§4 Test-hardening plan". The doc
wins on any conflict.
Work ONLY in the files you are told you own; other agents are concurrently
editing other files in this same working tree.
${CONSTRAINTS}
Your final message is consumed by an orchestrator script, not a human —
return only the structured data requested.
`
// ------------------------------------------------------------------ schemas
const RECON = {
type: 'object',
required: ['summary', 'facts'],
properties: {
summary: { type: 'string' },
facts: { type: 'array', items: { type: 'string' } },
blockers: { type: 'array', items: { type: 'string' } },
},
}
const SPEC = {
type: 'object',
required: ['rootConftest', 'supportPackage', 'loaderRewrites', 'fileMoves', 'newScenarios', 'productCodeEdits', 'ruffConfig', 'ciChanges', 'readmeDocs', 'ownership', 'notes'],
properties: {
rootConftest: { type: 'string', description: 'the complete new repo-root conftest.py design: dummy AWS env block, the moto BUILTIN_HANDLERS-before-handler-import registration with the carried :33 comment, and the single load_lambda_module(pipeline, name) signature keeping ONE sys.modules save/restore (state exactly why it does not shrink to nothing — template_parser bare name). How tests/conftest.py, _po_parser_support.py and _wo_parser_support.py reconcile against it (deleted / shimmed / re-pointed).' },
supportPackage: { type: 'string', description: 'tests/support/ package contents: the SUPERSET FakeTable field-by-field (PO update_item record + WO puts/keyed store), FakeDynamoResource, load_email, load_golden — explicitly pinning parse_float=Decimal is kept and that WO callers gain it without regressing (are WO goldens integer-only?). __init__.py exports.' },
loaderRewrites: { type: 'string', description: 'exact per-file edits to retire the bare-import strategy in _wo_parser_support.py and rewire both support files + tests/conftest.py onto load_lambda_module; the moto-before-handler ordering must survive; which fixtures (email_handler, ses_auth, po_handler, wo_handler) move where.' },
fileMoves: { type: 'string', description: 'the git mv list (test_po_merge.py, test_pad_zip.py -> PO tests dir), the stay-at-root set (test_parse_raw_email.py, test_ses_auth.py), the test_local.py deletion + its pytest.ini comment removal, and the import rewrites each moved file needs.' },
newScenarios: { type: 'string', description: 'per-module new test files (paths) and the scenarios each covers per constraint 5 (a-e): Bedrock transport errors both pipelines, SES-auth reject seam both pipelines, web_ui both functions, WO merge semantics, the small pins — and whether the reprocess synthetic-event pin already exists from Phase 7.' },
productCodeEdits: { type: 'string', description: 'the WO handler edit (invalid_status reason-code fix + Bedrock-error except-and-reraise metric-wrap with the exact emitted series proving NO double-count vs wo_stack alarm contract; the malformed_site_code dashboard-grep result) and the PO _validate_new_po_values / extract_new_po per-rule split (V4 dataclass carries summary_matches/price for V13). file:line for each.' },
ruffConfig: { type: 'string', description: 'the exact new ruff config (pyproject.toml or ruff.toml): which rule families (C901/PLR), the complexity ceilings (extract_new_po C901=35), lint scope incl scripts/, and the per-file-ignore for derived_fields.py with justification. The full list of newly-surfaced violations and their fix-or-noqa disposition.' },
ciChanges: { type: 'string', description: 'the exact .github/workflows/ci.yaml edits: --cov approach (explicit module list vs __init__.py, with the reason), where the fail-under lands and its value, keeping the AST bundle test + strict ci/ci check, scripts/ lint scope, no admin bypass.' },
readmeDocs: { type: 'string', description: 'README + docstring drift batch (findings 33-38): exact sections/lines to fix, matched to existing style.' },
ownership: { type: 'string', description: 'the DISJOINT file-ownership map for the 3 parallel implement agents (tests+support+root-conftest / product-code+newly-surfaced-lint / config+CI+README) — no path owned by two agents; note the ruff-config -> product-lint dependency and how the poll resolves it.' },
notes: { type: 'string' },
},
}
const IMPL = {
type: 'object',
required: ['filesChanged', 'summary', 'checksRun'],
properties: {
filesChanged: { type: 'array', items: { type: 'string' } },
summary: { type: 'string' },
checksRun: { type: 'string' },
blockers: { type: 'array', items: { type: 'string' } },
},
}
const CHECKS = {
type: 'object',
required: ['passed', 'details'],
properties: {
passed: { type: 'boolean' },
details: { type: 'string' },
scopeViolations: { type: 'array', items: { type: 'string' } },
},
}
const FINDINGS = {
type: 'object',
required: ['findings'],
properties: {
findings: {
type: 'array',
items: {
type: 'object',
required: ['title', 'severity', 'confirmed', 'evidence', 'fix'],
properties: {
title: { type: 'string' },
severity: { enum: ['critical', 'high', 'medium', 'low'] },
confirmed: { type: 'boolean' },
evidence: { type: 'string' },
fix: { type: 'string' },
},
},
},
},
}
// ------------------------------------------------------------------- setup
phase('Setup')
const setup = await agent(`
In ${REPO}:
1. SEQUENCING HARD-GATE — this is the LAST PR and it validates the NEW module
boundaries, so Phases 3 AND 5 must both be on ${BASE} (shared/ from Phase 3
is what the loader now resolves against; the decomposed handler siblings
from Phase 5 are the boundaries these suites exercise). git fetch origin,
then pick the base ref: origin/${BASE} if that remote ref exists, otherwise
the local branch ${BASE} (a stacked local-only base is expected and fine).
Verify on the base ref:
(a) Phase 3: lambdas/shared/ exists with its modules
(git ls-tree <baseref> -- lambdas/shared/ — must list ses_auth.py etc.);
(b) Phase 5: the decomposed PO handler siblings exist
(git ls-tree <baseref> -- lambdas/po/email_processor/ must include
extraction.py, enrichment.py, telemetry.py, persistence.py, prompts.py).
If EITHER is missing, STOP with a blocker naming the unmet phase and do
nothing else.
2. Verify clean working tree (untracked .coverage / .claude/ / the local
44 MB lambdas/po/email_processor/package/ dir are fine; any OTHER dirt =
blocker, never stash or discard).
3. git checkout ${BASE}; then git pull --ff-only ONLY if the branch has an
upstream (a local-only base skips the pull — not a blocker); then
git checkout -b ${BRANCH}
4. gh pr list --state open --json number,title,headRefName (overlap check).
Return facts: HEAD sha, per-phase gate evidence (the two git ls-tree outputs),
open PRs, blockers.
`, { label: 'setup:branch', model: 'haiku', schema: RECON })
if (!setup || (setup.blockers && setup.blockers.length)) {
return { aborted: 'setup blockers', blockers: setup ? setup.blockers : ['setup agent died'], facts: setup ? setup.facts : [] }
}
log(`Branch ${BRANCH} ready off ${BASE}. ${setup.summary}`)
// ------------------------------------------------------------------- recon
phase('Recon')
const recon = await parallel([
() => agent(`${PREAMBLE}
Read-only recon of the THREE loading idioms + the fake-Dynamo helpers this
phase collapses:
1. tests/conftest.py — quote the importlib loader, the _SIBLING_MODULES tuple
(ses_auth/template_parser/derived_fields), and the sys.modules save/restore
dance (lines ~70-85). Note that after Phase 5 the PO sibling set changed
(extraction/enrichment/telemetry/persistence/prompts) — quote the CURRENT
sibling list on this branch's base, it may already differ from the audit.
2. _po_parser_support.py — the independent importlib reimplementation, the
load-bearing moto-before-handler comment at line 33 (quote it verbatim, it
must be carried into the new root conftest), load_golden's parse_float=Decimal
(quote it), and its FakeTable (update_item only).
3. _wo_parser_support.py — the bare sys.path 'import handler' /
'from handler import parse_raw_email' strategy (lines 17-18, 39), and its
FakeTable (put_item + keyed store, lines 51-67) — the superset target.
4. pytest.ini — the testpaths (three roots) and the test_local.py exclusion
comment. Confirm test_local.py exists at repo root and what it imports at
collection.
Diff the two FakeTable classes field-by-field to define the superset. Confirm
whether WO's load_golden uses parse_float (it does not today) and whether any
WO golden is non-integer (would break under Decimal — pin it).
20-30 precise facts.`,
{ label: 'recon:loaders-support', model: 'sonnet', phase: 'Recon', schema: RECON }),
() => agent(`${PREAMBLE}
Read-only recon of the test INVENTORY + coverage gaps this phase must fill:
- Enumerate every test_*.py across all three roots (tests/,
lambdas/po/email_processor/tests/, lambdas/wo/email_processor/tests/) with a
one-line purpose each. Flag test_po_merge.py + test_pad_zip.py at root (move
targets) and test_parse_raw_email.py + test_ses_auth.py at root (stay).
- For the MISSING scenarios (§4), confirm the current gap and the seam each new
test hooks: (a) PO+WO Bedrock transport-error handling — quote where each
handler invokes Bedrock and where ai_fallback is emitted relative to the call
(PO pre-call; WO post-gate); (b) the SES-auth gate call site in each handler
(authenticate_inbound_email) and how existing dispatch tests monkeypatch it;
(c) web_ui — BOTH functions' auth entry points, whether they have any tests
or __init__.py today (they do not), how the ARN/Secrets-Manager fail-closed
path reads config (import-time vs call-time), the table-scan path a 401 must
avoid; (d) WO save_work_order merge semantics (table 'WorkOrders'); (e) the
small pins — the 64-char EMF clamp site, multi-record event loop, and whether
scripts/reprocess.py already has a synthetic-event contract test from Phase 7
(git log / grep — do NOT duplicate it).
- Confirm the golden corpora locations (expected/*.json, .eml dirs) that are
byte-frozen.
25-35 facts.`,
{ label: 'recon:inventory-gaps', model: 'sonnet', phase: 'Recon', schema: RECON }),
() => agent(`${PREAMBLE}
Read-only recon of the TWO product-code change surfaces (constraints 7-8):
1. WO handler (lambdas/wo/email_processor/handler.py): quote with file:line the
Bedrock invoke, the emit_parse_metric / ai_fallback emission, and the gate
that returns 'invalid_status' vs 'malformed_site_code'. Establish the CURRENT
emitted-metric series for a gate-rejected email and for a Bedrock-error email
— enough to prove the except-and-reraise wrap does NOT double-count. THEN
grep the whole repo (cdk/*.py, dashboards, metric filters, docs) for
'malformed_site_code' and 'invalid_status' and list every consumer — the
reason-code rename must not silently break a dashboard/alarm. Quote the
wo_stack "a rejected email emits nothing else" alarm/MathExpression it must
not violate.
2. PO template_parser.py: quote _validate_new_po_values (the ~294-line monolith,
its 'noqa: PLR09xx' suppressions) and extract_new_po (C901=35). Identify the
V4 anchor-frame dataclass and what fields it carries today vs the
summary_matches/price V13 needs. Confirm both currently pass lint ONLY
because no ruff config selects PLR/C901.
Return the exact emitted-series tables, the malformed_site_code consumer list,
and the validation-split surface. 20-30 facts.`,
{ label: 'recon:product-code', model: 'sonnet', phase: 'Recon', schema: RECON }),
() => agent(`${PREAMBLE}
Read-only recon of the CI / lint / coverage / README surface:
- .github/workflows/ci.yaml — quote the pytest invocation (does it pass --cov?),
the ruff steps, the required-check name (ci / ci), any admin-bypass config.
Confirm the Phase 0 AST bundle-consistency test (tests/test_bundle_consistency.py)
is or is not already in the standard run.
- Coverage honesty: confirm there is NO ruff config file today (so PLR/C901 are
unselected) and confirm web_ui/site_extractor lack __init__.py (so a naive
--cov=lambdas SILENTLY skips them). Prove the skip by reading the tree, not by
running.
- scripts/ — list the scripts and confirm they are outside the current lint
scope.
- README.md + docstrings — locate the drift items (findings 33-38 / X-series):
the false "Data is never corrupted" style claims, stale web_ui invocation
contract, missing derived-field/coverage docs, and any test-layout description
that this phase changes. Quote line numbers.
15-25 facts.`,
{ label: 'recon:ci-lint-readme', model: 'haiku', phase: 'Recon', schema: RECON }),
])
const reconOk = recon.filter(Boolean)
const reconBlockers = reconOk.flatMap(r => r.blockers || [])
log(`Recon complete: ${reconOk.length}/4 mappers, ${reconBlockers.length} blockers (all resolved in the main loop — see RESOLUTIONS)`)
// Main-loop resolutions for the first run's recon blockers — verified facts,
// checked directly against WO goldens, live CloudWatch, and the reusable CI
// workflow source. The remaining recon "blockers" were this phase's own work
// items misreported as blockers; none abort the run.
const RESOLUTIONS = `
## MAIN-LOOP RESOLUTIONS (verified AFTER recon flagged blockers — authoritative facts, supersede any recon hedge)
1. WO goldens under parse_float=Decimal: SAFE. All 55 golden files under
lambdas/wo/email_processor/tests/fixtures/expected/ were scanned
programmatically — ZERO float-typed JSON number values exist (decimal-looking
strings appear only inside comment_text STRING values, which parse_float never
touches). The superset load_golden keeping parse_float=Decimal cannot regress WO.
2. malformed_site_code consumers: NONE outside the repo. Live scan of the
deployment account 328440206208 (hosts both po-ingest and WorkorderIngestStack):
0 CloudWatch dashboards, 0 saved Logs Insights query definitions, 18 metric
filters (no match), 146 alarms (no match). seahaven-prod (011934824531) also
clean. Combined with recon's repo-scope grep, the invalid_status reason-code
rename has NO external consumer to break. Still note the rename in the commit
body per the deploy-then-merge outstanding gate.
3. Reusable CI workflow (Sea-Haven-Industries/.github ci-python-sam.yaml
@fd60e4c904, the pinned ref in ci.yaml): inputs.source-dirs threads ONLY into
'ruff check \${{ inputs.source-dirs }}' and 'ruff format --check' — so adding
'scripts' to source-dirs in THIS repo's ci.yaml genuinely puts scripts/ in lint
scope. Tests run as bare 'pytest' with NO args — so --cov and --cov-fail-under
MUST land via pytest.ini addopts (bare pytest picks addopts up); ci.yaml cannot
pass pytest flags. Dependency install is 'pip install pytest' plus
'pip install -r' for EVERY requirements.txt found outside ./.aws-sam — so
pytest-cov (and any other new test dep) must be added to a requirements.txt the
find loop reaches (e.g. a tests/requirements.txt; create it if absent). CI's
ruff is unpinned 'pip install ruff' — the new config must be valid on current ruff.
4. The remaining recon 'blockers' (README drift X1-X8, coverage honesty, absent
ruff config, test_local.py still present) are THIS PHASE'S OWN WORK ITEMS, not
blockers — implement them per the constraints.
`
const pack = reconOk.map(r => `## ${r.summary}\n${r.facts.join('\n')}`).join('\n\n') + '\n\n' + RESOLUTIONS
// -------------------------------------------------------------------- spec
phase('Spec')
const spec = await agent(`${PREAMBLE}
You are the SPEC agent — the single authority that pins every contested
decision BEFORE parallel implementation (parallel leaves cannot see each
other's choices). Using the recon pack below plus your own reads of the actual
files, produce the binding implementation spec:
- rootConftest: the complete new repo-root conftest.py — the dummy AWS env, the
moto BUILTIN_HANDLERS-before-handler-import registration WITH the carried :33
comment, and the single load_lambda_module(pipeline, name) keeping ONE
sys.modules save/restore. State explicitly why the dance does not shrink to
nothing (template_parser bare name) and how the three existing loaders
reconcile (which are deleted, which become thin shims, which re-point).
- supportPackage: tests/support/ — the SUPERSET FakeTable, FakeDynamoResource,
load_email, load_golden. PIN parse_float=Decimal kept in load_golden and
prove WO goldens survive it (recon's non-integer check).
- loaderRewrites: retire the bare-import strategy in _wo_parser_support.py;
rewire both support files + tests/conftest.py onto the new loader; moto
ordering survives; where each fixture lands.
- fileMoves: the git mv set, the stay-at-root set, the test_local.py deletion +
pytest.ini comment removal, per-file import rewrites.
- newScenarios: every new test file path + the scenarios per constraint 5 (a-e).
De-duplicate the reprocess synthetic-event pin against Phase 7 per recon.
- productCodeEdits: the WO handler invalid_status reason-code fix + the
Bedrock-error except-and-reraise metric-wrap with the exact emitted series
(compute by hand: gate-reject emits nothing else; Bedrock-error emits
ai_fallback + bedrock_error ONCE — NO double-count), plus the
malformed_site_code consumer disposition. The PO _validate_new_po_values /
extract_new_po per-rule split (V4 dataclass carries summary_matches/price).
file:line for each; nothing else in these two files changes.
- ruffConfig: the exact config, rule families, ceilings (extract_new_po
C901=35), scripts/ scope, the derived_fields.py per-file-ignore with
justification, and the FULL disposition of every newly-surfaced violation.
- ciChanges: the exact ci.yaml edits — --cov approach (explicit list vs
__init__.py, with reason), fail-under value + placement, AST test + strict
check kept, scripts/ in lint scope, no admin bypass.
- readmeDocs: the drift batch fixes matched to existing style.
- ownership: the DISJOINT 3-agent ownership map with NO shared path, and how
the ruff-config -> product-lint poll dependency resolves.
Recon pack:\n${pack}`,
{ label: 'spec:pin-consolidation', phase: 'Spec', schema: SPEC })
if (!spec) return { aborted: 'spec agent died — rerun workflow', reconBlockers }
const specBlock = `BINDING SPEC (from the spec agent — implement EXACTLY this):\n${JSON.stringify(spec, null, 2)}`
log('Spec pinned: root conftest, support package, loader rewrites, file moves, new scenarios, product-code edits, ruff config, CI changes, README, ownership')
// --------------------------------------------------------------- implement
phase('Implement')
const impl = await parallel([
() => agent(`${PREAMBLE}
YOU OWN: ALL test + support files — the repo-root conftest.py (NEW),
tests/ (including tests/support/ the new package, tests/conftest.py,
tests/test_bundle_consistency.py, the stay-at-root test_parse_raw_email.py /
test_ses_auth.py, and the test_local.py DELETION),
lambdas/po/email_processor/tests/ and lambdas/wo/email_processor/tests/.
You do NOT touch product code, cdk/, ci.yaml, ruff config, pytest.ini, or
README (other agents own those; pytest.ini's test_local.py comment is the
config agent's edit — coordinate only by not touching it).
Task per spec.rootConftest + spec.supportPackage + spec.loaderRewrites +
spec.fileMoves + spec.newScenarios:
1. Write the new repo-root conftest.py with the single load_lambda_module,
carrying the :33 moto comment VERBATIM and keeping ONE sys.modules
save/restore (it does NOT shrink to nothing).
2. Create tests/support/ (superset FakeTable, FakeDynamoResource, load_email,
load_golden with parse_float=Decimal). Retire _wo_parser_support.py's
bare-import strategy; rewire _po_parser_support.py + tests/conftest.py.
3. git mv test_po_merge.py + test_pad_zip.py into the PO tests dir; DELETE
test_local.py; fix all import lines.
4. Add every new scenario suite per constraint 5 (a-e) using ONLY the new
loader + support package. Do NOT edit any existing golden (constraint 9) —
a test that only passes after a golden edit is a bug in the test. WO's one
SES-stamped fixture is the only new fixture allowed.
Run before returning: pytest -q at repo root (all roots green against the OTHER
agents' edits — they land in parallel; the WO metric-wrap test and PO split may
be mid-flight, so poll by re-running up to ~10 min before reporting a blocker),
and confirm goldens are byte-unchanged (git diff --stat on expected/ dirs is
empty).
${specBlock}`,
{ label: 'impl:tests-support', model: 'opus', phase: 'Implement', schema: IMPL }),
() => agent(`${PREAMBLE}
YOU OWN: lambdas/wo/email_processor/handler.py and
lambdas/po/email_processor/template_parser.py ONLY, PLUS you are the designated
fixer for any NEWLY-SURFACED ruff PLR/C901 violation across product code and
scripts/ (fix-or-noqa-with-justification, NEVER a behavior change). You do NOT
touch tests, cdk/, ci.yaml, README, or derived_fields.py (its violations are
handled via the config agent's per-file-ignore, NOT an edit — constraint 8).
Task per spec.productCodeEdits:
1. WO handler: apply the invalid_status reason-code fix (only after confirming
spec captured the malformed_site_code consumer disposition) and the
Bedrock-error except-and-RERAISE metric-wrap. The wrap emits
ai_fallback + bedrock_error ONCE on a Bedrock error and NOTHING extra on a
gate reject — do the hand series-count in your summary. Handler event/return
contract UNCHANGED.
2. PO template_parser: the _validate_new_po_values / extract_new_po per-rule
split; the V4 anchor-frame dataclass carries summary_matches/price for V13.
3. The ruff config lands in a sibling agent's file. POLL for it (re-check every
~60s up to ~10 min); once present, run 'ruff check <config> .' and resolve
EVERY newly-surfaced violation in the files you own (and scripts/) — fix or
noqa with a written justification. If a violation lands in a file you do NOT
own and is not derived_fields.py, report it as a blocker for the fix loop.
Run before returning: ruff check + ruff format --check on the files you touched,
and a targeted pytest of the WO Bedrock-fallback + PO validation-gate suites
(poll for the test agent's new files up to ~10 min).
${specBlock}`,
{ label: 'impl:product-code', model: 'opus', phase: 'Implement', schema: IMPL }),
() => agent(`${PREAMBLE}
YOU OWN: the ruff config file (pyproject.toml or ruff.toml — spec picks which),
pytest.ini, .github/workflows/ci.yaml, and README.md ONLY. You do NOT touch
tests or product code.
Task per spec.ruffConfig + spec.ciChanges + spec.readmeDocs:
1. Create the ruff config ENABLING C901/PLR with the pinned ceilings
(extract_new_po C901=35), scripts/ in scope, and the derived_fields.py
per-file-ignore with a justification comment (constraint 8). Write this
FIRST so the product-code agent can poll for it.
2. pytest.ini: remove the test_local.py exclusion comment (the file is being
deleted by the test agent); keep the three testpaths roots.
3. ci.yaml: add --cov with the spec's approach (explicit module list OR
__init__.py note), a fail-under, keep the AST bundle-consistency test in the
standard run, add scripts/ to the lint scope, keep the strict 'ci / ci'
required check, no admin bypass.
4. README + docstring drift batch (findings 33-38): the false "Data is never
corrupted" claims, stale web_ui invocation contract, the new test layout
(single loader, tests/support/, moved files), coverage/ruff-floor note.
Run before returning: ruff check <your new config> . to confirm the config is
valid and to enumerate what it surfaces (report the list for the product-code
agent), and a yaml lint / dry parse of ci.yaml. Do NOT run the full suite (the
test agent owns that gate).
${specBlock}`,
{ label: 'impl:config-ci-readme', model: 'sonnet', phase: 'Implement', schema: IMPL }),
])
const implOk = impl.filter(Boolean)
const implBlockers = implOk.flatMap(r => r.blockers || [])
log(`Implement complete: ${implOk.length}/3 agents, blockers: ${implBlockers.length}`)
// ---------------------------------------------------- verify + fix loop
const EXPECTED_SCOPE = [
'conftest.py',
'tests/',
'lambdas/po/email_processor/tests/',
'lambdas/wo/email_processor/tests/',
'lambdas/wo/email_processor/handler.py',
'lambdas/po/email_processor/template_parser.py',
'pyproject.toml',
'ruff.toml',
'pytest.ini',
'.github/workflows/ci.yaml',
'README.md',
]
const mechanicalPrompt = `${PREAMBLE}
Independent re-verification — trust nothing self-reported. Run ALL gates,
quoting failures verbatim:
1. pytest -q at repo root — all three roots collected and GREEN under the NEW
single loader. Then prove the loader actually runs: pytest lambdas/po alone
and pytest lambdas/wo alone each collect and pass (the new root conftest
must load for a standalone run — the whole reason it moved to repo root).
2. GOLDENS UNCHANGED: git diff ${BASE}...HEAD -- '**/expected/*.json' and the
.eml corpora must be empty (constraint 9). Any golden edit = FAIL.
3. ruff check . && ruff format --check . under the NEW config, WITH scripts/ in
scope. Zero violations (every surfaced one fixed or noqa'd-with-justification;
derived_fields.py handled via per-file-ignore, its file byte-unchanged:
git diff ${BASE}...HEAD -- lambdas/po/email_processor/derived_fields.py empty).
4. cd cdk && npx cdk synth po-ingest -q && npx cdk synth workorder-ingest -q
(artifact-id selectors) — both synth clean (no CDK change this phase, but a
broken import would surface here).
5. COVERAGE HONESTY: run the CI --cov invocation and CONFIRM the coverage report
lists lambdas/po/web_ui, lambdas/wo/web_ui and lambdas/po/site_extractor with
NON-zero, NON-omitted lines (prove they are no longer silently skipped for a
missing __init__.py). Confirm the fail-under is present and would actually
fail if tripped (has teeth).
6. WO NO-DOUBLE-COUNT PIN: run the specific test that computes the WO
Bedrock-error emitted series by hand and assert it passes; then git grep for
the reason-code rename and confirm no dashboard/metric-filter consumer of
'malformed_site_code' was left dangling.
7. UNTOUCHABLES (each must output NOTHING): git diff ${BASE}...HEAD for
derived_fields.py, both ses_auth, both validate_ai_fallback gates,
lambdas/shared/, all cdk/*.py, site_extractor. The ONLY product diffs allowed
are lambdas/wo/email_processor/handler.py and
lambdas/po/email_processor/template_parser.py.
8. FILE MOVES + DELETION: test_po_merge.py + test_pad_zip.py are gone from
tests/ and present under lambdas/po/email_processor/tests/; test_local.py is
deleted; test_parse_raw_email.py + test_ses_auth.py still at root and green.
9. AST bundle-consistency test present in the standard run and green.
10. git status --porcelain scope check: every modified/added/deleted path under
${EXPECTED_SCOPE.join(', ')} plus tests/support/ (untracked
.coverage/.claude/package/ tolerated).
passed=true only if all green. YOU MAY NOT edit files.`
const lenses = [
{ key: 'loader-integrity', prompt: `${PREAMBLE}
ADVERSARIAL REVIEW — loader-integrity lens. ONE loader replaced three; attack
it. (1) Prove there is exactly ONE load path now and that the moto
BUILTIN_HANDLERS registration still happens BEFORE any handler import in EVERY
entry order (session start AND a standalone 'pytest lambdas/wo' run) — if a
handler's module-level boto3 client can be created before moto registers, the
moto-backed suites silently hit real AWS. (2) sys.modules isolation for
template_parser: run TWO cross-pipeline tests back to back (a PO golden then a
WO golden, and the reverse) and prove the save/restore still binds each handler
to its OWN template_parser — the dance did NOT shrink to nothing. (3) The file
moves (test_po_merge/test_pad_zip into the PO dir) must not break collection or
silently drop a test — count tests before/after. (4) Did retiring
_wo_parser_support.py's bare import leave any stale 'import handler' /
sys.path.insert that could re-introduce the collision? (5) load_golden's
parse_float=Decimal survived for PO (exact money) and did not corrupt WO.
confirmed=true only with file:line or a reproduced-failure.` },
{ key: 'wo-metric-contract', prompt: `${PREAMBLE}
ADVERSARIAL REVIEW — WO-metric-contract lens. The WO Bedrock-error metric-wrap
is the one behavior change; prove it does NOT break wo_stack's alarm contract.
Read the wrapped handler diff (git diff ${BASE}) and wo_stack's fallback-rate
MathExpression + the "a rejected email emits nothing else" alarm. By HAND,
enumerate the emitted metric series for: (a) a normal template parse, (b) a
gate-rejected email, (c) a Bedrock error (Throttling / missing-content /
empty-content / non-JSON). Prove NO series is double-counted — specifically that
the except-and-reraise emits ai_fallback + bedrock_error EXACTLY ONCE on a
Bedrock error and that a gate-rejected email still emits ONLY its rejected
datapoint and nothing else (a naive reorder would double-count — confirm this
was NOT done). Then the reason-code fix: confirm invalid_status no longer
collides with malformed_site_code and that no dashboard/metric-filter consumer
of the old code is left dangling. Confirm the handler event/return contract is
unchanged (no signature drift). confirmed=true only with the hand-computed
series tables as evidence.` },
{ key: 'coverage-honesty', prompt: `${PREAMBLE}
ADVERSARIAL REVIEW — coverage-honesty lens. The headline coverage historically
overstated reality because --cov=lambdas silently skipped web_ui/site_extractor
(missing __init__.py). Prove that is FIXED: run the new CI --cov invocation and
show lambdas/po/web_ui, lambdas/wo/web_ui, lambdas/po/site_extractor each appear
in the report with real measured lines (not omitted, not 0-of-0). Prove the
fail-under has TEETH — temporarily lower a threshold or add a trivially-uncovered
line on a scratch copy and show the job would fail (then discard). Prove the new
web_ui suites exercise the REAL auth module (fail-closed on unset ARN + on a
Secrets-Manager exception, 401 WITHOUT a table scan, non-ASCII token) and not a
mock that would pass with the gate deleted. Finally, confirm the ruff C901/PLR
config genuinely bites (extract_new_po at C901=35 is enforced, the previously
inert noqas are now live and justified, scripts/ is in scope, and
derived_fields.py is excluded via config not an edit). confirmed=true only with
command output as evidence.` },
]
let round = 0
let checks = null
let confirmed = []
while (round < 3) {
phase('Verify')
const results = await parallel([
() => agent(mechanicalPrompt, { label: `verify:mechanical-r${round}`, model: 'sonnet', phase: 'Verify', schema: CHECKS }),
...lenses.map(l => () =>
agent(l.prompt, { label: `verify:${l.key}-r${round}`, phase: 'Verify', schema: FINDINGS })),
])
checks = results[0]
confirmed = results.slice(1).filter(Boolean)
.flatMap(r => r.findings || [])
.filter(f => f.confirmed && f.severity !== 'low')
const green = checks && checks.passed
log(`Verify round ${round}: mechanical ${checks && checks.passed ? 'GREEN' : 'RED'}, confirmed findings: ${confirmed.length}`)
if (green && confirmed.length === 0) break
round += 1
if (round >= 3) break
phase('Fix')
await agent(`${PREAMBLE}
You are the fix agent — you may edit files under: ${EXPECTED_SCOPE.join(', ')}
and tests/support/. Fix EVERY item below minimally; the binding spec and the
10 pinned constraints still hold (a finding that conflicts with a constraint is
reported, not "fixed" — the constraint wins, esp. constraint 9's untouchable
goldens + derived_fields freeze, and constraint 7's no-double-count metric
wrap: do NOT convert it to a naive reorder to silence a finding). Re-run the
specific failing gate/test per fix.
MECHANICAL:\n${checks ? checks.details : '(agent died — rerun all gates)'}
CONFIRMED FINDINGS:\n${JSON.stringify(confirmed, null, 2)}
${specBlock}`,
{ label: `fix:round-${round}`, model: 'opus', phase: 'Fix', schema: IMPL })
}
const verifyClean = checks && checks.passed && confirmed.length === 0
if (!verifyClean) {
return {
status: 'NEEDS ATTENTION — verify not clean after 3 rounds; branch left uncommitted',
branch: BRANCH,
mechanical: checks,
unresolvedFindings: confirmed,
implBlockers,
reconBlockers,
spec,
}
}
// ----------------------------------------------------------------- package
phase('Package')
const commit = await agent(`${PREAMBLE.replace('do NOT commit, ', '')}
YOU are the commit agent:
1. Read ~/Documents/repositories/seahaven/engineering-handbook/commit-messages.md
and follow it exactly.
2. git add only paths under: ${EXPECTED_SCOPE.join(', ')}, tests/support/, and
.claude/workflows/phase-8-test-consolidation.js. NOT .coverage, NOT package/.
Verify the staged set with git status — the test_local.py DELETION and the
git mv of test_po_merge.py / test_pad_zip.py MUST be staged too.
3. ONE commit; write the message to /tmp/phase8-commit-msg.txt and use
git commit -F /tmp/phase8-commit-msg.txt (backticks in -m get eaten by zsh).
Suggested subject:
"test: consolidate test roots — one repo-root loader, shared support package, missing-scenario suites, enforced ruff/coverage floor (refactor phase 8)"
Body: the single repo-root conftest loader + moto-before-handler carry, the
tests/support/ superset (parse_float=Decimal kept), the file moves +
test_local.py deletion, the new scenario suites (Bedrock transport, SES-auth
seam, web_ui both, WO merge, small pins), the WO Bedrock-error metric-wrap
(except-and-reraise, no double-count) + invalid_status reason-code fix, the
PO _validate_new_po_values split, the enforced ruff C901/PLR config, and the
CI --cov + fail-under floor. NO AI attribution / Co-Authored-By lines.
4. Do NOT push. Return commit sha + shortstat in summary.`,
{ label: 'package:commit', model: 'sonnet', phase: 'Package', schema: IMPL })
return {
status: 'BUILT — committed locally, NOT pushed',
branch: BRANCH,
base: BASE,
commit: commit ? commit.summary : 'commit agent died — commit manually',
spec: { rootConftest: spec.rootConftest, productCodeEdits: spec.productCodeEdits, ruffConfig: spec.ruffConfig, ciChanges: spec.ciChanges, ownership: spec.ownership },
implementation: implOk.map(r => r.summary),
filesChanged: implOk.flatMap(r => r.filesChanged),
verifyRounds: round + 1,
blockers: implBlockers.concat(reconBlockers),
outstandingGates: [
'/sh-security-review RECOMMENDED before push — the WO Bedrock-error metric-wrap is a behavior change on the untrusted-input processing path, and web_ui auth is the mandatory-review surface (here only TESTS are added for it). Flag it for the WO metric change specifically; re-run after any post-review fix.',
'cross-family cross_review.py NOT required — no IAM/policy change and no handler-signature change (the WO metric-wrap keeps the event/return contract). Opt-in only if judgment says so.',
'deploy-then-merge for the WO handler metric/reason-code change: deploy from branch, smoke green, one live WO email, verify the ParseOutcome series still matches the alarm contract (NO double-count), and grep the dashboards for malformed_site_code BEFORE the reason-code rename ships.',
'LAST PHASE: the new suite / coverage fail-under / enforced ruff C901/PLR config become the new CI floor — do not weaken any existing gate to land a later change.',
],
}