Commit graph

3 commits

Author SHA1 Message Date
Adam Moussa
30112cc680
feat: extract lambdas/shared/ — single-source ses_auth, web_ui auth, email parsing, EMF emitter (refactor phase 3) (#111)
Some checks are pending
Deploy / deploy (push) Waiting to run
Four modules move into the handbook-mandated lambdas/shared/ location,
collapsing duplicated logic that had to be kept in sync by hand across
the PO and WO pipelines:

- ses_auth.py: the PO and WO copies were verified sha256-identical
  against the feature/phase-7-ops-recovery baseline before the move
  (no drift since the last audit). shared/ses_auth.py is the exact
  bytes of that one copy; both originals are git rm'd (the PO copy
  via rename, the WO copy as a straight delete). Bundling lands the
  module flat in /asset-output for both email processors, so the
  handlers keep `from ses_auth import authenticate_inbound_email`
  unchanged — zero handler diff for this move, which is what keeps
  fail-closed auth byte-identical through the change.

- web_ui_auth.py: extracts the byte-identical _get_auth_token /
  _header / is_authenticated block plus the four token-cache globals
  out of both web_ui handlers. The per-stack INFRA-74 comments stay
  in each handler as-is (deliberately drifted wording, stack-specific)
  rather than being unified into the shared module. Fail-closed
  semantics (unset ARN or Secrets Manager exception -> deny) are
  unchanged.

- email_parsing.py: parse_raw_email ships as the superset version that
  returns cc unconditionally. WO's output is bit-identical to before;
  PO simply ignores the cc field rather than being "cleaned up" to
  consume it. No second variant is kept.

- emf.py: a generic emitter parameterized by namespace, dimension
  sets, and properties. Every call site's emitted EMF envelope is
  unchanged, including the load-bearing
  [["ParseMethod"],["ParseMethod","TemplateId"]] dimension-set shape
  the alarms and metric filters depend on. Emission ordering is
  untouched: PO still emits ai_fallback before the Bedrock call, WO
  still emits its mutually-exclusive ai_fallback/ai_fallback_rejected
  after its gate. The deliberate-double-count comments survive.
  _emit_derived_agreement_metric was found living inside
  derived_fields.py, so per the DERIVED-FIELDS exception it is left
  as a third, unconverted copy (derived_fields.py and the shadow
  DerivedFieldAgreement telemetry stay untouchable while that bake
  runs) — a comment there points at shared/emf.py for the eventual
  follow-up.

Bundling: both email-processor cdk bundling commands gain a trailing
`cp shared/*.py /asset-output/` (they were already cp-only post-Phase
7, so no pip step or manylinux pin is reintroduced). Both web_ui
functions gain the same widened-root staging so web_ui_auth.py ships
beside their handler; site_extractor's from_asset is untouched.

Tests: PO_EXPECTED_TOP_LEVEL_MODULES gains the shared modules that now
ship, the AST sibling-import check resolves imports whose source now
lives under shared/, and the new shared cp line has its own
revert/mutation detection. _SIBLING_MODULES resolution and
_po_parser_support.py now load ses_auth/email_parsing/emf from
shared/; the two-copy ses_auth byte-identity fixture-hygiene test is
retired as obsolete now that there is one copy, and the ses_auth
fixture parameterization over two identical copies is dropped. The
sys.modules save/restore dance for template_parser (still duplicated
per-pipeline) is left in place.
2026-07-20 13:38:23 -04:00
seahaven-openswe[bot]
11bf0f5d12
fix(ses_auth): harden comment stripping and alarm evaluation window (#103)
Some checks are pending
Deploy / deploy (push) Waiting to run
SES-AR-01: treat ")" at depth 0 as an unmatched close, rejecting the
value as not well-formed so a ")(...)" pair cannot manufacture a depth-0
gap where a smuggled dkim=pass clause gets parsed.

SES-AR-02: emit a space when a comment is removed so comments act as
CFWS folding whitespace (RFC 5322). Without this, "dk(z)im=pass" would
become "dkim=pass" and an attacker comment could glue unrelated tokens.

Alarm: widen evaluation_periods from 3→6 (30-min window) with
datapoints_to_alarm still at 2, closing the sparse-outage residual where
rejections >10-15 min apart fail to place breaching datapoints in 2 of 3
consecutive periods.

Both ses_auth.py copies stay byte-identical. All 86 tests pass.

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-07-16 19:00:40 +00:00
Adam Moussa
7b9e26d79d
Add fail-closed SES sender authentication (INFRA-107) (#98)
Some checks are pending
Deploy / deploy (push) Waiting to run
* Add fail-closed SES sender authentication

The From header and any raw-MIME Authentication-Results copies are
attacker-forgeable, so a forged email to apm@int.seahaven.com or
amazon_po@int.seahaven.com could create or mutate a WO/PO (INFRA-107,
CRITICAL). Both S3-triggered email processors now authenticate the
sender against the Authentication-Results header SES itself prepends
at delivery: only the topmost header is consulted, its authserv-id
must be amazonses.com, and it must carry dkim=pass for a domain in
the per-pipeline ALLOWED_DKIM_DOMAINS env var (comma-separated, set
in CDK so ops can adjust without code changes).

Allowlists come from live traffic observed 2026-07-15 on both ingest
buckets: WO mail arrives via the apm@ Google Groups forward, which
re-signs as seahaven.com (the hxgnsmartcloud.com signature does not
survive the forward); PO mail passes for amazon.coupahost.com.
amazonses.com also passes on PO mail but is deliberately excluded --
every SES customer's outbound mail passes for it.

Every failure path (env var unset, header missing or unparseable,
verdict fail, unaligned domain) rejects the email: a structured
warning with the reason and S3 key is logged and the record skipped
without erroring the invocation, so rejected mail causes no Lambda
retries or DLQ messages. Handler signatures and event sources are
unchanged.

Refs: INFRA-107

* Harden AR parser per cross-family review

Cross-family (GPT-4.1) review findings: terminate the dkim result
token at end-of-clause, whitespace, or a comment so a value like
"dkim=pass-fake" can never be read as a pass; normalize trailing
dots off allowlist entries so "seahaven.com." matches; make the
compat32 parser policy explicit. Adds tests for result-token
boundaries, comments after the result, quoted domain values, and
folding inside a dkim clause.

Refs: INFRA-107

* Harden AR parsing and alarm on sender-auth rejects

The SES-stamped Authentication-Results value echoes attacker-controlled
SMTP-session tokens (envelope-from, helo, header.from) as their own
semicolon-delimited property clauses. A naive split(";") tore an RFC 5321
quoted-local-part MAIL FROM apart and manufactured a forged dkim=pass
clause, so a fully spoofed email was accepted on the genuinely
SES-stamped topmost header. Tokenise comment- and quoted-string-aware
(RFC 8601 / RFC 5322): strip CFWS comments, split clauses only on
semicolons outside a quoted-string, and fail closed on unbalanced
quotes/comments so a ';' inside a quoted pvalue can never start a clause.

Rejected mail returns normally (no error, no retry, no DLQ message), so a
signing-domain drift or a wrong allowlist would silently discard 100% of
legitimate mail while every alarm stayed green. Add a CloudWatch Logs
metric filter + alarm on the sender_auth_rejected warning to both stacks
so a false-reject storm pages instead of vanishing. This is also the
safety net for the WO seahaven.com allowlist assumption, which must be
validated against a live SES-stamped header (a plain Gmail auto-forward
re-signs under the sending Workspace domain, not seahaven.com).

Refs: INFRA-107

* chore: retrigger CI (no run recorded for 7c74ac1)

* Fix quoted-AUID DKIM domain spoof in sender auth

Resolve three confirmed /sh-security-review findings on the fail-closed
SES sender-authentication control.

HIGH: header.i/header.d domain extraction was not quoted-string aware.
An attacker with a valid DKIM key for their own domain could set an
RFC 6376-legal AUID such as i="@seahaven.com"@attacker.com; the naive
extractor stopped at the closing quote and returned seahaven.com,
accepting forged mail. Extraction now tokenises the clause with the same
quoted-string discipline already used for clause splitting: header.d
(the plain signing domain) is authoritative when present, otherwise the
header.i domain is the part after the AUID's LAST top-level "@", so a "@"
inside a quoted local-part is treated as signer-controlled label text and
yields the true signer (attacker.com), not seahaven.com.

LOW: the topmost-header parse ran outside evaluate_sender_authentication's
try/except, so an unexpected parser exception on crafted input could
propagate into the handler and Lambda async retries/DLQ. The parse now
fails CLOSED with an authentication_results_unparseable reason.

MEDIUM: the sender_auth_rejected alarm used Sum>=3 over 15 min, blind to
a low-volume total-reject outage (a trickle that never sums to 3). Both
stacks now alarm on >=1 reject per 5-min period with evaluation_periods=3
/ datapoints_to_alarm=2, so a sustained reject condition pages even at one
reject per period while a lone stray probe self-clears.

Refs: INFRA-107

* Load Lambda function dir on sys.path in tests

Rebasing INFRA-107 onto main folded #95's pytest suite into this
branch's tests. The unified conftest loads the PO/WO handlers by file
path, and handler.py now does `from ses_auth import
authenticate_inbound_email` -- a bare sibling import that resolves in
the Lambda only because the runtime puts each function's own directory
on sys.path. The shared load_handler now adds that directory so the
handler tests import correctly alongside the sender-auth tests.

Refs: INFRA-107

* Note #97 test files in README directory tree

The rebase onto main brought in #97's tests/requirements.txt and
tests/test_po_merge.py. List both in the directory tree so it matches
the tree on disk.

Refs: INFRA-107

* Document INFRA-107 forwarder-binding risk acceptance

Record the accepted risk that WO sender auth binds to the apm@ forward's
re-signing domain (seahaven.com) rather than the Hexagon originator; the
apm@ Google Group's restricted posting policy is the load-bearing control
(escalates to HIGH if the group is opened to external posting). Also
correct the sender-auth-rejected alarm docs to match the shipped config
(>=1 per 5-min, 2-of-3 datapoints, not the superseded >=3/15min) and
note the SES-AR-01/02 parser hardening follow-ups.

Refs: INFRA-107
2026-07-15 20:58:47 -04:00