open-swe/deploy/MIGRATION.md
Adam Moussa 860ce93ad7
docs: record final managed-deployment topology in MIGRATION.md (#77)
Add an authoritative "Final, verified topology" section (1a) and
reconcile the phased plan with the executed end-state, superseding the
spike-era values throughout.

Captures: two LangGraph Cloud deployments (dev/prod, one LangSmith
workspace) with their final URL hashes; one Vercel project open-swe-prod
with production + custom dev environments and per-env LANGGRAPH_BACKEND_URL;
the Nitro routeRules proxy mechanism (PR #76, superseding the vercel.json
rewrite and PR #75's build-vercel-output.mjs); two dev/prod-isolated
GitHub Apps; per-deployment user stores; the Bedrock IAM users; and the
seven hard-won operational gotchas.

Refs: #65 #74 #76
2026-06-30 15:00:28 -04:00

454 lines
45 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Open SWE — Migration Plan: Self-Hosted AWS → Managed LangGraph Cloud + Vercel
**Repo:** `Sea-Haven-Industries/open-swe` (private) · **AWS:** 328440206208 / us-east-1
**Author:** Adam Moussa · **Date:** 2026-06-29 (final topology added 2026-06-30) · **Status:** EXECUTED — managed cutover live; §§3–11 below are the original (now-historical) phased plan, **superseded by §1a for all current-state facts (URLs, project layout, env)**.
> **READ §1a FIRST.** The phased plan (§§2–11) and the Phase A/C spike notes capture how we got here and still hold for rationale, cost, and rollback. But the spike-era specifics they cite — the single `open-swe-dashboard` Vercel project, the `open-swe-dev-hosted-…`/`open-swe-v3-…` deployment URLs, the `ui/vercel.json` same-origin rewrite, the single GitHub App — are **stale**. §1a is the authoritative final topology and wins on every conflict.
---
## 1. Executive Summary
### Decision (settled — not re-litigated here)
Move the Open SWE deployment off the bespoke self-hosted AWS stack (stock `langgraph dev`, in-memory store, no durability — issue **#9**) onto the **managed runtime**:
- **Backend → LangGraph Cloud** (a.k.a. "LangSmith Deployment"): git-connected, auto-builds a revision on push, zero-downtime, revision rollback, durable Postgres-backed store/checkpointer.
- **UI → Vercel**: `ui/` SPA, atomic deploys, instant rollback, per-PR previews (net-new capability).
### Why
1. **Dissolves the licensing blocker.** Self-hosted `langgraph up` on the LangSmith **Plus** key is "Self-Hosted Lite" (1M node-exec/yr cap, Elastic License 2.0, production legally ambiguous). Managed **IS** the licensed product with **no node cap** — runs billed flat at $0.005, nodes not charged.
2. **Deletes the entire durable-runtime build** (RDS/Redis/Docker/AMI) **and most of the bespoke AWS CD** (S3 release pipeline, SSM deploy docs, packer, CDK box/ALB stacks).
3. **It's the upstream-canonical deployment** and the repo is **already wired for it**: `ui/vercel.json` has the same-origin `/dashboard/api/*` rewrite; `langgraph.json` is cloud-format (6 graphs + `http.app`).
### Cost
≈ **$160/mo incremental** for one always-on prod deployment:
- $0.0036/min prod uptime ≈ **$155/mo** (always-on)
- + **$0.005/run**
- + traces pay-as-you-go above 10k/mo
- The **$39/mo Plus seat is already paid** (not incremental)
- **Dev deployment is FREE** (1 included on Plus, preemptible)
A self-hosted `langgraph up` + RDS plan (already `/sh-plan-review`'d to APPROVE-after-revision this session) is the documented **FALLBACK** if managed is ever rejected (see §11).
---
## 1a. Final, verified topology (AUTHORITATIVE — supersedes spike-era values)
This is the live managed deployment as of 2026-06-30. Where any later section disagrees (old URLs, a single Vercel project, a single GitHub App, the `ui/vercel.json` rewrite), **this section wins**.
### Backend — managed LangGraph Cloud (two deployments, one LangSmith workspace)
Both deployments live in the **same LangSmith workspace**; the **same workspace API key authenticates both** (including the Store API — so per-deployment store writes use that one key with the per-deployment URL).
| Deployment | URL | Git connection |
|---|---|---|
| **dev** | `https://open-swe-dev-fb737aa219605c8bbdb30ecbb33f30c0.us.langgraph.app` | branch `dev` |
| **prod** | `https://open-swe-prod-d6c7bb63aaa651b6a1d92f9492b1d983.us.langgraph.app` | branch `main` (auto-deploys on push to `main`) |
> The dev deployment was **renamed `open-swe-dev`, deleted, and recreated** — which minted the **new URL hash** above. The spike-era `open-swe-v3-…` / `open-swe-dev-hosted-…` URLs are **dead/superseded**. Deleting + recreating a deployment is the one operation that changes the URL hash (otherwise stable across revisions) — when it happens, update **every** reference (Vercel env, GitHub App webhooks, OAuth callbacks, docs).
### UI — Vercel (ONE project, two environments)
**One** Vercel project `open-swe-prod` (team `sea-haven`, id `prj_OOh6yjXMp4ah3Ws3Y7XRQxjmMmQU`). The old separate `open-swe-dashboard` project was **DELETED**.
| Vercel environment | Branch | Backend | Custom domain |
|---|---|---|---|
| production | `main` | prod deployment URL | `openswe.seahaven.com` |
| custom **`dev`** (id `env_SMI23PULAJXk0GhwE0HLVhp5J3ZS`) | `dev` | dev deployment URL | `openswe-dev.seahaven.com` |
- A **per-environment** env var `LANGGRAPH_BACKEND_URL` (prod env = prod URL, dev env = dev URL) drives the `/dashboard/api/*` proxy.
- Project settings: `framework=null`, `outputDirectory` cleared, root directory `ui`.
- **Proxy mechanism (current, after PR #76):** **Nitro `routeRules`** in `ui/vite.config.ts` read `process.env.LANGGRAPH_BACKEND_URL` and Nitro's Vercel preset compiles them into `.vercel/output/config.json` (Build Output API) at build time — a CDN-level proxy (not redirect, so the `osw_session` cookie stays first-party). PR #75's hand-rolled `ui/scripts/build-vercel-output.mjs` was the broken first attempt and is **gone** (`ui/scripts/` no longer exists). **Never** add a manual script that `rm`s `.vercel/output` — Nitro's Vercel preset auto-emits it.
### DNS — Route 53 zone `seahaven.com` (`Z06652411XKH89KTZD3XA`)
- `openswe.seahaven.com` → CNAME `cname.vercel-dns.com` (prod env)
- `openswe-dev.seahaven.com` → CNAME to Vercel (dev env)
### GitHub Apps — TWO (dev/prod isolated; each its own webhook URL)
| App | app_id | install | client_id | org | members scope | repos |
|---|---|---|---|---|---|---|
| **prod** `seahaven-openswe` | `4146115` | `142615168` | `Iv23lil96pKQNNDUn5yp` | `Sea-Haven-Industries` | members:**write** | all |
| **dev** `seahaven-openswe-dev` | `4162963` | `143023302` | `Iv23licQwJvGAPJj1HJe` | `seahaven-open-swe-dev` | members:**read** | all |
- **Promotion App** `seahaven-promotion` (actor `4170147`) is the sole non-admin fast-forward-push bypass on the `main` ruleset `18238334` — its FF-push of `dev → main` is what triggers the managed prod build.
### Env per deployment (set in LangGraph Cloud config + Vercel env — NOT Secrets Manager)
| Var | dev | prod |
|---|---|---|
| `LANGGRAPH_URL` | own (dev) deployment URL | own (prod) deployment URL |
| `DASHBOARD_BASE_URL` / `DASHBOARD_API_BASE_URL` | `https://openswe-dev.seahaven.com` | `https://openswe.seahaven.com` |
| `VITE_DASHBOARD_API_BASE_URL` | empty (same-origin via Vercel proxy) | empty |
| `ALLOWED_GITHUB_ORGS` | dev org (`seahaven-open-swe-dev`) | `Sea-Haven-Industries` |
| `CONFIGURED_ADMINS` | `amoussa1229,adam@seahavenind.com` | `amoussa1229,adam@seahavenind.com` |
`DASHBOARD_BASE_URL` / `DASHBOARD_API_BASE_URL` **must include `https://`** (see gotcha 3). Secret **values** are still sourced from `open-swe-{dev,prod}/*` Secrets Manager + SSM (the remaining source of truth) and set into the LangGraph Cloud + Vercel env stores — the accepted `secrets-and-config.md` deviation.
### Bedrock IAM (PR #74, still OPEN)
Two IAM users `open-swe-dev-bedrock` + `open-swe-prod-bedrock`, each attached to customer-managed policy `open-swe-bedrock-invoke` (least-privilege `bedrock:InvokeModel[WithResponseStream]` on the `us.anthropic.claude-opus-4-8` inference-profile ARN + its 3 routed foundation-model ARNs in us-east-1/us-east-2/us-west-2). Default model `bedrock_converse:us.anthropic.claude-opus-4-8` + 3 Fireworks models. Static access keys live only in the deployment env (dev key → dev, prod key → prod).
### User store — per-deployment
**Each** managed deployment has its **own** Store. The GitHub→email mapping `amoussa1229 → adam@seahavenind.com` (namespace `["user_mappings"]`, key = lowercased login, record `{github_login, work_email, status:"active", source, created_at, updated_at}`) was written to **both** the dev and prod stores directly. New users need a mapping **per-deployment** (write each store directly, or use the dashboard admin User-mappings UI — the `work_email` field was added by PR #65 fix #4).
### AWS decommission (PR #64)
Self-host CDK stacks destroyed. Residual: `CDKToolkit` (shared, **preserved**); ~50 `RETAIN`'d Secrets Manager shells + 3 S3 asset buckets (**pending cleanup**); AWS **Bedrock** (live dependency, kept).
### Operational gotchas (hard-won — carry these into any runbook)
1. **Per-deployment store → seed user mappings per-deployment.** A missing mapping makes `process_github_issue` silently early-return ("No email mapping … skipping"): the webhook returns 200/accepted but produces **no reaction and no run**. Seed dev **and** prod.
2. **Org-login gate uses the App *installation* token**, so the App must be org-installed with **Members:read**. OAuth working ≠ membership check working — they use **separate creds** (CLIENT_ID/SECRET for OAuth vs APP_ID/INSTALLATION_ID/PRIVATE_KEY for the install token). A mangled multi-line `GITHUB_APP_PRIVATE_KEY` breaks the install token (and thus the gate) while OAuth still works.
3. **`DASHBOARD_API_BASE_URL` must be `https://`** — an `http://` value makes GitHub reject the OAuth callback with "redirect_uri not associated."
4. **`osw_oauth_state` cookie is host-only** — start login on the **same host** as `DASHBOARD_API_BASE_URL`, or you get "oauth state mismatch."
5. **Webhooks go DIRECT to the langgraph URL** (`/webhooks/*`). Vercel only proxies `/dashboard/api/*`. The app is **same-origin only** (no CORS).
6. **On Vercel CI, Nitro's Vercel preset auto-emits `.vercel/output`** — drive the proxy via Nitro `routeRules` from `LANGGRAPH_BACKEND_URL`; never a manual script that `rm`s `.vercel/output`.
7. **Deleting + recreating a LangGraph deployment mints a NEW URL hash** (otherwise stable across revisions) — update every reference (Vercel env, webhooks, OAuth, docs).
---
## 2. Architecture: Before → After
### Before (self-hosted AWS — LIVE as of 2026-06-29)
```
GitHub/Slack/Linear ──webhook──▶ hooks.seahaven.com ─┐
Browser (dashboard) ─────────────▶ openswe.seahaven.com ─┤
▼
shared seahaven-com ALB (:443, host+path rules)
▼
EC2 (prod i-08a729e50779c4b07 t4g.large ARM64, private subnet)
nginx :80 ──proxy /dashboard/api + /webhooks──▶ langgraph dev :2024 (loopback)
in-memory store (reseeded by seed_store.sh ExecStartPost)
▼
Config: Secrets Manager open-swe-prod/* + SSM /open-swe-prod/*
Release: GitHub Actions → S3 open-swe-prod-assets/releases/* → SSM doc deploy.sh
IaC: CDK OpenSweIamStack + OpenSweDevStack + OpenSweProdStack
Sandbox: LangSmith cloud (DEFAULT_SANDBOX_SNAPSHOT_ID + GitHub proxy)
```
### After (managed — FINAL, see §1a for exact values)
```
┌──────────────── DEV lane ────────────────┐ ┌──────────────── PROD lane ───────────────┐
GitHub(dev org)/Slack/Linear │ webhook → open-swe-dev-….us.langgraph.app │ │ webhook → open-swe-prod-….us.langgraph.app│ GitHub(SHI org)/Slack/Linear
App seahaven-openswe-dev ───┘ (DIRECT to langgraph URL, /webhooks/*) │ │ (DIRECT to langgraph URL, /webhooks/*) └─── App seahaven-openswe
▼ ▼
Browser ▶ openswe-dev.seahaven.com ─┐ ┌─▶ openswe.seahaven.com ◀ Browser
│ ONE Vercel project `open-swe-prod` (team sea-haven)
│ ├─ env `dev` (branch dev) → proxies /dashboard/api/* → dev langgraph URL
│ └─ env production (branch main) → proxies /dashboard/api/* → prod langgraph URL
└─ proxy compiled by Nitro routeRules from per-env LANGGRAPH_BACKEND_URL (PR #76)
▼
TWO LangGraph Cloud deployments (same LangSmith workspace; one workspace key auths both incl. Store)
├─ dev ← branch `dev` · prod ← branch `main` (push-to-main auto-deploys prod)
├─ each serves the graphs + the custom http.app (agent.webapp:app = webhooks + dashboard API + OAuth)
├─ each has its OWN durable Postgres store + checkpointer (issue #9 SOLVED) — user mappings per-deployment
└─ env/secrets in the Deployment + Vercel config (NOT Secrets Manager) — Adam accepted deviation
▼
Sandbox: LangSmith cloud (UNCHANGED — DEFAULT_SANDBOX_SNAPSHOT_ID + GitHub-App proxy)
Bedrock: IAM users open-swe-{dev,prod}-bedrock + policy open-swe-bedrock-invoke (static keys in deploy env)
```
**What changes shape:** runtime host (EC2 → managed PaaS), durability (in-memory → managed Postgres, one store **per deployment**), CD (bespoke S3/SSM/packer → git-connected auto-build), config home (Secrets Manager/SSM → Deployment+Vercel env), ingress topology (shared ALB + one App → two dev/prod-isolated GitHub Apps each hitting its own `*.langgraph.app` directly), UI proxy (`ui/vercel.json` rewrite → Nitro `routeRules`). **What stays:** the LangSmith sandbox plane, CI (lint/format/unit/Playwright), the app code itself, and Secrets Manager/SSM as the secret-**value** source of truth.
---
## 3. Phased Plan
### Phase A — Dev spike (MOSTLY DONE)
Goal: prove managed serves our custom app + durability, at $0, before committing prod $.
**Proven this session** (spike-era specifics — superseded by §1a; URLs/project below are DEAD):
- ✅ Dev backend deployed to LangGraph Cloud, connected to branch `dev`:
~~`https://open-swe-dev-hosted-e76c2b0e8a7955fe8ad3110a7a54e5d0.us.langgraph.app`~~ → final dev URL in §1a (`open-swe-dev-fb737aa…`; the deployment was later deleted + recreated)
(Bedrock + a Fireworks key set; AWS creds for Bedrock deferred — see §10 open decision).
- ✅ UI deployed to Vercel — team `sea-haven`, ~~project `open-swe-dashboard`, `https://open-swe-dashboard.vercel.app`~~ (DELETED); now the **single** project `open-swe-prod` with dev/prod environments (§1a); proxy is Nitro `routeRules`, not the `ui/vercel.json` rewrite.
- ✅ Managed serves the custom `http.app` (dashboard API + webhooks) — **no platform auth gate** in front of our routes (webhook returns 401 sig-enforced, so signatures still govern).
- ✅ Vercel same-origin rewrite → backend works.
- ✅ GitHub OAuth dashboard login end-to-end.
- ✅ **DURABILITY** — team-default model survived a revision redeploy (this is issue **#9**'s core acceptance goal, proven on managed).
**Remaining Phase A items (the GitHub `@openswe` run-trigger loop):**
- [ ] A1. Get an `@openswe` GitHub comment to dispatch a run end-to-end on dev. Blocked by the user-mapping + cache gotchas (see fixes #4 and #5 in §5).
- [ ] A2. Create the owner user-mapping in the managed Store (namespace `["user_mappings"]`, key = lowercased login `amoussa1229`, value `{github_login, work_email}`) — the Admin UI can't set `work_email` yet (fix #4).
- [ ] A3. Verify a freshly-added mapping is seen without a redeploy across managed's multi-replica autoscaling (fix #5).
- [ ] A4. Confirm `LANGGRAPH_URL` is set to the deployment URL on the dev deployment so `langgraph_client()` calls resolve (fix #1 — config now, code later).
### Phase B — Harden (the code fixes + env codification)
Land the 6 code fixes (§5), codify env/config, and resolve the Bedrock-auth decision **before** standing up prod.
- [ ] B1. Land code fixes #1–#6 (§5) as a PR into `dev` (or split into focused PRs). Re-run `make lint` + `make test`.
- [ ] B2. **Codify the env contract for managed.** Update `ui/vercel.json` rewrite to the *prod* deployment URL (Phase C) but keep dev pointing at dev. Add a documented LangGraph-Cloud env list to the repo (see §4) — NOT secrets, just the variable inventory + which are excluded.
- [ ] B3. **Resolve the Bedrock-auth decision** (§10 open decision): static AWS keys for a Bedrock-scoped IAM user in the deployment env, OR run the agent on Fireworks. This intersects the #62 model work.
- [x] B4. ~~Merge #62 first~~ — **DONE / N/A.** #62 (Bedrock + Fireworks) **IS** merged to `dev` and is what the dev deployment runs — verified on `origin/dev` (`options.py` → `DEFAULT_MODEL_ID = "bedrock_converse:us.anthropic.claude-opus-4-8"`) and confirmed by the live deployment's `git_ref_sha: a4ed19ba…` (= the #62 merge commit). The contrary note came from a STALE local checkout; no merge action needed.
- [ ] B5. Audit ALL in-process caches for the single-process → multi-replica assumption (generalization of fix #5; see §5).
- [ ] B6. Measure + reduce custom-app import time (fix #6) so the deployment isn't flagged unhealthy / slow to scale.
### Phase C — Prod deployment
- [ ] C1. Create a **prod LangGraph Cloud deployment** tracking branch `main` (the durable autoscaled 1→10 tier, not the free Dev tier). Record its `*.langgraph.app` URL.
- [x] C2. ~~Create the **prod Vercel project/target** (or promote the existing `open-swe-dashboard` to production)~~ — **DONE differently:** one project `open-swe-prod` with a production env (`main`) + a custom `dev` env (`dev`), each with its own `LANGGRAPH_BACKEND_URL`. See §1a.
- [ ] C3. Set the **prod env triad** (§4) on the prod deployment + Vercel:
- `LANGGRAPH_URL` = the prod `*.langgraph.app` URL
- `DASHBOARD_BASE_URL` + `DASHBOARD_API_BASE_URL` = the prod Vercel origin (with `https://` scheme — fix #2)
- [ ] C4. **Establish the prod approval gate** (§6): git-connected auto-deploy needs an explicit gate. Preserve the prod manual-approval that exists today (the GitHub `prod` Environment reviewer). Mechanism = protected `main` ruleset + the platform's "require manual promotion to production" if available (§10 open decision).
- [ ] C5. **Repoint webhooks + OAuth URLs** to prod:
- GitHub App `seahaven-openswe` webhook URL → prod backend webhook URL (`*.langgraph.app/webhooks/github` or `hooks.seahaven.com` — §10 custom-domain decision)
- Slack Event Subscriptions + Interactivity Request URLs → prod
- Linear webhook URL → prod (if/when wired)
- GitHub App OAuth callback → `<DASHBOARD_API_BASE_URL>/dashboard/api/auth/callback` (the Vercel prod origin; fixes #2, #3)
- [ ] C6. **Custom-domain decision** (§10): dashboard custom domain solved by Vercel; webhooks `hooks.seahaven.com` either re-point to `*.langgraph.app` directly or front via CNAME. Managed issues `*.langgraph.app` URLs ONLY (no documented custom-domain support on the backend).
### Phase D — Cutover + verify
- [ ] D1. Run the issue **#9 acceptance criteria on PROD**: durable store survives a revision redeploy; team defaults / user mappings persist; a paused/in-flight run survives a redeploy (the durability win self-host never had).
- [ ] D2. End-to-end prod smoke: `@openswe` GitHub comment → run dispatch → sandbox → draft PR → reply in source channel.
- [ ] D3. Dashboard prod smoke: GitHub OAuth login on the stable Vercel alias (fix #3); admin pages reachable for `CONFIGURED_ADMINS`.
- [ ] D4. Webhook sig-enforcement smoke: unsigned POST to each `/webhooks/*` returns 401; signed returns 200.
- [ ] D5. **Soak** for an agreed window (recommend 3–7 days) with the AWS prod stack still standing as the rollback target (§11) before any teardown.
### Phase E — AWS decommission (AFTER soak)
Only after managed prod is proven + soaked. See §7 for the precise retire-vs-keep list. Phased teardown:
- [ ] E1. Export the live ALB listener-rule / target-group config for `open-swe-prod` and `open-swe-dev` to JSON (the export is the rollback source — these rules were partly hand-built).
- [ ] E2. Disable the bespoke CD (build-artifacts / cd-infra workflows) so nothing re-deploys the EC2 boxes.
- [ ] E3. `cdk destroy OpenSweProdStack` then `OpenSweDevStack` (mind the **RETAIN** secret shells — they survive and hold the global `open-swe-<env>/*` names; force-delete only EMPTY shells once env is fully migrated; keep populated ones transitionally as the env source — §7).
- [ ] E4. Remove ALB rules/target groups, S3 assets buckets, SSM deploy docs, and the per-env CDK bootstrap qualifier `oswedev` (`CDKToolkit-oswedev`) — once nothing references them.
- [ ] E5. Decide DNS: repoint `openswe.seahaven.com` → Vercel; resolve `hooks.seahaven.com` → `*.langgraph.app` (or CNAME).
- [ ] E6. Leave `OpenSweIamStack` until last — the promotion App + any retained roles may still be referenced.
### Phase F — Docs / memory
- [ ] F1. Rework Confluence "AWS Architecture Map" (page **1540098**) + the open-swe child page (**26116098**): the RDS/EC2/ALB subgraph is replaced by an **external-services view** (LangGraph Cloud + Vercel + the retained GitHub App + LangSmith sandbox).
- [ ] F2. Update `README.md` + `INSTALLATION.md` (§10 is already canonical-correct; align Sea Haven specifics).
- [ ] F3. Update project memories `project_open_swe_migration.md` + `reference_open_swe_deployment.md` to mark the managed cutover and the AWS teardown.
---
## 4. Env / Config Reference
### The prod triad (per INSTALLATION.md §10, lines 630–658) — see §1a for the exact dev/prod values
| Var | Value (per deployment/env) | Notes |
|---|---|---|
| `LANGGRAPH_URL` | the **own** deployment URL (`https://...langgraph.app`) | **NOT** localhost. Drives `thread_ops.langgraph_url()` (fix #1). dev→dev URL, prod→prod URL. |
| `DASHBOARD_BASE_URL` | the own dashboard origin, **with `https://`** (`https://openswe-dev.seahaven.com` / `https://openswe.seahaven.com`) | |
| `DASHBOARD_API_BASE_URL` | same as above, **with `https://` scheme** | scheme required or OAuth `redirect_uri` is schemeless and GitHub rejects (fix #2) |
| `VITE_DASHBOARD_API_BASE_URL` | **empty** | same-origin mode; UI calls relative `/dashboard/api/*`, the Vercel Nitro proxy rewrites to backend |
| `LANGGRAPH_BACKEND_URL` | **Vercel env var**, per Vercel environment (dev env = dev URL, prod env = prod URL) | drives the Nitro `routeRules` `/dashboard/api/*` proxy at build time (PR #76) — required on Vercel builds |
| `ALLOWED_GITHUB_ORGS` | `Sea-Haven-Industries` (prod) / `seahaven-open-swe-dev` (dev) | org-login gate; checked via the App **installation** token (gotcha 2) |
GitHub App dashboard OAuth callback = `<DASHBOARD_API_BASE_URL>/dashboard/api/auth/callback` (the own dashboard origin — prod App on `openswe.seahaven.com`, dev App on `openswe-dev.seahaven.com`).
**UI proxy mechanism (final, PR #76):** the `/dashboard/api/*` proxy is **Nitro `routeRules`** in `ui/vite.config.ts` reading `process.env.LANGGRAPH_BACKEND_URL`, compiled by Nitro's Vercel preset into `.vercel/output/config.json`. This **supersedes** the spike-era `ui/vercel.json` same-origin rewrite and PR #75's hand-rolled `ui/scripts/build-vercel-output.mjs` (deleted). `ui/vercel.json` now only carries `framework:null` + `buildCommand: bun run build`.
### Where secrets/env live now
**LangGraph Cloud Deployment config + Vercel env** — NOT AWS Secrets Manager. This **deviates from the Sea Haven `secrets-and-config.md` "Secrets Manager for all sensitive" handbook rule** — **Adam ACCEPTED this deviation** (managed has no instance role / no fetch-config boot hook; the platform's own secret store is the mechanism).
### Sourcing env from the existing AWS stack (transitional)
Pull from the existing `open-swe-dev` Secrets Manager + SSM via a documented CLI dump:
- `aws secretsmanager list-secrets --filters Key=name,Values=open-swe-dev/` then per-name `get-secret-value` — **NOT** `batch-get-secret-value` (its pagination silently drops values past page 1; this bit the box twice — see memory).
- SSM: `aws ssm get-parameters-by-path --path /open-swe-dev/`.
### EXCLUDE when copying to managed
| Category | Vars |
|---|---|
| Box-/self-host-specific | `LANGGRAPH_URL` (set fresh to deployment URL), `LANGGRAPH_URL_PROD`, `LANGSMITH_ENDPOINT`, `LANGSMITH_ENDPOINT_PROD`, `LANGSMITH_URL_PROD`, `LANGSMITH_TENANT_ID_PROD`, `LANGSMITH_HOST_API_URL`, `LANGCHAIN_REVISION_ID` |
| Dropped providers (PR #62) | `OPENAI_API_KEY`, `GOOGLE_API_KEY`, `GROQ_API_KEY` (these 6 Secrets Manager shells deleted 2026-06-29, final purge 2026-07-06) |
| Unused sandbox providers | `DAYTONA_API_KEY`, `RUNLOOP_API_KEY` |
| Eval-only | `JUDGE_ANTHROPIC_API_KEY` (kept on AWS for `evals/reviewer/judge.py`, not needed in the runtime deployment) |
### KEY mappings on managed
- `LANGSMITH_API_KEY` = the **`LANGSMITH_API_KEY_PROD`** value (and `LANGCHAIN_API_KEY` = same).
- The full secret inventory is `SECRET_VARS` in `infra/lib/constructs/config-store.ts:46` (28 shells) — use it as the checklist of what to carry, minus the EXCLUDE rows above.
### Models (post-#62) + Bedrock auth
Post-#62 `SUPPORTED_MODELS` = `bedrock_converse:us.anthropic.claude-opus-4-8` (**DEFAULT**) + 3 Fireworks models; **no direct `anthropic:` option**.
✅ **#62 is merged to `dev` and deployed** — verified on `origin/dev` (`options.py` → `bedrock_converse` default) and the live deployment `git_ref_sha a4ed19ba`. (The pre-#62 values appear only on a stale LOCAL checkout — ignore.)
**Bedrock on managed has NO EC2 instance role** (the #62 IAM design used the EC2 instance role as the Bedrock principal — that breaks on managed). Options (OPEN DECISION, §10):
- (a) Static `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` / `AWS_REGION` for a **Bedrock-scoped IAM user** in the deployment env, or
- (b) Run the agent on **Fireworks** and avoid Bedrock entirely on managed.
### Bedrock IAM users (CREATED — resolves §10.1, discharges the §9 IAM gates)
Option (a) was chosen and executed live (2026-06-30, account **328440206208** / **us-east-1**). This is **click-ops IAM** — there is no remaining open-swe AWS IaC after the decommission (#64), so these are created with the CLI, not CDK.
- **Customer-managed policy `open-swe-bedrock-invoke`** (`arn:aws:iam::328440206208:policy/open-swe-bedrock-invoke`). Least-privilege: actions `bedrock:InvokeModel` + `bedrock:InvokeModelWithResponseStream` **only**, scoped to exactly the `us.anthropic.claude-opus-4-8` inference-profile ARN + its **three** routed foundation-model ARNs (us-east-1, us-east-2, us-west-2). **No wildcards, no other models.** (Supersedes the §10.1 plan to reuse the #62 instance-role policy — a fresh standalone policy was minted instead.)
- **Two IAM users**, each attached to that policy: **`open-swe-dev-bedrock`** and **`open-swe-prod-bedrock`**. Tagged `project=open-swe`, `managed-by=cli-migration`, `purpose=bedrock-invoke`.
- **Static access keys are minted separately by the owner** (`aws iam create-access-key`) — secret keys live **only** in the deployment env stores, never in this repo or memory. dev key → dev env; prod key → LangGraph Cloud prod config (`AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` / `AWS_REGION=us-east-1`).
- **Deviation rationale:** static long-lived keys are a deliberate departure from the Sea Haven OIDC norm because **managed LangGraph Cloud cannot assume an AWS role**. Mitigated by the tight least-privilege policy above.
- ✅ **Both mandatory gates ran and PASSED:**
- **GPT-4.1 IAM cross-review** (confirmed hit `gpt-4.1-2025-04-14`) — least-privilege confirmed.
- **`/sh-security-review`** — **0 critical/high**; two **accepted mediums** (the static-key deviation + no per-principal budget cap).
- **Recommended follow-ups:** shortest viable key-rotation cadence with recorded creation dates; an AWS Budgets / CloudWatch anomaly alarm on per-principal Bedrock `InvokeModel` volume; confirm CloudTrail captures these users.
---
## 5. Required Code Fixes (the 6 gotchas)
These are self-host → managed assumption breaks discovered on the spike. Each gets a config workaround now and a code fix to land in Phase B.
### Fix #1 — `langgraph_url()` localhost fallback
**File:** `agent/utils/thread_ops.py:38-41`
```python
def langgraph_url() -> str:
return os.environ.get("LANGGRAPH_URL") or os.environ.get(
"LANGGRAPH_URL_PROD", "http://localhost:2024"
)
```
On managed, every `langgraph_client()` call (thread sidebar, run creation, the same pattern in `agent/dashboard/user_mappings.py:43` `get_client()` with no URL) `ConnectError`s unless `LANGGRAPH_URL` is set to the deployment URL.
- **Config fix (now):** set `LANGGRAPH_URL` = the deployment URL on every deployment (Phase A4 / C3).
- **Code fix (Phase B):** default to the in-process / deployment URL rather than `http://localhost:2024` when running inside a managed deployment (e.g. honor the platform's own URL env, or fail loud instead of silently hitting localhost).
### Fix #2 — OAuth `DASHBOARD_API_BASE_URL` must include scheme
**Where:** OAuth redirect construction (dashboard `oauth.py` / `routes.py`), driven by `DASHBOARD_API_BASE_URL`.
A schemeless value yields a schemeless `redirect_uri` → GitHub rejects with "redirect_uri not associated."
- **Fix:** always set `DASHBOARD_API_BASE_URL` **with `https://`** (Phase C3). Code hardening (Phase B): assert/normalize a scheme at startup and fail loud if missing.
### Fix #3 — OAuth state cookie is host-only
**Cookie:** `osw_oauth_state` (host-only). Login must START on the SAME host as `DASHBOARD_API_BASE_URL` — the **stable Vercel alias**, never the immutable per-deploy URL — or you get "oauth state mismatch."
- **Fix:** pin `DASHBOARD_API_BASE_URL` + the login entry point to the stable alias / custom domain (Phase C3, D3). Document this so a future per-deploy preview URL isn't used for login.
### Fix #4 — Admin "User mappings" UI can't set `work_email`
**Files:** the Admin → User mappings UI (`ui/`) + `agent/dashboard/user_mappings.py` (`upsert_mapping` at line 257 takes `work_email` but the UI form doesn't expose it).
Mappings can't be fully created from the dashboard → must write the Store directly: namespace `["user_mappings"]`, key = **lowercased login**, value `{github_login, work_email}` (matching `_index_record` at `user_mappings.py:73-84`, which keys `_by_login` on `login.lower()`).
- **Fix (Phase B):** add the `work_email` field to the Admin mappings form so mappings are fully creatable from the UI.
### Fix #5 — GitHub webhook path doesn't refresh the user-mapping cache (multi-replica break)
**Files:** `agent/webapp.py:3052` (GitHub path) vs `agent/webapp.py:1091` (Slack path).
The Slack path refreshes before lookup:
```python
# agent/webapp.py:1089-1093 (Slack)
await refresh_user_mapping_cache()
...
```
The GitHub path does **not** — it calls `email = await email_for_login(github_login)` (`webapp.py:3052`, again at `:3331`) cold. The cache (`user_mappings.py` `_ensure_cache_loaded`, line 197) is **one-shot per process** (`_cache_loaded` flag). On self-host single-process this was fine; on managed's **multi-replica autoscaling**, a freshly-added mapping isn't seen by a replica whose cache loaded earlier — until restart.
- **Fix:** refresh-before-lookup on the GitHub path (mirror the Slack path), or add a TTL / cross-replica invalidation to the cache.
- **Generalize (Phase B5):** audit ALL in-process caches for the single-process → multi-replica assumption — `SANDBOX_BACKENDS` dict (`agent/utils/sandbox_state.py`), `_THREAD_RUN_LOCKS` (`thread_ops.py:18`), `_by_login`/`_by_email`/`_by_slack_id` (`user_mappings.py:67-69`). Sandbox affinity is already thread-keyed + persisted in thread metadata (`sandbox_id`), so it's the cache/lock state that needs the multi-replica review.
### Fix #6 — Slow custom-app import (~8s startup)
**Symptom:** "exceeded expected startup time" → risks the deployment being marked unhealthy / slow to scale out.
- **Fix (Phase B6):** lazy imports / reduce import-time work in `agent/webapp.py` and the graph factories. Profile with `FF_PROFILE_IMPORTS` (the import-profiling flag) to find the heavy modules.
---
## 6. CD / Ops Changes
### Retires (bespoke pipeline)
- GitHub Actions **build → S3 releases → SSM doc → `deploy.sh` → health-gate** (`.github/scripts/`, `deploy/ami/deploy.sh`)
- `roll-box.sh` / `publish-and-deploy.sh` / `rollback.sh`
- AMI baking (packer `deploy/ami/open-swe-base.pkr.hcl`) + `cdk.context.json` AMI pin
- `cd-infra.yml` CDK deploys + OIDC bootstrap qualifiers
- `fetch-config.sh` / `seed_store.sh` (boot-time config materialization + Store reseed)
### Replaced by git-connected PaaS
- **LangGraph Cloud** auto-builds a revision on push (first-party zero-downtime + revision rollback).
- **Vercel** builds / atomic-deploys / instant-rollback + per-PR previews (**net-new** capability the AWS stack never had).
### Stays
- **CI** (lint / format / unit / Playwright E2E) — still matters and still gates merges. `make lint`, `make test`.
### Gate re-homing
- The **dev → main promotion gate** (`check-dev-green.sh` + protected-`main` ruleset 18238334) **re-homes**: the **dev deployment tracks `dev`**, the **prod deployment tracks `main`**, so the gate governs exactly what reaches prod.
- Today's promotion machinery: `promote-dev-to-prod.yml` + the `seahaven-promotion` GitHub App (app_id 4170147, in ruleset 18238334 bypass_actors) FF-pushes `dev → main`. Under managed, a push to `main` is what triggers the prod build — so the promotion gate IS the prod deploy gate.
### Preserve the PROD manual-approval gate (LOAD-BEARING)
Today the `prod` GitHub Environment (required reviewer `amoussa1229`) is the manual approval. Git-connected auto-deploy removes the CD job that consulted that Environment, so the approval must be re-established explicitly:
- **Mechanism options (OPEN DECISION §10):** protected `main` (only the promotion App can FF-push, and that push is itself the gate) AND/OR the platform's "require manual promotion to production" if LangGraph Cloud exposes it.
- **Norm shift:** deploy-then-merge → **merge-to-deploy**. Lean on the dev deployment + Vercel previews to verify *before* promoting `dev → main`. (This inverts the Sea Haven handbook deploy-then-merge default — call it out in the README + handbook note.)
---
## 7. AWS Decommission — Retire vs Keep
### RETIRE (becomes vestigial once managed prod is live + soaked)
| Item | Path / resource |
|---|---|
| AMI bake + cloud-init | `deploy/ami/` (`open-swe-base.pkr.hcl`, `provision.sh`, `user-data.sh`, `deploy.sh`, `templates/`) |
| Self-host boot/config scripts | `deploy/seahaven/fetch-config.sh`, `seed_store.sh`, `put-config.sh`, `nginx/openswe.conf`, `systemd/open-swe.service`, `aegra/` (already deferred) |
| CDK box/ALB/AMI stacks | `infra/lib/open-swe-stack.ts`, `constructs/app-service.ts`, `assets-bucket.ts`, `ami-cache.ts`, `instance-role.ts`, `github-deploy-roles.ts` |
| EC2 instances | prod `i-08a729e50779c4b07` (t4g.large), dev `i-0a3bb8e0ddd36c29b` |
| S3 release buckets | `open-swe-dev-assets`, `open-swe-prod-assets` |
| SSM deploy docs | `open-swe-dev-deploy`, `open-swe-prod-deploy` |
| ALB plumbing | `open-swe-prod-tg` / `open-swe-dev-tg` target groups + the `seahaven-com` ALB host/path listener rules for openswe/hooks |
| Per-env bootstrap qualifier | dev `oswedev` (`CDKToolkit-oswedev`) |
| RDS / Redis durable plan | **never built** — fully dropped (managed provides durability) |
| Bespoke CD workflows | `cd-infra.yml`, `build-artifacts.yml`, `roll-box.sh`/`publish-and-deploy.sh`/`rollback.sh` |
### KEEP
| Item | Why |
|---|---|
| **GitHub App `seahaven-openswe`** (App 4146115 / Install 142615168) | unchanged auth + webhook source; only the webhook/OAuth URLs repoint |
| **LangSmith sandbox setup** incl. `DEFAULT_SANDBOX_SNAPSHOT_ID` (`dc36e509-d4d2-4efc-8a4e-61f74f3446ec`) | the only sandbox with working in-sandbox git/gh auth (`_configure_github_proxy`); managed default `SANDBOX_TYPE=langsmith` uses it |
| **Secrets Manager `open-swe-{dev,prod}/*` values** | the **source you copy env FROM**, at least transitionally — do NOT force-delete populated shells until managed env is fully cut over and verified |
| **DNS** | repoint `openswe.seahaven.com` → Vercel; decide `hooks.seahaven.com` → `*.langgraph.app` or a CNAME (§10) |
| **`seahaven-promotion` App** + main/dev rulesets (18238334 / 18238542) | still govern `dev → main` promotion = the prod deploy gate |
| **`evals/` + `JUDGE_ANTHROPIC_API_KEY`** | eval pipeline unaffected (separate from runtime) |
**Phased teardown rule:** remove AWS infra ONLY after managed prod is proven + soaked (Phase D5). The exported ALB JSON (E1) is the rollback source for the hand-built listener rules.
---
## 8. Cost
| Line item | Monthly |
|---|---|
| Prod deployment uptime ($0.0036/min, always-on) | ≈ $155 |
| Runs ($0.005/run) | usage-based |
| Traces above 10k/mo | pay-as-you-go |
| LangSmith **Plus** seat ($39) | already paid (not incremental) |
| Dev deployment | **$0** (1 free on Plus, preemptible) |
| **Incremental total** | **≈ $160/mo** |
Compare to the retired AWS stack: 2× EC2 (t4g.large prod + t4g.medium/large dev), 2× S3 buckets, ALB share, NAT egress, plus the **never-built** RDS/Redis durable tier the self-host path would have added. Managed nets out roughly cost-neutral-to-cheaper while *adding* durability + previews. **TODO: confirm current AWS run-rate for an apples-to-apples delta** (not verified in repo).
---
## 9. Sea Haven Gates & Obligations
- [ ] **`/sh-plan-review`** (GPT-4.1 `cross_reviewer`) on THIS plan **BEFORE prod cutover** (Phase C gate). ⚠️ `run.py` misroutes reviewer-framed prompts to the no-op `done` route — call `models.get_cross_reviewer()` directly (load orchestrator `.env`, `ChatOpenAI("gpt-4.1")`) per `feedback_orchestrator_usage`.
- [ ] **`/sh-security-review`** on the sensitive surface: this migration touches **auth/webhook signature verification** (the custom `http.app` is now publicly reachable on `*.langgraph.app` with no platform gate in front) and **secrets handling** (secrets move from Secrets Manager to the Deployment/Vercel env). Required, not opt-in — resolve confirmed critical/high before cutover.
- [ ] **GPT-4.1 cross-family review** if the Bedrock-auth fix changes IAM (new Bedrock-scoped IAM user + policy = IAM change → mandatory cross-review).
- [ ] **Confluence** — rework "AWS Architecture Map" (**1540098**) + open-swe child page (**26116098**): replace the RDS/EC2/ALB subgraph with an external-services view (LangGraph Cloud + Vercel + retained GitHub App + LangSmith sandbox). Do it in the same conversation as the cutover, not as a follow-up.
- [ ] **README + INSTALLATION** updated (INSTALLATION §10 already canonical; align Sea Haven specifics + the merge-to-deploy norm shift).
- [ ] **Memories** — update `project_open_swe_migration.md` + `reference_open_swe_deployment.md`.
---
## 10. Decisions — RESOLVED (2026-06-29, Adam)
1. **Bedrock auth on managed → STATIC AWS KEYS.** Create a **Bedrock-scoped IAM user**; set `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` / `AWS_REGION=us-east-1` in the deployment env; reuse the #62 instance-role Bedrock policy (`bedrock:InvokeModel[WithResponseStream]` on the `us.anthropic.claude-opus-4-8` inference-profile ARN + per-region foundation-model ARNs in us-east-1/2 + us-west-2). Caveat: long-lived keys → rotation reminder + keep the policy tightly scoped. **New IAM user+policy = mandatory GPT-4.1 IAM cross-review + `/sh-security-review` before merge.**
2. **Webhook domain → RAW `*.langgraph.app`** (the "easier" path — no CNAME/DNS/TLS). Point GitHub/Slack/Linear webhook URLs straight at the deployment. Dashboard gets a **Vercel custom domain** (e.g. `openswe.seahaven.com`) since that's trivial on Vercel.
3. **Prod approval gate → GATED PROMOTE-TO-MAIN WORKFLOW.** Re-home `promote-dev-to-prod.yml`: gate the promote job on the **`prod` GitHub Environment reviewer (`amoussa1229`)**; on approval it FF-pushes `dev → main`; the push triggers the managed prod build. Preserves today's manual-approval semantics — approval is on the promote job, the merge to `main` is the deploy trigger.
4. **Cost/trace monitoring → YES, in LangSmith** (usage/trace budget + alert in the LangSmith workspace).
5. **Dev/prod home → CONFIRMED:** prod = LangChain managed (LangGraph Cloud) + Vercel UI; dev = managed **free** Dev tier (same runtime as prod).
6. ~~#62 merge status~~ — RESOLVED (merged + deployed).
## 10b. Execution approach — 2-PR SPLIT (Adam's call, 2026-06-29, post-plan-review)
The plan-review BLOCKed a literal single PR and recommended a 3-PR minimum (isolate cache fix #5 + the promotion-gate workflow). Adam chose a **2-PR split** — isolating the prod-deploy-gate change (highest risk to prod), bundling the rest:
- **PR1 — promotion-gate workflow + ruleset.** Re-home `promote-dev-to-prod.yml` to the gated promote-to-main (decision 3), and lock `main` so ONLY the promote App can push (BLOCK 3). Isolated because it changes *how prod deploys* — must be independently reviewable/revertible.
- **PR2 — the 6 code fixes (§5) + `ui/vercel.json` prod repoint + README/INSTALLATION.** (Accepted deviation from the reviewer's 3-PR rec: the multi-replica cache fix #5 stays bundled in PR2 rather than isolated — Adam's call.)
- **NOT in either PR (rollback safety):** AWS **infra-code deletion** (`deploy/ami`, `deploy/seahaven`, `infra/` self-host stack) + `cdk destroy` stay a **POST-SOAK cleanup (Phase E)**.
- **Not PR content (operational, sequenced around the merges):** LangGraph Cloud prod deployment + env, Vercel prod + custom domain, webhook/OAuth repoint, Bedrock IAM user, LangSmith budget, Confluence/memory.
## 10c. Plan-review resolution (GPT-4.1 cross_reviewer, 2026-06-29) — verdict REQUEST CHANGES, all BLOCKs addressed
- **BLOCK 1 (single PR)** → addressed via the **2-PR split** above (conscious deviation: cache fix #5 not isolated — Adam accepted).
- **BLOCK 2 (raw LangGraph API public?)** → **EMPIRICALLY RESOLVED.** Unauthenticated probes of the dev deployment returned **403 "Missing authentication headers"** for `/threads/search`, `/assistants/search`, `/store/items`; `/ok`=200, `/dashboard/api/me`=401. The platform gates the raw control-plane API. **Keep as a pre-prod gate (Phase D4):** re-probe the PROD deployment URL before cutover.
- **BLOCK 3 (prod gate enforcement)** → PR1 must **lock `main`** so only the promote App can push (verify ruleset 18238334: no direct-push path, PR-required, promote App is the sole FF bypass). Any non-gated push to `main` would auto-deploy prod.
- **BLOCK 4 (secrets posture)** → add to Phase B/E: **rotate** all secrets after migrating them to the platform env stores and BEFORE deleting from AWS; **document who can read/write** the LangGraph Cloud + Vercel env (access audit); delete AWS shells only post-cutover; no dual-homed/stale secrets; record in memory.
- **FIX (rollback integrity)** → Phase E checklist: "no IaC/DNS/config the rollback needs is altered in PR1/PR2 or during cutover."
- **FIX (gates resolved-before-merge)** → make explicit: do NOT merge PR1/PR2 until the GPT-4.1 IAM cross-review (Bedrock IAM user) + `/sh-security-review` (auth/webhook/secrets surface) are **resolved with no critical/high**; Confluence (1540098 + 26116098) + memory updated in the cutover conversation.
- **NITs/QUESTIONs** → carried into §8/§10 TODOs (cost delta, `FF_PROFILE_IMPORTS`, platform feature availability, key-rotation owner, env-store access logging, cache race-review under autoscaling, atomic webhook repoint).
---
## 11. Rollback Story
**Pre-cutover:** the self-host `langgraph up` + RDS plan (RDS+Redis+Docker; `/sh-plan-review`'d to APPROVE-after-revision this session) is the documented **fallback** if managed is rejected before cutover. Its durable design wins (env-scoped RDS physical names; `Credentials.fromGeneratedSecret({secretName:"open-swe-<env>/rds-credentials"})` under the existing instance-role secret prefix → no new IAM; DESTROY-on-rollback RDS-managed secret; derive `DATABASE_URI` in fetch-config) are captured in the project memory.
**Post-cutover (managed is live, AWS still standing during soak):** if managed prod fails, **rollback = repoint webhooks + DNS back** to the AWS stack:
- GitHub App / Slack / Linear webhook URLs → `hooks.seahaven.com`
- `openswe.seahaven.com` DNS → the `seahaven-com` ALB (restore from the E1 export)
- `ui/vercel.json` rewrite → the AWS dashboard origin (or stop using Vercel)
- Do NOT run Phase E teardown until the soak passes — the AWS stack IS the rollback target.
**Revision-level rollback (within managed):** LangGraph Cloud revision rollback (backend) + Vercel instant rollback (UI) cover bad deploys without leaving the platform.
---
## TODOs flagged (couldn't verify in repo)
- ~~#62 merge state~~ — RESOLVED: merged to `dev` + deployed (the draft read a stale local checkout; `origin/dev` and the deployment `git_ref_sha a4ed19ba` confirm Bedrock+Fireworks).
- **Current AWS run-rate** for the cost delta (§8) — not derivable from repo.
- **LangGraph Cloud custom-domain + "manual promotion to production" feature availability** (§10.2, §10.3) — platform features, verify in the LangGraph Cloud console/docs at execution time.
- **`FF_PROFILE_IMPORTS`** exact flag name/usage (§5 fix #6) — referenced from session context; confirm the flag exists in `agent/` before relying on it.
- Exact webhook path on `*.langgraph.app` (whether the custom `http.app` mounts at root so `/webhooks/github` is reachable as-is) — proven reachable on the dev spike; re-verify the exact path on prod.