* Adopt upstream modular webhook skeleton (#1621) Apply the durable-interrupt-dispatch refactor: split the monolithic webapp.py into a thin routing layer plus per-source handlers in webhooks/{github,slack,linear}.py, and add completion.py, dispatch.py, and reconcile.py. Reconcile fork divergence by keeping the Bedrock/ Fireworks cross-provider fallback, the no-agent-attribution prompt policy, the dashboard-handoff re-export, and the Slack channel-info cache. ci_autofix is restored on the new dispatch model in a later commit. Refs: #80 * Port fork webhook security delta onto modular handlers Re-apply the fork's security customizations that #1621 did not carry: Linear webhook replay protection (freshness window on the signed webhookTimestamp), per-repo token-cache binding threaded through the thread token resolvers, the INTERNAL_BOT_LOGINS self-check in the review-finding-reply path, and a user-mapping cache refresh before email resolution on the issue and PR-comment paths (multi-replica staleness). Existing fork security tests pass unchanged. Refs: #80 * Restore CI auto-fix on the modular dispatch model Bring back ci_autofix.py and the ci_monitor graph that #1621 deleted, re-wiring the fork's security-reviewed PR-babysitting onto the new structure: the CI-event, autofix-toggle, and review-feedback handlers move into webhooks/github.py and the github_webhook router re-gains the check_run/check_suite/workflow_run/status routing plus the autofix command and actionable-review branches. Auto-fix runs now dispatch through dispatch_agent_run (durability + completion webhook) while keeping the deliberate batch-while-busy skip-rule via get_thread_active_status. Restore langgraph.json's ci_monitor entry and the fork autofix tests (dispatch mock + import paths re-pointed). Refs: #80 * Reformat and update docs for the modular webhook split Point CLAUDE.md and deploy/MIGRATION.md at the new webhooks/ modules and the dispatch/completion/reconcile contract, and mark the user-mapping cache-refresh fix as applied on the GitHub handlers. Refs: #80 * Restore reject backstop for autofix dispatch A burst of near-simultaneous CI events for one head SHA can slip past the busy-check before the dedupe SHA is recorded, so dispatch the autofix path with multitask_strategy=reject (dev's prior platform default) to drop duplicate concurrent creates instead of letting them interrupt each other. Also make the completion failure-reply dedup claim-then-post and drop the unreachable interrupted branch. --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
45 KiB
Open SWE — Migration Plan: Self-Hosted AWS → Managed LangGraph Cloud + Vercel
Repo: Sea-Haven-Industries/open-swe (private) · AWS: 328440206208 / us-east-1
Author: Adam Moussa · Date: 2026-06-29 (final topology added 2026-06-30) · Status: EXECUTED — managed cutover live; §§3–11 below are the original (now-historical) phased plan, superseded by §1a for all current-state facts (URLs, project layout, env).
READ §1a FIRST. The phased plan (§§2–11) and the Phase A/C spike notes capture how we got here and still hold for rationale, cost, and rollback. But the spike-era specifics they cite — the single
open-swe-dashboardVercel project, theopen-swe-dev-hosted-…/open-swe-v3-…deployment URLs, theui/vercel.jsonsame-origin rewrite, the single GitHub App — are stale. §1a is the authoritative final topology and wins on every conflict.
1. Executive Summary
Decision (settled — not re-litigated here)
Move the Open SWE deployment off the bespoke self-hosted AWS stack (stock langgraph dev, in-memory store, no durability — issue #9) onto the managed runtime:
- Backend → LangGraph Cloud (a.k.a. "LangSmith Deployment"): git-connected, auto-builds a revision on push, zero-downtime, revision rollback, durable Postgres-backed store/checkpointer.
- UI → Vercel:
ui/SPA, atomic deploys, instant rollback, per-PR previews (net-new capability).
Why
- Dissolves the licensing blocker. Self-hosted
langgraph upon the LangSmith Plus key is "Self-Hosted Lite" (1M node-exec/yr cap, Elastic License 2.0, production legally ambiguous). Managed IS the licensed product with no node cap — runs billed flat at $0.005, nodes not charged. - Deletes the entire durable-runtime build (RDS/Redis/Docker/AMI) and most of the bespoke AWS CD (S3 release pipeline, SSM deploy docs, packer, CDK box/ALB stacks).
- It's the upstream-canonical deployment and the repo is already wired for it:
ui/vercel.jsonhas the same-origin/dashboard/api/*rewrite;langgraph.jsonis cloud-format (6 graphs +http.app).
Cost
≈ $160/mo incremental for one always-on prod deployment:
- $0.0036/min prod uptime ≈ $155/mo (always-on)
-
- $0.005/run
-
- traces pay-as-you-go above 10k/mo
- The $39/mo Plus seat is already paid (not incremental)
- Dev deployment is FREE (1 included on Plus, preemptible)
A self-hosted langgraph up + RDS plan (already /sh-plan-review'd to APPROVE-after-revision this session) is the documented FALLBACK if managed is ever rejected (see §11).
1a. Final, verified topology (AUTHORITATIVE — supersedes spike-era values)
This is the live managed deployment as of 2026-06-30. Where any later section disagrees (old URLs, a single Vercel project, a single GitHub App, the ui/vercel.json rewrite), this section wins.
Backend — managed LangGraph Cloud (two deployments, one LangSmith workspace)
Both deployments live in the same LangSmith workspace; the same workspace API key authenticates both (including the Store API — so per-deployment store writes use that one key with the per-deployment URL).
| Deployment | URL | Git connection |
|---|---|---|
| dev | https://open-swe-dev-fb737aa219605c8bbdb30ecbb33f30c0.us.langgraph.app |
branch dev |
| prod | https://open-swe-prod-d6c7bb63aaa651b6a1d92f9492b1d983.us.langgraph.app |
branch main (auto-deploys on push to main) |
The dev deployment was renamed
open-swe-dev, deleted, and recreated — which minted the new URL hash above. The spike-eraopen-swe-v3-…/open-swe-dev-hosted-…URLs are dead/superseded. Deleting + recreating a deployment is the one operation that changes the URL hash (otherwise stable across revisions) — when it happens, update every reference (Vercel env, GitHub App webhooks, OAuth callbacks, docs).
UI — Vercel (ONE project, two environments)
One Vercel project open-swe-prod (team sea-haven, id prj_OOh6yjXMp4ah3Ws3Y7XRQxjmMmQU). The old separate open-swe-dashboard project was DELETED.
| Vercel environment | Branch | Backend | Custom domain |
|---|---|---|---|
| production | main |
prod deployment URL | openswe.seahaven.com |
custom dev (id env_SMI23PULAJXk0GhwE0HLVhp5J3ZS) |
dev |
dev deployment URL | openswe-dev.seahaven.com |
- A per-environment env var
LANGGRAPH_BACKEND_URL(prod env = prod URL, dev env = dev URL) drives the/dashboard/api/*proxy. - Project settings:
framework=null,outputDirectorycleared, root directoryui. - Proxy mechanism (current, after PR #76): Nitro
routeRulesinui/vite.config.tsreadprocess.env.LANGGRAPH_BACKEND_URLand Nitro's Vercel preset compiles them into.vercel/output/config.json(Build Output API) at build time — a CDN-level proxy (not redirect, so theosw_sessioncookie stays first-party). PR #75's hand-rolledui/scripts/build-vercel-output.mjswas the broken first attempt and is gone (ui/scripts/no longer exists). Never add a manual script thatrms.vercel/output— Nitro's Vercel preset auto-emits it.
DNS — Route 53 zone seahaven.com (Z06652411XKH89KTZD3XA)
openswe.seahaven.com→ CNAMEcname.vercel-dns.com(prod env)openswe-dev.seahaven.com→ CNAME to Vercel (dev env)
GitHub Apps — TWO (dev/prod isolated; each its own webhook URL)
| App | app_id | install | client_id | org | members scope | repos |
|---|---|---|---|---|---|---|
prod seahaven-openswe |
4146115 |
142615168 |
Iv23lil96pKQNNDUn5yp |
Sea-Haven-Industries |
members:write | all |
dev seahaven-openswe-dev |
4162963 |
143023302 |
Iv23licQwJvGAPJj1HJe |
seahaven-open-swe-dev |
members:read | all |
- Promotion App
seahaven-promotion(actor4170147) is the sole non-admin fast-forward-push bypass on themainruleset18238334— its FF-push ofdev → mainis what triggers the managed prod build.
Env per deployment (set in LangGraph Cloud config + Vercel env — NOT Secrets Manager)
| Var | dev | prod |
|---|---|---|
LANGGRAPH_URL |
own (dev) deployment URL | own (prod) deployment URL |
DASHBOARD_BASE_URL / DASHBOARD_API_BASE_URL |
https://openswe-dev.seahaven.com |
https://openswe.seahaven.com |
VITE_DASHBOARD_API_BASE_URL |
empty (same-origin via Vercel proxy) | empty |
ALLOWED_GITHUB_ORGS |
dev org (seahaven-open-swe-dev) |
Sea-Haven-Industries |
CONFIGURED_ADMINS |
amoussa1229,adam@seahavenind.com |
amoussa1229,adam@seahavenind.com |
DASHBOARD_BASE_URL / DASHBOARD_API_BASE_URL must include https:// (see gotcha 3). Secret values are still sourced from open-swe-{dev,prod}/* Secrets Manager + SSM (the remaining source of truth) and set into the LangGraph Cloud + Vercel env stores — the accepted secrets-and-config.md deviation.
Bedrock IAM (PR #74, still OPEN)
Two IAM users open-swe-dev-bedrock + open-swe-prod-bedrock, each attached to customer-managed policy open-swe-bedrock-invoke (least-privilege bedrock:InvokeModel[WithResponseStream] on the us.anthropic.claude-opus-4-8 inference-profile ARN + its 3 routed foundation-model ARNs in us-east-1/us-east-2/us-west-2). Default model bedrock_converse:us.anthropic.claude-opus-4-8 + 3 Fireworks models. Static access keys live only in the deployment env (dev key → dev, prod key → prod).
User store — per-deployment
Each managed deployment has its own Store. The GitHub→email mapping amoussa1229 → adam@seahavenind.com (namespace ["user_mappings"], key = lowercased login, record {github_login, work_email, status:"active", source, created_at, updated_at}) was written to both the dev and prod stores directly. New users need a mapping per-deployment (write each store directly, or use the dashboard admin User-mappings UI — the work_email field was added by PR #65 fix #4).
AWS decommission (PR #64)
Self-host CDK stacks destroyed. Residual: CDKToolkit (shared, preserved); ~50 RETAIN'd Secrets Manager shells + 3 S3 asset buckets (pending cleanup); AWS Bedrock (live dependency, kept).
Operational gotchas (hard-won — carry these into any runbook)
- Per-deployment store → seed user mappings per-deployment. A missing mapping makes
process_github_issuesilently early-return ("No email mapping … skipping"): the webhook returns 200/accepted but produces no reaction and no run. Seed dev and prod. - Org-login gate uses the App installation token, so the App must be org-installed with Members:read. OAuth working ≠ membership check working — they use separate creds (CLIENT_ID/SECRET for OAuth vs APP_ID/INSTALLATION_ID/PRIVATE_KEY for the install token). A mangled multi-line
GITHUB_APP_PRIVATE_KEYbreaks the install token (and thus the gate) while OAuth still works. DASHBOARD_API_BASE_URLmust behttps://— anhttp://value makes GitHub reject the OAuth callback with "redirect_uri not associated."osw_oauth_statecookie is host-only — start login on the same host asDASHBOARD_API_BASE_URL, or you get "oauth state mismatch."- Webhooks go DIRECT to the langgraph URL (
/webhooks/*). Vercel only proxies/dashboard/api/*. The app is same-origin only (no CORS). - On Vercel CI, Nitro's Vercel preset auto-emits
.vercel/output— drive the proxy via NitrorouteRulesfromLANGGRAPH_BACKEND_URL; never a manual script thatrms.vercel/output. - Deleting + recreating a LangGraph deployment mints a NEW URL hash (otherwise stable across revisions) — update every reference (Vercel env, webhooks, OAuth, docs).
2. Architecture: Before → After
Before (self-hosted AWS — LIVE as of 2026-06-29)
GitHub/Slack/Linear ──webhook──▶ hooks.seahaven.com ─┐
Browser (dashboard) ─────────────▶ openswe.seahaven.com ─┤
▼
shared seahaven-com ALB (:443, host+path rules)
▼
EC2 (prod i-08a729e50779c4b07 t4g.large ARM64, private subnet)
nginx :80 ──proxy /dashboard/api + /webhooks──▶ langgraph dev :2024 (loopback)
in-memory store (reseeded by seed_store.sh ExecStartPost)
▼
Config: Secrets Manager open-swe-prod/* + SSM /open-swe-prod/*
Release: GitHub Actions → S3 open-swe-prod-assets/releases/* → SSM doc deploy.sh
IaC: CDK OpenSweIamStack + OpenSweDevStack + OpenSweProdStack
Sandbox: LangSmith cloud (DEFAULT_SANDBOX_SNAPSHOT_ID + GitHub proxy)
After (managed — FINAL, see §1a for exact values)
┌──────────────── DEV lane ────────────────┐ ┌──────────────── PROD lane ───────────────┐
GitHub(dev org)/Slack/Linear │ webhook → open-swe-dev-….us.langgraph.app │ │ webhook → open-swe-prod-….us.langgraph.app│ GitHub(SHI org)/Slack/Linear
App seahaven-openswe-dev ───┘ (DIRECT to langgraph URL, /webhooks/*) │ │ (DIRECT to langgraph URL, /webhooks/*) └─── App seahaven-openswe
▼ ▼
Browser ▶ openswe-dev.seahaven.com ─┐ ┌─▶ openswe.seahaven.com ◀ Browser
│ ONE Vercel project `open-swe-prod` (team sea-haven)
│ ├─ env `dev` (branch dev) → proxies /dashboard/api/* → dev langgraph URL
│ └─ env production (branch main) → proxies /dashboard/api/* → prod langgraph URL
└─ proxy compiled by Nitro routeRules from per-env LANGGRAPH_BACKEND_URL (PR #76)
▼
TWO LangGraph Cloud deployments (same LangSmith workspace; one workspace key auths both incl. Store)
├─ dev ← branch `dev` · prod ← branch `main` (push-to-main auto-deploys prod)
├─ each serves the graphs + the custom http.app (agent.webapp:app = webhooks + dashboard API + OAuth)
├─ each has its OWN durable Postgres store + checkpointer (issue #9 SOLVED) — user mappings per-deployment
└─ env/secrets in the Deployment + Vercel config (NOT Secrets Manager) — Adam accepted deviation
▼
Sandbox: LangSmith cloud (UNCHANGED — DEFAULT_SANDBOX_SNAPSHOT_ID + GitHub-App proxy)
Bedrock: IAM users open-swe-{dev,prod}-bedrock + policy open-swe-bedrock-invoke (static keys in deploy env)
What changes shape: runtime host (EC2 → managed PaaS), durability (in-memory → managed Postgres, one store per deployment), CD (bespoke S3/SSM/packer → git-connected auto-build), config home (Secrets Manager/SSM → Deployment+Vercel env), ingress topology (shared ALB + one App → two dev/prod-isolated GitHub Apps each hitting its own *.langgraph.app directly), UI proxy (ui/vercel.json rewrite → Nitro routeRules). What stays: the LangSmith sandbox plane, CI (lint/format/unit/Playwright), the app code itself, and Secrets Manager/SSM as the secret-value source of truth.
3. Phased Plan
Phase A — Dev spike (MOSTLY DONE)
Goal: prove managed serves our custom app + durability, at 0, before committing prod .
Proven this session (spike-era specifics — superseded by §1a; URLs/project below are DEAD):
- ✅ Dev backend deployed to LangGraph Cloud, connected to branch
dev:→ final dev URL in §1a (https://open-swe-dev-hosted-e76c2b0e8a7955fe8ad3110a7a54e5d0.us.langgraph.appopen-swe-dev-fb737aa…; the deployment was later deleted + recreated) (Bedrock + a Fireworks key set; AWS creds for Bedrock deferred — see §10 open decision). - ✅ UI deployed to Vercel — team
sea-haven,project(DELETED); now the single projectopen-swe-dashboard,https://open-swe-dashboard.vercel.appopen-swe-prodwith dev/prod environments (§1a); proxy is NitrorouteRules, not theui/vercel.jsonrewrite. - ✅ Managed serves the custom
http.app(dashboard API + webhooks) — no platform auth gate in front of our routes (webhook returns 401 sig-enforced, so signatures still govern). - ✅ Vercel same-origin rewrite → backend works.
- ✅ GitHub OAuth dashboard login end-to-end.
- ✅ DURABILITY — team-default model survived a revision redeploy (this is issue #9's core acceptance goal, proven on managed).
Remaining Phase A items (the GitHub @openswe run-trigger loop):
- A1. Get an
@opensweGitHub comment to dispatch a run end-to-end on dev. Blocked by the user-mapping + cache gotchas (see fixes #4 and #5 in §5). - A2. Create the owner user-mapping in the managed Store (namespace
["user_mappings"], key = lowercased loginamoussa1229, value{github_login, work_email}) — the Admin UI can't setwork_emailyet (fix #4). - A3. Verify a freshly-added mapping is seen without a redeploy across managed's multi-replica autoscaling (fix #5).
- A4. Confirm
LANGGRAPH_URLis set to the deployment URL on the dev deployment solanggraph_client()calls resolve (fix #1 — config now, code later).
Phase B — Harden (the code fixes + env codification)
Land the 6 code fixes (§5), codify env/config, and resolve the Bedrock-auth decision before standing up prod.
- B1. Land code fixes #1–#6 (§5) as a PR into
dev(or split into focused PRs). Re-runmake lint+make test. - B2. Codify the env contract for managed. Update
ui/vercel.jsonrewrite to the prod deployment URL (Phase C) but keep dev pointing at dev. Add a documented LangGraph-Cloud env list to the repo (see §4) — NOT secrets, just the variable inventory + which are excluded. - B3. Resolve the Bedrock-auth decision (§10 open decision): static AWS keys for a Bedrock-scoped IAM user in the deployment env, OR run the agent on Fireworks. This intersects the #62 model work.
- B4.
Merge #62 first— DONE / N/A. #62 (Bedrock + Fireworks) IS merged todevand is what the dev deployment runs — verified onorigin/dev(options.py→DEFAULT_MODEL_ID = "bedrock_converse:us.anthropic.claude-opus-4-8") and confirmed by the live deployment'sgit_ref_sha: a4ed19ba…(= the #62 merge commit). The contrary note came from a STALE local checkout; no merge action needed. - B5. Audit ALL in-process caches for the single-process → multi-replica assumption (generalization of fix #5; see §5).
- B6. Measure + reduce custom-app import time (fix #6) so the deployment isn't flagged unhealthy / slow to scale.
Phase C — Prod deployment
- C1. Create a prod LangGraph Cloud deployment tracking branch
main(the durable autoscaled 1→10 tier, not the free Dev tier). Record its*.langgraph.appURL. - C2.
Create the prod Vercel project/target (or promote the existing— DONE differently: one projectopen-swe-dashboardto production)open-swe-prodwith a production env (main) + a customdevenv (dev), each with its ownLANGGRAPH_BACKEND_URL. See §1a. - C3. Set the prod env triad (§4) on the prod deployment + Vercel:
LANGGRAPH_URL= the prod*.langgraph.appURLDASHBOARD_BASE_URL+DASHBOARD_API_BASE_URL= the prod Vercel origin (withhttps://scheme — fix #2)
- C4. Establish the prod approval gate (§6): git-connected auto-deploy needs an explicit gate. Preserve the prod manual-approval that exists today (the GitHub
prodEnvironment reviewer). Mechanism = protectedmainruleset + the platform's "require manual promotion to production" if available (§10 open decision). - C5. Repoint webhooks + OAuth URLs to prod:
- GitHub App
seahaven-openswewebhook URL → prod backend webhook URL (*.langgraph.app/webhooks/githuborhooks.seahaven.com— §10 custom-domain decision) - Slack Event Subscriptions + Interactivity Request URLs → prod
- Linear webhook URL → prod (if/when wired)
- GitHub App OAuth callback →
<DASHBOARD_API_BASE_URL>/dashboard/api/auth/callback(the Vercel prod origin; fixes #2, #3)
- GitHub App
- C6. Custom-domain decision (§10): dashboard custom domain solved by Vercel; webhooks
hooks.seahaven.comeither re-point to*.langgraph.appdirectly or front via CNAME. Managed issues*.langgraph.appURLs ONLY (no documented custom-domain support on the backend).
Phase D — Cutover + verify
- D1. Run the issue #9 acceptance criteria on PROD: durable store survives a revision redeploy; team defaults / user mappings persist; a paused/in-flight run survives a redeploy (the durability win self-host never had).
- D2. End-to-end prod smoke:
@opensweGitHub comment → run dispatch → sandbox → draft PR → reply in source channel. - D3. Dashboard prod smoke: GitHub OAuth login on the stable Vercel alias (fix #3); admin pages reachable for
CONFIGURED_ADMINS. - D4. Webhook sig-enforcement smoke: unsigned POST to each
/webhooks/*returns 401; signed returns 200. - D5. Soak for an agreed window (recommend 3–7 days) with the AWS prod stack still standing as the rollback target (§11) before any teardown.
Phase E — AWS decommission (AFTER soak)
Only after managed prod is proven + soaked. See §7 for the precise retire-vs-keep list. Phased teardown:
- E1. Export the live ALB listener-rule / target-group config for
open-swe-prodandopen-swe-devto JSON (the export is the rollback source — these rules were partly hand-built). - E2. Disable the bespoke CD (build-artifacts / cd-infra workflows) so nothing re-deploys the EC2 boxes.
- E3.
cdk destroy OpenSweProdStackthenOpenSweDevStack(mind the RETAIN secret shells — they survive and hold the globalopen-swe-<env>/*names; force-delete only EMPTY shells once env is fully migrated; keep populated ones transitionally as the env source — §7). - E4. Remove ALB rules/target groups, S3 assets buckets, SSM deploy docs, and the per-env CDK bootstrap qualifier
oswedev(CDKToolkit-oswedev) — once nothing references them. - E5. Decide DNS: repoint
openswe.seahaven.com→ Vercel; resolvehooks.seahaven.com→*.langgraph.app(or CNAME). - E6. Leave
OpenSweIamStackuntil last — the promotion App + any retained roles may still be referenced.
Phase F — Docs / memory
- F1. Rework Confluence "AWS Architecture Map" (page 1540098) + the open-swe child page (26116098): the RDS/EC2/ALB subgraph is replaced by an external-services view (LangGraph Cloud + Vercel + the retained GitHub App + LangSmith sandbox).
- F2. Update
README.md+INSTALLATION.md(§10 is already canonical-correct; align Sea Haven specifics). - F3. Update project memories
project_open_swe_migration.md+reference_open_swe_deployment.mdto mark the managed cutover and the AWS teardown.
4. Env / Config Reference
The prod triad (per INSTALLATION.md §10, lines 630–658) — see §1a for the exact dev/prod values
| Var | Value (per deployment/env) | Notes |
|---|---|---|
LANGGRAPH_URL |
the own deployment URL (https://...langgraph.app) |
NOT localhost. Drives thread_ops.langgraph_url() (fix #1). dev→dev URL, prod→prod URL. |
DASHBOARD_BASE_URL |
the own dashboard origin, with https:// (https://openswe-dev.seahaven.com / https://openswe.seahaven.com) |
|
DASHBOARD_API_BASE_URL |
same as above, with https:// scheme |
scheme required or OAuth redirect_uri is schemeless and GitHub rejects (fix #2) |
VITE_DASHBOARD_API_BASE_URL |
empty | same-origin mode; UI calls relative /dashboard/api/*, the Vercel Nitro proxy rewrites to backend |
LANGGRAPH_BACKEND_URL |
Vercel env var, per Vercel environment (dev env = dev URL, prod env = prod URL) | drives the Nitro routeRules /dashboard/api/* proxy at build time (PR #76) — required on Vercel builds |
ALLOWED_GITHUB_ORGS |
Sea-Haven-Industries (prod) / seahaven-open-swe-dev (dev) |
org-login gate; checked via the App installation token (gotcha 2) |
GitHub App dashboard OAuth callback = <DASHBOARD_API_BASE_URL>/dashboard/api/auth/callback (the own dashboard origin — prod App on openswe.seahaven.com, dev App on openswe-dev.seahaven.com).
UI proxy mechanism (final, PR #76): the /dashboard/api/* proxy is Nitro routeRules in ui/vite.config.ts reading process.env.LANGGRAPH_BACKEND_URL, compiled by Nitro's Vercel preset into .vercel/output/config.json. This supersedes the spike-era ui/vercel.json same-origin rewrite and PR #75's hand-rolled ui/scripts/build-vercel-output.mjs (deleted). ui/vercel.json now only carries framework:null + buildCommand: bun run build.
Where secrets/env live now
LangGraph Cloud Deployment config + Vercel env — NOT AWS Secrets Manager. This deviates from the Sea Haven secrets-and-config.md "Secrets Manager for all sensitive" handbook rule — Adam ACCEPTED this deviation (managed has no instance role / no fetch-config boot hook; the platform's own secret store is the mechanism).
Sourcing env from the existing AWS stack (transitional)
Pull from the existing open-swe-dev Secrets Manager + SSM via a documented CLI dump:
aws secretsmanager list-secrets --filters Key=name,Values=open-swe-dev/then per-nameget-secret-value— NOTbatch-get-secret-value(its pagination silently drops values past page 1; this bit the box twice — see memory).- SSM:
aws ssm get-parameters-by-path --path /open-swe-dev/.
EXCLUDE when copying to managed
| Category | Vars |
|---|---|
| Box-/self-host-specific | LANGGRAPH_URL (set fresh to deployment URL), LANGGRAPH_URL_PROD, LANGSMITH_ENDPOINT, LANGSMITH_ENDPOINT_PROD, LANGSMITH_URL_PROD, LANGSMITH_TENANT_ID_PROD, LANGSMITH_HOST_API_URL, LANGCHAIN_REVISION_ID |
| Dropped providers (PR #62) | OPENAI_API_KEY, GOOGLE_API_KEY, GROQ_API_KEY (these 6 Secrets Manager shells deleted 2026-06-29, final purge 2026-07-06) |
| Unused sandbox providers | DAYTONA_API_KEY, RUNLOOP_API_KEY |
| Eval-only | JUDGE_ANTHROPIC_API_KEY (kept on AWS for evals/reviewer/judge.py, not needed in the runtime deployment) |
KEY mappings on managed
LANGSMITH_API_KEY= theLANGSMITH_API_KEY_PRODvalue (andLANGCHAIN_API_KEY= same).- The full secret inventory is
SECRET_VARSininfra/lib/constructs/config-store.ts:46(28 shells) — use it as the checklist of what to carry, minus the EXCLUDE rows above.
Models (post-#62) + Bedrock auth
Post-#62 SUPPORTED_MODELS = bedrock_converse:us.anthropic.claude-opus-4-8 (DEFAULT) + 3 Fireworks models; no direct anthropic: option.
✅ #62 is merged to dev and deployed — verified on origin/dev (options.py → bedrock_converse default) and the live deployment git_ref_sha a4ed19ba. (The pre-#62 values appear only on a stale LOCAL checkout — ignore.)
Bedrock on managed has NO EC2 instance role (the #62 IAM design used the EC2 instance role as the Bedrock principal — that breaks on managed). Options (OPEN DECISION, §10):
- (a) Static
AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY/AWS_REGIONfor a Bedrock-scoped IAM user in the deployment env, or - (b) Run the agent on Fireworks and avoid Bedrock entirely on managed.
Bedrock IAM users (CREATED — resolves §10.1, discharges the §9 IAM gates)
Option (a) was chosen and executed live (2026-06-30, account 328440206208 / us-east-1). This is click-ops IAM — there is no remaining open-swe AWS IaC after the decommission (#64), so these are created with the CLI, not CDK.
- Customer-managed policy
open-swe-bedrock-invoke(arn:aws:iam::328440206208:policy/open-swe-bedrock-invoke). Least-privilege: actionsbedrock:InvokeModel+bedrock:InvokeModelWithResponseStreamonly, scoped to exactly theus.anthropic.claude-opus-4-8inference-profile ARN + its three routed foundation-model ARNs (us-east-1, us-east-2, us-west-2). No wildcards, no other models. (Supersedes the §10.1 plan to reuse the #62 instance-role policy — a fresh standalone policy was minted instead.) - Two IAM users, each attached to that policy:
open-swe-dev-bedrockandopen-swe-prod-bedrock. Taggedproject=open-swe,managed-by=cli-migration,purpose=bedrock-invoke. - Static access keys are minted separately by the owner (
aws iam create-access-key) — secret keys live only in the deployment env stores, never in this repo or memory. dev key → dev env; prod key → LangGraph Cloud prod config (AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY/AWS_REGION=us-east-1). - Deviation rationale: static long-lived keys are a deliberate departure from the Sea Haven OIDC norm because managed LangGraph Cloud cannot assume an AWS role. Mitigated by the tight least-privilege policy above.
- ✅ Both mandatory gates ran and PASSED:
- GPT-4.1 IAM cross-review (confirmed hit
gpt-4.1-2025-04-14) — least-privilege confirmed. /sh-security-review— 0 critical/high; two accepted mediums (the static-key deviation + no per-principal budget cap).
- GPT-4.1 IAM cross-review (confirmed hit
- Recommended follow-ups: shortest viable key-rotation cadence with recorded creation dates; an AWS Budgets / CloudWatch anomaly alarm on per-principal Bedrock
InvokeModelvolume; confirm CloudTrail captures these users.
5. Required Code Fixes (the 6 gotchas)
These are self-host → managed assumption breaks discovered on the spike. Each gets a config workaround now and a code fix to land in Phase B.
Fix #1 — langgraph_url() localhost fallback
File: agent/utils/thread_ops.py:38-41
def langgraph_url() -> str:
return os.environ.get("LANGGRAPH_URL") or os.environ.get(
"LANGGRAPH_URL_PROD", "http://localhost:2024"
)
On managed, every langgraph_client() call (thread sidebar, run creation, the same pattern in agent/dashboard/user_mappings.py:43 get_client() with no URL) ConnectErrors unless LANGGRAPH_URL is set to the deployment URL.
- Config fix (now): set
LANGGRAPH_URL= the deployment URL on every deployment (Phase A4 / C3). - Code fix (Phase B): default to the in-process / deployment URL rather than
http://localhost:2024when running inside a managed deployment (e.g. honor the platform's own URL env, or fail loud instead of silently hitting localhost).
Fix #2 — OAuth DASHBOARD_API_BASE_URL must include scheme
Where: OAuth redirect construction (dashboard oauth.py / routes.py), driven by DASHBOARD_API_BASE_URL.
A schemeless value yields a schemeless redirect_uri → GitHub rejects with "redirect_uri not associated."
- Fix: always set
DASHBOARD_API_BASE_URLwithhttps://(Phase C3). Code hardening (Phase B): assert/normalize a scheme at startup and fail loud if missing.
Fix #3 — OAuth state cookie is host-only
Cookie: osw_oauth_state (host-only). Login must START on the SAME host as DASHBOARD_API_BASE_URL — the stable Vercel alias, never the immutable per-deploy URL — or you get "oauth state mismatch."
- Fix: pin
DASHBOARD_API_BASE_URL+ the login entry point to the stable alias / custom domain (Phase C3, D3). Document this so a future per-deploy preview URL isn't used for login.
Fix #4 — Admin "User mappings" UI can't set work_email
Files: the Admin → User mappings UI (ui/) + agent/dashboard/user_mappings.py (upsert_mapping at line 257 takes work_email but the UI form doesn't expose it).
Mappings can't be fully created from the dashboard → must write the Store directly: namespace ["user_mappings"], key = lowercased login, value {github_login, work_email} (matching _index_record at user_mappings.py:73-84, which keys _by_login on login.lower()).
- Fix (Phase B): add the
work_emailfield to the Admin mappings form so mappings are fully creatable from the UI.
Fix #5 — GitHub webhook path doesn't refresh the user-mapping cache (multi-replica break)
Files: agent/webhooks/github.py (GitHub handlers: process_github_pr_comment, process_github_issue) vs agent/webhooks/slack.py (process_slack_mention). (Pre-modular-refactor these all lived in agent/webapp.py.)
The Slack path refreshes before lookup:
# agent/webhooks/slack.py (process_slack_mention)
await webapp.refresh_user_mapping_cache()
...
The GitHub path historically did not — it called email = await email_for_login(github_login) cold. The cache (user_mappings.py _ensure_cache_loaded) is one-shot per process (_cache_loaded flag). On self-host single-process this was fine; on managed's multi-replica autoscaling, a freshly-added mapping isn't seen by a replica whose cache loaded earlier — until restart.
- Fix (applied): the GitHub issue and PR-comment handlers now call
webapp.refresh_user_mapping_cache()before email resolution, mirroring the Slack path. - Generalize (Phase B5): audit ALL in-process caches for the single-process → multi-replica assumption —
SANDBOX_BACKENDSdict (agent/utils/sandbox_state.py),_by_login/_by_email/_by_slack_id(user_mappings.py). Sandbox affinity is already thread-keyed + persisted in thread metadata (sandbox_id), so it's the cache state that needs the multi-replica review. (The legacy in-process thread lock has been removed: webhook triggers now serialize throughdispatch_agent_run'smultitask_strategy="interrupt"instead.)
Fix #6 — Slow custom-app import (~8s startup)
Symptom: "exceeded expected startup time" → risks the deployment being marked unhealthy / slow to scale out.
- Fix (Phase B6): lazy imports / reduce import-time work in
agent/webapp.py(+agent/webhooks/*.py) and the graph factories. Profile withFF_PROFILE_IMPORTS(the import-profiling flag) to find the heavy modules.
6. CD / Ops Changes
Retires (bespoke pipeline)
- GitHub Actions build → S3 releases → SSM doc →
deploy.sh→ health-gate (.github/scripts/,deploy/ami/deploy.sh) roll-box.sh/publish-and-deploy.sh/rollback.sh- AMI baking (packer
deploy/ami/open-swe-base.pkr.hcl) +cdk.context.jsonAMI pin cd-infra.ymlCDK deploys + OIDC bootstrap qualifiersfetch-config.sh/seed_store.sh(boot-time config materialization + Store reseed)
Replaced by git-connected PaaS
- LangGraph Cloud auto-builds a revision on push (first-party zero-downtime + revision rollback).
- Vercel builds / atomic-deploys / instant-rollback + per-PR previews (net-new capability the AWS stack never had).
Stays
- CI (lint / format / unit / Playwright E2E) — still matters and still gates merges.
make lint,make test.
Gate re-homing
- The dev → main promotion gate (
check-dev-green.sh+ protected-mainruleset 18238334) re-homes: the dev deployment tracksdev, the prod deployment tracksmain, so the gate governs exactly what reaches prod. - Today's promotion machinery:
promote-dev-to-prod.yml+ theseahaven-promotionGitHub App (app_id 4170147, in ruleset 18238334 bypass_actors) FF-pushesdev → main. Under managed, a push tomainis what triggers the prod build — so the promotion gate IS the prod deploy gate.
Preserve the PROD manual-approval gate (LOAD-BEARING)
Today the prod GitHub Environment (required reviewer amoussa1229) is the manual approval. Git-connected auto-deploy removes the CD job that consulted that Environment, so the approval must be re-established explicitly:
- Mechanism options (OPEN DECISION §10): protected
main(only the promotion App can FF-push, and that push is itself the gate) AND/OR the platform's "require manual promotion to production" if LangGraph Cloud exposes it. - Norm shift: deploy-then-merge → merge-to-deploy. Lean on the dev deployment + Vercel previews to verify before promoting
dev → main. (This inverts the Sea Haven handbook deploy-then-merge default — call it out in the README + handbook note.)
7. AWS Decommission — Retire vs Keep
RETIRE (becomes vestigial once managed prod is live + soaked)
| Item | Path / resource |
|---|---|
| AMI bake + cloud-init | deploy/ami/ (open-swe-base.pkr.hcl, provision.sh, user-data.sh, deploy.sh, templates/) |
| Self-host boot/config scripts | deploy/seahaven/fetch-config.sh, seed_store.sh, put-config.sh, nginx/openswe.conf, systemd/open-swe.service, aegra/ (already deferred) |
| CDK box/ALB/AMI stacks | infra/lib/open-swe-stack.ts, constructs/app-service.ts, assets-bucket.ts, ami-cache.ts, instance-role.ts, github-deploy-roles.ts |
| EC2 instances | prod i-08a729e50779c4b07 (t4g.large), dev i-0a3bb8e0ddd36c29b |
| S3 release buckets | open-swe-dev-assets, open-swe-prod-assets |
| SSM deploy docs | open-swe-dev-deploy, open-swe-prod-deploy |
| ALB plumbing | open-swe-prod-tg / open-swe-dev-tg target groups + the seahaven-com ALB host/path listener rules for openswe/hooks |
| Per-env bootstrap qualifier | dev oswedev (CDKToolkit-oswedev) |
| RDS / Redis durable plan | never built — fully dropped (managed provides durability) |
| Bespoke CD workflows | cd-infra.yml, build-artifacts.yml, roll-box.sh/publish-and-deploy.sh/rollback.sh |
KEEP
| Item | Why |
|---|---|
GitHub App seahaven-openswe (App 4146115 / Install 142615168) |
unchanged auth + webhook source; only the webhook/OAuth URLs repoint |
LangSmith sandbox setup incl. DEFAULT_SANDBOX_SNAPSHOT_ID (dc36e509-d4d2-4efc-8a4e-61f74f3446ec) |
the only sandbox with working in-sandbox git/gh auth (_configure_github_proxy); managed default SANDBOX_TYPE=langsmith uses it |
Secrets Manager open-swe-{dev,prod}/* values |
the source you copy env FROM, at least transitionally — do NOT force-delete populated shells until managed env is fully cut over and verified |
| DNS | repoint openswe.seahaven.com → Vercel; decide hooks.seahaven.com → *.langgraph.app or a CNAME (§10) |
seahaven-promotion App + main/dev rulesets (18238334 / 18238542) |
still govern dev → main promotion = the prod deploy gate |
evals/ + JUDGE_ANTHROPIC_API_KEY |
eval pipeline unaffected (separate from runtime) |
Phased teardown rule: remove AWS infra ONLY after managed prod is proven + soaked (Phase D5). The exported ALB JSON (E1) is the rollback source for the hand-built listener rules.
8. Cost
| Line item | Monthly |
|---|---|
| Prod deployment uptime ($0.0036/min, always-on) | ≈ $155 |
| Runs ($0.005/run) | usage-based |
| Traces above 10k/mo | pay-as-you-go |
| LangSmith Plus seat ($39) | already paid (not incremental) |
| Dev deployment | $0 (1 free on Plus, preemptible) |
| Incremental total | ≈ $160/mo |
Compare to the retired AWS stack: 2× EC2 (t4g.large prod + t4g.medium/large dev), 2× S3 buckets, ALB share, NAT egress, plus the never-built RDS/Redis durable tier the self-host path would have added. Managed nets out roughly cost-neutral-to-cheaper while adding durability + previews. TODO: confirm current AWS run-rate for an apples-to-apples delta (not verified in repo).
9. Sea Haven Gates & Obligations
/sh-plan-review(GPT-4.1cross_reviewer) on THIS plan BEFORE prod cutover (Phase C gate). ⚠️run.pymisroutes reviewer-framed prompts to the no-opdoneroute — callmodels.get_cross_reviewer()directly (load orchestrator.env,ChatOpenAI("gpt-4.1")) perfeedback_orchestrator_usage./sh-security-reviewon the sensitive surface: this migration touches auth/webhook signature verification (the customhttp.appis now publicly reachable on*.langgraph.appwith no platform gate in front) and secrets handling (secrets move from Secrets Manager to the Deployment/Vercel env). Required, not opt-in — resolve confirmed critical/high before cutover.- GPT-4.1 cross-family review if the Bedrock-auth fix changes IAM (new Bedrock-scoped IAM user + policy = IAM change → mandatory cross-review).
- Confluence — rework "AWS Architecture Map" (1540098) + open-swe child page (26116098): replace the RDS/EC2/ALB subgraph with an external-services view (LangGraph Cloud + Vercel + retained GitHub App + LangSmith sandbox). Do it in the same conversation as the cutover, not as a follow-up.
- README + INSTALLATION updated (INSTALLATION §10 already canonical; align Sea Haven specifics + the merge-to-deploy norm shift).
- Memories — update
project_open_swe_migration.md+reference_open_swe_deployment.md.
10. Decisions — RESOLVED (2026-06-29, Adam)
- Bedrock auth on managed → STATIC AWS KEYS. Create a Bedrock-scoped IAM user; set
AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY/AWS_REGION=us-east-1in the deployment env; reuse the #62 instance-role Bedrock policy (bedrock:InvokeModel[WithResponseStream]on theus.anthropic.claude-opus-4-8inference-profile ARN + per-region foundation-model ARNs in us-east-1/2 + us-west-2). Caveat: long-lived keys → rotation reminder + keep the policy tightly scoped. New IAM user+policy = mandatory GPT-4.1 IAM cross-review +/sh-security-reviewbefore merge. - Webhook domain → RAW
*.langgraph.app(the "easier" path — no CNAME/DNS/TLS). Point GitHub/Slack/Linear webhook URLs straight at the deployment. Dashboard gets a Vercel custom domain (e.g.openswe.seahaven.com) since that's trivial on Vercel. - Prod approval gate → GATED PROMOTE-TO-MAIN WORKFLOW. Re-home
promote-dev-to-prod.yml: gate the promote job on theprodGitHub Environment reviewer (amoussa1229); on approval it FF-pushesdev → main; the push triggers the managed prod build. Preserves today's manual-approval semantics — approval is on the promote job, the merge tomainis the deploy trigger. - Cost/trace monitoring → YES, in LangSmith (usage/trace budget + alert in the LangSmith workspace).
- Dev/prod home → CONFIRMED: prod = LangChain managed (LangGraph Cloud) + Vercel UI; dev = managed free Dev tier (same runtime as prod).
#62 merge status— RESOLVED (merged + deployed).
10b. Execution approach — 2-PR SPLIT (Adam's call, 2026-06-29, post-plan-review)
The plan-review BLOCKed a literal single PR and recommended a 3-PR minimum (isolate cache fix #5 + the promotion-gate workflow). Adam chose a 2-PR split — isolating the prod-deploy-gate change (highest risk to prod), bundling the rest:
- PR1 — promotion-gate workflow + ruleset. Re-home
promote-dev-to-prod.ymlto the gated promote-to-main (decision 3), and lockmainso ONLY the promote App can push (BLOCK 3). Isolated because it changes how prod deploys — must be independently reviewable/revertible. - PR2 — the 6 code fixes (§5) +
ui/vercel.jsonprod repoint + README/INSTALLATION. (Accepted deviation from the reviewer's 3-PR rec: the multi-replica cache fix #5 stays bundled in PR2 rather than isolated — Adam's call.) - NOT in either PR (rollback safety): AWS infra-code deletion (
deploy/ami,deploy/seahaven,infra/self-host stack) +cdk destroystay a POST-SOAK cleanup (Phase E). - Not PR content (operational, sequenced around the merges): LangGraph Cloud prod deployment + env, Vercel prod + custom domain, webhook/OAuth repoint, Bedrock IAM user, LangSmith budget, Confluence/memory.
10c. Plan-review resolution (GPT-4.1 cross_reviewer, 2026-06-29) — verdict REQUEST CHANGES, all BLOCKs addressed
- BLOCK 1 (single PR) → addressed via the 2-PR split above (conscious deviation: cache fix #5 not isolated — Adam accepted).
- BLOCK 2 (raw LangGraph API public?) → EMPIRICALLY RESOLVED. Unauthenticated probes of the dev deployment returned 403 "Missing authentication headers" for
/threads/search,/assistants/search,/store/items;/ok=200,/dashboard/api/me=401. The platform gates the raw control-plane API. Keep as a pre-prod gate (Phase D4): re-probe the PROD deployment URL before cutover. - BLOCK 3 (prod gate enforcement) → PR1 must lock
mainso only the promote App can push (verify ruleset 18238334: no direct-push path, PR-required, promote App is the sole FF bypass). Any non-gated push tomainwould auto-deploy prod. - BLOCK 4 (secrets posture) → add to Phase B/E: rotate all secrets after migrating them to the platform env stores and BEFORE deleting from AWS; document who can read/write the LangGraph Cloud + Vercel env (access audit); delete AWS shells only post-cutover; no dual-homed/stale secrets; record in memory.
- FIX (rollback integrity) → Phase E checklist: "no IaC/DNS/config the rollback needs is altered in PR1/PR2 or during cutover."
- FIX (gates resolved-before-merge) → make explicit: do NOT merge PR1/PR2 until the GPT-4.1 IAM cross-review (Bedrock IAM user) +
/sh-security-review(auth/webhook/secrets surface) are resolved with no critical/high; Confluence (1540098 + 26116098) + memory updated in the cutover conversation. - NITs/QUESTIONs → carried into §8/§10 TODOs (cost delta,
FF_PROFILE_IMPORTS, platform feature availability, key-rotation owner, env-store access logging, cache race-review under autoscaling, atomic webhook repoint).
11. Rollback Story
Pre-cutover: the self-host langgraph up + RDS plan (RDS+Redis+Docker; /sh-plan-review'd to APPROVE-after-revision this session) is the documented fallback if managed is rejected before cutover. Its durable design wins (env-scoped RDS physical names; Credentials.fromGeneratedSecret({secretName:"open-swe-<env>/rds-credentials"}) under the existing instance-role secret prefix → no new IAM; DESTROY-on-rollback RDS-managed secret; derive DATABASE_URI in fetch-config) are captured in the project memory.
Post-cutover (managed is live, AWS still standing during soak): if managed prod fails, rollback = repoint webhooks + DNS back to the AWS stack:
- GitHub App / Slack / Linear webhook URLs →
hooks.seahaven.com openswe.seahaven.comDNS → theseahaven-comALB (restore from the E1 export)ui/vercel.jsonrewrite → the AWS dashboard origin (or stop using Vercel)- Do NOT run Phase E teardown until the soak passes — the AWS stack IS the rollback target.
Revision-level rollback (within managed): LangGraph Cloud revision rollback (backend) + Vercel instant rollback (UI) cover bad deploys without leaving the platform.
TODOs flagged (couldn't verify in repo)
#62 merge state— RESOLVED: merged todev+ deployed (the draft read a stale local checkout;origin/devand the deploymentgit_ref_sha a4ed19baconfirm Bedrock+Fireworks).- Current AWS run-rate for the cost delta (§8) — not derivable from repo.
- LangGraph Cloud custom-domain + "manual promotion to production" feature availability (§10.2, §10.3) — platform features, verify in the LangGraph Cloud console/docs at execution time.
FF_PROFILE_IMPORTSexact flag name/usage (§5 fix #6) — referenced from session context; confirm the flag exists inagent/before relying on it.- Exact webhook path on
*.langgraph.app(whether the customhttp.appmounts at root so/webhooks/githubis reachable as-is) — proven reachable on the dev spike; re-verify the exact path on prod.