open-swe/deploy/MIGRATION.md
Adam Moussa 800c44f21b
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
docs: repoint plan-review gate to security-review cross_review.py (orchestrator archived) (#184)
2026-07-14 19:28:30 -04:00

45 KiB
Raw Blame History

Open SWE — Migration Plan: Self-Hosted AWS → Managed LangGraph Cloud + Vercel

Repo: Sea-Haven-Industries/open-swe (private) · AWS: 328440206208 / us-east-1 Author: Adam Moussa · Date: 2026-06-29 (final topology added 2026-06-30) · Status: EXECUTED — managed cutover live; §§3–11 below are the original (now-historical) phased plan, superseded by §1a for all current-state facts (URLs, project layout, env).

READ §1a FIRST. The phased plan (§§2–11) and the Phase A/C spike notes capture how we got here and still hold for rationale, cost, and rollback. But the spike-era specifics they cite — the single open-swe-dashboard Vercel project, the open-swe-dev-hosted-…/open-swe-v3-… deployment URLs, the ui/vercel.json same-origin rewrite, the single GitHub App — are stale. §1a is the authoritative final topology and wins on every conflict.


1. Executive Summary

Decision (settled — not re-litigated here)

Move the Open SWE deployment off the bespoke self-hosted AWS stack (stock langgraph dev, in-memory store, no durability — issue #9) onto the managed runtime:

  • Backend → LangGraph Cloud (a.k.a. "LangSmith Deployment"): git-connected, auto-builds a revision on push, zero-downtime, revision rollback, durable Postgres-backed store/checkpointer.
  • UI → Vercel: ui/ SPA, atomic deploys, instant rollback, per-PR previews (net-new capability).

Why

  1. Dissolves the licensing blocker. Self-hosted langgraph up on the LangSmith Plus key is "Self-Hosted Lite" (1M node-exec/yr cap, Elastic License 2.0, production legally ambiguous). Managed IS the licensed product with no node cap — runs billed flat at $0.005, nodes not charged.
  2. Deletes the entire durable-runtime build (RDS/Redis/Docker/AMI) and most of the bespoke AWS CD (S3 release pipeline, SSM deploy docs, packer, CDK box/ALB stacks).
  3. It's the upstream-canonical deployment and the repo is already wired for it: ui/vercel.json has the same-origin /dashboard/api/* rewrite; langgraph.json is cloud-format (6 graphs + http.app).

Cost

≈ $160/mo incremental for one always-on prod deployment:

  • $0.0036/min prod uptime ≈ $155/mo (always-on)
    • $0.005/run
    • traces pay-as-you-go above 10k/mo
  • The $39/mo Plus seat is already paid (not incremental)
  • Dev deployment is FREE (1 included on Plus, preemptible)

A self-hosted langgraph up + RDS plan (already /sh-plan-review'd to APPROVE-after-revision this session) is the documented FALLBACK if managed is ever rejected (see §11).


1a. Final, verified topology (AUTHORITATIVE — supersedes spike-era values)

This is the live managed deployment as of 2026-06-30. Where any later section disagrees (old URLs, a single Vercel project, a single GitHub App, the ui/vercel.json rewrite), this section wins.

Backend — managed LangGraph Cloud (two deployments, one LangSmith workspace)

Both deployments live in the same LangSmith workspace; the same workspace API key authenticates both (including the Store API — so per-deployment store writes use that one key with the per-deployment URL).

Deployment URL Git connection
dev https://open-swe-dev-fb737aa219605c8bbdb30ecbb33f30c0.us.langgraph.app branch dev
prod https://open-swe-prod-d6c7bb63aaa651b6a1d92f9492b1d983.us.langgraph.app branch main (auto-deploys on push to main)

The dev deployment was renamed open-swe-dev, deleted, and recreated — which minted the new URL hash above. The spike-era open-swe-v3-… / open-swe-dev-hosted-… URLs are dead/superseded. Deleting + recreating a deployment is the one operation that changes the URL hash (otherwise stable across revisions) — when it happens, update every reference (Vercel env, GitHub App webhooks, OAuth callbacks, docs).

UI — Vercel (ONE project, two environments)

One Vercel project open-swe-prod (team sea-haven, id prj_OOh6yjXMp4ah3Ws3Y7XRQxjmMmQU). The old separate open-swe-dashboard project was DELETED.

Vercel environment Branch Backend Custom domain
production main prod deployment URL openswe.seahaven.com
custom dev (id env_SMI23PULAJXk0GhwE0HLVhp5J3ZS) dev dev deployment URL openswe-dev.seahaven.com
  • A per-environment env var LANGGRAPH_BACKEND_URL (prod env = prod URL, dev env = dev URL) drives the /dashboard/api/* proxy.
  • Project settings: framework=null, outputDirectory cleared, root directory ui.
  • Proxy mechanism (current, after PR #76): Nitro routeRules in ui/vite.config.ts read process.env.LANGGRAPH_BACKEND_URL and Nitro's Vercel preset compiles them into .vercel/output/config.json (Build Output API) at build time — a CDN-level proxy (not redirect, so the osw_session cookie stays first-party). PR #75's hand-rolled ui/scripts/build-vercel-output.mjs was the broken first attempt and is gone (ui/scripts/ no longer exists). Never add a manual script that rms .vercel/output — Nitro's Vercel preset auto-emits it.

DNS — Route 53 zone seahaven.com (Z06652411XKH89KTZD3XA)

  • openswe.seahaven.com → CNAME cname.vercel-dns.com (prod env)
  • openswe-dev.seahaven.com → CNAME to Vercel (dev env)

GitHub Apps — TWO (dev/prod isolated; each its own webhook URL)

App app_id install client_id org members scope repos
prod seahaven-openswe 4146115 142615168 Iv23lil96pKQNNDUn5yp Sea-Haven-Industries members:write all
dev seahaven-openswe-dev 4162963 143023302 Iv23licQwJvGAPJj1HJe seahaven-open-swe-dev members:read all
  • Promotion App seahaven-promotion (actor 4170147) is the sole non-admin fast-forward-push bypass on the main ruleset 18238334 — its FF-push of dev → main is what triggers the managed prod build.

Env per deployment (set in LangGraph Cloud config + Vercel env — NOT Secrets Manager)

Var dev prod
LANGGRAPH_URL own (dev) deployment URL own (prod) deployment URL
DASHBOARD_BASE_URL / DASHBOARD_API_BASE_URL https://openswe-dev.seahaven.com https://openswe.seahaven.com
VITE_DASHBOARD_API_BASE_URL empty (same-origin via Vercel proxy) empty
ALLOWED_GITHUB_ORGS dev org (seahaven-open-swe-dev) Sea-Haven-Industries
CONFIGURED_ADMINS amoussa1229,adam@seahavenind.com amoussa1229,adam@seahavenind.com

DASHBOARD_BASE_URL / DASHBOARD_API_BASE_URL must include https:// (see gotcha 3). Secret values are still sourced from open-swe-{dev,prod}/* Secrets Manager + SSM (the remaining source of truth) and set into the LangGraph Cloud + Vercel env stores — the accepted secrets-and-config.md deviation.

Bedrock IAM (PR #74, still OPEN)

Two IAM users open-swe-dev-bedrock + open-swe-prod-bedrock, each attached to customer-managed policy open-swe-bedrock-invoke (least-privilege bedrock:InvokeModel[WithResponseStream] on the us.anthropic.claude-opus-4-8 inference-profile ARN + its 3 routed foundation-model ARNs in us-east-1/us-east-2/us-west-2). Default model bedrock_converse:us.anthropic.claude-opus-4-8 + 3 Fireworks models. Static access keys live only in the deployment env (dev key → dev, prod key → prod).

User store — per-deployment

Each managed deployment has its own Store. The GitHub→email mapping amoussa1229 → adam@seahavenind.com (namespace ["user_mappings"], key = lowercased login, record {github_login, work_email, status:"active", source, created_at, updated_at}) was written to both the dev and prod stores directly. New users need a mapping per-deployment (write each store directly, or use the dashboard admin User-mappings UI — the work_email field was added by PR #65 fix #4).

AWS decommission (PR #64)

Self-host CDK stacks destroyed. Residual: CDKToolkit (shared, preserved); ~50 RETAIN'd Secrets Manager shells + 3 S3 asset buckets (pending cleanup); AWS Bedrock (live dependency, kept).

Operational gotchas (hard-won — carry these into any runbook)

  1. Per-deployment store → seed user mappings per-deployment. A missing mapping makes process_github_issue silently early-return ("No email mapping … skipping"): the webhook returns 200/accepted but produces no reaction and no run. Seed dev and prod.
  2. Org-login gate uses the App installation token, so the App must be org-installed with Members:read. OAuth working ≠ membership check working — they use separate creds (CLIENT_ID/SECRET for OAuth vs APP_ID/INSTALLATION_ID/PRIVATE_KEY for the install token). A mangled multi-line GITHUB_APP_PRIVATE_KEY breaks the install token (and thus the gate) while OAuth still works.
  3. DASHBOARD_API_BASE_URL must be https:// — an http:// value makes GitHub reject the OAuth callback with "redirect_uri not associated."
  4. osw_oauth_state cookie is host-only — start login on the same host as DASHBOARD_API_BASE_URL, or you get "oauth state mismatch."
  5. Webhooks go DIRECT to the langgraph URL (/webhooks/*). Vercel only proxies /dashboard/api/*. The app is same-origin only (no CORS).
  6. On Vercel CI, Nitro's Vercel preset auto-emits .vercel/output — drive the proxy via Nitro routeRules from LANGGRAPH_BACKEND_URL; never a manual script that rms .vercel/output.
  7. Deleting + recreating a LangGraph deployment mints a NEW URL hash (otherwise stable across revisions) — update every reference (Vercel env, webhooks, OAuth, docs).

2. Architecture: Before → After

Before (self-hosted AWS — LIVE as of 2026-06-29)

GitHub/Slack/Linear ──webhook──▶ hooks.seahaven.com ─┐
Browser (dashboard) ─────────────▶ openswe.seahaven.com ─┤
                                                          ▼
                          shared seahaven-com ALB (:443, host+path rules)
                                                          ▼
                  EC2 (prod i-08a729e50779c4b07 t4g.large ARM64, private subnet)
                    nginx :80 ──proxy /dashboard/api + /webhooks──▶ langgraph dev :2024 (loopback)
                    in-memory store (reseeded by seed_store.sh ExecStartPost)
                                                          ▼
                  Config: Secrets Manager open-swe-prod/* + SSM /open-swe-prod/*
                  Release: GitHub Actions → S3 open-swe-prod-assets/releases/* → SSM doc deploy.sh
                  IaC: CDK OpenSweIamStack + OpenSweDevStack + OpenSweProdStack
                  Sandbox: LangSmith cloud (DEFAULT_SANDBOX_SNAPSHOT_ID + GitHub proxy)

After (managed — FINAL, see §1a for exact values)

                              ┌──────────────── DEV lane ────────────────┐   ┌──────────────── PROD lane ───────────────┐
GitHub(dev org)/Slack/Linear  │ webhook → open-swe-dev-….us.langgraph.app │   │ webhook → open-swe-prod-….us.langgraph.app│  GitHub(SHI org)/Slack/Linear
  App seahaven-openswe-dev ───┘    (DIRECT to langgraph URL, /webhooks/*)  │   │   (DIRECT to langgraph URL, /webhooks/*)  └─── App seahaven-openswe
                                                                          ▼   ▼
Browser ▶ openswe-dev.seahaven.com ─┐                          ┌─▶ openswe.seahaven.com ◀ Browser
                                    │   ONE Vercel project `open-swe-prod` (team sea-haven)
                                    │   ├─ env `dev`  (branch dev)  → proxies /dashboard/api/* → dev langgraph URL
                                    │   └─ env production (branch main) → proxies /dashboard/api/* → prod langgraph URL
                                    └─ proxy compiled by Nitro routeRules from per-env LANGGRAPH_BACKEND_URL (PR #76)
                                          ▼
            TWO LangGraph Cloud deployments (same LangSmith workspace; one workspace key auths both incl. Store)
              ├─ dev  ← branch `dev`     ·  prod ← branch `main` (push-to-main auto-deploys prod)
              ├─ each serves the graphs + the custom http.app (agent.webapp:app = webhooks + dashboard API + OAuth)
              ├─ each has its OWN durable Postgres store + checkpointer (issue #9 SOLVED) — user mappings per-deployment
              └─ env/secrets in the Deployment + Vercel config (NOT Secrets Manager) — Adam accepted deviation
                                          ▼
            Sandbox: LangSmith cloud (UNCHANGED — DEFAULT_SANDBOX_SNAPSHOT_ID + GitHub-App proxy)
            Bedrock: IAM users open-swe-{dev,prod}-bedrock + policy open-swe-bedrock-invoke (static keys in deploy env)

What changes shape: runtime host (EC2 → managed PaaS), durability (in-memory → managed Postgres, one store per deployment), CD (bespoke S3/SSM/packer → git-connected auto-build), config home (Secrets Manager/SSM → Deployment+Vercel env), ingress topology (shared ALB + one App → two dev/prod-isolated GitHub Apps each hitting its own *.langgraph.app directly), UI proxy (ui/vercel.json rewrite → Nitro routeRules). What stays: the LangSmith sandbox plane, CI (lint/format/unit/Playwright), the app code itself, and Secrets Manager/SSM as the secret-value source of truth.


3. Phased Plan

Phase A — Dev spike (MOSTLY DONE)

Goal: prove managed serves our custom app + durability, at 0, before committing prod .

Proven this session (spike-era specifics — superseded by §1a; URLs/project below are DEAD):

  • ✅ Dev backend deployed to LangGraph Cloud, connected to branch dev: https://open-swe-dev-hosted-e76c2b0e8a7955fe8ad3110a7a54e5d0.us.langgraph.app → final dev URL in §1a (open-swe-dev-fb737aa…; the deployment was later deleted + recreated) (Bedrock + a Fireworks key set; AWS creds for Bedrock deferred — see §10 open decision).
  • ✅ UI deployed to Vercel — team sea-haven, project open-swe-dashboard, https://open-swe-dashboard.vercel.app (DELETED); now the single project open-swe-prod with dev/prod environments (§1a); proxy is Nitro routeRules, not the ui/vercel.json rewrite.
  • ✅ Managed serves the custom http.app (dashboard API + webhooks) — no platform auth gate in front of our routes (webhook returns 401 sig-enforced, so signatures still govern).
  • ✅ Vercel same-origin rewrite → backend works.
  • ✅ GitHub OAuth dashboard login end-to-end.
  • ✅ DURABILITY — team-default model survived a revision redeploy (this is issue #9's core acceptance goal, proven on managed).

Remaining Phase A items (the GitHub @openswe run-trigger loop):

  • A1. Get an @openswe GitHub comment to dispatch a run end-to-end on dev. Blocked by the user-mapping + cache gotchas (see fixes #4 and #5 in §5).
  • A2. Create the owner user-mapping in the managed Store (namespace ["user_mappings"], key = lowercased login amoussa1229, value {github_login, work_email}) — the Admin UI can't set work_email yet (fix #4).
  • A3. Verify a freshly-added mapping is seen without a redeploy across managed's multi-replica autoscaling (fix #5).
  • A4. Confirm LANGGRAPH_URL is set to the deployment URL on the dev deployment so langgraph_client() calls resolve (fix #1 — config now, code later).

Phase B — Harden (the code fixes + env codification)

Land the 6 code fixes (§5), codify env/config, and resolve the Bedrock-auth decision before standing up prod.

  • B1. Land code fixes #1–#6 (§5) as a PR into dev (or split into focused PRs). Re-run make lint + make test.
  • B2. Codify the env contract for managed. Update ui/vercel.json rewrite to the prod deployment URL (Phase C) but keep dev pointing at dev. Add a documented LangGraph-Cloud env list to the repo (see §4) — NOT secrets, just the variable inventory + which are excluded.
  • B3. Resolve the Bedrock-auth decision (§10 open decision): static AWS keys for a Bedrock-scoped IAM user in the deployment env, OR run the agent on Fireworks. This intersects the #62 model work.
  • B4. Merge #62 first — DONE / N/A. #62 (Bedrock + Fireworks) IS merged to dev and is what the dev deployment runs — verified on origin/dev (options.py → DEFAULT_MODEL_ID = "bedrock_converse:us.anthropic.claude-opus-4-8") and confirmed by the live deployment's git_ref_sha: a4ed19ba… (= the #62 merge commit). The contrary note came from a STALE local checkout; no merge action needed.
  • B5. Audit ALL in-process caches for the single-process → multi-replica assumption (generalization of fix #5; see §5).
  • B6. Measure + reduce custom-app import time (fix #6) so the deployment isn't flagged unhealthy / slow to scale.

Phase C — Prod deployment

  • C1. Create a prod LangGraph Cloud deployment tracking branch main (the durable autoscaled 1→10 tier, not the free Dev tier). Record its *.langgraph.app URL.
  • C2. Create the prod Vercel project/target (or promote the existing open-swe-dashboard to production) — DONE differently: one project open-swe-prod with a production env (main) + a custom dev env (dev), each with its own LANGGRAPH_BACKEND_URL. See §1a.
  • C3. Set the prod env triad (§4) on the prod deployment + Vercel:
    • LANGGRAPH_URL = the prod *.langgraph.app URL
    • DASHBOARD_BASE_URL + DASHBOARD_API_BASE_URL = the prod Vercel origin (with https:// scheme — fix #2)
  • C4. Establish the prod approval gate (§6): git-connected auto-deploy needs an explicit gate. Preserve the prod manual-approval that exists today (the GitHub prod Environment reviewer). Mechanism = protected main ruleset + the platform's "require manual promotion to production" if available (§10 open decision).
  • C5. Repoint webhooks + OAuth URLs to prod:
    • GitHub App seahaven-openswe webhook URL → prod backend webhook URL (*.langgraph.app/webhooks/github or hooks.seahaven.com — §10 custom-domain decision)
    • Slack Event Subscriptions + Interactivity Request URLs → prod
    • Linear webhook URL → prod (if/when wired)
    • GitHub App OAuth callback → <DASHBOARD_API_BASE_URL>/dashboard/api/auth/callback (the Vercel prod origin; fixes #2, #3)
  • C6. Custom-domain decision (§10): dashboard custom domain solved by Vercel; webhooks hooks.seahaven.com either re-point to *.langgraph.app directly or front via CNAME. Managed issues *.langgraph.app URLs ONLY (no documented custom-domain support on the backend).

Phase D — Cutover + verify

  • D1. Run the issue #9 acceptance criteria on PROD: durable store survives a revision redeploy; team defaults / user mappings persist; a paused/in-flight run survives a redeploy (the durability win self-host never had).
  • D2. End-to-end prod smoke: @openswe GitHub comment → run dispatch → sandbox → draft PR → reply in source channel.
  • D3. Dashboard prod smoke: GitHub OAuth login on the stable Vercel alias (fix #3); admin pages reachable for CONFIGURED_ADMINS.
  • D4. Webhook sig-enforcement smoke: unsigned POST to each /webhooks/* returns 401; signed returns 200.
  • D5. Soak for an agreed window (recommend 3–7 days) with the AWS prod stack still standing as the rollback target (§11) before any teardown.

Phase E — AWS decommission (AFTER soak)

Only after managed prod is proven + soaked. See §7 for the precise retire-vs-keep list. Phased teardown:

  • E1. Export the live ALB listener-rule / target-group config for open-swe-prod and open-swe-dev to JSON (the export is the rollback source — these rules were partly hand-built).
  • E2. Disable the bespoke CD (build-artifacts / cd-infra workflows) so nothing re-deploys the EC2 boxes.
  • E3. cdk destroy OpenSweProdStack then OpenSweDevStack (mind the RETAIN secret shells — they survive and hold the global open-swe-<env>/* names; force-delete only EMPTY shells once env is fully migrated; keep populated ones transitionally as the env source — §7).
  • E4. Remove ALB rules/target groups, S3 assets buckets, SSM deploy docs, and the per-env CDK bootstrap qualifier oswedev (CDKToolkit-oswedev) — once nothing references them.
  • E5. Decide DNS: repoint openswe.seahaven.com → Vercel; resolve hooks.seahaven.com → *.langgraph.app (or CNAME).
  • E6. Leave OpenSweIamStack until last — the promotion App + any retained roles may still be referenced.

Phase F — Docs / memory

  • F1. Rework Confluence "AWS Architecture Map" (page 1540098) + the open-swe child page (26116098): the RDS/EC2/ALB subgraph is replaced by an external-services view (LangGraph Cloud + Vercel + the retained GitHub App + LangSmith sandbox).
  • F2. Update README.md + INSTALLATION.md (§10 is already canonical-correct; align Sea Haven specifics).
  • F3. Update project memories project_open_swe_migration.md + reference_open_swe_deployment.md to mark the managed cutover and the AWS teardown.

4. Env / Config Reference

The prod triad (per INSTALLATION.md §10, lines 630–658) — see §1a for the exact dev/prod values

Var Value (per deployment/env) Notes
LANGGRAPH_URL the own deployment URL (https://...langgraph.app) NOT localhost. Drives thread_ops.langgraph_url() (fix #1). dev→dev URL, prod→prod URL.
DASHBOARD_BASE_URL the own dashboard origin, with https:// (https://openswe-dev.seahaven.com / https://openswe.seahaven.com)
DASHBOARD_API_BASE_URL same as above, with https:// scheme scheme required or OAuth redirect_uri is schemeless and GitHub rejects (fix #2)
VITE_DASHBOARD_API_BASE_URL empty same-origin mode; UI calls relative /dashboard/api/*, the Vercel Nitro proxy rewrites to backend
LANGGRAPH_BACKEND_URL Vercel env var, per Vercel environment (dev env = dev URL, prod env = prod URL) drives the Nitro routeRules /dashboard/api/* proxy at build time (PR #76) — required on Vercel builds
ALLOWED_GITHUB_ORGS Sea-Haven-Industries (prod) / seahaven-open-swe-dev (dev) org-login gate; checked via the App installation token (gotcha 2)

GitHub App dashboard OAuth callback = <DASHBOARD_API_BASE_URL>/dashboard/api/auth/callback (the own dashboard origin — prod App on openswe.seahaven.com, dev App on openswe-dev.seahaven.com).

UI proxy mechanism (final, PR #76): the /dashboard/api/* proxy is Nitro routeRules in ui/vite.config.ts reading process.env.LANGGRAPH_BACKEND_URL, compiled by Nitro's Vercel preset into .vercel/output/config.json. This supersedes the spike-era ui/vercel.json same-origin rewrite and PR #75's hand-rolled ui/scripts/build-vercel-output.mjs (deleted). ui/vercel.json now only carries framework:null + buildCommand: bun run build.

Where secrets/env live now

LangGraph Cloud Deployment config + Vercel env — NOT AWS Secrets Manager. This deviates from the Sea Haven secrets-and-config.md "Secrets Manager for all sensitive" handbook rule — Adam ACCEPTED this deviation (managed has no instance role / no fetch-config boot hook; the platform's own secret store is the mechanism).

Sourcing env from the existing AWS stack (transitional)

Pull from the existing open-swe-dev Secrets Manager + SSM via a documented CLI dump:

  • aws secretsmanager list-secrets --filters Key=name,Values=open-swe-dev/ then per-name get-secret-value — NOT batch-get-secret-value (its pagination silently drops values past page 1; this bit the box twice — see memory).
  • SSM: aws ssm get-parameters-by-path --path /open-swe-dev/.

EXCLUDE when copying to managed

Category Vars
Box-/self-host-specific LANGGRAPH_URL (set fresh to deployment URL), LANGGRAPH_URL_PROD, LANGSMITH_ENDPOINT, LANGSMITH_ENDPOINT_PROD, LANGSMITH_URL_PROD, LANGSMITH_TENANT_ID_PROD, LANGSMITH_HOST_API_URL, LANGCHAIN_REVISION_ID
Dropped providers (PR #62) OPENAI_API_KEY, GOOGLE_API_KEY, GROQ_API_KEY (these 6 Secrets Manager shells deleted 2026-06-29, final purge 2026-07-06)
Unused sandbox providers DAYTONA_API_KEY, RUNLOOP_API_KEY
Eval-only JUDGE_ANTHROPIC_API_KEY (kept on AWS for evals/reviewer/judge.py, not needed in the runtime deployment)

KEY mappings on managed

  • LANGSMITH_API_KEY = the LANGSMITH_API_KEY_PROD value (and LANGCHAIN_API_KEY = same).
  • The full secret inventory is SECRET_VARS in infra/lib/constructs/config-store.ts:46 (28 shells) — use it as the checklist of what to carry, minus the EXCLUDE rows above.

Models (post-#62) + Bedrock auth

Post-#62 SUPPORTED_MODELS = bedrock_converse:us.anthropic.claude-opus-4-8 (DEFAULT) + 3 Fireworks models; no direct anthropic: option. ✅ #62 is merged to dev and deployed — verified on origin/dev (options.py → bedrock_converse default) and the live deployment git_ref_sha a4ed19ba. (The pre-#62 values appear only on a stale LOCAL checkout — ignore.)

Bedrock on managed has NO EC2 instance role (the #62 IAM design used the EC2 instance role as the Bedrock principal — that breaks on managed). Options (OPEN DECISION, §10):

  • (a) Static AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY / AWS_REGION for a Bedrock-scoped IAM user in the deployment env, or
  • (b) Run the agent on Fireworks and avoid Bedrock entirely on managed.

Bedrock IAM users (CREATED — resolves §10.1, discharges the §9 IAM gates)

Option (a) was chosen and executed live (2026-06-30, account 328440206208 / us-east-1). This is click-ops IAM — there is no remaining open-swe AWS IaC after the decommission (#64), so these are created with the CLI, not CDK.

  • Customer-managed policy open-swe-bedrock-invoke (arn:aws:iam::328440206208:policy/open-swe-bedrock-invoke). Least-privilege: actions bedrock:InvokeModel + bedrock:InvokeModelWithResponseStream only, scoped to exactly the us.anthropic.claude-opus-4-8 inference-profile ARN + its three routed foundation-model ARNs (us-east-1, us-east-2, us-west-2). No wildcards, no other models. (Supersedes the §10.1 plan to reuse the #62 instance-role policy — a fresh standalone policy was minted instead.)
  • Two IAM users, each attached to that policy: open-swe-dev-bedrock and open-swe-prod-bedrock. Tagged project=open-swe, managed-by=cli-migration, purpose=bedrock-invoke.
  • Static access keys are minted separately by the owner (aws iam create-access-key) — secret keys live only in the deployment env stores, never in this repo or memory. dev key → dev env; prod key → LangGraph Cloud prod config (AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY / AWS_REGION=us-east-1).
  • Deviation rationale: static long-lived keys are a deliberate departure from the Sea Haven OIDC norm because managed LangGraph Cloud cannot assume an AWS role. Mitigated by the tight least-privilege policy above.
  • ✅ Both mandatory gates ran and PASSED:
    • GPT-4.1 IAM cross-review (confirmed hit gpt-4.1-2025-04-14) — least-privilege confirmed.
    • /sh-security-review — 0 critical/high; two accepted mediums (the static-key deviation + no per-principal budget cap).
  • Recommended follow-ups: shortest viable key-rotation cadence with recorded creation dates; an AWS Budgets / CloudWatch anomaly alarm on per-principal Bedrock InvokeModel volume; confirm CloudTrail captures these users.

5. Required Code Fixes (the 6 gotchas)

These are self-host → managed assumption breaks discovered on the spike. Each gets a config workaround now and a code fix to land in Phase B.

Fix #1 — langgraph_url() localhost fallback

File: agent/utils/thread_ops.py:38-41

def langgraph_url() -> str:
    return os.environ.get("LANGGRAPH_URL") or os.environ.get(
        "LANGGRAPH_URL_PROD", "http://localhost:2024"
    )

On managed, every langgraph_client() call (thread sidebar, run creation, the same pattern in agent/dashboard/user_mappings.py:43 get_client() with no URL) ConnectErrors unless LANGGRAPH_URL is set to the deployment URL.

  • Config fix (now): set LANGGRAPH_URL = the deployment URL on every deployment (Phase A4 / C3).
  • Code fix (Phase B): default to the in-process / deployment URL rather than http://localhost:2024 when running inside a managed deployment (e.g. honor the platform's own URL env, or fail loud instead of silently hitting localhost).

Fix #2 — OAuth DASHBOARD_API_BASE_URL must include scheme

Where: OAuth redirect construction (dashboard oauth.py / routes.py), driven by DASHBOARD_API_BASE_URL. A schemeless value yields a schemeless redirect_uri → GitHub rejects with "redirect_uri not associated."

  • Fix: always set DASHBOARD_API_BASE_URL with https:// (Phase C3). Code hardening (Phase B): assert/normalize a scheme at startup and fail loud if missing.

Cookie: osw_oauth_state (host-only). Login must START on the SAME host as DASHBOARD_API_BASE_URL — the stable Vercel alias, never the immutable per-deploy URL — or you get "oauth state mismatch."

  • Fix: pin DASHBOARD_API_BASE_URL + the login entry point to the stable alias / custom domain (Phase C3, D3). Document this so a future per-deploy preview URL isn't used for login.

Fix #4 — Admin "User mappings" UI can't set work_email

Files: the Admin → User mappings UI (ui/) + agent/dashboard/user_mappings.py (upsert_mapping at line 257 takes work_email but the UI form doesn't expose it). Mappings can't be fully created from the dashboard → must write the Store directly: namespace ["user_mappings"], key = lowercased login, value {github_login, work_email} (matching _index_record at user_mappings.py:73-84, which keys _by_login on login.lower()).

  • Fix (Phase B): add the work_email field to the Admin mappings form so mappings are fully creatable from the UI.

Fix #5 — GitHub webhook path doesn't refresh the user-mapping cache (multi-replica break)

Files: agent/webhooks/github.py (GitHub handlers: process_github_pr_comment, process_github_issue) vs agent/webhooks/slack.py (process_slack_mention). (Pre-modular-refactor these all lived in agent/webapp.py.) The Slack path refreshes before lookup:

# agent/webhooks/slack.py (process_slack_mention)
await webapp.refresh_user_mapping_cache()
...

The GitHub path historically did not — it called email = await email_for_login(github_login) cold. The cache (user_mappings.py _ensure_cache_loaded) is one-shot per process (_cache_loaded flag). On self-host single-process this was fine; on managed's multi-replica autoscaling, a freshly-added mapping isn't seen by a replica whose cache loaded earlier — until restart.

  • Fix (applied): the GitHub issue and PR-comment handlers now call webapp.refresh_user_mapping_cache() before email resolution, mirroring the Slack path.
  • Generalize (Phase B5): audit ALL in-process caches for the single-process → multi-replica assumption — SANDBOX_BACKENDS dict (agent/utils/sandbox_state.py), _by_login/_by_email/_by_slack_id (user_mappings.py). Sandbox affinity is already thread-keyed + persisted in thread metadata (sandbox_id), so it's the cache state that needs the multi-replica review. (The legacy in-process thread lock has been removed: webhook triggers now serialize through dispatch_agent_run's multitask_strategy="interrupt" instead.)

Fix #6 — Slow custom-app import (~8s startup)

Symptom: "exceeded expected startup time" → risks the deployment being marked unhealthy / slow to scale out.

  • Fix (Phase B6): lazy imports / reduce import-time work in agent/webapp.py (+ agent/webhooks/*.py) and the graph factories. Profile with FF_PROFILE_IMPORTS (the import-profiling flag) to find the heavy modules.

6. CD / Ops Changes

Retires (bespoke pipeline)

  • GitHub Actions build → S3 releases → SSM doc → deploy.sh → health-gate (.github/scripts/, deploy/ami/deploy.sh)
  • roll-box.sh / publish-and-deploy.sh / rollback.sh
  • AMI baking (packer deploy/ami/open-swe-base.pkr.hcl) + cdk.context.json AMI pin
  • cd-infra.yml CDK deploys + OIDC bootstrap qualifiers
  • fetch-config.sh / seed_store.sh (boot-time config materialization + Store reseed)

Replaced by git-connected PaaS

  • LangGraph Cloud auto-builds a revision on push (first-party zero-downtime + revision rollback).
  • Vercel builds / atomic-deploys / instant-rollback + per-PR previews (net-new capability the AWS stack never had).

Stays

  • CI (lint / format / unit / Playwright E2E) — still matters and still gates merges. make lint, make test.

Gate re-homing

  • The dev → main promotion gate (check-dev-green.sh + protected-main ruleset 18238334) re-homes: the dev deployment tracks dev, the prod deployment tracks main, so the gate governs exactly what reaches prod.
  • Today's promotion machinery: promote-dev-to-prod.yml + the seahaven-promotion GitHub App (app_id 4170147, in ruleset 18238334 bypass_actors) FF-pushes dev → main. Under managed, a push to main is what triggers the prod build — so the promotion gate IS the prod deploy gate.

Preserve the PROD manual-approval gate (LOAD-BEARING)

Today the prod GitHub Environment (required reviewer amoussa1229) is the manual approval. Git-connected auto-deploy removes the CD job that consulted that Environment, so the approval must be re-established explicitly:

  • Mechanism options (OPEN DECISION §10): protected main (only the promotion App can FF-push, and that push is itself the gate) AND/OR the platform's "require manual promotion to production" if LangGraph Cloud exposes it.
  • Norm shift: deploy-then-merge → merge-to-deploy. Lean on the dev deployment + Vercel previews to verify before promoting dev → main. (This inverts the Sea Haven handbook deploy-then-merge default — call it out in the README + handbook note.)

7. AWS Decommission — Retire vs Keep

RETIRE (becomes vestigial once managed prod is live + soaked)

Item Path / resource
AMI bake + cloud-init deploy/ami/ (open-swe-base.pkr.hcl, provision.sh, user-data.sh, deploy.sh, templates/)
Self-host boot/config scripts deploy/seahaven/fetch-config.sh, seed_store.sh, put-config.sh, nginx/openswe.conf, systemd/open-swe.service, aegra/ (already deferred)
CDK box/ALB/AMI stacks infra/lib/open-swe-stack.ts, constructs/app-service.ts, assets-bucket.ts, ami-cache.ts, instance-role.ts, github-deploy-roles.ts
EC2 instances prod i-08a729e50779c4b07 (t4g.large), dev i-0a3bb8e0ddd36c29b
S3 release buckets open-swe-dev-assets, open-swe-prod-assets
SSM deploy docs open-swe-dev-deploy, open-swe-prod-deploy
ALB plumbing open-swe-prod-tg / open-swe-dev-tg target groups + the seahaven-com ALB host/path listener rules for openswe/hooks
Per-env bootstrap qualifier dev oswedev (CDKToolkit-oswedev)
RDS / Redis durable plan never built — fully dropped (managed provides durability)
Bespoke CD workflows cd-infra.yml, build-artifacts.yml, roll-box.sh/publish-and-deploy.sh/rollback.sh

KEEP

Item Why
GitHub App seahaven-openswe (App 4146115 / Install 142615168) unchanged auth + webhook source; only the webhook/OAuth URLs repoint
LangSmith sandbox setup incl. DEFAULT_SANDBOX_SNAPSHOT_ID (dc36e509-d4d2-4efc-8a4e-61f74f3446ec) the only sandbox with working in-sandbox git/gh auth (_configure_github_proxy); managed default SANDBOX_TYPE=langsmith uses it
Secrets Manager open-swe-{dev,prod}/* values the source you copy env FROM, at least transitionally — do NOT force-delete populated shells until managed env is fully cut over and verified
DNS repoint openswe.seahaven.com → Vercel; decide hooks.seahaven.com → *.langgraph.app or a CNAME (§10)
seahaven-promotion App + main/dev rulesets (18238334 / 18238542) still govern dev → main promotion = the prod deploy gate
evals/ + JUDGE_ANTHROPIC_API_KEY eval pipeline unaffected (separate from runtime)

Phased teardown rule: remove AWS infra ONLY after managed prod is proven + soaked (Phase D5). The exported ALB JSON (E1) is the rollback source for the hand-built listener rules.


8. Cost

Line item Monthly
Prod deployment uptime ($0.0036/min, always-on) ≈ $155
Runs ($0.005/run) usage-based
Traces above 10k/mo pay-as-you-go
LangSmith Plus seat ($39) already paid (not incremental)
Dev deployment $0 (1 free on Plus, preemptible)
Incremental total ≈ $160/mo

Compare to the retired AWS stack: 2× EC2 (t4g.large prod + t4g.medium/large dev), 2× S3 buckets, ALB share, NAT egress, plus the never-built RDS/Redis durable tier the self-host path would have added. Managed nets out roughly cost-neutral-to-cheaper while adding durability + previews. TODO: confirm current AWS run-rate for an apples-to-apples delta (not verified in repo).


9. Sea Haven Gates & Obligations

  • /sh-plan-review (GPT-4.1) on THIS plan BEFORE prod cutover (Phase C gate). Runs via python3 ~/Documents/repositories/seahaven/security-review/cross_review.py "<task>" (the orchestrator repo was archived 2026-07-14; the CLI has no router, so the old misroute workaround is obsolete).
  • /sh-security-review on the sensitive surface: this migration touches auth/webhook signature verification (the custom http.app is now publicly reachable on *.langgraph.app with no platform gate in front) and secrets handling (secrets move from Secrets Manager to the Deployment/Vercel env). Required, not opt-in — resolve confirmed critical/high before cutover.
  • GPT-4.1 cross-family review if the Bedrock-auth fix changes IAM (new Bedrock-scoped IAM user + policy = IAM change → mandatory cross-review).
  • Confluence — rework "AWS Architecture Map" (1540098) + open-swe child page (26116098): replace the RDS/EC2/ALB subgraph with an external-services view (LangGraph Cloud + Vercel + retained GitHub App + LangSmith sandbox). Do it in the same conversation as the cutover, not as a follow-up.
  • README + INSTALLATION updated (INSTALLATION §10 already canonical; align Sea Haven specifics + the merge-to-deploy norm shift).
  • Memories — update project_open_swe_migration.md + reference_open_swe_deployment.md.

10. Decisions — RESOLVED (2026-06-29, Adam)

  1. Bedrock auth on managed → STATIC AWS KEYS. Create a Bedrock-scoped IAM user; set AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY / AWS_REGION=us-east-1 in the deployment env; reuse the #62 instance-role Bedrock policy (bedrock:InvokeModel[WithResponseStream] on the us.anthropic.claude-opus-4-8 inference-profile ARN + per-region foundation-model ARNs in us-east-1/2 + us-west-2). Caveat: long-lived keys → rotation reminder + keep the policy tightly scoped. New IAM user+policy = mandatory GPT-4.1 IAM cross-review + /sh-security-review before merge.
  2. Webhook domain → RAW *.langgraph.app (the "easier" path — no CNAME/DNS/TLS). Point GitHub/Slack/Linear webhook URLs straight at the deployment. Dashboard gets a Vercel custom domain (e.g. openswe.seahaven.com) since that's trivial on Vercel.
  3. Prod approval gate → GATED PROMOTE-TO-MAIN WORKFLOW. Re-home promote-dev-to-prod.yml: gate the promote job on the prod GitHub Environment reviewer (amoussa1229); on approval it FF-pushes dev → main; the push triggers the managed prod build. Preserves today's manual-approval semantics — approval is on the promote job, the merge to main is the deploy trigger.
  4. Cost/trace monitoring → YES, in LangSmith (usage/trace budget + alert in the LangSmith workspace).
  5. Dev/prod home → CONFIRMED: prod = LangChain managed (LangGraph Cloud) + Vercel UI; dev = managed free Dev tier (same runtime as prod).
  6. #62 merge status — RESOLVED (merged + deployed).

10b. Execution approach — 2-PR SPLIT (Adam's call, 2026-06-29, post-plan-review)

The plan-review BLOCKed a literal single PR and recommended a 3-PR minimum (isolate cache fix #5 + the promotion-gate workflow). Adam chose a 2-PR split — isolating the prod-deploy-gate change (highest risk to prod), bundling the rest:

  • PR1 — promotion-gate workflow + ruleset. Re-home promote-dev-to-prod.yml to the gated promote-to-main (decision 3), and lock main so ONLY the promote App can push (BLOCK 3). Isolated because it changes how prod deploys — must be independently reviewable/revertible.
  • PR2 — the 6 code fixes (§5) + ui/vercel.json prod repoint + README/INSTALLATION. (Accepted deviation from the reviewer's 3-PR rec: the multi-replica cache fix #5 stays bundled in PR2 rather than isolated — Adam's call.)
  • NOT in either PR (rollback safety): AWS infra-code deletion (deploy/ami, deploy/seahaven, infra/ self-host stack) + cdk destroy stay a POST-SOAK cleanup (Phase E).
  • Not PR content (operational, sequenced around the merges): LangGraph Cloud prod deployment + env, Vercel prod + custom domain, webhook/OAuth repoint, Bedrock IAM user, LangSmith budget, Confluence/memory.

10c. Plan-review resolution (GPT-4.1 cross_reviewer, 2026-06-29) — verdict REQUEST CHANGES, all BLOCKs addressed

  • BLOCK 1 (single PR) → addressed via the 2-PR split above (conscious deviation: cache fix #5 not isolated — Adam accepted).
  • BLOCK 2 (raw LangGraph API public?) → EMPIRICALLY RESOLVED. Unauthenticated probes of the dev deployment returned 403 "Missing authentication headers" for /threads/search, /assistants/search, /store/items; /ok=200, /dashboard/api/me=401. The platform gates the raw control-plane API. Keep as a pre-prod gate (Phase D4): re-probe the PROD deployment URL before cutover.
  • BLOCK 3 (prod gate enforcement) → PR1 must lock main so only the promote App can push (verify ruleset 18238334: no direct-push path, PR-required, promote App is the sole FF bypass). Any non-gated push to main would auto-deploy prod.
  • BLOCK 4 (secrets posture) → add to Phase B/E: rotate all secrets after migrating them to the platform env stores and BEFORE deleting from AWS; document who can read/write the LangGraph Cloud + Vercel env (access audit); delete AWS shells only post-cutover; no dual-homed/stale secrets; record in memory.
  • FIX (rollback integrity) → Phase E checklist: "no IaC/DNS/config the rollback needs is altered in PR1/PR2 or during cutover."
  • FIX (gates resolved-before-merge) → make explicit: do NOT merge PR1/PR2 until the GPT-4.1 IAM cross-review (Bedrock IAM user) + /sh-security-review (auth/webhook/secrets surface) are resolved with no critical/high; Confluence (1540098 + 26116098) + memory updated in the cutover conversation.
  • NITs/QUESTIONs → carried into §8/§10 TODOs (cost delta, FF_PROFILE_IMPORTS, platform feature availability, key-rotation owner, env-store access logging, cache race-review under autoscaling, atomic webhook repoint).

11. Rollback Story

Pre-cutover: the self-host langgraph up + RDS plan (RDS+Redis+Docker; /sh-plan-review'd to APPROVE-after-revision this session) is the documented fallback if managed is rejected before cutover. Its durable design wins (env-scoped RDS physical names; Credentials.fromGeneratedSecret({secretName:"open-swe-<env>/rds-credentials"}) under the existing instance-role secret prefix → no new IAM; DESTROY-on-rollback RDS-managed secret; derive DATABASE_URI in fetch-config) are captured in the project memory.

Post-cutover (managed is live, AWS still standing during soak): if managed prod fails, rollback = repoint webhooks + DNS back to the AWS stack:

  • GitHub App / Slack / Linear webhook URLs → hooks.seahaven.com
  • openswe.seahaven.com DNS → the seahaven-com ALB (restore from the E1 export)
  • ui/vercel.json rewrite → the AWS dashboard origin (or stop using Vercel)
  • Do NOT run Phase E teardown until the soak passes — the AWS stack IS the rollback target.

Revision-level rollback (within managed): LangGraph Cloud revision rollback (backend) + Vercel instant rollback (UI) cover bad deploys without leaving the platform.


TODOs flagged (couldn't verify in repo)

  • #62 merge state — RESOLVED: merged to dev + deployed (the draft read a stale local checkout; origin/dev and the deployment git_ref_sha a4ed19ba confirm Bedrock+Fireworks).
  • Current AWS run-rate for the cost delta (§8) — not derivable from repo.
  • LangGraph Cloud custom-domain + "manual promotion to production" feature availability (§10.2, §10.3) — platform features, verify in the LangGraph Cloud console/docs at execution time.
  • FF_PROFILE_IMPORTS exact flag name/usage (§5 fix #6) — referenced from session context; confirm the flag exists in agent/ before relying on it.
  • Exact webhook path on *.langgraph.app (whether the custom http.app mounts at root so /webhooks/github is reachable as-is) — proven reachable on the dev spike; re-verify the exact path on prod.