Document the live execution of MIGRATION.md §10.1 (Bedrock auth via static keys): the customer-managed least-privilege policy open-swe-bedrock-invoke and the open-swe-dev-bedrock / open-swe-prod-bedrock IAM users. Captures the static-key deviation rationale (managed LangGraph Cloud cannot assume a role) and that both mandatory gates (GPT-4.1 IAM cross-review, /sh-security-review) passed with no critical/high.
35 KiB
Open SWE — Migration Plan: Self-Hosted AWS → Managed LangGraph Cloud + Vercel
Repo: Sea-Haven-Industries/open-swe (private) · AWS: 328440206208 / us-east-1
Author: Adam Moussa · Date: 2026-06-29 · Status: DRAFT — owes a /sh-plan-review before prod cutover (Phase C gate)
1. Executive Summary
Decision (settled — not re-litigated here)
Move the Open SWE deployment off the bespoke self-hosted AWS stack (stock langgraph dev, in-memory store, no durability — issue #9) onto the managed runtime:
- Backend → LangGraph Cloud (a.k.a. "LangSmith Deployment"): git-connected, auto-builds a revision on push, zero-downtime, revision rollback, durable Postgres-backed store/checkpointer.
- UI → Vercel:
ui/SPA, atomic deploys, instant rollback, per-PR previews (net-new capability).
Why
- Dissolves the licensing blocker. Self-hosted
langgraph upon the LangSmith Plus key is "Self-Hosted Lite" (1M node-exec/yr cap, Elastic License 2.0, production legally ambiguous). Managed IS the licensed product with no node cap — runs billed flat at $0.005, nodes not charged. - Deletes the entire durable-runtime build (RDS/Redis/Docker/AMI) and most of the bespoke AWS CD (S3 release pipeline, SSM deploy docs, packer, CDK box/ALB stacks).
- It's the upstream-canonical deployment and the repo is already wired for it:
ui/vercel.jsonhas the same-origin/dashboard/api/*rewrite;langgraph.jsonis cloud-format (6 graphs +http.app).
Cost
≈ $160/mo incremental for one always-on prod deployment:
- $0.0036/min prod uptime ≈ $155/mo (always-on)
-
- $0.005/run
-
- traces pay-as-you-go above 10k/mo
- The $39/mo Plus seat is already paid (not incremental)
- Dev deployment is FREE (1 included on Plus, preemptible)
A self-hosted langgraph up + RDS plan (already /sh-plan-review'd to APPROVE-after-revision this session) is the documented FALLBACK if managed is ever rejected (see §11).
2. Architecture: Before → After
Before (self-hosted AWS — LIVE as of 2026-06-29)
GitHub/Slack/Linear ──webhook──▶ hooks.seahaven.com ─┐
Browser (dashboard) ─────────────▶ openswe.seahaven.com ─┤
▼
shared seahaven-com ALB (:443, host+path rules)
▼
EC2 (prod i-08a729e50779c4b07 t4g.large ARM64, private subnet)
nginx :80 ──proxy /dashboard/api + /webhooks──▶ langgraph dev :2024 (loopback)
in-memory store (reseeded by seed_store.sh ExecStartPost)
▼
Config: Secrets Manager open-swe-prod/* + SSM /open-swe-prod/*
Release: GitHub Actions → S3 open-swe-prod-assets/releases/* → SSM doc deploy.sh
IaC: CDK OpenSweIamStack + OpenSweDevStack + OpenSweProdStack
Sandbox: LangSmith cloud (DEFAULT_SANDBOX_SNAPSHOT_ID + GitHub proxy)
After (managed)
GitHub/Slack/Linear ──webhook──▶ *.langgraph.app (or hooks.seahaven.com CNAME → TODO §10)
Browser (dashboard) ─────────────▶ open-swe-dashboard.vercel.app (stable alias / custom domain)
│ same-origin rewrite /dashboard/api/* (ui/vercel.json)
▼
LangGraph Cloud "Deployment" (managed, git-connected to `main` for prod / `dev` for dev)
├─ serves the 6 graphs (agent, reviewer, analyzer, chat, scheduler, ci_monitor)
├─ serves the custom http.app (agent.webapp:app = webhooks + dashboard API + OAuth)
├─ durable Postgres store + checkpointer (issue #9 SOLVED) — autoscaled 1→10 replicas
└─ env/secrets in the Deployment config (NOT Secrets Manager) — Adam accepted deviation
▼
Sandbox: LangSmith cloud (UNCHANGED — DEFAULT_SANDBOX_SNAPSHOT_ID + GitHub-App proxy)
Auth: GitHub App seahaven-openswe (UNCHANGED — App 4146115 / Install 142615168)
What changes shape: runtime host (EC2 → managed PaaS), durability (in-memory → managed Postgres), CD (bespoke S3/SSM/packer → git-connected auto-build), config home (Secrets Manager/SSM → Deployment+Vercel env). What stays: the GitHub App, the LangSmith sandbox plane, CI (lint/format/unit/Playwright), the app code itself.
3. Phased Plan
Phase A — Dev spike (MOSTLY DONE)
Goal: prove managed serves our custom app + durability, at 0, before committing prod .
Proven this session:
- ✅ Dev backend deployed to LangGraph Cloud, connected to branch
dev:https://open-swe-dev-hosted-e76c2b0e8a7955fe8ad3110a7a54e5d0.us.langgraph.app(Bedrock + a Fireworks key set; AWS creds for Bedrock deferred — see §10 open decision). - ✅ UI deployed to Vercel — team
sea-haven, projectopen-swe-dashboard,https://open-swe-dashboard.vercel.app;ui/vercel.jsonrewrite repointed at the dev deployment. - ✅ Managed serves the custom
http.app(dashboard API + webhooks) — no platform auth gate in front of our routes (webhook returns 401 sig-enforced, so signatures still govern). - ✅ Vercel same-origin rewrite → backend works.
- ✅ GitHub OAuth dashboard login end-to-end.
- ✅ DURABILITY — team-default model survived a revision redeploy (this is issue #9's core acceptance goal, proven on managed).
Remaining Phase A items (the GitHub @openswe run-trigger loop):
- A1. Get an
@opensweGitHub comment to dispatch a run end-to-end on dev. Blocked by the user-mapping + cache gotchas (see fixes #4 and #5 in §5). - A2. Create the owner user-mapping in the managed Store (namespace
["user_mappings"], key = lowercased loginamoussa1229, value{github_login, work_email}) — the Admin UI can't setwork_emailyet (fix #4). - A3. Verify a freshly-added mapping is seen without a redeploy across managed's multi-replica autoscaling (fix #5).
- A4. Confirm
LANGGRAPH_URLis set to the deployment URL on the dev deployment solanggraph_client()calls resolve (fix #1 — config now, code later).
Phase B — Harden (the code fixes + env codification)
Land the 6 code fixes (§5), codify env/config, and resolve the Bedrock-auth decision before standing up prod.
- B1. Land code fixes #1–#6 (§5) as a PR into
dev(or split into focused PRs). Re-runmake lint+make test. - B2. Codify the env contract for managed. Update
ui/vercel.jsonrewrite to the prod deployment URL (Phase C) but keep dev pointing at dev. Add a documented LangGraph-Cloud env list to the repo (see §4) — NOT secrets, just the variable inventory + which are excluded. - B3. Resolve the Bedrock-auth decision (§10 open decision): static AWS keys for a Bedrock-scoped IAM user in the deployment env, OR run the agent on Fireworks. This intersects the #62 model work.
- B4.
Merge #62 first— DONE / N/A. #62 (Bedrock + Fireworks) IS merged todevand is what the dev deployment runs — verified onorigin/dev(options.py→DEFAULT_MODEL_ID = "bedrock_converse:us.anthropic.claude-opus-4-8") and confirmed by the live deployment'sgit_ref_sha: a4ed19ba…(= the #62 merge commit). The contrary note came from a STALE local checkout; no merge action needed. - B5. Audit ALL in-process caches for the single-process → multi-replica assumption (generalization of fix #5; see §5).
- B6. Measure + reduce custom-app import time (fix #6) so the deployment isn't flagged unhealthy / slow to scale.
Phase C — Prod deployment
- C1. Create a prod LangGraph Cloud deployment tracking branch
main(the durable autoscaled 1→10 tier, not the free Dev tier). Record its*.langgraph.appURL. - C2. Create the prod Vercel project/target (or promote the existing
open-swe-dashboardto production); set its env (same-origin mode:VITE_DASHBOARD_API_BASE_URLempty). - C3. Set the prod env triad (§4) on the prod deployment + Vercel:
LANGGRAPH_URL= the prod*.langgraph.appURLDASHBOARD_BASE_URL+DASHBOARD_API_BASE_URL= the prod Vercel origin (withhttps://scheme — fix #2)
- C4. Establish the prod approval gate (§6): git-connected auto-deploy needs an explicit gate. Preserve the prod manual-approval that exists today (the GitHub
prodEnvironment reviewer). Mechanism = protectedmainruleset + the platform's "require manual promotion to production" if available (§10 open decision). - C5. Repoint webhooks + OAuth URLs to prod:
- GitHub App
seahaven-openswewebhook URL → prod backend webhook URL (*.langgraph.app/webhooks/githuborhooks.seahaven.com— §10 custom-domain decision) - Slack Event Subscriptions + Interactivity Request URLs → prod
- Linear webhook URL → prod (if/when wired)
- GitHub App OAuth callback →
<DASHBOARD_API_BASE_URL>/dashboard/api/auth/callback(the Vercel prod origin; fixes #2, #3)
- GitHub App
- C6. Custom-domain decision (§10): dashboard custom domain solved by Vercel; webhooks
hooks.seahaven.comeither re-point to*.langgraph.appdirectly or front via CNAME. Managed issues*.langgraph.appURLs ONLY (no documented custom-domain support on the backend).
Phase D — Cutover + verify
- D1. Run the issue #9 acceptance criteria on PROD: durable store survives a revision redeploy; team defaults / user mappings persist; a paused/in-flight run survives a redeploy (the durability win self-host never had).
- D2. End-to-end prod smoke:
@opensweGitHub comment → run dispatch → sandbox → draft PR → reply in source channel. - D3. Dashboard prod smoke: GitHub OAuth login on the stable Vercel alias (fix #3); admin pages reachable for
CONFIGURED_ADMINS. - D4. Webhook sig-enforcement smoke: unsigned POST to each
/webhooks/*returns 401; signed returns 200. - D5. Soak for an agreed window (recommend 3–7 days) with the AWS prod stack still standing as the rollback target (§11) before any teardown.
Phase E — AWS decommission (AFTER soak)
Only after managed prod is proven + soaked. See §7 for the precise retire-vs-keep list. Phased teardown:
- E1. Export the live ALB listener-rule / target-group config for
open-swe-prodandopen-swe-devto JSON (the export is the rollback source — these rules were partly hand-built). - E2. Disable the bespoke CD (build-artifacts / cd-infra workflows) so nothing re-deploys the EC2 boxes.
- E3.
cdk destroy OpenSweProdStackthenOpenSweDevStack(mind the RETAIN secret shells — they survive and hold the globalopen-swe-<env>/*names; force-delete only EMPTY shells once env is fully migrated; keep populated ones transitionally as the env source — §7). - E4. Remove ALB rules/target groups, S3 assets buckets, SSM deploy docs, and the per-env CDK bootstrap qualifier
oswedev(CDKToolkit-oswedev) — once nothing references them. - E5. Decide DNS: repoint
openswe.seahaven.com→ Vercel; resolvehooks.seahaven.com→*.langgraph.app(or CNAME). - E6. Leave
OpenSweIamStackuntil last — the promotion App + any retained roles may still be referenced.
Phase F — Docs / memory
- F1. Rework Confluence "AWS Architecture Map" (page 1540098) + the open-swe child page (26116098): the RDS/EC2/ALB subgraph is replaced by an external-services view (LangGraph Cloud + Vercel + the retained GitHub App + LangSmith sandbox).
- F2. Update
README.md+INSTALLATION.md(§10 is already canonical-correct; align Sea Haven specifics). - F3. Update project memories
project_open_swe_migration.md+reference_open_swe_deployment.mdto mark the managed cutover and the AWS teardown.
4. Env / Config Reference
The prod triad (per INSTALLATION.md §10, lines 630–658)
| Var | Prod value | Notes |
|---|---|---|
LANGGRAPH_URL |
the deployment URL (https://...langgraph.app) |
NOT localhost. Drives thread_ops.langgraph_url() (fix #1). |
DASHBOARD_BASE_URL |
the Vercel origin | same-origin rewrite mode |
DASHBOARD_API_BASE_URL |
the Vercel origin, with https:// scheme |
scheme required or OAuth redirect_uri is schemeless and GitHub rejects (fix #2) |
VITE_DASHBOARD_API_BASE_URL |
empty | same-origin mode; UI calls relative /dashboard/api/*, Vercel rewrites to backend |
GitHub App dashboard OAuth callback = <DASHBOARD_API_BASE_URL>/dashboard/api/auth/callback (the Vercel prod origin).
Where secrets/env live now
LangGraph Cloud Deployment config + Vercel env — NOT AWS Secrets Manager. This deviates from the Sea Haven secrets-and-config.md "Secrets Manager for all sensitive" handbook rule — Adam ACCEPTED this deviation (managed has no instance role / no fetch-config boot hook; the platform's own secret store is the mechanism).
Sourcing env from the existing AWS stack (transitional)
Pull from the existing open-swe-dev Secrets Manager + SSM via a documented CLI dump:
aws secretsmanager list-secrets --filters Key=name,Values=open-swe-dev/then per-nameget-secret-value— NOTbatch-get-secret-value(its pagination silently drops values past page 1; this bit the box twice — see memory).- SSM:
aws ssm get-parameters-by-path --path /open-swe-dev/.
EXCLUDE when copying to managed
| Category | Vars |
|---|---|
| Box-/self-host-specific | LANGGRAPH_URL (set fresh to deployment URL), LANGGRAPH_URL_PROD, LANGSMITH_ENDPOINT, LANGSMITH_ENDPOINT_PROD, LANGSMITH_URL_PROD, LANGSMITH_TENANT_ID_PROD, LANGSMITH_HOST_API_URL, LANGCHAIN_REVISION_ID |
| Dropped providers (PR #62) | OPENAI_API_KEY, GOOGLE_API_KEY, GROQ_API_KEY (these 6 Secrets Manager shells deleted 2026-06-29, final purge 2026-07-06) |
| Unused sandbox providers | DAYTONA_API_KEY, RUNLOOP_API_KEY |
| Eval-only | JUDGE_ANTHROPIC_API_KEY (kept on AWS for evals/reviewer/judge.py, not needed in the runtime deployment) |
KEY mappings on managed
LANGSMITH_API_KEY= theLANGSMITH_API_KEY_PRODvalue (andLANGCHAIN_API_KEY= same).- The full secret inventory is
SECRET_VARSininfra/lib/constructs/config-store.ts:46(28 shells) — use it as the checklist of what to carry, minus the EXCLUDE rows above.
Models (post-#62) + Bedrock auth
Post-#62 SUPPORTED_MODELS = bedrock_converse:us.anthropic.claude-opus-4-8 (DEFAULT) + 3 Fireworks models; no direct anthropic: option.
✅ #62 is merged to dev and deployed — verified on origin/dev (options.py → bedrock_converse default) and the live deployment git_ref_sha a4ed19ba. (The pre-#62 values appear only on a stale LOCAL checkout — ignore.)
Bedrock on managed has NO EC2 instance role (the #62 IAM design used the EC2 instance role as the Bedrock principal — that breaks on managed). Options (OPEN DECISION, §10):
- (a) Static
AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY/AWS_REGIONfor a Bedrock-scoped IAM user in the deployment env, or - (b) Run the agent on Fireworks and avoid Bedrock entirely on managed.
Bedrock IAM users (CREATED — resolves §10.1, discharges the §9 IAM gates)
Option (a) was chosen and executed live (2026-06-30, account 328440206208 / us-east-1). This is click-ops IAM — there is no remaining open-swe AWS IaC after the decommission (#64), so these are created with the CLI, not CDK.
- Customer-managed policy
open-swe-bedrock-invoke(arn:aws:iam::328440206208:policy/open-swe-bedrock-invoke). Least-privilege: actionsbedrock:InvokeModel+bedrock:InvokeModelWithResponseStreamonly, scoped to exactly theus.anthropic.claude-opus-4-8inference-profile ARN + its three routed foundation-model ARNs (us-east-1, us-east-2, us-west-2). No wildcards, no other models. (Supersedes the §10.1 plan to reuse the #62 instance-role policy — a fresh standalone policy was minted instead.) - Two IAM users, each attached to that policy:
open-swe-dev-bedrockandopen-swe-prod-bedrock. Taggedproject=open-swe,managed-by=cli-migration,purpose=bedrock-invoke. - Static access keys are minted separately by the owner (
aws iam create-access-key) — secret keys live only in the deployment env stores, never in this repo or memory. dev key → dev env; prod key → LangGraph Cloud prod config (AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY/AWS_REGION=us-east-1). - Deviation rationale: static long-lived keys are a deliberate departure from the Sea Haven OIDC norm because managed LangGraph Cloud cannot assume an AWS role. Mitigated by the tight least-privilege policy above.
- ✅ Both mandatory gates ran and PASSED:
- GPT-4.1 IAM cross-review (confirmed hit
gpt-4.1-2025-04-14) — least-privilege confirmed. /sh-security-review— 0 critical/high; two accepted mediums (the static-key deviation + no per-principal budget cap).
- GPT-4.1 IAM cross-review (confirmed hit
- Recommended follow-ups: shortest viable key-rotation cadence with recorded creation dates; an AWS Budgets / CloudWatch anomaly alarm on per-principal Bedrock
InvokeModelvolume; confirm CloudTrail captures these users.
5. Required Code Fixes (the 6 gotchas)
These are self-host → managed assumption breaks discovered on the spike. Each gets a config workaround now and a code fix to land in Phase B.
Fix #1 — langgraph_url() localhost fallback
File: agent/utils/thread_ops.py:38-41
def langgraph_url() -> str:
return os.environ.get("LANGGRAPH_URL") or os.environ.get(
"LANGGRAPH_URL_PROD", "http://localhost:2024"
)
On managed, every langgraph_client() call (thread sidebar, run creation, the same pattern in agent/dashboard/user_mappings.py:43 get_client() with no URL) ConnectErrors unless LANGGRAPH_URL is set to the deployment URL.
- Config fix (now): set
LANGGRAPH_URL= the deployment URL on every deployment (Phase A4 / C3). - Code fix (Phase B): default to the in-process / deployment URL rather than
http://localhost:2024when running inside a managed deployment (e.g. honor the platform's own URL env, or fail loud instead of silently hitting localhost).
Fix #2 — OAuth DASHBOARD_API_BASE_URL must include scheme
Where: OAuth redirect construction (dashboard oauth.py / routes.py), driven by DASHBOARD_API_BASE_URL.
A schemeless value yields a schemeless redirect_uri → GitHub rejects with "redirect_uri not associated."
- Fix: always set
DASHBOARD_API_BASE_URLwithhttps://(Phase C3). Code hardening (Phase B): assert/normalize a scheme at startup and fail loud if missing.
Fix #3 — OAuth state cookie is host-only
Cookie: osw_oauth_state (host-only). Login must START on the SAME host as DASHBOARD_API_BASE_URL — the stable Vercel alias, never the immutable per-deploy URL — or you get "oauth state mismatch."
- Fix: pin
DASHBOARD_API_BASE_URL+ the login entry point to the stable alias / custom domain (Phase C3, D3). Document this so a future per-deploy preview URL isn't used for login.
Fix #4 — Admin "User mappings" UI can't set work_email
Files: the Admin → User mappings UI (ui/) + agent/dashboard/user_mappings.py (upsert_mapping at line 257 takes work_email but the UI form doesn't expose it).
Mappings can't be fully created from the dashboard → must write the Store directly: namespace ["user_mappings"], key = lowercased login, value {github_login, work_email} (matching _index_record at user_mappings.py:73-84, which keys _by_login on login.lower()).
- Fix (Phase B): add the
work_emailfield to the Admin mappings form so mappings are fully creatable from the UI.
Fix #5 — GitHub webhook path doesn't refresh the user-mapping cache (multi-replica break)
Files: agent/webapp.py:3052 (GitHub path) vs agent/webapp.py:1091 (Slack path).
The Slack path refreshes before lookup:
# agent/webapp.py:1089-1093 (Slack)
await refresh_user_mapping_cache()
...
The GitHub path does not — it calls email = await email_for_login(github_login) (webapp.py:3052, again at :3331) cold. The cache (user_mappings.py _ensure_cache_loaded, line 197) is one-shot per process (_cache_loaded flag). On self-host single-process this was fine; on managed's multi-replica autoscaling, a freshly-added mapping isn't seen by a replica whose cache loaded earlier — until restart.
- Fix: refresh-before-lookup on the GitHub path (mirror the Slack path), or add a TTL / cross-replica invalidation to the cache.
- Generalize (Phase B5): audit ALL in-process caches for the single-process → multi-replica assumption —
SANDBOX_BACKENDSdict (agent/utils/sandbox_state.py),_THREAD_RUN_LOCKS(thread_ops.py:18),_by_login/_by_email/_by_slack_id(user_mappings.py:67-69). Sandbox affinity is already thread-keyed + persisted in thread metadata (sandbox_id), so it's the cache/lock state that needs the multi-replica review.
Fix #6 — Slow custom-app import (~8s startup)
Symptom: "exceeded expected startup time" → risks the deployment being marked unhealthy / slow to scale out.
- Fix (Phase B6): lazy imports / reduce import-time work in
agent/webapp.pyand the graph factories. Profile withFF_PROFILE_IMPORTS(the import-profiling flag) to find the heavy modules.
6. CD / Ops Changes
Retires (bespoke pipeline)
- GitHub Actions build → S3 releases → SSM doc →
deploy.sh→ health-gate (.github/scripts/,deploy/ami/deploy.sh) roll-box.sh/publish-and-deploy.sh/rollback.sh- AMI baking (packer
deploy/ami/open-swe-base.pkr.hcl) +cdk.context.jsonAMI pin cd-infra.ymlCDK deploys + OIDC bootstrap qualifiersfetch-config.sh/seed_store.sh(boot-time config materialization + Store reseed)
Replaced by git-connected PaaS
- LangGraph Cloud auto-builds a revision on push (first-party zero-downtime + revision rollback).
- Vercel builds / atomic-deploys / instant-rollback + per-PR previews (net-new capability the AWS stack never had).
Stays
- CI (lint / format / unit / Playwright E2E) — still matters and still gates merges.
make lint,make test.
Gate re-homing
- The dev → main promotion gate (
check-dev-green.sh+ protected-mainruleset 18238334) re-homes: the dev deployment tracksdev, the prod deployment tracksmain, so the gate governs exactly what reaches prod. - Today's promotion machinery:
promote-dev-to-prod.yml+ theseahaven-promotionGitHub App (app_id 4170147, in ruleset 18238334 bypass_actors) FF-pushesdev → main. Under managed, a push tomainis what triggers the prod build — so the promotion gate IS the prod deploy gate.
Preserve the PROD manual-approval gate (LOAD-BEARING)
Today the prod GitHub Environment (required reviewer amoussa1229) is the manual approval. Git-connected auto-deploy removes the CD job that consulted that Environment, so the approval must be re-established explicitly:
- Mechanism options (OPEN DECISION §10): protected
main(only the promotion App can FF-push, and that push is itself the gate) AND/OR the platform's "require manual promotion to production" if LangGraph Cloud exposes it. - Norm shift: deploy-then-merge → merge-to-deploy. Lean on the dev deployment + Vercel previews to verify before promoting
dev → main. (This inverts the Sea Haven handbook deploy-then-merge default — call it out in the README + handbook note.)
7. AWS Decommission — Retire vs Keep
RETIRE (becomes vestigial once managed prod is live + soaked)
| Item | Path / resource |
|---|---|
| AMI bake + cloud-init | deploy/ami/ (open-swe-base.pkr.hcl, provision.sh, user-data.sh, deploy.sh, templates/) |
| Self-host boot/config scripts | deploy/seahaven/fetch-config.sh, seed_store.sh, put-config.sh, nginx/openswe.conf, systemd/open-swe.service, aegra/ (already deferred) |
| CDK box/ALB/AMI stacks | infra/lib/open-swe-stack.ts, constructs/app-service.ts, assets-bucket.ts, ami-cache.ts, instance-role.ts, github-deploy-roles.ts |
| EC2 instances | prod i-08a729e50779c4b07 (t4g.large), dev i-0a3bb8e0ddd36c29b |
| S3 release buckets | open-swe-dev-assets, open-swe-prod-assets |
| SSM deploy docs | open-swe-dev-deploy, open-swe-prod-deploy |
| ALB plumbing | open-swe-prod-tg / open-swe-dev-tg target groups + the seahaven-com ALB host/path listener rules for openswe/hooks |
| Per-env bootstrap qualifier | dev oswedev (CDKToolkit-oswedev) |
| RDS / Redis durable plan | never built — fully dropped (managed provides durability) |
| Bespoke CD workflows | cd-infra.yml, build-artifacts.yml, roll-box.sh/publish-and-deploy.sh/rollback.sh |
KEEP
| Item | Why |
|---|---|
GitHub App seahaven-openswe (App 4146115 / Install 142615168) |
unchanged auth + webhook source; only the webhook/OAuth URLs repoint |
LangSmith sandbox setup incl. DEFAULT_SANDBOX_SNAPSHOT_ID (dc36e509-d4d2-4efc-8a4e-61f74f3446ec) |
the only sandbox with working in-sandbox git/gh auth (_configure_github_proxy); managed default SANDBOX_TYPE=langsmith uses it |
Secrets Manager open-swe-{dev,prod}/* values |
the source you copy env FROM, at least transitionally — do NOT force-delete populated shells until managed env is fully cut over and verified |
| DNS | repoint openswe.seahaven.com → Vercel; decide hooks.seahaven.com → *.langgraph.app or a CNAME (§10) |
seahaven-promotion App + main/dev rulesets (18238334 / 18238542) |
still govern dev → main promotion = the prod deploy gate |
evals/ + JUDGE_ANTHROPIC_API_KEY |
eval pipeline unaffected (separate from runtime) |
Phased teardown rule: remove AWS infra ONLY after managed prod is proven + soaked (Phase D5). The exported ALB JSON (E1) is the rollback source for the hand-built listener rules.
8. Cost
| Line item | Monthly |
|---|---|
| Prod deployment uptime ($0.0036/min, always-on) | ≈ $155 |
| Runs ($0.005/run) | usage-based |
| Traces above 10k/mo | pay-as-you-go |
| LangSmith Plus seat ($39) | already paid (not incremental) |
| Dev deployment | $0 (1 free on Plus, preemptible) |
| Incremental total | ≈ $160/mo |
Compare to the retired AWS stack: 2× EC2 (t4g.large prod + t4g.medium/large dev), 2× S3 buckets, ALB share, NAT egress, plus the never-built RDS/Redis durable tier the self-host path would have added. Managed nets out roughly cost-neutral-to-cheaper while adding durability + previews. TODO: confirm current AWS run-rate for an apples-to-apples delta (not verified in repo).
9. Sea Haven Gates & Obligations
/sh-plan-review(GPT-4.1cross_reviewer) on THIS plan BEFORE prod cutover (Phase C gate). ⚠️run.pymisroutes reviewer-framed prompts to the no-opdoneroute — callmodels.get_cross_reviewer()directly (load orchestrator.env,ChatOpenAI("gpt-4.1")) perfeedback_orchestrator_usage./sh-security-reviewon the sensitive surface: this migration touches auth/webhook signature verification (the customhttp.appis now publicly reachable on*.langgraph.appwith no platform gate in front) and secrets handling (secrets move from Secrets Manager to the Deployment/Vercel env). Required, not opt-in — resolve confirmed critical/high before cutover.- GPT-4.1 cross-family review if the Bedrock-auth fix changes IAM (new Bedrock-scoped IAM user + policy = IAM change → mandatory cross-review).
- Confluence — rework "AWS Architecture Map" (1540098) + open-swe child page (26116098): replace the RDS/EC2/ALB subgraph with an external-services view (LangGraph Cloud + Vercel + retained GitHub App + LangSmith sandbox). Do it in the same conversation as the cutover, not as a follow-up.
- README + INSTALLATION updated (INSTALLATION §10 already canonical; align Sea Haven specifics + the merge-to-deploy norm shift).
- Memories — update
project_open_swe_migration.md+reference_open_swe_deployment.md.
10. Decisions — RESOLVED (2026-06-29, Adam)
- Bedrock auth on managed → STATIC AWS KEYS. Create a Bedrock-scoped IAM user; set
AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY/AWS_REGION=us-east-1in the deployment env; reuse the #62 instance-role Bedrock policy (bedrock:InvokeModel[WithResponseStream]on theus.anthropic.claude-opus-4-8inference-profile ARN + per-region foundation-model ARNs in us-east-1/2 + us-west-2). Caveat: long-lived keys → rotation reminder + keep the policy tightly scoped. New IAM user+policy = mandatory GPT-4.1 IAM cross-review +/sh-security-reviewbefore merge. - Webhook domain → RAW
*.langgraph.app(the "easier" path — no CNAME/DNS/TLS). Point GitHub/Slack/Linear webhook URLs straight at the deployment. Dashboard gets a Vercel custom domain (e.g.openswe.seahaven.com) since that's trivial on Vercel. - Prod approval gate → GATED PROMOTE-TO-MAIN WORKFLOW. Re-home
promote-dev-to-prod.yml: gate the promote job on theprodGitHub Environment reviewer (amoussa1229); on approval it FF-pushesdev → main; the push triggers the managed prod build. Preserves today's manual-approval semantics — approval is on the promote job, the merge tomainis the deploy trigger. - Cost/trace monitoring → YES, in LangSmith (usage/trace budget + alert in the LangSmith workspace).
- Dev/prod home → CONFIRMED: prod = LangChain managed (LangGraph Cloud) + Vercel UI; dev = managed free Dev tier (same runtime as prod).
#62 merge status— RESOLVED (merged + deployed).
10b. Execution approach — 2-PR SPLIT (Adam's call, 2026-06-29, post-plan-review)
The plan-review BLOCKed a literal single PR and recommended a 3-PR minimum (isolate cache fix #5 + the promotion-gate workflow). Adam chose a 2-PR split — isolating the prod-deploy-gate change (highest risk to prod), bundling the rest:
- PR1 — promotion-gate workflow + ruleset. Re-home
promote-dev-to-prod.ymlto the gated promote-to-main (decision 3), and lockmainso ONLY the promote App can push (BLOCK 3). Isolated because it changes how prod deploys — must be independently reviewable/revertible. - PR2 — the 6 code fixes (§5) +
ui/vercel.jsonprod repoint + README/INSTALLATION. (Accepted deviation from the reviewer's 3-PR rec: the multi-replica cache fix #5 stays bundled in PR2 rather than isolated — Adam's call.) - NOT in either PR (rollback safety): AWS infra-code deletion (
deploy/ami,deploy/seahaven,infra/self-host stack) +cdk destroystay a POST-SOAK cleanup (Phase E). - Not PR content (operational, sequenced around the merges): LangGraph Cloud prod deployment + env, Vercel prod + custom domain, webhook/OAuth repoint, Bedrock IAM user, LangSmith budget, Confluence/memory.
10c. Plan-review resolution (GPT-4.1 cross_reviewer, 2026-06-29) — verdict REQUEST CHANGES, all BLOCKs addressed
- BLOCK 1 (single PR) → addressed via the 2-PR split above (conscious deviation: cache fix #5 not isolated — Adam accepted).
- BLOCK 2 (raw LangGraph API public?) → EMPIRICALLY RESOLVED. Unauthenticated probes of the dev deployment returned 403 "Missing authentication headers" for
/threads/search,/assistants/search,/store/items;/ok=200,/dashboard/api/me=401. The platform gates the raw control-plane API. Keep as a pre-prod gate (Phase D4): re-probe the PROD deployment URL before cutover. - BLOCK 3 (prod gate enforcement) → PR1 must lock
mainso only the promote App can push (verify ruleset 18238334: no direct-push path, PR-required, promote App is the sole FF bypass). Any non-gated push tomainwould auto-deploy prod. - BLOCK 4 (secrets posture) → add to Phase B/E: rotate all secrets after migrating them to the platform env stores and BEFORE deleting from AWS; document who can read/write the LangGraph Cloud + Vercel env (access audit); delete AWS shells only post-cutover; no dual-homed/stale secrets; record in memory.
- FIX (rollback integrity) → Phase E checklist: "no IaC/DNS/config the rollback needs is altered in PR1/PR2 or during cutover."
- FIX (gates resolved-before-merge) → make explicit: do NOT merge PR1/PR2 until the GPT-4.1 IAM cross-review (Bedrock IAM user) +
/sh-security-review(auth/webhook/secrets surface) are resolved with no critical/high; Confluence (1540098 + 26116098) + memory updated in the cutover conversation. - NITs/QUESTIONs → carried into §8/§10 TODOs (cost delta,
FF_PROFILE_IMPORTS, platform feature availability, key-rotation owner, env-store access logging, cache race-review under autoscaling, atomic webhook repoint).
11. Rollback Story
Pre-cutover: the self-host langgraph up + RDS plan (RDS+Redis+Docker; /sh-plan-review'd to APPROVE-after-revision this session) is the documented fallback if managed is rejected before cutover. Its durable design wins (env-scoped RDS physical names; Credentials.fromGeneratedSecret({secretName:"open-swe-<env>/rds-credentials"}) under the existing instance-role secret prefix → no new IAM; DESTROY-on-rollback RDS-managed secret; derive DATABASE_URI in fetch-config) are captured in the project memory.
Post-cutover (managed is live, AWS still standing during soak): if managed prod fails, rollback = repoint webhooks + DNS back to the AWS stack:
- GitHub App / Slack / Linear webhook URLs →
hooks.seahaven.com openswe.seahaven.comDNS → theseahaven-comALB (restore from the E1 export)ui/vercel.jsonrewrite → the AWS dashboard origin (or stop using Vercel)- Do NOT run Phase E teardown until the soak passes — the AWS stack IS the rollback target.
Revision-level rollback (within managed): LangGraph Cloud revision rollback (backend) + Vercel instant rollback (UI) cover bad deploys without leaving the platform.
TODOs flagged (couldn't verify in repo)
#62 merge state— RESOLVED: merged todev+ deployed (the draft read a stale local checkout;origin/devand the deploymentgit_ref_sha a4ed19baconfirm Bedrock+Fireworks).- Current AWS run-rate for the cost delta (§8) — not derivable from repo.
- LangGraph Cloud custom-domain + "manual promotion to production" feature availability (§10.2, §10.3) — platform features, verify in the LangGraph Cloud console/docs at execution time.
FF_PROFILE_IMPORTSexact flag name/usage (§5 fix #6) — referenced from session context; confirm the flag exists inagent/before relying on it.- Exact webhook path on
*.langgraph.app(whether the customhttp.appmounts at root so/webhooks/githubis reachable as-is) — proven reachable on the dev spike; re-verify the exact path on prod.