mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 20:53:15 +00:00
* Adopt upstream modular webhook skeleton (#1621) Apply the durable-interrupt-dispatch refactor: split the monolithic webapp.py into a thin routing layer plus per-source handlers in webhooks/{github,slack,linear}.py, and add completion.py, dispatch.py, and reconcile.py. Reconcile fork divergence by keeping the Bedrock/ Fireworks cross-provider fallback, the no-agent-attribution prompt policy, the dashboard-handoff re-export, and the Slack channel-info cache. ci_autofix is restored on the new dispatch model in a later commit. Refs: #80 * Port fork webhook security delta onto modular handlers Re-apply the fork's security customizations that #1621 did not carry: Linear webhook replay protection (freshness window on the signed webhookTimestamp), per-repo token-cache binding threaded through the thread token resolvers, the INTERNAL_BOT_LOGINS self-check in the review-finding-reply path, and a user-mapping cache refresh before email resolution on the issue and PR-comment paths (multi-replica staleness). Existing fork security tests pass unchanged. Refs: #80 * Restore CI auto-fix on the modular dispatch model Bring back ci_autofix.py and the ci_monitor graph that #1621 deleted, re-wiring the fork's security-reviewed PR-babysitting onto the new structure: the CI-event, autofix-toggle, and review-feedback handlers move into webhooks/github.py and the github_webhook router re-gains the check_run/check_suite/workflow_run/status routing plus the autofix command and actionable-review branches. Auto-fix runs now dispatch through dispatch_agent_run (durability + completion webhook) while keeping the deliberate batch-while-busy skip-rule via get_thread_active_status. Restore langgraph.json's ci_monitor entry and the fork autofix tests (dispatch mock + import paths re-pointed). Refs: #80 * Reformat and update docs for the modular webhook split Point CLAUDE.md and deploy/MIGRATION.md at the new webhooks/ modules and the dispatch/completion/reconcile contract, and mark the user-mapping cache-refresh fix as applied on the GitHub handlers. Refs: #80 * Restore reject backstop for autofix dispatch A burst of near-simultaneous CI events for one head SHA can slip past the busy-check before the dedupe SHA is recorded, so dispatch the autofix path with multitask_strategy=reject (dev's prior platform default) to drop duplicate concurrent creates instead of letting them interrupt each other. Also make the completion failure-reply dedup claim-then-post and drop the unreachable interrupted branch. --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
454 lines
45 KiB
Markdown
454 lines
45 KiB
Markdown
# Open SWE — Migration Plan: Self-Hosted AWS → Managed LangGraph Cloud + Vercel
|
||
|
||
**Repo:** `Sea-Haven-Industries/open-swe` (private) · **AWS:** 328440206208 / us-east-1
|
||
**Author:** Adam Moussa · **Date:** 2026-06-29 (final topology added 2026-06-30) · **Status:** EXECUTED — managed cutover live; §§3–11 below are the original (now-historical) phased plan, **superseded by §1a for all current-state facts (URLs, project layout, env)**.
|
||
|
||
> **READ §1a FIRST.** The phased plan (§§2–11) and the Phase A/C spike notes capture how we got here and still hold for rationale, cost, and rollback. But the spike-era specifics they cite — the single `open-swe-dashboard` Vercel project, the `open-swe-dev-hosted-…`/`open-swe-v3-…` deployment URLs, the `ui/vercel.json` same-origin rewrite, the single GitHub App — are **stale**. §1a is the authoritative final topology and wins on every conflict.
|
||
|
||
---
|
||
|
||
## 1. Executive Summary
|
||
|
||
### Decision (settled — not re-litigated here)
|
||
Move the Open SWE deployment off the bespoke self-hosted AWS stack (stock `langgraph dev`, in-memory store, no durability — issue **#9**) onto the **managed runtime**:
|
||
|
||
- **Backend → LangGraph Cloud** (a.k.a. "LangSmith Deployment"): git-connected, auto-builds a revision on push, zero-downtime, revision rollback, durable Postgres-backed store/checkpointer.
|
||
- **UI → Vercel**: `ui/` SPA, atomic deploys, instant rollback, per-PR previews (net-new capability).
|
||
|
||
### Why
|
||
1. **Dissolves the licensing blocker.** Self-hosted `langgraph up` on the LangSmith **Plus** key is "Self-Hosted Lite" (1M node-exec/yr cap, Elastic License 2.0, production legally ambiguous). Managed **IS** the licensed product with **no node cap** — runs billed flat at $0.005, nodes not charged.
|
||
2. **Deletes the entire durable-runtime build** (RDS/Redis/Docker/AMI) **and most of the bespoke AWS CD** (S3 release pipeline, SSM deploy docs, packer, CDK box/ALB stacks).
|
||
3. **It's the upstream-canonical deployment** and the repo is **already wired for it**: `ui/vercel.json` has the same-origin `/dashboard/api/*` rewrite; `langgraph.json` is cloud-format (6 graphs + `http.app`).
|
||
|
||
### Cost
|
||
≈ **$160/mo incremental** for one always-on prod deployment:
|
||
- $0.0036/min prod uptime ≈ **$155/mo** (always-on)
|
||
- + **$0.005/run**
|
||
- + traces pay-as-you-go above 10k/mo
|
||
- The **$39/mo Plus seat is already paid** (not incremental)
|
||
- **Dev deployment is FREE** (1 included on Plus, preemptible)
|
||
|
||
A self-hosted `langgraph up` + RDS plan (already `/sh-plan-review`'d to APPROVE-after-revision this session) is the documented **FALLBACK** if managed is ever rejected (see §11).
|
||
|
||
---
|
||
|
||
## 1a. Final, verified topology (AUTHORITATIVE — supersedes spike-era values)
|
||
|
||
This is the live managed deployment as of 2026-06-30. Where any later section disagrees (old URLs, a single Vercel project, a single GitHub App, the `ui/vercel.json` rewrite), **this section wins**.
|
||
|
||
### Backend — managed LangGraph Cloud (two deployments, one LangSmith workspace)
|
||
Both deployments live in the **same LangSmith workspace**; the **same workspace API key authenticates both** (including the Store API — so per-deployment store writes use that one key with the per-deployment URL).
|
||
|
||
| Deployment | URL | Git connection |
|
||
|---|---|---|
|
||
| **dev** | `https://open-swe-dev-fb737aa219605c8bbdb30ecbb33f30c0.us.langgraph.app` | branch `dev` |
|
||
| **prod** | `https://open-swe-prod-d6c7bb63aaa651b6a1d92f9492b1d983.us.langgraph.app` | branch `main` (auto-deploys on push to `main`) |
|
||
|
||
> The dev deployment was **renamed `open-swe-dev`, deleted, and recreated** — which minted the **new URL hash** above. The spike-era `open-swe-v3-…` / `open-swe-dev-hosted-…` URLs are **dead/superseded**. Deleting + recreating a deployment is the one operation that changes the URL hash (otherwise stable across revisions) — when it happens, update **every** reference (Vercel env, GitHub App webhooks, OAuth callbacks, docs).
|
||
|
||
### UI — Vercel (ONE project, two environments)
|
||
**One** Vercel project `open-swe-prod` (team `sea-haven`, id `prj_OOh6yjXMp4ah3Ws3Y7XRQxjmMmQU`). The old separate `open-swe-dashboard` project was **DELETED**.
|
||
|
||
| Vercel environment | Branch | Backend | Custom domain |
|
||
|---|---|---|---|
|
||
| production | `main` | prod deployment URL | `openswe.seahaven.com` |
|
||
| custom **`dev`** (id `env_SMI23PULAJXk0GhwE0HLVhp5J3ZS`) | `dev` | dev deployment URL | `openswe-dev.seahaven.com` |
|
||
|
||
- A **per-environment** env var `LANGGRAPH_BACKEND_URL` (prod env = prod URL, dev env = dev URL) drives the `/dashboard/api/*` proxy.
|
||
- Project settings: `framework=null`, `outputDirectory` cleared, root directory `ui`.
|
||
- **Proxy mechanism (current, after PR #76):** **Nitro `routeRules`** in `ui/vite.config.ts` read `process.env.LANGGRAPH_BACKEND_URL` and Nitro's Vercel preset compiles them into `.vercel/output/config.json` (Build Output API) at build time — a CDN-level proxy (not redirect, so the `osw_session` cookie stays first-party). PR #75's hand-rolled `ui/scripts/build-vercel-output.mjs` was the broken first attempt and is **gone** (`ui/scripts/` no longer exists). **Never** add a manual script that `rm`s `.vercel/output` — Nitro's Vercel preset auto-emits it.
|
||
|
||
### DNS — Route 53 zone `seahaven.com` (`Z06652411XKH89KTZD3XA`)
|
||
- `openswe.seahaven.com` → CNAME `cname.vercel-dns.com` (prod env)
|
||
- `openswe-dev.seahaven.com` → CNAME to Vercel (dev env)
|
||
|
||
### GitHub Apps — TWO (dev/prod isolated; each its own webhook URL)
|
||
| App | app_id | install | client_id | org | members scope | repos |
|
||
|---|---|---|---|---|---|---|
|
||
| **prod** `seahaven-openswe` | `4146115` | `142615168` | `Iv23lil96pKQNNDUn5yp` | `Sea-Haven-Industries` | members:**write** | all |
|
||
| **dev** `seahaven-openswe-dev` | `4162963` | `143023302` | `Iv23licQwJvGAPJj1HJe` | `seahaven-open-swe-dev` | members:**read** | all |
|
||
|
||
- **Promotion App** `seahaven-promotion` (actor `4170147`) is the sole non-admin fast-forward-push bypass on the `main` ruleset `18238334` — its FF-push of `dev → main` is what triggers the managed prod build.
|
||
|
||
### Env per deployment (set in LangGraph Cloud config + Vercel env — NOT Secrets Manager)
|
||
| Var | dev | prod |
|
||
|---|---|---|
|
||
| `LANGGRAPH_URL` | own (dev) deployment URL | own (prod) deployment URL |
|
||
| `DASHBOARD_BASE_URL` / `DASHBOARD_API_BASE_URL` | `https://openswe-dev.seahaven.com` | `https://openswe.seahaven.com` |
|
||
| `VITE_DASHBOARD_API_BASE_URL` | empty (same-origin via Vercel proxy) | empty |
|
||
| `ALLOWED_GITHUB_ORGS` | dev org (`seahaven-open-swe-dev`) | `Sea-Haven-Industries` |
|
||
| `CONFIGURED_ADMINS` | `amoussa1229,adam@seahavenind.com` | `amoussa1229,adam@seahavenind.com` |
|
||
|
||
`DASHBOARD_BASE_URL` / `DASHBOARD_API_BASE_URL` **must include `https://`** (see gotcha 3). Secret **values** are still sourced from `open-swe-{dev,prod}/*` Secrets Manager + SSM (the remaining source of truth) and set into the LangGraph Cloud + Vercel env stores — the accepted `secrets-and-config.md` deviation.
|
||
|
||
### Bedrock IAM (PR #74, still OPEN)
|
||
Two IAM users `open-swe-dev-bedrock` + `open-swe-prod-bedrock`, each attached to customer-managed policy `open-swe-bedrock-invoke` (least-privilege `bedrock:InvokeModel[WithResponseStream]` on the `us.anthropic.claude-opus-4-8` inference-profile ARN + its 3 routed foundation-model ARNs in us-east-1/us-east-2/us-west-2). Default model `bedrock_converse:us.anthropic.claude-opus-4-8` + 3 Fireworks models. Static access keys live only in the deployment env (dev key → dev, prod key → prod).
|
||
|
||
### User store — per-deployment
|
||
**Each** managed deployment has its **own** Store. The GitHub→email mapping `amoussa1229 → adam@seahavenind.com` (namespace `["user_mappings"]`, key = lowercased login, record `{github_login, work_email, status:"active", source, created_at, updated_at}`) was written to **both** the dev and prod stores directly. New users need a mapping **per-deployment** (write each store directly, or use the dashboard admin User-mappings UI — the `work_email` field was added by PR #65 fix #4).
|
||
|
||
### AWS decommission (PR #64)
|
||
Self-host CDK stacks destroyed. Residual: `CDKToolkit` (shared, **preserved**); ~50 `RETAIN`'d Secrets Manager shells + 3 S3 asset buckets (**pending cleanup**); AWS **Bedrock** (live dependency, kept).
|
||
|
||
### Operational gotchas (hard-won — carry these into any runbook)
|
||
1. **Per-deployment store → seed user mappings per-deployment.** A missing mapping makes `process_github_issue` silently early-return ("No email mapping … skipping"): the webhook returns 200/accepted but produces **no reaction and no run**. Seed dev **and** prod.
|
||
2. **Org-login gate uses the App *installation* token**, so the App must be org-installed with **Members:read**. OAuth working ≠ membership check working — they use **separate creds** (CLIENT_ID/SECRET for OAuth vs APP_ID/INSTALLATION_ID/PRIVATE_KEY for the install token). A mangled multi-line `GITHUB_APP_PRIVATE_KEY` breaks the install token (and thus the gate) while OAuth still works.
|
||
3. **`DASHBOARD_API_BASE_URL` must be `https://`** — an `http://` value makes GitHub reject the OAuth callback with "redirect_uri not associated."
|
||
4. **`osw_oauth_state` cookie is host-only** — start login on the **same host** as `DASHBOARD_API_BASE_URL`, or you get "oauth state mismatch."
|
||
5. **Webhooks go DIRECT to the langgraph URL** (`/webhooks/*`). Vercel only proxies `/dashboard/api/*`. The app is **same-origin only** (no CORS).
|
||
6. **On Vercel CI, Nitro's Vercel preset auto-emits `.vercel/output`** — drive the proxy via Nitro `routeRules` from `LANGGRAPH_BACKEND_URL`; never a manual script that `rm`s `.vercel/output`.
|
||
7. **Deleting + recreating a LangGraph deployment mints a NEW URL hash** (otherwise stable across revisions) — update every reference (Vercel env, webhooks, OAuth, docs).
|
||
|
||
---
|
||
|
||
## 2. Architecture: Before → After
|
||
|
||
### Before (self-hosted AWS — LIVE as of 2026-06-29)
|
||
```
|
||
GitHub/Slack/Linear ──webhook──▶ hooks.seahaven.com ─┐
|
||
Browser (dashboard) ─────────────▶ openswe.seahaven.com ─┤
|
||
▼
|
||
shared seahaven-com ALB (:443, host+path rules)
|
||
▼
|
||
EC2 (prod i-08a729e50779c4b07 t4g.large ARM64, private subnet)
|
||
nginx :80 ──proxy /dashboard/api + /webhooks──▶ langgraph dev :2024 (loopback)
|
||
in-memory store (reseeded by seed_store.sh ExecStartPost)
|
||
▼
|
||
Config: Secrets Manager open-swe-prod/* + SSM /open-swe-prod/*
|
||
Release: GitHub Actions → S3 open-swe-prod-assets/releases/* → SSM doc deploy.sh
|
||
IaC: CDK OpenSweIamStack + OpenSweDevStack + OpenSweProdStack
|
||
Sandbox: LangSmith cloud (DEFAULT_SANDBOX_SNAPSHOT_ID + GitHub proxy)
|
||
```
|
||
|
||
### After (managed — FINAL, see §1a for exact values)
|
||
```
|
||
┌──────────────── DEV lane ────────────────┐ ┌──────────────── PROD lane ───────────────┐
|
||
GitHub(dev org)/Slack/Linear │ webhook → open-swe-dev-….us.langgraph.app │ │ webhook → open-swe-prod-….us.langgraph.app│ GitHub(SHI org)/Slack/Linear
|
||
App seahaven-openswe-dev ───┘ (DIRECT to langgraph URL, /webhooks/*) │ │ (DIRECT to langgraph URL, /webhooks/*) └─── App seahaven-openswe
|
||
▼ ▼
|
||
Browser ▶ openswe-dev.seahaven.com ─┐ ┌─▶ openswe.seahaven.com ◀ Browser
|
||
│ ONE Vercel project `open-swe-prod` (team sea-haven)
|
||
│ ├─ env `dev` (branch dev) → proxies /dashboard/api/* → dev langgraph URL
|
||
│ └─ env production (branch main) → proxies /dashboard/api/* → prod langgraph URL
|
||
└─ proxy compiled by Nitro routeRules from per-env LANGGRAPH_BACKEND_URL (PR #76)
|
||
▼
|
||
TWO LangGraph Cloud deployments (same LangSmith workspace; one workspace key auths both incl. Store)
|
||
├─ dev ← branch `dev` · prod ← branch `main` (push-to-main auto-deploys prod)
|
||
├─ each serves the graphs + the custom http.app (agent.webapp:app = webhooks + dashboard API + OAuth)
|
||
├─ each has its OWN durable Postgres store + checkpointer (issue #9 SOLVED) — user mappings per-deployment
|
||
└─ env/secrets in the Deployment + Vercel config (NOT Secrets Manager) — Adam accepted deviation
|
||
▼
|
||
Sandbox: LangSmith cloud (UNCHANGED — DEFAULT_SANDBOX_SNAPSHOT_ID + GitHub-App proxy)
|
||
Bedrock: IAM users open-swe-{dev,prod}-bedrock + policy open-swe-bedrock-invoke (static keys in deploy env)
|
||
```
|
||
|
||
**What changes shape:** runtime host (EC2 → managed PaaS), durability (in-memory → managed Postgres, one store **per deployment**), CD (bespoke S3/SSM/packer → git-connected auto-build), config home (Secrets Manager/SSM → Deployment+Vercel env), ingress topology (shared ALB + one App → two dev/prod-isolated GitHub Apps each hitting its own `*.langgraph.app` directly), UI proxy (`ui/vercel.json` rewrite → Nitro `routeRules`). **What stays:** the LangSmith sandbox plane, CI (lint/format/unit/Playwright), the app code itself, and Secrets Manager/SSM as the secret-**value** source of truth.
|
||
|
||
---
|
||
|
||
## 3. Phased Plan
|
||
|
||
### Phase A — Dev spike (MOSTLY DONE)
|
||
Goal: prove managed serves our custom app + durability, at $0, before committing prod $.
|
||
|
||
**Proven this session** (spike-era specifics — superseded by §1a; URLs/project below are DEAD):
|
||
- ✅ Dev backend deployed to LangGraph Cloud, connected to branch `dev`:
|
||
~~`https://open-swe-dev-hosted-e76c2b0e8a7955fe8ad3110a7a54e5d0.us.langgraph.app`~~ → final dev URL in §1a (`open-swe-dev-fb737aa…`; the deployment was later deleted + recreated)
|
||
(Bedrock + a Fireworks key set; AWS creds for Bedrock deferred — see §10 open decision).
|
||
- ✅ UI deployed to Vercel — team `sea-haven`, ~~project `open-swe-dashboard`, `https://open-swe-dashboard.vercel.app`~~ (DELETED); now the **single** project `open-swe-prod` with dev/prod environments (§1a); proxy is Nitro `routeRules`, not the `ui/vercel.json` rewrite.
|
||
- ✅ Managed serves the custom `http.app` (dashboard API + webhooks) — **no platform auth gate** in front of our routes (webhook returns 401 sig-enforced, so signatures still govern).
|
||
- ✅ Vercel same-origin rewrite → backend works.
|
||
- ✅ GitHub OAuth dashboard login end-to-end.
|
||
- ✅ **DURABILITY** — team-default model survived a revision redeploy (this is issue **#9**'s core acceptance goal, proven on managed).
|
||
|
||
**Remaining Phase A items (the GitHub `@openswe` run-trigger loop):**
|
||
- [ ] A1. Get an `@openswe` GitHub comment to dispatch a run end-to-end on dev. Blocked by the user-mapping + cache gotchas (see fixes #4 and #5 in §5).
|
||
- [ ] A2. Create the owner user-mapping in the managed Store (namespace `["user_mappings"]`, key = lowercased login `amoussa1229`, value `{github_login, work_email}`) — the Admin UI can't set `work_email` yet (fix #4).
|
||
- [ ] A3. Verify a freshly-added mapping is seen without a redeploy across managed's multi-replica autoscaling (fix #5).
|
||
- [ ] A4. Confirm `LANGGRAPH_URL` is set to the deployment URL on the dev deployment so `langgraph_client()` calls resolve (fix #1 — config now, code later).
|
||
|
||
### Phase B — Harden (the code fixes + env codification)
|
||
Land the 6 code fixes (§5), codify env/config, and resolve the Bedrock-auth decision **before** standing up prod.
|
||
|
||
- [ ] B1. Land code fixes #1–#6 (§5) as a PR into `dev` (or split into focused PRs). Re-run `make lint` + `make test`.
|
||
- [ ] B2. **Codify the env contract for managed.** Update `ui/vercel.json` rewrite to the *prod* deployment URL (Phase C) but keep dev pointing at dev. Add a documented LangGraph-Cloud env list to the repo (see §4) — NOT secrets, just the variable inventory + which are excluded.
|
||
- [ ] B3. **Resolve the Bedrock-auth decision** (§10 open decision): static AWS keys for a Bedrock-scoped IAM user in the deployment env, OR run the agent on Fireworks. This intersects the #62 model work.
|
||
- [x] B4. ~~Merge #62 first~~ — **DONE / N/A.** #62 (Bedrock + Fireworks) **IS** merged to `dev` and is what the dev deployment runs — verified on `origin/dev` (`options.py` → `DEFAULT_MODEL_ID = "bedrock_converse:us.anthropic.claude-opus-4-8"`) and confirmed by the live deployment's `git_ref_sha: a4ed19ba…` (= the #62 merge commit). The contrary note came from a STALE local checkout; no merge action needed.
|
||
- [ ] B5. Audit ALL in-process caches for the single-process → multi-replica assumption (generalization of fix #5; see §5).
|
||
- [ ] B6. Measure + reduce custom-app import time (fix #6) so the deployment isn't flagged unhealthy / slow to scale.
|
||
|
||
### Phase C — Prod deployment
|
||
- [ ] C1. Create a **prod LangGraph Cloud deployment** tracking branch `main` (the durable autoscaled 1→10 tier, not the free Dev tier). Record its `*.langgraph.app` URL.
|
||
- [x] C2. ~~Create the **prod Vercel project/target** (or promote the existing `open-swe-dashboard` to production)~~ — **DONE differently:** one project `open-swe-prod` with a production env (`main`) + a custom `dev` env (`dev`), each with its own `LANGGRAPH_BACKEND_URL`. See §1a.
|
||
- [ ] C3. Set the **prod env triad** (§4) on the prod deployment + Vercel:
|
||
- `LANGGRAPH_URL` = the prod `*.langgraph.app` URL
|
||
- `DASHBOARD_BASE_URL` + `DASHBOARD_API_BASE_URL` = the prod Vercel origin (with `https://` scheme — fix #2)
|
||
- [ ] C4. **Establish the prod approval gate** (§6): git-connected auto-deploy needs an explicit gate. Preserve the prod manual-approval that exists today (the GitHub `prod` Environment reviewer). Mechanism = protected `main` ruleset + the platform's "require manual promotion to production" if available (§10 open decision).
|
||
- [ ] C5. **Repoint webhooks + OAuth URLs** to prod:
|
||
- GitHub App `seahaven-openswe` webhook URL → prod backend webhook URL (`*.langgraph.app/webhooks/github` or `hooks.seahaven.com` — §10 custom-domain decision)
|
||
- Slack Event Subscriptions + Interactivity Request URLs → prod
|
||
- Linear webhook URL → prod (if/when wired)
|
||
- GitHub App OAuth callback → `<DASHBOARD_API_BASE_URL>/dashboard/api/auth/callback` (the Vercel prod origin; fixes #2, #3)
|
||
- [ ] C6. **Custom-domain decision** (§10): dashboard custom domain solved by Vercel; webhooks `hooks.seahaven.com` either re-point to `*.langgraph.app` directly or front via CNAME. Managed issues `*.langgraph.app` URLs ONLY (no documented custom-domain support on the backend).
|
||
|
||
### Phase D — Cutover + verify
|
||
- [ ] D1. Run the issue **#9 acceptance criteria on PROD**: durable store survives a revision redeploy; team defaults / user mappings persist; a paused/in-flight run survives a redeploy (the durability win self-host never had).
|
||
- [ ] D2. End-to-end prod smoke: `@openswe` GitHub comment → run dispatch → sandbox → draft PR → reply in source channel.
|
||
- [ ] D3. Dashboard prod smoke: GitHub OAuth login on the stable Vercel alias (fix #3); admin pages reachable for `CONFIGURED_ADMINS`.
|
||
- [ ] D4. Webhook sig-enforcement smoke: unsigned POST to each `/webhooks/*` returns 401; signed returns 200.
|
||
- [ ] D5. **Soak** for an agreed window (recommend 3–7 days) with the AWS prod stack still standing as the rollback target (§11) before any teardown.
|
||
|
||
### Phase E — AWS decommission (AFTER soak)
|
||
Only after managed prod is proven + soaked. See §7 for the precise retire-vs-keep list. Phased teardown:
|
||
- [ ] E1. Export the live ALB listener-rule / target-group config for `open-swe-prod` and `open-swe-dev` to JSON (the export is the rollback source — these rules were partly hand-built).
|
||
- [ ] E2. Disable the bespoke CD (build-artifacts / cd-infra workflows) so nothing re-deploys the EC2 boxes.
|
||
- [ ] E3. `cdk destroy OpenSweProdStack` then `OpenSweDevStack` (mind the **RETAIN** secret shells — they survive and hold the global `open-swe-<env>/*` names; force-delete only EMPTY shells once env is fully migrated; keep populated ones transitionally as the env source — §7).
|
||
- [ ] E4. Remove ALB rules/target groups, S3 assets buckets, SSM deploy docs, and the per-env CDK bootstrap qualifier `oswedev` (`CDKToolkit-oswedev`) — once nothing references them.
|
||
- [ ] E5. Decide DNS: repoint `openswe.seahaven.com` → Vercel; resolve `hooks.seahaven.com` → `*.langgraph.app` (or CNAME).
|
||
- [ ] E6. Leave `OpenSweIamStack` until last — the promotion App + any retained roles may still be referenced.
|
||
|
||
### Phase F — Docs / memory
|
||
- [ ] F1. Rework Confluence "AWS Architecture Map" (page **1540098**) + the open-swe child page (**26116098**): the RDS/EC2/ALB subgraph is replaced by an **external-services view** (LangGraph Cloud + Vercel + the retained GitHub App + LangSmith sandbox).
|
||
- [ ] F2. Update `README.md` + `INSTALLATION.md` (§10 is already canonical-correct; align Sea Haven specifics).
|
||
- [ ] F3. Update project memories `project_open_swe_migration.md` + `reference_open_swe_deployment.md` to mark the managed cutover and the AWS teardown.
|
||
|
||
---
|
||
|
||
## 4. Env / Config Reference
|
||
|
||
### The prod triad (per INSTALLATION.md §10, lines 630–658) — see §1a for the exact dev/prod values
|
||
| Var | Value (per deployment/env) | Notes |
|
||
|---|---|---|
|
||
| `LANGGRAPH_URL` | the **own** deployment URL (`https://...langgraph.app`) | **NOT** localhost. Drives `thread_ops.langgraph_url()` (fix #1). dev→dev URL, prod→prod URL. |
|
||
| `DASHBOARD_BASE_URL` | the own dashboard origin, **with `https://`** (`https://openswe-dev.seahaven.com` / `https://openswe.seahaven.com`) | |
|
||
| `DASHBOARD_API_BASE_URL` | same as above, **with `https://` scheme** | scheme required or OAuth `redirect_uri` is schemeless and GitHub rejects (fix #2) |
|
||
| `VITE_DASHBOARD_API_BASE_URL` | **empty** | same-origin mode; UI calls relative `/dashboard/api/*`, the Vercel Nitro proxy rewrites to backend |
|
||
| `LANGGRAPH_BACKEND_URL` | **Vercel env var**, per Vercel environment (dev env = dev URL, prod env = prod URL) | drives the Nitro `routeRules` `/dashboard/api/*` proxy at build time (PR #76) — required on Vercel builds |
|
||
| `ALLOWED_GITHUB_ORGS` | `Sea-Haven-Industries` (prod) / `seahaven-open-swe-dev` (dev) | org-login gate; checked via the App **installation** token (gotcha 2) |
|
||
|
||
GitHub App dashboard OAuth callback = `<DASHBOARD_API_BASE_URL>/dashboard/api/auth/callback` (the own dashboard origin — prod App on `openswe.seahaven.com`, dev App on `openswe-dev.seahaven.com`).
|
||
|
||
**UI proxy mechanism (final, PR #76):** the `/dashboard/api/*` proxy is **Nitro `routeRules`** in `ui/vite.config.ts` reading `process.env.LANGGRAPH_BACKEND_URL`, compiled by Nitro's Vercel preset into `.vercel/output/config.json`. This **supersedes** the spike-era `ui/vercel.json` same-origin rewrite and PR #75's hand-rolled `ui/scripts/build-vercel-output.mjs` (deleted). `ui/vercel.json` now only carries `framework:null` + `buildCommand: bun run build`.
|
||
|
||
### Where secrets/env live now
|
||
**LangGraph Cloud Deployment config + Vercel env** — NOT AWS Secrets Manager. This **deviates from the Sea Haven `secrets-and-config.md` "Secrets Manager for all sensitive" handbook rule** — **Adam ACCEPTED this deviation** (managed has no instance role / no fetch-config boot hook; the platform's own secret store is the mechanism).
|
||
|
||
### Sourcing env from the existing AWS stack (transitional)
|
||
Pull from the existing `open-swe-dev` Secrets Manager + SSM via a documented CLI dump:
|
||
- `aws secretsmanager list-secrets --filters Key=name,Values=open-swe-dev/` then per-name `get-secret-value` — **NOT** `batch-get-secret-value` (its pagination silently drops values past page 1; this bit the box twice — see memory).
|
||
- SSM: `aws ssm get-parameters-by-path --path /open-swe-dev/`.
|
||
|
||
### EXCLUDE when copying to managed
|
||
| Category | Vars |
|
||
|---|---|
|
||
| Box-/self-host-specific | `LANGGRAPH_URL` (set fresh to deployment URL), `LANGGRAPH_URL_PROD`, `LANGSMITH_ENDPOINT`, `LANGSMITH_ENDPOINT_PROD`, `LANGSMITH_URL_PROD`, `LANGSMITH_TENANT_ID_PROD`, `LANGSMITH_HOST_API_URL`, `LANGCHAIN_REVISION_ID` |
|
||
| Dropped providers (PR #62) | `OPENAI_API_KEY`, `GOOGLE_API_KEY`, `GROQ_API_KEY` (these 6 Secrets Manager shells deleted 2026-06-29, final purge 2026-07-06) |
|
||
| Unused sandbox providers | `DAYTONA_API_KEY`, `RUNLOOP_API_KEY` |
|
||
| Eval-only | `JUDGE_ANTHROPIC_API_KEY` (kept on AWS for `evals/reviewer/judge.py`, not needed in the runtime deployment) |
|
||
|
||
### KEY mappings on managed
|
||
- `LANGSMITH_API_KEY` = the **`LANGSMITH_API_KEY_PROD`** value (and `LANGCHAIN_API_KEY` = same).
|
||
- The full secret inventory is `SECRET_VARS` in `infra/lib/constructs/config-store.ts:46` (28 shells) — use it as the checklist of what to carry, minus the EXCLUDE rows above.
|
||
|
||
### Models (post-#62) + Bedrock auth
|
||
Post-#62 `SUPPORTED_MODELS` = `bedrock_converse:us.anthropic.claude-opus-4-8` (**DEFAULT**) + 3 Fireworks models; **no direct `anthropic:` option**.
|
||
✅ **#62 is merged to `dev` and deployed** — verified on `origin/dev` (`options.py` → `bedrock_converse` default) and the live deployment `git_ref_sha a4ed19ba`. (The pre-#62 values appear only on a stale LOCAL checkout — ignore.)
|
||
|
||
**Bedrock on managed has NO EC2 instance role** (the #62 IAM design used the EC2 instance role as the Bedrock principal — that breaks on managed). Options (OPEN DECISION, §10):
|
||
- (a) Static `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` / `AWS_REGION` for a **Bedrock-scoped IAM user** in the deployment env, or
|
||
- (b) Run the agent on **Fireworks** and avoid Bedrock entirely on managed.
|
||
|
||
### Bedrock IAM users (CREATED — resolves §10.1, discharges the §9 IAM gates)
|
||
Option (a) was chosen and executed live (2026-06-30, account **328440206208** / **us-east-1**). This is **click-ops IAM** — there is no remaining open-swe AWS IaC after the decommission (#64), so these are created with the CLI, not CDK.
|
||
|
||
- **Customer-managed policy `open-swe-bedrock-invoke`** (`arn:aws:iam::328440206208:policy/open-swe-bedrock-invoke`). Least-privilege: actions `bedrock:InvokeModel` + `bedrock:InvokeModelWithResponseStream` **only**, scoped to exactly the `us.anthropic.claude-opus-4-8` inference-profile ARN + its **three** routed foundation-model ARNs (us-east-1, us-east-2, us-west-2). **No wildcards, no other models.** (Supersedes the §10.1 plan to reuse the #62 instance-role policy — a fresh standalone policy was minted instead.)
|
||
- **Two IAM users**, each attached to that policy: **`open-swe-dev-bedrock`** and **`open-swe-prod-bedrock`**. Tagged `project=open-swe`, `managed-by=cli-migration`, `purpose=bedrock-invoke`.
|
||
- **Static access keys are minted separately by the owner** (`aws iam create-access-key`) — secret keys live **only** in the deployment env stores, never in this repo or memory. dev key → dev env; prod key → LangGraph Cloud prod config (`AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` / `AWS_REGION=us-east-1`).
|
||
- **Deviation rationale:** static long-lived keys are a deliberate departure from the Sea Haven OIDC norm because **managed LangGraph Cloud cannot assume an AWS role**. Mitigated by the tight least-privilege policy above.
|
||
- ✅ **Both mandatory gates ran and PASSED:**
|
||
- **GPT-4.1 IAM cross-review** (confirmed hit `gpt-4.1-2025-04-14`) — least-privilege confirmed.
|
||
- **`/sh-security-review`** — **0 critical/high**; two **accepted mediums** (the static-key deviation + no per-principal budget cap).
|
||
- **Recommended follow-ups:** shortest viable key-rotation cadence with recorded creation dates; an AWS Budgets / CloudWatch anomaly alarm on per-principal Bedrock `InvokeModel` volume; confirm CloudTrail captures these users.
|
||
|
||
---
|
||
|
||
## 5. Required Code Fixes (the 6 gotchas)
|
||
|
||
These are self-host → managed assumption breaks discovered on the spike. Each gets a config workaround now and a code fix to land in Phase B.
|
||
|
||
### Fix #1 — `langgraph_url()` localhost fallback
|
||
**File:** `agent/utils/thread_ops.py:38-41`
|
||
```python
|
||
def langgraph_url() -> str:
|
||
return os.environ.get("LANGGRAPH_URL") or os.environ.get(
|
||
"LANGGRAPH_URL_PROD", "http://localhost:2024"
|
||
)
|
||
```
|
||
On managed, every `langgraph_client()` call (thread sidebar, run creation, the same pattern in `agent/dashboard/user_mappings.py:43` `get_client()` with no URL) `ConnectError`s unless `LANGGRAPH_URL` is set to the deployment URL.
|
||
- **Config fix (now):** set `LANGGRAPH_URL` = the deployment URL on every deployment (Phase A4 / C3).
|
||
- **Code fix (Phase B):** default to the in-process / deployment URL rather than `http://localhost:2024` when running inside a managed deployment (e.g. honor the platform's own URL env, or fail loud instead of silently hitting localhost).
|
||
|
||
### Fix #2 — OAuth `DASHBOARD_API_BASE_URL` must include scheme
|
||
**Where:** OAuth redirect construction (dashboard `oauth.py` / `routes.py`), driven by `DASHBOARD_API_BASE_URL`.
|
||
A schemeless value yields a schemeless `redirect_uri` → GitHub rejects with "redirect_uri not associated."
|
||
- **Fix:** always set `DASHBOARD_API_BASE_URL` **with `https://`** (Phase C3). Code hardening (Phase B): assert/normalize a scheme at startup and fail loud if missing.
|
||
|
||
### Fix #3 — OAuth state cookie is host-only
|
||
**Cookie:** `osw_oauth_state` (host-only). Login must START on the SAME host as `DASHBOARD_API_BASE_URL` — the **stable Vercel alias**, never the immutable per-deploy URL — or you get "oauth state mismatch."
|
||
- **Fix:** pin `DASHBOARD_API_BASE_URL` + the login entry point to the stable alias / custom domain (Phase C3, D3). Document this so a future per-deploy preview URL isn't used for login.
|
||
|
||
### Fix #4 — Admin "User mappings" UI can't set `work_email`
|
||
**Files:** the Admin → User mappings UI (`ui/`) + `agent/dashboard/user_mappings.py` (`upsert_mapping` at line 257 takes `work_email` but the UI form doesn't expose it).
|
||
Mappings can't be fully created from the dashboard → must write the Store directly: namespace `["user_mappings"]`, key = **lowercased login**, value `{github_login, work_email}` (matching `_index_record` at `user_mappings.py:73-84`, which keys `_by_login` on `login.lower()`).
|
||
- **Fix (Phase B):** add the `work_email` field to the Admin mappings form so mappings are fully creatable from the UI.
|
||
|
||
### Fix #5 — GitHub webhook path doesn't refresh the user-mapping cache (multi-replica break)
|
||
**Files:** `agent/webhooks/github.py` (GitHub handlers: `process_github_pr_comment`, `process_github_issue`) vs `agent/webhooks/slack.py` (`process_slack_mention`). (Pre-modular-refactor these all lived in `agent/webapp.py`.)
|
||
The Slack path refreshes before lookup:
|
||
```python
|
||
# agent/webhooks/slack.py (process_slack_mention)
|
||
await webapp.refresh_user_mapping_cache()
|
||
...
|
||
```
|
||
The GitHub path historically did **not** — it called `email = await email_for_login(github_login)` cold. The cache (`user_mappings.py` `_ensure_cache_loaded`) is **one-shot per process** (`_cache_loaded` flag). On self-host single-process this was fine; on managed's **multi-replica autoscaling**, a freshly-added mapping isn't seen by a replica whose cache loaded earlier — until restart.
|
||
- **Fix (applied):** the GitHub issue and PR-comment handlers now call `webapp.refresh_user_mapping_cache()` before email resolution, mirroring the Slack path.
|
||
- **Generalize (Phase B5):** audit ALL in-process caches for the single-process → multi-replica assumption — `SANDBOX_BACKENDS` dict (`agent/utils/sandbox_state.py`), `_by_login`/`_by_email`/`_by_slack_id` (`user_mappings.py`). Sandbox affinity is already thread-keyed + persisted in thread metadata (`sandbox_id`), so it's the cache state that needs the multi-replica review. (The legacy in-process thread lock has been removed: webhook triggers now serialize through `dispatch_agent_run`'s `multitask_strategy="interrupt"` instead.)
|
||
|
||
### Fix #6 — Slow custom-app import (~8s startup)
|
||
**Symptom:** "exceeded expected startup time" → risks the deployment being marked unhealthy / slow to scale out.
|
||
- **Fix (Phase B6):** lazy imports / reduce import-time work in `agent/webapp.py` (+ `agent/webhooks/*.py`) and the graph factories. Profile with `FF_PROFILE_IMPORTS` (the import-profiling flag) to find the heavy modules.
|
||
|
||
---
|
||
|
||
## 6. CD / Ops Changes
|
||
|
||
### Retires (bespoke pipeline)
|
||
- GitHub Actions **build → S3 releases → SSM doc → `deploy.sh` → health-gate** (`.github/scripts/`, `deploy/ami/deploy.sh`)
|
||
- `roll-box.sh` / `publish-and-deploy.sh` / `rollback.sh`
|
||
- AMI baking (packer `deploy/ami/open-swe-base.pkr.hcl`) + `cdk.context.json` AMI pin
|
||
- `cd-infra.yml` CDK deploys + OIDC bootstrap qualifiers
|
||
- `fetch-config.sh` / `seed_store.sh` (boot-time config materialization + Store reseed)
|
||
|
||
### Replaced by git-connected PaaS
|
||
- **LangGraph Cloud** auto-builds a revision on push (first-party zero-downtime + revision rollback).
|
||
- **Vercel** builds / atomic-deploys / instant-rollback + per-PR previews (**net-new** capability the AWS stack never had).
|
||
|
||
### Stays
|
||
- **CI** (lint / format / unit / Playwright E2E) — still matters and still gates merges. `make lint`, `make test`.
|
||
|
||
### Gate re-homing
|
||
- The **dev → main promotion gate** (`check-dev-green.sh` + protected-`main` ruleset 18238334) **re-homes**: the **dev deployment tracks `dev`**, the **prod deployment tracks `main`**, so the gate governs exactly what reaches prod.
|
||
- Today's promotion machinery: `promote-dev-to-prod.yml` + the `seahaven-promotion` GitHub App (app_id 4170147, in ruleset 18238334 bypass_actors) FF-pushes `dev → main`. Under managed, a push to `main` is what triggers the prod build — so the promotion gate IS the prod deploy gate.
|
||
|
||
### Preserve the PROD manual-approval gate (LOAD-BEARING)
|
||
Today the `prod` GitHub Environment (required reviewer `amoussa1229`) is the manual approval. Git-connected auto-deploy removes the CD job that consulted that Environment, so the approval must be re-established explicitly:
|
||
- **Mechanism options (OPEN DECISION §10):** protected `main` (only the promotion App can FF-push, and that push is itself the gate) AND/OR the platform's "require manual promotion to production" if LangGraph Cloud exposes it.
|
||
- **Norm shift:** deploy-then-merge → **merge-to-deploy**. Lean on the dev deployment + Vercel previews to verify *before* promoting `dev → main`. (This inverts the Sea Haven handbook deploy-then-merge default — call it out in the README + handbook note.)
|
||
|
||
---
|
||
|
||
## 7. AWS Decommission — Retire vs Keep
|
||
|
||
### RETIRE (becomes vestigial once managed prod is live + soaked)
|
||
| Item | Path / resource |
|
||
|---|---|
|
||
| AMI bake + cloud-init | `deploy/ami/` (`open-swe-base.pkr.hcl`, `provision.sh`, `user-data.sh`, `deploy.sh`, `templates/`) |
|
||
| Self-host boot/config scripts | `deploy/seahaven/fetch-config.sh`, `seed_store.sh`, `put-config.sh`, `nginx/openswe.conf`, `systemd/open-swe.service`, `aegra/` (already deferred) |
|
||
| CDK box/ALB/AMI stacks | `infra/lib/open-swe-stack.ts`, `constructs/app-service.ts`, `assets-bucket.ts`, `ami-cache.ts`, `instance-role.ts`, `github-deploy-roles.ts` |
|
||
| EC2 instances | prod `i-08a729e50779c4b07` (t4g.large), dev `i-0a3bb8e0ddd36c29b` |
|
||
| S3 release buckets | `open-swe-dev-assets`, `open-swe-prod-assets` |
|
||
| SSM deploy docs | `open-swe-dev-deploy`, `open-swe-prod-deploy` |
|
||
| ALB plumbing | `open-swe-prod-tg` / `open-swe-dev-tg` target groups + the `seahaven-com` ALB host/path listener rules for openswe/hooks |
|
||
| Per-env bootstrap qualifier | dev `oswedev` (`CDKToolkit-oswedev`) |
|
||
| RDS / Redis durable plan | **never built** — fully dropped (managed provides durability) |
|
||
| Bespoke CD workflows | `cd-infra.yml`, `build-artifacts.yml`, `roll-box.sh`/`publish-and-deploy.sh`/`rollback.sh` |
|
||
|
||
### KEEP
|
||
| Item | Why |
|
||
|---|---|
|
||
| **GitHub App `seahaven-openswe`** (App 4146115 / Install 142615168) | unchanged auth + webhook source; only the webhook/OAuth URLs repoint |
|
||
| **LangSmith sandbox setup** incl. `DEFAULT_SANDBOX_SNAPSHOT_ID` (`dc36e509-d4d2-4efc-8a4e-61f74f3446ec`) | the only sandbox with working in-sandbox git/gh auth (`_configure_github_proxy`); managed default `SANDBOX_TYPE=langsmith` uses it |
|
||
| **Secrets Manager `open-swe-{dev,prod}/*` values** | the **source you copy env FROM**, at least transitionally — do NOT force-delete populated shells until managed env is fully cut over and verified |
|
||
| **DNS** | repoint `openswe.seahaven.com` → Vercel; decide `hooks.seahaven.com` → `*.langgraph.app` or a CNAME (§10) |
|
||
| **`seahaven-promotion` App** + main/dev rulesets (18238334 / 18238542) | still govern `dev → main` promotion = the prod deploy gate |
|
||
| **`evals/` + `JUDGE_ANTHROPIC_API_KEY`** | eval pipeline unaffected (separate from runtime) |
|
||
|
||
**Phased teardown rule:** remove AWS infra ONLY after managed prod is proven + soaked (Phase D5). The exported ALB JSON (E1) is the rollback source for the hand-built listener rules.
|
||
|
||
---
|
||
|
||
## 8. Cost
|
||
|
||
| Line item | Monthly |
|
||
|---|---|
|
||
| Prod deployment uptime ($0.0036/min, always-on) | ≈ $155 |
|
||
| Runs ($0.005/run) | usage-based |
|
||
| Traces above 10k/mo | pay-as-you-go |
|
||
| LangSmith **Plus** seat ($39) | already paid (not incremental) |
|
||
| Dev deployment | **$0** (1 free on Plus, preemptible) |
|
||
| **Incremental total** | **≈ $160/mo** |
|
||
|
||
Compare to the retired AWS stack: 2× EC2 (t4g.large prod + t4g.medium/large dev), 2× S3 buckets, ALB share, NAT egress, plus the **never-built** RDS/Redis durable tier the self-host path would have added. Managed nets out roughly cost-neutral-to-cheaper while *adding* durability + previews. **TODO: confirm current AWS run-rate for an apples-to-apples delta** (not verified in repo).
|
||
|
||
---
|
||
|
||
## 9. Sea Haven Gates & Obligations
|
||
|
||
- [ ] **`/sh-plan-review`** (GPT-4.1 `cross_reviewer`) on THIS plan **BEFORE prod cutover** (Phase C gate). ⚠️ `run.py` misroutes reviewer-framed prompts to the no-op `done` route — call `models.get_cross_reviewer()` directly (load orchestrator `.env`, `ChatOpenAI("gpt-4.1")`) per `feedback_orchestrator_usage`.
|
||
- [ ] **`/sh-security-review`** on the sensitive surface: this migration touches **auth/webhook signature verification** (the custom `http.app` is now publicly reachable on `*.langgraph.app` with no platform gate in front) and **secrets handling** (secrets move from Secrets Manager to the Deployment/Vercel env). Required, not opt-in — resolve confirmed critical/high before cutover.
|
||
- [ ] **GPT-4.1 cross-family review** if the Bedrock-auth fix changes IAM (new Bedrock-scoped IAM user + policy = IAM change → mandatory cross-review).
|
||
- [ ] **Confluence** — rework "AWS Architecture Map" (**1540098**) + open-swe child page (**26116098**): replace the RDS/EC2/ALB subgraph with an external-services view (LangGraph Cloud + Vercel + retained GitHub App + LangSmith sandbox). Do it in the same conversation as the cutover, not as a follow-up.
|
||
- [ ] **README + INSTALLATION** updated (INSTALLATION §10 already canonical; align Sea Haven specifics + the merge-to-deploy norm shift).
|
||
- [ ] **Memories** — update `project_open_swe_migration.md` + `reference_open_swe_deployment.md`.
|
||
|
||
---
|
||
|
||
## 10. Decisions — RESOLVED (2026-06-29, Adam)
|
||
|
||
1. **Bedrock auth on managed → STATIC AWS KEYS.** Create a **Bedrock-scoped IAM user**; set `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` / `AWS_REGION=us-east-1` in the deployment env; reuse the #62 instance-role Bedrock policy (`bedrock:InvokeModel[WithResponseStream]` on the `us.anthropic.claude-opus-4-8` inference-profile ARN + per-region foundation-model ARNs in us-east-1/2 + us-west-2). Caveat: long-lived keys → rotation reminder + keep the policy tightly scoped. **New IAM user+policy = mandatory GPT-4.1 IAM cross-review + `/sh-security-review` before merge.**
|
||
2. **Webhook domain → RAW `*.langgraph.app`** (the "easier" path — no CNAME/DNS/TLS). Point GitHub/Slack/Linear webhook URLs straight at the deployment. Dashboard gets a **Vercel custom domain** (e.g. `openswe.seahaven.com`) since that's trivial on Vercel.
|
||
3. **Prod approval gate → GATED PROMOTE-TO-MAIN WORKFLOW.** Re-home `promote-dev-to-prod.yml`: gate the promote job on the **`prod` GitHub Environment reviewer (`amoussa1229`)**; on approval it FF-pushes `dev → main`; the push triggers the managed prod build. Preserves today's manual-approval semantics — approval is on the promote job, the merge to `main` is the deploy trigger.
|
||
4. **Cost/trace monitoring → YES, in LangSmith** (usage/trace budget + alert in the LangSmith workspace).
|
||
5. **Dev/prod home → CONFIRMED:** prod = LangChain managed (LangGraph Cloud) + Vercel UI; dev = managed **free** Dev tier (same runtime as prod).
|
||
6. ~~#62 merge status~~ — RESOLVED (merged + deployed).
|
||
|
||
## 10b. Execution approach — 2-PR SPLIT (Adam's call, 2026-06-29, post-plan-review)
|
||
|
||
The plan-review BLOCKed a literal single PR and recommended a 3-PR minimum (isolate cache fix #5 + the promotion-gate workflow). Adam chose a **2-PR split** — isolating the prod-deploy-gate change (highest risk to prod), bundling the rest:
|
||
- **PR1 — promotion-gate workflow + ruleset.** Re-home `promote-dev-to-prod.yml` to the gated promote-to-main (decision 3), and lock `main` so ONLY the promote App can push (BLOCK 3). Isolated because it changes *how prod deploys* — must be independently reviewable/revertible.
|
||
- **PR2 — the 6 code fixes (§5) + `ui/vercel.json` prod repoint + README/INSTALLATION.** (Accepted deviation from the reviewer's 3-PR rec: the multi-replica cache fix #5 stays bundled in PR2 rather than isolated — Adam's call.)
|
||
- **NOT in either PR (rollback safety):** AWS **infra-code deletion** (`deploy/ami`, `deploy/seahaven`, `infra/` self-host stack) + `cdk destroy` stay a **POST-SOAK cleanup (Phase E)**.
|
||
- **Not PR content (operational, sequenced around the merges):** LangGraph Cloud prod deployment + env, Vercel prod + custom domain, webhook/OAuth repoint, Bedrock IAM user, LangSmith budget, Confluence/memory.
|
||
|
||
## 10c. Plan-review resolution (GPT-4.1 cross_reviewer, 2026-06-29) — verdict REQUEST CHANGES, all BLOCKs addressed
|
||
|
||
- **BLOCK 1 (single PR)** → addressed via the **2-PR split** above (conscious deviation: cache fix #5 not isolated — Adam accepted).
|
||
- **BLOCK 2 (raw LangGraph API public?)** → **EMPIRICALLY RESOLVED.** Unauthenticated probes of the dev deployment returned **403 "Missing authentication headers"** for `/threads/search`, `/assistants/search`, `/store/items`; `/ok`=200, `/dashboard/api/me`=401. The platform gates the raw control-plane API. **Keep as a pre-prod gate (Phase D4):** re-probe the PROD deployment URL before cutover.
|
||
- **BLOCK 3 (prod gate enforcement)** → PR1 must **lock `main`** so only the promote App can push (verify ruleset 18238334: no direct-push path, PR-required, promote App is the sole FF bypass). Any non-gated push to `main` would auto-deploy prod.
|
||
- **BLOCK 4 (secrets posture)** → add to Phase B/E: **rotate** all secrets after migrating them to the platform env stores and BEFORE deleting from AWS; **document who can read/write** the LangGraph Cloud + Vercel env (access audit); delete AWS shells only post-cutover; no dual-homed/stale secrets; record in memory.
|
||
- **FIX (rollback integrity)** → Phase E checklist: "no IaC/DNS/config the rollback needs is altered in PR1/PR2 or during cutover."
|
||
- **FIX (gates resolved-before-merge)** → make explicit: do NOT merge PR1/PR2 until the GPT-4.1 IAM cross-review (Bedrock IAM user) + `/sh-security-review` (auth/webhook/secrets surface) are **resolved with no critical/high**; Confluence (1540098 + 26116098) + memory updated in the cutover conversation.
|
||
- **NITs/QUESTIONs** → carried into §8/§10 TODOs (cost delta, `FF_PROFILE_IMPORTS`, platform feature availability, key-rotation owner, env-store access logging, cache race-review under autoscaling, atomic webhook repoint).
|
||
|
||
---
|
||
|
||
## 11. Rollback Story
|
||
|
||
**Pre-cutover:** the self-host `langgraph up` + RDS plan (RDS+Redis+Docker; `/sh-plan-review`'d to APPROVE-after-revision this session) is the documented **fallback** if managed is rejected before cutover. Its durable design wins (env-scoped RDS physical names; `Credentials.fromGeneratedSecret({secretName:"open-swe-<env>/rds-credentials"})` under the existing instance-role secret prefix → no new IAM; DESTROY-on-rollback RDS-managed secret; derive `DATABASE_URI` in fetch-config) are captured in the project memory.
|
||
|
||
**Post-cutover (managed is live, AWS still standing during soak):** if managed prod fails, **rollback = repoint webhooks + DNS back** to the AWS stack:
|
||
- GitHub App / Slack / Linear webhook URLs → `hooks.seahaven.com`
|
||
- `openswe.seahaven.com` DNS → the `seahaven-com` ALB (restore from the E1 export)
|
||
- `ui/vercel.json` rewrite → the AWS dashboard origin (or stop using Vercel)
|
||
- Do NOT run Phase E teardown until the soak passes — the AWS stack IS the rollback target.
|
||
|
||
**Revision-level rollback (within managed):** LangGraph Cloud revision rollback (backend) + Vercel instant rollback (UI) cover bad deploys without leaving the platform.
|
||
|
||
---
|
||
|
||
## TODOs flagged (couldn't verify in repo)
|
||
- ~~#62 merge state~~ — RESOLVED: merged to `dev` + deployed (the draft read a stale local checkout; `origin/dev` and the deployment `git_ref_sha a4ed19ba` confirm Bedrock+Fireworks).
|
||
- **Current AWS run-rate** for the cost delta (§8) — not derivable from repo.
|
||
- **LangGraph Cloud custom-domain + "manual promotion to production" feature availability** (§10.2, §10.3) — platform features, verify in the LangGraph Cloud console/docs at execution time.
|
||
- **`FF_PROFILE_IMPORTS`** exact flag name/usage (§5 fix #6) — referenced from session context; confirm the flag exists in `agent/` before relying on it.
|
||
- Exact webhook path on `*.langgraph.app` (whether the custom `http.app` mounts at root so `/webhooks/github` is reachable as-is) — proven reachable on the dev spike; re-verify the exact path on prod.
|