# Sea Haven — Open SWE deployment runbook How this fork is deployed at Sea Haven. **PROD is LIVE as of 2026-06-29.** The runtime is the **stock LangGraph dev server** (not the Aegra path — see [Aegra](#aegra-deferred)), running on a self-hosted ARM64 EC2 box behind the shared `seahaven-com` ALB. Account `328440206208`, region `us-east-1`. Internal addresses, ARNs, snapshot ids, and account-scoped values that are sensitive are shown as ``; the real values live in the private IT docs (Confluence "AWS Architecture Map", id 1540098) and in AWS — **do not commit them to this public fork.** This is the canonical deploy runbook. The CDK details live in [`infra/README.md`](../../infra/README.md); the rotation procedure in [`ROTATION.md`](ROTATION.md). ## Live prod facts (2026-06-29) | | | |---|---| | Dashboard | `https://openswe.seahaven.com` | | Webhooks | `https://hooks.seahaven.com/webhooks/*` | | Ingress | shared internet-facing ALB `app/seahaven-com` → target group `open-swe-prod-tg` → EC2 `i-08a729e50779c4b07` (`t4g.large`, ARM64) on nginx `:80` | | Backend | `langgraph dev` bound to `127.0.0.1:2024` (loopback only); nginx is the sole ingress | | CDK stacks | `open-swe-iam` (OIDC roles) · `open-swe-dev` · `open-swe-prod` | Dev mirrors prod with `-dev` hosts (`openswe-dev.seahaven.com` / `hooks-dev.seahaven.com`), a `t4g.medium` box, and no GitHub-App/Slack/webhook integration (it is a deployment-validation env, not a live-triggered agent). > The retired on-prem `*.seahavenind.com` ALB routing and DNS were removed on > 2026-06-29; prod is now live exclusively on `*.seahaven.com`. ## Hosting model ``` GitHub / Slack ──▶ hooks.seahaven.com ──┐ │ (shared ALB :443, host+path rules) Browser ─────────▶ openswe.seahaven.com ──┤ ▼ ALB app/seahaven-com ──▶ open-swe-prod-tg ──▶ EC2 box :80 (nginx) ├─ nginx — SPA + scoped proxy │ /dashboard/api/* and /webhooks/* └─ langgraph dev 127.0.0.1:2024 └─▶ LangSmith cloud sandbox (build/git/PR) ``` - A **single** VPC and a **single** internet-facing ALB (`app/seahaven-com`) are shared with the on-prem `seahaven-site` stack. open-swe **imports** the VPC, ALB SG, `:443` listener, and `seahaven.com` zone — it never owns/mutates them; it only adds its own instance SG, a standalone ALB-egress rule, two listener rules, a target group, and Route53 aliases. - The EC2 box is in a **private** subnet (us-east-1a, same AZ as the single NAT for in-AZ egress). It is reachable **only** from the shared ALB SG on `:80`. - **nginx is the security boundary.** It serves the static dashboard SPA and proxies exactly two prefixes to `:2024` — `/dashboard/api/*` and `/webhooks/*`. The unauthenticated LangGraph API (`/threads`, `/runs`, `/assistants`, `/store`) is never proxied; those paths return the SPA shell. `:2024` is loopback-only and never network-reachable, even inside the SG. - **Webhooks** ride listener rules below the on-prem host-agnostic `/webhooks/*` rule (priority 2 dev / 3 prod, host-scoped to the open-swe hosts) so they reach the open-swe box and never steal an on-prem host's webhooks. The box holds **no durable state of its own**: secrets/config are materialized to a tmpfs `.env` at boot, the app artifact is pulled from S3, and the in-memory LangGraph store is re-seeded on every start. Replacement is tolerated; there is no RETAIN volume. --- ## Deploy pipeline (end to end) Two independent CD lanes, both OIDC-only (no static keys), both with a manual approval gate on prod via the GitHub **`prod` Environment** (required reviewer: Adam). The `environment: prod` declaration both fires the approval gate and makes the OIDC subject `…:environment:prod`, which is the only subject the prod deploy roles trust — so a dev-branch token can never reach prod. ### (a) Infra CD — `cd-infra.yml` Deploys the CDK stacks. Path-filtered to `infra/**`. ``` push to dev → Infra CI (tsc + jest + cdk synth) → cdk deploy OpenSweDevStack (AUTO, CI-green-gated) push to main → Infra CI → cdk deploy OpenSweProdStack (manual approval: env "prod") ``` - Roles: `githubdeploy-open-swe-infra-{dev,prod}` (in the `open-swe-iam` stack; set as repo variables `AWS_DEPLOY_ROLE_INFRA_{DEV,PROD}`). - It targets **one stack explicitly per env** (`cdk deploy OpenSweDevStack` / `OpenSweProdStack`), not `cdk deploy --all`, so a single-env push can never deploy the other env or the shared IAM stack. - The shared `open-swe-iam` stack (owns both envs' OIDC deploy roles) is **not** deployed by CD — it is a privileged, human-gated apply. **Stack order on a clean account:** `open-swe-iam` first (creates the OIDC roles; set the repo deploy-role variables and configure the `prod` Environment reviewer from its outputs), then `open-swe-dev`, then `open-swe-prod`. ### (b) Seed the config store — `put-config.sh ` Run **after** `cdk deploy open-swe-` and **before** the box first boots. CDK creates the value-less Secrets Manager shells (`open-swe-/`) and the IaC-managed SSM params (`/open-swe-/`); `put-config.sh` populates the secret values plus the out-of-band SSM params that cannot live in IaC. ```bash deploy/seahaven/put-config.sh # set each value inline, via OPENSWE_PUT_, or from a vault deploy/seahaven/fetch-config.sh # (on the box) fail-fast verify before first start ``` `put-config.sh` ships `` placeholders only — **no real secret values are committed**. It does not touch the IaC-managed SSM params (CDK owns those). **13 prod boot-required vars** — `fetch-config.sh` fail-fasts (refuses to write a partial `.env`, the unit does not start) if any are missing/empty: - **9 secrets** (Secrets Manager `open-swe-prod/`): `DASHBOARD_JWT_SECRET`, `TOKEN_ENCRYPTION_KEY`, `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `LANGSMITH_API_KEY_PROD`, `GITHUB_APP_PRIVATE_KEY`, `GITHUB_APP_CLIENT_SECRET`, `GITHUB_WEBHOOK_SECRET`, `SLACK_SIGNING_SECRET`. - **4 SSM params** (`/open-swe-prod/`): `DEFAULT_SANDBOX_SNAPSHOT_ID`, `GITHUB_APP_ID`, `GITHUB_APP_INSTALLATION_ID`, `GITHUB_APP_CLIENT_ID`. (`ANTHROPIC_API_KEY` + `OPENAI_API_KEY` are required because that is the seeded cross-family pair; the active set follows `REQUIRED_PROVIDER_KEYS`. The `LANGSMITH_API_KEY_PROD` + `DEFAULT_SANDBOX_SNAPSHOT_ID` pair is required because `SANDBOX_TYPE=langsmith`.) Dev boots without the GitHub-App / Slack / webhook secrets — it has no such integration. `fetch-config.sh` runs as an `ExecStartPre=+` hook (root, only long enough to write the `openswe`-owned `0600` tmpfs `.env`), reads all `/open-swe-/*` SSM params + all `open-swe-/*` secrets via the instance role, and forces `DEFAULT_REPO_OWNER` away from the upstream `langchain-ai` org. ### (c) App artifact deploy — `build-artifacts.yml` Builds the release and rolls the box. Path-filtered to `agent/**`, `ui/**`, `deploy/**`, `langgraph.json`, `pyproject.toml`, `uv.lock`. ``` push to dev → build SPA + package → open-swe-dev-assets/releases/ → SSM open-swe-dev-deploy (AUTO) push to main → build SPA + package → open-swe-prod-assets/releases/ → SSM open-swe-prod-deploy (manual approval: env "prod") ``` 1. The dashboard SPA is built **on the runner** (`bun run build` → vite → `ui/.output/public`) — the box is small, so the memory-heavy build runs in CI. 2. `package-artifacts.sh` produces two tarballs: `spa.tar.gz` (built SPA) and `app.tar.gz` (Python source tree — no `ui/`, no `.venv`). 3. Both are uploaded to S3 `open-swe--assets` under `releases//` (immutable, auditable) and mirrored to `releases/latest/` (what the box pulls). 4. CI fires the `open-swe--deploy` SSM document (tag-scoped to `project=open-swe,env=`), which runs `/opt/open-swe/bin/deploy.sh` on the box: pull the release from S3, build a native-ARM64 venv with `uv sync --frozen --no-dev`, extract the SPA to the nginx web root, `systemctl restart open-swe.service`, reload nginx, then gate on `systemctl is-active --quiet open-swe.service` (a non-active unit exits the deploy non-zero). Roles: `githubdeploy-open-swe-app-{dev,prod}` (repo variables `AWS_DEPLOY_ROLE_APP_{DEV,PROD}`) — tag-scoped `ssm:SendCommand` on the deploy document only (not the generic `AWS-RunShellScript`) + write to the env's S3 bucket. Secrets/config are **not** fetched by `deploy.sh`; the `systemctl restart`'s `ExecStartPre=fetch-config.sh` re-materializes the `.env` on every restart, so a bad config surfaces as a failed unit. ### (d) dev → main promotion + rollback **Promotion — `promote-dev-to-prod.yml`** (nightly cron `0 8 * * *` + manual dispatch): mints a GitHub App installation token (a bypass actor on the `main` ruleset), gates on **every check-run on the dev HEAD commit being completed and passing**, then **fast-forward-only** pushes `dev` → `main`. A diverged `main` fails loudly rather than force-updating. The push to `main` is what triggers the prod lanes of `cd-infra.yml` / `build-artifacts.yml` (each still behind the `prod` Environment approval). Re-gating via a PR on `main` would be redundant since the commit already passed every check on dev. **Rollback — `rollback.yml`** (manual dispatch, `env` + optional `sha`): re-points `releases/latest/` at a prior release and re-fires the `open-swe--deploy` SSM document — same fire/wait/gate path as a forward deploy, no rebuild. ``` env=dev, sha blank → restore open-swe-dev-assets/releases/last-good/ (AUTO) env=prod, sha blank → restore open-swe-prod-assets/releases/last-good/ (manual approval: env "prod") sha= → restore that exact releases// instead ``` It reuses the existing `githubdeploy-open-swe-app-` role (no new IAM). --- ## On-box layout (reference) | Path | What | |---|---| | `open-swe.service` (systemd) | `langgraph dev --host 127.0.0.1 --port 2024 --no-browser --no-reload` as the unprivileged `openswe` user. In-memory runtime. | | `fetch-config.sh` | `ExecStartPre=+` — materializes the tmpfs `.env` from Secrets Manager + SSM, fail-fast. | | `seed_store.sh` | `ExecStartPost` — re-seeds `team_settings/default` + `user_mappings` (the in-memory store loses them on every restart). | | nginx | SPA from `/var/www/open-swe`, proxy `/dashboard/api/` + `/webhooks/` → `127.0.0.1:2024`, `/healthz` → 200. | | `deploy.sh` | the release procedure run on first boot (non-fatal) and by every SSM deploy. | | CloudWatch logs | `/open-swe//{app,user-data,nginx-access,nginx-error}` at 30-day retention. | The live systemd unit + nginx site are the **AMI templates** (`deploy/ami/templates/open-swe.service`, `open-swe.nginx.conf`), rendered at first boot by `deploy/ami/user-data.sh`. The AMI is the baked `open-swe-base-arm64` image (Ubuntu 24.04 + uv/py3.12 + nginx + CW agent), pinned by exact id in `infra/lib/constructs/ami-cache.ts`. There is intentionally **no** on-box swapfile — the OOM-prone SPA build now runs in CI, not on the box. > `deploy/seahaven/{nginx/openswe.conf,systemd/open-swe.service}` are the > **retired on-prem VM** variants (run as `adam` from a home dir, bound `0.0.0.0`, > Postgres-backed). They are kept only for on-prem-contrast reference and are not > used by the AWS deployment. ### Models Model selection is **store-driven**, not env. The `team_settings/default` store doc wins (then per-user profile, then per-thread); `LLM_MODEL_ID` is only a seed-time fallback. Defaults seeded by `seed_store.sh`: - builder: `bedrock_converse:us.anthropic.claude-opus-4-8` (effort `high`) - reviewer: `bedrock_converse:us.anthropic.claude-opus-4-8` (effort `high`) — set `SEED_REVIEWER_MODEL` (or change it in the UI) to a Fireworks model if you want a cross-family reviewer. Only ids present in `SUPPORTED_MODELS` (`agent/dashboard/options.py`) are valid; OpenAI/Google models were removed in the Bedrock/Fireworks migration. - the `analyzer` graph is hardcoded to the code default and ignores team settings. ## Triggering Mention **`@openswe`** (or `@open-swe` / `@seahaven-openswe`) in a GitHub issue or PR comment, a Linear comment, or a Slack thread. The commenter must have a `user_mappings` entry (seeded by `seed_store.sh` from `CONFIGURED_ADMINS` / `SEED_USER_MAPPINGS`) or the run is skipped. Live integration endpoints (set in each provider's app config): | Integration | URL | |---|---| | GitHub webhook | `https://hooks.seahaven.com/webhooks/github` | | Slack events | `https://hooks.seahaven.com/webhooks/slack` (+ `/webhooks/slack/interactivity`) | | Linear webhook | `https://hooks.seahaven.com/webhooks/linear` | | GitHub OAuth callback | `https://openswe.seahaven.com/dashboard/api/auth/callback` | --- ## Troubleshooting **RETAIN secret-shell orphan on stack re-create.** The Secrets Manager shells use `DeletionPolicy: Retain` + a fixed `open-swe-/` name. If a stack's first create rolls back (or on a teardown/rebuild, a secret logical-id refactor, or standing up a new env), the empty shells survive and keep their global names, so every later create fails `AlreadyExists` — and a plain `delete-secret` does not free the name (it stays reserved for the 7–30 day recovery window). Before re-creating the stack, **force-delete the empty orphans** (only shells with no value version — never a populated secret). Hit on prod 2026-06-29 (PR #51 deploy failure). Full recovery command + rationale: [`infra/README.md`](../../infra/README.md) (PR #52). **`langgraph dev` won't start after a deploy.** `fetch-config.sh` fail-fasts on a missing/empty required var and prints the offending variable **names** (never values) to the unit journal. Confirm the 13 prod boot-required vars are populated (`put-config.sh prod`), then `systemctl restart open-swe.service`. **ALB target unhealthy.** The TG health check is `GET /healthz` on nginx `:80` (static 200). nginx starts before the app on first boot, so an unhealthy target usually means the box can't reach the ALB SG on `:80` (the standalone ALB-egress rule) rather than an app fault. --- ## Aegra (deferred) `aegra/aegra.json` + `aegra/aegra_entry.py` are the self-hosted-runtime alternative (Apache-2.0, avoids the LangGraph-Platform Elastic license). Not active on the stock deployment. To use: place both at the repo root, run `aegra serve` (:2026), and point `LANGGRAPH_URL` at `:2026`. Aegra gives a Postgres-backed durable store/checkpointer, which removes the need for `seed_store.sh` and survives restarts (paused HITL interrupts persist).