Prod went live 2026-06-29 on self-hosted AWS EC2 behind the shared seahaven-com ALB, superseding the on-prem VM model the runbook described. - Rewrite deploy/seahaven/DEPLOYMENT.md as the canonical end-to-end runbook: infra CD (CDK stacks + OIDC roles + prod approval gate), config seeding (put-config.sh, the 13 boot-required prod vars, fetch-config fail-fast), app artifact deploy (S3 + SSM roll + is-active gate), promotion/rollback, live prod facts, and a RETAIN secret-shell troubleshooting entry that cross-references infra/README.md. - Correct retired *.seahavenind.com hosts to *.seahaven.com throughout and document the live GitHub/Slack/Linear webhook + OAuth endpoints. - Add a concise Deployment section to README pointing at the runbook. - Fix the stale host in the retired on-prem nginx/openswe.conf and mark it superseded by the AMI template.
15 KiB
Sea Haven — Open SWE deployment runbook
How this fork is deployed at Sea Haven. PROD is LIVE as of 2026-06-29. The
runtime is the stock LangGraph dev server (not the Aegra path — see
Aegra), running on a self-hosted ARM64 EC2 box behind the
shared seahaven-com ALB.
Account 328440206208, region us-east-1. Internal addresses, ARNs, snapshot
ids, and account-scoped values that are sensitive are shown as <PLACEHOLDERS>;
the real values live in the private IT docs (Confluence "AWS Architecture Map",
id 1540098) and in AWS — do not commit them to this public fork.
This is the canonical deploy runbook. The CDK details live in
infra/README.md; the rotation procedure in
ROTATION.md.
Live prod facts (2026-06-29)
| Dashboard | https://openswe.seahaven.com |
| Webhooks | https://hooks.seahaven.com/webhooks/* |
| Ingress | shared internet-facing ALB app/seahaven-com → target group open-swe-prod-tg → EC2 i-08a729e50779c4b07 (t4g.large, ARM64) on nginx :80 |
| Backend | langgraph dev bound to 127.0.0.1:2024 (loopback only); nginx is the sole ingress |
| CDK stacks | open-swe-iam (OIDC roles) · open-swe-dev · open-swe-prod |
Dev mirrors prod with -dev hosts (openswe-dev.seahaven.com /
hooks-dev.seahaven.com), a t4g.medium box, and no GitHub-App/Slack/webhook
integration (it is a deployment-validation env, not a live-triggered agent).
The retired on-prem
*.seahavenind.comALB routing and DNS were removed on 2026-06-29; prod is now live exclusively on*.seahaven.com.
Hosting model
GitHub / Slack ──▶ hooks.seahaven.com ──┐
│ (shared ALB :443, host+path rules)
Browser ─────────▶ openswe.seahaven.com ──┤
▼
ALB app/seahaven-com ──▶ open-swe-prod-tg ──▶ EC2 box :80 (nginx)
├─ nginx — SPA + scoped proxy
│ /dashboard/api/* and /webhooks/*
└─ langgraph dev 127.0.0.1:2024
└─▶ LangSmith cloud sandbox (build/git/PR)
- A single VPC and a single internet-facing ALB (
app/seahaven-com) are shared with the on-premseahaven-sitestack. open-swe imports the VPC, ALB SG,:443listener, andseahaven.comzone — it never owns/mutates them; it only adds its own instance SG, a standalone ALB-egress rule, two listener rules, a target group, and Route53 aliases. - The EC2 box is in a private subnet (us-east-1a, same AZ as the single NAT
for in-AZ egress). It is reachable only from the shared ALB SG on
:80. - nginx is the security boundary. It serves the static dashboard SPA and
proxies exactly two prefixes to
:2024—/dashboard/api/*and/webhooks/*. The unauthenticated LangGraph API (/threads,/runs,/assistants,/store) is never proxied; those paths return the SPA shell.:2024is loopback-only and never network-reachable, even inside the SG. - Webhooks ride listener rules below the on-prem host-agnostic
/webhooks/*rule (priority 2 dev / 3 prod, host-scoped to the open-swe hosts) so they reach the open-swe box and never steal an on-prem host's webhooks.
The box holds no durable state of its own: secrets/config are materialized to
a tmpfs .env at boot, the app artifact is pulled from S3, and the in-memory
LangGraph store is re-seeded on every start. Replacement is tolerated; there is no
RETAIN volume.
Deploy pipeline (end to end)
Two independent CD lanes, both OIDC-only (no static keys), both with a manual
approval gate on prod via the GitHub prod Environment (required reviewer:
Adam). The environment: prod declaration both fires the approval gate and makes
the OIDC subject …:environment:prod, which is the only subject the prod deploy
roles trust — so a dev-branch token can never reach prod.
(a) Infra CD — cd-infra.yml
Deploys the CDK stacks. Path-filtered to infra/**.
push to dev → Infra CI (tsc + jest + cdk synth) → cdk deploy OpenSweDevStack (AUTO, CI-green-gated)
push to main → Infra CI → cdk deploy OpenSweProdStack (manual approval: env "prod")
- Roles:
githubdeploy-open-swe-infra-{dev,prod}(in theopen-swe-iamstack; set as repo variablesAWS_DEPLOY_ROLE_INFRA_{DEV,PROD}). - It targets one stack explicitly per env (
cdk deploy OpenSweDevStack/OpenSweProdStack), notcdk deploy --all, so a single-env push can never deploy the other env or the shared IAM stack. - The shared
open-swe-iamstack (owns both envs' OIDC deploy roles) is not deployed by CD — it is a privileged, human-gated apply.
Stack order on a clean account: open-swe-iam first (creates the OIDC roles;
set the repo deploy-role variables and configure the prod Environment reviewer
from its outputs), then open-swe-dev, then open-swe-prod.
(b) Seed the config store — put-config.sh <env>
Run after cdk deploy open-swe-<env> and before the box first boots. CDK
creates the value-less Secrets Manager shells (open-swe-<env>/<VAR>) and the
IaC-managed SSM params (/open-swe-<env>/<VAR>); put-config.sh populates the
secret values plus the out-of-band SSM params that cannot live in IaC.
deploy/seahaven/put-config.sh <dev|prod> # set each value inline, via OPENSWE_PUT_<VAR>, or from a vault
deploy/seahaven/fetch-config.sh <dev|prod> # (on the box) fail-fast verify before first start
put-config.sh ships <FILL> placeholders only — no real secret values are
committed. It does not touch the IaC-managed SSM params (CDK owns those).
13 prod boot-required vars — fetch-config.sh fail-fasts (refuses to write a
partial .env, the unit does not start) if any are missing/empty:
- 9 secrets (Secrets Manager
open-swe-prod/<VAR>):DASHBOARD_JWT_SECRET,TOKEN_ENCRYPTION_KEY,ANTHROPIC_API_KEY,OPENAI_API_KEY,LANGSMITH_API_KEY_PROD,GITHUB_APP_PRIVATE_KEY,GITHUB_APP_CLIENT_SECRET,GITHUB_WEBHOOK_SECRET,SLACK_SIGNING_SECRET. - 4 SSM params (
/open-swe-prod/<VAR>):DEFAULT_SANDBOX_SNAPSHOT_ID,GITHUB_APP_ID,GITHUB_APP_INSTALLATION_ID,GITHUB_APP_CLIENT_ID.
(ANTHROPIC_API_KEY + OPENAI_API_KEY are required because that is the seeded
cross-family pair; the active set follows REQUIRED_PROVIDER_KEYS. The
LANGSMITH_API_KEY_PROD + DEFAULT_SANDBOX_SNAPSHOT_ID pair is required because
SANDBOX_TYPE=langsmith.) Dev boots without the GitHub-App / Slack / webhook
secrets — it has no such integration.
fetch-config.sh runs as an ExecStartPre=+ hook (root, only long enough to
write the openswe-owned 0600 tmpfs .env), reads all /open-swe-<env>/* SSM
params + all open-swe-<env>/* secrets via the instance role, and forces
DEFAULT_REPO_OWNER away from the upstream langchain-ai org.
(c) App artifact deploy — build-artifacts.yml
Builds the release and rolls the box. Path-filtered to agent/**, ui/**,
deploy/**, langgraph.json, pyproject.toml, uv.lock.
push to dev → build SPA + package → open-swe-dev-assets/releases/ → SSM open-swe-dev-deploy (AUTO)
push to main → build SPA + package → open-swe-prod-assets/releases/ → SSM open-swe-prod-deploy (manual approval: env "prod")
- The dashboard SPA is built on the runner (
bun run build→ vite →ui/.output/public) — the box is small, so the memory-heavy build runs in CI. package-artifacts.shproduces two tarballs:spa.tar.gz(built SPA) andapp.tar.gz(Python source tree — noui/, no.venv).- Both are uploaded to S3
open-swe-<env>-assetsunderreleases/<sha>/(immutable, auditable) and mirrored toreleases/latest/(what the box pulls). - CI fires the
open-swe-<env>-deploySSM document (tag-scoped toproject=open-swe,env=<env>), which runs/opt/open-swe/bin/deploy.shon the box: pull the release from S3, build a native-ARM64 venv withuv sync --frozen --no-dev, extract the SPA to the nginx web root,systemctl restart open-swe.service, reload nginx, then gate onsystemctl is-active --quiet open-swe.service(a non-active unit exits the deploy non-zero).
Roles: githubdeploy-open-swe-app-{dev,prod} (repo variables
AWS_DEPLOY_ROLE_APP_{DEV,PROD}) — tag-scoped ssm:SendCommand on the deploy
document only (not the generic AWS-RunShellScript) + write to the env's S3
bucket.
Secrets/config are not fetched by deploy.sh; the systemctl restart's
ExecStartPre=fetch-config.sh re-materializes the .env on every restart, so a
bad config surfaces as a failed unit.
(d) dev → main promotion + rollback
Promotion — promote-dev-to-prod.yml (nightly cron 0 8 * * * + manual
dispatch): mints a GitHub App installation token (a bypass actor on the main
ruleset), gates on every check-run on the dev HEAD commit being completed and
passing, then fast-forward-only pushes dev → main. A diverged main
fails loudly rather than force-updating. The push to main is what triggers the
prod lanes of cd-infra.yml / build-artifacts.yml (each still behind the prod
Environment approval). Re-gating via a PR on main would be redundant since the
commit already passed every check on dev.
Rollback — rollback.yml (manual dispatch, env + optional sha): re-points
releases/latest/ at a prior release and re-fires the open-swe-<env>-deploy SSM
document — same fire/wait/gate path as a forward deploy, no rebuild.
env=dev, sha blank → restore open-swe-dev-assets/releases/last-good/ (AUTO)
env=prod, sha blank → restore open-swe-prod-assets/releases/last-good/ (manual approval: env "prod")
sha=<commit> → restore that exact releases/<sha>/ instead
It reuses the existing githubdeploy-open-swe-app-<env> role (no new IAM).
On-box layout (reference)
| Path | What |
|---|---|
open-swe.service (systemd) |
langgraph dev --host 127.0.0.1 --port 2024 --no-browser --no-reload as the unprivileged openswe user. In-memory runtime. |
fetch-config.sh |
ExecStartPre=+ — materializes the tmpfs .env from Secrets Manager + SSM, fail-fast. |
seed_store.sh |
ExecStartPost — re-seeds team_settings/default + user_mappings (the in-memory store loses them on every restart). |
| nginx | SPA from /var/www/open-swe, proxy /dashboard/api/ + /webhooks/ → 127.0.0.1:2024, /healthz → 200. |
deploy.sh |
the release procedure run on first boot (non-fatal) and by every SSM deploy. |
| CloudWatch logs | /open-swe/<env>/{app,user-data,nginx-access,nginx-error} at 30-day retention. |
The live systemd unit + nginx site are the AMI templates
(deploy/ami/templates/open-swe.service, open-swe.nginx.conf), rendered at
first boot by deploy/ami/user-data.sh. The AMI is the baked
open-swe-base-arm64 image (Ubuntu 24.04 + uv/py3.12 + nginx + CW agent), pinned
by exact id in infra/lib/constructs/ami-cache.ts. There is intentionally no
on-box swapfile — the OOM-prone SPA build now runs in CI, not on the box.
deploy/seahaven/{nginx/openswe.conf,systemd/open-swe.service}are the retired on-prem VM variants (run asadamfrom a home dir, bound0.0.0.0, Postgres-backed). They are kept only for on-prem-contrast reference and are not used by the AWS deployment.
Models
Model selection is store-driven, not env. The team_settings/default store
doc wins (then per-user profile, then per-thread); LLM_MODEL_ID is only a
seed-time fallback. Defaults seeded by seed_store.sh:
- builder:
anthropic:claude-opus-4-8(efforthigh) - reviewer (cross-family):
openai:gpt-5.5(efforthigh) —openai:gpt-4.1is not in this fork'sSUPPORTED_MODELS(agent/dashboard/options.py); add it there first if you need 4.1. - the
analyzergraph is hardcoded to the code default and ignores team settings.
Triggering
Mention @openswe (or @open-swe / @seahaven-openswe) in a GitHub issue or
PR comment, a Linear comment, or a Slack thread. The commenter must have a
user_mappings entry (seeded by seed_store.sh from CONFIGURED_ADMINS /
SEED_USER_MAPPINGS) or the run is skipped.
Live integration endpoints (set in each provider's app config):
| Integration | URL |
|---|---|
| GitHub webhook | https://hooks.seahaven.com/webhooks/github |
| Slack events | https://hooks.seahaven.com/webhooks/slack (+ /webhooks/slack/interactivity) |
| Linear webhook | https://hooks.seahaven.com/webhooks/linear |
| GitHub OAuth callback | https://openswe.seahaven.com/dashboard/api/auth/callback |
Troubleshooting
RETAIN secret-shell orphan on stack re-create. The Secrets Manager shells use
DeletionPolicy: Retain + a fixed open-swe-<env>/<VAR> name. If a stack's first
create rolls back (or on a teardown/rebuild, a secret logical-id refactor, or
standing up a new env), the empty shells survive and keep their global names, so
every later create fails AlreadyExists — and a plain delete-secret does not
free the name (it stays reserved for the 7–30 day recovery window). Before
re-creating the stack, force-delete the empty orphans (only shells with no
value version — never a populated secret). Hit on prod 2026-06-29 (PR #51 deploy
failure). Full recovery command + rationale:
infra/README.md (PR #52).
langgraph dev won't start after a deploy. fetch-config.sh fail-fasts on a
missing/empty required var and prints the offending variable names (never
values) to the unit journal. Confirm the 13 prod boot-required vars are populated
(put-config.sh prod), then systemctl restart open-swe.service.
ALB target unhealthy. The TG health check is GET /healthz on nginx :80
(static 200). nginx starts before the app on first boot, so an unhealthy target
usually means the box can't reach the ALB SG on :80 (the standalone ALB-egress
rule) rather than an app fault.
Aegra (deferred)
aegra/aegra.json + aegra/aegra_entry.py are the self-hosted-runtime
alternative (Apache-2.0, avoids the LangGraph-Platform Elastic license). Not
active on the stock deployment. To use: place both at the repo root, run
aegra serve (:2026), and point LANGGRAPH_URL at :2026. Aegra gives a
Postgres-backed durable store/checkpointer, which removes the need for
seed_store.sh and survives restarts (paused HITL interrupts persist).