feat(deploy): add Sea Haven self-hosted deployment capture

Captures the stock-LangGraph deployment of this fork at Sea Haven:
- systemd/open-swe.service: langgraph dev (:2024) + store seed ExecStartPost
- seed_store.sh: re-seeds team_settings + user_mappings (in-memory store
  resets on restart); env-parameterized, no secrets
- nginx/openswe.conf: dashboard SPA + scoped /dashboard/api proxy (security
  boundary; agent API not exposed)
- aegra/: deferred self-hosted-runtime alternative (not active on stock)
- DEPLOYMENT.md: full runbook (models, build, ingress, OAuth callback)

Secrets and internal infra identifiers are intentionally excluded (public
fork); real values live in private IT docs.
This commit is contained in:
Adam Moussa 2026-06-25 17:27:08 -04:00
parent f6c215fff7
commit cf57ba7aa5
6 changed files with 252 additions and 0 deletions

View file

@ -0,0 +1,99 @@
# Sea Haven — Open SWE self-hosted deployment
How this fork is deployed at Sea Haven. The runtime is the **stock LangGraph dev
server** (not the Aegra path — see [Aegra](#aegra-deferred)). Internal addresses,
ARNs, and account IDs are shown as `<PLACEHOLDERS>`; the real values live in the
private IT docs (Confluence "AWS Architecture Map") — **do not commit them to this
public fork.**
## Topology
```
GitHub / Slack ──▶ hooks.seahavenind.com ──┐
│ (public ALB :443, host+path rule)
Browser ─────────▶ openswe.seahavenind.com ─┤
▼
AWS ALB ──(Site-to-Site VPN)──▶ on-prem VM
├─ nginx :80 (dashboard SPA + /dashboard/api proxy)
└─ langgraph dev :2024 (3+ graphs + FastAPI webapp)
└─▶ LangSmith cloud sandbox (build/git/PR)
```
- The VM is **internet-closed**; all inbound rides the existing ALB over the VPN.
- **Webhooks** (`hooks.seahavenind.com`) → ALB listener rule scoped to `/webhooks/*`
only → VM `:2024`. The unauthenticated LangGraph API (`/threads`, `/runs`,
`/assistants`, `/store`) is never path-forwarded.
- **Dashboard** (`openswe.seahavenind.com`) → ALB → VM `:80` (nginx). nginx is the
security boundary: it serves the static SPA and proxies **only** `/dashboard/api/`
to `:2024`; the agent API is not reachable through it.
## VM components
| Component | What |
|---|---|
| `langgraph dev` | systemd `open-swe.service` — `--host 0.0.0.0 --port 2024 --no-browser --no-reload`. In-memory runtime. |
| Store seeding | `seed_store.sh` as `ExecStartPost` (re-seeds team settings + user mappings, which the in-memory store loses on restart). |
| nginx | `nginx/openswe.conf` — SPA from `/var/www/openswe`, proxy `/dashboard/api/` → `:2024`. |
| Postgres | present (was for the Aegra path); unused by the stock in-memory runtime. |
| swap | 8 GB swapfile — **required**: the dashboard (`ui/`) Nitro build OOMs on an 8 GB box without it. |
### Models
Model selection is **store-driven**, not env. `LLM_MODEL_ID` is effectively dead for
runtime selection; the `team_settings/default` store doc wins (then per-user profile,
then per-thread). Defaults seeded by `seed_store.sh`:
- builder: `anthropic:claude-opus-4-8` (effort `high`)
- reviewer (cross-family): `openai:gpt-5.5` (effort `high`) — `openai:gpt-4.1` is **not**
in this fork's `SUPPORTED_MODELS` (`agent/dashboard/options.py`); a raw value is
silently rewritten to gpt-5.5. Add it to `SUPPORTED_MODELS` first if you need 4.1.
- `analyzer` graph is hardcoded to the code default and ignores team settings.
## Build & deploy the dashboard (`ui/`)
`ui/` is a **TanStack Start + Nitro** app (build with `bun`, not plain Vite):
```bash
cd ui
export PATH="$HOME/.bun/bin:$PATH"
export NODE_OPTIONS=--max-old-space-size=6144 # + the 8 GB swapfile, or the build OOMs
bun install
bun run build # -> .output/public (static SPA, _shell.html)
sudo cp -r .output/public/. /var/www/openswe/ # served by nginx
```
Served as a static SPA (per `ui/vercel.json`); the Nitro `.output/server` is unused.
## Install / wire-up checklist
1. App config in `.env` (gitignored — never commit): LLM keys, GitHub App creds,
`LANGSMITH_API_KEY*` + `DEFAULT_SANDBOX_SNAPSHOT_ID` (`SANDBOX_TYPE=langsmith` — the
only sandbox provider with working in-sandbox git/gh auth), `LANGGRAPH_URL=http://127.0.0.1:2024`,
dashboard vars (`DASHBOARD_JWT_SECRET`, `CONFIGURED_ADMINS`, `DASHBOARD_*_URL=https://openswe.seahavenind.com`).
2. `systemd/open-swe.service` → `/etc/systemd/system/`, `seed_store.sh` on the VM with
`OPENSWE_*` env exported (owner login/email, default repo, model ids).
3. `nginx/openswe.conf` → `/etc/nginx/sites-available/openswe`, symlink into
`sites-enabled`, remove the default site, `nginx -t && systemctl reload nginx`.
4. AWS (real IDs in Confluence): IP target groups → `<VM_LAN_IP>:2024` and `:80`;
ALB SG **egress** rules to those ports (the ALB SG is allow-listed — health checks
time out without them); `:443` listener rules for the two hostnames; Route53 ALIAS
records → ALB. Webhook rule must stay path-scoped to `/webhooks/*`.
5. GitHub App: webhook URL `https://hooks.seahavenind.com/webhooks/github`; subscribe to
the events the install guide lists (Issue comment, PR review×2, check_run/suite,
workflow_run, status) — add **Issues** too if you want issue-title/body triggers.
6. **GitHub App OAuth callback (manual, UI-only — not API-settable):**
`https://openswe.seahavenind.com/dashboard/api/auth/callback` — without it, dashboard
login fails with a `redirect_uri` mismatch.
## Triggering
Start a task by mentioning **`@openswe`** in a GitHub issue comment (the documented
intake — the `open-swe` *label* path needs the `Issues` event subscription). The
commenter must have a `user_mappings` entry or the run is skipped.
## Aegra (deferred)
`aegra/aegra.json` + `aegra/aegra_entry.py` are the self-hosted-runtime alternative
(Apache-2.0, avoids the LangGraph-Platform Elastic license). Not active on the stock
deployment. To use: place both at the repo root, run `aegra serve` (:2026), and point
`LANGGRAPH_URL` at `:2026`. Aegra gives a Postgres-backed durable store/checkpointer,
which removes the need for `seed_store.sh` and survives restarts (paused HITL
interrupts persist).

View file

@ -0,0 +1,11 @@
{
"dependencies": ["."],
"graphs": {
"agent": "./aegra_entry.py:agent_graph",
"reviewer": "./aegra_entry.py:reviewer_graph",
"analyzer": "./aegra_entry.py:analyzer_graph"
},
"http": {
"app": "./aegra_entry.py:webapp_app"
}
}

View file

@ -0,0 +1,19 @@
"""Aegra entrypoint for Open SWE graphs.
Aegra loads graph files standalone via importlib.spec_from_file_location, which
gives them a synthetic module name and breaks agent/*.py's package-relative
imports (e.g. `from .dashboard.admin import ...`). Re-exporting the graphs here
through the installed ``agent`` package (absolute imports) restores correct
``__package__`` resolution, so the relative imports inside the agent modules work.
"""
from agent.server import traced_agent as agent_graph
from agent.reviewer import traced_reviewer_agent as reviewer_graph
from agent.analyzer import traced_analyzer as analyzer_graph
# Open SWE's FastAPI webapp (GitHub/Slack webhooks + dashboard API), mounted by
# Aegra via the "http" key in aegra.json. Absolute import for the same reason as
# the graphs above (relative imports break under Aegra's standalone file loader).
from agent.webapp import app as webapp_app
__all__ = ["agent_graph", "reviewer_graph", "analyzer_graph", "webapp_app"]

View file

@ -0,0 +1,32 @@
# Open SWE dashboard frontend (TanStack Start SPA) + scoped API proxy.
# nginx is the security boundary: ONLY /dashboard/api/* reaches the backend;
# the unauthenticated LangGraph agent API (/threads,/runs,/assistants,/store) is NOT proxied.
server {
listen 80 default_server;
listen [::]:80 default_server;
server_name openswe.seahavenind.com;
root /var/www/openswe;
index _shell.html;
# ALB health check
location = /healthz { default_type text/plain; return 200 "ok\n"; }
# Dashboard API + OAuth callback -> backend webapp on :2024 (the ONLY proxied path)
location /dashboard/api/ {
proxy_pass http://127.0.0.1:2024;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto https;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_read_timeout 300s;
}
# Static assets + SPA shell fallback (client-side routing)
location / {
try_files $uri $uri/ /_shell.html;
}
}

74
deploy/seahaven/seed_store.sh Executable file
View file

@ -0,0 +1,74 @@
#!/usr/bin/env bash
# Seed the LangGraph store after a (re)start.
#
# The stock `langgraph dev` server uses an IN-MEMORY store, so anything written
# to it (team model settings, user mappings) is lost on every restart. This
# script idempotently re-PUTs that state and is wired as a systemd
# ExecStartPost on the open-swe.service unit so it runs after each start.
#
# Replace this with Postgres-backed durability (Aegra / `langgraph up`) to make
# the store survive restarts and drop this script.
#
# Configuration comes from the environment (set these in the service env or a
# sourced file alongside the app .env) — no real values are committed here:
# OPENSWE_BASE_URL default http://127.0.0.1:2024
# OPENSWE_AGENT_MODEL default anthropic:claude-opus-4-8
# OPENSWE_AGENT_EFFORT default high
# OPENSWE_REVIEWER_MODEL default openai:gpt-5.5
# OPENSWE_REVIEWER_EFFORT default high
# OPENSWE_DEFAULT_REPO e.g. your-org/your-pilot-repo (required)
# OPENSWE_OWNER_LOGIN GitHub login of the triggering owner (required)
# OPENSWE_OWNER_EMAIL work email mapped to that login (required)
set -euo pipefail
BASE="${OPENSWE_BASE_URL:-http://127.0.0.1:2024}"
AGENT_MODEL="${OPENSWE_AGENT_MODEL:-anthropic:claude-opus-4-8}"
AGENT_EFFORT="${OPENSWE_AGENT_EFFORT:-high}"
REVIEWER_MODEL="${OPENSWE_REVIEWER_MODEL:-openai:gpt-5.5}"
REVIEWER_EFFORT="${OPENSWE_REVIEWER_EFFORT:-high}"
DEFAULT_REPO="${OPENSWE_DEFAULT_REPO:?set OPENSWE_DEFAULT_REPO=owner/repo}"
OWNER_LOGIN="${OPENSWE_OWNER_LOGIN:?set OPENSWE_OWNER_LOGIN=github-login}"
OWNER_EMAIL="${OPENSWE_OWNER_EMAIL:?set OPENSWE_OWNER_EMAIL=work-email}"
NOW="$(date -u +%Y-%m-%dT%H:%M:%S+00:00)"
# Wait for the server to accept requests (up to ~60s).
for _ in $(seq 1 30); do
[ "$(curl -s -o /dev/null -w '%{http_code}' "$BASE/ok" || true)" = "200" ] && break
sleep 2
done
# 1) team_settings/default — builder + reviewer models (NOT read from env by the
# app; the store value wins over LLM_MODEL_ID). gpt-4.1 is NOT in this fork's
# SUPPORTED_MODELS, so the reviewer uses gpt-5.5 (cross-family vs the builder).
curl -s -X PUT "$BASE/store/items" -H "Content-Type: application/json" -d @- <<JSON
{"namespace":["team_settings"],"key":"default","value":{
"review_draft_prs": false,
"pr_summaries": true,
"review_trace_links": true,
"org_guidelines": null,
"default_agent_model": "$AGENT_MODEL",
"default_agent_reasoning_effort": "$AGENT_EFFORT",
"default_agent_subagent_model": "$AGENT_MODEL",
"default_agent_subagent_reasoning_effort": "$AGENT_EFFORT",
"default_repo": "$DEFAULT_REPO",
"default_reviewer_model": "$REVIEWER_MODEL",
"default_reviewer_reasoning_effort": "$REVIEWER_EFFORT",
"default_reviewer_subagent_model": "$REVIEWER_MODEL",
"default_reviewer_subagent_reasoning_effort": "$REVIEWER_EFFORT",
"default_grouping_model": null,
"default_grouping_reasoning_effort": null,
"default_chat_model": null,
"default_chat_reasoning_effort": null,
"updated_at": "$NOW"
}}
JSON
# 2) user_mappings/<login> — required, or the @openswe trigger ignores the commenter.
curl -s -X PUT "$BASE/store/items" -H "Content-Type: application/json" -d @- <<JSON
{"namespace":["user_mappings"],"key":"$OWNER_LOGIN","value":{
"github_login":"$OWNER_LOGIN","work_email":"$OWNER_EMAIL","slack_user_id":null,
"source":"slack_oauth","status":"active","created_at":"$NOW","updated_at":"$NOW"
}}
JSON
echo "seed_store: done at $NOW"

View file

@ -0,0 +1,17 @@
[Unit]
Description=Open SWE stock LangGraph dev server (graphs + webapp, :2024)
After=network-online.target postgresql.service
Wants=network-online.target
[Service]
Type=simple
User=adam
WorkingDirectory=/home/adam/open-swe
ExecStart=/home/adam/open-swe/.venv/bin/langgraph dev --host 0.0.0.0 --port 2024 --no-browser --no-reload
ExecStartPost=/home/adam/open-swe/seed_store.sh
Restart=on-failure
RestartSec=5
TimeoutStartSec=120
[Install]
WantedBy=multi-user.target