open-swe/deploy/seahaven/DEPLOYMENT.md
Adam Moussa 434d0ad80c
feat(deploy): AWS-sourced fetch-config + seed_store + rotation docs (PR#7) (#8)
PR#7 of the AWS migration. deploy/seahaven/: fetch-config.sh materializes a
service-user-owned 0600 tmpfs .env from Secrets Manager + SSM (fail-fast);
seed_store.sh reseeds the in-memory store; ROTATION.md.

Incorporates T5 /sh-security-review fixes:
- seed_store no longer bash-sources the .env (closes the SH-INJ-001 RCE); uses a
  non-eval reader, jq --arg JSON bodies, and a loopback-pinned BASE.
- fetch-config: .env owned by the openswe service user (app no longer runs as
  root); DEFAULT_REPO_OWNER hard-pinned; key-identifier validation + flat-namespace
  collision detection; dropped SSM --recursive.
shellcheck + bash -n clean.
2026-06-26 15:06:41 -04:00

118 lines
7.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Sea Haven — Open SWE self-hosted deployment
How this fork is deployed at Sea Haven. The runtime is the **stock LangGraph dev
server** (not the Aegra path — see [Aegra](#aegra-deferred)). Internal addresses,
ARNs, and account IDs are shown as `<PLACEHOLDERS>`; the real values live in the
private IT docs (Confluence "AWS Architecture Map") — **do not commit them to this
public fork.**
## Topology
```
GitHub / Slack ──▶ hooks.seahavenind.com ──┐
│ (public ALB :443, host+path rule)
Browser ─────────▶ openswe.seahavenind.com ─┤
▼
AWS ALB ──(Site-to-Site VPN)──▶ on-prem VM
├─ nginx :80 (dashboard SPA + /dashboard/api proxy)
└─ langgraph dev :2024 (3+ graphs + FastAPI webapp)
└─▶ LangSmith cloud sandbox (build/git/PR)
```
- The VM is **internet-closed**; all inbound rides the existing ALB over the VPN.
- **Webhooks** (`hooks.seahavenind.com`) → ALB listener rule scoped to `/webhooks/*`
only → VM `:2024`. The unauthenticated LangGraph API (`/threads`, `/runs`,
`/assistants`, `/store`) is never path-forwarded.
- **Dashboard** (`openswe.seahavenind.com`) → ALB → VM `:80` (nginx). nginx is the
security boundary: it serves the static SPA and proxies **only** `/dashboard/api/`
to `:2024`; the agent API is not reachable through it.
## VM components
| Component | What |
|---|---|
| `langgraph dev` | systemd `open-swe.service` — `--host 0.0.0.0 --port 2024 --no-browser --no-reload`. In-memory runtime. |
| Store seeding | `seed_store.sh` as `ExecStartPost` (re-seeds team settings + user mappings, which the in-memory store loses on restart). |
| nginx | `nginx/openswe.conf` — SPA from `/var/www/openswe`, proxy `/dashboard/api/` → `:2024`. |
| Postgres | present (was for the Aegra path); unused by the stock in-memory runtime. |
| swap | 8 GB swapfile — **required**: the dashboard (`ui/`) Nitro build OOMs on an 8 GB box without it. |
### Models
Model selection is **store-driven**, not env. `LLM_MODEL_ID` is effectively dead for
runtime selection; the `team_settings/default` store doc wins (then per-user profile,
then per-thread). Defaults seeded by `seed_store.sh`:
- builder: `anthropic:claude-opus-4-8` (effort `high`)
- reviewer (cross-family): `openai:gpt-5.5` (effort `high`) — `openai:gpt-4.1` is **not**
in this fork's `SUPPORTED_MODELS` (`agent/dashboard/options.py`); a raw value is
silently rewritten to gpt-5.5. Add it to `SUPPORTED_MODELS` first if you need 4.1.
- `analyzer` graph is hardcoded to the code default and ignores team settings.
## Build & deploy the dashboard (`ui/`)
`ui/` is a **TanStack Start + Nitro** app (build with `bun`, not plain Vite):
```bash
cd ui
export PATH="$HOME/.bun/bin:$PATH"
export NODE_OPTIONS=--max-old-space-size=6144 # + the 8 GB swapfile, or the build OOMs
bun install
bun run build # -> .output/public (static SPA, _shell.html)
sudo cp -r .output/public/. /var/www/openswe/ # served by nginx
```
Served as a static SPA (per `ui/vercel.json`); the Nitro `.output/server` is unused.
## Install / wire-up checklist
1. App config in `.env` (gitignored — never commit): LLM keys, GitHub App creds,
`LANGSMITH_API_KEY*` + `DEFAULT_SANDBOX_SNAPSHOT_ID` (`SANDBOX_TYPE=langsmith` — the
only sandbox provider with working in-sandbox git/gh auth), `LANGGRAPH_URL=http://127.0.0.1:2024`,
dashboard vars (`DASHBOARD_JWT_SECRET`, `CONFIGURED_ADMINS`, `DASHBOARD_*_URL=https://openswe.seahavenind.com`).
2. `systemd/open-swe.service` → `/etc/systemd/system/`, `seed_store.sh` on the VM with
`OPENSWE_*` env exported (owner login/email, default repo, model ids).
3. `nginx/openswe.conf` → `/etc/nginx/sites-available/openswe`, symlink into
`sites-enabled`, remove the default site, `nginx -t && systemctl reload nginx`.
4. AWS (real IDs in Confluence): IP target groups → `<VM_LAN_IP>:2024` and `:80`;
ALB SG **egress** rules to those ports (the ALB SG is allow-listed — health checks
time out without them); `:443` listener rules for the two hostnames; Route53 ALIAS
records → ALB. Webhook rule must stay path-scoped to `/webhooks/*`.
5. GitHub App: webhook URL `https://hooks.seahavenind.com/webhooks/github`; subscribe to
the events the install guide lists (Issue comment, PR review×2, check_run/suite,
workflow_run, status) — add **Issues** too if you want issue-title/body triggers.
6. **GitHub App OAuth callback (manual, UI-only — not API-settable):**
`https://openswe.seahavenind.com/dashboard/api/auth/callback` — without it, dashboard
login fails with a `redirect_uri` mismatch.
## Triggering
Start a task by mentioning **`@openswe`** in a GitHub issue comment (the documented
intake — the `open-swe` *label* path needs the `Issues` event subscription). The
commenter must have a `user_mappings` entry or the run is skipped.
## AWS variant — env-sourced config (.env from Secrets Manager + SSM)
On the AWS lift-and-shift, the `.env` is **not** staged by hand. A boot hook pulls
config from AWS via the EC2 instance role and materializes a root-only `.env` on a
tmpfs before the service starts. Same stock runtime; only how `.env` is produced
changes.
| File | Role |
|---|---|
| `fetch-config.sh <dev\|prod>` | `ExecStartPre=+` hook. Reads all SSM params `/open-swe-<env>/*` + all secrets `open-swe-<env>/*`, FAIL-FAST on any missing required var, writes `/run/open-swe/.env` (tmpfs, root:root 0600), symlinks `<APP_DIR>/.env` → it. Forces `DEFAULT_REPO_OWNER`=`Sea-Haven-Industries` (never upstream `langchain-ai`). |
| `seed_store.sh <dev\|prod>` | `ExecStartPost` hook (unchanged role). Now sources the materialized `.env`, so `default_repo`/models/user-mappings come from AWS config instead of hardcoded `OPENSWE_*`. Still re-seeds the in-memory store on **every** restart. |
| `ROTATION.md` | How to rotate any secret/param (update AWS → `systemctl restart`) + per-secret notes (`TOKEN_ENCRYPTION_KEY` overlap list, GitHub App PEM, webhook signing secrets). |
Because the `.env` is root-only, the AWS `open-swe.service` runs **as root** (the
on-prem unit's `User=adam` cannot read it). See the header of `fetch-config.sh` for
the exact unit snippet and the tmpfs / `RUN_DEDICATED_TMPFS` options. Naming/secret
placement follows the T9 inventory: 29 secrets → Secrets Manager `open-swe-<env>/*`,
53 config → SSM `/open-swe-<env>/*`, each keyed by the literal env-var name.
## Aegra (deferred)
`aegra/aegra.json` + `aegra/aegra_entry.py` are the self-hosted-runtime alternative
(Apache-2.0, avoids the LangGraph-Platform Elastic license). Not active on the stock
deployment. To use: place both at the repo root, run `aegra serve` (:2026), and point
`LANGGRAPH_URL` at `:2026`. Aegra gives a Postgres-backed durable store/checkpointer,
which removes the need for `seed_store.sh` and survives restarts (paused HITL
interrupts persist).