From ce9e5fb51bde5db77b51c48edb408ffa34b85e8d Mon Sep 17 00:00:00 2001 From: Adam Moussa <166072409+amoussa1229@users.noreply.github.com> Date: Fri, 26 Jun 2026 15:06:36 -0400 Subject: [PATCH] feat(deploy): EC2 AMI recipe (Packer) + cloud-init/user-data (PR#2) (#7) PR#2 of the AWS migration. deploy/ami/: Packer template (Ubuntu 24.04 arm64, uv+py3.12, nginx, awscli v2, CW agent; no swapfile), provisioning-only user-data (userDataCausesReplacement rationale), systemd unit + nginx + CW templates. Incorporates T5 /sh-security-review fixes: langgraph binds 127.0.0.1 (not 0.0.0.0); nginx is the sole ingress proxying only /dashboard/api/ + /webhooks/; ExecStartPre runs fetch-config as root (+) and passes the env arg; the app runs as the unprivileged openswe user reading an openswe-owned 0600 .env. packer validate clean. --- deploy/ami/README.md | 167 ++++++++++++++++++ deploy/ami/open-swe-base.pkr.hcl | 133 ++++++++++++++ deploy/ami/scripts/provision.sh | 105 +++++++++++ .../templates/amazon-cloudwatch-agent.json | 55 ++++++ deploy/ami/templates/open-swe.nginx.conf | 56 ++++++ deploy/ami/templates/open-swe.service | 54 ++++++ deploy/ami/user-data.sh | 128 ++++++++++++++ 7 files changed, 698 insertions(+) create mode 100644 deploy/ami/README.md create mode 100644 deploy/ami/open-swe-base.pkr.hcl create mode 100755 deploy/ami/scripts/provision.sh create mode 100644 deploy/ami/templates/amazon-cloudwatch-agent.json create mode 100644 deploy/ami/templates/open-swe.nginx.conf create mode 100644 deploy/ami/templates/open-swe.service create mode 100755 deploy/ami/user-data.sh diff --git a/deploy/ami/README.md b/deploy/ami/README.md new file mode 100644 index 00000000..e328533b --- /dev/null +++ b/deploy/ami/README.md @@ -0,0 +1,167 @@ +# Open SWE base AMI (T8) + +Packer recipe + first-boot user-data for the single EC2 instance per env +(`open-swe-dev` / `open-swe-prod`) in the Open SWE → AWS migration. Builds an +**ARM64 (Graviton) Ubuntu 24.04 LTS** base AMI and provisions the box on first +boot with the stock `langgraph dev` runtime, nginx, and the CloudWatch agent. + +The architecture is locked in the repo `TODO.md` ("Architecture (locked)"): ONE +EC2 ARM64 (~t4g.large) instance per env, `seahaven-vpc` **private subnet + NAT**, +inbound **only from the ALB SG**. Runtime is **stock `langgraph dev`** (in-memory +store, `--no-reload`) + nginx + systemd. The box has **no git auth** — it pulls +its deploy artifact from S3 via the instance role. The SPA build runs in GitHub +Actions (T7), **not** on the box, so the old 8 GB-swapfile OOM hack is gone. + +## Files + +| Path | Purpose | +|---|---| +| `open-swe-base.pkr.hcl` | Packer template (HCL2). Latest Canonical 24.04 arm64 source → base AMI. | +| `scripts/provision.sh` | Packer provisioner. System packages, uv+py3.12, node+bun, service user, stages templates. | +| `user-data.sh` | First-boot provisioning (S3 artifact pull, render templates, CW agent, start services). | +| `templates/open-swe.service` | systemd unit TEMPLATE (`@@tokens@@` rendered at boot). | +| `templates/open-swe.nginx.conf` | nginx site TEMPLATE (dashboard SPA + scoped `/dashboard/api/` proxy). | +| `templates/amazon-cloudwatch-agent.json` | CW agent config TEMPLATE — **30-day log retention**. | + +`deploy/seahaven/fetch-config.sh` and `deploy/seahaven/seed_store.sh` are owned by +the parallel T10 work and ship **inside the app artifact**; this AMI wires them in +but does not author them (see "Integration contract" below). + +## Build the AMI + +```bash +cd deploy/ami +packer init . +packer fmt -check . +packer validate -var aws_region=us-east-1 open-swe-base.pkr.hcl +packer build open-swe-base.pkr.hcl +``` + +Builds in account **328440206208 / us-east-1**. Source = latest Canonical Ubuntu +24.04 (Noble) **arm64** AMI (`source_ami_filter`, owner `099720109477`). Build host +is `t4g.medium` (ARM64). The output AMI is tagged: + +``` +Name=open-swe-base-arm64 Purpose=open-swe-runtime-base ManagedBy=packer +``` + +Pinned versions live in the template `variable` defaults (`uv_version`, +`python_version`, `node_major`, the CW-agent / awscli URLs) and the +`required_plugins` block (`amazon` 1.3.6) — bump deliberately. + +## AMI → `cdk.context.json` pinning contract + +The CDK stacks in `/infra` (owned by T3/T12) consume the AMI **by id, pinned in the +committed `infra/cdk.context.json`** — they never resolve "latest" at synth time. +This is the EBS/AMI-fix discipline: an uncached `MachineImage.lookup` resolves a new +AMI on every deploy and silently triggers instance replacement. + +Contract (CDK side does the wiring; this is the handshake): + +1. `packer build` prints the new AMI id (and tags it `open-swe-base-arm64`). +2. CDK looks the AMI up with **`cachedInContext: true`** (e.g. + `MachineImage.lookup({ name: "open-swe-base-arm64-*", owners: ["328440206208"], cachedInContext: true })`), + which writes the resolved id into `infra/cdk.context.json`. +3. **`infra/cdk.context.json` is committed.** From then on every synth/deploy uses + the pinned id — no surprise replacement when a newer AMI exists. +4. To adopt a new AMI: `cdk context --reset ` (or edit the pinned + value), commit the change, and review the cdk-diff — the PR will show + "requires replacement", which is the intended, visible signal. + +Record the built AMI id in project memory (`project_open_swe_migration`) per the +"memory updated for AMI id" build criterion. + +## `userDataCausesReplacement` rationale + +`user-data.sh` is **provisioning-only** — it runs once at first boot and never +carries durable runtime config. CDK sets **`userDataCausesReplacement: true`** so +that any change to it is a deliberate, diff-visible instance replacement rather than +a no-op edit that drifts from the running box. Durable runtime config is fetched +**fresh on every service start** by `fetch-config.sh` (ExecStartPre) — changing a +secret or SSM value needs only a `systemctl restart open-swe.service`, not a +replacement. + +## EBS discipline (binding — `feedback_inline_ebs_volumes`) + +**The box holds no durable state of its own:** + +| State | Lives in | On replacement | +|---|---|---| +| secrets / config | Secrets Manager + SSM → tmpfs `.env` | re-fetched at boot | +| app code + SPA | S3 `open-swe--assets` | re-pulled at boot | +| store (team_settings, user_mappings) | reseeded by `seed_store.sh` | re-seeded at boot | +| logs | CloudWatch (30-day) — **not** a CFN resource in the stack | survive replacement | + +→ **No local-only durable state ⇒ no standalone RETAIN volume is needed.** The root +volume is disposable; there is intentionally no inline data `blockDevices` to lose. + +**Even so, snapshot before any replacing deploy.** Per the operational guard, before +merging/deploying any change that REPLACES the instance (`userDataCausesReplacement`, +AMI bump, instance-type change): + +1. Enumerate the instance's volumes and assert **"no local-only durable state"** + (the table above is the checklist). +2. Take an **EBS snapshot of the root volume and WAIT for `state=completed`** before + letting the deploy proceed. Keep it as insurance; delete after a grace period. +3. Confirm the CloudWatch log groups are **not** CFN-managed in the stack so history + survives; re-verify history after the new instance is healthy. + +cdk-diff-on-PR must flag "requires replacement" at review time (T8/T12 acceptance +criterion). This is the enforced version — not just an assertion in the runbook. + +## Integration contract (T10 — `fetch-config.sh` + `seed_store.sh`) + +Both ship in the app artifact under `deploy/seahaven/` and are wired into the unit: + +- **`fetch-config.sh`** (ExecStartPre, runs as `openswe`): reads `/etc/open-swe/boot.env` + (`OPENSWE_ENV`, `AWS_REGION`, `SECRETS_PREFIX=open-swe-`, `SSM_PREFIX=/open-swe-`, + `ENV_FILE=/run/open-swe/.env`), pulls Secrets Manager `open-swe-/*` + SSM + `/open-swe-/*`, and writes: + - `/run/open-swe/.env` (**0600, tmpfs**, secret-bearing app env incl. the multiline + GitHub App PEM) — loaded by langgraph/dotenv via the `${APP_DIR}/.env` symlink. + - `/run/open-swe/seed.env` (**0600, tmpfs**, simple `OPENSWE_*` vars only: + `OPENSWE_DEFAULT_REPO`, `OPENSWE_OWNER_LOGIN`, `OPENSWE_OWNER_EMAIL`, model ids) — + loaded by systemd `EnvironmentFile` so `seed_store.sh` (ExecStartPost) has them. + - It must **fail-fast** (non-zero exit) if any required value is missing, so the + unit never starts half-configured. +- **`seed_store.sh`** (ExecStartPost): existing script, reseeds `team_settings/default` + + `user_mappings/` into the in-memory store after each start. + +## Smoke-boot checklist (after first boot) + +SSM Session Manager onto the instance (no public SSH — private subnet) and verify: + +- [ ] `cloud-init status --wait` → `done`; `/var/log/open-swe-user-data.log` ends with + "user-data done" and shows the S3 pulls + service starts. +- [ ] `systemctl is-active open-swe.service` → `active`. (If it failed, check + `ExecStartPre`/`fetch-config.sh` — fail-fast means missing config = failed unit.) +- [ ] **fetch-config fail-fast works:** `/run/open-swe/.env` exists, owner `openswe`, + mode `0600`, on tmpfs (`findmnt /run/open-swe`); `seed.env` present. +- [ ] `curl -fsS http://127.0.0.1:2024/ok` → `200` (raw LangGraph health). +- [ ] `systemctl is-active nginx` → `active`; `curl -fsS http://127.0.0.1/healthz` → + `200`; `curl -s http://127.0.0.1/threads` returns the SPA shell, **not** JSON + (proves the agent API is not proxied — the security boundary holds). +- [ ] `seed_store: done` in the journal / app.log (store reseeded). +- [ ] CloudWatch: log groups `/open-swe//{app,user-data,nginx-access,nginx-error}` + exist with **30-day** retention and are receiving events. +- [ ] **No swapfile** (`swapon --show` empty) — the on-box SPA build is gone. +- [ ] From the ALB only: dashboard host serves the SPA; `hooks` host reaches + `/webhooks/*` on :2024 and nothing else (raw API paths hit the ALB default, not + the box). + +## Assumptions + +- **Artifact bucket** `open-swe--assets` (T7), with objects + `${ARTIFACT_PREFIX}/app.tar.gz` (Python app incl. `deploy/seahaven/` and a prebuilt + arm64 `.venv`) and `${ARTIFACT_PREFIX}/spa.tar.gz` (built SPA → `/var/www/open-swe`). + `ARTIFACT_PREFIX` defaults to `releases/latest`; CDK renders the concrete value. +- **Instance role** (defined in `/infra`, least-privilege per T4/T12) grants: + `s3:GetObject` on `open-swe--assets/*`; `secretsmanager:GetSecretValue` on + `open-swe-/*`; `ssm:GetParameter(s)`/`GetParametersByPath` on `/open-swe-/*`; + `logs:*` for the CW agent log groups + `cloudwatch:PutMetricData`; SSM Session + Manager (`ssm:UpdateInstanceInformation`, `ssmmessages:*`) for shell access. +- **CDK substitutes** the `@@OPENSWE_ENV@@`, `@@ASSETS_BUCKET@@`, `@@SERVER_NAME@@`, + `@@ARTIFACT_PREFIX@@` tokens in `user-data.sh` when rendering the launch template. +- `:2024` binds `0.0.0.0` so the ALB hooks target group can reach `/webhooks/*`; it is + reachable only from the ALB SG (private subnet, SG-scoped inbound). The raw API is + never internet-exposed — the ALB hooks rule is path-scoped to `/webhooks/*`. diff --git a/deploy/ami/open-swe-base.pkr.hcl b/deploy/ami/open-swe-base.pkr.hcl new file mode 100644 index 00000000..d5c30737 --- /dev/null +++ b/deploy/ami/open-swe-base.pkr.hcl @@ -0,0 +1,133 @@ +# Open SWE base AMI — ARM64 (Graviton) Ubuntu 24.04 LTS. +# +# Builds the immutable base image for the single EC2 instance per env +# (open-swe-dev / open-swe-prod) in seahaven-vpc. The image bakes the runtime +# (uv + Python 3.12, nginx, awscli v2, CloudWatch agent) and the service-user / +# systemd / nginx TEMPLATES. It bakes NO secrets and NO env-specific values — +# those are materialized at first boot by user-data + deploy/seahaven/fetch-config.sh +# (Secrets Manager + SSM -> root-only tmpfs .env, fail-fast). +# +# Build: packer init . && packer build open-swe-base.pkr.hcl +# The resulting AMI id is pinned in infra/cdk.context.json (CDK cachedInContext:true); +# see README.md "AMI -> cdk.context.json pinning contract". + +packer { + required_version = ">= 1.11.0, < 2.0.0" + required_plugins { + amazon = { + source = "github.com/hashicorp/amazon" + version = "1.3.6" + } + } +} + +variable "aws_region" { + type = string + default = "us-east-1" +} + +variable "instance_type" { + type = string + default = "t4g.medium" # ARM64 (Graviton) build host; runtime instances are ~t4g.large +} + +variable "ami_name_prefix" { + type = string + default = "open-swe-base-arm64" +} + +# Versions baked into the image. Pin and bump deliberately. +variable "python_version" { + type = string + default = "3.12" +} + +variable "node_major" { + type = string + default = "24" +} + +variable "uv_version" { + type = string + default = "0.11.24" +} + +variable "cloudwatch_agent_deb_url" { + type = string + default = "https://amazoncloudwatch-agent.s3.amazonaws.com/ubuntu/arm64/latest/amazon-cloudwatch-agent.deb" +} + +variable "awscli_zip_url" { + type = string + default = "https://awscli.amazonaws.com/awscli-exe-linux-aarch64.zip" +} + +locals { + timestamp = formatdate("YYYYMMDD-hhmmss", timestamp()) +} + +# Latest Canonical Ubuntu 24.04 (Noble) arm64 server image. +source "amazon-ebs" "open-swe" { + region = var.aws_region + instance_type = var.instance_type + ssh_username = "ubuntu" + + ami_name = "${var.ami_name_prefix}-${local.timestamp}" + ami_description = "Open SWE base — Ubuntu 24.04 arm64 + uv/py3.12 + nginx + CW agent (templates only, no secrets)" + + source_ami_filter { + filters = { + name = "ubuntu/images/hvm-ssd*/ubuntu-noble-24.04-arm64-server-*" + architecture = "arm64" + root-device-type = "ebs" + virtualization-type = "hvm" + } + owners = ["099720109477"] # Canonical + most_recent = true + } + + # IMDSv2 required on the build host. + metadata_options { + http_endpoint = "enabled" + http_tokens = "required" + http_put_response_hop_limit = 1 + } + + # gp3 root, encrypted. Runtime root size is set by CDK; this is just the build host. + launch_block_device_mappings { + device_name = "/dev/sda1" + volume_size = 20 + volume_type = "gp3" + encrypted = true + delete_on_termination = true + } + + tags = { + Name = "open-swe-base-arm64" + Purpose = "open-swe-runtime-base" + ManagedBy = "packer" + } +} + +build { + name = "open-swe-base" + sources = ["source.amazon-ebs.open-swe"] + + # Stage the boot-time templates and helper scripts into the image. + provisioner "file" { + source = "${path.root}/templates/" + destination = "/tmp/open-swe-templates/" + } + + provisioner "shell" { + environment_vars = [ + "PYTHON_VERSION=${var.python_version}", + "NODE_MAJOR=${var.node_major}", + "UV_VERSION=${var.uv_version}", + "CLOUDWATCH_AGENT_DEB_URL=${var.cloudwatch_agent_deb_url}", + "AWSCLI_ZIP_URL=${var.awscli_zip_url}", + ] + execute_command = "chmod +x {{ .Path }}; sudo -E bash '{{ .Path }}'" + script = "${path.root}/scripts/provision.sh" + } +} diff --git a/deploy/ami/scripts/provision.sh b/deploy/ami/scripts/provision.sh new file mode 100755 index 00000000..b64a2043 --- /dev/null +++ b/deploy/ami/scripts/provision.sh @@ -0,0 +1,105 @@ +#!/usr/bin/env bash +# Packer provisioner for the Open SWE base AMI (ARM64 Ubuntu 24.04). +# +# Bakes the runtime + boot-time templates ONLY. No secrets, no env-specific +# values. Everything env-specific is materialized at first boot by user-data.sh +# + deploy/seahaven/fetch-config.sh. +set -euo pipefail + +PYTHON_VERSION="${PYTHON_VERSION:-3.12}" +NODE_MAJOR="${NODE_MAJOR:-24}" +UV_VERSION="${UV_VERSION:-0.11.24}" +CLOUDWATCH_AGENT_DEB_URL="${CLOUDWATCH_AGENT_DEB_URL:?}" +AWSCLI_ZIP_URL="${AWSCLI_ZIP_URL:?}" + +# Layout (must match user-data.sh and the templates). +SERVICE_USER="openswe" +APP_DIR="/opt/open-swe/app" +SERVICE_HOME="/opt/open-swe" +WWW_ROOT="/var/www/open-swe" +TEMPLATE_DIR="/opt/open-swe/templates" +LOG_DIR="/var/log/open-swe" +UV_BIN="/usr/local/bin/uv" + +export DEBIAN_FRONTEND=noninteractive + +echo "==> apt base packages" +apt-get update -y +apt-get upgrade -y +apt-get install -y --no-install-recommends \ + nginx jq curl unzip ca-certificates gnupg lsb-release \ + build-essential pkg-config git acl + +echo "==> awscli v2 (aarch64)" +tmp="$(mktemp -d)" +curl -fsSL "$AWSCLI_ZIP_URL" -o "$tmp/awscliv2.zip" +unzip -q "$tmp/awscliv2.zip" -d "$tmp" +"$tmp/aws/install" --update +rm -rf "$tmp" +aws --version + +echo "==> CloudWatch agent (arm64)" +tmp="$(mktemp -d)" +curl -fsSL "$CLOUDWATCH_AGENT_DEB_URL" -o "$tmp/amazon-cloudwatch-agent.deb" +dpkg -i -E "$tmp/amazon-cloudwatch-agent.deb" +rm -rf "$tmp" +# Do NOT enable/start the agent during the build; user-data fetches its config +# (with env-specific log-group names + 30-day retention) and starts it at boot. +systemctl disable amazon-cloudwatch-agent.service || true + +echo "==> uv ${UV_VERSION} + Python ${PYTHON_VERSION} (system-wide)" +export UV_INSTALL_DIR=/usr/local/bin +curl -fsSL "https://astral.sh/uv/${UV_VERSION}/install.sh" | env UV_NO_MODIFY_PATH=1 sh +"$UV_BIN" --version +# Pre-install the interpreter so the box never reaches out at boot to build a venv. +UV_PYTHON_INSTALL_DIR=/opt/uv/python "$UV_BIN" python install "$PYTHON_VERSION" + +echo "==> node ${NODE_MAJOR} + bun (build-time UI tooling only; the SPA is built in CI)" +curl -fsSL "https://deb.nodesource.com/setup_${NODE_MAJOR}.x" | bash - +apt-get install -y --no-install-recommends nodejs +node --version +# bun installed system-wide; used only if any UI tooling must run on-box. The +# production SPA build runs in GitHub Actions -> S3 (no on-box build, no swapfile). +export BUN_INSTALL=/usr/local +curl -fsSL https://bun.sh/install | bash +/usr/local/bin/bun --version || true + +echo "==> non-login service user '${SERVICE_USER}'" +if ! id "$SERVICE_USER" >/dev/null 2>&1; then + useradd --system --create-home --home-dir "$SERVICE_HOME" \ + --shell /usr/sbin/nologin "$SERVICE_USER" +fi + +echo "==> directories" +install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0755 "$SERVICE_HOME" "$APP_DIR" +install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0755 "$WWW_ROOT" +install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0750 "$LOG_DIR" +install -d -o root -g root -m 0755 "$TEMPLATE_DIR" + +echo "==> stage boot-time templates into the image" +cp /tmp/open-swe-templates/* "$TEMPLATE_DIR/" +chown root:root "$TEMPLATE_DIR"/* +chmod 0644 "$TEMPLATE_DIR"/* +rm -rf /tmp/open-swe-templates + +echo "==> tmpfs for the runtime .env (root/owner-only, noexec/nosuid/nodev)" +# /run is already tmpfs on Ubuntu; this is an explicit, deliberately-small mount +# scoped to the service user so the materialized .env never touches disk. +if ! grep -q '/run/open-swe' /etc/fstab; then + cat >>/etc/fstab < disable nginx default site (open-swe site is installed at boot)" +rm -f /etc/nginx/sites-enabled/default +systemctl enable nginx + +echo "==> harden: no password auth, IMDSv2 already enforced by launch template" +# (sshd is not exposed publicly — instance is in a private subnet, SG inbound = ALB only.) + +echo "==> clean apt caches" +apt-get clean +rm -rf /var/lib/apt/lists/* + +echo "==> provision complete" diff --git a/deploy/ami/templates/amazon-cloudwatch-agent.json b/deploy/ami/templates/amazon-cloudwatch-agent.json new file mode 100644 index 00000000..079450cc --- /dev/null +++ b/deploy/ami/templates/amazon-cloudwatch-agent.json @@ -0,0 +1,55 @@ +{ + "agent": { + "metrics_collection_interval": 60, + "run_as_user": "root" + }, + "metrics": { + "namespace": "open-swe/@@OPENSWE_ENV@@", + "append_dimensions": { + "InstanceId": "${aws:InstanceId}" + }, + "metrics_collected": { + "mem": { "measurement": ["mem_used_percent"] }, + "disk": { + "measurement": ["used_percent"], + "resources": ["/"] + } + } + }, + "logs": { + "logs_collected": { + "files": { + "collect_list": [ + { + "file_path": "/var/log/open-swe/app.log", + "log_group_name": "/open-swe/@@OPENSWE_ENV@@/app", + "log_stream_name": "{instance_id}", + "retention_in_days": 30, + "timezone": "UTC" + }, + { + "file_path": "/var/log/open-swe-user-data.log", + "log_group_name": "/open-swe/@@OPENSWE_ENV@@/user-data", + "log_stream_name": "{instance_id}", + "retention_in_days": 30, + "timezone": "UTC" + }, + { + "file_path": "/var/log/nginx/access.log", + "log_group_name": "/open-swe/@@OPENSWE_ENV@@/nginx-access", + "log_stream_name": "{instance_id}", + "retention_in_days": 30, + "timezone": "UTC" + }, + { + "file_path": "/var/log/nginx/error.log", + "log_group_name": "/open-swe/@@OPENSWE_ENV@@/nginx-error", + "log_stream_name": "{instance_id}", + "retention_in_days": 30, + "timezone": "UTC" + } + ] + } + } + } +} diff --git a/deploy/ami/templates/open-swe.nginx.conf b/deploy/ami/templates/open-swe.nginx.conf new file mode 100644 index 00000000..ba219646 --- /dev/null +++ b/deploy/ami/templates/open-swe.nginx.conf @@ -0,0 +1,56 @@ +# Open SWE dashboard frontend (TanStack Start SPA) + scoped API proxy. +# TEMPLATE: tokens (@@...@@) are rendered at first boot by user-data.sh. +# +# nginx is the SOLE ingress and security boundary (T5 OSWE-IAC-03): the backend +# binds 127.0.0.1:2024 and is NOT network-reachable. nginx proxies exactly two +# prefixes to it — /dashboard/api/* and /webhooks/* — and nothing else. The +# unauthenticated LangGraph agent API (/threads, /runs, /assistants, /store) is +# NEVER proxied; those paths return the SPA shell. +# +# Both ALB target groups (dashboard host + hooks host) point at this nginx :80, +# not at :2024 directly, so there is no path to the raw control plane even from +# inside the SG. Webhook signature verification still happens in the app (the raw +# body + GitHub/Slack/Linear signature headers are passed through unmodified). +server { + listen 80 default_server; + listen [::]:80 default_server; + server_name @@SERVER_NAME@@; + + root @@WWW_ROOT@@; + index _shell.html; + + # ALB target-group health check (dashboard TG). + location = /healthz { default_type text/plain; return 200 "ok\n"; } + + # Dashboard API + OAuth callback -> backend webapp. + location /dashboard/api/ { + proxy_pass http://@@BACKEND_ADDR@@; + proxy_http_version 1.1; + proxy_set_header Host $host; + proxy_set_header X-Real-IP $remote_addr; + proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; + proxy_set_header X-Forwarded-Proto https; + proxy_set_header Upgrade $http_upgrade; + proxy_set_header Connection "upgrade"; + proxy_read_timeout 300s; + } + + # Inbound webhooks (GitHub/Slack/Linear) -> backend webapp. Routed through + # nginx so :2024 stays loopback-only (T5 OSWE-IAC-03). The raw request body + + # signature headers pass through unmodified for in-app signature verification. + location /webhooks/ { + proxy_pass http://@@BACKEND_ADDR@@; + proxy_http_version 1.1; + proxy_set_header Host $host; + proxy_set_header X-Real-IP $remote_addr; + proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; + proxy_set_header X-Forwarded-Proto https; + proxy_request_buffering off; + proxy_read_timeout 300s; + } + + # Static assets + SPA shell fallback (client-side routing). + location / { + try_files $uri $uri/ /_shell.html; + } +} diff --git a/deploy/ami/templates/open-swe.service b/deploy/ami/templates/open-swe.service new file mode 100644 index 00000000..f3fd8b64 --- /dev/null +++ b/deploy/ami/templates/open-swe.service @@ -0,0 +1,54 @@ +# Open SWE — stock LangGraph dev server (3+ graphs + FastAPI webapp, :2024). +# TEMPLATE: tokens (@@...@@) are rendered at first boot by user-data.sh. +# In-memory runtime (--no-reload) + ExecStartPost reseed; no Aegra/Postgres. +# +# Boot contract: +# ExecStartPre = fetch-config.sh -> runs as root (`+`) ONLY to materialize +# the SERVICE-USER-owned tmpfs .env (@@ENV_FILE@@) from Secrets +# Manager + SSM and chown it to @@SERVICE_USER@@, fail-fast (the +# unit does NOT start if config can't be fetched). +# ExecStart = langgraph dev (as @@SERVICE_USER@@, bound to 127.0.0.1 — nginx +# is the sole ingress; never binds 0.0.0.0). +# ExecStartPost = seed_store.sh -> reseeds team_settings + user_mappings +# that the in-memory store loses on every restart. +[Unit] +Description=Open SWE stock LangGraph dev server (graphs + webapp, :@@PORT@@) +After=network-online.target +Wants=network-online.target +RequiresMountsFor=/run/open-swe + +[Service] +Type=simple +User=@@SERVICE_USER@@ +Group=@@SERVICE_USER@@ +WorkingDirectory=@@APP_DIR@@ + +# No EnvironmentFile: the secret-bearing app .env (@@ENV_FILE@@) is loaded by +# langgraph/dotenv (so the multiline GitHub App PEM never hits systemd's env +# parser), and seed_store.sh reads the same .env directly (without sourcing it). +# +# ExecStartPre runs as root (`+`) so it can chown the tmpfs .env to the service +# user; the env arg (@@OPENSWE_ENV@@) selects the SSM/Secrets prefix (T5 BOOT-01). +ExecStartPre=+@@FETCH_CONFIG@@ @@OPENSWE_ENV@@ +# Bind 127.0.0.1 only — nginx proxies dashboard + webhooks; :@@PORT@@ is never +# directly network-reachable (T5 OSWE-IAC-03). +ExecStart=@@VENV@@/bin/langgraph dev --host 127.0.0.1 --port @@PORT@@ --no-browser --no-reload +ExecStartPost=@@SEED_STORE@@ @@OPENSWE_ENV@@ + +# App logs to a file CloudWatch collects (30-day retention set in the CW config). +StandardOutput=append:/var/log/open-swe/app.log +StandardError=append:/var/log/open-swe/app.log + +Restart=on-failure +RestartSec=5 +TimeoutStartSec=180 + +# Hardening — the box holds no durable state of its own. +NoNewPrivileges=true +ProtectSystem=full +ProtectHome=true +PrivateTmp=true +ReadWritePaths=/var/log/open-swe /var/www/open-swe /run/open-swe @@APP_DIR@@ + +[Install] +WantedBy=multi-user.target diff --git a/deploy/ami/user-data.sh b/deploy/ami/user-data.sh new file mode 100755 index 00000000..b204be19 --- /dev/null +++ b/deploy/ami/user-data.sh @@ -0,0 +1,128 @@ +#!/usr/bin/env bash +# Open SWE EC2 user-data — PROVISIONING-ONLY (runs once, at first boot). +# +# This is the rationale for `userDataCausesReplacement: true` in CDK: user-data +# does FIRST-BOOT provisioning, never durable runtime config. Editing it is a +# deliberate instance replacement. Durable runtime config is fetched fresh on +# every service start by deploy/seahaven/fetch-config.sh (ExecStartPre). +# +# The box holds NO durable state of its own: +# - secrets/config -> Secrets Manager + SSM, materialized to a tmpfs .env at boot +# - app artifact -> pulled from S3 (open-swe--assets) via the instance role +# - store state -> reseeded by seed_store.sh (ExecStartPost) on every start +# => there is no RETAIN volume to protect; replacement is tolerated. The EBS +# discipline (snapshot root + wait state=completed BEFORE any replacing deploy) +# is the safety net, not durable on-box state. See README "EBS discipline". +# +# Tokens (@@...@@) are substituted by CDK when it renders this script into the +# launch template. region is read from IMDSv2 as a fallback. + +set -euo pipefail +exec > >(tee -a /var/log/open-swe-user-data.log) 2>&1 +echo "==> open-swe user-data start $(date -u +%FT%TZ)" + +# --- CDK-rendered values ----------------------------------------------------- +OPENSWE_ENV="@@OPENSWE_ENV@@" # dev | prod +ASSETS_BUCKET="@@ASSETS_BUCKET@@" # open-swe--assets +SERVER_NAME="@@SERVER_NAME@@" # openswe[-dev].seahaven.com +ARTIFACT_PREFIX="@@ARTIFACT_PREFIX@@" # e.g. releases/latest + +# --- fixed layout (must match provision.sh + templates) ---------------------- +SERVICE_USER="openswe" +APP_DIR="/opt/open-swe/app" +VENV="${APP_DIR}/.venv" +WWW_ROOT="/var/www/open-swe" +TEMPLATE_DIR="/opt/open-swe/templates" +ENV_FILE="/run/open-swe/.env" +PORT="2024" +FETCH_CONFIG="${APP_DIR}/deploy/seahaven/fetch-config.sh" +SEED_STORE="${APP_DIR}/deploy/seahaven/seed_store.sh" + +# region from IMDSv2 +TOKEN="$(curl -fsS -X PUT "http://169.254.169.254/latest/api/token" \ + -H "X-aws-ec2-metadata-token-ttl-seconds: 300" || true)" +AWS_REGION="$(curl -fsS -H "X-aws-ec2-metadata-token: ${TOKEN}" \ + http://169.254.169.254/latest/meta-data/placement/region || echo us-east-1)" +export AWS_DEFAULT_REGION="$AWS_REGION" +echo "env=${OPENSWE_ENV} region=${AWS_REGION} bucket=${ASSETS_BUCKET} host=${SERVER_NAME}" + +# --- boot.env: non-secret pointers fetch-config.sh reads --------------------- +install -d -o root -g root -m 0755 /etc/open-swe +cat >/etc/open-swe/boot.env < fetch app artifact from s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/" +tmp="$(mktemp -d)" +aws s3 cp "s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/app.tar.gz" "${tmp}/app.tar.gz" +aws s3 cp "s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/spa.tar.gz" "${tmp}/spa.tar.gz" + +install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0755 "$APP_DIR" "$WWW_ROOT" +tar -xzf "${tmp}/app.tar.gz" -C "$APP_DIR" +tar -xzf "${tmp}/spa.tar.gz" -C "$WWW_ROOT" +chown -R "$SERVICE_USER":"$SERVICE_USER" "$APP_DIR" "$WWW_ROOT" +rm -rf "$tmp" + +# The app reads ./.env from WorkingDirectory (langgraph.json "env": ".env"). +# Point it at the tmpfs file fetch-config.sh materializes. +ln -sfn "$ENV_FILE" "${APP_DIR}/.env" + +# --- render + install the systemd unit --------------------------------------- +echo "==> install systemd unit" +sed \ + -e "s|@@SERVICE_USER@@|${SERVICE_USER}|g" \ + -e "s|@@APP_DIR@@|${APP_DIR}|g" \ + -e "s|@@VENV@@|${VENV}|g" \ + -e "s|@@PORT@@|${PORT}|g" \ + -e "s|@@ENV_FILE@@|${ENV_FILE}|g" \ + -e "s|@@OPENSWE_ENV@@|${OPENSWE_ENV}|g" \ + -e "s|@@FETCH_CONFIG@@|${FETCH_CONFIG}|g" \ + -e "s|@@SEED_STORE@@|${SEED_STORE}|g" \ + "${TEMPLATE_DIR}/open-swe.service" >/etc/systemd/system/open-swe.service +systemctl daemon-reload + +# --- render + install the nginx site ----------------------------------------- +echo "==> install nginx site" +sed \ + -e "s|@@SERVER_NAME@@|${SERVER_NAME}|g" \ + -e "s|@@WWW_ROOT@@|${WWW_ROOT}|g" \ + -e "s|@@BACKEND_ADDR@@|127.0.0.1:${PORT}|g" \ + "${TEMPLATE_DIR}/open-swe.nginx.conf" >/etc/nginx/sites-available/open-swe +ln -sfn /etc/nginx/sites-available/open-swe /etc/nginx/sites-enabled/open-swe +rm -f /etc/nginx/sites-enabled/default +nginx -t + +# --- CloudWatch agent: 30-day log retention ---------------------------------- +echo "==> configure CloudWatch agent (30-day retention)" +sed -e "s|@@OPENSWE_ENV@@|${OPENSWE_ENV}|g" \ + "${TEMPLATE_DIR}/amazon-cloudwatch-agent.json" \ + >/opt/aws/amazon-cloudwatch-agent/etc/open-swe-cw.json +/opt/aws/amazon-cloudwatch-agent/bin/amazon-cloudwatch-agent-ctl \ + -a fetch-config -m ec2 -s \ + -c file:/opt/aws/amazon-cloudwatch-agent/etc/open-swe-cw.json + +# --- start services ---------------------------------------------------------- +# NOTE: intentionally NO swapfile here. The 8 GB-swapfile OOM hack existed only +# for the on-box Nitro SPA build, which now runs in GitHub Actions -> S3. +echo "==> start nginx + open-swe.service" +systemctl enable --now nginx +systemctl reload nginx +# open-swe.service ExecStartPre=fetch-config.sh fails-fast if config is missing, +# so a bad secrets/SSM setup surfaces as a failed unit (not a half-up box). +systemctl enable --now open-swe.service + +echo "==> open-swe user-data done $(date -u +%FT%TZ)"