feat(deploy): EC2 AMI recipe (Packer) + cloud-init/user-data (PR#2) (#7)

PR#2 of the AWS migration. deploy/ami/: Packer template (Ubuntu 24.04 arm64,
uv+py3.12, nginx, awscli v2, CW agent; no swapfile), provisioning-only user-data
(userDataCausesReplacement rationale), systemd unit + nginx + CW templates.

Incorporates T5 /sh-security-review fixes: langgraph binds 127.0.0.1 (not 0.0.0.0);
nginx is the sole ingress proxying only /dashboard/api/ + /webhooks/; ExecStartPre
runs fetch-config as root (+) and passes the env arg; the app runs as the
unprivileged openswe user reading an openswe-owned 0600 .env. packer validate clean.
This commit is contained in:
Adam Moussa 2026-06-26 15:06:36 -04:00 • committed by GitHub
parent 92fd886076
commit ce9e5fb51b
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
7 changed files with 698 additions and 0 deletions

167
deploy/ami/README.md Normal file
View file

@ -0,0 +1,167 @@
# Open SWE base AMI (T8)
Packer recipe + first-boot user-data for the single EC2 instance per env
(`open-swe-dev` / `open-swe-prod`) in the Open SWE → AWS migration. Builds an
**ARM64 (Graviton) Ubuntu 24.04 LTS** base AMI and provisions the box on first
boot with the stock `langgraph dev` runtime, nginx, and the CloudWatch agent.
The architecture is locked in the repo `TODO.md` ("Architecture (locked)"): ONE
EC2 ARM64 (~t4g.large) instance per env, `seahaven-vpc` **private subnet + NAT**,
inbound **only from the ALB SG**. Runtime is **stock `langgraph dev`** (in-memory
store, `--no-reload`) + nginx + systemd. The box has **no git auth** — it pulls
its deploy artifact from S3 via the instance role. The SPA build runs in GitHub
Actions (T7), **not** on the box, so the old 8 GB-swapfile OOM hack is gone.
## Files
| Path | Purpose |
|---|---|
| `open-swe-base.pkr.hcl` | Packer template (HCL2). Latest Canonical 24.04 arm64 source → base AMI. |
| `scripts/provision.sh` | Packer provisioner. System packages, uv+py3.12, node+bun, service user, stages templates. |
| `user-data.sh` | First-boot provisioning (S3 artifact pull, render templates, CW agent, start services). |
| `templates/open-swe.service` | systemd unit TEMPLATE (`@@tokens@@` rendered at boot). |
| `templates/open-swe.nginx.conf` | nginx site TEMPLATE (dashboard SPA + scoped `/dashboard/api/` proxy). |
| `templates/amazon-cloudwatch-agent.json` | CW agent config TEMPLATE — **30-day log retention**. |
`deploy/seahaven/fetch-config.sh` and `deploy/seahaven/seed_store.sh` are owned by
the parallel T10 work and ship **inside the app artifact**; this AMI wires them in
but does not author them (see "Integration contract" below).
## Build the AMI
```bash
cd deploy/ami
packer init .
packer fmt -check .
packer validate -var aws_region=us-east-1 open-swe-base.pkr.hcl
packer build open-swe-base.pkr.hcl
```
Builds in account **328440206208 / us-east-1**. Source = latest Canonical Ubuntu
24.04 (Noble) **arm64** AMI (`source_ami_filter`, owner `099720109477`). Build host
is `t4g.medium` (ARM64). The output AMI is tagged:
```
Name=open-swe-base-arm64 Purpose=open-swe-runtime-base ManagedBy=packer
```
Pinned versions live in the template `variable` defaults (`uv_version`,
`python_version`, `node_major`, the CW-agent / awscli URLs) and the
`required_plugins` block (`amazon` 1.3.6) — bump deliberately.
## AMI → `cdk.context.json` pinning contract
The CDK stacks in `/infra` (owned by T3/T12) consume the AMI **by id, pinned in the
committed `infra/cdk.context.json`** — they never resolve "latest" at synth time.
This is the EBS/AMI-fix discipline: an uncached `MachineImage.lookup` resolves a new
AMI on every deploy and silently triggers instance replacement.
Contract (CDK side does the wiring; this is the handshake):
1. `packer build` prints the new AMI id (and tags it `open-swe-base-arm64`).
2. CDK looks the AMI up with **`cachedInContext: true`** (e.g.
`MachineImage.lookup({ name: "open-swe-base-arm64-*", owners: ["328440206208"], cachedInContext: true })`),
which writes the resolved id into `infra/cdk.context.json`.
3. **`infra/cdk.context.json` is committed.** From then on every synth/deploy uses
the pinned id — no surprise replacement when a newer AMI exists.
4. To adopt a new AMI: `cdk context --reset <ami-lookup-key>` (or edit the pinned
value), commit the change, and review the cdk-diff — the PR will show
"requires replacement", which is the intended, visible signal.
Record the built AMI id in project memory (`project_open_swe_migration`) per the
"memory updated for AMI id" build criterion.
## `userDataCausesReplacement` rationale
`user-data.sh` is **provisioning-only** — it runs once at first boot and never
carries durable runtime config. CDK sets **`userDataCausesReplacement: true`** so
that any change to it is a deliberate, diff-visible instance replacement rather than
a no-op edit that drifts from the running box. Durable runtime config is fetched
**fresh on every service start** by `fetch-config.sh` (ExecStartPre) — changing a
secret or SSM value needs only a `systemctl restart open-swe.service`, not a
replacement.
## EBS discipline (binding — `feedback_inline_ebs_volumes`)
**The box holds no durable state of its own:**
| State | Lives in | On replacement |
|---|---|---|
| secrets / config | Secrets Manager + SSM → tmpfs `.env` | re-fetched at boot |
| app code + SPA | S3 `open-swe-<env>-assets` | re-pulled at boot |
| store (team_settings, user_mappings) | reseeded by `seed_store.sh` | re-seeded at boot |
| logs | CloudWatch (30-day) — **not** a CFN resource in the stack | survive replacement |
→ **No local-only durable state ⇒ no standalone RETAIN volume is needed.** The root
volume is disposable; there is intentionally no inline data `blockDevices` to lose.
**Even so, snapshot before any replacing deploy.** Per the operational guard, before
merging/deploying any change that REPLACES the instance (`userDataCausesReplacement`,
AMI bump, instance-type change):
1. Enumerate the instance's volumes and assert **"no local-only durable state"**
(the table above is the checklist).
2. Take an **EBS snapshot of the root volume and WAIT for `state=completed`** before
letting the deploy proceed. Keep it as insurance; delete after a grace period.
3. Confirm the CloudWatch log groups are **not** CFN-managed in the stack so history
survives; re-verify history after the new instance is healthy.
cdk-diff-on-PR must flag "requires replacement" at review time (T8/T12 acceptance
criterion). This is the enforced version — not just an assertion in the runbook.
## Integration contract (T10 — `fetch-config.sh` + `seed_store.sh`)
Both ship in the app artifact under `deploy/seahaven/` and are wired into the unit:
- **`fetch-config.sh`** (ExecStartPre, runs as `openswe`): reads `/etc/open-swe/boot.env`
(`OPENSWE_ENV`, `AWS_REGION`, `SECRETS_PREFIX=open-swe-<env>`, `SSM_PREFIX=/open-swe-<env>`,
`ENV_FILE=/run/open-swe/.env`), pulls Secrets Manager `open-swe-<env>/*` + SSM
`/open-swe-<env>/*`, and writes:
- `/run/open-swe/.env` (**0600, tmpfs**, secret-bearing app env incl. the multiline
GitHub App PEM) — loaded by langgraph/dotenv via the `${APP_DIR}/.env` symlink.
- `/run/open-swe/seed.env` (**0600, tmpfs**, simple `OPENSWE_*` vars only:
`OPENSWE_DEFAULT_REPO`, `OPENSWE_OWNER_LOGIN`, `OPENSWE_OWNER_EMAIL`, model ids) —
loaded by systemd `EnvironmentFile` so `seed_store.sh` (ExecStartPost) has them.
- It must **fail-fast** (non-zero exit) if any required value is missing, so the
unit never starts half-configured.
- **`seed_store.sh`** (ExecStartPost): existing script, reseeds `team_settings/default`
+ `user_mappings/<login>` into the in-memory store after each start.
## Smoke-boot checklist (after first boot)
SSM Session Manager onto the instance (no public SSH — private subnet) and verify:
- [ ] `cloud-init status --wait` → `done`; `/var/log/open-swe-user-data.log` ends with
"user-data done" and shows the S3 pulls + service starts.
- [ ] `systemctl is-active open-swe.service` → `active`. (If it failed, check
`ExecStartPre`/`fetch-config.sh` — fail-fast means missing config = failed unit.)
- [ ] **fetch-config fail-fast works:** `/run/open-swe/.env` exists, owner `openswe`,
mode `0600`, on tmpfs (`findmnt /run/open-swe`); `seed.env` present.
- [ ] `curl -fsS http://127.0.0.1:2024/ok` → `200` (raw LangGraph health).
- [ ] `systemctl is-active nginx` → `active`; `curl -fsS http://127.0.0.1/healthz` →
`200`; `curl -s http://127.0.0.1/threads` returns the SPA shell, **not** JSON
(proves the agent API is not proxied — the security boundary holds).
- [ ] `seed_store: done` in the journal / app.log (store reseeded).
- [ ] CloudWatch: log groups `/open-swe/<env>/{app,user-data,nginx-access,nginx-error}`
exist with **30-day** retention and are receiving events.
- [ ] **No swapfile** (`swapon --show` empty) — the on-box SPA build is gone.
- [ ] From the ALB only: dashboard host serves the SPA; `hooks` host reaches
`/webhooks/*` on :2024 and nothing else (raw API paths hit the ALB default, not
the box).
## Assumptions
- **Artifact bucket** `open-swe-<env>-assets` (T7), with objects
`${ARTIFACT_PREFIX}/app.tar.gz` (Python app incl. `deploy/seahaven/` and a prebuilt
arm64 `.venv`) and `${ARTIFACT_PREFIX}/spa.tar.gz` (built SPA → `/var/www/open-swe`).
`ARTIFACT_PREFIX` defaults to `releases/latest`; CDK renders the concrete value.
- **Instance role** (defined in `/infra`, least-privilege per T4/T12) grants:
`s3:GetObject` on `open-swe-<env>-assets/*`; `secretsmanager:GetSecretValue` on
`open-swe-<env>/*`; `ssm:GetParameter(s)`/`GetParametersByPath` on `/open-swe-<env>/*`;
`logs:*` for the CW agent log groups + `cloudwatch:PutMetricData`; SSM Session
Manager (`ssm:UpdateInstanceInformation`, `ssmmessages:*`) for shell access.
- **CDK substitutes** the `@@OPENSWE_ENV@@`, `@@ASSETS_BUCKET@@`, `@@SERVER_NAME@@`,
`@@ARTIFACT_PREFIX@@` tokens in `user-data.sh` when rendering the launch template.
- `:2024` binds `0.0.0.0` so the ALB hooks target group can reach `/webhooks/*`; it is
reachable only from the ALB SG (private subnet, SG-scoped inbound). The raw API is
never internet-exposed — the ALB hooks rule is path-scoped to `/webhooks/*`.

View file

@ -0,0 +1,133 @@
# Open SWE base AMI — ARM64 (Graviton) Ubuntu 24.04 LTS.
#
# Builds the immutable base image for the single EC2 instance per env
# (open-swe-dev / open-swe-prod) in seahaven-vpc. The image bakes the runtime
# (uv + Python 3.12, nginx, awscli v2, CloudWatch agent) and the service-user /
# systemd / nginx TEMPLATES. It bakes NO secrets and NO env-specific values —
# those are materialized at first boot by user-data + deploy/seahaven/fetch-config.sh
# (Secrets Manager + SSM -> root-only tmpfs .env, fail-fast).
#
# Build: packer init . && packer build open-swe-base.pkr.hcl
# The resulting AMI id is pinned in infra/cdk.context.json (CDK cachedInContext:true);
# see README.md "AMI -> cdk.context.json pinning contract".
packer {
required_version = ">= 1.11.0, < 2.0.0"
required_plugins {
amazon = {
source = "github.com/hashicorp/amazon"
version = "1.3.6"
}
}
}
variable "aws_region" {
type = string
default = "us-east-1"
}
variable "instance_type" {
type = string
default = "t4g.medium" # ARM64 (Graviton) build host; runtime instances are ~t4g.large
}
variable "ami_name_prefix" {
type = string
default = "open-swe-base-arm64"
}
# Versions baked into the image. Pin and bump deliberately.
variable "python_version" {
type = string
default = "3.12"
}
variable "node_major" {
type = string
default = "24"
}
variable "uv_version" {
type = string
default = "0.11.24"
}
variable "cloudwatch_agent_deb_url" {
type = string
default = "https://amazoncloudwatch-agent.s3.amazonaws.com/ubuntu/arm64/latest/amazon-cloudwatch-agent.deb"
}
variable "awscli_zip_url" {
type = string
default = "https://awscli.amazonaws.com/awscli-exe-linux-aarch64.zip"
}
locals {
timestamp = formatdate("YYYYMMDD-hhmmss", timestamp())
}
# Latest Canonical Ubuntu 24.04 (Noble) arm64 server image.
source "amazon-ebs" "open-swe" {
region = var.aws_region
instance_type = var.instance_type
ssh_username = "ubuntu"
ami_name = "${var.ami_name_prefix}-${local.timestamp}"
ami_description = "Open SWE base — Ubuntu 24.04 arm64 + uv/py3.12 + nginx + CW agent (templates only, no secrets)"
source_ami_filter {
filters = {
name = "ubuntu/images/hvm-ssd*/ubuntu-noble-24.04-arm64-server-*"
architecture = "arm64"
root-device-type = "ebs"
virtualization-type = "hvm"
}
owners = ["099720109477"] # Canonical
most_recent = true
}
# IMDSv2 required on the build host.
metadata_options {
http_endpoint = "enabled"
http_tokens = "required"
http_put_response_hop_limit = 1
}
# gp3 root, encrypted. Runtime root size is set by CDK; this is just the build host.
launch_block_device_mappings {
device_name = "/dev/sda1"
volume_size = 20
volume_type = "gp3"
encrypted = true
delete_on_termination = true
}
tags = {
Name = "open-swe-base-arm64"
Purpose = "open-swe-runtime-base"
ManagedBy = "packer"
}
}
build {
name = "open-swe-base"
sources = ["source.amazon-ebs.open-swe"]
# Stage the boot-time templates and helper scripts into the image.
provisioner "file" {
source = "${path.root}/templates/"
destination = "/tmp/open-swe-templates/"
}
provisioner "shell" {
environment_vars = [
"PYTHON_VERSION=${var.python_version}",
"NODE_MAJOR=${var.node_major}",
"UV_VERSION=${var.uv_version}",
"CLOUDWATCH_AGENT_DEB_URL=${var.cloudwatch_agent_deb_url}",
"AWSCLI_ZIP_URL=${var.awscli_zip_url}",
]
execute_command = "chmod +x {{ .Path }}; sudo -E bash '{{ .Path }}'"
script = "${path.root}/scripts/provision.sh"
}
}

105
deploy/ami/scripts/provision.sh Executable file
View file

@ -0,0 +1,105 @@
#!/usr/bin/env bash
# Packer provisioner for the Open SWE base AMI (ARM64 Ubuntu 24.04).
#
# Bakes the runtime + boot-time templates ONLY. No secrets, no env-specific
# values. Everything env-specific is materialized at first boot by user-data.sh
# + deploy/seahaven/fetch-config.sh.
set -euo pipefail
PYTHON_VERSION="${PYTHON_VERSION:-3.12}"
NODE_MAJOR="${NODE_MAJOR:-24}"
UV_VERSION="${UV_VERSION:-0.11.24}"
CLOUDWATCH_AGENT_DEB_URL="${CLOUDWATCH_AGENT_DEB_URL:?}"
AWSCLI_ZIP_URL="${AWSCLI_ZIP_URL:?}"
# Layout (must match user-data.sh and the templates).
SERVICE_USER="openswe"
APP_DIR="/opt/open-swe/app"
SERVICE_HOME="/opt/open-swe"
WWW_ROOT="/var/www/open-swe"
TEMPLATE_DIR="/opt/open-swe/templates"
LOG_DIR="/var/log/open-swe"
UV_BIN="/usr/local/bin/uv"
export DEBIAN_FRONTEND=noninteractive
echo "==> apt base packages"
apt-get update -y
apt-get upgrade -y
apt-get install -y --no-install-recommends \
nginx jq curl unzip ca-certificates gnupg lsb-release \
build-essential pkg-config git acl
echo "==> awscli v2 (aarch64)"
tmp="$(mktemp -d)"
curl -fsSL "$AWSCLI_ZIP_URL" -o "$tmp/awscliv2.zip"
unzip -q "$tmp/awscliv2.zip" -d "$tmp"
"$tmp/aws/install" --update
rm -rf "$tmp"
aws --version
echo "==> CloudWatch agent (arm64)"
tmp="$(mktemp -d)"
curl -fsSL "$CLOUDWATCH_AGENT_DEB_URL" -o "$tmp/amazon-cloudwatch-agent.deb"
dpkg -i -E "$tmp/amazon-cloudwatch-agent.deb"
rm -rf "$tmp"
# Do NOT enable/start the agent during the build; user-data fetches its config
# (with env-specific log-group names + 30-day retention) and starts it at boot.
systemctl disable amazon-cloudwatch-agent.service || true
echo "==> uv ${UV_VERSION} + Python ${PYTHON_VERSION} (system-wide)"
export UV_INSTALL_DIR=/usr/local/bin
curl -fsSL "https://astral.sh/uv/${UV_VERSION}/install.sh" | env UV_NO_MODIFY_PATH=1 sh
"$UV_BIN" --version
# Pre-install the interpreter so the box never reaches out at boot to build a venv.
UV_PYTHON_INSTALL_DIR=/opt/uv/python "$UV_BIN" python install "$PYTHON_VERSION"
echo "==> node ${NODE_MAJOR} + bun (build-time UI tooling only; the SPA is built in CI)"
curl -fsSL "https://deb.nodesource.com/setup_${NODE_MAJOR}.x" | bash -
apt-get install -y --no-install-recommends nodejs
node --version
# bun installed system-wide; used only if any UI tooling must run on-box. The
# production SPA build runs in GitHub Actions -> S3 (no on-box build, no swapfile).
export BUN_INSTALL=/usr/local
curl -fsSL https://bun.sh/install | bash
/usr/local/bin/bun --version || true
echo "==> non-login service user '${SERVICE_USER}'"
if ! id "$SERVICE_USER" >/dev/null 2>&1; then
useradd --system --create-home --home-dir "$SERVICE_HOME" \
--shell /usr/sbin/nologin "$SERVICE_USER"
fi
echo "==> directories"
install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0755 "$SERVICE_HOME" "$APP_DIR"
install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0755 "$WWW_ROOT"
install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0750 "$LOG_DIR"
install -d -o root -g root -m 0755 "$TEMPLATE_DIR"
echo "==> stage boot-time templates into the image"
cp /tmp/open-swe-templates/* "$TEMPLATE_DIR/"
chown root:root "$TEMPLATE_DIR"/*
chmod 0644 "$TEMPLATE_DIR"/*
rm -rf /tmp/open-swe-templates
echo "==> tmpfs for the runtime .env (root/owner-only, noexec/nosuid/nodev)"
# /run is already tmpfs on Ubuntu; this is an explicit, deliberately-small mount
# scoped to the service user so the materialized .env never touches disk.
if ! grep -q '/run/open-swe' /etc/fstab; then
cat >>/etc/fstab <<EOF
tmpfs /run/open-swe tmpfs rw,nosuid,nodev,noexec,mode=0700,uid=${SERVICE_USER},gid=${SERVICE_USER},size=8m 0 0
EOF
fi
echo "==> disable nginx default site (open-swe site is installed at boot)"
rm -f /etc/nginx/sites-enabled/default
systemctl enable nginx
echo "==> harden: no password auth, IMDSv2 already enforced by launch template"
# (sshd is not exposed publicly — instance is in a private subnet, SG inbound = ALB only.)
echo "==> clean apt caches"
apt-get clean
rm -rf /var/lib/apt/lists/*
echo "==> provision complete"

View file

@ -0,0 +1,55 @@
{
"agent": {
"metrics_collection_interval": 60,
"run_as_user": "root"
},
"metrics": {
"namespace": "open-swe/@@OPENSWE_ENV@@",
"append_dimensions": {
"InstanceId": "${aws:InstanceId}"
},
"metrics_collected": {
"mem": { "measurement": ["mem_used_percent"] },
"disk": {
"measurement": ["used_percent"],
"resources": ["/"]
}
}
},
"logs": {
"logs_collected": {
"files": {
"collect_list": [
{
"file_path": "/var/log/open-swe/app.log",
"log_group_name": "/open-swe/@@OPENSWE_ENV@@/app",
"log_stream_name": "{instance_id}",
"retention_in_days": 30,
"timezone": "UTC"
},
{
"file_path": "/var/log/open-swe-user-data.log",
"log_group_name": "/open-swe/@@OPENSWE_ENV@@/user-data",
"log_stream_name": "{instance_id}",
"retention_in_days": 30,
"timezone": "UTC"
},
{
"file_path": "/var/log/nginx/access.log",
"log_group_name": "/open-swe/@@OPENSWE_ENV@@/nginx-access",
"log_stream_name": "{instance_id}",
"retention_in_days": 30,
"timezone": "UTC"
},
{
"file_path": "/var/log/nginx/error.log",
"log_group_name": "/open-swe/@@OPENSWE_ENV@@/nginx-error",
"log_stream_name": "{instance_id}",
"retention_in_days": 30,
"timezone": "UTC"
}
]
}
}
}
}

View file

@ -0,0 +1,56 @@
# Open SWE dashboard frontend (TanStack Start SPA) + scoped API proxy.
# TEMPLATE: tokens (@@...@@) are rendered at first boot by user-data.sh.
#
# nginx is the SOLE ingress and security boundary (T5 OSWE-IAC-03): the backend
# binds 127.0.0.1:2024 and is NOT network-reachable. nginx proxies exactly two
# prefixes to it — /dashboard/api/* and /webhooks/* — and nothing else. The
# unauthenticated LangGraph agent API (/threads, /runs, /assistants, /store) is
# NEVER proxied; those paths return the SPA shell.
#
# Both ALB target groups (dashboard host + hooks host) point at this nginx :80,
# not at :2024 directly, so there is no path to the raw control plane even from
# inside the SG. Webhook signature verification still happens in the app (the raw
# body + GitHub/Slack/Linear signature headers are passed through unmodified).
server {
listen 80 default_server;
listen [::]:80 default_server;
server_name @@SERVER_NAME@@;
root @@WWW_ROOT@@;
index _shell.html;
# ALB target-group health check (dashboard TG).
location = /healthz { default_type text/plain; return 200 "ok\n"; }
# Dashboard API + OAuth callback -> backend webapp.
location /dashboard/api/ {
proxy_pass http://@@BACKEND_ADDR@@;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto https;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_read_timeout 300s;
}
# Inbound webhooks (GitHub/Slack/Linear) -> backend webapp. Routed through
# nginx so :2024 stays loopback-only (T5 OSWE-IAC-03). The raw request body +
# signature headers pass through unmodified for in-app signature verification.
location /webhooks/ {
proxy_pass http://@@BACKEND_ADDR@@;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto https;
proxy_request_buffering off;
proxy_read_timeout 300s;
}
# Static assets + SPA shell fallback (client-side routing).
location / {
try_files $uri $uri/ /_shell.html;
}
}

View file

@ -0,0 +1,54 @@
# Open SWE — stock LangGraph dev server (3+ graphs + FastAPI webapp, :2024).
# TEMPLATE: tokens (@@...@@) are rendered at first boot by user-data.sh.
# In-memory runtime (--no-reload) + ExecStartPost reseed; no Aegra/Postgres.
#
# Boot contract:
# ExecStartPre = fetch-config.sh <env> -> runs as root (`+`) ONLY to materialize
# the SERVICE-USER-owned tmpfs .env (@@ENV_FILE@@) from Secrets
# Manager + SSM and chown it to @@SERVICE_USER@@, fail-fast (the
# unit does NOT start if config can't be fetched).
# ExecStart = langgraph dev (as @@SERVICE_USER@@, bound to 127.0.0.1 — nginx
# is the sole ingress; never binds 0.0.0.0).
# ExecStartPost = seed_store.sh <env> -> reseeds team_settings + user_mappings
# that the in-memory store loses on every restart.
[Unit]
Description=Open SWE stock LangGraph dev server (graphs + webapp, :@@PORT@@)
After=network-online.target
Wants=network-online.target
RequiresMountsFor=/run/open-swe
[Service]
Type=simple
User=@@SERVICE_USER@@
Group=@@SERVICE_USER@@
WorkingDirectory=@@APP_DIR@@
# No EnvironmentFile: the secret-bearing app .env (@@ENV_FILE@@) is loaded by
# langgraph/dotenv (so the multiline GitHub App PEM never hits systemd's env
# parser), and seed_store.sh reads the same .env directly (without sourcing it).
#
# ExecStartPre runs as root (`+`) so it can chown the tmpfs .env to the service
# user; the env arg (@@OPENSWE_ENV@@) selects the SSM/Secrets prefix (T5 BOOT-01).
ExecStartPre=+@@FETCH_CONFIG@@ @@OPENSWE_ENV@@
# Bind 127.0.0.1 only — nginx proxies dashboard + webhooks; :@@PORT@@ is never
# directly network-reachable (T5 OSWE-IAC-03).
ExecStart=@@VENV@@/bin/langgraph dev --host 127.0.0.1 --port @@PORT@@ --no-browser --no-reload
ExecStartPost=@@SEED_STORE@@ @@OPENSWE_ENV@@
# App logs to a file CloudWatch collects (30-day retention set in the CW config).
StandardOutput=append:/var/log/open-swe/app.log
StandardError=append:/var/log/open-swe/app.log
Restart=on-failure
RestartSec=5
TimeoutStartSec=180
# Hardening — the box holds no durable state of its own.
NoNewPrivileges=true
ProtectSystem=full
ProtectHome=true
PrivateTmp=true
ReadWritePaths=/var/log/open-swe /var/www/open-swe /run/open-swe @@APP_DIR@@
[Install]
WantedBy=multi-user.target

128
deploy/ami/user-data.sh Executable file
View file

@ -0,0 +1,128 @@
#!/usr/bin/env bash
# Open SWE EC2 user-data — PROVISIONING-ONLY (runs once, at first boot).
#
# This is the rationale for `userDataCausesReplacement: true` in CDK: user-data
# does FIRST-BOOT provisioning, never durable runtime config. Editing it is a
# deliberate instance replacement. Durable runtime config is fetched fresh on
# every service start by deploy/seahaven/fetch-config.sh (ExecStartPre).
#
# The box holds NO durable state of its own:
# - secrets/config -> Secrets Manager + SSM, materialized to a tmpfs .env at boot
# - app artifact -> pulled from S3 (open-swe-<env>-assets) via the instance role
# - store state -> reseeded by seed_store.sh (ExecStartPost) on every start
# => there is no RETAIN volume to protect; replacement is tolerated. The EBS
# discipline (snapshot root + wait state=completed BEFORE any replacing deploy)
# is the safety net, not durable on-box state. See README "EBS discipline".
#
# Tokens (@@...@@) are substituted by CDK when it renders this script into the
# launch template. region is read from IMDSv2 as a fallback.
set -euo pipefail
exec > >(tee -a /var/log/open-swe-user-data.log) 2>&1
echo "==> open-swe user-data start $(date -u +%FT%TZ)"
# --- CDK-rendered values -----------------------------------------------------
OPENSWE_ENV="@@OPENSWE_ENV@@" # dev | prod
ASSETS_BUCKET="@@ASSETS_BUCKET@@" # open-swe-<env>-assets
SERVER_NAME="@@SERVER_NAME@@" # openswe[-dev].seahaven.com
ARTIFACT_PREFIX="@@ARTIFACT_PREFIX@@" # e.g. releases/latest
# --- fixed layout (must match provision.sh + templates) ----------------------
SERVICE_USER="openswe"
APP_DIR="/opt/open-swe/app"
VENV="${APP_DIR}/.venv"
WWW_ROOT="/var/www/open-swe"
TEMPLATE_DIR="/opt/open-swe/templates"
ENV_FILE="/run/open-swe/.env"
PORT="2024"
FETCH_CONFIG="${APP_DIR}/deploy/seahaven/fetch-config.sh"
SEED_STORE="${APP_DIR}/deploy/seahaven/seed_store.sh"
# region from IMDSv2
TOKEN="$(curl -fsS -X PUT "http://169.254.169.254/latest/api/token" \
-H "X-aws-ec2-metadata-token-ttl-seconds: 300" || true)"
AWS_REGION="$(curl -fsS -H "X-aws-ec2-metadata-token: ${TOKEN}" \
http://169.254.169.254/latest/meta-data/placement/region || echo us-east-1)"
export AWS_DEFAULT_REGION="$AWS_REGION"
echo "env=${OPENSWE_ENV} region=${AWS_REGION} bucket=${ASSETS_BUCKET} host=${SERVER_NAME}"
# --- boot.env: non-secret pointers fetch-config.sh reads ---------------------
install -d -o root -g root -m 0755 /etc/open-swe
cat >/etc/open-swe/boot.env <<EOF
OPENSWE_ENV=${OPENSWE_ENV}
AWS_REGION=${AWS_REGION}
ASSETS_BUCKET=${ASSETS_BUCKET}
ENV_FILE=${ENV_FILE}
SECRETS_PREFIX=open-swe-${OPENSWE_ENV}
SSM_PREFIX=/open-swe-${OPENSWE_ENV}
EOF
chmod 0644 /etc/open-swe/boot.env
# --- ensure the tmpfs for the materialized .env is mounted -------------------
# (baked into /etc/fstab by the AMI; mount it now in case it isn't yet.)
install -d -o root -g root -m 0755 /run/open-swe || true
mountpoint -q /run/open-swe || mount /run/open-swe || mount -t tmpfs \
-o rw,nosuid,nodev,noexec,mode=0700,uid=${SERVICE_USER},gid=${SERVICE_USER},size=8m \
tmpfs /run/open-swe
# --- pull the app artifact from S3 (instance role; no git auth on box) -------
echo "==> fetch app artifact from s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/"
tmp="$(mktemp -d)"
aws s3 cp "s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/app.tar.gz" "${tmp}/app.tar.gz"
aws s3 cp "s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/spa.tar.gz" "${tmp}/spa.tar.gz"
install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0755 "$APP_DIR" "$WWW_ROOT"
tar -xzf "${tmp}/app.tar.gz" -C "$APP_DIR"
tar -xzf "${tmp}/spa.tar.gz" -C "$WWW_ROOT"
chown -R "$SERVICE_USER":"$SERVICE_USER" "$APP_DIR" "$WWW_ROOT"
rm -rf "$tmp"
# The app reads ./.env from WorkingDirectory (langgraph.json "env": ".env").
# Point it at the tmpfs file fetch-config.sh materializes.
ln -sfn "$ENV_FILE" "${APP_DIR}/.env"
# --- render + install the systemd unit ---------------------------------------
echo "==> install systemd unit"
sed \
-e "s|@@SERVICE_USER@@|${SERVICE_USER}|g" \
-e "s|@@APP_DIR@@|${APP_DIR}|g" \
-e "s|@@VENV@@|${VENV}|g" \
-e "s|@@PORT@@|${PORT}|g" \
-e "s|@@ENV_FILE@@|${ENV_FILE}|g" \
-e "s|@@OPENSWE_ENV@@|${OPENSWE_ENV}|g" \
-e "s|@@FETCH_CONFIG@@|${FETCH_CONFIG}|g" \
-e "s|@@SEED_STORE@@|${SEED_STORE}|g" \
"${TEMPLATE_DIR}/open-swe.service" >/etc/systemd/system/open-swe.service
systemctl daemon-reload
# --- render + install the nginx site -----------------------------------------
echo "==> install nginx site"
sed \
-e "s|@@SERVER_NAME@@|${SERVER_NAME}|g" \
-e "s|@@WWW_ROOT@@|${WWW_ROOT}|g" \
-e "s|@@BACKEND_ADDR@@|127.0.0.1:${PORT}|g" \
"${TEMPLATE_DIR}/open-swe.nginx.conf" >/etc/nginx/sites-available/open-swe
ln -sfn /etc/nginx/sites-available/open-swe /etc/nginx/sites-enabled/open-swe
rm -f /etc/nginx/sites-enabled/default
nginx -t
# --- CloudWatch agent: 30-day log retention ----------------------------------
echo "==> configure CloudWatch agent (30-day retention)"
sed -e "s|@@OPENSWE_ENV@@|${OPENSWE_ENV}|g" \
"${TEMPLATE_DIR}/amazon-cloudwatch-agent.json" \
>/opt/aws/amazon-cloudwatch-agent/etc/open-swe-cw.json
/opt/aws/amazon-cloudwatch-agent/bin/amazon-cloudwatch-agent-ctl \
-a fetch-config -m ec2 -s \
-c file:/opt/aws/amazon-cloudwatch-agent/etc/open-swe-cw.json
# --- start services ----------------------------------------------------------
# NOTE: intentionally NO swapfile here. The 8 GB-swapfile OOM hack existed only
# for the on-box Nitro SPA build, which now runs in GitHub Actions -> S3.
echo "==> start nginx + open-swe.service"
systemctl enable --now nginx
systemctl reload nginx
# open-swe.service ExecStartPre=fetch-config.sh fails-fast if config is missing,
# so a bad secrets/SSM setup surfaces as a failed unit (not a half-up box).
systemctl enable --now open-swe.service
echo "==> open-swe user-data done $(date -u +%FT%TZ)"