open-swe/deploy/ami/README.md
Adam Moussa ce9e5fb51b
feat(deploy): EC2 AMI recipe (Packer) + cloud-init/user-data (PR#2) (#7)
PR#2 of the AWS migration. deploy/ami/: Packer template (Ubuntu 24.04 arm64,
uv+py3.12, nginx, awscli v2, CW agent; no swapfile), provisioning-only user-data
(userDataCausesReplacement rationale), systemd unit + nginx + CW templates.

Incorporates T5 /sh-security-review fixes: langgraph binds 127.0.0.1 (not 0.0.0.0);
nginx is the sole ingress proxying only /dashboard/api/ + /webhooks/; ExecStartPre
runs fetch-config as root (+) and passes the env arg; the app runs as the
unprivileged openswe user reading an openswe-owned 0600 .env. packer validate clean.
2026-06-26 15:06:36 -04:00

167 lines
9.2 KiB
Markdown

# Open SWE base AMI (T8)
Packer recipe + first-boot user-data for the single EC2 instance per env
(`open-swe-dev` / `open-swe-prod`) in the Open SWE → AWS migration. Builds an
**ARM64 (Graviton) Ubuntu 24.04 LTS** base AMI and provisions the box on first
boot with the stock `langgraph dev` runtime, nginx, and the CloudWatch agent.
The architecture is locked in the repo `TODO.md` ("Architecture (locked)"): ONE
EC2 ARM64 (~t4g.large) instance per env, `seahaven-vpc` **private subnet + NAT**,
inbound **only from the ALB SG**. Runtime is **stock `langgraph dev`** (in-memory
store, `--no-reload`) + nginx + systemd. The box has **no git auth** — it pulls
its deploy artifact from S3 via the instance role. The SPA build runs in GitHub
Actions (T7), **not** on the box, so the old 8 GB-swapfile OOM hack is gone.
## Files
| Path | Purpose |
|---|---|
| `open-swe-base.pkr.hcl` | Packer template (HCL2). Latest Canonical 24.04 arm64 source → base AMI. |
| `scripts/provision.sh` | Packer provisioner. System packages, uv+py3.12, node+bun, service user, stages templates. |
| `user-data.sh` | First-boot provisioning (S3 artifact pull, render templates, CW agent, start services). |
| `templates/open-swe.service` | systemd unit TEMPLATE (`@@tokens@@` rendered at boot). |
| `templates/open-swe.nginx.conf` | nginx site TEMPLATE (dashboard SPA + scoped `/dashboard/api/` proxy). |
| `templates/amazon-cloudwatch-agent.json` | CW agent config TEMPLATE — **30-day log retention**. |
`deploy/seahaven/fetch-config.sh` and `deploy/seahaven/seed_store.sh` are owned by
the parallel T10 work and ship **inside the app artifact**; this AMI wires them in
but does not author them (see "Integration contract" below).
## Build the AMI
```bash
cd deploy/ami
packer init .
packer fmt -check .
packer validate -var aws_region=us-east-1 open-swe-base.pkr.hcl
packer build open-swe-base.pkr.hcl
```
Builds in account **328440206208 / us-east-1**. Source = latest Canonical Ubuntu
24.04 (Noble) **arm64** AMI (`source_ami_filter`, owner `099720109477`). Build host
is `t4g.medium` (ARM64). The output AMI is tagged:
```
Name=open-swe-base-arm64 Purpose=open-swe-runtime-base ManagedBy=packer
```
Pinned versions live in the template `variable` defaults (`uv_version`,
`python_version`, `node_major`, the CW-agent / awscli URLs) and the
`required_plugins` block (`amazon` 1.3.6) — bump deliberately.
## AMI → `cdk.context.json` pinning contract
The CDK stacks in `/infra` (owned by T3/T12) consume the AMI **by id, pinned in the
committed `infra/cdk.context.json`** — they never resolve "latest" at synth time.
This is the EBS/AMI-fix discipline: an uncached `MachineImage.lookup` resolves a new
AMI on every deploy and silently triggers instance replacement.
Contract (CDK side does the wiring; this is the handshake):
1. `packer build` prints the new AMI id (and tags it `open-swe-base-arm64`).
2. CDK looks the AMI up with **`cachedInContext: true`** (e.g.
`MachineImage.lookup({ name: "open-swe-base-arm64-*", owners: ["328440206208"], cachedInContext: true })`),
which writes the resolved id into `infra/cdk.context.json`.
3. **`infra/cdk.context.json` is committed.** From then on every synth/deploy uses
the pinned id — no surprise replacement when a newer AMI exists.
4. To adopt a new AMI: `cdk context --reset <ami-lookup-key>` (or edit the pinned
value), commit the change, and review the cdk-diff — the PR will show
"requires replacement", which is the intended, visible signal.
Record the built AMI id in project memory (`project_open_swe_migration`) per the
"memory updated for AMI id" build criterion.
## `userDataCausesReplacement` rationale
`user-data.sh` is **provisioning-only** — it runs once at first boot and never
carries durable runtime config. CDK sets **`userDataCausesReplacement: true`** so
that any change to it is a deliberate, diff-visible instance replacement rather than
a no-op edit that drifts from the running box. Durable runtime config is fetched
**fresh on every service start** by `fetch-config.sh` (ExecStartPre) — changing a
secret or SSM value needs only a `systemctl restart open-swe.service`, not a
replacement.
## EBS discipline (binding — `feedback_inline_ebs_volumes`)
**The box holds no durable state of its own:**
| State | Lives in | On replacement |
|---|---|---|
| secrets / config | Secrets Manager + SSM → tmpfs `.env` | re-fetched at boot |
| app code + SPA | S3 `open-swe-<env>-assets` | re-pulled at boot |
| store (team_settings, user_mappings) | reseeded by `seed_store.sh` | re-seeded at boot |
| logs | CloudWatch (30-day) — **not** a CFN resource in the stack | survive replacement |
→ **No local-only durable state ⇒ no standalone RETAIN volume is needed.** The root
volume is disposable; there is intentionally no inline data `blockDevices` to lose.
**Even so, snapshot before any replacing deploy.** Per the operational guard, before
merging/deploying any change that REPLACES the instance (`userDataCausesReplacement`,
AMI bump, instance-type change):
1. Enumerate the instance's volumes and assert **"no local-only durable state"**
(the table above is the checklist).
2. Take an **EBS snapshot of the root volume and WAIT for `state=completed`** before
letting the deploy proceed. Keep it as insurance; delete after a grace period.
3. Confirm the CloudWatch log groups are **not** CFN-managed in the stack so history
survives; re-verify history after the new instance is healthy.
cdk-diff-on-PR must flag "requires replacement" at review time (T8/T12 acceptance
criterion). This is the enforced version — not just an assertion in the runbook.
## Integration contract (T10 — `fetch-config.sh` + `seed_store.sh`)
Both ship in the app artifact under `deploy/seahaven/` and are wired into the unit:
- **`fetch-config.sh`** (ExecStartPre, runs as `openswe`): reads `/etc/open-swe/boot.env`
(`OPENSWE_ENV`, `AWS_REGION`, `SECRETS_PREFIX=open-swe-<env>`, `SSM_PREFIX=/open-swe-<env>`,
`ENV_FILE=/run/open-swe/.env`), pulls Secrets Manager `open-swe-<env>/*` + SSM
`/open-swe-<env>/*`, and writes:
- `/run/open-swe/.env` (**0600, tmpfs**, secret-bearing app env incl. the multiline
GitHub App PEM) — loaded by langgraph/dotenv via the `${APP_DIR}/.env` symlink.
- `/run/open-swe/seed.env` (**0600, tmpfs**, simple `OPENSWE_*` vars only:
`OPENSWE_DEFAULT_REPO`, `OPENSWE_OWNER_LOGIN`, `OPENSWE_OWNER_EMAIL`, model ids) —
loaded by systemd `EnvironmentFile` so `seed_store.sh` (ExecStartPost) has them.
- It must **fail-fast** (non-zero exit) if any required value is missing, so the
unit never starts half-configured.
- **`seed_store.sh`** (ExecStartPost): existing script, reseeds `team_settings/default`
+ `user_mappings/<login>` into the in-memory store after each start.
## Smoke-boot checklist (after first boot)
SSM Session Manager onto the instance (no public SSH — private subnet) and verify:
- [ ] `cloud-init status --wait` → `done`; `/var/log/open-swe-user-data.log` ends with
"user-data done" and shows the S3 pulls + service starts.
- [ ] `systemctl is-active open-swe.service` → `active`. (If it failed, check
`ExecStartPre`/`fetch-config.sh` — fail-fast means missing config = failed unit.)
- [ ] **fetch-config fail-fast works:** `/run/open-swe/.env` exists, owner `openswe`,
mode `0600`, on tmpfs (`findmnt /run/open-swe`); `seed.env` present.
- [ ] `curl -fsS http://127.0.0.1:2024/ok` → `200` (raw LangGraph health).
- [ ] `systemctl is-active nginx` → `active`; `curl -fsS http://127.0.0.1/healthz` →
`200`; `curl -s http://127.0.0.1/threads` returns the SPA shell, **not** JSON
(proves the agent API is not proxied — the security boundary holds).
- [ ] `seed_store: done` in the journal / app.log (store reseeded).
- [ ] CloudWatch: log groups `/open-swe/<env>/{app,user-data,nginx-access,nginx-error}`
exist with **30-day** retention and are receiving events.
- [ ] **No swapfile** (`swapon --show` empty) — the on-box SPA build is gone.
- [ ] From the ALB only: dashboard host serves the SPA; `hooks` host reaches
`/webhooks/*` on :2024 and nothing else (raw API paths hit the ALB default, not
the box).
## Assumptions
- **Artifact bucket** `open-swe-<env>-assets` (T7), with objects
`${ARTIFACT_PREFIX}/app.tar.gz` (Python app incl. `deploy/seahaven/` and a prebuilt
arm64 `.venv`) and `${ARTIFACT_PREFIX}/spa.tar.gz` (built SPA → `/var/www/open-swe`).
`ARTIFACT_PREFIX` defaults to `releases/latest`; CDK renders the concrete value.
- **Instance role** (defined in `/infra`, least-privilege per T4/T12) grants:
`s3:GetObject` on `open-swe-<env>-assets/*`; `secretsmanager:GetSecretValue` on
`open-swe-<env>/*`; `ssm:GetParameter(s)`/`GetParametersByPath` on `/open-swe-<env>/*`;
`logs:*` for the CW agent log groups + `cloudwatch:PutMetricData`; SSM Session
Manager (`ssm:UpdateInstanceInformation`, `ssmmessages:*`) for shell access.
- **CDK substitutes** the `@@OPENSWE_ENV@@`, `@@ASSETS_BUCKET@@`, `@@SERVER_NAME@@`,
`@@ARTIFACT_PREFIX@@` tokens in `user-data.sh` when rendering the launch template.
- `:2024` binds `0.0.0.0` so the ALB hooks target group can reach `/webhooks/*`; it is
reachable only from the ALB SG (private subnet, SG-scoped inbound). The raw API is
never internet-exposed — the ALB hooks rule is path-scoped to `/webhooks/*`.