open-swe/deploy/ami
Adam Moussa fcbdfb67aa
feat: open-swe dev/prod compute + ALB ingress (T12) (#14)
AppService construct wires the per-env EC2 box and its internet path. The
seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the
on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and
never owned/mutated; open-swe only ADDS its own resources.

Per env (open-swe-stack.ts → AppService):
- ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a,
  in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination
  (no RETAIN volume — replacement-tolerant; see ami-cache.ts).
  userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh.
- Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress
  via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress
  (the imported, on-prem-owned SG is never mutated).
- Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays
  loopback). Health check GET /healthz.
- Two rules on the imported :443 listener, both → the TG:
    * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/*
    * site     (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api)
  Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5)
  or it would steal every webhook — first-match-by-ascending-priority.
- Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB.
- 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config).

Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed
medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large
GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on
/dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene);
XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed
critical/high.

Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked
open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean.
Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
..
scripts feat(deploy): EC2 AMI recipe (Packer) + cloud-init/user-data (PR#2) (#7) 2026-06-26 15:06:36 -04:00
templates feat: open-swe dev/prod compute + ALB ingress (T12) (#14) 2026-06-26 16:10:23 -04:00
open-swe-base.pkr.hcl feat(deploy): EC2 AMI recipe (Packer) + cloud-init/user-data (PR#2) (#7) 2026-06-26 15:06:36 -04:00
README.md feat(deploy): EC2 AMI recipe (Packer) + cloud-init/user-data (PR#2) (#7) 2026-06-26 15:06:36 -04:00
user-data.sh feat(deploy): EC2 AMI recipe (Packer) + cloud-init/user-data (PR#2) (#7) 2026-06-26 15:06:36 -04:00

Open SWE base AMI (T8)

Packer recipe + first-boot user-data for the single EC2 instance per env (open-swe-dev / open-swe-prod) in the Open SWE → AWS migration. Builds an ARM64 (Graviton) Ubuntu 24.04 LTS base AMI and provisions the box on first boot with the stock langgraph dev runtime, nginx, and the CloudWatch agent.

The architecture is locked in the repo TODO.md ("Architecture (locked)"): ONE EC2 ARM64 (~t4g.large) instance per env, seahaven-vpc private subnet + NAT, inbound only from the ALB SG. Runtime is stock langgraph dev (in-memory store, --no-reload) + nginx + systemd. The box has no git auth — it pulls its deploy artifact from S3 via the instance role. The SPA build runs in GitHub Actions (T7), not on the box, so the old 8 GB-swapfile OOM hack is gone.

Files

Path Purpose
open-swe-base.pkr.hcl Packer template (HCL2). Latest Canonical 24.04 arm64 source → base AMI.
scripts/provision.sh Packer provisioner. System packages, uv+py3.12, node+bun, service user, stages templates.
user-data.sh First-boot provisioning (S3 artifact pull, render templates, CW agent, start services).
templates/open-swe.service systemd unit TEMPLATE (@@tokens@@ rendered at boot).
templates/open-swe.nginx.conf nginx site TEMPLATE (dashboard SPA + scoped /dashboard/api/ proxy).
templates/amazon-cloudwatch-agent.json CW agent config TEMPLATE — 30-day log retention.

deploy/seahaven/fetch-config.sh and deploy/seahaven/seed_store.sh are owned by the parallel T10 work and ship inside the app artifact; this AMI wires them in but does not author them (see "Integration contract" below).

Build the AMI

cd deploy/ami
packer init .
packer fmt -check .
packer validate -var aws_region=us-east-1 open-swe-base.pkr.hcl
packer build open-swe-base.pkr.hcl

Builds in account 328440206208 / us-east-1. Source = latest Canonical Ubuntu 24.04 (Noble) arm64 AMI (source_ami_filter, owner 099720109477). Build host is t4g.medium (ARM64). The output AMI is tagged:

Name=open-swe-base-arm64   Purpose=open-swe-runtime-base   ManagedBy=packer

Pinned versions live in the template variable defaults (uv_version, python_version, node_major, the CW-agent / awscli URLs) and the required_plugins block (amazon 1.3.6) — bump deliberately.

AMI → cdk.context.json pinning contract

The CDK stacks in /infra (owned by T3/T12) consume the AMI by id, pinned in the committed infra/cdk.context.json — they never resolve "latest" at synth time. This is the EBS/AMI-fix discipline: an uncached MachineImage.lookup resolves a new AMI on every deploy and silently triggers instance replacement.

Contract (CDK side does the wiring; this is the handshake):

  1. packer build prints the new AMI id (and tags it open-swe-base-arm64).
  2. CDK looks the AMI up with cachedInContext: true (e.g. MachineImage.lookup({ name: "open-swe-base-arm64-*", owners: ["328440206208"], cachedInContext: true })), which writes the resolved id into infra/cdk.context.json.
  3. infra/cdk.context.json is committed. From then on every synth/deploy uses the pinned id — no surprise replacement when a newer AMI exists.
  4. To adopt a new AMI: cdk context --reset <ami-lookup-key> (or edit the pinned value), commit the change, and review the cdk-diff — the PR will show "requires replacement", which is the intended, visible signal.

Record the built AMI id in project memory (project_open_swe_migration) per the "memory updated for AMI id" build criterion.

userDataCausesReplacement rationale

user-data.sh is provisioning-only — it runs once at first boot and never carries durable runtime config. CDK sets userDataCausesReplacement: true so that any change to it is a deliberate, diff-visible instance replacement rather than a no-op edit that drifts from the running box. Durable runtime config is fetched fresh on every service start by fetch-config.sh (ExecStartPre) — changing a secret or SSM value needs only a systemctl restart open-swe.service, not a replacement.

EBS discipline (binding — feedback_inline_ebs_volumes)

The box holds no durable state of its own:

State Lives in On replacement
secrets / config Secrets Manager + SSM → tmpfs .env re-fetched at boot
app code + SPA S3 open-swe-<env>-assets re-pulled at boot
store (team_settings, user_mappings) reseeded by seed_store.sh re-seeded at boot
logs CloudWatch (30-day) — not a CFN resource in the stack survive replacement

→ No local-only durable state ⇒ no standalone RETAIN volume is needed. The root volume is disposable; there is intentionally no inline data blockDevices to lose.

Even so, snapshot before any replacing deploy. Per the operational guard, before merging/deploying any change that REPLACES the instance (userDataCausesReplacement, AMI bump, instance-type change):

  1. Enumerate the instance's volumes and assert "no local-only durable state" (the table above is the checklist).
  2. Take an EBS snapshot of the root volume and WAIT for state=completed before letting the deploy proceed. Keep it as insurance; delete after a grace period.
  3. Confirm the CloudWatch log groups are not CFN-managed in the stack so history survives; re-verify history after the new instance is healthy.

cdk-diff-on-PR must flag "requires replacement" at review time (T8/T12 acceptance criterion). This is the enforced version — not just an assertion in the runbook.

Integration contract (T10 — fetch-config.sh + seed_store.sh)

Both ship in the app artifact under deploy/seahaven/ and are wired into the unit:

  • fetch-config.sh (ExecStartPre, runs as openswe): reads /etc/open-swe/boot.env (OPENSWE_ENV, AWS_REGION, SECRETS_PREFIX=open-swe-<env>, SSM_PREFIX=/open-swe-<env>, ENV_FILE=/run/open-swe/.env), pulls Secrets Manager open-swe-<env>/* + SSM /open-swe-<env>/*, and writes:
    • /run/open-swe/.env (0600, tmpfs, secret-bearing app env incl. the multiline GitHub App PEM) — loaded by langgraph/dotenv via the ${APP_DIR}/.env symlink.
    • /run/open-swe/seed.env (0600, tmpfs, simple OPENSWE_* vars only: OPENSWE_DEFAULT_REPO, OPENSWE_OWNER_LOGIN, OPENSWE_OWNER_EMAIL, model ids) — loaded by systemd EnvironmentFile so seed_store.sh (ExecStartPost) has them.
    • It must fail-fast (non-zero exit) if any required value is missing, so the unit never starts half-configured.
  • seed_store.sh (ExecStartPost): existing script, reseeds team_settings/default
    • user_mappings/<login> into the in-memory store after each start.

Smoke-boot checklist (after first boot)

SSM Session Manager onto the instance (no public SSH — private subnet) and verify:

  • cloud-init status --wait → done; /var/log/open-swe-user-data.log ends with "user-data done" and shows the S3 pulls + service starts.
  • systemctl is-active open-swe.service → active. (If it failed, check ExecStartPre/fetch-config.sh — fail-fast means missing config = failed unit.)
  • fetch-config fail-fast works: /run/open-swe/.env exists, owner openswe, mode 0600, on tmpfs (findmnt /run/open-swe); seed.env present.
  • curl -fsS http://127.0.0.1:2024/ok → 200 (raw LangGraph health).
  • systemctl is-active nginx → active; curl -fsS http://127.0.0.1/healthz → 200; curl -s http://127.0.0.1/threads returns the SPA shell, not JSON (proves the agent API is not proxied — the security boundary holds).
  • seed_store: done in the journal / app.log (store reseeded).
  • CloudWatch: log groups /open-swe/<env>/{app,user-data,nginx-access,nginx-error} exist with 30-day retention and are receiving events.
  • No swapfile (swapon --show empty) — the on-box SPA build is gone.
  • From the ALB only: dashboard host serves the SPA; hooks host reaches /webhooks/* on :2024 and nothing else (raw API paths hit the ALB default, not the box).

Assumptions

  • Artifact bucket open-swe-<env>-assets (T7), with objects ${ARTIFACT_PREFIX}/app.tar.gz (Python app incl. deploy/seahaven/ and a prebuilt arm64 .venv) and ${ARTIFACT_PREFIX}/spa.tar.gz (built SPA → /var/www/open-swe). ARTIFACT_PREFIX defaults to releases/latest; CDK renders the concrete value.
  • Instance role (defined in /infra, least-privilege per T4/T12) grants: s3:GetObject on open-swe-<env>-assets/*; secretsmanager:GetSecretValue on open-swe-<env>/*; ssm:GetParameter(s)/GetParametersByPath on /open-swe-<env>/*; logs:* for the CW agent log groups + cloudwatch:PutMetricData; SSM Session Manager (ssm:UpdateInstanceInformation, ssmmessages:*) for shell access.
  • CDK substitutes the @@OPENSWE_ENV@@, @@ASSETS_BUCKET@@, @@SERVER_NAME@@, @@ARTIFACT_PREFIX@@ tokens in user-data.sh when rendering the launch template.
  • :2024 binds 0.0.0.0 so the ALB hooks target group can reach /webhooks/*; it is reachable only from the ALB SG (private subnet, SG-scoped inbound). The raw API is never internet-exposed — the ALB hooks rule is path-scoped to /webhooks/*.