AppService construct wires the per-env EC2 box and its internet path. The
seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the
on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and
never owned/mutated; open-swe only ADDS its own resources.
Per env (open-swe-stack.ts → AppService):
- ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a,
in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination
(no RETAIN volume — replacement-tolerant; see ami-cache.ts).
userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh.
- Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress
via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress
(the imported, on-prem-owned SG is never mutated).
- Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays
loopback). Health check GET /healthz.
- Two rules on the imported :443 listener, both → the TG:
* webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/*
* site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api)
Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5)
or it would steal every webhook — first-match-by-ascending-priority.
- Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB.
- 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config).
Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed
medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large
GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on
/dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene);
XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed
critical/high.
Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked
open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean.
Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
|
||
|---|---|---|
| .. | ||
| scripts | ||
| templates | ||
| open-swe-base.pkr.hcl | ||
| README.md | ||
| user-data.sh | ||
Open SWE base AMI (T8)
Packer recipe + first-boot user-data for the single EC2 instance per env
(open-swe-dev / open-swe-prod) in the Open SWE → AWS migration. Builds an
ARM64 (Graviton) Ubuntu 24.04 LTS base AMI and provisions the box on first
boot with the stock langgraph dev runtime, nginx, and the CloudWatch agent.
The architecture is locked in the repo TODO.md ("Architecture (locked)"): ONE
EC2 ARM64 (~t4g.large) instance per env, seahaven-vpc private subnet + NAT,
inbound only from the ALB SG. Runtime is stock langgraph dev (in-memory
store, --no-reload) + nginx + systemd. The box has no git auth — it pulls
its deploy artifact from S3 via the instance role. The SPA build runs in GitHub
Actions (T7), not on the box, so the old 8 GB-swapfile OOM hack is gone.
Files
| Path | Purpose |
|---|---|
open-swe-base.pkr.hcl |
Packer template (HCL2). Latest Canonical 24.04 arm64 source → base AMI. |
scripts/provision.sh |
Packer provisioner. System packages, uv+py3.12, node+bun, service user, stages templates. |
user-data.sh |
First-boot provisioning (S3 artifact pull, render templates, CW agent, start services). |
templates/open-swe.service |
systemd unit TEMPLATE (@@tokens@@ rendered at boot). |
templates/open-swe.nginx.conf |
nginx site TEMPLATE (dashboard SPA + scoped /dashboard/api/ proxy). |
templates/amazon-cloudwatch-agent.json |
CW agent config TEMPLATE — 30-day log retention. |
deploy/seahaven/fetch-config.sh and deploy/seahaven/seed_store.sh are owned by
the parallel T10 work and ship inside the app artifact; this AMI wires them in
but does not author them (see "Integration contract" below).
Build the AMI
cd deploy/ami
packer init .
packer fmt -check .
packer validate -var aws_region=us-east-1 open-swe-base.pkr.hcl
packer build open-swe-base.pkr.hcl
Builds in account 328440206208 / us-east-1. Source = latest Canonical Ubuntu
24.04 (Noble) arm64 AMI (source_ami_filter, owner 099720109477). Build host
is t4g.medium (ARM64). The output AMI is tagged:
Name=open-swe-base-arm64 Purpose=open-swe-runtime-base ManagedBy=packer
Pinned versions live in the template variable defaults (uv_version,
python_version, node_major, the CW-agent / awscli URLs) and the
required_plugins block (amazon 1.3.6) — bump deliberately.
AMI → cdk.context.json pinning contract
The CDK stacks in /infra (owned by T3/T12) consume the AMI by id, pinned in the
committed infra/cdk.context.json — they never resolve "latest" at synth time.
This is the EBS/AMI-fix discipline: an uncached MachineImage.lookup resolves a new
AMI on every deploy and silently triggers instance replacement.
Contract (CDK side does the wiring; this is the handshake):
packer buildprints the new AMI id (and tags itopen-swe-base-arm64).- CDK looks the AMI up with
cachedInContext: true(e.g.MachineImage.lookup({ name: "open-swe-base-arm64-*", owners: ["328440206208"], cachedInContext: true })), which writes the resolved id intoinfra/cdk.context.json. infra/cdk.context.jsonis committed. From then on every synth/deploy uses the pinned id — no surprise replacement when a newer AMI exists.- To adopt a new AMI:
cdk context --reset <ami-lookup-key>(or edit the pinned value), commit the change, and review the cdk-diff — the PR will show "requires replacement", which is the intended, visible signal.
Record the built AMI id in project memory (project_open_swe_migration) per the
"memory updated for AMI id" build criterion.
userDataCausesReplacement rationale
user-data.sh is provisioning-only — it runs once at first boot and never
carries durable runtime config. CDK sets userDataCausesReplacement: true so
that any change to it is a deliberate, diff-visible instance replacement rather than
a no-op edit that drifts from the running box. Durable runtime config is fetched
fresh on every service start by fetch-config.sh (ExecStartPre) — changing a
secret or SSM value needs only a systemctl restart open-swe.service, not a
replacement.
EBS discipline (binding — feedback_inline_ebs_volumes)
The box holds no durable state of its own:
| State | Lives in | On replacement |
|---|---|---|
| secrets / config | Secrets Manager + SSM → tmpfs .env |
re-fetched at boot |
| app code + SPA | S3 open-swe-<env>-assets |
re-pulled at boot |
| store (team_settings, user_mappings) | reseeded by seed_store.sh |
re-seeded at boot |
| logs | CloudWatch (30-day) — not a CFN resource in the stack | survive replacement |
→ No local-only durable state ⇒ no standalone RETAIN volume is needed. The root
volume is disposable; there is intentionally no inline data blockDevices to lose.
Even so, snapshot before any replacing deploy. Per the operational guard, before
merging/deploying any change that REPLACES the instance (userDataCausesReplacement,
AMI bump, instance-type change):
- Enumerate the instance's volumes and assert "no local-only durable state" (the table above is the checklist).
- Take an EBS snapshot of the root volume and WAIT for
state=completedbefore letting the deploy proceed. Keep it as insurance; delete after a grace period. - Confirm the CloudWatch log groups are not CFN-managed in the stack so history survives; re-verify history after the new instance is healthy.
cdk-diff-on-PR must flag "requires replacement" at review time (T8/T12 acceptance criterion). This is the enforced version — not just an assertion in the runbook.
Integration contract (T10 — fetch-config.sh + seed_store.sh)
Both ship in the app artifact under deploy/seahaven/ and are wired into the unit:
fetch-config.sh(ExecStartPre, runs asopenswe): reads/etc/open-swe/boot.env(OPENSWE_ENV,AWS_REGION,SECRETS_PREFIX=open-swe-<env>,SSM_PREFIX=/open-swe-<env>,ENV_FILE=/run/open-swe/.env), pulls Secrets Manageropen-swe-<env>/*+ SSM/open-swe-<env>/*, and writes:/run/open-swe/.env(0600, tmpfs, secret-bearing app env incl. the multiline GitHub App PEM) — loaded by langgraph/dotenv via the${APP_DIR}/.envsymlink./run/open-swe/seed.env(0600, tmpfs, simpleOPENSWE_*vars only:OPENSWE_DEFAULT_REPO,OPENSWE_OWNER_LOGIN,OPENSWE_OWNER_EMAIL, model ids) — loaded by systemdEnvironmentFilesoseed_store.sh(ExecStartPost) has them.- It must fail-fast (non-zero exit) if any required value is missing, so the unit never starts half-configured.
seed_store.sh(ExecStartPost): existing script, reseedsteam_settings/defaultuser_mappings/<login>into the in-memory store after each start.
Smoke-boot checklist (after first boot)
SSM Session Manager onto the instance (no public SSH — private subnet) and verify:
cloud-init status --wait→done;/var/log/open-swe-user-data.logends with "user-data done" and shows the S3 pulls + service starts.systemctl is-active open-swe.service→active. (If it failed, checkExecStartPre/fetch-config.sh— fail-fast means missing config = failed unit.)- fetch-config fail-fast works:
/run/open-swe/.envexists, owneropenswe, mode0600, on tmpfs (findmnt /run/open-swe);seed.envpresent. curl -fsS http://127.0.0.1:2024/ok→200(raw LangGraph health).systemctl is-active nginx→active;curl -fsS http://127.0.0.1/healthz→200;curl -s http://127.0.0.1/threadsreturns the SPA shell, not JSON (proves the agent API is not proxied — the security boundary holds).seed_store: donein the journal / app.log (store reseeded).- CloudWatch: log groups
/open-swe/<env>/{app,user-data,nginx-access,nginx-error}exist with 30-day retention and are receiving events. - No swapfile (
swapon --showempty) — the on-box SPA build is gone. - From the ALB only: dashboard host serves the SPA;
hookshost reaches/webhooks/*on :2024 and nothing else (raw API paths hit the ALB default, not the box).
Assumptions
- Artifact bucket
open-swe-<env>-assets(T7), with objects${ARTIFACT_PREFIX}/app.tar.gz(Python app incl.deploy/seahaven/and a prebuilt arm64.venv) and${ARTIFACT_PREFIX}/spa.tar.gz(built SPA →/var/www/open-swe).ARTIFACT_PREFIXdefaults toreleases/latest; CDK renders the concrete value. - Instance role (defined in
/infra, least-privilege per T4/T12) grants:s3:GetObjectonopen-swe-<env>-assets/*;secretsmanager:GetSecretValueonopen-swe-<env>/*;ssm:GetParameter(s)/GetParametersByPathon/open-swe-<env>/*;logs:*for the CW agent log groups +cloudwatch:PutMetricData; SSM Session Manager (ssm:UpdateInstanceInformation,ssmmessages:*) for shell access. - CDK substitutes the
@@OPENSWE_ENV@@,@@ASSETS_BUCKET@@,@@SERVER_NAME@@,@@ARTIFACT_PREFIX@@tokens inuser-data.shwhen rendering the launch template. :2024binds0.0.0.0so the ALB hooks target group can reach/webhooks/*; it is reachable only from the ALB SG (private subnet, SG-scoped inbound). The raw API is never internet-exposed — the ALB hooks rule is path-scoped to/webhooks/*.