open-swe/infra/lib/open-swe-stack.ts

76 lines
3.1 KiB
TypeScript
Raw Normal View History

import * as cdk from "aws-cdk-lib";
import { Construct } from "constructs";
import { EnvName, prefix } from "./config";
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
import { AppService } from "./constructs/app-service";
import { ConfigStore } from "./constructs/config-store";
import { InstanceRole } from "./constructs/instance-role";
feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14).
2026-06-26 16:31:58 -04:00
import { BAKED_OPEN_SWE_AMI_ID } from "./constructs/ami-cache";
export interface OpenSweStackProps extends cdk.StackProps {
/** open-swe environment — drives the `open-swe-<env>-*` resource naming. */
readonly envName: EnvName;
}
/**
* Per-env open-swe stack (`open-swe-dev` / `open-swe-prod`). Resource names are
* prefixed `open-swe-<env>-*`.
*
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
* Composes: the per-env least-privilege instance role (T6), the Secrets/SSM
* config store (T11), and the compute + ingress wiring (T12, AppService — EC2
* box, instance SG, target group, imported-listener rules, Route53 aliases,
* 30-day log groups). The shared VPC and ALB are imported, never owned. Synth is
* offline (AMI is the cdk.context.json-pinned placeholder until T12-deploy).
*/
export class OpenSweStack extends cdk.Stack {
public readonly instanceRole: InstanceRole;
public readonly configStore: ConfigStore;
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
public readonly appService: AppService;
constructor(scope: Construct, id: string, props: OpenSweStackProps) {
super(scope, id, props);
const envName = props.envName;
const p = prefix(envName);
cdk.Tags.of(this).add("project", "open-swe");
cdk.Tags.of(this).add("env", envName);
cdk.Tags.of(this).add("ManagedBy", "cdk");
// Per-env least-privilege EC2 instance role (open-swe-<env>-instance-role).
this.instanceRole = new InstanceRole(this, "Instance", envName);
// Secrets Manager + SSM Parameter Store shells the boot hook reads
// (deploy/seahaven/fetch-config.sh). Secret shells are value-less and
// populated out-of-band; IaC-managed SSM params carry real derivable values.
// The instance role already grants read on open-swe-<env>/* + /open-swe-<env>/*.
this.configStore = new ConfigStore(this, "Config", { envName });
feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14).
2026-06-26 16:31:58 -04:00
// Surface the baked open-swe base AMI id the box runs on (pinned by id in
// ami-cache.ts; refreshed by a deliberate packer rebuild → replacement).
new cdk.CfnOutput(this, "BakedAmiId", {
value: BAKED_OPEN_SWE_AMI_ID,
description: "Baked open-swe-base-arm64 AMI id consumed by the EC2 instance.",
});
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
// T12: compute + ingress. Imports the shared seahaven-vpc + ALB and adds the
// env's EC2 box, instance SG, target group, listener rules, DNS, log groups.
this.appService = new AppService(this, "App", {
envName,
instanceRole: this.instanceRole.role,
});
new cdk.CfnOutput(this, "InstanceRoleArn", {
value: this.instanceRole.role.roleArn,
description: `${p} EC2 instance role ARN.`,
});
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
new cdk.CfnOutput(this, "InstanceId", {
value: this.appService.instance.instanceId,
description: `${p} EC2 instance id.`,
});
new cdk.CfnOutput(this, "TargetGroupArn", {
value: this.appService.targetGroup.targetGroupArn,
description: `${p} ALB target group ARN (→ instance:80 nginx).`,
});
}
}