mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 10:23:14 +00:00
feat: open-swe dev/prod compute + ALB ingress (T12) (#14)
AppService construct wires the per-env EC2 box and its internet path. The
seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the
on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and
never owned/mutated; open-swe only ADDS its own resources.
Per env (open-swe-stack.ts → AppService):
- ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a,
in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination
(no RETAIN volume — replacement-tolerant; see ami-cache.ts).
userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh.
- Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress
via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress
(the imported, on-prem-owned SG is never mutated).
- Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays
loopback). Health check GET /healthz.
- Two rules on the imported :443 listener, both → the TG:
* webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/*
* site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api)
Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5)
or it would steal every webhook — first-match-by-ascending-priority.
- Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB.
- 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config).
Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed
medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large
GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on
/dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene);
XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed
critical/high.
Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked
open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean.
Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
This commit is contained in:
parent
aef6b26912
commit
fcbdfb67aa
4 changed files with 390 additions and 13 deletions
|
|
@ -26,6 +26,7 @@ server {
|
|||
location /dashboard/api/ {
|
||||
proxy_pass http://@@BACKEND_ADDR@@;
|
||||
proxy_http_version 1.1;
|
||||
client_max_body_size 10m;
|
||||
proxy_set_header Host $host;
|
||||
proxy_set_header X-Real-IP $remote_addr;
|
||||
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
|
||||
|
|
@ -41,6 +42,11 @@ server {
|
|||
location /webhooks/ {
|
||||
proxy_pass http://@@BACKEND_ADDR@@;
|
||||
proxy_http_version 1.1;
|
||||
# GitHub permits webhook payloads up to 25 MB; nginx's 1 MB default would
|
||||
# 413 large push/PR events at the edge BEFORE in-app signature verification
|
||||
# runs, silently dropping them (OSWE-T12-01). proxy_request_buffering off
|
||||
# does not relax the size cap — set it explicitly.
|
||||
client_max_body_size 25m;
|
||||
proxy_set_header Host $host;
|
||||
proxy_set_header X-Real-IP $remote_addr;
|
||||
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
|
||||
|
|
|
|||
|
|
@ -13,13 +13,14 @@ infra/
|
|||
├── lib/
|
||||
│ ├── config.ts # account/region/org constants, env type, OIDC trust subjects
|
||||
│ ├── open-swe-iam-stack.ts # account-level: shared OIDC deploy roles
|
||||
│ ├── open-swe-stack.ts # per-env stack (instance role + AMI cache wiring)
|
||||
│ ├── open-swe-stack.ts # per-env stack (instance role + config store + AppService)
|
||||
│ ├── aspects/
|
||||
│ │ └── kebab-naming-aspect.ts # fails synth on any non-kebab-case explicit name
|
||||
│ └── constructs/
|
||||
│ ├── github-deploy-roles.ts # githubdeploy-open-swe-infra + githubdeploy-open-swe-app
|
||||
│ ├── instance-role.ts # open-swe-<env>-instance-role (least-privilege)
|
||||
│ ├── config-store.ts # Secrets Manager + SSM Parameter Store shells (T11)
|
||||
│ ├── app-service.ts # EC2 box + imported-ALB ingress + Route53 + logs (T12)
|
||||
│ └── ami-cache.ts # cached ARM64 AL2023 helper + EBS/AMI discipline docs
|
||||
├── test/
|
||||
│ └── kebab-naming-aspect.test.ts # jest: Aspect passes conforming names, flags bad ones
|
||||
|
|
@ -36,8 +37,8 @@ infra/
|
|||
| Stack name (kebab) | Construct | Contents |
|
||||
|---|---|---|
|
||||
| `open-swe-iam` | `OpenSweIamStack` | Account-level shared GitHub OIDC deploy roles (singletons). |
|
||||
| `open-swe-dev` | `OpenSweStack` (`envName: dev`) | `open-swe-dev-instance-role` + AMI-cache wiring. EC2/ALB/etc. land at T12. |
|
||||
| `open-swe-prod` | `OpenSweStack` (`envName: prod`) | `open-swe-prod-instance-role` + AMI-cache wiring. |
|
||||
| `open-swe-dev` | `OpenSweStack` (`envName: dev`) | `open-swe-dev-instance-role`, config store (T11), and the EC2 box + ALB ingress (T12, `AppService`). |
|
||||
| `open-swe-prod` | `OpenSweStack` (`envName: prod`) | `open-swe-prod-instance-role`, config store, and the EC2 box + ALB ingress. |
|
||||
|
||||
Account `328440206208`, region `us-east-1`. Stack names are set explicitly so CDK
|
||||
never defaults to PascalCase; resource names follow `open-swe-<env>-*`.
|
||||
|
|
@ -139,7 +140,64 @@ deploy/seahaven/fetch-config.sh <dev|prod> # (on the box) fail-fast verify befor
|
|||
provide each value inline, via `OPENSWE_PUT_<VAR>` env vars, or from a vault. It does
|
||||
NOT touch the IaC-managed params (CDK owns those — editing them here would drift).
|
||||
|
||||
## AMI cache discipline (EBS-fix plumbing — stub for T12)
|
||||
## Compute + ingress (`AppService` — T12)
|
||||
|
||||
`AppService` (`lib/constructs/app-service.ts`, one per env from `OpenSweStack`)
|
||||
builds the box and its path to the internet. A **single** internet-facing ALB
|
||||
(`app/seahaven-com`) and a **single** VPC are shared with the on-prem
|
||||
`seahaven-site` stack, so open-swe **imports** the VPC, the ALB security group
|
||||
(`sg-0b0301deed193258a`), the `:443` listener, and the public `seahaven.com`
|
||||
zone — and never owns/mutates them. It **adds**:
|
||||
|
||||
- **One ARM64 EC2 box** (`open-swe-<env>-box`, `t4g.medium` dev / `t4g.large`
|
||||
prod) in **private1 (us-east-1a)** — same AZ as the single NAT for in-AZ egress.
|
||||
`requireImdsv2`, gp3 **encrypted** root, `deleteOnTermination` (no RETAIN
|
||||
volume — see below). `userDataCausesReplacement: true`; user-data is rendered
|
||||
from `deploy/ami/user-data.sh`.
|
||||
- **A standalone instance SG** reachable **only** from the shared ALB SG on `:80`
|
||||
(nginx). Egress open (NAT). The ALB SG is opened to the box via a **standalone
|
||||
`CfnSecurityGroupEgress`** so the imported (on-prem-owned) SG is never mutated.
|
||||
- **A target group → instance `:80`** (nginx is the sole ingress; the LangGraph
|
||||
control plane stays on loopback `:2024`). Health check `GET /healthz`.
|
||||
- **Two listener rules** on the imported `:443` listener, both → the same TG:
|
||||
- **Webhooks** (priority **2** dev / **3** prod): `host ∈ {openswe-<env>, hooks-<env>}.seahaven.com` **AND** path `/webhooks/*`.
|
||||
- **Site** (priority **10** dev / **11** prod): `host = openswe-<env>.seahaven.com` (dashboard SPA + `/dashboard/api/`).
|
||||
- **Route53 alias records** `openswe[-dev]` + `hooks[-dev]` → the shared ALB.
|
||||
- **Four CloudWatch log groups** (`/open-swe/<env>/{app,user-data,nginx-access,nginx-error}`) at **30-day** retention (IaC-owned; mirrors the CW-agent config).
|
||||
|
||||
### Listener-rule ordering (load-bearing)
|
||||
|
||||
The shared listener already has a **host-agnostic** `/webhooks/*` PATH rule at
|
||||
**priority 5** (on-prem). ALB rules are first-match by ascending priority, so the
|
||||
open-swe webhook rule **must** sit below 5 or every `…/webhooks/*` request (any
|
||||
host) is forwarded to the on-prem target first. Hence priority 2/3. The rule ANDs
|
||||
a host condition, so it does **not** steal the on-prem hosts' webhooks. The
|
||||
dashboard "site" rule carries no path that collides with rule 5, so it sits at
|
||||
10/11.
|
||||
|
||||
**Cross-stack coordination (T13 review).** The `seahaven-site` (on-prem) and
|
||||
`open-swe` stacks both add resources to the *same imported* listener and ALB SG.
|
||||
This is safe: each stack owns only the resources it declares (its own logical
|
||||
ids), so an on-prem deploy can't delete open-swe's rules/egress and vice-versa,
|
||||
and the standalone `CfnSecurityGroupEgress` never mutates the shared SG's own
|
||||
definition (the pattern on-prem itself uses). The one shared namespace that needs
|
||||
care is **listener-rule priority** (globally unique per listener; a collision is
|
||||
a fail-*safe* deploy error, not silent drift). Ownership — keep disjoint:
|
||||
`seahaven-site` = **4-7 + default**; `open-swe` = **2, 3, 10, 11**. open-swe's
|
||||
webhook rules are host-scoped to its own `*.seahaven.com` hosts, so they never
|
||||
match an on-prem `seahavenind.com` host.
|
||||
|
||||
### Security review (T5/T12 `/sh-security-review`)
|
||||
|
||||
The T12 surface was run through the detector-fan-out + proof-or-kill verifier.
|
||||
One **confirmed medium** (OSWE-T12-01: nginx's 1 MB default `client_max_body_size`
|
||||
would 413 large GitHub webhooks before in-app signature verification) is fixed in
|
||||
`open-swe.nginx.conf` (`25m` on `/webhooks/`, `10m` on `/dashboard/api/`). The
|
||||
hooks hostname is scoped to `/webhooks/*` only (OSWE-T12-02 hygiene). An
|
||||
X-Forwarded-For spoof candidate was **killed** — no code trusts the leftmost XFF.
|
||||
No confirmed critical/high; no block.
|
||||
|
||||
## AMI cache discipline (EBS-fix plumbing — consumed by T12 `AppService`)
|
||||
|
||||
`cachedArm64AmazonLinux2023()` (in `lib/constructs/ami-cache.ts`) returns an
|
||||
ARM64 Amazon Linux 2023 image with `cachedInContext: true`, so the resolved AMI
|
||||
|
|
@ -147,7 +205,7 @@ id is pinned in the committed `cdk.context.json`. Without the pin, every deploy
|
|||
could pick up a newer AL2023 release → AMI change → **EC2 instance replacement**
|
||||
(the file-share data-loss root cause — memory `feedback_inline_ebs_volumes`).
|
||||
|
||||
Design intent documented in code for T12 to plug into:
|
||||
Design intent (now consumed by `AppService`):
|
||||
|
||||
- `userDataCausesReplacement: true` is the **deliberate** choice — user-data is
|
||||
provisioning-only and carries no durable state.
|
||||
|
|
@ -155,9 +213,9 @@ Design intent documented in code for T12 to plug into:
|
|||
store is rebuilt on every boot from S3 + Secrets Manager / SSM, so there is
|
||||
intentionally no standalone `ec2.Volume` + `removalPolicy.RETAIN`. The goal is
|
||||
replacement-*tolerance*, not avoidance.
|
||||
- **Snapshot-before-replace** still applies operationally at T12: snapshot the
|
||||
root volume and wait `state=completed` before any replacing deploy, and re-verify
|
||||
"no local-only durable state" first.
|
||||
- **Snapshot-before-replace** still applies operationally: before any replacing
|
||||
deploy snapshot the root volume and wait `state=completed`, and re-verify "no
|
||||
local-only durable state" first.
|
||||
|
||||
Refresh the AMI pin deliberately:
|
||||
|
||||
|
|
@ -168,7 +226,9 @@ cdk synth # review the diff — it WILL show "requires replacement"
|
|||
|
||||
> The committed `cdk.context.json` ships a dummy-but-valid-shaped AMI id
|
||||
> (`ami-00000000000000000`) so `cdk synth` resolves the cache locally without any
|
||||
> live AWS call. Replace it with the real resolved id when T8/T12 build the AMI.
|
||||
> live AWS call. **Before the first real deploy**, repoint `AppService` to the
|
||||
> baked `open-swe-base-arm64` AMI and pin its real id — the placeholder is
|
||||
> intentionally un-bootable on stock AL2023.
|
||||
|
||||
## Commands
|
||||
|
||||
|
|
|
|||
293
infra/lib/constructs/app-service.ts
Normal file
293
infra/lib/constructs/app-service.ts
Normal file
|
|
@ -0,0 +1,293 @@
|
|||
import * as fs from "fs";
|
||||
import * as path from "path";
|
||||
import * as cdk from "aws-cdk-lib";
|
||||
import * as ec2 from "aws-cdk-lib/aws-ec2";
|
||||
import * as elbv2 from "aws-cdk-lib/aws-elasticloadbalancingv2";
|
||||
import * as elbTargets from "aws-cdk-lib/aws-elasticloadbalancingv2-targets";
|
||||
import * as logs from "aws-cdk-lib/aws-logs";
|
||||
import * as route53 from "aws-cdk-lib/aws-route53";
|
||||
import * as iam from "aws-cdk-lib/aws-iam";
|
||||
import { Construct } from "constructs";
|
||||
import { EnvName, prefix } from "../config";
|
||||
import { cachedArm64AmazonLinux2023 } from "./ami-cache";
|
||||
|
||||
/**
|
||||
* Shared seahaven-vpc + internet-facing ALB facts (read-only recon 2026-06-26;
|
||||
* scratchpad/T12-infra-facts.md). A SINGLE VPC and a SINGLE shared ALB front
|
||||
* both the on-prem `seahaven-site` stack and open-swe. We IMPORT every one of
|
||||
* these and NEVER own them — open-swe only ADDS its own instance SG, a standalone
|
||||
* ALB-egress rule, listener rules, a target group, and DNS records.
|
||||
*
|
||||
* ── Cross-stack coordination on the SHARED listener + ALB SG (T13 review) ──
|
||||
* Two CDK stacks (seahaven-site, open-swe) add resources to the same imported
|
||||
* `:443` listener and ALB SG. This is safe because each stack owns ONLY the
|
||||
* resources it declares (its own logical ids): an on-prem `cdk deploy` computes a
|
||||
* changeset over its own template and cannot delete rules/egress it never
|
||||
* declared. The standalone-egress pattern is what on-prem itself uses
|
||||
* (sgr-0c57812752a3bca13), so it does not mutate the shared SG's own definition.
|
||||
*
|
||||
* The ONE shared namespace that REQUIRES coordination is listener-rule PRIORITY
|
||||
* (globally unique per listener; a collision is a fail-SAFE deploy error, not
|
||||
* silent drift). Ownership map — keep these disjoint when editing either stack:
|
||||
* - seahaven-site (on-prem): priorities 4-7 + default.
|
||||
* - open-swe: priorities 2, 3 (webhooks) and 10, 11 (site).
|
||||
* open-swe's webhook rules are HOST-scoped to its own *.seahaven.com hosts, so
|
||||
* they never match (let alone "steal") any seahavenind.com / on-prem host.
|
||||
*/
|
||||
const SHARED = {
|
||||
vpcId: "vpc-0d3d4b67bd0cf8a68",
|
||||
availabilityZones: ["us-east-1a", "us-east-1b"],
|
||||
// Private subnets host the EC2 box. The single NAT gateway lives in 1a, so the
|
||||
// box is pinned to private1 (1a) for in-AZ NAT egress (no cross-AZ data $).
|
||||
privateSubnetIds: ["subnet-04e38c507e96f1926", "subnet-0a0b4fc6f296dfba5"],
|
||||
instanceSubnetId: "subnet-04e38c507e96f1926",
|
||||
instanceAz: "us-east-1a",
|
||||
albDnsName: "seahaven-com-1856441924.us-east-1.elb.amazonaws.com",
|
||||
albCanonicalHostedZoneId: "Z35SXDOTRQ7X7K",
|
||||
albSecurityGroupId: "sg-0b0301deed193258a",
|
||||
httpsListenerArn:
|
||||
"arn:aws:elasticloadbalancing:us-east-1:328440206208:listener/app/seahaven-com/222c3257354ab559/bab8bcf0da0e2927",
|
||||
publicZoneId: "Z06652411XKH89KTZD3XA",
|
||||
publicZoneName: "seahaven.com",
|
||||
} as const;
|
||||
|
||||
/**
|
||||
* Per-env public hostnames, listener-rule priorities, and instance size.
|
||||
*
|
||||
* ── Listener-rule ordering hazard (load-bearing) ──
|
||||
* The shared listener already has a HOST-AGNOSTIC `/webhooks/*` PATH rule at
|
||||
* priority 5 (the on-prem seahaven-site stack owns it). ALB rules are first-match
|
||||
* by ASCENDING priority, so a `…/webhooks/*` request to our host would match
|
||||
* rule 5 (priority 5) and be forwarded to the on-prem target BEFORE any host rule
|
||||
* at 10+. Therefore our webhook rule MUST sit below priority 5. The catch-all
|
||||
* "site" rule (dashboard SPA + /dashboard/api/) carries no path that collides
|
||||
* with rule 5, so it can sit at any free higher number (10/11). Free priorities
|
||||
* confirmed by recon: 1-3 and 8+ (4=forgejo, 5=/webhooks/*, 6/7=seahavenind).
|
||||
*/
|
||||
const ENV_NET: Record<
|
||||
EnvName,
|
||||
{
|
||||
dashboardHost: string;
|
||||
hooksHost: string;
|
||||
webhookPriority: number;
|
||||
sitePriority: number;
|
||||
instanceType: string;
|
||||
}
|
||||
> = {
|
||||
dev: {
|
||||
dashboardHost: "openswe-dev.seahaven.com",
|
||||
hooksHost: "hooks-dev.seahaven.com",
|
||||
webhookPriority: 2,
|
||||
sitePriority: 10,
|
||||
instanceType: "t4g.medium",
|
||||
},
|
||||
prod: {
|
||||
dashboardHost: "openswe.seahaven.com",
|
||||
hooksHost: "hooks.seahaven.com",
|
||||
webhookPriority: 3,
|
||||
sitePriority: 11,
|
||||
instanceType: "t4g.large",
|
||||
},
|
||||
};
|
||||
|
||||
export interface AppServiceProps {
|
||||
readonly envName: EnvName;
|
||||
/** Least-privilege EC2 instance role (per-env; from InstanceRole). */
|
||||
readonly instanceRole: iam.IRole;
|
||||
/** S3 artifact key prefix the box pulls app.tar.gz / spa.tar.gz from. */
|
||||
readonly artifactPrefix?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* The open-swe compute + ingress wiring for one env (T12):
|
||||
* - one ARM64 EC2 box in private1 (1a), replacement-tolerant (no RETAIN volume),
|
||||
* - a standalone instance SG reachable ONLY from the shared ALB SG on :80,
|
||||
* - a target group → instance:80 (nginx is the sole ingress; :2024 stays loopback),
|
||||
* - two listener rules on the imported :443 listener (webhooks below the on-prem
|
||||
* path rule; site catch-all above it), both → the same TG,
|
||||
* - Route53 alias records for both hostnames → the shared ALB,
|
||||
* - IaC-owned CloudWatch log groups at 30-day retention.
|
||||
*
|
||||
* Everything ALB/VPC/zone-side is IMPORTED. Synth is offline: the AMI is the
|
||||
* cdk.context.json-pinned AL2023 ARM64 placeholder until the baked
|
||||
* open-swe-base-arm64 id is pinned before the first real deploy.
|
||||
*/
|
||||
export class AppService extends Construct {
|
||||
public readonly instance: ec2.Instance;
|
||||
public readonly targetGroup: elbv2.ApplicationTargetGroup;
|
||||
|
||||
constructor(scope: Construct, id: string, props: AppServiceProps) {
|
||||
super(scope, id);
|
||||
const env = props.envName;
|
||||
const p = prefix(env);
|
||||
const net = ENV_NET[env];
|
||||
const artifactPrefix = props.artifactPrefix ?? "releases/latest";
|
||||
|
||||
// Import the shared VPC with explicit attributes (no fromLookup → offline synth).
|
||||
const vpc = ec2.Vpc.fromVpcAttributes(this, "Vpc", {
|
||||
vpcId: SHARED.vpcId,
|
||||
availabilityZones: [...SHARED.availabilityZones],
|
||||
privateSubnetIds: [...SHARED.privateSubnetIds],
|
||||
});
|
||||
|
||||
// Standalone instance SG. Egress open (NAT path); ingress only from the ALB SG.
|
||||
const instanceSg = new ec2.SecurityGroup(this, "InstanceSg", {
|
||||
vpc,
|
||||
securityGroupName: `${p}-instance-sg`,
|
||||
description: `${p} instance SG — ingress only from the shared ALB SG on :80; egress via NAT.`,
|
||||
allowAllOutbound: true,
|
||||
});
|
||||
instanceSg.addIngressRule(
|
||||
ec2.Peer.securityGroupId(SHARED.albSecurityGroupId),
|
||||
ec2.Port.tcp(80),
|
||||
`${p}: shared ALB SG → nginx :80`,
|
||||
);
|
||||
// Open the IMPORTED ALB SG to our instance via a STANDALONE egress rule, so we
|
||||
// never mutate the ALB SG's own (on-prem-owned) definition.
|
||||
new ec2.CfnSecurityGroupEgress(this, "AlbToInstanceEgress", {
|
||||
groupId: SHARED.albSecurityGroupId,
|
||||
ipProtocol: "tcp",
|
||||
fromPort: 80,
|
||||
toPort: 80,
|
||||
destinationSecurityGroupId: instanceSg.securityGroupId,
|
||||
description: `${p}: ALB → instance nginx :80`,
|
||||
});
|
||||
|
||||
// Render the provisioning script's @@tokens@@ into the instance user-data.
|
||||
// userDataCausesReplacement makes a bootstrap change roll a fresh box (the box
|
||||
// holds no durable state — see ami-cache.ts / user-data.sh).
|
||||
const userDataPath = path.join(__dirname, "..", "..", "..", "deploy", "ami", "user-data.sh");
|
||||
const userData = ec2.UserData.custom(
|
||||
fs
|
||||
.readFileSync(userDataPath, "utf8")
|
||||
.replace(/@@OPENSWE_ENV@@/g, env)
|
||||
.replace(/@@ASSETS_BUCKET@@/g, `${p}-assets`)
|
||||
.replace(/@@SERVER_NAME@@/g, net.dashboardHost)
|
||||
.replace(/@@ARTIFACT_PREFIX@@/g, artifactPrefix),
|
||||
);
|
||||
|
||||
this.instance = new ec2.Instance(this, "Instance", {
|
||||
vpc,
|
||||
vpcSubnets: {
|
||||
subnets: [
|
||||
ec2.Subnet.fromSubnetAttributes(this, "InstanceSubnet", {
|
||||
subnetId: SHARED.instanceSubnetId,
|
||||
availabilityZone: SHARED.instanceAz,
|
||||
}),
|
||||
],
|
||||
},
|
||||
instanceType: new ec2.InstanceType(net.instanceType),
|
||||
// TODO(T12-deploy): repoint to the custom open-swe-base-arm64 AMI baked by
|
||||
// packer (deploy/ami). The AL2023 ARM64 cache is a synth-time placeholder
|
||||
// (pinned in cdk.context.json) so the box DEFINITION synths offline; the
|
||||
// baked AMI id is pinned (cdk context) before the first real deploy.
|
||||
// user-data.sh assumes the baked layout (/opt/open-swe, openswe user, nginx,
|
||||
// CloudWatch agent) — it is NOT runnable on the stock AL2023 placeholder.
|
||||
machineImage: cachedArm64AmazonLinux2023(),
|
||||
role: props.instanceRole,
|
||||
securityGroup: instanceSg,
|
||||
userData,
|
||||
userDataCausesReplacement: true,
|
||||
requireImdsv2: true,
|
||||
instanceName: `${p}-box`,
|
||||
blockDevices: [
|
||||
{
|
||||
deviceName: "/dev/xvda",
|
||||
// gp3 encrypted root; deleteOnTermination (no durable on-box state →
|
||||
// intentionally NO standalone RETAIN volume; see ami-cache.ts).
|
||||
volume: ec2.BlockDeviceVolume.ebs(30, {
|
||||
volumeType: ec2.EbsDeviceVolumeType.GP3,
|
||||
encrypted: true,
|
||||
deleteOnTermination: true,
|
||||
}),
|
||||
},
|
||||
],
|
||||
});
|
||||
|
||||
// Target group → instance:80 (nginx). Health check hits nginx's /healthz
|
||||
// (returns 200; the dashboard TG health path defined in open-swe.nginx.conf).
|
||||
this.targetGroup = new elbv2.ApplicationTargetGroup(this, "Tg", {
|
||||
vpc,
|
||||
targetGroupName: `${p}-tg`,
|
||||
port: 80,
|
||||
protocol: elbv2.ApplicationProtocol.HTTP,
|
||||
targetType: elbv2.TargetType.INSTANCE,
|
||||
targets: [new elbTargets.InstanceTarget(this.instance)],
|
||||
deregistrationDelay: cdk.Duration.seconds(15),
|
||||
healthCheck: {
|
||||
path: "/healthz",
|
||||
healthyHttpCodes: "200",
|
||||
interval: cdk.Duration.seconds(30),
|
||||
timeout: cdk.Duration.seconds(5),
|
||||
healthyThresholdCount: 2,
|
||||
unhealthyThresholdCount: 3,
|
||||
},
|
||||
});
|
||||
|
||||
// Import the shared :443 listener (with its ALB SG) and ADD our two rules.
|
||||
const albSg = ec2.SecurityGroup.fromSecurityGroupId(this, "AlbSg", SHARED.albSecurityGroupId, {
|
||||
mutable: false,
|
||||
});
|
||||
const listener = elbv2.ApplicationListener.fromApplicationListenerAttributes(this, "HttpsListener", {
|
||||
listenerArn: SHARED.httpsListenerArn,
|
||||
securityGroup: albSg,
|
||||
});
|
||||
|
||||
// (1) Webhooks — accepted on EITHER host (integrations may target either), and
|
||||
// MUST be below the on-prem path-only rule 5 (see ENV_NET note).
|
||||
new elbv2.ApplicationListenerRule(this, "WebhooksRule", {
|
||||
listener,
|
||||
priority: net.webhookPriority,
|
||||
conditions: [
|
||||
elbv2.ListenerCondition.hostHeaders([net.dashboardHost, net.hooksHost]),
|
||||
elbv2.ListenerCondition.pathPatterns(["/webhooks/*"]),
|
||||
],
|
||||
action: elbv2.ListenerAction.forward([this.targetGroup]),
|
||||
});
|
||||
// (2) Dashboard SPA + /dashboard/api/ (OAuth) — DASHBOARD host ONLY. The hooks
|
||||
// host intentionally serves nothing but /webhooks/* (rule 1), so the OAuth /
|
||||
// dashboard surface stays single-origin (OSWE-T12-02). Non-webhook paths on the
|
||||
// hooks host fall through to the on-prem default.
|
||||
new elbv2.ApplicationListenerRule(this, "SiteRule", {
|
||||
listener,
|
||||
priority: net.sitePriority,
|
||||
conditions: [elbv2.ListenerCondition.hostHeaders([net.dashboardHost])],
|
||||
action: elbv2.ListenerAction.forward([this.targetGroup]),
|
||||
});
|
||||
|
||||
// Route53 ALIAS records → the shared ALB, for both hostnames.
|
||||
const zone = route53.HostedZone.fromHostedZoneAttributes(this, "PublicZone", {
|
||||
hostedZoneId: SHARED.publicZoneId,
|
||||
zoneName: SHARED.publicZoneName,
|
||||
});
|
||||
const albAlias: route53.IAliasRecordTarget = {
|
||||
bind: () => ({
|
||||
dnsName: SHARED.albDnsName,
|
||||
hostedZoneId: SHARED.albCanonicalHostedZoneId,
|
||||
}),
|
||||
};
|
||||
for (const [label, host] of [
|
||||
["Dashboard", net.dashboardHost],
|
||||
["Hooks", net.hooksHost],
|
||||
] as const) {
|
||||
new route53.ARecord(this, `${label}Alias`, {
|
||||
zone,
|
||||
recordName: host,
|
||||
target: route53.RecordTarget.fromAlias(albAlias),
|
||||
comment: `${p} ${label.toLowerCase()} → shared seahaven-com ALB`,
|
||||
});
|
||||
}
|
||||
|
||||
// IaC-owned CloudWatch log groups at 30-day retention. Names mirror the
|
||||
// CloudWatch-agent config (deploy/ami/templates/amazon-cloudwatch-agent.json);
|
||||
// owning them here makes retention declarative rather than agent-set. Logs are
|
||||
// not durable state → DESTROY on stack delete.
|
||||
for (const suffix of ["app", "user-data", "nginx-access", "nginx-error"]) {
|
||||
new logs.LogGroup(this, `Log-${suffix}`, {
|
||||
logGroupName: `/open-swe/${env}/${suffix}`,
|
||||
retention: logs.RetentionDays.ONE_MONTH,
|
||||
removalPolicy: cdk.RemovalPolicy.DESTROY,
|
||||
});
|
||||
}
|
||||
}
|
||||
}
|
||||
|
|
@ -1,6 +1,7 @@
|
|||
import * as cdk from "aws-cdk-lib";
|
||||
import { Construct } from "constructs";
|
||||
import { EnvName, prefix } from "./config";
|
||||
import { AppService } from "./constructs/app-service";
|
||||
import { ConfigStore } from "./constructs/config-store";
|
||||
import { InstanceRole } from "./constructs/instance-role";
|
||||
import {
|
||||
|
|
@ -17,14 +18,16 @@ export interface OpenSweStackProps extends cdk.StackProps {
|
|||
* Per-env open-swe stack (`open-swe-dev` / `open-swe-prod`). Resource names are
|
||||
* prefixed `open-swe-<env>-*`.
|
||||
*
|
||||
* T3 scope: the per-env EC2 instance role + the wired-but-not-yet-instantiated
|
||||
* AMI cache helper. The EC2 instance, ALB target groups, listener rules, SG,
|
||||
* Route53 and NAT come at T12 — this stack is the synth-able shell they plug
|
||||
* into.
|
||||
* Composes: the per-env least-privilege instance role (T6), the Secrets/SSM
|
||||
* config store (T11), and the compute + ingress wiring (T12, AppService — EC2
|
||||
* box, instance SG, target group, imported-listener rules, Route53 aliases,
|
||||
* 30-day log groups). The shared VPC and ALB are imported, never owned. Synth is
|
||||
* offline (AMI is the cdk.context.json-pinned placeholder until T12-deploy).
|
||||
*/
|
||||
export class OpenSweStack extends cdk.Stack {
|
||||
public readonly instanceRole: InstanceRole;
|
||||
public readonly configStore: ConfigStore;
|
||||
public readonly appService: AppService;
|
||||
|
||||
constructor(scope: Construct, id: string, props: OpenSweStackProps) {
|
||||
super(scope, id, props);
|
||||
|
|
@ -58,9 +61,24 @@ export class OpenSweStack extends cdk.Stack {
|
|||
});
|
||||
}
|
||||
|
||||
// T12: compute + ingress. Imports the shared seahaven-vpc + ALB and adds the
|
||||
// env's EC2 box, instance SG, target group, listener rules, DNS, log groups.
|
||||
this.appService = new AppService(this, "App", {
|
||||
envName,
|
||||
instanceRole: this.instanceRole.role,
|
||||
});
|
||||
|
||||
new cdk.CfnOutput(this, "InstanceRoleArn", {
|
||||
value: this.instanceRole.role.roleArn,
|
||||
description: `${p} EC2 instance role ARN.`,
|
||||
});
|
||||
new cdk.CfnOutput(this, "InstanceId", {
|
||||
value: this.appService.instance.instanceId,
|
||||
description: `${p} EC2 instance id.`,
|
||||
});
|
||||
new cdk.CfnOutput(this, "TargetGroupArn", {
|
||||
value: this.appService.targetGroup.targetGroupArn,
|
||||
description: `${p} ALB target group ARN (→ instance:80 nginx).`,
|
||||
});
|
||||
}
|
||||
}
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue