mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 15:03:16 +00:00
* feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3)
Packer-build the custom base image and repoint AppService off the AL2023
placeholder onto it.
deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real
`packer build` (the config had only ever been `packer validate`'d at T8):
- the file provisioner failed uploading the templates dir ('scp: …: Is a
directory') — a trailing-slash contents-upload needs the dest dir to exist;
added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the
dest trailing slash.
- the shell provisioner's custom execute_command omitted {{ .Vars }}, so the
environment_vars never reached provision.sh (which runs under set -u and
aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}.
infra:
- ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26
from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by
exact id via MachineImage.genericLinux (offline, deterministic). Dropped the
now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement
discipline docs.
- app-service.ts: machineImage → bakedOpenSweArm64().
- open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard).
- cdk.context.json → {} (AMI is a static id pin; no context lookups remain).
- README: Baked AMI + EBS-replacement-discipline section.
tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI.
NOTE: held — do NOT merge until the open-swe-dev secret values are populated
(put-config.sh). The infra CD is live, so merging this to dev auto-deploys
OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts →
unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14).
* fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021
Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects
non-ASCII in the AMI Description attribute, so packer registered then
DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error.
Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available).
Re-pinned BAKED_OPEN_SWE_AMI_ID.
* fix(deploy): GitHub App + Slack required for prod only, not dev
Per the migration decision: do NOT create/duplicate a separate dev GitHub App or
Slack app — only prod owns the single shared app. So fetch-config.sh no longer
hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/
CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only
block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET.
Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active
provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env
(boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is
unchanged (prod still requires everything).
* feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19)
Make the dev/prod box deployable end-to-end: a real artifact pipeline and a
re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy.
Infra (T7):
- assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access,
SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent +
abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput.
- app-service.ts: open-swe-<env>-deploy SSM document that runs the baked
/opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the
baked open-swe-base-arm64 AMI (folds in the held #16).
IAM (app deploy role — cross-review gated):
- github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to
open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic
AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document
is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox.
Boot/deploy (T19):
- deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull
app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv
at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx.
- user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB
target is healthy even before the first release); deploy.sh is base64-rendered
by CDK into user-data (a normal reviewable repo file, not a heredoc) and the
first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy).
CI (T7+T19):
- build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite ->
ui/.output/public -> spa.tar.gz), package the Python source via git archive
(app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via
the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy.
push dev -> dev (auto); push main -> prod (env "prod" approval gate).
Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK,
deploy.sh base64 round-trips exact.
* harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard
Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one
confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01,
the account-wide CDK cfn-exec residual already documented in config.ts; recorded in
.security-review/suppressions.json with justification + flagged for the per-env
bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified;
these are the cheap defense-in-depth fixes worth taking regardless:
- deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root
never honors an archive's uid/mode → no setuid/foreign-owned file can land); and
treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy
failure (set -e stays loud once a release exists).
- publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount),
not CommandInvocations[0], so a partial failure across the brief 2-instance
replacement window can't be reported as success.
- instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app
role's write scope) instead of the whole bucket.
- package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz
(defense in depth over .gitignore; scoped to data extensions so *_credentials.py
source is not a false positive — verified against the real tree).
Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the
CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable
releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release
to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz).
shellcheck/tsc/jest(16) clean; both stacks synth offline.
* fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard
The instance-SG GroupDescription + ingress/egress rule descriptions carried an
em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects
non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"),
so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing
from #14; same class as the AMI-description ASCII bug.)
- app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the
ingress/egress rule descriptions, and the Route53 comment.
- test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup
GroupDescription + rule descriptions are pure ASCII, so this fails the build
instead of a deploy next time.
jest 18/18; tsc clean.
* fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`)
The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule*
descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`,
which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the
ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII
only" to the exact EC2 allowed charset so it catches `>` (and `<`) too.
jest 18/18; tsc clean.
* fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit
The base64 deploy.sh embedded in user-data pushed the encoded boot script to
27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with
"Encoded User data is limited to 25600 bytes". Strip full-line comments + blank
lines from deploy.sh before base64-embedding it (repo file keeps comments; only
the on-box copy is minified; the script is opaque base64 so user-data heredocs are
unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a
synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded.
jest 19/19; minified deploy.sh passes bash -n + shellcheck.
157 lines
6.8 KiB
TypeScript
157 lines
6.8 KiB
TypeScript
import * as iam from "aws-cdk-lib/aws-iam";
|
|
import { Construct } from "constructs";
|
|
import {
|
|
ACCOUNT,
|
|
EnvName,
|
|
GITHUB_OIDC_PROVIDER_ARN,
|
|
REGION,
|
|
oidcSubject,
|
|
} from "../config";
|
|
|
|
/**
|
|
* Per-ENV GitHub Actions OIDC deploy roles. Created ONCE per env in the
|
|
* dedicated `open-swe-iam` stack. Two roles per env, per the locked architecture's
|
|
* "dual OIDC roles":
|
|
*
|
|
* - githubdeploy-open-swe-infra-<env> → CFN/IAM (CDK) deploys of that env's stack
|
|
* - githubdeploy-open-swe-app-<env> → app deploys (env-tag-scoped SSM + S3 read)
|
|
*
|
|
* T5 OSWE-IAC-01/02 fix: roles are split per env and the trust subject is
|
|
* env-scoped (dev = dev branch ref; prod = the GitHub `prod` Environment subject,
|
|
* so the manual-approval gate is IAM-enforced). A dev-branch token therefore
|
|
* cannot SendCommand to the prod box nor assume a prod deploy role.
|
|
*
|
|
* Reviewed at T4 (GPT-4.1 IAM cross-review) + T5 (/sh-security-review) and
|
|
* deployed FIRST (BLOCK#3 "OIDC-role-first" ordering) before any other infra or
|
|
* secrets CI step.
|
|
*/
|
|
export class GithubDeployRoles extends Construct {
|
|
public readonly infraRole: iam.Role;
|
|
public readonly appRole: iam.Role;
|
|
|
|
constructor(scope: Construct, id: string, envName: EnvName) {
|
|
super(scope, id);
|
|
|
|
// The provider already exists account-wide — reference, never re-create.
|
|
const provider = iam.OpenIdConnectProvider.fromOpenIdConnectProviderArn(
|
|
this,
|
|
"GithubOidcProvider",
|
|
GITHUB_OIDC_PROVIDER_ARN,
|
|
);
|
|
|
|
// T4 BLOCK#2 + T5 IAC-01/02: exact env-scoped subject via StringEquals (no
|
|
// StringLike, no `*`). prod = environment:prod (manual-approval gate),
|
|
// dev = the dev branch ref.
|
|
const trust = new iam.WebIdentityPrincipal(provider.openIdConnectProviderArn, {
|
|
StringEquals: {
|
|
"token.actions.githubusercontent.com:aud": "sts.amazonaws.com",
|
|
"token.actions.githubusercontent.com:sub": oidcSubject(envName),
|
|
},
|
|
});
|
|
|
|
// ---- githubdeploy-open-swe-infra-<env> --------------------------------
|
|
this.infraRole = new iam.Role(this, "InfraDeployRole", {
|
|
roleName: `githubdeploy-open-swe-infra-${envName}`,
|
|
assumedBy: trust,
|
|
description: `GitHub OIDC role for CDK deploys of the open-swe-${envName} infra stack (assumes CDK bootstrap roles).`,
|
|
});
|
|
|
|
// Org-standard CDK deploy pattern (mirrors githubdeploy-seahaven-account-
|
|
// baseline / -forgejo / -apm-wo-analysis): the deploy role only needs to
|
|
// assume the CDK bootstrap roles. The actual CloudFormation + IAM + resource
|
|
// permissions are exercised by the bootstrap `cfn-exec-role`, whose scope is
|
|
// owned by the CDKToolkit stack — NOT granted directly here.
|
|
//
|
|
// T4 BLOCK#1: GPT-4.1 flagged the `cdk-hnb659fds-*` wildcard and recommended
|
|
// enumerating the four exact ARNs. ACCEPTED EXCEPTION (Adam, 2026-06-26): kept
|
|
// as the verified org-wide convention (githubdeploy-seahaven-account-baseline
|
|
// uses the identical wildcard). Only `cdk bootstrap` creates roles with this
|
|
// prefix, so practical escalation risk is low.
|
|
// T5 residual (OSWE-IAC-02): the single account-wide cfn-exec-role means the
|
|
// dev infra role can technically deploy any stack; per-env trust gates WHO can
|
|
// assume, and the prod role requires the environment:prod approval. Per-env
|
|
// bootstrap qualifiers would close the residual fully (future hardening).
|
|
this.infraRole.addToPolicy(
|
|
new iam.PolicyStatement({
|
|
sid: "AssumeCdkBootstrapRoles",
|
|
actions: ["sts:AssumeRole"],
|
|
resources: [`arn:aws:iam::${ACCOUNT}:role/cdk-hnb659fds-*`],
|
|
}),
|
|
);
|
|
|
|
// ---- githubdeploy-open-swe-app-<env> ----------------------------------
|
|
this.appRole = new iam.Role(this, "AppDeployRole", {
|
|
roleName: `githubdeploy-open-swe-app-${envName}`,
|
|
assumedBy: trust,
|
|
description: `GitHub OIDC role for open-swe-${envName} app deploys: env-tag-scoped ssm:SendCommand + read of the ${envName} S3 artifact bucket.`,
|
|
});
|
|
|
|
// T5 OSWE-IAC-01 fix: SendCommand only to instances tagged project=open-swe
|
|
// AND env=<this env> (a SINGLE value, not {dev,prod}). The dev app role can
|
|
// never command the prod box and vice versa — env isolation in IAM.
|
|
this.appRole.addToPolicy(
|
|
new iam.PolicyStatement({
|
|
sid: "SsmSendCommandTagScoped",
|
|
actions: ["ssm:SendCommand"],
|
|
resources: [`arn:aws:ec2:${REGION}:${ACCOUNT}:instance/*`],
|
|
conditions: {
|
|
StringEquals: {
|
|
"ssm:resourceTag/project": "open-swe",
|
|
"ssm:resourceTag/env": envName,
|
|
},
|
|
},
|
|
}),
|
|
);
|
|
|
|
// SendCommand also has to reference the command document. Scope to this env's
|
|
// open-swe deploy document ONLY.
|
|
// T4 BLOCK#3 (CLOSED at T19): GPT-4.1 flagged AWS-RunShellScript as an
|
|
// arbitrary-shell escalation path. The dedicated `open-swe-${envName}-deploy`
|
|
// SSM document (app-service.ts) now runs the fixed, parameter-less command
|
|
// `bash /opt/open-swe/bin/deploy.sh`, so AWS-RunShellScript is dropped here:
|
|
// this role can run ONLY that one document, and only on its own env's box
|
|
// (tag-scoped by the SsmSendCommandTagScoped statement above).
|
|
this.appRole.addToPolicy(
|
|
new iam.PolicyStatement({
|
|
sid: "SsmSendCommandDocuments",
|
|
actions: ["ssm:SendCommand"],
|
|
resources: [`arn:aws:ssm:${REGION}:${ACCOUNT}:document/open-swe-${envName}-deploy`],
|
|
}),
|
|
);
|
|
|
|
// Poll command results. These read actions do not support resource-level
|
|
// scoping, so `*` is required by the API (T4 FIX: API limitation, documented).
|
|
this.appRole.addToPolicy(
|
|
new iam.PolicyStatement({
|
|
sid: "SsmReadCommandStatus",
|
|
actions: [
|
|
"ssm:GetCommandInvocation",
|
|
"ssm:ListCommands",
|
|
"ssm:ListCommandInvocations",
|
|
],
|
|
resources: ["*"],
|
|
}),
|
|
);
|
|
|
|
// Read+WRITE access to THIS env's artifact bucket only (T19): the
|
|
// build-artifacts workflow uploads app.tar.gz / spa.tar.gz under releases/*,
|
|
// then fires the deploy document so the box pulls them via its instance role.
|
|
// Object actions are scoped to releases/* (the only prefix CI writes), and to
|
|
// THIS env's bucket — a dev token can never write the prod bucket. No
|
|
// bucket-level mutation (no PutBucket*/Delete bucket) — that stays with CDK.
|
|
this.appRole.addToPolicy(
|
|
new iam.PolicyStatement({
|
|
sid: "ReadWriteArtifactObjects",
|
|
actions: ["s3:GetObject", "s3:PutObject", "s3:DeleteObject"],
|
|
resources: [`arn:aws:s3:::open-swe-${envName}-assets/releases/*`],
|
|
}),
|
|
);
|
|
this.appRole.addToPolicy(
|
|
new iam.PolicyStatement({
|
|
sid: "ListArtifactBucket",
|
|
actions: ["s3:ListBucket", "s3:GetBucketLocation"],
|
|
resources: [`arn:aws:s3:::open-swe-${envName}-assets`],
|
|
}),
|
|
);
|
|
}
|
|
}
|