mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 19:43:15 +00:00
7 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7f60324f0c
|
chore: decommission self-hosted AWS LangGraph stack (#64)
* chore: decommission self-hosted AWS LangGraph stack Removes the now-dead self-host IaC and AWS-only CI/CD after destroying the dev + prod CloudFormation stacks (open-swe-dev, open-swe-prod, open-swe-iam, and the dev-exclusive CDKToolkit-oswedev bootstrap) in account 328440206208, us-east-1. The deployment is now managed (LangGraph Cloud + Vercel). - remove infra/ (CDK app: app + IAM stacks, constructs, aspects, tests) - remove deploy/ami (Packer AMI build) and deploy/seahaven (boot/config scripts, DEPLOYMENT/ROTATION runbooks) - remove AWS-only workflows: cd-infra, ci-infra, build-artifacts, rollback - README: rewrite the Deployment section to the managed LangGraph Cloud + Vercel view; drop dead links to infra/ and deploy/seahaven Preserved: the shared default CDKToolkit bootstrap and promote-dev-to-prod.yml. The RETAIN'd Secrets Manager shells and open-swe-<env>-assets S3 buckets survive cdk destroy by design (orphaned) and need a separate deliberate cleanup. * chore: clean up dangling references left by the AWS decommission Folds in the FIX-level items from the #64 review gates (GPT-4.1 cross-review + /sh-security-review), none of which were blockers: - delete orphaned .github/scripts/{package-artifacts,publish-and-deploy,roll-box, rollback}.sh — their only callers were the removed AWS deploy workflows - drop the deleted /infra dir from dependabot.yml npm directories (was producing a recurring Dependabot config error) - remove the stale OSWE-IAC-SECRETS-LIST-01 suppression (referenced the deleted infra/lib/constructs/instance-role.ts) - repoint the README promotion link to promote-to-main.yml (renamed in #63) The promote-dev-to-prod.yml comment in check-dev-green.sh is intentionally left to #63, which rewrites that same line. |
||
|
|
430a1cdff9
|
ci: re-home prod promotion into a gated promote-to-main workflow (PR1: managed-LGC migration) (#63)
* ci: re-home prod promotion into a gated promote-to-main workflow Migrate the prod-deploy gate to managed LangGraph Cloud (git-connected to `main`) + Vercel. Under managed, a push to `main` auto-deploys prod, so the dev -> main fast-forward IS the prod deploy trigger -- the bespoke AWS CD step is obsolete and already gone from this workflow. Re-home `promote-dev-to-prod.yml` -> `promote-to-main.yml`: - gate the promote job on the `prod` GitHub Environment (required reviewer amoussa1229), restoring the manual prod-approval that the retired AWS CD job used to carry; - drop the nightly auto-promote cron -- a scheduled auto-promotion conflicts with a manual approval gate now that the push deploys prod; promotion is workflow_dispatch only; - keep the dev-HEAD-fully-green precondition and the seahaven-promotion App fast-forward push (sole non-admin bypass actor on `main` ruleset 18238334). Update the companion check-dev-green.sh filename reference. * ci: refuse promote-to-main dispatch from any ref other than dev Defense-in-depth atop the already-pinned `ref: dev` checkout: workflow_dispatch runs the workflow definition from the launched ref, so reject a non-dev dispatch before the App token is minted. Surfaced by the GPT-4.1 cross-review of #63. |
||
|
|
9acf071ae4
|
fix: make dev->main promotion push succeed via bypass-actor App token (#49)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
The nightly promote fast-forwards main to a fully-green dev HEAD, but the push (as github-actions[bot]) is rejected by the main ruleset: it requires PRs + a status check and the default token is not a bypass actor, so a direct ref push can never land regardless of fast-forwardability. The prior comment claiming protection 'only rejects non-FF' was wrong. Mint a GitHub App installation token (actions/create-github-app-token, SHA-pinned) and push with it; the App must be added to the main ruleset's bypass actors out-of-band. The promoted commit already passed every check on dev (gated by check-dev-green.sh), so re-gating it via a PR on main is redundant. Also fix a gate self-poison: a stale failed 'promote' check-run from a prior run on the same dev HEAD blocked every subsequent gate run (it was excluded only by the current run_id). Exclude prior promote check-runs too, scoped to name=='promote' AND a /actions/runs/ details_url so an external app cannot hide a real failing check by naming it 'promote'; the positive REQUIRED_CHECKS allow-list stays authoritative. Gates: GPT-4.1 cross-review APPROVE (no security regression). Unit-tested: stale promote ignored -> PASS; real failure / external promote / missing required check -> BLOCK. shellcheck clean (also fixed a pre-existing SC2295 on the run_id match). |
||
|
|
3af0bd5e16
|
ci: align workflows with Sea Haven CI/CD handbook (#29)
* Align workflows with Sea Haven CI/CD handbook Bring the workflow suite in line with the handbook: bump actions/checkout to v7 (Node 24 runtime, already standardized), kebab-case the two snake_case workflow filenames, and add the org-standard Labeler caller and Dependency Review gate so vulnerable or disallowed-license deps and unlabeled PRs are caught automatically. File renames only — job/check display names are unchanged, so the promotion gate's REQUIRED_CHECKS and branch-protection required checks are unaffected. Refs: INFRA-115 * Drop Agent prefix from CI workflow + job names The handbook names workflows for what they do (CI, Deploy, Labeler), not the component they run, matching .github and afterhours-shift-manager. Rename the suite to CI and its jobs to Lint / Format check / Unit tests, and keep the promotion gate's REQUIRED_CHECKS in sync. Refs: INFRA-115 --------- Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com> |
||
|
|
3224d7abf4
|
ci: gate dev→prod promotion on green checks + add rollback safety net (#28)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
Agent CI / Agent lint (push) Waiting to run
Agent CI / Agent format check (push) Waiting to run
Agent CI / Agent unit tests (push) Waiting to run
Agent CI / Playwright E2E (push) Waiting to run
T20 CD safety nets. Two gaps closed before the first real prod deploy: 1. Promotion gate. promote_dev_to_prod.yml previously fast-forwarded dev→main unconditionally. It now hard-gates on check-dev-green.sh: every check-run on the dev HEAD must be completed+passing AND the Agent CI suite (lint/format/unit/E2E) must be present+success, or the promotion blocks (fails safe on a missing/renamed check). The promote run excludes its OWN check-run by run-id (unforgeable), never by the mutable name "promote", so a colliding red check cannot hide. Fields are read with a 0x1F separator so an empty conclusion (every in_progress check) cannot shift columns. ci.yml now also runs on push:dev so dev HEAD actually carries that signal (a PR check alone can be admin-merged past). 2. Rollback + last-good. publish-and-deploy.sh advances releases/last-good/ only after a successful roll (deploy.sh gates on `systemctl is-active`), and makes releases/latest/ transactional — reverting to the prior release if the roll fails so a replaced box never self-deploys a broken release. New rollback.yml + rollback.sh re-point latest at last-good (or an explicit sha) and re-fire the deploy; prod is gated by the `prod` Environment approval, same as a deploy. The shared fire/wait/aggregate-gate logic is factored into roll-box.sh (used by both forward and backward rolls). Least-privilege: drop the unused s3:DeleteObject from the app deploy role — publish/rollback/deploy only Get+Put (S3-to-S3 copy), and the rollback fallback now depends on immutable release history staying intact. Lifecycle expiry (not CI) handles old-version cleanup. Gate logic unit-tested (7 cases + jq round-trip). IAM change + release-safety control cross-reviewed by GPT-4.1: APPROVE, no blocks. Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi |
||
|
|
9e2b215f08
|
fix(ci): avoid tar|grep -q SIGPIPE false-failure in package-artifacts (#20)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
`tar -tzf app.tar.gz | grep -qx` under set -o pipefail fails the pipeline when grep -q matches and exits early (SIGPIPEs tar -> 'write error' -> non-zero), a false 'missing agent/server.py'. List the archive once into a var, then grep the var. Same fix for the secret-guard pipe (which was also silently broken). |
||
|
|
404b3f6f75
|
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18)
* feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3)
Packer-build the custom base image and repoint AppService off the AL2023
placeholder onto it.
deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real
`packer build` (the config had only ever been `packer validate`'d at T8):
- the file provisioner failed uploading the templates dir ('scp: …: Is a
directory') — a trailing-slash contents-upload needs the dest dir to exist;
added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the
dest trailing slash.
- the shell provisioner's custom execute_command omitted {{ .Vars }}, so the
environment_vars never reached provision.sh (which runs under set -u and
aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}.
infra:
- ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26
from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by
exact id via MachineImage.genericLinux (offline, deterministic). Dropped the
now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement
discipline docs.
- app-service.ts: machineImage → bakedOpenSweArm64().
- open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard).
- cdk.context.json → {} (AMI is a static id pin; no context lookups remain).
- README: Baked AMI + EBS-replacement-discipline section.
tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI.
NOTE: held — do NOT merge until the open-swe-dev secret values are populated
(put-config.sh). The infra CD is live, so merging this to dev auto-deploys
OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts →
unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14).
* fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021
Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects
non-ASCII in the AMI Description attribute, so packer registered then
DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error.
Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available).
Re-pinned BAKED_OPEN_SWE_AMI_ID.
* fix(deploy): GitHub App + Slack required for prod only, not dev
Per the migration decision: do NOT create/duplicate a separate dev GitHub App or
Slack app — only prod owns the single shared app. So fetch-config.sh no longer
hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/
CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only
block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET.
Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active
provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env
(boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is
unchanged (prod still requires everything).
* feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19)
Make the dev/prod box deployable end-to-end: a real artifact pipeline and a
re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy.
Infra (T7):
- assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access,
SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent +
abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput.
- app-service.ts: open-swe-<env>-deploy SSM document that runs the baked
/opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the
baked open-swe-base-arm64 AMI (folds in the held #16).
IAM (app deploy role — cross-review gated):
- github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to
open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic
AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document
is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox.
Boot/deploy (T19):
- deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull
app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv
at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx.
- user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB
target is healthy even before the first release); deploy.sh is base64-rendered
by CDK into user-data (a normal reviewable repo file, not a heredoc) and the
first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy).
CI (T7+T19):
- build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite ->
ui/.output/public -> spa.tar.gz), package the Python source via git archive
(app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via
the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy.
push dev -> dev (auto); push main -> prod (env "prod" approval gate).
Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK,
deploy.sh base64 round-trips exact.
* harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard
Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one
confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01,
the account-wide CDK cfn-exec residual already documented in config.ts; recorded in
.security-review/suppressions.json with justification + flagged for the per-env
bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified;
these are the cheap defense-in-depth fixes worth taking regardless:
- deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root
never honors an archive's uid/mode → no setuid/foreign-owned file can land); and
treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy
failure (set -e stays loud once a release exists).
- publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount),
not CommandInvocations[0], so a partial failure across the brief 2-instance
replacement window can't be reported as success.
- instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app
role's write scope) instead of the whole bucket.
- package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz
(defense in depth over .gitignore; scoped to data extensions so *_credentials.py
source is not a false positive — verified against the real tree).
Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the
CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable
releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release
to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz).
shellcheck/tsc/jest(16) clean; both stacks synth offline.
* fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard
The instance-SG GroupDescription + ingress/egress rule descriptions carried an
em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects
non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"),
so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing
from #14; same class as the AMI-description ASCII bug.)
- app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the
ingress/egress rule descriptions, and the Route53 comment.
- test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup
GroupDescription + rule descriptions are pure ASCII, so this fails the build
instead of a deploy next time.
jest 18/18; tsc clean.
* fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`)
The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule*
descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`,
which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the
ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII
only" to the exact EC2 allowed charset so it catches `>` (and `<`) too.
jest 18/18; tsc clean.
* fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit
The base64 deploy.sh embedded in user-data pushed the encoded boot script to
27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with
"Encoded User data is limited to 25600 bytes". Strip full-line comments + blank
lines from deploy.sh before base64-embedding it (repo file keeps comments; only
the on-box copy is minified; the script is opaque base64 so user-data heredocs are
unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a
synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded.
jest 19/19; minified deploy.sh passes bash -n + shellcheck.
|