mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 13:53:15 +00:00
18 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
860ce93ad7
|
docs: record final managed-deployment topology in MIGRATION.md (#77)
Add an authoritative "Final, verified topology" section (1a) and reconcile the phased plan with the executed end-state, superseding the spike-era values throughout. Captures: two LangGraph Cloud deployments (dev/prod, one LangSmith workspace) with their final URL hashes; one Vercel project open-swe-prod with production + custom dev environments and per-env LANGGRAPH_BACKEND_URL; the Nitro routeRules proxy mechanism (PR #76, superseding the vercel.json rewrite and PR #75's build-vercel-output.mjs); two dev/prod-isolated GitHub Apps; per-deployment user stores; the Bedrock IAM users; and the seven hard-won operational gotchas. Refs: #65 #74 #76 |
||
|
|
8441bbb2d8
|
docs(migration): record the two Bedrock IAM users + scoped invoke policy (#74)
Document the live execution of MIGRATION.md §10.1 (Bedrock auth via static keys): the customer-managed least-privilege policy open-swe-bedrock-invoke and the open-swe-dev-bedrock / open-swe-prod-bedrock IAM users. Captures the static-key deviation rationale (managed LangGraph Cloud cannot assume a role) and that both mandatory gates (GPT-4.1 IAM cross-review, /sh-security-review) passed with no critical/high. |
||
|
|
a30ce2ab40
|
feat: managed LangGraph Cloud + Vercel migration (PR2 — code fixes + docs) (#65)
* fix(dashboard): managed-cloud OAuth hardening + admin user-mapping endpoint Prepare the dashboard backend for the managed LangGraph Cloud + Vercel runtime, where the API is HTTPS and cross-site from the UI. - OAuth redirect_uri (#2): coerce a schemeless DASHBOARD_API_BASE_URL to https:// in _api_base_url() so GitHub stops rejecting login with "redirect_uri not associated with this application". _cookie_security() now treats a schemeless (managed) value as Secure; SameSite=None too, consistent with the coerced scheme. - OAuth state cookie (#3): document that osw_oauth_state is host-only by design (a Domain cookie is unsafe across *.vercel.app, a public suffix), so login must always start on the stable alias to avoid "oauth state mismatch". Operational contract; no behavioral change. - Admin user mappings (#4): add POST /admin/user-mappings so an admin can set the github_login -> work_email link from the dashboard instead of a raw Store write. New "admin" MappingSource provenance value. * fix(webapp): refresh user-mapping cache on GitHub webhook paths On managed LangGraph Cloud the backend runs multiple replicas, so the per-process GitHub<->work-email mapping cache can be stale on the replica handling a webhook (a mapping created on another replica is invisible until refresh). process_github_pr_comment and process_github_issue now refresh the cache from the durable Store before resolving the author's email, matching the existing Slack mention path (process_slack_mention). * perf(webapp): defer deepagents import to speed custom-app cold start The custom FastAPI app (agent.webapp:app, the langgraph.json http.app) pulled deepagents -> langchain_anthropic -> anthropic into its import graph via dashboard.routes, only to build skill/chat seed files. Defer those create_file_data imports into the functions that use them. Removes deepagents/langchain_anthropic/anthropic from app import entirely and roughly halves module-import wall time (~0.6-0.8s -> ~0.35s warm; larger cold-start saving since native anthropic init is skipped). Behavior identical. (reviewer_diff already imports deepagents under TYPE_CHECKING.) * feat(ui): set work_email user mappings from the admin dashboard Add an "Add / update" form to the admin User mappings section and the adminUpsertUserMapping API client method, wiring the new POST /admin/user-mappings endpoint. Admins can now create or update a github_login -> work_email mapping directly instead of waiting for the user to self-connect Slack. * docs: document managed LangGraph Cloud + Vercel deployment - INSTALLATION §10: add the managed production env triad (LANGGRAPH_URL, DASHBOARD_BASE_URL + DASHBOARD_API_BASE_URL with https://, empty VITE_DASHBOARD_API_BASE_URL for same-origin), the stable-alias login and vercel.json stable-deployment-URL requirements, multi-replica cache note, plus redirect_uri-scheme and oauth-state-mismatch troubleshooting. Refresh the langgraph.json snippet to all six graphs. - README: reframe deployment around the managed migration; link the plan. - deploy/MIGRATION.md: import the self-hosted -> managed migration plan. |
||
|
|
7f60324f0c
|
chore: decommission self-hosted AWS LangGraph stack (#64)
* chore: decommission self-hosted AWS LangGraph stack Removes the now-dead self-host IaC and AWS-only CI/CD after destroying the dev + prod CloudFormation stacks (open-swe-dev, open-swe-prod, open-swe-iam, and the dev-exclusive CDKToolkit-oswedev bootstrap) in account 328440206208, us-east-1. The deployment is now managed (LangGraph Cloud + Vercel). - remove infra/ (CDK app: app + IAM stacks, constructs, aspects, tests) - remove deploy/ami (Packer AMI build) and deploy/seahaven (boot/config scripts, DEPLOYMENT/ROTATION runbooks) - remove AWS-only workflows: cd-infra, ci-infra, build-artifacts, rollback - README: rewrite the Deployment section to the managed LangGraph Cloud + Vercel view; drop dead links to infra/ and deploy/seahaven Preserved: the shared default CDKToolkit bootstrap and promote-dev-to-prod.yml. The RETAIN'd Secrets Manager shells and open-swe-<env>-assets S3 buckets survive cdk destroy by design (orphaned) and need a separate deliberate cleanup. * chore: clean up dangling references left by the AWS decommission Folds in the FIX-level items from the #64 review gates (GPT-4.1 cross-review + /sh-security-review), none of which were blockers: - delete orphaned .github/scripts/{package-artifacts,publish-and-deploy,roll-box, rollback}.sh — their only callers were the removed AWS deploy workflows - drop the deleted /infra dir from dependabot.yml npm directories (was producing a recurring Dependabot config error) - remove the stale OSWE-IAC-SECRETS-LIST-01 suppression (referenced the deleted infra/lib/constructs/instance-role.ts) - repoint the README promotion link to promote-to-main.yml (renamed in #63) The promote-dev-to-prod.yml comment in check-dev-green.sh is intentionally left to #63, which rewrites that same line. |
||
|
|
a4ed19ba61
|
feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62)
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)
Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.
- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module
* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8
The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.
Map profile effort to additional_model_request_fields:
{thinking: {type: adaptive, display: summarized},
output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.
* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids
Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
(set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.
Surfaced by the cross-family review + verified against deploy/.
* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip
From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
error code only, so the role ARN + account id in the raw botocore message never
reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
(Converse emits reasoning_content, not thinking) so the middleware is not a no-op
on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)
* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids
Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
each routed region (us-east-1/2, us-west-2). The model runs in the server process
on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
seed_store.sh's default via pick precedence, so the seed-script fix alone was
insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
ids to the Bedrock id (config.toml's model_id was an active, now-broken value).
AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.
* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)
Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
(eval judge only — Bedrock builder/reviewer auth via the host IAM role).
REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
|
||
|
|
9444fd7677
|
docs: document live AWS prod deploy for Open SWE (#53)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Prod went live 2026-06-29 on self-hosted AWS EC2 behind the shared seahaven-com ALB, superseding the on-prem VM model the runbook described. - Rewrite deploy/seahaven/DEPLOYMENT.md as the canonical end-to-end runbook: infra CD (CDK stacks + OIDC roles + prod approval gate), config seeding (put-config.sh, the 13 boot-required prod vars, fetch-config fail-fast), app artifact deploy (S3 + SSM roll + is-active gate), promotion/rollback, live prod facts, and a RETAIN secret-shell troubleshooting entry that cross-references infra/README.md. - Correct retired *.seahavenind.com hosts to *.seahaven.com throughout and document the live GitHub/Slack/Linear webhook + OAuth endpoints. - Add a concise Deployment section to README pointing at the runbook. - Fix the stale host in the retired on-prem nginx/openswe.conf and mark it superseded by the AMI template. |
||
|
|
faae9a685b
|
Scope secrets fetch to --secret-id-list; drop ListSecrets grant (#48)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Switch fetch-config.sh from a name-prefix batch-get-secret-value --filters scan to an explicit --secret-id-list (the 28 SECRET_VARS, chunked at the 20/call cap). An id-list batch authorizes per-secret ARN, so the instance role's BatchGetSecretValue moves from Resource:* to the open-swe-<env>/* prefix and the account-wide ListSecrets grant is dropped entirely. The box can no longer enumerate secret names account-wide; cross-env value isolation is unchanged (GetSecretValue was already prefix-scoped). Resolves OSWE-IAC-SECRETS-LIST-01. Also capture each chunk response into a variable and consume the producer via command substitution so a failed AWS call aborts under set -e instead of being swallowed by process substitution and misreported as a missing required var. Refs: OSWE-IAC-SECRETS-LIST-01 |
||
|
|
5e30bc6be2
|
Point dev agent at seahaven-open-swe-dev org + un-pin the owner guard (#27)
Dev now runs against the dedicated seahaven-open-swe-dev org (repo openswe-dev-sandbox), isolated from the real Sea Haven org. Two changes: - config-store.ts: iacManagedSsm repo targeting is now per-env — dev = seahaven-open-swe-dev/openswe-dev-sandbox, prod stays Sea-Haven-Industries/open-swe-pilot. ALLOWED_GITHUB_ORGS tracks the env owner. Adds a dev-only SEED_USER_MAPPINGS param so the triggering GitHub login (amoussa1229) resolves and @openswe comments aren't skipped. - fetch-config.sh: replace the unconditional hard-pin to Sea-Haven-Industries with an owner GUARD that HONORS the configured owner (OPENSWE_REPO_OWNER override, else the SSM value) but forces a per-env safe org when the normalized owner is blank or the upstream langchain-ai. Normalization (lowercase, strip whitespace, first path segment, drop dots) catches langchain-ai/<repo>, langchain-ai., and case variants without over-blocking legit orgs (e.g. langchain-ai-fork). Fallback org is per-env so dev can't fall back into the real org. Reviews: GPT-4.1 cross-review APPROVE (round 1 found a path/dot bypass -> hardened, round 2 clean); /sh-security-review authz one LOW (env-invariant fallback) -> fixed. Positive org allowlist still enforced by the app via ALLOWED_GITHUB_ORGS. |
||
|
|
f2633f9bd0
|
fix: don't crash-loop the box when no user mapping is configured (#25)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
seed_store.sh (ExecStartPost) exited 1 when it couldn't resolve a user_mappings entry, which — with Type=simple — fails the whole unit and crash-loops the box. This contradicts the script's own OSWE-SEED-03 precedent (the server-not-ready path exits 0 specifically to avoid a restart loop). A missing user mapping is a seeding gap, not an unhealthy server: the @openswe trigger just won't resolve a commenter, which a deployment-validation env (dev) does not need. Make it non-fatal — warn and skip the user_mappings seed; team_settings is still seeded and the service starts. The empty MAPPINGS array makes the seed loop a no-op. Set SEED_USER_MAPPINGS / CONFIGURED_ADMINS / OPENSWE_OWNER_LOGIN+EMAIL to seed it. Verified on the dev box: service active, langgraph bound to 127.0.0.1:2024 (loopback), /ok 200, /healthz 200. |
||
|
|
a2877d46ff
|
fix: paginate batch-get-secret-value in fetch-config (secrets dropped past page 1) (#24)
fetch-config materialized only the FIRST page of Secrets Manager results: the AWS CLI does NOT auto-paginate batch-get-secret-value (the script's comment claiming it does was wrong; --no-cli-pager only disables the output pager, not API pagination). With 28 secret shells under open-swe-<env>/ the first page returned ~10 items, stranding the rest on later pages. Required secrets that landed past page 1 (DASHBOARD_JWT_SECRET, TOKEN_ENCRYPTION_KEY, LANGSMITH_API_KEY_PROD) were silently dropped, tripping the FAIL-FAST 'missing required var' guard and crash-looping open-swe.service. Follow NextToken across pages (new batch_get_secrets_tsv helper). Verified against the live open-swe-dev secrets: now loads all 5 populated secrets (was 2). Same per-record base64 / exact-prefix / accept_var / emit_var hardening — only the page loop is new. |
||
|
|
5fa132205b
|
fix(deploy): use %%...%% for CDK user-data tokens (don't collide with @@ sed) (#21)
The systemd unit booted with a literal `@@OPENSWE_ENV@@` (fetch-config.sh got the token, not "dev" -> exit 2 -> crash-loop) because user-data.sh is double-templated: CDK substitutes @@tokens@@ AND user-data seds @@tokens@@ into the baked systemd/nginx files. CDK's `.replace(/@@OPENSWE_ENV@@/g, "dev")` clobbered the sed PATTERN (`s|@@OPENSWE_ENV@@|...|` -> `s|dev|...|`, a no-op), so the unit's token never got replaced. Same collision hit @@SERVER_NAME@@ (masked by nginx default_server). Fix: CDK tokens move to a DISTINCT delimiter %%...%% (rendered in app-service.ts); the @@...@@ tokens stay for the baked-template seds. No AMI rebuild (templates unchanged). Add a guard test asserting no unresolved %%CDK%% token survives in the synthesized user-data. Also add .github/scripts/** to build-artifacts paths so script-only changes trigger a publish. jest 20/20; tsc + shellcheck clean; rendered user-data: OPENSWE_ENV="dev", SERVER_NAME="openswe-dev.seahaven.com", @@ sed patterns preserved, 16872 B. |
||
|
|
404b3f6f75
|
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18)
* feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3)
Packer-build the custom base image and repoint AppService off the AL2023
placeholder onto it.
deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real
`packer build` (the config had only ever been `packer validate`'d at T8):
- the file provisioner failed uploading the templates dir ('scp: …: Is a
directory') — a trailing-slash contents-upload needs the dest dir to exist;
added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the
dest trailing slash.
- the shell provisioner's custom execute_command omitted {{ .Vars }}, so the
environment_vars never reached provision.sh (which runs under set -u and
aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}.
infra:
- ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26
from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by
exact id via MachineImage.genericLinux (offline, deterministic). Dropped the
now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement
discipline docs.
- app-service.ts: machineImage → bakedOpenSweArm64().
- open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard).
- cdk.context.json → {} (AMI is a static id pin; no context lookups remain).
- README: Baked AMI + EBS-replacement-discipline section.
tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI.
NOTE: held — do NOT merge until the open-swe-dev secret values are populated
(put-config.sh). The infra CD is live, so merging this to dev auto-deploys
OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts →
unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14).
* fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021
Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects
non-ASCII in the AMI Description attribute, so packer registered then
DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error.
Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available).
Re-pinned BAKED_OPEN_SWE_AMI_ID.
* fix(deploy): GitHub App + Slack required for prod only, not dev
Per the migration decision: do NOT create/duplicate a separate dev GitHub App or
Slack app — only prod owns the single shared app. So fetch-config.sh no longer
hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/
CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only
block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET.
Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active
provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env
(boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is
unchanged (prod still requires everything).
* feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19)
Make the dev/prod box deployable end-to-end: a real artifact pipeline and a
re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy.
Infra (T7):
- assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access,
SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent +
abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput.
- app-service.ts: open-swe-<env>-deploy SSM document that runs the baked
/opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the
baked open-swe-base-arm64 AMI (folds in the held #16).
IAM (app deploy role — cross-review gated):
- github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to
open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic
AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document
is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox.
Boot/deploy (T19):
- deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull
app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv
at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx.
- user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB
target is healthy even before the first release); deploy.sh is base64-rendered
by CDK into user-data (a normal reviewable repo file, not a heredoc) and the
first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy).
CI (T7+T19):
- build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite ->
ui/.output/public -> spa.tar.gz), package the Python source via git archive
(app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via
the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy.
push dev -> dev (auto); push main -> prod (env "prod" approval gate).
Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK,
deploy.sh base64 round-trips exact.
* harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard
Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one
confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01,
the account-wide CDK cfn-exec residual already documented in config.ts; recorded in
.security-review/suppressions.json with justification + flagged for the per-env
bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified;
these are the cheap defense-in-depth fixes worth taking regardless:
- deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root
never honors an archive's uid/mode → no setuid/foreign-owned file can land); and
treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy
failure (set -e stays loud once a release exists).
- publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount),
not CommandInvocations[0], so a partial failure across the brief 2-instance
replacement window can't be reported as success.
- instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app
role's write scope) instead of the whole bucket.
- package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz
(defense in depth over .gitignore; scoped to data extensions so *_credentials.py
source is not a false positive — verified against the real tree).
Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the
CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable
releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release
to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz).
shellcheck/tsc/jest(16) clean; both stacks synth offline.
* fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard
The instance-SG GroupDescription + ingress/egress rule descriptions carried an
em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects
non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"),
so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing
from #14; same class as the AMI-description ASCII bug.)
- app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the
ingress/egress rule descriptions, and the Route53 comment.
- test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup
GroupDescription + rule descriptions are pure ASCII, so this fails the build
instead of a deploy next time.
jest 18/18; tsc clean.
* fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`)
The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule*
descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`,
which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the
ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII
only" to the exact EC2 allowed charset so it catches `>` (and `<`) too.
jest 18/18; tsc clean.
* fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit
The base64 deploy.sh embedded in user-data pushed the encoded boot script to
27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with
"Encoded User data is limited to 25600 bytes". Strip full-line comments + blank
lines from deploy.sh before base64-embedding it (repo file keeps comments; only
the on-box copy is minified; the script is opaque base64 so user-data heredocs are
unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a
synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded.
jest 19/19; minified deploy.sh passes bash -n + shellcheck.
|
||
|
|
fcbdfb67aa
|
feat: open-swe dev/prod compute + ALB ingress (T12) (#14)
AppService construct wires the per-env EC2 box and its internet path. The
seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the
on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and
never owned/mutated; open-swe only ADDS its own resources.
Per env (open-swe-stack.ts → AppService):
- ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a,
in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination
(no RETAIN volume — replacement-tolerant; see ami-cache.ts).
userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh.
- Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress
via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress
(the imported, on-prem-owned SG is never mutated).
- Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays
loopback). Health check GET /healthz.
- Two rules on the imported :443 listener, both → the TG:
* webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/*
* site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api)
Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5)
or it would steal every webhook — first-match-by-ascending-priority.
- Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB.
- 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config).
Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed
medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large
GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on
/dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene);
XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed
critical/high.
Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked
open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean.
Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
|
||
|
|
aef6b26912
|
feat: Secrets Manager + SSM config store for open-swe (T11) (#10)
Create the per-env config surface the EC2 box reads at boot via
fetch-config.sh / seed_store.sh:
- ConfigStore construct (infra/lib/constructs/config-store.ts):
- 28 value-LESS Secrets Manager shells open-swe-<env>/<VAR>
(RemovalPolicy.RETAIN, no SecretString/generateSecretString — real
values are set out-of-band by put-config.sh, never in IaC/state).
- 8 IaC-managed SSM params /open-swe-<env>/<VAR> with real,
stable/derivable values (SANDBOX_TYPE, DEFAULT_REPO_OWNER/NAME,
ALLOWED_GITHUB_ORGS, DASHBOARD_*_URL/ORIGINS, LLM_MODEL_ID).
- OUT_OF_BAND_SSM documents the ~30 params CDK intentionally does NOT
own (operationally-variable / env-specific-unknown).
- Wire ConfigStore into OpenSwe<Env>Stack.
- KebabNamingAspect: exempt Secrets Manager + SSM names, which carry the
literal UPPER_SNAKE env-var segment (open-swe-dev/DASHBOARD_JWT_SECRET).
- deploy/seahaven/put-config.sh: out-of-band populator (placeholders only,
OPENSWE_PUT_<VAR> env indirection; no real values committed).
Synth-only; not deployed. Instance-role read grants on open-swe-<env>/*
already exist from T6 — no IAM/trust changes here.
|
||
|
|
434d0ad80c
|
feat(deploy): AWS-sourced fetch-config + seed_store + rotation docs (PR#7) (#8)
PR#7 of the AWS migration. deploy/seahaven/: fetch-config.sh materializes a service-user-owned 0600 tmpfs .env from Secrets Manager + SSM (fail-fast); seed_store.sh reseeds the in-memory store; ROTATION.md. Incorporates T5 /sh-security-review fixes: - seed_store no longer bash-sources the .env (closes the SH-INJ-001 RCE); uses a non-eval reader, jq --arg JSON bodies, and a loopback-pinned BASE. - fetch-config: .env owned by the openswe service user (app no longer runs as root); DEFAULT_REPO_OWNER hard-pinned; key-identifier validation + flat-namespace collision detection; dropped SSM --recursive. shellcheck + bash -n clean. |
||
|
|
ce9e5fb51b
|
feat(deploy): EC2 AMI recipe (Packer) + cloud-init/user-data (PR#2) (#7)
PR#2 of the AWS migration. deploy/ami/: Packer template (Ubuntu 24.04 arm64, uv+py3.12, nginx, awscli v2, CW agent; no swapfile), provisioning-only user-data (userDataCausesReplacement rationale), systemd unit + nginx + CW templates. Incorporates T5 /sh-security-review fixes: langgraph binds 127.0.0.1 (not 0.0.0.0); nginx is the sole ingress proxying only /dashboard/api/ + /webhooks/; ExecStartPre runs fetch-config as root (+) and passes the env arg; the app runs as the unprivileged openswe user reading an openswe-owned 0600 .env. packer validate clean. |
||
|
|
35bc9e22f4
|
ci: promote dev → prod via fast-forward (not force-push from main) (#2)
* ci: promote dev to prod via fast-forward instead of force-pushing main Repoints the daily promotion workflow to feed prod from the dev integration trunk (prod <- dev) rather than force-pushing the upstream mirror (main). Uses a fast-forward push so a release can never rewrite prod history. * fix: sort imports in aegra_entry.py to unblock dev CI The Sea Haven deployment commit left aegra_entry.py with an unsorted import block (ruff I001), which fails the required 'Agent lint' check and blocks all merges into dev. Apply the import-sort autofix. |
||
| cf57ba7aa5 |
feat(deploy): add Sea Haven self-hosted deployment capture
Captures the stock-LangGraph deployment of this fork at Sea Haven: - systemd/open-swe.service: langgraph dev (:2024) + store seed ExecStartPost - seed_store.sh: re-seeds team_settings + user_mappings (in-memory store resets on restart); env-parameterized, no secrets - nginx/openswe.conf: dashboard SPA + scoped /dashboard/api proxy (security boundary; agent API not exposed) - aegra/: deferred self-hosted-runtime alternative (not active on stock) - DEPLOYMENT.md: full runbook (models, build, ingress, OAuth callback) Secrets and internal infra identifiers are intentionally excluded (public fork); real values live in private IT docs. |