* Add 10 Fireworks models to selectable set
Surface additional Fireworks-served models in the profile editor so
they can be chosen per-thread, per-profile, and as team defaults. Each
entry carries its recommended efforts and image support; only MiniMax
M3 is multimodal.
Refs: #78
* Suppress reasoning_effort on non-reasoning models
Instruct-only Fireworks ids (kimi-k2-instruct-0905, mistral-large-3-fp8,
qwen3-30b-a3b-instruct-2507) don't reason, so sending reasoning_effort
either 400s (unusable at default effort) or is a silent no-op. Add a
per-model reasoning flag (default True) and omit the param entirely for
ids marked non-reasoning.
Refs: #78
* Gate out 7 undeployed Fireworks models
Account serverless probe returned 404 for 7 of the 10 proposed ids, so
only minimax-m3, gpt-oss-120b, and deepseek-v4-flash are callable. Keep
those 3 and drop the rest. All 3 survivors are reasoning-capable, so the
per-model reasoning-effort suppression added earlier is no longer needed
and is reverted.
Refs: #78
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Add an authoritative "Final, verified topology" section (1a) and
reconcile the phased plan with the executed end-state, superseding the
spike-era values throughout.
Captures: two LangGraph Cloud deployments (dev/prod, one LangSmith
workspace) with their final URL hashes; one Vercel project open-swe-prod
with production + custom dev environments and per-env LANGGRAPH_BACKEND_URL;
the Nitro routeRules proxy mechanism (PR #76, superseding the vercel.json
rewrite and PR #75's build-vercel-output.mjs); two dev/prod-isolated
GitHub Apps; per-deployment user stores; the Bedrock IAM users; and the
seven hard-won operational gotchas.
Refs: #65#74#76
Document the live execution of MIGRATION.md §10.1 (Bedrock auth via static
keys): the customer-managed least-privilege policy open-swe-bedrock-invoke
and the open-swe-dev-bedrock / open-swe-prod-bedrock IAM users. Captures the
static-key deviation rationale (managed LangGraph Cloud cannot assume a role)
and that both mandatory gates (GPT-4.1 IAM cross-review, /sh-security-review)
passed with no critical/high.
The Vercel deploy 404'd at `/` because PR #75's `vercel-build` script
(`scripts/build-vercel-output.mjs`) ran `rm -rf .vercel/output` and rebuilt
it from `.output/public`. On Vercel CI, Nitro's Vercel preset auto-activates
(from the VERCEL env var) and emits the Build Output API layout to
`.vercel/output` itself during `vite build`; the script then clobbered that
correct output with a static-only config that could not resolve the SPA
`_shell.html` fallback or the server functions, so production returned
`404: NOT_FOUND` even though the build was READY.
Stop fighting Nitro and drive the proxy through it instead:
- Remove `scripts/build-vercel-output.mjs` and the `vercel-build` script;
restore the plain `vite build` for both local and Vercel builds.
- Add an env-driven Nitro `routeRules` proxy in `vite.config.ts`. Nitro's
Vercel preset compiles a plain external-URL `proxy` rule into a CDN-level
rewrite in `.vercel/output/config.json` at build time, reading
LANGGRAPH_BACKEND_URL (the per-project Vercel env var). Proxy, not redirect,
so the osw_session cookie stays first-party (same-origin). The build fails
loudly if LANGGRAPH_BACKEND_URL is missing on Vercel; the rule is omitted
for plain local/off-Vercel builds (which use the E2E_HARNESS mock proxy).
- `vercel.json`: `buildCommand` back to `bun run build`; drop the
`.output/public` outputDirectory so Vercel serves Nitro's `.vercel/output`.
Validated with `LANGGRAPH_BACKEND_URL=... NITRO_PRESET=vercel bun run build`:
the generated `.vercel/output/config.json` contains the
`/dashboard/api/(.*)` -> backend rewrite ahead of `handle: filesystem` and
the SPA catch-all, `_shell.html` and the 699 hashed assets are emitted, and
typecheck passes. Supersedes the broken approach in #75.
Generate the /dashboard/api/* proxy destination at build time from
LANGGRAPH_BACKEND_URL via the Vercel Build Output API, instead of hardcoding
it in ui/vercel.json. This lets the dev and main branches stay byte-identical
(required by the fast-forward-only prod promotion) while each Vercel project
proxies to its own LangGraph backend purely via its own env var — so both the
dev and prod dashboards can be git-linked.
- ui/vercel.json: drop the hardcoded rewrites; buildCommand -> bun run vercel-build
- ui/package.json: add vercel-build = vite build && node scripts/build-vercel-output.mjs
- ui/scripts/build-vercel-output.mjs: emit .vercel/output/config.json with
[ proxy /dashboard/api/* -> $LANGGRAPH_BACKEND_URL, filesystem, SPA fallback ];
throws if LANGGRAPH_BACKEND_URL is unset so a misconfig fails the build.
Keeps the app same-origin (no CORS/auth change); preserves the _shell.html SPA
fallback. Set LANGGRAPH_BACKEND_URL per Vercel project (All Environments).
* fix(dashboard): managed-cloud OAuth hardening + admin user-mapping endpoint
Prepare the dashboard backend for the managed LangGraph Cloud + Vercel
runtime, where the API is HTTPS and cross-site from the UI.
- OAuth redirect_uri (#2): coerce a schemeless DASHBOARD_API_BASE_URL to
https:// in _api_base_url() so GitHub stops rejecting login with
"redirect_uri not associated with this application". _cookie_security()
now treats a schemeless (managed) value as Secure; SameSite=None too,
consistent with the coerced scheme.
- OAuth state cookie (#3): document that osw_oauth_state is host-only by
design (a Domain cookie is unsafe across *.vercel.app, a public suffix),
so login must always start on the stable alias to avoid "oauth state
mismatch". Operational contract; no behavioral change.
- Admin user mappings (#4): add POST /admin/user-mappings so an admin can
set the github_login -> work_email link from the dashboard instead of a
raw Store write. New "admin" MappingSource provenance value.
* fix(webapp): refresh user-mapping cache on GitHub webhook paths
On managed LangGraph Cloud the backend runs multiple replicas, so the
per-process GitHub<->work-email mapping cache can be stale on the replica
handling a webhook (a mapping created on another replica is invisible
until refresh). process_github_pr_comment and process_github_issue now
refresh the cache from the durable Store before resolving the author's
email, matching the existing Slack mention path (process_slack_mention).
* perf(webapp): defer deepagents import to speed custom-app cold start
The custom FastAPI app (agent.webapp:app, the langgraph.json http.app)
pulled deepagents -> langchain_anthropic -> anthropic into its import
graph via dashboard.routes, only to build skill/chat seed files. Defer
those create_file_data imports into the functions that use them. Removes
deepagents/langchain_anthropic/anthropic from app import entirely and
roughly halves module-import wall time (~0.6-0.8s -> ~0.35s warm; larger
cold-start saving since native anthropic init is skipped). Behavior
identical. (reviewer_diff already imports deepagents under TYPE_CHECKING.)
* feat(ui): set work_email user mappings from the admin dashboard
Add an "Add / update" form to the admin User mappings section and the
adminUpsertUserMapping API client method, wiring the new
POST /admin/user-mappings endpoint. Admins can now create or update a
github_login -> work_email mapping directly instead of waiting for the
user to self-connect Slack.
* docs: document managed LangGraph Cloud + Vercel deployment
- INSTALLATION §10: add the managed production env triad (LANGGRAPH_URL,
DASHBOARD_BASE_URL + DASHBOARD_API_BASE_URL with https://, empty
VITE_DASHBOARD_API_BASE_URL for same-origin), the stable-alias login
and vercel.json stable-deployment-URL requirements, multi-replica cache
note, plus redirect_uri-scheme and oauth-state-mismatch troubleshooting.
Refresh the langgraph.json snippet to all six graphs.
- README: reframe deployment around the managed migration; link the plan.
- deploy/MIGRATION.md: import the self-hosted -> managed migration plan.
* chore: decommission self-hosted AWS LangGraph stack
Removes the now-dead self-host IaC and AWS-only CI/CD after destroying the
dev + prod CloudFormation stacks (open-swe-dev, open-swe-prod, open-swe-iam,
and the dev-exclusive CDKToolkit-oswedev bootstrap) in account 328440206208,
us-east-1. The deployment is now managed (LangGraph Cloud + Vercel).
- remove infra/ (CDK app: app + IAM stacks, constructs, aspects, tests)
- remove deploy/ami (Packer AMI build) and deploy/seahaven (boot/config
scripts, DEPLOYMENT/ROTATION runbooks)
- remove AWS-only workflows: cd-infra, ci-infra, build-artifacts, rollback
- README: rewrite the Deployment section to the managed LangGraph Cloud +
Vercel view; drop dead links to infra/ and deploy/seahaven
Preserved: the shared default CDKToolkit bootstrap and promote-dev-to-prod.yml.
The RETAIN'd Secrets Manager shells and open-swe-<env>-assets S3 buckets
survive cdk destroy by design (orphaned) and need a separate deliberate cleanup.
* chore: clean up dangling references left by the AWS decommission
Folds in the FIX-level items from the #64 review gates (GPT-4.1 cross-review +
/sh-security-review), none of which were blockers:
- delete orphaned .github/scripts/{package-artifacts,publish-and-deploy,roll-box,
rollback}.sh — their only callers were the removed AWS deploy workflows
- drop the deleted /infra dir from dependabot.yml npm directories (was producing
a recurring Dependabot config error)
- remove the stale OSWE-IAC-SECRETS-LIST-01 suppression (referenced the deleted
infra/lib/constructs/instance-role.ts)
- repoint the README promotion link to promote-to-main.yml (renamed in #63)
The promote-dev-to-prod.yml comment in check-dev-green.sh is intentionally left
to #63, which rewrites that same line.
* ci: re-home prod promotion into a gated promote-to-main workflow
Migrate the prod-deploy gate to managed LangGraph Cloud (git-connected to
`main`) + Vercel. Under managed, a push to `main` auto-deploys prod, so the
dev -> main fast-forward IS the prod deploy trigger -- the bespoke AWS CD step
is obsolete and already gone from this workflow.
Re-home `promote-dev-to-prod.yml` -> `promote-to-main.yml`:
- gate the promote job on the `prod` GitHub Environment (required reviewer
amoussa1229), restoring the manual prod-approval that the retired AWS CD job
used to carry;
- drop the nightly auto-promote cron -- a scheduled auto-promotion conflicts
with a manual approval gate now that the push deploys prod; promotion is
workflow_dispatch only;
- keep the dev-HEAD-fully-green precondition and the seahaven-promotion App
fast-forward push (sole non-admin bypass actor on `main` ruleset 18238334).
Update the companion check-dev-green.sh filename reference.
* ci: refuse promote-to-main dispatch from any ref other than dev
Defense-in-depth atop the already-pinned `ref: dev` checkout: workflow_dispatch
runs the workflow definition from the launched ref, so reject a non-dev dispatch
before the App token is minted. Surfaced by the GPT-4.1 cross-review of #63.
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)
Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.
- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module
* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8
The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.
Map profile effort to additional_model_request_fields:
{thinking: {type: adaptive, display: summarized},
output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.
* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids
Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
(set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.
Surfaced by the cross-family review + verified against deploy/.
* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip
From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
error code only, so the role ARN + account id in the raw botocore message never
reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
(Converse emits reasoning_content, not thinking) so the middleware is not a no-op
on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)
* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids
Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
each routed region (us-east-1/2, us-west-2). The model runs in the server process
on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
seed_store.sh's default via pick precedence, so the seed-script fix alone was
insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
ids to the Bedrock id (config.toml's model_id was an active, now-broken value).
AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.
* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)
Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
(eval judge only — Bedrock builder/reviewer auth via the host IAM role).
REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
* feat: default Slack/dashboard/schedule PRs + commits to the app identity (#57)
Slack/dashboard/schedule runs now author PRs and run git/gh operations as the
GitHub App seahaven-openswe[bot] by default (matching GitHub-issue runs), so the
self-review 422 is impossible by construction rather than guarded in the prompt.
A profile flag author_prs_as_user restores per-user attribution.
- open_pull_request._resolve_pr_author_token + auth.resolve_github_token: default
to the installation token for these sources; per-user only when opted in.
- authorship: commit identity -> seahaven-openswe[bot] (numeric noreply;
accepted Vercel-resolution risk, documented inline).
- self-trigger safety: INTERNAL_BOT_LOGINS + webapp/reviewer_reconcile/reply
markers recognize seahaven-openswe[bot] (bot-authored events are now ours).
Supersedes the prompt-only guard in #58.
* fix: author commits as the app bot in the default path (SH-IDSPLIT-01)
Security review found the commit identity was NOT actually unified to the bot:
resolve_triggering_user_identity got a 403 from the installation token and fell
back to configurable['github_login'], so commits were still authored as the
triggering user (commit=user, push+PR=bot — a three-way split that missed the
stated goal). Now gate the triggering-user identity resolution on the same
default-bot decision as the token: slack/dashboard/schedule default to the app
bot identity unless author_prs_as_user is set.
* docs(security): record AUTHZ-SLACK-BOT-DEFAULT-001 as an accepted residual (#59)
Single-user deployment; bounded by App-on-pilot + ALLOWED_GITHUB_REPOS lock.
Revisit (add a per-user gate) before expanding users or the App installation.
Adds repo-local suppressions for 5 verified-FP gitleaks findings that block
pushes (forcing --no-verify), all in committed history / docs / CI fixtures:
- .env.ci (dev-only e2e values, intentionally committed)
- .env.example (placeholders)
- .github/ci/fake_github_app_key.pem (throwaway CI test key)
- INSTALLATION.md (example GITHUB_APP_PRIVATE_KEY .env block)
- README.md (prose mis-matched by the generic-api-key heuristic)
Repo-local (not machine-level) so they load in git worktrees too. Also drops
the now-obsolete OSWE-IAC-AUDIT-01 suppression (B-1, fixed in #55).
* fix: isolate dev CDK deploys on their own bootstrap qualifier (B-1)
Dev synthesizes against the oswedev qualifier and the dev infra deploy role
is scoped to cdk-oswedev-* — it can no longer assume the default hnb659fds
bootstrap roles whose admin cfn-exec-role deploys prod, closing the cross-env
escalation (OSWE-IAC-01). Prod stays on the default qualifier.
* test: assert per-env bootstrap qualifier isolation + document (B-1)
* fix: enforce a replay window on Linear webhooks (AUTHZ-001)
verify_linear_signature accepted any correctly-signed body with no freshness
check, so a captured request could be replayed indefinitely. Parse the
signed webhookTimestamp (Unix ms) and reject requests outside a 60s window,
failing closed when the field is missing or malformed — mirroring the Slack
verifier.
* fix: stop leaking upstream auth-error bodies into user comments
get_github_token_for_user folded the raw upstream response text into the
error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the
full body server-side only and return a generic "GitHub auth failed (status
<code>)". Also document the accepted shared-installation-token blast radius on
the bot-token-only path (AUTHZ-003).
* fix: bind sandbox and token caches to repo to prevent thread-id collision
A PR head-branch name is attacker-controllable and get_thread_id_from_branch
derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01).
The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on
thread_id alone, and a cached sandbox was reused after only an echo-ping, so a
different repo's webhook could bind to another thread's sandbox or token.
Without changing the persistent thread-id scheme:
- Persist the bound repo (owner/name) in thread metadata on sandbox creation and
refuse to reuse a sandbox whose bound repo does not match the current event
(SandboxRepoMismatchError); the in-memory proxy also carries the binding.
- Bind the GitHub-token cache entries to their repo and evict on a cross-repo
read so a colliding thread_id cannot be served another repo's token.
- Thread repo through the reviewer and the webhook token resolvers.
* fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04)
The instance role and the GitHub deploy app role granted s3:ListBucket on the
whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts)
only ever lists under releases/, so add a StringLike s3:prefix=releases/*
condition. GetBucketLocation has no s3:prefix in its request context, so it
moves to its own unconditioned statement. Also document the accepted F-2
cross-env existence-oracle residual on BatchGetSecretValue.
* chore: suppress test-fixture credential false positive; document AUTHZ-002
Add a machine-level suppression for the fake Datadog key in the
test_team_credentials encryption-roundtrip fixture (CWE-798, not a real
credential). Clarify that the within-org thread-write path is intentional by
design (AUTHZ-002) — comment only, no behavior change.
* fix: casefold repo-binding keys to avoid spurious cross-repo mismatch
GitHub owner/name are case-insensitive. Casefold the owner/name key on both the
write (binding) and read (compare) sides — repo_cache_key and the metadata
bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate
same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2).
* fix: stop leaking upstream auth body in unexpected-result branch
The 2xx-but-missing-token/url branch echoed the parsed upstream response body
into the user-facing error. Return a generic message and log response_data
server-side only, mirroring the existing HTTPStatusError fix (Gap 4).
* fix: fail closed for unbound-legacy sandboxes and catch repo mismatch
Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no
recorded bound_repo (a pre-binding legacy thread, post-deploy) previously
reconnected-and-served the sandbox to the current repo, then rebound it. Now
fail closed: drop the stale id and recreate a fresh sandbox bound to this repo,
logging a reconnect-with-missing-binding event. A sandbox is never served to a
repo unless its binding is known and matches; new threads bind on first run
unchanged.
Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints,
log it for alarming, and surface a clean sanitized error instead of letting an
opaque deep-stack exception crash-loop the worker.
* chore: suppress test-fixture credential false positive in token-TTL tests
Add a machine-level suppression for the fake "ghp_secret" GitHub token used by
the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and
not a valid PAT; scoped to the unit test only.
Prod went live 2026-06-29 on self-hosted AWS EC2 behind the shared
seahaven-com ALB, superseding the on-prem VM model the runbook described.
- Rewrite deploy/seahaven/DEPLOYMENT.md as the canonical end-to-end runbook:
infra CD (CDK stacks + OIDC roles + prod approval gate), config seeding
(put-config.sh, the 13 boot-required prod vars, fetch-config fail-fast),
app artifact deploy (S3 + SSM roll + is-active gate), promotion/rollback,
live prod facts, and a RETAIN secret-shell troubleshooting entry that
cross-references infra/README.md.
- Correct retired *.seahavenind.com hosts to *.seahaven.com throughout and
document the live GitHub/Slack/Linear webhook + OAuth endpoints.
- Add a concise Deployment section to README pointing at the runbook.
- Fix the stale host in the retired on-prem nginx/openswe.conf and mark it
superseded by the AMI template.
RETAIN + a fixed secret name means a failed FIRST create leaves empty
secret shells behind when the stack rolls back. The shells keep the
global `open-swe-<env>/<VAR>` names, so every later create fails with
`AlreadyExists`, and a plain delete-secret keeps the name reserved for
the recovery window rather than freeing it.
Record the trap and the force-delete recovery (only for empty shells)
in the config-store construct and the infra README so the next
teardown/rebuild, secret logical-id change, or new-env stand-up does
not rediscover it the hard way.
Prod's first deploy hit this on 2026-06-29: 28 orphaned shells from an
earlier failed create reserved the names and had to be force-deleted
before the stack would create.
requireImdsv2:true makes CDK auto-create a launch template named from the
construct id ('Instance' -> 'InstanceLaunchTemplate') with no env qualifier,
so OpenSweDevStack and OpenSweProdStack both render
LaunchTemplateName: InstanceLaunchTemplate. dev created it first (the live
dev box runs on it); the prod first-deploy then failed with
InvalidLaunchTemplateName.AlreadyExistsException and the whole stack rolled
back.
Force a per-env LT name (open-swe-<env>-lt) via an aspect (the LT is created
at synth time by the requireImdsv2 handling, not in the constructor), and
rename the instance's launch-template REFERENCE in lockstep so CFN still
resolves it. synth-verified: dev=open-swe-dev-lt, prod=open-swe-prod-lt on
both the LT resource and the instance reference; version GetAtt preserved.
NOTE: deploying this renames dev's LT -> one-time dev box replacement
(stateless; boots from the baked AMI + pulls releases/latest). prod then
creates open-swe-prod-lt cleanly.
#48 (OSWE-IAC-SECRETS-LIST-01) scoped secretsmanager:BatchGetSecretValue
to the open-swe-<env>/* ARN on the theory that an explicit --secret-id-list
batch authorizes per-secret. That is FALSE: BatchGetSecretValue is a
collection action AWS authorizes against the account (*), regardless of
--filters vs --secret-id-list. A prefix-scoped grant AccessDenies the whole
call. The dev box passed right after #48 only because the prior broad grant
had not finished propagating; once it lapsed, fetch-config got AccessDenied
-> loaded 0 secrets -> FAIL-FAST -> open-swe.service crash-loop. Verified on
the live dev box (i-0af4e03e8bf70e6c3): the exact call returned
'not authorized to perform: secretsmanager:BatchGetSecretValue'; restoring
the * grant recovered it.
Move BatchGetSecretValue back to Resource:* (its own statement); keep
GetSecretValue + DescribeSecret prefix-scoped (those gate VALUE access, so
cross-env isolation holds). The surviving win from #48: --secret-id-list
needs no name filter, so ListSecrets stays dropped -> no account-wide name
enumeration. fetch-config.sh is unchanged (--secret-id-list is correct).
The /sh-security-review finding OSWE-IAC-IAM-01 called this out and was
wrongly refuted; the reference_secretsmanager_batch_get memory was wrong.
The nightly promote fast-forwards main to a fully-green dev HEAD, but the
push (as github-actions[bot]) is rejected by the main ruleset: it requires
PRs + a status check and the default token is not a bypass actor, so a
direct ref push can never land regardless of fast-forwardability. The
prior comment claiming protection 'only rejects non-FF' was wrong.
Mint a GitHub App installation token (actions/create-github-app-token,
SHA-pinned) and push with it; the App must be added to the main ruleset's
bypass actors out-of-band. The promoted commit already passed every check
on dev (gated by check-dev-green.sh), so re-gating it via a PR on main is
redundant.
Also fix a gate self-poison: a stale failed 'promote' check-run from a
prior run on the same dev HEAD blocked every subsequent gate run (it was
excluded only by the current run_id). Exclude prior promote check-runs
too, scoped to name=='promote' AND a /actions/runs/ details_url so an
external app cannot hide a real failing check by naming it 'promote'; the
positive REQUIRED_CHECKS allow-list stays authoritative.
Gates: GPT-4.1 cross-review APPROVE (no security regression). Unit-tested:
stale promote ignored -> PASS; real failure / external promote / missing
required check -> BLOCK. shellcheck clean (also fixed a pre-existing
SC2295 on the run_id match).
Switch fetch-config.sh from a name-prefix batch-get-secret-value
--filters scan to an explicit --secret-id-list (the 28 SECRET_VARS,
chunked at the 20/call cap). An id-list batch authorizes per-secret
ARN, so the instance role's BatchGetSecretValue moves from Resource:*
to the open-swe-<env>/* prefix and the account-wide ListSecrets grant
is dropped entirely. The box can no longer enumerate secret names
account-wide; cross-env value isolation is unchanged (GetSecretValue
was already prefix-scoped). Resolves OSWE-IAC-SECRETS-LIST-01.
Also capture each chunk response into a variable and consume the
producer via command substitution so a failed AWS call aborts under
set -e instead of being swallowed by process substitution and
misreported as a missing required var.
Refs: OSWE-IAC-SECRETS-LIST-01
PyJWT <2.13.0 accepts a public-key JWK as an HMAC secret, letting an
attacker forge HS256 tokens when mixed key families are allowed
(GHSA high-sev Dependabot alert). 2.13.0 rejects the mismatch.
Direct dependency; uv.lock re-resolves 2.12.1 -> 2.13.0.
Hardcoding the Sea Haven no-type-prefix PR title made every PR fail
semantic-PR-title gates (this repo's PR Title Lint, upstream open-swe),
forcing manual retitling. Make the title rule detect a conventional-commit
gate and conform, falling back to the imperative style otherwise. Also add
Closes/Refs issue-linking guidance and the default-branch auto-close caveat.
Refs: #41
Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com>
Switch all ecosystems to a weekly schedule and drop the redundant
patterns: ["*"] selectors. Remove the major group so patch+minor bumps
stay grouped into one PR per ecosystem while each major lands as its own
PR, matching the org standard.
Add npm coverage for the JS/TS surface (root, /infra CDK, /tests/e2e
Playwright, /ui dashboard) and assign updates to amoussa1229. Delete the
stale ui/yarn.lock so Dependabot tracks ui/bun.lock cleanly.
Refs: #34
Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com>
Codify the box-only #4 customizations into Git so the AWS deployment
(which deploys from this repo) actually applies them — previously only
the retired sh-openswe box had them.
- prompt.py: branch names feature|bug|hotfix/<kebab> (optional <KEY->);
imperative PR titles with no conventional-commit type: prefix; PR body
Summary/Validation/Tests/Notes; handbook commit format. Rewrite the
collaboration template from an attribution MANDATE to a PROHIBITION —
no Co-authored-by bot trailer, no "Made by [Open SWE]" footer, no
agent/AI notes on any artifact.
- github_comments.py: add @seahaven-openswe (the deployed App slug) to
the mention triggers.
- authorship.py: remove the now-unused attribution helpers
(build_pr_attribution_footer, add_bot_coauthor_trailer,
add_pr_collaboration_note, PR_ATTRIBUTION_*). Keep OPEN_SWE_BOT_* —
server.py still uses them for the sandbox git identity.
- Flip the attribution unit tests to assert the no-attribution behavior;
drop tests for the removed helpers.
Commits stay authored as the triggering user for now — flipping
authorship to the bot account depends on the Vercel preview-deploy
constraint and is deferred to #11.
Refs: #4#11
Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi
* Align workflows with Sea Haven CI/CD handbook
Bring the workflow suite in line with the handbook: bump
actions/checkout to v7 (Node 24 runtime, already standardized),
kebab-case the two snake_case workflow filenames, and add the
org-standard Labeler caller and Dependency Review gate so vulnerable
or disallowed-license deps and unlabeled PRs are caught automatically.
File renames only — job/check display names are unchanged, so the
promotion gate's REQUIRED_CHECKS and branch-protection required
checks are unaffected.
Refs: INFRA-115
* Drop Agent prefix from CI workflow + job names
The handbook names workflows for what they do (CI, Deploy, Labeler),
not the component they run, matching .github and afterhours-shift-manager.
Rename the suite to CI and its jobs to Lint / Format check / Unit tests,
and keep the promotion gate's REQUIRED_CHECKS in sync.
Refs: INFRA-115
---------
Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com>
T20 CD safety nets. Two gaps closed before the first real prod deploy:
1. Promotion gate. promote_dev_to_prod.yml previously fast-forwarded
dev→main unconditionally. It now hard-gates on check-dev-green.sh:
every check-run on the dev HEAD must be completed+passing AND the
Agent CI suite (lint/format/unit/E2E) must be present+success, or the
promotion blocks (fails safe on a missing/renamed check). The promote
run excludes its OWN check-run by run-id (unforgeable), never by the
mutable name "promote", so a colliding red check cannot hide. Fields
are read with a 0x1F separator so an empty conclusion (every
in_progress check) cannot shift columns. ci.yml now also runs on
push:dev so dev HEAD actually carries that signal (a PR check alone
can be admin-merged past).
2. Rollback + last-good. publish-and-deploy.sh advances
releases/last-good/ only after a successful roll (deploy.sh gates on
`systemctl is-active`), and makes releases/latest/ transactional —
reverting to the prior release if the roll fails so a replaced box
never self-deploys a broken release. New rollback.yml + rollback.sh
re-point latest at last-good (or an explicit sha) and re-fire the
deploy; prod is gated by the `prod` Environment approval, same as a
deploy. The shared fire/wait/aggregate-gate logic is factored into
roll-box.sh (used by both forward and backward rolls).
Least-privilege: drop the unused s3:DeleteObject from the app deploy
role — publish/rollback/deploy only Get+Put (S3-to-S3 copy), and the
rollback fallback now depends on immutable release history staying
intact. Lifecycle expiry (not CI) handles old-version cleanup.
Gate logic unit-tested (7 cases + jq round-trip). IAM change +
release-safety control cross-reviewed by GPT-4.1: APPROVE, no blocks.
Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi
Dev now runs against the dedicated seahaven-open-swe-dev org (repo openswe-dev-sandbox),
isolated from the real Sea Haven org. Two changes:
- config-store.ts: iacManagedSsm repo targeting is now per-env — dev =
seahaven-open-swe-dev/openswe-dev-sandbox, prod stays Sea-Haven-Industries/open-swe-pilot.
ALLOWED_GITHUB_ORGS tracks the env owner. Adds a dev-only SEED_USER_MAPPINGS param so the
triggering GitHub login (amoussa1229) resolves and @openswe comments aren't skipped.
- fetch-config.sh: replace the unconditional hard-pin to Sea-Haven-Industries with an owner
GUARD that HONORS the configured owner (OPENSWE_REPO_OWNER override, else the SSM value) but
forces a per-env safe org when the normalized owner is blank or the upstream langchain-ai.
Normalization (lowercase, strip whitespace, first path segment, drop dots) catches
langchain-ai/<repo>, langchain-ai., and case variants without over-blocking legit orgs
(e.g. langchain-ai-fork). Fallback org is per-env so dev can't fall back into the real org.
Reviews: GPT-4.1 cross-review APPROVE (round 1 found a path/dot bypass -> hardened, round 2 clean);
/sh-security-review authz one LOW (env-invariant fallback) -> fixed. Positive org allowlist still
enforced by the app via ALLOWED_GITHUB_ORGS.
cd-infra reported failure on every successful deploy: the 'Stack outputs' step ran
'aws cloudformation describe-stacks' with the githubdeploy-open-swe-infra-<env> role,
which intentionally lacks cloudformation:DescribeStacks. The cdk deploy itself succeeds
(it reads outputs via the bootstrap cfn-exec role it assumes). Switch to
'cdk deploy --outputs-file cdk-outputs.json' + cat — no extra IAM grant, and the job
goes green on actual deploy success instead of masking real failures behind a red run.
seed_store.sh (ExecStartPost) exited 1 when it couldn't resolve a user_mappings
entry, which — with Type=simple — fails the whole unit and crash-loops the box. This
contradicts the script's own OSWE-SEED-03 precedent (the server-not-ready path exits
0 specifically to avoid a restart loop). A missing user mapping is a seeding gap, not
an unhealthy server: the @openswe trigger just won't resolve a commenter, which a
deployment-validation env (dev) does not need.
Make it non-fatal — warn and skip the user_mappings seed; team_settings is still
seeded and the service starts. The empty MAPPINGS array makes the seed loop a no-op.
Set SEED_USER_MAPPINGS / CONFIGURED_ADMINS / OPENSWE_OWNER_LOGIN+EMAIL to seed it.
Verified on the dev box: service active, langgraph bound to 127.0.0.1:2024 (loopback),
/ok 200, /healthz 200.
fetch-config materialized only the FIRST page of Secrets Manager results: the AWS CLI
does NOT auto-paginate batch-get-secret-value (the script's comment claiming it does
was wrong; --no-cli-pager only disables the output pager, not API pagination). With 28
secret shells under open-swe-<env>/ the first page returned ~10 items, stranding the
rest on later pages. Required secrets that landed past page 1 (DASHBOARD_JWT_SECRET,
TOKEN_ENCRYPTION_KEY, LANGSMITH_API_KEY_PROD) were silently dropped, tripping the
FAIL-FAST 'missing required var' guard and crash-looping open-swe.service.
Follow NextToken across pages (new batch_get_secrets_tsv helper). Verified against the
live open-swe-dev secrets: now loads all 5 populated secrets (was 2). Same per-record
base64 / exact-prefix / accept_var / emit_var hardening — only the page loop is new.
The prior fix scoped secretsmanager:BatchGetSecretValue to the env-prefixed secret
ARN, but the live box still got AccessDenied: batch-get-secret-value invoked WITH a
name --filters is a COLLECTION call that AWS authorizes against * (a per-secret ARN
does not satisfy it). Split the statement:
- GetSecretValue + DescribeSecret stay PREFIX-scoped (secret:open-swe-<env>/*) — this
is what gates which secret VALUES the box can read (checked per-secret in the batch).
- BatchGetSecretValue + ListSecrets move to a * operation-level statement (the filtered
collection call + the list action; neither is resource-scopable for this usage).
VALUE isolation preserved (dev box still cannot read prod secret values); only secret
NAME/metadata enumeration is widened. GPT-4.1 IAM cross-review: BLOCK none, FIX none.
Suppression OSWE-IAC-SECRETS-LIST-01 updated; future hardening (explicit --secret-id-list
to drop both * grants) tracked there.
* fix(infra): grant instance role BatchGetSecretValue + ListSecrets for .env materialization
fetch-config.sh materializes the box's .env via
`secretsmanager batch-get-secret-value --filters Key=name,Values=open-swe-<env>/`,
but the instance role only granted GetSecretValue/DescribeSecret. BatchGetSecretValue
is a distinct IAM action, so the call was AccessDenied and open-swe.service
crash-looped (no .env written -> ExecStartPre exit 1).
- Add secretsmanager:BatchGetSecretValue to the prefix-scoped ReadSecrets statement.
- Add secretsmanager:ListSecrets on * (required by the name-prefix filtered batch
call; the API has no resource-level scoping for the list action — fits the role's
stated exception). Secret VALUES stay prefix-scoped; only names are enumerable.
Reviews: GPT-4.1 IAM cross-review BLOCK=none; /sh-security-review iac-iam one LOW
metadata residual (no critical/high), recorded as OSWE-IAC-SECRETS-LIST-01.
Refs T7/T19 dev bring-up.
* ci: lift Node heap cap for Playwright E2E build (vite OOM)
The E2E job's Playwright globalSetup runs the real `bun run build`, whose vite
bundle exceeds Node's default ~2 GB heap and OOMs (JavaScript heap out of memory) —
the same failure fixed for build-artifacts.yml in #19. Set
NODE_OPTIONS=--max-old-space-size=8192 on the Run E2E step.
The systemd unit booted with a literal `@@OPENSWE_ENV@@` (fetch-config.sh got the
token, not "dev" -> exit 2 -> crash-loop) because user-data.sh is double-templated:
CDK substitutes @@tokens@@ AND user-data seds @@tokens@@ into the baked
systemd/nginx files. CDK's `.replace(/@@OPENSWE_ENV@@/g, "dev")` clobbered the sed
PATTERN (`s|@@OPENSWE_ENV@@|...|` -> `s|dev|...|`, a no-op), so the unit's token
never got replaced. Same collision hit @@SERVER_NAME@@ (masked by nginx
default_server).
Fix: CDK tokens move to a DISTINCT delimiter %%...%% (rendered in app-service.ts);
the @@...@@ tokens stay for the baked-template seds. No AMI rebuild (templates
unchanged). Add a guard test asserting no unresolved %%CDK%% token survives in the
synthesized user-data. Also add .github/scripts/** to build-artifacts paths so
script-only changes trigger a publish.
jest 20/20; tsc + shellcheck clean; rendered user-data: OPENSWE_ENV="dev",
SERVER_NAME="openswe-dev.seahaven.com", @@ sed patterns preserved, 16872 B.
`tar -tzf app.tar.gz | grep -qx` under set -o pipefail fails the pipeline when
grep -q matches and exits early (SIGPIPEs tar -> 'write error' -> non-zero), a
false 'missing agent/server.py'. List the archive once into a var, then grep the
var. Same fix for the secret-guard pipe (which was also silently broken).
The build-artifacts SPA build hit Node's default ~2 GB heap cap and aborted
(JS heap OOM, exit 134) — the same memory-hungry vite build that needed an 8 GB
swapfile on-box. The runner has ~16 GB, so set NODE_OPTIONS=--max-old-space-size=8192
on both build steps.
* feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3)
Packer-build the custom base image and repoint AppService off the AL2023
placeholder onto it.
deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real
`packer build` (the config had only ever been `packer validate`'d at T8):
- the file provisioner failed uploading the templates dir ('scp: …: Is a
directory') — a trailing-slash contents-upload needs the dest dir to exist;
added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the
dest trailing slash.
- the shell provisioner's custom execute_command omitted {{ .Vars }}, so the
environment_vars never reached provision.sh (which runs under set -u and
aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}.
infra:
- ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26
from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by
exact id via MachineImage.genericLinux (offline, deterministic). Dropped the
now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement
discipline docs.
- app-service.ts: machineImage → bakedOpenSweArm64().
- open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard).
- cdk.context.json → {} (AMI is a static id pin; no context lookups remain).
- README: Baked AMI + EBS-replacement-discipline section.
tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI.
NOTE: held — do NOT merge until the open-swe-dev secret values are populated
(put-config.sh). The infra CD is live, so merging this to dev auto-deploys
OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts →
unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14).
* fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021
Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects
non-ASCII in the AMI Description attribute, so packer registered then
DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error.
Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available).
Re-pinned BAKED_OPEN_SWE_AMI_ID.
* fix(deploy): GitHub App + Slack required for prod only, not dev
Per the migration decision: do NOT create/duplicate a separate dev GitHub App or
Slack app — only prod owns the single shared app. So fetch-config.sh no longer
hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/
CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only
block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET.
Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active
provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env
(boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is
unchanged (prod still requires everything).
* feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19)
Make the dev/prod box deployable end-to-end: a real artifact pipeline and a
re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy.
Infra (T7):
- assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access,
SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent +
abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput.
- app-service.ts: open-swe-<env>-deploy SSM document that runs the baked
/opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the
baked open-swe-base-arm64 AMI (folds in the held #16).
IAM (app deploy role — cross-review gated):
- github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to
open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic
AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document
is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox.
Boot/deploy (T19):
- deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull
app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv
at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx.
- user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB
target is healthy even before the first release); deploy.sh is base64-rendered
by CDK into user-data (a normal reviewable repo file, not a heredoc) and the
first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy).
CI (T7+T19):
- build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite ->
ui/.output/public -> spa.tar.gz), package the Python source via git archive
(app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via
the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy.
push dev -> dev (auto); push main -> prod (env "prod" approval gate).
Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK,
deploy.sh base64 round-trips exact.
* harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard
Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one
confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01,
the account-wide CDK cfn-exec residual already documented in config.ts; recorded in
.security-review/suppressions.json with justification + flagged for the per-env
bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified;
these are the cheap defense-in-depth fixes worth taking regardless:
- deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root
never honors an archive's uid/mode → no setuid/foreign-owned file can land); and
treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy
failure (set -e stays loud once a release exists).
- publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount),
not CommandInvocations[0], so a partial failure across the brief 2-instance
replacement window can't be reported as success.
- instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app
role's write scope) instead of the whole bucket.
- package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz
(defense in depth over .gitignore; scoped to data extensions so *_credentials.py
source is not a false positive — verified against the real tree).
Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the
CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable
releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release
to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz).
shellcheck/tsc/jest(16) clean; both stacks synth offline.
* fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard
The instance-SG GroupDescription + ingress/egress rule descriptions carried an
em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects
non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"),
so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing
from #14; same class as the AMI-description ASCII bug.)
- app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the
ingress/egress rule descriptions, and the Route53 comment.
- test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup
GroupDescription + rule descriptions are pure ASCII, so this fails the build
instead of a deploy next time.
jest 18/18; tsc clean.
* fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`)
The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule*
descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`,
which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the
ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII
only" to the exact EC2 allowed charset so it catches `>` (and `<`) too.
jest 18/18; tsc clean.
* fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit
The base64 deploy.sh embedded in user-data pushed the encoded boot script to
27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with
"Encoded User data is limited to 25600 bytes". Strip full-line comments + blank
lines from deploy.sh before base64-embedding it (repo file keeps comments; only
the on-box copy is minified; the script is opaque base64 so user-data heredocs are
unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a
synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded.
jest 19/19; minified deploy.sh passes bash -n + shellcheck.
* ci: path-filtered infra CI/CD with dual OIDC roles + prod approval gate (T18)
Add the /infra half of the combined-repo pipeline (the Python agent keeps ci.yml):
- ci-infra.yml — PR check on infra/** : tsc + jest + cdk synth via the org
reusable ci-typescript-cdk.yaml (working-directory: infra).
- cd-infra.yml — push to dev/main on infra/** (or dispatch):
* job 'ci' (reusable) is the CI-green precondition (deploy needs: ci).
* deploy-dev (ref=dev, NO environment) → cdk deploy OpenSweDevStack,
assuming githubdeploy-open-swe-infra-dev (OIDC sub ref:refs/heads/dev). AUTO.
* deploy-prod (ref=main, environment: prod) → cdk deploy OpenSweProdStack,
assuming githubdeploy-open-swe-infra-prod (OIDC sub environment:prod). The
'prod' Environment's required reviewer is the manual-approval gate.
Deliberately self-contained (NOT the reusable cd-cdk.yaml) because that runs
'cdk deploy --all' — from a single-env push it would deploy the other env + the
shared IAM stack, breaking the per-env boundary. CD targets one stack per env;
the shared open-swe-iam stack is human-gated (T6), never deployed by CD.
Infra CI is enforced at the DEPLOY boundary (deploy jobs need ci), not as a
branch-protection required check — path-filtering a required check would deadlock
app-only PRs. Documented in infra/README.md along with the post-T6 prerequisites
(repo vars AWS_DEPLOY_ROLE_INFRA_{DEV,PROD}; a 'prod' Environment w/ reviewer).
Not active until the IAM roles are applied (T6) — assuming a nonexistent role
just fails closed. App-side CD (S3 artifact + SSM) is T19.
* fix(infra): commit jest.config.js (was ignored by *.js → infra CI used Babel)
The infra/.gitignore *.js rule (for compiled CDK output) silently swept up the
hand-authored jest.config.js, so it was never committed. Local jest passed (file
present in the working tree) but CI's fresh checkout lacked it → jest fell back to
the default Babel transform → 'Cannot use import statement outside a module' on the
TypeScript test. Surfaced now because T18 is the first workflow to run infra jest
in CI. Negate the ignore for this one file and commit it.
AppService construct wires the per-env EC2 box and its internet path. The
seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the
on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and
never owned/mutated; open-swe only ADDS its own resources.
Per env (open-swe-stack.ts → AppService):
- ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a,
in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination
(no RETAIN volume — replacement-tolerant; see ami-cache.ts).
userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh.
- Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress
via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress
(the imported, on-prem-owned SG is never mutated).
- Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays
loopback). Health check GET /healthz.
- Two rules on the imported :443 listener, both → the TG:
* webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/*
* site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api)
Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5)
or it would steal every webhook — first-match-by-ascending-priority.
- Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB.
- 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config).
Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed
medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large
GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on
/dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene);
XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed
critical/high.
Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked
open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean.
Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
Create the per-env config surface the EC2 box reads at boot via
fetch-config.sh / seed_store.sh:
- ConfigStore construct (infra/lib/constructs/config-store.ts):
- 28 value-LESS Secrets Manager shells open-swe-<env>/<VAR>
(RemovalPolicy.RETAIN, no SecretString/generateSecretString — real
values are set out-of-band by put-config.sh, never in IaC/state).
- 8 IaC-managed SSM params /open-swe-<env>/<VAR> with real,
stable/derivable values (SANDBOX_TYPE, DEFAULT_REPO_OWNER/NAME,
ALLOWED_GITHUB_ORGS, DASHBOARD_*_URL/ORIGINS, LLM_MODEL_ID).
- OUT_OF_BAND_SSM documents the ~30 params CDK intentionally does NOT
own (operationally-variable / env-specific-unknown).
- Wire ConfigStore into OpenSwe<Env>Stack.
- KebabNamingAspect: exempt Secrets Manager + SSM names, which carry the
literal UPPER_SNAKE env-var segment (open-swe-dev/DASHBOARD_JWT_SECRET).
- deploy/seahaven/put-config.sh: out-of-band populator (placeholders only,
OPENSWE_PUT_<VAR> env indirection; no real values committed).
Synth-only; not deployed. Instance-role read grants on open-swe-<env>/*
already exist from T6 — no IAM/trust changes here.
PR#7 of the AWS migration. deploy/seahaven/: fetch-config.sh materializes a
service-user-owned 0600 tmpfs .env from Secrets Manager + SSM (fail-fast);
seed_store.sh reseeds the in-memory store; ROTATION.md.
Incorporates T5 /sh-security-review fixes:
- seed_store no longer bash-sources the .env (closes the SH-INJ-001 RCE); uses a
non-eval reader, jq --arg JSON bodies, and a loopback-pinned BASE.
- fetch-config: .env owned by the openswe service user (app no longer runs as
root); DEFAULT_REPO_OWNER hard-pinned; key-identifier validation + flat-namespace
collision detection; dropped SSM --recursive.
shellcheck + bash -n clean.
PR#2 of the AWS migration. deploy/ami/: Packer template (Ubuntu 24.04 arm64,
uv+py3.12, nginx, awscli v2, CW agent; no swapfile), provisioning-only user-data
(userDataCausesReplacement rationale), systemd unit + nginx + CW templates.
Incorporates T5 /sh-security-review fixes: langgraph binds 127.0.0.1 (not 0.0.0.0);
nginx is the sole ingress proxying only /dashboard/api/ + /webhooks/; ExecStartPre
runs fetch-config as root (+) and passes the env arg; the app runs as the
unprivileged openswe user reading an openswe-owned 0600 .env. packer validate clean.
Aligns the promotion target with the locked migration architecture
(main = PROD, dev = integration). Plain ref push is fast-forward-only;
branch protection on main rejects non-FF pushes, so a diverged main
fails loudly instead of force-moving prod.
* ci: promote dev to prod via fast-forward instead of force-pushing main
Repoints the daily promotion workflow to feed prod from the dev integration
trunk (prod <- dev) rather than force-pushing the upstream mirror (main).
Uses a fast-forward push so a release can never rewrite prod history.
* fix: sort imports in aegra_entry.py to unblock dev CI
The Sea Haven deployment commit left aegra_entry.py with an unsorted import
block (ruff I001), which fails the required 'Agent lint' check and blocks all
merges into dev. Apply the import-sort autofix.
Captures the stock-LangGraph deployment of this fork at Sea Haven:
- systemd/open-swe.service: langgraph dev (:2024) + store seed ExecStartPost
- seed_store.sh: re-seeds team_settings + user_mappings (in-memory store
resets on restart); env-parameterized, no secrets
- nginx/openswe.conf: dashboard SPA + scoped /dashboard/api proxy (security
boundary; agent API not exposed)
- aegra/: deferred self-hosted-runtime alternative (not active on stock)
- DEPLOYMENT.md: full runbook (models, build, ingress, OAuth callback)
Secrets and internal infra identifiers are intentionally excluded (public
fork); real values live in private IT docs.
Add a "Start from a template" gallery to the Automations page so users can
scaffold common scheduled agent runs (PR review digest, issue triage,
dependency checks, flaky-test tracking, release notes, docs freshness,
security audit) instead of writing every automation from scratch. Selecting
a template opens the new-automation editor prefilled with its instructions
and a default schedule, which the user can tweak before saving.
Also adds an isDescribableCron helper so template schedules render as
human-readable labels rather than raw cron in the editor.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>