Commit graph

19 commits

Author SHA1 Message Date
seahaven-openswe[bot]
2b01652754
refactor: adopt modular webhook architecture (#1621) + port fork customizations (#85)
* Adopt upstream modular webhook skeleton (#1621)

Apply the durable-interrupt-dispatch refactor: split the monolithic
webapp.py into a thin routing layer plus per-source handlers in
webhooks/{github,slack,linear}.py, and add completion.py, dispatch.py,
and reconcile.py. Reconcile fork divergence by keeping the Bedrock/
Fireworks cross-provider fallback, the no-agent-attribution prompt
policy, the dashboard-handoff re-export, and the Slack channel-info
cache. ci_autofix is restored on the new dispatch model in a later
commit.

Refs: #80

* Port fork webhook security delta onto modular handlers

Re-apply the fork's security customizations that #1621 did not carry:
Linear webhook replay protection (freshness window on the signed
webhookTimestamp), per-repo token-cache binding threaded through the
thread token resolvers, the INTERNAL_BOT_LOGINS self-check in the
review-finding-reply path, and a user-mapping cache refresh before
email resolution on the issue and PR-comment paths (multi-replica
staleness). Existing fork security tests pass unchanged.

Refs: #80

* Restore CI auto-fix on the modular dispatch model

Bring back ci_autofix.py and the ci_monitor graph that #1621 deleted,
re-wiring the fork's security-reviewed PR-babysitting onto the new
structure: the CI-event, autofix-toggle, and review-feedback handlers
move into webhooks/github.py and the github_webhook router re-gains the
check_run/check_suite/workflow_run/status routing plus the autofix
command and actionable-review branches. Auto-fix runs now dispatch
through dispatch_agent_run (durability + completion webhook) while
keeping the deliberate batch-while-busy skip-rule via
get_thread_active_status. Restore langgraph.json's ci_monitor entry and
the fork autofix tests (dispatch mock + import paths re-pointed).

Refs: #80

* Reformat and update docs for the modular webhook split

Point CLAUDE.md and deploy/MIGRATION.md at the new webhooks/ modules
and the dispatch/completion/reconcile contract, and mark the
user-mapping cache-refresh fix as applied on the GitHub handlers.

Refs: #80

* Restore reject backstop for autofix dispatch

A burst of near-simultaneous CI events for one head SHA can slip past
the busy-check before the dedupe SHA is recorded, so dispatch the
autofix path with multitask_strategy=reject (dev's prior platform
default) to drop duplicate concurrent creates instead of letting them
interrupt each other. Also make the completion failure-reply dedup
claim-then-post and drop the unreachable interrupted branch.

---------

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-06-30 18:46:46 -04:00
Adam Moussa
860ce93ad7
docs: record final managed-deployment topology in MIGRATION.md (#77)
Add an authoritative "Final, verified topology" section (1a) and
reconcile the phased plan with the executed end-state, superseding the
spike-era values throughout.

Captures: two LangGraph Cloud deployments (dev/prod, one LangSmith
workspace) with their final URL hashes; one Vercel project open-swe-prod
with production + custom dev environments and per-env LANGGRAPH_BACKEND_URL;
the Nitro routeRules proxy mechanism (PR #76, superseding the vercel.json
rewrite and PR #75's build-vercel-output.mjs); two dev/prod-isolated
GitHub Apps; per-deployment user stores; the Bedrock IAM users; and the
seven hard-won operational gotchas.

Refs: #65 #74 #76
2026-06-30 15:00:28 -04:00
Adam Moussa
8441bbb2d8
docs(migration): record the two Bedrock IAM users + scoped invoke policy (#74)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Document the live execution of MIGRATION.md §10.1 (Bedrock auth via static
keys): the customer-managed least-privilege policy open-swe-bedrock-invoke
and the open-swe-dev-bedrock / open-swe-prod-bedrock IAM users. Captures the
static-key deviation rationale (managed LangGraph Cloud cannot assume a role)
and that both mandatory gates (GPT-4.1 IAM cross-review, /sh-security-review)
passed with no critical/high.
2026-06-30 14:38:47 -04:00
Adam Moussa
a30ce2ab40
feat: managed LangGraph Cloud + Vercel migration (PR2 — code fixes + docs) (#65)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
* fix(dashboard): managed-cloud OAuth hardening + admin user-mapping endpoint

Prepare the dashboard backend for the managed LangGraph Cloud + Vercel
runtime, where the API is HTTPS and cross-site from the UI.

- OAuth redirect_uri (#2): coerce a schemeless DASHBOARD_API_BASE_URL to
  https:// in _api_base_url() so GitHub stops rejecting login with
  "redirect_uri not associated with this application". _cookie_security()
  now treats a schemeless (managed) value as Secure; SameSite=None too,
  consistent with the coerced scheme.
- OAuth state cookie (#3): document that osw_oauth_state is host-only by
  design (a Domain cookie is unsafe across *.vercel.app, a public suffix),
  so login must always start on the stable alias to avoid "oauth state
  mismatch". Operational contract; no behavioral change.
- Admin user mappings (#4): add POST /admin/user-mappings so an admin can
  set the github_login -> work_email link from the dashboard instead of a
  raw Store write. New "admin" MappingSource provenance value.

* fix(webapp): refresh user-mapping cache on GitHub webhook paths

On managed LangGraph Cloud the backend runs multiple replicas, so the
per-process GitHub<->work-email mapping cache can be stale on the replica
handling a webhook (a mapping created on another replica is invisible
until refresh). process_github_pr_comment and process_github_issue now
refresh the cache from the durable Store before resolving the author's
email, matching the existing Slack mention path (process_slack_mention).

* perf(webapp): defer deepagents import to speed custom-app cold start

The custom FastAPI app (agent.webapp:app, the langgraph.json http.app)
pulled deepagents -> langchain_anthropic -> anthropic into its import
graph via dashboard.routes, only to build skill/chat seed files. Defer
those create_file_data imports into the functions that use them. Removes
deepagents/langchain_anthropic/anthropic from app import entirely and
roughly halves module-import wall time (~0.6-0.8s -> ~0.35s warm; larger
cold-start saving since native anthropic init is skipped). Behavior
identical. (reviewer_diff already imports deepagents under TYPE_CHECKING.)

* feat(ui): set work_email user mappings from the admin dashboard

Add an "Add / update" form to the admin User mappings section and the
adminUpsertUserMapping API client method, wiring the new
POST /admin/user-mappings endpoint. Admins can now create or update a
github_login -> work_email mapping directly instead of waiting for the
user to self-connect Slack.

* docs: document managed LangGraph Cloud + Vercel deployment

- INSTALLATION §10: add the managed production env triad (LANGGRAPH_URL,
  DASHBOARD_BASE_URL + DASHBOARD_API_BASE_URL with https://, empty
  VITE_DASHBOARD_API_BASE_URL for same-origin), the stable-alias login
  and vercel.json stable-deployment-URL requirements, multi-replica cache
  note, plus redirect_uri-scheme and oauth-state-mismatch troubleshooting.
  Refresh the langgraph.json snippet to all six graphs.
- README: reframe deployment around the managed migration; link the plan.
- deploy/MIGRATION.md: import the self-hosted -> managed migration plan.
2026-06-29 19:58:38 -04:00
Adam Moussa
7f60324f0c
chore: decommission self-hosted AWS LangGraph stack (#64)
* chore: decommission self-hosted AWS LangGraph stack

Removes the now-dead self-host IaC and AWS-only CI/CD after destroying the
dev + prod CloudFormation stacks (open-swe-dev, open-swe-prod, open-swe-iam,
and the dev-exclusive CDKToolkit-oswedev bootstrap) in account 328440206208,
us-east-1. The deployment is now managed (LangGraph Cloud + Vercel).

- remove infra/ (CDK app: app + IAM stacks, constructs, aspects, tests)
- remove deploy/ami (Packer AMI build) and deploy/seahaven (boot/config
  scripts, DEPLOYMENT/ROTATION runbooks)
- remove AWS-only workflows: cd-infra, ci-infra, build-artifacts, rollback
- README: rewrite the Deployment section to the managed LangGraph Cloud +
  Vercel view; drop dead links to infra/ and deploy/seahaven

Preserved: the shared default CDKToolkit bootstrap and promote-dev-to-prod.yml.
The RETAIN'd Secrets Manager shells and open-swe-<env>-assets S3 buckets
survive cdk destroy by design (orphaned) and need a separate deliberate cleanup.

* chore: clean up dangling references left by the AWS decommission

Folds in the FIX-level items from the #64 review gates (GPT-4.1 cross-review +
/sh-security-review), none of which were blockers:

- delete orphaned .github/scripts/{package-artifacts,publish-and-deploy,roll-box,
  rollback}.sh — their only callers were the removed AWS deploy workflows
- drop the deleted /infra dir from dependabot.yml npm directories (was producing
  a recurring Dependabot config error)
- remove the stale OSWE-IAC-SECRETS-LIST-01 suppression (referenced the deleted
  infra/lib/constructs/instance-role.ts)
- repoint the README promotion link to promote-to-main.yml (renamed in #63)

The promote-dev-to-prod.yml comment in check-dev-green.sh is intentionally left
to #63, which rewrites that same line.
2026-06-29 19:54:38 -04:00
Adam Moussa
a4ed19ba61
feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62)
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)

Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.

- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
  region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
  FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module

* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8

The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.

Map profile effort to additional_model_request_fields:
  {thinking: {type: adaptive, display: summarized},
   output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.

* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids

Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
  (set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
  FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
  otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.

Surfaced by the cross-family review + verified against deploy/.

* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip

From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
  validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
  error code only, so the role ARN + account id in the raw botocore message never
  reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
  (Converse emits reasoning_content, not thinking) so the middleware is not a no-op
  on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)

* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids

Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
  us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
  each routed region (us-east-1/2, us-west-2). The model runs in the server process
  on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
  for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
  mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
  bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
  seed_store.sh's default via pick precedence, so the seed-script fix alone was
  insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
  ids to the Bedrock id (config.toml's model_id was an active, now-broken value).

AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.

* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)

Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
  (eval judge only — Bedrock builder/reviewer auth via the host IAM role).

REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
2026-06-29 15:57:19 -04:00
Adam Moussa
9444fd7677
docs: document live AWS prod deploy for Open SWE (#53)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Prod went live 2026-06-29 on self-hosted AWS EC2 behind the shared
seahaven-com ALB, superseding the on-prem VM model the runbook described.

- Rewrite deploy/seahaven/DEPLOYMENT.md as the canonical end-to-end runbook:
  infra CD (CDK stacks + OIDC roles + prod approval gate), config seeding
  (put-config.sh, the 13 boot-required prod vars, fetch-config fail-fast),
  app artifact deploy (S3 + SSM roll + is-active gate), promotion/rollback,
  live prod facts, and a RETAIN secret-shell troubleshooting entry that
  cross-references infra/README.md.
- Correct retired *.seahavenind.com hosts to *.seahaven.com throughout and
  document the live GitHub/Slack/Linear webhook + OAuth endpoints.
- Add a concise Deployment section to README pointing at the runbook.
- Fix the stale host in the retired on-prem nginx/openswe.conf and mark it
  superseded by the AMI template.
2026-06-29 11:12:44 -04:00
Adam Moussa
faae9a685b
Scope secrets fetch to --secret-id-list; drop ListSecrets grant (#48)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Switch fetch-config.sh from a name-prefix batch-get-secret-value
--filters scan to an explicit --secret-id-list (the 28 SECRET_VARS,
chunked at the 20/call cap). An id-list batch authorizes per-secret
ARN, so the instance role's BatchGetSecretValue moves from Resource:*
to the open-swe-<env>/* prefix and the account-wide ListSecrets grant
is dropped entirely. The box can no longer enumerate secret names
account-wide; cross-env value isolation is unchanged (GetSecretValue
was already prefix-scoped). Resolves OSWE-IAC-SECRETS-LIST-01.

Also capture each chunk response into a variable and consume the
producer via command substitution so a failed AWS call aborts under
set -e instead of being swallowed by process substitution and
misreported as a missing required var.

Refs: OSWE-IAC-SECRETS-LIST-01
2026-06-28 20:45:17 -04:00
Adam Moussa
5e30bc6be2
Point dev agent at seahaven-open-swe-dev org + un-pin the owner guard (#27)
Dev now runs against the dedicated seahaven-open-swe-dev org (repo openswe-dev-sandbox),
isolated from the real Sea Haven org. Two changes:

- config-store.ts: iacManagedSsm repo targeting is now per-env — dev =
  seahaven-open-swe-dev/openswe-dev-sandbox, prod stays Sea-Haven-Industries/open-swe-pilot.
  ALLOWED_GITHUB_ORGS tracks the env owner. Adds a dev-only SEED_USER_MAPPINGS param so the
  triggering GitHub login (amoussa1229) resolves and @openswe comments aren't skipped.

- fetch-config.sh: replace the unconditional hard-pin to Sea-Haven-Industries with an owner
  GUARD that HONORS the configured owner (OPENSWE_REPO_OWNER override, else the SSM value) but
  forces a per-env safe org when the normalized owner is blank or the upstream langchain-ai.
  Normalization (lowercase, strip whitespace, first path segment, drop dots) catches
  langchain-ai/<repo>, langchain-ai., and case variants without over-blocking legit orgs
  (e.g. langchain-ai-fork). Fallback org is per-env so dev can't fall back into the real org.

Reviews: GPT-4.1 cross-review APPROVE (round 1 found a path/dot bypass -> hardened, round 2 clean);
/sh-security-review authz one LOW (env-invariant fallback) -> fixed. Positive org allowlist still
enforced by the app via ALLOWED_GITHUB_ORGS.
2026-06-27 19:29:03 -04:00
Adam Moussa
f2633f9bd0
fix: don't crash-loop the box when no user mapping is configured (#25)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
seed_store.sh (ExecStartPost) exited 1 when it couldn't resolve a user_mappings
entry, which — with Type=simple — fails the whole unit and crash-loops the box. This
contradicts the script's own OSWE-SEED-03 precedent (the server-not-ready path exits
0 specifically to avoid a restart loop). A missing user mapping is a seeding gap, not
an unhealthy server: the @openswe trigger just won't resolve a commenter, which a
deployment-validation env (dev) does not need.

Make it non-fatal — warn and skip the user_mappings seed; team_settings is still
seeded and the service starts. The empty MAPPINGS array makes the seed loop a no-op.
Set SEED_USER_MAPPINGS / CONFIGURED_ADMINS / OPENSWE_OWNER_LOGIN+EMAIL to seed it.

Verified on the dev box: service active, langgraph bound to 127.0.0.1:2024 (loopback),
/ok 200, /healthz 200.
2026-06-26 20:06:53 -04:00
Adam Moussa
a2877d46ff
fix: paginate batch-get-secret-value in fetch-config (secrets dropped past page 1) (#24)
fetch-config materialized only the FIRST page of Secrets Manager results: the AWS CLI
does NOT auto-paginate batch-get-secret-value (the script's comment claiming it does
was wrong; --no-cli-pager only disables the output pager, not API pagination). With 28
secret shells under open-swe-<env>/ the first page returned ~10 items, stranding the
rest on later pages. Required secrets that landed past page 1 (DASHBOARD_JWT_SECRET,
TOKEN_ENCRYPTION_KEY, LANGSMITH_API_KEY_PROD) were silently dropped, tripping the
FAIL-FAST 'missing required var' guard and crash-looping open-swe.service.

Follow NextToken across pages (new batch_get_secrets_tsv helper). Verified against the
live open-swe-dev secrets: now loads all 5 populated secrets (was 2). Same per-record
base64 / exact-prefix / accept_var / emit_var hardening — only the page loop is new.
2026-06-26 19:55:01 -04:00
Adam Moussa
5fa132205b
fix(deploy): use %%...%% for CDK user-data tokens (don't collide with @@ sed) (#21)
The systemd unit booted with a literal `@@OPENSWE_ENV@@` (fetch-config.sh got the
token, not "dev" -> exit 2 -> crash-loop) because user-data.sh is double-templated:
CDK substitutes @@tokens@@ AND user-data seds @@tokens@@ into the baked
systemd/nginx files. CDK's `.replace(/@@OPENSWE_ENV@@/g, "dev")` clobbered the sed
PATTERN (`s|@@OPENSWE_ENV@@|...|` -> `s|dev|...|`, a no-op), so the unit's token
never got replaced. Same collision hit @@SERVER_NAME@@ (masked by nginx
default_server).

Fix: CDK tokens move to a DISTINCT delimiter %%...%% (rendered in app-service.ts);
the @@...@@ tokens stay for the baked-template seds. No AMI rebuild (templates
unchanged). Add a guard test asserting no unresolved %%CDK%% token survives in the
synthesized user-data. Also add .github/scripts/** to build-artifacts paths so
script-only changes trigger a publish.

jest 20/20; tsc + shellcheck clean; rendered user-data: OPENSWE_ENV="dev",
SERVER_NAME="openswe-dev.seahaven.com", @@ sed patterns preserved, 16872 B.
2026-06-26 19:06:37 -04:00
Adam Moussa
404b3f6f75
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18)
* feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3)

Packer-build the custom base image and repoint AppService off the AL2023
placeholder onto it.

deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real
`packer build` (the config had only ever been `packer validate`'d at T8):
  - the file provisioner failed uploading the templates dir ('scp: …: Is a
    directory') — a trailing-slash contents-upload needs the dest dir to exist;
    added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the
    dest trailing slash.
  - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the
    environment_vars never reached provision.sh (which runs under set -u and
    aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}.

infra:
  - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26
    from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by
    exact id via MachineImage.genericLinux (offline, deterministic). Dropped the
    now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement
    discipline docs.
  - app-service.ts: machineImage → bakedOpenSweArm64().
  - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard).
  - cdk.context.json → {} (AMI is a static id pin; no context lookups remain).
  - README: Baked AMI + EBS-replacement-discipline section.

tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI.

NOTE: held — do NOT merge until the open-swe-dev secret values are populated
(put-config.sh). The infra CD is live, so merging this to dev auto-deploys
OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts →
unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14).

* fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021

Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects
non-ASCII in the AMI Description attribute, so packer registered then
DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error.
Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available).
Re-pinned BAKED_OPEN_SWE_AMI_ID.

* fix(deploy): GitHub App + Slack required for prod only, not dev

Per the migration decision: do NOT create/duplicate a separate dev GitHub App or
Slack app — only prod owns the single shared app. So fetch-config.sh no longer
hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/
CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only
block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET.

Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active
provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env
(boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is
unchanged (prod still requires everything).

* feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19)

Make the dev/prod box deployable end-to-end: a real artifact pipeline and a
re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy.

Infra (T7):
- assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access,
  SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent +
  abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput.
- app-service.ts: open-swe-<env>-deploy SSM document that runs the baked
  /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the
  baked open-swe-base-arm64 AMI (folds in the held #16).

IAM (app deploy role — cross-review gated):
- github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to
  open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic
  AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document
  is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox.

Boot/deploy (T19):
- deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull
  app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv
  at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx.
- user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB
  target is healthy even before the first release); deploy.sh is base64-rendered
  by CDK into user-data (a normal reviewable repo file, not a heredoc) and the
  first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy).

CI (T7+T19):
- build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite ->
  ui/.output/public -> spa.tar.gz), package the Python source via git archive
  (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via
  the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy.
  push dev -> dev (auto); push main -> prod (env "prod" approval gate).

Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK,
deploy.sh base64 round-trips exact.

* harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard

Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one
confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01,
the account-wide CDK cfn-exec residual already documented in config.ts; recorded in
.security-review/suppressions.json with justification + flagged for the per-env
bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified;
these are the cheap defense-in-depth fixes worth taking regardless:

- deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root
  never honors an archive's uid/mode → no setuid/foreign-owned file can land); and
  treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy
  failure (set -e stays loud once a release exists).
- publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount),
  not CommandInvocations[0], so a partial failure across the brief 2-instance
  replacement window can't be reported as success.
- instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app
  role's write scope) instead of the whole bucket.
- package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz
  (defense in depth over .gitignore; scoped to data extensions so *_credentials.py
  source is not a false positive — verified against the real tree).

Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the
CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable
releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release
to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz).

shellcheck/tsc/jest(16) clean; both stacks synth offline.

* fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard

The instance-SG GroupDescription + ingress/egress rule descriptions carried an
em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects
non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"),
so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing
from #14; same class as the AMI-description ASCII bug.)

- app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the
  ingress/egress rule descriptions, and the Route53 comment.
- test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup
  GroupDescription + rule descriptions are pure ASCII, so this fails the build
  instead of a deploy next time.

jest 18/18; tsc clean.

* fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`)

The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule*
descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`,
which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the
ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII
only" to the exact EC2 allowed charset so it catches `>` (and `<`) too.

jest 18/18; tsc clean.

* fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit

The base64 deploy.sh embedded in user-data pushed the encoded boot script to
27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with
"Encoded User data is limited to 25600 bytes". Strip full-line comments + blank
lines from deploy.sh before base64-embedding it (repo file keeps comments; only
the on-box copy is minified; the script is opaque base64 so user-data heredocs are
unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a
synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded.

jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
Adam Moussa
fcbdfb67aa
feat: open-swe dev/prod compute + ALB ingress (T12) (#14)
AppService construct wires the per-env EC2 box and its internet path. The
seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the
on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and
never owned/mutated; open-swe only ADDS its own resources.

Per env (open-swe-stack.ts → AppService):
- ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a,
  in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination
  (no RETAIN volume — replacement-tolerant; see ami-cache.ts).
  userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh.
- Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress
  via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress
  (the imported, on-prem-owned SG is never mutated).
- Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays
  loopback). Health check GET /healthz.
- Two rules on the imported :443 listener, both → the TG:
    * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/*
    * site     (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api)
  Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5)
  or it would steal every webhook — first-match-by-ascending-priority.
- Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB.
- 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config).

Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed
medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large
GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on
/dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene);
XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed
critical/high.

Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked
open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean.
Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
Adam Moussa
aef6b26912
feat: Secrets Manager + SSM config store for open-swe (T11) (#10)
Create the per-env config surface the EC2 box reads at boot via
fetch-config.sh / seed_store.sh:

- ConfigStore construct (infra/lib/constructs/config-store.ts):
  - 28 value-LESS Secrets Manager shells  open-swe-<env>/<VAR>
    (RemovalPolicy.RETAIN, no SecretString/generateSecretString — real
    values are set out-of-band by put-config.sh, never in IaC/state).
  - 8 IaC-managed SSM params  /open-swe-<env>/<VAR>  with real,
    stable/derivable values (SANDBOX_TYPE, DEFAULT_REPO_OWNER/NAME,
    ALLOWED_GITHUB_ORGS, DASHBOARD_*_URL/ORIGINS, LLM_MODEL_ID).
  - OUT_OF_BAND_SSM documents the ~30 params CDK intentionally does NOT
    own (operationally-variable / env-specific-unknown).
- Wire ConfigStore into OpenSwe<Env>Stack.
- KebabNamingAspect: exempt Secrets Manager + SSM names, which carry the
  literal UPPER_SNAKE env-var segment (open-swe-dev/DASHBOARD_JWT_SECRET).
- deploy/seahaven/put-config.sh: out-of-band populator (placeholders only,
  OPENSWE_PUT_<VAR> env indirection; no real values committed).

Synth-only; not deployed. Instance-role read grants on open-swe-<env>/*
already exist from T6 — no IAM/trust changes here.
2026-06-26 16:07:23 -04:00
Adam Moussa
434d0ad80c
feat(deploy): AWS-sourced fetch-config + seed_store + rotation docs (PR#7) (#8)
PR#7 of the AWS migration. deploy/seahaven/: fetch-config.sh materializes a
service-user-owned 0600 tmpfs .env from Secrets Manager + SSM (fail-fast);
seed_store.sh reseeds the in-memory store; ROTATION.md.

Incorporates T5 /sh-security-review fixes:
- seed_store no longer bash-sources the .env (closes the SH-INJ-001 RCE); uses a
  non-eval reader, jq --arg JSON bodies, and a loopback-pinned BASE.
- fetch-config: .env owned by the openswe service user (app no longer runs as
  root); DEFAULT_REPO_OWNER hard-pinned; key-identifier validation + flat-namespace
  collision detection; dropped SSM --recursive.
shellcheck + bash -n clean.
2026-06-26 15:06:41 -04:00
Adam Moussa
ce9e5fb51b
feat(deploy): EC2 AMI recipe (Packer) + cloud-init/user-data (PR#2) (#7)
PR#2 of the AWS migration. deploy/ami/: Packer template (Ubuntu 24.04 arm64,
uv+py3.12, nginx, awscli v2, CW agent; no swapfile), provisioning-only user-data
(userDataCausesReplacement rationale), systemd unit + nginx + CW templates.

Incorporates T5 /sh-security-review fixes: langgraph binds 127.0.0.1 (not 0.0.0.0);
nginx is the sole ingress proxying only /dashboard/api/ + /webhooks/; ExecStartPre
runs fetch-config as root (+) and passes the env arg; the app runs as the
unprivileged openswe user reading an openswe-owned 0600 .env. packer validate clean.
2026-06-26 15:06:36 -04:00
Adam Moussa
35bc9e22f4
ci: promote dev → prod via fast-forward (not force-push from main) (#2)
* ci: promote dev to prod via fast-forward instead of force-pushing main

Repoints the daily promotion workflow to feed prod from the dev integration
trunk (prod <- dev) rather than force-pushing the upstream mirror (main).
Uses a fast-forward push so a release can never rewrite prod history.

* fix: sort imports in aegra_entry.py to unblock dev CI

The Sea Haven deployment commit left aegra_entry.py with an unsorted import
block (ruff I001), which fails the required 'Agent lint' check and blocks all
merges into dev. Apply the import-sort autofix.
2026-06-26 11:20:10 -04:00
cf57ba7aa5 feat(deploy): add Sea Haven self-hosted deployment capture
Captures the stock-LangGraph deployment of this fork at Sea Haven:
- systemd/open-swe.service: langgraph dev (:2024) + store seed ExecStartPost
- seed_store.sh: re-seeds team_settings + user_mappings (in-memory store
  resets on restart); env-parameterized, no secrets
- nginx/openswe.conf: dashboard SPA + scoped /dashboard/api proxy (security
  boundary; agent API not exposed)
- aegra/: deferred self-hosted-runtime alternative (not active on stock)
- DEPLOYMENT.md: full runbook (models, build, ingress, OAuth callback)

Secrets and internal infra identifiers are intentionally excluded (public
fork); real values live in private IT docs.
2026-06-25 17:28:34 -04:00