Commit graph

1034 commits

Author SHA1 Message Date
Adam Moussa
7f60324f0c
chore: decommission self-hosted AWS LangGraph stack (#64)
* chore: decommission self-hosted AWS LangGraph stack

Removes the now-dead self-host IaC and AWS-only CI/CD after destroying the
dev + prod CloudFormation stacks (open-swe-dev, open-swe-prod, open-swe-iam,
and the dev-exclusive CDKToolkit-oswedev bootstrap) in account 328440206208,
us-east-1. The deployment is now managed (LangGraph Cloud + Vercel).

- remove infra/ (CDK app: app + IAM stacks, constructs, aspects, tests)
- remove deploy/ami (Packer AMI build) and deploy/seahaven (boot/config
  scripts, DEPLOYMENT/ROTATION runbooks)
- remove AWS-only workflows: cd-infra, ci-infra, build-artifacts, rollback
- README: rewrite the Deployment section to the managed LangGraph Cloud +
  Vercel view; drop dead links to infra/ and deploy/seahaven

Preserved: the shared default CDKToolkit bootstrap and promote-dev-to-prod.yml.
The RETAIN'd Secrets Manager shells and open-swe-<env>-assets S3 buckets
survive cdk destroy by design (orphaned) and need a separate deliberate cleanup.

* chore: clean up dangling references left by the AWS decommission

Folds in the FIX-level items from the #64 review gates (GPT-4.1 cross-review +
/sh-security-review), none of which were blockers:

- delete orphaned .github/scripts/{package-artifacts,publish-and-deploy,roll-box,
  rollback}.sh — their only callers were the removed AWS deploy workflows
- drop the deleted /infra dir from dependabot.yml npm directories (was producing
  a recurring Dependabot config error)
- remove the stale OSWE-IAC-SECRETS-LIST-01 suppression (referenced the deleted
  infra/lib/constructs/instance-role.ts)
- repoint the README promotion link to promote-to-main.yml (renamed in #63)

The promote-dev-to-prod.yml comment in check-dev-green.sh is intentionally left
to #63, which rewrites that same line.
2026-06-29 19:54:38 -04:00
Adam Moussa
430a1cdff9
ci: re-home prod promotion into a gated promote-to-main workflow (PR1: managed-LGC migration) (#63)
* ci: re-home prod promotion into a gated promote-to-main workflow

Migrate the prod-deploy gate to managed LangGraph Cloud (git-connected to
`main`) + Vercel. Under managed, a push to `main` auto-deploys prod, so the
dev -> main fast-forward IS the prod deploy trigger -- the bespoke AWS CD step
is obsolete and already gone from this workflow.

Re-home `promote-dev-to-prod.yml` -> `promote-to-main.yml`:
- gate the promote job on the `prod` GitHub Environment (required reviewer
  amoussa1229), restoring the manual prod-approval that the retired AWS CD job
  used to carry;
- drop the nightly auto-promote cron -- a scheduled auto-promotion conflicts
  with a manual approval gate now that the push deploys prod; promotion is
  workflow_dispatch only;
- keep the dev-HEAD-fully-green precondition and the seahaven-promotion App
  fast-forward push (sole non-admin bypass actor on `main` ruleset 18238334).

Update the companion check-dev-green.sh filename reference.

* ci: refuse promote-to-main dispatch from any ref other than dev

Defense-in-depth atop the already-pinned `ref: dev` checkout: workflow_dispatch
runs the workflow definition from the launched ref, so reject a non-dev dispatch
before the App token is minted. Surfaced by the GPT-4.1 cross-review of #63.
2026-06-29 19:49:21 -04:00
Adam Moussa
a4ed19ba61
feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62)
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)

Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.

- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
  region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
  FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module

* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8

The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.

Map profile effort to additional_model_request_fields:
  {thinking: {type: adaptive, display: summarized},
   output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.

* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids

Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
  (set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
  FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
  otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.

Surfaced by the cross-family review + verified against deploy/.

* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip

From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
  validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
  error code only, so the role ARN + account id in the raw botocore message never
  reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
  (Converse emits reasoning_content, not thinking) so the middleware is not a no-op
  on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)

* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids

Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
  us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
  each routed region (us-east-1/2, us-west-2). The model runs in the server process
  on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
  for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
  mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
  bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
  seed_store.sh's default via pick precedence, so the seed-script fix alone was
  insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
  ids to the Bedrock id (config.toml's model_id was an active, now-broken value).

AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.

* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)

Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
  (eval judge only — Bedrock builder/reviewer auth via the host IAM role).

REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
2026-06-29 15:57:19 -04:00
Adam Moussa
8a9974c3c4
feat: author Slack/dashboard/schedule commits + PRs as the app by default (#57) (#60)
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: default Slack/dashboard/schedule PRs + commits to the app identity (#57)

Slack/dashboard/schedule runs now author PRs and run git/gh operations as the
GitHub App seahaven-openswe[bot] by default (matching GitHub-issue runs), so the
self-review 422 is impossible by construction rather than guarded in the prompt.
A profile flag author_prs_as_user restores per-user attribution.

- open_pull_request._resolve_pr_author_token + auth.resolve_github_token: default
  to the installation token for these sources; per-user only when opted in.
- authorship: commit identity -> seahaven-openswe[bot] (numeric noreply;
  accepted Vercel-resolution risk, documented inline).
- self-trigger safety: INTERNAL_BOT_LOGINS + webapp/reviewer_reconcile/reply
  markers recognize seahaven-openswe[bot] (bot-authored events are now ours).

Supersedes the prompt-only guard in #58.

* fix: author commits as the app bot in the default path (SH-IDSPLIT-01)

Security review found the commit identity was NOT actually unified to the bot:
resolve_triggering_user_identity got a 403 from the installation token and fell
back to configurable['github_login'], so commits were still authored as the
triggering user (commit=user, push+PR=bot — a three-way split that missed the
stated goal). Now gate the triggering-user identity resolution on the same
default-bot decision as the token: slack/dashboard/schedule default to the app
bot identity unless author_prs_as_user is set.

* docs(security): record AUTHZ-SLACK-BOT-DEFAULT-001 as an accepted residual (#59)

Single-user deployment; bounded by App-on-pilot + ALLOWED_GITHUB_REPOS lock.
Revisit (add a per-user gate) before expanding users or the App installation.
2026-06-29 14:22:33 -04:00
Adam Moussa
134963647b
chore(security): suppress pre-existing history scanner false-positives (#56)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Adds repo-local suppressions for 5 verified-FP gitleaks findings that block
pushes (forcing --no-verify), all in committed history / docs / CI fixtures:
- .env.ci (dev-only e2e values, intentionally committed)
- .env.example (placeholders)
- .github/ci/fake_github_app_key.pem (throwaway CI test key)
- INSTALLATION.md (example GITHUB_APP_PRIVATE_KEY .env block)
- README.md (prose mis-matched by the generic-api-key heuristic)
Repo-local (not machine-level) so they load in git worktrees too. Also drops
the now-obsolete OSWE-IAC-AUDIT-01 suppression (B-1, fixed in #55).
2026-06-29 12:54:40 -04:00
Adam Moussa
e9499e49b8
fix: isolate dev CDK deploys on their own bootstrap qualifier (B-1/OSWE-IAC-01) (#55)
* fix: isolate dev CDK deploys on their own bootstrap qualifier (B-1)

Dev synthesizes against the oswedev qualifier and the dev infra deploy role
is scoped to cdk-oswedev-* — it can no longer assume the default hnb659fds
bootstrap roles whose admin cfn-exec-role deploys prod, closing the cross-env
escalation (OSWE-IAC-01). Prod stays on the default qualifier.

* test: assert per-env bootstrap qualifier isolation + document (B-1)
2026-06-29 12:39:23 -04:00
Adam Moussa
a33aaec495
fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54)
* fix: enforce a replay window on Linear webhooks (AUTHZ-001)

verify_linear_signature accepted any correctly-signed body with no freshness
check, so a captured request could be replayed indefinitely. Parse the
signed webhookTimestamp (Unix ms) and reject requests outside a 60s window,
failing closed when the field is missing or malformed — mirroring the Slack
verifier.

* fix: stop leaking upstream auth-error bodies into user comments

get_github_token_for_user folded the raw upstream response text into the
error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the
full body server-side only and return a generic "GitHub auth failed (status
<code>)". Also document the accepted shared-installation-token blast radius on
the bot-token-only path (AUTHZ-003).

* fix: bind sandbox and token caches to repo to prevent thread-id collision

A PR head-branch name is attacker-controllable and get_thread_id_from_branch
derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01).
The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on
thread_id alone, and a cached sandbox was reused after only an echo-ping, so a
different repo's webhook could bind to another thread's sandbox or token.

Without changing the persistent thread-id scheme:
- Persist the bound repo (owner/name) in thread metadata on sandbox creation and
  refuse to reuse a sandbox whose bound repo does not match the current event
  (SandboxRepoMismatchError); the in-memory proxy also carries the binding.
- Bind the GitHub-token cache entries to their repo and evict on a cross-repo
  read so a colliding thread_id cannot be served another repo's token.
- Thread repo through the reviewer and the webhook token resolvers.

* fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04)

The instance role and the GitHub deploy app role granted s3:ListBucket on the
whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts)
only ever lists under releases/, so add a StringLike s3:prefix=releases/*
condition. GetBucketLocation has no s3:prefix in its request context, so it
moves to its own unconditioned statement. Also document the accepted F-2
cross-env existence-oracle residual on BatchGetSecretValue.

* chore: suppress test-fixture credential false positive; document AUTHZ-002

Add a machine-level suppression for the fake Datadog key in the
test_team_credentials encryption-roundtrip fixture (CWE-798, not a real
credential). Clarify that the within-org thread-write path is intentional by
design (AUTHZ-002) — comment only, no behavior change.

* fix: casefold repo-binding keys to avoid spurious cross-repo mismatch

GitHub owner/name are case-insensitive. Casefold the owner/name key on both the
write (binding) and read (compare) sides — repo_cache_key and the metadata
bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate
same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2).

* fix: stop leaking upstream auth body in unexpected-result branch

The 2xx-but-missing-token/url branch echoed the parsed upstream response body
into the user-facing error. Return a generic message and log response_data
server-side only, mirroring the existing HTTPStatusError fix (Gap 4).

* fix: fail closed for unbound-legacy sandboxes and catch repo mismatch

Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no
recorded bound_repo (a pre-binding legacy thread, post-deploy) previously
reconnected-and-served the sandbox to the current repo, then rebound it. Now
fail closed: drop the stale id and recreate a fresh sandbox bound to this repo,
logging a reconnect-with-missing-binding event. A sandbox is never served to a
repo unless its binding is known and matches; new threads bind on first run
unchanged.

Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints,
log it for alarming, and surface a clean sanitized error instead of letting an
opaque deep-stack exception crash-loop the worker.

* chore: suppress test-fixture credential false positive in token-TTL tests

Add a machine-level suppression for the fake "ghp_secret" GitHub token used by
the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and
not a valid PAT; scoped to the unit test only.
2026-06-29 12:21:19 -04:00
Adam Moussa
9444fd7677
docs: document live AWS prod deploy for Open SWE (#53)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Prod went live 2026-06-29 on self-hosted AWS EC2 behind the shared
seahaven-com ALB, superseding the on-prem VM model the runbook described.

- Rewrite deploy/seahaven/DEPLOYMENT.md as the canonical end-to-end runbook:
  infra CD (CDK stacks + OIDC roles + prod approval gate), config seeding
  (put-config.sh, the 13 boot-required prod vars, fetch-config fail-fast),
  app artifact deploy (S3 + SSM roll + is-active gate), promotion/rollback,
  live prod facts, and a RETAIN secret-shell troubleshooting entry that
  cross-references infra/README.md.
- Correct retired *.seahavenind.com hosts to *.seahaven.com throughout and
  document the live GitHub/Slack/Linear webhook + OAuth endpoints.
- Add a concise Deployment section to README pointing at the runbook.
- Fix the stale host in the retired on-prem nginx/openswe.conf and mark it
  superseded by the AMI template.
2026-06-29 11:12:44 -04:00
Adam Moussa
f379fbdaa9
docs: document RETAIN secret-shell orphan gotcha (#52)
Some checks are pending
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
RETAIN + a fixed secret name means a failed FIRST create leaves empty
secret shells behind when the stack rolls back. The shells keep the
global `open-swe-<env>/<VAR>` names, so every later create fails with
`AlreadyExists`, and a plain delete-secret keeps the name reserved for
the recovery window rather than freeing it.

Record the trap and the force-delete recovery (only for empty shells)
in the config-store construct and the infra README so the next
teardown/rebuild, secret logical-id change, or new-env stand-up does
not rediscover it the hard way.

Prod's first deploy hit this on 2026-06-29: 28 orphaned shells from an
earlier failed create reserved the names and had to be force-deleted
before the stack would create.
2026-06-29 10:51:49 -04:00
Adam Moussa
86b4859589
fix: env-scope the EC2 launch template name (unblocks prod deploy) (#51)
Some checks are pending
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
requireImdsv2:true makes CDK auto-create a launch template named from the
construct id ('Instance' -> 'InstanceLaunchTemplate') with no env qualifier,
so OpenSweDevStack and OpenSweProdStack both render
LaunchTemplateName: InstanceLaunchTemplate. dev created it first (the live
dev box runs on it); the prod first-deploy then failed with
InvalidLaunchTemplateName.AlreadyExistsException and the whole stack rolled
back.

Force a per-env LT name (open-swe-<env>-lt) via an aspect (the LT is created
at synth time by the requireImdsv2 handling, not in the constructor), and
rename the instance's launch-template REFERENCE in lockstep so CFN still
resolves it. synth-verified: dev=open-swe-dev-lt, prod=open-swe-prod-lt on
both the LT resource and the instance reference; version GetAtt preserved.

NOTE: deploying this renames dev's LT -> one-time dev box replacement
(stateless; boots from the baked AMI + pulls releases/latest). prod then
creates open-swe-prod-lt cleanly.
2026-06-29 00:51:16 -04:00
Adam Moussa
3c69dd9de6
fix: BatchGetSecretValue must be granted on * (corrects #48, fixes dev crash-loop) (#50)
Some checks are pending
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
#48 (OSWE-IAC-SECRETS-LIST-01) scoped secretsmanager:BatchGetSecretValue
to the open-swe-<env>/* ARN on the theory that an explicit --secret-id-list
batch authorizes per-secret. That is FALSE: BatchGetSecretValue is a
collection action AWS authorizes against the account (*), regardless of
--filters vs --secret-id-list. A prefix-scoped grant AccessDenies the whole
call. The dev box passed right after #48 only because the prior broad grant
had not finished propagating; once it lapsed, fetch-config got AccessDenied
-> loaded 0 secrets -> FAIL-FAST -> open-swe.service crash-loop. Verified on
the live dev box (i-0af4e03e8bf70e6c3): the exact call returned
'not authorized to perform: secretsmanager:BatchGetSecretValue'; restoring
the * grant recovered it.

Move BatchGetSecretValue back to Resource:* (its own statement); keep
GetSecretValue + DescribeSecret prefix-scoped (those gate VALUE access, so
cross-env isolation holds). The surviving win from #48: --secret-id-list
needs no name filter, so ListSecrets stays dropped -> no account-wide name
enumeration. fetch-config.sh is unchanged (--secret-id-list is correct).

The /sh-security-review finding OSWE-IAC-IAM-01 called this out and was
wrongly refuted; the reference_secretsmanager_batch_get memory was wrong.
2026-06-28 22:08:02 -04:00
Adam Moussa
9acf071ae4
fix: make dev->main promotion push succeed via bypass-actor App token (#49)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
The nightly promote fast-forwards main to a fully-green dev HEAD, but the
push (as github-actions[bot]) is rejected by the main ruleset: it requires
PRs + a status check and the default token is not a bypass actor, so a
direct ref push can never land regardless of fast-forwardability. The
prior comment claiming protection 'only rejects non-FF' was wrong.

Mint a GitHub App installation token (actions/create-github-app-token,
SHA-pinned) and push with it; the App must be added to the main ruleset's
bypass actors out-of-band. The promoted commit already passed every check
on dev (gated by check-dev-green.sh), so re-gating it via a PR on main is
redundant.

Also fix a gate self-poison: a stale failed 'promote' check-run from a
prior run on the same dev HEAD blocked every subsequent gate run (it was
excluded only by the current run_id). Exclude prior promote check-runs
too, scoped to name=='promote' AND a /actions/runs/ details_url so an
external app cannot hide a real failing check by naming it 'promote'; the
positive REQUIRED_CHECKS allow-list stays authoritative.

Gates: GPT-4.1 cross-review APPROVE (no security regression). Unit-tested:
stale promote ignored -> PASS; real failure / external promote / missing
required check -> BLOCK. shellcheck clean (also fixed a pre-existing
SC2295 on the run_id match).
2026-06-28 21:42:50 -04:00
Adam Moussa
faae9a685b
Scope secrets fetch to --secret-id-list; drop ListSecrets grant (#48)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Switch fetch-config.sh from a name-prefix batch-get-secret-value
--filters scan to an explicit --secret-id-list (the 28 SECRET_VARS,
chunked at the 20/call cap). An id-list batch authorizes per-secret
ARN, so the instance role's BatchGetSecretValue moves from Resource:*
to the open-swe-<env>/* prefix and the account-wide ListSecrets grant
is dropped entirely. The box can no longer enumerate secret names
account-wide; cross-env value isolation is unchanged (GetSecretValue
was already prefix-scoped). Resolves OSWE-IAC-SECRETS-LIST-01.

Also capture each chunk response into a variable and consume the
producer via command substitution so a failed AWS call aborts under
set -e instead of being swallowed by process substitution and
misreported as a missing required var.

Refs: OSWE-IAC-SECRETS-LIST-01
2026-06-28 20:45:17 -04:00
Adam Moussa
52cc696431
fix(deps): bump PyJWT to >=2.13.0 to close HS256 forgery (#47)
PyJWT <2.13.0 accepts a public-key JWK as an HMAC secret, letting an
attacker forge HS256 tokens when mixed key families are allowed
(GHSA high-sev Dependabot alert). 2.13.0 rejects the mismatch.

Direct dependency; uv.lock re-resolves 2.12.1 -> 2.13.0.
2026-06-28 20:41:27 -04:00
seahaven-openswe[bot]
f79f49c755
Make PR title repo-aware and auto-link issues (#46)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Hardcoding the Sea Haven no-type-prefix PR title made every PR fail
semantic-PR-title gates (this repo's PR Title Lint, upstream open-swe),
forcing manual retitling. Make the title rule detect a conventional-commit
gate and conform, falling back to the imperative style otherwise. Also add
Closes/Refs issue-linking guidance and the default-branch auto-close caveat.

Refs: #41

Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com>
2026-06-27 22:56:01 -04:00
dependabot[bot]
2533402f6b
chore(deps): update langgraph-cli[inmem] requirement (#38)
Updates the requirements on [langgraph-cli[inmem]](https://github.com/langchain-ai/langgraph) to permit the latest version.
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/cli==0.4.27...cli==0.4.30)

---
updated-dependencies:
- dependency-name: langgraph-cli[inmem]
  dependency-version: 0.4.30
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-28 02:40:34 +00:00
dependabot[bot]
df96fe49b5
chore(deps): bump astral-sh/setup-uv (#32)
Bumps the minor-and-patch group with 1 update in the / directory: [astral-sh/setup-uv](https://github.com/astral-sh/setup-uv).


Updates `astral-sh/setup-uv` from 8.1.0 to 8.2.0
- [Release notes](https://github.com/astral-sh/setup-uv/releases)
- [Commits](08807647e7...fac544c07d)

---
updated-dependencies:
- dependency-name: astral-sh/setup-uv
  dependency-version: 8.2.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
  dependency-group: minor-and-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-27 22:35:48 -04:00
dependabot[bot]
485674d941
chore(deps): bump python in the minor-and-patch group across 1 directory (#31)
Bumps the minor-and-patch group with 1 update in the / directory: python.


Updates `python` from 3.14.5-slim-trixie to 3.14.6-slim-trixie

---
updated-dependencies:
- dependency-name: python
  dependency-version: 3.14.6-slim-trixie
  dependency-type: direct:production
  update-type: version-update:semver-patch
  dependency-group: minor-and-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-27 22:35:45 -04:00
seahaven-openswe[bot]
719741f279
Align dependabot with Sea Haven conventions (#39)
Switch all ecosystems to a weekly schedule and drop the redundant
patterns: ["*"] selectors. Remove the major group so patch+minor bumps
stay grouped into one PR per ecosystem while each major lands as its own
PR, matching the org standard.

Add npm coverage for the JS/TS surface (root, /infra CDK, /tests/e2e
Playwright, /ui dashboard) and assign updates to amoussa1229. Delete the
stale ui/yarn.lock so Dependabot tracks ui/bun.lock cleanly.

Refs: #34

Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com>
2026-06-27 22:20:31 -04:00
Adam Moussa
f3db9f02e3
Adopt Sea Haven agent conventions, no attribution (#30)
Codify the box-only #4 customizations into Git so the AWS deployment
(which deploys from this repo) actually applies them — previously only
the retired sh-openswe box had them.

- prompt.py: branch names feature|bug|hotfix/<kebab> (optional <KEY->);
  imperative PR titles with no conventional-commit type: prefix; PR body
  Summary/Validation/Tests/Notes; handbook commit format. Rewrite the
  collaboration template from an attribution MANDATE to a PROHIBITION —
  no Co-authored-by bot trailer, no "Made by [Open SWE]" footer, no
  agent/AI notes on any artifact.
- github_comments.py: add @seahaven-openswe (the deployed App slug) to
  the mention triggers.
- authorship.py: remove the now-unused attribution helpers
  (build_pr_attribution_footer, add_bot_coauthor_trailer,
  add_pr_collaboration_note, PR_ATTRIBUTION_*). Keep OPEN_SWE_BOT_* —
  server.py still uses them for the sandbox git identity.
- Flip the attribution unit tests to assert the no-attribution behavior;
  drop tests for the removed helpers.

Commits stay authored as the triggering user for now — flipping
authorship to the bot account depends on the Vercel preview-deploy
constraint and is deferred to #11.

Refs: #4 #11

Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi
2026-06-27 22:08:45 -04:00
seahaven-openswe[bot]
3af0bd5e16
ci: align workflows with Sea Haven CI/CD handbook (#29)
* Align workflows with Sea Haven CI/CD handbook

Bring the workflow suite in line with the handbook: bump
actions/checkout to v7 (Node 24 runtime, already standardized),
kebab-case the two snake_case workflow filenames, and add the
org-standard Labeler caller and Dependency Review gate so vulnerable
or disallowed-license deps and unlabeled PRs are caught automatically.

File renames only — job/check display names are unchanged, so the
promotion gate's REQUIRED_CHECKS and branch-protection required
checks are unaffected.

Refs: INFRA-115

* Drop Agent prefix from CI workflow + job names

The handbook names workflows for what they do (CI, Deploy, Labeler),
not the component they run, matching .github and afterhours-shift-manager.
Rename the suite to CI and its jobs to Lint / Format check / Unit tests,
and keep the promotion gate's REQUIRED_CHECKS in sync.

Refs: INFRA-115

---------

Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com>
2026-06-27 21:47:25 -04:00
Adam Moussa
3224d7abf4
ci: gate dev→prod promotion on green checks + add rollback safety net (#28)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
Agent CI / Agent lint (push) Waiting to run
Agent CI / Agent format check (push) Waiting to run
Agent CI / Agent unit tests (push) Waiting to run
Agent CI / Playwright E2E (push) Waiting to run
T20 CD safety nets. Two gaps closed before the first real prod deploy:

1. Promotion gate. promote_dev_to_prod.yml previously fast-forwarded
   dev→main unconditionally. It now hard-gates on check-dev-green.sh:
   every check-run on the dev HEAD must be completed+passing AND the
   Agent CI suite (lint/format/unit/E2E) must be present+success, or the
   promotion blocks (fails safe on a missing/renamed check). The promote
   run excludes its OWN check-run by run-id (unforgeable), never by the
   mutable name "promote", so a colliding red check cannot hide. Fields
   are read with a 0x1F separator so an empty conclusion (every
   in_progress check) cannot shift columns. ci.yml now also runs on
   push:dev so dev HEAD actually carries that signal (a PR check alone
   can be admin-merged past).

2. Rollback + last-good. publish-and-deploy.sh advances
   releases/last-good/ only after a successful roll (deploy.sh gates on
   `systemctl is-active`), and makes releases/latest/ transactional —
   reverting to the prior release if the roll fails so a replaced box
   never self-deploys a broken release. New rollback.yml + rollback.sh
   re-point latest at last-good (or an explicit sha) and re-fire the
   deploy; prod is gated by the `prod` Environment approval, same as a
   deploy. The shared fire/wait/aggregate-gate logic is factored into
   roll-box.sh (used by both forward and backward rolls).

Least-privilege: drop the unused s3:DeleteObject from the app deploy
role — publish/rollback/deploy only Get+Put (S3-to-S3 copy), and the
rollback fallback now depends on immutable release history staying
intact. Lifecycle expiry (not CI) handles old-version cleanup.

Gate logic unit-tested (7 cases + jq round-trip). IAM change +
release-safety control cross-reviewed by GPT-4.1: APPROVE, no blocks.

Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi
2026-06-27 20:21:59 -04:00
Adam Moussa
5e30bc6be2
Point dev agent at seahaven-open-swe-dev org + un-pin the owner guard (#27)
Dev now runs against the dedicated seahaven-open-swe-dev org (repo openswe-dev-sandbox),
isolated from the real Sea Haven org. Two changes:

- config-store.ts: iacManagedSsm repo targeting is now per-env — dev =
  seahaven-open-swe-dev/openswe-dev-sandbox, prod stays Sea-Haven-Industries/open-swe-pilot.
  ALLOWED_GITHUB_ORGS tracks the env owner. Adds a dev-only SEED_USER_MAPPINGS param so the
  triggering GitHub login (amoussa1229) resolves and @openswe comments aren't skipped.

- fetch-config.sh: replace the unconditional hard-pin to Sea-Haven-Industries with an owner
  GUARD that HONORS the configured owner (OPENSWE_REPO_OWNER override, else the SSM value) but
  forces a per-env safe org when the normalized owner is blank or the upstream langchain-ai.
  Normalization (lowercase, strip whitespace, first path segment, drop dots) catches
  langchain-ai/<repo>, langchain-ai., and case variants without over-blocking legit orgs
  (e.g. langchain-ai-fork). Fallback org is per-env so dev can't fall back into the real org.

Reviews: GPT-4.1 cross-review APPROVE (round 1 found a path/dot bypass -> hardened, round 2 clean);
/sh-security-review authz one LOW (env-invariant fallback) -> fixed. Positive org allowlist still
enforced by the app via ALLOWED_GITHUB_ORGS.
2026-06-27 19:29:03 -04:00
Adam Moussa
082768ff51
ci: print stack outputs from CDK --outputs-file (drop describe-stacks) (#26)
Some checks are pending
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
cd-infra reported failure on every successful deploy: the 'Stack outputs' step ran
'aws cloudformation describe-stacks' with the githubdeploy-open-swe-infra-<env> role,
which intentionally lacks cloudformation:DescribeStacks. The cdk deploy itself succeeds
(it reads outputs via the bootstrap cfn-exec role it assumes). Switch to
'cdk deploy --outputs-file cdk-outputs.json' + cat — no extra IAM grant, and the job
goes green on actual deploy success instead of masking real failures behind a red run.
2026-06-26 20:17:17 -04:00
Adam Moussa
f2633f9bd0
fix: don't crash-loop the box when no user mapping is configured (#25)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
seed_store.sh (ExecStartPost) exited 1 when it couldn't resolve a user_mappings
entry, which — with Type=simple — fails the whole unit and crash-loops the box. This
contradicts the script's own OSWE-SEED-03 precedent (the server-not-ready path exits
0 specifically to avoid a restart loop). A missing user mapping is a seeding gap, not
an unhealthy server: the @openswe trigger just won't resolve a commenter, which a
deployment-validation env (dev) does not need.

Make it non-fatal — warn and skip the user_mappings seed; team_settings is still
seeded and the service starts. The empty MAPPINGS array makes the seed loop a no-op.
Set SEED_USER_MAPPINGS / CONFIGURED_ADMINS / OPENSWE_OWNER_LOGIN+EMAIL to seed it.

Verified on the dev box: service active, langgraph bound to 127.0.0.1:2024 (loopback),
/ok 200, /healthz 200.
2026-06-26 20:06:53 -04:00
Adam Moussa
a2877d46ff
fix: paginate batch-get-secret-value in fetch-config (secrets dropped past page 1) (#24)
fetch-config materialized only the FIRST page of Secrets Manager results: the AWS CLI
does NOT auto-paginate batch-get-secret-value (the script's comment claiming it does
was wrong; --no-cli-pager only disables the output pager, not API pagination). With 28
secret shells under open-swe-<env>/ the first page returned ~10 items, stranding the
rest on later pages. Required secrets that landed past page 1 (DASHBOARD_JWT_SECRET,
TOKEN_ENCRYPTION_KEY, LANGSMITH_API_KEY_PROD) were silently dropped, tripping the
FAIL-FAST 'missing required var' guard and crash-looping open-swe.service.

Follow NextToken across pages (new batch_get_secrets_tsv helper). Verified against the
live open-swe-dev secrets: now loads all 5 populated secrets (was 2). Same per-record
base64 / exact-prefix / accept_var / emit_var hardening — only the page loop is new.
2026-06-26 19:55:01 -04:00
Adam Moussa
cdcb6625b9
fix: BatchGetSecretValue must be on * for the filtered batch call (#23)
The prior fix scoped secretsmanager:BatchGetSecretValue to the env-prefixed secret
ARN, but the live box still got AccessDenied: batch-get-secret-value invoked WITH a
name --filters is a COLLECTION call that AWS authorizes against * (a per-secret ARN
does not satisfy it). Split the statement:

- GetSecretValue + DescribeSecret stay PREFIX-scoped (secret:open-swe-<env>/*) — this
  is what gates which secret VALUES the box can read (checked per-secret in the batch).
- BatchGetSecretValue + ListSecrets move to a * operation-level statement (the filtered
  collection call + the list action; neither is resource-scopable for this usage).

VALUE isolation preserved (dev box still cannot read prod secret values); only secret
NAME/metadata enumeration is widened. GPT-4.1 IAM cross-review: BLOCK none, FIX none.
Suppression OSWE-IAC-SECRETS-LIST-01 updated; future hardening (explicit --secret-id-list
to drop both * grants) tracked there.
2026-06-26 19:44:12 -04:00
Adam Moussa
cfbdcda9b7
fix: instance role BatchGetSecretValue + ListSecrets for .env materialization (#22)
* fix(infra): grant instance role BatchGetSecretValue + ListSecrets for .env materialization

fetch-config.sh materializes the box's .env via
`secretsmanager batch-get-secret-value --filters Key=name,Values=open-swe-<env>/`,
but the instance role only granted GetSecretValue/DescribeSecret. BatchGetSecretValue
is a distinct IAM action, so the call was AccessDenied and open-swe.service
crash-looped (no .env written -> ExecStartPre exit 1).

- Add secretsmanager:BatchGetSecretValue to the prefix-scoped ReadSecrets statement.
- Add secretsmanager:ListSecrets on * (required by the name-prefix filtered batch
  call; the API has no resource-level scoping for the list action — fits the role's
  stated exception). Secret VALUES stay prefix-scoped; only names are enumerable.

Reviews: GPT-4.1 IAM cross-review BLOCK=none; /sh-security-review iac-iam one LOW
metadata residual (no critical/high), recorded as OSWE-IAC-SECRETS-LIST-01.

Refs T7/T19 dev bring-up.

* ci: lift Node heap cap for Playwright E2E build (vite OOM)

The E2E job's Playwright globalSetup runs the real `bun run build`, whose vite
bundle exceeds Node's default ~2 GB heap and OOMs (JavaScript heap out of memory) —
the same failure fixed for build-artifacts.yml in #19. Set
NODE_OPTIONS=--max-old-space-size=8192 on the Run E2E step.
2026-06-26 19:29:50 -04:00
Adam Moussa
5fa132205b
fix(deploy): use %%...%% for CDK user-data tokens (don't collide with @@ sed) (#21)
The systemd unit booted with a literal `@@OPENSWE_ENV@@` (fetch-config.sh got the
token, not "dev" -> exit 2 -> crash-loop) because user-data.sh is double-templated:
CDK substitutes @@tokens@@ AND user-data seds @@tokens@@ into the baked
systemd/nginx files. CDK's `.replace(/@@OPENSWE_ENV@@/g, "dev")` clobbered the sed
PATTERN (`s|@@OPENSWE_ENV@@|...|` -> `s|dev|...|`, a no-op), so the unit's token
never got replaced. Same collision hit @@SERVER_NAME@@ (masked by nginx
default_server).

Fix: CDK tokens move to a DISTINCT delimiter %%...%% (rendered in app-service.ts);
the @@...@@ tokens stay for the baked-template seds. No AMI rebuild (templates
unchanged). Add a guard test asserting no unresolved %%CDK%% token survives in the
synthesized user-data. Also add .github/scripts/** to build-artifacts paths so
script-only changes trigger a publish.

jest 20/20; tsc + shellcheck clean; rendered user-data: OPENSWE_ENV="dev",
SERVER_NAME="openswe-dev.seahaven.com", @@ sed patterns preserved, 16872 B.
2026-06-26 19:06:37 -04:00
Adam Moussa
9e2b215f08
fix(ci): avoid tar|grep -q SIGPIPE false-failure in package-artifacts (#20)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
`tar -tzf app.tar.gz | grep -qx` under set -o pipefail fails the pipeline when
grep -q matches and exits early (SIGPIPEs tar -> 'write error' -> non-zero), a
false 'missing agent/server.py'. List the archive once into a var, then grep the
var. Same fix for the secret-guard pipe (which was also silently broken).
2026-06-26 18:57:59 -04:00
Adam Moussa
2765b65acd
fix(ci): raise Node heap for the SPA build (vite OOM'd at ~2 GB) (#19)
The build-artifacts SPA build hit Node's default ~2 GB heap cap and aborted
(JS heap OOM, exit 134) — the same memory-hungry vite build that needed an 8 GB
swapfile on-box. The runner has ~16 GB, so set NODE_OPTIONS=--max-old-space-size=8192
on both build steps.
2026-06-26 18:54:10 -04:00
Adam Moussa
404b3f6f75
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18)
* feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3)

Packer-build the custom base image and repoint AppService off the AL2023
placeholder onto it.

deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real
`packer build` (the config had only ever been `packer validate`'d at T8):
  - the file provisioner failed uploading the templates dir ('scp: …: Is a
    directory') — a trailing-slash contents-upload needs the dest dir to exist;
    added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the
    dest trailing slash.
  - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the
    environment_vars never reached provision.sh (which runs under set -u and
    aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}.

infra:
  - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26
    from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by
    exact id via MachineImage.genericLinux (offline, deterministic). Dropped the
    now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement
    discipline docs.
  - app-service.ts: machineImage → bakedOpenSweArm64().
  - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard).
  - cdk.context.json → {} (AMI is a static id pin; no context lookups remain).
  - README: Baked AMI + EBS-replacement-discipline section.

tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI.

NOTE: held — do NOT merge until the open-swe-dev secret values are populated
(put-config.sh). The infra CD is live, so merging this to dev auto-deploys
OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts →
unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14).

* fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021

Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects
non-ASCII in the AMI Description attribute, so packer registered then
DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error.
Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available).
Re-pinned BAKED_OPEN_SWE_AMI_ID.

* fix(deploy): GitHub App + Slack required for prod only, not dev

Per the migration decision: do NOT create/duplicate a separate dev GitHub App or
Slack app — only prod owns the single shared app. So fetch-config.sh no longer
hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/
CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only
block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET.

Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active
provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env
(boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is
unchanged (prod still requires everything).

* feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19)

Make the dev/prod box deployable end-to-end: a real artifact pipeline and a
re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy.

Infra (T7):
- assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access,
  SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent +
  abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput.
- app-service.ts: open-swe-<env>-deploy SSM document that runs the baked
  /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the
  baked open-swe-base-arm64 AMI (folds in the held #16).

IAM (app deploy role — cross-review gated):
- github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to
  open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic
  AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document
  is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox.

Boot/deploy (T19):
- deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull
  app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv
  at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx.
- user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB
  target is healthy even before the first release); deploy.sh is base64-rendered
  by CDK into user-data (a normal reviewable repo file, not a heredoc) and the
  first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy).

CI (T7+T19):
- build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite ->
  ui/.output/public -> spa.tar.gz), package the Python source via git archive
  (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via
  the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy.
  push dev -> dev (auto); push main -> prod (env "prod" approval gate).

Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK,
deploy.sh base64 round-trips exact.

* harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard

Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one
confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01,
the account-wide CDK cfn-exec residual already documented in config.ts; recorded in
.security-review/suppressions.json with justification + flagged for the per-env
bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified;
these are the cheap defense-in-depth fixes worth taking regardless:

- deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root
  never honors an archive's uid/mode → no setuid/foreign-owned file can land); and
  treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy
  failure (set -e stays loud once a release exists).
- publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount),
  not CommandInvocations[0], so a partial failure across the brief 2-instance
  replacement window can't be reported as success.
- instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app
  role's write scope) instead of the whole bucket.
- package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz
  (defense in depth over .gitignore; scoped to data extensions so *_credentials.py
  source is not a false positive — verified against the real tree).

Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the
CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable
releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release
to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz).

shellcheck/tsc/jest(16) clean; both stacks synth offline.

* fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard

The instance-SG GroupDescription + ingress/egress rule descriptions carried an
em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects
non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"),
so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing
from #14; same class as the AMI-description ASCII bug.)

- app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the
  ingress/egress rule descriptions, and the Route53 comment.
- test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup
  GroupDescription + rule descriptions are pure ASCII, so this fails the build
  instead of a deploy next time.

jest 18/18; tsc clean.

* fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`)

The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule*
descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`,
which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the
ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII
only" to the exact EC2 allowed charset so it catches `>` (and `<`) too.

jest 18/18; tsc clean.

* fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit

The base64 deploy.sh embedded in user-data pushed the encoded boot script to
27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with
"Encoded User data is limited to 25600 bytes". Strip full-line comments + blank
lines from deploy.sh before base64-embedding it (repo file keeps comments; only
the on-box copy is minified; the script is opaque base64 so user-data heredocs are
unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a
synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded.

jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
Adam Moussa
70319cac0d
ci: path-filtered infra CI/CD + dual OIDC roles + prod approval gate (T18) (#15)
Some checks are pending
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
* ci: path-filtered infra CI/CD with dual OIDC roles + prod approval gate (T18)

Add the /infra half of the combined-repo pipeline (the Python agent keeps ci.yml):

- ci-infra.yml — PR check on infra/** : tsc + jest + cdk synth via the org
  reusable ci-typescript-cdk.yaml (working-directory: infra).
- cd-infra.yml — push to dev/main on infra/** (or dispatch):
    * job 'ci' (reusable) is the CI-green precondition (deploy needs: ci).
    * deploy-dev  (ref=dev, NO environment)  → cdk deploy OpenSweDevStack,
      assuming githubdeploy-open-swe-infra-dev  (OIDC sub ref:refs/heads/dev). AUTO.
    * deploy-prod (ref=main, environment: prod) → cdk deploy OpenSweProdStack,
      assuming githubdeploy-open-swe-infra-prod (OIDC sub environment:prod). The
      'prod' Environment's required reviewer is the manual-approval gate.

Deliberately self-contained (NOT the reusable cd-cdk.yaml) because that runs
'cdk deploy --all' — from a single-env push it would deploy the other env + the
shared IAM stack, breaking the per-env boundary. CD targets one stack per env;
the shared open-swe-iam stack is human-gated (T6), never deployed by CD.

Infra CI is enforced at the DEPLOY boundary (deploy jobs need ci), not as a
branch-protection required check — path-filtering a required check would deadlock
app-only PRs. Documented in infra/README.md along with the post-T6 prerequisites
(repo vars AWS_DEPLOY_ROLE_INFRA_{DEV,PROD}; a 'prod' Environment w/ reviewer).

Not active until the IAM roles are applied (T6) — assuming a nonexistent role
just fails closed. App-side CD (S3 artifact + SSM) is T19.

* fix(infra): commit jest.config.js (was ignored by *.js → infra CI used Babel)

The infra/.gitignore *.js rule (for compiled CDK output) silently swept up the
hand-authored jest.config.js, so it was never committed. Local jest passed (file
present in the working tree) but CI's fresh checkout lacked it → jest fell back to
the default Babel transform → 'Cannot use import statement outside a module' on the
TypeScript test. Surfaced now because T18 is the first workflow to run infra jest
in CI. Negate the ignore for this one file and commit it.
2026-06-26 16:11:22 -04:00
Adam Moussa
fcbdfb67aa
feat: open-swe dev/prod compute + ALB ingress (T12) (#14)
AppService construct wires the per-env EC2 box and its internet path. The
seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the
on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and
never owned/mutated; open-swe only ADDS its own resources.

Per env (open-swe-stack.ts → AppService):
- ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a,
  in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination
  (no RETAIN volume — replacement-tolerant; see ami-cache.ts).
  userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh.
- Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress
  via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress
  (the imported, on-prem-owned SG is never mutated).
- Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays
  loopback). Health check GET /healthz.
- Two rules on the imported :443 listener, both → the TG:
    * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/*
    * site     (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api)
  Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5)
  or it would steal every webhook — first-match-by-ascending-priority.
- Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB.
- 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config).

Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed
medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large
GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on
/dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene);
XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed
critical/high.

Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked
open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean.
Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
Adam Moussa
aef6b26912
feat: Secrets Manager + SSM config store for open-swe (T11) (#10)
Create the per-env config surface the EC2 box reads at boot via
fetch-config.sh / seed_store.sh:

- ConfigStore construct (infra/lib/constructs/config-store.ts):
  - 28 value-LESS Secrets Manager shells  open-swe-<env>/<VAR>
    (RemovalPolicy.RETAIN, no SecretString/generateSecretString — real
    values are set out-of-band by put-config.sh, never in IaC/state).
  - 8 IaC-managed SSM params  /open-swe-<env>/<VAR>  with real,
    stable/derivable values (SANDBOX_TYPE, DEFAULT_REPO_OWNER/NAME,
    ALLOWED_GITHUB_ORGS, DASHBOARD_*_URL/ORIGINS, LLM_MODEL_ID).
  - OUT_OF_BAND_SSM documents the ~30 params CDK intentionally does NOT
    own (operationally-variable / env-specific-unknown).
- Wire ConfigStore into OpenSwe<Env>Stack.
- KebabNamingAspect: exempt Secrets Manager + SSM names, which carry the
  literal UPPER_SNAKE env-var segment (open-swe-dev/DASHBOARD_JWT_SECRET).
- deploy/seahaven/put-config.sh: out-of-band populator (placeholders only,
  OPENSWE_PUT_<VAR> env indirection; no real values committed).

Synth-only; not deployed. Instance-role read grants on open-swe-<env>/*
already exist from T6 — no IAM/trust changes here.
2026-06-26 16:07:23 -04:00
Adam Moussa
434d0ad80c
feat(deploy): AWS-sourced fetch-config + seed_store + rotation docs (PR#7) (#8)
PR#7 of the AWS migration. deploy/seahaven/: fetch-config.sh materializes a
service-user-owned 0600 tmpfs .env from Secrets Manager + SSM (fail-fast);
seed_store.sh reseeds the in-memory store; ROTATION.md.

Incorporates T5 /sh-security-review fixes:
- seed_store no longer bash-sources the .env (closes the SH-INJ-001 RCE); uses a
  non-eval reader, jq --arg JSON bodies, and a loopback-pinned BASE.
- fetch-config: .env owned by the openswe service user (app no longer runs as
  root); DEFAULT_REPO_OWNER hard-pinned; key-identifier validation + flat-namespace
  collision detection; dropped SSM --recursive.
shellcheck + bash -n clean.
2026-06-26 15:06:41 -04:00
Adam Moussa
ce9e5fb51b
feat(deploy): EC2 AMI recipe (Packer) + cloud-init/user-data (PR#2) (#7)
PR#2 of the AWS migration. deploy/ami/: Packer template (Ubuntu 24.04 arm64,
uv+py3.12, nginx, awscli v2, CW agent; no swapfile), provisioning-only user-data
(userDataCausesReplacement rationale), systemd unit + nginx + CW templates.

Incorporates T5 /sh-security-review fixes: langgraph binds 127.0.0.1 (not 0.0.0.0);
nginx is the sole ingress proxying only /dashboard/api/ + /webhooks/; ExecStartPre
runs fetch-config as root (+) and passes the env arg; the app runs as the
unprivileged openswe user reading an openswe-owned 0600 .env. packer validate clean.
2026-06-26 15:06:36 -04:00
Adam Moussa
92fd886076
feat(infra): add /infra CDK scaffold + per-env OIDC/instance IAM role defs (#6)
PR#1 of the AWS migration. CDK TypeScript app under /infra: stacks open-swe-iam
(four per-env GitHub-OIDC deploy roles) + open-swe-dev/-prod (per-env EC2 instance
role + AMI cache). aws-cdk-lib pinned exact 2.260.0; kebab naming Aspect + 15 tests.

IAM is synth-only (NOT deployed). Cleared the Phase-1 security gates:
- T4 GPT-4.1 IAM cross-review (StringEquals trust; cdk-hnb659fds-* wildcard kept as
  org convention; AWS-RunShellScript timeboxed to T19).
- T5 /sh-security-review: deploy roles split PER-ENV with env-scoped OIDC trust
  (dev=branch ref+tag dev, prod=environment:prod+tag prod) so a dev token cannot
  reach prod; re-verified block:false.
2026-06-26 15:06:30 -04:00
Adam Moussa
5a5818f4c2
ci: promote dev → main (prod), not a separate prod branch (#3)
Aligns the promotion target with the locked migration architecture
(main = PROD, dev = integration). Plain ref push is fast-forward-only;
branch protection on main rejects non-FF pushes, so a diverged main
fails loudly instead of force-moving prod.
2026-06-26 13:49:02 -04:00
Adam Moussa
35bc9e22f4
ci: promote dev → prod via fast-forward (not force-push from main) (#2)
* ci: promote dev to prod via fast-forward instead of force-pushing main

Repoints the daily promotion workflow to feed prod from the dev integration
trunk (prod <- dev) rather than force-pushing the upstream mirror (main).
Uses a fast-forward push so a release can never rewrite prod history.

* fix: sort imports in aegra_entry.py to unblock dev CI

The Sea Haven deployment commit left aegra_entry.py with an unsorted import
block (ruff I001), which fails the required 'Agent lint' check and blocks all
merges into dev. Apply the import-sort autofix.
2026-06-26 11:20:10 -04:00
cf57ba7aa5 feat(deploy): add Sea Haven self-hosted deployment capture
Captures the stock-LangGraph deployment of this fork at Sea Haven:
- systemd/open-swe.service: langgraph dev (:2024) + store seed ExecStartPost
- seed_store.sh: re-seeds team_settings + user_mappings (in-memory store
  resets on restart); env-parameterized, no secrets
- nginx/openswe.conf: dashboard SPA + scoped /dashboard/api proxy (security
  boundary; agent API not exposed)
- aegra/: deferred self-hosted-runtime alternative (not active on stock)
- DEPLOYMENT.md: full runbook (models, build, ingress, OAuth callback)

Secrets and internal infra identifiers are intentionally excluded (public
fork); real values live in private IT docs.
2026-06-25 17:28:34 -04:00
Johannes du Plessis
f6c215fff7
feat: add automation template gallery (#1608)
Add a "Start from a template" gallery to the Automations page so users can
scaffold common scheduled agent runs (PR review digest, issue triage,
dependency checks, flaky-test tracking, release notes, docs freshness,
security audit) instead of writing every automation from scratch. Selecting
a template opens the new-automation editor prefilled with its instructions
and a default schedule, which the user can tweak before saving.

Also adds an isDescribableCron helper so template schedules render as
human-readable labels rather than raw cron in the editor.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-25 10:36:34 -07:00
Johannes du Plessis
fb467e84fd
fix: gate Open SWE in docs-plz Slack channel (#1605)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-24 13:00:09 -07:00
Johannes du Plessis
39681102d6
fix: repair orphaned tool calls before model calls (#1604)
When a run is cancelled or the sandbox dies mid-tool-call, LangGraph persists
the AIMessage tool_call but never the matching ToolMessage. The next run sends
the provider an orphaned tool_use (Anthropic 400: "tool_use ids were found
without tool_result blocks"), permanently wedging the thread on every retry.

Add RepairOrphanedToolCallsMiddleware, which inserts a synthetic error
ToolMessage immediately after any tool_call lacking a result so the agent can
retry instead of dying. Wired into the agent and reviewer graphs.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-24 12:58:10 -07:00
Ramon Nogueira
534f402f48
feat(open-swe): copy-plan-as-markdown button + fix premature "ready" banner (#1603)
- Add a "Copy markdown" button to the plan header that copies the whole plan.
  Cross-browser: async Clipboard API in secure contexts, hidden-textarea +
  execCommand fallback for older Safari/Firefox and non-HTTPS origins.
- Fix the conversation banner: it showed "A plan is ready for your review" for
  every non-approved/cancelled status, including "planning" — so it claimed the
  plan was ready the instant plan mode began (the agent shares the link early
  to follow along), then the plan page correctly said it was still being
  written. Now: planning → "writing a plan", revising → "revising", ready →
  "ready for your review".
2026-06-24 15:07:25 -04:00
Ramon Nogueira
714914b14c
feat(open-swe): disable "Request changes" until the plan has a comment (#1602)
Requesting changes hands the reviewer comments to the agent, so it's
meaningless with none. Disable the button (with a hint tooltip) until at
least one comment exists; approve is unaffected.
2026-06-23 16:02:08 -07:00
Johannes du Plessis
9370a8c7f4
feat: inline PR comments in the reviews UI (#1600)
* feat: inline PR comments in the reviews UI

Click the diff gutter "+" on a line to open an inline comment composer
(rendered like the finding card via a Pierre annotation); submitting
posts a real inline PR review comment as the signed-in user through a
new POST /reviews/{owner}/{repo}/{number}/comments. The "+" press-drag →
"Add to Chat" selection path is unchanged.

* feat: GitHub-parity comment box, PR comments dropdown, collapse nav

- Comment composer now mirrors GitHub's box: Write/Preview tabs (markdown
  rendered via the existing Markdown component) and a markdown toolbar
  (heading, bold, italic, quote, code, link, bulleted/numbered/task list).
- Surface other people's inline PR comments in a Devin-style dropdown in the
  review header (search + link to the thread on GitHub). New
  GET /reviews/{owner}/{repo}/{number}/comments lists them and flags the
  reviewer's own (marker-bearing) comments so they're filtered out.
- Collapse the global nav by default on a review detail page, restoring the
  prior preference on leave.

* feat: bigger comment-toolbar icons; open dropdown comments inline

- Enlarge the markdown toolbar glyphs (Phosphor) in the comment composer —
  they were rendering at 10px.
- Clicking a comment in the PR comments dropdown now opens it inline in the
  diff as a read-only finding-style card (InlineComment), scrolling its line
  into view, instead of navigating to GitHub. Falls back to GitHub when the
  comment's file/line isn't in the current diff.

* fix: drive "Add to Chat" from native text selection

The gutter "+" is now comment-only; wiring its click to the composer
conflicted with its old double-duty as the drag-to-select handle, which
broke selection → "Add to Chat". Switch to Devin's model: disable Pierre's
interactive line selection and instead map a native text highlight in the
diff to a line range (via the data-line / data-line-type attributes Pierre
stamps on each line, read from the diff's open shadow root) to show the
"Add to Chat" popup. ⌘L and the existing attachment/popup path are unchanged.

* feat: gutter "+" drag selects a range for multi-line comments

Re-enable Pierre's gutter line selection so dragging the "+" down the
gutter comments across a range (click still comments on a single line);
onLineSelectionEnd routes the range to the composer. Native code-text
selection still drives "Add to Chat" — Pierre only line-selects from the
gutter, and onLineSelectionEnd bails when a native text selection is
present, so a code highlight never opens the composer.

* fix: keep the range highlighted while its comment composer is open

Previously opening the composer cleared the selection, so the lines being
commented on lost their highlight. Drive the controlled selection from the
open comment draft's range so the rows stay highlighted until the composer
is closed.

* fix: address PR review — paginate comments, fall back for outdated ones

- list_review_comments now pages through all PR review comments (bounded by
  _MAX_REVIEW_COMMENT_PAGES) instead of returning only the first 100, so older
  comments still show in the dropdown.
- Surface GitHub's outdated flag (position == null) as is_outdated; opening such
  a comment (or one whose line isn't in the diff) now opens it on GitHub instead
  of silently rendering nothing, plus a timeout fallback if the annotation never
  mounts (e.g. collapsed context).
2026-06-23 16:01:46 -07:00
Ramon Nogueira
ca9280d25c
refactor(open-swe): plain HTTP comments instead of Yjs/BlockNote collab (#1601)
* feat: add plan mode for read-only research and planning

Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: enforce plan-mode read-only at tool layer and disable subagents

Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: harden plan-mode shell guard against wrapped mutations

Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow

- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline

Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.

* test(open-swe): add Playwright E2E for the Slack → PR → web handoff

Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.

- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).

Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.

* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright

The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.

Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.

* test(open-swe): record Playwright trace + video on every E2E run

Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.

* feat(plan-mode): collaborative plan review with BlockNote + Yjs

When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.

- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
  the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
  snapshots; plan content/status store; plan REST API (get/approve/reject,
  owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
  mounted under the agents shell, with a "Review plan" banner in the thread view
  and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
  flow, including cross-user comment sync and owner-only approval.

* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)

- plan-collab WS: authorize per-thread before joining a room (same read gate as
  the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
  opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
  resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
  dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
  unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
  before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.

Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).

* style: ruff format plan_collab.py

* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS

- Slack "Approve & Implement" now verifies the clicking user is the plan
  requester (owner, via the stored triggering_user_id) before implementing —
  matching the dashboard API's owner-only approval. Non-owners are pointed to
  Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
  allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
  the REST require_same_origin CSRF defense.

* fix(plan-mode): enter plan mode only via the model + local mock dev harness

Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.

- enter_plan_mode returns a terminating ToolMessage, fixing the missing
  ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
  remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
  call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
  Alice/Bob mock users, and a GitHub login picker.

* docs(plan-mode): drop stale references to removed profile/team defaults

The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.

* feat(plan-mode): let any reviewer edit the plan, not just comment

Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.

* test(plan-mode): assert plan-mode entry via the tool's success message

plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.

* refactor(plan-mode): replace Yjs/BlockNote collab with plain HTTP comments

Drop the realtime collaborative editor (it can't work behind Vercel's
rewrite — WebSocket upgrades aren't proxied to the external LangGraph
backend) in favor of a simple whole-document comments API over plain HTTP.

Backend:
- Remove the Yjs WebSocket server (plan_collab.py), its lifespan, and the
  collab router; drop pycrdt / pycrdt-websocket deps.
- plan_store: replace the Yjs snapshot with comment CRUD (one store item per
  comment under ["plan","comments",thread_id]).
- plan_api: add GET/POST/DELETE comment endpoints; approve/reject now read
  comments server-side and format them for the follow-up run (no longer
  client-harvested). Comment delete is author-or-owner; approve stays owner-only.

Frontend:
- PlanReview renders the plan markdown read-only and shows a comments panel
  (list + add, polled every 4s for cross-user visibility).
- Drop @blocknote/*, y-websocket, yjs; lib/plan exposes get/add/deletePlanComment.

Tests: unit tests for the comments API + route registration; e2e drives the
HTTP comment UI (owner + collaborator, cross-user visibility, owner-only approve,
PR echoes the harvested feedback).

* fix(open-swe): clear stale plan comments on republish; fail loud on store errors

Address reviewer feedback:
- Clear comments when a revised plan is published (save_plan_content) so
  feedback on the prior revision doesn't resurface and get re-fed to the agent.
- list_plan_comments gains raise_on_error; approve/reject read comments before
  mutating state and propagate store failures (500) instead of silently
  dispatching the follow-up run with no feedback.

---------

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 22:12:42 +00:00
Johannes du Plessis
4c7d45273f
perf: cut review-chat time-to-first-token (#1598)
* perf: cut review-chat time-to-first-token

The sandbox-less PR review chat paid several blocking network round-trips
before the first token on every message. Cache GitHub App installation
tokens in-process (per scope, until ~10m before expiry, above the proxy's
5m refresh window) so the chat graph factory and proxy stop re-minting one
each turn. Also drop the duplicate thread-metadata read in the commands
proxy and replace the heavy per-message get_review staleness check with a
single lightweight PR head-SHA lookup.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: keep review chat alive when reseed fails

Address review: a moved PR head now triggers _build_pr_context (and thus
get_review). For an existing chat, fall back to the last seeded context on
HTTPException instead of failing the command; fresh chats still surface the
error since they have no prior context.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:32:10 -07:00
Caroline di Vittorio
63e4c2baac
feat: relocate review chunks and embed review in git panel (#1590)
* feat: relocate review chunks and embed review in git panel

Move the AI-sorted review chunk navigator out of the global left sidebar
and into the review page's own main panel (sidebar keeps the thread list),
and surface the review main body inside the agent thread's Git > Review
sub-tab with an expand link to the full review page.

The review main body + side panel + helpers are extracted into a shared
ReviewMainBody component with "full" and "embedded" variants.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: move ReviewTab into its own file

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:27:51 -07:00