mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 10:23:14 +00:00
1068 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2a226e2cd1
|
fix(ui): pin nitro to patched 3.0.260603-beta (#116)
Resolves Dependabot GHSA-9phm-9p8f-hw5m (open redirect via protocol-relative URL in wildcard route rules) and GHSA-5w89-w975-hf9q (proxy scope bypass via percent-encoded path traversal in routeRules). nitro was pinned to "latest", which Dependabot can't resolve to a fixed version, so the alerts stayed open even though the lockfile already resolved 3.0.260603-beta (newer than the 3.0.260429-beta patch line). Pin it to an exact version — nitro ships a date-stamped beta channel where caret ranges behave unpredictably — so installs are reproducible and both alerts close. bun.lock also reconciles @pierre/trees beta.4 -> beta.5, which the manifest already declared but the committed lockfile was stale on. |
||
|
|
b65c3a07db
|
feat: distill Sea Haven conventions into agent prompt, reviewer, and fork docs (#113)
* Add Dependabot ignore for @types/node semver-major bumps Prevent Dependabot from proposing wrong-direction @types/node major bumps (e.g. 24 -> 26). /ui runs on Node 24 on Vercel; a too-new types major still compiles but describes APIs absent at runtime. Refs: #110 * feat(agent): seed all-repos custom instructions in default_prompt.md Distill the universally-applicable Sea Haven authoring conventions into the team-default Custom Instructions the main agent gets on every repo: secrets/ config placement, keep-docs-in-sync, verify-before-push, re-run-real-gates after delegating, confirm-a-convention-before-adopting, and house writing style. Toolchain references are generalized (not tied to a specific stack). * feat(reviewer): seed Sea Haven review baseline as org-guidelines default Bake DEFAULT_ORG_REVIEW_GUIDELINES (severity model, secrets, security surface, tests, naming, deferred-work-needs-an-issue) and default org_guidelines to it in _default_settings(). The reviewer now applies the Sea Haven baseline on every repo until a workspace admin overrides it with a non-empty value via the dashboard. Stack-agnostic and well under the 10k-char cap. * refactor(prompt): consolidate duplicated COMMIT_PR_SECTION + add fork-sync runbook COMMIT_PR_SECTION had two overlapping passes with a contradictory PR-title rule (a fixed 'type: description' form vs the repo-aware detection). Collapse into one numbered sequence (lint -> commit -> push/PR -> notify), keep the authoritative repo-aware title rule, and drop the duplicate notify step. All IMPORTANT directives (force-push ban, workflow-approval, autonomy, 403 handling) are preserved verbatim. Add a fork-maintenance runbook to CLAUDE.md distilling the durable upstream-sync methodology (conflict triage, deferred-refactor resolution rule, the silent re-import/wiring hazards, test-impl-same-side, layered CI). * refactor(prompt): adopt conventional-commit style Flip the Sea Haven authoring convention baked into the agent prompt from imperative/no-prefix to conventional-commit style: - Commit subjects and the no-gate PR-title default now use type(scope): description with the allowed type set (feat, fix, docs, style, refactor, perf, test, build, ci, chore, revert, release). - Branch prefixes expanded to feature/, fix/, hotfix/, chore/, docs/, refactor/, release/ (kebab-case description). - The repo-aware gate detection is preserved: a repo's own title gate still wins and may narrow the allowed types/scopes. Updated test_github_comment_prompts.py to assert the new convention. |
||
|
|
34c5da7328
|
Add Dependabot ignore for @types/node semver-major bumps (#112)
Prevent Dependabot from proposing wrong-direction @types/node major bumps (e.g. 24 -> 26). /ui runs on Node 24 on Vercel; a too-new types major still compiles but describes APIs absent at runtime. Refs: #110 Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> |
||
|
|
f43606309e
|
fix(ui): clear eslint errors from dep bumps + track .env.example (#111)
* fix(ui): resolve eslint errors surfaced by dep bumps Clear the 6 eslint errors introduced by the typescript-eslint / TS 6.0 dependency bumps: - sidebarPrefs.ts, vite.config.ts: inline type imports -> top-level type-only imports; sort import members (autofix) - cloud-agents.tsx: drop optional chain on session.data (already narrowed non-null by the earlier login guard) - vite.config.ts: drop optional chain on proxyHead (typed Buffer from the upgrade event, always present) No behavior change. eslint and tsc --noEmit both clean. * docs(env): track .env.example and document deployment vars - Un-ignore .env.example (!.env.example) so the template stays tracked; .env / .env.* secret files remain ignored. - Add LLM_MODEL_ID, AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY (used by the Bedrock/managed deployment but previously undocumented). - Mark LANGCHAIN_TRACING_V2 and LANGCHAIN_PROJECT as reserved on LangGraph Platform (the platform sets them; deploys are rejected if provided). |
||
|
|
232f22f9d0
|
chore(deps): bump @pierre/trees from 1.0.0-beta.4 to 1.0.0-beta.5 in /ui (#106)
Bumps @pierre/trees from 1.0.0-beta.4 to 1.0.0-beta.5. --- updated-dependencies: - dependency-name: "@pierre/trees" dependency-version: 1.0.0-beta.5 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
b129ba6654
|
chore(deps): bump the minor-and-patch group with 3 updates (#104)
Bumps the minor-and-patch group with 3 updates: [fastapi](https://github.com/fastapi/fastapi), [langsmith](https://github.com/langchain-ai/langsmith-sdk) and exa-py. Updates `fastapi` from 0.138.2 to 0.139.0 - [Release notes](https://github.com/fastapi/fastapi/releases) - [Commits](https://github.com/fastapi/fastapi/compare/0.138.2...0.139.0) Updates `langsmith` from 0.9.4 to 0.9.6 - [Release notes](https://github.com/langchain-ai/langsmith-sdk/releases) - [Commits](https://github.com/langchain-ai/langsmith-sdk/compare/v0.9.4...v0.9.6) Updates `exa-py` from 2.15.0 to 2.16.0 --- updated-dependencies: - dependency-name: fastapi dependency-version: 0.139.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: minor-and-patch - dependency-name: langsmith dependency-version: 0.9.6 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: minor-and-patch - dependency-name: exa-py dependency-version: 2.16.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: minor-and-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
1bea7ae2d2
|
feat: Add dashboard UI for workflow push approvals (#103) | ||
|
|
a7ddd79e98
|
feat: prompt-mode PWA update toast instead of autoUpdate reload (#102) | ||
|
|
47a9a1ba89
|
feat: Loosen workflow push approval fingerprint to repo/branch/files identity (#101) | ||
|
|
421290d066
|
fix: abort approved workflow push when proxy token elevation fails [closes #97] (#100)
* Abort approved workflow pushes when proxy elevation fails When _run_with_workflow_token cannot elevate the sandbox proxy token to workflows:write, it now returns a clear ToolMessage error instead of running the push over the base token and getting a raw GitHub remote rejection. Refs: #97 * Fix workflow push guard crash and non-langsmith regression - Thread the ToolCallRequest into _run_with_workflow_token so the WorkflowPushElevationFailed ToolMessage is stamped with the real tool_call_id instead of an empty id that crashes the Anthropic API. - Only perform the elevation/abort path on SANDBOX_TYPE=langsmith; other providers run the approved push directly. - Add tests covering the real tool_call_id, single refresh call, and non-langsmith approved pushes. Refs: #97 --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> |
||
|
|
a52ebed77c
|
chore: align docs, templates, and metadata with Sea Haven handbook [closes SH-93] (#95)
* Align repo docs, templates, and metadata with Sea Haven handbook - Replace dynamic upstream-fork badges with static License, Python, and TypeScript badges; add the GitHub-native CI badge for dev. - Repoint SECURITY.md contact to adam@seahavenind.com. - Add .github/CODEOWNERS assigning reviews to @amoussa1229. - Add PR template using the handbook structure. - Add issue templates for bug, feature, and task plus a security link. - Remove emoji from the README per CONTRIBUTING.md style. Refs: SH-93 * Restore Linear 👀 acknowledgement and update metadata contacts - README.md: bring back the documented 👀 reaction on Linear issues. - README.md: refresh the TypeScript badge to 6.0+ to match ui/package.json. - SECURITY.md + issue template: switch security contact to role alias. Refs: SH-93 --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> Co-authored-by: Adam Moussa <adam@seahavenind.com> |
||
|
|
2335ac59c0
|
fix: switch PWA service worker to autoUpdate to avoid stale bundles after deploy (#94)
* Standardize remaining Node pins on Node 24 Refs: #40 - Bump sandbox Dockerfile to Node 24.18.0-1nodesource1 + setup_24.x - Move Playwright E2E job to node-version: "24" (quoted) - Update INSTALLATION.md to recommend Node 24 * Switch PWA service worker to autoUpdate to avoid stale bundles after deploy VitePWA's registerType: 'prompt' kept the new service worker waiting indefinitely because no update-prompt UX was wired. Returning visitors kept receiving the old precached app shell and hashed chunks. 'autoUpdate' uses skipWaiting + clientsClaim so a deploy takes over on the next navigation. Refs: #86 * Update PWA registration comment to match autoUpdate behavior The previous comment described prompt-mode behavior, but the SW is now configured with autoUpdate. Document the reload tradeoff explicitly so the comment no longer contradicts the VitePWA setting. Refs: #86 --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> |
||
|
|
98fa659303
|
feat: Standardize remaining Node pins on Node 24 (#92)
* Standardize remaining Node pins on Node 24 Refs: #40 - Bump sandbox Dockerfile to Node 24.18.0-1nodesource1 + setup_24.x - Move Playwright E2E job to node-version: "24" (quoted) - Update INSTALLATION.md to recommend Node 24 * Address review feedback on Node 24 standardization Refs: #40 - Align @types/node with Node 24 major (^24.0.0) - Add root .nvmrc declaring Node 24 - Harden INSTALLATION.md to require Node 24 for local UI dev * ci: add Docker build smoke job to validate sandbox image * Address follow-up review feedback on Node 24 standardization Refs: #40 - Fix INSTALLATION.md to reference bun and ui/bun.lock - Regenerate ui/bun.lock so @types/node ^24.0.0 is reflected --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> Co-authored-by: Adam Moussa <adam@seahavenind.com> |
||
|
|
93c2eba47f
|
chore(deps-dev): bump @vitejs/plugin-react from 5.2.0 to 6.0.3 in /ui (#70)
Bumps [@vitejs/plugin-react](https://github.com/vitejs/vite-plugin-react/tree/HEAD/packages/plugin-react) from 5.2.0 to 6.0.3. - [Release notes](https://github.com/vitejs/vite-plugin-react/releases) - [Changelog](https://github.com/vitejs/vite-plugin-react/blob/main/packages/plugin-react/CHANGELOG.md) - [Commits](https://github.com/vitejs/vite-plugin-react/commits/plugin-react@6.0.3/packages/plugin-react) --- updated-dependencies: - dependency-name: "@vitejs/plugin-react" dependency-version: 6.0.3 dependency-type: direct:development update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
68e02e58e0
|
Bump minor-level dependencies and fix FastAPI route test (#91)
Apply all minor-version bumps from the Dependabot group: - fastapi 0.136.3 -> 0.138.2 - uvicorn 0.48.0 -> 0.49.0 - langchain-openai 1.2.2 -> 1.3.3 - exa-py 2.13.0 -> 2.15.0 - langchain-mcp-adapters 0.2.2 -> 0.3.0 - pytest 9.0.3 -> 9.1.1 FastAPI 0.137+ changed included routers to be stored as _IncludedRouter wrappers instead of flattened routes. Update test_plan_routes_registered to collect paths from the original router so the plan routes remain discoverable. Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> |
||
|
|
901e66ba17
|
Bump patch-level dependencies from group update (#89)
Apply only the patch-version bumps from the minor-and-patch group PR, leaving the minor bumps (fastapi, uvicorn, langchain-openai, exa-py, langchain-mcp-adapters, pytest) on dev. Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> |
||
|
|
735471afc8
|
chore(deps-dev): bump vite from 7.3.6 to 8.1.2 in /ui (#71)
Bumps [vite](https://github.com/vitejs/vite/tree/HEAD/packages/vite) from 7.3.6 to 8.1.2. - [Release notes](https://github.com/vitejs/vite/releases) - [Changelog](https://github.com/vitejs/vite/blob/main/packages/vite/CHANGELOG.md) - [Commits](https://github.com/vitejs/vite/commits/v8.1.2/packages/vite) --- updated-dependencies: - dependency-name: vite dependency-version: 8.1.0 dependency-type: direct:development update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
a55bac8a2b
|
chore(deps-dev): bump typescript from 5.9.3 to 6.0.3 in /ui (#73)
Bumps [typescript](https://github.com/microsoft/TypeScript) from 5.9.3 to 6.0.3. - [Release notes](https://github.com/microsoft/TypeScript/releases) - [Commits](https://github.com/microsoft/TypeScript/compare/v5.9.3...v6.0.3) --- updated-dependencies: - dependency-name: typescript dependency-version: 6.0.3 dependency-type: direct:development update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
9eceacb6dc
|
chore(deps-dev): bump @types/node from 22.20.0 to 26.0.1 in /ui (#72)
Bumps [@types/node](https://github.com/DefinitelyTyped/DefinitelyTyped/tree/HEAD/types/node) from 22.20.0 to 26.0.1. - [Release notes](https://github.com/DefinitelyTyped/DefinitelyTyped/releases) - [Commits](https://github.com/DefinitelyTyped/DefinitelyTyped/commits/HEAD/types/node) --- updated-dependencies: - dependency-name: "@types/node" dependency-version: 26.0.1 dependency-type: direct:development update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
e9271c946f
|
ci: allow deps-dev, ui, deploy, ci, infra scopes in PR title lint (#88)
* ci: allow deps-dev, ui, deploy, ci scopes in PR title lint Dependabot dev-dependency PRs use the deps-dev scope and humans commonly use ui/deploy/ci, none of which were in the allowed list, so lint-pr-title failed on otherwise-valid PRs. * ci: allow infra scope in PR title lint |
||
|
|
7ac4babe71
|
chore(deps): bump actions/upload-artifact from 4 to 7 (#68)
Bumps [actions/upload-artifact](https://github.com/actions/upload-artifact) from 4 to 7. - [Release notes](https://github.com/actions/upload-artifact/releases) - [Commits](https://github.com/actions/upload-artifact/compare/v4...v7) --- updated-dependencies: - dependency-name: actions/upload-artifact dependency-version: '7' dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
c9e6df76d1
|
chore(deps): bump actions/dependency-review-action from 4 to 5 (#67)
Bumps [actions/dependency-review-action](https://github.com/actions/dependency-review-action) from 4 to 5. - [Release notes](https://github.com/actions/dependency-review-action/releases) - [Commits](https://github.com/actions/dependency-review-action/compare/v4...v5) --- updated-dependencies: - dependency-name: actions/dependency-review-action dependency-version: '5' dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
f9b02a64d6
|
chore(deps): bump actions/setup-node from 4 to 6 (#66)
Bumps [actions/setup-node](https://github.com/actions/setup-node) from 4 to 6. - [Release notes](https://github.com/actions/setup-node/releases) - [Commits](https://github.com/actions/setup-node/compare/v4...v6) --- updated-dependencies: - dependency-name: actions/setup-node dependency-version: '6' dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
2b01652754
|
refactor: adopt modular webhook architecture (#1621) + port fork customizations (#85)
* Adopt upstream modular webhook skeleton (#1621) Apply the durable-interrupt-dispatch refactor: split the monolithic webapp.py into a thin routing layer plus per-source handlers in webhooks/{github,slack,linear}.py, and add completion.py, dispatch.py, and reconcile.py. Reconcile fork divergence by keeping the Bedrock/ Fireworks cross-provider fallback, the no-agent-attribution prompt policy, the dashboard-handoff re-export, and the Slack channel-info cache. ci_autofix is restored on the new dispatch model in a later commit. Refs: #80 * Port fork webhook security delta onto modular handlers Re-apply the fork's security customizations that #1621 did not carry: Linear webhook replay protection (freshness window on the signed webhookTimestamp), per-repo token-cache binding threaded through the thread token resolvers, the INTERNAL_BOT_LOGINS self-check in the review-finding-reply path, and a user-mapping cache refresh before email resolution on the issue and PR-comment paths (multi-replica staleness). Existing fork security tests pass unchanged. Refs: #80 * Restore CI auto-fix on the modular dispatch model Bring back ci_autofix.py and the ci_monitor graph that #1621 deleted, re-wiring the fork's security-reviewed PR-babysitting onto the new structure: the CI-event, autofix-toggle, and review-feedback handlers move into webhooks/github.py and the github_webhook router re-gains the check_run/check_suite/workflow_run/status routing plus the autofix command and actionable-review branches. Auto-fix runs now dispatch through dispatch_agent_run (durability + completion webhook) while keeping the deliberate batch-while-busy skip-rule via get_thread_active_status. Restore langgraph.json's ci_monitor entry and the fork autofix tests (dispatch mock + import paths re-pointed). Refs: #80 * Reformat and update docs for the modular webhook split Point CLAUDE.md and deploy/MIGRATION.md at the new webhooks/ modules and the dispatch/completion/reconcile contract, and mark the user-mapping cache-refresh fix as applied on the GitHub handlers. Refs: #80 * Restore reject backstop for autofix dispatch A burst of near-simultaneous CI events for one head SHA can slip past the busy-check before the dedupe SHA is recorded, so dispatch the autofix path with multitask_strategy=reject (dev's prior platform default) to drop duplicate concurrent creates instead of letting them interrupt each other. Also make the completion failure-reply dedup claim-then-post and drop the unreachable interrupted branch. --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> |
||
|
|
9f7a1cc481
|
feat: post scheduled-run report to a configured Slack channel (#83)
* Post scheduled-run reports to a Slack channel Scheduled runs previously had no source channel and finished silently in the dashboard. Allow an automation to post its final report to a configured Slack channel as the bot, reusing existing Slack plumbing instead of the deferred run-completion webhook. Refs: #82 * Reconcile Slack report feature with dev merge Dev refactored slack_thread_reply to async and already added post_slack_top_level_message_with_ts; drop the duplicate definition and await the tool in the feature's tests. Refs: #82 * Harden scheduled Slack report channel posting The no-thread_ts top-level path fired for every slack_thread_reply call during a scheduled run, spraying disconnected messages and dead interactive buttons into the report channel. Cap top-level posts at one per run and drop options/plan_approval blocks in that mode, so the mechanism (not just the prompt) enforces a single clean report. Also tighten the channel-ID regex to require a leading letter and document why top-level posts store no run mapping. Refs: #82 --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> |
||
|
|
1f060f2a1d
|
chore: sync upstream/main, defer #1621 modular webhooks (#81)
* chore: bake sfw binary into sandbox image (#1611) sfw only ships a launcher that fetches its real binary at first run and does a daily update check against api.github.com/repos/SocketDev/sfw-free. Both fail in the sandbox (restricted egress; the proxy injects the GitHub App installation token, which lacks access to that repo), so `sfw yarn install` errors with "could not fetch its binary". Pin sfw 2.0.6, warm + verify the binary cache at build, and set SFW_SKIP_UPDATE_CHECK=1 so runs use the baked binary offline. * feat: editable plan mode + fix review-plan banner overlap (#1610) * feat: editable plan mode + fix review-plan banner overlap Lets the thread owner edit the plan markdown by hand from the plan-review page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id} endpoint that re-publishes the plan and mirrors it into the sandbox plan.md, so approve hands the edited plan to the agent as the source of truth. Also fixes the collapsed git-panel's floating expand button covering the "Review plan ->" banner by reserving space for it. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: abort plan approval when the published plan read fails get_plan_content() swallowed store errors and returned None, so a transient failure during approve would still mark the plan approved and dispatch the generic fallback text — silently dropping an owner's edited plan. Read the plan strictly (raise_on_error=True) so approval aborts instead, matching the comment read. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show message timestamps (#1609) * feat: show message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: suppress fallback message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: stable message + tool-call hover timestamps Stamp a stable client-side arrival time per message and tool call (keyed by id, persisted to localStorage). Messages render the timestamp inline; tool rows reveal a dim timestamp chip on hover. Real backend created_at still takes precedence when present. * fix: hide client-stamped message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add PR trace resolution (#1612) * feat: add PR trace resolution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: inject reviewer trace context as JSON Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: address review on PR trace resolution Use the documented LangSmith metadata filter syntax (and(eq(metadata_key,...), eq(metadata_value,...))) instead of has(metadata, '{...}'), which does not match runs — _list_thread_runs was silently returning nothing. Bound full-text searches to a 90-day window so they don't hit LangSmith's large-window rate limit. Also folds in the best-effort branch->head-sha resolver (dropping the weighted scoring/threshold + repo/file evidence + GitHub hydration), sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint. The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session were removed; resolution now runs deterministically from the trusted run config with no model-controlled pr_url or thread_id. * fix: scope branch trace search to the repo Branch names like fix-tests aren't unique across repos (or older PRs) in a shared tracing project, so an unscoped branch hit could resolve to an unrelated thread and write its runs into the reviewer sandbox. Require the repo slug to co-occur with the branch in matched runs; the full head SHA stays unscoped since it is globally unique. Addresses open-swe review on PR #1612. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: include plan links in PR descriptions (#1613) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: gate workflow pushes with approval (#1614) * feat: gate workflow pushes with approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve proxy refresh test compatibility Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: bind workflow approvals to pushed ref Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: recover thread work as patch (#1615) * feat: recover thread work as patch Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: search sandbox cwd for recovery patches Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: omit plan link in PR description when no plan exists (#1618) Plan links in PR descriptions were always built from the thread id, so runs that never produced a plan linked to an empty plan-review page. Now the plan content store is consulted first; the link is only added when a plan with non-empty markdown actually exists. A transient store failure degrades gracefully (no link) rather than blocking PR creation. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add filter & grouping menu to agents threads sidebar (#1617) Add a Cursor-style control to the agents sidebar that groups (None/Date/ Status/Project), filters (ownership, status, source, pull request, model, repo, include-resolved), and compacts the threads list. All client-side over already-fetched sidebar threads; preferences persist in localStorage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: update langsmith sdk to 0.9.3 (#1616) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: clickable shared PR header in git panel and reviews (#1620) * feat: clickable shared PR header in git panel and reviews Replace the standalone "View PR" button in the agent git panel with a clickable PR title, matching the reviews view. Extract a shared PrHeader component reused by both the git panel and the review main body. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: drop PrHeader wrapper, use shared component directly The review-side PrHeader was just a thin adapter mapping detail -> the shared component's props. Inline it at the call site and use the shared PrHeader directly so there's a single component. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: durable interrupt dispatch + completion webhook (#1621) * wip(rebuild): core reliability spine - remove PR-babysitting (ci_autofix + ci_monitor graph + webhook wiring) - dispatch core: agent/dispatch.py with multitask_strategy=interrupt + durability=sync + completion webhook; reroute all webhook + plan triggers; drop the racy in-process lock + is_thread_active busy-check - completion webhook: agent/completion.py + /webhooks/run-complete loopback route for failure/timeout replies (idempotent) Co-authored-by: open-swe[bot] * feat(rebuild): async tools, reconcile, shared http timeouts, assembly tuning Parallel batch on top of the reliability spine: - async-ify all 24 tools (drop asyncio.run; requests->httpx); re-implement the http_request/fetch_url SSRF + DNS-rebinding defense httpx-natively and harden the IP check to 'not is_global' (+ IPv4-mapped unwrap) - reconcile.py: stale pending-run sweep (threads.search -> per-thread runs.list -> cancel_many), wired into the scheduler graph via task='reconcile' - shared DEFAULT_HTTP_TIMEOUT (agent/utils/http.py) on every bare httpx.AsyncClient() across utils/dashboard/webapp/middleware - run budget: MODEL_CALL_RECURSION_LIMIT 5000->250 - fix stale OpenAI->Anthropic fallback id (claude-opus-4-5 -> 4-8) - drop redundant custom repair middleware (deepagents auto-adds PatchToolCalls) - confirm tool-result eviction + summarization auto-wired via backend - slim system prompt ~8% (full harness-profile rewrite deferred) Co-authored-by: open-swe[bot] * feat(rebuild): harness-profile prompt + split webhooks out of webapp - prompt.py: own the system prompt via a registered harness profile (OPEN_SWE_SHARED_BASE, kept neutral so the read-only reviewer/analyzer that share it stay safe), registered across all 4 providers; per-thread values stay in construct_system_prompt. Assembled main-agent prompt ~6.8k -> ~3.1k tokens (~55% smaller); de-duped PR/commit/suite/force-push guidance; dropped ALL-CAPS markers. - webapp.py 3325 -> 1890 LOC: moved 14 per-source handlers into agent/webhooks/{linear,slack,github}.py; webapp re-exports them for the routes + tests; moved handlers reach shared helpers via the webapp namespace to preserve the test suite's monkeypatch targets. Full suite: 1168 passing, lint clean. Co-authored-by: open-swe[bot] * Restore MODEL_CALL_RECURSION_LIMIT to 5000 for long-running tasks Reverts the 250 cap from the run-budget change — long-running tasks legitimately need many model calls. The notify_step_limit_reached safety net still fires if a run does hit the cap, so runs end with a signal either way. Co-authored-by: open-swe[bot] * fix: address PR review (auth, SSRF, interrupted status, redirect headers) - completion.py: drop `interrupted` from failure statuses — with multitask_strategy=interrupt a follow-up ends the prior run as interrupted, which is healthy, not a failure to report. [open-swe] - /webhooks/run-complete: shared-secret auth — dispatch appends ?token= when RUN_COMPLETE_WEBHOOK_SECRET is set; route verifies via hmac.compare_digest. [corridor-security] - SSRF: extract the URL validator to agent/utils/url_safety.py and apply it before server-side image fetches in multimodal.fetch_image_block. [corridor-security] - http_request: preserve caller headers/extensions across redirect hops instead of dropping them on the first hop. [open-swe] Co-authored-by: open-swe[bot] * chore: remove REBUILD_PLAN.md (planning doc, not needed in the repo) Co-authored-by: open-swe[bot] * fix: fail closed on run-complete webhook auth when secret unset Corridor follow-up: verify_run_complete_token returns False (not True) when RUN_COMPLETE_WEBHOOK_SECRET is unset, so the public route is never unauthenticated. Logs a startup warning when the secret is absent, and dispatch skips registering the webhook when there's no secret (no rejected callbacks). Co-authored-by: open-swe[bot] --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: restore forced tool call to prevent premature run stops (#1622) Restore the ensure_no_empty_msg middleware and the always-call-a-tool system-prompt instruction that #1535 removed. When the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion), the middleware re-injects a no_op / confirming_completion tool call so the run continues instead of ending mid-task. Shipping to test whether it fixes runs that stop halfway through. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore(deps): bump langgraph-checkpoint from 4.1.0 to 4.1.1 (#1619) Bumps [langgraph-checkpoint](https://github.com/langchain-ai/langgraph) from 4.1.0 to 4.1.1. - [Release notes](https://github.com/langchain-ai/langgraph/releases) - [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1) --- updated-dependencies: - dependency-name: langgraph-checkpoint dependency-version: 4.1.1 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * fix: post reviewer resolution notes verbatim (#1624) * fix: post reviewer resolution notes verbatim Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: stabilize dashboard follow-up e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make e2e attribution marker durable Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: only echo found e2e attribution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: check live dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * hotfix: stop prompting agent/reviewer to wrap installs in sfw (#1625) Installs hung when prefixed with sfw inside the sandbox (trace 019f0608 stalled on a pending `sfw npm install` execute, never returned). Strip the Socket Firewall guidance from the agent and reviewer prompts so installs run through the project's package manager directly. sfw stays in the Docker image; nothing invokes it now. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make plan view mobile friendly (#1636) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: fall back to vision model for image threads (#1626) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: surface Slack thread errors (#1627) * fix: surface Slack thread errors Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: don't set failure_reply_posted on Slack preprocessing errors The preprocessing error handler was setting failure_reply_posted=True, the same idempotency flag handle_run_completion checks to suppress duplicate run-failure replies. Since preprocessing failures happen before any run exists but the flag persists on the thread, a subsequent run failure on the same thread would be silently ignored. The preprocessing handler already posts its own Slack reply, so the run-completion idempotency flag should not be set here. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: avoid recapping Slack replies (#1629) * chore: avoid recapping Slack replies Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: simplify Slack reply prompt wording Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: update Slack trace reply on web handoff (#1630) * fix: update Slack trace reply on web handoff Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: trigger web handoff on dashboard starts Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: format web handoff as contextual fragment Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve trace_message_ts when overwriting Slack run mapping When store_slack_run_mapping is called without trace_message_ts (e.g. on follow-up Slack mentions), it was unconditionally overwriting the thread-level mapping and clobbering the timestamp captured from the initial trace reply. After that, _notify_slack_web_handoff could not find the original message, so a subsequent move to Web silently skipped the Slack trace update. Now, when trace_message_ts is not passed, the existing thread mapping is read first and its trace_message_ts is preserved. * style: ruff format --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * fix: pre-bundle shiki/@pierre deps to stop dev dynamic-import failures (#1643) * fix(ui): pre-bundle shiki/@pierre deps to stop dev dynamic-import failures shiki lazy-imports a grammar per language and these libs only live inside lazy route components, so Vite's startup scanner never sees them. They get discovered on first thread navigation, triggering a dep re-optimize + force-reload that aborts the in-flight route-chunk import, surfacing as "Failed to fetch dynamically imported module: .../$threadId.tsx". Pre-bundle them (and the github themes + common code-block languages) via optimizeDeps.include so the optimize happens once at startup. Dev-only; production bundles are unaffected. * fix: pre-bundle canonical shiki docker/make langs instead of aliases --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show queued dashboard follow-ups (#1631) * feat: show queued dashboard follow-ups Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: de-dupe queued follow-ups while streaming --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: notify Slack on plan approval (#1632) * feat: notify Slack on plan approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: post Slack approval notice after successful dispatch Move the _maybe_post_plan_approved_to_slack call until after _dispatch_followup succeeds so the Slack thread is not told implementation is beginning before the LangGraph run is created. Addresses PR review comment. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: include Slack channel context in prompts (#1633) Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> * chore: keep plan guidance high-level (#1634) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: publish plans from sandbox files (#1635) * feat: publish plans from sandbox files Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: avoid fixed plan filenames Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: virtualize local sandbox file paths Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve plan_file_path across set_plan_status set_plan_status was rewriting the content record with only markdown and status, dropping plan_file_path. After a reject, the owner's dashboard edit would mirror to a different file than the agent's original, and the next save_plan could republish the stale file. Preserve plan_file_path when updating status. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: return to thread after plan approval (#1637) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Slack breakout thread tool (#1638) * feat: add Slack breakout thread tool Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: make fake LLM scripts declarative Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: exclude slack_start_new_thread from plan mode The breakout tool can dispatch a fresh agent run that starts outside the current plan-mode state, bypassing the approval flow. Add it to PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating tools while planning. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: require bun for ui agent work (#1639) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: request actions read for sandbox logs (#1642) * fix: request actions read for sandbox logs Request optional Actions read permission for sandbox proxy tokens, with fallback for installations that have not approved it yet. Update setup docs and prompt guidance for safe GitHub Actions log usage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: restore actions:read scope after workflow push After an approved workflow push, the guard was restoring the proxy with BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS, which excludes the actions: read scope this PR adds. Restore with RUNTIME_PROXY_TOKEN_PERMISSIONS (which includes actions: read) and fall back to BASE if the install hasn't granted Actions read — mirroring the pattern in _create_sandbox_with_proxy. Addresses review comment on PR #1642. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: widen split review diffs (#1647) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: install missing deps before verification (#1646) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm (#1645) * chore: require pnpm for ui agent work Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm Replace Bun and Yarn lockfiles with pnpm lockfile and update UI/Vercel commands to use pnpm. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * ci: use corepack for ui pnpm e2e build Run pnpm through Corepack in the E2E global setup so CI can use the pinned package manager without a separate pnpm install step. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Sonnet 5 to model picker (#1651) * chore: update Sonnet examples to Sonnet 5 Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: add Sonnet 5 to model picker Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * Remove dead breakout-thread e2e scenario after dropping the tool The merge resolution deferred upstream's Slack breakout-thread tool (slack_start_new_thread, #1638) since it depends on the #1621 dispatch module, but the e2e harness still scripted it. Removing the tool name from fake_llm.py's _tool_step call left a malformed scenario, crashing the langgraph-dev web server at import (TypeError: _tool_step() missing 'call_id') and failing Playwright E2E. Drop the "breakout" script scenario, its _is_breakout_request helper + ScriptRule, and the corresponding full_flow.spec.ts test. * Revert upstream pnpm switch; keep bun for the UI build The merge auto-adopted upstream's pnpm switch (#1645) in tests/e2e/ global-setup.ts and ui/package.json, but our fork builds the UI with bun (vercel.json + the E2E workflow's setup-bun). That left the Playwright globalSetup running `corepack pnpm install --frozen-lockfile` with no pnpm-lock.yaml, failing E2E at UI build time. Revert global-setup.ts and ui/package.json to the dev (bun) baseline, drop the merge-added ui/pnpm-lock.yaml, and remove the re-added ui/AGENTS.md (our fork had deleted it). * Align plan-review e2e + UI with the HEAD (pre-#1635) backend The merge left a split plan vertical: the backend save_plan/plan_api are HEAD (we deferred the editable-plan/sandbox-publish features #1610/#1635/ #1637 per #80), but the plan UI and e2e harness were upstream's. The fake_llm scenario called save_plan(plan_file_path=...) — upstream's file-based #1635 contract — while HEAD save_plan takes plan_markdown, so the plan never saved and PlanReview never rendered (E2E failure on the plan-review locator). Pass plan_markdown to save_plan, and revert PlanReview.tsx / plan.ts / $threadId_.plan.tsx / plan_review.spec.ts to the dev baseline so the whole plan flow (save -> render -> approve -> implement) is consistent with the HEAD backend. --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Ankush Gola <9536492+agola11@users.noreply.github.com> Co-authored-by: Mukil Loganathan <mukil@langchain.dev> |
||
|
|
eb98ff4c30
|
feat: add 3 verified Fireworks models to selectable set (#79)
* Add 10 Fireworks models to selectable set Surface additional Fireworks-served models in the profile editor so they can be chosen per-thread, per-profile, and as team defaults. Each entry carries its recommended efforts and image support; only MiniMax M3 is multimodal. Refs: #78 * Suppress reasoning_effort on non-reasoning models Instruct-only Fireworks ids (kimi-k2-instruct-0905, mistral-large-3-fp8, qwen3-30b-a3b-instruct-2507) don't reason, so sending reasoning_effort either 400s (unusable at default effort) or is a silent no-op. Add a per-model reasoning flag (default True) and omit the param entirely for ids marked non-reasoning. Refs: #78 * Gate out 7 undeployed Fireworks models Account serverless probe returned 404 for 7 of the 10 proposed ids, so only minimax-m3, gpt-oss-120b, and deepseek-v4-flash are callable. Keep those 3 and drop the rest. All 3 survivors are reasoning-capable, so the per-model reasoning-effort suppression added earlier is no longer needed and is reverted. Refs: #78 --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> |
||
|
|
860ce93ad7
|
docs: record final managed-deployment topology in MIGRATION.md (#77)
Add an authoritative "Final, verified topology" section (1a) and reconcile the phased plan with the executed end-state, superseding the spike-era values throughout. Captures: two LangGraph Cloud deployments (dev/prod, one LangSmith workspace) with their final URL hashes; one Vercel project open-swe-prod with production + custom dev environments and per-env LANGGRAPH_BACKEND_URL; the Nitro routeRules proxy mechanism (PR #76, superseding the vercel.json rewrite and PR #75's build-vercel-output.mjs); two dev/prod-isolated GitHub Apps; per-deployment user stores; the Bedrock IAM users; and the seven hard-won operational gotchas. Refs: #65 #74 #76 |
||
|
|
8441bbb2d8
|
docs(migration): record the two Bedrock IAM users + scoped invoke policy (#74)
Document the live execution of MIGRATION.md §10.1 (Bedrock auth via static keys): the customer-managed least-privilege policy open-swe-bedrock-invoke and the open-swe-dev-bedrock / open-swe-prod-bedrock IAM users. Captures the static-key deviation rationale (managed LangGraph Cloud cannot assume a role) and that both mandatory gates (GPT-4.1 IAM cross-review, /sh-security-review) passed with no critical/high. |
||
|
|
49fd48d32f
|
chore(deps): bump the minor-and-patch group across 2 directories with 4 updates (#69)
Bumps the minor-and-patch group with 1 update in the /tests/e2e directory: [@playwright/test](https://github.com/microsoft/playwright). Bumps the minor-and-patch group with 3 updates in the /ui directory: [monaco-editor](https://github.com/microsoft/monaco-editor), [@tanstack/devtools-vite](https://github.com/TanStack/devtools/tree/HEAD/packages/devtools-vite) and [prettier-plugin-tailwindcss](https://github.com/tailwindlabs/prettier-plugin-tailwindcss). Updates `@playwright/test` from 1.61.0 to 1.61.1 - [Release notes](https://github.com/microsoft/playwright/releases) - [Commits](https://github.com/microsoft/playwright/compare/v1.61.0...v1.61.1) Updates `monaco-editor` from 0.52.2 to 0.55.1 - [Release notes](https://github.com/microsoft/monaco-editor/releases) - [Changelog](https://github.com/microsoft/monaco-editor/blob/main/CHANGELOG.md) - [Commits](https://github.com/microsoft/monaco-editor/compare/v0.52.2...v0.55.1) Updates `@tanstack/devtools-vite` from 0.6.1 to 0.8.1 - [Release notes](https://github.com/TanStack/devtools/releases) - [Changelog](https://github.com/TanStack/devtools/blob/main/packages/devtools-vite/CHANGELOG.md) - [Commits](https://github.com/TanStack/devtools/commits/@tanstack/devtools-vite@0.8.1/packages/devtools-vite) Updates `prettier-plugin-tailwindcss` from 0.7.4 to 0.8.0 - [Release notes](https://github.com/tailwindlabs/prettier-plugin-tailwindcss/releases) - [Changelog](https://github.com/tailwindlabs/prettier-plugin-tailwindcss/blob/main/CHANGELOG.md) - [Commits](https://github.com/tailwindlabs/prettier-plugin-tailwindcss/compare/v0.7.4...v0.8.0) --- updated-dependencies: - dependency-name: "@playwright/test" dependency-version: 1.61.1 dependency-type: direct:development update-type: version-update:semver-patch dependency-group: minor-and-patch - dependency-name: "@tanstack/devtools-vite" dependency-version: 0.8.1 dependency-type: direct:development update-type: version-update:semver-minor dependency-group: minor-and-patch - dependency-name: monaco-editor dependency-version: 0.55.1 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: minor-and-patch - dependency-name: prettier-plugin-tailwindcss dependency-version: 0.8.0 dependency-type: direct:development update-type: version-update:semver-minor dependency-group: minor-and-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
5caaf885a4
|
chore(deps): bump fireworks-ai from 1.2.0a75 to 1.2.0a85 (#37)
Bumps [fireworks-ai](https://fireworks.ai) from 1.2.0a75 to 1.2.0a85. --- updated-dependencies: - dependency-name: fireworks-ai dependency-version: 1.2.0a85 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
c75e76c05e
|
fix(ui): drive Vercel dashboard-API proxy through Nitro routeRules (#76)
The Vercel deploy 404'd at `/` because PR #75's `vercel-build` script (`scripts/build-vercel-output.mjs`) ran `rm -rf .vercel/output` and rebuilt it from `.output/public`. On Vercel CI, Nitro's Vercel preset auto-activates (from the VERCEL env var) and emits the Build Output API layout to `.vercel/output` itself during `vite build`; the script then clobbered that correct output with a static-only config that could not resolve the SPA `_shell.html` fallback or the server functions, so production returned `404: NOT_FOUND` even though the build was READY. Stop fighting Nitro and drive the proxy through it instead: - Remove `scripts/build-vercel-output.mjs` and the `vercel-build` script; restore the plain `vite build` for both local and Vercel builds. - Add an env-driven Nitro `routeRules` proxy in `vite.config.ts`. Nitro's Vercel preset compiles a plain external-URL `proxy` rule into a CDN-level rewrite in `.vercel/output/config.json` at build time, reading LANGGRAPH_BACKEND_URL (the per-project Vercel env var). Proxy, not redirect, so the osw_session cookie stays first-party (same-origin). The build fails loudly if LANGGRAPH_BACKEND_URL is missing on Vercel; the rule is omitted for plain local/off-Vercel builds (which use the E2E_HARNESS mock proxy). - `vercel.json`: `buildCommand` back to `bun run build`; drop the `.output/public` outputDirectory so Vercel serves Nitro's `.vercel/output`. Validated with `LANGGRAPH_BACKEND_URL=... NITRO_PRESET=vercel bun run build`: the generated `.vercel/output/config.json` contains the `/dashboard/api/(.*)` -> backend rewrite ahead of `handle: filesystem` and the SPA catch-all, `_shell.html` and the 699 hashed assets are emitted, and typecheck passes. Supersedes the broken approach in #75. |
||
|
|
076e20d7e6
|
feat(ui): derive Vercel proxy backend from per-project env var (#75)
Generate the /dashboard/api/* proxy destination at build time from LANGGRAPH_BACKEND_URL via the Vercel Build Output API, instead of hardcoding it in ui/vercel.json. This lets the dev and main branches stay byte-identical (required by the fast-forward-only prod promotion) while each Vercel project proxies to its own LangGraph backend purely via its own env var — so both the dev and prod dashboards can be git-linked. - ui/vercel.json: drop the hardcoded rewrites; buildCommand -> bun run vercel-build - ui/package.json: add vercel-build = vite build && node scripts/build-vercel-output.mjs - ui/scripts/build-vercel-output.mjs: emit .vercel/output/config.json with [ proxy /dashboard/api/* -> $LANGGRAPH_BACKEND_URL, filesystem, SPA fallback ]; throws if LANGGRAPH_BACKEND_URL is unset so a misconfig fails the build. Keeps the app same-origin (no CORS/auth change); preserves the _shell.html SPA fallback. Set LANGGRAPH_BACKEND_URL per Vercel project (All Environments). |
||
|
|
a30ce2ab40
|
feat: managed LangGraph Cloud + Vercel migration (PR2 — code fixes + docs) (#65)
* fix(dashboard): managed-cloud OAuth hardening + admin user-mapping endpoint Prepare the dashboard backend for the managed LangGraph Cloud + Vercel runtime, where the API is HTTPS and cross-site from the UI. - OAuth redirect_uri (#2): coerce a schemeless DASHBOARD_API_BASE_URL to https:// in _api_base_url() so GitHub stops rejecting login with "redirect_uri not associated with this application". _cookie_security() now treats a schemeless (managed) value as Secure; SameSite=None too, consistent with the coerced scheme. - OAuth state cookie (#3): document that osw_oauth_state is host-only by design (a Domain cookie is unsafe across *.vercel.app, a public suffix), so login must always start on the stable alias to avoid "oauth state mismatch". Operational contract; no behavioral change. - Admin user mappings (#4): add POST /admin/user-mappings so an admin can set the github_login -> work_email link from the dashboard instead of a raw Store write. New "admin" MappingSource provenance value. * fix(webapp): refresh user-mapping cache on GitHub webhook paths On managed LangGraph Cloud the backend runs multiple replicas, so the per-process GitHub<->work-email mapping cache can be stale on the replica handling a webhook (a mapping created on another replica is invisible until refresh). process_github_pr_comment and process_github_issue now refresh the cache from the durable Store before resolving the author's email, matching the existing Slack mention path (process_slack_mention). * perf(webapp): defer deepagents import to speed custom-app cold start The custom FastAPI app (agent.webapp:app, the langgraph.json http.app) pulled deepagents -> langchain_anthropic -> anthropic into its import graph via dashboard.routes, only to build skill/chat seed files. Defer those create_file_data imports into the functions that use them. Removes deepagents/langchain_anthropic/anthropic from app import entirely and roughly halves module-import wall time (~0.6-0.8s -> ~0.35s warm; larger cold-start saving since native anthropic init is skipped). Behavior identical. (reviewer_diff already imports deepagents under TYPE_CHECKING.) * feat(ui): set work_email user mappings from the admin dashboard Add an "Add / update" form to the admin User mappings section and the adminUpsertUserMapping API client method, wiring the new POST /admin/user-mappings endpoint. Admins can now create or update a github_login -> work_email mapping directly instead of waiting for the user to self-connect Slack. * docs: document managed LangGraph Cloud + Vercel deployment - INSTALLATION §10: add the managed production env triad (LANGGRAPH_URL, DASHBOARD_BASE_URL + DASHBOARD_API_BASE_URL with https://, empty VITE_DASHBOARD_API_BASE_URL for same-origin), the stable-alias login and vercel.json stable-deployment-URL requirements, multi-replica cache note, plus redirect_uri-scheme and oauth-state-mismatch troubleshooting. Refresh the langgraph.json snippet to all six graphs. - README: reframe deployment around the managed migration; link the plan. - deploy/MIGRATION.md: import the self-hosted -> managed migration plan. |
||
|
|
7f60324f0c
|
chore: decommission self-hosted AWS LangGraph stack (#64)
* chore: decommission self-hosted AWS LangGraph stack Removes the now-dead self-host IaC and AWS-only CI/CD after destroying the dev + prod CloudFormation stacks (open-swe-dev, open-swe-prod, open-swe-iam, and the dev-exclusive CDKToolkit-oswedev bootstrap) in account 328440206208, us-east-1. The deployment is now managed (LangGraph Cloud + Vercel). - remove infra/ (CDK app: app + IAM stacks, constructs, aspects, tests) - remove deploy/ami (Packer AMI build) and deploy/seahaven (boot/config scripts, DEPLOYMENT/ROTATION runbooks) - remove AWS-only workflows: cd-infra, ci-infra, build-artifacts, rollback - README: rewrite the Deployment section to the managed LangGraph Cloud + Vercel view; drop dead links to infra/ and deploy/seahaven Preserved: the shared default CDKToolkit bootstrap and promote-dev-to-prod.yml. The RETAIN'd Secrets Manager shells and open-swe-<env>-assets S3 buckets survive cdk destroy by design (orphaned) and need a separate deliberate cleanup. * chore: clean up dangling references left by the AWS decommission Folds in the FIX-level items from the #64 review gates (GPT-4.1 cross-review + /sh-security-review), none of which were blockers: - delete orphaned .github/scripts/{package-artifacts,publish-and-deploy,roll-box, rollback}.sh — their only callers were the removed AWS deploy workflows - drop the deleted /infra dir from dependabot.yml npm directories (was producing a recurring Dependabot config error) - remove the stale OSWE-IAC-SECRETS-LIST-01 suppression (referenced the deleted infra/lib/constructs/instance-role.ts) - repoint the README promotion link to promote-to-main.yml (renamed in #63) The promote-dev-to-prod.yml comment in check-dev-green.sh is intentionally left to #63, which rewrites that same line. |
||
|
|
430a1cdff9
|
ci: re-home prod promotion into a gated promote-to-main workflow (PR1: managed-LGC migration) (#63)
* ci: re-home prod promotion into a gated promote-to-main workflow Migrate the prod-deploy gate to managed LangGraph Cloud (git-connected to `main`) + Vercel. Under managed, a push to `main` auto-deploys prod, so the dev -> main fast-forward IS the prod deploy trigger -- the bespoke AWS CD step is obsolete and already gone from this workflow. Re-home `promote-dev-to-prod.yml` -> `promote-to-main.yml`: - gate the promote job on the `prod` GitHub Environment (required reviewer amoussa1229), restoring the manual prod-approval that the retired AWS CD job used to carry; - drop the nightly auto-promote cron -- a scheduled auto-promotion conflicts with a manual approval gate now that the push deploys prod; promotion is workflow_dispatch only; - keep the dev-HEAD-fully-green precondition and the seahaven-promotion App fast-forward push (sole non-admin bypass actor on `main` ruleset 18238334). Update the companion check-dev-green.sh filename reference. * ci: refuse promote-to-main dispatch from any ref other than dev Defense-in-depth atop the already-pinned `ref: dev` checkout: workflow_dispatch runs the workflow definition from the launched ref, so reject a non-dev dispatch before the App token is minted. Surfaced by the GPT-4.1 cross-review of #63. |
||
|
|
a4ed19ba61
|
feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62)
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)
Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.
- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module
* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8
The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.
Map profile effort to additional_model_request_fields:
{thinking: {type: adaptive, display: summarized},
output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.
* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids
Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
(set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.
Surfaced by the cross-family review + verified against deploy/.
* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip
From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
error code only, so the role ARN + account id in the raw botocore message never
reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
(Converse emits reasoning_content, not thinking) so the middleware is not a no-op
on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)
* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids
Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
each routed region (us-east-1/2, us-west-2). The model runs in the server process
on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
seed_store.sh's default via pick precedence, so the seed-script fix alone was
insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
ids to the Bedrock id (config.toml's model_id was an active, now-broken value).
AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.
* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)
Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
(eval judge only — Bedrock builder/reviewer auth via the host IAM role).
REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
|
||
|
|
8a9974c3c4
|
feat: author Slack/dashboard/schedule commits + PRs as the app by default (#57) (#60)
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: default Slack/dashboard/schedule PRs + commits to the app identity (#57) Slack/dashboard/schedule runs now author PRs and run git/gh operations as the GitHub App seahaven-openswe[bot] by default (matching GitHub-issue runs), so the self-review 422 is impossible by construction rather than guarded in the prompt. A profile flag author_prs_as_user restores per-user attribution. - open_pull_request._resolve_pr_author_token + auth.resolve_github_token: default to the installation token for these sources; per-user only when opted in. - authorship: commit identity -> seahaven-openswe[bot] (numeric noreply; accepted Vercel-resolution risk, documented inline). - self-trigger safety: INTERNAL_BOT_LOGINS + webapp/reviewer_reconcile/reply markers recognize seahaven-openswe[bot] (bot-authored events are now ours). Supersedes the prompt-only guard in #58. * fix: author commits as the app bot in the default path (SH-IDSPLIT-01) Security review found the commit identity was NOT actually unified to the bot: resolve_triggering_user_identity got a 403 from the installation token and fell back to configurable['github_login'], so commits were still authored as the triggering user (commit=user, push+PR=bot — a three-way split that missed the stated goal). Now gate the triggering-user identity resolution on the same default-bot decision as the token: slack/dashboard/schedule default to the app bot identity unless author_prs_as_user is set. * docs(security): record AUTHZ-SLACK-BOT-DEFAULT-001 as an accepted residual (#59) Single-user deployment; bounded by App-on-pilot + ALLOWED_GITHUB_REPOS lock. Revisit (add a per-user gate) before expanding users or the App installation. |
||
|
|
134963647b
|
chore(security): suppress pre-existing history scanner false-positives (#56)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Adds repo-local suppressions for 5 verified-FP gitleaks findings that block pushes (forcing --no-verify), all in committed history / docs / CI fixtures: - .env.ci (dev-only e2e values, intentionally committed) - .env.example (placeholders) - .github/ci/fake_github_app_key.pem (throwaway CI test key) - INSTALLATION.md (example GITHUB_APP_PRIVATE_KEY .env block) - README.md (prose mis-matched by the generic-api-key heuristic) Repo-local (not machine-level) so they load in git worktrees too. Also drops the now-obsolete OSWE-IAC-AUDIT-01 suppression (B-1, fixed in #55). |
||
|
|
e9499e49b8
|
fix: isolate dev CDK deploys on their own bootstrap qualifier (B-1/OSWE-IAC-01) (#55)
* fix: isolate dev CDK deploys on their own bootstrap qualifier (B-1) Dev synthesizes against the oswedev qualifier and the dev infra deploy role is scoped to cdk-oswedev-* — it can no longer assume the default hnb659fds bootstrap roles whose admin cfn-exec-role deploys prod, closing the cross-env escalation (OSWE-IAC-01). Prod stays on the default qualifier. * test: assert per-env bootstrap qualifier isolation + document (B-1) |
||
|
|
a33aaec495
|
fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54)
* fix: enforce a replay window on Linear webhooks (AUTHZ-001) verify_linear_signature accepted any correctly-signed body with no freshness check, so a captured request could be replayed indefinitely. Parse the signed webhookTimestamp (Unix ms) and reject requests outside a 60s window, failing closed when the field is missing or malformed — mirroring the Slack verifier. * fix: stop leaking upstream auth-error bodies into user comments get_github_token_for_user folded the raw upstream response text into the error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the full body server-side only and return a generic "GitHub auth failed (status <code>)". Also document the accepted shared-installation-token blast radius on the bot-token-only path (AUTHZ-003). * fix: bind sandbox and token caches to repo to prevent thread-id collision A PR head-branch name is attacker-controllable and get_thread_id_from_branch derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01). The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on thread_id alone, and a cached sandbox was reused after only an echo-ping, so a different repo's webhook could bind to another thread's sandbox or token. Without changing the persistent thread-id scheme: - Persist the bound repo (owner/name) in thread metadata on sandbox creation and refuse to reuse a sandbox whose bound repo does not match the current event (SandboxRepoMismatchError); the in-memory proxy also carries the binding. - Bind the GitHub-token cache entries to their repo and evict on a cross-repo read so a colliding thread_id cannot be served another repo's token. - Thread repo through the reviewer and the webhook token resolvers. * fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04) The instance role and the GitHub deploy app role granted s3:ListBucket on the whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts) only ever lists under releases/, so add a StringLike s3:prefix=releases/* condition. GetBucketLocation has no s3:prefix in its request context, so it moves to its own unconditioned statement. Also document the accepted F-2 cross-env existence-oracle residual on BatchGetSecretValue. * chore: suppress test-fixture credential false positive; document AUTHZ-002 Add a machine-level suppression for the fake Datadog key in the test_team_credentials encryption-roundtrip fixture (CWE-798, not a real credential). Clarify that the within-org thread-write path is intentional by design (AUTHZ-002) — comment only, no behavior change. * fix: casefold repo-binding keys to avoid spurious cross-repo mismatch GitHub owner/name are case-insensitive. Casefold the owner/name key on both the write (binding) and read (compare) sides — repo_cache_key and the metadata bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2). * fix: stop leaking upstream auth body in unexpected-result branch The 2xx-but-missing-token/url branch echoed the parsed upstream response body into the user-facing error. Return a generic message and log response_data server-side only, mirroring the existing HTTPStatusError fix (Gap 4). * fix: fail closed for unbound-legacy sandboxes and catch repo mismatch Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no recorded bound_repo (a pre-binding legacy thread, post-deploy) previously reconnected-and-served the sandbox to the current repo, then rebound it. Now fail closed: drop the stale id and recreate a fresh sandbox bound to this repo, logging a reconnect-with-missing-binding event. A sandbox is never served to a repo unless its binding is known and matches; new threads bind on first run unchanged. Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints, log it for alarming, and surface a clean sanitized error instead of letting an opaque deep-stack exception crash-loop the worker. * chore: suppress test-fixture credential false positive in token-TTL tests Add a machine-level suppression for the fake "ghp_secret" GitHub token used by the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and not a valid PAT; scoped to the unit test only. |
||
|
|
9444fd7677
|
docs: document live AWS prod deploy for Open SWE (#53)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Prod went live 2026-06-29 on self-hosted AWS EC2 behind the shared seahaven-com ALB, superseding the on-prem VM model the runbook described. - Rewrite deploy/seahaven/DEPLOYMENT.md as the canonical end-to-end runbook: infra CD (CDK stacks + OIDC roles + prod approval gate), config seeding (put-config.sh, the 13 boot-required prod vars, fetch-config fail-fast), app artifact deploy (S3 + SSM roll + is-active gate), promotion/rollback, live prod facts, and a RETAIN secret-shell troubleshooting entry that cross-references infra/README.md. - Correct retired *.seahavenind.com hosts to *.seahaven.com throughout and document the live GitHub/Slack/Linear webhook + OAuth endpoints. - Add a concise Deployment section to README pointing at the runbook. - Fix the stale host in the retired on-prem nginx/openswe.conf and mark it superseded by the AMI template. |
||
|
|
f379fbdaa9
|
docs: document RETAIN secret-shell orphan gotcha (#52)
Some checks are pending
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
RETAIN + a fixed secret name means a failed FIRST create leaves empty secret shells behind when the stack rolls back. The shells keep the global `open-swe-<env>/<VAR>` names, so every later create fails with `AlreadyExists`, and a plain delete-secret keeps the name reserved for the recovery window rather than freeing it. Record the trap and the force-delete recovery (only for empty shells) in the config-store construct and the infra README so the next teardown/rebuild, secret logical-id change, or new-env stand-up does not rediscover it the hard way. Prod's first deploy hit this on 2026-06-29: 28 orphaned shells from an earlier failed create reserved the names and had to be force-deleted before the stack would create. |
||
|
|
86b4859589
|
fix: env-scope the EC2 launch template name (unblocks prod deploy) (#51)
Some checks are pending
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
requireImdsv2:true makes CDK auto-create a launch template named from the
construct id ('Instance' -> 'InstanceLaunchTemplate') with no env qualifier,
so OpenSweDevStack and OpenSweProdStack both render
LaunchTemplateName: InstanceLaunchTemplate. dev created it first (the live
dev box runs on it); the prod first-deploy then failed with
InvalidLaunchTemplateName.AlreadyExistsException and the whole stack rolled
back.
Force a per-env LT name (open-swe-<env>-lt) via an aspect (the LT is created
at synth time by the requireImdsv2 handling, not in the constructor), and
rename the instance's launch-template REFERENCE in lockstep so CFN still
resolves it. synth-verified: dev=open-swe-dev-lt, prod=open-swe-prod-lt on
both the LT resource and the instance reference; version GetAtt preserved.
NOTE: deploying this renames dev's LT -> one-time dev box replacement
(stateless; boots from the baked AMI + pulls releases/latest). prod then
creates open-swe-prod-lt cleanly.
|
||
|
|
3c69dd9de6
|
fix: BatchGetSecretValue must be granted on * (corrects #48, fixes dev crash-loop) (#50)
Some checks are pending
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
#48 (OSWE-IAC-SECRETS-LIST-01) scoped secretsmanager:BatchGetSecretValue to the open-swe-<env>/* ARN on the theory that an explicit --secret-id-list batch authorizes per-secret. That is FALSE: BatchGetSecretValue is a collection action AWS authorizes against the account (*), regardless of --filters vs --secret-id-list. A prefix-scoped grant AccessDenies the whole call. The dev box passed right after #48 only because the prior broad grant had not finished propagating; once it lapsed, fetch-config got AccessDenied -> loaded 0 secrets -> FAIL-FAST -> open-swe.service crash-loop. Verified on the live dev box (i-0af4e03e8bf70e6c3): the exact call returned 'not authorized to perform: secretsmanager:BatchGetSecretValue'; restoring the * grant recovered it. Move BatchGetSecretValue back to Resource:* (its own statement); keep GetSecretValue + DescribeSecret prefix-scoped (those gate VALUE access, so cross-env isolation holds). The surviving win from #48: --secret-id-list needs no name filter, so ListSecrets stays dropped -> no account-wide name enumeration. fetch-config.sh is unchanged (--secret-id-list is correct). The /sh-security-review finding OSWE-IAC-IAM-01 called this out and was wrongly refuted; the reference_secretsmanager_batch_get memory was wrong. |
||
|
|
9acf071ae4
|
fix: make dev->main promotion push succeed via bypass-actor App token (#49)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
The nightly promote fast-forwards main to a fully-green dev HEAD, but the push (as github-actions[bot]) is rejected by the main ruleset: it requires PRs + a status check and the default token is not a bypass actor, so a direct ref push can never land regardless of fast-forwardability. The prior comment claiming protection 'only rejects non-FF' was wrong. Mint a GitHub App installation token (actions/create-github-app-token, SHA-pinned) and push with it; the App must be added to the main ruleset's bypass actors out-of-band. The promoted commit already passed every check on dev (gated by check-dev-green.sh), so re-gating it via a PR on main is redundant. Also fix a gate self-poison: a stale failed 'promote' check-run from a prior run on the same dev HEAD blocked every subsequent gate run (it was excluded only by the current run_id). Exclude prior promote check-runs too, scoped to name=='promote' AND a /actions/runs/ details_url so an external app cannot hide a real failing check by naming it 'promote'; the positive REQUIRED_CHECKS allow-list stays authoritative. Gates: GPT-4.1 cross-review APPROVE (no security regression). Unit-tested: stale promote ignored -> PASS; real failure / external promote / missing required check -> BLOCK. shellcheck clean (also fixed a pre-existing SC2295 on the run_id match). |
||
|
|
faae9a685b
|
Scope secrets fetch to --secret-id-list; drop ListSecrets grant (#48)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Switch fetch-config.sh from a name-prefix batch-get-secret-value --filters scan to an explicit --secret-id-list (the 28 SECRET_VARS, chunked at the 20/call cap). An id-list batch authorizes per-secret ARN, so the instance role's BatchGetSecretValue moves from Resource:* to the open-swe-<env>/* prefix and the account-wide ListSecrets grant is dropped entirely. The box can no longer enumerate secret names account-wide; cross-env value isolation is unchanged (GetSecretValue was already prefix-scoped). Resolves OSWE-IAC-SECRETS-LIST-01. Also capture each chunk response into a variable and consume the producer via command substitution so a failed AWS call aborts under set -e instead of being swallowed by process substitution and misreported as a missing required var. Refs: OSWE-IAC-SECRETS-LIST-01 |
||
|
|
52cc696431
|
fix(deps): bump PyJWT to >=2.13.0 to close HS256 forgery (#47)
PyJWT <2.13.0 accepts a public-key JWK as an HMAC secret, letting an attacker forge HS256 tokens when mixed key families are allowed (GHSA high-sev Dependabot alert). 2.13.0 rejects the mismatch. Direct dependency; uv.lock re-resolves 2.12.1 -> 2.13.0. |
||
|
|
f79f49c755
|
Make PR title repo-aware and auto-link issues (#46)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Hardcoding the Sea Haven no-type-prefix PR title made every PR fail semantic-PR-title gates (this repo's PR Title Lint, upstream open-swe), forcing manual retitling. Make the title rule detect a conventional-commit gate and conform, falling back to the imperative style otherwise. Also add Closes/Refs issue-linking guidance and the default-branch auto-close caveat. Refs: #41 Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com> |
||
|
|
2533402f6b
|
chore(deps): update langgraph-cli[inmem] requirement (#38)
Updates the requirements on [langgraph-cli[inmem]](https://github.com/langchain-ai/langgraph) to permit the latest version. - [Release notes](https://github.com/langchain-ai/langgraph/releases) - [Commits](https://github.com/langchain-ai/langgraph/compare/cli==0.4.27...cli==0.4.30) --- updated-dependencies: - dependency-name: langgraph-cli[inmem] dependency-version: 0.4.30 dependency-type: direct:production ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |