An Open-Source Asynchronous Coding Agent
Find a file
Adam Moussa 589cd236c6
chore: cherry-pick clean upstream fixes + cherry-pick runbook (#117)
* fix: make plan view mobile friendly (#1636)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 7ee3e05724)

* fix: return to thread after plan approval (#1637)

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit f32e492ab4)

* feat: reviews block agenda, sticky headers, accurate diff scroll (#1653)

Rework the AI-sorted blocks experience on the PR reviews page into a
Google-Docs-style outline: the left sidebar is now a clean number+title
agenda with scroll-spy highlighting of the active block; each block shows
its title + description (sticky) above its diff; and diff rows are pinned to
a uniform height so scroll-to lands precisely via the virtualizer's own
geometry instead of an estimate-driven correction loop.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 0b76afdc955e33805c7623d1502a75a9c7c9c1b7)

* fix: jump + ResizeObserver settle for review scroll-to (#1655)

Replace smooth-scroll plus frame-count correction loops on the PR
reviews page with an instant jump that re-asserts its target via a
ResizeObserver (the real "layout settled" signal). Block/file
navigation and finding/comment centering now land deterministically as
off-screen cards mount, files expand, and annotation cards measure,
instead of racing a smooth-scroll animation against height
reconciliation. Holds bail on user wheel/touch input and after a short
ceiling, and a new navigation cancels the previous hold.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
(cherry picked from commit 7530653bba7774d66a54b8bef0d2bbc25f519942)

* fix: purge expired thread_wakeup crons (#1656)

* fix: purge expired thread_wakeup crons

One-shot wakeup crons set an end_time that stops re-firing but the cron
row is never deleted, so dead rows accumulate (86 in prod). Add a purge
that deletes thread_wakeup crons past their end_time, called
opportunistically before scheduling a new wakeup, plus a one-time
backfill script. Conservative: matches only kind=thread_wakeup with a
past end_time.

* chore: retrigger Open SWE review

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 9e5a1924ef306269322c31342a1831e57831cfee)

* fix: add top padding to sticky review block header (#1660)

* fix: add top padding to sticky review block header

The sticky per-block header on the reviews page had padding below but
none above, so the block number badge sat glued against the top edge
when pinned. Add matching top padding for breathing room.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: use py-2 shorthand for review block header padding

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 23bd4a63fc5ba0fe853babf79ed33feb866cc8b2)

* fix: use global tokens for sidebar filter popover border (#1661)

The filter popover renders via base-ui Menu.Portal into document.body,
outside the .agents-ui container where the --ui-* CSS variables are
scoped. As a result border-[var(--ui-border)] resolved to an undefined
variable and border-color fell back to currentColor, producing a strong
near-black border (separators/hover/labels were similarly off).

Switch the portaled popup styling to the same global shadcn tokens the
theme/settings popover (SidebarUserMenu) already uses (border-border,
bg-border, bg-muted, text-muted-foreground). These are defined at :root
so they resolve inside portals too, and match the settings popover.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 63eb9a08209f683016abf01cdcc548bc5905f158)

* fix: preserve dashboard redirect after login (#1668)

* fix: preserve dashboard redirect after login

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: cover plan login redirect in e2e

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit bc7ce59169b5350da7286164afb83a7b037b528d)

* Disable React StrictMode (#1654)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 6575c327a3ac2b107a6e79a04fa61168d779dbf0)

* docs(upstream-sync): add cherry-pick runbook

Repo-specific runbook for bringing upstream (langchain-ai/open-swe) commits
into the fork: triage-sync discovery, the git cp workflow, the triage ledger,
themed-branch layout, and conflict/regression handling.

---------

Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com>
2026-07-03 11:48:40 -04:00
.githooks feat(infra): upstream-sync triage ledger, cherry-pick hooks, and git cp (#118) 2026-07-03 11:46:00 -04:00
.github Add Dependabot ignore for @types/node semver-major bumps (#112) 2026-07-02 18:15:22 -04:00
.security-review chore: decommission self-hosted AWS LangGraph stack (#64) 2026-06-29 19:54:38 -04:00
.vscode Brace/07 16/fixes (#431) 2025-07-16 13:17:35 -07:00
agent chore: cherry-pick clean upstream fixes + cherry-pick runbook (#117) 2026-07-03 11:48:40 -04:00
deploy refactor: adopt modular webhook architecture (#1621) + port fork customizations (#85) 2026-06-30 18:46:46 -04:00
docs/upstream-sync chore: cherry-pick clean upstream fixes + cherry-pick runbook (#117) 2026-07-03 11:48:40 -04:00
evals/reviewer feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62) 2026-06-29 15:57:19 -04:00
scripts chore: cherry-pick clean upstream fixes + cherry-pick runbook (#117) 2026-07-03 11:48:40 -04:00
static fix: Add back logo to readme (#1068) 2026-03-17 10:41:41 -07:00
tests chore: cherry-pick clean upstream fixes + cherry-pick runbook (#117) 2026-07-03 11:48:40 -04:00
ui chore: cherry-pick clean upstream fixes + cherry-pick runbook (#117) 2026-07-03 11:48:40 -04:00
.codespellignore init commit 2025-05-21 14:47:56 -07:00
.dockerignore feat: move github workflows to gh cli (#1238) 2026-05-04 18:03:53 -07:00
.env.example fix(ui): clear eslint errors from dep bumps + track .env.example (#111) 2026-07-02 17:31:07 -04:00
.gitignore fix(ui): clear eslint errors from dep bumps + track .env.example (#111) 2026-07-02 17:31:07 -04:00
.nvmrc feat: Standardize remaining Node pins on Node 24 (#92) 2026-07-01 13:50:48 -04:00
AGENTS.md chore: sync upstream/main, defer #1621 modular webhooks (#81) 2026-06-30 16:45:19 -04:00
CLAUDE.md feat: distill Sea Haven conventions into agent prompt, reviewer, and fork docs (#113) 2026-07-02 18:29:15 -04:00
CUSTOMIZATION.md chore: sync upstream/main, defer #1621 modular webhooks (#81) 2026-06-30 16:45:19 -04:00
default_prompt.md feat: distill Sea Haven conventions into agent prompt, reviewer, and fork docs (#113) 2026-07-02 18:29:15 -04:00
Dockerfile feat: Standardize remaining Node pins on Node 24 (#92) 2026-07-01 13:50:48 -04:00
INSTALLATION.md feat: Standardize remaining Node pins on Node 24 (#92) 2026-07-01 13:50:48 -04:00
langgraph.json refactor: adopt modular webhook architecture (#1621) + port fork customizations (#85) 2026-06-30 18:46:46 -04:00
LICENSE feat: Monorepo (#22) 2025-05-26 13:02:51 -07:00
Makefile feat(infra): upstream-sync triage ledger, cherry-pick hooks, and git cp (#118) 2026-07-03 11:46:00 -04:00
package.json feat: plan mode with model-driven entry and collaborative review (#1580) 2026-06-23 12:06:58 -07:00
pyproject.toml chore(deps): bump the minor-and-patch group with 3 updates (#104) 2026-07-02 15:24:01 -04:00
README.md chore: align docs, templates, and metadata with Sea Haven handbook [closes SH-93] (#95) 2026-07-01 14:38:39 -04:00
SECURITY.md chore: align docs, templates, and metadata with Sea Haven handbook [closes SH-93] (#95) 2026-07-01 14:38:39 -04:00
uv.lock chore(deps): bump the minor-and-patch group with 3 updates (#104) 2026-07-02 15:24:01 -04:00

Open-source framework for building your org's internal coding agent.

CI License: MIT Python 3.11+ TypeScript 6.0+ Built on LangGraph Built on Deep Agents

Elite engineering orgs like Stripe, Ramp, and Coinbase are building their own internal coding agents — Slackbots, CLIs, and web apps that meet engineers where they already work. These agents are connected to internal systems with the right context, permissioning, and safety boundaries to operate with minimal human oversight.

Open SWE is the open-source version of this pattern. Built on LangGraph and Deep Agents, it gives you the same architecture those companies built internally: cloud sandboxes, Slack and Linear invocation, subagent orchestration, and automatic PR creation — ready to customize for your own codebase and workflows.

Note

Read the announcement blog post here


Architecture

Open SWE makes the same core architectural decisions as the best internal coding agents. Here's how it maps to the patterns described in this overview of Stripe's Minions, Ramp's Inspect, and Coinbase's Cloudbot:

1. Agent Harness — Composed on Deep Agents

Rather than forking an existing agent or building from scratch, Open SWE composes on the Deep Agents framework — similar to how Ramp built on top of OpenCode. This gives you an upgrade path (pull in upstream improvements) while letting you customize the orchestration, tools, and middleware for your org.

create_deep_agent(
    model="openai:gpt-5.5",
    system_prompt=construct_system_prompt(...),
    tools=[http_request, fetch_url, linear_comment, slack_thread_reply],
    backend=sandbox_backend,
    middleware=[ToolErrorMiddleware(), check_message_queue_before_model, ...],
)

2. Sandbox — Isolated Cloud Environments

Every task runs in its own isolated cloud sandbox — a remote Linux environment with full shell access. The repo is cloned in, the agent gets full permissions, and the blast radius of any mistake is fully contained. No production access, no confirmation prompts.

Open SWE supports multiple sandbox providers out of the box — Modal, Daytona, Runloop, and LangSmith — and you can plug in your own. See the Customization Guide for details.

This follows the principle all three companies converge on: isolate first, then give full permissions inside the boundary.

  • Each thread gets a persistent sandbox (reused across follow-up messages)
  • Sandboxes auto-recreate if they become unreachable
  • Multiple tasks run in parallel — each in its own sandbox, no queuing

3. Tools — Curated, Not Accumulated

Stripe's key insight: tool curation matters more than tool quantity. Open SWE follows this principle with a small, focused toolset:

Tool Purpose
execute Shell commands in the sandbox
fetch_url Fetch web pages as markdown
http_request API calls (GET, POST, etc.)
linear_comment Post updates to Linear tickets
slack_thread_reply Reply in Slack threads

GitHub operations are performed with GH_TOKEN=dummy gh inside the sandbox, backed by the LangSmith proxy. Plus the built-in Deep Agents tools: read_file, write_file, edit_file, ls, glob, grep, write_todos, and task (subagent spawning).

Optional observability tools (server-side): Admins can connect Datadog and LangSmith from team settings (Admin → Observability credentials). When connected, the agent gains Datadog tools (via Datadog's hosted MCP server, default toolsets=core) and read-only LangSmith tools (langsmith_get_trace, langsmith_list_runs). These run in the LangGraph server process using credentials encrypted at rest — the sandbox never holds Datadog or LangSmith keys. They are loaded only for runs triggered by an authorized user (admins, plus any emails in OBSERVABILITY_AUTHORIZED_EMAILS), so a prompt-injected run from an untrusted contributor cannot reach team observability data. Use scoped, read-oriented keys regardless: observability data (logs, traces) is attacker-influenced content that can carry prompt injection, and the agent has network egress — the same residual-risk class as web_search / fetch_url.

Optional Corridor guardrails (server-side MCP): Set CORRIDOR_API_TOKEN (or CORRIDOR_MCP_TOKEN / CORRIDOR_TOKEN) to load Corridor's hosted MCP server for each agent run. Open SWE exposes only Corridor's analyzePlan tool. CORRIDOR_MCP_URL defaults to https://app.corridor.dev/api/mcp; if set explicitly, Open SWE only accepts the same HTTPS host and /api/mcp path. Tokens are sent via Authorization: Bearer ... from the LangGraph server process and are never placed in the sandbox. A legacy ?token=... URL is accepted and normalized into the header form.

4. Context Engineering — AGENTS.md + Source Context

Open SWE gathers context from two sources:

  • AGENTS.md — If the repo contains an AGENTS.md file at the root, it's read from the sandbox and injected into the system prompt. This is your repo-level equivalent of Stripe's rule files: encoding conventions, testing requirements, and architectural decisions that every agent run should follow.
  • Source context — The full Linear issue (title, description, comments) or Slack thread history is assembled and passed to the agent, so it starts with rich context rather than discovering everything through tool calls.

5. Orchestration — Subagents + Middleware

Open SWE's orchestration has two layers:

Subagents: The Deep Agents framework natively supports spawning child agents via the task tool. The main agent can fan out independent subtasks to isolated subagents — each with its own middleware stack, todo list, and file operations. This is similar to Ramp's child sessions for parallel work.

Middleware: Deterministic middleware hooks run around the agent loop:

  • check_message_queue_before_model — Injects follow-up messages (Linear comments or Slack messages that arrive mid-run) before the next model call. You can message the agent while it's working and it'll pick up your input at its next step.
  • notify_step_limit_reached — After-agent hook that posts a Slack reply when the agent hits the model-call limit, so users get a clear signal instead of silence.
  • ToolErrorMiddleware — Catches and handles tool errors gracefully.

6. Invocation — Slack, Linear, and GitHub

All three companies in the article converge on Slack as the primary invocation surface. Open SWE does the same:

  • Slack — Mention the bot in any thread. Supports repo:owner/name syntax to specify which repo to work on. The agent replies in-thread with status updates and PR links.
  • Linear — Comment @openswe on any issue. The agent reacts with 👀 to acknowledge, reads the full issue context, and posts results back as comments.
  • GitHub — Tag @openswe in PR comments on agent-created PRs to have it address review feedback and push fixes to the same branch.

Each invocation creates a deterministic thread ID, so follow-up messages on the same issue or thread route to the same running agent.

Trigger tags (Sea Haven fork): a mention is a case-insensitive substring match on the comment body — @openswe, @open-swe, @openswe-dev, or @seahaven-openswe (the deployed App slug). GitHub won't linkify @seahaven-openswe (App [bot] accounts aren't user-mentionable), but the text still fires a run.

Engineering conventions & attribution (Sea Haven fork): the main agent's system prompt is tuned to the Sea Haven engineering handbook — branch names are feature|bug|hotfix/<kebab-desc> (optional resolvable <KEY>- prefix), PR bodies use ## Summary / Validation / Tests / Notes, and commit messages follow the handbook format (≤50-char imperative subject, why over what). The PR title rule is repo-aware: when the target repo enforces a conventional-commit title (an amannn/action-semantic-pull-request workflow, a commitlint config, or a documented requirement in AGENTS.md / CONTRIBUTING.md), the agent emits a conforming type(scope): … title that reads the action's allowed types/scopes — this lets it pass gates like this repo's own PR Title Lint and upstream langchain-ai/open-swe without manual retitling; otherwise it falls back to the Sea Haven imperative style with no type: prefix. PRs that resolve a GitHub issue auto-link it in the body (Closes #<n> for full fixes, Refs #<n>/Part of #<n> for partial work, Closes owner/repo#<n> cross-repo); because the Sea Haven flow targets dev rather than the default branch, the issue closes when dev is promoted, not at dev-merge. No agent/AI attribution is added to any artifact — no Co-authored-by bot trailer, no Made by [Open SWE] footer, no "generated by an agent" notes. Commits are currently authored as the triggering user (the upstream behavior, which keeps Vercel preview deploys resolvable); flipping authorship to the bot account is tracked separately in issue #11 pending the Vercel-resolvability decision.

7. Validation — Prompt-Driven

The agent is instructed to run linters, formatters, and tests before committing, and is responsible end-to-end for committing, pushing, opening/updating the draft PR, and replying in the source channel. This is an area where you can extend Open SWE for your org: add deterministic CI checks, visual verification, or review gates as additional middleware. See the Customization Guide for how.


Comparison

Decision Open SWE Stripe (Minions) Ramp (Inspect) Coinbase (Cloudbot)
Harness Composed (Deep Agents/LangGraph) Forked (Goose) Composed (OpenCode) Built from scratch
Sandbox Pluggable (Modal, Daytona, Runloop, etc.) AWS EC2 devboxes (pre-warmed) Modal containers (pre-warmed) In-house
Tools ~15, curated ~500, curated per-agent OpenCode SDK + extensions MCPs + custom Skills
Context AGENTS.md + issue/thread Rule files + pre-hydration OpenCode built-in Linear-first + MCPs
Orchestration Subagents + middleware Blueprints (deterministic + agentic) Sessions + child sessions Three modes
Invocation Slack, Linear, GitHub Slack + embedded buttons Slack + web + Chrome extension Slack-native
Validation Prompt-driven 3-layer (local + CI + 1 retry) Visual DOM verification Agent councils + auto-merge

Features

  • Trigger from Linear, Slack, or GitHub — mention @openswe in a comment to kick off a task
  • Instant acknowledgement — acknowledges the moment it picks up your message
  • Message it while it's running — send follow-up messages mid-task and it'll pick them up before its next step
  • Run multiple tasks in parallel — each task runs in its own isolated cloud sandbox
  • GitHub OAuth built-in — authenticates with your GitHub account automatically
  • Opens PRs automatically — commits changes and opens a draft PR when done, linked back to your ticket
  • Subagent support — the agent can spawn child agents for parallel subtasks
  • Web dashboard — a companion app (in ui/) for GitHub login, per-user model/profile settings, team defaults, enabled-repo and review-style management, user mappings, and an Agents chat UI

Getting Started

  • Installation Guide — local dev (backend + dashboard), GitHub App creation, LangSmith, Linear/Slack/GitHub triggers, and production deployment
  • Customization Guide — swap the sandbox, model, tools, triggers, system prompt, and middleware for your org

Deployment (Sea Haven fork)

This fork runs on a managed deployment: the backend (all three graphs + the FastAPI webapp) runs on LangGraph Cloud / Platform, and the ui/ dashboard deploys to Vercel. Configuration and secrets live in the LangGraph deployment config and Vercel environment variables. Promotion from dev to prod (main) is handled by .github/workflows/promote-to-main.yml.

See INSTALLATION.md § 10 "Production deployment" for the full backend + dashboard setup.

The earlier self-hosted AWS stack (CDK under infra/, an ARM64 EC2 box + nginx behind the shared ALB, and the cd-infra / build-artifacts release pipelines) was decommissioned in favor of the managed deployment above.

License

MIT