open-swe/README.md

148 lines
9.5 KiB
Markdown
Raw Normal View History

<div align="center">
<a href="https://github.com/langchain-ai/open-swe">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="static/dark.svg">
<source media="(prefers-color-scheme: light)" srcset="static/light.svg">
<img alt="Open SWE Logo" src="static/dark.svg" width="35%">
</picture>
</a>
</div>
<div align="center">
<h3>Open-source framework for building your org's internal coding agent.</h3>
</div>
<div align="center">
<a href="https://opensource.org/licenses/MIT" target="_blank"><img src="https://img.shields.io/github/license/langchain-ai/open-swe" alt="License"></a>
<a href="https://github.com/langchain-ai/open-swe/stargazers" target="_blank"><img src="https://img.shields.io/github/stars/langchain-ai/open-swe" alt="GitHub Stars"></a>
<a href="https://github.com/langchain-ai/langgraph" target="_blank"><img src="https://img.shields.io/badge/Built%20on-LangGraph-blue" alt="Built on LangGraph"></a>
<a href="https://github.com/langchain-ai/deepagents" target="_blank"><img src="https://img.shields.io/badge/Built%20on-Deep%20Agents-blue" alt="Built on Deep Agents"></a>
<a href="https://x.com/langchain" target="_blank"><img src="https://img.shields.io/twitter/url/https/twitter.com/langchain.svg?style=social&label=Follow%20%40LangChain" alt="Twitter / X"></a>
</div>
<br>
2026-03-07 13:24:10 -08:00
Elite engineering orgs like Stripe, Ramp, and Coinbase are building their own internal coding agents — Slackbots, CLIs, and web apps that meet engineers where they already work. These agents are connected to internal systems with the right context, permissioning, and safety boundaries to operate with minimal human oversight.
2026-03-07 13:24:10 -08:00
Open SWE is the open-source version of this pattern. Built on [LangGraph](https://langchain-ai.github.io/langgraph/) and [Deep Agents](https://github.com/langchain-ai/deepagents), it gives you the same architecture those companies built internally: cloud sandboxes, Slack and Linear invocation, subagent orchestration, and automatic PR creation — ready to customize for your own codebase and workflows.
2025-08-04 17:02:24 -07:00
> [!NOTE]
> 💬 Read the **announcement blog post [here](https://blog.langchain.com/open-swe-an-open-source-framework-for-internal-coding-agents/)**
2026-03-07 13:24:10 -08:00
---
2026-03-07 13:24:10 -08:00
## Architecture
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
Open SWE makes the same core architectural decisions as the best internal coding agents. Here's how it maps to the patterns described in [this overview](https://x.com/kishan_dahya/status/2028971339974099317) of Stripe's Minions, Ramp's Inspect, and Coinbase's Cloudbot:
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
### 1. Agent Harness — Composed on Deep Agents
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
Rather than forking an existing agent or building from scratch, Open SWE **composes** on the [Deep Agents](https://github.com/langchain-ai/deepagents) framework — similar to how Ramp built on top of OpenCode. This gives you an upgrade path (pull in upstream improvements) while letting you customize the orchestration, tools, and middleware for your org.
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
```python
create_deep_agent(
model="openai:gpt-5.5",
feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] (#1159) * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * feat: authenticate git operations via sandbox proxy instead of credential files * removing logger.info * formatting and linting * fix: resolve lint errors in server.py (imports, unused vars, undefined names) * feat: use opaque proxy headers for GitHub auth in sandbox * linting formatting and test changes * linting * Delete .claude directory * Delete tests/evals directory * fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests * fix: restore authorship, branch_name support, and installation token for PR creation * linitng * fix: move installation token fetch before commit, clean up dead proxy validation code * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] * fix: address review feedback — restore agents_md, add git user config, lint fixes * fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config * fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config * linting * linting * feat: add installation token auth to list_repos GitHub API call * agents.md update * linting * fix: address PR review feedback — shell precedence bug in prompt, remove dead code * linting * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Apply suggestion from @bracesproul Co-authored-by: Brace Sproul <braceasproul@gmail.com> * fix: address PR review feedback — restore {working_dir} in prompt, remove clone code block * fix:Extract check_or_recreate_sandbox utility from inline sandbox health check * fix: address PR review feedback — async list_repos, restore template name, fix prompt colon * fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox * linting * yogesh/ope-21-stop-auto-cloning * Update agent/tools/list_repos.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * Update agent/prompt.py Co-authored-by: Brace Sproul <braceasproul@gmail.com> * feat: address PR review — list_repos uses GitHub API only, PR trigger includes org/repo * linting * feat: address PR review feedback — list_repos pagination, simpler return, sandbox health check * feat: support listing repos for personal user accounts via is_organization flag --------- Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2026-04-10 17:04:55 -07:00
system_prompt=construct_system_prompt(...),
tools=[http_request, fetch_url, linear_comment, slack_thread_reply],
2026-03-07 13:24:10 -08:00
backend=sandbox_backend,
middleware=[ToolErrorMiddleware(), check_message_queue_before_model, ...],
)
```
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
### 2. Sandbox — Isolated Cloud Environments
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
Every task runs in its own **isolated cloud sandbox** — a remote Linux environment with full shell access. The repo is cloned in, the agent gets full permissions, and the blast radius of any mistake is fully contained. No production access, no confirmation prompts.
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
Open SWE supports multiple sandbox providers out of the box — [Modal](https://modal.com/), [Daytona](https://www.daytona.io/), [Runloop](https://www.runloop.ai/), and [LangSmith](https://smith.langchain.com/) — and you can plug in your own. See the [Customization Guide](CUSTOMIZATION.md#1-sandbox) for details.
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
This follows the principle all three companies converge on: **isolate first, then give full permissions inside the boundary.**
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
- Each thread gets a persistent sandbox (reused across follow-up messages)
- Sandboxes auto-recreate if they become unreachable
- Multiple tasks run in parallel — each in its own sandbox, no queuing
2026-03-06 14:47:05 -08:00
2026-03-07 13:24:10 -08:00
### 3. Tools — Curated, Not Accumulated
2026-03-06 14:47:05 -08:00
2026-03-07 13:24:10 -08:00
Stripe's key insight: *tool curation matters more than tool quantity.* Open SWE follows this principle with a small, focused toolset:
2026-03-06 15:05:42 -08:00
2026-03-07 13:24:10 -08:00
| Tool | Purpose |
|---|---|
| `execute` | Shell commands in the sandbox |
| `fetch_url` | Fetch web pages as markdown |
| `http_request` | API calls (GET, POST, etc.) |
| `linear_comment` | Post updates to Linear tickets |
| `slack_thread_reply` | Reply in Slack threads |
2026-03-06 14:47:05 -08:00
GitHub operations are performed with `GH_TOKEN=dummy gh` inside the sandbox, backed by the LangSmith proxy. Plus the built-in Deep Agents tools: `read_file`, `write_file`, `edit_file`, `ls`, `glob`, `grep`, `write_todos`, and `task` (subagent spawning).
2026-03-06 14:47:05 -08:00
2026-03-07 13:24:10 -08:00
### 4. Context Engineering — AGENTS.md + Source Context
2026-03-06 14:47:05 -08:00
2026-03-07 13:24:10 -08:00
Open SWE gathers context from two sources:
2026-03-06 14:47:05 -08:00
2026-03-07 13:24:10 -08:00
- **`AGENTS.md`** — If the repo contains an `AGENTS.md` file at the root, it's read from the sandbox and injected into the system prompt. This is your repo-level equivalent of Stripe's rule files: encoding conventions, testing requirements, and architectural decisions that every agent run should follow.
- **Source context** — The full Linear issue (title, description, comments) or Slack thread history is assembled and passed to the agent, so it starts with rich context rather than discovering everything through tool calls.
2026-03-06 14:47:05 -08:00
2026-03-07 13:24:10 -08:00
### 5. Orchestration — Subagents + Middleware
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
Open SWE's orchestration has two layers:
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
**Subagents:** The Deep Agents framework natively supports spawning child agents via the `task` tool. The main agent can fan out independent subtasks to isolated subagents — each with its own middleware stack, todo list, and file operations. This is similar to Ramp's child sessions for parallel work.
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
**Middleware:** Deterministic middleware hooks run around the agent loop:
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
- **`check_message_queue_before_model`** — Injects follow-up messages (Linear comments or Slack messages that arrive mid-run) before the next model call. You can message the agent while it's working and it'll pick up your input at its next step.
- **`notify_step_limit_reached`** — After-agent hook that posts a Slack reply when the agent hits the model-call limit, so users get a clear signal instead of silence.
2026-03-07 13:24:10 -08:00
- **`ToolErrorMiddleware`** — Catches and handles tool errors gracefully.
2026-03-07 13:24:10 -08:00
### 6. Invocation — Slack, Linear, and GitHub
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
All three companies in the article converge on **Slack as the primary invocation surface**. Open SWE does the same:
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
- **Slack** — Mention the bot in any thread. Supports `repo:owner/name` syntax to specify which repo to work on. The agent replies in-thread with status updates and PR links.
- **Linear** — Comment `@openswe` on any issue. The agent reads the full issue context, reacts with 👀 to acknowledge, and posts results back as comments.
- **GitHub** — Tag `@openswe` in PR comments on agent-created PRs to have it address review feedback and push fixes to the same branch.
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
Each invocation creates a deterministic thread ID, so follow-up messages on the same issue or thread route to the same running agent.
2026-03-06 14:23:29 -08:00
### 7. Validation — Prompt-Driven
2026-03-06 14:23:29 -08:00
The agent is instructed to run linters, formatters, and tests before committing, and is responsible end-to-end for committing, pushing, opening/updating the draft PR, and replying in the source channel.
2026-03-07 13:24:10 -08:00
This is an area where you can extend Open SWE for your org: add deterministic CI checks, visual verification, or review gates as additional middleware. See the [Customization Guide](CUSTOMIZATION.md#6-middleware) for how.
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
---
## Comparison
| Decision | Open SWE | Stripe (Minions) | Ramp (Inspect) | Coinbase (Cloudbot) |
|---|---|---|---|---|
| **Harness** | Composed (Deep Agents/LangGraph) | Forked (Goose) | Composed (OpenCode) | Built from scratch |
| **Sandbox** | Pluggable (Modal, Daytona, Runloop, etc.) | AWS EC2 devboxes (pre-warmed) | Modal containers (pre-warmed) | In-house |
| **Tools** | ~15, curated | ~500, curated per-agent | OpenCode SDK + extensions | MCPs + custom Skills |
| **Context** | AGENTS.md + issue/thread | Rule files + pre-hydration | OpenCode built-in | Linear-first + MCPs |
| **Orchestration** | Subagents + middleware | Blueprints (deterministic + agentic) | Sessions + child sessions | Three modes |
| **Invocation** | Slack, Linear, GitHub | Slack + embedded buttons | Slack + web + Chrome extension | Slack-native |
| **Validation** | Prompt-driven | 3-layer (local + CI + 1 retry) | Visual DOM verification | Agent councils + auto-merge |
2026-03-06 14:23:29 -08:00
2026-03-07 13:24:10 -08:00
---
## Features
2026-03-06 14:36:26 -08:00
2026-03-07 13:24:10 -08:00
- **Trigger from Linear, Slack, or GitHub** — mention `@openswe` in a comment to kick off a task
- **Instant acknowledgement** — reacts with 👀 the moment it picks up your message
- **Message it while it's running** — send follow-up messages mid-task and it'll pick them up before its next step
- **Run multiple tasks in parallel** — each task runs in its own isolated cloud sandbox
- **GitHub OAuth built-in** — authenticates with your GitHub account automatically
- **Opens PRs automatically** — commits changes and opens a draft PR when done, linked back to your ticket
- **Subagent support** — the agent can spawn child agents for parallel subtasks
2026-03-06 14:36:26 -08:00
---
2026-03-07 13:24:10 -08:00
## Getting Started
- **[Installation Guide](INSTALLATION.md)** — GitHub App creation, LangSmith, Linear/Slack/GitHub triggers, and production deployment
- **[Customization Guide](CUSTOMIZATION.md)** — swap the sandbox, model, tools, triggers, system prompt, and middleware for your org
2026-03-07 13:24:10 -08:00
## License
2026-03-07 13:24:10 -08:00
MIT