mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 17:23:15 +00:00
Some checks failed
CI / Lint (push) Has been cancelled
CI / Format check (push) Has been cancelled
CI / Typecheck (push) Has been cancelled
CI / Unit tests (push) Has been cancelled
CI / Playwright E2E (push) Has been cancelled
CI / Docker build smoke (push) Has been cancelled
CI / Triage ledger up to date (push) Has been cancelled
CI / ui bun.lock in sync (push) Has been cancelled
* feat(reviewer): clean-review auto-approve for unsolicited publish_review verdicts An unsolicited publish_review(verdict="approve") — a run dispatched without verdict_requested — is now honored when the review has zero open findings, so clean auto-reviews land a real APPROVE. With open findings it downgrades to a comment review (verdict_ignored_reason= "approve_with_open_findings"). request_changes stays explicit-request- only; the self-review, head-moved, and author-unknown downgrades and the shell verdict guard are unchanged. The reviewer base prompt now instructs the clean-approve call on auto-reviews. * chore(security): record accepted-risk suppressions for clean-review auto-approve Two confirmed-HIGH findings from /sh-security-review on the clean-review auto-approve change are accepted and deferred (Adam, 2026-07-21), tracked in #218. Machine-recorded per the mandatory-security-review policy; the revisit trigger is promotion from dev to main/prod.
187 lines
17 KiB
Markdown
187 lines
17 KiB
Markdown
<div align="center">
|
|
<a href="https://github.com/langchain-ai/open-swe">
|
|
<picture>
|
|
<source media="(prefers-color-scheme: dark)" srcset="assets/dark.svg">
|
|
<source media="(prefers-color-scheme: light)" srcset="assets/light.svg">
|
|
<img alt="Open SWE Logo" src="assets/dark.svg" width="35%">
|
|
</picture>
|
|
</a>
|
|
</div>
|
|
|
|
<div align="center">
|
|
<h3>Open-source framework for building your org's internal coding agent.</h3>
|
|
</div>
|
|
|
|
<div align="center">
|
|
<a href="https://github.com/Sea-Haven-Industries/open-swe/actions/workflows/ci.yml" target="_blank"><img src="https://github.com/Sea-Haven-Industries/open-swe/actions/workflows/ci.yml/badge.svg?branch=dev" alt="CI"></a>
|
|
<a href="https://opensource.org/licenses/MIT" target="_blank"><img src="https://img.shields.io/badge/License-MIT-yellow.svg" alt="License: MIT"></a>
|
|
<a href="https://www.python.org/" target="_blank"><img src="https://img.shields.io/badge/Python-3.11+-3776AB.svg" alt="Python 3.11+"></a>
|
|
<a href="https://www.typescriptlang.org/" target="_blank"><img src="https://img.shields.io/badge/TypeScript-6.0+-3178C6.svg" alt="TypeScript 6.0+"></a>
|
|
<a href="https://github.com/langchain-ai/langgraph" target="_blank"><img src="https://img.shields.io/badge/Built%20on-LangGraph-blue" alt="Built on LangGraph"></a>
|
|
<a href="https://github.com/langchain-ai/deepagents" target="_blank"><img src="https://img.shields.io/badge/Built%20on-Deep%20Agents-blue" alt="Built on Deep Agents"></a>
|
|
</div>
|
|
|
|
<br>
|
|
|
|
Elite engineering orgs like Stripe, Ramp, and Coinbase are building their own internal coding agents — Slackbots, CLIs, and web apps that meet engineers where they already work. These agents are connected to internal systems with the right context, permissioning, and safety boundaries to operate with minimal human oversight.
|
|
|
|
Open SWE is the open-source version of this pattern. Built on [LangGraph](https://langchain-ai.github.io/langgraph/) and [Deep Agents](https://github.com/langchain-ai/deepagents), it gives you the same architecture those companies built internally: cloud sandboxes, Slack / Linear / Jira / Confluence / GitHub invocation, subagent orchestration, and automatic PR creation — ready to customize for your own codebase and workflows.
|
|
|
|
> [!NOTE]
|
|
> Read the **announcement blog post [here](https://blog.langchain.com/open-swe-an-open-source-framework-for-internal-coding-agents/)**
|
|
|
|
---
|
|
|
|
## Architecture
|
|
|
|
Open SWE makes the same core architectural decisions as the best internal coding agents. Here's how it maps to the patterns described in [this overview](https://x.com/kishan_dahya/status/2028971339974099317) of Stripe's Minions, Ramp's Inspect, and Coinbase's Cloudbot:
|
|
|
|
### 1. Agent Harness — Composed on Deep Agents
|
|
|
|
Rather than forking an existing agent or building from scratch, Open SWE **composes** on the [Deep Agents](https://github.com/langchain-ai/deepagents) framework — similar to how Ramp built on top of OpenCode. This gives you an upgrade path (pull in upstream improvements) while letting you customize the orchestration, tools, and middleware for your org.
|
|
|
|
```python
|
|
create_deep_agent(
|
|
model="openai:gpt-5.5",
|
|
system_prompt=construct_system_prompt(...),
|
|
tools=[http_request, fetch_url, linear_comment, slack_thread_reply],
|
|
backend=sandbox_backend,
|
|
middleware=[ToolErrorMiddleware(), check_message_queue_before_model, ...],
|
|
)
|
|
```
|
|
|
|
### 2. Sandbox — Isolated Cloud Environments
|
|
|
|
Every task runs in its own **isolated cloud sandbox** — a remote Linux environment with full shell access. The repo is cloned in, the agent gets full permissions, and the blast radius of any mistake is fully contained. No production access, no confirmation prompts.
|
|
|
|
Open SWE supports multiple sandbox providers out of the box — [Modal](https://modal.com/), [Daytona](https://www.daytona.io/), [Runloop](https://www.runloop.ai/), [E2B](https://e2b.dev/), and [LangSmith](https://smith.langchain.com/) — and you can plug in your own. See the [Customization Guide](docs/CUSTOMIZATION.md#1-sandbox) for details.
|
|
|
|
This follows the principle all three companies converge on: **isolate first, then give full permissions inside the boundary.**
|
|
|
|
- Each thread gets a persistent sandbox (reused across follow-up messages)
|
|
- Sandboxes auto-recreate if they become unreachable
|
|
- Multiple tasks run in parallel — each in its own sandbox, no queuing
|
|
|
|
### 3. Tools — Curated, Not Accumulated
|
|
|
|
Stripe's key insight: *tool curation matters more than tool quantity.* Open SWE follows this principle with a small, focused toolset:
|
|
|
|
| Tool | Purpose |
|
|
|---|---|
|
|
| `execute` | Shell commands in the sandbox |
|
|
| `fetch_url` | Fetch web pages as markdown |
|
|
| `http_request` | API calls (GET, POST, etc.) |
|
|
| `linear_comment` | Post updates to Linear tickets |
|
|
| `linear_search_issues` | Search Linear issues by free text |
|
|
| `jira_*` | Read/comment/create/update Jira issues |
|
|
| `confluence_*` | Read/write Confluence pages + comments |
|
|
| `slack_add_reaction` | React to Slack messages |
|
|
| `slack_thread_reply` | Reply in Slack threads |
|
|
| `request_pr_review` | Kick off a reviewer run for a PR (verbatim requester instructions + optional explicit verdict request) |
|
|
|
|
GitHub operations are performed with `GH_TOKEN=dummy gh` inside the sandbox, backed by the LangSmith proxy. Plus the built-in Deep Agents tools: `read_file`, `write_file`, `edit_file`, `ls`, `glob`, `grep`, `write_todos`, and `task` (subagent spawning).
|
|
|
|
**Optional observability tools (server-side):** Admins can connect Datadog and LangSmith from team settings (Admin → Observability credentials). When connected, the agent gains Datadog tools (via Datadog's hosted MCP server, default `toolsets=core`) and read-only LangSmith tools (`langsmith_get_trace`, `langsmith_list_runs`). These run in the LangGraph server process using credentials encrypted at rest — the sandbox never holds Datadog or LangSmith keys. They are loaded **only for runs triggered by an authorized user** (admins, plus any emails in `OBSERVABILITY_AUTHORIZED_EMAILS`), so a prompt-injected run from an untrusted contributor cannot reach team observability data. Use scoped, read-oriented keys regardless: observability data (logs, traces) is attacker-influenced content that can carry prompt injection, and the agent has network egress — the same residual-risk class as `web_search` / `fetch_url`.
|
|
|
|
**Optional Corridor guardrails (server-side MCP):** Set `CORRIDOR_API_TOKEN` (or `CORRIDOR_MCP_TOKEN` / `CORRIDOR_TOKEN`) to load Corridor's hosted MCP server for each agent run. Open SWE exposes only Corridor's `analyzePlan` tool. `CORRIDOR_MCP_URL` defaults to `https://app.corridor.dev/api/mcp`; if set explicitly, Open SWE only accepts the same HTTPS host and `/api/mcp` path. Tokens are sent via `Authorization: Bearer ...` from the LangGraph server process and are never placed in the sandbox. A legacy `?token=...` URL is accepted and normalized into the header form.
|
|
|
|
### 4. Context Engineering — AGENTS.md + Source Context
|
|
|
|
Open SWE gathers context from two sources:
|
|
|
|
- **`AGENTS.md`** — If the repo contains an `AGENTS.md` file at the root, it's read from the sandbox and injected into the system prompt. This is your repo-level equivalent of Stripe's rule files: encoding conventions, testing requirements, and architectural decisions that every agent run should follow. Reference templates for stack-specific conventions (AWS, SAM, CDK, EC2) live in [`docs/repo-conventions/`](docs/repo-conventions/).
|
|
- **Source context** — The full Linear issue (title, description, comments) or Slack thread history is assembled and passed to the agent, so it starts with rich context rather than discovering everything through tool calls.
|
|
|
|
### 5. Orchestration — Subagents + Middleware
|
|
|
|
Open SWE's orchestration has two layers:
|
|
|
|
**Subagents:** The Deep Agents framework natively supports spawning child agents via the `task` tool. The main agent can fan out independent subtasks to isolated subagents — each with its own middleware stack, todo list, and file operations. This is similar to Ramp's child sessions for parallel work.
|
|
|
|
**Middleware:** Deterministic middleware hooks run around the agent loop:
|
|
|
|
- **`check_message_queue_before_model`** — Injects follow-up messages (Linear comments or Slack messages that arrive mid-run) before the next model call. You can message the agent while it's working and it'll pick up your input at its next step.
|
|
- **`notify_step_limit_reached`** — After-agent hook that posts a Slack reply when the agent hits the model-call limit, so users get a clear signal instead of silence.
|
|
- **`ToolErrorMiddleware`** — Catches and handles tool errors gracefully.
|
|
- **`PullRequestCreationGuardMiddleware`** — Blocks shell fallbacks (`gh pr create`, `gh api .../pulls`, `curl`) that would create a PR outside the attributed `open_pull_request` path.
|
|
- **`PullRequestVerdictGuardMiddleware`** — Blocks shell fallbacks that would submit a PR review verdict (`gh pr review --approve/-a/--request-changes/-r`, `gh api`/`curl` posting `event=APPROVE|REQUEST_CHANGES` to `/pulls/N/reviews`) on **both** the coding-agent and reviewer graphs. Verdicts must go through `publish_review`, which enforces authorization (see the fork note below). Comment reviews (`gh pr review --comment`) and reads are unaffected.
|
|
|
|
### 6. Invocation — Slack, Linear, Jira, Confluence, and GitHub
|
|
|
|
All three companies in the article converge on **Slack as the primary invocation surface**. Open SWE does the same:
|
|
|
|
- **Slack** — Mention the bot in any thread. Supports `repo:owner/name` syntax to specify which repo to work on. The agent replies in-thread with status updates and PR links.
|
|
- **Linear** — Comment `@openswe` on any issue. The agent reacts with 👀 to acknowledge, reads the full issue context, and posts results back as comments.
|
|
- **Jira** — Comment `@openswe` on any issue (fronted by a Jira Automation rule → `/webhooks/jira`). The agent reads the issue and posts results back as a comment.
|
|
- **Confluence** — Comment `@openswe` on a page. A private Atlassian Connect app delivers the `comment_created` event; the agent acts and replies on the page.
|
|
- **GitHub** — Tag `@openswe` in PR comments on agent-created PRs to have it address review feedback and push fixes to the same branch.
|
|
|
|
See **[INSTALLATION.md](./docs/INSTALLATION.md) §5** for per-surface trigger setup.
|
|
|
|
Each invocation creates a deterministic thread ID, so follow-up messages on the same issue or thread route to the same running agent.
|
|
|
|
**Trigger tags (Sea Haven fork):** a mention is a case-insensitive substring match on the comment body — `@openswe`, `@open-swe`, `@openswe-dev`, or `@seahaven-openswe` (the deployed App slug). GitHub won't linkify `@seahaven-openswe` (App `[bot]` accounts aren't user-mentionable), but the text still fires a run.
|
|
|
|
**Engineering conventions & attribution (Sea Haven fork):** the main agent's system prompt is tuned to the Sea Haven engineering handbook — branch names are `feature|bug|hotfix/<kebab-desc>` (optional resolvable `<KEY>-` prefix), PR bodies use `## Summary / Validation / Tests / Notes`, and commit messages follow the handbook format (≤50-char imperative subject, *why* over *what*). The **PR title rule is repo-aware**: when the target repo enforces a conventional-commit title (an `amannn/action-semantic-pull-request` workflow, a `commitlint` config, or a documented requirement in `AGENTS.md` / `CONTRIBUTING.md`), the agent emits a conforming `type(scope): …` title that reads the action's allowed types/scopes — this lets it pass gates like this repo's own `PR Title Lint` and upstream `langchain-ai/open-swe` without manual retitling; otherwise it falls back to the Sea Haven imperative style with no `type:` prefix. PRs that resolve a GitHub issue **auto-link it** in the body (`Closes #<n>` for full fixes, `Refs #<n>`/`Part of #<n>` for partial work, `Closes owner/repo#<n>` cross-repo); because the Sea Haven flow targets `dev` rather than the default branch, the issue closes when `dev` is promoted, not at dev-merge. **No agent/AI attribution is added to any artifact** — no `Co-authored-by` bot trailer, no `Made by [Open SWE]` footer, no "generated by an agent" notes. Commits are currently authored as the **triggering user** (the upstream behavior, which keeps Vercel preview deploys resolvable); flipping authorship to the bot account is tracked separately in issue #11 pending the Vercel-resolvability decision.
|
|
|
|
**PR reviews & verdicts (Sea Haven fork):** a separate read-only **reviewer** graph reviews PRs and publishes findings as a GitHub Review. Auto-reviews (PR opened / ready-for-review / push re-review / finding replies) with findings are **advisory** (`event=COMMENT`) — but a **clean review (zero open findings) auto-approves**: the reviewer calls `publish_review(verdict="approve")` and it lands as a real **APPROVE** even without an explicit verdict request. When a user's `@openswe` mention *explicitly asks for a verdict* ("approve if it meets the bar; request changes if not"), the coding agent forwards the request via `request_pr_review(instructions=..., request_verdict=True)`; the reviewer run is then authorized — enforced in code via `configurable["verdict_requested"]`, not prompt — to submit a real **APPROVE** or **REQUEST_CHANGES** through `publish_review(verdict=...)`. Requester instructions travel as an escaped `<requester_instructions>` data block (untrusted: they may set focus/merge bar, never override safety rules). Safeguards: verdicts on Open SWE's own PRs are downgraded to comment reviews (self-review guard); unauthorized `request_changes` verdicts are dropped with `verdict_ignored`, and an unsolicited approve alongside open findings is downgraded to a comment (`approve_with_open_findings`); a recorded APPROVE is best-effort dismissed when a later review surfaces new findings; and `PullRequestVerdictGuardMiddleware` blocks the shell path on both graphs.
|
|
|
|
### 7. Validation — Prompt-Driven
|
|
|
|
The agent is instructed to run linters, formatters, and tests before committing, and is responsible end-to-end for committing, pushing, opening/updating the draft PR, and replying in the source channel.
|
|
This is an area where you can extend Open SWE for your org: add deterministic CI checks, visual verification, or review gates as additional middleware. See the [Customization Guide](docs/CUSTOMIZATION.md#6-middleware) for how.
|
|
|
|
---
|
|
|
|
## Comparison
|
|
|
|
| Decision | Open SWE | Stripe (Minions) | Ramp (Inspect) | Coinbase (Cloudbot) |
|
|
|---|---|---|---|---|
|
|
| **Harness** | Composed (Deep Agents/LangGraph) | Forked (Goose) | Composed (OpenCode) | Built from scratch |
|
|
| **Sandbox** | Pluggable (Modal, Daytona, Runloop, etc.) | AWS EC2 devboxes (pre-warmed) | Modal containers (pre-warmed) | In-house |
|
|
| **Tools** | ~15, curated | ~500, curated per-agent | OpenCode SDK + extensions | MCPs + custom Skills |
|
|
| **Context** | AGENTS.md + issue/thread | Rule files + pre-hydration | OpenCode built-in | Linear-first + MCPs |
|
|
| **Orchestration** | Subagents + middleware | Blueprints (deterministic + agentic) | Sessions + child sessions | Three modes |
|
|
| **Invocation** | Slack, Linear, GitHub | Slack + embedded buttons | Slack + web + Chrome extension | Slack-native |
|
|
| **Validation** | Prompt-driven | 3-layer (local + CI + 1 retry) | Visual DOM verification | Agent councils + auto-merge |
|
|
|
|
---
|
|
|
|
## Features
|
|
|
|
- **Trigger from Linear, Slack, or GitHub** — mention `@openswe` in a comment to kick off a task
|
|
- **Instant acknowledgement** — acknowledges the moment it picks up your message
|
|
- **Message it while it's running** — send follow-up messages mid-task and it'll pick them up before its next step
|
|
- **Run multiple tasks in parallel** — each task runs in its own isolated cloud sandbox
|
|
- **GitHub OAuth built-in** — authenticates with your GitHub account automatically
|
|
- **Opens PRs automatically** — commits changes and opens a draft PR when done, linked back to your ticket
|
|
- **Subagent support** — the agent can spawn child agents for parallel subtasks
|
|
- **Web dashboard** — a companion app (in `ui/`) for GitHub login, per-user model/profile settings, team defaults, enabled-repo and review-style management, user mappings, and an Agents chat UI
|
|
|
|
---
|
|
|
|
## Getting Started
|
|
|
|
- **[Installation Guide](docs/INSTALLATION.md)** — local dev (backend + dashboard), GitHub App creation, LangSmith, Linear/Slack/GitHub triggers, and production deployment
|
|
- **[Customization Guide](docs/CUSTOMIZATION.md)** — swap the sandbox, model, tools, triggers, system prompt, and middleware for your org
|
|
|
|
## Deployment (Sea Haven fork)
|
|
|
|
This fork runs on a **managed deployment**: the backend (all three graphs + the
|
|
FastAPI webapp) runs on
|
|
[LangGraph Cloud / Platform](https://langchain-ai.github.io/langgraph/cloud/),
|
|
and the `ui/` dashboard deploys to [Vercel](https://vercel.com/). Configuration
|
|
and secrets live in the LangGraph deployment config and Vercel environment
|
|
variables. Promotion from `dev` to `prod` (`main`) is handled by
|
|
[`.github/workflows/promote-to-main.yml`](.github/workflows/promote-to-main.yml).
|
|
|
|
See **[INSTALLATION.md § 10 "Production deployment"](docs/INSTALLATION.md#10-production-deployment)**
|
|
for the full backend + dashboard setup.
|
|
|
|
> The earlier self-hosted AWS stack (CDK under `infra/`, an ARM64 EC2 box + nginx
|
|
> behind the shared ALB, and the `cd-infra` / `build-artifacts` release pipelines)
|
|
> was **decommissioned** in favor of the managed deployment above.
|
|
|
|
## License
|
|
|
|
MIT
|