An Open-Source Asynchronous Coding Agent
Find a file
Johannes du Plessis 32ec9b485f
feat: restructure Open SWE Review tab + wire create_prs (#1319)
* feat(dashboard): restructure Open SWE Review tab + wire create_prs

Restructures the dashboard around two related changes the reviewer settings
have been asking for:

- Wire profile.create_prs. Defaults to true (opt-out); when off the system
  prompt gets a `Pull Request Policy Override` section telling the agent
  to push the branch and notify with the branch URL instead of opening a
  PR. Removes the noop Slack Notifications / Allow Artifacts / First Name
  / Last Name controls and their schema fields.
- Repositories opt-in for Open SWE Review. New per-team enabled list
  stored in the LangGraph Store (`["enabled_review_repos"]`). Every
  reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review`
  which AND-combines the existing env allowlist with the dashboard list.
  Default is empty (opt-in) — admins enable repos per-installation from
  the new Repositories page nested under Open SWE Review.
- Open SWE Review tab now mirrors the Cursor "rules" pattern: main page
  shows installation rows + a Rules entry; both drill into nested pages
  (/review/repositories/$owner and /review/styles) with a back link.
- Adds the new logo/favicon assets shipped from sidebar + html head.

Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults
`is_review_repo_enabled` to True for existing allowlist tests.

* fix(dashboard): make main content scroll independently of the sidebar

Outer flex container was min-h-svh, so it grew with main's content and the
whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden
so the sidebar stays put and only <main> scrolls.

* fix(dashboard): make disabled repo toggles obviously disabled

Switch's disabled state used opacity-50 against a muted background, so
the not-admin state looked nearly identical to the off state. Bump to
opacity-40 + grayscale, and wrap each repo toggle in a span carrying a
native hover tooltip explaining why it's disabled.

* fix(switch): handle base-ui's data-disabled state

base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute)
when disabled, so Tailwind's disabled: variant never matches and the
button keeps its cursor-pointer + clickable look. Mirror the styling
under the data-[disabled] variant and add pointer-events-none so the
disabled state is both visible and actually unclickable.

* feat(dashboard): paginate per-installation repository list

20 repos per page with Prev / page X of Y / Next controls at the bottom.
Pager only renders when there are more than 20 repos. Page resets to 0
when navigating between installations.

* feat(dashboard): global default model selectors for Agent + Reviewer

Adds team-wide default model + reasoning effort for both agents in the
Admin tab so operators can switch models without redeploying.

Resolution chain:
  Agent:    hardcoded -> LLM_MODEL_ID env -> team default -> user profile
  Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable

Team defaults live in team_settings and are validated against the
SUPPORTED_MODELS allowlist + the model's supported reasoning efforts.
'Inherit from env' clears the override and falls back to LLM_MODEL_ID.

* refactor(models): drop LLM_MODEL_ID env in favour of the team default

The team default is now the single source of truth for the runtime model
choice; per-user (agent) and per-call configurable (reviewer) selections
still win on top. When no admin has touched the team default, it surfaces
the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the
admin UI's dropdown is always pre-populated with a sensible value.

The Admin UI loses the 'Inherit from env' option since there is no longer
an env layer to inherit from.

* chore(models): set hardcoded fallback to gpt-5.5 medium

Decouple the team-default boot value (gpt-5.5 / medium) from each model's
ProfileForm-suggested default_effort so we can change one without nudging
the other. The Opus xhigh default for new user profiles is unchanged.

* feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings

- Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new
  description copy that matches the screenshot. Legacy stored values
  fall back to 'every_push' on read so the UI never shows an unknown
  selection.
- Add a 'Coming soon' badge + greyed-out + disabled state on the
  controls that don't have runtime consumers yet: Trigger Mode,
  Autofix Mode, Autofix Severity Threshold, and Automatically fix CI
  failures. SettingsRow grew a comingSoon prop to keep this consistent.
- My Settings drops the noop PR Preferences section and adds a Sign
  Out button. preferred_pr_destination is removed from the profile
  schema; old records get the field popped on next write.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
.github ci: nightly promote of main to prod for LangGraph deploys (#1285) 2026-05-09 01:15:26 +00:00
.vscode Brace/07 16/fixes (#431) 2025-07-16 13:17:35 -07:00
agent feat: restructure Open SWE Review tab + wire create_prs (#1319) 2026-05-21 09:17:07 -07:00
evals/reviewer feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00
scripts feat: move github workflows to gh cli (#1238) 2026-05-04 18:03:53 -07:00
static fix: Add back logo to readme (#1068) 2026-03-17 10:41:41 -07:00
tests feat: restructure Open SWE Review tab + wire create_prs (#1319) 2026-05-21 09:17:07 -07:00
ui feat: restructure Open SWE Review tab + wire create_prs (#1319) 2026-05-21 09:17:07 -07:00
.codespellignore init commit 2025-05-21 14:47:56 -07:00
.dockerignore feat: move github workflows to gh cli (#1238) 2026-05-04 18:03:53 -07:00
.gitignore feat(open-swe): Default to GPT-5.5 medium reasoning (#1224) 2026-04-28 15:03:21 -07:00
AGENTS.md feat: move github workflows to gh cli (#1238) 2026-05-04 18:03:53 -07:00
CLAUDE.md feat: move github workflows to gh cli (#1238) 2026-05-04 18:03:53 -07:00
CUSTOMIZATION.md feat: add idle TTL and delete-after-stop sandbox lifecycle controls (#1265) 2026-05-08 00:30:14 -04:00
default_prompt.md feat: add configurable default prompt file for org-level agent instructions [close OPE-36] (#1187) 2026-04-15 15:18:14 -07:00
Dockerfile feat: move github workflows to gh cli (#1238) 2026-05-04 18:03:53 -07:00
INSTALLATION.md feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00
langgraph.json feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00
LICENSE feat: Monorepo (#22) 2025-05-26 13:02:51 -07:00
Makefile feat: add reviewer graph + eval target wiring (#1241) 2026-05-06 10:15:58 -07:00
pyproject.toml feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00
README.md feat: move github workflows to gh cli (#1238) 2026-05-04 18:03:53 -07:00
REVIEWER_DESIGN.md feat: implement reviewer findings, publish_review, and watch mode (#1253) 2026-05-07 14:48:43 -07:00
SECURITY.md fix: Security stuff (#441) 2025-07-17 12:19:20 -07:00
uv.lock feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00

Open-source framework for building your org's internal coding agent.

License GitHub Stars Built on LangGraph Built on Deep Agents Twitter / X

Elite engineering orgs like Stripe, Ramp, and Coinbase are building their own internal coding agents — Slackbots, CLIs, and web apps that meet engineers where they already work. These agents are connected to internal systems with the right context, permissioning, and safety boundaries to operate with minimal human oversight.

Open SWE is the open-source version of this pattern. Built on LangGraph and Deep Agents, it gives you the same architecture those companies built internally: cloud sandboxes, Slack and Linear invocation, subagent orchestration, and automatic PR creation — ready to customize for your own codebase and workflows.

Note

💬 Read the announcement blog post here


Architecture

Open SWE makes the same core architectural decisions as the best internal coding agents. Here's how it maps to the patterns described in this overview of Stripe's Minions, Ramp's Inspect, and Coinbase's Cloudbot:

1. Agent Harness — Composed on Deep Agents

Rather than forking an existing agent or building from scratch, Open SWE composes on the Deep Agents framework — similar to how Ramp built on top of OpenCode. This gives you an upgrade path (pull in upstream improvements) while letting you customize the orchestration, tools, and middleware for your org.

create_deep_agent(
    model="openai:gpt-5.5",
    system_prompt=construct_system_prompt(...),
    tools=[http_request, fetch_url, linear_comment, slack_thread_reply],
    backend=sandbox_backend,
    middleware=[ToolErrorMiddleware(), check_message_queue_before_model, ...],
)

2. Sandbox — Isolated Cloud Environments

Every task runs in its own isolated cloud sandbox — a remote Linux environment with full shell access. The repo is cloned in, the agent gets full permissions, and the blast radius of any mistake is fully contained. No production access, no confirmation prompts.

Open SWE supports multiple sandbox providers out of the box — Modal, Daytona, Runloop, and LangSmith — and you can plug in your own. See the Customization Guide for details.

This follows the principle all three companies converge on: isolate first, then give full permissions inside the boundary.

  • Each thread gets a persistent sandbox (reused across follow-up messages)
  • Sandboxes auto-recreate if they become unreachable
  • Multiple tasks run in parallel — each in its own sandbox, no queuing

3. Tools — Curated, Not Accumulated

Stripe's key insight: tool curation matters more than tool quantity. Open SWE follows this principle with a small, focused toolset:

Tool Purpose
execute Shell commands in the sandbox
fetch_url Fetch web pages as markdown
http_request API calls (GET, POST, etc.)
linear_comment Post updates to Linear tickets
slack_thread_reply Reply in Slack threads

GitHub operations are performed with GH_TOKEN=dummy gh inside the sandbox, backed by the LangSmith proxy. Plus the built-in Deep Agents tools: read_file, write_file, edit_file, ls, glob, grep, write_todos, and task (subagent spawning).

4. Context Engineering — AGENTS.md + Source Context

Open SWE gathers context from two sources:

  • AGENTS.md — If the repo contains an AGENTS.md file at the root, it's read from the sandbox and injected into the system prompt. This is your repo-level equivalent of Stripe's rule files: encoding conventions, testing requirements, and architectural decisions that every agent run should follow.
  • Source context — The full Linear issue (title, description, comments) or Slack thread history is assembled and passed to the agent, so it starts with rich context rather than discovering everything through tool calls.

5. Orchestration — Subagents + Middleware

Open SWE's orchestration has two layers:

Subagents: The Deep Agents framework natively supports spawning child agents via the task tool. The main agent can fan out independent subtasks to isolated subagents — each with its own middleware stack, todo list, and file operations. This is similar to Ramp's child sessions for parallel work.

Middleware: Deterministic middleware hooks run around the agent loop:

  • check_message_queue_before_model — Injects follow-up messages (Linear comments or Slack messages that arrive mid-run) before the next model call. You can message the agent while it's working and it'll pick up your input at its next step.
  • notify_step_limit_reached — After-agent hook that posts a Slack reply when the agent hits the model-call limit, so users get a clear signal instead of silence.
  • ToolErrorMiddleware — Catches and handles tool errors gracefully.

6. Invocation — Slack, Linear, and GitHub

All three companies in the article converge on Slack as the primary invocation surface. Open SWE does the same:

  • Slack — Mention the bot in any thread. Supports repo:owner/name syntax to specify which repo to work on. The agent replies in-thread with status updates and PR links.
  • Linear — Comment @openswe on any issue. The agent reads the full issue context, reacts with 👀 to acknowledge, and posts results back as comments.
  • GitHub — Tag @openswe in PR comments on agent-created PRs to have it address review feedback and push fixes to the same branch.

Each invocation creates a deterministic thread ID, so follow-up messages on the same issue or thread route to the same running agent.

7. Validation — Prompt-Driven

The agent is instructed to run linters, formatters, and tests before committing, and is responsible end-to-end for committing, pushing, opening/updating the draft PR, and replying in the source channel. This is an area where you can extend Open SWE for your org: add deterministic CI checks, visual verification, or review gates as additional middleware. See the Customization Guide for how.


Comparison

Decision Open SWE Stripe (Minions) Ramp (Inspect) Coinbase (Cloudbot)
Harness Composed (Deep Agents/LangGraph) Forked (Goose) Composed (OpenCode) Built from scratch
Sandbox Pluggable (Modal, Daytona, Runloop, etc.) AWS EC2 devboxes (pre-warmed) Modal containers (pre-warmed) In-house
Tools ~15, curated ~500, curated per-agent OpenCode SDK + extensions MCPs + custom Skills
Context AGENTS.md + issue/thread Rule files + pre-hydration OpenCode built-in Linear-first + MCPs
Orchestration Subagents + middleware Blueprints (deterministic + agentic) Sessions + child sessions Three modes
Invocation Slack, Linear, GitHub Slack + embedded buttons Slack + web + Chrome extension Slack-native
Validation Prompt-driven 3-layer (local + CI + 1 retry) Visual DOM verification Agent councils + auto-merge

Features

  • Trigger from Linear, Slack, or GitHub — mention @openswe in a comment to kick off a task
  • Instant acknowledgement — reacts with 👀 the moment it picks up your message
  • Message it while it's running — send follow-up messages mid-task and it'll pick them up before its next step
  • Run multiple tasks in parallel — each task runs in its own isolated cloud sandbox
  • GitHub OAuth built-in — authenticates with your GitHub account automatically
  • Opens PRs automatically — commits changes and opens a draft PR when done, linked back to your ticket
  • Subagent support — the agent can spawn child agents for parallel subtasks

Getting Started

  • Installation Guide — GitHub App creation, LangSmith, Linear/Slack/GitHub triggers, and production deployment
  • Customization Guide — swap the sandbox, model, tools, triggers, system prompt, and middleware for your org

License

MIT