|
Some checks failed
CI / Lint (push) Has been cancelled
CI / Format check (push) Has been cancelled
CI / Typecheck (push) Has been cancelled
CI / Unit tests (push) Has been cancelled
CI / Playwright E2E (push) Has been cancelled
CI / Docker build smoke (push) Has been cancelled
CI / Triage ledger up to date (push) Has been cancelled
CI / ui bun.lock in sync (push) Has been cancelled
* refactor: move docs/resources/assets to domain layout
Part of the domain-reorg adoption (build plan step C1): fork content,
upstream layout. Moves INSTALLATION.md/CUSTOMIZATION.md under docs/,
static/ under assets/, and default_prompt.md under agent/resources/
(packaged via agent/resources/__init__.py), then switches prompt.py's
loader to importlib.resources with an explicit DEFAULT_PROMPT_PATH
override, matching upstream's hunk. README and CUSTOMIZATION.md links
updated for the new paths; wheel build verified to still ship
agent/resources/default_prompt.md.
* refactor: consolidate reviewer modules into agent/review/
Part of the domain-reorg adoption (build plan step C2): fork content,
upstream layout. Nine 1:1 module moves (reviewer_diff/eval_store/
findings/groups/publish/reconcile/trace_context + review_style_
collector/guidance) into agent/review/, with internal relative
imports re-wired to the new package depth. agent/review/__init__.py
mirrors upstream's thin re-export shim (one of the 21 verified "A"
structural adds).
Rewrote the 38 grep hits across importer files (agent/{analyzer,
ci_autofix,reviewer,webapp}.py, agent/dashboard/*, agent/middleware/
settle_review_check.py, agent/tools/*, agent/utils/github_feedback.py,
agent/webhooks/github.py, evals/reviewer/*, and the reviewer test
suite) to point at agent.review.*; 4 of the 38 hits were name
collisions (list_reviewer_findings, reviewer_outcomes,
_reviewer_thread_id, reviewer_thread_id — not the moved modules) and
were left untouched. tests/test_github_checks.py's module-alias
import (`from agent import reviewer_publish`) follows upstream's own
`from agent.review import publish as reviewer_publish` pattern so
downstream `reviewer_publish.*` call sites needed no changes.
agent/reviewer.py and agent/webapp.py stay in place per the hard
rule (fork content, import-only rewire) and are not part of this
package.
Gates: ruff check + ruff format --check, pytest --co -q (1637
collected), full unit suite (1637 passed), and the reviewer/findings
suite in isolation (pytest -k "review or finding", 421 passed).
* refactor: adopt graphs/runtime/providers shims + retarget langgraph.json
Part of the domain-reorg adoption (build plan step C3): fork content,
upstream layout. Adds agent/graphs/{agent,analyzer,chat,reviewer,
scheduler}.py as thin re-export shims delegating to the existing fork
graph factories (agent.server/analyzer/chat/reviewer/scheduler), plus
agent/providers/__init__.py re-exporting agent.utils.model's
make_model/provider_model_kwargs/fallback_model_id_for surface —
verbatim upstream content, verified each import resolves against fork
modules with no name changes needed.
agent/runtime/{constants,execution}.py deviate from upstream's
verbatim shim bodies: rather than duplicating DEFAULT_LLM_MODEL_ID/
DEFAULT_LLM_MAX_TOKENS/DEFAULT_RECURSION_LIMIT/MODEL_CALL_RECURSION_LIMIT
and graph_loaded_for_execution's logic (upstream's shims assume
agent/server.py already had these extracted into runtime/ modules,
which is out of this commit's scope — server.py is untouched), they
import the fork's existing agent.server attributes directly. This
keeps the values/logic single-sourced instead of forking a second
copy that could drift.
agent/runtime/sandbox.py's delegation targets also differ from
upstream: fork's sandbox lifecycle helpers are private
(_get_cached_sandbox_backend, _configure_git_identity,
_recreate_sandbox in agent/server.py) since the fork's sync
4-case `__creating__` sentinel design (AGENTS.md) never made them
public. get_cached_sandbox_backend() also drops upstream's
caller-supplied `reconnect` callback parameter — fork's
_get_cached_sandbox_backend is a plain cache lookup; reconnection is
handled internally by ensure_sandbox_for_thread/
check_or_recreate_sandbox, not via a passed-in callback. No other
signature changes.
Added fork-only agent/graphs/ci_monitor.py (delegates to
agent.ci_monitor:get_ci_monitor) for symmetry, since upstream deleted
its ci-autofix cluster and has no equivalent shim. langgraph.json's
five stock graph entrypoints plus the fork-only ci_monitor now all
point at agent.graphs.<name>; http.app stays agent.webapp:app
(unchanged, per plan).
Deliberately NOT included (owned by build plan step C4, the FastAPI
split, gated on /sh-security-review): agent/api/{__init__,app,
health}.py, agent/webhooks/common.py, and the three
agent/webhooks/{github,linear,slack}_routes.py files. Those aren't
thin structural shims like the 21-file list implies in isolation —
they carry the fork's actual webhook dispatch/verify logic split out
of the still-monolithic webapp.py, which hasn't happened yet.
Building them now against upstream's placeholder content would ship
incomplete auth surface that C4 would just discard and redo.
Pinned oven-sh/setup-bun's bun-version to 1.3.14 (the version
installed locally; ui/ has no .bun-version file or package.json
engines/packageManager field pinning one) across all three CI jobs
that install bun, removing the latest-resolution flake.
Gates: ruff check + ruff format --check (clean), pytest --co -q
(1637 collected, no import errors), a direct import smoke-test of
every new module's public symbols, and a make dev boot check —
langgraph dev registered all six graphs (agent, reviewer, analyzer,
chat, scheduler, ci_monitor) each importing from agent.graphs.*, and
loaded the custom app from agent.webapp:app, before the process was
killed. (The subsequent lifespan failure, "DEFAULT_SANDBOX_SNAPSHOT_ID
must be set when SANDBOX_TYPE=langsmith", is expected with no .env
secrets configured in this environment and unrelated to this commit.)
* refactor: split webapp.py into api/ + per-source webhook routes
Plan step C4 (docs/upstream-sync/domain-reorg/reorg-build-plan.md, approved
decisions 1-2): split the 2,590-line agent/webapp.py monolith into
agent/webhooks/common.py (shared verify/dispatch helpers), agent/api/app.py
(composition), agent/api/health.py (/health + /webhooks/run-complete), and
per-source {github,linear,slack,jira,confluence}_routes.py. Atlassian
Connect lifecycle + descriptor routes (/connect/*) fold into
confluence_routes.py; webapp.py becomes the upstream-shaped compatibility
shim (from .api.app import app). langgraph.json http.app stays
agent.webapp:app via the shim.
Fork content, upstream layout: linear/slack route files verified
content-identical to upstream 8356eb34 and taken verbatim; github_routes is
upstream + the fork's CI auto-fix trigger wiring; jira/confluence routes are
fork-only, transformed to the same common.X / service.X module-attribute
style. All signature verification (GitHub HMAC, Slack, Linear
timestamp-freshness, verify_jira_secret + opt-in HMAC/timestamp/IP
allowlist, Connect JWT/qsh), token-attribution gating, TID-COLLIDE-01 repo
binding, _is_repo_auto_review_enabled gates, and public-repo org gate move
unchanged.
Handlers rewired from webapp.X to common.X; test monkeypatch sites across
26 files + conftest.py + e2e/harness.py retargeted to
webhook_common/handler/route modules per upstream's pattern. Residual
agent.webapp importers: only the shim, langgraph.json http.app, Makefile
uvicorn target, and docs (doc-path updates land in C7).
Gates: ruff check + format, pytest --co, full unit (1637 passed), full
Playwright E2E vs real langgraph dev (9/9), residual-importer sweep.
* refactor: move tests into tests/<domain>/ layout
Applies the plan's C5 step: git mv every test per the domain-reorg
move-map (movemap-m50.txt) into tests/{agent,analyzer,auth,dashboard,
github,middleware,models,reviewer,sandbox,slack,tools,webhooks}/, plus
the 13 fork-only placements from the scoping report §2c (Atlassian
webhook tests -> tests/webhooks/, test_atlassian_connect.py and
test_auth_error_leak.py -> tests/auth/, jira/confluence util tests ->
tests/tools/, test_repo_binding_isolation.py -> tests/sandbox/,
bot-identity/autofix tests -> tests/github/).
Path-only move: the only content edits are parents[1] -> parents[2]
fixes in test_e2b_integration.py and test_daytona_integration.py,
required because their __file__-relative ROOT path gained one more
directory level in the move.
Monkeypatch retargets for these files were already completed in C4;
none remained outstanding here.
* refactor: move ui/src/{components,lib}/agents into features/ layout
Part of the domain-reorg adoption (build plan step C6): fork content,
upstream layout. git mv's all 80 pure-rename files from the move-map
(agent-scoped components/lib -> ui/src/features/{agents,automations,
reviews,settings}/...), including ported/ splitting into
features/agents/experiments/ (still-unintegrated desktop-host ports)
and features/agents/components/chat/ (CloudPromptBar, CodeBlock,
DiffView, Logo, Markdown, ReplyCard, ShellCommand, ToolExecution --
files actually wired into the live dashboard). Fork-diverged files
(PlanReview.tsx, WorkflowApprovalCard.tsx, AgentsSidebar.tsx,
SidebarFilterMenu.tsx, AgentGitPanel.tsx, AutomationEditor.tsx,
DiffView.tsx, CloudPromptBar.tsx) keep fork content -- verified via
diff that only import lines changed.
Rewrote @/components/agents and @/lib/agents imports across all 54
importer files plus the moved files' own internal imports (grep-driven,
85-entry alias map covering every old->new path pair). Two hazards
caught only by the build gate (not tsc, since both files sit under the
new experiments/** tsconfig exclude): ui/src/lib/notifications.ts had
a relative `./agents/types` import (no `@/` alias) that the grep missed;
and features/agents/experiments/index.ts's barrel re-export of the
messages module needed to switch from an `@/` alias to a relative
import (`../components/messages`) after landing inside the newly
tsconfig-excluded experiments/ directory -- vite-tsconfig-paths failed
to resolve it at build time even though tsc stayed silent.
AgentPromptBar.tsx (the 2-line CloudPromptBar re-export shim) moves to
features/agents/components/ with its export path retargeted at
CloudPromptBar's new chat/ location -- a D/A pair, not a content loss,
since the import-path edit drops it below git's rename-similarity
threshold. ported/index.ts is the same D/A story for the same reason.
Swapped the 8-file ported/ exclude lists in tsconfig.json and
eslint.config.js for the single `src/features/agents/experiments/**`
glob (upstream's simplification, no behavior change). Deleted
ui/pnpm-workspace.yaml (lockfile hygiene -- bun stays the toolchain).
Gates: bunx tsc --noEmit (clean), bun run build (clean after the two
notifications.ts / experiments-index.ts fixes above), bun run test
(33/33), and the full Playwright E2E suite against the real langgraph
dev server + built dashboard (9/9).
* docs: land reorg in ledger, fix path refs, ship plan artifacts (C7)
C7 of the domain-reorg adoption (build plan docs/upstream-sync/domain-reorg/
reorg-build-plan.md, step C7):
- CLAUDE.md / AGENTS.md: retarget architecture path references to the new
layout — graph entrypoints via agent.graphs.* shims, the agent/api/ +
agent/webhooks/*_routes.py FastAPI split (webapp.py now a shim),
agent/review/ package, tests/<domain>/ test paths, and the new-graph/
test conventions. README.md already pointed at docs/ (C1) — no change.
- triage.jsonl: flip 8356eb34 (#1726) to landed on branch
refactor/domain-reorg-adoption; add re-triage notes to the 8 unblocked
rows (#1732/#1761/#1744/#1748/#1742 clean, #1736/#1758/#1760 near-clean).
triage.md regenerated via make triage-render.
- Ship the plan, scoping report, move-map artifacts, and the
domain-reorg-adoption workflow so the exercise is reproducible.
Memory + Confluence handled out-of-band (not in this commit): project
memory updated with the new module layout; Confluence check found no IT
page documents the repo module map — no Confluence change required.
|
||
|---|---|---|
| .claude/workflows | ||
| .githooks | ||
| .github | ||
| .security-review | ||
| .vscode | ||
| agent | ||
| assets | ||
| deploy | ||
| docs | ||
| evals/reviewer | ||
| scripts | ||
| tests | ||
| ui | ||
| .codespellignore | ||
| .dockerignore | ||
| .env.example | ||
| .gitignore | ||
| .nvmrc | ||
| AGENTS.md | ||
| CLAUDE.md | ||
| Dockerfile | ||
| langgraph.json | ||
| LICENSE | ||
| Makefile | ||
| package.json | ||
| pyproject.toml | ||
| README.md | ||
| SECURITY.md | ||
| uv.lock | ||
Open-source framework for building your org's internal coding agent.
Elite engineering orgs like Stripe, Ramp, and Coinbase are building their own internal coding agents — Slackbots, CLIs, and web apps that meet engineers where they already work. These agents are connected to internal systems with the right context, permissioning, and safety boundaries to operate with minimal human oversight.
Open SWE is the open-source version of this pattern. Built on LangGraph and Deep Agents, it gives you the same architecture those companies built internally: cloud sandboxes, Slack / Linear / Jira / Confluence / GitHub invocation, subagent orchestration, and automatic PR creation — ready to customize for your own codebase and workflows.
Note
Read the announcement blog post here
Architecture
Open SWE makes the same core architectural decisions as the best internal coding agents. Here's how it maps to the patterns described in this overview of Stripe's Minions, Ramp's Inspect, and Coinbase's Cloudbot:
1. Agent Harness — Composed on Deep Agents
Rather than forking an existing agent or building from scratch, Open SWE composes on the Deep Agents framework — similar to how Ramp built on top of OpenCode. This gives you an upgrade path (pull in upstream improvements) while letting you customize the orchestration, tools, and middleware for your org.
create_deep_agent(
model="openai:gpt-5.5",
system_prompt=construct_system_prompt(...),
tools=[http_request, fetch_url, linear_comment, slack_thread_reply],
backend=sandbox_backend,
middleware=[ToolErrorMiddleware(), check_message_queue_before_model, ...],
)
2. Sandbox — Isolated Cloud Environments
Every task runs in its own isolated cloud sandbox — a remote Linux environment with full shell access. The repo is cloned in, the agent gets full permissions, and the blast radius of any mistake is fully contained. No production access, no confirmation prompts.
Open SWE supports multiple sandbox providers out of the box — Modal, Daytona, Runloop, E2B, and LangSmith — and you can plug in your own. See the Customization Guide for details.
This follows the principle all three companies converge on: isolate first, then give full permissions inside the boundary.
- Each thread gets a persistent sandbox (reused across follow-up messages)
- Sandboxes auto-recreate if they become unreachable
- Multiple tasks run in parallel — each in its own sandbox, no queuing
3. Tools — Curated, Not Accumulated
Stripe's key insight: tool curation matters more than tool quantity. Open SWE follows this principle with a small, focused toolset:
| Tool | Purpose |
|---|---|
execute |
Shell commands in the sandbox |
fetch_url |
Fetch web pages as markdown |
http_request |
API calls (GET, POST, etc.) |
linear_comment |
Post updates to Linear tickets |
jira_* |
Read/comment/create/update Jira issues |
confluence_* |
Read/write Confluence pages + comments |
slack_add_reaction |
React to Slack messages |
slack_thread_reply |
Reply in Slack threads |
GitHub operations are performed with GH_TOKEN=dummy gh inside the sandbox, backed by the LangSmith proxy. Plus the built-in Deep Agents tools: read_file, write_file, edit_file, ls, glob, grep, write_todos, and task (subagent spawning).
Optional observability tools (server-side): Admins can connect Datadog and LangSmith from team settings (Admin → Observability credentials). When connected, the agent gains Datadog tools (via Datadog's hosted MCP server, default toolsets=core) and read-only LangSmith tools (langsmith_get_trace, langsmith_list_runs). These run in the LangGraph server process using credentials encrypted at rest — the sandbox never holds Datadog or LangSmith keys. They are loaded only for runs triggered by an authorized user (admins, plus any emails in OBSERVABILITY_AUTHORIZED_EMAILS), so a prompt-injected run from an untrusted contributor cannot reach team observability data. Use scoped, read-oriented keys regardless: observability data (logs, traces) is attacker-influenced content that can carry prompt injection, and the agent has network egress — the same residual-risk class as web_search / fetch_url.
Optional Corridor guardrails (server-side MCP): Set CORRIDOR_API_TOKEN (or CORRIDOR_MCP_TOKEN / CORRIDOR_TOKEN) to load Corridor's hosted MCP server for each agent run. Open SWE exposes only Corridor's analyzePlan tool. CORRIDOR_MCP_URL defaults to https://app.corridor.dev/api/mcp; if set explicitly, Open SWE only accepts the same HTTPS host and /api/mcp path. Tokens are sent via Authorization: Bearer ... from the LangGraph server process and are never placed in the sandbox. A legacy ?token=... URL is accepted and normalized into the header form.
4. Context Engineering — AGENTS.md + Source Context
Open SWE gathers context from two sources:
AGENTS.md— If the repo contains anAGENTS.mdfile at the root, it's read from the sandbox and injected into the system prompt. This is your repo-level equivalent of Stripe's rule files: encoding conventions, testing requirements, and architectural decisions that every agent run should follow. Reference templates for stack-specific conventions (AWS, SAM, CDK, EC2) live indocs/repo-conventions/.- Source context — The full Linear issue (title, description, comments) or Slack thread history is assembled and passed to the agent, so it starts with rich context rather than discovering everything through tool calls.
5. Orchestration — Subagents + Middleware
Open SWE's orchestration has two layers:
Subagents: The Deep Agents framework natively supports spawning child agents via the task tool. The main agent can fan out independent subtasks to isolated subagents — each with its own middleware stack, todo list, and file operations. This is similar to Ramp's child sessions for parallel work.
Middleware: Deterministic middleware hooks run around the agent loop:
check_message_queue_before_model— Injects follow-up messages (Linear comments or Slack messages that arrive mid-run) before the next model call. You can message the agent while it's working and it'll pick up your input at its next step.notify_step_limit_reached— After-agent hook that posts a Slack reply when the agent hits the model-call limit, so users get a clear signal instead of silence.ToolErrorMiddleware— Catches and handles tool errors gracefully.
6. Invocation — Slack, Linear, Jira, Confluence, and GitHub
All three companies in the article converge on Slack as the primary invocation surface. Open SWE does the same:
- Slack — Mention the bot in any thread. Supports
repo:owner/namesyntax to specify which repo to work on. The agent replies in-thread with status updates and PR links. - Linear — Comment
@opensweon any issue. The agent reacts with 👀 to acknowledge, reads the full issue context, and posts results back as comments. - Jira — Comment
@opensweon any issue (fronted by a Jira Automation rule →/webhooks/jira). The agent reads the issue and posts results back as a comment. - Confluence — Comment
@opensweon a page. A private Atlassian Connect app delivers thecomment_createdevent; the agent acts and replies on the page. - GitHub — Tag
@openswein PR comments on agent-created PRs to have it address review feedback and push fixes to the same branch.
See INSTALLATION.md §5 for per-surface trigger setup.
Each invocation creates a deterministic thread ID, so follow-up messages on the same issue or thread route to the same running agent.
Trigger tags (Sea Haven fork): a mention is a case-insensitive substring match on the comment body — @openswe, @open-swe, @openswe-dev, or @seahaven-openswe (the deployed App slug). GitHub won't linkify @seahaven-openswe (App [bot] accounts aren't user-mentionable), but the text still fires a run.
Engineering conventions & attribution (Sea Haven fork): the main agent's system prompt is tuned to the Sea Haven engineering handbook — branch names are feature|bug|hotfix/<kebab-desc> (optional resolvable <KEY>- prefix), PR bodies use ## Summary / Validation / Tests / Notes, and commit messages follow the handbook format (≤50-char imperative subject, why over what). The PR title rule is repo-aware: when the target repo enforces a conventional-commit title (an amannn/action-semantic-pull-request workflow, a commitlint config, or a documented requirement in AGENTS.md / CONTRIBUTING.md), the agent emits a conforming type(scope): … title that reads the action's allowed types/scopes — this lets it pass gates like this repo's own PR Title Lint and upstream langchain-ai/open-swe without manual retitling; otherwise it falls back to the Sea Haven imperative style with no type: prefix. PRs that resolve a GitHub issue auto-link it in the body (Closes #<n> for full fixes, Refs #<n>/Part of #<n> for partial work, Closes owner/repo#<n> cross-repo); because the Sea Haven flow targets dev rather than the default branch, the issue closes when dev is promoted, not at dev-merge. No agent/AI attribution is added to any artifact — no Co-authored-by bot trailer, no Made by [Open SWE] footer, no "generated by an agent" notes. Commits are currently authored as the triggering user (the upstream behavior, which keeps Vercel preview deploys resolvable); flipping authorship to the bot account is tracked separately in issue #11 pending the Vercel-resolvability decision.
7. Validation — Prompt-Driven
The agent is instructed to run linters, formatters, and tests before committing, and is responsible end-to-end for committing, pushing, opening/updating the draft PR, and replying in the source channel. This is an area where you can extend Open SWE for your org: add deterministic CI checks, visual verification, or review gates as additional middleware. See the Customization Guide for how.
Comparison
| Decision | Open SWE | Stripe (Minions) | Ramp (Inspect) | Coinbase (Cloudbot) |
|---|---|---|---|---|
| Harness | Composed (Deep Agents/LangGraph) | Forked (Goose) | Composed (OpenCode) | Built from scratch |
| Sandbox | Pluggable (Modal, Daytona, Runloop, etc.) | AWS EC2 devboxes (pre-warmed) | Modal containers (pre-warmed) | In-house |
| Tools | ~15, curated | ~500, curated per-agent | OpenCode SDK + extensions | MCPs + custom Skills |
| Context | AGENTS.md + issue/thread | Rule files + pre-hydration | OpenCode built-in | Linear-first + MCPs |
| Orchestration | Subagents + middleware | Blueprints (deterministic + agentic) | Sessions + child sessions | Three modes |
| Invocation | Slack, Linear, GitHub | Slack + embedded buttons | Slack + web + Chrome extension | Slack-native |
| Validation | Prompt-driven | 3-layer (local + CI + 1 retry) | Visual DOM verification | Agent councils + auto-merge |
Features
- Trigger from Linear, Slack, or GitHub — mention
@openswein a comment to kick off a task - Instant acknowledgement — acknowledges the moment it picks up your message
- Message it while it's running — send follow-up messages mid-task and it'll pick them up before its next step
- Run multiple tasks in parallel — each task runs in its own isolated cloud sandbox
- GitHub OAuth built-in — authenticates with your GitHub account automatically
- Opens PRs automatically — commits changes and opens a draft PR when done, linked back to your ticket
- Subagent support — the agent can spawn child agents for parallel subtasks
- Web dashboard — a companion app (in
ui/) for GitHub login, per-user model/profile settings, team defaults, enabled-repo and review-style management, user mappings, and an Agents chat UI
Getting Started
- Installation Guide — local dev (backend + dashboard), GitHub App creation, LangSmith, Linear/Slack/GitHub triggers, and production deployment
- Customization Guide — swap the sandbox, model, tools, triggers, system prompt, and middleware for your org
Deployment (Sea Haven fork)
This fork runs on a managed deployment: the backend (all three graphs + the
FastAPI webapp) runs on
LangGraph Cloud / Platform,
and the ui/ dashboard deploys to Vercel. Configuration
and secrets live in the LangGraph deployment config and Vercel environment
variables. Promotion from dev to prod (main) is handled by
.github/workflows/promote-to-main.yml.
See INSTALLATION.md § 10 "Production deployment" for the full backend + dashboard setup.
The earlier self-hosted AWS stack (CDK under
infra/, an ARM64 EC2 box + nginx behind the shared ALB, and thecd-infra/build-artifactsrelease pipelines) was decommissioned in favor of the managed deployment above.
License
MIT