* fix: enforce a replay window on Linear webhooks (AUTHZ-001) verify_linear_signature accepted any correctly-signed body with no freshness check, so a captured request could be replayed indefinitely. Parse the signed webhookTimestamp (Unix ms) and reject requests outside a 60s window, failing closed when the field is missing or malformed — mirroring the Slack verifier. * fix: stop leaking upstream auth-error bodies into user comments get_github_token_for_user folded the raw upstream response text into the error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the full body server-side only and return a generic "GitHub auth failed (status <code>)". Also document the accepted shared-installation-token blast radius on the bot-token-only path (AUTHZ-003). * fix: bind sandbox and token caches to repo to prevent thread-id collision A PR head-branch name is attacker-controllable and get_thread_id_from_branch derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01). The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on thread_id alone, and a cached sandbox was reused after only an echo-ping, so a different repo's webhook could bind to another thread's sandbox or token. Without changing the persistent thread-id scheme: - Persist the bound repo (owner/name) in thread metadata on sandbox creation and refuse to reuse a sandbox whose bound repo does not match the current event (SandboxRepoMismatchError); the in-memory proxy also carries the binding. - Bind the GitHub-token cache entries to their repo and evict on a cross-repo read so a colliding thread_id cannot be served another repo's token. - Thread repo through the reviewer and the webhook token resolvers. * fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04) The instance role and the GitHub deploy app role granted s3:ListBucket on the whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts) only ever lists under releases/, so add a StringLike s3:prefix=releases/* condition. GetBucketLocation has no s3:prefix in its request context, so it moves to its own unconditioned statement. Also document the accepted F-2 cross-env existence-oracle residual on BatchGetSecretValue. * chore: suppress test-fixture credential false positive; document AUTHZ-002 Add a machine-level suppression for the fake Datadog key in the test_team_credentials encryption-roundtrip fixture (CWE-798, not a real credential). Clarify that the within-org thread-write path is intentional by design (AUTHZ-002) — comment only, no behavior change. * fix: casefold repo-binding keys to avoid spurious cross-repo mismatch GitHub owner/name are case-insensitive. Casefold the owner/name key on both the write (binding) and read (compare) sides — repo_cache_key and the metadata bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2). * fix: stop leaking upstream auth body in unexpected-result branch The 2xx-but-missing-token/url branch echoed the parsed upstream response body into the user-facing error. Return a generic message and log response_data server-side only, mirroring the existing HTTPStatusError fix (Gap 4). * fix: fail closed for unbound-legacy sandboxes and catch repo mismatch Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no recorded bound_repo (a pre-binding legacy thread, post-deploy) previously reconnected-and-served the sandbox to the current repo, then rebound it. Now fail closed: drop the stale id and recreate a fresh sandbox bound to this repo, logging a reconnect-with-missing-binding event. A sandbox is never served to a repo unless its binding is known and matches; new threads bind on first run unchanged. Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints, log it for alarming, and surface a clean sanitized error instead of letting an opaque deep-stack exception crash-loop the worker. * chore: suppress test-fixture credential false positive in token-TTL tests Add a machine-level suppression for the fake "ghp_secret" GitHub token used by the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and not a valid PAT; scoped to the unit test only. |
||
|---|---|---|
| .github | ||
| .security-review | ||
| .vscode | ||
| agent | ||
| deploy | ||
| evals/reviewer | ||
| infra | ||
| scripts | ||
| static | ||
| tests | ||
| ui | ||
| .codespellignore | ||
| .dockerignore | ||
| .gitignore | ||
| AGENTS.md | ||
| CLAUDE.md | ||
| CUSTOMIZATION.md | ||
| default_prompt.md | ||
| Dockerfile | ||
| INSTALLATION.md | ||
| langgraph.json | ||
| LICENSE | ||
| Makefile | ||
| package.json | ||
| pyproject.toml | ||
| README.md | ||
| SECURITY.md | ||
| uv.lock | ||
Open-source framework for building your org's internal coding agent.
Elite engineering orgs like Stripe, Ramp, and Coinbase are building their own internal coding agents — Slackbots, CLIs, and web apps that meet engineers where they already work. These agents are connected to internal systems with the right context, permissioning, and safety boundaries to operate with minimal human oversight.
Open SWE is the open-source version of this pattern. Built on LangGraph and Deep Agents, it gives you the same architecture those companies built internally: cloud sandboxes, Slack and Linear invocation, subagent orchestration, and automatic PR creation — ready to customize for your own codebase and workflows.
Note
💬 Read the announcement blog post here
Architecture
Open SWE makes the same core architectural decisions as the best internal coding agents. Here's how it maps to the patterns described in this overview of Stripe's Minions, Ramp's Inspect, and Coinbase's Cloudbot:
1. Agent Harness — Composed on Deep Agents
Rather than forking an existing agent or building from scratch, Open SWE composes on the Deep Agents framework — similar to how Ramp built on top of OpenCode. This gives you an upgrade path (pull in upstream improvements) while letting you customize the orchestration, tools, and middleware for your org.
create_deep_agent(
model="openai:gpt-5.5",
system_prompt=construct_system_prompt(...),
tools=[http_request, fetch_url, linear_comment, slack_thread_reply],
backend=sandbox_backend,
middleware=[ToolErrorMiddleware(), check_message_queue_before_model, ...],
)
2. Sandbox — Isolated Cloud Environments
Every task runs in its own isolated cloud sandbox — a remote Linux environment with full shell access. The repo is cloned in, the agent gets full permissions, and the blast radius of any mistake is fully contained. No production access, no confirmation prompts.
Open SWE supports multiple sandbox providers out of the box — Modal, Daytona, Runloop, and LangSmith — and you can plug in your own. See the Customization Guide for details.
This follows the principle all three companies converge on: isolate first, then give full permissions inside the boundary.
- Each thread gets a persistent sandbox (reused across follow-up messages)
- Sandboxes auto-recreate if they become unreachable
- Multiple tasks run in parallel — each in its own sandbox, no queuing
3. Tools — Curated, Not Accumulated
Stripe's key insight: tool curation matters more than tool quantity. Open SWE follows this principle with a small, focused toolset:
| Tool | Purpose |
|---|---|
execute |
Shell commands in the sandbox |
fetch_url |
Fetch web pages as markdown |
http_request |
API calls (GET, POST, etc.) |
linear_comment |
Post updates to Linear tickets |
slack_thread_reply |
Reply in Slack threads |
GitHub operations are performed with GH_TOKEN=dummy gh inside the sandbox, backed by the LangSmith proxy. Plus the built-in Deep Agents tools: read_file, write_file, edit_file, ls, glob, grep, write_todos, and task (subagent spawning).
Optional observability tools (server-side): Admins can connect Datadog and LangSmith from team settings (Admin → Observability credentials). When connected, the agent gains Datadog tools (via Datadog's hosted MCP server, default toolsets=core) and read-only LangSmith tools (langsmith_get_trace, langsmith_list_runs). These run in the LangGraph server process using credentials encrypted at rest — the sandbox never holds Datadog or LangSmith keys. They are loaded only for runs triggered by an authorized user (admins, plus any emails in OBSERVABILITY_AUTHORIZED_EMAILS), so a prompt-injected run from an untrusted contributor cannot reach team observability data. Use scoped, read-oriented keys regardless: observability data (logs, traces) is attacker-influenced content that can carry prompt injection, and the agent has network egress — the same residual-risk class as web_search / fetch_url.
Optional Corridor guardrails (server-side MCP): Set CORRIDOR_API_TOKEN (or CORRIDOR_MCP_TOKEN / CORRIDOR_TOKEN) to load Corridor's hosted MCP server for each agent run. Open SWE exposes only Corridor's analyzePlan tool. CORRIDOR_MCP_URL defaults to https://app.corridor.dev/api/mcp; if set explicitly, Open SWE only accepts the same HTTPS host and /api/mcp path. Tokens are sent via Authorization: Bearer ... from the LangGraph server process and are never placed in the sandbox. A legacy ?token=... URL is accepted and normalized into the header form.
4. Context Engineering — AGENTS.md + Source Context
Open SWE gathers context from two sources:
AGENTS.md— If the repo contains anAGENTS.mdfile at the root, it's read from the sandbox and injected into the system prompt. This is your repo-level equivalent of Stripe's rule files: encoding conventions, testing requirements, and architectural decisions that every agent run should follow.- Source context — The full Linear issue (title, description, comments) or Slack thread history is assembled and passed to the agent, so it starts with rich context rather than discovering everything through tool calls.
5. Orchestration — Subagents + Middleware
Open SWE's orchestration has two layers:
Subagents: The Deep Agents framework natively supports spawning child agents via the task tool. The main agent can fan out independent subtasks to isolated subagents — each with its own middleware stack, todo list, and file operations. This is similar to Ramp's child sessions for parallel work.
Middleware: Deterministic middleware hooks run around the agent loop:
check_message_queue_before_model— Injects follow-up messages (Linear comments or Slack messages that arrive mid-run) before the next model call. You can message the agent while it's working and it'll pick up your input at its next step.notify_step_limit_reached— After-agent hook that posts a Slack reply when the agent hits the model-call limit, so users get a clear signal instead of silence.ToolErrorMiddleware— Catches and handles tool errors gracefully.
6. Invocation — Slack, Linear, and GitHub
All three companies in the article converge on Slack as the primary invocation surface. Open SWE does the same:
- Slack — Mention the bot in any thread. Supports
repo:owner/namesyntax to specify which repo to work on. The agent replies in-thread with status updates and PR links. - Linear — Comment
@opensweon any issue. The agent reads the full issue context, reacts with 👀 to acknowledge, and posts results back as comments. - GitHub — Tag
@openswein PR comments on agent-created PRs to have it address review feedback and push fixes to the same branch.
Each invocation creates a deterministic thread ID, so follow-up messages on the same issue or thread route to the same running agent.
Trigger tags (Sea Haven fork): a mention is a case-insensitive substring match on the comment body — @openswe, @open-swe, @openswe-dev, or @seahaven-openswe (the deployed App slug). GitHub won't linkify @seahaven-openswe (App [bot] accounts aren't user-mentionable), but the text still fires a run.
Engineering conventions & attribution (Sea Haven fork): the main agent's system prompt is tuned to the Sea Haven engineering handbook — branch names are feature|bug|hotfix/<kebab-desc> (optional resolvable <KEY>- prefix), PR bodies use ## Summary / Validation / Tests / Notes, and commit messages follow the handbook format (≤50-char imperative subject, why over what). The PR title rule is repo-aware: when the target repo enforces a conventional-commit title (an amannn/action-semantic-pull-request workflow, a commitlint config, or a documented requirement in AGENTS.md / CONTRIBUTING.md), the agent emits a conforming type(scope): … title that reads the action's allowed types/scopes — this lets it pass gates like this repo's own PR Title Lint and upstream langchain-ai/open-swe without manual retitling; otherwise it falls back to the Sea Haven imperative style with no type: prefix. PRs that resolve a GitHub issue auto-link it in the body (Closes #<n> for full fixes, Refs #<n>/Part of #<n> for partial work, Closes owner/repo#<n> cross-repo); because the Sea Haven flow targets dev rather than the default branch, the issue closes when dev is promoted, not at dev-merge. No agent/AI attribution is added to any artifact — no Co-authored-by bot trailer, no Made by [Open SWE] footer, no "generated by an agent" notes. Commits are currently authored as the triggering user (the upstream behavior, which keeps Vercel preview deploys resolvable); flipping authorship to the bot account is tracked separately in issue #11 pending the Vercel-resolvability decision.
7. Validation — Prompt-Driven
The agent is instructed to run linters, formatters, and tests before committing, and is responsible end-to-end for committing, pushing, opening/updating the draft PR, and replying in the source channel. This is an area where you can extend Open SWE for your org: add deterministic CI checks, visual verification, or review gates as additional middleware. See the Customization Guide for how.
Comparison
| Decision | Open SWE | Stripe (Minions) | Ramp (Inspect) | Coinbase (Cloudbot) |
|---|---|---|---|---|
| Harness | Composed (Deep Agents/LangGraph) | Forked (Goose) | Composed (OpenCode) | Built from scratch |
| Sandbox | Pluggable (Modal, Daytona, Runloop, etc.) | AWS EC2 devboxes (pre-warmed) | Modal containers (pre-warmed) | In-house |
| Tools | ~15, curated | ~500, curated per-agent | OpenCode SDK + extensions | MCPs + custom Skills |
| Context | AGENTS.md + issue/thread | Rule files + pre-hydration | OpenCode built-in | Linear-first + MCPs |
| Orchestration | Subagents + middleware | Blueprints (deterministic + agentic) | Sessions + child sessions | Three modes |
| Invocation | Slack, Linear, GitHub | Slack + embedded buttons | Slack + web + Chrome extension | Slack-native |
| Validation | Prompt-driven | 3-layer (local + CI + 1 retry) | Visual DOM verification | Agent councils + auto-merge |
Features
- Trigger from Linear, Slack, or GitHub — mention
@openswein a comment to kick off a task - Instant acknowledgement — reacts with 👀 the moment it picks up your message
- Message it while it's running — send follow-up messages mid-task and it'll pick them up before its next step
- Run multiple tasks in parallel — each task runs in its own isolated cloud sandbox
- GitHub OAuth built-in — authenticates with your GitHub account automatically
- Opens PRs automatically — commits changes and opens a draft PR when done, linked back to your ticket
- Subagent support — the agent can spawn child agents for parallel subtasks
- Web dashboard — a companion app (in
ui/) for GitHub login, per-user model/profile settings, team defaults, enabled-repo and review-style management, user mappings, and an Agents chat UI
Getting Started
- Installation Guide — local dev (backend + dashboard), GitHub App creation, LangSmith, Linear/Slack/GitHub triggers, and production deployment
- Customization Guide — swap the sandbox, model, tools, triggers, system prompt, and middleware for your org
Deployment (Sea Haven fork)
This fork is self-hosted on AWS and live in production. Each env (dev /
prod) runs the stock langgraph dev server (all three graphs + the FastAPI
webapp) bound to loopback 127.0.0.1:2024 on a single ARM64 EC2 box, fronted by
nginx (the sole ingress) behind the shared seahaven-com ALB. CDK (infra/)
owns the per-env stacks; GitHub Actions handle CDK deploys (cd-infra.yml) and
app-artifact releases to S3 rolled onto the box via an SSM document
(build-artifacts.yml), with dev auto-deploying and prod gated behind a
manual GitHub Environment approval.
deploy/seahaven/DEPLOYMENT.md is the canonical
deploy runbook — full end-to-end pipeline, config seeding, promotion/rollback,
and live prod facts. CDK specifics live in infra/README.md.
License
MIT