open-swe/CLAUDE.md
2026-08-04 11:33:43 -04:00

22 KiB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. Sea Haven org governance (Jira routing, PR conventions, security gates, CI/SHA pin standards) is defined in the root AGENTS.md; this file covers repo-specific guidance only.

Project

Open SWE is an open-source coding-agent framework built on LangGraph + Deep Agents (deepagents.create_deep_agent). It runs as a LangGraph app: each thread spawns its own isolated cloud sandbox, and the agent is invoked from Slack, Linear, Jira, Confluence, or GitHub (PR comments, plus auto-review on opened / ready-for-review).

A separate reviewer graph runs read-only code reviews on PRs, and a review-style analyzer graph learns per-repo review style from historical PRs.

Commands

Dependencies are managed with uv. Tests use pytest (asyncio_mode = "auto"). Lint/format is ruff (line-length 100, target py311). requires-python = ">=3.11"; langgraph.json pins the runtime to 3.12.

make install            # uv pip install -e .
make dev                # uv run langgraph dev — serves all six graphs + the FastAPI app from langgraph.json
make run                # uvicorn agent.webapp:app --reload --port 8000 (FastAPI only, no LangGraph runtime)
make test               # uv run pytest -vvv tests/
make test TEST_FILE=tests/github/test_open_pull_request.py    # single test file
uv run pytest -vvv tests/github/test_open_pull_request.py::test_name  # single test
make lint               # ruff check + ruff format --diff
make format             # ruff format + ruff check --fix

langgraph.json declares six graph entrypoints and the FastAPI app, all served together by langgraph dev. Since the domain reorg, every graph entrypoint targets a thin agent.graphs.* re-export shim (agent/graphs/<name>.py) rather than the factory module directly; the shims delegate to the unmoved factories (agent/server.py, agent/reviewer.py, agent/analyzer.py, agent/chat.py, agent/scheduler.py, agent/ci_monitor.py):

Graph Entrypoint Purpose
agent agent.graphs.agent:traced_agent (shim → agent.server:get_agent) Main coding agent (Slack/Linear/Jira/Confluence/GitHub-triggered).
reviewer agent.graphs.reviewer:traced_reviewer_agent (shim → agent.reviewer:get_reviewer_agent) Read-only PR reviewer. Findings model + publish_review.
analyzer agent.graphs.analyzer:traced_analyzer (shim → agent.analyzer:get_analyzer) Learns per-repo reviewer style from historical PRs and this reviewer's own finding outcomes.
chat agent.graphs.chat:traced_chat_agent Dashboard Agents chat graph.
scheduler agent.graphs.scheduler:get_scheduler Reconcile sweep for stragglers.
ci_monitor agent.graphs.ci_monitor:get_ci_monitor Fork-only polling fallback for the CI auto-fix flow.

The FastAPI app is agent.webapp:app — now a compatibility shim that re-exports agent.api.app:app (see the FastAPI split under "Entrypoints").

Architecture

Entrypoints

  • agent/server.py → get_agent(config) — main graph factory. Called per-thread. Resolves the GitHub token, gets-or-creates the sandbox for the thread, resolves the team/profile/per-thread model + effort, then constructs a fresh create_deep_agent(...) with the curated tool list and middleware stack. The agent itself is stateless — all per-thread state lives in the sandbox + thread metadata.
  • agent/reviewer.py → get_reviewer_agent(config) — reviewer graph factory. Shares ensure_sandbox_for_thread with the main agent but wires a reviewer-only toolset (add_finding, update_finding, list_findings, publish_review, web_search, fetch_url, http_request) and a different system prompt that pins the single-evolving-findings model and the diff-anchored bar for filing a finding. Read-only: no commit/push/PR-opening tools. Its supporting modules live in the agent/review/ package (domain reorg): findings.py, publish.py, reconcile.py, trace_context.py, diff.py, groups.py, eval_store.py, style_collector.py, style_guidance.py.
  • agent/analyzer.py → get_analyzer(config) — small graph that emits a per-repo style prompt via the save_review_style_prompt tool, consumed by the reviewer as a "repository-specific review style" appendix. It runs in one of two modes (analyzer_mode in configurable): bootstrap (cold-start: crawl historical PR reviews) and continual (nightly: refine using this reviewer's own finding outcomes via read_finding_outcomes). Each mode's procedure lives in a deepagents skill (agent/skills/bootstrap-repo-analysis/, agent/skills/continual-learning/) served as virtual files via a CompositeBackend /skills/ route + StateBackend (seeded into the run's files channel by the launcher — never written to the sandbox). Launchers and the per-repo nightly cron live in agent/dashboard/review_style_jobs.py and agent/dashboard/analyzer_cron.py; the cron is registered when bootstrap completes.
  • FastAPI layer (agent/api/ + agent/webhooks/) — the domain reorg split the fork's former 2,590-line agent/webapp.py monolith into a per-source layout; agent/webapp.py is now a 4-line compatibility shim (from .api.app import app). The pieces:
    • agent/api/app.py composes the FastAPI app and mounts every router; agent/api/health.py owns /health and /webhooks/run-complete.
    • agent/webhooks/common.py holds the shared verify/dispatch helpers and constants (signature verification, thread-id derivation, the dispatch entry). Route modules and handlers reach these via module-attribute access (common.X) so the ~240 monkeypatch.setattr test sites retarget cleanly.
    • Per-source route modules agent/webhooks/{github,linear,slack,jira,confluence}_routes.py define the HTTP routes (GitHub, Linear, Slack, Jira, and the Confluence Atlassian Connect /connect/* + /webhooks/jira routes; the fork-only jira_routes.py/confluence_routes.py mirror upstream's github_routes.py pattern, and the Connect lifecycle/descriptor routes fold into confluence_routes.py). The handler modules agent/webhooks/{github,slack,linear,jira,confluence}.py carry the per-source logic.
    • Each webhook resolves a deterministic thread_id (so follow-up messages route to the same agent run) and triggers a run through the single durable dispatch contract in agent/dispatch.py (dispatch_agent_run: multitask_strategy="interrupt" + durability="sync" + completion webhook); agent/completion.py posts a failure reply if a run dies, and agent/reconcile.py (a scheduler-graph sweep) catches stragglers. The GitHub handler also auto-reviews PRs on opened / ready_for_review and drives the CI auto-fix flow (agent/ci_autofix.py).
    • Atlassian triggers. The Jira trigger is a Jira Automation rule POSTing to /webhooks/jira with a shared-secret header (verify_jira_secret; optional HMAC-body+timestamp via JIRA_WEBHOOK_REQUIRE_SIGNATURE), since Jira Cloud has no native webhook signing. The Confluence trigger is a private Atlassian Connect app (agent/utils/atlassian_connect.py): the comment_created webhook is HS256-JWT-verified against the per-tenant stored sharedSecret with a hand-rolled qsh (query-string-hash) check; signed-install is on, so install/uninstall lifecycle callbacks are RS256-verified against Atlassian's published keys (no trust-on-first-use). Install secrets are stored encrypted in the LangGraph store, keyed by clientKey. Both Atlassian webhook bodies are treated as pointers only — the triggering comment's real author/text is re-fetched server-side via the Basic-auth service account before anything security-relevant is derived, and attribution is gated on an active user mapping (mirrors the GitHub-login / token-attribution flow). Descriptor served at GET /connect/atlassian-connect.json.
  • agent/dashboard/ — router mounted under the FastAPI app at startup (app.include_router(dashboard_router)). Owns GitHub OAuth, per-user profiles, admin endpoints, team defaults, enabled-repo lists, review-style management, and the Agents chat thread API used by the UI in ui/.

Sandbox lifecycle (the tricky part)

SANDBOX_BACKENDS (in agent/utils/sandbox_state.py) is an in-process dict keyed by thread_id. Thread metadata persists sandbox_id across processes. ensure_sandbox_for_thread handles four cases:

  1. Sandbox cached in memory → ping it (echo ok); recreate on SandboxClientError. Healthy reused sandboxes also get a GitHub-proxy refresh (recreate on failure).
  2. Metadata says __creating__ and no cache → poll until ready (_wait_for_sandbox_id).
  3. No sandbox at all → set __creating__ sentinel, create one, persist the real id.
  4. Metadata has an id but no cache → reconnect; fall back to recreate on failure.

For SANDBOX_TYPE=langsmith (default), every sandbox creation/refresh also calls _configure_github_proxy with a fresh GitHub App installation token (get_github_app_installation_token). The proxy injects Basic auth for github.com git traffic and Bearer auth for api.github.com so sandbox commands can use GH_TOKEN=dummy gh ... without storing real tokens in the sandbox. Other providers (modal, daytona, runloop, local) skip the proxy step. Provider is selected via SANDBOX_TYPE; factory is agent/utils/sandbox.py:create_sandbox (SANDBOX_FACTORIES maps each provider name to a creator in agent/integrations/).

Every run re-applies git config --global user.name/email for the bot identity, because reused/reconnected sandboxes can lose --global config and Vercel preview deploys reject commits whose author email doesn't resolve to a GitHub account.

Middleware stack (order matters)

Configured in agent/server.py:get_agent, runs around every model call (in this order):

  1. SanitizeToolInputsMiddleware — strips/normalizes tool inputs before they reach tools.
  2. ModelCallLimitMiddleware (from langchain.agents.middleware) — caps model calls at MODEL_CALL_RECURSION_LIMIT (~half of DEFAULT_RECURSION_LIMIT); exit_behavior="end".
  3. ToolErrorMiddleware — catches tool exceptions and surfaces them as tool messages.
  4. check_message_queue_before_model — pulls Linear comments / Slack messages that arrived mid-run from the thread queue and injects them as user messages before the next LLM call. This is what makes "message the agent while it's working" work.
  5. SlackAssistantStatusMiddleware — keeps the Slack "assistant is typing"-style status up to date around model calls.
  6. ensure_no_empty_msg — after-model hook; when the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion) it re-injects a synthetic no_op / confirming_completion tool call so the run continues instead of ending prematurely.
  7. notify_step_limit_reached — after-agent hook that posts a Slack reply when the agent hits the step limit, so the user gets a clear signal instead of silence.
  8. SandboxCircuitBreakerMiddleware — trips the agent out of repeated sandbox failures instead of looping.
  9. ModelFallbackMiddleware (optional, last) — added only when LLM_FALLBACK_MODEL_ID or the per-model default fallback differs from the primary model.

The system prompt instructs the agent to call a tool every turn, and ensure_no_empty_msg re-injects a tool call when it doesn't — together these keep runs from stopping partway through a task.

Other middleware exists in agent/middleware/ (ExcludeToolsMiddleware) but isn't wired into the default agent. The reviewer uses a leaner stack (see reviewer.py:get_reviewer_agent for the authoritative order), including SanitizeToolInputsMiddleware, ModelCallLimitMiddleware, ToolErrorMiddleware, PullRequestVerdictGuardMiddleware, SlackAssistantStatusMiddleware, and settle_review_check_on_exit.

Verdict gating (Sea Haven fork): Verdict authority is set by dispatch, not inferred by the reviewer model. verdict_requested=True is reserved for explicit human paths, including a direct @openswe review command and the main agent's request_pr_review tool when the user explicitly asks for a verdict. verdict_authorized=True is reserved for automatic paths after deterministic policy and trust checks. Automatic verdicts default off at both the team and per-repository levels; either opt-in enables policy evaluation, but only a non-fork PR whose author is an internal bot or an active member of the repository organization receives automatic authority. Automatic verdicts must match authoritative finding state: approve requires zero blocking findings and request_changes requires one or more. Explicit human verdict requests are not subject to that automatic consistency rule, but both paths still re-check the current PR head, GitHub's recorded review state, and self-review constraints. A self-review is downgraded to a comment review; its Open SWE Review check remains finding-aware and can still fail when blocking findings exist. PullRequestVerdictGuardMiddleware is wired into both the coding and reviewer graphs, and all agent shell verdict paths are blocked: gh pr review verdict flags plus direct gh api/curl APPROVE or REQUEST_CHANGES submissions. Comment reviews and reads remain allowed. Verdicts must go through publish_review(verdict=...), which also best-effort dismisses a recorded stale approval when a later review surfaces new findings.

There is intentionally no after-agent safety net that opens a PR for the agent. The agent itself is responsible for committing, pushing, opening/updating the draft PR, and replying in the source channel — all via GH_TOKEN=dummy gh and slack_thread_reply / linear_comment.

Tools

All tools live in agent/tools/ and are flat-imported via agent/tools/__init__.py. The set is intentionally small and curated — see README "Tools — Curated, Not Accumulated".

Wired into get_agent: http_request, fetch_url, web_search, linear_comment, linear_create_issue, linear_delete_issue, linear_get_issue, linear_get_issue_comments, linear_list_teams, linear_search_issues, linear_update_issue, jira_comment, jira_create_issue, jira_get_issue, jira_get_issue_comments, jira_list_projects, jira_update_issue, confluence_get_page, confluence_create_page, confluence_update_page, confluence_comment, confluence_search, request_pr_review, schedule_thread_wakeup, slack_add_reaction, slack_read_thread_messages, slack_thread_reply.

Jira uses a service-account REST client (agent/utils/jira.py, Basic auth) with ADF↔markdown conversion (agent/utils/adf.py); Confluence likewise (agent/utils/confluence.py, XHTML storage-format). Both are dark-safe: unset env returns a clean error.

Reviewer-only tools (in agent/reviewer.py): add_finding, update_finding, list_findings, publish_review (accepts verdict="approve"|"request_changes", honored only on verdict-authorized runs). request_pr_review (main agent) forwards the user's instructions verbatim and sets request_verdict=True only on an explicit ask. The review-style analyzer uses save_review_style (exported as save_review_style_prompt).

Built-in deepagents tools (read_file, write_file, edit_file, ls, glob, grep, execute, write_todos, task for subagent spawning, …) are added by create_deep_agent itself; don't duplicate them.

Models, profiles, and team defaults

Model + reasoning effort are resolved per run in this precedence (highest wins):

  1. Per-thread config (agent_model_id + agent_effort in configurable) — set by webhooks/UI.
  2. Per-user dashboard profile override (agent/dashboard/agent_overrides.py:load_profile), keyed by resolved GitHub login.
  3. Team default model (agent/dashboard/team_settings.py:get_team_default_model("agent")).

Supported model IDs and per-model effort/reasoning rules live in agent/dashboard/options.py. Profile flags also drive run behavior — e.g. profile_create_prs enables the opt-in Always Create PRs policy. Model construction goes through agent/utils/model.py (make_model, provider_model_kwargs, fallback_model_id_for).

Auth

  • GitHub: dual-mode. User OAuth tokens are encrypted-at-rest in thread metadata (agent/encryption.py, utils/auth.py:resolve_github_token). When no user token is available, falls back to a GitHub App installation token (utils/github_app.py). The installation token is also what configures the LangSmith sandbox's GitHub proxy.
  • Webhooks: GitHub signatures verified in utils/github_comments.py:verify_github_signature; Slack/Linear handled in their respective utils.
  • Dashboard / UI: GitHub OAuth login lives in agent/dashboard/oauth.py and routes.py (/auth/login, /auth/callback, /auth/logout, /me).

Thread-id derivation

Webhooks compute deterministic thread ids so the same Linear issue / Slack thread / PR routes back to the same running agent. See utils/github_comments.py:get_thread_id_from_branch and the equivalents in utils/linear.py / utils/slack.py. Reviewer threads have their own deterministic ids and are tagged with REVIEWER_THREAD_KIND metadata so the FastAPI side can find them.

Conventions

  • Tests are unit-only by default and organized by domain under tests/<domain>/ (tests/agent/, tests/auth/, tests/github/, tests/webhooks/, tests/reviewer/, tests/sandbox/, tests/middleware/, tests/models/, tests/slack/, tests/tools/, tests/dashboard/, tests/analyzer/); tests/conftest.py and the tests/e2e/ harness stay at the top level. Integration tests would go under tests/integration_tests/ (currently empty — make integration_tests no-ops if missing).
  • New sandbox providers: add a module under agent/integrations/ and wire it into SANDBOX_FACTORIES in agent/utils/sandbox.py. See CUSTOMIZATION.md.
  • New tools: add to agent/tools/, export from agent/tools/__init__.py, add to the tools=[...] list in server.py:get_agent (or reviewer.py for reviewer-only tools).
  • New middleware: add to agent/middleware/, export from agent/middleware/__init__.py, add to the middleware=[...] list in server.py:get_agent — order is significant (see the stack above).
  • New dashboard endpoints: add to agent/dashboard/routes.py. The router is auto-mounted on the FastAPI app.
  • New graphs: add an agent/graphs/<name>.py re-export shim that delegates to the factory module, then register the shim entrypoint (agent.graphs.<name>:<symbol>) in langgraph.json under graphs.
  • Minimal-to-no code comments — only when the why isn't obvious from the code.

Fork maintenance — syncing upstream/main

This is a long-lived fork of langchain-ai/open-swe with Sea Haven customizations woven into upstream-owned files (notably agent/prompt.py prompt constants, the agent/api/ + agent/webhooks/ FastAPI layer — the former agent/webapp.py monolith, now split and left as a shim — and the tool/middleware wiring). Merging upstream is a triage exercise, not a fast-forward. When you want upstream's clean changes but must defer a large structural refactor (and its entangled features), work in this order:

  1. Triage before resolving. Merge-base is git merge-base HEAD upstream/main. The truthful conflict set is the combined merge, git merge-tree --write-tree --name-only HEAD upstream/main — a per-commit probe against each commit's parent overstates conflicts (a file a refactor merely added shows up as a phantom modify/delete). Decide keep-baseline vs adopt-refactor before resolving, and surface the choice to a human for any auth/webhook/IAM surface.
  2. Chase the cascade, not just the textual conflicts. The hard part is the non-conflicting files the refactor also touched. Get the refactor's file set (git diff-tree --no-commit-id --name-status -r <refactor-sha>) and cross-reference the files this fork modified (git diff --name-only <fork-base> HEAD). Files in both = hand-resolve; files only the refactor touched = mechanical.
  3. Deferring a refactor: default every refactor-touched file to upstream, except the deleted-module cluster, which stays at your baseline (HEAD) — and move its paired tests to the same side. A file goes to HEAD when its upstream version imports a module the refactor deleted, or kept code needs an old API. Bring back files the refactor deleted but you still use with git checkout HEAD -- <file>. Iterate pytest --co -q to chase import breaks one module at a time.
  4. Two silent hazards. (a) A thin upstream router ends with from .webhooks.slack import process_slack_mention; merged alongside your monolith's local def process_slack_mention, Python rebinds the name at import, so upstream's handler runs and silently drops your fixes — delete those re-import lines. (b) A new tool/middleware importing a deleted module crashes the whole graph at import — if you defer the feature, delete the tool file and all its wiring (server.py tool list, tools/__init__.py, prompt guidance, e2e harness, its test).
  5. Keep test + impl on the same side — a test at upstream and its impl at HEAD (or vice-versa) yields async-vs-sync or contract drift. Keep the whole vertical (backend + UI + e2e spec + fixtures) on one side.
  6. Run CI in layers — ruff/tsc (syntax/types) → pytest --co (import-time breaks) → unit tests (contract mismatches) → E2E (Playwright + the real LangGraph dev server), which is the only layer that catches import-time crashes in tool/middleware wiring and frontend↔backend contract drift. "Unit green" is not "done" for a structural merge.
  7. Tooling-switch fallout — a package-manager/build-tool switch (upstream pnpm, this fork keeps bun) auto-merges into build scripts, CI, the packageManager field, and lockfiles even when you reject it for the product build. After merging, sweep those and never ship two lockfiles.

Validate on a throwaway branch with granular commits (one per cascade class) and let each CI layer prove out before promoting.