* feat(reviewer): clean-review auto-approve for unsolicited publish_review verdicts An unsolicited publish_review(verdict="approve") — a run dispatched without verdict_requested — is now honored when the review has zero open findings, so clean auto-reviews land a real APPROVE. With open findings it downgrades to a comment review (verdict_ignored_reason= "approve_with_open_findings"). request_changes stays explicit-request- only; the self-review, head-moved, and author-unknown downgrades and the shell verdict guard are unchanged. The reviewer base prompt now instructs the clean-approve call on auto-reviews. * chore(security): record accepted-risk suppressions for clean-review auto-approve Two confirmed-HIGH findings from /sh-security-review on the clean-review auto-approve change are accepted and deferred (Adam, 2026-07-21), tracked in #218. Machine-recorded per the mandatory-security-review policy; the revisit trigger is promotion from dev to main/prod.
21 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Project
Open SWE is an open-source coding-agent framework built on LangGraph + Deep Agents (deepagents.create_deep_agent). It runs as a LangGraph app: each thread spawns its own isolated cloud sandbox, and the agent is invoked from Slack, Linear, Jira, Confluence, or GitHub (PR comments, plus auto-review on opened / ready-for-review).
A separate reviewer graph runs read-only code reviews on PRs, and a review-style analyzer graph learns per-repo review style from historical PRs.
Commands
Dependencies are managed with uv. Tests use pytest (asyncio_mode = "auto"). Lint/format is ruff (line-length 100, target py311). requires-python = ">=3.11"; langgraph.json pins the runtime to 3.12.
make install # uv pip install -e .
make dev # uv run langgraph dev — serves all six graphs + the FastAPI app from langgraph.json
make run # uvicorn agent.webapp:app --reload --port 8000 (FastAPI only, no LangGraph runtime)
make test # uv run pytest -vvv tests/
make test TEST_FILE=tests/github/test_open_pull_request.py # single test file
uv run pytest -vvv tests/github/test_open_pull_request.py::test_name # single test
make lint # ruff check + ruff format --diff
make format # ruff format + ruff check --fix
langgraph.json declares six graph entrypoints and the FastAPI app, all served together by langgraph dev. Since the domain reorg, every graph entrypoint targets a thin agent.graphs.* re-export shim (agent/graphs/<name>.py) rather than the factory module directly; the shims delegate to the unmoved factories (agent/server.py, agent/reviewer.py, agent/analyzer.py, agent/chat.py, agent/scheduler.py, agent/ci_monitor.py):
| Graph | Entrypoint | Purpose |
|---|---|---|
agent |
agent.graphs.agent:traced_agent (shim → agent.server:get_agent) |
Main coding agent (Slack/Linear/Jira/Confluence/GitHub-triggered). |
reviewer |
agent.graphs.reviewer:traced_reviewer_agent (shim → agent.reviewer:get_reviewer_agent) |
Read-only PR reviewer. Findings model + publish_review. |
analyzer |
agent.graphs.analyzer:traced_analyzer (shim → agent.analyzer:get_analyzer) |
Learns per-repo reviewer style from historical PRs and this reviewer's own finding outcomes. |
chat |
agent.graphs.chat:traced_chat_agent |
Dashboard Agents chat graph. |
scheduler |
agent.graphs.scheduler:get_scheduler |
Reconcile sweep for stragglers. |
ci_monitor |
agent.graphs.ci_monitor:get_ci_monitor |
Fork-only polling fallback for the CI auto-fix flow. |
The FastAPI app is agent.webapp:app — now a compatibility shim that re-exports agent.api.app:app (see the FastAPI split under "Entrypoints").
Architecture
Entrypoints
agent/server.py→get_agent(config)— main graph factory. Called per-thread. Resolves the GitHub token, gets-or-creates the sandbox for the thread, resolves the team/profile/per-thread model + effort, then constructs a freshcreate_deep_agent(...)with the curated tool list and middleware stack. The agent itself is stateless — all per-thread state lives in the sandbox + thread metadata.agent/reviewer.py→get_reviewer_agent(config)— reviewer graph factory. Sharesensure_sandbox_for_threadwith the main agent but wires a reviewer-only toolset (add_finding,update_finding,list_findings,publish_review,web_search,fetch_url,http_request) and a different system prompt that pins the single-evolving-findings model and the diff-anchored bar for filing a finding. Read-only: no commit/push/PR-opening tools. Its supporting modules live in theagent/review/package (domain reorg):findings.py,publish.py,reconcile.py,trace_context.py,diff.py,groups.py,eval_store.py,style_collector.py,style_guidance.py.agent/analyzer.py→get_analyzer(config)— small graph that emits a per-repo style prompt via thesave_review_style_prompttool, consumed by the reviewer as a "repository-specific review style" appendix. It runs in one of two modes (analyzer_modeinconfigurable): bootstrap (cold-start: crawl historical PR reviews) and continual (nightly: refine using this reviewer's own finding outcomes viaread_finding_outcomes). Each mode's procedure lives in a deepagents skill (agent/skills/bootstrap-repo-analysis/,agent/skills/continual-learning/) served as virtual files via aCompositeBackend/skills/route +StateBackend(seeded into the run'sfileschannel by the launcher — never written to the sandbox). Launchers and the per-repo nightly cron live inagent/dashboard/review_style_jobs.pyandagent/dashboard/analyzer_cron.py; the cron is registered when bootstrap completes.- FastAPI layer (
agent/api/+agent/webhooks/) — the domain reorg split the fork's former 2,590-lineagent/webapp.pymonolith into a per-source layout;agent/webapp.pyis now a 4-line compatibility shim (from .api.app import app). The pieces:agent/api/app.pycomposes the FastAPI app and mounts every router;agent/api/health.pyowns/healthand/webhooks/run-complete.agent/webhooks/common.pyholds the shared verify/dispatch helpers and constants (signature verification, thread-id derivation, the dispatch entry). Route modules and handlers reach these via module-attribute access (common.X) so the ~240monkeypatch.setattrtest sites retarget cleanly.- Per-source route modules
agent/webhooks/{github,linear,slack,jira,confluence}_routes.pydefine the HTTP routes (GitHub, Linear, Slack, Jira, and the Confluence Atlassian Connect/connect/*+/webhooks/jiraroutes; the fork-onlyjira_routes.py/confluence_routes.pymirror upstream'sgithub_routes.pypattern, and the Connect lifecycle/descriptor routes fold intoconfluence_routes.py). The handler modulesagent/webhooks/{github,slack,linear,jira,confluence}.pycarry the per-source logic. - Each webhook resolves a deterministic
thread_id(so follow-up messages route to the same agent run) and triggers a run through the single durable dispatch contract inagent/dispatch.py(dispatch_agent_run:multitask_strategy="interrupt"+durability="sync"+ completion webhook);agent/completion.pyposts a failure reply if a run dies, andagent/reconcile.py(ascheduler-graph sweep) catches stragglers. The GitHub handler also auto-reviews PRs onopened/ready_for_reviewand drives the CI auto-fix flow (agent/ci_autofix.py). - Atlassian triggers. The Jira trigger is a Jira Automation rule POSTing to
/webhooks/jirawith a shared-secret header (verify_jira_secret; optional HMAC-body+timestamp viaJIRA_WEBHOOK_REQUIRE_SIGNATURE), since Jira Cloud has no native webhook signing. The Confluence trigger is a private Atlassian Connect app (agent/utils/atlassian_connect.py): thecomment_createdwebhook is HS256-JWT-verified against the per-tenant storedsharedSecretwith a hand-rolledqsh(query-string-hash) check;signed-installis on, so install/uninstall lifecycle callbacks are RS256-verified against Atlassian's published keys (no trust-on-first-use). Install secrets are stored encrypted in the LangGraph store, keyed byclientKey. Both Atlassian webhook bodies are treated as pointers only — the triggering comment's real author/text is re-fetched server-side via the Basic-auth service account before anything security-relevant is derived, and attribution is gated on an active user mapping (mirrors the GitHub-login / token-attribution flow). Descriptor served atGET /connect/atlassian-connect.json.
agent/dashboard/—routermounted under the FastAPI app at startup (app.include_router(dashboard_router)). Owns GitHub OAuth, per-user profiles, admin endpoints, team defaults, enabled-repo lists, review-style management, and the Agents chat thread API used by the UI inui/.
Sandbox lifecycle (the tricky part)
SANDBOX_BACKENDS (in agent/utils/sandbox_state.py) is an in-process dict keyed by thread_id. Thread metadata persists sandbox_id across processes. ensure_sandbox_for_thread handles four cases:
- Sandbox cached in memory → ping it (
echo ok); recreate onSandboxClientError. Healthy reused sandboxes also get a GitHub-proxy refresh (recreate on failure). - Metadata says
__creating__and no cache → poll until ready (_wait_for_sandbox_id). - No sandbox at all → set
__creating__sentinel, create one, persist the real id. - Metadata has an id but no cache → reconnect; fall back to recreate on failure.
For SANDBOX_TYPE=langsmith (default), every sandbox creation/refresh also calls _configure_github_proxy with a fresh GitHub App installation token (get_github_app_installation_token). The proxy injects Basic auth for github.com git traffic and Bearer auth for api.github.com so sandbox commands can use GH_TOKEN=dummy gh ... without storing real tokens in the sandbox. Other providers (modal, daytona, runloop, local) skip the proxy step. Provider is selected via SANDBOX_TYPE; factory is agent/utils/sandbox.py:create_sandbox (SANDBOX_FACTORIES maps each provider name to a creator in agent/integrations/).
Every run re-applies git config --global user.name/email for the bot identity, because reused/reconnected sandboxes can lose --global config and Vercel preview deploys reject commits whose author email doesn't resolve to a GitHub account.
Middleware stack (order matters)
Configured in agent/server.py:get_agent, runs around every model call (in this order):
SanitizeToolInputsMiddleware— strips/normalizes tool inputs before they reach tools.ModelCallLimitMiddleware(fromlangchain.agents.middleware) — caps model calls atMODEL_CALL_RECURSION_LIMIT(~half ofDEFAULT_RECURSION_LIMIT);exit_behavior="end".ToolErrorMiddleware— catches tool exceptions and surfaces them as tool messages.check_message_queue_before_model— pulls Linear comments / Slack messages that arrived mid-run from the thread queue and injects them as user messages before the next LLM call. This is what makes "message the agent while it's working" work.SlackAssistantStatusMiddleware— keeps the Slack "assistant is typing"-style status up to date around model calls.ensure_no_empty_msg— after-model hook; when the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion) it re-injects a syntheticno_op/confirming_completiontool call so the run continues instead of ending prematurely.notify_step_limit_reached— after-agent hook that posts a Slack reply when the agent hits the step limit, so the user gets a clear signal instead of silence.SandboxCircuitBreakerMiddleware— trips the agent out of repeated sandbox failures instead of looping.ModelFallbackMiddleware(optional, last) — added only whenLLM_FALLBACK_MODEL_IDor the per-model default fallback differs from the primary model.
The system prompt instructs the agent to call a tool every turn, and ensure_no_empty_msg re-injects a tool call when it doesn't — together these keep runs from stopping partway through a task.
Other middleware exists in agent/middleware/ (ExcludeToolsMiddleware) but isn't wired into the default agent. The reviewer uses a leaner stack (see reviewer.py:get_reviewer_agent for the authoritative order), including SanitizeToolInputsMiddleware, ModelCallLimitMiddleware, ToolErrorMiddleware, PullRequestVerdictGuardMiddleware, SlackAssistantStatusMiddleware, and settle_review_check_on_exit.
Verdict gating (Sea Haven fork): PullRequestVerdictGuardMiddleware (agent/middleware/pr_verdict_guard.py) is wired into BOTH graphs (server.py:get_agent after PullRequestCreationGuardMiddleware; reviewer.py:get_reviewer_agent after ToolErrorMiddleware). It blocks shell-path review verdicts (gh pr review --approve/-a/--request-changes/-r, gh api/curl posting event=APPROVE|REQUEST_CHANGES to /pulls/N/reviews); comment reviews and reads pass through. Verdicts go exclusively through publish_review(verdict=...): request_changes is honored only when the run's configurable["verdict_requested"] is True — set solely by the explicit-mention dispatch path (trigger_pr_review_from_ref(request_verdict=True) → _build_reviewer_configurable); auto-review dispatches never set it. approve is additionally honored on non-verdict-requested runs when the review is clean (zero open findings — the clean-review auto-approve); with open findings it downgrades to a comment (verdict_ignored_reason="approve_with_open_findings"). The tool layer also downgrades self-reviews (PR author in INTERNAL_BOT_LOGINS) to comment reviews and best-effort dismisses a recorded stale APPROVE when a later publish surfaces new findings.
There is intentionally no after-agent safety net that opens a PR for the agent. The agent itself is responsible for committing, pushing, opening/updating the draft PR, and replying in the source channel — all via GH_TOKEN=dummy gh and slack_thread_reply / linear_comment.
Tools
All tools live in agent/tools/ and are flat-imported via agent/tools/__init__.py. The set is intentionally small and curated — see README "Tools — Curated, Not Accumulated".
Wired into get_agent:
http_request, fetch_url, web_search, linear_comment, linear_create_issue, linear_delete_issue, linear_get_issue, linear_get_issue_comments, linear_list_teams, linear_search_issues, linear_update_issue, jira_comment, jira_create_issue, jira_get_issue, jira_get_issue_comments, jira_list_projects, jira_update_issue, confluence_get_page, confluence_create_page, confluence_update_page, confluence_comment, confluence_search, request_pr_review, schedule_thread_wakeup, slack_add_reaction, slack_read_thread_messages, slack_thread_reply.
Jira uses a service-account REST client (agent/utils/jira.py, Basic auth) with ADF↔markdown conversion (agent/utils/adf.py); Confluence likewise (agent/utils/confluence.py, XHTML storage-format). Both are dark-safe: unset env returns a clean error.
Reviewer-only tools (in agent/reviewer.py): add_finding, update_finding, list_findings, publish_review (accepts verdict="approve"|"request_changes", honored only on verdict-authorized runs). request_pr_review (main agent) forwards the user's instructions verbatim and sets request_verdict=True only on an explicit ask. The review-style analyzer uses save_review_style (exported as save_review_style_prompt).
Built-in deepagents tools (read_file, write_file, edit_file, ls, glob, grep, execute, write_todos, task for subagent spawning, …) are added by create_deep_agent itself; don't duplicate them.
Models, profiles, and team defaults
Model + reasoning effort are resolved per run in this precedence (highest wins):
- Per-thread config (
agent_model_id+agent_effortinconfigurable) — set by webhooks/UI. - Per-user dashboard profile override (
agent/dashboard/agent_overrides.py:load_profile), keyed by resolved GitHub login. - Team default model (
agent/dashboard/team_settings.py:get_team_default_model("agent")).
Supported model IDs and per-model effort/reasoning rules live in agent/dashboard/options.py. Profile flags also drive run behavior — e.g. profile_create_prs enables the opt-in Always Create PRs policy. Model construction goes through agent/utils/model.py (make_model, provider_model_kwargs, fallback_model_id_for).
Auth
- GitHub: dual-mode. User OAuth tokens are encrypted-at-rest in thread metadata (
agent/encryption.py,utils/auth.py:resolve_github_token). When no user token is available, falls back to a GitHub App installation token (utils/github_app.py). The installation token is also what configures the LangSmith sandbox's GitHub proxy. - Webhooks: GitHub signatures verified in
utils/github_comments.py:verify_github_signature; Slack/Linear handled in their respective utils. - Dashboard / UI: GitHub OAuth login lives in
agent/dashboard/oauth.pyandroutes.py(/auth/login,/auth/callback,/auth/logout,/me).
Thread-id derivation
Webhooks compute deterministic thread ids so the same Linear issue / Slack thread / PR routes back to the same running agent. See utils/github_comments.py:get_thread_id_from_branch and the equivalents in utils/linear.py / utils/slack.py. Reviewer threads have their own deterministic ids and are tagged with REVIEWER_THREAD_KIND metadata so the FastAPI side can find them.
Conventions
- Tests are unit-only by default and organized by domain under
tests/<domain>/(tests/agent/,tests/auth/,tests/github/,tests/webhooks/,tests/reviewer/,tests/sandbox/,tests/middleware/,tests/models/,tests/slack/,tests/tools/,tests/dashboard/,tests/analyzer/);tests/conftest.pyand thetests/e2e/harness stay at the top level. Integration tests would go undertests/integration_tests/(currently empty —make integration_testsno-ops if missing). - New sandbox providers: add a module under
agent/integrations/and wire it intoSANDBOX_FACTORIESinagent/utils/sandbox.py. SeeCUSTOMIZATION.md. - New tools: add to
agent/tools/, export fromagent/tools/__init__.py, add to thetools=[...]list inserver.py:get_agent(orreviewer.pyfor reviewer-only tools). - New middleware: add to
agent/middleware/, export fromagent/middleware/__init__.py, add to themiddleware=[...]list inserver.py:get_agent— order is significant (see the stack above). - New dashboard endpoints: add to
agent/dashboard/routes.py. The router is auto-mounted on the FastAPI app. - New graphs: add an
agent/graphs/<name>.pyre-export shim that delegates to the factory module, then register the shim entrypoint (agent.graphs.<name>:<symbol>) inlanggraph.jsonundergraphs. - Minimal-to-no code comments — only when the why isn't obvious from the code.
Fork maintenance — syncing upstream/main
This is a long-lived fork of langchain-ai/open-swe with Sea Haven customizations woven into upstream-owned files (notably agent/prompt.py prompt constants, the agent/api/ + agent/webhooks/ FastAPI layer — the former agent/webapp.py monolith, now split and left as a shim — and the tool/middleware wiring). Merging upstream is a triage exercise, not a fast-forward. When you want upstream's clean changes but must defer a large structural refactor (and its entangled features), work in this order:
- Triage before resolving. Merge-base is
git merge-base HEAD upstream/main. The truthful conflict set is the combined merge,git merge-tree --write-tree --name-only HEAD upstream/main— a per-commit probe against each commit's parent overstates conflicts (a file a refactor merely added shows up as a phantommodify/delete). Decide keep-baseline vs adopt-refactor before resolving, and surface the choice to a human for any auth/webhook/IAM surface. - Chase the cascade, not just the textual conflicts. The hard part is the non-conflicting files the refactor also touched. Get the refactor's file set (
git diff-tree --no-commit-id --name-status -r <refactor-sha>) and cross-reference the files this fork modified (git diff --name-only <fork-base> HEAD). Files in both = hand-resolve; files only the refactor touched = mechanical. - Deferring a refactor: default every refactor-touched file to upstream, except the deleted-module cluster, which stays at your baseline (HEAD) — and move its paired tests to the same side. A file goes to HEAD when its upstream version imports a module the refactor deleted, or kept code needs an old API. Bring back files the refactor deleted but you still use with
git checkout HEAD -- <file>. Iteratepytest --co -qto chase import breaks one module at a time. - Two silent hazards. (a) A thin upstream router ends with
from .webhooks.slack import process_slack_mention; merged alongside your monolith's localdef process_slack_mention, Python rebinds the name at import, so upstream's handler runs and silently drops your fixes — delete those re-import lines. (b) A new tool/middleware importing a deleted module crashes the whole graph at import — if you defer the feature, delete the tool file and all its wiring (server.pytool list,tools/__init__.py, prompt guidance, e2e harness, its test). - Keep test + impl on the same side — a test at upstream and its impl at HEAD (or vice-versa) yields async-vs-sync or contract drift. Keep the whole vertical (backend + UI + e2e spec + fixtures) on one side.
- Run CI in layers —
ruff/tsc(syntax/types) →pytest --co(import-time breaks) → unit tests (contract mismatches) → E2E (Playwright + the real LangGraph dev server), which is the only layer that catches import-time crashes in tool/middleware wiring and frontend↔backend contract drift. "Unit green" is not "done" for a structural merge. - Tooling-switch fallout — a package-manager/build-tool switch (upstream
pnpm, this fork keepsbun) auto-merges into build scripts, CI, thepackageManagerfield, and lockfiles even when you reject it for the product build. After merging, sweep those and never ship two lockfiles.
Validate on a throwaway branch with granular commits (one per cascade class) and let each CI layer prove out before promoting.