mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 09:13:14 +00:00
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat(open-swe): add Jira tool plane (Phase 1)
Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear
tools:
- utils/jira.py: service-account REST client (Basic auth) with get/
create/update issue, comments, list projects, trace comment; issue and
comment bodies normalized to markdown.
- utils/adf.py: minimal ADF <-> markdown conversion (read paths convert
Jira ADF to markdown; agent comments convert prose to ADF).
- tools/jira_{comment,get_issue,get_issue_comments,create_issue,
update_issue,list_projects}.py wired into the tool registry and the
main agent tool list.
- tests/test_jira_utils.py: ADF conversion + mocked-transport util tests.
Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env
returns a clean error, so this is safe to land dark. Trigger plane,
prompt guidance, and config plumbing follow in Phase 2.
* feat(open-swe): add Confluence tool plane (Phase 3)
Curated Confluence Cloud REST toolset for the agent, mirroring the Jira
tools:
- utils/confluence.py: service-account REST client (Basic auth) with
get/create/update page, add comment, CQL search. Page bodies are XHTML
storage format (not ADF), with minimal storage<->text converters;
update_page reads the current version and bumps it, as Confluence
requires.
- tools/confluence_{get_page,create_page,update_page,comment,search}.py
registered in the tool registry.
- tests/test_confluence_utils.py: converter + mocked-transport tests
including the version-bump path.
Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN;
unset env returns a clean error. Activation in the agent tool list lands
with the Phase 2 server.py wiring.
* feat(open-swe): add Jira trigger plane (Phase 2)
Make an @openswe comment on a Jira issue spawn an agent run, mirroring
the Linear trigger plane:
- webhooks/jira.py: process_jira_issue clones process_linear_issue —
deterministic thread id, full-issue fetch, actor accountId->email
attribution feeding resolve_login_from_email_async (PRs open as the
human), multimodal image handling, source="jira" + jira_issue config.
- webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time
X-Automation-Webhook-Token check, fails closed), repo-resolution
cascade, get_repo_config_from_jira_mapping.
- utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder
entry — real project->repo mappings still needed).
- utils/jira.py: get_user_email (accountId -> email) for attribution.
- completion.py: source=="jira" failure-reply branch.
- prompt.py: Jira-triggered notify guidance + Refs:/branch key from
{jira_project_key}-{jira_issue_number}.
- server.py: read jira_issue config + pass jira key to the system
prompt; also activates the Phase 3 Confluence tools in the agent list.
Jira Automation lacks native webhook HMAC signing, so trust is a shared
secret header (decision D2); replay protection is weaker than Linear's
HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the
outstanding gate/hardening before push.
* fix(open-swe): harden Jira webhook trust (sh-security-review)
Resolves findings from the Phase 2 security review (detector fan-out +
proof-or-kill verifier). The unsigned Jira Automation webhook body was
trusted for identity, comment content, repo routing, and issue
existence; a JIRA_WEBHOOK_SECRET holder could forge those fields.
- Corroborate against the real Jira record: the webhook body is now only
a pointer (issue_key + required comment_id). The triggering comment's
author and text are re-fetched server-side via get_comment/fetch_jira_
comment, and identity, the @openswe check, prompt text, and project
key are derived from that authoritative record — never payload author/
body fields. An uncorroborated comment is rejected. (closes the
account-id impersonation, unsigned-body prompt injection, and
fabricated-issue findings)
- Validate issue_key against the Jira key format and percent-encode all
untrusted path segments (_seg) so a crafted key can't traverse to a
different Jira REST endpoint or inject query params. (closes the path-
traversal / query-injection findings)
- Route source=="jira" through the bot-token-default / author_prs_as_
user opt-in path in resolve_github_token, matching Linear, instead of
unconditionally resolving a per-user OAuth token from a payload email.
- Gate attribution on an active user mapping (is_login_mapped) so a
pending/unconfirmed mapping can't drive PR authorship.
Adds regression tests: server-corroboration wins over payload, malformed
issue_key rejected, uncorroborated comment rejected, path-segment
encoding, project-key derivation, active-mapping gate.
Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/
REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body +
timestamp on the Automation payload to close the residual replay gap.
* harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist
Folds the two deployment-hardening items from the Phase 2 security review
into code (all opt-in / default-off, so existing and upstream deployments
are unaffected):
- JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must
carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by
JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by
verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear
HMAC+freshness model). Closes the static-token model's replay/forgery
gap when enabled.
- JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's
direct client IP (verify_jira_source_ip). Documented as direct-peer
only; behind a proxy/LB, allowlist Atlassian's ranges at that layer.
- REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail
CLOSED instead of the back-compat allow-all, plus a startup fail-open
warning. Applies to all channels for consistency.
Documents all new vars (and a Jira section) in .env.example. Adds tests
for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and
the fail-closed allowlist.
* feat(open-swe): Confluence Atlassian Connect trigger (Phase 4)
Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian
Connect app. Designed and adversarially verified with the ultracode
workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus
skeptics on the implemented crypto).
- utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's
official test vector), PyJWT HS256 webhook verifier with alg-pinning,
issuer binding, and qsh-verified-last ordering; RS256 signed-install
lifecycle verifier against Atlassian's published keys; installation
store keyed by clientKey with the sharedSecret encrypted at rest
(TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned).
- webhooks/confluence.py: install/uninstall lifecycle + comment handler.
The JWT-signed webhook body is only a pointer; the comment's real
author/text/container are re-fetched server-side via the Basic-auth
service account (Phase-2 corroboration lesson), with active-only login
attribution and the repo allowlist.
- utils/confluence.py: get_comment / get_user_email (path-encoded).
- webapp.py: GET /connect/atlassian-connect.json (served dynamically),
POST /connect/{installed,uninstalled,webhook/comment-created}, the
space->repo resolver, thread-id, and fetch helpers.
- completion.py: source=="confluence" failure-reply branch.
Security: the sh-security-review verify pass confirmed one HIGH — the
symmetric signed-install=false first-install was trust-on-first-use gated
only by the public Confluence hostname (webhook-auth bypass). Fixed by
switching to signed-install=true + RS256 verification of lifecycle
callbacks, which cryptographically authenticates the first install. All
other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall
DoS/corroboration/injection) were defeated; residuals are deployment
config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover
bodies; comment-trigger prompt injection, shared with all sources).
New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN,
CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets
require the durable Postgres LangGraph store in prod.
Outstanding before push: /sh-security-review on the real diff and the
GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory.
* docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs
- prompt.py: Confluence-triggered runs notify via confluence_comment on
the triggering page; add Confluence to the shared-base source list.
- CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian
triggers (Jira Automation shared-secret webhook; Confluence Connect app
with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the
server-side corroboration + encrypted install store.
Phase 5 also verified the trigger surface end-to-end against a running
uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed
without valid auth) and recorded the integration in project memory.
* fix(open-swe): resolve /sh-security-review findings on the Atlassian surface
Formal sh-security-review (detector fan-out + verifier) over the Phase-4
Connect surface (esp. the new RS256 signed-install code, unseen by the
earlier adversarial verify) and the Phase-2 opt-in hardening.
CRITICAL — cross-tenant install (origin validation, CWE-346): signed-
install proves the caller is *an* Atlassian tenant, not *ours*, and the
descriptor is served publicly, so any attacker could install the app on
their own Confluence site and drive agent runs against our allowlisted
repos. The baseUrl body field is attacker-controlled and cannot bind the
tenant; only the signature-verified clientKey (JWT iss) can. Added a
MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in
process_install after signature+iss verification.
HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment
ids are per-instance, so generate_thread_id_from_confluence_comment now
salts the hash with the verified clientKey (plumbed from the webhook JWT
iss) to prevent thread hijack across tenants.
HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page
interpolated page_id into the REST path unencoded (update_page on a
mutating PUT with no params= backstop). Now _seg()-encoded, matching the
rest of the module.
MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no
bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID
guard mirroring the Linear botActor / Jira comment_author_is_bot checks.
LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as
opt-in defense-in-depth (the clientKey allowlist is the real gate).
Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp,
kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption,
constant-time comparisons, and the Phase-2 hardening. New regression
tests for each fix; full suite green (1602).
* harden(open-swe): GPT-4.1 cross-family review follow-ups
Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no
critical/high issues and confirmed the auth boundary is fail-closed and
correct. Two low-cost defense-in-depth items applied:
- Validate the signed-install JWT 'kid' against a strict charset before
the public-key fetch, so a malformed kid fails fast with no network
call (on top of the existing fixed host + percent-encoding).
- Make JWT nbf verification explicit (verify_nbf) on both the RS256
lifecycle and HS256 webhook decodes.
Other suggestions triaged as already-handled (aud cross-app replay is
blocked by the per-tenant iss->secret lookup; documented static-token/IP/
baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra
(Fernet rotation via MultiFernet; rate limiting at the gateway).
* docs(open-swe): document Jira + Confluence in installation & customization guides
- INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret,
service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect
app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note,
CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars +
REQUIRE_REPO_ALLOWLIST.
- CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction
note covers all four sources.
- AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth).
- README.md: invocation section, tools table, and overview line.
135 lines
17 KiB
Markdown
135 lines
17 KiB
Markdown
# CLAUDE.md
|
|
|
|
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
|
|
|
|
## Project
|
|
|
|
Open SWE is an open-source coding-agent framework built on **LangGraph** + **Deep Agents** (`deepagents.create_deep_agent`). It runs as a LangGraph app: each thread spawns its own isolated cloud sandbox, and the agent is invoked from Slack, Linear, Jira, Confluence, or GitHub (PR comments, plus auto-review on opened / ready-for-review).
|
|
|
|
A separate **reviewer** graph runs read-only code reviews on PRs, and a **review-style analyzer** graph learns per-repo review style from historical PRs.
|
|
|
|
## Commands
|
|
|
|
Dependencies are managed with **uv**. Tests use pytest (`asyncio_mode = "auto"`). Lint/format is **ruff** (line-length 100, target py311). `requires-python = ">=3.11"`; `langgraph.json` pins the runtime to 3.12.
|
|
|
|
```bash
|
|
make install # uv pip install -e .
|
|
make dev # uv run langgraph dev — serves all three graphs + the FastAPI app from langgraph.json
|
|
make run # uvicorn agent.webapp:app --reload --port 8000 (FastAPI only, no LangGraph runtime)
|
|
make test # uv run pytest -vvv tests/
|
|
make test TEST_FILE=tests/test_open_pr_middleware.py # single test file
|
|
uv run pytest -vvv tests/test_open_pr_middleware.py::test_name # single test
|
|
make lint # ruff check + ruff format --diff
|
|
make format # ruff format + ruff check --fix
|
|
```
|
|
|
|
`langgraph.json` declares three graph entrypoints and the FastAPI app, all served together by `langgraph dev`:
|
|
|
|
| Graph | Entrypoint | Purpose |
|
|
|---|---|---|
|
|
| `agent` | `agent.server:traced_agent` (wraps `get_agent`) | Main coding agent (Slack/Linear/Jira/Confluence/GitHub-triggered). |
|
|
| `reviewer` | `agent.reviewer:traced_reviewer_agent` (wraps `get_reviewer_agent`) | Read-only PR reviewer. Findings model + `publish_review`. |
|
|
| `analyzer` | `agent.analyzer:traced_analyzer` (wraps `get_analyzer`) | Learns per-repo reviewer style from historical PRs and this reviewer's own finding outcomes. |
|
|
|
|
The FastAPI app is `agent.webapp:app`.
|
|
|
|
## Architecture
|
|
|
|
### Entrypoints
|
|
|
|
- **`agent/server.py` → `get_agent(config)`** — main graph factory. Called per-thread. Resolves the GitHub token, gets-or-creates the sandbox for the thread, resolves the team/profile/per-thread model + effort, then constructs a fresh `create_deep_agent(...)` with the curated tool list and middleware stack. The agent itself is stateless — all per-thread state lives in the sandbox + thread metadata.
|
|
- **`agent/reviewer.py` → `get_reviewer_agent(config)`** — reviewer graph factory. Shares `ensure_sandbox_for_thread` with the main agent but wires a reviewer-only toolset (`add_finding`, `update_finding`, `list_findings`, `publish_review`, `web_search`, `fetch_url`, `http_request`) and a different system prompt that pins the single-evolving-findings model and the diff-anchored bar for filing a finding. Read-only: no commit/push/PR-opening tools.
|
|
- **`agent/analyzer.py` → `get_analyzer(config)`** — small graph that emits a per-repo style prompt via the `save_review_style_prompt` tool, consumed by the reviewer as a "repository-specific review style" appendix. It runs in one of two modes (`analyzer_mode` in `configurable`): **bootstrap** (cold-start: crawl historical PR reviews) and **continual** (nightly: refine using this reviewer's own finding outcomes via `read_finding_outcomes`). Each mode's procedure lives in a deepagents **skill** (`agent/skills/bootstrap-repo-analysis/`, `agent/skills/continual-learning/`) served as virtual files via a `CompositeBackend` `/skills/` route + `StateBackend` (seeded into the run's `files` channel by the launcher — never written to the sandbox). Launchers and the per-repo nightly cron live in `agent/dashboard/review_style_jobs.py` and `agent/dashboard/analyzer_cron.py`; the cron is registered when bootstrap completes.
|
|
- **`agent/webapp.py`** — thin FastAPI routing layer mounted alongside the LangGraph server. Defines the webhook routes (GitHub, Linear, Slack, Jira, Confluence Connect `/connect/*`) plus `/webhooks/run-complete`, and keeps the shared helpers/constants; the per-source handlers live in **`agent/webhooks/{github,slack,linear,jira,confluence}.py`** (re-exported from `webapp` so existing call sites and tests keep working). Each webhook resolves a deterministic `thread_id` (so follow-up messages route to the same agent run) and triggers a run through the single durable dispatch contract in **`agent/dispatch.py`** (`dispatch_agent_run`: `multitask_strategy="interrupt"` + `durability="sync"` + completion webhook); `agent/completion.py` posts a failure reply if a run dies, and `agent/reconcile.py` (a `scheduler`-graph sweep) catches stragglers. The GitHub handler also auto-reviews PRs on `opened` / `ready_for_review` and drives the CI auto-fix flow (`agent/ci_autofix.py`).
|
|
- **Atlassian triggers.** The Jira trigger is a Jira **Automation** rule POSTing to `/webhooks/jira` with a shared-secret header (`verify_jira_secret`; optional HMAC-body+timestamp via `JIRA_WEBHOOK_REQUIRE_SIGNATURE`), since Jira Cloud has no native webhook signing. The Confluence trigger is a private **Atlassian Connect app** (`agent/utils/atlassian_connect.py`): the `comment_created` webhook is HS256-JWT-verified against the per-tenant stored `sharedSecret` with a hand-rolled `qsh` (query-string-hash) check; `signed-install` is on, so install/uninstall lifecycle callbacks are RS256-verified against Atlassian's published keys (no trust-on-first-use). Install secrets are stored **encrypted** in the LangGraph store, keyed by `clientKey`. Both Atlassian webhook bodies are treated as pointers only — the triggering comment's real author/text is re-fetched server-side via the Basic-auth service account before anything security-relevant is derived, and attribution is gated on an active user mapping (mirrors the GitHub-login / token-attribution flow). Descriptor served at `GET /connect/atlassian-connect.json`.
|
|
- **`agent/dashboard/`** — `router` mounted under the FastAPI app at startup (`app.include_router(dashboard_router)`). Owns GitHub OAuth, per-user profiles, admin endpoints, team defaults, enabled-repo lists, review-style management, and the Agents chat thread API used by the UI in `ui/`.
|
|
|
|
### Sandbox lifecycle (the tricky part)
|
|
|
|
`SANDBOX_BACKENDS` (in `agent/utils/sandbox_state.py`) is an in-process dict keyed by `thread_id`. Thread metadata persists `sandbox_id` across processes. `ensure_sandbox_for_thread` handles four cases:
|
|
|
|
1. Sandbox cached in memory → ping it (`echo ok`); recreate on `SandboxClientError`. Healthy reused sandboxes also get a GitHub-proxy refresh (recreate on failure).
|
|
2. Metadata says `__creating__` and no cache → poll until ready (`_wait_for_sandbox_id`).
|
|
3. No sandbox at all → set `__creating__` sentinel, create one, persist the real id.
|
|
4. Metadata has an id but no cache → reconnect; fall back to recreate on failure.
|
|
|
|
For `SANDBOX_TYPE=langsmith` (default), every sandbox creation/refresh also calls `_configure_github_proxy` with a fresh GitHub App installation token (`get_github_app_installation_token`). The proxy injects Basic auth for `github.com` git traffic and Bearer auth for `api.github.com` so sandbox commands can use `GH_TOKEN=dummy gh ...` without storing real tokens in the sandbox. Other providers (modal, daytona, runloop, local) skip the proxy step. Provider is selected via `SANDBOX_TYPE`; factory is `agent/utils/sandbox.py:create_sandbox` (`SANDBOX_FACTORIES` maps each provider name to a creator in `agent/integrations/`).
|
|
|
|
Every run re-applies `git config --global user.name/email` for the bot identity, because reused/reconnected sandboxes can lose `--global` config and Vercel preview deploys reject commits whose author email doesn't resolve to a GitHub account.
|
|
|
|
### Middleware stack (order matters)
|
|
|
|
Configured in `agent/server.py:get_agent`, runs around every model call (in this order):
|
|
|
|
1. `SanitizeToolInputsMiddleware` — strips/normalizes tool inputs before they reach tools.
|
|
2. `ModelCallLimitMiddleware` (from `langchain.agents.middleware`) — caps model calls at `MODEL_CALL_RECURSION_LIMIT` (~half of `DEFAULT_RECURSION_LIMIT`); `exit_behavior="end"`.
|
|
3. `ToolErrorMiddleware` — catches tool exceptions and surfaces them as tool messages.
|
|
4. `check_message_queue_before_model` — pulls Linear comments / Slack messages that arrived mid-run from the thread queue and injects them as user messages before the next LLM call. This is what makes "message the agent while it's working" work.
|
|
5. `SlackAssistantStatusMiddleware` — keeps the Slack "assistant is typing"-style status up to date around model calls.
|
|
6. `ensure_no_empty_msg` — after-model hook; when the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion) it re-injects a synthetic `no_op` / `confirming_completion` tool call so the run continues instead of ending prematurely.
|
|
7. `notify_step_limit_reached` — after-agent hook that posts a Slack reply when the agent hits the step limit, so the user gets a clear signal instead of silence.
|
|
8. `SandboxCircuitBreakerMiddleware` — trips the agent out of repeated sandbox failures instead of looping.
|
|
9. `ModelFallbackMiddleware` (optional, last) — added only when `LLM_FALLBACK_MODEL_ID` or the per-model default fallback differs from the primary model.
|
|
|
|
The system prompt instructs the agent to call a tool every turn, and `ensure_no_empty_msg` re-injects a tool call when it doesn't — together these keep runs from stopping partway through a task.
|
|
|
|
Other middleware exists in `agent/middleware/` (`ExcludeToolsMiddleware`) but isn't wired into the default agent. The reviewer uses a leaner stack: `SanitizeToolInputsMiddleware`, `ModelCallLimitMiddleware`, `ToolErrorMiddleware`, `SlackAssistantStatusMiddleware`.
|
|
|
|
There is intentionally no after-agent safety net that opens a PR for the agent. The agent itself is responsible for committing, pushing, opening/updating the draft PR, and replying in the source channel — all via `GH_TOKEN=dummy gh` and `slack_thread_reply` / `linear_comment`.
|
|
|
|
### Tools
|
|
|
|
All tools live in `agent/tools/` and are flat-imported via `agent/tools/__init__.py`. The set is intentionally small and curated — see README "Tools — Curated, Not Accumulated".
|
|
|
|
Wired into `get_agent`:
|
|
`http_request`, `fetch_url`, `web_search`, `linear_comment`, `linear_create_issue`, `linear_delete_issue`, `linear_get_issue`, `linear_get_issue_comments`, `linear_list_teams`, `linear_update_issue`, `jira_comment`, `jira_create_issue`, `jira_get_issue`, `jira_get_issue_comments`, `jira_list_projects`, `jira_update_issue`, `confluence_get_page`, `confluence_create_page`, `confluence_update_page`, `confluence_comment`, `confluence_search`, `request_pr_review`, `schedule_thread_wakeup`, `slack_add_reaction`, `slack_read_thread_messages`, `slack_thread_reply`.
|
|
|
|
Jira uses a service-account REST client (`agent/utils/jira.py`, Basic auth) with ADF↔markdown conversion (`agent/utils/adf.py`); Confluence likewise (`agent/utils/confluence.py`, XHTML storage-format). Both are dark-safe: unset env returns a clean error.
|
|
|
|
Reviewer-only tools (in `agent/reviewer.py`): `add_finding`, `update_finding`, `list_findings`, `publish_review`. The review-style analyzer uses `save_review_style` (exported as `save_review_style_prompt`).
|
|
|
|
Built-in deepagents tools (`read_file`, `write_file`, `edit_file`, `ls`, `glob`, `grep`, `execute`, `write_todos`, `task` for subagent spawning, …) are added by `create_deep_agent` itself; don't duplicate them.
|
|
|
|
### Models, profiles, and team defaults
|
|
|
|
Model + reasoning effort are resolved per run in this precedence (highest wins):
|
|
|
|
1. Per-thread config (`agent_model_id` + `agent_effort` in `configurable`) — set by webhooks/UI.
|
|
2. Per-user dashboard profile override (`agent/dashboard/agent_overrides.py:load_profile`), keyed by resolved GitHub login.
|
|
3. Team default model (`agent/dashboard/team_settings.py:get_team_default_model("agent")`).
|
|
|
|
Supported model IDs and per-model effort/reasoning rules live in `agent/dashboard/options.py`. Profile flags also drive run behavior — e.g. `profile_create_prs` enables the opt-in Always Create PRs policy. Model construction goes through `agent/utils/model.py` (`make_model`, `provider_model_kwargs`, `fallback_model_id_for`).
|
|
|
|
### Auth
|
|
|
|
- **GitHub**: dual-mode. User OAuth tokens are encrypted-at-rest in thread metadata (`agent/encryption.py`, `utils/auth.py:resolve_github_token`). When no user token is available, falls back to a GitHub App installation token (`utils/github_app.py`). The installation token is also what configures the LangSmith sandbox's GitHub proxy.
|
|
- **Webhooks**: GitHub signatures verified in `utils/github_comments.py:verify_github_signature`; Slack/Linear handled in their respective utils.
|
|
- **Dashboard / UI**: GitHub OAuth login lives in `agent/dashboard/oauth.py` and `routes.py` (`/auth/login`, `/auth/callback`, `/auth/logout`, `/me`).
|
|
|
|
### Thread-id derivation
|
|
|
|
Webhooks compute deterministic thread ids so the same Linear issue / Slack thread / PR routes back to the same running agent. See `utils/github_comments.py:get_thread_id_from_branch` and the equivalents in `utils/linear.py` / `utils/slack.py`. Reviewer threads have their own deterministic ids and are tagged with `REVIEWER_THREAD_KIND` metadata so the FastAPI side can find them.
|
|
|
|
## Conventions
|
|
|
|
- Tests are unit-only by default (`tests/`). Integration tests would go under `tests/integration_tests/` (currently empty — `make integration_tests` no-ops if missing).
|
|
- New sandbox providers: add a module under `agent/integrations/` and wire it into `SANDBOX_FACTORIES` in `agent/utils/sandbox.py`. See `CUSTOMIZATION.md`.
|
|
- New tools: add to `agent/tools/`, export from `agent/tools/__init__.py`, add to the `tools=[...]` list in `server.py:get_agent` (or `reviewer.py` for reviewer-only tools).
|
|
- New middleware: add to `agent/middleware/`, export from `agent/middleware/__init__.py`, add to the `middleware=[...]` list in `server.py:get_agent` — order is significant (see the stack above).
|
|
- New dashboard endpoints: add to `agent/dashboard/routes.py`. The router is auto-mounted on the FastAPI app.
|
|
- New graphs: register the entrypoint in `langgraph.json` under `graphs`.
|
|
- Minimal-to-no code comments — only when the *why* isn't obvious from the code.
|
|
|
|
## Fork maintenance — syncing `upstream/main`
|
|
|
|
This is a long-lived fork of `langchain-ai/open-swe` with Sea Haven customizations woven into upstream-owned files (notably `agent/prompt.py` prompt constants, `agent/webapp.py`, and the tool/middleware wiring). Merging upstream is a triage exercise, not a fast-forward. When you want upstream's clean changes but must **defer a large structural refactor** (and its entangled features), work in this order:
|
|
|
|
1. **Triage before resolving.** Merge-base is `git merge-base HEAD upstream/main`. The truthful conflict set is the combined merge, `git merge-tree --write-tree --name-only HEAD upstream/main` — a per-commit probe against each commit's parent *overstates* conflicts (a file a refactor merely added shows up as a phantom `modify/delete`). Decide keep-baseline vs adopt-refactor **before** resolving, and surface the choice to a human for any auth/webhook/IAM surface.
|
|
2. **Chase the cascade, not just the textual conflicts.** The hard part is the non-conflicting files the refactor also touched. Get the refactor's file set (`git diff-tree --no-commit-id --name-status -r <refactor-sha>`) and cross-reference the files this fork modified (`git diff --name-only <fork-base> HEAD`). Files in both = hand-resolve; files only the refactor touched = mechanical.
|
|
3. **Deferring a refactor:** default every refactor-touched file to **upstream**, except the deleted-module cluster, which stays at your **baseline (HEAD)** — and move its **paired tests to the same side**. A file goes to HEAD when its upstream version imports a module the refactor deleted, or kept code needs an old API. Bring back files the refactor deleted but you still use with `git checkout HEAD -- <file>`. Iterate `pytest --co -q` to chase import breaks one module at a time.
|
|
4. **Two silent hazards.** (a) A thin upstream router ends with `from .webhooks.slack import process_slack_mention`; merged alongside your monolith's *local* `def process_slack_mention`, Python rebinds the name at import, so **upstream's handler runs and silently drops your fixes** — delete those re-import lines. (b) A new tool/middleware importing a deleted module crashes the whole graph at import — if you defer the feature, delete the tool file **and** all its wiring (`server.py` tool list, `tools/__init__.py`, prompt guidance, e2e harness, its test).
|
|
5. **Keep test + impl on the same side** — a test at upstream and its impl at HEAD (or vice-versa) yields async-vs-sync or contract drift. Keep the whole vertical (backend + UI + e2e spec + fixtures) on one side.
|
|
6. **Run CI in layers** — `ruff`/`tsc` (syntax/types) → `pytest --co` (import-time breaks) → unit tests (contract mismatches) → **E2E (Playwright + the real LangGraph dev server)**, which is the only layer that catches import-time crashes in tool/middleware *wiring* and frontend↔backend contract drift. "Unit green" is not "done" for a structural merge.
|
|
7. **Tooling-switch fallout** — a package-manager/build-tool switch (upstream `pnpm`, this fork keeps `bun`) auto-merges into build scripts, CI, the `packageManager` field, and lockfiles even when you reject it for the product build. After merging, sweep those and never ship two lockfiles.
|
|
|
|
Validate on a throwaway branch with granular commits (one per cascade class) and let each CI layer prove out before promoting.
|