2026-04-30 20:06:58 -07:00
# AGENTS.md
This file provides guidance to Coding Agents when working with code in this repository.
## Project
feat: Jira + Confluence integration (tools + triggers) (#182)
* feat(open-swe): add Jira tool plane (Phase 1)
Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear
tools:
- utils/jira.py: service-account REST client (Basic auth) with get/
create/update issue, comments, list projects, trace comment; issue and
comment bodies normalized to markdown.
- utils/adf.py: minimal ADF <-> markdown conversion (read paths convert
Jira ADF to markdown; agent comments convert prose to ADF).
- tools/jira_{comment,get_issue,get_issue_comments,create_issue,
update_issue,list_projects}.py wired into the tool registry and the
main agent tool list.
- tests/test_jira_utils.py: ADF conversion + mocked-transport util tests.
Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env
returns a clean error, so this is safe to land dark. Trigger plane,
prompt guidance, and config plumbing follow in Phase 2.
* feat(open-swe): add Confluence tool plane (Phase 3)
Curated Confluence Cloud REST toolset for the agent, mirroring the Jira
tools:
- utils/confluence.py: service-account REST client (Basic auth) with
get/create/update page, add comment, CQL search. Page bodies are XHTML
storage format (not ADF), with minimal storage<->text converters;
update_page reads the current version and bumps it, as Confluence
requires.
- tools/confluence_{get_page,create_page,update_page,comment,search}.py
registered in the tool registry.
- tests/test_confluence_utils.py: converter + mocked-transport tests
including the version-bump path.
Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN;
unset env returns a clean error. Activation in the agent tool list lands
with the Phase 2 server.py wiring.
* feat(open-swe): add Jira trigger plane (Phase 2)
Make an @openswe comment on a Jira issue spawn an agent run, mirroring
the Linear trigger plane:
- webhooks/jira.py: process_jira_issue clones process_linear_issue —
deterministic thread id, full-issue fetch, actor accountId->email
attribution feeding resolve_login_from_email_async (PRs open as the
human), multimodal image handling, source="jira" + jira_issue config.
- webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time
X-Automation-Webhook-Token check, fails closed), repo-resolution
cascade, get_repo_config_from_jira_mapping.
- utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder
entry — real project->repo mappings still needed).
- utils/jira.py: get_user_email (accountId -> email) for attribution.
- completion.py: source=="jira" failure-reply branch.
- prompt.py: Jira-triggered notify guidance + Refs:/branch key from
{jira_project_key}-{jira_issue_number}.
- server.py: read jira_issue config + pass jira key to the system
prompt; also activates the Phase 3 Confluence tools in the agent list.
Jira Automation lacks native webhook HMAC signing, so trust is a shared
secret header (decision D2); replay protection is weaker than Linear's
HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the
outstanding gate/hardening before push.
* fix(open-swe): harden Jira webhook trust (sh-security-review)
Resolves findings from the Phase 2 security review (detector fan-out +
proof-or-kill verifier). The unsigned Jira Automation webhook body was
trusted for identity, comment content, repo routing, and issue
existence; a JIRA_WEBHOOK_SECRET holder could forge those fields.
- Corroborate against the real Jira record: the webhook body is now only
a pointer (issue_key + required comment_id). The triggering comment's
author and text are re-fetched server-side via get_comment/fetch_jira_
comment, and identity, the @openswe check, prompt text, and project
key are derived from that authoritative record — never payload author/
body fields. An uncorroborated comment is rejected. (closes the
account-id impersonation, unsigned-body prompt injection, and
fabricated-issue findings)
- Validate issue_key against the Jira key format and percent-encode all
untrusted path segments (_seg) so a crafted key can't traverse to a
different Jira REST endpoint or inject query params. (closes the path-
traversal / query-injection findings)
- Route source=="jira" through the bot-token-default / author_prs_as_
user opt-in path in resolve_github_token, matching Linear, instead of
unconditionally resolving a per-user OAuth token from a payload email.
- Gate attribution on an active user mapping (is_login_mapped) so a
pending/unconfirmed mapping can't drive PR authorship.
Adds regression tests: server-corroboration wins over payload, malformed
issue_key rejected, uncorroborated comment rejected, path-segment
encoding, project-key derivation, active-mapping gate.
Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/
REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body +
timestamp on the Automation payload to close the residual replay gap.
* harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist
Folds the two deployment-hardening items from the Phase 2 security review
into code (all opt-in / default-off, so existing and upstream deployments
are unaffected):
- JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must
carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by
JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by
verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear
HMAC+freshness model). Closes the static-token model's replay/forgery
gap when enabled.
- JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's
direct client IP (verify_jira_source_ip). Documented as direct-peer
only; behind a proxy/LB, allowlist Atlassian's ranges at that layer.
- REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail
CLOSED instead of the back-compat allow-all, plus a startup fail-open
warning. Applies to all channels for consistency.
Documents all new vars (and a Jira section) in .env.example. Adds tests
for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and
the fail-closed allowlist.
* feat(open-swe): Confluence Atlassian Connect trigger (Phase 4)
Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian
Connect app. Designed and adversarially verified with the ultracode
workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus
skeptics on the implemented crypto).
- utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's
official test vector), PyJWT HS256 webhook verifier with alg-pinning,
issuer binding, and qsh-verified-last ordering; RS256 signed-install
lifecycle verifier against Atlassian's published keys; installation
store keyed by clientKey with the sharedSecret encrypted at rest
(TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned).
- webhooks/confluence.py: install/uninstall lifecycle + comment handler.
The JWT-signed webhook body is only a pointer; the comment's real
author/text/container are re-fetched server-side via the Basic-auth
service account (Phase-2 corroboration lesson), with active-only login
attribution and the repo allowlist.
- utils/confluence.py: get_comment / get_user_email (path-encoded).
- webapp.py: GET /connect/atlassian-connect.json (served dynamically),
POST /connect/{installed,uninstalled,webhook/comment-created}, the
space->repo resolver, thread-id, and fetch helpers.
- completion.py: source=="confluence" failure-reply branch.
Security: the sh-security-review verify pass confirmed one HIGH — the
symmetric signed-install=false first-install was trust-on-first-use gated
only by the public Confluence hostname (webhook-auth bypass). Fixed by
switching to signed-install=true + RS256 verification of lifecycle
callbacks, which cryptographically authenticates the first install. All
other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall
DoS/corroboration/injection) were defeated; residuals are deployment
config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover
bodies; comment-trigger prompt injection, shared with all sources).
New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN,
CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets
require the durable Postgres LangGraph store in prod.
Outstanding before push: /sh-security-review on the real diff and the
GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory.
* docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs
- prompt.py: Confluence-triggered runs notify via confluence_comment on
the triggering page; add Confluence to the shared-base source list.
- CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian
triggers (Jira Automation shared-secret webhook; Confluence Connect app
with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the
server-side corroboration + encrypted install store.
Phase 5 also verified the trigger surface end-to-end against a running
uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed
without valid auth) and recorded the integration in project memory.
* fix(open-swe): resolve /sh-security-review findings on the Atlassian surface
Formal sh-security-review (detector fan-out + verifier) over the Phase-4
Connect surface (esp. the new RS256 signed-install code, unseen by the
earlier adversarial verify) and the Phase-2 opt-in hardening.
CRITICAL — cross-tenant install (origin validation, CWE-346): signed-
install proves the caller is *an* Atlassian tenant, not *ours*, and the
descriptor is served publicly, so any attacker could install the app on
their own Confluence site and drive agent runs against our allowlisted
repos. The baseUrl body field is attacker-controlled and cannot bind the
tenant; only the signature-verified clientKey (JWT iss) can. Added a
MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in
process_install after signature+iss verification.
HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment
ids are per-instance, so generate_thread_id_from_confluence_comment now
salts the hash with the verified clientKey (plumbed from the webhook JWT
iss) to prevent thread hijack across tenants.
HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page
interpolated page_id into the REST path unencoded (update_page on a
mutating PUT with no params= backstop). Now _seg()-encoded, matching the
rest of the module.
MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no
bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID
guard mirroring the Linear botActor / Jira comment_author_is_bot checks.
LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as
opt-in defense-in-depth (the clientKey allowlist is the real gate).
Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp,
kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption,
constant-time comparisons, and the Phase-2 hardening. New regression
tests for each fix; full suite green (1602).
* harden(open-swe): GPT-4.1 cross-family review follow-ups
Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no
critical/high issues and confirmed the auth boundary is fail-closed and
correct. Two low-cost defense-in-depth items applied:
- Validate the signed-install JWT 'kid' against a strict charset before
the public-key fetch, so a malformed kid fails fast with no network
call (on top of the existing fixed host + percent-encoding).
- Make JWT nbf verification explicit (verify_nbf) on both the RS256
lifecycle and HS256 webhook decodes.
Other suggestions triaged as already-handled (aud cross-app replay is
blocked by the per-tenant iss->secret lookup; documented static-token/IP/
baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra
(Fernet rotation via MultiFernet; rate limiting at the gateway).
* docs(open-swe): document Jira + Confluence in installation & customization guides
- INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret,
service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect
app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note,
CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars +
REQUIRE_REPO_ALLOWLIST.
- CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction
note covers all four sources.
- AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth).
- README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
Open SWE is an open-source coding-agent framework built on **LangGraph** + **Deep Agents** (`deepagents.create_deep_agent` ). It runs as a LangGraph app: each thread spawns its own isolated cloud sandbox, and the agent is invoked from Slack, Linear, Jira, Confluence, or GitHub (PR comments, plus auto-review on opened / ready-for-review).
2026-05-22 14:33:46 -07:00
A separate **reviewer** graph runs read-only code reviews on PRs, and a **review-style analyzer** graph learns per-repo review style from historical PRs.
2026-04-30 20:06:58 -07:00
## Commands
2026-05-22 14:33:46 -07:00
Dependencies are managed with **uv** . Tests use pytest (`asyncio_mode = "auto"` ). Lint/format is **ruff** (line-length 100, target py311). `requires-python = ">=3.11"` ; `langgraph.json` pins the runtime to 3.12.
2026-04-30 20:06:58 -07:00
```bash
make install # uv pip install -e .
docs: land reorg in ledger, fix path refs, ship plan artifacts (C7)
C7 of the domain-reorg adoption (build plan docs/upstream-sync/domain-reorg/
reorg-build-plan.md, step C7):
- CLAUDE.md / AGENTS.md: retarget architecture path references to the new
layout — graph entrypoints via agent.graphs.* shims, the agent/api/ +
agent/webhooks/*_routes.py FastAPI split (webapp.py now a shim),
agent/review/ package, tests/<domain>/ test paths, and the new-graph/
test conventions. README.md already pointed at docs/ (C1) — no change.
- triage.jsonl: flip 8356eb34 (#1726) to landed on branch
refactor/domain-reorg-adoption; add re-triage notes to the 8 unblocked
rows (#1732/#1761/#1744/#1748/#1742 clean, #1736/#1758/#1760 near-clean).
triage.md regenerated via make triage-render.
- Ship the plan, scoping report, move-map artifacts, and the
domain-reorg-adoption workflow so the exercise is reproducible.
Memory + Confluence handled out-of-band (not in this commit): project
memory updated with the new module layout; Confluence check found no IT
page documents the repo module map — no Confluence change required.
2026-07-17 15:01:45 -04:00
make dev # uv run langgraph dev — serves all six graphs + the FastAPI app from langgraph.json
2026-05-22 14:33:46 -07:00
make run # uvicorn agent.webapp:app --reload --port 8000 (FastAPI only, no LangGraph runtime)
2026-04-30 20:06:58 -07:00
make test # uv run pytest -vvv tests/
docs: land reorg in ledger, fix path refs, ship plan artifacts (C7)
C7 of the domain-reorg adoption (build plan docs/upstream-sync/domain-reorg/
reorg-build-plan.md, step C7):
- CLAUDE.md / AGENTS.md: retarget architecture path references to the new
layout — graph entrypoints via agent.graphs.* shims, the agent/api/ +
agent/webhooks/*_routes.py FastAPI split (webapp.py now a shim),
agent/review/ package, tests/<domain>/ test paths, and the new-graph/
test conventions. README.md already pointed at docs/ (C1) — no change.
- triage.jsonl: flip 8356eb34 (#1726) to landed on branch
refactor/domain-reorg-adoption; add re-triage notes to the 8 unblocked
rows (#1732/#1761/#1744/#1748/#1742 clean, #1736/#1758/#1760 near-clean).
triage.md regenerated via make triage-render.
- Ship the plan, scoping report, move-map artifacts, and the
domain-reorg-adoption workflow so the exercise is reproducible.
Memory + Confluence handled out-of-band (not in this commit): project
memory updated with the new module layout; Confluence check found no IT
page documents the repo module map — no Confluence change required.
2026-07-17 15:01:45 -04:00
make test TEST_FILE=tests/github/test_open_pull_request.py # single test file
uv run pytest -vvv tests/github/test_open_pull_request.py::test_name # single test
2026-04-30 20:06:58 -07:00
make lint # ruff check + ruff format --diff
make format # ruff format + ruff check --fix
```
docs: land reorg in ledger, fix path refs, ship plan artifacts (C7)
C7 of the domain-reorg adoption (build plan docs/upstream-sync/domain-reorg/
reorg-build-plan.md, step C7):
- CLAUDE.md / AGENTS.md: retarget architecture path references to the new
layout — graph entrypoints via agent.graphs.* shims, the agent/api/ +
agent/webhooks/*_routes.py FastAPI split (webapp.py now a shim),
agent/review/ package, tests/<domain>/ test paths, and the new-graph/
test conventions. README.md already pointed at docs/ (C1) — no change.
- triage.jsonl: flip 8356eb34 (#1726) to landed on branch
refactor/domain-reorg-adoption; add re-triage notes to the 8 unblocked
rows (#1732/#1761/#1744/#1748/#1742 clean, #1736/#1758/#1760 near-clean).
triage.md regenerated via make triage-render.
- Ship the plan, scoping report, move-map artifacts, and the
domain-reorg-adoption workflow so the exercise is reproducible.
Memory + Confluence handled out-of-band (not in this commit): project
memory updated with the new module layout; Confluence check found no IT
page documents the repo module map — no Confluence change required.
2026-07-17 15:01:45 -04:00
`langgraph.json` declares six graph entrypoints and the FastAPI app, all served together by `langgraph dev` . Every graph entrypoint targets a thin `agent.graphs.*` re-export shim that delegates to the unmoved factory module (domain reorg):
2026-05-22 14:33:46 -07:00
| Graph | Entrypoint | Purpose |
|---|---|---|
docs: land reorg in ledger, fix path refs, ship plan artifacts (C7)
C7 of the domain-reorg adoption (build plan docs/upstream-sync/domain-reorg/
reorg-build-plan.md, step C7):
- CLAUDE.md / AGENTS.md: retarget architecture path references to the new
layout — graph entrypoints via agent.graphs.* shims, the agent/api/ +
agent/webhooks/*_routes.py FastAPI split (webapp.py now a shim),
agent/review/ package, tests/<domain>/ test paths, and the new-graph/
test conventions. README.md already pointed at docs/ (C1) — no change.
- triage.jsonl: flip 8356eb34 (#1726) to landed on branch
refactor/domain-reorg-adoption; add re-triage notes to the 8 unblocked
rows (#1732/#1761/#1744/#1748/#1742 clean, #1736/#1758/#1760 near-clean).
triage.md regenerated via make triage-render.
- Ship the plan, scoping report, move-map artifacts, and the
domain-reorg-adoption workflow so the exercise is reproducible.
Memory + Confluence handled out-of-band (not in this commit): project
memory updated with the new module layout; Confluence check found no IT
page documents the repo module map — no Confluence change required.
2026-07-17 15:01:45 -04:00
| `agent` | `agent.graphs.agent:traced_agent` (shim → `agent.server:get_agent` ) | Main coding agent (Slack/Linear/Jira/Confluence/GitHub-triggered). |
| `reviewer` | `agent.graphs.reviewer:traced_reviewer_agent` (shim → `agent.reviewer:get_reviewer_agent` ) | Read-only PR reviewer. Findings model + `publish_review` . |
| `analyzer` | `agent.graphs.analyzer:traced_analyzer` (shim → `agent.analyzer:get_analyzer` ) | Learns per-repo reviewer style from historical PRs and this reviewer's own finding outcomes. |
| `chat` | `agent.graphs.chat:traced_chat_agent` | Dashboard Agents chat graph. |
| `scheduler` | `agent.graphs.scheduler:get_scheduler` | Reconcile sweep for stragglers. |
| `ci_monitor` | `agent.graphs.ci_monitor:get_ci_monitor` (shim → `agent.ci_monitor:get_ci_monitor` ) | Polling fallback for CI auto-fix: each tick sweeps open agent-authored PRs for failing checks / merge conflicts via `agent.ci_autofix.sweep_open_prs` . |
2026-05-22 14:33:46 -07:00
docs: land reorg in ledger, fix path refs, ship plan artifacts (C7)
C7 of the domain-reorg adoption (build plan docs/upstream-sync/domain-reorg/
reorg-build-plan.md, step C7):
- CLAUDE.md / AGENTS.md: retarget architecture path references to the new
layout — graph entrypoints via agent.graphs.* shims, the agent/api/ +
agent/webhooks/*_routes.py FastAPI split (webapp.py now a shim),
agent/review/ package, tests/<domain>/ test paths, and the new-graph/
test conventions. README.md already pointed at docs/ (C1) — no change.
- triage.jsonl: flip 8356eb34 (#1726) to landed on branch
refactor/domain-reorg-adoption; add re-triage notes to the 8 unblocked
rows (#1732/#1761/#1744/#1748/#1742 clean, #1736/#1758/#1760 near-clean).
triage.md regenerated via make triage-render.
- Ship the plan, scoping report, move-map artifacts, and the
domain-reorg-adoption workflow so the exercise is reproducible.
Memory + Confluence handled out-of-band (not in this commit): project
memory updated with the new module layout; Confluence check found no IT
page documents the repo module map — no Confluence change required.
2026-07-17 15:01:45 -04:00
The FastAPI app is `agent.webapp:app` — now a compatibility shim re-exporting `agent.api.app:app` .
2026-04-30 20:06:58 -07:00
feat: activate PR babysitting UI toggles for autofix and trigger mode (#1561)
* feat: activate PR babysitting UI toggles for autofix and trigger mode
Remove the "coming soon" gating on the Autofix Mode, Autofix Severity
Threshold, and Trigger Mode controls in the review settings page so
admins can enable CI auto-fix and review-comment resolution on PRs
that Open SWE opens. The backend (ci_autofix.py, webapp.py webhook
routing) was already fully wired — only the UI was disabled.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: simplify autofix to on/off toggle, remove severity threshold
Replace the four-level AutofixMode (off/low/medium/high) and the
autofix_severity_threshold setting with a single boolean
autofix_enabled toggle. The severity threshold was leftover from the
reviewer finding-severity model and does not apply to CI autofix;
the agent should fix any failing CI and resolve any comments on PRs
it opens.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: move autofix toggle to per-user profile, remove team-level setting
The autofix toggle is now per-user (auto_fix_ci in the user profile)
instead of team-level (admin-only). This uses the existing auto_fix_ci
field that was already in ProfileUpdate but never wired up.
Changes:
- ci_autofix.py: check per-user auto_fix_ci profile flag after
resolving the agent thread's github_login, instead of checking
team-level autofix_enabled before knowing the PR
- webapp.py: removed early is_autofix_enabled() webhook gates; the
per-user check now happens in ci_autofix.py once the thread is found
- team_settings.py: removed autofix_enabled field, is_autofix_enabled()
- cloud-agents.tsx: enabled the auto_fix_ci toggle (was comingSoon)
- review.tsx: removed the admin-level autofix switch
- Updated tests and AGENTS.md
The agent graph (not the reviewer) is what gets dispatched - this was
already correct in ci_autofix.py line 223: client.runs.create(
thread_id, "agent", ...).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: batch PR babysitting events
Remove the leftover trigger-mode gate from PR babysitting and batch new CI/review events while an agent run is already active so the running agent can handle the latest PR state before finishing. Also moves review-feedback permission checks behind the per-user opt-out and applies the auto-fix profile gate to merge-conflict babysitting.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: consume batched babysitting events
Teach the agent queue middleware to turn pending PR babysitting metadata into an injected instruction for the active run, so batched CI/review events are not dropped while still avoiding duplicate run creation.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review findings in PR babysitting batching
- Route batched events through the LangGraph store (read in-process by the
message-queue middleware) instead of a per-model-call threads.get on every
agent thread.
- Only record an attempt / mark the head SHA handled on a real dispatch, not
on a batch, so an event isn't permanently dropped if the in-flight run ends
before consuming it.
- Carry the reviewer's comment through batched review feedback instead of
replacing it with a generic re-check nudge.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 14:12:04 -07:00
CI auto-fix ("PR babysitting") lives in `agent/ci_autofix.py` : when a CI check fails (webhook `check_run` / `check_suite` / `workflow_run` / `status` ) or a reviewer leaves actionable feedback on a PR Open SWE opened, it locates the originating agent thread (by `pr_url` metadata) and dispatches a confidence-gated fix run on the `agent` graph. Gated by the per-user `auto_fix_ci` profile flag, the enabled-repos opt-in, and a per-PR `@open-swe autofix on|off` toggle (`agent/dashboard/autofix_state.py` ). Skip-rules (base-branch failures, human commits, same-head dedupe, batching while runs are active, loop cap) all live in `ci_autofix.py` .
2026-06-15 13:53:50 -07:00
2026-04-30 20:06:58 -07:00
## Architecture
2026-05-22 14:33:46 -07:00
### Entrypoints
2026-04-30 20:06:58 -07:00
2026-05-22 14:33:46 -07:00
- **`agent/server.py` → `get_agent(config)` ** — main graph factory. Called per-thread. Resolves the GitHub token, gets-or-creates the sandbox for the thread, resolves the team/profile/per-thread model + effort, then constructs a fresh `create_deep_agent(...)` with the curated tool list and middleware stack. The agent itself is stateless — all per-thread state lives in the sandbox + thread metadata.
docs: land reorg in ledger, fix path refs, ship plan artifacts (C7)
C7 of the domain-reorg adoption (build plan docs/upstream-sync/domain-reorg/
reorg-build-plan.md, step C7):
- CLAUDE.md / AGENTS.md: retarget architecture path references to the new
layout — graph entrypoints via agent.graphs.* shims, the agent/api/ +
agent/webhooks/*_routes.py FastAPI split (webapp.py now a shim),
agent/review/ package, tests/<domain>/ test paths, and the new-graph/
test conventions. README.md already pointed at docs/ (C1) — no change.
- triage.jsonl: flip 8356eb34 (#1726) to landed on branch
refactor/domain-reorg-adoption; add re-triage notes to the 8 unblocked
rows (#1732/#1761/#1744/#1748/#1742 clean, #1736/#1758/#1760 near-clean).
triage.md regenerated via make triage-render.
- Ship the plan, scoping report, move-map artifacts, and the
domain-reorg-adoption workflow so the exercise is reproducible.
Memory + Confluence handled out-of-band (not in this commit): project
memory updated with the new module layout; Confluence check found no IT
page documents the repo module map — no Confluence change required.
2026-07-17 15:01:45 -04:00
- **`agent/reviewer.py` → `get_reviewer_agent(config)` ** — reviewer graph factory. Shares `ensure_sandbox_for_thread` with the main agent but wires a reviewer-only toolset (`add_finding` , `update_finding` , `list_findings` , `publish_review` , `web_search` , `fetch_url` , `http_request` ) and a different system prompt that pins the single-evolving-findings model and the diff-anchored bar for filing a finding. Read-only: no commit/push/PR-opening tools. Its supporting modules live in the `agent/review/` package (domain reorg): `findings.py` , `publish.py` , `reconcile.py` , `trace_context.py` , `diff.py` , `groups.py` , `eval_store.py` , `style_collector.py` , `style_guidance.py` .
feat: outcomes dataset + bootstrap/continual split via skills (#1365)
* fix: reset stale sandbox creation sentinel
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* fix: treat SANDBOX_CREATING as a timestamped cross-process lock
Only reset the sentinel when proven stale (older than the creation
timeout); otherwise wait for the worker that holds the lock so a
concurrent run does not create a duplicate sandbox.
* feat(analyzer): outcomes dataset + bootstrap/continual split via skills
Rename the review_style_analyzer graph to `analyzer` and split it into two
modes, plus capture reviewer finding outcomes for continual learning.
- Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false
positive), and GitHub/Slack thumbs findings into a single LangSmith dataset
(openswe-reviewer-outcomes), keyed deterministically per finding+source.
Emit points wired into update_finding, resolve_finding_thread, and the
GitHub/Slack reaction handlers.
- Two playbooks delivered as deepagents skills (bootstrap-repo-analysis,
continual-learning), served as virtual files via a CompositeBackend /skills/
route + StateBackend (seeded into the run files channel at invoke time, never
written to the sandbox). Mode is set by the launcher; continual runs fall
back to the GitHub App installation token.
- Split launcher into start_bootstrap_analysis + start_continual_run; register
a per-repo nightly continual-learning cron when bootstrap completes.
- New read_finding_outcomes tool feeds confirmed/dismissed findings back to the
continual playbook.
Tests for outcome label mapping, skills helper, and cron idempotency.
* fix(analyzer): anchor continual cron runs to a real thread_id
The nightly continual-learning cron is threadless, and get_analyzer
early-returns an empty agent when configurable.thread_id is missing — so
every cron-launched run no-op'd before reading outcomes or saving a refined
prompt. Include the repo's deterministic analyzer thread_id in the continual
run configurable so the run executes; the threadless run carries no message
history, so nightly runs don't accumulate context.
* refactor(analyzer): move cron lifecycle calls out of the review-styles store
Drop the inline `analyzer_cron` imports from review_styles.py (added only to
dodge a circular import) by relocating the cron-trigger calls to the layer
above the store: registration to the save_review_style tool (after a prompt is
saved) and removal to the dashboard delete route. review_styles.py is now a
pure store again with top-level imports only.
* refactor: hoist reviewer_outcomes imports to module level
Move the two inline emit_finding_status_outcome imports introduced in this PR
(update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes
only depends on langsmith, so there is no circular import to avoid.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
- **`agent/analyzer.py` → `get_analyzer(config)` ** — small graph that emits a per-repo style prompt via the `save_review_style_prompt` tool, consumed by the reviewer as a "repository-specific review style" appendix. It runs in one of two modes (`analyzer_mode` in `configurable` ): **bootstrap** (cold-start: crawl historical PR reviews) and **continual** (nightly: refine using this reviewer's own finding outcomes via `read_finding_outcomes` ). Each mode's procedure lives in a deepagents **skill** (`agent/skills/bootstrap-repo-analysis/` , `agent/skills/continual-learning/` ) served as virtual files via a `CompositeBackend` `/skills/` route + `StateBackend` (seeded into the run's `files` channel by the launcher — never written to the sandbox). Launchers and the per-repo nightly cron live in `agent/dashboard/review_style_jobs.py` and `agent/dashboard/analyzer_cron.py` ; the cron is registered when bootstrap completes.
docs: land reorg in ledger, fix path refs, ship plan artifacts (C7)
C7 of the domain-reorg adoption (build plan docs/upstream-sync/domain-reorg/
reorg-build-plan.md, step C7):
- CLAUDE.md / AGENTS.md: retarget architecture path references to the new
layout — graph entrypoints via agent.graphs.* shims, the agent/api/ +
agent/webhooks/*_routes.py FastAPI split (webapp.py now a shim),
agent/review/ package, tests/<domain>/ test paths, and the new-graph/
test conventions. README.md already pointed at docs/ (C1) — no change.
- triage.jsonl: flip 8356eb34 (#1726) to landed on branch
refactor/domain-reorg-adoption; add re-triage notes to the 8 unblocked
rows (#1732/#1761/#1744/#1748/#1742 clean, #1736/#1758/#1760 near-clean).
triage.md regenerated via make triage-render.
- Ship the plan, scoping report, move-map artifacts, and the
domain-reorg-adoption workflow so the exercise is reproducible.
Memory + Confluence handled out-of-band (not in this commit): project
memory updated with the new module layout; Confluence check found no IT
page documents the repo module map — no Confluence change required.
2026-07-17 15:01:45 -04:00
- **FastAPI layer (`agent/api/` + `agent/webhooks/` )** — custom FastAPI routes mounted alongside the LangGraph server. The domain reorg split the fork's former `agent/webapp.py` monolith into `agent/api/app.py` (app composition + router mounts), `agent/api/health.py` (`/health` , `/webhooks/run-complete` ), shared helpers in `agent/webhooks/common.py` , and per-source route modules `agent/webhooks/{github,linear,slack,jira,confluence}_routes.py` ; `agent/webapp.py` remains a compatibility shim re-exporting `app` . Webhooks land in those route modules (GitHub, Linear, Slack, Jira, and the Confluence Atlassian Connect `/connect/*` routes, which fold into `confluence_routes.py` ). Each webhook resolves a deterministic `thread_id` (so follow-up messages route to the same agent run) and triggers/streams a run via the `langgraph_sdk` client. Also auto-reviews PRs on `opened` / `ready_for_review` events when the repo+author opt in. The Atlassian triggers (`agent/webhooks/{jira,confluence}.py` , `agent/utils/atlassian_connect.py` ) verify webhook trust — a Jira Automation shared secret / optional HMAC, and a Confluence Connect HS256-JWT + `qsh` with RS256 signed-install — then re-fetch the triggering comment server-side before deriving identity.
2026-05-22 14:33:46 -07:00
- **`agent/dashboard/` ** — `router` mounted under the FastAPI app at startup (`app.include_router(dashboard_router)` ). Owns GitHub OAuth, per-user profiles, admin endpoints, team defaults, enabled-repo lists, review-style management, and the Agents chat thread API used by the UI in `ui/` .
2026-04-30 20:06:58 -07:00
### Sandbox lifecycle (the tricky part)
2026-05-22 14:33:46 -07:00
`SANDBOX_BACKENDS` (in `agent/utils/sandbox_state.py` ) is an in-process dict keyed by `thread_id` . Thread metadata persists `sandbox_id` across processes. `ensure_sandbox_for_thread` handles four cases:
2026-04-30 20:06:58 -07:00
2026-05-22 14:33:46 -07:00
1. Sandbox cached in memory → ping it (`echo ok` ); recreate on `SandboxClientError` . Healthy reused sandboxes also get a GitHub-proxy refresh (recreate on failure).
feat: outcomes dataset + bootstrap/continual split via skills (#1365)
* fix: reset stale sandbox creation sentinel
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* fix: treat SANDBOX_CREATING as a timestamped cross-process lock
Only reset the sentinel when proven stale (older than the creation
timeout); otherwise wait for the worker that holds the lock so a
concurrent run does not create a duplicate sandbox.
* feat(analyzer): outcomes dataset + bootstrap/continual split via skills
Rename the review_style_analyzer graph to `analyzer` and split it into two
modes, plus capture reviewer finding outcomes for continual learning.
- Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false
positive), and GitHub/Slack thumbs findings into a single LangSmith dataset
(openswe-reviewer-outcomes), keyed deterministically per finding+source.
Emit points wired into update_finding, resolve_finding_thread, and the
GitHub/Slack reaction handlers.
- Two playbooks delivered as deepagents skills (bootstrap-repo-analysis,
continual-learning), served as virtual files via a CompositeBackend /skills/
route + StateBackend (seeded into the run files channel at invoke time, never
written to the sandbox). Mode is set by the launcher; continual runs fall
back to the GitHub App installation token.
- Split launcher into start_bootstrap_analysis + start_continual_run; register
a per-repo nightly continual-learning cron when bootstrap completes.
- New read_finding_outcomes tool feeds confirmed/dismissed findings back to the
continual playbook.
Tests for outcome label mapping, skills helper, and cron idempotency.
* fix(analyzer): anchor continual cron runs to a real thread_id
The nightly continual-learning cron is threadless, and get_analyzer
early-returns an empty agent when configurable.thread_id is missing — so
every cron-launched run no-op'd before reading outcomes or saving a refined
prompt. Include the repo's deterministic analyzer thread_id in the continual
run configurable so the run executes; the threadless run carries no message
history, so nightly runs don't accumulate context.
* refactor(analyzer): move cron lifecycle calls out of the review-styles store
Drop the inline `analyzer_cron` imports from review_styles.py (added only to
dodge a circular import) by relocating the cron-trigger calls to the layer
above the store: registration to the save_review_style tool (after a prompt is
saved) and removal to the dashboard delete route. review_styles.py is now a
pure store again with top-level imports only.
* refactor: hoist reviewer_outcomes imports to module level
Move the two inline emit_finding_status_outcome imports introduced in this PR
(update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes
only depends on langsmith, so there is no circular import to avoid.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
2. Metadata says `__creating__` and no cache → reset stale metadata so a fresh sandbox can be created.
2026-05-22 14:33:46 -07:00
3. No sandbox at all → set `__creating__` sentinel, create one, persist the real id.
2026-04-30 20:06:58 -07:00
4. Metadata has an id but no cache → reconnect; fall back to recreate on failure.
2026-05-22 14:33:46 -07:00
For `SANDBOX_TYPE=langsmith` (default), every sandbox creation/refresh also calls `_configure_github_proxy` with a fresh GitHub App installation token (`get_github_app_installation_token` ). The proxy injects Basic auth for `github.com` git traffic and Bearer auth for `api.github.com` so sandbox commands can use `GH_TOKEN=dummy gh ...` without storing real tokens in the sandbox. Other providers (modal, daytona, runloop, local) skip the proxy step. Provider is selected via `SANDBOX_TYPE` ; factory is `agent/utils/sandbox.py:create_sandbox` (`SANDBOX_FACTORIES` maps each provider name to a creator in `agent/integrations/` ).
Every run re-applies `git config --global user.name/email` for the bot identity, because reused/reconnected sandboxes can lose `--global` config and Vercel preview deploys reject commits whose author email doesn't resolve to a GitHub account.
2026-04-30 20:06:58 -07:00
### Middleware stack (order matters)
2026-05-22 14:33:46 -07:00
Configured in `agent/server.py:get_agent` , runs around every model call (in this order):
1. `SanitizeToolInputsMiddleware` — strips/normalizes tool inputs before they reach tools.
2. `ModelCallLimitMiddleware` (from `langchain.agents.middleware` ) — caps model calls at `MODEL_CALL_RECURSION_LIMIT` (~half of `DEFAULT_RECURSION_LIMIT` ); `exit_behavior="end"` .
3. `ToolErrorMiddleware` — catches tool exceptions and surfaces them as tool messages.
2026-07-08 18:52:01 -04:00
4. `SubdirAgentsReadMiddleware` — appends applicable ancestor `AGENTS.md` instructions to `read_file` results once per run, so scoped rules are visible before edits.
5. `check_message_queue_before_model` — pulls Linear comments / Slack messages that arrived mid-run from the thread queue and injects them as user messages before the next LLM call. This is what makes "message the agent while it's working" work.
6. `SlackAssistantStatusMiddleware` — keeps the Slack "assistant is typing"-style status up to date around model calls.
7. `ensure_no_empty_msg` — after-model hook; when the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion) it re-injects a synthetic `no_op` / `confirming_completion` tool call so the run continues instead of ending prematurely.
8. `notify_step_limit_reached` — after-agent hook that posts a Slack reply when the agent hits the step limit, so the user gets a clear signal instead of silence.
9. `SandboxCircuitBreakerMiddleware` — trips the agent out of repeated sandbox failures instead of looping.
10. `ModelFallbackMiddleware` (optional) — added only when `LLM_FALLBACK_MODEL_ID` or the per-model default fallback differs from the primary model.
11. `SanitizeThinkingBlocksMiddleware` — strips malformed empty Anthropic thinking blocks immediately before provider calls.
2026-06-15 14:54:01 -07:00
chore: sync upstream/main, defer #1621 modular webhooks (#81)
* chore: bake sfw binary into sandbox image (#1611)
sfw only ships a launcher that fetches its real binary at first run and does a
daily update check against api.github.com/repos/SocketDev/sfw-free. Both fail in
the sandbox (restricted egress; the proxy injects the GitHub App installation
token, which lacks access to that repo), so `sfw yarn install` errors with
"could not fetch its binary". Pin sfw 2.0.6, warm + verify the binary cache at
build, and set SFW_SKIP_UPDATE_CHECK=1 so runs use the baked binary offline.
* feat: editable plan mode + fix review-plan banner overlap (#1610)
* feat: editable plan mode + fix review-plan banner overlap
Lets the thread owner edit the plan markdown by hand from the plan-review
page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id}
endpoint that re-publishes the plan and mirrors it into the sandbox
plan.md, so approve hands the edited plan to the agent as the source of
truth. Also fixes the collapsed git-panel's floating expand button
covering the "Review plan ->" banner by reserving space for it.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: abort plan approval when the published plan read fails
get_plan_content() swallowed store errors and returned None, so a
transient failure during approve would still mark the plan approved and
dispatch the generic fallback text — silently dropping an owner's edited
plan. Read the plan strictly (raise_on_error=True) so approval aborts
instead, matching the comment read.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: show message timestamps (#1609)
* feat: show message timestamps
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: suppress fallback message timestamps
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: stable message + tool-call hover timestamps
Stamp a stable client-side arrival time per message and tool call (keyed
by id, persisted to localStorage). Messages render the timestamp inline;
tool rows reveal a dim timestamp chip on hover. Real backend created_at
still takes precedence when present.
* fix: hide client-stamped message timestamps
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add PR trace resolution (#1612)
* feat: add PR trace resolution
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: inject reviewer trace context as JSON
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review on PR trace resolution
Use the documented LangSmith metadata filter syntax
(and(eq(metadata_key,...), eq(metadata_value,...))) instead of
has(metadata, '{...}'), which does not match runs — _list_thread_runs
was silently returning nothing. Bound full-text searches to a 90-day
window so they don't hit LangSmith's large-window rate limit.
Also folds in the best-effort branch->head-sha resolver (dropping the
weighted scoring/threshold + repo/file evidence + GitHub hydration),
sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint.
The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session
were removed; resolution now runs deterministically from the trusted run
config with no model-controlled pr_url or thread_id.
* fix: scope branch trace search to the repo
Branch names like fix-tests aren't unique across repos (or older PRs) in
a shared tracing project, so an unscoped branch hit could resolve to an
unrelated thread and write its runs into the reviewer sandbox. Require
the repo slug to co-occur with the branch in matched runs; the full head
SHA stays unscoped since it is globally unique. Addresses open-swe review
on PR #1612.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: include plan links in PR descriptions (#1613)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: gate workflow pushes with approval (#1614)
* feat: gate workflow pushes with approval
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: preserve proxy refresh test compatibility
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: bind workflow approvals to pushed ref
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: recover thread work as patch (#1615)
* feat: recover thread work as patch
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: search sandbox cwd for recovery patches
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: omit plan link in PR description when no plan exists (#1618)
Plan links in PR descriptions were always built from the thread id, so
runs that never produced a plan linked to an empty plan-review page.
Now the plan content store is consulted first; the link is only added
when a plan with non-empty markdown actually exists. A transient store
failure degrades gracefully (no link) rather than blocking PR creation.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add filter & grouping menu to agents threads sidebar (#1617)
Add a Cursor-style control to the agents sidebar that groups (None/Date/
Status/Project), filters (ownership, status, source, pull request, model,
repo, include-resolved), and compacts the threads list. All client-side over
already-fetched sidebar threads; preferences persist in localStorage.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: update langsmith sdk to 0.9.3 (#1616)
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: clickable shared PR header in git panel and reviews (#1620)
* feat: clickable shared PR header in git panel and reviews
Replace the standalone "View PR" button in the agent git panel with a
clickable PR title, matching the reviews view. Extract a shared PrHeader
component reused by both the git panel and the review main body.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor: drop PrHeader wrapper, use shared component directly
The review-side PrHeader was just a thin adapter mapping detail -> the
shared component's props. Inline it at the call site and use the shared
PrHeader directly so there's a single component.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor: durable interrupt dispatch + completion webhook (#1621)
* wip(rebuild): core reliability spine
- remove PR-babysitting (ci_autofix + ci_monitor graph + webhook wiring)
- dispatch core: agent/dispatch.py with multitask_strategy=interrupt +
durability=sync + completion webhook; reroute all webhook + plan triggers;
drop the racy in-process lock + is_thread_active busy-check
- completion webhook: agent/completion.py + /webhooks/run-complete loopback
route for failure/timeout replies (idempotent)
Co-authored-by: open-swe[bot]
* feat(rebuild): async tools, reconcile, shared http timeouts, assembly tuning
Parallel batch on top of the reliability spine:
- async-ify all 24 tools (drop asyncio.run; requests->httpx); re-implement the
http_request/fetch_url SSRF + DNS-rebinding defense httpx-natively and harden
the IP check to 'not is_global' (+ IPv4-mapped unwrap)
- reconcile.py: stale pending-run sweep (threads.search -> per-thread runs.list
-> cancel_many), wired into the scheduler graph via task='reconcile'
- shared DEFAULT_HTTP_TIMEOUT (agent/utils/http.py) on every bare
httpx.AsyncClient() across utils/dashboard/webapp/middleware
- run budget: MODEL_CALL_RECURSION_LIMIT 5000->250
- fix stale OpenAI->Anthropic fallback id (claude-opus-4-5 -> 4-8)
- drop redundant custom repair middleware (deepagents auto-adds PatchToolCalls)
- confirm tool-result eviction + summarization auto-wired via backend
- slim system prompt ~8% (full harness-profile rewrite deferred)
Co-authored-by: open-swe[bot]
* feat(rebuild): harness-profile prompt + split webhooks out of webapp
- prompt.py: own the system prompt via a registered harness profile
(OPEN_SWE_SHARED_BASE, kept neutral so the read-only reviewer/analyzer that
share it stay safe), registered across all 4 providers; per-thread values
stay in construct_system_prompt. Assembled main-agent prompt ~6.8k -> ~3.1k
tokens (~55% smaller); de-duped PR/commit/suite/force-push guidance; dropped
ALL-CAPS markers.
- webapp.py 3325 -> 1890 LOC: moved 14 per-source handlers into
agent/webhooks/{linear,slack,github}.py; webapp re-exports them for the
routes + tests; moved handlers reach shared helpers via the webapp namespace
to preserve the test suite's monkeypatch targets.
Full suite: 1168 passing, lint clean.
Co-authored-by: open-swe[bot]
* Restore MODEL_CALL_RECURSION_LIMIT to 5000 for long-running tasks
Reverts the 250 cap from the run-budget change — long-running tasks legitimately
need many model calls. The notify_step_limit_reached safety net still fires if a
run does hit the cap, so runs end with a signal either way.
Co-authored-by: open-swe[bot]
* fix: address PR review (auth, SSRF, interrupted status, redirect headers)
- completion.py: drop `interrupted` from failure statuses — with
multitask_strategy=interrupt a follow-up ends the prior run as interrupted,
which is healthy, not a failure to report. [open-swe]
- /webhooks/run-complete: shared-secret auth — dispatch appends ?token= when
RUN_COMPLETE_WEBHOOK_SECRET is set; route verifies via hmac.compare_digest.
[corridor-security]
- SSRF: extract the URL validator to agent/utils/url_safety.py and apply it
before server-side image fetches in multimodal.fetch_image_block.
[corridor-security]
- http_request: preserve caller headers/extensions across redirect hops instead
of dropping them on the first hop. [open-swe]
Co-authored-by: open-swe[bot]
* chore: remove REBUILD_PLAN.md (planning doc, not needed in the repo)
Co-authored-by: open-swe[bot]
* fix: fail closed on run-complete webhook auth when secret unset
Corridor follow-up: verify_run_complete_token returns False (not True) when
RUN_COMPLETE_WEBHOOK_SECRET is unset, so the public route is never
unauthenticated. Logs a startup warning when the secret is absent, and dispatch
skips registering the webhook when there's no secret (no rejected callbacks).
Co-authored-by: open-swe[bot]
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: restore forced tool call to prevent premature run stops (#1622)
Restore the ensure_no_empty_msg middleware and the always-call-a-tool system-prompt instruction that #1535 removed. When the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion), the middleware re-injects a no_op / confirming_completion tool call so the run continues instead of ending mid-task.
Shipping to test whether it fixes runs that stop halfway through.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore(deps): bump langgraph-checkpoint from 4.1.0 to 4.1.1 (#1619)
Bumps [langgraph-checkpoint](https://github.com/langchain-ai/langgraph) from 4.1.0 to 4.1.1.
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1)
---
updated-dependencies:
- dependency-name: langgraph-checkpoint
dependency-version: 4.1.1
dependency-type: indirect
...
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
* fix: post reviewer resolution notes verbatim (#1624)
* fix: post reviewer resolution notes verbatim
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: stabilize dashboard follow-up e2e
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: preserve dashboard attribution in e2e
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: make e2e attribution marker durable
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: only echo found e2e attribution
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: check live dashboard attribution in e2e
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* hotfix: stop prompting agent/reviewer to wrap installs in sfw (#1625)
Installs hung when prefixed with sfw inside the sandbox (trace 019f0608
stalled on a pending `sfw npm install` execute, never returned). Strip the
Socket Firewall guidance from the agent and reviewer prompts so installs run
through the project's package manager directly. sfw stays in the Docker image;
nothing invokes it now.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: make plan view mobile friendly (#1636)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: fall back to vision model for image threads (#1626)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: surface Slack thread errors (#1627)
* fix: surface Slack thread errors
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: don't set failure_reply_posted on Slack preprocessing errors
The preprocessing error handler was setting failure_reply_posted=True,
the same idempotency flag handle_run_completion checks to suppress
duplicate run-failure replies. Since preprocessing failures happen
before any run exists but the flag persists on the thread, a subsequent
run failure on the same thread would be silently ignored.
The preprocessing handler already posts its own Slack reply, so the
run-completion idempotency flag should not be set here.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: avoid recapping Slack replies (#1629)
* chore: avoid recapping Slack replies
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: simplify Slack reply prompt wording
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: update Slack trace reply on web handoff (#1630)
* fix: update Slack trace reply on web handoff
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: trigger web handoff on dashboard starts
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: format web handoff as contextual fragment
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: preserve trace_message_ts when overwriting Slack run mapping
When store_slack_run_mapping is called without trace_message_ts (e.g. on
follow-up Slack mentions), it was unconditionally overwriting the
thread-level mapping and clobbering the timestamp captured from the
initial trace reply. After that, _notify_slack_web_handoff could not find
the original message, so a subsequent move to Web silently skipped the
Slack trace update.
Now, when trace_message_ts is not passed, the existing thread mapping is
read first and its trace_message_ts is preserved.
* style: ruff format
---------
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* fix: pre-bundle shiki/@pierre deps to stop dev dynamic-import failures (#1643)
* fix(ui): pre-bundle shiki/@pierre deps to stop dev dynamic-import failures
shiki lazy-imports a grammar per language and these libs only live inside
lazy route components, so Vite's startup scanner never sees them. They get
discovered on first thread navigation, triggering a dep re-optimize +
force-reload that aborts the in-flight route-chunk import, surfacing as
"Failed to fetch dynamically imported module: .../$threadId.tsx".
Pre-bundle them (and the github themes + common code-block languages) via
optimizeDeps.include so the optimize happens once at startup. Dev-only;
production bundles are unaffected.
* fix: pre-bundle canonical shiki docker/make langs instead of aliases
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: show queued dashboard follow-ups (#1631)
* feat: show queued dashboard follow-ups
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: de-dupe queued follow-ups while streaming
---------
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* feat: notify Slack on plan approval (#1632)
* feat: notify Slack on plan approval
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: post Slack approval notice after successful dispatch
Move the _maybe_post_plan_approved_to_slack call until after
_dispatch_followup succeeds so the Slack thread is not told
implementation is beginning before the LangGraph run is created.
Addresses PR review comment.
---------
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* feat: include Slack channel context in prompts (#1633)
Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
* chore: keep plan guidance high-level (#1634)
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: publish plans from sandbox files (#1635)
* feat: publish plans from sandbox files
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: avoid fixed plan filenames
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: virtualize local sandbox file paths
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: preserve plan_file_path across set_plan_status
set_plan_status was rewriting the content record with only markdown
and status, dropping plan_file_path. After a reject, the owner's
dashboard edit would mirror to a different file than the agent's
original, and the next save_plan could republish the stale file.
Preserve plan_file_path when updating status.
---------
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: return to thread after plan approval (#1637)
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add Slack breakout thread tool (#1638)
* feat: add Slack breakout thread tool
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: make fake LLM scripts declarative
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: exclude slack_start_new_thread from plan mode
The breakout tool can dispatch a fresh agent run that starts outside the
current plan-mode state, bypassing the approval flow. Add it to
PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating
tools while planning.
---------
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: require bun for ui agent work (#1639)
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: request actions read for sandbox logs (#1642)
* fix: request actions read for sandbox logs
Request optional Actions read permission for sandbox proxy tokens, with fallback for installations that have not approved it yet. Update setup docs and prompt guidance for safe GitHub Actions log usage.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: restore actions:read scope after workflow push
After an approved workflow push, the guard was restoring the proxy with
BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS, which excludes the actions: read
scope this PR adds. Restore with RUNTIME_PROXY_TOKEN_PERMISSIONS (which
includes actions: read) and fall back to BASE if the install hasn't
granted Actions read — mirroring the pattern in _create_sandbox_with_proxy.
Addresses review comment on PR #1642.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: widen split review diffs (#1647)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: install missing deps before verification (#1646)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: switch ui to pnpm (#1645)
* chore: require pnpm for ui agent work
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: switch ui to pnpm
Replace Bun and Yarn lockfiles with pnpm lockfile and update UI/Vercel commands to use pnpm.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* ci: use corepack for ui pnpm e2e build
Run pnpm through Corepack in the E2E global setup so CI can use the pinned package manager without a separate pnpm install step.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add Sonnet 5 to model picker (#1651)
* chore: update Sonnet examples to Sonnet 5
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: add Sonnet 5 to model picker
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* Remove dead breakout-thread e2e scenario after dropping the tool
The merge resolution deferred upstream's Slack breakout-thread tool
(slack_start_new_thread, #1638) since it depends on the #1621 dispatch
module, but the e2e harness still scripted it. Removing the tool name
from fake_llm.py's _tool_step call left a malformed scenario, crashing
the langgraph-dev web server at import (TypeError: _tool_step() missing
'call_id') and failing Playwright E2E.
Drop the "breakout" script scenario, its _is_breakout_request helper +
ScriptRule, and the corresponding full_flow.spec.ts test.
* Revert upstream pnpm switch; keep bun for the UI build
The merge auto-adopted upstream's pnpm switch (#1645) in tests/e2e/
global-setup.ts and ui/package.json, but our fork builds the UI with
bun (vercel.json + the E2E workflow's setup-bun). That left the
Playwright globalSetup running `corepack pnpm install --frozen-lockfile`
with no pnpm-lock.yaml, failing E2E at UI build time.
Revert global-setup.ts and ui/package.json to the dev (bun) baseline,
drop the merge-added ui/pnpm-lock.yaml, and remove the re-added
ui/AGENTS.md (our fork had deleted it).
* Align plan-review e2e + UI with the HEAD (pre-#1635) backend
The merge left a split plan vertical: the backend save_plan/plan_api are
HEAD (we deferred the editable-plan/sandbox-publish features #1610/#1635/
#1637 per #80), but the plan UI and e2e harness were upstream's. The
fake_llm scenario called save_plan(plan_file_path=...) — upstream's
file-based #1635 contract — while HEAD save_plan takes plan_markdown,
so the plan never saved and PlanReview never rendered (E2E failure on
the plan-review locator).
Pass plan_markdown to save_plan, and revert PlanReview.tsx / plan.ts /
$threadId_.plan.tsx / plan_review.spec.ts to the dev baseline so the
whole plan flow (save -> render -> approve -> implement) is consistent
with the HEAD backend.
---------
Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Ankush Gola <9536492+agola11@users.noreply.github.com>
Co-authored-by: Mukil Loganathan <mukil@langchain.dev>
2026-06-30 16:45:19 -04:00
The system prompt instructs the agent to call a tool every turn, and `ensure_no_empty_msg` re-injects a tool call when it doesn't — together these keep runs from stopping partway through a task.
2026-04-30 20:06:58 -07:00
2026-05-28 16:15:13 -07:00
Other middleware exists in `agent/middleware/` (`ExcludeToolsMiddleware` ) but isn't wired into the default agent. The reviewer uses a leaner stack: `SanitizeToolInputsMiddleware` , `ModelCallLimitMiddleware` , `ToolErrorMiddleware` , `SlackAssistantStatusMiddleware` , `SanitizeThinkingBlocksMiddleware` .
2026-05-04 18:03:53 -07:00
There is intentionally no after-agent safety net that opens a PR for the agent. The agent itself is responsible for committing, pushing, opening/updating the draft PR, and replying in the source channel — all via `GH_TOKEN=dummy gh` and `slack_thread_reply` / `linear_comment` .
2026-04-30 20:06:58 -07:00
### Tools
2026-05-22 14:33:46 -07:00
All tools live in `agent/tools/` and are flat-imported via `agent/tools/__init__.py` . The set is intentionally small and curated — see README "Tools — Curated, Not Accumulated".
Wired into `get_agent` :
2026-07-17 16:22:38 -04:00
`http_request` , `fetch_url` , `web_search` , `linear_comment` , `linear_create_issue` , `linear_delete_issue` , `linear_get_issue` , `linear_get_issue_comments` , `linear_list_teams` , `linear_search_issues` , `linear_update_issue` , `jira_comment` , `jira_create_issue` , `jira_get_issue` , `jira_get_issue_comments` , `jira_list_projects` , `jira_update_issue` , `confluence_get_page` , `confluence_create_page` , `confluence_update_page` , `confluence_comment` , `confluence_search` , `request_pr_review` , `schedule_thread_wakeup` , `slack_add_reaction` , `slack_read_thread_messages` , `slack_thread_reply` .
2026-05-22 14:33:46 -07:00
Reviewer-only tools (in `agent/reviewer.py` ): `add_finding` , `update_finding` , `list_findings` , `publish_review` . The review-style analyzer uses `save_review_style` (exported as `save_review_style_prompt` ).
Built-in deepagents tools (`read_file` , `write_file` , `edit_file` , `ls` , `glob` , `grep` , `execute` , `write_todos` , `task` for subagent spawning, …) are added by `create_deep_agent` itself; don't duplicate them.
### Models, profiles, and team defaults
Model + reasoning effort are resolved per run in this precedence (highest wins):
1. Per-thread config (`agent_model_id` + `agent_effort` in `configurable` ) — set by webhooks/UI.
2. Per-user dashboard profile override (`agent/dashboard/agent_overrides.py:load_profile` ), keyed by resolved GitHub login.
3. Team default model (`agent/dashboard/team_settings.py:get_team_default_model("agent")` ).
2026-05-26 13:30:18 -07:00
Supported model IDs and per-model effort/reasoning rules live in `agent/dashboard/options.py` . Profile flags also drive run behavior — e.g. `profile_create_prs` enables the opt-in Always Create PRs policy. Model construction goes through `agent/utils/model.py` (`make_model` , `provider_model_kwargs` , `fallback_model_id_for` ).
2026-04-30 20:06:58 -07:00
### Auth
2026-06-04 09:33:51 -07:00
- **GitHub**: dual-mode. User OAuth tokens are encrypted at rest in the dashboard OAuth store and cached only in process during a run (`utils/auth.py:resolve_github_token` , `utils/github_token.py` ). When no user token is available, falls back to a GitHub App installation token (`utils/github_app.py` ). The installation token is also what configures the LangSmith sandbox's GitHub proxy.
feat: Jira + Confluence integration (tools + triggers) (#182)
* feat(open-swe): add Jira tool plane (Phase 1)
Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear
tools:
- utils/jira.py: service-account REST client (Basic auth) with get/
create/update issue, comments, list projects, trace comment; issue and
comment bodies normalized to markdown.
- utils/adf.py: minimal ADF <-> markdown conversion (read paths convert
Jira ADF to markdown; agent comments convert prose to ADF).
- tools/jira_{comment,get_issue,get_issue_comments,create_issue,
update_issue,list_projects}.py wired into the tool registry and the
main agent tool list.
- tests/test_jira_utils.py: ADF conversion + mocked-transport util tests.
Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env
returns a clean error, so this is safe to land dark. Trigger plane,
prompt guidance, and config plumbing follow in Phase 2.
* feat(open-swe): add Confluence tool plane (Phase 3)
Curated Confluence Cloud REST toolset for the agent, mirroring the Jira
tools:
- utils/confluence.py: service-account REST client (Basic auth) with
get/create/update page, add comment, CQL search. Page bodies are XHTML
storage format (not ADF), with minimal storage<->text converters;
update_page reads the current version and bumps it, as Confluence
requires.
- tools/confluence_{get_page,create_page,update_page,comment,search}.py
registered in the tool registry.
- tests/test_confluence_utils.py: converter + mocked-transport tests
including the version-bump path.
Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN;
unset env returns a clean error. Activation in the agent tool list lands
with the Phase 2 server.py wiring.
* feat(open-swe): add Jira trigger plane (Phase 2)
Make an @openswe comment on a Jira issue spawn an agent run, mirroring
the Linear trigger plane:
- webhooks/jira.py: process_jira_issue clones process_linear_issue —
deterministic thread id, full-issue fetch, actor accountId->email
attribution feeding resolve_login_from_email_async (PRs open as the
human), multimodal image handling, source="jira" + jira_issue config.
- webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time
X-Automation-Webhook-Token check, fails closed), repo-resolution
cascade, get_repo_config_from_jira_mapping.
- utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder
entry — real project->repo mappings still needed).
- utils/jira.py: get_user_email (accountId -> email) for attribution.
- completion.py: source=="jira" failure-reply branch.
- prompt.py: Jira-triggered notify guidance + Refs:/branch key from
{jira_project_key}-{jira_issue_number}.
- server.py: read jira_issue config + pass jira key to the system
prompt; also activates the Phase 3 Confluence tools in the agent list.
Jira Automation lacks native webhook HMAC signing, so trust is a shared
secret header (decision D2); replay protection is weaker than Linear's
HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the
outstanding gate/hardening before push.
* fix(open-swe): harden Jira webhook trust (sh-security-review)
Resolves findings from the Phase 2 security review (detector fan-out +
proof-or-kill verifier). The unsigned Jira Automation webhook body was
trusted for identity, comment content, repo routing, and issue
existence; a JIRA_WEBHOOK_SECRET holder could forge those fields.
- Corroborate against the real Jira record: the webhook body is now only
a pointer (issue_key + required comment_id). The triggering comment's
author and text are re-fetched server-side via get_comment/fetch_jira_
comment, and identity, the @openswe check, prompt text, and project
key are derived from that authoritative record — never payload author/
body fields. An uncorroborated comment is rejected. (closes the
account-id impersonation, unsigned-body prompt injection, and
fabricated-issue findings)
- Validate issue_key against the Jira key format and percent-encode all
untrusted path segments (_seg) so a crafted key can't traverse to a
different Jira REST endpoint or inject query params. (closes the path-
traversal / query-injection findings)
- Route source=="jira" through the bot-token-default / author_prs_as_
user opt-in path in resolve_github_token, matching Linear, instead of
unconditionally resolving a per-user OAuth token from a payload email.
- Gate attribution on an active user mapping (is_login_mapped) so a
pending/unconfirmed mapping can't drive PR authorship.
Adds regression tests: server-corroboration wins over payload, malformed
issue_key rejected, uncorroborated comment rejected, path-segment
encoding, project-key derivation, active-mapping gate.
Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/
REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body +
timestamp on the Automation payload to close the residual replay gap.
* harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist
Folds the two deployment-hardening items from the Phase 2 security review
into code (all opt-in / default-off, so existing and upstream deployments
are unaffected):
- JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must
carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by
JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by
verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear
HMAC+freshness model). Closes the static-token model's replay/forgery
gap when enabled.
- JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's
direct client IP (verify_jira_source_ip). Documented as direct-peer
only; behind a proxy/LB, allowlist Atlassian's ranges at that layer.
- REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail
CLOSED instead of the back-compat allow-all, plus a startup fail-open
warning. Applies to all channels for consistency.
Documents all new vars (and a Jira section) in .env.example. Adds tests
for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and
the fail-closed allowlist.
* feat(open-swe): Confluence Atlassian Connect trigger (Phase 4)
Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian
Connect app. Designed and adversarially verified with the ultracode
workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus
skeptics on the implemented crypto).
- utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's
official test vector), PyJWT HS256 webhook verifier with alg-pinning,
issuer binding, and qsh-verified-last ordering; RS256 signed-install
lifecycle verifier against Atlassian's published keys; installation
store keyed by clientKey with the sharedSecret encrypted at rest
(TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned).
- webhooks/confluence.py: install/uninstall lifecycle + comment handler.
The JWT-signed webhook body is only a pointer; the comment's real
author/text/container are re-fetched server-side via the Basic-auth
service account (Phase-2 corroboration lesson), with active-only login
attribution and the repo allowlist.
- utils/confluence.py: get_comment / get_user_email (path-encoded).
- webapp.py: GET /connect/atlassian-connect.json (served dynamically),
POST /connect/{installed,uninstalled,webhook/comment-created}, the
space->repo resolver, thread-id, and fetch helpers.
- completion.py: source=="confluence" failure-reply branch.
Security: the sh-security-review verify pass confirmed one HIGH — the
symmetric signed-install=false first-install was trust-on-first-use gated
only by the public Confluence hostname (webhook-auth bypass). Fixed by
switching to signed-install=true + RS256 verification of lifecycle
callbacks, which cryptographically authenticates the first install. All
other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall
DoS/corroboration/injection) were defeated; residuals are deployment
config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover
bodies; comment-trigger prompt injection, shared with all sources).
New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN,
CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets
require the durable Postgres LangGraph store in prod.
Outstanding before push: /sh-security-review on the real diff and the
GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory.
* docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs
- prompt.py: Confluence-triggered runs notify via confluence_comment on
the triggering page; add Confluence to the shared-base source list.
- CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian
triggers (Jira Automation shared-secret webhook; Confluence Connect app
with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the
server-side corroboration + encrypted install store.
Phase 5 also verified the trigger surface end-to-end against a running
uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed
without valid auth) and recorded the integration in project memory.
* fix(open-swe): resolve /sh-security-review findings on the Atlassian surface
Formal sh-security-review (detector fan-out + verifier) over the Phase-4
Connect surface (esp. the new RS256 signed-install code, unseen by the
earlier adversarial verify) and the Phase-2 opt-in hardening.
CRITICAL — cross-tenant install (origin validation, CWE-346): signed-
install proves the caller is *an* Atlassian tenant, not *ours*, and the
descriptor is served publicly, so any attacker could install the app on
their own Confluence site and drive agent runs against our allowlisted
repos. The baseUrl body field is attacker-controlled and cannot bind the
tenant; only the signature-verified clientKey (JWT iss) can. Added a
MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in
process_install after signature+iss verification.
HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment
ids are per-instance, so generate_thread_id_from_confluence_comment now
salts the hash with the verified clientKey (plumbed from the webhook JWT
iss) to prevent thread hijack across tenants.
HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page
interpolated page_id into the REST path unencoded (update_page on a
mutating PUT with no params= backstop). Now _seg()-encoded, matching the
rest of the module.
MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no
bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID
guard mirroring the Linear botActor / Jira comment_author_is_bot checks.
LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as
opt-in defense-in-depth (the clientKey allowlist is the real gate).
Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp,
kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption,
constant-time comparisons, and the Phase-2 hardening. New regression
tests for each fix; full suite green (1602).
* harden(open-swe): GPT-4.1 cross-family review follow-ups
Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no
critical/high issues and confirmed the auth boundary is fail-closed and
correct. Two low-cost defense-in-depth items applied:
- Validate the signed-install JWT 'kid' against a strict charset before
the public-key fetch, so a malformed kid fails fast with no network
call (on top of the existing fixed host + percent-encoding).
- Make JWT nbf verification explicit (verify_nbf) on both the RS256
lifecycle and HS256 webhook decodes.
Other suggestions triaged as already-handled (aud cross-app replay is
blocked by the per-tenant iss->secret lookup; documented static-token/IP/
baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra
(Fernet rotation via MultiFernet; rate limiting at the gateway).
* docs(open-swe): document Jira + Confluence in installation & customization guides
- INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret,
service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect
app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note,
CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars +
REQUIRE_REPO_ALLOWLIST.
- CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction
note covers all four sources.
- AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth).
- README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
- **Webhooks**: GitHub signatures verified in `utils/github_comments.py:verify_github_signature` ; Slack/Linear handled in their respective utils; Jira via a shared-secret header (`verify_jira_secret` , optional HMAC/timestamp); Confluence via Atlassian Connect JWT + `qsh` and RS256 signed-install (`utils/atlassian_connect.py` ), with install secrets stored encrypted.
2026-05-22 14:33:46 -07:00
- **Dashboard / UI**: GitHub OAuth login lives in `agent/dashboard/oauth.py` and `routes.py` (`/auth/login` , `/auth/callback` , `/auth/logout` , `/me` ).
2026-04-30 20:06:58 -07:00
### Thread-id derivation
2026-05-22 14:33:46 -07:00
Webhooks compute deterministic thread ids so the same Linear issue / Slack thread / PR routes back to the same running agent. See `utils/github_comments.py:get_thread_id_from_branch` and the equivalents in `utils/linear.py` / `utils/slack.py` . Reviewer threads have their own deterministic ids and are tagged with `REVIEWER_THREAD_KIND` metadata so the FastAPI side can find them.
2026-04-30 20:06:58 -07:00
## Conventions
docs: land reorg in ledger, fix path refs, ship plan artifacts (C7)
C7 of the domain-reorg adoption (build plan docs/upstream-sync/domain-reorg/
reorg-build-plan.md, step C7):
- CLAUDE.md / AGENTS.md: retarget architecture path references to the new
layout — graph entrypoints via agent.graphs.* shims, the agent/api/ +
agent/webhooks/*_routes.py FastAPI split (webapp.py now a shim),
agent/review/ package, tests/<domain>/ test paths, and the new-graph/
test conventions. README.md already pointed at docs/ (C1) — no change.
- triage.jsonl: flip 8356eb34 (#1726) to landed on branch
refactor/domain-reorg-adoption; add re-triage notes to the 8 unblocked
rows (#1732/#1761/#1744/#1748/#1742 clean, #1736/#1758/#1760 near-clean).
triage.md regenerated via make triage-render.
- Ship the plan, scoping report, move-map artifacts, and the
domain-reorg-adoption workflow so the exercise is reproducible.
Memory + Confluence handled out-of-band (not in this commit): project
memory updated with the new module layout; Confluence check found no IT
page documents the repo module map — no Confluence change required.
2026-07-17 15:01:45 -04:00
- Tests are unit-only by default and organized by domain under `tests/<domain>/` (`tests/agent/` , `tests/auth/` , `tests/github/` , `tests/webhooks/` , `tests/reviewer/` , `tests/sandbox/` , `tests/middleware/` , `tests/models/` , `tests/slack/` , `tests/tools/` , `tests/dashboard/` , `tests/analyzer/` ); `tests/conftest.py` and the `tests/e2e/` harness stay at the top level. Integration tests would go under `tests/integration_tests/` (currently empty — `make integration_tests` no-ops if missing).
2026-05-22 14:33:46 -07:00
- New sandbox providers: add a module under `agent/integrations/` and wire it into `SANDBOX_FACTORIES` in `agent/utils/sandbox.py` . See `CUSTOMIZATION.md` .
- New tools: add to `agent/tools/` , export from `agent/tools/__init__.py` , add to the `tools=[...]` list in `server.py:get_agent` (or `reviewer.py` for reviewer-only tools).
- New middleware: add to `agent/middleware/` , export from `agent/middleware/__init__.py` , add to the `middleware=[...]` list in `server.py:get_agent` — order is significant (see the stack above).
- New dashboard endpoints: add to `agent/dashboard/routes.py` . The router is auto-mounted on the FastAPI app.
docs: land reorg in ledger, fix path refs, ship plan artifacts (C7)
C7 of the domain-reorg adoption (build plan docs/upstream-sync/domain-reorg/
reorg-build-plan.md, step C7):
- CLAUDE.md / AGENTS.md: retarget architecture path references to the new
layout — graph entrypoints via agent.graphs.* shims, the agent/api/ +
agent/webhooks/*_routes.py FastAPI split (webapp.py now a shim),
agent/review/ package, tests/<domain>/ test paths, and the new-graph/
test conventions. README.md already pointed at docs/ (C1) — no change.
- triage.jsonl: flip 8356eb34 (#1726) to landed on branch
refactor/domain-reorg-adoption; add re-triage notes to the 8 unblocked
rows (#1732/#1761/#1744/#1748/#1742 clean, #1736/#1758/#1760 near-clean).
triage.md regenerated via make triage-render.
- Ship the plan, scoping report, move-map artifacts, and the
domain-reorg-adoption workflow so the exercise is reproducible.
Memory + Confluence handled out-of-band (not in this commit): project
memory updated with the new module layout; Confluence check found no IT
page documents the repo module map — no Confluence change required.
2026-07-17 15:01:45 -04:00
- New graphs: add an `agent/graphs/<name>.py` re-export shim that delegates to the factory module, then register the shim entrypoint (`agent.graphs.<name>:<symbol>` ) in `langgraph.json` under `graphs` .
2026-05-22 14:33:46 -07:00
- Minimal-to-no code comments — only when the *why* isn't obvious from the code.