Commit graph

81 commits

Author SHA1 Message Date
Johannes du Plessis
f1baceb8ab
feat: add Gemini 3.5 Flash provider (#1420) 2026-06-04 17:38:54 -07:00
Johannes du Plessis
a26d5d32cb
feat: Add Fireworks model provider (#1415)
* feat: add Fireworks model provider with accurate reasoning levels

Adds Kimi K2.6, DeepSeek V4 Pro, Nemotron 3 Ultra, GLM 5.1 via fireworks: provider, mapping per-model reasoning_effort. Surfaces each model's real effort range (incl. openai none, deepseek xhigh/max).

* fix: drop serverless-unsupported Nemotron 3 Ultra, document FIREWORKS_API_KEY

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-04 21:59:47 +00:00
Johannes du Plessis
8b18e7e95d
fix: Reduce graph load and dashboard repo failures (#1412)
* fix: reduce graph load and dashboard repo failures

* fix: surface repo listing timeouts
2026-06-04 13:54:06 -07:00
Johannes du Plessis
60426e8b10
fix: reviewer resolves GitHub App token at run start, not via cross-process cache (#1409)
Reviewer/push runs execute in a worker process, but the GitHub token was
cached by the webhook handler in the API server process — a different
process — so the worker's in-process cache was always cold. resolve_github_token
then failed (User not authenticated / Unknown source: github_push) and only the
app-token fallback kept reviews working, noisily.

The reviewer always acts as the GitHub App (open-swe[bot]), so resolve the
installation token directly at run start, scoped to the repo. This also bypasses
org SAML enforcement that blocks user OAuth tokens. Drop the now-dead
cross-process cache writes in the webhook reviewer-dispatch handlers, and stop
leave_failure_comment raising on the github_push source.
2026-06-04 12:11:23 -07:00
Johannes du Plessis
c1e46b938e
feat: add Slack Block Kit reply options (#1407)
* feat: add Slack Block Kit reply options

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* docs: document Slack interactivity setup

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-04 10:26:09 -07:00
Brace Sproul
2070a770c2
fix: stop storing GitHub tokens in metadata (#1405)
* fix: stop storing GitHub tokens in metadata

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* fix: bound in-process GitHub token cache with 24h TTL + sweep

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
Co-authored-by: open-swe[bot] <johannes@langchain.dev>
2026-06-04 09:33:51 -07:00
Johannes du Plessis
8913926dba
fix: remove manual review trigger surfaces (#1396)
Open SWE Review now runs from automated PR triggers, so drop the old Slack/GitHub review keyword entrypoints and keep PR comments on the regular agent path.
2026-06-03 12:33:22 -07:00
Mukil Loganathan
f4b27c0fee
feat: add Slack Open in Web link (#1392)
* feat: add Slack Open in Web link

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: skip web link for Slack reviewer runs

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: accept include_dashboard_link kwarg in Slack reviewer test double

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <johannes@langchain.dev>
2026-06-03 11:50:37 -07:00
open-swe[bot]
ae946d1aa5 chore: attribute commit co-author to open-swe[bot], not open-swe user
The Co-authored-by trailer and bot git identity used
open-swe@users.noreply.github.com, which resolves to the separate
open-swe *user* account rather than the open-swe[bot] GitHub App.
Switch OPEN_SWE_BOT_EMAIL to the bot's noreply address
(215916821+open-swe[bot]@users.noreply.github.com) so co-author credit
and the fallback author identity point at the bot.

Drive the prompt trailer and sandbox git config from the constant
instead of hardcoding the address.
2026-06-03 11:19:24 -07:00
Johannes du Plessis
0a2e682364
fix: scope public reviewer tokens (#1389)
* fix: scope public reviewer tokens

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: simplify reviewer token wiring; fix push re-scope + red test

- Remove the redundant _check_or_recreate_sandbox_for_proxy /
  _refresh_github_proxy_or_recreate_for_proxy wrappers and call the
  underlying functions directly (they already default the token to None).
- process_github_push_event: re-scope the GitHub App token when the push
  payload lacked repo privacy/id but PR metadata reveals a public repo, so
  reviewer.py never proxies a full-installation token for a public PR.
- Clarify the two-token sequence in trigger_pr_review_from_ref.
- Fix pre-existing failing test test_proxy_refresh_failure_recreates_sandbox
  and add coverage for _reviewer_token_for_repo + push-event scoping.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-03 09:25:17 -07:00
Johannes du Plessis
c8a997c125
fix: reliable, safe Slack account-connect prompt + first-login Slack dialog (#1383)
* fix: deliver Slack account-link prompt as a visible threaded reply

Blocked Slack users got no prompt at all. Prod logs show chat.postEphemeral
returns ok, but ephemeral messages are silently dropped in Slack's assistant
threads (where Open SWE runs), so the user sees nothing. Post the prompt as a
normal threaded reply instead — the same channel the agent uses to reply.

* fix: deliver Slack auth-failure prompt as a visible threaded reply

leave_failure_comment() tried an ephemeral message first and only fell back
to a thread reply on failure. Ephemeral messages succeed (ok) but are dropped
in Slack's assistant threads, so the fallback never fired and the user saw no
auth-failure prompt. Post the visible threaded reply directly, matching the
account-link prompt fix.

* fix: prompt blocked Slack users with a generic, token-free dashboard link

Addresses the review findings that posting the per-user account-link token /
auth URL in a visible thread lets any channel member bind their GitHub account
to the triggering user's Slack identity.

Drop the per-user signed link entirely. Both the account-link prompt
(_post_account_link_prompt) and the runtime auth-failure prompt
(leave_failure_comment) now post a plain dashboard settings link
(build_settings_url) as a visible threaded reply. The user signs in with GitHub
from their own session and connects Slack via verified OIDC on the settings
page — no secret in the thread, nothing to hijack, and no DM machinery.

* feat: nudge first-time users to connect Slack from the dashboard home

Show a Connect Slack banner on the agents landing page whenever Slack OAuth is
enabled and the user hasn't linked Slack yet. A first-time user (no Slack
mapping) sees it immediately after signing in; it disappears once connected.

* feat: prompt first-time users to connect Slack via a dialog

Replace the inline Connect Slack card on the agents home with a modal dialog
(Base UI). It opens automatically once the mapping query resolves to
"not connected" and closes itself once Slack is linked; "Maybe later" dismisses
it for the session. No new dependency — uses the design system's Base UI.

* copy: frame Slack connect as resolving the user's GitHub account

Drop 'act/reply on your behalf' wording across the connect-Slack dialog, the
Slack thread prompts (blocked + auth-failure), and the settings description.
Connecting Slack lets Open SWE resolve the user's GitHub account when they tag
it in Slack.
2026-06-02 20:55:07 -07:00
Johannes du Plessis
046388a4e9
feat: simplify PR attribution footer to "Made by Open SWE" (#1382)
Replace the double-attribution PR body footer (_Opened collaboratively by
{user} and open-swe._) with a single Cursor-style footer linking to the
project. Since PRs are now opened as the triggering user, the user no longer
needs to be named in the footer. The commit Co-authored-by trailer is kept.

Legacy footers are migrated on PR updates.
2026-06-02 19:52:27 -07:00
Johannes du Plessis
cb4c643e43
feat: open Slack-triggered PRs as the triggering user (#1375)
* feat: open Slack-triggered PRs as the triggering user

Route the Slack per-user GitHub token through the dashboard OAuth store
(the backend the self-service link prompt populates) and block runs that
lack a valid user token, prompting the user to (re-)link. Per-user OAuth
now wins over bot-token-only mode for mapped Slack/dashboard users.

Flip commit/PR authorship across all sources: the triggering user is the
commit author (via repo-local git identity using their resolvable GitHub
noreply email) and open-swe[bot] is the Co-authored-by collaborator.

* fix: address PR review — shell-escape commit identity, fix token cache impersonation

- Shell-escape the triggering user's name/email with shlex.quote before
  embedding them in the repo-setup `git config` command, so a name like
  O'Connor (or a crafted one) can't break or inject into the command.
- Stop consulting the shared thread-metadata token cache in
  _resolve_dashboard_user_token. Slack thread ids are shared across the
  conversation, so a cached token from a prior triggering user could be
  returned for the current github_login. Always resolve by login from the
  dashboard OAuth store instead.

* feat: dashboard self-service user mapping + UI cleanup

- Add session-scoped GET/PUT /dashboard/api/my-mapping so users can set their
  own work email / Slack member ID (keyed by their GitHub login, source=self).
- Slack account-link prompt now redirects to Profile Settings after auth.
- Rename "My Settings" -> "Profile Settings" and "Cloud Agents" -> "Open SWE
  Agent"; remove the Integrations tab/section (folded out, low value for now)
  and redirect /integrations to Profile Settings.
- Add a "User mapping" section to Profile Settings (work email used by Slack
  and Linear, optional Slack member ID).
- Make dashboard auth cookies scheme-aware: Secure;SameSite=None over HTTPS,
  non-Secure;SameSite=Lax over http://localhost so local login works.

* feat: self-service Slack account linking via Sign in with Slack (OIDC)

Replace the spoofable manual work-email/Slack-ID form with a verified
"Sign in with Slack" flow so a logged-in GitHub user can only ever link
their own Slack identity.

- New agent/dashboard/slack_oauth.py: OIDC authorize URL, code exchange,
  userInfo identity parse, optional workspace gate, configured check.
- routes.py: session-gated GET /slack/login and /slack/callback that upsert
  the mapping from Slack-verified user_id + email (source=slack_oauth).
  Remove the spoofable PUT /my-mapping; expose slack_oauth_enabled on /me.
- UI: drop the editable inputs; add a Connect Slack button + status to the
  User mapping section.

Admin-managed mappings are unaffected and still resolve at trigger time.
2026-06-02 15:04:20 -07:00
open-swe[bot]
111f00ea9a
fix: include GitHub username in collaboration footer (#1374)
* fix: include GitHub username in collaboration footer

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: clarify legacy footer replacement

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-06-02 09:27:51 -07:00
Johannes du Plessis
c41185a3ca
chore: remove legacy mapping + admin profiles section, page user mappings (#1371)
* chore: remove legacy mapping + admin profiles section, page user mappings

- Delete the hardcoded github_user_email_map.py and the one-time
  POST /admin/user-mappings/import endpoint + Import legacy mapping UI
  button (the Store is now the sole source of truth post-import).
- Remove the per-user profiles admin section and its GET/PUT
  /admin/profiles endpoints + ProfileForm component; users still manage
  their own profile via My Settings.
- Page the user mappings list: /admin/user-mappings now takes
  page/page_size and returns {items,total,page,page_size}; the admin UI
  shows 20 rows per page with Previous/Next controls.

* Address review: fix stale import-button doc + clamp mappings page on shrink
2026-06-01 22:10:55 +00:00
Johannes du Plessis
427bfe4f56
feat: Store-backed GitHub/Slack user mapping (self-service + admin) (#1369)
* Replace hardcoded GitHub-email map with Store-backed user mapping

Move the static GITHUB_USER_EMAIL_MAP to a Store-backed bidirectional
mapping (GitHub login <-> work email <-> optional Slack ID) with an
in-process cache, self-service onboarding, and admin management.

- agent/dashboard/user_mappings.py: Store CRUD + login/email/slack-id
  indexes, sync cache readers for hot paths, async fallthrough, and a
  bulk_import that preserves existing richer records.
- Migrate all read sites (auth.py, agent_overrides.py, authorship.py,
  github_comments.py, webapp.py x2) off the dict.
- Unmapped Slack tags now run on the GitHub App installation token
  (use_installation_token_fallback) and get an ephemeral "link your
  GitHub account" prompt carrying the Slack id + email via a signed
  account-link token threaded through the OAuth state.
- OAuth callback completes a self-service (org-gated) mapping from that
  token, falling back to the verified GitHub email.
- Admin CRUD endpoints + one-time legacy import; dashboard UI section.
- Legacy dict retained only as the import payload (no longer read).

Tests: mapping store, account-link round-trip + completion, mapped vs
unmapped Slack flows; existing trust-gate tests updated to prime cache.

* Address review: cold-cache email resolution + stale alias de-indexing

- agent_overrides: add resolve_login_from_email_async that falls through to
  the Store on a cold cache; use it at the async repo-resolution call sites
  (Slack repo config, Linear comment, owner-metadata) so a mapped user still
  resolves to their GitHub login + dashboard default_repo on a fresh worker.
- user_mappings.upsert_mapping: de-index the existing login before re-indexing
  so a changed email/Slack id no longer leaves stale aliases resolving to the
  login in-process.
- Tests for both fixes; update Slack repo-config test to patch the async resolver.
2026-06-01 14:37:19 -07:00
Johannes du Plessis
bed3eefbcb
fix: Lock dashboard login to GitHub org members (#1367)
* Lock dashboard login to GitHub org members

Add an org-membership gate to the dashboard OAuth callback. After
resolving the GitHub login, enforce_org_login_gate(login) checks the
existing ALLOWED_GITHUB_ORGS allowlist before issuing a session.

- Reuses ALLOWED_GITHUB_ORGS (no new config knob) and
  is_user_active_org_member (installation-token check, so no extra
  OAuth scope and private memberships are visible).
- Fail-open when unset/blank so existing deployments keep working;
  fail-closed on API errors.
- Gate runs before the session cookie/token is persisted.

Adds unit tests and documents the behavior in INSTALLATION.md.

* docs: document Organization Members permission required for org login gate
2026-06-01 20:56:30 +00:00
Johannes du Plessis
4a55145bb1
feat: outcomes dataset + bootstrap/continual split via skills (#1365)
* fix: reset stale sandbox creation sentinel

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: treat SANDBOX_CREATING as a timestamped cross-process lock

Only reset the sentinel when proven stale (older than the creation
timeout); otherwise wait for the worker that holds the lock so a
concurrent run does not create a duplicate sandbox.

* feat(analyzer): outcomes dataset + bootstrap/continual split via skills

Rename the review_style_analyzer graph to `analyzer` and split it into two
modes, plus capture reviewer finding outcomes for continual learning.

- Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false
  positive), and GitHub/Slack thumbs findings into a single LangSmith dataset
  (openswe-reviewer-outcomes), keyed deterministically per finding+source.
  Emit points wired into update_finding, resolve_finding_thread, and the
  GitHub/Slack reaction handlers.
- Two playbooks delivered as deepagents skills (bootstrap-repo-analysis,
  continual-learning), served as virtual files via a CompositeBackend /skills/
  route + StateBackend (seeded into the run files channel at invoke time, never
  written to the sandbox). Mode is set by the launcher; continual runs fall
  back to the GitHub App installation token.
- Split launcher into start_bootstrap_analysis + start_continual_run; register
  a per-repo nightly continual-learning cron when bootstrap completes.
- New read_finding_outcomes tool feeds confirmed/dismissed findings back to the
  continual playbook.

Tests for outcome label mapping, skills helper, and cron idempotency.

* fix(analyzer): anchor continual cron runs to a real thread_id

The nightly continual-learning cron is threadless, and get_analyzer
early-returns an empty agent when configurable.thread_id is missing — so
every cron-launched run no-op'd before reading outcomes or saving a refined
prompt. Include the repo's deterministic analyzer thread_id in the continual
run configurable so the run executes; the threadless run carries no message
history, so nightly runs don't accumulate context.

* refactor(analyzer): move cron lifecycle calls out of the review-styles store

Drop the inline `analyzer_cron` imports from review_styles.py (added only to
dodge a circular import) by relocating the cron-trigger calls to the layer
above the store: registration to the save_review_style tool (after a prompt is
saved) and removal to the dashboard delete route. review_styles.py is now a
pure store again with top-level imports only.

* refactor: hoist reviewer_outcomes imports to module level

Move the two inline emit_finding_status_outcome imports introduced in this PR
(update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes
only depends on langsmith, so there is no circular import to avoid.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
open-swe[bot]
eb92947806
fix: propagate Slack reply errors (#1358)
* fix: propagate Slack reply errors

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: classify Slack 429 as rate_limited with retry-after

Slack chat.postMessage rate limiting returns HTTP 429, which raise_for_status
turned into a generic http_error and told the agent to retry immediately,
ignoring Slack's retry window. Special-case 429 (threading Retry-After) and
normalize the ratelimited body code before the generic HTTP path so the
existing rate_limited hint actually fires.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-28 23:11:42 +00:00
open-swe[bot]
dd4ca9181e
fix: route PR review requests through agent tool (#1340)
* fix: route PR review requests through agent tool

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: align public repo gate test with agent routing

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: allow app-token PR review requests

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-27 10:29:43 -07:00
Johannes du Plessis
5702a9d452
feat: reconcile reviewer comment lifecycle (#1332)
* feat: reconcile reviewer comment lifecycle

Track GitHub review threads for reviewer findings so re-reviews can resolve or reply to existing comments, and collect thumbs feedback on new review comments in LangSmith.

* fix: clarify reviewer comment lifecycle

* chore: apply reviewer formatting

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-26 16:24:34 -07:00
open-swe[bot]
e347aed851
feat: make PR creation policy opt-in (#1334)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-26 13:30:18 -07:00
Johannes du Plessis
ccea80b887
feat: inline AGENTS.md into reviewer system prompt (#1328)
* feat(reviewer): inline AGENTS.md into reviewer system prompt

Fetches AGENTS.md from the PR's head_sha via the GitHub contents API
during reviewer setup and inlines it as a "Repository conventions"
block in the system prompt. Mirrors how the per-repo review style
prompt is already wired.

The main agent has long had a mandatory step to read AGENTS.md after
cloning, but the reviewer often skips cloning entirely (it can
`gh pr diff` directly), so it never saw the file. Loading it
deterministically means the reviewer judges findings against the
project's own conventions instead of relying on the model to fetch
the file itself.

* fix(reviewer): fetch AGENTS.md from base_sha, not head_sha

The reviewer inlines AGENTS.md into its system prompt. Reading from
head_sha means a PR author can edit AGENTS.md in the same PR being
reviewed and smuggle instructions like "ignore all bugs" / "publish
no findings" into the reviewer's prompt. Switch to base_sha (the
target branch's pre-PR state, which is trusted) and update the prompt
text to reflect that the contents come from the base, not the head.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-22 14:58:27 -07:00
Johannes du Plessis
e87085139b
feat: add Agents chat UI for cloud threads (#1323)
* feat(ui): add Agents chat UI ported from open-swe-app

Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(dashboard): wire Agents UI to LangGraph thread APIs

Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(dashboard): single agent reply per turn in Agents UI

Use UUID thread IDs LangGraph accepts, skip confirming_completion for
dashboard threads, and merge adapter agent messages so duplicate bubbles
do not render.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(ui): polish Agents UI with floating prompt and layout cleanup

Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar
from open-swe-app, and refine chat layout so messages scroll behind the input.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent): patch deepagents reducer for None messages on checkpoint replay

LangGraph thread state could 500 when cancelled runs left messages as None.
Apply the reducer guard before graph import, fall back to metadata in the
dashboard API, and adjust Agents prompt bar layout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(ui): unify sidebar user menu and clean up Agents UI navigation

Extract SidebarUserMenu so the dashboard and Agents sidebars render the
same profile button, drop the redundant Agents nav row in favor of the
existing Back to Agents link, add the open-swe logo header to the Agents
sidebar, flatten the New Agent button, and cap the home screen run list
to keep the prompt input in view.

* feat(ui): resizable/collapsible sidebar shared across dashboard and Agents

Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars
share a persisted width (default 260px, drag to resize, 200-420 range)
and a collapse toggle that hides the panel and surfaces a floating
reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on-
hover thread delete control in the Agents sidebar.

* feat(ui): instant user message and busy indicator on Agents transition

Stash submitted prompts in sessionStorage, pre-populate the new thread
detail cache, and merge pending prompts into the rendered message list
so the Agents page renders the user bubble plus the existing thinking
spinner immediately instead of flashing a skeleton and "Agent is
starting" while the run boots.

* feat(ui): token-stream agent replies in the Agents thread view

Opt the LangGraph runs into messages-tuple streaming and forward those
events through the existing SSE channel. The frontend now applies
AIMessageChunk deltas directly to the cached thread (cancelling any
in-flight refetch first so optimistic tokens are not clobbered) and
keeps positional pending prompts so the user bubble stays in the right
place while the agent streams its reply.

* fix(dashboard): await threads.join_stream before iterating

threads.join_stream is async def returning an AsyncIterator, so it must
be awaited before async for. The SSE endpoint was raising
TypeError: 'async for' requires an object with __aiter__ method, got
coroutine on every connection.

* fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns

Setting stream_mode=["values","messages-tuple","updates"] on
runs.create forces langchain_anthropic into streaming, and on the
second model call (after tool execution) its serialized thinking
blocks come back malformed, so Anthropic rejects the request with
'messages.1.content.0.thinking.thinking: Field required'. Revert to
the default stream_mode so claude-opus thinking + tool use runs to
completion. The frontend keeps the messages-event handler in place
as a no-op fallback for when streaming is re-enabled.

* feat(agents): per-thread model picker wired through to the run

Add optional model_id/effort to the create-thread and send-message
request bodies, forward them as agent_model_id/agent_effort in the
LangGraph run configurable, and record the resolved choice in thread
metadata so the UI can show the model the run is actually using.
get_agent now picks the per-thread override last (highest priority over
team default + profile override) and falls back gracefully when it is
absent or unsupported.

The frontend prompt bar becomes a controlled component fed by a
shared useModelOptions hook (options + profile -> defaultSelection).
AgentsHome seeds the picker from the user's profile default; the
thread view seeds from the thread's recorded model/effort and lets
each follow-up retarget the run.

* refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar

Drop the absolute-positioned send button, restore the original
px-4 py-3.5 min-h-[106px] flex-col container, and move the model
picker into a mt-auto pt-2 footer row so the placeholder text and
the model selector share the same horizontal padding.

* chore: fix lint/format CI failures

Remove unused imports and reformat two files flagged by ruff.

* fix(tests): stop messages-reducer patch tests from polluting the suite

Restore agent modules after reducer patch tests and import LangSmithSandbox
from agent.server in proxy refresh tests so isinstance checks stay valid.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 18:15:59 +00:00
Johannes du Plessis
82852f9eda
feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312)
* feat: tune reviewer for precision — web/wiki tools + recalibrated prompt

Reviewer agent now has web_search, fetch_url, and http_request alongside the
finding tools, so it can verify library semantics and consult the DeepWiki
auto-generated wiki for public repos (https://deepwiki.com/<owner>/<repo>)
before flagging cross-file or architectural concerns.

Prompt rewritten to push precision over recall:
- explicit severity ladder pushing reviews toward bimodal high/low instead of
  defaulting to medium
- ≤200-char description target (gold set averages ~186 chars; we were at ~436)
- mandatory docs / wiki / code lookup before flagging concurrency, security,
  or perf — the three categories that dominated false positives
- "do not flag" list covering compiler/linter-catchable nits, speculative
  claims without a concrete attacker/interleaving/scale, style preferences
  the codebase doesn't share, and test-quality nits on non-test diffs
- smart file-selection guidance for large PRs (deprioritize generated /
  vendored / pure-rename hunks)

Eval config switched to openai:gpt-5.5 + high reasoning effort for the next
benchmark run.

* trim prompt

* subagent prompting

* confidence ratings

* added medium

* enforce confidence threshold

* .

* reviewer: precision-tuned prompt + drop confidence gate

Rewrites the reviewer system prompt around a defensibility bar (anchor +
failure mode + maintainer wouldn't say "not a bug"), an explicit do-not-file
list (style nits, speculation, scope-policing, same-bug fan-out), and a
checklist of 10 bug archetypes drawn from a per-PR audit of the eval golden
set. The audit showed 145 FPs in the last eval split ~28% speculative, ~26%
style-nit, ~31% real-but-unscored (mostly same-archetype fan-out); the new
prompt targets each class directly.

Confidence is still recorded on every finding for post-hoc calibration but
no longer gates publication — the audit showed the gate was a no-op (agent
self-rated 65% of findings "high" regardless), and the prompt's defensibility
bar is the actual discipline. Drops CONFIDENCE_ORDER, CONFIDENCE_THRESHOLD,
the confidence_threshold kwarg on filter_findings_for_publish, the
confidence_filtered score_mode, and the min_confidence kwarg on the eval
target's _extract_comments — all dead once the gate is gone.

Also removes the "informational" severity tier from the Severity enum,
SEVERITY_ORDER, and all validators / tests / docstrings. It was reserved for
FYI observations the dataset never rewards.

* benchmax

* adding google provider

* slight steering

* tuning

* more tuning

* fix

* cleanup

* reducing overfitting

* Add per-repo review style profiles and inject them into the reviewer.

Dashboard users can analyze historical PR review feedback per repository,
edit the resulting style guide, and have it loaded from LangGraph Store at
reviewer runtime (including Martian eval runs) keyed by owner/name.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix review style job errors leaking exception details to clients.

Return generic dashboard messages while logging full stack traces server-side.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-20 18:35:00 +00:00
Johannes du Plessis
834efbc33c
feat: Adds ability to run evals against deployment (#1311)
* feat: tighten reviewer eval workflow

Require the reviewer to verify and dedupe findings before recording them, and make benchmark runs safe to execute against deployed reviewer graphs without posting GitHub reviews.

* chore: move reviewer eval settings to config

Load reviewer benchmark settings from the default eval config file so deployed eval runs do not require a wide CLI surface.

* feat: allow reviewer eval model overrides

Pass reviewer model and reasoning effort from the eval config into reviewer runs so isolated benchmark deployments can test Opus 4.7 high thinking.

* fix: use adaptive thinking for Opus 4.7

Switch Opus 4.7 model overrides to Anthropic adaptive thinking with effort instead of the deprecated budgeted thinking payload rejected by the API.

* refactor: use latest Anthropic effort API

Remove legacy Anthropic budget-token thinking support and route Anthropic efforts through adaptive thinking plus effort.

* revert prompting
2026-05-18 15:47:13 -07:00
Johannes du Plessis
88856a04fa
feat: open-swe dashboard for per-user profile config (#1302)
* feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints

Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering:
- GitHub App OAuth login → JWT cookie session (cross-domain ready)
- profile CRUD against LangGraph Store with model+effort validation
- admin gate via CONFIGURED_ADMINS
- /repos via /user/installations using the user's encrypted OAuth token

CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the
Vercel-hosted frontend can call the LangSmith deployment with credentials.

* feat: apply dashboard profile model/effort overrides in get_agent

Look up the triggering user's GitHub login from config (direct field or
GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store,
and apply default_model + reasoning_effort to make_model when both are
valid. Effort 'max' is captured on the profile but not yet wired through —
the OpenAI Reasoning Literal doesn't accept it.

* feat: ui/ TanStack Start dashboard for profile config

Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template,
base-ui primitives, Tailwind v4). Three routes:

- /login   — Sign in with GitHub (links to /dashboard/api/auth/login)
- /profile — Edit default model, reasoning effort, default repo
- /admin   — Admin-only: list users and edit other profiles

API client (src/lib/api.ts) uses credentials: include so the osw_session
cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL
points at the LangSmith deployment.

Effort options re-render when the model changes; 'max' on Opus 4.7 is
captured on the profile but ignored downstream until anthropic reasoning
is wired through make_model.

* feat: searchable Combobox for default repo picker

Replaces the Select with a base-ui Combobox so users can filter by typing,
the popup is wider than the trigger so full owner/repo names are readable,
and the list caps at max-h-80 to stay on screen.

* fix: address review comments + wire default_repo and Anthropic thinking

Security/correctness fixes from PR review:

* Open redirect: validate `redirect_to` in `/auth/login` against
  `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it
  into the state JWT. Anything off-allowlist falls back to the dashboard
  base URL. (PR #1302 r3250054386)

* Login CSRF: bind the OAuth `state` to the requesting browser. At
  `/auth/login` we generate a fresh nonce, set it as a short-lived
  HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and
  embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback`
  we require the cookie nonce to hash-match the state JWT's nonce_hash
  (constant-time compare). (PR #1302 r3250054395)

* RMW race in profile vs token writes: split storage into two
  namespaces — `["profiles"]` for user-editable settings and
  `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now
  only writes its own namespace so an in-flight profile save can no
  longer clobber a fresh token from a concurrent re-login (and vice
  versa). (PR #1302 r3250054393)

* /repos pagination: follow `Link: rel="next"` for both
  `/user/installations` and per-installation `/repositories` with
  per_page=100, capped at 1000 items. (PR #1302 r3250054401)

Feature wires:

* default_repo: applied as a fallback in `get_slack_repo_config` (after
  explicit-repo / thread metadata, before the env defaults) and in the
  Linear webhook (after comment-body extraction, before team mapping).
  Both paths resolve the triggering user's GitHub login via
  GITHUB_USER_EMAIL_MAP and read the profile's default_repo.

* Anthropic "thinking" effort: `make_model` now accepts a `thinking`
  kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max}
  to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is
  anthropic. OpenAI path still ignores "max" since the Literal doesn't
  accept it.
2026-05-15 11:23:53 -07:00
open-swe[bot]
d2c62e5fc6
chore: remove slack assistants api feature flag (#1295)
The SLACK_ASSISTANTS_API_ENABLED env flag and its gating function
are removed. set_slack_assistant_status now always proceeds when a
bot token and channel/thread are provided.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-12 00:41:51 +00:00
Johannes du Plessis
5c7c78406c
fix: keep sandbox backend stable across recovery (#1294)w
Use a per-thread proxy so in-flight tools continue through the latest recreated sandbox instead of holding a stale backend reference.
2026-05-11 16:03:38 -07:00
open-swe[bot]
85343fab63
feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322] (#1280)
* feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322]

Persist github_token_expires_at alongside github_token_encrypted, treat
expired cache entries as missing so we re-resolve before kicking off
runs, and invalidate the cached ciphertext on a downstream 401 so the
next invocation gets a fresh token instead of replaying a revoked one.

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* webapp: forward installation-token expiry to reviewer cache writes

The three reviewer-thread persist sites in webapp.py were calling
get_github_app_installation_token() (no expiry) and persist_encrypted_github_token
without expires_at, so cached App tokens were treated as never-expiring even
though they actually expire in ~1 hour.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 22:57:01 +00:00
open-swe[bot]
6a984d86c4
feat: add 10 more tips to the trace-reply rotation (#1283)
* feat: add 10 more tips to the trace-reply rotation

The TRACE_REPLY_TIPS pool only had 9 tips, so users mostly saw the same
ones. Added 10 more grounded in actual features (review command, image
attachments, GitHub-issue triggers, persistent sandboxes, OAuth fallback,
etc.) so the rotation surfaces more of what open-swe can actually do.

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* swap out 3 tips per review feedback

Replaced the suggestion-blocks, sandbox-auto-recovery, and OAuth-fallback
tips with three more practically useful ones: cross-posted Slack message
resolution, the web_search tool, and the Linear ticket-management tools.

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-08 22:52:29 +00:00
Johannes du Plessis
094b2df939
feat: cross-provider model fallback on transient errors (#1281)
When the primary model raises a transient provider error (5xx, 429,
connection/timeout) the request is retried once against a fallback
model from the other provider. Anthropic primaries fall back to
OpenAI and vice versa. Also bumps the SDK max_retries from the
default 2 to 6 so quick blips stay on the primary and keep prompt
caching warm.

Triggered by 529 OverloadedError traces that ended runs silently
with no Slack/Linear/PR reply.
2026-05-08 15:35:13 -07:00
Johannes du Plessis
51bdde93b9
slack: move feedback-reaction ask out of every reply, into the tip rotation (#1279)
Appending "Please react with 👍 or 👎..." to every Slack completion
message ended up being repetitive noise. Drop the prompt-side instruction
and add the same ask as one of the rotating tips on the trace reply, so
users still see it occasionally without it cluttering each final summary.
2026-05-08 21:59:14 +00:00
Johannes du Plessis
b541f6d08a
slack: slim down the trace reply (drop greeting, suppress unfurl) (#1278)
Now that assistant.threads.setStatus carries the "is thinking…"
indicator, the per-run greeting phrase + auto-unfurl on the LangSmith
link are just visual noise stacking on top of it. The trace reply now
posts only `<url|View trace>` + a tip, with link unfurling off so the
smith.langchain.com card no longer appears.

Adds an `unfurl_links`/`unfurl_media` knob to post_slack_thread_reply_with_ts
(default-on to preserve behaviour for every other caller).
2026-05-08 14:51:28 -07:00
Johannes du Plessis
3f9dbb6597
feat: add Slack reaction feedback to LangSmith (#1231)
* feat: add Slack reaction feedback to LangSmith

Record Slack reaction feedback against explicitly mapped LangGraph runs so user ratings are idempotent and tied to the message they reacted to.

* Address review feedback on Slack reaction → LangSmith feedback

- langsmith.py: drop lru_cache on _build_langsmith_feedback_clients so
  rotated keys / late env hydration are picked up; dedupe by (key, url)
  tuple instead of key alone so the same key pointing at different
  endpoints (cloud + self-hosted) builds both clients.
- langsmith.py: treat LangSmithNotFoundError on delete_feedback as
  success — out-of-order or redelivered reaction_removed events would
  otherwise loop forever on Slack's retry policy.
- slack_feedback.py: include channel_id in _feedback_key so the same
  message_ts in two channels can't collide on the same feedback id.
- slack_feedback.py: treat conflicting +/- reactions from one user as
  ambiguous (clear feedback) instead of averaging to a misleading 0.5.
- slack_feedback.py + slack.py + webapp.py: gate reaction handling to
  the user who triggered the run (stored in the slack_run_map mapping
  alongside run_id). Prevents bystanders in shared channels from
  polluting eval feedback.
2026-05-08 14:24:12 -07:00
Johannes du Plessis
d151a118ae
Use Claude Code style spinner vocab for Slack loading messages (#1276) 2026-05-08 14:23:21 -07:00
Johannes du Plessis
d38f17ffe8
fix Slack assistant status endpoint (#1272) 2026-05-08 13:07:32 -07:00
open-swe[bot]
dc9a0b98da
feat: gate @open-swe mentions on public repos to org members (#1273)
Adds a webhook-level check so only members of $PUBLIC_REPO_ORG_GATE
(e.g. langchain-ai) can trigger Open SWE via mentions or review
requests on public repositories. Private repos remain governed by the
existing org/repo allowlists. Internal bots bypass the gate.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-08 11:38:29 -07:00
open-swe[bot]
96f97710ad
feat: add optional Slack Assistants API typing status indicator (#1269)
* feat: add optional Slack Assistants API typing status indicator

Mirrors OpenClaw's pragmatic approach: instead of rebuilding around
assistant_thread_started events, just opt into assistants.threads.setStatus
to show 'is thinking…' while the agent is working, and clear it when
post_slack_thread_reply lands. Gated behind SLACK_ASSISTANTS_API_ENABLED so
it can be toggled without touching code.

* fix(slack): drop redundant clear, add status heartbeat across model calls

- Slack auto-clears the typing indicator on bot post; remove the explicit
  assistants.threads.setStatus("") call from post_slack_thread_reply.
- The indicator expires after ~2 minutes; add a before_model middleware
  that refreshes it on every model tick so it stays visible across long
  agent runs. Reuses the existing slack_thread.{channel_id,thread_ts}
  configurable already plumbed for notify_step_limit.
- chat:write is sufficient on the bot token (assistant:write is on the
  way out per Slack docs); no scope or app-config change required.

* feat(slack): contextual status text + rotating loading_messages

- set_slack_assistant_status now accepts an optional loading_messages list
  (capped at 10 per Slack's API), surfaced via the assistants.threads.setStatus
  payload so Slack rotates through them client-side.
- The heartbeat middleware derives a contextual status from the last
  assistant message's tool calls (e.g. "searching the codebase…" after
  grep, "running commands…" after execute), falling back to the default
  "is thinking…" when no tool calls or unknown tool name.
- Adds a curated DEFAULT_LOADING_MESSAGES list passed alongside the
  contextual status on each refresh.

* fix slack assistant status lifecycle

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 10:21:55 -07:00
open-swe[bot]
5a845ba99f
feat: add Tip section to Slack trace-reply initial message (#1268)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-08 09:25:06 -07:00
Johannes du Plessis
7e1746f420
feat(open-swe): trigger reviewer from @open-swe review PR comment (#1259)
* feat(open-swe): trigger reviewer agent from `@open-swe review` PR comment

Mirrors the Slack `@open-swe review` flow on GitHub: a comment containing
`@open-swe review` (optionally followed by a PR URL) on a PR triggers the
reviewer agent. Without a URL it reviews the commenting PR; with a URL it
targets that PR. Works for `issue_comment`, `pull_request_review_comment`,
and `pull_request_review` events, gated by the existing reviewer repo
allowlist and reusing `trigger_pr_review_from_ref`.

* fix(open-swe): require URL after `@open-swe review`, don't swallow trailing text

The previous regex matched any non-whitespace token after `review`, including
across newlines. Comments like `@open-swe review\nthanks!` parsed as
`(True, "thanks!")`, which then failed PR-URL parsing and was silently
dropped — the user got no review and the comment never reached the regular
PR-comment handler.

Restrict the optional URL token to `https?://\S+` so non-URL trailing text
falls through to `process_github_pr_comment` instead of being eaten by the
review-command branch. Adds regression tests for the multiline and
trailing-word cases.
2026-05-07 17:04:35 -07:00
open-swe[bot]
2fb719822f
feat: add "Running to the roar!" trace reply phrase (#1252)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-07 13:18:54 -07:00
open-swe[bot]
5a01aff1c7
feat: only post Slack 'Working on it!' on first thread mention (#1250)
* feat: only post Slack 'Working on it!' on first thread mention

* feat: randomize Slack trace reply phrase

Pick from a small list of friendly phrases instead of always saying
'Working on it!' so the bot feels less robotic. Explicit messages (e.g.
'Taking a look...' from PR review path) are unaffected.

* adjust phrases

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-07 12:37:46 -07:00
Johannes du Plessis
1319347dd9
feat: route Slack PR review requests (#1245)
* feat: route Slack PR review requests

Add a lightweight Slack review command path that starts the reviewer graph directly and gives the core agent a handoff tool when review requests are misrouted.

* fix: harden Slack PR review routing

* fix: validate Slack PR review URLs

* fix: preserve malformed GitHub review routing
2026-05-06 17:14:43 -07:00
Johannes du Plessis
96774f20ae
feat: move github workflows to gh cli (#1238)
* feat: move github workflows to gh cli

Use LangSmith proxy auth to support gh-driven GitHub workflows while removing custom GitHub wrapper tools.

* docker ignore + snapshot and docker image updates

* updated image and instructions

* removing open_pr if needed after agent call
2026-05-04 18:03:53 -07:00
Johannes du Plessis
13f5d8a1c9
fix: preserve existing PR descriptions (#1237)
* fix: preserve existing PR descriptions

* test: update existing PR label expectations
2026-05-04 11:43:32 -07:00
Johannes du Plessis
30530bf4d5
fix: preserve existing PR titles (#1236)
Keep automatic existing-PR updates from retitling pull requests, while documenting the explicit edit path for intentional title changes.
2026-05-04 09:35:08 -07:00
Brace Sproul
3405d145ac
feat: add edit_pull_request tool for editing PR titles/descriptions (#1063)
* feat: add edit_pull_request tool for editing PR titles and descriptions

* fix: patch auth flow in open PR middleware tests

* fix: support app token for editing PRs

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 18:12:41 -07:00
langsmith-forge[bot]
bd97678a5e
fix: prevent futile retry loop when commit_and_open_pr fails (#1210)
* fix: prevent futile retry loop when commit_and_open_pr fails with git/API errors

- Root cause: when git checkout or GitHub PR API fails, the tool returned a generic {"success": false} error with no signal that retrying is futile, causing the agent to loop 9-13+ times until hitting the 1000-step recursion limit
- Change: (1) git_checkout_branch now returns (bool, str) so the actual git error output is surfaced in the tool response; (2) checkout and PR creation failures now include "fatal": true and an explicit "Do not retry" message; (3) prompt.py COMMIT_PR_SECTION adds an explicit instruction to stop on fatal errors
- Verified: 109 unit tests pass, no regressions

* fix: skip PR safety net on fatal commit failures

* style(open_pr): ruff-format fatal retry skip condition

---------

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 22:05:22 +00:00
langsmith-forge[bot]
8095f93364
fix: update existing PR title and body when create returns 422 (#1163)
- Root cause: create_github_pr found an existing PR on 422 but never
  PATCHed it, so callers like commit_and_open_pr could not update the
  PR body (e.g. adding "Closes AB-1159") on subsequent invocations.
- Change: after _find_existing_pr succeeds, call new _update_github_pr
  helper which PATCHes /repos/{owner}/{repo}/pulls/{number} with the
  requested title and body before returning pr_existing=True.
- Verified: self-evident API call addition; proof in production traces.

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-01 14:02:39 -07:00