Commit graph

78 commits

Author SHA1 Message Date
Johannes du Plessis
e87085139b
feat: add Agents chat UI for cloud threads (#1323)
* feat(ui): add Agents chat UI ported from open-swe-app

Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(dashboard): wire Agents UI to LangGraph thread APIs

Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(dashboard): single agent reply per turn in Agents UI

Use UUID thread IDs LangGraph accepts, skip confirming_completion for
dashboard threads, and merge adapter agent messages so duplicate bubbles
do not render.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(ui): polish Agents UI with floating prompt and layout cleanup

Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar
from open-swe-app, and refine chat layout so messages scroll behind the input.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent): patch deepagents reducer for None messages on checkpoint replay

LangGraph thread state could 500 when cancelled runs left messages as None.
Apply the reducer guard before graph import, fall back to metadata in the
dashboard API, and adjust Agents prompt bar layout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(ui): unify sidebar user menu and clean up Agents UI navigation

Extract SidebarUserMenu so the dashboard and Agents sidebars render the
same profile button, drop the redundant Agents nav row in favor of the
existing Back to Agents link, add the open-swe logo header to the Agents
sidebar, flatten the New Agent button, and cap the home screen run list
to keep the prompt input in view.

* feat(ui): resizable/collapsible sidebar shared across dashboard and Agents

Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars
share a persisted width (default 260px, drag to resize, 200-420 range)
and a collapse toggle that hides the panel and surfaces a floating
reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on-
hover thread delete control in the Agents sidebar.

* feat(ui): instant user message and busy indicator on Agents transition

Stash submitted prompts in sessionStorage, pre-populate the new thread
detail cache, and merge pending prompts into the rendered message list
so the Agents page renders the user bubble plus the existing thinking
spinner immediately instead of flashing a skeleton and "Agent is
starting" while the run boots.

* feat(ui): token-stream agent replies in the Agents thread view

Opt the LangGraph runs into messages-tuple streaming and forward those
events through the existing SSE channel. The frontend now applies
AIMessageChunk deltas directly to the cached thread (cancelling any
in-flight refetch first so optimistic tokens are not clobbered) and
keeps positional pending prompts so the user bubble stays in the right
place while the agent streams its reply.

* fix(dashboard): await threads.join_stream before iterating

threads.join_stream is async def returning an AsyncIterator, so it must
be awaited before async for. The SSE endpoint was raising
TypeError: 'async for' requires an object with __aiter__ method, got
coroutine on every connection.

* fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns

Setting stream_mode=["values","messages-tuple","updates"] on
runs.create forces langchain_anthropic into streaming, and on the
second model call (after tool execution) its serialized thinking
blocks come back malformed, so Anthropic rejects the request with
'messages.1.content.0.thinking.thinking: Field required'. Revert to
the default stream_mode so claude-opus thinking + tool use runs to
completion. The frontend keeps the messages-event handler in place
as a no-op fallback for when streaming is re-enabled.

* feat(agents): per-thread model picker wired through to the run

Add optional model_id/effort to the create-thread and send-message
request bodies, forward them as agent_model_id/agent_effort in the
LangGraph run configurable, and record the resolved choice in thread
metadata so the UI can show the model the run is actually using.
get_agent now picks the per-thread override last (highest priority over
team default + profile override) and falls back gracefully when it is
absent or unsupported.

The frontend prompt bar becomes a controlled component fed by a
shared useModelOptions hook (options + profile -> defaultSelection).
AgentsHome seeds the picker from the user's profile default; the
thread view seeds from the thread's recorded model/effort and lets
each follow-up retarget the run.

* refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar

Drop the absolute-positioned send button, restore the original
px-4 py-3.5 min-h-[106px] flex-col container, and move the model
picker into a mt-auto pt-2 footer row so the placeholder text and
the model selector share the same horizontal padding.

* chore: fix lint/format CI failures

Remove unused imports and reformat two files flagged by ruff.

* fix(tests): stop messages-reducer patch tests from polluting the suite

Restore agent modules after reducer patch tests and import LangSmithSandbox
from agent.server in proxy refresh tests so isinstance checks stay valid.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 18:15:59 +00:00
Johannes du Plessis
faa27479c3
fix: review style prompts UX, stale runs, and OAuth refresh (#1321)
* fix(dashboard): review style prompts UX, stale runs, and OAuth refresh

Reconcile stuck "running" analysis state, add cancel/remove for style
profiles, stack the review styles UI vertically, and auto-refresh expiring
GitHub user tokens with proactive and 401-triggered rotation.

* fix(dashboard): address review feedback on style prompts PR

Restore GitHub repo access checks on create/save with token refresh,
and preserve running status when LangGraph sync fails transiently.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 17:44:04 +00:00
Johannes du Plessis
32ec9b485f
feat: restructure Open SWE Review tab + wire create_prs (#1319)
* feat(dashboard): restructure Open SWE Review tab + wire create_prs

Restructures the dashboard around two related changes the reviewer settings
have been asking for:

- Wire profile.create_prs. Defaults to true (opt-out); when off the system
  prompt gets a `Pull Request Policy Override` section telling the agent
  to push the branch and notify with the branch URL instead of opening a
  PR. Removes the noop Slack Notifications / Allow Artifacts / First Name
  / Last Name controls and their schema fields.
- Repositories opt-in for Open SWE Review. New per-team enabled list
  stored in the LangGraph Store (`["enabled_review_repos"]`). Every
  reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review`
  which AND-combines the existing env allowlist with the dashboard list.
  Default is empty (opt-in) — admins enable repos per-installation from
  the new Repositories page nested under Open SWE Review.
- Open SWE Review tab now mirrors the Cursor "rules" pattern: main page
  shows installation rows + a Rules entry; both drill into nested pages
  (/review/repositories/$owner and /review/styles) with a back link.
- Adds the new logo/favicon assets shipped from sidebar + html head.

Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults
`is_review_repo_enabled` to True for existing allowlist tests.

* fix(dashboard): make main content scroll independently of the sidebar

Outer flex container was min-h-svh, so it grew with main's content and the
whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden
so the sidebar stays put and only <main> scrolls.

* fix(dashboard): make disabled repo toggles obviously disabled

Switch's disabled state used opacity-50 against a muted background, so
the not-admin state looked nearly identical to the off state. Bump to
opacity-40 + grayscale, and wrap each repo toggle in a span carrying a
native hover tooltip explaining why it's disabled.

* fix(switch): handle base-ui's data-disabled state

base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute)
when disabled, so Tailwind's disabled: variant never matches and the
button keeps its cursor-pointer + clickable look. Mirror the styling
under the data-[disabled] variant and add pointer-events-none so the
disabled state is both visible and actually unclickable.

* feat(dashboard): paginate per-installation repository list

20 repos per page with Prev / page X of Y / Next controls at the bottom.
Pager only renders when there are more than 20 repos. Page resets to 0
when navigating between installations.

* feat(dashboard): global default model selectors for Agent + Reviewer

Adds team-wide default model + reasoning effort for both agents in the
Admin tab so operators can switch models without redeploying.

Resolution chain:
  Agent:    hardcoded -> LLM_MODEL_ID env -> team default -> user profile
  Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable

Team defaults live in team_settings and are validated against the
SUPPORTED_MODELS allowlist + the model's supported reasoning efforts.
'Inherit from env' clears the override and falls back to LLM_MODEL_ID.

* refactor(models): drop LLM_MODEL_ID env in favour of the team default

The team default is now the single source of truth for the runtime model
choice; per-user (agent) and per-call configurable (reviewer) selections
still win on top. When no admin has touched the team default, it surfaces
the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the
admin UI's dropdown is always pre-populated with a sensible value.

The Admin UI loses the 'Inherit from env' option since there is no longer
an env layer to inherit from.

* chore(models): set hardcoded fallback to gpt-5.5 medium

Decouple the team-default boot value (gpt-5.5 / medium) from each model's
ProfileForm-suggested default_effort so we can change one without nudging
the other. The Opus xhigh default for new user profiles is unchanged.

* feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings

- Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new
  description copy that matches the screenshot. Legacy stored values
  fall back to 'every_push' on read so the UI never shows an unknown
  selection.
- Add a 'Coming soon' badge + greyed-out + disabled state on the
  controls that don't have runtime consumers yet: Trigger Mode,
  Autofix Mode, Autofix Severity Threshold, and Automatically fix CI
  failures. SettingsRow grew a comingSoon prop to keep this consistent.
- My Settings drops the noop PR Preferences section and adds a Sign
  Out button. preferred_pr_destination is removed from the profile
  schema; old records get the field popped on next write.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
Johannes du Plessis
82852f9eda
feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312)
* feat: tune reviewer for precision — web/wiki tools + recalibrated prompt

Reviewer agent now has web_search, fetch_url, and http_request alongside the
finding tools, so it can verify library semantics and consult the DeepWiki
auto-generated wiki for public repos (https://deepwiki.com/<owner>/<repo>)
before flagging cross-file or architectural concerns.

Prompt rewritten to push precision over recall:
- explicit severity ladder pushing reviews toward bimodal high/low instead of
  defaulting to medium
- ≤200-char description target (gold set averages ~186 chars; we were at ~436)
- mandatory docs / wiki / code lookup before flagging concurrency, security,
  or perf — the three categories that dominated false positives
- "do not flag" list covering compiler/linter-catchable nits, speculative
  claims without a concrete attacker/interleaving/scale, style preferences
  the codebase doesn't share, and test-quality nits on non-test diffs
- smart file-selection guidance for large PRs (deprioritize generated /
  vendored / pure-rename hunks)

Eval config switched to openai:gpt-5.5 + high reasoning effort for the next
benchmark run.

* trim prompt

* subagent prompting

* confidence ratings

* added medium

* enforce confidence threshold

* .

* reviewer: precision-tuned prompt + drop confidence gate

Rewrites the reviewer system prompt around a defensibility bar (anchor +
failure mode + maintainer wouldn't say "not a bug"), an explicit do-not-file
list (style nits, speculation, scope-policing, same-bug fan-out), and a
checklist of 10 bug archetypes drawn from a per-PR audit of the eval golden
set. The audit showed 145 FPs in the last eval split ~28% speculative, ~26%
style-nit, ~31% real-but-unscored (mostly same-archetype fan-out); the new
prompt targets each class directly.

Confidence is still recorded on every finding for post-hoc calibration but
no longer gates publication — the audit showed the gate was a no-op (agent
self-rated 65% of findings "high" regardless), and the prompt's defensibility
bar is the actual discipline. Drops CONFIDENCE_ORDER, CONFIDENCE_THRESHOLD,
the confidence_threshold kwarg on filter_findings_for_publish, the
confidence_filtered score_mode, and the min_confidence kwarg on the eval
target's _extract_comments — all dead once the gate is gone.

Also removes the "informational" severity tier from the Severity enum,
SEVERITY_ORDER, and all validators / tests / docstrings. It was reserved for
FYI observations the dataset never rewards.

* benchmax

* adding google provider

* slight steering

* tuning

* more tuning

* fix

* cleanup

* reducing overfitting

* Add per-repo review style profiles and inject them into the reviewer.

Dashboard users can analyze historical PR review feedback per repository,
edit the resulting style guide, and have it loaded from LangGraph Store at
reviewer runtime (including Martian eval runs) keyed by owner/name.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix review style job errors leaking exception details to clients.

Return generic dashboard messages while logging full stack traces server-side.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-20 18:35:00 +00:00
Johannes du Plessis
834efbc33c
feat: Adds ability to run evals against deployment (#1311)
* feat: tighten reviewer eval workflow

Require the reviewer to verify and dedupe findings before recording them, and make benchmark runs safe to execute against deployed reviewer graphs without posting GitHub reviews.

* chore: move reviewer eval settings to config

Load reviewer benchmark settings from the default eval config file so deployed eval runs do not require a wide CLI surface.

* feat: allow reviewer eval model overrides

Pass reviewer model and reasoning effort from the eval config into reviewer runs so isolated benchmark deployments can test Opus 4.7 high thinking.

* fix: use adaptive thinking for Opus 4.7

Switch Opus 4.7 model overrides to Anthropic adaptive thinking with effort instead of the deprecated budgeted thinking payload rejected by the API.

* refactor: use latest Anthropic effort API

Remove legacy Anthropic budget-token thinking support and route Anthropic efforts through adaptive thinking plus effort.

* revert prompting
2026-05-18 15:47:13 -07:00
Johannes du Plessis
f662ad6587
fix: stop inferring Slack repos from message text (#1306)
Avoid treating regex-parsed Slack text as authoritative routing or allowlist input so repository mentions remain model context and access is bounded by installation permissions.
2026-05-15 22:36:44 +00:00
Johannes du Plessis
88856a04fa
feat: open-swe dashboard for per-user profile config (#1302)
* feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints

Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering:
- GitHub App OAuth login → JWT cookie session (cross-domain ready)
- profile CRUD against LangGraph Store with model+effort validation
- admin gate via CONFIGURED_ADMINS
- /repos via /user/installations using the user's encrypted OAuth token

CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the
Vercel-hosted frontend can call the LangSmith deployment with credentials.

* feat: apply dashboard profile model/effort overrides in get_agent

Look up the triggering user's GitHub login from config (direct field or
GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store,
and apply default_model + reasoning_effort to make_model when both are
valid. Effort 'max' is captured on the profile but not yet wired through —
the OpenAI Reasoning Literal doesn't accept it.

* feat: ui/ TanStack Start dashboard for profile config

Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template,
base-ui primitives, Tailwind v4). Three routes:

- /login   — Sign in with GitHub (links to /dashboard/api/auth/login)
- /profile — Edit default model, reasoning effort, default repo
- /admin   — Admin-only: list users and edit other profiles

API client (src/lib/api.ts) uses credentials: include so the osw_session
cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL
points at the LangSmith deployment.

Effort options re-render when the model changes; 'max' on Opus 4.7 is
captured on the profile but ignored downstream until anthropic reasoning
is wired through make_model.

* feat: searchable Combobox for default repo picker

Replaces the Select with a base-ui Combobox so users can filter by typing,
the popup is wider than the trigger so full owner/repo names are readable,
and the list caps at max-h-80 to stay on screen.

* fix: address review comments + wire default_repo and Anthropic thinking

Security/correctness fixes from PR review:

* Open redirect: validate `redirect_to` in `/auth/login` against
  `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it
  into the state JWT. Anything off-allowlist falls back to the dashboard
  base URL. (PR #1302 r3250054386)

* Login CSRF: bind the OAuth `state` to the requesting browser. At
  `/auth/login` we generate a fresh nonce, set it as a short-lived
  HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and
  embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback`
  we require the cookie nonce to hash-match the state JWT's nonce_hash
  (constant-time compare). (PR #1302 r3250054395)

* RMW race in profile vs token writes: split storage into two
  namespaces — `["profiles"]` for user-editable settings and
  `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now
  only writes its own namespace so an in-flight profile save can no
  longer clobber a fresh token from a concurrent re-login (and vice
  versa). (PR #1302 r3250054393)

* /repos pagination: follow `Link: rel="next"` for both
  `/user/installations` and per-installation `/repositories` with
  per_page=100, capped at 1000 items. (PR #1302 r3250054401)

Feature wires:

* default_repo: applied as a fallback in `get_slack_repo_config` (after
  explicit-repo / thread metadata, before the env defaults) and in the
  Linear webhook (after comment-body extraction, before team mapping).
  Both paths resolve the triggering user's GitHub login via
  GITHUB_USER_EMAIL_MAP and read the profile's default_repo.

* Anthropic "thinking" effort: `make_model` now accepts a `thinking`
  kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max}
  to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is
  anthropic. OpenAI path still ignores "max" since the Literal doesn't
  accept it.
2026-05-15 11:23:53 -07:00
langsmith-forge[bot]
74f5df5df8
fix: publish_review tool returns generic "Failed to POST PR review" without GitHub API status/body, agent retries with no signal (#1299)
* fix(reviewer): surface HTTP status and body excerpt for non-dict GitHub PR review responses

When post_pull_request_review received a non-dict body, it returned
None and publish_review surfaced a generic 'Failed to POST PR review'
string with no signal for the agent to adapt — leading to blind
retries with permuted cap/severity_threshold args.

Now the non-dict-body path mirrors the existing HTTPStatusError /
HTTPError paths: it returns {'_error': 'HTTP <status>: non-dict
response body: <excerpt>'} so the user-facing tool can include the
underlying detail. The bare-None branch in publish_review.py is kept
as a defensive guard with a clearer message.

* ci: apply ruff format to reviewer_publish.py

Collapse the multi-line return dict into a single line so it matches the
output of `ruff format`, unblocking the Agent lint / format-check CI jobs.

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

---------

Co-authored-by: LangSmith Issues Agent <issues-agent@langsmith.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-12 13:55:39 -07:00
open-swe[bot]
d2c62e5fc6
chore: remove slack assistants api feature flag (#1295)
The SLACK_ASSISTANTS_API_ENABLED env flag and its gating function
are removed. set_slack_assistant_status now always proceeds when a
bot token and channel/thread are provided.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-12 00:41:51 +00:00
Johannes du Plessis
5c7c78406c
fix: keep sandbox backend stable across recovery (#1294)w
Use a per-thread proxy so in-flight tools continue through the latest recreated sandbox instead of holding a stale backend reference.
2026-05-11 16:03:38 -07:00
open-swe[bot]
ed7974f62d
feat: remove eyes reaction on Slack invocation (#1289)
Drops the add_slack_reaction call in process_slack_mention; the Slack
assistant status indicator and the agent's first reply still signal
acknowledgement. Linear and GitHub eyes reactions are unchanged.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
2026-05-11 15:07:02 -07:00
open-swe[bot]
85343fab63
feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322] (#1280)
* feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322]

Persist github_token_expires_at alongside github_token_encrypted, treat
expired cache entries as missing so we re-resolve before kicking off
runs, and invalidate the cached ciphertext on a downstream 401 so the
next invocation gets a fresh token instead of replaying a revoked one.

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* webapp: forward installation-token expiry to reviewer cache writes

The three reviewer-thread persist sites in webapp.py were calling
get_github_app_installation_token() (no expiry) and persist_encrypted_github_token
without expires_at, so cached App tokens were treated as never-expiring even
though they actually expire in ~1 hour.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 22:57:01 +00:00
Johannes du Plessis
094b2df939
feat: cross-provider model fallback on transient errors (#1281)
When the primary model raises a transient provider error (5xx, 429,
connection/timeout) the request is retried once against a fallback
model from the other provider. Anthropic primaries fall back to
OpenAI and vice versa. Also bumps the SDK max_retries from the
default 2 to 6 so quick blips stay on the primary and keep prompt
caching warm.

Triggered by 529 OverloadedError traces that ended runs silently
with no Slack/Linear/PR reply.
2026-05-08 15:35:13 -07:00
open-swe[bot]
a9331e78e6
fix: harden http_request SSRF guard against DNS rebinding [closes AB-2321] (#1277)
* fix: harden http_request SSRF guard against DNS rebinding [closes AB-2321]

Pin DNS resolution per request hop so urllib3's connection-time lookup
cannot rebind to a private IP after _is_url_safe validated a public one.

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: scope DNS pin to urllib3 with reference-counted install

Address review feedback: the previous version permanently overwrote the
process-global socket.getaddrinfo on first use. Now the patch targets
urllib3.util.connection.create_connection (much narrower blast radius),
and is installed/uninstalled via reference count so no global mutation
persists once no http_request calls are in flight.

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: forward timeout and socket_options through pinned create_connection

urllib3 calls create_connection with timeout positional and socket_options
as a keyword. The previous wrapper only read kwargs, silently dropping the
caller's connect timeout (so a slow validated IP could hang) and TCP
options like TCP_NODELAY. Accept both positionally and forward them to
the underlying socket.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 22:06:25 +00:00
Johannes du Plessis
b541f6d08a
slack: slim down the trace reply (drop greeting, suppress unfurl) (#1278)
Now that assistant.threads.setStatus carries the "is thinking…"
indicator, the per-run greeting phrase + auto-unfurl on the LangSmith
link are just visual noise stacking on top of it. The trace reply now
posts only `<url|View trace>` + a tip, with link unfurling off so the
smith.langchain.com card no longer appears.

Adds an `unfurl_links`/`unfurl_media` knob to post_slack_thread_reply_with_ts
(default-on to preserve behaviour for every other caller).
2026-05-08 14:51:28 -07:00
open-swe[bot]
148efeb269
feat: support TOKEN_ENCRYPTION_KEY rotation via MultiFernet [closes AB-2323] (#1275)
Threat model T6: Fernet token-encryption key had no rotation path. Rotating
TOKEN_ENCRYPTION_KEY immediately invalidated every github_token_encrypted
value in thread metadata, forcing re-auth or bot-token re-resolution.

Switch to cryptography.fernet.MultiFernet and parse TOKEN_ENCRYPTION_KEY as a
comma- or newline-separated ordered list (most-recent-first). New writes
encrypt under the first key; reads try every key. Deployers can prepend a new
key, let active threads roll over, then drop the old key.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-08 14:29:23 -07:00
Johannes du Plessis
3f9dbb6597
feat: add Slack reaction feedback to LangSmith (#1231)
* feat: add Slack reaction feedback to LangSmith

Record Slack reaction feedback against explicitly mapped LangGraph runs so user ratings are idempotent and tied to the message they reacted to.

* Address review feedback on Slack reaction → LangSmith feedback

- langsmith.py: drop lru_cache on _build_langsmith_feedback_clients so
  rotated keys / late env hydration are picked up; dedupe by (key, url)
  tuple instead of key alone so the same key pointing at different
  endpoints (cloud + self-hosted) builds both clients.
- langsmith.py: treat LangSmithNotFoundError on delete_feedback as
  success — out-of-order or redelivered reaction_removed events would
  otherwise loop forever on Slack's retry policy.
- slack_feedback.py: include channel_id in _feedback_key so the same
  message_ts in two channels can't collide on the same feedback id.
- slack_feedback.py: treat conflicting +/- reactions from one user as
  ambiguous (clear feedback) instead of averaging to a misleading 0.5.
- slack_feedback.py + slack.py + webapp.py: gate reaction handling to
  the user who triggered the run (stored in the slack_run_map mapping
  alongside run_id). Prevents bystanders in shared channels from
  polluting eval feedback.
2026-05-08 14:24:12 -07:00
Brace Sproul
1f8a83d4f8
feat: support allowed repos in addition to allowed orgs for webhook filtering (#1092)
* feat: add ALLOWED_GITHUB_REPOS env var for owner/repo-level webhook filtering

Previously only org-level filtering was supported via ALLOWED_GITHUB_ORGS.
This adds ALLOWED_GITHUB_REPOS for finer-grained control, allowing specific
owner/repo pairs to be allowlisted independently of org membership.

* fix: update test patches for _is_repo_org_allowed rename

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 13:26:30 -07:00
open-swe[bot]
37c194dc6d
feat: give agent self-awareness of its own repo (langchain-ai/open-swe) (#1266)
* feat: tell agent its source code lives at langchain-ai/open-swe

* fix: scope self-reference to only when user talks to the agent about itself

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-08 13:13:35 -07:00
Johannes du Plessis
d38f17ffe8
fix Slack assistant status endpoint (#1272) 2026-05-08 13:07:32 -07:00
Johannes du Plessis
743b2b9ba4
fix: recover from mid-run sandbox death (#1274)
* fix: recover from mid-run sandbox death

Recreate dead sandboxes during tool execution and stop repeated unrecoverable timeout loops with a user-facing notification.

* fix: count repeated sandbox recreations

Treat consecutive sandbox recreations as an unrecovered failure streak so outages cannot loop until the model-call limit.
2026-05-08 12:55:36 -07:00
open-swe[bot]
dc9a0b98da
feat: gate @open-swe mentions on public repos to org members (#1273)
Adds a webhook-level check so only members of $PUBLIC_REPO_ORG_GATE
(e.g. langchain-ai) can trigger Open SWE via mentions or review
requests on public repositories. Private repos remain governed by the
existing org/repo allowlists. Internal bots bypass the gate.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-08 11:38:29 -07:00
Johannes du Plessis
5d1020e8d7
fix: restore Co-authored-by trailer for triggering user (#1271)
The gh-cli migration removed agent/middleware/open_pr.py and
agent/tools/commit_and_open_pr.py — the only callers of
add_user_coauthor_trailer / add_pr_collaboration_note. Since then the
agent has been driving commits and PRs entirely via gh, with no
attribution back to the Slack/Linear/GitHub user who triggered the run.

Resolve the triggering user's identity in get_agent (reusing the
existing authorship helpers) and inject a Collaborative Attribution
section into the system prompt with the exact trailer and PR-body note
to use. The section is only rendered when an identity is resolvable, so
runs without a known triggering user are unchanged.
2026-05-08 10:42:15 -07:00
open-swe[bot]
96f97710ad
feat: add optional Slack Assistants API typing status indicator (#1269)
* feat: add optional Slack Assistants API typing status indicator

Mirrors OpenClaw's pragmatic approach: instead of rebuilding around
assistant_thread_started events, just opt into assistants.threads.setStatus
to show 'is thinking…' while the agent is working, and clear it when
post_slack_thread_reply lands. Gated behind SLACK_ASSISTANTS_API_ENABLED so
it can be toggled without touching code.

* fix(slack): drop redundant clear, add status heartbeat across model calls

- Slack auto-clears the typing indicator on bot post; remove the explicit
  assistants.threads.setStatus("") call from post_slack_thread_reply.
- The indicator expires after ~2 minutes; add a before_model middleware
  that refreshes it on every model tick so it stays visible across long
  agent runs. Reuses the existing slack_thread.{channel_id,thread_ts}
  configurable already plumbed for notify_step_limit.
- chat:write is sufficient on the bot token (assistant:write is on the
  way out per Slack docs); no scope or app-config change required.

* feat(slack): contextual status text + rotating loading_messages

- set_slack_assistant_status now accepts an optional loading_messages list
  (capped at 10 per Slack's API), surfaced via the assistants.threads.setStatus
  payload so Slack rotates through them client-side.
- The heartbeat middleware derives a contextual status from the last
  assistant message's tool calls (e.g. "searching the codebase…" after
  grep, "running commands…" after execute), falling back to the default
  "is thinking…" when no tool calls or unknown tool name.
- Adds a curated DEFAULT_LOADING_MESSAGES list passed alongside the
  contextual status on each refresh.

* fix slack assistant status lifecycle

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 10:21:55 -07:00
open-swe[bot]
5a845ba99f
feat: add Tip section to Slack trace-reply initial message (#1268)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-08 09:25:06 -07:00
open-swe[bot]
c4d9ae4a67
feat: add idle TTL and delete-after-stop sandbox lifecycle controls (#1265)
* feat: add idle TTL and delete-after-stop sandbox lifecycle controls

* bump langsmith>=0.8.3, lower default idle TTL to 10 min

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-08 00:30:14 -04:00
Johannes du Plessis
bdea1fb6da
feat(open-swe): always anchor findings to a single line (#1264)
* feat(open-swe): always anchor findings to a single line

GitHub renders multi-line review comment ranges as walls of context above
the comment, which buries the point. Match Devin Review's behavior and
always collapse `end_line` to `start_line` so each finding is anchored to
the single most relevant line.

Removes the (now unused) `MAX_FINDING_RANGE_LINES` cap and
`clip_finding_range` helper.

* fix(open-swe): collapse end_line before diff-range check
2026-05-07 20:26:59 -07:00
Johannes du Plessis
089db06135
feat(open-swe): collapse oversized finding ranges to start line (#1263)
Anchoring a finding to a 25-line range (e.g. an entire function) makes
GitHub render the whole block as context above the comment, burying the
review text. Cap ranges at 10 lines and collapse to the start line when
exceeded; small ranges (≤10 lines) still render as a multi-line anchor.
Updated the reviewer prompt to anchor tightly rather than range-select
the whole function.
2026-05-07 19:59:25 -07:00
Johannes du Plessis
0212a60f10
feat(open-swe): cap reviewer suggestions at 4 lines (#1262)
* feat(open-swe): cap reviewer suggestions at 4 lines

Long suggestion blocks read as the reviewer rewriting the code rather
than flagging an issue, which clutters PR comments. Steer the reviewer
toward description-only findings for non-trivial fixes, and enforce the
cap in `add_finding` / `update_finding` so suggestions over 4 lines are
dropped (the finding itself still publishes).

* fix(open-swe): don't clobber prior suggestion on over-cap update

`update_finding` was setting `suggestion=None` whenever clip_suggestion
dropped the input, which silently wiped any existing suggestion on the
finding. Distinguish the three cases: empty string clears, valid value
sets, over-cap value is rejected without touching the stored field.
2026-05-07 19:27:16 -07:00
Johannes du Plessis
7e1746f420
feat(open-swe): trigger reviewer from @open-swe review PR comment (#1259)
* feat(open-swe): trigger reviewer agent from `@open-swe review` PR comment

Mirrors the Slack `@open-swe review` flow on GitHub: a comment containing
`@open-swe review` (optionally followed by a PR URL) on a PR triggers the
reviewer agent. Without a URL it reviews the commenting PR; with a URL it
targets that PR. Works for `issue_comment`, `pull_request_review_comment`,
and `pull_request_review` events, gated by the existing reviewer repo
allowlist and reusing `trigger_pr_review_from_ref`.

* fix(open-swe): require URL after `@open-swe review`, don't swallow trailing text

The previous regex matched any non-whitespace token after `review`, including
across newlines. Comments like `@open-swe review\nthanks!` parsed as
`(True, "thanks!")`, which then failed PR-URL parsing and was silently
dropped — the user got no review and the comment never reached the regular
PR-comment handler.

Restrict the optional URL token to `https?://\S+` so non-URL trailing text
falls through to `process_github_pr_comment` instead of being eaten by the
review-command branch. Adds regression tests for the multiline and
trailing-word cases.
2026-05-07 17:04:35 -07:00
Johannes du Plessis
892347041f
feat(open-swe): post review summary to Slack on first review (#1258)
When a Slack user kicks off a PR review with `@open-swe review <pr-url>`,
the reviewer agent now posts a one-line summary back to the Slack thread
when it finishes — either "No issues found" or "found N potential
issue(s)" with a link to the GitHub review.

The reviewer agent has no Slack tools by design, so the summary is sent
host-side from `publish_review` after the GitHub review POST succeeds.
The Slack channel/thread_ts is persisted on reviewer thread metadata at
trigger time and read back on publish. Re-reviews triggered by push
events stay silent in Slack to avoid noise on the original thread.
2026-05-07 16:11:36 -07:00
Johannes du Plessis
88a55057ea
fix(open-swe): host-format the review summary, drop agent prose (#1257)
The reviewer agent was writing a 1–2 sentence "top-level take" as the
review body, which produced noisy paragraph-style summaries on PRs
("Reviewed the PR. The new ALLOWED_GITHUB_REPOS allowlist…"). Devin's
review comment is just a one-liner (`✅ No Issues Found` or
`**Devin Review** found N potential issue.`), and was preferred in the
internal A/B vs Graphite.

Drop the `summary` parameter from `publish_review`; render a fixed,
host-formatted body in `render_review_body` instead. Update the reviewer
prompt to forbid prose summaries.
2026-05-07 15:41:40 -07:00
Johannes du Plessis
03d9f23645
fix(open-swe): always post a review summary, even with no findings (#1256)
* fix(reviewer): log every push/close early-return so 'silent ignore' is debuggable

Pushes to PRs that haven't had a first review fall through the watch
handler because the reviewer thread doesn't have kind=reviewer set.
Without log lines on the early-return paths, this scenario was
indistinguishable from 'webhook reached the handler at all' in the
hosted log stream.

Now every early-return logs at info or debug:
- info when a real PR exists but the reviewer thread isn't set up
  (with a hint pointing at the trigger paths the user can use)
- info when the repo isn't in the reviewer allowlist
- debug for benign skips (non-branch refs, branch deletions,
  already-reviewed head_sha)

* fix(reviewer): always post a summary review, even with no findings

The publish_review tool gated POSTing on `inline_comments or summary`,
so when the agent called publish_review() with no args on a clean PR
the result returned `success: true` but no GitHub review was posted —
the user got silence instead of a "no issues found" comment.

- Drop the gate so publish_review always POSTs.
- Friendlier no-findings render: `**No issues found.**` when the
  findings list is empty, vs. `**No issues at or above \`<sev>\`
  severity.**` with hidden count when only sub-threshold findings
  exist. Agent summary renders below.
- Prompt now requires the agent to always pass a `summary` so the
  body is meaningful; calls out specifically not to skip on a clean PR.
2026-05-07 22:16:20 +00:00
Johannes du Plessis
378b95266e
feat: implement reviewer findings, publish_review, and watch mode (#1253)
* feat: implement reviewer findings, publish_review, and watch mode

Build out the reviewer agent end-to-end against the design in
REVIEWER_DESIGN.md:

- Findings as first-class state on the reviewer thread metadata
  (`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line
  ranges, suggestion text for ```suggestion blocks, github_review_comment_id
  for cross-run reconciliation, diff_hunk for UI rendering. Thread-level
  metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a
  future frontend can list reviewer threads via the langgraph SDK.
- Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff,
  compute_diff_line_set for in-diff validation, extract_diff_hunk for
  caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA
  diffs against the prepped repo.
- Tools: `add_finding` (validates against the diff line set so out-of-diff
  ranges fail at creation, not at GitHub-publish), `update_finding`,
  `list_findings`, `publish_review`. The reviewer agent's tool list is
  swapped from `[]` (direct shell `gh api` calls) to these four.
- Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`):
  one POST /reviews call with body + inline comments + ```suggestion blocks,
  per-comment IDs stored back on findings, GraphQL `resolveReviewThread`
  fired for findings transitioning open->resolved on a re-review.
- Reviewer graph: deterministic clone-or-fetch + checkout in the factory
  before the agent's first model call (warm- and cold-path symmetric);
  computed diff and in-diff line set passed via runnable config; system
  prompt rewritten for the single-evolving-findings model, severity ladder,
  in-diff-only discipline, and watch-mode reconciliation flow.
- Watch mode in webapp.py: `push` event + `pull_request` closed/reopened
  added to supported events. New `process_github_push_event` resolves the
  open PR for the pushed branch, gates on the reviewer thread's `watch`
  flag, builds a re-review configurable, and triggers a run on the same
  canonical thread. `process_github_pr_close` toggles watch on
  closed/reopened. `set_reviewer_thread_metadata` is called on first
  review to install `kind=reviewer` + PR identity + watch=True.
- Eval harness: target.py now extracts `add_finding` calls (mapped to the
  legacy {file, line, body, severity} shape the judge expects) and passes
  the right configurable so the prep step has base/head SHAs.
- Tests: new unit suites for findings helpers, diff parsing, finding tools,
  publish rendering + GraphQL resolve, and watch-mode webhook handlers
  (push triggers re-review only when watching, idempotent on unchanged
  head SHA, PR close disables watch). Updated existing reviewer-webhook
  tests to mock `set_reviewer_thread_metadata`.
- REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context
  into REVIEWER_DESIGN.md.

* fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL

Address PR #1253 review findings:

- compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag
  (`option no-prefix takes no value` — every prep run was failing
  silently and the agent saw an empty diff).
- compute_diff_in_sandbox grew a `merge_base` flag. First-review path
  now uses three-dot `base...head` (the merge-base diff GitHub renders
  on Files-changed) so we don't pick up changes that landed on the base
  branch after the PR diverged. Re-review delta keeps two-dot
  `last_reviewed_sha..head` since that's exactly the new commits.
- publish_review skips findings that already carry
  `github_review_comment_id`. Without this, watched re-reviews
  re-posted every previously surfaced finding, and only the most-recent
  duplicate's id would later resolve when the issue got addressed.
- fetch_review_comments URL now includes `{pull_number}` —
  `/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments`
  is the canonical endpoint; the old form 404s, so comment ids were
  never stored and watch-mode resolution couldn't run.

Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix`
flag in the executed command, and that publish_review does not re-post
findings whose `github_review_comment_id` is set.

* fix(reviewer): default publish cap from 15 to 4

A clean PR with one critical issue padded out by three lower-severity
findings is fine; fifteen is review spam. The agent can override per
call when a PR genuinely warrants more.
2026-05-07 14:48:43 -07:00
open-swe[bot]
5a01aff1c7
feat: only post Slack 'Working on it!' on first thread mention (#1250)
* feat: only post Slack 'Working on it!' on first thread mention

* feat: randomize Slack trace reply phrase

Pick from a small list of friendly phrases instead of always saying
'Working on it!' so the bot feels less robotic. Explicit messages (e.g.
'Taking a look...' from PR review path) are unaffected.

* adjust phrases

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-07 12:37:46 -07:00
Johannes du Plessis
bb836b58bc
fix: drop multitask_strategy=enqueue and forward image_urls on slack busy path (#1251)
The Slack mention path already routes mid-run messages through the
store-based queue + check_message_queue_before_model middleware (the
same path Linear uses), so multitask_strategy="enqueue" on runs.create
was a leftover no-op — that line is only reached when is_thread_active
returned False. The busy-path payload was also hardcoding image_urls=[]
which silently dropped any images attached to mid-run Slack mentions;
forward the resolved image_urls so the middleware can rebuild image
blocks like Linear does.
2026-05-07 11:55:14 -07:00
Johannes du Plessis
88d3659d00
start sandbox before proxy refresh (#1249) 2026-05-07 11:26:25 -07:00
Johannes du Plessis
da74342da4
add logging (#1248) 2026-05-07 10:19:53 -07:00
Johannes du Plessis
6ded489751
fix: reuse reviewer thread github token (#1247) 2026-05-06 18:09:39 -07:00
Johannes du Plessis
d060b8a863
fix: reduce GitHub webhook ignore log noise (#1246) 2026-05-07 01:01:33 +00:00
Johannes du Plessis
1319347dd9
feat: route Slack PR review requests (#1245)
* feat: route Slack PR review requests

Add a lightweight Slack review command path that starts the reviewer graph directly and gives the core agent a handoff tool when review requests are misrouted.

* fix: harden Slack PR review routing

* fix: validate Slack PR review URLs

* fix: preserve malformed GitHub review routing
2026-05-06 17:14:43 -07:00
Johannes du Plessis
65f6b4636b
feat(open-swe): remove github_comment and enable reviewer trigger (#1244)
* Use gh for reviewer inline comments

* Trigger reviewer on PR review requests

* Format reviewer webhook test

* Add GitHub repo allowlist

* Separate reviewer repo allowlist

* only implement allowlist for reviewer
2026-05-06 16:14:38 -07:00
Johannes du Plessis
96774f20ae
feat: move github workflows to gh cli (#1238)
* feat: move github workflows to gh cli

Use LangSmith proxy auth to support gh-driven GitHub workflows while removing custom GitHub wrapper tools.

* docker ignore + snapshot and docker image updates

* updated image and instructions

* removing open_pr if needed after agent call
2026-05-04 18:03:53 -07:00
Johannes du Plessis
13f5d8a1c9
fix: preserve existing PR descriptions (#1237)
* fix: preserve existing PR descriptions

* test: update existing PR label expectations
2026-05-04 11:43:32 -07:00
Johannes du Plessis
30530bf4d5
fix: preserve existing PR titles (#1236)
Keep automatic existing-PR updates from retitling pull requests, while documenting the explicit edit path for intentional title changes.
2026-05-04 09:35:08 -07:00
Brace Sproul
3405d145ac
feat: add edit_pull_request tool for editing PR titles/descriptions (#1063)
* feat: add edit_pull_request tool for editing PR titles and descriptions

* fix: patch auth flow in open PR middleware tests

* fix: support app token for editing PRs

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 18:12:41 -07:00
fuyua9
226e6c8263
fix(daytona): make sandbox snapshot configurable (#1220) 2026-05-01 22:51:34 +00:00
langsmith-forge[bot]
bd97678a5e
fix: prevent futile retry loop when commit_and_open_pr fails (#1210)
* fix: prevent futile retry loop when commit_and_open_pr fails with git/API errors

- Root cause: when git checkout or GitHub PR API fails, the tool returned a generic {"success": false} error with no signal that retrying is futile, causing the agent to loop 9-13+ times until hitting the 1000-step recursion limit
- Change: (1) git_checkout_branch now returns (bool, str) so the actual git error output is surfaced in the tool response; (2) checkout and PR creation failures now include "fatal": true and an explicit "Do not retry" message; (3) prompt.py COMMIT_PR_SECTION adds an explicit instruction to stop on fatal errors
- Verified: 109 unit tests pass, no regressions

* fix: skip PR safety net on fatal commit failures

* style(open_pr): ruff-format fatal retry skip condition

---------

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 22:05:22 +00:00
langsmith-forge[bot]
48965a82a6
fix: coerce malformed integer strings in read_file offset/limit params (#1216)
- Root cause: LLM occasionally generates strings like '1, 80' or '170, "limit": 60'
  for integer fields, causing a Pydantic ValidationError and wasting an LLM turn
- Change: add SanitizeToolInputsMiddleware in agent/middleware/sanitize_tool_inputs.py
  that extracts the leading integer from any string value in offset/limit before
  the call reaches Pydantic validation; registered before ToolErrorMiddleware in server.py
- Verified: 14 unit tests covering all three production trace patterns pass

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 14:29:48 -07:00
langsmith-forge[bot]
e5bc27a0ad
fix: notify users via Slack when agent hits model call step limit (#1204)
* fix: notify users via Slack when agent hits model call step limit

- Root cause: GraphRecursionError at 1000 steps bypassed all @after_agent
  middleware including open_pr_if_needed, leaving users with no notification
- Change: Added ModelCallLimitMiddleware(run_limit=60) to intercept gracefully
  before the hard recursion limit, and added notify_step_limit_reached
  @after_agent middleware to post a Slack thread reply when the limit fires
- Verified: 107 existing tests pass, no regressions

* fix: harden step-limit Slack notification

Ensure the step-limit notification runs after the PR safety net and cover the new middleware behavior with focused unit tests.

---------

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 14:24:25 -07:00