Commit graph

267 commits

Author SHA1 Message Date
Johannes du Plessis
39a26e16b5
fix: optimize agent thread lists (#1570)
* fix: optimize agent thread lists

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: refresh missing thread run status

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-18 10:09:22 -07:00
Suraj Bayas
4030001ebf
feat: validate LLM API keys on startup (#1438)
* feat: validate LLM API keys on startup

* fix: correct relative import for options module

* refactor: move imports to top of file

* style: fix linting and formatting issues

* refactor: scope LLM validation to local dev and rename function

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: open-swe[bot] <johannes@langchain.dev>
2026-06-18 09:24:46 -07:00
Johannes du Plessis
c07434a221
fix: allow read-only cross-user access to agent threads via Open in Web links (#1568)
The thread detail endpoint already returned metadata for non-owners, but
the transcript hydration endpoints (state, stream/events, history, pr-diff)
all asserted ownership and 404-ed. This caused the UI to redirect non-owners
back to /agents when they clicked an "Open in Web" link shared in Slack.

Dashboard login is already gated by ALLOWED_GITHUB_ORGS, so any logged-in
user is a trusted org member. This commit:
- Adds _thread_is_readable / _assert_thread_readable helpers that grant
  read access to any surfaced-source thread for authenticated users
- Relaxes read endpoints (state, stream/events, history, pr-diff, SSE
  stream) to use readable checks instead of ownership checks
- Keeps write endpoints (send message, cancel, delete, resolve, run
  commands) owner-only
- Adds an isOwner field to the thread summary so the frontend can render
  a read-only mode (hides the prompt bar, resolve/delete buttons)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-18 09:10:09 -07:00
Johannes du Plessis
98b824bd54
feat: add size caps for PR diff, fetch_url, Slack threads, pagination, message queue [closes OPE-51] (#1567)
* feat: add size caps for PR diff, fetch_url, Slack threads, pagination, message queue

Per-source byte/token caps with explicit truncation markers to prevent
unbounded payloads from blowing up LLM context/memory.

- reviewer_diff.py: cap PR diff at 200K chars with head+tail truncation
- fetch_url.py: cap markdownify output at 100K chars
- slack.py: cap thread message fetch at 500 messages
- github_comments.py: cap _fetch_paginated at 50 pages
- thread_ops.py: cap queued messages at 100 (drop oldest)

Closes OPE-51

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: compute diff line set from full diff, keep most recent Slack messages

Address PR review comments:

1. Truncated diffs rejected valid findings: fetch_pr_diff now returns the
   full diff; truncate_diff is called separately in reviewer.py so the
   line set used for add_finding/publish_review validation is computed
   from the complete diff, not the truncated prompt text.

2. Slack cap dropped recent thread context: fetch_slack_thread_messages
   now keeps the most recent SLACK_THREAD_MAX_MESSAGES messages (was
   keeping the oldest). The tool surfaces a truncation marker in the
   formatted output so the LLM knows the thread was truncated.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 15:40:59 -07:00
Johannes du Plessis
7dd758f845
feat: shared GitHub HTTP helper with retries, rate-limit handling [closes OPE-45] (#1565)
* feat: shared GitHub HTTP helper with retries, rate-limit handling, and sane timeouts

Introduces agent/utils/github_http.py — a single place for GitHub API HTTP
calls with 30s/10s-connect timeouts (vs httpx's 5s default), exponential
backoff with jitter, Retry-After header support, and 429/secondary-rate-limit
detection. Migrates the reviewer publish path (reviewer_publish.py,
reviewer_diff.py, github_checks.py, github_ci.py) from one-shot
httpx.AsyncClient() calls to the shared helper.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: don't retry transport errors on non-idempotent GitHub writes

POST/DELETE/PATCH can create side effects server-side even when the client
gets a timeout or connection reset. Only retry transport errors for
idempotent methods (GET, HEAD, PUT, DELETE). 429/5xx status codes are still
retried for all methods since the server explicitly did not process the
request.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: don't retry 502/504 on non-idempotent GitHub writes

502 (bad gateway) and 504 (gateway timeout) are ambiguous — the upstream
may have processed the write before the gateway returned an error. Only
retry these for idempotent methods. 429 and 503 are still retried for all
methods since the server explicitly did not process the request.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 14:45:57 -07:00
Johannes du Plessis
8e39f62122
feat: activate PR babysitting UI toggles for autofix and trigger mode (#1561)
* feat: activate PR babysitting UI toggles for autofix and trigger mode

Remove the "coming soon" gating on the Autofix Mode, Autofix Severity
Threshold, and Trigger Mode controls in the review settings page so
admins can enable CI auto-fix and review-comment resolution on PRs
that Open SWE opens. The backend (ci_autofix.py, webapp.py webhook
routing) was already fully wired — only the UI was disabled.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: simplify autofix to on/off toggle, remove severity threshold

Replace the four-level AutofixMode (off/low/medium/high) and the
autofix_severity_threshold setting with a single boolean
autofix_enabled toggle. The severity threshold was leftover from the
reviewer finding-severity model and does not apply to CI autofix;
the agent should fix any failing CI and resolve any comments on PRs
it opens.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: move autofix toggle to per-user profile, remove team-level setting

The autofix toggle is now per-user (auto_fix_ci in the user profile)
instead of team-level (admin-only). This uses the existing auto_fix_ci
field that was already in ProfileUpdate but never wired up.

Changes:
- ci_autofix.py: check per-user auto_fix_ci profile flag after
  resolving the agent thread's github_login, instead of checking
  team-level autofix_enabled before knowing the PR
- webapp.py: removed early is_autofix_enabled() webhook gates; the
  per-user check now happens in ci_autofix.py once the thread is found
- team_settings.py: removed autofix_enabled field, is_autofix_enabled()
- cloud-agents.tsx: enabled the auto_fix_ci toggle (was comingSoon)
- review.tsx: removed the admin-level autofix switch
- Updated tests and AGENTS.md

The agent graph (not the reviewer) is what gets dispatched - this was
already correct in ci_autofix.py line 223: client.runs.create(
thread_id, "agent", ...).

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: batch PR babysitting events

Remove the leftover trigger-mode gate from PR babysitting and batch new CI/review events while an agent run is already active so the running agent can handle the latest PR state before finishing. Also moves review-feedback permission checks behind the per-user opt-out and applies the auto-fix profile gate to merge-conflict babysitting.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: consume batched babysitting events

Teach the agent queue middleware to turn pending PR babysitting metadata into an injected instruction for the active run, so batched CI/review events are not dropped while still avoiding duplicate run creation.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: address review findings in PR babysitting batching

- Route batched events through the LangGraph store (read in-process by the
  message-queue middleware) instead of a per-model-call threads.get on every
  agent thread.
- Only record an attempt / mark the head SHA handled on a real dispatch, not
  on a batch, so an event isn't permanently dropped if the in-flight run ends
  before consuming it.
- Carry the reviewer's comment through batched review feedback instead of
  replacing it with a generic re-check nudge.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 14:12:04 -07:00
Johannes du Plessis
60b7f4677a
feat: add user-scoped Currents.dev API key for e2e test investigation (#1566)
* feat: add user-scoped Currents.dev API key for e2e test investigation

Allow each user to configure their own Currents.dev API key on the
Profile Settings page. The key is encrypted at rest in a per-user
LangGraph Store namespace and feeds server-side read-only tools that
query the Currents REST API (runs, instances, projects, test results)
so agent runs can inspect e2e test failures including screenshots and
DOM snapshots.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: add pagination cursors to currents_list_project_runs

Address review feedback: forward starting_after/ending_before cursor
parameters to /projects/{projectId}/runs so the agent can paginate
beyond the first 50 results.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 13:59:55 -07:00
Johannes du Plessis
efe07486e2
feat: weight merged PRs above LOC in usage leaderboard sorting (#1563)
Move merged_prs to the primary sort key in the agent usage leaderboard,
ahead of agent_loc, prs_opened, and agent_runs. Update the UI description
to reflect the new ranking order.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 10:59:42 -07:00
Johannes du Plessis
b19804536c
feat: handle images sent to non-vision models in Slack, Linear, and web UI (#1560)
* feat: handle images sent to non-vision models in Slack, Linear, and web UI

Add vision capability checks across all image input paths. When a user
sends images to a text-only model (e.g. GLM 5.2, DeepSeek V4 Pro), the
images are now skipped and a warning is injected into the prompt instead
of sending unsupported content to the model.

- Slack: resolve model at webhook time, skip image fetch + add warning
- Linear: same pattern as Slack
- Queued message middleware: read resolved model from thread metadata,
  strip images from queued payloads for text-only models
- Web UI: disable submit + show inline warning when images are attached
  to a non-vision model selection
- Shared: resolve_agent_model_id helper + vision_not_supported_warning

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: mock resolve_agent_model_id in Slack mention test

The test_process_slack_mention_queues_active_thread_message test was
missing a mock for the new resolve_agent_model_id call added to the
Slack webhook handler, causing a TypeError when image URLs triggered
the model resolution path.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: include vision warning in queued payload for text-only models

Update the prompt variable (not just content_blocks) before clearing
image_urls so the queued payload also carries the warning text when a
Slack/Linear follow-up arrives while the thread is busy.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 09:12:52 -07:00
Johannes du Plessis
46e8982b29
feat: add max effort level for GLM 5.2 (#1562)
* feat: add max effort level for GLM 5.2

GLM 5.2 supports a 'max' thinking effort level (recommended for coding
tasks per Z.ai/Fireworks docs). Add it to the model's effort list so it
surfaces in the profile editor and maps to reasoning_effort=max.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: restrict GLM 5.2 efforts to none, high, max

GLM 5.2 only supports non-thinking (none), high, and max effort levels.
Remove low and medium which the model does not support.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 09:10:22 -07:00
Johannes du Plessis
0927f2dd9c
feat: Run reviewer eval in a GitHub Action; dashboard becomes read-only (#1556)
* Run reviewer eval in a GitHub Action; make dashboard a read-only progress view

The dashboard launched the eval as a subprocess inside the serving deployment
worker, so a container recycle killed long runs and discarded results that had
already completed server-side. Move the harness to a workflow_dispatch Action
(run on prod). run_eval now publishes status/progress/log-tail to the LangGraph
store record the dashboard reads, so /admin/evals stays a live view; a killed
Action surfaces as failed via the stale-heartbeat reconcile.

* reviewer_eval workflow: pass inputs via env, no shell interpolation

Addresses the reviewer finding: workflow_dispatch string inputs were
interpolated into the run: block (limit unquoted), allowing shell injection in
a job holding LANGSMITH/ANTHROPIC keys. Pass inputs through env and reference
quoted "$VARS"; validate limit is numeric and build its flag in bash.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 19:38:36 -07:00
Johannes du Plessis
e58b609b2f
fix: Simplify review explanation: full-width, plain prose, no diff links (#1547)
* Simplify review explanation: full-width, plain prose, no diff links

* Update _build_prompt test for plain diff fences

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 15:32:24 -07:00
Johannes du Plessis
77bc120583
feat: add GLM 5.2 model option (#1543)
* feat: add GLM 5.2 model option

Add GLM 5.2 to the supported model list so it's selectable in the
profile editor and as a team/per-thread model.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: remove GLM 5.1 model option

Remove GLM 5.1 from the supported model list now that GLM 5.2 is available.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 15:31:10 -07:00
Johannes du Plessis
801f93b4de
feat: AI-sorted PR review view with diff grouping (#1544)
* feat: AI-sorted PR review view with diff grouping

Group a PR's changed files into a logical top-to-bottom walkthrough via a best-effort structured-output LLM pass, kicked off concurrently with the reviewer run (~0 added latency) and persisted on reviewer thread metadata. Results render in the review UI behind an AI sorted / file tree toggle that persists across PRs; the view falls back to the file tree when groups are absent or stale.

Adds a grouping-model team default (inherits the Reviewer subagent model when unset) and the admin RolePicker for it.

* feat(reviews): richer AI-sorted explanations + sidebar polish

Sidebar group rows get Devin-style spacing (dividers, padding), a per-group file list (click a file to jump to its diff), inline-code chips in titles, and an accent Read explanation link.

Group explanations are now rich markdown: the grouping prompt feeds per-hunk line ranges and asks for inline code, a short code block, and [path:line](#loc=...) references. The Markdown renderer turns those #loc= links into in-page buttons that scroll the diff to the hunk and highlight the range, reusing the existing selectedLines path.

* fix(reviews): drop stale diff groups from the AI-sorted view

When groups were generated for a previous head, a persisted "ai" view in localStorage still rendered the outdated walkthrough. groupedView now returns null on diff_groups_stale, so the file-tree fallback is used and the view toggle hides until fresh groups arrive.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 13:59:37 -07:00
Johannes du Plessis
e3025dee77
Reviewer eval admin: configurable runs + stacked form layout (#1540)
Drive dashboard-triggered reviewer eval runs with per-run model, effort,
score mode, severity threshold, cap, limit, and concurrency overrides, plus
per-example start/finish/error logging in the eval target.

Rework the admin eval form from the label-left/control-right SettingsRow
(which crushed the description column when packing 3-4 wide inputs) into
stacked field groups with captioned inputs in a responsive grid.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 10:12:35 -07:00
Caroline di Vittorio
554d754585
feat: link PR attribution footer to the originating thread (#1539)
The "Made by Open SWE" PR footer linked to the generic dashboard
homepage. Point it at the dashboard thread that generated the PR
(/agents/<thread_id>), falling back to the homepage when no thread
id is available.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 09:50:18 -07:00
Johannes du Plessis
911c835c2a
feat: chat with your PR on the review page (#1534)
* feat: chat with your PR on the review page

Add a sandbox-less `chat` graph that answers questions about a single PR
from its diff, the published review findings, and read-only GitHub access.

- agent/chat.py: deepagents graph, no sandbox (default StateBackend, file
  mutation + execute tools excluded). PR context is seeded as virtual files
  under /pr/; a repo-scoped App token is resolved in-graph.
- tools: read_repo_file, search_repo_code, list_review_findings.
- dashboard/review_chat_api.py + routes: per-user chat thread, LangGraph
  stream/commands/state/history proxy pinned to the chat assistant, seeds
  diff/findings/overview on first run. Gated by repo access.
- UI: Chat tab wired to a chat-scoped StreamProvider (replaces Coming Soon).

* feat: admin setting for review-chat default model

Add a 'Open SWE Review Chat' default to team settings (default_chat_model /
default_chat_reasoning_effort). get_team_default_model("chat") inherits the
Agent default when unset; the chat graph resolves through it. Admin RolePicker
gains an 'Agent default' inherit option that clears the override.

* feat: multi-conversation review chat (tabs, new chat, history)

Replace the single per-PR chat thread with multiple per-user conversations:
- threads minted client-side; first message persists with a title derived
  from the prompt.
- list + delete endpoints; chat panel gets a tab strip (history), new-chat
  (+), close (x), refresh, an intro greeting, and suggested prompts.
- get_review_chat now returns availability only (ids are client-minted).

* ui fixes

* ui: review-chat history dropdown, full-width AI replies, resizable side panel

* fix(review-chat): enforce per-user thread ownership on proxy endpoints; reseed PR context on head change

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 17:17:30 -07:00
Johannes du Plessis
ddd7190864
feat: let the agent stop naturally without forced tool calls (#1535)
Remove the hardcoded "call a tool every turn" instruction from the system
prompt and delete the ensure_no_empty_msg middleware that re-injected no_op /
confirming_completion tool calls. The agent now ends its turn naturally when
the model emits a final message with no tool call, which avoids needlessly
extending trajectories (and token spend) on tasks that are already complete.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 14:54:01 -07:00
Johannes du Plessis
7397ff93ba
feat: CI auto-fix and PR babysitting for agent PRs (#1530)
* feat: CI auto-fix and PR babysitting for agent PRs

Watch CI failures and review feedback on PRs Open SWE opened, then dispatch
confidence-gated fix runs on the originating agent thread. Adds CI webhook
ingestion (check_run/check_suite/workflow_run/status), a per-PR @open-swe
autofix on|off toggle, auto-response to review comments, and a polling
ci_monitor graph that also flags merge conflicts. Gated by the existing
autofix_mode/trigger_mode settings, the enabled-repos opt-in, base-branch and
human-commit skip rules, dedupe, and a per-PR attempt cap.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: address review feedback on CI auto-fix

- Security: gate the no-mention review-feedback path on author trust —
  require a trusted author_association (OWNER/MEMBER/COLLABORATOR) plus a
  GitHub write/maintain/admin permission check before dispatching a
  write-capable agent run, preventing privilege escalation from
  read/triage/outside reviewers.
- Auth: reuse the originating PR thread's source + login/email when
  dispatching fix runs so the GitHub-token resolver authenticates them in
  non-bot-token deployments (bespoke github_ci source failed to resolve).
- Docs: document the Commit statuses: Read-only permission required for the
  Status webhook event.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 13:53:50 -07:00
Johannes du Plessis
be2fe7131b
feat: live reviewer eval logs on a dedicated admin page (#1527)
* feat: stream reviewer eval logs on a dedicated admin page

Stream the eval subprocess output into a rolling log_tail and persist it
during the run (was only captured at exit), so the live output is visible
while the eval runs. Move the eval runner off the admin page onto its own
/admin/evals page (linked like Review Style Prompts) with a live log viewer.

* chore: drop unrelated SSR-register drift from generated route tree

* fix(ui): pre-bundle workbox-window to stop dev re-optimize reload

The PWA service worker (devOptions.enabled) pulls workbox-window, which
Vite discovers after first render and re-optimizes, forcing a reload that
cancels in-flight code-split route imports (Failed to fetch dynamically
imported module). Pre-bundling it via optimizeDeps.include avoids the
mid-session reload.

* fix(ui): suppress html hydration warning for pre-hydration theme script

The inline theme script sets class="dark"/color-scheme on <html> before
React hydrates, so the prerendered HTML never matches. suppressHydrationWarning
on <html> silences the (expected) one-level attribute mismatch.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 11:29:16 -07:00
Johannes du Plessis
eb18b07b20
feat: trigger reviewer evals from the admin page (#1524)
* feat: trigger reviewer evals from the admin page

Add an admin-only "Reviewer eval" section + endpoints that launch the
reviewer benchmark as an isolated subprocess against the running
deployment, with live status and the LangSmith experiment link. Route
eval traces to a dedicated open-swe-evals project so they stay out of
the production tracing project.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: reconcile reviewer eval status via heartbeat, not local process

The persisted record is shared across workers but _PROCS is process-local.
The owning worker now refreshes a heartbeat while the subprocess runs, and
status is only reconciled to failed once the heartbeat is stale, so a poll on
a worker without the local handle no longer kills a live run (and a duplicate
start is rejected across workers).

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 09:48:12 -07:00
Johannes du Plessis
7f52fce2ea
Remove Kimi K2.6 from supported model selector (#1522)
Kimi K2.7 is now the supported version, so drop the older K2.6 entry.
Update the related fireworks provider test to use K2.7 instead.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 09:39:08 -07:00
Johannes du Plessis
889b2663d9
fix: point View trace link at the correct tracing project (#1521)
* fix: resolve trace URL project id by tracing project name

Graphs were split into separate LangSmith tracing projects
(open-swe-agent, open-swe-review) but the "View trace" link still used
a single fixed project-id env var pointing at the old combined project.
Resolve the project id from the tracing project name so agent and
reviewer links point at their respective projects, falling back to the
env var when resolution is unavailable.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* docs: document per-graph tracing projects for trace links

View trace links now resolve project IDs from the open-swe-agent /
open-swe-review project names. Document this so fresh deployments create
the right projects instead of relying on a single project ID.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-13 09:37:15 -07:00
Johannes du Plessis
38dcc37040
feat: add Kimi K2.7 model option (#1515)
* feat: add Kimi K2.7 model option

Surface Moonshot's newly released Kimi K2.7 in the model picker, routed
through Fireworks following the existing kimi-k2pN convention, with the
same none/low/medium/high reasoning efforts (default high) as K2.6.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: correct Kimi K2.7 Fireworks model id to kimi-k2p7-code

Fireworks now hosts the K2.7 release as
accounts/fireworks/models/kimi-k2p7-code (the `kimi-k2p7` slug 404s).
Point the model option and its test at the live model id.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: drop unsupported "none" effort for Kimi K2.7

Kimi K2.7 Code is a thinking-only model — Moonshot documents that it does
not support non-thinking mode (passing thinking={"type":"disabled"} or
reasoning_effort="none" errors). Removing "none" from the advertised
efforts so the picker can't send an invalid Fireworks request.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-13 08:10:04 -07:00
Johannes du Plessis
bd5d5b24d4
fix: point reviewer "Open in Web" link to the review page (#1519)
The top-level review comment's "Open in Web" link pointed at the
agent thread (/agents/{thread_id}). Point it at the dashboard review
detail page (/agents/reviews/{owner}/{repo}/{number}) instead.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-12 14:15:58 -07:00
Johannes du Plessis
67239bb2c3
perf: speed up Reviews list (My PRs + All PRs) (#1518)
* perf: speed up Reviews list (My PRs + All PRs)

Push the "My PRs" author filter into the threads.search metadata
(pr.author containment) instead of scanning up to 1000 reviewer threads
in Python, and replace the per-repo GitHub access check (an N+1 of
sequential GET /repos calls) with a single per-login accessible-repo set,
cached for 60s. Detail endpoints still re-validate access live.

Frontend: prefetch the inactive tab and adjacent page on hover/focus so
tab switches and pagination are instant.

* fix: don't cache repo access for the reviews list

The /reviews list is an authorization boundary for private PR metadata
(repo/PR titles, branches, authors, finding counts). A cross-request TTL
cache on the accessible-repo set could surface that metadata for up to
60s after a user lost repo access. Resolve the set fresh per request
instead — still a fixed, repo-count-independent burst of GitHub calls
(no per-repo N+1), with no staleness.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-12 14:14:50 -07:00
Caroline di Vittorio
e8bb6b497b
feat: mark threads as resolved to hide from sidebar (#1500)
* feat: mark threads resolved to hide from sidebar [closes resolve-threads]

Add a resolved flag (thread metadata) so users can clear finished
threads from the sidebar without deleting them. Resolved threads move to
a collapsible "Resolved" group (capped at 20 with a "Show all" link) and
are fully searchable/filterable + paginated on a new /agents/threads
page driven by URL query params. Sending a new message auto-unresolves.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: refresh paginated agent thread lists

Invalidate all agent thread list queries when thread lifecycle events change sidebar data, keeping the new paginated sidebar cache in sync.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-06-12 12:01:42 -07:00
Johannes du Plessis
09d5d00e59
feat: render Reviews page diffs with pierre MultiFileDiff (#1517)
* feat: render Reviews page diffs with pierre MultiFileDiff

The Reviews detail page used a hand-rolled hunk renderer with no syntax
highlighting. Switch it to the same pierre MultiFileDiff + theming the agent
chat git panel uses.

- review-diff API now returns full original/modified file contents instead of
  hunks, via a shared build_pr_diff_files helper extracted from thread_api
- findings render as right-anchored markers (pierre line annotations); focus
  highlight uses selectedLines; floating finding card still anchors to the marker

* feat: auto-collapse a review diff card when marked as viewed

* fix: URL-encode file path in Contents API fetch

Filenames containing reserved URL characters (#, ?) were truncated, so those
files rendered as empty/unrenderable. quote(path, safe='/') preserves the path
separators while escaping the rest.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-12 11:03:51 -07:00
Johannes du Plessis
258b5b4034
feat: disable out-of-diff reviewer findings (#1516)
* feat: disable out-of-diff reviewer findings

PR reviews were surfacing findings about code outside the PR's changed
lines, which read as random/off-topic noise. Reject out-of-diff findings
at add_finding and stop surfacing them in publish_review so only findings
anchored to changed lines reach the PR.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: align reviewer bar with out-of-diff rejection

The filing bar still permitted proven regressions in files absent from
the diff, but add_finding now rejects those, so a concrete regression
could be silently dropped after a failed tool call. Update the bar and
the "Do NOT file" list so the prompt only directs in-diff findings.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-12 10:19:30 -07:00
Johannes du Plessis
aa98d195a8
feat: route graphs to separate LangSmith tracing projects (#1508)
* feat: route graphs to separate LangSmith tracing projects

* docs: point graph entrypoint references at traced wrappers

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 17:57:16 -07:00
Johannes du Plessis
3e105f5027
fix: Reviews tab — anchored finding card, paginated list, file tree truncation (#1507)
* fix: Reviews tab — anchored finding card, paginated list, file tree truncation

- Finding card now tracks the diff anchor while scrolling instead of staying
  frozen in the viewport; auto-hides when its diff card collapses (including
  collapse via mark-as-viewed) and on click outside
- Checks section capped with max height + scroll
- /reviews paginated (page size 20) with has_more, filtered to the current
  user's PRs by default with an All toggle; PR author login now stored in
  reviewer thread metadata
- File tree truncation marker overlapped filenames because the sidebar bg
  was transparent; use the opaque sidebar color

* fix: finding card tracks anchor 1:1 while scrolling

Drop the vertical viewport clamp — it pinned the card at the clamp
boundary while the highlighted lines kept scrolling, breaking the
attachment.

* fix: anchor finding card with Base UI popover

Replace manual fixed-position tracking (laggy: setState per scroll
frame) with a Popover anchored to the finding's diff row. Floating UI
tracks the anchor outside React renders, so the card moves 1:1 with
the content and scrolls out of view with it. Unanchored findings keep
the fixed top-right card.

* fix: lock finding card to diff scroll

Replace the Base UI popover (async repositioning, paints a frame behind
native scroll) with a card absolutely positioned inside the scroll
container, so it scrolls with the diff in the same compositor frame.
Scroll moves to the ReviewBody root, side panel becomes sticky.
Position recomputes only on layout shifts via ResizeObserver.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 17:52:20 -07:00
Johannes du Plessis
bfcb714c5c
feat: PR review page (#1495)
* feat: PR review page

* feat: manual review trigger card on admin page

* feat: finding focus mode with floating card and dimmed backdrop

* fix: dim only info panel, subtle block highlight for focused finding

* refactor: move PR reviews into agents page with file-tree sidebar

* fix: address review feedback

- paginate list_reviews until enough accessible records collected
- include rename/metadata-only files in diff API with empty hunks
- remount ReviewBody on head_sha change so viewed-files state resets

* fix: review tree background + finding card anchoring

* fix: remove no-op PR size chip and dead check links

* fix: parallel diff fetch, complete reReview type, shared github_headers

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 16:11:10 -07:00
Johannes du Plessis
194e9cd60e
feat: link Slack thread/Linear ticket in PRs and append ticket to title (#1504)
* feat: link Slack thread/Linear ticket in PRs and append ticket to title

When opening PRs, include any referenced Slack thread or Linear ticket
in the description and append the resolvable ticket number to the title.
Adds a Slack permalink to the webhook prompt so the agent has a
ready-to-link thread URL.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: use [closes <TICKET>] format in PR title suffix

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: deterministically append source refs to private-repo PRs only

Move the Slack/Linear source-reference linking out of the prompt and into
open_pull_request, gated to private repos so private Slack thread URLs and
Linear identifiers are never published to a public PR. Reverts the prompt
instruction and webhook permalink injection in favor of this server-side append.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 15:52:40 -07:00
Johannes du Plessis
ce3af2d9b2
fix: reviewer can silently review a stale checkout on reused sandboxes (#1503)
prepare_review_repo is best-effort, but the system prompt unconditionally
told the agent the repo is checked out at the PR head. On a reused sandbox,
a dirty worktree (or a transient fetch failure) makes 'git checkout <sha>'
fail; prep returns False, the old checkout stays in place, and the reviewer
confidently reads stale code — observed as 'I rechecked the current head'
replies quoting pre-push code.

- checkout with --force so leftover worktree state can't block it, verify
  HEAD matches the requested sha, tolerate 'git fetch --all' failures
- when prep fails, the prompt now warns the checkout may be stale and tells
  the agent to re-fetch/checkout (or fall back to API file contents) before
  trusting local files

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 13:42:10 -07:00
Johannes du Plessis
b8035a4758
fix: stream proxy preflight + optimistic thread staleTime (#1502)
* fix: preflight stream proxy before SSE starts, let optimistic thread survive refetch

* fix: seed sidebar thread details as stale so mark-viewed fetch still fires

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 13:07:24 -07:00
Johannes du Plessis
8fa98398fa
fix: settle incomplete review check as neutral, not failure (#1501)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 13:03:34 -07:00
Johannes du Plessis
e0678e8c01
feat: track PR lifecycle state per thread for sidebar (#1492)
* feat: track PR lifecycle state per thread for sidebar

Persist a PR's draft/open/merged/closed state on the agent thread and keep
it in sync as PR webhooks fire, so the Agents UI can show per-thread PR
status the way Cursor does. open_pull_request now records the initial
draft/open state and pr_title; thread summaries expose diffStats; and the
PR webhook refreshes pr_state on close/reopen/draft/ready transitions.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: consolidate PR state mapping into shared derive_pr_state helper

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 12:21:59 -07:00
Johannes du Plessis
9a2f68e99d
fix: auto-recover from expired GitHub refresh tokens (#1491)
* fix: auto-recover from expired GitHub refresh tokens

When a user's GitHub OAuth refresh token was permanently dead (revoked or
expired), token refresh failed but get_valid_access_token still handed back
the known-stale access token, so dashboard GitHub calls kept 401ing until the
user manually logged out and back in.

Now we distinguish unrecoverable refresh failures (bad_refresh_token /
unauthorized_client) from transient ones: on an unrecoverable failure we drop
the dead stored authorization and return None, so callers prompt a clean
re-login. Transient failures still fall back to the stored token.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: don't delete fresh re-auth when stale refresh fails

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 12:05:27 -07:00
Johannes du Plessis
f08177aed7
fix: harden origin parsing and home prompt submit failure (#1499)
- _origin_of: treat invalid ports as invalid origin instead of letting
  urlparse ValueError turn CSRF rejections into 500s
- AgentsHome: reset submitting/draft when stream.submit rejects before a
  thread id is minted, so the prompt isn't left disabled

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 11:54:25 -07:00
Caroline di Vittorio
8f983e9e40
feat: add View PR link in git panel header (#1498)
* feat: add View PR link in git panel header

Surface a clickable "View PR" link in the agent git panel header so
users can jump straight from a chat thread to its pull request instead
of hunting for the URL in the conversation. Uses the pr.url already
captured in thread metadata.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: use shared buttonVariants for View PR link

Style the View PR link with the shared buttonVariants (outline/sm)
helper instead of hand-rolled classes, matching the existing
anchor-as-button pattern used in login.tsx and AutomationsList.tsx.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: show PR diff in git panel diff tab

Adds /threads/{id}/pr-diff endpoint fetching PR files + contents from GitHub;
git panel prefers it and falls back to the live message-derived diff.

* feat: auto-expand git panel when a PR lands mid-run

Panel still defaults to collapsed and remembers manual toggles; the
auto-expand is ephemeral and not written to localStorage.

* fix: git panel always starts collapsed

Drop localStorage persistence of collapsed state; panel opens only via
manual toggle or when a PR lands mid-run, and re-collapses on thread switch.

* fix: address PR review findings

Require user OAuth token for pr-diff (no app-token fallback) so GitHub
enforces current repo access; render placeholder for binary/oversized files.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 11:48:53 -07:00
Johannes du Plessis
da20e57bcc
fix: refresh sandbox GitHub proxy token before mid-run expiry (#1496)
* fix: refresh sandbox GitHub proxy token before mid-run expiry

GitHub App installation tokens expire after exactly 1 hour. The LangSmith
sandbox proxy was configured once at run start with a snapshot of that
token, so runs longer than ~1h hit 401s on every gh/git call. Record the
proxy token's expiry per thread and add a before-model hook that
re-configures the proxy with a fresh token when it nears expiry.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve repo-scoped proxy token on mid-run refresh

Reviewer runs mint a repository-scoped installation token. Record the
repo scope per thread alongside the expiry so the before-model refresh
re-mints a token with the same scope instead of an installation-wide
token, avoiding privilege expansion on long reviewer runs.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: update passthrough stub for github_proxy_repositories param

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 10:59:21 -07:00
Christian Bromann
abf354bb05
feat(open-swe): stream agent chat via @langchain/react v2 protocol (#1475)
* feat(dashboard): stream agent chat via @langchain/react v2 protocol

Replace the bespoke SSE + React Query polling path with LangGraph’s
v2 event stream through credentialed dashboard proxies. Run starts go
through stream commands; mid-run follow-ups still queue via /messages.

* fix import path

* fix tests after rebase

* format

* PR feedback

* improved model fallback

* fix image handling

* embrace sdk

* cleanup

* cr

* more cleanup

* fix cors

* harden security

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 09:54:35 -07:00
Johannes du Plessis
8c7b789122
feat: add right-click open trace context menu for threads (#1490)
* feat: add right-click open trace context menu for threads

Add a context menu to sidebar thread rows exposing an Open trace action
that links to the LangSmith thread trace, surfaced via a new traceUrl
field on the thread summary.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: drop stray yarn artifacts from PR

Remove accidentally committed ui/.yarnrc.yml and ui/.yarn/install-state.gz
that broke immutable installs against the v1 lockfile, and gitignore yarn
artifacts to prevent recurrence.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: reuse thread_id var, drop redundant danger-color fallback

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 15:40:43 -07:00
Johannes du Plessis
259ec57186
fix: keep review check on follow-up commits, drop neutral conclusion (#1486)
* fix: keep review check visible on follow-up commits and stop using neutral

The "Open SWE Review" check completed as `neutral` whenever findings were
surfaced, which GitHub renders as a confusing "neutral check" group. Always
complete it as `success` (informational/non-blocking), matching Devin and
Corridor — the finding count stays in the title and findings post as comments.

Also create a fresh check run on the new head SHA in the push re-review path:
GitHub only shows check runs on a PR's current head, so the check vanished
after a follow-up push (and the stale id settled on an outdated commit).

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: surface settled review check when push leaves diff unchanged

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 13:33:34 -07:00
Johannes du Plessis
e8770e1004
feat: report Open SWE Review as a PR check run (#1484)
* feat: report Open SWE Review as a PR check run

Auto-review dispatch now creates an in-progress 'Open SWE Review' check run
on the PR head SHA; publish_review completes it (neutral with findings,
success when clean). An after-agent hook fails the check if the run dies
before publishing. Requires the GitHub App's Checks: Read & write permission;
all calls are best-effort so a missing permission never breaks reviews.

* fix: address review feedback on check-run settling

Keep review_check_run_id when the completion PATCH fails so a later
publish or the after-agent hook can retry instead of hanging the check;
count out-of-diff findings toward the check conclusion.

* fix: retry failed check completion with the real publish conclusion

A transient PATCH failure after a successful publish previously left the
check id for the after-agent hook, which settled it as 'failure'. Persist
the intended result as review_check_pending_result and have the hook
prefer it over the generic failure fallback.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 11:53:48 -07:00
Johannes du Plessis
370644ae8b
chore: hide Fable 5 model from supported options (#1483)
Fable 5 is currently unusable with our API key. Remove it from the
selectable model list; it can be re-added later.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 11:18:12 -07:00
Johannes du Plessis
f539962c73
feat: server-side Datadog/LangSmith observability tools + team creds [closes OPE-54] (#1476)
* feat: server-side Datadog/LangSmith observability tools + team creds

Add team-wide observability credential settings (Datadog DD_SITE/API/APP
keys, LangSmith API key) stored encrypted server-side, with an admin
dashboard section to connect/disconnect each provider. When connected,
get_agent loads read-only observability tools server-side: Datadog via its
hosted MCP server (langchain-mcp-adapters, toolsets=core) and LangSmith
read tools (langsmith_get_trace, langsmith_list_runs). Credentials live in
the LangGraph server process and are never exposed to the sandbox.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: address review on observability tools

Address PR review feedback:
- Authorize observability tools per triggering user (admins + the
  OBSERVABILITY_AUTHORIZED_EMAILS allowlist) so prompt-injected runs from
  untrusted contributors can't reach team Datadog/LangSmith data.
- Use the documented Datadog MCP auth headers DD_API_KEY / DD_APPLICATION_KEY.
- Store each provider's credentials under its own store key to avoid a
  read-modify-write race dropping the other provider on concurrent saves.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: async email resolution in observability authorization gate

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 11:07:42 -07:00
Johannes du Plessis
8461979b0d
fix: honest publish_review reporting + structured thread-not-found errors (#1481)
* fix: honest publish_review reporting + structured thread-not-found errors

- Document skipped_empty_re_review and dry_run in the publish_review
  docstring and add a closing-summary contract to the reviewer prompt so
  the agent never claims a review was published when review_id is null.
- Raise ReviewerThreadMissingError from replace_findings on SDK
  NotFoundError; add_finding/update_finding/publish_review return a
  structured do-not-retry result instead of raising, so the agent reports
  the blocker after one failure instead of retrying 10-30 times.

* fix: translate thread 404s across all reviewer tool boundaries

get_thread_metadata now raises ReviewerThreadMissingError instead of
swallowing a missing thread as {} (which produced misleading 'No finding
found' results), set_reviewer_thread_metadata translates the SDK 404 the
same way, and every reviewer tool entrypoint (add/update/list findings,
publish_review incl. eval dry-run, resolve/reply thread) returns the
structured do-not-retry result.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 10:44:13 -07:00
Johannes du Plessis
7cd882bb67
fix: make review publish idempotent (partial-failure recovery) (#1477)
Publish had partial-failure windows that double-posted summaries or
corrupted findings state. This hardens the recovery paths:

- open_swe_review_exists is now tri-state (True/False/None). On a
  pagination/API failure it returns None ("unknown") instead of False,
  and the empty-summary dedup keys off the durable last_reviewed_sha
  before consulting GitHub, so a transient failure never double-posts a
  "no issues found" summary.
- Comment-id backfill matches strictly on the embedded open-swe marker;
  the colliding (path, line, body) fallback is gone, so similar findings
  no longer share a comment id and break resolve-on-fix.
- Review-id and comment-id stamping collapse into one guarded
  read-modify-write (re-reads latest before writing), removing the
  half-stamped intermediate states the prior multi-write flow left open.
- New mutate_findings primitive centralizes read-modify-write so finding
  updates operate on the freshest persisted list and skip no-op writes.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 10:17:31 -07:00
Johannes du Plessis
8f596a421d
feat: prep reviewer repo at init and load repo skills (#1480)
* feat: prep reviewer repo at init and load repo skills

Clone + checkout the PR head during reviewer agent init so SkillsMiddleware
can discover the repo's .agents/skills and .claude/skills from disk at its
one-shot scan, and so the LLM no longer narrates the clone mid-run.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: load reviewer skills from trusted base sha and fetch PR pull ref

Address review findings: skills are now extracted from the PR base sha via
git archive into a dir outside the checkout (prevents PR-authored SKILL.md
prompt injection), and repo prep fetches refs/pull/<n>/head with a strict
checkout so fork PRs fail loudly instead of silently reviewing the default
branch.

* fix: drop ref from skill-extraction log to satisfy CodeQL

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 09:54:32 -07:00