Commit graph

82 commits

Author SHA1 Message Date
Johannes du Plessis
39a26e16b5
fix: optimize agent thread lists (#1570)
* fix: optimize agent thread lists

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: refresh missing thread run status

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-18 10:09:22 -07:00
Johannes du Plessis
c07434a221
fix: allow read-only cross-user access to agent threads via Open in Web links (#1568)
The thread detail endpoint already returned metadata for non-owners, but
the transcript hydration endpoints (state, stream/events, history, pr-diff)
all asserted ownership and 404-ed. This caused the UI to redirect non-owners
back to /agents when they clicked an "Open in Web" link shared in Slack.

Dashboard login is already gated by ALLOWED_GITHUB_ORGS, so any logged-in
user is a trusted org member. This commit:
- Adds _thread_is_readable / _assert_thread_readable helpers that grant
  read access to any surfaced-source thread for authenticated users
- Relaxes read endpoints (state, stream/events, history, pr-diff, SSE
  stream) to use readable checks instead of ownership checks
- Keeps write endpoints (send message, cancel, delete, resolve, run
  commands) owner-only
- Adds an isOwner field to the thread summary so the frontend can render
  a read-only mode (hides the prompt bar, resolve/delete buttons)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-18 09:10:09 -07:00
Johannes du Plessis
8e39f62122
feat: activate PR babysitting UI toggles for autofix and trigger mode (#1561)
* feat: activate PR babysitting UI toggles for autofix and trigger mode

Remove the "coming soon" gating on the Autofix Mode, Autofix Severity
Threshold, and Trigger Mode controls in the review settings page so
admins can enable CI auto-fix and review-comment resolution on PRs
that Open SWE opens. The backend (ci_autofix.py, webapp.py webhook
routing) was already fully wired — only the UI was disabled.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: simplify autofix to on/off toggle, remove severity threshold

Replace the four-level AutofixMode (off/low/medium/high) and the
autofix_severity_threshold setting with a single boolean
autofix_enabled toggle. The severity threshold was leftover from the
reviewer finding-severity model and does not apply to CI autofix;
the agent should fix any failing CI and resolve any comments on PRs
it opens.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: move autofix toggle to per-user profile, remove team-level setting

The autofix toggle is now per-user (auto_fix_ci in the user profile)
instead of team-level (admin-only). This uses the existing auto_fix_ci
field that was already in ProfileUpdate but never wired up.

Changes:
- ci_autofix.py: check per-user auto_fix_ci profile flag after
  resolving the agent thread's github_login, instead of checking
  team-level autofix_enabled before knowing the PR
- webapp.py: removed early is_autofix_enabled() webhook gates; the
  per-user check now happens in ci_autofix.py once the thread is found
- team_settings.py: removed autofix_enabled field, is_autofix_enabled()
- cloud-agents.tsx: enabled the auto_fix_ci toggle (was comingSoon)
- review.tsx: removed the admin-level autofix switch
- Updated tests and AGENTS.md

The agent graph (not the reviewer) is what gets dispatched - this was
already correct in ci_autofix.py line 223: client.runs.create(
thread_id, "agent", ...).

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: batch PR babysitting events

Remove the leftover trigger-mode gate from PR babysitting and batch new CI/review events while an agent run is already active so the running agent can handle the latest PR state before finishing. Also moves review-feedback permission checks behind the per-user opt-out and applies the auto-fix profile gate to merge-conflict babysitting.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: consume batched babysitting events

Teach the agent queue middleware to turn pending PR babysitting metadata into an injected instruction for the active run, so batched CI/review events are not dropped while still avoiding duplicate run creation.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: address review findings in PR babysitting batching

- Route batched events through the LangGraph store (read in-process by the
  message-queue middleware) instead of a per-model-call threads.get on every
  agent thread.
- Only record an attempt / mark the head SHA handled on a real dispatch, not
  on a batch, so an event isn't permanently dropped if the in-flight run ends
  before consuming it.
- Carry the reviewer's comment through batched review feedback instead of
  replacing it with a generic re-check nudge.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 14:12:04 -07:00
Johannes du Plessis
60b7f4677a
feat: add user-scoped Currents.dev API key for e2e test investigation (#1566)
* feat: add user-scoped Currents.dev API key for e2e test investigation

Allow each user to configure their own Currents.dev API key on the
Profile Settings page. The key is encrypted at rest in a per-user
LangGraph Store namespace and feeds server-side read-only tools that
query the Currents REST API (runs, instances, projects, test results)
so agent runs can inspect e2e test failures including screenshots and
DOM snapshots.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: add pagination cursors to currents_list_project_runs

Address review feedback: forward starting_after/ending_before cursor
parameters to /projects/{projectId}/runs so the agent can paginate
beyond the first 50 results.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 13:59:55 -07:00
Johannes du Plessis
efe07486e2
feat: weight merged PRs above LOC in usage leaderboard sorting (#1563)
Move merged_prs to the primary sort key in the agent usage leaderboard,
ahead of agent_loc, prs_opened, and agent_runs. Update the UI description
to reflect the new ranking order.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 10:59:42 -07:00
Johannes du Plessis
b19804536c
feat: handle images sent to non-vision models in Slack, Linear, and web UI (#1560)
* feat: handle images sent to non-vision models in Slack, Linear, and web UI

Add vision capability checks across all image input paths. When a user
sends images to a text-only model (e.g. GLM 5.2, DeepSeek V4 Pro), the
images are now skipped and a warning is injected into the prompt instead
of sending unsupported content to the model.

- Slack: resolve model at webhook time, skip image fetch + add warning
- Linear: same pattern as Slack
- Queued message middleware: read resolved model from thread metadata,
  strip images from queued payloads for text-only models
- Web UI: disable submit + show inline warning when images are attached
  to a non-vision model selection
- Shared: resolve_agent_model_id helper + vision_not_supported_warning

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: mock resolve_agent_model_id in Slack mention test

The test_process_slack_mention_queues_active_thread_message test was
missing a mock for the new resolve_agent_model_id call added to the
Slack webhook handler, causing a TypeError when image URLs triggered
the model resolution path.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: include vision warning in queued payload for text-only models

Update the prompt variable (not just content_blocks) before clearing
image_urls so the queued payload also carries the warning text when a
Slack/Linear follow-up arrives while the thread is busy.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 09:12:52 -07:00
Johannes du Plessis
46e8982b29
feat: add max effort level for GLM 5.2 (#1562)
* feat: add max effort level for GLM 5.2

GLM 5.2 supports a 'max' thinking effort level (recommended for coding
tasks per Z.ai/Fireworks docs). Add it to the model's effort list so it
surfaces in the profile editor and maps to reasoning_effort=max.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: restrict GLM 5.2 efforts to none, high, max

GLM 5.2 only supports non-thinking (none), high, and max effort levels.
Remove low and medium which the model does not support.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 09:10:22 -07:00
Johannes du Plessis
0927f2dd9c
feat: Run reviewer eval in a GitHub Action; dashboard becomes read-only (#1556)
* Run reviewer eval in a GitHub Action; make dashboard a read-only progress view

The dashboard launched the eval as a subprocess inside the serving deployment
worker, so a container recycle killed long runs and discarded results that had
already completed server-side. Move the harness to a workflow_dispatch Action
(run on prod). run_eval now publishes status/progress/log-tail to the LangGraph
store record the dashboard reads, so /admin/evals stays a live view; a killed
Action surfaces as failed via the stale-heartbeat reconcile.

* reviewer_eval workflow: pass inputs via env, no shell interpolation

Addresses the reviewer finding: workflow_dispatch string inputs were
interpolated into the run: block (limit unquoted), allowing shell injection in
a job holding LANGSMITH/ANTHROPIC keys. Pass inputs through env and reference
quoted "$VARS"; validate limit is numeric and build its flag in bash.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 19:38:36 -07:00
Johannes du Plessis
77bc120583
feat: add GLM 5.2 model option (#1543)
* feat: add GLM 5.2 model option

Add GLM 5.2 to the supported model list so it's selectable in the
profile editor and as a team/per-thread model.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: remove GLM 5.1 model option

Remove GLM 5.1 from the supported model list now that GLM 5.2 is available.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 15:31:10 -07:00
Johannes du Plessis
801f93b4de
feat: AI-sorted PR review view with diff grouping (#1544)
* feat: AI-sorted PR review view with diff grouping

Group a PR's changed files into a logical top-to-bottom walkthrough via a best-effort structured-output LLM pass, kicked off concurrently with the reviewer run (~0 added latency) and persisted on reviewer thread metadata. Results render in the review UI behind an AI sorted / file tree toggle that persists across PRs; the view falls back to the file tree when groups are absent or stale.

Adds a grouping-model team default (inherits the Reviewer subagent model when unset) and the admin RolePicker for it.

* feat(reviews): richer AI-sorted explanations + sidebar polish

Sidebar group rows get Devin-style spacing (dividers, padding), a per-group file list (click a file to jump to its diff), inline-code chips in titles, and an accent Read explanation link.

Group explanations are now rich markdown: the grouping prompt feeds per-hunk line ranges and asks for inline code, a short code block, and [path:line](#loc=...) references. The Markdown renderer turns those #loc= links into in-page buttons that scroll the diff to the hunk and highlight the range, reusing the existing selectedLines path.

* fix(reviews): drop stale diff groups from the AI-sorted view

When groups were generated for a previous head, a persisted "ai" view in localStorage still rendered the outdated walkthrough. groupedView now returns null on diff_groups_stale, so the file-tree fallback is used and the view toggle hides until fresh groups arrive.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 13:59:37 -07:00
Johannes du Plessis
e3025dee77
Reviewer eval admin: configurable runs + stacked form layout (#1540)
Drive dashboard-triggered reviewer eval runs with per-run model, effort,
score mode, severity threshold, cap, limit, and concurrency overrides, plus
per-example start/finish/error logging in the eval target.

Rework the admin eval form from the label-left/control-right SettingsRow
(which crushed the description column when packing 3-4 wide inputs) into
stacked field groups with captioned inputs in a responsive grid.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 10:12:35 -07:00
Johannes du Plessis
911c835c2a
feat: chat with your PR on the review page (#1534)
* feat: chat with your PR on the review page

Add a sandbox-less `chat` graph that answers questions about a single PR
from its diff, the published review findings, and read-only GitHub access.

- agent/chat.py: deepagents graph, no sandbox (default StateBackend, file
  mutation + execute tools excluded). PR context is seeded as virtual files
  under /pr/; a repo-scoped App token is resolved in-graph.
- tools: read_repo_file, search_repo_code, list_review_findings.
- dashboard/review_chat_api.py + routes: per-user chat thread, LangGraph
  stream/commands/state/history proxy pinned to the chat assistant, seeds
  diff/findings/overview on first run. Gated by repo access.
- UI: Chat tab wired to a chat-scoped StreamProvider (replaces Coming Soon).

* feat: admin setting for review-chat default model

Add a 'Open SWE Review Chat' default to team settings (default_chat_model /
default_chat_reasoning_effort). get_team_default_model("chat") inherits the
Agent default when unset; the chat graph resolves through it. Admin RolePicker
gains an 'Agent default' inherit option that clears the override.

* feat: multi-conversation review chat (tabs, new chat, history)

Replace the single per-PR chat thread with multiple per-user conversations:
- threads minted client-side; first message persists with a title derived
  from the prompt.
- list + delete endpoints; chat panel gets a tab strip (history), new-chat
  (+), close (x), refresh, an intro greeting, and suggested prompts.
- get_review_chat now returns availability only (ids are client-minted).

* ui fixes

* ui: review-chat history dropdown, full-width AI replies, resizable side panel

* fix(review-chat): enforce per-user thread ownership on proxy endpoints; reseed PR context on head change

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 17:17:30 -07:00
Johannes du Plessis
7397ff93ba
feat: CI auto-fix and PR babysitting for agent PRs (#1530)
* feat: CI auto-fix and PR babysitting for agent PRs

Watch CI failures and review feedback on PRs Open SWE opened, then dispatch
confidence-gated fix runs on the originating agent thread. Adds CI webhook
ingestion (check_run/check_suite/workflow_run/status), a per-PR @open-swe
autofix on|off toggle, auto-response to review comments, and a polling
ci_monitor graph that also flags merge conflicts. Gated by the existing
autofix_mode/trigger_mode settings, the enabled-repos opt-in, base-branch and
human-commit skip rules, dedupe, and a per-PR attempt cap.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: address review feedback on CI auto-fix

- Security: gate the no-mention review-feedback path on author trust —
  require a trusted author_association (OWNER/MEMBER/COLLABORATOR) plus a
  GitHub write/maintain/admin permission check before dispatching a
  write-capable agent run, preventing privilege escalation from
  read/triage/outside reviewers.
- Auth: reuse the originating PR thread's source + login/email when
  dispatching fix runs so the GitHub-token resolver authenticates them in
  non-bot-token deployments (bespoke github_ci source failed to resolve).
- Docs: document the Commit statuses: Read-only permission required for the
  Status webhook event.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 13:53:50 -07:00
Johannes du Plessis
be2fe7131b
feat: live reviewer eval logs on a dedicated admin page (#1527)
* feat: stream reviewer eval logs on a dedicated admin page

Stream the eval subprocess output into a rolling log_tail and persist it
during the run (was only captured at exit), so the live output is visible
while the eval runs. Move the eval runner off the admin page onto its own
/admin/evals page (linked like Review Style Prompts) with a live log viewer.

* chore: drop unrelated SSR-register drift from generated route tree

* fix(ui): pre-bundle workbox-window to stop dev re-optimize reload

The PWA service worker (devOptions.enabled) pulls workbox-window, which
Vite discovers after first render and re-optimizes, forcing a reload that
cancels in-flight code-split route imports (Failed to fetch dynamically
imported module). Pre-bundling it via optimizeDeps.include avoids the
mid-session reload.

* fix(ui): suppress html hydration warning for pre-hydration theme script

The inline theme script sets class="dark"/color-scheme on <html> before
React hydrates, so the prerendered HTML never matches. suppressHydrationWarning
on <html> silences the (expected) one-level attribute mismatch.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 11:29:16 -07:00
Johannes du Plessis
eb18b07b20
feat: trigger reviewer evals from the admin page (#1524)
* feat: trigger reviewer evals from the admin page

Add an admin-only "Reviewer eval" section + endpoints that launch the
reviewer benchmark as an isolated subprocess against the running
deployment, with live status and the LangSmith experiment link. Route
eval traces to a dedicated open-swe-evals project so they stay out of
the production tracing project.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: reconcile reviewer eval status via heartbeat, not local process

The persisted record is shared across workers but _PROCS is process-local.
The owning worker now refreshes a heartbeat while the subprocess runs, and
status is only reconciled to failed once the heartbeat is stale, so a poll on
a worker without the local handle no longer kills a live run (and a duplicate
start is rejected across workers).

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 09:48:12 -07:00
Johannes du Plessis
7f52fce2ea
Remove Kimi K2.6 from supported model selector (#1522)
Kimi K2.7 is now the supported version, so drop the older K2.6 entry.
Update the related fireworks provider test to use K2.7 instead.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 09:39:08 -07:00
Johannes du Plessis
38dcc37040
feat: add Kimi K2.7 model option (#1515)
* feat: add Kimi K2.7 model option

Surface Moonshot's newly released Kimi K2.7 in the model picker, routed
through Fireworks following the existing kimi-k2pN convention, with the
same none/low/medium/high reasoning efforts (default high) as K2.6.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: correct Kimi K2.7 Fireworks model id to kimi-k2p7-code

Fireworks now hosts the K2.7 release as
accounts/fireworks/models/kimi-k2p7-code (the `kimi-k2p7` slug 404s).
Point the model option and its test at the live model id.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: drop unsupported "none" effort for Kimi K2.7

Kimi K2.7 Code is a thinking-only model — Moonshot documents that it does
not support non-thinking mode (passing thinking={"type":"disabled"} or
reasoning_effort="none" errors). Removing "none" from the advertised
efforts so the picker can't send an invalid Fireworks request.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-13 08:10:04 -07:00
Johannes du Plessis
67239bb2c3
perf: speed up Reviews list (My PRs + All PRs) (#1518)
* perf: speed up Reviews list (My PRs + All PRs)

Push the "My PRs" author filter into the threads.search metadata
(pr.author containment) instead of scanning up to 1000 reviewer threads
in Python, and replace the per-repo GitHub access check (an N+1 of
sequential GET /repos calls) with a single per-login accessible-repo set,
cached for 60s. Detail endpoints still re-validate access live.

Frontend: prefetch the inactive tab and adjacent page on hover/focus so
tab switches and pagination are instant.

* fix: don't cache repo access for the reviews list

The /reviews list is an authorization boundary for private PR metadata
(repo/PR titles, branches, authors, finding counts). A cross-request TTL
cache on the accessible-repo set could surface that metadata for up to
60s after a user lost repo access. Resolve the set fresh per request
instead — still a fixed, repo-count-independent burst of GitHub calls
(no per-repo N+1), with no staleness.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-12 14:14:50 -07:00
Caroline di Vittorio
e8bb6b497b
feat: mark threads as resolved to hide from sidebar (#1500)
* feat: mark threads resolved to hide from sidebar [closes resolve-threads]

Add a resolved flag (thread metadata) so users can clear finished
threads from the sidebar without deleting them. Resolved threads move to
a collapsible "Resolved" group (capped at 20 with a "Show all" link) and
are fully searchable/filterable + paginated on a new /agents/threads
page driven by URL query params. Sending a new message auto-unresolves.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: refresh paginated agent thread lists

Invalidate all agent thread list queries when thread lifecycle events change sidebar data, keeping the new paginated sidebar cache in sync.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-06-12 12:01:42 -07:00
Johannes du Plessis
09d5d00e59
feat: render Reviews page diffs with pierre MultiFileDiff (#1517)
* feat: render Reviews page diffs with pierre MultiFileDiff

The Reviews detail page used a hand-rolled hunk renderer with no syntax
highlighting. Switch it to the same pierre MultiFileDiff + theming the agent
chat git panel uses.

- review-diff API now returns full original/modified file contents instead of
  hunks, via a shared build_pr_diff_files helper extracted from thread_api
- findings render as right-anchored markers (pierre line annotations); focus
  highlight uses selectedLines; floating finding card still anchors to the marker

* feat: auto-collapse a review diff card when marked as viewed

* fix: URL-encode file path in Contents API fetch

Filenames containing reserved URL characters (#, ?) were truncated, so those
files rendered as empty/unrenderable. quote(path, safe='/') preserves the path
separators while escaping the rest.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-12 11:03:51 -07:00
Johannes du Plessis
3e105f5027
fix: Reviews tab — anchored finding card, paginated list, file tree truncation (#1507)
* fix: Reviews tab — anchored finding card, paginated list, file tree truncation

- Finding card now tracks the diff anchor while scrolling instead of staying
  frozen in the viewport; auto-hides when its diff card collapses (including
  collapse via mark-as-viewed) and on click outside
- Checks section capped with max height + scroll
- /reviews paginated (page size 20) with has_more, filtered to the current
  user's PRs by default with an All toggle; PR author login now stored in
  reviewer thread metadata
- File tree truncation marker overlapped filenames because the sidebar bg
  was transparent; use the opaque sidebar color

* fix: finding card tracks anchor 1:1 while scrolling

Drop the vertical viewport clamp — it pinned the card at the clamp
boundary while the highlighted lines kept scrolling, breaking the
attachment.

* fix: anchor finding card with Base UI popover

Replace manual fixed-position tracking (laggy: setState per scroll
frame) with a Popover anchored to the finding's diff row. Floating UI
tracks the anchor outside React renders, so the card moves 1:1 with
the content and scrolls out of view with it. Unanchored findings keep
the fixed top-right card.

* fix: lock finding card to diff scroll

Replace the Base UI popover (async repositioning, paints a frame behind
native scroll) with a card absolutely positioned inside the scroll
container, so it scrolls with the diff in the same compositor frame.
Scroll moves to the ReviewBody root, side panel becomes sticky.
Position recomputes only on layout shifts via ResizeObserver.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 17:52:20 -07:00
Johannes du Plessis
bfcb714c5c
feat: PR review page (#1495)
* feat: PR review page

* feat: manual review trigger card on admin page

* feat: finding focus mode with floating card and dimmed backdrop

* fix: dim only info panel, subtle block highlight for focused finding

* refactor: move PR reviews into agents page with file-tree sidebar

* fix: address review feedback

- paginate list_reviews until enough accessible records collected
- include rename/metadata-only files in diff API with empty hunks
- remount ReviewBody on head_sha change so viewed-files state resets

* fix: review tree background + finding card anchoring

* fix: remove no-op PR size chip and dead check links

* fix: parallel diff fetch, complete reReview type, shared github_headers

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 16:11:10 -07:00
Johannes du Plessis
b8035a4758
fix: stream proxy preflight + optimistic thread staleTime (#1502)
* fix: preflight stream proxy before SSE starts, let optimistic thread survive refetch

* fix: seed sidebar thread details as stale so mark-viewed fetch still fires

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 13:07:24 -07:00
Johannes du Plessis
e0678e8c01
feat: track PR lifecycle state per thread for sidebar (#1492)
* feat: track PR lifecycle state per thread for sidebar

Persist a PR's draft/open/merged/closed state on the agent thread and keep
it in sync as PR webhooks fire, so the Agents UI can show per-thread PR
status the way Cursor does. open_pull_request now records the initial
draft/open state and pr_title; thread summaries expose diffStats; and the
PR webhook refreshes pr_state on close/reopen/draft/ready transitions.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: consolidate PR state mapping into shared derive_pr_state helper

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 12:21:59 -07:00
Johannes du Plessis
9a2f68e99d
fix: auto-recover from expired GitHub refresh tokens (#1491)
* fix: auto-recover from expired GitHub refresh tokens

When a user's GitHub OAuth refresh token was permanently dead (revoked or
expired), token refresh failed but get_valid_access_token still handed back
the known-stale access token, so dashboard GitHub calls kept 401ing until the
user manually logged out and back in.

Now we distinguish unrecoverable refresh failures (bad_refresh_token /
unauthorized_client) from transient ones: on an unrecoverable failure we drop
the dead stored authorization and return None, so callers prompt a clean
re-login. Transient failures still fall back to the stored token.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: don't delete fresh re-auth when stale refresh fails

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 12:05:27 -07:00
Johannes du Plessis
f08177aed7
fix: harden origin parsing and home prompt submit failure (#1499)
- _origin_of: treat invalid ports as invalid origin instead of letting
  urlparse ValueError turn CSRF rejections into 500s
- AgentsHome: reset submitting/draft when stream.submit rejects before a
  thread id is minted, so the prompt isn't left disabled

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 11:54:25 -07:00
Caroline di Vittorio
8f983e9e40
feat: add View PR link in git panel header (#1498)
* feat: add View PR link in git panel header

Surface a clickable "View PR" link in the agent git panel header so
users can jump straight from a chat thread to its pull request instead
of hunting for the URL in the conversation. Uses the pr.url already
captured in thread metadata.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: use shared buttonVariants for View PR link

Style the View PR link with the shared buttonVariants (outline/sm)
helper instead of hand-rolled classes, matching the existing
anchor-as-button pattern used in login.tsx and AutomationsList.tsx.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: show PR diff in git panel diff tab

Adds /threads/{id}/pr-diff endpoint fetching PR files + contents from GitHub;
git panel prefers it and falls back to the live message-derived diff.

* feat: auto-expand git panel when a PR lands mid-run

Panel still defaults to collapsed and remembers manual toggles; the
auto-expand is ephemeral and not written to localStorage.

* fix: git panel always starts collapsed

Drop localStorage persistence of collapsed state; panel opens only via
manual toggle or when a PR lands mid-run, and re-collapses on thread switch.

* fix: address PR review findings

Require user OAuth token for pr-diff (no app-token fallback) so GitHub
enforces current repo access; render placeholder for binary/oversized files.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 11:48:53 -07:00
Christian Bromann
abf354bb05
feat(open-swe): stream agent chat via @langchain/react v2 protocol (#1475)
* feat(dashboard): stream agent chat via @langchain/react v2 protocol

Replace the bespoke SSE + React Query polling path with LangGraph’s
v2 event stream through credentialed dashboard proxies. Run starts go
through stream commands; mid-run follow-ups still queue via /messages.

* fix import path

* fix tests after rebase

* format

* PR feedback

* improved model fallback

* fix image handling

* embrace sdk

* cleanup

* cr

* more cleanup

* fix cors

* harden security

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 09:54:35 -07:00
Johannes du Plessis
8c7b789122
feat: add right-click open trace context menu for threads (#1490)
* feat: add right-click open trace context menu for threads

Add a context menu to sidebar thread rows exposing an Open trace action
that links to the LangSmith thread trace, surfaced via a new traceUrl
field on the thread summary.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: drop stray yarn artifacts from PR

Remove accidentally committed ui/.yarnrc.yml and ui/.yarn/install-state.gz
that broke immutable installs against the v1 lockfile, and gitignore yarn
artifacts to prevent recurrence.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: reuse thread_id var, drop redundant danger-color fallback

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 15:40:43 -07:00
Johannes du Plessis
370644ae8b
chore: hide Fable 5 model from supported options (#1483)
Fable 5 is currently unusable with our API key. Remove it from the
selectable model list; it can be re-added later.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 11:18:12 -07:00
Johannes du Plessis
f539962c73
feat: server-side Datadog/LangSmith observability tools + team creds [closes OPE-54] (#1476)
* feat: server-side Datadog/LangSmith observability tools + team creds

Add team-wide observability credential settings (Datadog DD_SITE/API/APP
keys, LangSmith API key) stored encrypted server-side, with an admin
dashboard section to connect/disconnect each provider. When connected,
get_agent loads read-only observability tools server-side: Datadog via its
hosted MCP server (langchain-mcp-adapters, toolsets=core) and LangSmith
read tools (langsmith_get_trace, langsmith_list_runs). Credentials live in
the LangGraph server process and are never exposed to the sandbox.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: address review on observability tools

Address PR review feedback:
- Authorize observability tools per triggering user (admins + the
  OBSERVABILITY_AUTHORIZED_EMAILS allowlist) so prompt-injected runs from
  untrusted contributors can't reach team Datadog/LangSmith data.
- Use the documented Datadog MCP auth headers DD_API_KEY / DD_APPLICATION_KEY.
- Store each provider's credentials under its own store key to avoid a
  read-modify-write race dropping the other provider on concurrent saves.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: async email resolution in observability authorization gate

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 11:07:42 -07:00
Johannes du Plessis
722279e277
fix: restore org-wide read access to agent threads (#1474)
The viewed-marker feature (#1441) reintroduced an owner-only gate on
thread reads, regressing #1425 which made all threads readable by any
org user. Reads no longer assert ownership; viewed markers are only
written for the thread owner.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-09 16:41:07 -07:00
Johannes du Plessis
da8f003933
feat: per-repo custom instructions for the coding agent (#1460)
* feat: per-repo custom instructions for the coding agent

Adds per-repository custom instructions for the main coding agent,
mirroring the reviewer's per-repo style prompts. Instructions are stored
in the LangGraph Store, managed via dashboard API + UI (Monaco editor),
and appended to the agent's system prompt for runs targeting that repo.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: wire agent instructions route into generated route tree

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: enforce repo access on instruction routes

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-09 15:46:29 -07:00
Johannes du Plessis
c7acf2ff27
Revert out-of-process usage-snapshot builder (#1473)
* Revert "fix: schedule usage snapshot runs with explicit empty input (#1470)"

This reverts commit 49aa638c58.

* Revert "feat: out-of-process usage-snapshot builder (Phase 1) (#1468)"

This reverts commit 798cc6edf1.
2026-06-09 13:19:11 -07:00
Johannes du Plessis
49aa638c58
fix: schedule usage snapshot runs with explicit empty input (#1470) 2026-06-09 12:53:15 -07:00
Johannes du Plessis
798cc6edf1
feat: out-of-process usage-snapshot builder (Phase 1) (#1468)
* feat: out-of-process usage-snapshot builder (Phase 1)

Move all usage-tab compute off the run-serving HTTP process (the #1434 bug
class). Read path is now a pure cache read with a typed computing placeholder
on cold miss; a dedicated usage_snapshot graph + global ~10min cron rebuild
every period's snapshot out-of-process, wrapped in asyncio.timeout and gated by
USAGE_SNAPSHOT_CRON_ENABLED. Lifespan makes one fire-and-forget scheduling call,
never a retained/looping task.

* fix: address PR review on usage-snapshot builder

- admin guard on /admin/usage/rebuild (was any logged-in user)
- cap lifespan loopback calls with asyncio.timeout(5) so a startup hang
  can't block boot
- memoize cron registration so steady-state requests skip the loopback
  check; reap duplicate crons from concurrent-replica races
- thread the computing flag through the usage payload too
2026-06-09 11:11:06 -07:00
Johannes du Plessis
f349311d20
feat: add Claude Fable 5 as a supported model (#1467)
* feat: add Claude Fable 5 as a supported model

Add Anthropic's claude-fable-5 (Mythos-class, released June 9 2026) to
the supported model list so it can be selected in the profile editor.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: order Fable 5 after Opus 4.8 so stale-Anthropic fallback is preserved

provider_fallback_pair picks the first same-provider model, so placing
Fable 5 first redirected stale Opus selections to it. Keep Opus 4.8 first.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-09 10:59:26 -07:00
Johannes du Plessis
c7b32e34e6
Revert "fix: precompute usage tab caches (#1434)" (#1463)
This reverts commit 10dfc6d2b0.
2026-06-09 09:21:23 -07:00
Johannes du Plessis
5430672edb
fix: revert dashboard stop run controls (#1461)
Revert the dashboard stop-button behavior from #1433 because it allows duplicate submissions while optimistic prompts are still pending.
2026-06-08 21:47:53 -07:00
Johannes du Plessis
8fc06dac8a
fix: prevent thread prefetches from marking threads viewed (#1456)
* fix: prevent thread prefetches from marking threads viewed

Sidebar prefetches now request thread details with mark_viewed=false so
loading /agents no longer clears the unread/finished indicator for
threads the user has not opened.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: always refetch active thread on mount to mark viewed

Prefetches cache details fetched with mark_viewed=false under the same
query key as the active thread. Forcing refetchOnMount="always" ensures
opening a thread issues a mark_viewed=true request, so last_viewed_* is
recorded even when prefetched data is still fresh.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-08 14:41:18 -07:00
Johannes du Plessis
ec41f138fa
fix: update sidebar thread activity indicators (#1441)
* fix: update sidebar thread activity indicators

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: improve sidebar thread organization

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-08 10:38:31 -07:00
Johannes du Plessis
5faf190954
fix: reject images for text-only models (#1439)
* fix: reject image uploads for text-only models

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: validate queued images against active model

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-06 13:09:31 -07:00
Johannes du Plessis
072c0158ff
feat: support dashboard chat images (#1435)
* feat: support dashboard chat images

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: keep pending image prompts visible

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-06 11:50:24 -07:00
Johannes du Plessis
1f619eae0f
feat: add stop button to cancel running agent from web UI (#1433)
* feat: add stop button to cancel running agent from web UI

* fix: handle stopped agent runs

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-06 08:12:30 -07:00
Johannes du Plessis
10dfc6d2b0
fix: precompute usage tab caches (#1434)
* fix: precompute usage tab caches

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: dedupe usage cache refreshes

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-06 08:11:53 -07:00
Johannes du Plessis
e128f2d6dd
feat: cache usage stats and add reviewer metrics (#1432)
* feat: cache usage dashboard stats

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: adjust usage nav placement

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: paginate reviewer usage stats

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-05 14:33:22 -07:00
Johannes du Plessis
f3512841fc
feat: inject org-wide guidelines into reviewer prompt (#1431)
Adds an admin-managed, org-wide review guidelines field to team settings
that the reviewer injects into every PR review across all repos, alongside
the existing per-repo style prompt and AGENTS.md context. Repo-specific
rules take precedence when they conflict.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-05 13:49:30 -07:00
Johannes du Plessis
449cb5d1a8
fix: make default repository configurable (#1429)
* fix: make default repository configurable

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve dashboard repo-less runs

* fix: distinguish explicit repo-less dashboard runs

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-05 13:48:47 -07:00
Ramon Nogueira
bbe449244a
feat: make web threads readable by any org user (#1425)
* feat: make web threads viewable by any org user

Reads (get/stream) now allow any logged-in org user to view a thread
whose source is surfaced (dashboard/github/slack/linear/schedule),
instead of requiring ownership. Internal reviewer/analyzer threads stay
hidden via the same source filter. Writes (send/cancel/delete) remain
owner-only.

Adds GET /threads?all=true to list every surfaced thread; the default
list stays per-user.

* feat: make all web threads readable by any org user

Drop the per-thread source/owner gate on reads: any logged-in org user
can now view and stream any thread, including reviewer/analyzer threads.
GET /threads?all=true returns every thread regardless of source. Writes
(send/cancel/delete) stay owner-only. All routes remain behind the
session + org-login gate.
2026-06-05 08:28:08 -07:00
Johannes du Plessis
6895ddcedc
feat: add scheduled web agents (#1422)
* feat: add scheduled web agents

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* fix: secure scheduled agent repositories

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* feat: rebuild scheduled agents as Automations tab

Scrap the inline ScheduledAgentsPanel and replace it with a dedicated
Automations tab: sidebar nav entry, list view with stat cards + empty
state, and a full editor (name, Active toggle, repo, scheduled trigger
picker, agent instructions + model).

* fix: clear collapsed-sidebar button on mobile in Automations

* fix: allow clearing automation repo on update

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-05 02:20:24 +00:00