* feat: add PR trace resolution
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: inject reviewer trace context as JSON
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review on PR trace resolution
Use the documented LangSmith metadata filter syntax
(and(eq(metadata_key,...), eq(metadata_value,...))) instead of
has(metadata, '{...}'), which does not match runs — _list_thread_runs
was silently returning nothing. Bound full-text searches to a 90-day
window so they don't hit LangSmith's large-window rate limit.
Also folds in the best-effort branch->head-sha resolver (dropping the
weighted scoring/threshold + repo/file evidence + GitHub hydration),
sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint.
The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session
were removed; resolution now runs deterministically from the trusted run
config with no model-controlled pr_url or thread_id.
* fix: scope branch trace search to the repo
Branch names like fix-tests aren't unique across repos (or older PRs) in
a shared tracing project, so an unscoped branch hit could resolve to an
unrelated thread and write its runs into the reviewer sandbox. Require
the repo slug to co-occur with the branch in matched runs; the full head
SHA stays unscoped since it is globally unique. Addresses open-swe review
on PR #1612.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Add a "Start from a template" gallery to the Automations page so users can
scaffold common scheduled agent runs (PR review digest, issue triage,
dependency checks, flaky-test tracking, release notes, docs freshness,
security audit) instead of writing every automation from scratch. Selecting
a template opens the new-automation editor prefilled with its instructions
and a default schedule, which the user can tweak before saving.
Also adds an isDescribableCron helper so template schedules render as
human-readable labels rather than raw cron in the editor.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: inline PR comments in the reviews UI
Click the diff gutter "+" on a line to open an inline comment composer
(rendered like the finding card via a Pierre annotation); submitting
posts a real inline PR review comment as the signed-in user through a
new POST /reviews/{owner}/{repo}/{number}/comments. The "+" press-drag →
"Add to Chat" selection path is unchanged.
* feat: GitHub-parity comment box, PR comments dropdown, collapse nav
- Comment composer now mirrors GitHub's box: Write/Preview tabs (markdown
rendered via the existing Markdown component) and a markdown toolbar
(heading, bold, italic, quote, code, link, bulleted/numbered/task list).
- Surface other people's inline PR comments in a Devin-style dropdown in the
review header (search + link to the thread on GitHub). New
GET /reviews/{owner}/{repo}/{number}/comments lists them and flags the
reviewer's own (marker-bearing) comments so they're filtered out.
- Collapse the global nav by default on a review detail page, restoring the
prior preference on leave.
* feat: bigger comment-toolbar icons; open dropdown comments inline
- Enlarge the markdown toolbar glyphs (Phosphor) in the comment composer —
they were rendering at 10px.
- Clicking a comment in the PR comments dropdown now opens it inline in the
diff as a read-only finding-style card (InlineComment), scrolling its line
into view, instead of navigating to GitHub. Falls back to GitHub when the
comment's file/line isn't in the current diff.
* fix: drive "Add to Chat" from native text selection
The gutter "+" is now comment-only; wiring its click to the composer
conflicted with its old double-duty as the drag-to-select handle, which
broke selection → "Add to Chat". Switch to Devin's model: disable Pierre's
interactive line selection and instead map a native text highlight in the
diff to a line range (via the data-line / data-line-type attributes Pierre
stamps on each line, read from the diff's open shadow root) to show the
"Add to Chat" popup. ⌘L and the existing attachment/popup path are unchanged.
* feat: gutter "+" drag selects a range for multi-line comments
Re-enable Pierre's gutter line selection so dragging the "+" down the
gutter comments across a range (click still comments on a single line);
onLineSelectionEnd routes the range to the composer. Native code-text
selection still drives "Add to Chat" — Pierre only line-selects from the
gutter, and onLineSelectionEnd bails when a native text selection is
present, so a code highlight never opens the composer.
* fix: keep the range highlighted while its comment composer is open
Previously opening the composer cleared the selection, so the lines being
commented on lost their highlight. Drive the controlled selection from the
open comment draft's range so the rows stay highlighted until the composer
is closed.
* fix: address PR review — paginate comments, fall back for outdated ones
- list_review_comments now pages through all PR review comments (bounded by
_MAX_REVIEW_COMMENT_PAGES) instead of returning only the first 100, so older
comments still show in the dropdown.
- Surface GitHub's outdated flag (position == null) as is_outdated; opening such
a comment (or one whose line isn't in the diff) now opens it on GitHub instead
of silently rendering nothing, plus a timeout fallback if the annotation never
mounts (e.g. collapsed context).
* feat: relocate review chunks and embed review in git panel
Move the AI-sorted review chunk navigator out of the global left sidebar
and into the review page's own main panel (sidebar keeps the thread list),
and surface the review main body inside the agent thread's Git > Review
sub-tab with an expand link to the full review page.
The review main body + side panel + helpers are extracted into a shared
ReviewMainBody component with "full" and "embedded" variants.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: move ReviewTab into its own file
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: repo-scoped dynamic sandbox snapshots
Let admins build a per-repo sandbox image from a custom Dockerfile so runs
targeting that repo boot from a snapshot with its deps pre-baked. Snapshot
selection is purely additive: repos without a `ready` repo-scoped snapshot
always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID.
Backend adds a repo_snapshots store module (Dockerfile + build status keyed by
owner/name), threads the resolved repo through the LangSmith sandbox creation
path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a
throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The
UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor +
build status/logs).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden repo snapshot builds
Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins
cannot accidentally build a repo snapshot from a bare Python image that lacks
Open SWE's sandbox tools. Allow stale building records to be retried by tracking
build_started_at and treating old or missing timestamps as stale.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: document repo snapshot base image config
Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert
missing base-image configuration into a handled dashboard API error so admins see
a clear configuration message instead of an unhandled template-generation error.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add plan mode for read-only research and planning
Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce plan-mode read-only at tool layer and disable subagents
Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden plan-mode shell guard against wrapped mutations
Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow
- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline
Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.
* test(open-swe): add Playwright E2E for the Slack → PR → web handoff
Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.
- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).
Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.
* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright
The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.
Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.
* test(open-swe): record Playwright trace + video on every E2E run
Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.
* feat(plan-mode): collaborative plan review with BlockNote + Yjs
When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.
- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
snapshots; plan content/status store; plan REST API (get/approve/reject,
owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
mounted under the agents shell, with a "Review plan" banner in the thread view
and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
flow, including cross-user comment sync and owner-only approval.
* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)
- plan-collab WS: authorize per-thread before joining a room (same read gate as
the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.
Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).
* style: ruff format plan_collab.py
* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS
- Slack "Approve & Implement" now verifies the clicking user is the plan
requester (owner, via the stored triggering_user_id) before implementing —
matching the dashboard API's owner-only approval. Non-owners are pointed to
Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
the REST require_same_origin CSRF defense.
* fix(plan-mode): enter plan mode only via the model + local mock dev harness
Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.
- enter_plan_mode returns a terminating ToolMessage, fixing the missing
ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
Alice/Bob mock users, and a GitHub login picker.
* docs(plan-mode): drop stale references to removed profile/team defaults
The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.
* feat(plan-mode): let any reviewer edit the plan, not just comment
Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.
* test(plan-mode): assert plan-mode entry via the tool's success message
plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.
---------
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: render GitHub-hosted images in PR descriptions on reviews page
PR description images hosted on GitHub (user-attachment uploads and
*.githubusercontent.com) render broken on the reviews page because
private-repo attachments require GitHub auth the browser session lacks.
Add an authenticated backend image proxy (host-allowlisted to guard
against SSRF) and route those image URLs through it from the reviews UI.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden PR image proxy (IDOR, SVG XSS, unbounded buffering)
Address review findings on the review-page image proxy:
- IDOR: the proxy fetched any *.githubusercontent.com URL with the App
installation token, gated only by route-param repo access, so a user
authorized for one repo could read images from another private repo the
App can see. Bind the URL to the authorized PR — only proxy URLs that
appear in that PR's body.
- SVG XSS: served any image/* inline from the API origin, including
image/svg+xml which can run script. Restrict to safe raster types and
add X-Content-Type-Options: nosniff + a locked-down CSP.
- DoS: enforced the size cap only after buffering the full response.
Stream and abort once the cap is exceeded.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Replace the floating finding cards with inline annotations in the PR diff.
An anchored finding shows a collapsed header at its line and expands the full
details in place; non-anchored findings expand inline in the side-panel row.
Expand state is lifted into a context so it survives Pierre's annotation
windowing and the side panel can drive it.
Removes the anchored/fallback floating cards and their rAF/ResizeObserver
scroll-tracking. Also stops the line-selection highlight from bleeding behind
the inline annotation row.
* fix: reviewer reviews full diff; fix review UI scroll + dark-mode composer
- Drop the 200K-char PR diff truncation. The head/tail slice silently dropped
whole files out of the middle; the reviewer's large context window reviews
the complete diff. The grouping pass uses the full diff too. Other #1567 caps
(fetch_url, Slack, pagination, message queue) are kept.
- Sidebar file/group click now lands flush at the top under virtualization:
re-assert alignment after the Virtualizer's mid-scroll height reconciliation
settles, instead of trusting scrollIntoView's estimate-based target.
- Review chat composer textarea no longer shows a lighter square in dark mode
(override the Textarea's dark:bg-input default with dark:bg-transparent).
* fix: anchored finding card sits flush in the diff gutter, no side-panel overlap
The card was positioned next to the anchor and clamped to window.innerWidth, so
when the anchor sat near the diff/side-panel boundary the 412px card spilled
across into the side panel. Pin it flush to the diff column's right edge and
clamp its width to the column so it never overlaps the side panel and shrinks to
fit a narrow column.
* fix: anchored finding card sits over the side panel, not the diff
Flip the horizontal anchor: place the card flush against the diff column's right
edge and extend it rightward over the side panel, instead of leftward over the
diff. Width still caps at the preferred size and shrinks to fit a narrow panel.
* fix: anchor finding card to diff content edge + small-screen fallback
- Anchor to the centered diff content's right edge instead of the full-width
scroller, removing the centering-whitespace gap so the card sits closer to the
diff.
- Below xl the side panel is hidden and the diff fills the viewport, leaving no
gutter; positioning against the content edge would collapse the card to ~0px.
Fall back to overlaying the diff flush-right when there's no usable gutter
(addresses Open SWE review finding f_f6ea77527c).
* fix: anchor finding card next to the annotation, close to the hunk
Anchor the card's left to the finding's annotation (anchorRect.right) instead of
the diff content's right edge, so it sits right by the hunk and overlaps the diff
edge rather than parking out in the gutter. Extends right over the side panel,
clamped to stay on-screen (also keeps it readable below xl with no side panel).
* fix: finding card width shrinks to fit a small side panel
Size the card to the room to the right of the annotation (capped at the
preferred width) so it narrows as the side panel shrinks instead of overflowing.
Below a readable minimum, hold that width and shift left over the diff.
* feat: reviews page split view, add-to-chat, virtualization + perf
- Virtualize the diff (Pierre Virtualizer + worker pool), mirroring the agent
chat panel, so large PRs window rows instead of materializing every line.
- Split chat and diff into independent scroll containers and make chat
auto-scroll fully contained, so typing/streaming no longer moves the diff.
- Memoize FileDiffCard with stable callbacks so focusing a finding re-renders
only the affected card.
- Add a persisted unified/split diff toggle.
- Add highlight-to-chat: select lines (drag / shift-click) + gutter "+" to drop
a file:line snippet into the chat composer.
- Rebuild sidebar group rows: whole card scrolls to the group (incl. the
expanded explanation), Read explanation stays a separate toggle, memoized.
* feat: add-to-chat uses attachment pills + selection popup / ⌘L
Replace the raw-snippet injection with a Cursor-style flow:
- Selecting lines shows a floating "Add to Chat ⌘L" popup at the pointer; ⌘L
adds the current selection without it. Removes the auto-adding gutter "+".
- "Add to chat" now creates a removable attachment pill in the composer (and a
pill in the sent message bubble) instead of pasting raw text. The code is
still serialized into the message content so the model receives it as context.
* feat: restore gutter + drag-handle for line selection
Re-enable Pierre's gutter '+' as a click-and-drag line selector (with the
highlight growing as you drag) — the affordance that was lost when the
auto-adding gutter button was removed. It no longer auto-adds: the commit flows
through onLineSelected to the 'Add to Chat' popup / ⌘L.
* fix: live selection highlight while dragging + reposition Add to Chat popup
- Feed onLineSelectionChange into the controlled selection so rows highlight
live as you drag, not just on release (Pierre only paints the controlled
selection when the prop updates). Popup now fires on onLineSelectionEnd.
- Anchor the popup's bottom-left to the drag handle (drop horizontal centering)
so it no longer overlaps the '+' button.
* fix: anchor Add to Chat popup to gutter handle + click-away to unselect
- Position the popup from the gutter '+' handle's rect (in the diff shadow DOM,
placed on the selection's bottom line) instead of the pointer-release point,
which landed inconsistently. Falls back to the pointer if not found.
- Clear the line selection (and popup) on any outside pointer-down.
* fix: sidebar-collapse header overlap + PR review comments
- Lift useSidebarLayout to AgentsShell (single source), share collapsed via
context, and pad the reviews header left when the sidebar is collapsed so the
fixed collapse toggle no longer overlaps the header content.
- add-to-chat: collect each diff side separately so a selection that spans a
deletion->addition no longer pastes wrong-file lines (PR comment).
- chat: clear attachments after sending via a suggested prompt, so an attached
snippet isn't silently resent on the next message (PR comment).
* fix: anchored finding card positions to the right of the diff again
The virtualization refactor moved the card inside the main-width Virtualizer
scroller, so it clamped over the diff. Render it in the outer container as a
viewport-fixed card clamped to window width (right gutter / over the side panel,
like prod) and track the finding as the diff scrolls (rAF-throttled), hiding
when the finding scrolls out of view.
* feat: activate PR babysitting UI toggles for autofix and trigger mode
Remove the "coming soon" gating on the Autofix Mode, Autofix Severity
Threshold, and Trigger Mode controls in the review settings page so
admins can enable CI auto-fix and review-comment resolution on PRs
that Open SWE opens. The backend (ci_autofix.py, webapp.py webhook
routing) was already fully wired — only the UI was disabled.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: simplify autofix to on/off toggle, remove severity threshold
Replace the four-level AutofixMode (off/low/medium/high) and the
autofix_severity_threshold setting with a single boolean
autofix_enabled toggle. The severity threshold was leftover from the
reviewer finding-severity model and does not apply to CI autofix;
the agent should fix any failing CI and resolve any comments on PRs
it opens.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: move autofix toggle to per-user profile, remove team-level setting
The autofix toggle is now per-user (auto_fix_ci in the user profile)
instead of team-level (admin-only). This uses the existing auto_fix_ci
field that was already in ProfileUpdate but never wired up.
Changes:
- ci_autofix.py: check per-user auto_fix_ci profile flag after
resolving the agent thread's github_login, instead of checking
team-level autofix_enabled before knowing the PR
- webapp.py: removed early is_autofix_enabled() webhook gates; the
per-user check now happens in ci_autofix.py once the thread is found
- team_settings.py: removed autofix_enabled field, is_autofix_enabled()
- cloud-agents.tsx: enabled the auto_fix_ci toggle (was comingSoon)
- review.tsx: removed the admin-level autofix switch
- Updated tests and AGENTS.md
The agent graph (not the reviewer) is what gets dispatched - this was
already correct in ci_autofix.py line 223: client.runs.create(
thread_id, "agent", ...).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: batch PR babysitting events
Remove the leftover trigger-mode gate from PR babysitting and batch new CI/review events while an agent run is already active so the running agent can handle the latest PR state before finishing. Also moves review-feedback permission checks behind the per-user opt-out and applies the auto-fix profile gate to merge-conflict babysitting.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: consume batched babysitting events
Teach the agent queue middleware to turn pending PR babysitting metadata into an injected instruction for the active run, so batched CI/review events are not dropped while still avoiding duplicate run creation.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review findings in PR babysitting batching
- Route batched events through the LangGraph store (read in-process by the
message-queue middleware) instead of a per-model-call threads.get on every
agent thread.
- Only record an attempt / mark the head SHA handled on a real dispatch, not
on a batch, so an event isn't permanently dropped if the in-flight run ends
before consuming it.
- Carry the reviewer's comment through batched review feedback instead of
replacing it with a generic re-check nudge.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add user-scoped Currents.dev API key for e2e test investigation
Allow each user to configure their own Currents.dev API key on the
Profile Settings page. The key is encrypted at rest in a per-user
LangGraph Store namespace and feeds server-side read-only tools that
query the Currents REST API (runs, instances, projects, test results)
so agent runs can inspect e2e test failures including screenshots and
DOM snapshots.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: add pagination cursors to currents_list_project_runs
Address review feedback: forward starting_after/ending_before cursor
parameters to /projects/{projectId}/runs so the agent can paginate
beyond the first 50 results.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Move merged_prs to the primary sort key in the agent usage leaderboard,
ahead of agent_loc, prs_opened, and agent_runs. Update the UI description
to reflect the new ranking order.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add browser notifications for agent run completion
Request notification permission when a user starts their first agent
run from the home page. A toggle in Profile Settings lets users
enable/disable desktop notifications. When a run transitions from
running to a terminal state (finished/error/interrupted), a browser
notification is fired — suppressed for the thread the user is
currently viewing.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: notify for active thread when page is in background tab
Only suppress the notification for the active thread when the page is
actually visible. If the user switched to another browser tab, the
notification fires even for the thread they have open.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* Run reviewer eval in a GitHub Action; make dashboard a read-only progress view
The dashboard launched the eval as a subprocess inside the serving deployment
worker, so a container recycle killed long runs and discarded results that had
already completed server-side. Move the harness to a workflow_dispatch Action
(run on prod). run_eval now publishes status/progress/log-tail to the LangGraph
store record the dashboard reads, so /admin/evals stays a live view; a killed
Action surfaces as failed via the stale-heartbeat reconcile.
* reviewer_eval workflow: pass inputs via env, no shell interpolation
Addresses the reviewer finding: workflow_dispatch string inputs were
interpolated into the run: block (limit unquoted), allowing shell injection in
a job holding LANGSMITH/ANTHROPIC keys. Pass inputs through env and reference
quoted "$VARS"; validate limit is numeric and build its flag in bash.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Lighter, more consistent type across the dashboard and agent UI:
unify title sizes, drop semibold/bold in chrome to medium, harmonize
stray text-sm to text-xs, normalize chat markdown headings, and enable
antialiased font smoothing (fixes heavy text on dark backgrounds).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: AI-sorted PR review view with diff grouping
Group a PR's changed files into a logical top-to-bottom walkthrough via a best-effort structured-output LLM pass, kicked off concurrently with the reviewer run (~0 added latency) and persisted on reviewer thread metadata. Results render in the review UI behind an AI sorted / file tree toggle that persists across PRs; the view falls back to the file tree when groups are absent or stale.
Adds a grouping-model team default (inherits the Reviewer subagent model when unset) and the admin RolePicker for it.
* feat(reviews): richer AI-sorted explanations + sidebar polish
Sidebar group rows get Devin-style spacing (dividers, padding), a per-group file list (click a file to jump to its diff), inline-code chips in titles, and an accent Read explanation link.
Group explanations are now rich markdown: the grouping prompt feeds per-hunk line ranges and asks for inline code, a short code block, and [path:line](#loc=...) references. The Markdown renderer turns those #loc= links into in-page buttons that scroll the diff to the hunk and highlight the range, reusing the existing selectedLines path.
* fix(reviews): drop stale diff groups from the AI-sorted view
When groups were generated for a previous head, a persisted "ai" view in localStorage still rendered the outdated walkthrough. groupedView now returns null on diff_groups_stale, so the file-tree fallback is used and the view toggle hides until fresh groups arrive.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Drive dashboard-triggered reviewer eval runs with per-run model, effort,
score mode, severity threshold, cap, limit, and concurrency overrides, plus
per-example start/finish/error logging in the eval target.
Rework the admin eval form from the label-left/control-right SettingsRow
(which crushed the description column when packing 3-4 wide inputs) into
stacked field groups with captioned inputs in a responsive grid.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Clears ~83 pre-existing eslint errors in the UI: auto-fixable rules
(array-type, import/order, sort-imports, type-only imports, redundant
assertions/conditions) plus manual fixes for unnecessary conditions and
banned @ts-nocheck directives. The 8 ported components excluded from
tsconfig are now also ignored by eslint so type-aware linting no longer
fails to parse them.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: chat with your PR on the review page
Add a sandbox-less `chat` graph that answers questions about a single PR
from its diff, the published review findings, and read-only GitHub access.
- agent/chat.py: deepagents graph, no sandbox (default StateBackend, file
mutation + execute tools excluded). PR context is seeded as virtual files
under /pr/; a repo-scoped App token is resolved in-graph.
- tools: read_repo_file, search_repo_code, list_review_findings.
- dashboard/review_chat_api.py + routes: per-user chat thread, LangGraph
stream/commands/state/history proxy pinned to the chat assistant, seeds
diff/findings/overview on first run. Gated by repo access.
- UI: Chat tab wired to a chat-scoped StreamProvider (replaces Coming Soon).
* feat: admin setting for review-chat default model
Add a 'Open SWE Review Chat' default to team settings (default_chat_model /
default_chat_reasoning_effort). get_team_default_model("chat") inherits the
Agent default when unset; the chat graph resolves through it. Admin RolePicker
gains an 'Agent default' inherit option that clears the override.
* feat: multi-conversation review chat (tabs, new chat, history)
Replace the single per-PR chat thread with multiple per-user conversations:
- threads minted client-side; first message persists with a title derived
from the prompt.
- list + delete endpoints; chat panel gets a tab strip (history), new-chat
(+), close (x), refresh, an intro greeting, and suggested prompts.
- get_review_chat now returns availability only (ids are client-minted).
* ui fixes
* ui: review-chat history dropdown, full-width AI replies, resizable side panel
* fix(review-chat): enforce per-user thread ownership on proxy endpoints; reseed PR context on head change
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: stream reviewer eval logs on a dedicated admin page
Stream the eval subprocess output into a rolling log_tail and persist it
during the run (was only captured at exit), so the live output is visible
while the eval runs. Move the eval runner off the admin page onto its own
/admin/evals page (linked like Review Style Prompts) with a live log viewer.
* chore: drop unrelated SSR-register drift from generated route tree
* fix(ui): pre-bundle workbox-window to stop dev re-optimize reload
The PWA service worker (devOptions.enabled) pulls workbox-window, which
Vite discovers after first render and re-optimizes, forcing a reload that
cancels in-flight code-split route imports (Failed to fetch dynamically
imported module). Pre-bundling it via optimizeDeps.include avoids the
mid-session reload.
* fix(ui): suppress html hydration warning for pre-hydration theme script
The inline theme script sets class="dark"/color-scheme on <html> before
React hydrates, so the prerendered HTML never matches. suppressHydrationWarning
on <html> silences the (expected) one-level attribute mismatch.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: trigger reviewer evals from the admin page
Add an admin-only "Reviewer eval" section + endpoints that launch the
reviewer benchmark as an isolated subprocess against the running
deployment, with live status and the LangSmith experiment link. Route
eval traces to a dedicated open-swe-evals project so they stay out of
the production tracing project.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: reconcile reviewer eval status via heartbeat, not local process
The persisted record is shared across workers but _PROCS is process-local.
The owning worker now refreshes a heartbeat while the subprocess runs, and
status is only reconciled to failed once the heartbeat is stale, so a poll on
a worker without the local handle no longer kills a live run (and a duplicate
start is rejected across workers).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* perf: speed up Reviews list (My PRs + All PRs)
Push the "My PRs" author filter into the threads.search metadata
(pr.author containment) instead of scanning up to 1000 reviewer threads
in Python, and replace the per-repo GitHub access check (an N+1 of
sequential GET /repos calls) with a single per-login accessible-repo set,
cached for 60s. Detail endpoints still re-validate access live.
Frontend: prefetch the inactive tab and adjacent page on hover/focus so
tab switches and pagination are instant.
* fix: don't cache repo access for the reviews list
The /reviews list is an authorization boundary for private PR metadata
(repo/PR titles, branches, authors, finding counts). A cross-request TTL
cache on the accessible-repo set could surface that metadata for up to
60s after a user lost repo access. Resolve the set fresh per request
instead — still a fixed, repo-count-independent burst of GitHub calls
(no per-repo N+1), with no staleness.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: mark threads resolved to hide from sidebar [closes resolve-threads]
Add a resolved flag (thread metadata) so users can clear finished
threads from the sidebar without deleting them. Resolved threads move to
a collapsible "Resolved" group (capped at 20 with a "Show all" link) and
are fully searchable/filterable + paginated on a new /agents/threads
page driven by URL query params. Sending a new message auto-unresolves.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: refresh paginated agent thread lists
Invalidate all agent thread list queries when thread lifecycle events change sidebar data, keeping the new paginated sidebar cache in sync.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* feat: render Reviews page diffs with pierre MultiFileDiff
The Reviews detail page used a hand-rolled hunk renderer with no syntax
highlighting. Switch it to the same pierre MultiFileDiff + theming the agent
chat git panel uses.
- review-diff API now returns full original/modified file contents instead of
hunks, via a shared build_pr_diff_files helper extracted from thread_api
- findings render as right-anchored markers (pierre line annotations); focus
highlight uses selectedLines; floating finding card still anchors to the marker
* feat: auto-collapse a review diff card when marked as viewed
* fix: URL-encode file path in Contents API fetch
Filenames containing reserved URL characters (#, ?) were truncated, so those
files rendered as empty/unrenderable. quote(path, safe='/') preserves the path
separators while escaping the rest.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: Reviews tab — anchored finding card, paginated list, file tree truncation
- Finding card now tracks the diff anchor while scrolling instead of staying
frozen in the viewport; auto-hides when its diff card collapses (including
collapse via mark-as-viewed) and on click outside
- Checks section capped with max height + scroll
- /reviews paginated (page size 20) with has_more, filtered to the current
user's PRs by default with an All toggle; PR author login now stored in
reviewer thread metadata
- File tree truncation marker overlapped filenames because the sidebar bg
was transparent; use the opaque sidebar color
* fix: finding card tracks anchor 1:1 while scrolling
Drop the vertical viewport clamp — it pinned the card at the clamp
boundary while the highlighted lines kept scrolling, breaking the
attachment.
* fix: anchor finding card with Base UI popover
Replace manual fixed-position tracking (laggy: setState per scroll
frame) with a Popover anchored to the finding's diff row. Floating UI
tracks the anchor outside React renders, so the card moves 1:1 with
the content and scrolls out of view with it. Unanchored findings keep
the fixed top-right card.
* fix: lock finding card to diff scroll
Replace the Base UI popover (async repositioning, paints a frame behind
native scroll) with a card absolutely positioned inside the scroll
container, so it scrolls with the diff in the same compositor frame.
Scroll moves to the ReviewBody root, side panel becomes sticky.
Position recomputes only on layout shifts via ResizeObserver.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: server-side Datadog/LangSmith observability tools + team creds
Add team-wide observability credential settings (Datadog DD_SITE/API/APP
keys, LangSmith API key) stored encrypted server-side, with an admin
dashboard section to connect/disconnect each provider. When connected,
get_agent loads read-only observability tools server-side: Datadog via its
hosted MCP server (langchain-mcp-adapters, toolsets=core) and LangSmith
read tools (langsmith_get_trace, langsmith_list_runs). Credentials live in
the LangGraph server process and are never exposed to the sandbox.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review on observability tools
Address PR review feedback:
- Authorize observability tools per triggering user (admins + the
OBSERVABILITY_AUTHORIZED_EMAILS allowlist) so prompt-injected runs from
untrusted contributors can't reach team Datadog/LangSmith data.
- Use the documented Datadog MCP auth headers DD_API_KEY / DD_APPLICATION_KEY.
- Store each provider's credentials under its own store key to avoid a
read-modify-write race dropping the other provider on concurrent saves.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: async email resolution in observability authorization gate
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: per-repo custom instructions for the coding agent
Adds per-repository custom instructions for the main coding agent,
mirroring the reviewer's per-repo style prompts. Instructions are stored
in the LangGraph Store, managed via dashboard API + UI (Monaco editor),
and appended to the agent's system prompt for runs targeting that repo.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: wire agent instructions route into generated route tree
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce repo access on instruction routes
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: let usage table expand to fill available width
The usage page was constrained to max-w-3xl while the leaderboard table
needs a 760px minimum for its 7 columns, so it always overflowed and
forced horizontal scrolling. Add an optional maxWidthClassName prop to
AppShell and widen the usage page to max-w-5xl so the table expands when
space allows, keeping overflow-x-auto as a narrow-screen fallback.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor: use cn-merged className override on AppShell
Replace the single-purpose maxWidthClassName prop with a general
className prop merged via cn (twMerge), matching the shadcn pattern. This
lets callers override any of the content-container classes ad hoc rather
than adding a new prop per override. Usage page passes className=max-w-5xl.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* feat: add dark mode support to web UI
Add a persisted, system-aware theme with a Light/Dark/System toggle in
the sidebar user menu. An inline head script applies the stored theme
before paint to avoid a flash, and the agents-ui surface gets dark
overrides for its --ui-* variables.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: use static theme init script to satisfy CodeQL
Build the no-FOUC theme script from a fully static string literal
instead of interpolating THEME_STORAGE_KEY, so no value flows into
code construction (CodeQL "Improper code sanitization").
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: avoid theme flip on hydration for persisted preferences
Only apply a theme after the stored preference is read, instead of
eagerly applying the default "system" theme on first effect. This
prevented a brief flip to the OS theme (then back) on load for users
with a saved preference that differs from their OS setting.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Adds an admin-managed, org-wide review guidelines field to team settings
that the reviewer injects into every PR review across all repos, alongside
the existing per-repo style prompt and AGENTS.md context. Repo-specific
rules take precedence when they conflict.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add scheduled web agents
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* fix: secure scheduled agent repositories
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* feat: rebuild scheduled agents as Automations tab
Scrap the inline ScheduledAgentsPanel and replace it with a dedicated
Automations tab: sidebar nav entry, list view with stat cards + empty
state, and a full editor (name, Active toggle, repo, scheduled trigger
picker, agent instructions + model).
* fix: clear collapsed-sidebar button on mobile in Automations
* fix: allow clearing automation repo on update
---------
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* fix: tidy agents home, mobile logo, and PWA dev/icon issues
- Remove recent-runs cards from the new agent page
- Scale the ASCII logo to fit narrow viewports instead of overflowing
- Serve manifest in dev and skip SW registration in dev to clear console errors
- Use a flat, opaque apple-touch-icon so iOS stops adding a black border
- Ignore generated ui/dev-dist
* feat: add send button and prevent iOS focus-zoom on chat input
- Add a circular send button (accent, spinner while sending) to the prompt bar
- Bump textarea to 16px on mobile so iOS doesn't zoom the viewport on focus
* fix: polish prompt bar and center logo
- Center the ASCII logo at all widths
- Prevent iOS focus-zoom via viewport maximum-scale instead of bumping the
input to 16px, so mobile text stays the intended size
- Give the model picker a chip/chevron treatment to anchor it next to the send button
* fix: simplify model picker to plain text + chevron
* fix: tighten prompt bar padding and shrink send button
---------
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* fix: stabilize agent defaults and repo selector
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* fix: align default repo selector styling, restore text fallback
Match the repo selector to sibling settings controls (h-7, bg-input/20,
text-xs). Fall back to a text input when the repo list is empty so a
default repo can still be entered.
* fix: make repo selector dropdown more compact
---------
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* fix: tag Slack threads with stored identity so they surface in web
process_slack_mention gated the run on mapped_login (resolved from the stable
Slack user id), but upsert_agent_thread_owner_metadata independently re-resolved
the GitHub login from the Slack profile email. When that email differs from the
user's mapping email (e.g. a personal vs work address), the lookup returned None,
so github_login was never stamped on the thread and the thread never surfaced in
the web Agents UI (which searches by github_login / triggering_user_email).
Resolve the GitHub user from the store via the Slack id, pass that login through
to the owner metadata, and use the mapping's stored work email (falling back to
the Slack profile email for unmapped users) for both the run config and the
thread tagging, so Slack-started threads reliably appear in web.
* fix: stamp github_login on Slack threads so they surface in web
process_slack_mention gated the run on mapped_login (resolved from the stable
Slack user id) but upsert_agent_thread_owner_metadata re-resolved the login from
the Slack profile email; when that email isn't the user's mapping email the
lookup returns None and github_login is never stamped, so the thread is invisible
in the web Agents UI (which searches by github_login / triggering_user_email).
Pass the already-resolved mapped_login through to the owner metadata. The
dashboard match keys on the stable GitHub login, so this is sufficient; the
triggering email stays the live Slack profile value.
* fix: preserve Slack email during account mapping
* fix: require Slack OIDC for email mappings
* chore: format Slack OIDC mapping cleanup
* fix: deliver Slack account-link prompt as a visible threaded reply
Blocked Slack users got no prompt at all. Prod logs show chat.postEphemeral
returns ok, but ephemeral messages are silently dropped in Slack's assistant
threads (where Open SWE runs), so the user sees nothing. Post the prompt as a
normal threaded reply instead — the same channel the agent uses to reply.
* fix: deliver Slack auth-failure prompt as a visible threaded reply
leave_failure_comment() tried an ephemeral message first and only fell back
to a thread reply on failure. Ephemeral messages succeed (ok) but are dropped
in Slack's assistant threads, so the fallback never fired and the user saw no
auth-failure prompt. Post the visible threaded reply directly, matching the
account-link prompt fix.
* fix: prompt blocked Slack users with a generic, token-free dashboard link
Addresses the review findings that posting the per-user account-link token /
auth URL in a visible thread lets any channel member bind their GitHub account
to the triggering user's Slack identity.
Drop the per-user signed link entirely. Both the account-link prompt
(_post_account_link_prompt) and the runtime auth-failure prompt
(leave_failure_comment) now post a plain dashboard settings link
(build_settings_url) as a visible threaded reply. The user signs in with GitHub
from their own session and connects Slack via verified OIDC on the settings
page — no secret in the thread, nothing to hijack, and no DM machinery.
* feat: nudge first-time users to connect Slack from the dashboard home
Show a Connect Slack banner on the agents landing page whenever Slack OAuth is
enabled and the user hasn't linked Slack yet. A first-time user (no Slack
mapping) sees it immediately after signing in; it disappears once connected.
* feat: prompt first-time users to connect Slack via a dialog
Replace the inline Connect Slack card on the agents home with a modal dialog
(Base UI). It opens automatically once the mapping query resolves to
"not connected" and closes itself once Slack is linked; "Maybe later" dismisses
it for the session. No new dependency — uses the design system's Base UI.
* copy: frame Slack connect as resolving the user's GitHub account
Drop 'act/reply on your behalf' wording across the connect-Slack dialog, the
Slack thread prompts (blocked + auth-failure), and the settings description.
Connecting Slack lets Open SWE resolve the user's GitHub account when they tag
it in Slack.