Commit graph

94 commits

Author SHA1 Message Date
Johannes du Plessis
f6c215fff7
feat: add automation template gallery (#1608)
Add a "Start from a template" gallery to the Automations page so users can
scaffold common scheduled agent runs (PR review digest, issue triage,
dependency checks, flaky-test tracking, release notes, docs freshness,
security audit) instead of writing every automation from scratch. Selecting
a template opens the new-automation editor prefilled with its instructions
and a default schedule, which the user can tweak before saving.

Also adds an isDescribableCron helper so template schedules render as
human-readable labels rather than raw cron in the editor.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-25 10:36:34 -07:00
Ramon Nogueira
534f402f48
feat(open-swe): copy-plan-as-markdown button + fix premature "ready" banner (#1603)
- Add a "Copy markdown" button to the plan header that copies the whole plan.
  Cross-browser: async Clipboard API in secure contexts, hidden-textarea +
  execCommand fallback for older Safari/Firefox and non-HTTPS origins.
- Fix the conversation banner: it showed "A plan is ready for your review" for
  every non-approved/cancelled status, including "planning" — so it claimed the
  plan was ready the instant plan mode began (the agent shares the link early
  to follow along), then the plan page correctly said it was still being
  written. Now: planning → "writing a plan", revising → "revising", ready →
  "ready for your review".
2026-06-24 15:07:25 -04:00
Ramon Nogueira
714914b14c
feat(open-swe): disable "Request changes" until the plan has a comment (#1602)
Requesting changes hands the reviewer comments to the agent, so it's
meaningless with none. Disable the button (with a hint tooltip) until at
least one comment exists; approve is unaffected.
2026-06-23 16:02:08 -07:00
Johannes du Plessis
9370a8c7f4
feat: inline PR comments in the reviews UI (#1600)
* feat: inline PR comments in the reviews UI

Click the diff gutter "+" on a line to open an inline comment composer
(rendered like the finding card via a Pierre annotation); submitting
posts a real inline PR review comment as the signed-in user through a
new POST /reviews/{owner}/{repo}/{number}/comments. The "+" press-drag →
"Add to Chat" selection path is unchanged.

* feat: GitHub-parity comment box, PR comments dropdown, collapse nav

- Comment composer now mirrors GitHub's box: Write/Preview tabs (markdown
  rendered via the existing Markdown component) and a markdown toolbar
  (heading, bold, italic, quote, code, link, bulleted/numbered/task list).
- Surface other people's inline PR comments in a Devin-style dropdown in the
  review header (search + link to the thread on GitHub). New
  GET /reviews/{owner}/{repo}/{number}/comments lists them and flags the
  reviewer's own (marker-bearing) comments so they're filtered out.
- Collapse the global nav by default on a review detail page, restoring the
  prior preference on leave.

* feat: bigger comment-toolbar icons; open dropdown comments inline

- Enlarge the markdown toolbar glyphs (Phosphor) in the comment composer —
  they were rendering at 10px.
- Clicking a comment in the PR comments dropdown now opens it inline in the
  diff as a read-only finding-style card (InlineComment), scrolling its line
  into view, instead of navigating to GitHub. Falls back to GitHub when the
  comment's file/line isn't in the current diff.

* fix: drive "Add to Chat" from native text selection

The gutter "+" is now comment-only; wiring its click to the composer
conflicted with its old double-duty as the drag-to-select handle, which
broke selection → "Add to Chat". Switch to Devin's model: disable Pierre's
interactive line selection and instead map a native text highlight in the
diff to a line range (via the data-line / data-line-type attributes Pierre
stamps on each line, read from the diff's open shadow root) to show the
"Add to Chat" popup. ⌘L and the existing attachment/popup path are unchanged.

* feat: gutter "+" drag selects a range for multi-line comments

Re-enable Pierre's gutter line selection so dragging the "+" down the
gutter comments across a range (click still comments on a single line);
onLineSelectionEnd routes the range to the composer. Native code-text
selection still drives "Add to Chat" — Pierre only line-selects from the
gutter, and onLineSelectionEnd bails when a native text selection is
present, so a code highlight never opens the composer.

* fix: keep the range highlighted while its comment composer is open

Previously opening the composer cleared the selection, so the lines being
commented on lost their highlight. Drive the controlled selection from the
open comment draft's range so the rows stay highlighted until the composer
is closed.

* fix: address PR review — paginate comments, fall back for outdated ones

- list_review_comments now pages through all PR review comments (bounded by
  _MAX_REVIEW_COMMENT_PAGES) instead of returning only the first 100, so older
  comments still show in the dropdown.
- Surface GitHub's outdated flag (position == null) as is_outdated; opening such
  a comment (or one whose line isn't in the diff) now opens it on GitHub instead
  of silently rendering nothing, plus a timeout fallback if the annotation never
  mounts (e.g. collapsed context).
2026-06-23 16:01:46 -07:00
Ramon Nogueira
ca9280d25c
refactor(open-swe): plain HTTP comments instead of Yjs/BlockNote collab (#1601)
* feat: add plan mode for read-only research and planning

Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: enforce plan-mode read-only at tool layer and disable subagents

Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: harden plan-mode shell guard against wrapped mutations

Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow

- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline

Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.

* test(open-swe): add Playwright E2E for the Slack → PR → web handoff

Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.

- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).

Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.

* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright

The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.

Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.

* test(open-swe): record Playwright trace + video on every E2E run

Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.

* feat(plan-mode): collaborative plan review with BlockNote + Yjs

When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.

- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
  the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
  snapshots; plan content/status store; plan REST API (get/approve/reject,
  owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
  mounted under the agents shell, with a "Review plan" banner in the thread view
  and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
  flow, including cross-user comment sync and owner-only approval.

* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)

- plan-collab WS: authorize per-thread before joining a room (same read gate as
  the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
  opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
  resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
  dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
  unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
  before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.

Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).

* style: ruff format plan_collab.py

* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS

- Slack "Approve & Implement" now verifies the clicking user is the plan
  requester (owner, via the stored triggering_user_id) before implementing —
  matching the dashboard API's owner-only approval. Non-owners are pointed to
  Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
  allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
  the REST require_same_origin CSRF defense.

* fix(plan-mode): enter plan mode only via the model + local mock dev harness

Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.

- enter_plan_mode returns a terminating ToolMessage, fixing the missing
  ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
  remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
  call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
  Alice/Bob mock users, and a GitHub login picker.

* docs(plan-mode): drop stale references to removed profile/team defaults

The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.

* feat(plan-mode): let any reviewer edit the plan, not just comment

Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.

* test(plan-mode): assert plan-mode entry via the tool's success message

plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.

* refactor(plan-mode): replace Yjs/BlockNote collab with plain HTTP comments

Drop the realtime collaborative editor (it can't work behind Vercel's
rewrite — WebSocket upgrades aren't proxied to the external LangGraph
backend) in favor of a simple whole-document comments API over plain HTTP.

Backend:
- Remove the Yjs WebSocket server (plan_collab.py), its lifespan, and the
  collab router; drop pycrdt / pycrdt-websocket deps.
- plan_store: replace the Yjs snapshot with comment CRUD (one store item per
  comment under ["plan","comments",thread_id]).
- plan_api: add GET/POST/DELETE comment endpoints; approve/reject now read
  comments server-side and format them for the follow-up run (no longer
  client-harvested). Comment delete is author-or-owner; approve stays owner-only.

Frontend:
- PlanReview renders the plan markdown read-only and shows a comments panel
  (list + add, polled every 4s for cross-user visibility).
- Drop @blocknote/*, y-websocket, yjs; lib/plan exposes get/add/deletePlanComment.

Tests: unit tests for the comments API + route registration; e2e drives the
HTTP comment UI (owner + collaborator, cross-user visibility, owner-only approve,
PR echoes the harvested feedback).

* fix(open-swe): clear stale plan comments on republish; fail loud on store errors

Address reviewer feedback:
- Clear comments when a revised plan is published (save_plan_content) so
  feedback on the prior revision doesn't resurface and get re-fed to the agent.
- list_plan_comments gains raise_on_error; approve/reject read comments before
  mutating state and propagate store failures (500) instead of silently
  dispatching the follow-up run with no feedback.

---------

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 22:12:42 +00:00
Caroline di Vittorio
63e4c2baac
feat: relocate review chunks and embed review in git panel (#1590)
* feat: relocate review chunks and embed review in git panel

Move the AI-sorted review chunk navigator out of the global left sidebar
and into the review page's own main panel (sidebar keeps the thread list),
and surface the review main body inside the agent thread's Git > Review
sub-tab with an expand link to the full review page.

The review main body + side panel + helpers are extracted into a shared
ReviewMainBody component with "full" and "embedded" variants.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: move ReviewTab into its own file

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:27:51 -07:00
Johannes du Plessis
3992d3ef5d
feat: repo-scoped dynamic sandbox snapshots (#1595)
* feat: repo-scoped dynamic sandbox snapshots

Let admins build a per-repo sandbox image from a custom Dockerfile so runs
targeting that repo boot from a snapshot with its deps pre-baked. Snapshot
selection is purely additive: repos without a `ready` repo-scoped snapshot
always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID.

Backend adds a repo_snapshots store module (Dockerfile + build status keyed by
owner/name), threads the resolved repo through the LangSmith sandbox creation
path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a
throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The
UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor +
build status/logs).

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: harden repo snapshot builds

Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins
cannot accidentally build a repo snapshot from a bare Python image that lacks
Open SWE's sandbox tools. Allow stale building records to be retried by tracking
build_started_at and treating old or missing timestamps as stale.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: document repo snapshot base image config

Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert
missing base-image configuration into a handled dashboard API error so admins see
a clear configuration message instead of an unhandled template-generation error.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
Johannes du Plessis
60663800e5
fix: prevent reviews page crash on binary/large file diffs (#1597)
The PR reviews page builds a Pierre FileContents cache key for every
file in an unconditional useMemo. Binary/oversized/added/removed blobs
arrive with null originalContent/modifiedContent (pr_diff.py flags them
unrenderable), so fileContentsCacheKey dereferenced null via
contents.length and crashed the whole route — escaping the markdown
error boundary, which was unrelated. Coerce null/undefined contents to
"" in the cache-key helper so it can never throw.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:21:20 -07:00
Ramon Nogueira
3a0e2b4672
feat: plan mode with model-driven entry and collaborative review (#1580)
* feat: add plan mode for read-only research and planning

Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: enforce plan-mode read-only at tool layer and disable subagents

Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: harden plan-mode shell guard against wrapped mutations

Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow

- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline

Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.

* test(open-swe): add Playwright E2E for the Slack → PR → web handoff

Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.

- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).

Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.

* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright

The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.

Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.

* test(open-swe): record Playwright trace + video on every E2E run

Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.

* feat(plan-mode): collaborative plan review with BlockNote + Yjs

When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.

- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
  the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
  snapshots; plan content/status store; plan REST API (get/approve/reject,
  owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
  mounted under the agents shell, with a "Review plan" banner in the thread view
  and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
  flow, including cross-user comment sync and owner-only approval.

* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)

- plan-collab WS: authorize per-thread before joining a room (same read gate as
  the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
  opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
  resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
  dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
  unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
  before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.

Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).

* style: ruff format plan_collab.py

* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS

- Slack "Approve & Implement" now verifies the clicking user is the plan
  requester (owner, via the stored triggering_user_id) before implementing —
  matching the dashboard API's owner-only approval. Non-owners are pointed to
  Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
  allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
  the REST require_same_origin CSRF defense.

* fix(plan-mode): enter plan mode only via the model + local mock dev harness

Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.

- enter_plan_mode returns a terminating ToolMessage, fixing the missing
  ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
  remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
  call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
  Alice/Bob mock users, and a GitHub login picker.

* docs(plan-mode): drop stale references to removed profile/team defaults

The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.

* feat(plan-mode): let any reviewer edit the plan, not just comment

Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.

* test(plan-mode): assert plan-mode entry via the tool's success message

plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.

---------

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:06:58 -07:00
Ramon Nogueira
6117780982
feat(open-swe): let any org member post to a thread, with attribution (#1594)
* feat(dashboard): let any org member post to a thread, with attribution

Posting to an Agents chat thread from the web UI was restricted to the
thread owner. Open it to any authenticated org member (login is already
org-gated by OAuth) on both write paths — the queued follow-up
(send_dashboard_message) and the idle-thread run.start
(_enrich_run_start_command). Non-owner messages are prefixed with the
poster's verified GitHub login (@login:) so the agent and owner can tell
who sent them. Thread management (cancel/delete/resolve) stays owner-only,
and the UI now shows the composer to non-owners.

* fix(dashboard): keep non-run.start commands owner-only

Non-owner posting is allowed only via the attributed run.start path. Other
write commands (e.g. input.respond) carry unattributed user input, so the
commands proxy keeps them owner-only instead of readable-by-any-org-member.

* docs(e2e): drop per-test details from the E2E README
2026-06-23 11:06:53 -07:00
Johannes du Plessis
0bff510aae
fix: render GitHub-hosted images in PR descriptions on reviews page (#1589)
* fix: render GitHub-hosted images in PR descriptions on reviews page

PR description images hosted on GitHub (user-attachment uploads and
*.githubusercontent.com) render broken on the reviews page because
private-repo attachments require GitHub auth the browser session lacks.
Add an authenticated backend image proxy (host-allowlisted to guard
against SSRF) and route those image URLs through it from the reviews UI.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: harden PR image proxy (IDOR, SVG XSS, unbounded buffering)

Address review findings on the review-page image proxy:

- IDOR: the proxy fetched any *.githubusercontent.com URL with the App
  installation token, gated only by route-param repo access, so a user
  authorized for one repo could read images from another private repo the
  App can see. Bind the URL to the authorized PR — only proxy URLs that
  appear in that PR's body.
- SVG XSS: served any image/* inline from the API origin, including
  image/svg+xml which can run script. Restrict to safe raster types and
  add X-Content-Type-Options: nosniff + a locked-down CSP.
- DoS: enforced the size cap only after buffering the full response.
  Stream and abort once the cap is exceeded.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-22 14:01:06 -07:00
Johannes du Plessis
609e5551c4
fix: contain markdown render crashes with an error boundary (#1586)
Streamdown bundles Mermaid and renders ```mermaid blocks itself; a
diagram it can't parse throws during render and, with no boundary,
white-screens the whole review page. Wrap the markdown render so any
crash falls back to the raw text instead.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-22 12:54:05 -07:00
Johannes du Plessis
855d0fcb42
feat: render review findings as inline annotations (#1581)
Replace the floating finding cards with inline annotations in the PR diff.
An anchored finding shows a collapsed header at its line and expands the full
details in place; non-anchored findings expand inline in the side-panel row.
Expand state is lifted into a context so it survives Pierre's annotation
windowing and the side panel can drive it.

Removes the anchored/fallback floating cards and their rAF/ResizeObserver
scroll-tracking. Also stops the line-selection highlight from bleeding behind
the inline annotation row.
2026-06-19 20:24:05 -07:00
Johannes du Plessis
75e594c861
fix: reviewer reviews full diff; fix review UI scroll + dark-mode composer (#1575)
* fix: reviewer reviews full diff; fix review UI scroll + dark-mode composer

- Drop the 200K-char PR diff truncation. The head/tail slice silently dropped
  whole files out of the middle; the reviewer's large context window reviews
  the complete diff. The grouping pass uses the full diff too. Other #1567 caps
  (fetch_url, Slack, pagination, message queue) are kept.
- Sidebar file/group click now lands flush at the top under virtualization:
  re-assert alignment after the Virtualizer's mid-scroll height reconciliation
  settles, instead of trusting scrollIntoView's estimate-based target.
- Review chat composer textarea no longer shows a lighter square in dark mode
  (override the Textarea's dark:bg-input default with dark:bg-transparent).

* fix: anchored finding card sits flush in the diff gutter, no side-panel overlap

The card was positioned next to the anchor and clamped to window.innerWidth, so
when the anchor sat near the diff/side-panel boundary the 412px card spilled
across into the side panel. Pin it flush to the diff column's right edge and
clamp its width to the column so it never overlaps the side panel and shrinks to
fit a narrow column.

* fix: anchored finding card sits over the side panel, not the diff

Flip the horizontal anchor: place the card flush against the diff column's right
edge and extend it rightward over the side panel, instead of leftward over the
diff. Width still caps at the preferred size and shrinks to fit a narrow panel.

* fix: anchor finding card to diff content edge + small-screen fallback

- Anchor to the centered diff content's right edge instead of the full-width
  scroller, removing the centering-whitespace gap so the card sits closer to the
  diff.
- Below xl the side panel is hidden and the diff fills the viewport, leaving no
  gutter; positioning against the content edge would collapse the card to ~0px.
  Fall back to overlaying the diff flush-right when there's no usable gutter
  (addresses Open SWE review finding f_f6ea77527c).

* fix: anchor finding card next to the annotation, close to the hunk

Anchor the card's left to the finding's annotation (anchorRect.right) instead of
the diff content's right edge, so it sits right by the hunk and overlaps the diff
edge rather than parking out in the gutter. Extends right over the side panel,
clamped to stay on-screen (also keeps it readable below xl with no side panel).

* fix: finding card width shrinks to fit a small side panel

Size the card to the room to the right of the annotation (capped at the
preferred width) so it narrows as the side panel shrinks instead of overflowing.
Below a readable minimum, hold that width and shift left over the diff.
2026-06-18 19:24:52 -07:00
Johannes du Plessis
ecf0898f51
feat: split view, add-to-chat, virtualization + scroll/grouping perf (#1574)
* feat: reviews page split view, add-to-chat, virtualization + perf

- Virtualize the diff (Pierre Virtualizer + worker pool), mirroring the agent
  chat panel, so large PRs window rows instead of materializing every line.
- Split chat and diff into independent scroll containers and make chat
  auto-scroll fully contained, so typing/streaming no longer moves the diff.
- Memoize FileDiffCard with stable callbacks so focusing a finding re-renders
  only the affected card.
- Add a persisted unified/split diff toggle.
- Add highlight-to-chat: select lines (drag / shift-click) + gutter "+" to drop
  a file:line snippet into the chat composer.
- Rebuild sidebar group rows: whole card scrolls to the group (incl. the
  expanded explanation), Read explanation stays a separate toggle, memoized.

* feat: add-to-chat uses attachment pills + selection popup / ⌘L

Replace the raw-snippet injection with a Cursor-style flow:
- Selecting lines shows a floating "Add to Chat ⌘L" popup at the pointer; ⌘L
  adds the current selection without it. Removes the auto-adding gutter "+".
- "Add to chat" now creates a removable attachment pill in the composer (and a
  pill in the sent message bubble) instead of pasting raw text. The code is
  still serialized into the message content so the model receives it as context.

* feat: restore gutter + drag-handle for line selection

Re-enable Pierre's gutter '+' as a click-and-drag line selector (with the
highlight growing as you drag) — the affordance that was lost when the
auto-adding gutter button was removed. It no longer auto-adds: the commit flows
through onLineSelected to the 'Add to Chat' popup / ⌘L.

* fix: live selection highlight while dragging + reposition Add to Chat popup

- Feed onLineSelectionChange into the controlled selection so rows highlight
  live as you drag, not just on release (Pierre only paints the controlled
  selection when the prop updates). Popup now fires on onLineSelectionEnd.
- Anchor the popup's bottom-left to the drag handle (drop horizontal centering)
  so it no longer overlaps the '+' button.

* fix: anchor Add to Chat popup to gutter handle + click-away to unselect

- Position the popup from the gutter '+' handle's rect (in the diff shadow DOM,
  placed on the selection's bottom line) instead of the pointer-release point,
  which landed inconsistently. Falls back to the pointer if not found.
- Clear the line selection (and popup) on any outside pointer-down.

* fix: sidebar-collapse header overlap + PR review comments

- Lift useSidebarLayout to AgentsShell (single source), share collapsed via
  context, and pad the reviews header left when the sidebar is collapsed so the
  fixed collapse toggle no longer overlaps the header content.
- add-to-chat: collect each diff side separately so a selection that spans a
  deletion->addition no longer pastes wrong-file lines (PR comment).
- chat: clear attachments after sending via a suggested prompt, so an attached
  snippet isn't silently resent on the next message (PR comment).

* fix: anchored finding card positions to the right of the diff again

The virtualization refactor moved the card inside the main-width Virtualizer
scroller, so it clamped over the diff. Render it in the outer container as a
viewport-fixed card clamped to window width (right gutter / over the side panel,
like prod) and track the finding as the diff scrolls (rAF-throttled), hiding
when the finding scrolls out of view.
2026-06-18 16:54:07 -07:00
Johannes du Plessis
39a26e16b5
fix: optimize agent thread lists (#1570)
* fix: optimize agent thread lists

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: refresh missing thread run status

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-18 10:09:22 -07:00
Johannes du Plessis
c07434a221
fix: allow read-only cross-user access to agent threads via Open in Web links (#1568)
The thread detail endpoint already returned metadata for non-owners, but
the transcript hydration endpoints (state, stream/events, history, pr-diff)
all asserted ownership and 404-ed. This caused the UI to redirect non-owners
back to /agents when they clicked an "Open in Web" link shared in Slack.

Dashboard login is already gated by ALLOWED_GITHUB_ORGS, so any logged-in
user is a trusted org member. This commit:
- Adds _thread_is_readable / _assert_thread_readable helpers that grant
  read access to any surfaced-source thread for authenticated users
- Relaxes read endpoints (state, stream/events, history, pr-diff, SSE
  stream) to use readable checks instead of ownership checks
- Keeps write endpoints (send message, cancel, delete, resolve, run
  commands) owner-only
- Adds an isOwner field to the thread summary so the frontend can render
  a read-only mode (hides the prompt bar, resolve/delete buttons)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-18 09:10:09 -07:00
Johannes du Plessis
e4d737e18c
fix: optimize agent git diff rendering (#1564)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 13:19:53 -07:00
Johannes du Plessis
d2c232cc54
feat: add browser notifications for agent run completion (#1558)
* feat: add browser notifications for agent run completion

Request notification permission when a user starts their first agent
run from the home page. A toggle in Profile Settings lets users
enable/disable desktop notifications. When a run transitions from
running to a terminal state (finished/error/interrupted), a browser
notification is fired — suppressed for the thread the user is
currently viewing.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: notify for active thread when page is in background tab

Only suppress the notification for the active thread when the page is
actually visible. If the user switched to another browser tab, the
notification fires even for the thread they have open.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 09:24:44 -07:00
Johannes du Plessis
0b6806c4ba
fix: make git panel a full-screen overlay on mobile widths (#1559)
At mobile widths the resizable git panel collided with the chat column's
360px min-width, overflowing the viewport. Treat mobile (<768px) like the
sidebar: the git panel becomes a full-screen overlay the user navigates to
and back from, while desktop keeps the inline resizable panel.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 09:24:28 -07:00
Johannes du Plessis
b19804536c
feat: handle images sent to non-vision models in Slack, Linear, and web UI (#1560)
* feat: handle images sent to non-vision models in Slack, Linear, and web UI

Add vision capability checks across all image input paths. When a user
sends images to a text-only model (e.g. GLM 5.2, DeepSeek V4 Pro), the
images are now skipped and a warning is injected into the prompt instead
of sending unsupported content to the model.

- Slack: resolve model at webhook time, skip image fetch + add warning
- Linear: same pattern as Slack
- Queued message middleware: read resolved model from thread metadata,
  strip images from queued payloads for text-only models
- Web UI: disable submit + show inline warning when images are attached
  to a non-vision model selection
- Shared: resolve_agent_model_id helper + vision_not_supported_warning

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: mock resolve_agent_model_id in Slack mention test

The test_process_slack_mention_queues_active_thread_message test was
missing a mock for the new resolve_agent_model_id call added to the
Slack webhook handler, causing a TypeError when image URLs triggered
the model resolution path.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: include vision warning in queued payload for text-only models

Update the prompt variable (not just content_blocks) before clearing
image_urls so the queued payload also carries the warning text when a
Slack/Linear follow-up arrives while the thread is busy.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 09:12:52 -07:00
Johannes du Plessis
3945757592
feat(ui): reintroduce subtle blue accents on greyscale frontend (#1555)
Restore a subtle blue --ui-accent token (driving send/stop buttons and
accent spots), the blue git status indicators for renamed/copied files,
and color the thread-list diff stat badge. Keeps the overall greyscale
aesthetic with tasteful blue/green/red accents.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 19:41:19 -07:00
Johannes du Plessis
bac1591888
Neutralize frontend: blue accents → neutral aesthetic (#1553)
Strip the blue/cool cast from the agent + reviewer chat and shared UI.
Neutralize the --ui-* (agents.css) and shadcn (styles.css) color tokens
to pure neutral grays (chroma/hue zeroed, lightness preserved); primary
and filled buttons become inverse-foreground neutral solids. Replaces a
few hardcoded blue/cyan spots (send button, prompt bar bg, context ring,
loop dot, git renamed status) with neutral tokens. Semantic green/amber/
red are kept.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 15:49:33 -07:00
Johannes du Plessis
e58b609b2f
fix: Simplify review explanation: full-width, plain prose, no diff links (#1547)
* Simplify review explanation: full-width, plain prose, no diff links

* Update _build_prompt test for plain diff fences

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 15:32:24 -07:00
Johannes du Plessis
876836bdfe
Normalize frontend typography + enable font smoothing (#1546)
Lighter, more consistent type across the dashboard and agent UI:
unify title sizes, drop semibold/bold in chrome to medium, harmonize
stray text-sm to text-xs, normalize chat markdown headings, and enable
antialiased font smoothing (fixes heavy text on dark backgrounds).

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 14:56:11 -07:00
Johannes du Plessis
801f93b4de
feat: AI-sorted PR review view with diff grouping (#1544)
* feat: AI-sorted PR review view with diff grouping

Group a PR's changed files into a logical top-to-bottom walkthrough via a best-effort structured-output LLM pass, kicked off concurrently with the reviewer run (~0 added latency) and persisted on reviewer thread metadata. Results render in the review UI behind an AI sorted / file tree toggle that persists across PRs; the view falls back to the file tree when groups are absent or stale.

Adds a grouping-model team default (inherits the Reviewer subagent model when unset) and the admin RolePicker for it.

* feat(reviews): richer AI-sorted explanations + sidebar polish

Sidebar group rows get Devin-style spacing (dividers, padding), a per-group file list (click a file to jump to its diff), inline-code chips in titles, and an accent Read explanation link.

Group explanations are now rich markdown: the grouping prompt feeds per-hunk line ranges and asks for inline code, a short code block, and [path:line](#loc=...) references. The Markdown renderer turns those #loc= links into in-page buttons that scroll the diff to the hunk and highlight the range, reusing the existing selectedLines path.

* fix(reviews): drop stale diff groups from the AI-sorted view

When groups were generated for a previous head, a persisted "ai" view in localStorage still rendered the outdated walkthrough. groupedView now returns null on diff_groups_stale, so the file-tree fallback is used and the view toggle hides until fresh groups arrive.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 13:59:37 -07:00
Johannes du Plessis
272ffc78a6
feat: flatten reviews file tree and refine pierre tree styling (#1541)
* feat(ui): flatten reviews file tree and refine pierre tree styling

Switch the reviews-page file tree off forced full expansion so it renders
as a compact, flattened directory tree (single-child chains merged into one
row) like the pierre "flattened directories" preset. Expand only the
selected file's ancestors so the active diff stays revealed. Refine the
shared tree theme: subtle accent-tinted hover/selection and a complete
git-status palette (renamed/untracked/ignored) plus search input fg.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix(ui): neutralize pierre tree filenames and selection styling

Use neutral grey/white filename colors instead of git-status accents,
apply unsafe CSS overrides for selected-row contrast, and switch to
complete icons in both agent and review file trees.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 10:54:30 -07:00
Johannes du Plessis
3021bbe4b5
feat: add default model selection to onboarding dialog (#1542)
Extend the first-run Slack connect modal into a two-step onboarding
dialog that first prompts new users to pick a default agent model
(persisted to their profile) before connecting Slack, so model choice
happens up front during onboarding alongside Slack authentication.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 10:33:11 -07:00
Johannes du Plessis
0108d764d9
fix(ui): resolve eslint errors in dashboard components (#1538)
Clears ~83 pre-existing eslint errors in the UI: auto-fixable rules
(array-type, import/order, sort-imports, type-only imports, redundant
assertions/conditions) plus manual fixes for unnecessary conditions and
banned @ts-nocheck directives. The 8 ported components excluded from
tsconfig are now also ignored by eslint so type-aware linting no longer
fails to parse them.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 17:29:38 -07:00
Caroline di Vittorio
3297e799de
fix: keep chat prompt bar editable while a run streams (#1533)
* fix: keep prompt bar editable while a run streams

The follow-up submit path awaited stream.submit, which resolves only when the
run finishes. That kept the react-query mutation isPending for the whole run,
so AgentThreadView disabled the prompt bar textarea the entire time the agent
was streaming - the user saw the "queue next" placeholder but couldn't click
in or type. Fire the run without awaiting its full lifecycle (matching the
new-thread path in AgentsHome) so the input stays editable for queueing.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor(ui): guard double-submit internally instead of via disabled prop

Stop threading a run-lifecycle signal (the send mutation's isPending) into the
prompt bar to gate the input. Instead, the prompt bar owns a synchronous
double-submit guard (submittingRef) plus a short-lived isSubmitting state
scoped to the in-flight send. The textarea now stays editable while a run
streams (so follow-ups can be queued), and onSubmit is awaitable so the guard
tracks the actual send request rather than the run.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix(ui): surface failed run-start instead of leaving thread running

When stream.submit rejects (e.g. 401 expired token or 409 active-run
race), the fire-and-forget catch was a no-op while onSuccess had already
optimistically set status: running, leaving the thread stuck in a busy
state with no surfaced error. Clear the busy state and mark the thread
errored on submit failure.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-06-15 17:29:17 -07:00
Johannes du Plessis
911c835c2a
feat: chat with your PR on the review page (#1534)
* feat: chat with your PR on the review page

Add a sandbox-less `chat` graph that answers questions about a single PR
from its diff, the published review findings, and read-only GitHub access.

- agent/chat.py: deepagents graph, no sandbox (default StateBackend, file
  mutation + execute tools excluded). PR context is seeded as virtual files
  under /pr/; a repo-scoped App token is resolved in-graph.
- tools: read_repo_file, search_repo_code, list_review_findings.
- dashboard/review_chat_api.py + routes: per-user chat thread, LangGraph
  stream/commands/state/history proxy pinned to the chat assistant, seeds
  diff/findings/overview on first run. Gated by repo access.
- UI: Chat tab wired to a chat-scoped StreamProvider (replaces Coming Soon).

* feat: admin setting for review-chat default model

Add a 'Open SWE Review Chat' default to team settings (default_chat_model /
default_chat_reasoning_effort). get_team_default_model("chat") inherits the
Agent default when unset; the chat graph resolves through it. Admin RolePicker
gains an 'Agent default' inherit option that clears the override.

* feat: multi-conversation review chat (tabs, new chat, history)

Replace the single per-PR chat thread with multiple per-user conversations:
- threads minted client-side; first message persists with a title derived
  from the prompt.
- list + delete endpoints; chat panel gets a tab strip (history), new-chat
  (+), close (x), refresh, an intro greeting, and suggested prompts.
- get_review_chat now returns availability only (ids are client-minted).

* ui fixes

* ui: review-chat history dropdown, full-width AI replies, resizable side panel

* fix(review-chat): enforce per-user thread ownership on proxy endpoints; reseed PR context on head change

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 17:17:30 -07:00
Caroline di Vittorio
698613e6db
fix: keep sidebar thread rows at a fixed height on hover (#1536)
Hover swaps the PR icon/badge (which has its own padding) for the
resolve/delete action buttons, changing the row's natural height. Pin
the row to a fixed height so the swap no longer shifts its size.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 16:15:27 -07:00
Caroline di Vittorio
dc70ca1f2f
feat(ui): soften bash tool output scroll indicator to a fade (#1532)
Replace the hard inset box-shadow scroll indicators on bash/shell tool
output with a mask-image gradient so the output softly fades at the
scroll edges, matching the user message bubble treatment.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 20:39:52 +00:00
Caroline di Vittorio
147a393ed0
feat(ui): soften long-message scroll indicator to a fade (#1531)
Replace the hard inset box-shadow scroll indicators on long user
messages with a mask-image gradient so the text softly fades at the
scroll edges instead of showing a heavy shadow.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 13:32:05 -07:00
Johannes du Plessis
3deb3ef4e0
fix: show subagent current status instead of full tool-call list (#1528)
The subagent card listed every nested tool call, ballooning the card as
the subagent ran. Show a single status line (current activity + step
count) until a richer activity UI lands.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 10:53:28 -07:00
Johannes du Plessis
dfa3fc13b9
fix: enforce minimum chat width so the git panel can't squish it (#1526)
The git panel can now be dragged toward full window width, but its max was
reserved against window.innerWidth (ignoring the sidebar), letting the panel
squeeze the chat below a usable width. Clamp the panel against the actual
container width and give the chat column a matching min-width floor.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 10:52:45 -07:00
Johannes du Plessis
acae8b9419
feat: let agent git panel resize to nearly full window width (#1525)
The git panel was capped at a fixed 720px max width, which wasted space on
wide/ultrawide screens. Make the max width dynamic relative to the viewport
(window width minus a minimum chat width) so it can be dragged up to ~50/50
or wider, and re-clamp on window resize.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 09:55:19 -07:00
Johannes du Plessis
2b181a77a6
fix: add bottom padding to shell command card before output streams (#1523)
The bash command card had only pb-1 on the header block, leaving the
content cramped against the card's bottom edge until output streamed in
(which adds its own pb-2). Apply pb-3 when there is no output yet so the
card looks uniformly padded in the pending/in-progress state.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 09:38:30 -07:00
Caroline di Vittorio
e8bb6b497b
feat: mark threads as resolved to hide from sidebar (#1500)
* feat: mark threads resolved to hide from sidebar [closes resolve-threads]

Add a resolved flag (thread metadata) so users can clear finished
threads from the sidebar without deleting them. Resolved threads move to
a collapsible "Resolved" group (capped at 20 with a "Show all" link) and
are fully searchable/filterable + paginated on a new /agents/threads
page driven by URL query params. Sending a new message auto-unresolves.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: refresh paginated agent thread lists

Invalidate all agent thread list queries when thread lifecycle events change sidebar data, keeping the new paginated sidebar cache in sync.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-06-12 12:01:42 -07:00
Johannes du Plessis
09d5d00e59
feat: render Reviews page diffs with pierre MultiFileDiff (#1517)
* feat: render Reviews page diffs with pierre MultiFileDiff

The Reviews detail page used a hand-rolled hunk renderer with no syntax
highlighting. Switch it to the same pierre MultiFileDiff + theming the agent
chat git panel uses.

- review-diff API now returns full original/modified file contents instead of
  hunks, via a shared build_pr_diff_files helper extracted from thread_api
- findings render as right-anchored markers (pierre line annotations); focus
  highlight uses selectedLines; floating finding card still anchors to the marker

* feat: auto-collapse a review diff card when marked as viewed

* fix: URL-encode file path in Contents API fetch

Filenames containing reserved URL characters (#, ?) were truncated, so those
files rendered as empty/unrenderable. quote(path, safe='/') preserves the path
separators while escaping the rest.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-12 11:03:51 -07:00
Brendan Whiting
7dac89a3b4
feat: add UI hotkey wiring and Cmd/Ctrl+B sidebar toggle (#1514)
* feat: add UI hotkey wiring and Cmd/Ctrl+B sidebar toggle

Add a reusable useHotkey hook for registering global keyboard
shortcuts (with 'mod' resolving to Cmd on macOS, Ctrl elsewhere),
and wire up the first shortcut: mod+b toggles the left sidebar.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: ignore key repeat for sidebar toggle hotkey

Holding Cmd/Ctrl+B fired repeated keydown events, toggling the sidebar
multiple times. Add an ignoreRepeat option to useHotkey and enable it for
the sidebar toggle so one held press toggles once.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: Brendan Whiting <16016903+bwhiting2356@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-06-12 10:53:28 -07:00
Johannes du Plessis
73b3ba7930
feat: collapse finished agent work into a Worked for X dropdown (#1509)
Once an agent turn finishes, its reasoning/tool/exploration steps collapse
behind a single "Worked for …" toggle so the chat shows only the final reply,
mirroring Devin's transcript. Work stays expanded live and auto-collapses on
completion; the final reply text/cards remain visible.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-12 10:08:12 -07:00
Brendan Whiting
9259a52b2f
fix: add cursor-pointer to sidebar expand/collapse buttons (#1511)
The sidebar expand and collapse toggle buttons were missing the
cursor-pointer style, so hovering didn't show a pointer cursor.

Co-authored-by: Brendan Whiting <16016903+bwhiting2356@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-12 09:49:26 -07:00
Brendan Whiting
639058a8b6
feat: support pasting images from clipboard into prompt bar (#1512)
The prompt bar only accepted images via drag-and-drop or the file picker.
Pasting an image from the clipboard inserted the file name as text instead
of attaching the image. Add an onPaste handler that extracts image files
from the clipboard and attaches them as pending images.

Co-authored-by: Brendan Whiting <16016903+bwhiting2356@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-12 09:48:48 -07:00
Johannes du Plessis
3e105f5027
fix: Reviews tab — anchored finding card, paginated list, file tree truncation (#1507)
* fix: Reviews tab — anchored finding card, paginated list, file tree truncation

- Finding card now tracks the diff anchor while scrolling instead of staying
  frozen in the viewport; auto-hides when its diff card collapses (including
  collapse via mark-as-viewed) and on click outside
- Checks section capped with max height + scroll
- /reviews paginated (page size 20) with has_more, filtered to the current
  user's PRs by default with an All toggle; PR author login now stored in
  reviewer thread metadata
- File tree truncation marker overlapped filenames because the sidebar bg
  was transparent; use the opaque sidebar color

* fix: finding card tracks anchor 1:1 while scrolling

Drop the vertical viewport clamp — it pinned the card at the clamp
boundary while the highlighted lines kept scrolling, breaking the
attachment.

* fix: anchor finding card with Base UI popover

Replace manual fixed-position tracking (laggy: setState per scroll
frame) with a Popover anchored to the finding's diff row. Floating UI
tracks the anchor outside React renders, so the card moves 1:1 with
the content and scrolls out of view with it. Unanchored findings keep
the fixed top-right card.

* fix: lock finding card to diff scroll

Replace the Base UI popover (async repositioning, paints a frame behind
native scroll) with a card absolutely positioned inside the scroll
container, so it scrolls with the diff in the same compositor frame.
Scroll moves to the ReviewBody root, side panel becomes sticky.
Position recomputes only on layout shifts via ResizeObserver.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 17:52:20 -07:00
Caroline di Vittorio
3724014d80
feat: persist git panel open/closed state to localStorage (#1506)
* feat: persist git panel open/closed state to localStorage

Make the agent thread git panel remember whether the user opened or
closed it. The collapsed state now reads from and writes to
localStorage, so it stays open or closed across thread navigation and
page reloads instead of always re-collapsing per thread.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: use plain function and named constants for git panel collapse

Drop the unnecessary useCallback around setCollapsed and replace the
"1"/"0" string literals with COLLAPSED_STATE_TRUE / COLLAPSED_STATE_FALSE
constants.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 17:09:28 -07:00
Caroline di Vittorio
d10cc15e82
fix: use gray code-box background for changed-files diff card (#1505)
The turn changed-files diff card used the blue accent-bubble color
that matches user messages; switch it to the gray code-bubble used by
code blocks so it visually reads as agent output.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 17:09:03 -07:00
Johannes du Plessis
bfcb714c5c
feat: PR review page (#1495)
* feat: PR review page

* feat: manual review trigger card on admin page

* feat: finding focus mode with floating card and dimmed backdrop

* fix: dim only info panel, subtle block highlight for focused finding

* refactor: move PR reviews into agents page with file-tree sidebar

* fix: address review feedback

- paginate list_reviews until enough accessible records collected
- include rename/metadata-only files in diff API with empty hunks
- remount ReviewBody on head_sha change so viewed-files state resets

* fix: review tree background + finding card anchoring

* fix: remove no-op PR size chip and dead check links

* fix: parallel diff fetch, complete reReview type, shared github_headers

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 16:11:10 -07:00
Johannes du Plessis
e0678e8c01
feat: track PR lifecycle state per thread for sidebar (#1492)
* feat: track PR lifecycle state per thread for sidebar

Persist a PR's draft/open/merged/closed state on the agent thread and keep
it in sync as PR webhooks fire, so the Agents UI can show per-thread PR
status the way Cursor does. open_pull_request now records the initial
draft/open state and pr_title; thread summaries expose diffStats; and the
PR webhook refreshes pr_state on close/reopen/draft/ready transitions.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: consolidate PR state mapping into shared derive_pr_state helper

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 12:21:59 -07:00
Johannes du Plessis
f08177aed7
fix: harden origin parsing and home prompt submit failure (#1499)
- _origin_of: treat invalid ports as invalid origin instead of letting
  urlparse ValueError turn CSRF rejections into 500s
- AgentsHome: reset submitting/draft when stream.submit rejects before a
  thread id is minted, so the prompt isn't left disabled

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 11:54:25 -07:00