* feat: editable plan mode + fix review-plan banner overlap
Lets the thread owner edit the plan markdown by hand from the plan-review
page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id}
endpoint that re-publishes the plan and mirrors it into the sandbox
plan.md, so approve hands the edited plan to the agent as the source of
truth. Also fixes the collapsed git-panel's floating expand button
covering the "Review plan ->" banner by reserving space for it.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: abort plan approval when the published plan read fails
get_plan_content() swallowed store errors and returned None, so a
transient failure during approve would still mark the plan approved and
dispatch the generic fallback text — silently dropping an owner's edited
plan. Read the plan strictly (raise_on_error=True) so approval aborts
instead, matching the comment read.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
sfw only ships a launcher that fetches its real binary at first run and does a
daily update check against api.github.com/repos/SocketDev/sfw-free. Both fail in
the sandbox (restricted egress; the proxy injects the GitHub App installation
token, which lacks access to that repo), so `sfw yarn install` errors with
"could not fetch its binary". Pin sfw 2.0.6, warm + verify the binary cache at
build, and set SFW_SKIP_UPDATE_CHECK=1 so runs use the baked binary offline.
Add a "Start from a template" gallery to the Automations page so users can
scaffold common scheduled agent runs (PR review digest, issue triage,
dependency checks, flaky-test tracking, release notes, docs freshness,
security audit) instead of writing every automation from scratch. Selecting
a template opens the new-automation editor prefilled with its instructions
and a default schedule, which the user can tweak before saving.
Also adds an isDescribableCron helper so template schedules render as
human-readable labels rather than raw cron in the editor.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
When a run is cancelled or the sandbox dies mid-tool-call, LangGraph persists
the AIMessage tool_call but never the matching ToolMessage. The next run sends
the provider an orphaned tool_use (Anthropic 400: "tool_use ids were found
without tool_result blocks"), permanently wedging the thread on every retry.
Add RepairOrphanedToolCallsMiddleware, which inserts a synthetic error
ToolMessage immediately after any tool_call lacking a result so the agent can
retry instead of dying. Wired into the agent and reviewer graphs.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
- Add a "Copy markdown" button to the plan header that copies the whole plan.
Cross-browser: async Clipboard API in secure contexts, hidden-textarea +
execCommand fallback for older Safari/Firefox and non-HTTPS origins.
- Fix the conversation banner: it showed "A plan is ready for your review" for
every non-approved/cancelled status, including "planning" — so it claimed the
plan was ready the instant plan mode began (the agent shares the link early
to follow along), then the plan page correctly said it was still being
written. Now: planning → "writing a plan", revising → "revising", ready →
"ready for your review".
Requesting changes hands the reviewer comments to the agent, so it's
meaningless with none. Disable the button (with a hint tooltip) until at
least one comment exists; approve is unaffected.
* feat: inline PR comments in the reviews UI
Click the diff gutter "+" on a line to open an inline comment composer
(rendered like the finding card via a Pierre annotation); submitting
posts a real inline PR review comment as the signed-in user through a
new POST /reviews/{owner}/{repo}/{number}/comments. The "+" press-drag →
"Add to Chat" selection path is unchanged.
* feat: GitHub-parity comment box, PR comments dropdown, collapse nav
- Comment composer now mirrors GitHub's box: Write/Preview tabs (markdown
rendered via the existing Markdown component) and a markdown toolbar
(heading, bold, italic, quote, code, link, bulleted/numbered/task list).
- Surface other people's inline PR comments in a Devin-style dropdown in the
review header (search + link to the thread on GitHub). New
GET /reviews/{owner}/{repo}/{number}/comments lists them and flags the
reviewer's own (marker-bearing) comments so they're filtered out.
- Collapse the global nav by default on a review detail page, restoring the
prior preference on leave.
* feat: bigger comment-toolbar icons; open dropdown comments inline
- Enlarge the markdown toolbar glyphs (Phosphor) in the comment composer —
they were rendering at 10px.
- Clicking a comment in the PR comments dropdown now opens it inline in the
diff as a read-only finding-style card (InlineComment), scrolling its line
into view, instead of navigating to GitHub. Falls back to GitHub when the
comment's file/line isn't in the current diff.
* fix: drive "Add to Chat" from native text selection
The gutter "+" is now comment-only; wiring its click to the composer
conflicted with its old double-duty as the drag-to-select handle, which
broke selection → "Add to Chat". Switch to Devin's model: disable Pierre's
interactive line selection and instead map a native text highlight in the
diff to a line range (via the data-line / data-line-type attributes Pierre
stamps on each line, read from the diff's open shadow root) to show the
"Add to Chat" popup. ⌘L and the existing attachment/popup path are unchanged.
* feat: gutter "+" drag selects a range for multi-line comments
Re-enable Pierre's gutter line selection so dragging the "+" down the
gutter comments across a range (click still comments on a single line);
onLineSelectionEnd routes the range to the composer. Native code-text
selection still drives "Add to Chat" — Pierre only line-selects from the
gutter, and onLineSelectionEnd bails when a native text selection is
present, so a code highlight never opens the composer.
* fix: keep the range highlighted while its comment composer is open
Previously opening the composer cleared the selection, so the lines being
commented on lost their highlight. Drive the controlled selection from the
open comment draft's range so the rows stay highlighted until the composer
is closed.
* fix: address PR review — paginate comments, fall back for outdated ones
- list_review_comments now pages through all PR review comments (bounded by
_MAX_REVIEW_COMMENT_PAGES) instead of returning only the first 100, so older
comments still show in the dropdown.
- Surface GitHub's outdated flag (position == null) as is_outdated; opening such
a comment (or one whose line isn't in the diff) now opens it on GitHub instead
of silently rendering nothing, plus a timeout fallback if the annotation never
mounts (e.g. collapsed context).
* feat: add plan mode for read-only research and planning
Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce plan-mode read-only at tool layer and disable subagents
Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden plan-mode shell guard against wrapped mutations
Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow
- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline
Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.
* test(open-swe): add Playwright E2E for the Slack → PR → web handoff
Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.
- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).
Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.
* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright
The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.
Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.
* test(open-swe): record Playwright trace + video on every E2E run
Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.
* feat(plan-mode): collaborative plan review with BlockNote + Yjs
When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.
- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
snapshots; plan content/status store; plan REST API (get/approve/reject,
owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
mounted under the agents shell, with a "Review plan" banner in the thread view
and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
flow, including cross-user comment sync and owner-only approval.
* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)
- plan-collab WS: authorize per-thread before joining a room (same read gate as
the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.
Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).
* style: ruff format plan_collab.py
* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS
- Slack "Approve & Implement" now verifies the clicking user is the plan
requester (owner, via the stored triggering_user_id) before implementing —
matching the dashboard API's owner-only approval. Non-owners are pointed to
Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
the REST require_same_origin CSRF defense.
* fix(plan-mode): enter plan mode only via the model + local mock dev harness
Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.
- enter_plan_mode returns a terminating ToolMessage, fixing the missing
ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
Alice/Bob mock users, and a GitHub login picker.
* docs(plan-mode): drop stale references to removed profile/team defaults
The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.
* feat(plan-mode): let any reviewer edit the plan, not just comment
Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.
* test(plan-mode): assert plan-mode entry via the tool's success message
plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.
* refactor(plan-mode): replace Yjs/BlockNote collab with plain HTTP comments
Drop the realtime collaborative editor (it can't work behind Vercel's
rewrite — WebSocket upgrades aren't proxied to the external LangGraph
backend) in favor of a simple whole-document comments API over plain HTTP.
Backend:
- Remove the Yjs WebSocket server (plan_collab.py), its lifespan, and the
collab router; drop pycrdt / pycrdt-websocket deps.
- plan_store: replace the Yjs snapshot with comment CRUD (one store item per
comment under ["plan","comments",thread_id]).
- plan_api: add GET/POST/DELETE comment endpoints; approve/reject now read
comments server-side and format them for the follow-up run (no longer
client-harvested). Comment delete is author-or-owner; approve stays owner-only.
Frontend:
- PlanReview renders the plan markdown read-only and shows a comments panel
(list + add, polled every 4s for cross-user visibility).
- Drop @blocknote/*, y-websocket, yjs; lib/plan exposes get/add/deletePlanComment.
Tests: unit tests for the comments API + route registration; e2e drives the
HTTP comment UI (owner + collaborator, cross-user visibility, owner-only approve,
PR echoes the harvested feedback).
* fix(open-swe): clear stale plan comments on republish; fail loud on store errors
Address reviewer feedback:
- Clear comments when a revised plan is published (save_plan_content) so
feedback on the prior revision doesn't resurface and get re-fed to the agent.
- list_plan_comments gains raise_on_error; approve/reject read comments before
mutating state and propagate store failures (500) instead of silently
dispatching the follow-up run with no feedback.
---------
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* perf: cut review-chat time-to-first-token
The sandbox-less PR review chat paid several blocking network round-trips
before the first token on every message. Cache GitHub App installation
tokens in-process (per scope, until ~10m before expiry, above the proxy's
5m refresh window) so the chat graph factory and proxy stop re-minting one
each turn. Also drop the duplicate thread-metadata read in the commands
proxy and replace the heavy per-message get_review staleness check with a
single lightweight PR head-SHA lookup.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: keep review chat alive when reseed fails
Address review: a moved PR head now triggers _build_pr_context (and thus
get_review). For an existing chat, fall back to the last seeded context on
HTTPException instead of failing the command; fresh chats still surface the
error since they have no prior context.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: relocate review chunks and embed review in git panel
Move the AI-sorted review chunk navigator out of the global left sidebar
and into the review page's own main panel (sidebar keeps the thread list),
and surface the review main body inside the agent thread's Git > Review
sub-tab with an expand link to the full review page.
The review main body + side panel + helpers are extracted into a shared
ReviewMainBody component with "full" and "embedded" variants.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: move ReviewTab into its own file
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: repo-scoped dynamic sandbox snapshots
Let admins build a per-repo sandbox image from a custom Dockerfile so runs
targeting that repo boot from a snapshot with its deps pre-baked. Snapshot
selection is purely additive: repos without a `ready` repo-scoped snapshot
always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID.
Backend adds a repo_snapshots store module (Dockerfile + build status keyed by
owner/name), threads the resolved repo through the LangSmith sandbox creation
path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a
throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The
UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor +
build status/logs).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden repo snapshot builds
Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins
cannot accidentally build a repo snapshot from a bare Python image that lacks
Open SWE's sandbox tools. Allow stale building records to be retried by tracking
build_started_at and treating old or missing timestamps as stale.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: document repo snapshot base image config
Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert
missing base-image configuration into a handled dashboard API error so admins see
a clear configuration message instead of an unhandled template-generation error.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
The PR reviews page builds a Pierre FileContents cache key for every
file in an unconditional useMemo. Binary/oversized/added/removed blobs
arrive with null originalContent/modifiedContent (pr_diff.py flags them
unrenderable), so fileContentsCacheKey dereferenced null via
contents.length and crashed the whole route — escaping the markdown
error boundary, which was unrelated. Coerce null/undefined contents to
"" in the cache-key helper so it can never throw.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add plan mode for read-only research and planning
Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce plan-mode read-only at tool layer and disable subagents
Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden plan-mode shell guard against wrapped mutations
Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow
- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline
Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.
* test(open-swe): add Playwright E2E for the Slack → PR → web handoff
Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.
- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).
Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.
* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright
The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.
Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.
* test(open-swe): record Playwright trace + video on every E2E run
Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.
* feat(plan-mode): collaborative plan review with BlockNote + Yjs
When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.
- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
snapshots; plan content/status store; plan REST API (get/approve/reject,
owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
mounted under the agents shell, with a "Review plan" banner in the thread view
and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
flow, including cross-user comment sync and owner-only approval.
* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)
- plan-collab WS: authorize per-thread before joining a room (same read gate as
the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.
Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).
* style: ruff format plan_collab.py
* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS
- Slack "Approve & Implement" now verifies the clicking user is the plan
requester (owner, via the stored triggering_user_id) before implementing —
matching the dashboard API's owner-only approval. Non-owners are pointed to
Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
the REST require_same_origin CSRF defense.
* fix(plan-mode): enter plan mode only via the model + local mock dev harness
Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.
- enter_plan_mode returns a terminating ToolMessage, fixing the missing
ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
Alice/Bob mock users, and a GitHub login picker.
* docs(plan-mode): drop stale references to removed profile/team defaults
The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.
* feat(plan-mode): let any reviewer edit the plan, not just comment
Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.
* test(plan-mode): assert plan-mode entry via the tool's success message
plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.
---------
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat(dashboard): let any org member post to a thread, with attribution
Posting to an Agents chat thread from the web UI was restricted to the
thread owner. Open it to any authenticated org member (login is already
org-gated by OAuth) on both write paths — the queued follow-up
(send_dashboard_message) and the idle-thread run.start
(_enrich_run_start_command). Non-owner messages are prefixed with the
poster's verified GitHub login (@login:) so the agent and owner can tell
who sent them. Thread management (cancel/delete/resolve) stays owner-only,
and the UI now shows the composer to non-owners.
* fix(dashboard): keep non-run.start commands owner-only
Non-owner posting is allowed only via the attributed run.start path. Other
write commands (e.g. input.respond) carry unattributed user input, so the
commands proxy keeps them owner-only instead of readable-by-any-org-member.
* docs(e2e): drop per-test details from the E2E README
* feat: add schedule_thread_wakeup tool for self-polling
Add a new agent tool that schedules a one-shot re-trigger of the
current thread after a configurable delay (1–1440 minutes). Uses a
LangGraph cron with end_time to fire exactly once, then retire.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: prevent early thread wakeups
Round scheduled wakeup times up to the next whole minute so cron minute precision cannot fire before the requested delay.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: render GitHub-hosted images in PR descriptions on reviews page
PR description images hosted on GitHub (user-attachment uploads and
*.githubusercontent.com) render broken on the reviews page because
private-repo attachments require GitHub auth the browser session lacks.
Add an authenticated backend image proxy (host-allowlisted to guard
against SSRF) and route those image URLs through it from the reviews UI.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden PR image proxy (IDOR, SVG XSS, unbounded buffering)
Address review findings on the review-page image proxy:
- IDOR: the proxy fetched any *.githubusercontent.com URL with the App
installation token, gated only by route-param repo access, so a user
authorized for one repo could read images from another private repo the
App can see. Bind the URL to the authorized PR — only proxy URLs that
appear in that PR's body.
- SVG XSS: served any image/* inline from the API origin, including
image/svg+xml which can run script. Restrict to safe raster types and
add X-Content-Type-Options: nosniff + a locked-down CSP.
- DoS: enforced the size cap only after buffering the full response.
Stream and abort once the cap is exceeded.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test(open-swe): add Playwright E2E for the Slack → PR → web handoff
Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.
- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).
Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.
* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright
The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.
Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.
* test(open-swe): record Playwright trace + video on every E2E run
Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.
Streamdown bundles Mermaid and renders ```mermaid blocks itself; a
diagram it can't parse throws during render and, with no boundary,
white-screens the whole review page. Wrap the markdown render so any
crash falls back to the raw text instead.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Replace the floating finding cards with inline annotations in the PR diff.
An anchored finding shows a collapsed header at its line and expands the full
details in place; non-anchored findings expand inline in the side-panel row.
Expand state is lifted into a context so it survives Pierre's annotation
windowing and the side panel can drive it.
Removes the anchored/fallback floating cards and their rAF/ResizeObserver
scroll-tracking. Also stops the line-selection highlight from bleeding behind
the inline annotation row.
* fix(prompt): allow agent to pause and ask before adding a dependency
Brace's review on #1577 noted that the agent can in fact pause to ask
mid-task: post a Slack message (or PR-description note) and end the
turn without a tool call, then resume when the user replies. Reword
the DEPENDENCY_SECTION guidance so it tells the agent to use that
mechanism instead of claiming it cannot pause.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Update agent/prompt.py
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* fix(prompt): restore dependency pause guidance assertion
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* fix: reviewer reviews full diff; fix review UI scroll + dark-mode composer
- Drop the 200K-char PR diff truncation. The head/tail slice silently dropped
whole files out of the middle; the reviewer's large context window reviews
the complete diff. The grouping pass uses the full diff too. Other #1567 caps
(fetch_url, Slack, pagination, message queue) are kept.
- Sidebar file/group click now lands flush at the top under virtualization:
re-assert alignment after the Virtualizer's mid-scroll height reconciliation
settles, instead of trusting scrollIntoView's estimate-based target.
- Review chat composer textarea no longer shows a lighter square in dark mode
(override the Textarea's dark:bg-input default with dark:bg-transparent).
* fix: anchored finding card sits flush in the diff gutter, no side-panel overlap
The card was positioned next to the anchor and clamped to window.innerWidth, so
when the anchor sat near the diff/side-panel boundary the 412px card spilled
across into the side panel. Pin it flush to the diff column's right edge and
clamp its width to the column so it never overlaps the side panel and shrinks to
fit a narrow column.
* fix: anchored finding card sits over the side panel, not the diff
Flip the horizontal anchor: place the card flush against the diff column's right
edge and extend it rightward over the side panel, instead of leftward over the
diff. Width still caps at the preferred size and shrinks to fit a narrow panel.
* fix: anchor finding card to diff content edge + small-screen fallback
- Anchor to the centered diff content's right edge instead of the full-width
scroller, removing the centering-whitespace gap so the card sits closer to the
diff.
- Below xl the side panel is hidden and the diff fills the viewport, leaving no
gutter; positioning against the content edge would collapse the card to ~0px.
Fall back to overlaying the diff flush-right when there's no usable gutter
(addresses Open SWE review finding f_f6ea77527c).
* fix: anchor finding card next to the annotation, close to the hunk
Anchor the card's left to the finding's annotation (anchorRect.right) instead of
the diff content's right edge, so it sits right by the hunk and overlaps the diff
edge rather than parking out in the gutter. Extends right over the side panel,
clamped to stay on-screen (also keeps it readable below xl with no side panel).
* fix: finding card width shrinks to fit a small side panel
Size the card to the room to the right of the annotation (capped at the
preferred width) so it narrows as the side panel shrinks instead of overflowing.
Below a readable minimum, hold that width and shift left over the diff.
* feat: reviews page split view, add-to-chat, virtualization + perf
- Virtualize the diff (Pierre Virtualizer + worker pool), mirroring the agent
chat panel, so large PRs window rows instead of materializing every line.
- Split chat and diff into independent scroll containers and make chat
auto-scroll fully contained, so typing/streaming no longer moves the diff.
- Memoize FileDiffCard with stable callbacks so focusing a finding re-renders
only the affected card.
- Add a persisted unified/split diff toggle.
- Add highlight-to-chat: select lines (drag / shift-click) + gutter "+" to drop
a file:line snippet into the chat composer.
- Rebuild sidebar group rows: whole card scrolls to the group (incl. the
expanded explanation), Read explanation stays a separate toggle, memoized.
* feat: add-to-chat uses attachment pills + selection popup / ⌘L
Replace the raw-snippet injection with a Cursor-style flow:
- Selecting lines shows a floating "Add to Chat ⌘L" popup at the pointer; ⌘L
adds the current selection without it. Removes the auto-adding gutter "+".
- "Add to chat" now creates a removable attachment pill in the composer (and a
pill in the sent message bubble) instead of pasting raw text. The code is
still serialized into the message content so the model receives it as context.
* feat: restore gutter + drag-handle for line selection
Re-enable Pierre's gutter '+' as a click-and-drag line selector (with the
highlight growing as you drag) — the affordance that was lost when the
auto-adding gutter button was removed. It no longer auto-adds: the commit flows
through onLineSelected to the 'Add to Chat' popup / ⌘L.
* fix: live selection highlight while dragging + reposition Add to Chat popup
- Feed onLineSelectionChange into the controlled selection so rows highlight
live as you drag, not just on release (Pierre only paints the controlled
selection when the prop updates). Popup now fires on onLineSelectionEnd.
- Anchor the popup's bottom-left to the drag handle (drop horizontal centering)
so it no longer overlaps the '+' button.
* fix: anchor Add to Chat popup to gutter handle + click-away to unselect
- Position the popup from the gutter '+' handle's rect (in the diff shadow DOM,
placed on the selection's bottom line) instead of the pointer-release point,
which landed inconsistently. Falls back to the pointer if not found.
- Clear the line selection (and popup) on any outside pointer-down.
* fix: sidebar-collapse header overlap + PR review comments
- Lift useSidebarLayout to AgentsShell (single source), share collapsed via
context, and pad the reviews header left when the sidebar is collapsed so the
fixed collapse toggle no longer overlaps the header content.
- add-to-chat: collect each diff side separately so a selection that spans a
deletion->addition no longer pastes wrong-file lines (PR comment).
- chat: clear attachments after sending via a suggested prompt, so an attached
snippet isn't silently resent on the next message (PR comment).
* fix: anchored finding card positions to the right of the diff again
The virtualization refactor moved the card inside the main-width Virtualizer
scroller, so it clamped over the diff. Render it in the outer container as a
viewport-fixed card clamped to window width (right gutter / over the side panel,
like prod) and track the finding as the diff scrolls (rAF-throttled), hiding
when the finding scrolls out of view.
Instead of silently swallowing findings below the severity threshold,
the review summary now mentions them with a count and links to the
web app where they can be viewed. For example, if 2 low-severity
findings are filtered out, the PR comment says "No issues found" and
"2 additional findings can be viewed in the web app." with the
existing [Open in Web] link.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: reviewer enforces AGENTS.md/CLAUDE.md repo rules as mandatory pass
The reviewer already fetched AGENTS.md but treated violations as optional
candidate findings. Now the reviewer runs a dedicated compliance pass that
checks every changed hunk against each rule in AGENTS.md (or CLAUDE.md as
fallback), treating violations as mandatory findings rather than style nits.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: oversized AGENTS.md returns None instead of falling back to CLAUDE.md
Only a 404 (file absent) triggers fallback to CLAUDE.md. Oversize,
HTTP errors, and unexpected status codes now return None immediately
so the reviewer does not enforce stale rules from a secondary file.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: validate LLM API keys on startup
* fix: correct relative import for options module
* refactor: move imports to top of file
* style: fix linting and formatting issues
* refactor: scope LLM validation to local dev and rename function
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: open-swe[bot] <johannes@langchain.dev>
The thread detail endpoint already returned metadata for non-owners, but
the transcript hydration endpoints (state, stream/events, history, pr-diff)
all asserted ownership and 404-ed. This caused the UI to redirect non-owners
back to /agents when they clicked an "Open in Web" link shared in Slack.
Dashboard login is already gated by ALLOWED_GITHUB_ORGS, so any logged-in
user is a trusted org member. This commit:
- Adds _thread_is_readable / _assert_thread_readable helpers that grant
read access to any surfaced-source thread for authenticated users
- Relaxes read endpoints (state, stream/events, history, pr-diff, SSE
stream) to use readable checks instead of ownership checks
- Keeps write endpoints (send message, cancel, delete, resolve, run
commands) owner-only
- Adds an isOwner field to the thread summary so the frontend can render
a read-only mode (hides the prompt bar, resolve/delete buttons)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add size caps for PR diff, fetch_url, Slack threads, pagination, message queue
Per-source byte/token caps with explicit truncation markers to prevent
unbounded payloads from blowing up LLM context/memory.
- reviewer_diff.py: cap PR diff at 200K chars with head+tail truncation
- fetch_url.py: cap markdownify output at 100K chars
- slack.py: cap thread message fetch at 500 messages
- github_comments.py: cap _fetch_paginated at 50 pages
- thread_ops.py: cap queued messages at 100 (drop oldest)
Closes OPE-51
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: compute diff line set from full diff, keep most recent Slack messages
Address PR review comments:
1. Truncated diffs rejected valid findings: fetch_pr_diff now returns the
full diff; truncate_diff is called separately in reviewer.py so the
line set used for add_finding/publish_review validation is computed
from the complete diff, not the truncated prompt text.
2. Slack cap dropped recent thread context: fetch_slack_thread_messages
now keeps the most recent SLACK_THREAD_MAX_MESSAGES messages (was
keeping the oldest). The tool surfaces a truncation marker in the
formatted output so the LLM knows the thread was truncated.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: shared GitHub HTTP helper with retries, rate-limit handling, and sane timeouts
Introduces agent/utils/github_http.py — a single place for GitHub API HTTP
calls with 30s/10s-connect timeouts (vs httpx's 5s default), exponential
backoff with jitter, Retry-After header support, and 429/secondary-rate-limit
detection. Migrates the reviewer publish path (reviewer_publish.py,
reviewer_diff.py, github_checks.py, github_ci.py) from one-shot
httpx.AsyncClient() calls to the shared helper.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: don't retry transport errors on non-idempotent GitHub writes
POST/DELETE/PATCH can create side effects server-side even when the client
gets a timeout or connection reset. Only retry transport errors for
idempotent methods (GET, HEAD, PUT, DELETE). 429/5xx status codes are still
retried for all methods since the server explicitly did not process the
request.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: don't retry 502/504 on non-idempotent GitHub writes
502 (bad gateway) and 504 (gateway timeout) are ambiguous — the upstream
may have processed the write before the gateway returned an error. Only
retry these for idempotent methods. 429 and 503 are still retried for all
methods since the server explicitly did not process the request.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: activate PR babysitting UI toggles for autofix and trigger mode
Remove the "coming soon" gating on the Autofix Mode, Autofix Severity
Threshold, and Trigger Mode controls in the review settings page so
admins can enable CI auto-fix and review-comment resolution on PRs
that Open SWE opens. The backend (ci_autofix.py, webapp.py webhook
routing) was already fully wired — only the UI was disabled.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: simplify autofix to on/off toggle, remove severity threshold
Replace the four-level AutofixMode (off/low/medium/high) and the
autofix_severity_threshold setting with a single boolean
autofix_enabled toggle. The severity threshold was leftover from the
reviewer finding-severity model and does not apply to CI autofix;
the agent should fix any failing CI and resolve any comments on PRs
it opens.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: move autofix toggle to per-user profile, remove team-level setting
The autofix toggle is now per-user (auto_fix_ci in the user profile)
instead of team-level (admin-only). This uses the existing auto_fix_ci
field that was already in ProfileUpdate but never wired up.
Changes:
- ci_autofix.py: check per-user auto_fix_ci profile flag after
resolving the agent thread's github_login, instead of checking
team-level autofix_enabled before knowing the PR
- webapp.py: removed early is_autofix_enabled() webhook gates; the
per-user check now happens in ci_autofix.py once the thread is found
- team_settings.py: removed autofix_enabled field, is_autofix_enabled()
- cloud-agents.tsx: enabled the auto_fix_ci toggle (was comingSoon)
- review.tsx: removed the admin-level autofix switch
- Updated tests and AGENTS.md
The agent graph (not the reviewer) is what gets dispatched - this was
already correct in ci_autofix.py line 223: client.runs.create(
thread_id, "agent", ...).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: batch PR babysitting events
Remove the leftover trigger-mode gate from PR babysitting and batch new CI/review events while an agent run is already active so the running agent can handle the latest PR state before finishing. Also moves review-feedback permission checks behind the per-user opt-out and applies the auto-fix profile gate to merge-conflict babysitting.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: consume batched babysitting events
Teach the agent queue middleware to turn pending PR babysitting metadata into an injected instruction for the active run, so batched CI/review events are not dropped while still avoiding duplicate run creation.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review findings in PR babysitting batching
- Route batched events through the LangGraph store (read in-process by the
message-queue middleware) instead of a per-model-call threads.get on every
agent thread.
- Only record an attempt / mark the head SHA handled on a real dispatch, not
on a batch, so an event isn't permanently dropped if the in-flight run ends
before consuming it.
- Carry the reviewer's comment through batched review feedback instead of
replacing it with a generic re-check nudge.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add user-scoped Currents.dev API key for e2e test investigation
Allow each user to configure their own Currents.dev API key on the
Profile Settings page. The key is encrypted at rest in a per-user
LangGraph Store namespace and feeds server-side read-only tools that
query the Currents REST API (runs, instances, projects, test results)
so agent runs can inspect e2e test failures including screenshots and
DOM snapshots.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: add pagination cursors to currents_list_project_runs
Address review feedback: forward starting_after/ending_before cursor
parameters to /projects/{projectId}/runs so the agent can paginate
beyond the first 50 results.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Move merged_prs to the primary sort key in the agent usage leaderboard,
ahead of agent_loc, prs_opened, and agent_runs. Update the UI description
to reflect the new ranking order.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add browser notifications for agent run completion
Request notification permission when a user starts their first agent
run from the home page. A toggle in Profile Settings lets users
enable/disable desktop notifications. When a run transitions from
running to a terminal state (finished/error/interrupted), a browser
notification is fired — suppressed for the thread the user is
currently viewing.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: notify for active thread when page is in background tab
Only suppress the notification for the active thread when the page is
actually visible. If the user switched to another browser tab, the
notification fires even for the thread they have open.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
At mobile widths the resizable git panel collided with the chat column's
360px min-width, overflowing the viewport. Treat mobile (<768px) like the
sidebar: the git panel becomes a full-screen overlay the user navigates to
and back from, while desktop keeps the inline resizable panel.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: handle images sent to non-vision models in Slack, Linear, and web UI
Add vision capability checks across all image input paths. When a user
sends images to a text-only model (e.g. GLM 5.2, DeepSeek V4 Pro), the
images are now skipped and a warning is injected into the prompt instead
of sending unsupported content to the model.
- Slack: resolve model at webhook time, skip image fetch + add warning
- Linear: same pattern as Slack
- Queued message middleware: read resolved model from thread metadata,
strip images from queued payloads for text-only models
- Web UI: disable submit + show inline warning when images are attached
to a non-vision model selection
- Shared: resolve_agent_model_id helper + vision_not_supported_warning
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: mock resolve_agent_model_id in Slack mention test
The test_process_slack_mention_queues_active_thread_message test was
missing a mock for the new resolve_agent_model_id call added to the
Slack webhook handler, causing a TypeError when image URLs triggered
the model resolution path.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: include vision warning in queued payload for text-only models
Update the prompt variable (not just content_blocks) before clearing
image_urls so the queued payload also carries the warning text when a
Slack/Linear follow-up arrives while the thread is busy.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>