* feat: publish plans from sandbox files
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: avoid fixed plan filenames
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: virtualize local sandbox file paths
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: preserve plan_file_path across set_plan_status
set_plan_status was rewriting the content record with only markdown
and status, dropping plan_file_path. After a reject, the owner's
dashboard edit would mirror to a different file than the agent's
original, and the next save_plan could republish the stale file.
Preserve plan_file_path when updating status.
---------
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: notify Slack on plan approval
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: post Slack approval notice after successful dispatch
Move the _maybe_post_plan_approved_to_slack call until after
_dispatch_followup succeeds so the Slack thread is not told
implementation is beginning before the LangGraph run is created.
Addresses PR review comment.
---------
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* fix: update Slack trace reply on web handoff
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: trigger web handoff on dashboard starts
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: format web handoff as contextual fragment
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: preserve trace_message_ts when overwriting Slack run mapping
When store_slack_run_mapping is called without trace_message_ts (e.g. on
follow-up Slack mentions), it was unconditionally overwriting the
thread-level mapping and clobbering the timestamp captured from the
initial trace reply. After that, _notify_slack_web_handoff could not find
the original message, so a subsequent move to Web silently skipped the
Slack trace update.
Now, when trace_message_ts is not passed, the existing thread mapping is
read first and its trace_message_ts is preserved.
* style: ruff format
---------
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* wip(rebuild): core reliability spine
- remove PR-babysitting (ci_autofix + ci_monitor graph + webhook wiring)
- dispatch core: agent/dispatch.py with multitask_strategy=interrupt +
durability=sync + completion webhook; reroute all webhook + plan triggers;
drop the racy in-process lock + is_thread_active busy-check
- completion webhook: agent/completion.py + /webhooks/run-complete loopback
route for failure/timeout replies (idempotent)
Co-authored-by: open-swe[bot]
* feat(rebuild): async tools, reconcile, shared http timeouts, assembly tuning
Parallel batch on top of the reliability spine:
- async-ify all 24 tools (drop asyncio.run; requests->httpx); re-implement the
http_request/fetch_url SSRF + DNS-rebinding defense httpx-natively and harden
the IP check to 'not is_global' (+ IPv4-mapped unwrap)
- reconcile.py: stale pending-run sweep (threads.search -> per-thread runs.list
-> cancel_many), wired into the scheduler graph via task='reconcile'
- shared DEFAULT_HTTP_TIMEOUT (agent/utils/http.py) on every bare
httpx.AsyncClient() across utils/dashboard/webapp/middleware
- run budget: MODEL_CALL_RECURSION_LIMIT 5000->250
- fix stale OpenAI->Anthropic fallback id (claude-opus-4-5 -> 4-8)
- drop redundant custom repair middleware (deepagents auto-adds PatchToolCalls)
- confirm tool-result eviction + summarization auto-wired via backend
- slim system prompt ~8% (full harness-profile rewrite deferred)
Co-authored-by: open-swe[bot]
* feat(rebuild): harness-profile prompt + split webhooks out of webapp
- prompt.py: own the system prompt via a registered harness profile
(OPEN_SWE_SHARED_BASE, kept neutral so the read-only reviewer/analyzer that
share it stay safe), registered across all 4 providers; per-thread values
stay in construct_system_prompt. Assembled main-agent prompt ~6.8k -> ~3.1k
tokens (~55% smaller); de-duped PR/commit/suite/force-push guidance; dropped
ALL-CAPS markers.
- webapp.py 3325 -> 1890 LOC: moved 14 per-source handlers into
agent/webhooks/{linear,slack,github}.py; webapp re-exports them for the
routes + tests; moved handlers reach shared helpers via the webapp namespace
to preserve the test suite's monkeypatch targets.
Full suite: 1168 passing, lint clean.
Co-authored-by: open-swe[bot]
* Restore MODEL_CALL_RECURSION_LIMIT to 5000 for long-running tasks
Reverts the 250 cap from the run-budget change — long-running tasks legitimately
need many model calls. The notify_step_limit_reached safety net still fires if a
run does hit the cap, so runs end with a signal either way.
Co-authored-by: open-swe[bot]
* fix: address PR review (auth, SSRF, interrupted status, redirect headers)
- completion.py: drop `interrupted` from failure statuses — with
multitask_strategy=interrupt a follow-up ends the prior run as interrupted,
which is healthy, not a failure to report. [open-swe]
- /webhooks/run-complete: shared-secret auth — dispatch appends ?token= when
RUN_COMPLETE_WEBHOOK_SECRET is set; route verifies via hmac.compare_digest.
[corridor-security]
- SSRF: extract the URL validator to agent/utils/url_safety.py and apply it
before server-side image fetches in multimodal.fetch_image_block.
[corridor-security]
- http_request: preserve caller headers/extensions across redirect hops instead
of dropping them on the first hop. [open-swe]
Co-authored-by: open-swe[bot]
* chore: remove REBUILD_PLAN.md (planning doc, not needed in the repo)
Co-authored-by: open-swe[bot]
* fix: fail closed on run-complete webhook auth when secret unset
Corridor follow-up: verify_run_complete_token returns False (not True) when
RUN_COMPLETE_WEBHOOK_SECRET is unset, so the public route is never
unauthenticated. Logs a startup warning when the secret is absent, and dispatch
skips registering the webhook when there's no secret (no rejected callbacks).
Co-authored-by: open-swe[bot]
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add PR trace resolution
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: inject reviewer trace context as JSON
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review on PR trace resolution
Use the documented LangSmith metadata filter syntax
(and(eq(metadata_key,...), eq(metadata_value,...))) instead of
has(metadata, '{...}'), which does not match runs — _list_thread_runs
was silently returning nothing. Bound full-text searches to a 90-day
window so they don't hit LangSmith's large-window rate limit.
Also folds in the best-effort branch->head-sha resolver (dropping the
weighted scoring/threshold + repo/file evidence + GitHub hydration),
sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint.
The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session
were removed; resolution now runs deterministically from the trusted run
config with no model-controlled pr_url or thread_id.
* fix: scope branch trace search to the repo
Branch names like fix-tests aren't unique across repos (or older PRs) in
a shared tracing project, so an unscoped branch hit could resolve to an
unrelated thread and write its runs into the reviewer sandbox. Require
the repo slug to co-occur with the branch in matched runs; the full head
SHA stays unscoped since it is globally unique. Addresses open-swe review
on PR #1612.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: editable plan mode + fix review-plan banner overlap
Lets the thread owner edit the plan markdown by hand from the plan-review
page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id}
endpoint that re-publishes the plan and mirrors it into the sandbox
plan.md, so approve hands the edited plan to the agent as the source of
truth. Also fixes the collapsed git-panel's floating expand button
covering the "Review plan ->" banner by reserving space for it.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: abort plan approval when the published plan read fails
get_plan_content() swallowed store errors and returned None, so a
transient failure during approve would still mark the plan approved and
dispatch the generic fallback text — silently dropping an owner's edited
plan. Read the plan strictly (raise_on_error=True) so approval aborts
instead, matching the comment read.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: inline PR comments in the reviews UI
Click the diff gutter "+" on a line to open an inline comment composer
(rendered like the finding card via a Pierre annotation); submitting
posts a real inline PR review comment as the signed-in user through a
new POST /reviews/{owner}/{repo}/{number}/comments. The "+" press-drag →
"Add to Chat" selection path is unchanged.
* feat: GitHub-parity comment box, PR comments dropdown, collapse nav
- Comment composer now mirrors GitHub's box: Write/Preview tabs (markdown
rendered via the existing Markdown component) and a markdown toolbar
(heading, bold, italic, quote, code, link, bulleted/numbered/task list).
- Surface other people's inline PR comments in a Devin-style dropdown in the
review header (search + link to the thread on GitHub). New
GET /reviews/{owner}/{repo}/{number}/comments lists them and flags the
reviewer's own (marker-bearing) comments so they're filtered out.
- Collapse the global nav by default on a review detail page, restoring the
prior preference on leave.
* feat: bigger comment-toolbar icons; open dropdown comments inline
- Enlarge the markdown toolbar glyphs (Phosphor) in the comment composer —
they were rendering at 10px.
- Clicking a comment in the PR comments dropdown now opens it inline in the
diff as a read-only finding-style card (InlineComment), scrolling its line
into view, instead of navigating to GitHub. Falls back to GitHub when the
comment's file/line isn't in the current diff.
* fix: drive "Add to Chat" from native text selection
The gutter "+" is now comment-only; wiring its click to the composer
conflicted with its old double-duty as the drag-to-select handle, which
broke selection → "Add to Chat". Switch to Devin's model: disable Pierre's
interactive line selection and instead map a native text highlight in the
diff to a line range (via the data-line / data-line-type attributes Pierre
stamps on each line, read from the diff's open shadow root) to show the
"Add to Chat" popup. ⌘L and the existing attachment/popup path are unchanged.
* feat: gutter "+" drag selects a range for multi-line comments
Re-enable Pierre's gutter line selection so dragging the "+" down the
gutter comments across a range (click still comments on a single line);
onLineSelectionEnd routes the range to the composer. Native code-text
selection still drives "Add to Chat" — Pierre only line-selects from the
gutter, and onLineSelectionEnd bails when a native text selection is
present, so a code highlight never opens the composer.
* fix: keep the range highlighted while its comment composer is open
Previously opening the composer cleared the selection, so the lines being
commented on lost their highlight. Drive the controlled selection from the
open comment draft's range so the rows stay highlighted until the composer
is closed.
* fix: address PR review — paginate comments, fall back for outdated ones
- list_review_comments now pages through all PR review comments (bounded by
_MAX_REVIEW_COMMENT_PAGES) instead of returning only the first 100, so older
comments still show in the dropdown.
- Surface GitHub's outdated flag (position == null) as is_outdated; opening such
a comment (or one whose line isn't in the diff) now opens it on GitHub instead
of silently rendering nothing, plus a timeout fallback if the annotation never
mounts (e.g. collapsed context).
* feat: add plan mode for read-only research and planning
Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce plan-mode read-only at tool layer and disable subagents
Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden plan-mode shell guard against wrapped mutations
Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow
- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline
Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.
* test(open-swe): add Playwright E2E for the Slack → PR → web handoff
Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.
- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).
Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.
* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright
The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.
Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.
* test(open-swe): record Playwright trace + video on every E2E run
Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.
* feat(plan-mode): collaborative plan review with BlockNote + Yjs
When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.
- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
snapshots; plan content/status store; plan REST API (get/approve/reject,
owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
mounted under the agents shell, with a "Review plan" banner in the thread view
and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
flow, including cross-user comment sync and owner-only approval.
* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)
- plan-collab WS: authorize per-thread before joining a room (same read gate as
the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.
Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).
* style: ruff format plan_collab.py
* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS
- Slack "Approve & Implement" now verifies the clicking user is the plan
requester (owner, via the stored triggering_user_id) before implementing —
matching the dashboard API's owner-only approval. Non-owners are pointed to
Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
the REST require_same_origin CSRF defense.
* fix(plan-mode): enter plan mode only via the model + local mock dev harness
Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.
- enter_plan_mode returns a terminating ToolMessage, fixing the missing
ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
Alice/Bob mock users, and a GitHub login picker.
* docs(plan-mode): drop stale references to removed profile/team defaults
The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.
* feat(plan-mode): let any reviewer edit the plan, not just comment
Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.
* test(plan-mode): assert plan-mode entry via the tool's success message
plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.
* refactor(plan-mode): replace Yjs/BlockNote collab with plain HTTP comments
Drop the realtime collaborative editor (it can't work behind Vercel's
rewrite — WebSocket upgrades aren't proxied to the external LangGraph
backend) in favor of a simple whole-document comments API over plain HTTP.
Backend:
- Remove the Yjs WebSocket server (plan_collab.py), its lifespan, and the
collab router; drop pycrdt / pycrdt-websocket deps.
- plan_store: replace the Yjs snapshot with comment CRUD (one store item per
comment under ["plan","comments",thread_id]).
- plan_api: add GET/POST/DELETE comment endpoints; approve/reject now read
comments server-side and format them for the follow-up run (no longer
client-harvested). Comment delete is author-or-owner; approve stays owner-only.
Frontend:
- PlanReview renders the plan markdown read-only and shows a comments panel
(list + add, polled every 4s for cross-user visibility).
- Drop @blocknote/*, y-websocket, yjs; lib/plan exposes get/add/deletePlanComment.
Tests: unit tests for the comments API + route registration; e2e drives the
HTTP comment UI (owner + collaborator, cross-user visibility, owner-only approve,
PR echoes the harvested feedback).
* fix(open-swe): clear stale plan comments on republish; fail loud on store errors
Address reviewer feedback:
- Clear comments when a revised plan is published (save_plan_content) so
feedback on the prior revision doesn't resurface and get re-fed to the agent.
- list_plan_comments gains raise_on_error; approve/reject read comments before
mutating state and propagate store failures (500) instead of silently
dispatching the follow-up run with no feedback.
---------
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* perf: cut review-chat time-to-first-token
The sandbox-less PR review chat paid several blocking network round-trips
before the first token on every message. Cache GitHub App installation
tokens in-process (per scope, until ~10m before expiry, above the proxy's
5m refresh window) so the chat graph factory and proxy stop re-minting one
each turn. Also drop the duplicate thread-metadata read in the commands
proxy and replace the heavy per-message get_review staleness check with a
single lightweight PR head-SHA lookup.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: keep review chat alive when reseed fails
Address review: a moved PR head now triggers _build_pr_context (and thus
get_review). For an existing chat, fall back to the last seeded context on
HTTPException instead of failing the command; fresh chats still surface the
error since they have no prior context.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: repo-scoped dynamic sandbox snapshots
Let admins build a per-repo sandbox image from a custom Dockerfile so runs
targeting that repo boot from a snapshot with its deps pre-baked. Snapshot
selection is purely additive: repos without a `ready` repo-scoped snapshot
always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID.
Backend adds a repo_snapshots store module (Dockerfile + build status keyed by
owner/name), threads the resolved repo through the LangSmith sandbox creation
path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a
throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The
UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor +
build status/logs).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden repo snapshot builds
Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins
cannot accidentally build a repo snapshot from a bare Python image that lacks
Open SWE's sandbox tools. Allow stale building records to be retried by tracking
build_started_at and treating old or missing timestamps as stale.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: document repo snapshot base image config
Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert
missing base-image configuration into a handled dashboard API error so admins see
a clear configuration message instead of an unhandled template-generation error.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add plan mode for read-only research and planning
Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce plan-mode read-only at tool layer and disable subagents
Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden plan-mode shell guard against wrapped mutations
Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow
- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline
Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.
* test(open-swe): add Playwright E2E for the Slack → PR → web handoff
Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.
- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).
Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.
* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright
The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.
Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.
* test(open-swe): record Playwright trace + video on every E2E run
Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.
* feat(plan-mode): collaborative plan review with BlockNote + Yjs
When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.
- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
snapshots; plan content/status store; plan REST API (get/approve/reject,
owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
mounted under the agents shell, with a "Review plan" banner in the thread view
and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
flow, including cross-user comment sync and owner-only approval.
* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)
- plan-collab WS: authorize per-thread before joining a room (same read gate as
the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.
Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).
* style: ruff format plan_collab.py
* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS
- Slack "Approve & Implement" now verifies the clicking user is the plan
requester (owner, via the stored triggering_user_id) before implementing —
matching the dashboard API's owner-only approval. Non-owners are pointed to
Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
the REST require_same_origin CSRF defense.
* fix(plan-mode): enter plan mode only via the model + local mock dev harness
Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.
- enter_plan_mode returns a terminating ToolMessage, fixing the missing
ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
Alice/Bob mock users, and a GitHub login picker.
* docs(plan-mode): drop stale references to removed profile/team defaults
The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.
* feat(plan-mode): let any reviewer edit the plan, not just comment
Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.
* test(plan-mode): assert plan-mode entry via the tool's success message
plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.
---------
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat(dashboard): let any org member post to a thread, with attribution
Posting to an Agents chat thread from the web UI was restricted to the
thread owner. Open it to any authenticated org member (login is already
org-gated by OAuth) on both write paths — the queued follow-up
(send_dashboard_message) and the idle-thread run.start
(_enrich_run_start_command). Non-owner messages are prefixed with the
poster's verified GitHub login (@login:) so the agent and owner can tell
who sent them. Thread management (cancel/delete/resolve) stays owner-only,
and the UI now shows the composer to non-owners.
* fix(dashboard): keep non-run.start commands owner-only
Non-owner posting is allowed only via the attributed run.start path. Other
write commands (e.g. input.respond) carry unattributed user input, so the
commands proxy keeps them owner-only instead of readable-by-any-org-member.
* docs(e2e): drop per-test details from the E2E README
* fix: render GitHub-hosted images in PR descriptions on reviews page
PR description images hosted on GitHub (user-attachment uploads and
*.githubusercontent.com) render broken on the reviews page because
private-repo attachments require GitHub auth the browser session lacks.
Add an authenticated backend image proxy (host-allowlisted to guard
against SSRF) and route those image URLs through it from the reviews UI.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden PR image proxy (IDOR, SVG XSS, unbounded buffering)
Address review findings on the review-page image proxy:
- IDOR: the proxy fetched any *.githubusercontent.com URL with the App
installation token, gated only by route-param repo access, so a user
authorized for one repo could read images from another private repo the
App can see. Bind the URL to the authorized PR — only proxy URLs that
appear in that PR's body.
- SVG XSS: served any image/* inline from the API origin, including
image/svg+xml which can run script. Restrict to safe raster types and
add X-Content-Type-Options: nosniff + a locked-down CSP.
- DoS: enforced the size cap only after buffering the full response.
Stream and abort once the cap is exceeded.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
The thread detail endpoint already returned metadata for non-owners, but
the transcript hydration endpoints (state, stream/events, history, pr-diff)
all asserted ownership and 404-ed. This caused the UI to redirect non-owners
back to /agents when they clicked an "Open in Web" link shared in Slack.
Dashboard login is already gated by ALLOWED_GITHUB_ORGS, so any logged-in
user is a trusted org member. This commit:
- Adds _thread_is_readable / _assert_thread_readable helpers that grant
read access to any surfaced-source thread for authenticated users
- Relaxes read endpoints (state, stream/events, history, pr-diff, SSE
stream) to use readable checks instead of ownership checks
- Keeps write endpoints (send message, cancel, delete, resolve, run
commands) owner-only
- Adds an isOwner field to the thread summary so the frontend can render
a read-only mode (hides the prompt bar, resolve/delete buttons)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: activate PR babysitting UI toggles for autofix and trigger mode
Remove the "coming soon" gating on the Autofix Mode, Autofix Severity
Threshold, and Trigger Mode controls in the review settings page so
admins can enable CI auto-fix and review-comment resolution on PRs
that Open SWE opens. The backend (ci_autofix.py, webapp.py webhook
routing) was already fully wired — only the UI was disabled.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: simplify autofix to on/off toggle, remove severity threshold
Replace the four-level AutofixMode (off/low/medium/high) and the
autofix_severity_threshold setting with a single boolean
autofix_enabled toggle. The severity threshold was leftover from the
reviewer finding-severity model and does not apply to CI autofix;
the agent should fix any failing CI and resolve any comments on PRs
it opens.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: move autofix toggle to per-user profile, remove team-level setting
The autofix toggle is now per-user (auto_fix_ci in the user profile)
instead of team-level (admin-only). This uses the existing auto_fix_ci
field that was already in ProfileUpdate but never wired up.
Changes:
- ci_autofix.py: check per-user auto_fix_ci profile flag after
resolving the agent thread's github_login, instead of checking
team-level autofix_enabled before knowing the PR
- webapp.py: removed early is_autofix_enabled() webhook gates; the
per-user check now happens in ci_autofix.py once the thread is found
- team_settings.py: removed autofix_enabled field, is_autofix_enabled()
- cloud-agents.tsx: enabled the auto_fix_ci toggle (was comingSoon)
- review.tsx: removed the admin-level autofix switch
- Updated tests and AGENTS.md
The agent graph (not the reviewer) is what gets dispatched - this was
already correct in ci_autofix.py line 223: client.runs.create(
thread_id, "agent", ...).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: batch PR babysitting events
Remove the leftover trigger-mode gate from PR babysitting and batch new CI/review events while an agent run is already active so the running agent can handle the latest PR state before finishing. Also moves review-feedback permission checks behind the per-user opt-out and applies the auto-fix profile gate to merge-conflict babysitting.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: consume batched babysitting events
Teach the agent queue middleware to turn pending PR babysitting metadata into an injected instruction for the active run, so batched CI/review events are not dropped while still avoiding duplicate run creation.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review findings in PR babysitting batching
- Route batched events through the LangGraph store (read in-process by the
message-queue middleware) instead of a per-model-call threads.get on every
agent thread.
- Only record an attempt / mark the head SHA handled on a real dispatch, not
on a batch, so an event isn't permanently dropped if the in-flight run ends
before consuming it.
- Carry the reviewer's comment through batched review feedback instead of
replacing it with a generic re-check nudge.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add user-scoped Currents.dev API key for e2e test investigation
Allow each user to configure their own Currents.dev API key on the
Profile Settings page. The key is encrypted at rest in a per-user
LangGraph Store namespace and feeds server-side read-only tools that
query the Currents REST API (runs, instances, projects, test results)
so agent runs can inspect e2e test failures including screenshots and
DOM snapshots.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: add pagination cursors to currents_list_project_runs
Address review feedback: forward starting_after/ending_before cursor
parameters to /projects/{projectId}/runs so the agent can paginate
beyond the first 50 results.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Move merged_prs to the primary sort key in the agent usage leaderboard,
ahead of agent_loc, prs_opened, and agent_runs. Update the UI description
to reflect the new ranking order.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: handle images sent to non-vision models in Slack, Linear, and web UI
Add vision capability checks across all image input paths. When a user
sends images to a text-only model (e.g. GLM 5.2, DeepSeek V4 Pro), the
images are now skipped and a warning is injected into the prompt instead
of sending unsupported content to the model.
- Slack: resolve model at webhook time, skip image fetch + add warning
- Linear: same pattern as Slack
- Queued message middleware: read resolved model from thread metadata,
strip images from queued payloads for text-only models
- Web UI: disable submit + show inline warning when images are attached
to a non-vision model selection
- Shared: resolve_agent_model_id helper + vision_not_supported_warning
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: mock resolve_agent_model_id in Slack mention test
The test_process_slack_mention_queues_active_thread_message test was
missing a mock for the new resolve_agent_model_id call added to the
Slack webhook handler, causing a TypeError when image URLs triggered
the model resolution path.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: include vision warning in queued payload for text-only models
Update the prompt variable (not just content_blocks) before clearing
image_urls so the queued payload also carries the warning text when a
Slack/Linear follow-up arrives while the thread is busy.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add max effort level for GLM 5.2
GLM 5.2 supports a 'max' thinking effort level (recommended for coding
tasks per Z.ai/Fireworks docs). Add it to the model's effort list so it
surfaces in the profile editor and maps to reasoning_effort=max.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: restrict GLM 5.2 efforts to none, high, max
GLM 5.2 only supports non-thinking (none), high, and max effort levels.
Remove low and medium which the model does not support.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* Run reviewer eval in a GitHub Action; make dashboard a read-only progress view
The dashboard launched the eval as a subprocess inside the serving deployment
worker, so a container recycle killed long runs and discarded results that had
already completed server-side. Move the harness to a workflow_dispatch Action
(run on prod). run_eval now publishes status/progress/log-tail to the LangGraph
store record the dashboard reads, so /admin/evals stays a live view; a killed
Action surfaces as failed via the stale-heartbeat reconcile.
* reviewer_eval workflow: pass inputs via env, no shell interpolation
Addresses the reviewer finding: workflow_dispatch string inputs were
interpolated into the run: block (limit unquoted), allowing shell injection in
a job holding LANGSMITH/ANTHROPIC keys. Pass inputs through env and reference
quoted "$VARS"; validate limit is numeric and build its flag in bash.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add GLM 5.2 model option
Add GLM 5.2 to the supported model list so it's selectable in the
profile editor and as a team/per-thread model.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: remove GLM 5.1 model option
Remove GLM 5.1 from the supported model list now that GLM 5.2 is available.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: AI-sorted PR review view with diff grouping
Group a PR's changed files into a logical top-to-bottom walkthrough via a best-effort structured-output LLM pass, kicked off concurrently with the reviewer run (~0 added latency) and persisted on reviewer thread metadata. Results render in the review UI behind an AI sorted / file tree toggle that persists across PRs; the view falls back to the file tree when groups are absent or stale.
Adds a grouping-model team default (inherits the Reviewer subagent model when unset) and the admin RolePicker for it.
* feat(reviews): richer AI-sorted explanations + sidebar polish
Sidebar group rows get Devin-style spacing (dividers, padding), a per-group file list (click a file to jump to its diff), inline-code chips in titles, and an accent Read explanation link.
Group explanations are now rich markdown: the grouping prompt feeds per-hunk line ranges and asks for inline code, a short code block, and [path:line](#loc=...) references. The Markdown renderer turns those #loc= links into in-page buttons that scroll the diff to the hunk and highlight the range, reusing the existing selectedLines path.
* fix(reviews): drop stale diff groups from the AI-sorted view
When groups were generated for a previous head, a persisted "ai" view in localStorage still rendered the outdated walkthrough. groupedView now returns null on diff_groups_stale, so the file-tree fallback is used and the view toggle hides until fresh groups arrive.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Drive dashboard-triggered reviewer eval runs with per-run model, effort,
score mode, severity threshold, cap, limit, and concurrency overrides, plus
per-example start/finish/error logging in the eval target.
Rework the admin eval form from the label-left/control-right SettingsRow
(which crushed the description column when packing 3-4 wide inputs) into
stacked field groups with captioned inputs in a responsive grid.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: chat with your PR on the review page
Add a sandbox-less `chat` graph that answers questions about a single PR
from its diff, the published review findings, and read-only GitHub access.
- agent/chat.py: deepagents graph, no sandbox (default StateBackend, file
mutation + execute tools excluded). PR context is seeded as virtual files
under /pr/; a repo-scoped App token is resolved in-graph.
- tools: read_repo_file, search_repo_code, list_review_findings.
- dashboard/review_chat_api.py + routes: per-user chat thread, LangGraph
stream/commands/state/history proxy pinned to the chat assistant, seeds
diff/findings/overview on first run. Gated by repo access.
- UI: Chat tab wired to a chat-scoped StreamProvider (replaces Coming Soon).
* feat: admin setting for review-chat default model
Add a 'Open SWE Review Chat' default to team settings (default_chat_model /
default_chat_reasoning_effort). get_team_default_model("chat") inherits the
Agent default when unset; the chat graph resolves through it. Admin RolePicker
gains an 'Agent default' inherit option that clears the override.
* feat: multi-conversation review chat (tabs, new chat, history)
Replace the single per-PR chat thread with multiple per-user conversations:
- threads minted client-side; first message persists with a title derived
from the prompt.
- list + delete endpoints; chat panel gets a tab strip (history), new-chat
(+), close (x), refresh, an intro greeting, and suggested prompts.
- get_review_chat now returns availability only (ids are client-minted).
* ui fixes
* ui: review-chat history dropdown, full-width AI replies, resizable side panel
* fix(review-chat): enforce per-user thread ownership on proxy endpoints; reseed PR context on head change
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: CI auto-fix and PR babysitting for agent PRs
Watch CI failures and review feedback on PRs Open SWE opened, then dispatch
confidence-gated fix runs on the originating agent thread. Adds CI webhook
ingestion (check_run/check_suite/workflow_run/status), a per-PR @open-swe
autofix on|off toggle, auto-response to review comments, and a polling
ci_monitor graph that also flags merge conflicts. Gated by the existing
autofix_mode/trigger_mode settings, the enabled-repos opt-in, base-branch and
human-commit skip rules, dedupe, and a per-PR attempt cap.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review feedback on CI auto-fix
- Security: gate the no-mention review-feedback path on author trust —
require a trusted author_association (OWNER/MEMBER/COLLABORATOR) plus a
GitHub write/maintain/admin permission check before dispatching a
write-capable agent run, preventing privilege escalation from
read/triage/outside reviewers.
- Auth: reuse the originating PR thread's source + login/email when
dispatching fix runs so the GitHub-token resolver authenticates them in
non-bot-token deployments (bespoke github_ci source failed to resolve).
- Docs: document the Commit statuses: Read-only permission required for the
Status webhook event.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: stream reviewer eval logs on a dedicated admin page
Stream the eval subprocess output into a rolling log_tail and persist it
during the run (was only captured at exit), so the live output is visible
while the eval runs. Move the eval runner off the admin page onto its own
/admin/evals page (linked like Review Style Prompts) with a live log viewer.
* chore: drop unrelated SSR-register drift from generated route tree
* fix(ui): pre-bundle workbox-window to stop dev re-optimize reload
The PWA service worker (devOptions.enabled) pulls workbox-window, which
Vite discovers after first render and re-optimizes, forcing a reload that
cancels in-flight code-split route imports (Failed to fetch dynamically
imported module). Pre-bundling it via optimizeDeps.include avoids the
mid-session reload.
* fix(ui): suppress html hydration warning for pre-hydration theme script
The inline theme script sets class="dark"/color-scheme on <html> before
React hydrates, so the prerendered HTML never matches. suppressHydrationWarning
on <html> silences the (expected) one-level attribute mismatch.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: trigger reviewer evals from the admin page
Add an admin-only "Reviewer eval" section + endpoints that launch the
reviewer benchmark as an isolated subprocess against the running
deployment, with live status and the LangSmith experiment link. Route
eval traces to a dedicated open-swe-evals project so they stay out of
the production tracing project.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: reconcile reviewer eval status via heartbeat, not local process
The persisted record is shared across workers but _PROCS is process-local.
The owning worker now refreshes a heartbeat while the subprocess runs, and
status is only reconciled to failed once the heartbeat is stale, so a poll on
a worker without the local handle no longer kills a live run (and a duplicate
start is rejected across workers).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Kimi K2.7 is now the supported version, so drop the older K2.6 entry.
Update the related fireworks provider test to use K2.7 instead.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add Kimi K2.7 model option
Surface Moonshot's newly released Kimi K2.7 in the model picker, routed
through Fireworks following the existing kimi-k2pN convention, with the
same none/low/medium/high reasoning efforts (default high) as K2.6.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: correct Kimi K2.7 Fireworks model id to kimi-k2p7-code
Fireworks now hosts the K2.7 release as
accounts/fireworks/models/kimi-k2p7-code (the `kimi-k2p7` slug 404s).
Point the model option and its test at the live model id.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: drop unsupported "none" effort for Kimi K2.7
Kimi K2.7 Code is a thinking-only model — Moonshot documents that it does
not support non-thinking mode (passing thinking={"type":"disabled"} or
reasoning_effort="none" errors). Removing "none" from the advertised
efforts so the picker can't send an invalid Fireworks request.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* perf: speed up Reviews list (My PRs + All PRs)
Push the "My PRs" author filter into the threads.search metadata
(pr.author containment) instead of scanning up to 1000 reviewer threads
in Python, and replace the per-repo GitHub access check (an N+1 of
sequential GET /repos calls) with a single per-login accessible-repo set,
cached for 60s. Detail endpoints still re-validate access live.
Frontend: prefetch the inactive tab and adjacent page on hover/focus so
tab switches and pagination are instant.
* fix: don't cache repo access for the reviews list
The /reviews list is an authorization boundary for private PR metadata
(repo/PR titles, branches, authors, finding counts). A cross-request TTL
cache on the accessible-repo set could surface that metadata for up to
60s after a user lost repo access. Resolve the set fresh per request
instead — still a fixed, repo-count-independent burst of GitHub calls
(no per-repo N+1), with no staleness.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: mark threads resolved to hide from sidebar [closes resolve-threads]
Add a resolved flag (thread metadata) so users can clear finished
threads from the sidebar without deleting them. Resolved threads move to
a collapsible "Resolved" group (capped at 20 with a "Show all" link) and
are fully searchable/filterable + paginated on a new /agents/threads
page driven by URL query params. Sending a new message auto-unresolves.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: refresh paginated agent thread lists
Invalidate all agent thread list queries when thread lifecycle events change sidebar data, keeping the new paginated sidebar cache in sync.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* feat: render Reviews page diffs with pierre MultiFileDiff
The Reviews detail page used a hand-rolled hunk renderer with no syntax
highlighting. Switch it to the same pierre MultiFileDiff + theming the agent
chat git panel uses.
- review-diff API now returns full original/modified file contents instead of
hunks, via a shared build_pr_diff_files helper extracted from thread_api
- findings render as right-anchored markers (pierre line annotations); focus
highlight uses selectedLines; floating finding card still anchors to the marker
* feat: auto-collapse a review diff card when marked as viewed
* fix: URL-encode file path in Contents API fetch
Filenames containing reserved URL characters (#, ?) were truncated, so those
files rendered as empty/unrenderable. quote(path, safe='/') preserves the path
separators while escaping the rest.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: Reviews tab — anchored finding card, paginated list, file tree truncation
- Finding card now tracks the diff anchor while scrolling instead of staying
frozen in the viewport; auto-hides when its diff card collapses (including
collapse via mark-as-viewed) and on click outside
- Checks section capped with max height + scroll
- /reviews paginated (page size 20) with has_more, filtered to the current
user's PRs by default with an All toggle; PR author login now stored in
reviewer thread metadata
- File tree truncation marker overlapped filenames because the sidebar bg
was transparent; use the opaque sidebar color
* fix: finding card tracks anchor 1:1 while scrolling
Drop the vertical viewport clamp — it pinned the card at the clamp
boundary while the highlighted lines kept scrolling, breaking the
attachment.
* fix: anchor finding card with Base UI popover
Replace manual fixed-position tracking (laggy: setState per scroll
frame) with a Popover anchored to the finding's diff row. Floating UI
tracks the anchor outside React renders, so the card moves 1:1 with
the content and scrolls out of view with it. Unanchored findings keep
the fixed top-right card.
* fix: lock finding card to diff scroll
Replace the Base UI popover (async repositioning, paints a frame behind
native scroll) with a card absolutely positioned inside the scroll
container, so it scrolls with the diff in the same compositor frame.
Scroll moves to the ReviewBody root, side panel becomes sticky.
Position recomputes only on layout shifts via ResizeObserver.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: preflight stream proxy before SSE starts, let optimistic thread survive refetch
* fix: seed sidebar thread details as stale so mark-viewed fetch still fires
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: track PR lifecycle state per thread for sidebar
Persist a PR's draft/open/merged/closed state on the agent thread and keep
it in sync as PR webhooks fire, so the Agents UI can show per-thread PR
status the way Cursor does. open_pull_request now records the initial
draft/open state and pr_title; thread summaries expose diffStats; and the
PR webhook refreshes pr_state on close/reopen/draft/ready transitions.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor: consolidate PR state mapping into shared derive_pr_state helper
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: auto-recover from expired GitHub refresh tokens
When a user's GitHub OAuth refresh token was permanently dead (revoked or
expired), token refresh failed but get_valid_access_token still handed back
the known-stale access token, so dashboard GitHub calls kept 401ing until the
user manually logged out and back in.
Now we distinguish unrecoverable refresh failures (bad_refresh_token /
unauthorized_client) from transient ones: on an unrecoverable failure we drop
the dead stored authorization and return None, so callers prompt a clean
re-login. Transient failures still fall back to the stored token.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: don't delete fresh re-auth when stale refresh fails
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
- _origin_of: treat invalid ports as invalid origin instead of letting
urlparse ValueError turn CSRF rejections into 500s
- AgentsHome: reset submitting/draft when stream.submit rejects before a
thread id is minted, so the prompt isn't left disabled
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add View PR link in git panel header
Surface a clickable "View PR" link in the agent git panel header so
users can jump straight from a chat thread to its pull request instead
of hunting for the URL in the conversation. Uses the pr.url already
captured in thread metadata.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor: use shared buttonVariants for View PR link
Style the View PR link with the shared buttonVariants (outline/sm)
helper instead of hand-rolled classes, matching the existing
anchor-as-button pattern used in login.tsx and AutomationsList.tsx.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: show PR diff in git panel diff tab
Adds /threads/{id}/pr-diff endpoint fetching PR files + contents from GitHub;
git panel prefers it and falls back to the live message-derived diff.
* feat: auto-expand git panel when a PR lands mid-run
Panel still defaults to collapsed and remembers manual toggles; the
auto-expand is ephemeral and not written to localStorage.
* fix: git panel always starts collapsed
Drop localStorage persistence of collapsed state; panel opens only via
manual toggle or when a PR lands mid-run, and re-collapses on thread switch.
* fix: address PR review findings
Require user OAuth token for pr-diff (no app-token fallback) so GitHub
enforces current repo access; render placeholder for binary/oversized files.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add right-click open trace context menu for threads
Add a context menu to sidebar thread rows exposing an Open trace action
that links to the LangSmith thread trace, surfaced via a new traceUrl
field on the thread summary.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: drop stray yarn artifacts from PR
Remove accidentally committed ui/.yarnrc.yml and ui/.yarn/install-state.gz
that broke immutable installs against the v1 lockfile, and gitignore yarn
artifacts to prevent recurrence.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: reuse thread_id var, drop redundant danger-color fallback
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Fable 5 is currently unusable with our API key. Remove it from the
selectable model list; it can be re-added later.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: server-side Datadog/LangSmith observability tools + team creds
Add team-wide observability credential settings (Datadog DD_SITE/API/APP
keys, LangSmith API key) stored encrypted server-side, with an admin
dashboard section to connect/disconnect each provider. When connected,
get_agent loads read-only observability tools server-side: Datadog via its
hosted MCP server (langchain-mcp-adapters, toolsets=core) and LangSmith
read tools (langsmith_get_trace, langsmith_list_runs). Credentials live in
the LangGraph server process and are never exposed to the sandbox.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review on observability tools
Address PR review feedback:
- Authorize observability tools per triggering user (admins + the
OBSERVABILITY_AUTHORIZED_EMAILS allowlist) so prompt-injected runs from
untrusted contributors can't reach team Datadog/LangSmith data.
- Use the documented Datadog MCP auth headers DD_API_KEY / DD_APPLICATION_KEY.
- Store each provider's credentials under its own store key to avoid a
read-modify-write race dropping the other provider on concurrent saves.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: async email resolution in observability authorization gate
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
The viewed-marker feature (#1441) reintroduced an owner-only gate on
thread reads, regressing #1425 which made all threads readable by any
org user. Reads no longer assert ownership; viewed markers are only
written for the thread owner.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>