Commit graph

30 commits

Author SHA1 Message Date
Ramon Nogueira
3a0e2b4672
feat: plan mode with model-driven entry and collaborative review (#1580)
* feat: add plan mode for read-only research and planning

Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: enforce plan-mode read-only at tool layer and disable subagents

Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: harden plan-mode shell guard against wrapped mutations

Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow

- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline

Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.

* test(open-swe): add Playwright E2E for the Slack → PR → web handoff

Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.

- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).

Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.

* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright

The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.

Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.

* test(open-swe): record Playwright trace + video on every E2E run

Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.

* feat(plan-mode): collaborative plan review with BlockNote + Yjs

When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.

- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
  the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
  snapshots; plan content/status store; plan REST API (get/approve/reject,
  owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
  mounted under the agents shell, with a "Review plan" banner in the thread view
  and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
  flow, including cross-user comment sync and owner-only approval.

* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)

- plan-collab WS: authorize per-thread before joining a room (same read gate as
  the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
  opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
  resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
  dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
  unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
  before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.

Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).

* style: ruff format plan_collab.py

* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS

- Slack "Approve & Implement" now verifies the clicking user is the plan
  requester (owner, via the stored triggering_user_id) before implementing —
  matching the dashboard API's owner-only approval. Non-owners are pointed to
  Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
  allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
  the REST require_same_origin CSRF defense.

* fix(plan-mode): enter plan mode only via the model + local mock dev harness

Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.

- enter_plan_mode returns a terminating ToolMessage, fixing the missing
  ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
  remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
  call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
  Alice/Bob mock users, and a GitHub login picker.

* docs(plan-mode): drop stale references to removed profile/team defaults

The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.

* feat(plan-mode): let any reviewer edit the plan, not just comment

Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.

* test(plan-mode): assert plan-mode entry via the tool's success message

plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.

---------

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:06:58 -07:00
Johannes du Plessis
0bff510aae
fix: render GitHub-hosted images in PR descriptions on reviews page (#1589)
* fix: render GitHub-hosted images in PR descriptions on reviews page

PR description images hosted on GitHub (user-attachment uploads and
*.githubusercontent.com) render broken on the reviews page because
private-repo attachments require GitHub auth the browser session lacks.
Add an authenticated backend image proxy (host-allowlisted to guard
against SSRF) and route those image URLs through it from the reviews UI.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: harden PR image proxy (IDOR, SVG XSS, unbounded buffering)

Address review findings on the review-page image proxy:

- IDOR: the proxy fetched any *.githubusercontent.com URL with the App
  installation token, gated only by route-param repo access, so a user
  authorized for one repo could read images from another private repo the
  App can see. Bind the URL to the authorized PR — only proxy URLs that
  appear in that PR's body.
- SVG XSS: served any image/* inline from the API origin, including
  image/svg+xml which can run script. Restrict to safe raster types and
  add X-Content-Type-Options: nosniff + a locked-down CSP.
- DoS: enforced the size cap only after buffering the full response.
  Stream and abort once the cap is exceeded.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-22 14:01:06 -07:00
Johannes du Plessis
609e5551c4
fix: contain markdown render crashes with an error boundary (#1586)
Streamdown bundles Mermaid and renders ```mermaid blocks itself; a
diagram it can't parse throws during render and, with no boundary,
white-screens the whole review page. Wrap the markdown render so any
crash falls back to the raw text instead.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-22 12:54:05 -07:00
Johannes du Plessis
b19804536c
feat: handle images sent to non-vision models in Slack, Linear, and web UI (#1560)
* feat: handle images sent to non-vision models in Slack, Linear, and web UI

Add vision capability checks across all image input paths. When a user
sends images to a text-only model (e.g. GLM 5.2, DeepSeek V4 Pro), the
images are now skipped and a warning is injected into the prompt instead
of sending unsupported content to the model.

- Slack: resolve model at webhook time, skip image fetch + add warning
- Linear: same pattern as Slack
- Queued message middleware: read resolved model from thread metadata,
  strip images from queued payloads for text-only models
- Web UI: disable submit + show inline warning when images are attached
  to a non-vision model selection
- Shared: resolve_agent_model_id helper + vision_not_supported_warning

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: mock resolve_agent_model_id in Slack mention test

The test_process_slack_mention_queues_active_thread_message test was
missing a mock for the new resolve_agent_model_id call added to the
Slack webhook handler, causing a TypeError when image URLs triggered
the model resolution path.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: include vision warning in queued payload for text-only models

Update the prompt variable (not just content_blocks) before clearing
image_urls so the queued payload also carries the warning text when a
Slack/Linear follow-up arrives while the thread is busy.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 09:12:52 -07:00
Johannes du Plessis
3945757592
feat(ui): reintroduce subtle blue accents on greyscale frontend (#1555)
Restore a subtle blue --ui-accent token (driving send/stop buttons and
accent spots), the blue git status indicators for renamed/copied files,
and color the thread-list diff stat badge. Keeps the overall greyscale
aesthetic with tasteful blue/green/red accents.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 19:41:19 -07:00
Johannes du Plessis
bac1591888
Neutralize frontend: blue accents → neutral aesthetic (#1553)
Strip the blue/cool cast from the agent + reviewer chat and shared UI.
Neutralize the --ui-* (agents.css) and shadcn (styles.css) color tokens
to pure neutral grays (chroma/hue zeroed, lightness preserved); primary
and filled buttons become inverse-foreground neutral solids. Replaces a
few hardcoded blue/cyan spots (send button, prompt bar bg, context ring,
loop dot, git renamed status) with neutral tokens. Semantic green/amber/
red are kept.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 15:49:33 -07:00
Johannes du Plessis
e58b609b2f
fix: Simplify review explanation: full-width, plain prose, no diff links (#1547)
* Simplify review explanation: full-width, plain prose, no diff links

* Update _build_prompt test for plain diff fences

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 15:32:24 -07:00
Johannes du Plessis
876836bdfe
Normalize frontend typography + enable font smoothing (#1546)
Lighter, more consistent type across the dashboard and agent UI:
unify title sizes, drop semibold/bold in chrome to medium, harmonize
stray text-sm to text-xs, normalize chat markdown headings, and enable
antialiased font smoothing (fixes heavy text on dark backgrounds).

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 14:56:11 -07:00
Johannes du Plessis
801f93b4de
feat: AI-sorted PR review view with diff grouping (#1544)
* feat: AI-sorted PR review view with diff grouping

Group a PR's changed files into a logical top-to-bottom walkthrough via a best-effort structured-output LLM pass, kicked off concurrently with the reviewer run (~0 added latency) and persisted on reviewer thread metadata. Results render in the review UI behind an AI sorted / file tree toggle that persists across PRs; the view falls back to the file tree when groups are absent or stale.

Adds a grouping-model team default (inherits the Reviewer subagent model when unset) and the admin RolePicker for it.

* feat(reviews): richer AI-sorted explanations + sidebar polish

Sidebar group rows get Devin-style spacing (dividers, padding), a per-group file list (click a file to jump to its diff), inline-code chips in titles, and an accent Read explanation link.

Group explanations are now rich markdown: the grouping prompt feeds per-hunk line ranges and asks for inline code, a short code block, and [path:line](#loc=...) references. The Markdown renderer turns those #loc= links into in-page buttons that scroll the diff to the hunk and highlight the range, reusing the existing selectedLines path.

* fix(reviews): drop stale diff groups from the AI-sorted view

When groups were generated for a previous head, a persisted "ai" view in localStorage still rendered the outdated walkthrough. groupedView now returns null on diff_groups_stale, so the file-tree fallback is used and the view toggle hides until fresh groups arrive.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 13:59:37 -07:00
Johannes du Plessis
0108d764d9
fix(ui): resolve eslint errors in dashboard components (#1538)
Clears ~83 pre-existing eslint errors in the UI: auto-fixable rules
(array-type, import/order, sort-imports, type-only imports, redundant
assertions/conditions) plus manual fixes for unnecessary conditions and
banned @ts-nocheck directives. The 8 ported components excluded from
tsconfig are now also ignored by eslint so type-aware linting no longer
fails to parse them.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 17:29:38 -07:00
Caroline di Vittorio
3297e799de
fix: keep chat prompt bar editable while a run streams (#1533)
* fix: keep prompt bar editable while a run streams

The follow-up submit path awaited stream.submit, which resolves only when the
run finishes. That kept the react-query mutation isPending for the whole run,
so AgentThreadView disabled the prompt bar textarea the entire time the agent
was streaming - the user saw the "queue next" placeholder but couldn't click
in or type. Fire the run without awaiting its full lifecycle (matching the
new-thread path in AgentsHome) so the input stays editable for queueing.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor(ui): guard double-submit internally instead of via disabled prop

Stop threading a run-lifecycle signal (the send mutation's isPending) into the
prompt bar to gate the input. Instead, the prompt bar owns a synchronous
double-submit guard (submittingRef) plus a short-lived isSubmitting state
scoped to the in-flight send. The textarea now stays editable while a run
streams (so follow-ups can be queued), and onSubmit is awaitable so the guard
tracks the actual send request rather than the run.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix(ui): surface failed run-start instead of leaving thread running

When stream.submit rejects (e.g. 401 expired token or 409 active-run
race), the fire-and-forget catch was a no-op while onSuccess had already
optimistically set status: running, leaving the thread stuck in a busy
state with no surfaced error. Clear the busy state and mark the thread
errored on submit failure.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-06-15 17:29:17 -07:00
Caroline di Vittorio
dc70ca1f2f
feat(ui): soften bash tool output scroll indicator to a fade (#1532)
Replace the hard inset box-shadow scroll indicators on bash/shell tool
output with a mask-image gradient so the output softly fades at the
scroll edges, matching the user message bubble treatment.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 20:39:52 +00:00
Johannes du Plessis
2b181a77a6
fix: add bottom padding to shell command card before output streams (#1523)
The bash command card had only pb-1 on the header block, leaving the
content cramped against the card's bottom edge until output streamed in
(which adds its own pb-2). Apply pb-3 when there is no output yet so the
card looks uniformly padded in the pending/in-progress state.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 09:38:30 -07:00
Caroline di Vittorio
e8bb6b497b
feat: mark threads as resolved to hide from sidebar (#1500)
* feat: mark threads resolved to hide from sidebar [closes resolve-threads]

Add a resolved flag (thread metadata) so users can clear finished
threads from the sidebar without deleting them. Resolved threads move to
a collapsible "Resolved" group (capped at 20 with a "Show all" link) and
are fully searchable/filterable + paginated on a new /agents/threads
page driven by URL query params. Sending a new message auto-unresolves.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: refresh paginated agent thread lists

Invalidate all agent thread list queries when thread lifecycle events change sidebar data, keeping the new paginated sidebar cache in sync.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-06-12 12:01:42 -07:00
Brendan Whiting
639058a8b6
feat: support pasting images from clipboard into prompt bar (#1512)
The prompt bar only accepted images via drag-and-drop or the file picker.
Pasting an image from the clipboard inserted the file name as text instead
of attaching the image. Add an onPaste handler that extracts image files
from the clipboard and attaches them as pending images.

Co-authored-by: Brendan Whiting <16016903+bwhiting2356@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-12 09:48:48 -07:00
Christian Bromann
abf354bb05
feat(open-swe): stream agent chat via @langchain/react v2 protocol (#1475)
* feat(dashboard): stream agent chat via @langchain/react v2 protocol

Replace the bespoke SSE + React Query polling path with LangGraph’s
v2 event stream through credentialed dashboard proxies. Run starts go
through stream commands; mid-run follow-ups still queue via /messages.

* fix import path

* fix tests after rebase

* format

* PR feedback

* improved model fallback

* fix image handling

* embrace sdk

* cleanup

* cr

* more cleanup

* fix cors

* harden security

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 09:54:35 -07:00
Caroline di Vittorio
5515fff44c
fix: distinguish agent code/shell blocks from user message bubbles (#1497)
* fix: distinguish agent code/shell blocks from user message bubbles

Code blocks and shell command output shared the same accent bubble
background as user message bubbles, making them hard to tell apart at
full width. Give them a distinct neutral gray background plus a subtle
border.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: apply neutral code bubble to all agent output cards + add second accent

Extend the neutral code-bubble background to the remaining agent output
cards (files-changed card, reply preview, edited-file diff, todo list,
streaming-turn cards) so none of them share the user message bubble
background. Introduce a violet --ui-accent-2 used on code/shell header
labels so the blocks pop.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 09:44:10 -07:00
Johannes du Plessis
4cabdb2351
feat: add pierre file tree to full-screen git panel (#1487)
* feat: add pierre file tree to full-screen git panel

Render the dashboard git panel with real changed-file data using pierre
diffs, add a full-screen toggle, and show a pierre FileTree explorer on the
right when expanded.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: always show git panel as collapsible resizable card

- git panel renders on every thread with collapse to a floating expand
  button and drag-resize, persisted to localStorage (matches left sidebar)
- panel content wrapped in a rounded card with Git/Desktop/Terminal tabs
  above it; Desktop/Terminal and Review/Commits show Coming Soon
- send button morphs into a stop button during an active run

* fix: follow app theme for pierre diff rendering

diffOptions used themeType "system" so @pierre/diffs followed the OS
color scheme instead of the app's .dark class, producing unreadable
dark-on-light diffs when the two disagreed. Add useDiffOptions() which
resolves themeType from the app theme, and use it at all diff render
sites.

* fix: align git panel card bottom gutter with prompt bar

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 15:13:12 -07:00
Johannes du Plessis
1135e9342e
feat: add cancel run button to agent UI (#1485)
* feat: add cancel run button to agent UI

Wire the existing cancel-thread mutation into the agent thread view by
adding a stop button to the prompt bar that appears while a run is active.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve thread messages when cancel response omits them

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 13:27:04 -07:00
Johannes du Plessis
9475afdf2e
fix: theme-aware syntax highlighting in chat code blocks (#1479)
Inline chat code blocks hardcoded the github-dark Shiki theme, so in light
mode dark-theme token colors rendered on a light bubble with poor contrast.
Resolve the Shiki theme from the active light/dark mode and cache tokens per
theme. Adds a reactive useResolvedTheme hook so blocks update live on toggle.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 09:58:20 -07:00
Johannes du Plessis
5430672edb
fix: revert dashboard stop run controls (#1461)
Revert the dashboard stop-button behavior from #1433 because it allows duplicate submissions while optimistic prompts are still pending.
2026-06-08 21:47:53 -07:00
Johannes du Plessis
b0cfa2cfad
feat: show sandbox setup status (#1455)
* feat: show sandbox setup status

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: avoid sandbox setup status for queued follow-ups

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-08 14:35:21 -07:00
Johannes du Plessis
072c0158ff
feat: support dashboard chat images (#1435)
* feat: support dashboard chat images

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: keep pending image prompts visible

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-06 11:50:24 -07:00
Johannes du Plessis
1f619eae0f
feat: add stop button to cancel running agent from web UI (#1433)
* feat: add stop button to cancel running agent from web UI

* fix: handle stopped agent runs

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-06 08:12:30 -07:00
Johannes du Plessis
6f330f777d
fix: tidy agents home + mobile logo + PWA dev/icon issues (#1419)
* fix: tidy agents home, mobile logo, and PWA dev/icon issues

- Remove recent-runs cards from the new agent page
- Scale the ASCII logo to fit narrow viewports instead of overflowing
- Serve manifest in dev and skip SW registration in dev to clear console errors
- Use a flat, opaque apple-touch-icon so iOS stops adding a black border
- Ignore generated ui/dev-dist

* feat: add send button and prevent iOS focus-zoom on chat input

- Add a circular send button (accent, spinner while sending) to the prompt bar
- Bump textarea to 16px on mobile so iOS doesn't zoom the viewport on focus

* fix: polish prompt bar and center logo

- Center the ASCII logo at all widths
- Prevent iOS focus-zoom via viewport maximum-scale instead of bumping the
  input to 16px, so mobile text stays the intended size
- Give the model picker a chip/chevron treatment to anchor it next to the send button

* fix: simplify model picker to plain text + chevron

* fix: tighten prompt bar padding and shrink send button

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-04 16:47:35 -07:00
Johannes du Plessis
a75e9b027b
fix: stabilize agent selector defaults (#1414)
* fix: stabilize agent defaults and repo selector

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* fix: align default repo selector styling, restore text fallback

Match the repo selector to sibling settings controls (h-7, bg-input/20,
text-xs). Fall back to a text input when the repo list is empty so a
default repo can still be entered.

* fix: make repo selector dropdown more compact

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-04 21:06:25 +00:00
Johannes du Plessis
c1e46b938e
feat: add Slack Block Kit reply options (#1407)
* feat: add Slack Block Kit reply options

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* docs: document Slack interactivity setup

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-04 10:26:09 -07:00
Johannes du Plessis
27677c4949
feat(dashboard): render Slack/Linear replies as a card in chat (#1399)
Slack thread replies and Linear comments showed only the bare tool name in the dashboard chat. Map them to dedicated 'slack'/'linear' toolKinds and render the message body in a ReplyCard, so Open-in-Web shows what the agent actually posted.
2026-06-03 22:05:47 +00:00
Johannes du Plessis
dee78e7e84
fix: make repository optional when starting a dashboard run (#1397)
* fix: make repository optional when starting a dashboard run

The agent infers and clones the target repo from the task itself, so a
default repo is never actually required to run — but the dashboard 400'd
("no default repository configured") when a user had none set.

Treat repo as optional: _resolve_repo_config returns {} instead of
raising, repo metadata/config are only written when a repo is present,
and the "missing repository metadata" gate on follow-up messages is
dropped. UI hides the repo chip when absent.

* feat: add repo picker to the run prompt bar

Adds an optional, searchable repository selector next to the model picker
on the Agents home prompt bar (Cursor-style). It pre-fills the user's saved
default repo and can be cleared to "No repository" for a repo-less run.

Because the picker now resolves the default on the client, the create
endpoint honors the request value verbatim: _resolve_repo_config just parses
what's sent ({} when empty) instead of falling back to the saved default,
so an explicit "No repository" is respected.

* refactor: match Cursor layout for repo placement

Move the repo selector out of the prompt-box footer to a pill row above
the input (folder + caret, dropdown opens downward); the model picker
stays inside the box. In the thread view, show the thread title and repo
in a header at the top of the chat, with the follow-up input pinned to
the bottom as before.
2026-06-03 14:24:08 -07:00
Johannes du Plessis
e87085139b
feat: add Agents chat UI for cloud threads (#1323)
* feat(ui): add Agents chat UI ported from open-swe-app

Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(dashboard): wire Agents UI to LangGraph thread APIs

Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(dashboard): single agent reply per turn in Agents UI

Use UUID thread IDs LangGraph accepts, skip confirming_completion for
dashboard threads, and merge adapter agent messages so duplicate bubbles
do not render.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(ui): polish Agents UI with floating prompt and layout cleanup

Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar
from open-swe-app, and refine chat layout so messages scroll behind the input.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent): patch deepagents reducer for None messages on checkpoint replay

LangGraph thread state could 500 when cancelled runs left messages as None.
Apply the reducer guard before graph import, fall back to metadata in the
dashboard API, and adjust Agents prompt bar layout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(ui): unify sidebar user menu and clean up Agents UI navigation

Extract SidebarUserMenu so the dashboard and Agents sidebars render the
same profile button, drop the redundant Agents nav row in favor of the
existing Back to Agents link, add the open-swe logo header to the Agents
sidebar, flatten the New Agent button, and cap the home screen run list
to keep the prompt input in view.

* feat(ui): resizable/collapsible sidebar shared across dashboard and Agents

Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars
share a persisted width (default 260px, drag to resize, 200-420 range)
and a collapse toggle that hides the panel and surfaces a floating
reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on-
hover thread delete control in the Agents sidebar.

* feat(ui): instant user message and busy indicator on Agents transition

Stash submitted prompts in sessionStorage, pre-populate the new thread
detail cache, and merge pending prompts into the rendered message list
so the Agents page renders the user bubble plus the existing thinking
spinner immediately instead of flashing a skeleton and "Agent is
starting" while the run boots.

* feat(ui): token-stream agent replies in the Agents thread view

Opt the LangGraph runs into messages-tuple streaming and forward those
events through the existing SSE channel. The frontend now applies
AIMessageChunk deltas directly to the cached thread (cancelling any
in-flight refetch first so optimistic tokens are not clobbered) and
keeps positional pending prompts so the user bubble stays in the right
place while the agent streams its reply.

* fix(dashboard): await threads.join_stream before iterating

threads.join_stream is async def returning an AsyncIterator, so it must
be awaited before async for. The SSE endpoint was raising
TypeError: 'async for' requires an object with __aiter__ method, got
coroutine on every connection.

* fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns

Setting stream_mode=["values","messages-tuple","updates"] on
runs.create forces langchain_anthropic into streaming, and on the
second model call (after tool execution) its serialized thinking
blocks come back malformed, so Anthropic rejects the request with
'messages.1.content.0.thinking.thinking: Field required'. Revert to
the default stream_mode so claude-opus thinking + tool use runs to
completion. The frontend keeps the messages-event handler in place
as a no-op fallback for when streaming is re-enabled.

* feat(agents): per-thread model picker wired through to the run

Add optional model_id/effort to the create-thread and send-message
request bodies, forward them as agent_model_id/agent_effort in the
LangGraph run configurable, and record the resolved choice in thread
metadata so the UI can show the model the run is actually using.
get_agent now picks the per-thread override last (highest priority over
team default + profile override) and falls back gracefully when it is
absent or unsupported.

The frontend prompt bar becomes a controlled component fed by a
shared useModelOptions hook (options + profile -> defaultSelection).
AgentsHome seeds the picker from the user's profile default; the
thread view seeds from the thread's recorded model/effort and lets
each follow-up retarget the run.

* refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar

Drop the absolute-positioned send button, restore the original
px-4 py-3.5 min-h-[106px] flex-col container, and move the model
picker into a mt-auto pt-2 footer row so the placeholder text and
the model selector share the same horizontal padding.

* chore: fix lint/format CI failures

Remove unused imports and reformat two files flagged by ruff.

* fix(tests): stop messages-reducer patch tests from polluting the suite

Restore agent modules after reducer patch tests and import LangSmithSandbox
from agent.server in proxy refresh tests so isinstance checks stay valid.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 18:15:59 +00:00